Skip to content
OTFotf
All posts

How to test whether coding agents discover the skills you need

D
DaveAuthor
8 min read
How to test whether coding agents discover the skills you need

A coding skill can be clear, correct, and still fail to help if an agent never finds it. Testing a skill by naming it directly only answers whether the agent can use it after it has been loaded. It does not show whether an agent will discover it from the kind of request a teammate or customer actually makes.

Expo’s October 8, 2026 evaluation write-up describes a practical way to test that gap: give coding agents realistic app-building tasks, inspect their traces, and measure whether they load relevant skills without being told a skill’s name. The team’s experiment is useful as a method, not a universal performance promise. Its results came from one plugin, one model setup, and ten app-building tasks.

If you maintain a set of instructions for coding agents, the useful question is not only “Does this skill work when invoked?” It is also “Will the agent find the right entry point during ordinary work, and what evidence shows it followed the guidance?”

Separate discovery, use, and result

Expo’s evaluation harness tracked three different signals:

  • Skill triggering: Did the agent load a skill expected to help with the task?
  • Skill uptake: Is there evidence in the agent’s work that it followed the loaded guidance?
  • App behavior: Did the resulting app build and perform the behaviors described in the task?

Keep these measures separate in your own evaluation. A missing skill trigger means the guidance never entered the agent’s context. A trigger without uptake means the agent read the instruction but did not use it. A task can also succeed without a skill: the agent may already know how to do the work. That can be a good result for the user, even if it means the skill did not contribute to that run.

This distinction is why “the final answer looked right” is weak evidence of skill discovery. It cannot tell you whether the agent found your instructions, followed them, or completed the task from its existing knowledge. Read the trace alongside the output.

A coding agent follows a natural task prompt to a clear entry skill, then to focused guidance while a trace records each step

For a second perspective on evaluating skill contribution after a skill is loaded, see our guide to measuring plugin lift in Claude Code evals. That answers a related but different question: whether a skill changes task outcomes compared with a run without it. Discovery testing asks whether an ordinary request leads the agent to the skill in the first place.

Start with ordinary task prompts

Create a small set of tasks that reflect work people actually ask an agent to do. Use broad requests such as “Add tabs and a settings screen to this app” or “Debug why this development build keeps crashing.” Do not mention the skill name, command, or desired lookup in the prompt. If the prompt says “load the navigation skill,” it tests obedience to an explicit instruction, not discovery.

Before each run, write down which skills should be relevant and what observable evidence would count as their use. Expo’s harness annotated scenarios with the skills expected to help, evidence that the guidance was followed, and the product behaviors to check. This keeps the scoring from being rewritten after you see what the agent did.

Use a mix of task shapes. Include a new project, a change to an existing project, and a debugging task. A small but varied suite can reveal whether discovery works only for a narrow prompt pattern. Keep each task stable across runs so that you can compare changes to your skill catalog rather than changes in the request.

Then inspect the trace, not just the final response. Record which skills loaded, in what order, and whether the agent moved from a general entry point to a specialized instruction. Separately inspect the edited files or app behavior for specific recommendations your skill contains. A skill-trigger check alone is not proof that the advice changed the implementation.

11 production screens. Login, database, payments — all wired.

The SaaS Dashboard Kit ships everything already connected. Nothing to set up. Live demo at saas.otf-kit.dev.

See the live demo

Give a large collection one clear starting point

When a collection has many specialized instructions, the agent has to choose among them while also solving the task. Expo added expo-overview as a starting point that routes broad requests to more focused guidance. In three example scenarios, the entry point loaded first and the traces then showed relevant follow-on skills.

That is a design pattern worth testing: give the collection one obvious entry point, make its purpose clear in the title and opening description, and include a compact map from common goals to specialized instructions. The entry point should help the agent choose; it should not become a copy of every detailed skill.

Do not assume that adding a router solves every discovery miss. Expo reported that some relevant downstream skills still did not load, even after the entry point was present. The trace can show where selection stopped: perhaps the router did not direct the agent clearly enough, or the agent continued from its own knowledge instead. Treat each miss as evidence for the next targeted change rather than rewriting every skill at once.

For installation and supported-agent setup, use our separate Expo Skills installation guide. This article focuses on evaluating discovery after the skills are available to an agent.

Test under the catalog conditions people use

A skill may be visible in a clean environment and harder to discover in a crowded one. Expo’s post describes a Codex CLI catalog test where a fixed token allowance was shared across available skills. When the catalog exceeded that allowance, descriptions were shortened from their beginnings; with more crowding, descriptions and eventually skills could disappear. In one environment with 175 installed skills, the Expo entries remained listed but only around the first 60 characters of each description were visible.

Those are observations from Expo’s tested version and setup, not a guarantee about every agent or release. The practical implication is to test your own catalog in the environment your team uses. Include the other plugins and instructions people normally install, then check what the agent can actually see before it starts a task. A clean-room test is still useful for comparing two versions of your own collection, but it cannot answer whether the skills remain discoverable amid the user’s full set of tools.

A clean skill catalog keeps descriptions visible, while a crowded catalog clips the opening cues an agent needs

Make each skill’s title and first sentence informative enough to survive truncation. Put its purpose and the situations that should trigger it early. Then rerun the same natural prompts in the crowded environment and compare traces with the clean environment. If discovery falls, inspect the visible catalog before adding more prose to the skill body.

Read the percentages in context

Expo compared the collection before and after adding its entry point. The evaluation used Claude Code with sonnet-5-high and ten greenfield mobile-app task descriptions. Each task ran three times in each condition; the prompt mentioned the platform but did not name any skill.

In the baseline, 9% of sessions loaded any Expo skill. After adding the entry point, 55% loaded at least one specialized downstream skill. Expo also reported that relevant-skill recall rose from 2% to 18%, meaning agents still missed many skills that the team had marked as relevant to a task. These are Expo’s measurements on its own task set and plugin. They are not a general rate you should expect from your collection.

The gap between 55% downstream loading and 18% relevant-skill recall matters. An entry point made the first discovery step more common, but it did not ensure that every useful instruction was found. That is why the next run should examine both the first skill and the route to later skills.

Turn discovery into a repeatable check

A small evaluation loop can fit into normal skill maintenance:

  1. Choose realistic tasks and record the expected skills before running them.
  2. Run the tasks without naming skills in the prompts.
  3. Inspect traces for triggering and order; inspect the work product for uptake and behavior.
  4. Repeat with the full catalog installed, not only a clean test environment.
  5. Change one part of the collection, such as an entry-point description, then rerun the same tasks.
  6. Keep the traces and results with the change so the next maintainer can see what improved and what still fails.

A single successful run is a lead, not enough evidence to generalize. Agent behavior varies. Repeat tasks enough to see whether a change consistently shifts discovery, and keep the prompts and environment fixed when comparing versions. Avoid optimizing for “skill loaded” as an end in itself: the purpose is to make useful guidance available when it can improve the work, while accepting that an agent may already solve some tasks well without it.

The durable lesson from Expo’s experiment is straightforward: evaluate the conditions that determine whether guidance reaches the agent, not only the guidance after it arrives. Natural prompts, trace inspection, a clear entry point, and a realistic catalog give skill authors a way to find discovery failures before they become a user’s debugging session.

Sources

ai-toolsagentsarchitecture
OTF SaaS Dashboard Kit

Ship the product, not the setup.

  • 11 production screens — auth, billing, team, analytics, settings
  • Real database, payments, and login — all wired on day 1
  • AI configs pre-tuned so your agent extends instead of regenerates
Need more than components?

Full-stack kits.
Pay once, own the code.

Auth, database, and payments already connected — so you ship product, not setup. Or take the delivered kits in the Bundle.

Everything Bundle — $149See full pricing

Get the free AI configs pack

Pre-tuned AI configs for Cursor, Claude, and Lovable — drop them in and your AI tool instantly understands your project.

No spam. Unsubscribe any time.

Prefer the free SDK? Star it on GitHub →