|
| 1 | +--- |
| 2 | +name: build-agent-eval-scaffold |
| 3 | +description: "Creates a realistic project scaffold — a test fixture — for rubric-based effectiveness evaluation of a Claude Code agent. Given a plugin:agent identifier and an optional project description, analyzes the target agent's definition file to learn what inputs it expects and what signals it looks for, then interviews the user in three phases (technology and shape, signals to plant, file plan) before writing the scaffold to tests/test-suites/{agent}/scaffolds/{name}/. Use when creating, building, or setting up a test scaffold, fixture project, or sample codebase to evaluate an agent's output quality against. Does not create tests.json entries or rubric files — use write-acil-evals or write-agent-eval-rubric for those. Does not create skill scaffolds — use build-skill-eval-scaffold for skills." |
| 4 | +argument-hint: "[plugin:agent] [optional description] e.g. example-plugin:gap-analyzer for a rails 7 project with postgres" |
| 5 | +allowed-tools: Read, Write, Glob, Grep, Bash(mkdir *) |
| 6 | +--- |
| 7 | + |
| 8 | +Build a realistic project fixture that the test harness runs the target agent against, so an LLM judge can score the agent's output against a rubric. The fixture only works as an eval when it contains signals the target agent is designed to find and reads like code a real developer wrote, so this workflow analyzes the target agent first, then interviews the user in three phases before writing any files. |
| 9 | + |
| 10 | +## Constraints |
| 11 | + |
| 12 | +These apply to every scaffold and shape both the file plan (Step 5) and generation (Step 6): |
| 13 | + |
| 14 | +- Always write the scaffold to `tests/test-suites/{agent}/scaffolds/{name}/`, relative to the repository root, BECAUSE the harness discovers scaffolds at that path when it builds the Test Sandbox. |
| 15 | +- Never write a `.git` directory BECAUSE the harness auto-initializes a git repo with `git init` and commits all scaffold files itself. |
| 16 | +- Never write lock files (`package-lock.json`, `Gemfile.lock`, `go.sum`) or dependency directories (`node_modules`, `vendor`, `__pycache__`) unless one is itself a planted signal BECAUSE they add hundreds of generated lines the target agent never inspects and bury the signals that matter. |
| 17 | +- Never mark a signal with comments like `BUG HERE` or `INTENTIONAL ISSUE` BECAUSE a signal the agent can find by reading a comment measures nothing about its analysis; signals must require the same work a real codebase would. |
| 18 | +- Always keep every file syntactically valid for its language (it should parse or compile apart from intentional logic bugs) BECAUSE a file that fails to parse makes the agent report the syntax error instead of the planted signal. |
| 19 | +- Keep interview turns compact: present the list, then the question. Put explanation in the "Why" line of each signal rather than in surrounding prose. |
| 20 | + |
| 21 | +## Step 1: Parse arguments |
| 22 | + |
| 23 | +Parse the user's input into two parts: |
| 24 | + |
| 25 | +- **`plugin:agent`** (required) — the first token, which contains a colon. Split on the colon to extract the plugin name (before the colon) and the agent name (after the colon). |
| 26 | +- **Description** (optional) — everything after the first token. This describes the technology and shape of the scaffold project (e.g., "for a rails 7 project with postgres as the database"). |
| 27 | + |
| 28 | +If no argument was provided, or the first token contains no colon, ask the user which `plugin:agent` to create a scaffold for and stop until they answer. |
| 29 | + |
| 30 | +Validate that the agent exists by confirming the file `{plugin}/agents/{agent}.md` exists in the repository root. If it does not exist, tell the user and ask them to correct the input. |
| 31 | + |
| 32 | +## Step 2: Analyze target agent |
| 33 | + |
| 34 | +Read the agent's definition file at `{plugin}/agents/{agent}.md`. Agents are self-contained markdown files with YAML frontmatter and a prompt body. They do not have `references/` directories, `scripts/`, or context injection commands. |
| 35 | + |
| 36 | +Analyze the full agent definition to identify: |
| 37 | + |
| 38 | +- What inputs and environment the agent expects (source code files, config files, documentation, specific project structure) |
| 39 | +- What outputs the agent produces (analysis reports, gap assessments, architectural reviews, etc.) |
| 40 | +- What signals the agent looks for (bugs, security flaws, architectural patterns, missing documentation, implementation gaps) |
| 41 | +- What tools the agent uses (Read, Glob, Grep, Bash commands — these reveal what file types and patterns the agent inspects) |
| 42 | +- Whether the agent operates on project files at all, or whether it operates on external state (GitHub PRs, CI pipelines, conversation context) |
| 43 | + |
| 44 | +**Graceful skip:** If the analysis reveals the agent does not operate on project files — for example, it queries GitHub APIs (`gh` commands), operates on pull request state, or generates content from conversation context rather than analyzing files in a working directory — inform the user that a file scaffold would not be useful for this agent type and stop. Explain what the agent operates on instead. |
| 45 | + |
| 46 | +## Step 3: Interview Phase 1 — Analysis and project shape |
| 47 | + |
| 48 | +Present the following in one message: |
| 49 | + |
| 50 | +1. **Agent purpose** — a one-sentence summary of what the agent does |
| 51 | +2. **Expected inputs** — what file types, config files, and project structure the agent expects to find in a working directory |
| 52 | +3. **Signal categories** — the kinds of issues, patterns, or signals the agent detects (grouped into categories) |
| 53 | +4. **Environment requirements** — any specific project conventions the scaffold needs to follow (e.g., needs a Gemfile for Ruby, needs a go.mod for Go) |
| 54 | +5. **Existing scaffolds** — use Glob on `tests/test-suites/{agent}/scaffolds/*/` and list any found by name so the user can avoid duplicating one; if the suite directory does not exist yet, note that it will be created |
| 55 | +6. **Proposed project** — if a description was provided in the arguments, restate the tech stack and project shape it implies; otherwise propose one or two options that fit the agent's expected environment from Step 2 |
| 56 | +7. **Proposed scaffold name** — derive it from the description: kebab-case with a `-project` suffix (e.g., "rails 7 with postgres" becomes `rails-postgres-project`, "python flask API" becomes `python-flask-project`). If `tests/test-suites/{agent}/scaffolds/{name}/` already exists, say so and ask whether to overwrite it or choose a different name. |
| 57 | + |
| 58 | +Ask the user to confirm or correct the analysis, the tech stack, and the scaffold name in one reply. Wait for their answer, and carry any corrections into the remaining steps. |
| 59 | + |
| 60 | +## Step 4: Interview Phase 2 — Signals to plant |
| 61 | + |
| 62 | +Based on the Step 2 analysis, suggest specific signals to plant in the scaffold. Present them as a numbered list where each entry includes: |
| 63 | + |
| 64 | +- **What** — the signal itself (e.g., "SQL injection via string interpolation in a database query") |
| 65 | +- **Where** — where it would live in the scaffold (e.g., "in a database access layer module") |
| 66 | +- **Why** — why this matters for the target agent (e.g., "the agent specifically checks for parameterized queries vs. string interpolation") |
| 67 | + |
| 68 | +Draw suggestions from the agent definition — what the agent's instructions explicitly check for and what analysis patterns it follows. |
| 69 | + |
| 70 | +Present the list and ask the user to: |
| 71 | + |
| 72 | +- Approve the list as-is |
| 73 | +- Remove signals that are not relevant to their testing goals |
| 74 | +- Modify signals to be more specific to their project context |
| 75 | +- Add signals not covered by the analysis |
| 76 | + |
| 77 | +Wait for the user to finalize the signal list before proceeding. |
| 78 | + |
| 79 | +## Step 5: Interview Phase 3 — File plan |
| 80 | + |
| 81 | +Present a complete file plan as a structured list. For each file include: |
| 82 | + |
| 83 | +- **Path** — relative to the scaffold root directory |
| 84 | +- **Description** — what the file contains and its role in the project |
| 85 | +- **Signals** — which signals from Step 4 this file carries (reference by number), or "none" for clean files |
| 86 | + |
| 87 | +Design the file plan against these guidelines: |
| 88 | + |
| 89 | +- Include the standard config files for the tech stack (package.json, Gemfile, go.mod, requirements.txt, etc.) BECAUSE the agent uses them to detect the language and framework. |
| 90 | +- Keep the plan focused — enough structure to feel like a real project, but only files the agent would actually inspect or that establish necessary context — BECAUSE every extra file costs the eval run tokens without adding signal. |
| 91 | +- Plan source files at roughly 50-150 lines each BECAUSE shorter files read as stubs and longer ones hide signals in bulk. |
| 92 | +- Spread signals across multiple source files in a realistic directory structure, and include some files with no intentional signals, BECAUSE real projects mix clean and problematic code; an agent shown only problem files is never tested on telling the two apart. |
| 93 | + |
| 94 | +Present the file plan and wait for the user to approve it. The user may request additions, removals, or modifications to the file plan. |
| 95 | + |
| 96 | +## Step 6: Generate scaffold |
| 97 | + |
| 98 | +After the user approves the file plan: |
| 99 | + |
| 100 | +1. Create the directory structure using `mkdir -p tests/test-suites/{agent}/scaffolds/{name}/` including any subdirectories needed for the planned files. |
| 101 | + |
| 102 | +2. Write each file using the Write tool. For each file, generate realistic content that matches the tech stack and project context, and include its planned signals the way a real developer would have written them — the Constraints above govern what the content may and may not contain. |
| 103 | + |
| 104 | +3. Report the outcome: the scaffold path first, then the complete list of files created with their paths relative to the repository root. |
0 commit comments