Skip to content

Commit 3c9bf09

Browse files
committed
feat: package eval-authoring skills as a plugin and align them with skill-building guidance
Move build-skill-eval-scaffold and build-agent-eval-scaffold out of .claude/skills/ into a top-level eval-authoring plugin with a .claude-plugin/plugin.json manifest, and add a root .claude-plugin/marketplace.json (skills-test-harness) that serves it from a relative path. Rework both SKILL.md files against the han-plugin-builder guidance: - Lead with the goal and reasons; front-load scaffold constraints in Always/Never ... BECAUSE form instead of leaving them in the last step - Merge the separate analysis-confirmation pause into the technology and shape phase, so the workflow has the three interview gates its description promises instead of five - Resolve subagent_type values as namespaced plugin:agent and look up the agent in its own plugin, not only the target skill's plugin - Handle a first argument with no colon; add trigger breadth (fixture project, sample codebase) to the descriptions while staying well under the 1024-character cap Update the two docs pages' Workflow sections to match the six-step flow.
1 parent 1193b4d commit 3c9bf09

6 files changed

Lines changed: 271 additions & 14 deletions

File tree

‎.claude-plugin/marketplace.json‎

Lines changed: 28 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,28 @@
1+
{
2+
"name": "skills-test-harness",
3+
"description": "Plugins that ship with the skills-test-harness: eval-authoring skills that generate test scaffolds and eval suites for Claude Code skills and agents.",
4+
"version": "1.0.0",
5+
"owner": {
6+
"name": "Test Double"
7+
},
8+
"plugins": [
9+
{
10+
"name": "eval-authoring",
11+
"displayName": "Eval Authoring",
12+
"source": "./eval-authoring",
13+
"description": "Skills that author eval suites for the skills-test-harness. Home of build-skill-eval-scaffold and build-agent-eval-scaffold, which interview you and generate realistic project scaffolds for rubric (effectiveness) evaluation of a Claude Code skill or agent. Depends on nothing; run it from the harness repo where the target plugin:skill or plugin:agent lives.",
14+
"version": "1.0.0",
15+
"defaultEnabled": true,
16+
"author": {
17+
"name": "Test Double"
18+
},
19+
"homepage": "https://github.com/testdouble/skills-test-harness#eval-authoring-skills",
20+
"repository": "https://github.com/testdouble/skills-test-harness",
21+
"license": "MIT",
22+
"keywords": ["evals", "testing", "scaffold", "rubric", "skills", "agents", "claude-code"],
23+
"category": "developer-tools",
24+
"tags": ["eval-authoring", "test-scaffolds", "llm-judge"],
25+
"strict": true
26+
}
27+
]
28+
}

‎docs/build-agent-eval-scaffold.md‎

Lines changed: 5 additions & 6 deletions
Original file line numberDiff line numberDiff line change
@@ -55,15 +55,14 @@ The following are excluded from scaffolds:
5555

5656
## Workflow
5757

58-
The skill walks through a 7-step process:
58+
The skill walks through a 6-step process with three interview pauses:
5959

6060
1. **Parse arguments** — extract the `plugin:agent` identifier and optional project description
6161
2. **Analyze target agent** — read the agent's definition file to understand what inputs, signals, and environment the agent expects
62-
3. **Present analysis summary** — show agent purpose, expected inputs, signal categories, environment requirements, and any existing scaffolds
63-
4. **Interview: Technology and shape** — confirm the tech stack and derive a kebab-case scaffold name with `-project` suffix
64-
5. **Interview: Signals to plant** — suggest specific signals based on the agent analysis; the user approves, removes, modifies, or adds signals
65-
6. **Interview: File plan** — present a complete file plan with paths, descriptions, and signal assignments for each file
66-
7. **Generate scaffold** — create directories and write all files with realistic content
62+
3. **Interview: Analysis and project shape** — present the agent's purpose, expected inputs, signal categories, environment requirements, and any existing scaffolds alongside the proposed tech stack and a kebab-case scaffold name with `-project` suffix; the user confirms or corrects all of it in one reply
63+
4. **Interview: Signals to plant** — suggest specific signals based on the agent analysis; the user approves, removes, modifies, or adds signals
64+
5. **Interview: File plan** — present a complete file plan with paths, descriptions, and signal assignments for each file
65+
6. **Generate scaffold** — create directories and write all files with realistic content
6766

6867
## Agent Analysis
6968

‎docs/build-skill-eval-scaffold.md‎

Lines changed: 7 additions & 8 deletions
Original file line numberDiff line numberDiff line change
@@ -55,15 +55,14 @@ The following are excluded from scaffolds:
5555

5656
## Workflow
5757

58-
The skill walks through a 7-step process:
58+
The skill walks through a 6-step process with three interview pauses:
5959

6060
1. **Parse arguments** — extract the `plugin:skill` identifier and optional project description
61-
2. **Analyze target skill** — read the skill's SKILL.md, reference files, and agent definitions to understand what inputs, signals, and environment the skill expects
62-
3. **Present analysis summary** — show skill purpose, expected inputs, signal categories, environment requirements, and any existing scaffolds
63-
4. **Interview: Technology and shape** — confirm the tech stack and derive a kebab-case scaffold name with `-project` suffix
64-
5. **Interview: Signals to plant** — suggest specific signals based on the skill analysis; the user approves, removes, modifies, or adds signals
65-
6. **Interview: File plan** — present a complete file plan with paths, descriptions, and signal assignments for each file
66-
7. **Generate scaffold** — create directories and write all files with realistic content
61+
2. **Analyze target skill** — read the skill's SKILL.md, reference files, and dispatched agent definitions to understand what inputs, signals, and environment the skill expects
62+
3. **Interview: Analysis and project shape** — present the skill's purpose, expected inputs, signal categories, environment requirements, and any existing scaffolds alongside the proposed tech stack and a kebab-case scaffold name with `-project` suffix; the user confirms or corrects all of it in one reply
63+
4. **Interview: Signals to plant** — suggest specific signals based on the skill analysis; the user approves, removes, modifies, or adds signals
64+
5. **Interview: File plan** — present a complete file plan with paths, descriptions, and signal assignments for each file
65+
6. **Generate scaffold** — create directories and write all files with realistic content
6766

6867
## Skill Analysis
6968

@@ -83,7 +82,7 @@ Files under `{plugin}/skills/{skill}/references/` contain templates, checklists,
8382

8483
### Agent definitions
8584

86-
Agent definitions referenced by the skill (via `subagent_type` in `Agent` tool calls) describe specific analysis focuses — structural coupling, security vulnerabilities, concurrency patterns — that inform what signals should be planted.
85+
Agent definitions referenced by the skill (via `subagent_type` in `Agent` tool calls) describe specific analysis focuses — structural coupling, security vulnerabilities, concurrency patterns — that inform what signals should be planted. `subagent_type` values are namespaced `plugin:agent`, and the agent's plugin is often not the skill's own, so the skill resolves each one to `{agent-plugin}/agents/{agent}.md` at the repository root.
8786

8887
### Graceful skip
8988

Lines changed: 15 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,15 @@
1+
{
2+
"$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json",
3+
"name": "eval-authoring",
4+
"displayName": "Eval Authoring",
5+
"version": "1.0.0",
6+
"description": "Skills that author eval suites for the skills-test-harness. Home of build-skill-eval-scaffold and build-agent-eval-scaffold, which interview you and generate realistic project scaffolds for rubric (effectiveness) evaluation of a Claude Code skill or agent. Depends on nothing; run it from the harness repo where the target plugin:skill or plugin:agent lives.",
7+
"author": {
8+
"name": "Test Double",
9+
"url": "https://github.com/testdouble"
10+
},
11+
"homepage": "https://github.com/testdouble/skills-test-harness#eval-authoring-skills",
12+
"repository": "https://github.com/testdouble/skills-test-harness",
13+
"license": "MIT",
14+
"keywords": ["evals", "testing", "scaffold", "rubric", "skills", "agents", "claude-code"]
15+
}
Lines changed: 104 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,104 @@
1+
---
2+
name: build-agent-eval-scaffold
3+
description: "Creates a realistic project scaffold — a test fixture — for rubric-based effectiveness evaluation of a Claude Code agent. Given a plugin:agent identifier and an optional project description, analyzes the target agent's definition file to learn what inputs it expects and what signals it looks for, then interviews the user in three phases (technology and shape, signals to plant, file plan) before writing the scaffold to tests/test-suites/{agent}/scaffolds/{name}/. Use when creating, building, or setting up a test scaffold, fixture project, or sample codebase to evaluate an agent's output quality against. Does not create tests.json entries or rubric files — use write-acil-evals or write-agent-eval-rubric for those. Does not create skill scaffolds — use build-skill-eval-scaffold for skills."
4+
argument-hint: "[plugin:agent] [optional description] e.g. example-plugin:gap-analyzer for a rails 7 project with postgres"
5+
allowed-tools: Read, Write, Glob, Grep, Bash(mkdir *)
6+
---
7+
8+
Build a realistic project fixture that the test harness runs the target agent against, so an LLM judge can score the agent's output against a rubric. The fixture only works as an eval when it contains signals the target agent is designed to find and reads like code a real developer wrote, so this workflow analyzes the target agent first, then interviews the user in three phases before writing any files.
9+
10+
## Constraints
11+
12+
These apply to every scaffold and shape both the file plan (Step 5) and generation (Step 6):
13+
14+
- Always write the scaffold to `tests/test-suites/{agent}/scaffolds/{name}/`, relative to the repository root, BECAUSE the harness discovers scaffolds at that path when it builds the Test Sandbox.
15+
- Never write a `.git` directory BECAUSE the harness auto-initializes a git repo with `git init` and commits all scaffold files itself.
16+
- Never write lock files (`package-lock.json`, `Gemfile.lock`, `go.sum`) or dependency directories (`node_modules`, `vendor`, `__pycache__`) unless one is itself a planted signal BECAUSE they add hundreds of generated lines the target agent never inspects and bury the signals that matter.
17+
- Never mark a signal with comments like `BUG HERE` or `INTENTIONAL ISSUE` BECAUSE a signal the agent can find by reading a comment measures nothing about its analysis; signals must require the same work a real codebase would.
18+
- Always keep every file syntactically valid for its language (it should parse or compile apart from intentional logic bugs) BECAUSE a file that fails to parse makes the agent report the syntax error instead of the planted signal.
19+
- Keep interview turns compact: present the list, then the question. Put explanation in the "Why" line of each signal rather than in surrounding prose.
20+
21+
## Step 1: Parse arguments
22+
23+
Parse the user's input into two parts:
24+
25+
- **`plugin:agent`** (required) — the first token, which contains a colon. Split on the colon to extract the plugin name (before the colon) and the agent name (after the colon).
26+
- **Description** (optional) — everything after the first token. This describes the technology and shape of the scaffold project (e.g., "for a rails 7 project with postgres as the database").
27+
28+
If no argument was provided, or the first token contains no colon, ask the user which `plugin:agent` to create a scaffold for and stop until they answer.
29+
30+
Validate that the agent exists by confirming the file `{plugin}/agents/{agent}.md` exists in the repository root. If it does not exist, tell the user and ask them to correct the input.
31+
32+
## Step 2: Analyze target agent
33+
34+
Read the agent's definition file at `{plugin}/agents/{agent}.md`. Agents are self-contained markdown files with YAML frontmatter and a prompt body. They do not have `references/` directories, `scripts/`, or context injection commands.
35+
36+
Analyze the full agent definition to identify:
37+
38+
- What inputs and environment the agent expects (source code files, config files, documentation, specific project structure)
39+
- What outputs the agent produces (analysis reports, gap assessments, architectural reviews, etc.)
40+
- What signals the agent looks for (bugs, security flaws, architectural patterns, missing documentation, implementation gaps)
41+
- What tools the agent uses (Read, Glob, Grep, Bash commands — these reveal what file types and patterns the agent inspects)
42+
- Whether the agent operates on project files at all, or whether it operates on external state (GitHub PRs, CI pipelines, conversation context)
43+
44+
**Graceful skip:** If the analysis reveals the agent does not operate on project files — for example, it queries GitHub APIs (`gh` commands), operates on pull request state, or generates content from conversation context rather than analyzing files in a working directory — inform the user that a file scaffold would not be useful for this agent type and stop. Explain what the agent operates on instead.
45+
46+
## Step 3: Interview Phase 1 — Analysis and project shape
47+
48+
Present the following in one message:
49+
50+
1. **Agent purpose** — a one-sentence summary of what the agent does
51+
2. **Expected inputs** — what file types, config files, and project structure the agent expects to find in a working directory
52+
3. **Signal categories** — the kinds of issues, patterns, or signals the agent detects (grouped into categories)
53+
4. **Environment requirements** — any specific project conventions the scaffold needs to follow (e.g., needs a Gemfile for Ruby, needs a go.mod for Go)
54+
5. **Existing scaffolds** — use Glob on `tests/test-suites/{agent}/scaffolds/*/` and list any found by name so the user can avoid duplicating one; if the suite directory does not exist yet, note that it will be created
55+
6. **Proposed project** — if a description was provided in the arguments, restate the tech stack and project shape it implies; otherwise propose one or two options that fit the agent's expected environment from Step 2
56+
7. **Proposed scaffold name** — derive it from the description: kebab-case with a `-project` suffix (e.g., "rails 7 with postgres" becomes `rails-postgres-project`, "python flask API" becomes `python-flask-project`). If `tests/test-suites/{agent}/scaffolds/{name}/` already exists, say so and ask whether to overwrite it or choose a different name.
57+
58+
Ask the user to confirm or correct the analysis, the tech stack, and the scaffold name in one reply. Wait for their answer, and carry any corrections into the remaining steps.
59+
60+
## Step 4: Interview Phase 2 — Signals to plant
61+
62+
Based on the Step 2 analysis, suggest specific signals to plant in the scaffold. Present them as a numbered list where each entry includes:
63+
64+
- **What** — the signal itself (e.g., "SQL injection via string interpolation in a database query")
65+
- **Where** — where it would live in the scaffold (e.g., "in a database access layer module")
66+
- **Why** — why this matters for the target agent (e.g., "the agent specifically checks for parameterized queries vs. string interpolation")
67+
68+
Draw suggestions from the agent definition — what the agent's instructions explicitly check for and what analysis patterns it follows.
69+
70+
Present the list and ask the user to:
71+
72+
- Approve the list as-is
73+
- Remove signals that are not relevant to their testing goals
74+
- Modify signals to be more specific to their project context
75+
- Add signals not covered by the analysis
76+
77+
Wait for the user to finalize the signal list before proceeding.
78+
79+
## Step 5: Interview Phase 3 — File plan
80+
81+
Present a complete file plan as a structured list. For each file include:
82+
83+
- **Path** — relative to the scaffold root directory
84+
- **Description** — what the file contains and its role in the project
85+
- **Signals** — which signals from Step 4 this file carries (reference by number), or "none" for clean files
86+
87+
Design the file plan against these guidelines:
88+
89+
- Include the standard config files for the tech stack (package.json, Gemfile, go.mod, requirements.txt, etc.) BECAUSE the agent uses them to detect the language and framework.
90+
- Keep the plan focused — enough structure to feel like a real project, but only files the agent would actually inspect or that establish necessary context — BECAUSE every extra file costs the eval run tokens without adding signal.
91+
- Plan source files at roughly 50-150 lines each BECAUSE shorter files read as stubs and longer ones hide signals in bulk.
92+
- Spread signals across multiple source files in a realistic directory structure, and include some files with no intentional signals, BECAUSE real projects mix clean and problematic code; an agent shown only problem files is never tested on telling the two apart.
93+
94+
Present the file plan and wait for the user to approve it. The user may request additions, removals, or modifications to the file plan.
95+
96+
## Step 6: Generate scaffold
97+
98+
After the user approves the file plan:
99+
100+
1. Create the directory structure using `mkdir -p tests/test-suites/{agent}/scaffolds/{name}/` including any subdirectories needed for the planned files.
101+
102+
2. Write each file using the Write tool. For each file, generate realistic content that matches the tech stack and project context, and include its planned signals the way a real developer would have written them — the Constraints above govern what the content may and may not contain.
103+
104+
3. Report the outcome: the scaffold path first, then the complete list of files created with their paths relative to the repository root.

0 commit comments

Comments
 (0)