Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 3 additions & 3 deletions .github/workflows/harness.yml
Original file line number Diff line number Diff line change
Expand Up @@ -18,19 +18,19 @@ jobs:

- uses: actions/setup-node@v4
with:
node-version: '22' # cowork-harness >= 3.10.0 declares engines node >= 22
node-version: '22' # cowork-harness >= 4.2.1 declares engines node >= 22

- name: Install cowork-harness (pinned)
run: |
npm i --prefix "$RUNNER_TEMP/ch" cowork-harness@3.10.0
npm i --prefix "$RUNNER_TEMP/ch" cowork-harness@4.2.1
echo "$RUNNER_TEMP/ch/node_modules/.bin" >> "$GITHUB_PATH"

- name: Verify cowork-harness version
run: |
V=$(cowork-harness --version | tail -1)
echo "cowork-harness $V"
# The pin lives TWICE in this job — the npm install above and this assert. Move both.
case "$V" in 3.10.*) ;; *) echo "expected 3.10.x, got $V"; exit 1;; esac
case "$V" in 4.2.*) ;; *) echo "expected 4.2.x, got $V"; exit 1;; esac

- name: Lint skill (strict)
run: |
Expand Down
78 changes: 78 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,84 @@

All notable changes to this project will be documented in this file.

## [0.17.0] - 2026-10-03

### Changed
- **SKILL.md now fits under its own compaction cap.** It was 47,992 characters, about 2.4× the
19,900-character limit it teaches, so after compaction the eval loop, packaging and environment
routing were cut. It is now about 18,470. The eval/improve loop moved to
`references/running-evals.md` and validation and packaging to `references/validate-and-package.md`,
both with tables of contents and mostly moved verbatim. SKILL.md keeps environment routing (moved
near the top), the workspace rules, a five-step eval skeleton with its two can't-miss rules (spawn
with-skill and baseline runs in the same turn; save timing from each task notification), and a
validate, package and deliver section with the commands the one-pass path needs.
- **Commands in references use `<this-skill-dir>`.** `${CLAUDE_SKILL_DIR}` arrives literally in a
reference file, so SKILL.md resolves the skill directory once (the token or a `find` fallback) and
the references refer to it. This also fixes the eval-loop commands that ran `python -m scripts.*`
without changing to the skill's directory.
- **`references/environments.md` has a table of contents and subsections** under *Sandboxed
sessions*; no text was reworded.
- **`official-guide-patterns.md` is split under the agent's read cap.** At 88,635 bytes it was over the
60,000-byte limit for reading a file whole, so the model saw a partial view. *Advanced Skill Authoring
Features* and *Runtime Mechanics & Gotchas* moved verbatim to `references/advanced-features.md`.
- **That reference no longer presents this project's findings as Anthropic's.** It was titled "Official
Skill-Building Patterns" while also carrying this project's measurements and code reading. It is
retitled "Anthropic's Guidance and This Project's Additions", and a provenance note lists the sections
that are ours, to be read as true of the version they name rather than as documented behaviour.

### Added
- **"Answering a question about skills."** A narrow question about skill mechanics now gets a direct
answer from the matching reference, not the build loop.
- **`quick_validate` rejects a skill that would package more than one `SKILL.md`.** claude.ai and the
Skills API reject such uploads (from the skill-creator built into claude.ai). A `SKILL.md` under
`tests/`, `evals/`, a cache or behind a symlink doesn't count; the exclusions are shared with
`package_skill`.
- **Description optimization says to export a credential first**, to check `isolated` and `canary`,
that exit 4 means nothing was measured, and that isolation leaves out the user's other skills, so
include near-miss queries a neighbouring skill should win.
- **Pick slash commands from the `/` menu in the Claude app.** One typed in full can be refused with
"Unknown skill: <name>." before anything is sent (measured on macOS Desktop; anthropics/claude-code
#94309 on Windows).
- **Same-name organization plugins.** When updating a plugin the organization also ships under the
same name, the `/` menu shows two identical entries, and in local sessions both loaded the
organization's copy; switching the organization's copy off for the user under Customize made theirs
load.

### Fixed
- **The grader's save location contradicted itself.** It now writes `grading.json` to the run
directory, where the benchmark reads it; a missing path is reported instead of asked for. The
comparator drafts its rubric before forming a view of either output, matching its step order.
- **Static-mode feedback** is no longer described as auto-saving to `feedback.json` (that is server
mode); and the eval viewer goes in front of the user before your own analysis in every environment,
not only in sandboxed sessions.
- **The reason not to delete from the outputs directory.** From Desktop 2.16120.0 the outputs
directory allows deletes outside bridge sessions; a connected folder in a local session still
refuses them until approved. The rule (build once, overwrite in place) stays.
- **A path the shell printed is for the shell.** Stanza B now says to use the path `find` prints only
in shell commands: in a local session (the default host loop) it is the VM's path, which the file
tools refuse.
- **Older-build and scope wording.** The "directory does not exist" fallback is attributed to older
local Desktop builds (current ones rewrite the path); the `bin/` launcher note says the cloud case is
untested; the read-only skill directory is scoped to sandboxed sessions (in Claude Code it is
writable but replaced on update); workspace trust is scoped to project skills.
- **Small corrections:** the `schemas.md` example delta (+50%), `run_summary`'s improve-mode keys,
`check_portability`'s `--strict` description, `run_loop`'s exit 4, a stale "falls back to a
character count" reason in the grader, and the Save skill button shown only when the user's
organization allows skill creation.

### Development
- **cowork-harness 4.2.1** in CI (from 3.10.0); the cassette is re-recorded against its
desktop-2.16120.0 baseline.
- **`harness/eval-scenarios/`**: eight Q&A scenarios for `cowork-harness eval`, plus `invoked/` variants
that tell the model to use the skill. Used to A/B this restructure against the previous text
(Sonnet 5.5 agent and judge, 5 reps per arm, about 150 runs): with the skill invoked, no claim
dropped, including after the reference split, and one wording gap the A/B surfaced (the eval
skeleton no longer mentioned the benchmark) was fixed and re-checked. Without the instruction to use it, Sonnet invoked the skill on only about a third of
these questions in a Cowork-like session, a triggering gap for a later release.
- **Tests:** SKILL.md size asserts, pins for routing and the eval-loop pointer, pointer-resolution
tests (named files exist, section pointers resolve, no `${CLAUDE_SKILL_DIR}` in references or agent
prompts), and the shipped lint baseline drops to three rules.

## [0.16.2] - 2026-10-03

### Fixed
Expand Down
18 changes: 9 additions & 9 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,25 +13,25 @@ The skill that builds skills. Draft one and ship it in a single pass, or run eva

| | Official `skill-creator` | `skill-creator-plus` |
|---|---|---|
| Best practices guide | — | 620+ line Anthropic patterns reference |
| Best practices guide | — | Patterns reference from Anthropic's guidance, plus this project's measured additions (labelled) |
| Script vs. Instruct guidance | — | Decision framework for when to bundle scripts vs. use instructions |
| Structure validation | — | Frontmatter, naming and length caps against the [agentskills.io](https://agentskills.io/specification) spec |
| Structure validation | Allowed keys, name format, length caps | Also type-checks Claude-specific fields, matches the name to the folder, and checks the 1,536-char listing cap, against the [agentskills.io](https://agentskills.io/specification) spec |
| Cross-runtime linting | — | 14 rules for what breaks outside Claude Code — sub-agent, `claude` CLI and browser dependencies, third-party imports, file delivery, compaction |
| Cloud and local session guidance | — | Runtime-specific guidance for the Claude app's cloud and local sessions ([below](#authoring-for-cowork)) |
| Claude Code runtime docs | — | Listing budget + per-skill degradation order, permission semantics, truncation caps, live-reload behavior — verified against 2.1.222–2.1.251 (listing budget re-verified at 2.1.280) |
| Eval viewer in cloud and local sessions | Silent fail on submit | Copyable JSON textarea (fixed) |
| Description optimizer | Requires separate `ANTHROPIC_API_KEY`; drifts toward 1,024-char bloat | Uses your existing `claude` session; length-aware selection + plateau early-stop |
| Description optimizer | Runs `claude -p`; keeps the top score, no length preference | Runs `claude -p`, isolated from your installed skills when a credential is exported; length-aware selection + plateau early-stop; refuses to score a run it couldn't measure |
| Benchmarking script | Silent empty results | Fixed directory handling |
| Skill type taxonomy | 3 broad categories | 3 + 9 Anthropic internal types |
| Skill type taxonomy | — | 3 categories from Anthropic's guide + 9 types used inside Anthropic |
| Pre-packaging checklist | — | Official checklist built in |

## Why This Over the Built-in?

Anthropic ships a `skill-creator` plugin. It's good, but several parts are broken or missing:

- **Best practices guide included** — 620+ lines of patterns, structural templates, troubleshooting guide, and checklists extracted from Anthropic's [Complete Guide to Building Skills for Claude](https://resources.anthropic.com/hubfs/The-Complete-Guide-to-Building-Skill-for-Claude.pdf) and Thariq's [Lessons from Building Claude Code Skills](https://x.com/trq212/status/2024574133011673516). Includes a "Script vs. Instruct" decision framework for when to bundle pre-made scripts vs. keep logic as instructions — covering context window efficiency, reliability, and auditability. The built-in doesn't ship any of this.
- **A skill that works here can break there, silently** — the `claude` CLI is on PATH in Claude Code and was found in the Claude app's cloud and local sessions, but not in its chat runtime (Chat in the older Chat/Cowork picker), so gate a `claude -p` step on `command -v claude`; the Claude app's chat runtime has no sub-agent tool, so parallel eval runs have to collapse to serial there (a Claude app conversation running in a cloud session does have one — route by the tool list, not the product); an agent in a cloud or local session can't serve a local HTTP server and open it, so the eval viewer needs a static build; a third-party import there costs an install on every run, and egress is org-configurable, so a locked-down org can refuse it; and file delivery differs per surface — in a cloud session, writing a file is not delivering it. Structure validation can't see any of that, so there are two checks: `quick_validate` for structure, and a 14-rule portability linter for runtime assumptions — including the two ways compaction loses a skill. Both are stdlib-only, because they have to run inside the sandboxes they lint. Rule ids, flags and exit codes are in [`docs/DEVELOPMENT.md`](docs/DEVELOPMENT.md).
- **Claude Code runtime docs** — the mechanics most skill authors hit the hard way, read out of the shipping binary rather than inherited from a blog post: overflow of the shared listing budget drops descriptions **per skill, least-recently-used first**, packing first-fit, so full and name-only entries coexist and a long description can lose to a shorter one; `allowed-tools` **grants** permission rather than requesting it; and compaction is a CHARACTER gate, not the documented token one, losing content two different ways — truncation keeps the first 19,900 characters and leaves a marker, while the combined cross-skill cap **zeroes** a skill outright with no marker and no entry. Verified against Claude Code 2.1.222–2.1.251; the listing budget re-verified at 2.1.280. Where the public docs and the binary disagree, [the reference](skill-creator-plus/skills/skill-creator-plus/references/official-guide-patterns.md) says so and shows which one shipped.
- **Best practices guide included** — patterns, structural templates, troubleshooting guide, and checklists drawn from Anthropic's [Complete Guide to Building Skills for Claude](https://resources.anthropic.com/hubfs/The-Complete-Guide-to-Building-Skill-for-Claude.pdf) and Thariq's [Lessons from Building Claude Code Skills](https://x.com/trq212/status/2024574133011673516). Includes a "Script vs. Instruct" decision framework for when to bundle pre-made scripts vs. keep logic as instructions — covering context window efficiency, reliability, and auditability. This project's own measured runtime additions sit alongside, labelled as such. The built-in doesn't ship any of this.
- **A skill that works here can break there, silently** — the `claude` CLI is on PATH in Claude Code and was found in the Claude app's cloud and local sessions, but not in its chat runtime (Chat in the older Chat/Cowork picker; one check each), so gate a `claude -p` step on `command -v claude`; the Claude app's chat runtime typically has no sub-agent tool, so parallel eval runs have to collapse to serial there (a Claude app conversation running in a cloud session does have one — route by the tool list, not the product); an agent in a cloud or local session can't serve a local HTTP server and open it, so the eval viewer needs a static build; a third-party import there costs an install on every run, and egress is org-configurable, so a locked-down org can refuse it; and file delivery differs per surface — in a cloud session, writing a file is not delivering it. Structure validation can't see any of that, so there are two checks: `quick_validate` for structure, and a 14-rule portability linter for runtime assumptions — including the two ways compaction loses a skill. Both are stdlib-only, because they have to run inside the sandboxes they lint. Rule ids, flags and exit codes are in [`docs/DEVELOPMENT.md`](docs/DEVELOPMENT.md).
- **Claude Code runtime docs** — the mechanics most skill authors hit the hard way, read out of the shipping binary rather than inherited from a blog post: overflow of the shared listing budget drops descriptions **per skill, least-recently-used first**, packing first-fit, so full and name-only entries coexist and a long description can lose to a shorter one; `allowed-tools` **grants** permission rather than requesting it; and compaction is a CHARACTER gate, not the documented token one, losing content two different ways — truncation keeps the first 19,900 characters and leaves a marker, while the combined cross-skill cap **zeroes** a skill outright with no marker and no entry. Verified against Claude Code 2.1.222–2.1.251; the listing budget re-verified at 2.1.280. Where the public docs and the binary disagree, the references ([patterns](skill-creator-plus/skills/skill-creator-plus/references/official-guide-patterns.md), [advanced features](skill-creator-plus/skills/skill-creator-plus/references/advanced-features.md)) say so and show which one shipped.

The eval viewer, description optimizer and benchmarking fixes are in the table above; see the [CHANGELOG](CHANGELOG.md) for the full list.

Expand All @@ -54,7 +54,7 @@ The skill takes it from there — intent capture, drafting, test cases, and deli

You can also just ask it a question. Mechanics are in scope on their own — frontmatter fields, path variables, size limits, directory layout, or what breaks across runtimes — however small the question.

> **Note:** If you also have Anthropic's built-in `skill-creator` installed, Claude may pick that one instead. Either uninstall the built-in, or use `/skill-creator-plus:skill-creator-plus` to invoke this version explicitly.
> **Note:** If Anthropic's `skill-creator` is also available (installed in Claude Code, or built into your claude.ai account), Claude may pick that one instead. Remove or turn off the other one where your app allows it, or in Claude Code use `/skill-creator-plus:skill-creator-plus` to invoke it explicitly.

## How It Works

Expand Down Expand Up @@ -88,7 +88,7 @@ Then repeat. Order is flexible, and an existing draft can join at step 1.
<a id="authoring-for-cowork"></a>
## Authoring for cloud and local sessions

Cloud and local sessions break assumptions that hold everywhere else, and it breaks them *quietly* — the write succeeds, the tool reports success, and the file is somewhere nobody will look. The skill knows about:
Cloud and local sessions break assumptions that hold everywhere else, and they break them *quietly* — the write succeeds, the tool reports success, and the file is somewhere nobody will look. The skill knows about:

- **Where the workspace must live.** The skill directory is a read-only plugin mount, so the workspace can't sit beside it. The agent falls back to the session scratchpad — which in a cloud session is reclaimed at session end, destroying the skill it just built.
- **File tools and the shell don't share a working directory.** No single relative path is correct for both, so the identifier and the base have to stay apart.
Expand Down
2 changes: 1 addition & 1 deletion VERSION
Original file line number Diff line number Diff line change
@@ -1 +1 @@
0.16.2
0.17.0
Loading
Loading