Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 18 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,24 @@

All notable changes to this project will be documented in this file.

## [0.16.2] - 2026-10-03

### Fixed
- **The eval viewer's Download button failed for text outputs on claude.ai.** Opened from a
delivered file on the web and in Claude Desktop, it raised an `atob` error and saved nothing, while
binary outputs downloaded after the app's trust prompt. Text downloads now use a base64 `data:` URI
like the others (not yet re-tested in the app).
- **What the viewer's downloads do in the Claude app.** `environments.md` now says the Download links
don't work when the viewer is published as an Artifact (the app blocks file downloads there; the
outputs still show in the page), and that from a delivered viewer, binary downloads worked after a
confirmation.

### Changed
- **Where the `claude` CLI is.** Measured once each: found on PATH in cloud and local sessions (a
nested `claude -p` ran in a cloud session), and not on the chat runtime's PATH (Chat in the older
Chat/Cowork picker). The `claude-cli-dependency` message, `SKILL.md` and the README say so again,
now backed by a measurement; the advice to gate on `command -v claude` is unchanged.

## [0.16.1] - 2026-10-03

### Fixed
Expand Down
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,7 +30,7 @@ The skill that builds skills. Draft one and ship it in a single pass, or run eva
Anthropic ships a `skill-creator` plugin. It's good, but several parts are broken or missing:

- **Best practices guide included** — 620+ lines of patterns, structural templates, troubleshooting guide, and checklists extracted from Anthropic's [Complete Guide to Building Skills for Claude](https://resources.anthropic.com/hubfs/The-Complete-Guide-to-Building-Skill-for-Claude.pdf) and Thariq's [Lessons from Building Claude Code Skills](https://x.com/trq212/status/2024574133011673516). Includes a "Script vs. Instruct" decision framework for when to bundle pre-made scripts vs. keep logic as instructions — covering context window efficiency, reliability, and auditability. The built-in doesn't ship any of this.
- **A skill that works here can break there, silently** — the `claude` CLI is on PATH in Claude Code and was found in a local session's shell, but other runtimes are unchecked, so gate a `claude -p` step on `command -v claude`; the Claude app's chat runtime has no sub-agent tool, so parallel eval runs have to collapse to serial there (a Claude app conversation running in a cloud session does have one — route by the tool list, not the product); an agent in a cloud or local session can't serve a local HTTP server and open it, so the eval viewer needs a static build; a third-party import there costs an install on every run, and egress is org-configurable, so a locked-down org can refuse it; and file delivery differs per surface — in a cloud session, writing a file is not delivering it. Structure validation can't see any of that, so there are two checks: `quick_validate` for structure, and a 14-rule portability linter for runtime assumptions — including the two ways compaction loses a skill. Both are stdlib-only, because they have to run inside the sandboxes they lint. Rule ids, flags and exit codes are in [`docs/DEVELOPMENT.md`](docs/DEVELOPMENT.md).
- **A skill that works here can break there, silently** — the `claude` CLI is on PATH in Claude Code and was found in the Claude app's cloud and local sessions, but not in its chat runtime (Chat in the older Chat/Cowork picker), so gate a `claude -p` step on `command -v claude`; the Claude app's chat runtime has no sub-agent tool, so parallel eval runs have to collapse to serial there (a Claude app conversation running in a cloud session does have one — route by the tool list, not the product); an agent in a cloud or local session can't serve a local HTTP server and open it, so the eval viewer needs a static build; a third-party import there costs an install on every run, and egress is org-configurable, so a locked-down org can refuse it; and file delivery differs per surface — in a cloud session, writing a file is not delivering it. Structure validation can't see any of that, so there are two checks: `quick_validate` for structure, and a 14-rule portability linter for runtime assumptions — including the two ways compaction loses a skill. Both are stdlib-only, because they have to run inside the sandboxes they lint. Rule ids, flags and exit codes are in [`docs/DEVELOPMENT.md`](docs/DEVELOPMENT.md).
- **Claude Code runtime docs** — the mechanics most skill authors hit the hard way, read out of the shipping binary rather than inherited from a blog post: overflow of the shared listing budget drops descriptions **per skill, least-recently-used first**, packing first-fit, so full and name-only entries coexist and a long description can lose to a shorter one; `allowed-tools` **grants** permission rather than requesting it; and compaction is a CHARACTER gate, not the documented token one, losing content two different ways — truncation keeps the first 19,900 characters and leaves a marker, while the combined cross-skill cap **zeroes** a skill outright with no marker and no entry. Verified against Claude Code 2.1.222–2.1.251; the listing budget re-verified at 2.1.280. Where the public docs and the binary disagree, [the reference](skill-creator-plus/skills/skill-creator-plus/references/official-guide-patterns.md) says so and shows which one shipped.

The eval viewer, description optimizer and benchmarking fixes are in the table above; see the [CHANGELOG](CHANGELOG.md) for the full list.
Expand Down
2 changes: 1 addition & 1 deletion VERSION
Original file line number Diff line number Diff line change
@@ -1 +1 @@
0.16.1
0.16.2
Loading
Loading