Skip to content

Release 0.17.0 - #18

Merged
yaniv-golan merged 11 commits into
mainfrom
release/0.17.0
Oct 3, 2026
Merged

yaniv-golan merged 11 commits into
mainfrom
release/0.17.0

Conversation

@yaniv-golan

Copy link
Copy Markdown
Owner

Release 0.17.0. See the [0.17.0] section of CHANGELOG.md for the full list.

Highlights:

  • SKILL.md restructured to 18,483 chars, under the 19,900-char post-compaction cap; the eval loop and validate/package steps moved to references/running-evals.md and references/validate-and-package.md, with routing near the top and a path for answering questions about skills.
  • official-guide-patterns.md split under the agent's whole-file read cap (advanced-features.md), and labelled for which parts are Anthropic's guidance and which are this project's additions.
  • Fixes from a comparison with Anthropic's skill-creator: nested SKILL.md rejected at validation, grader/analyzer save paths, credential and exit-4 guidance for the description optimizer, schema corrections.
  • A documentation correctness pass over the references, the maintainer docs and the README.
  • Harness: CI on cowork-harness 4.2.1; answer-quality eval scenarios; no-trigger cassette re-recorded.

Verified locally: 190 tests, quick_validate, check_portability --target all, lint-skill --strict, analyze-skill --strict, scenario lint, record --dry-run, verify-cassettes.

…name org plugins

From claude-code-internals v2.62.0/v2.62.1 (Desktop 2.19675.0), reviewed
by that project before applying.

- SKILL.md: the outputs directory allows deletes from Desktop 2.16120.0
  outside bridge sessions (code-read, measured live); a connected folder
  in a local session still refuses them until approved. Keep the rule,
  restate the reason, "can fail" rather than "fails".
- Stanza B: the path `find` prints is for shell commands only; in a local
  session (default host loop) it is a VM path the file tools refuse. A
  model reused such a path even with the host path in the skill text
  (L122 "printed-root trap").
- A typed slash command can be refused with "Unknown skill: <name>."
  unless picked from the / menu (measured on macOS Desktop; upstream
  anthropics/claude-code#94309).
- Same-name organization plugin: the / menu shows two identical entries,
  and in local sessions both loaded the organization's copy (n=2, Team);
  switching the org copy off for the user made the private one load (2
  local, 1 cloud). Upstream anthropics/claude-code#99174.
…ema and help nits

- quick_validate rejects a skill that would package more than one SKILL.md
  (claude.ai and the Skills API reject the upload; adopted from the
  skill-creator built into claude.ai). Exclusions are shared with
  package_skill through a new scripts/_packaging_rules.py, so a SKILL.md
  under tests/, evals/, caches or behind a symlink doesn't count.
- grader.md: write grading.json to the run directory, matching Step 8;
  a missing path is reported, not asked for (one-shot sub-agents can't
  ask); drop the stale character-count fallback reason. Same for
  analyzer.md.
- comparator.md: the rubric is drafted from the task before forming a
  view of either output, matching the step order.
- description-optimization.md: export a credential first, check
  isolated/canary, exit 4 means nothing measured; isolation hides
  competing skills.
- schemas.md: +50% in the delta example, the skill's name, old_skill/
  new_skill in run_summary. check_portability docstring: --strict vs
  --strict-advisories. run_loop epilog: exit 4.
SKILL.md was 47,992 chars, about 2.4x the cap it teaches, so after
compaction the eval loop, validate/package and environment routing were
cut. It is now 18,146.

- New references/running-evals.md (Steps 1-5, viewer, feedback,
  improving, iteration, blind comparison, test-case detail) and
  references/validate-and-package.md (checklist, validators, harness
  static checks, packaging, .skill layout, delivery), each with a TOC.
  Text moved, not reworded, apart from the fixes below.
- SKILL.md keeps: routing (moved up), a new "Answering a question about
  skills" section, the workspace paragraph with every pinned fact, a
  five-step eval skeleton with the same-turn spawn and timing rules
  inline, and a validate/package/deliver stub with the one-pass commands,
  two-step delivery and the save-skill sentence (plus "shown only when
  the user's org allows skill creation").
- References never use ${CLAUDE_SKILL_DIR} (it arrives literally there):
  SKILL.md resolves <this-skill-dir> once, with the find fallback, and
  every reference command uses it. This also fixes the Step 4 commands
  that ran `python -m scripts.*` with no cd.
- Static-mode feedback is marked server-only where it described
  feedback.json; the viewer goes in front of the user before your own
  analysis in every environment; templates use <eval-name>.
- environments.md: headings only (a TOC and seven H3s under Sandboxed
  sessions), no text reworded.
- Tests: whole-file size asserts, pins for routing and the running-evals
  gate, pointer-resolution tests (paths exist, section pointers resolve,
  no CLAUDE_SKILL_DIR in references/agents), baseline rule set now 3
  (compaction-truncation-risk gone).

Findings from comparing with Anthropic's skill-creator and its review
of this skill; plan reviewed adversarially before implementation.
Eight Q&A scenarios for cowork-harness eval: script paths, workspace
location, deliver and save, SKILL.md size, inline commands, setup and
arguments, viewer without a display, and the eval loop. Each claim is
about what the agent says (the judge sees no tool calls) and holds for
both the pre-restructure and restructured content. Not part of CI replay.
… eval-surfaced wording

- official-guide-patterns.md was 88,635 bytes, over the agent's 60,000-byte
  whole-file read cap (a Read returns a partial view and must page;
  flagged by cowork-harness 4.2.1 lint-skill). "Advanced Skill Authoring
  Features" and "Runtime Mechanics & Gotchas" move verbatim to a new
  references/advanced-features.md (33 KB, with contents); the rest is
  56.6 KB. Cross-file pointers updated; the pointer tests cover SKILL.md.
- claude-code-internals review of the restructured SKILL.md: the
  "directory does not exist" fallback is attributed to older local
  Desktop builds (current ones rewrite the path), also in
  validate-and-package.md and stanza B; the bin/ launcher note says the
  cloud case is untested; the read-only skill directory is scoped to
  sandboxed sessions (in Claude Code it is writable but replaced on
  update); workspace trust is scoped to project skills.
- The eval skeleton names the benchmark again (pass rate, time, tokens):
  the 0.17.0 A/B showed answers dropping that detail (5/5 -> 2/5), and
  after this fix it was 5/5.
CI installs and asserts cowork-harness 4.2.1 (from 3.10.0); docs and the
harness README follow. eval-scenarios/invoked/ holds five scenarios that
tell the model to use the skill: in the A/B, Sonnet answered the plain
versions without invoking it, so they compared nothing about the text.
…rence

official-guide-patterns.md was titled "Official Skill-Building Patterns"
and cited only Anthropic's guide and Thariq's lessons, but also carried
this project's measurements and code reading (Declare at authoring time,
what Claude-specific frontmatter does, the body skeleton, Designing
Scripts for Agent Use, SKILL.md Size, the Claude-specific description
addenda, the listing-overflow troubleshooting entry). A model read
those as Anthropic's documented guidance. Retitle it, add a provenance
note naming those sections, and describe it accordingly in SKILL.md.
Moving them out is backlog item 4.
Bump the version, date the changelog, and re-record the no-trigger
cassette under cowork-harness 4.2.1 (baseline desktop-2.16120.0).
Upstream's validator, optimizer and taxonomy rows were out of date; the
patterns reference is no longer described as purely Anthropic's; the
chat-runtime and claude CLI claims carry the hedges the skill's own docs use.
- environments.md: cowork-harness floor is 4.2.1 (4.0.0 refuses an unpinned
  model); drop a pointer to README text that no longer exists
- running-evals.md: drop the claim that a user-saved feedback.json can't be read
- analyzer.md, advanced-features.md: cd to the skill dir before python -m
- schemas.md: benchmark.json and comparison.json paths match the scripts;
  improve mode is with_skill/old_skill; drop history.json, which nothing writes
- references are not compaction-capped, but are read-capped
- official-guide-patterns.md: provenance list covers four more passages; fix
  pointers left by the advanced-features split; cwd carryover is unreliable,
  not absent; drop a dead anchor
- exit 4 also covers a run hijacked by an installed copy
- validate-and-package.md: list all 14 portability rules
- harness docs: one cassette is committed; one scenario is hostloop; the
  hash ignores tests/ and eval-viewer/
@yaniv-golan
yaniv-golan merged commit bf1aea9 into main Oct 3, 2026
2 checks passed
@yaniv-golan
yaniv-golan deleted the release/0.17.0 branch October 3, 2026 19:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant