Repository navigation
Release 0.17.0 - #18
Merged
Merged
Conversation
…name org plugins From claude-code-internals v2.62.0/v2.62.1 (Desktop 2.19675.0), reviewed by that project before applying. - SKILL.md: the outputs directory allows deletes from Desktop 2.16120.0 outside bridge sessions (code-read, measured live); a connected folder in a local session still refuses them until approved. Keep the rule, restate the reason, "can fail" rather than "fails". - Stanza B: the path `find` prints is for shell commands only; in a local session (default host loop) it is a VM path the file tools refuse. A model reused such a path even with the host path in the skill text (L122 "printed-root trap"). - A typed slash command can be refused with "Unknown skill: <name>." unless picked from the / menu (measured on macOS Desktop; upstream anthropics/claude-code#94309). - Same-name organization plugin: the / menu shows two identical entries, and in local sessions both loaded the organization's copy (n=2, Team); switching the org copy off for the user made the private one load (2 local, 1 cloud). Upstream anthropics/claude-code#99174.
…ema and help nits - quick_validate rejects a skill that would package more than one SKILL.md (claude.ai and the Skills API reject the upload; adopted from the skill-creator built into claude.ai). Exclusions are shared with package_skill through a new scripts/_packaging_rules.py, so a SKILL.md under tests/, evals/, caches or behind a symlink doesn't count. - grader.md: write grading.json to the run directory, matching Step 8; a missing path is reported, not asked for (one-shot sub-agents can't ask); drop the stale character-count fallback reason. Same for analyzer.md. - comparator.md: the rubric is drafted from the task before forming a view of either output, matching the step order. - description-optimization.md: export a credential first, check isolated/canary, exit 4 means nothing measured; isolation hides competing skills. - schemas.md: +50% in the delta example, the skill's name, old_skill/ new_skill in run_summary. check_portability docstring: --strict vs --strict-advisories. run_loop epilog: exit 4.
SKILL.md was 47,992 chars, about 2.4x the cap it teaches, so after
compaction the eval loop, validate/package and environment routing were
cut. It is now 18,146.
- New references/running-evals.md (Steps 1-5, viewer, feedback,
improving, iteration, blind comparison, test-case detail) and
references/validate-and-package.md (checklist, validators, harness
static checks, packaging, .skill layout, delivery), each with a TOC.
Text moved, not reworded, apart from the fixes below.
- SKILL.md keeps: routing (moved up), a new "Answering a question about
skills" section, the workspace paragraph with every pinned fact, a
five-step eval skeleton with the same-turn spawn and timing rules
inline, and a validate/package/deliver stub with the one-pass commands,
two-step delivery and the save-skill sentence (plus "shown only when
the user's org allows skill creation").
- References never use ${CLAUDE_SKILL_DIR} (it arrives literally there):
SKILL.md resolves <this-skill-dir> once, with the find fallback, and
every reference command uses it. This also fixes the Step 4 commands
that ran `python -m scripts.*` with no cd.
- Static-mode feedback is marked server-only where it described
feedback.json; the viewer goes in front of the user before your own
analysis in every environment; templates use <eval-name>.
- environments.md: headings only (a TOC and seven H3s under Sandboxed
sessions), no text reworded.
- Tests: whole-file size asserts, pins for routing and the running-evals
gate, pointer-resolution tests (paths exist, section pointers resolve,
no CLAUDE_SKILL_DIR in references/agents), baseline rule set now 3
(compaction-truncation-risk gone).
Findings from comparing with Anthropic's skill-creator and its review
of this skill; plan reviewed adversarially before implementation.
Eight Q&A scenarios for cowork-harness eval: script paths, workspace location, deliver and save, SKILL.md size, inline commands, setup and arguments, viewer without a display, and the eval loop. Each claim is about what the agent says (the judge sees no tool calls) and holds for both the pre-restructure and restructured content. Not part of CI replay.
… eval-surfaced wording - official-guide-patterns.md was 88,635 bytes, over the agent's 60,000-byte whole-file read cap (a Read returns a partial view and must page; flagged by cowork-harness 4.2.1 lint-skill). "Advanced Skill Authoring Features" and "Runtime Mechanics & Gotchas" move verbatim to a new references/advanced-features.md (33 KB, with contents); the rest is 56.6 KB. Cross-file pointers updated; the pointer tests cover SKILL.md. - claude-code-internals review of the restructured SKILL.md: the "directory does not exist" fallback is attributed to older local Desktop builds (current ones rewrite the path), also in validate-and-package.md and stanza B; the bin/ launcher note says the cloud case is untested; the read-only skill directory is scoped to sandboxed sessions (in Claude Code it is writable but replaced on update); workspace trust is scoped to project skills. - The eval skeleton names the benchmark again (pass rate, time, tokens): the 0.17.0 A/B showed answers dropping that detail (5/5 -> 2/5), and after this fix it was 5/5.
CI installs and asserts cowork-harness 4.2.1 (from 3.10.0); docs and the harness README follow. eval-scenarios/invoked/ holds five scenarios that tell the model to use the skill: in the A/B, Sonnet answered the plain versions without invoking it, so they compared nothing about the text.
…rence official-guide-patterns.md was titled "Official Skill-Building Patterns" and cited only Anthropic's guide and Thariq's lessons, but also carried this project's measurements and code reading (Declare at authoring time, what Claude-specific frontmatter does, the body skeleton, Designing Scripts for Agent Use, SKILL.md Size, the Claude-specific description addenda, the listing-overflow troubleshooting entry). A model read those as Anthropic's documented guidance. Retitle it, add a provenance note naming those sections, and describe it accordingly in SKILL.md. Moving them out is backlog item 4.
Bump the version, date the changelog, and re-record the no-trigger cassette under cowork-harness 4.2.1 (baseline desktop-2.16120.0).
Upstream's validator, optimizer and taxonomy rows were out of date; the patterns reference is no longer described as purely Anthropic's; the chat-runtime and claude CLI claims carry the hedges the skill's own docs use.
- environments.md: cowork-harness floor is 4.2.1 (4.0.0 refuses an unpinned model); drop a pointer to README text that no longer exists - running-evals.md: drop the claim that a user-saved feedback.json can't be read - analyzer.md, advanced-features.md: cd to the skill dir before python -m - schemas.md: benchmark.json and comparison.json paths match the scripts; improve mode is with_skill/old_skill; drop history.json, which nothing writes - references are not compaction-capped, but are read-capped - official-guide-patterns.md: provenance list covers four more passages; fix pointers left by the advanced-features split; cwd carryover is unreliable, not absent; drop a dead anchor - exit 4 also covers a run hijacked by an installed copy - validate-and-package.md: list all 14 portability rules - harness docs: one cassette is committed; one scenario is hostloop; the hash ignores tests/ and eval-viewer/
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Release 0.17.0. See the
[0.17.0]section of CHANGELOG.md for the full list.Highlights:
references/running-evals.mdandreferences/validate-and-package.md, with routing near the top and a path for answering questions about skills.official-guide-patterns.mdsplit under the agent's whole-file read cap (advanced-features.md), and labelled for which parts are Anthropic's guidance and which are this project's additions.SKILL.mdrejected at validation, grader/analyzer save paths, credential and exit-4 guidance for the description optimizer, schema corrections.no-triggercassette re-recorded.Verified locally: 190 tests,
quick_validate,check_portability --target all,lint-skill --strict,analyze-skill --strict, scenario lint,record --dry-run,verify-cassettes.