Skip to content
Merged
Show file tree
Hide file tree
Changes from 2 commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 4 additions & 4 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,8 +30,8 @@
- Test: `make test` (Vitest, unit + integration)
- Build: `make build` (Bun compile + Vite 8)
- Dev server: `make dev`
- Packages: `packages/cli` (Yargs CLI), `packages/execution` (test-run, test-eval, SCIL/ACIL orchestration), `packages/data` (DuckDB), `packages/web` (Hono + React 18 + Tailwind v4), `packages/test-fixtures`, `packages/docker-integration` (Docker sandbox API)
- See [`docs/docker-integration.md`](docs/docker-integration.md) for Docker sandbox architecture, API reference, and consumer patterns
- Packages: `packages/cli` (Yargs CLI), `packages/execution` (test-run, test-eval, SCIL/ACIL orchestration), `packages/data` (DuckDB), `packages/web` (Hono + React 18 + Tailwind v4), `packages/test-fixtures`, `packages/sandbox-integration` (Test Sandbox API)
- See [`docs/sandbox-integration.md`](docs/sandbox-integration.md) for Test Sandbox architecture, API reference, and consumer patterns
- See [`docs/test-harness-architecture.md`](docs/test-harness-architecture.md) for system architecture, package boundaries, data flow, and dependency graph
- See [`docs/execution.md`](docs/execution.md) for the execution package: test-run pipeline, test-eval, SCIL/ACIL loops, error hierarchy, and path config
- See [`docs/cli.md`](docs/cli.md) for the CLI package: thin Yargs wrapper, command definitions, path resolution
Expand All @@ -41,14 +41,14 @@
- See [`docs/data.md`](docs/data.md) for the shared data layer: types, config parsing, JSONL I/O, DuckDB analytics, stream parsing, and SCIL/ACIL utilities
- See [`docs/web.md`](docs/web.md) for the web dashboard: Hono API server, React SPA, test run and SCIL views, per-test analytics
- See [`docs/evals.md`](docs/evals.md) for the evaluation engine: boolean evals, LLM judge scoring, rubric parsing, and the `evaluateTestRun` orchestrator
- See [`docs/docker-integration-package.md`](docs/docker-integration-package.md) for the Docker integration package deep-dive: full public API, error handling matrix, and testing patterns
- See [`docs/sandbox-integration-package.md`](docs/sandbox-integration-package.md) for the sandbox integration package deep-dive: full public API, error handling matrix, and testing patterns

### Guides and Configuration

- See [`docs/scil-evals-guide.md`](docs/scil-evals-guide.md) for building and running SCIL trigger accuracy evals
- See [`docs/rubric-evals-guide.md`](docs/rubric-evals-guide.md) for building and running LLM-judge quality evals
- See [`docs/test-suite-reference.md`](docs/test-suite-reference.md) for the full tests.json field reference
- See [`docs/test-scaffolding.md`](docs/test-scaffolding.md) for how scaffolds provide project context in the Docker sandbox
- See [`docs/test-scaffolding.md`](docs/test-scaffolding.md) for how scaffolds provide project context in the Test Sandbox
- See [`docs/skill-call-improvement-loop.md`](docs/skill-call-improvement-loop.md) for SCIL mechanics: holdout splits, scoring, improvement prompt
- See [`docs/agent-call-improvement-loop.md`](docs/agent-call-improvement-loop.md) for ACIL mechanics: agent detection, temp plugin isolation, holdout splits, scoring
- See [`docs/llm-judge.md`](docs/llm-judge.md) for judge mechanics: prompt construction, scoring, output format
Expand Down
3 changes: 2 additions & 1 deletion Makefile
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@ sandbox-setup: build
./harness sandbox-setup

sandbox-clean:
docker sandbox rm claude-skills-harness
sbx rm --force claude-skills-harness

dev:
bun install
Expand All @@ -31,6 +31,7 @@ build:
--external '@duckdb/node-bindings-win32-x64'
DUCKDB_DIR=$$(find node_modules/.bun -maxdepth 6 -name "duckdb.node" -path "*node-bindings-$(DUCKDB_PLATFORM)*" 2>/dev/null | head -1 | xargs dirname) && \
rm -rf node_modules/@duckdb/node-bindings-$(DUCKDB_PLATFORM) && \
mkdir -p node_modules/@duckdb && \
ln -sf $(TESTS_DIR)/$$DUCKDB_DIR node_modules/@duckdb/node-bindings-$(DUCKDB_PLATFORM) && \
cp $$DUCKDB_DIR/libduckdb.dylib $(TESTS_DIR)/libduckdb.dylib

Expand Down
18 changes: 10 additions & 8 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,15 +8,16 @@ things you can improve independently:
- **Effectiveness** — once invoked, does the output meet quality criteria you
define? An LLM judge scores it against a rubric.

You write evals, run them inside a Docker sandbox, and track trigger accuracy,
You write evals, run them inside a Test Sandbox, and track trigger accuracy,
output quality, and cost over time through a dashboard and analytics. **Skills**
and **agents** each have their own trigger-accuracy and effectiveness paths —
pick the one that matches what you're improving.

## Prerequisites

- **Docker** — the harness runs Claude Code in a Docker sandbox. Install
[Docker Desktop](https://www.docker.com/products/docker-desktop/).
- **Docker Sandboxes (`sbx`)** — the harness runs Claude Code inside Docker
Sandboxes via the standalone `sbx` CLI. Install `sbx` using Docker's
instructions, then run `sbx login` before creating the harness sandbox.
Comment thread
robsdudeson marked this conversation as resolved.
Outdated
- **Bun** — the CLI and web app are built with Bun. Install from
[bun.sh](https://bun.sh).

Expand All @@ -31,10 +32,11 @@ All commands run from the **repository root**.
make build
```

2. **Create the Docker sandbox and log in.** This creates a persistent sandbox
and opens Claude Code so you can authenticate:
2. **Create the Test Sandbox and log in.** Authenticate the Sandbox CLI, then
create a persistent sandbox and open Claude Code so you can authenticate:

```bash
sbx login
./harness sandbox-setup
```

Expand Down Expand Up @@ -84,7 +86,7 @@ reference material — reach for them when a guide points you here.
### Configuration Reference

- [Test Suite Reference](docs/test-suite-reference.md) — full `tests.json` field reference: test types, expectation types, validation
- [Test Scaffolding](docs/test-scaffolding.md) — how scaffolds provide project context inside the Docker sandbox
- [Test Scaffolding](docs/test-scaffolding.md) — how scaffolds provide project context inside the Test Sandbox

### Eval Authoring Skills

Expand All @@ -108,7 +110,7 @@ Claude Code skills that generate eval suites for you:
### Architecture

- [Test Harness Architecture](docs/test-harness-architecture.md) — system architecture, package boundaries, data flow, and dependency graph
- [Docker Integration](docs/docker-integration.md) — Docker sandbox architecture, API, lifecycle, and consumer patterns
- [Sandbox Integration](docs/sandbox-integration.md) — Test Sandbox architecture, API, lifecycle, and consumer patterns
- [Project Discovery](docs/project-discovery.md) — generated project attributes: languages, frameworks, tooling, commands

### Package Documentation (contributor)
Expand All @@ -118,7 +120,7 @@ Claude Code skills that generate eval suites for you:
- [Data](docs/data.md) — shared data layer: types, config parsing, JSONL I/O, DuckDB analytics, SCIL utilities
- [Evals](docs/evals.md) — evaluation engine: boolean evals, LLM judge scoring, rubric parsing, orchestrator
- [Claude Integration](docs/claude-integration.md) — Claude CLI wrapper API, argument construction, sandbox delegation
- [Docker Integration Package](docs/docker-integration-package.md) — Docker sandbox API: full public interface, error handling, testing patterns
- [Sandbox Integration Package](docs/sandbox-integration-package.md) — Test Sandbox API: full public interface, error handling, testing patterns
- [Web](docs/web.md) — Hono API server, React SPA, test run and SCIL views, per-test analytics
- [Bun Helpers](docs/bun-helpers.md) — cross-runtime path resolution utilities
- [Test Fixtures](docs/test-fixtures.md) — shared fixture data, loadFixtures utility, analytics JSONL scenarios
Expand Down
12 changes: 6 additions & 6 deletions bun.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

20 changes: 10 additions & 10 deletions docs/adrs/20260326084800-skip-permissions-in-test-sandbox.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@

## Context

The test harness runs Claude Code in `--print` mode (non-interactive) inside an isolated Docker sandbox. In this mode, Claude cannot prompt the user for permission approval. When a prompt test invokes a skill via the `Skill` tool, Claude Code denies the call because `Skill` is not in any auto-approved tools list. Claude then falls back to performing the task manually — bypassing the skill entirely — which causes `skill-call` expectations to fail consistently.
The test harness runs Claude Code in `--print` mode (non-interactive) inside an isolated Test Sandbox. In this mode, Claude cannot prompt the user for permission approval. When a prompt test invokes a skill via the `Skill` tool, Claude Code denies the call because `Skill` is not in any auto-approved tools list. Claude then falls back to performing the task manually — bypassing the skill entirely — which causes `skill-call` expectations to fail consistently.

This was discovered through the code-review test suite, where the prompt `run a /code-review on lib/example.rb` correctly triggered a `Skill` tool call, but the call returned `is_error: true` with the `Skill` tool listed in `permission_denials`. Claude's thinking confirmed the fallback: "Let me read the file first and then do the code review manually."

Expand All @@ -18,7 +18,7 @@ The problem affects all prompt-type tests that expect skill invocation, making i
## Decision Drivers

- Tests must run non-interactively — no human is present to approve permission prompts
- The Docker sandbox already provides process and filesystem isolation
- The Test Sandbox already provides process and filesystem isolation
- Permission denials are infrastructure artifacts, not skill quality signals — they mask real test results
- The test harness is ephemeral and disposable; it is not a production workload
- Test results must be deterministic and reproducible across runs
Expand All @@ -32,11 +32,11 @@ The problem affects all prompt-type tests that expect skill invocation, making i
- Con: May be insufficient — once a skill executes, its body invokes additional tools (`Bash(git *)`, `Read`, `Grep`, `Glob`, `Agent`). While the skill's `allowed-tools` frontmatter should cover these, untested permission interactions could cause silent failures deeper in execution
- Con: Requires incremental flag additions as new permission issues are discovered, creating ongoing maintenance burden

2. **`--dangerously-skip-permissions`** — Skip all permission checks entirely in the Docker sandbox.
2. **`--dangerously-skip-permissions`** — Skip all permission checks entirely in the Test Sandbox.

- Pro: Eliminates the entire class of permission-related test failures
- Pro: Simple, deterministic, and zero-maintenance — no classifier behavior or incremental flag additions to reason about
- Pro: The Docker sandbox already provides the isolation boundary that permissions would otherwise enforce
- Pro: The Test Sandbox already provides the isolation boundary that permissions would otherwise enforce
- Con: No permission guardrails at all within the sandbox; a misbehaving test could modify or delete sandbox contents unchecked
- Con: The flag name itself signals risk, which could cause concern during code review

Expand All @@ -50,7 +50,7 @@ The problem affects all prompt-type tests that expect skill invocation, making i

## Decision

We will use **`--dangerously-skip-permissions`** because the Docker sandbox already provides the isolation boundary that makes permission checks redundant in this context. The test harness exists to evaluate skill quality — not to test Claude Code's permission system. Stripping permissions eliminates an entire class of infrastructure-artifact failures and keeps test results focused on what matters: whether skills trigger correctly and produce quality output.
We will use **`--dangerously-skip-permissions`** because the Test Sandbox already provides the isolation boundary that makes permission checks redundant in this context. The test harness exists to evaluate skill quality — not to test Claude Code's permission system. Stripping permissions eliminates an entire class of infrastructure-artifact failures and keeps test results focused on what matters: whether skills trigger correctly and produce quality output.

The flag will be added to the Claude CLI arguments in both the prompt test runner and the skill-call test runner.

Expand All @@ -69,7 +69,7 @@ The flag will be added to the Claude CLI arguments in both the prompt test runne

**Neutral:**

- The Docker sandbox isolation model is unchanged — it already prevents test actions from affecting the host system regardless of permission settings
- The Test Sandbox isolation model is unchanged — it already prevents test actions from affecting the host system regardless of permission settings

## Notes

Expand All @@ -79,13 +79,13 @@ The flag will be added to the Claude CLI arguments in both the prompt test runne
|------|---------|
| `tests/packages/cli/src/test-runners/prompt/index.ts` | Prompt test runner — builds Claude CLI args (lines 74-79) |
| `tests/packages/cli/src/test-runners/skill-call/index.ts` | Skill-call test runner — builds Claude CLI args (lines 77-82) |
| `tests/packages/docker-integration/src/sandbox.ts` | Docker sandbox executor — passes args to `docker sandbox exec` |
| `tests/packages/docker-integration/sandbox-run.sh` | Shell script inside sandbox that invokes `claude "$@"` |
| `tests/packages/cli/src/commands/sandbox-setup.ts` | Sandbox creation and OAuth setup (delegates to `@testdouble/docker-integration`) |
| `tests/packages/sandbox-integration/src/sandbox.ts` | Test Sandbox executor — passes args to `sbx exec` |
| `tests/packages/sandbox-integration/sandbox-run.sh` | Shell script inside sandbox that invokes `claude "$@"` |
| `tests/packages/cli/src/commands/sandbox-setup.ts` | Sandbox creation and OAuth setup (delegates to `@testdouble/sandbox-integration`) |

**Related documentation:**

- [Claude Code Auto Mode](https://www.anthropic.com/engineering/claude-code-auto-mode) — Anthropic's guidance on permission handling in automated environments
- [Claude Code CLI Reference](https://code.claude.com/docs/en/cli-reference) — Full CLI flag documentation including `--dangerously-skip-permissions` and `--allowedTools`
- [Claude Code Permission Modes](https://code.claude.com/docs/en/permission-modes) — Permission mode documentation
- [Docker Integration](../docker-integration.md) — Full API reference for the sandbox execution package
- [Sandbox Integration](../sandbox-integration.md) — Full API reference for the sandbox execution package
15 changes: 15 additions & 0 deletions docs/adrs/20260515000000-migrate-sandbox-cli-to-sbx.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
# Migrate sandbox integration to sbx

The harness will hard-cut over from the deprecated `docker sandbox` subcommand to Docker's standalone `sbx` CLI for managing the **Test Sandbox**. We will rename `@testdouble/docker-integration` to `@testdouble/sandbox-integration` and `DockerError` to `SandboxError` at the same time, while keeping user-facing harness commands such as `sandbox-setup`, `shell`, and `clean` stable. We are not keeping a `docker sandbox` fallback because sandbox management is centralized behind a small package API, dual command support would add stale branching, and first-time setup should teach the current Test Sandboxes workflow (`sbx login` then `./harness sandbox-setup`). The integration layer will also convert common `sbx` failures into tailored harness errors for missing CLI, auth/readiness failure, and missing named sandbox.

## Considered Options

- **Hard cutover to `sbx`** — simpler implementation, current upstream command surface, no stale compatibility path.
- **Compatibility adapter** — lower short-term risk for users with the old CLI, but adds branching around a retired command and keeps outdated terminology in the codebase.

## Consequences

- Developers must install and authenticate `sbx` before running the harness sandbox setup.
- The integration package name now describes the harness boundary rather than the retired Docker CLI shape.
- Existing harness command names remain stable for users and scripts.
- Missing `sbx` installation or authentication failures produce explicit harness errors instead of raw process failures.
6 changes: 3 additions & 3 deletions docs/agent-call-improvement-loop.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,17 +10,17 @@ Agent descriptions determine when Claude delegates tasks to custom agents — th

ACIL runs a loop over `agent-call` type tests in a test suite:

1. **Evaluate** — run each test case against the current agent description in a Docker container, recording whether the agent was delegated to as expected
1. **Evaluate** — run each test case against the current agent description in a Test Sandbox, recording whether the agent was delegated to as expected
2. **Score** — compute trigger accuracy across all test cases
3. **Improve** — send the failures and history to Claude in a Docker container and ask for an improved description, using phase-specific instructions (see [Divergent-Convergent Phases](#divergent-convergent-phases) below)
3. **Improve** — send the failures and history to Claude in a Test Sandbox and ask for an improved description, using phase-specific instructions (see [Divergent-Convergent Phases](#divergent-convergent-phases) below)
4. **Repeat** — loop up to `--max-iterations` times, tracking the best description found
5. **Apply** — write the best description back to the agent `.md` file, either automatically (`--apply`) or after prompting

At the end of every iteration, ACIL prints a progress summary. When the loop exits, it shows a table of all iterations with accuracy scores and highlights the best result.

## Prerequisites

`acil` uses the same Docker sandbox as `test-run`. Build the harness and set up the sandbox before running:
`acil` uses the same Test Sandbox as `test-run`. Build the harness and set up the sandbox before running:

```bash
make build
Expand Down
2 changes: 1 addition & 1 deletion docs/build-agent-eval-scaffold.md
Original file line number Diff line number Diff line change
Expand Up @@ -110,7 +110,7 @@ After generating the scaffold, you can:
## References

- [Building Rubric Evals](rubric-evals-guide.md) — step-by-step guide covering the full workflow from scaffolds to rubric evaluation
- [Test Scaffolding](test-scaffolding.md) — how scaffolds provide project context inside the Docker sandbox
- [Test Scaffolding](test-scaffolding.md) — how scaffolds provide project context inside the Test Sandbox
- [Test Suite Reference](test-suite-reference.md) — full tests.json field reference
- [Writing Agent Eval Rubrics](write-agent-eval-rubric.md) — the `/write-agent-eval-rubric` skill: workflow, criteria categories, output format
- [Building Skill Eval Scaffolds](build-skill-eval-scaffold.md) — the equivalent skill for skill-based scaffold generation
Expand Down
Loading
Loading