A worked end-to-end example of using Trailblaze to test a real web
application. Drives live en.wikipedia.org through the Playwright
Native driver. The trailmap ships:
- 9 scripted tools (in TypeScript) covering search, language switching, random article, main-page section verification, banner dismissal, article structure assertions, plus a composition example.
- 30 trails under
trails/wikipedia/exercising those tools- the built-in
web_*toolset, withtagsfor selective runs and oneskip:example for opting out gracefully.
- the built-in
- A target-scoped system prompt that teaches the LLM when to reach for
the scripted tools instead of inline-expanding to
web_*primitives. - CLI-first runtime — every trail runs via
trailblaze run …, the same path an end user takes. There is no JUnit harness.
If you're building a target trailmap for your own site, this is the reference shape to copy.
You're already using Playwright under the hood. Trailblaze adds two things:
- Natural-language trails. A test like "search Wikipedia for Albert Einstein and verify the article rendered" stays readable to a PM, QA, or support engineer who doesn't write Playwright. The LLM resolves it against the live DOM each run.
- Per-target scripted tools. When you DO want determinism for hot paths, you write a TypeScript tool once, advertise its task pattern to the LLM via the system prompt, and let trails compose it by name. The selector knowledge stays in one place; the trail just describes intent.
In practice you'll write both: NL trails for breadth, scripted tools for the 5-10 workflows you run every build.
Prerequisites
- The
trailblazeCLI on yourPATH. See Getting Started for the install steps. Paths in the commands below are relative to wherever you cloned this example tree — the simplest setup is to cloneblock/trailblazeandcdto the repo root. - The Playwright Native driver ships with Trailblaze — no separate
npm install/npx playwright installstep needed for this example; the first run downloads the Chromium build into the workspace. - Network access to
en.wikipedia.org. This trailmap drives the live site; proxy + corporate firewall users should confirm reachability first.
Run all commands below from the OSS repo root (the parent of
examples/andtrails/). The$PWD/examples/...path and thetrails/wikipedia/...invocation both anchor there.
# 1. From the repo root, point the daemon at this trailmap's config dir.
# This is what registers the wikipedia_web_* tools when the daemon boots.
export TRAILBLAZE_CONFIG_DIR=$PWD/examples/wikipedia/trails/config
# 2. Stop any existing daemon (so it re-reads the config dir) and start
# fresh. Source builds rebuild here automatically.
trailblaze app --stop
trailblaze app --headless & disown
# 3. Confirm the scripted tools registered. You should see 9 entries:
trailblaze toolbox --device web --target wikipedia --search wikipedia_
# 4. Run a single trail end-to-end.
trailblaze run trails/wikipedia/test-search-einstein \
--device web/playwright-nativeThe first run takes ~10s (browser launch + 2 LLM rounds). Expect
Results: 1 passed, 0 failed. If you got that, you're set up. Skip to
Anatomy of a tool to see how the wiring works.
The --device string. Step 3 uses web (the platform); step 4 uses
web/playwright-native (platform + driver). The shorter form picks the
default driver for the platform; the longer form pins one driver — useful
in CI / scripts where you want determinism. trailblaze device list
shows the full set of <platform>/<driver> pairs reachable from your
machine.
If the run hangs or reports Daemon unreachable after 30 consecutive poll failures, see Known issues — it's a known CLI flake on
long LLM rounds. The daemon itself is fine; re-run the trail.
Every scripted tool is a single .ts file under
trails/config/trailmaps/wikipedia/tools/ — no sibling YAML, no separate
descriptor to maintain. The trailmap loader auto-discovers every .ts
that exports a trailblaze.tool(...) declaration; the names listed under
target.tools: in trailmap.yaml decide which of those are advertised to
the agent. Here's wikipedia_web_openArticle broken into the parts that
matter. (@trailblaze/scripting is the SDK the daemon synthesizes into the
workspace at trails/.trailblaze/sdk/ — no npm install needed.)
// wikipedia_web_openArticle.ts
import { trailblaze } from "@trailblaze/scripting";
import { SELECTORS, articleUrl, nonEmptyString } from "./wikipedia_shared";
export interface OpenArticleArgs {
/** Title of the Wikipedia article to open. Defaults to "Wikipedia". */
title?: string;
}
/**
* Open a Wikipedia article by title. Use this whenever the task is to
* navigate to a specific article — e.g. "open the Albert Einstein
* article". Asserts the destination's #firstHeading is visible.
*/
export const wikipedia_web_openArticle = trailblaze.tool<OpenArticleArgs>(
{ supportedPlatforms: ["web"], requiresContext: true }, // ← spec object: gates + hints
async (input, ctx) => { // ← typed handler
const title = nonEmptyString(input.title, "Wikipedia");
await ctx.tools.web_navigate({ action: "GOTO", url: articleUrl(title) });
await ctx.tools.web_verifyElementVisible({ ref: SELECTORS.firstHeading });
return `Opened "${title}".`; // ← string return surfaces to the LLM as the tool result
},
);What the framework derives from this one file:
| From | Becomes |
|---|---|
Export name wikipedia_web_openArticle |
The dispatchable tool name + an entry on ctx.tools for sibling tools. |
TSDoc above export const |
The LLM-facing description. |
The <OpenArticleArgs> type parameter |
The tool's input JSON Schema (with per-field TSDoc carried through). |
The spec object { supportedPlatforms, requiresContext } |
Runtime registration gates and metadata hints. |
| The handler body | The runtime behavior. |
Key conventions to copy:
- Selectors live in
wikipedia_shared.tsas a frozenSELECTORSconstant — never inlined in tool bodies. If Wikipedia changes the#firstHeadingid tomorrow, you change one constant, not 9 files. - Helpers are shared too —
nonEmptyString,tryOrFalse,elementIsVisible,ensureOn. Every tool imports from./wikipedia_shared. - All dispatch goes through
ctx.tools.X(args)— typesafe;tscflags typos and bad arg shapes at compile time. The exportedToolContexttype deliberately omits a genericcallTool, so unknown tool names surface as compile errors, not runtime "tool not registered" failures. - Return a short human-readable string. The LLM sees this as the
tool result; keep it informative ("Opened X" beats "ok"). For typed
structured returns, declare a second type parameter:
trailblaze.tool<MyInput, MyOutput>(spec, handler). - Throw on failure. Don't swallow — the framework treats a thrown
Erroras the tool failing, which is what you want for the agent to see and react to.
One file, then restart the daemon. Concrete steps:
TOOL=wikipedia_web_myNewTool
TRAILMAP=examples/wikipedia/trails/config/trailmaps/wikipedia
# 1. Body — what the tool does. The TSDoc on the exported `const` is what
# the LLM reads as the tool's description; the spec object's
# `supportedPlatforms` / `requiresContext` flags carry runtime hints.
cat > $TRAILMAP/tools/${TOOL}.ts <<'TS'
import { trailblaze } from "@trailblaze/scripting";
export interface MyNewToolArgs {
/** <describe> */
someArg?: string;
}
/**
* <One paragraph in plain English, including the task patterns the LLM
* should match against. NO "USE THIS TOOL" — describe what it DOES.>
*/
export const wikipedia_web_myNewTool = trailblaze.tool<MyNewToolArgs>(
{ supportedPlatforms: ["web"], requiresContext: true },
async (input, ctx) => {
// ... do work via ctx.tools.X(args) ...
return "Did the thing.";
},
);
TS
# 2. Register in trailmap.yaml — add `wikipedia_web_myNewTool` under target.tools.
$EDITOR $TRAILMAP/trailmap.yaml
# 3. Restart the daemon so the new tool is discovered + bindings regenerate.
trailblaze app --stop && trailblaze app --headless & disown
# 4. Confirm registration + open the .ts file in your IDE; `ctx.tools.`
# autocomplete should now include your new tool.
trailblaze toolbox --device web --target wikipedia --search myNewToolFor reference, here's what trailmap.yaml's target.tools: list looks like — bare
export names, one per line:
# trails/config/trailmaps/wikipedia/trailmap.yaml (excerpt)
id: wikipedia
target:
display_name: Wikipedia (en)
system_prompt_file: wikipedia-system-prompt.md
tools:
- wikipedia_web_openMainPage
- wikipedia_web_searchAndOpenFirstResult
- wikipedia_web_myNewTool # ← your new tool, by export name
# ... others ...
platforms:
web:
drivers: [playwright-native, playwright-electron]
tool_sets: [web_core, web_verification, memory]That's the whole loop. If autocomplete doesn't pick up the new tool, run
trailblaze check --workspace examples/wikipedia/trails/config
manually — that's the per-trailmap trailblaze-client.d.ts regeneration step the daemon
fires on restart.
The TSDoc on each exported const is read by the LLM at session start
as the only way it learns what your tool does. Two non-obvious rules:
- Don't say "USE THIS TOOL FOR X". The LLM picks tools by matching its understanding of the task against the description. Telling it to "use" the tool reduces the description to a single keyword and loses the surrounding context. Describe what the tool does and include the task patterns it matches — "Search Wikipedia for X / look up Y on Wikipedia / find articles about Z" — not "USE THIS WHEN SEARCHING".
- Match real user phrasing. If a trail says "search Wikipedia for Python" but the tool description only mentions "query the search endpoint", the LLM may not connect them. Include the synonyms (search, look up, find articles about, etc.) that real prompts will use.
You can sanity-check picking by running a trail in --verbose mode and
watching which tool the agent chose for each step.
A trail directory's trail.yaml is the source of truth — a natural-language
step the agent resolves against the live page. Here's
test-search-einstein/trail.yaml
verbatim:
config:
target: wikipedia
devices:
web:
driver: PLAYWRIGHT_NATIVE
tags:
- smoke
- search
title: 'Wikipedia: Search for Albert Einstein'
trail:
- step: Search Wikipedia for "Albert Einstein" and verify the resulting article page shows the heading "Albert Einstein".The agent sees the trailmap's target.tools: (which includes the scripted
wikipedia_web_searchAndOpenFirstResult and wikipedia_web_verifyArticleStructure)
and the system prompt at wikipedia-system-prompt.md nudges it toward those
tools when a step matches their task patterns. Run it:
trailblaze run trails/wikipedia/test-search-einstein --device web/playwright-nativeAfter a passing run, the CLI folds a fresh recording into the trail's
trail.yaml — a recording: block nested under each step, keyed by device
classifier (web) — that's the artifact you commit for deterministic replay.
One caveat for this example specifically: scripted-tool calls
(wikipedia_web_*) currently don't dispatch from a saved web recording —
the Playwright agent rejects them at replay (tracked under
Known issues). 5/30 trails here carry recordings — the ones
that exercise pure web_* built-ins; the rest stay NL.
Each trail directory has a trail.yaml (the source of truth). If a step
carries an embedded recording: block for the device under test, the CLI
replays those tools deterministically instead of going through the LLM.
Use the decision tree:
| Property of the trail | NL (no recording) | Recorded (embedded recording:) |
|---|---|---|
| Should pass on any reasonable Wikipedia state | ✅ preferred | risky — recording fixes one DOM |
| You want zero LLM cost on every run | ❌ ~$0.02-0.05/run | ✅ free after recording |
| The DOM under test changes daily (e.g. featured content) | ✅ stays robust | ❌ will drift |
| You need stable CI signal across hundreds of runs | depends on flakiness | ✅ if the selectors are stable |
The trail composes scripted (wikipedia_web_*) tools |
✅ works fine | ⚠ see "known issues" below |
The 5 recorded trails in this repo are the ones that hit structural
anchors stable across days (test-article-shakespeare,
test-language-switch-spanish, the 3 main-page section trails).
To re-record a trail: delete the step's recording: block(s) from its
trail.yaml and re-run it. The CLI folds a fresh recording back into the
trail file on a passing run.
Every trail's config: block carries a tags: list. --tags filters by them.
# Smoke pass (fast, deterministic, must-pass on every build):
trailblaze run trails/wikipedia --device web --tags smoke
# Anything in the i18n or article suites:
trailblaze run trails/wikipedia --device web --tags i18n,article
# Long-tail nightly:
trailblaze run trails/wikipedia --device web --tags slowConventions used in this example:
| Tag | Meaning |
|---|---|
smoke |
Fast, deterministic, must-pass on every run. |
main-page |
Asserts something about wiki/Main_Page. |
nav |
Exercises in-page navigation primitives (Random article, hamburger menu). |
search |
Drives the header search box. |
article |
Loads a specific article and asserts its shape. |
i18n |
Crosses language subdomains. |
slow |
Multi-step assertions; skip during local iteration. |
flaky |
Known to be timing-sensitive; usually paired with a skip: reason. |
A trail's config: block can carry a skip: map — keyed by device classifier
— to opt out until the reason is removed. Listing every device the trail
targets skips it entirely; keying a single device skips just that one while the
trail still runs on the others. Better than commenting trails out or
maintaining a separate exclusion list — the reason is committed alongside the
trail and shows up in --verbose output.
config:
title: "Wikipedia: Autocomplete suggestions visible while typing search"
tags: [search, flaky]
devices:
web:
driver: PLAYWRIGHT_NATIVE
skip:
web: "Autocomplete popup is timing-sensitive; remove `skip:` once the suggestion-list probe lands."
trail:
- step: …To run a skipped trail anyway, delete the device's entry from skip: (or set
its reason to an empty string). The CLI treats blank values as "not skipped".
A few trails are specifically chosen as teaching examples. Read these when you're authoring new tools or trails for your own target:
| Pattern | Trail | What it shows |
|---|---|---|
Simple NL → built-in web_* tools |
test-main-page-loads |
Smallest possible trail; agent uses web_* directly. |
| NL → scripted tool (LLM picks it) | test-search-einstein |
System prompt nudges the agent toward searchAndOpenFirstResult without naming it. |
| Direct scripted-tool dispatch | test-custom-search |
Trail prompts name the tool explicitly + pass typed args. Maximum determinism short of a recording. |
| Composition (tool calls tool) | test-search-multi-topic |
Exercises wikipedia_web_searchAndVerify, which internally calls two other scripted tools. |
| Data-driven shape | test-search-multi-topic |
Same workflow across multiple queries — copy-paste the prompt block + change the data. |
| Recorded deterministic replay | test-article-shakespeare |
trail.yaml with an embedded recording: block — replay path skips the LLM entirely. |
| Conditional UI (banner present?) | test-main-page-featured-article |
Exercises dismissBannerIfPresent, which no-ops cleanly when no banner is shown. |
| Branch coverage on a feature flag | test-article-short-no-refs + test-article-references-section |
Pair covers both branches of verifyArticleStructure's requireReferences flag. |
| Graceful skip | test-search-autocomplete |
Shows the per-device skip: config map. |
Trailmap tools are TypeScript only — the bundler rejects .js / .mjs / .cjs
sources so authors can't fall back to a subprocess runtime that would lose
the typesafe contract.
trailblaze check --workspace examples/wikipedia/trails/configThat writes two artifacts (both gitignored as derived output):
- Workspace SDK at
trails/.trailblaze/sdk/— the@trailblaze/scriptingsource the per-trailmaptsconfig.jsonresolves through path mapping. Ships curated runtime globals (URL,fetch,AbortController,console). - Per-trailmap
trailblaze-client.d.tsat<trailmap>/tools/trailblaze-client.d.ts— exhaustive bindings for every tool the runtime registry knows about (framework built-ins + trailmap-local scripted tools + Kotlin tools surfaced throughtool_sets:declared intrailmap.yaml).
Open any .ts file in tools/ and hover input / ctx to confirm the
types resolve. Mistype a tool name or pass the wrong arg shape and tsc
flags it at compile time.
To run the bundled tsc against your trailmap:
trailblaze check wikipediaPair every .ts tool with a sibling *.test.ts file and run them through the mock
client + mock context from @trailblaze/scripting/testing. Tests execute without a
daemon or a device:
trailblaze check wikipedia(Per-trailmap unit tests run as the third phase of trailblaze check, after materialize +
tsc.)
This trailmap ships scripted tools but no .test.ts files yet — for working examples of
single-tool sequences, multi-step workflows, and client.stub(...) recovery-branch
tests, see the playwright-native trailmap's .test.ts
samples and the
canonical authoring tutorial's
Testing your tool section.
.ts files route through the daemon's QuickJS in-process runtime. No
subprocess fork, no Node APIs available beyond the curated globals the SDK
ships. Sub-millisecond dispatch; every tool call goes through
ctx.tools.X(args).
Daemon won't start — check trailblaze app --status, then look at
~/.trailblaze/daemon.log for the last few hundred lines. The daemon-<pid>.pid
file under ~/.trailblaze/ is the canonical location; stale entries from
previous crashes are safe to delete.
Tool not registering / "tool not found at dispatch" —
- Confirm the file is named exactly
<toolName>.tsundertools/and exports atrailblaze.tool(...)declaration whose export name matches. - Confirm the bare tool name is listed under
target.tools:intrailmap.yaml. Names omitted from that list are loaded by the bundler but not advertised to the agent for this target. - Restart the daemon after editing anything under
trails/config/trailmaps/<trailmap>/. Trailmaps are discovered at boot.
IDE shows red squiggles on ctx.tools.X(args) — run
trailblaze check --workspace <path-to-your-config> to regenerate the per-trailmap
trailblaze-client.d.ts. If that fails, check trailblaze app --status — the
daemon writes the SDK + trailblaze-client.d.ts on every restart.
"Daemon unreachable after 30 consecutive poll failures" — known CLI flake on long agent rounds. See Known issues below.
Tool calls that worked live now fail on replay — wikipedia_web_*
scripted tools currently don't dispatch from an embedded web recording.
See Known issues below.
These are tracked framework gaps. The example works around them today; the canonical shape will improve when these land.
- Scripted-tool replay-dispatch gap. When an embedded
webrecording captures a call towikipedia_web_*, the Playwright agent rejects it at replay asOtherTrailblazeTool(the dispatcher only handlesweb_*built-ins). Until fixed, only trails that exercise pureweb_*tools get useful recordings. That's why 25/30 trails here stay NL-only. - CLI poll-timeout on long LLM rounds.
trailblaze run …can returnFAILED: Daemon unreachable after 30 consecutive poll failureswhile the daemon is still healthy and the session is making progress. The session artifacts underlogs/<session>/are still written; you can replay or inspect them after. Workaround: re-run; not deterministic.
examples/wikipedia/
├── README.md ← you are here
├── build.gradle.kts ← only runs `trailblaze.bundle`
└── trails/config/
├── trailblaze.yaml ← workspace anchor
└── trailmaps/wikipedia/
├── trailmap.yaml ← target manifest
├── wikipedia-system-prompt.md ← LLM tool-selection guidance
└── tools/ ← scripted tool sources
├── tsconfig.json ← framework-managed (extends workspace base)
├── trailblaze-client.d.ts ← framework-managed (typed `ctx.tools` bindings)
├── wikipedia_shared.ts ← helpers reused by every tool
└── wikipedia_web_*.ts ← one .ts per scripted tool
trails/wikipedia/
├── test-main-page-*/ ← main-page sections + smoke
├── test-search-*/ ← search flows (NL + scripted)
├── test-article-*/ ← article-shape assertions
├── test-language-*/ ← i18n / language switcher
├── test-random-article/ ← Random article navigation
├── test-back-navigation/ ← Browser-history nav
├── test-footer-privacy-link/ ← Footer scrolling assertion
├── test-wiki-logo-visible/ ← Header logo presence
└── test-custom-*/ ← typed scripted-tool composition
CI runs the same commands a developer does — there's no separate harness.
The benchmark pipeline scripts under scripts/.../runway/ set up the
daemon, fan out via trailblaze run, and summarize results. Look at
those if you're wiring a similar trailmap into your own CI.