Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
2 changes: 1 addition & 1 deletion .trailblaze-sync
Original file line number Diff line number Diff line change
@@ -1 +1 @@
1921ffb740992b42bab05cc9000ad12e1fb89e76
7a8bb63d6cda9175bb3b7aa872f05062881eed80
36 changes: 17 additions & 19 deletions docs/CLI.md

Large diffs are not rendered by default.

47 changes: 41 additions & 6 deletions docs/configuration.md

Large diffs are not rendered by default.

103 changes: 103 additions & 0 deletions docs/devlog/2026-09-27-strings-on-every-capture.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,103 @@
---
title: "Strings on Every Capture"
type: decision
date: 2026-09-27
---

# Strings on Every Capture

## Summary

Every capture log now carries the text on screen. The strings are read from the view tree when
the log is written, on device and on the host, so every session has them without an export step.
Each string points at its capture and its box on it. That powers the report's Strings tab, crops of
each string cut out of its screenshot, the survey SDK's `screenText`, and `strings diff`.

## What Changed

- `TrailblazeLog.withVisibleStrings()` (models `commonMain`) fills a `visibleStrings` field on
`AgentDriverLog`, `TrailblazeSnapshotLog` and `TrailblazeLlmRequestLog`. It runs in the logging
rule's emitter and again in `LogsRepo.saveLogToDisk` as a catch-all. It is idempotent and never
throws: a capture with no strings is better than a capture lost to a parse error.
- Each string has `text`, `source`, `bounds` (`[left, top, right, bottom]` in device pixels) and
`visible`. `visible` is left out on disk when true, so a reader must treat a missing value as
true.
- `visible-strings.ndjson` is still written, now as version 2. It is derived from the logs.
Each line names its capture with `captureId` and carries the image separately in
`screenshot` (plus `captureUrl` on the farm).
- In the report, the Strings tab lists every string with Visible / Not visible / All filters.
By screen mode draws each string's box on its screenshot, and each row shows the string cropped
out of its screenshot.

## Key Decisions

**Record at capture time, not in a pass afterwards.** Before this, strings existed only when
someone ran the export over a finished session. Putting them on the log means every reader (the
report, surveys, diffs, anything downstream) gets them from the same place for every session,
including CI runs. Nearly all of the value came from this decision.

**Name the capture, and keep the image in its own field.** A string is only trustworthy if you
can see where it was: which frame, and which box on it. A capture with a screenshot is named by
that screenshot, which every reader already has. A capture without one (screenshots turned off,
or a failed capture) gets an id stamped on its log once, as it's emitted, and never rewritten.
The id is stored, not derived, so every reader sees the same one. Log file names and timestamps
were ruled out: farm runs rename log files and none of the readers see them, and several logs of
one screen carry different timestamps. `captureId` is opaque, and the image is a separate
`screenshot` field, so a video frame can later be another field without changing the key. This
also replaced a capture counter that was named `stepIndex` but was never a trail step. Each
Timeline row lists the ids of its screenshot-less captures, so the Strings tab links those to the
exact row and dispatch too.

**Judge volatile text when it's read, not when it's written.** Clocks, balances and counters (a
digit and no words except AM/PM) are recognised by `VolatileText.looksVolatile` in Kotlin and
`looksVolatile` in TypeScript. The two must stay in step. Because nothing is stored, a better
rule reaches every session already recorded.

**A crop is exactly what the device drew.** Crops are a CSS background: the whole screenshot,
scaled and shifted so only the box shows. Each screenshot is written once, in one style rule,
not once per row. Crops have square corners. They come only from a place where the string was
visible, because a box that was scrolled away cuts out blank space and a covered box cuts out
whatever covers it.

## What We Learned

**Renaming fields in an export breaks readers you don't own.** Version 2 dropped `stepIndex`,
`repeatOfStepIndex`, `logType` and the stored `volatile` flag, and changed `bounds` to corners.
That was tidier, but a comparison tool outside this repo read those fields and showed no strings
until it was updated. Adding the new fields next to the old ones would have delivered the same
value with no break. Next time a file has readers outside the repo, add fields and don't rename
or remove them.

**Pair by position only across the whole stream.** A reader that buckets captures by trail step
sees a slice of the stream. Numbering captures within the slice pairs them with the wrong
capture, so a capture's number must be its position in the whole stream.

## Future Work

- **Trail step.** Add the trail step each capture happened in as a new field. Several captures
share a step. That answers "which step first showed this string" and lets a comparison group
by real steps rather than inferring them from timestamps.
- **Frames from video — built in the browser.** A capture with no screenshot now shows the
recording's frame at its time, taken by the report page and never stored, and labeled "From
video". Capture times are moved onto the host clock the recording uses, and `bounds` scale by
the frame's size like a screenshot's. `trailblaze run` and `trailblaze strings extract` now
also save that frame beside the screenshots as `<captureId>.webp`, by the same rule, and name it
in the strings export's `frame` field; a report that carries the file shows it without reading
the recording. Still open:
- Multi-device sessions. A capture log doesn't name its device, so when two recordings cover
its instant no frame is saved and only the page can take one.
- Daemon sessions. They never write the strings export, so they save no frames.
- Linked clips. A recording linked from a host that doesn't allow reading its pixels gives no
frames; embedded recordings always do.
- Quality. Lossy encoding blurs small text, and a frame may catch the screen mid-animation.
- **Log the screen an action decided on (outside the Maestro path).** On the drivers that don't
run through Maestro, a step's capture is taken before the action starts, and the polls that
follow (an assert waiting for its text, a tap waiting for its target) are discarded. So an
assert that passed after waiting logs a screen that may not show the text, and a tap that
waited draws its marker on a screen taken before the one it matched. Taking the step's capture
from the poll that matched fixes both without adding captures. The Maestro path already does
this: it captures right before the gesture, after the element was found, and again once an
assertion passes. Before changing the others, check every reader that assumes a step's capture
is the "before" screen.
- **Stable pairing across runs.** Pairing screens by position breaks when a list is re-sorted.
Recording an element id per string at capture would let two runs pair by element instead.
Original file line number Diff line number Diff line change
@@ -0,0 +1,95 @@
---
title: "Write Tools in TypeScript: a Scripted Tool Costs a Fifth of a Millisecond"
type: decision
date: 2026-09-29
---

# Write Tools in TypeScript: a Scripted Tool Costs a Fifth of a Millisecond

## Summary

Write new tools in TypeScript, not Kotlin. A TypeScript tool runs in-process in QuickJS, and a call
costs about 0.2 ms more than the same tool written in Kotlin. A device tap or assertion takes
hundreds of milliseconds, so the difference never shows up in a run. What you get for it is a
tool that ships with its trailmap and doesn't need a framework build.

## The numbers

`QuickJsToolDispatchBenchmark` (in `trailblaze-quickjs-tools`) times the real dispatch path. The
TypeScript tools are written with the SDK (`trailblaze.tool(...)`) and bundled the way production
bundles them, so the SDK's per-call wrapper is included. Every tool body is empty, so each row is
overhead. Median of 5,000 calls across two quiet runs on an M-series Mac:

```text
Kotlin tool, dispatched by name ██████ 0.01 ms
TypeScript tool ██████████████ 0.20 ms
TypeScript + 1 Kotlin call ██████████████ 0.21 ms
TypeScript + 10 Kotlin calls ███████████████ 0.31 ms
TypeScript + 10 TS helpers ██████████████ 0.19 ms
Start /usr/bin/true (once) ██████████████████████████ 20 ms
Start node -e 0 (once) █████████████████████████████ 60 ms
One device tap (typical) █████████████████████████████████ hundreds of ms
| | | | | |
ms per call, log scale 0.01 0.1 1 10 100 1000
```

Every row but the device tap is measured; the tap is there for scale. The two process starts are
paid once by a subprocess runtime, not on every call.

- **Entering and leaving QuickJS is the only real cost:** about 0.19 ms per tool call over the same
Kotlin tool. About 0.02 ms of that is the SDK wrapper.
- **Composing Kotlin tools from TypeScript is cheap:** about 0.01 ms per `ctx.tools.<name>(...)` call.
- **Helper functions are free.** They're plain JavaScript calls inside the engine.
- **Session start is small too.** Loading a tool bundle into its engine takes 0.9 ms, and each
scripted tool gets its own engine, so 50 tools cost about 45 ms once.
- **In-process means no process to start.** A subprocess tool runtime pays 20–60 ms once to start,
then an inter-process round trip on every call, which this benchmark doesn't measure.

To reproduce:

```bash
./gradlew :trailblaze-quickjs-tools:jvmTest --tests '*QuickJsToolDispatchBenchmark*' \
-Dtrailblaze.benchmark=true -i
```

It measures the JVM host. On-device QuickJS wasn't measured.

## Tips

- **Write it in TypeScript.** Use Kotlin only when the tool *is* a primitive the sandbox lacks: a
socket, a binary HTTP body, sensitive memory. Even then, keep the Kotlin to that primitive and
write the rest in TypeScript on top of it.
- **Register a tool only when trails or many callers need it.** Every registered tool is listed in
every session's tool registry. Plumbing that one or two tools share is an exported function they
import.
- **Host work goes through `exec`, HTTP through `fetch`.** Between them they cover most of what
Kotlin tools used to be written for.
- **Keep the framework to primitives.** Helpers belong in the trailmap that uses them
([Framework provides primitives; helpers are not framework](2026-06-22-framework-primitives-helpers-compose.md)).

## Example: read an app's saved preference

Reading a value an iOS simulator app saved looks like a job for a dedicated Kotlin tool. It's a
TypeScript helper built on `exec`:

```typescript
import type { ToolContext } from "@trailblaze/scripting";

/** Reads one key an iOS simulator app saved to its preferences. Throws if the app or key is missing. */
export async function readIosAppPreference(ctx: ToolContext, appId: string, key: string): Promise<string> {
const udid = ctx.device?.instanceId;
if (!udid) throw new Error("This session has no simulator.");
const container = String(await ctx.tools.exec({
argv: ["xcrun", "simctl", "get_app_container", udid, appId, "data"],
})).trim();
const plist = `${container}/Library/Preferences/${appId}.plist`;
return String(await ctx.tools.exec({
argv: ["/usr/libexec/PlistBuddy", "-c", `Print :${key}`, plist],
})).trim();
}
```

The tool that needs the value imports the helper. Nothing new is registered, and the framework
gains no surface. `exec` fails on a non-zero exit, so a missing app or key throws. It also returns
stderr merged with stdout, so a production version should check the output's shape before
trusting it.
82 changes: 82 additions & 0 deletions docs/devlog/2026-09-29-what-a-screenshot-costs-per-capture.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,82 @@
---
title: "What a Screenshot Costs per Capture"
type: devlog
date: 2026-09-29
---

# What a Screenshot Costs per Capture

## Summary

Taking the screenshot is the most expensive part of a capture on Android and web, and the second
most expensive on iOS: roughly 100–175 ms of every capture. The session recording already holds the
same pixels, and a frame pulled from it after the run matches the real screenshot (mean SSIM 0.989
over 51 captures, see [Strings on Every Capture](2026-09-27-strings-on-every-capture.md)). So when a
run records video, a capture that only needs its picture for the report can skip the screenshot and
get its frame from the video after the run.

## The numbers

Time a capture spends reading the screen, with and without the screenshot:

| Platform (driver) | Screen read, no screenshot | Screenshot | Screenshot's share |
| :--- | ---: | ---: | ---: |
| Android (on-device accessibility) | 103 ms | +175 ms | ~63% |
| iOS (XCTest runner) | 161 ms | +100 ms, plus host encode | ≥38% |
| Web (Playwright) | 13 ms | +114 ms, plus host encode | ~90% |

All figures are medians.

- **Android** is an A/B of the runner's own `GetScreenStateRequest`, called directly on the device
and alternating `includeScreenshot` false and true (40 calls each). The emulator was a private
API 33 one at 1080×2400. A capture took 103 ms without the screenshot and 277 ms with it, which
includes scaling, WebP encoding (about 32 KB), and the RPC reply. A second pass, taken while the
laptop was under heavier load, gave 126 → 368 ms (+242 ms).
- **iOS** comes from `takeScreenshot` and `contentDescriptor` spans recorded in real runs over two
weeks. p90 is 164 ms for the screenshot and 227 ms for the tree. The host decode, scale and encode
that follows is not traced. An earlier measurement on a simulator put the decode at about 34 ms.
- **Web** comes from `screenshot` and `ariaSnapshot` spans in real runs over two weeks, with p90 of
170 ms. That figure is the raw PNG only. Scaling and encoding come after it and are not traced.

Over a whole run: the 3-minute session used as the frame proof had 51 captures. At these rates that
is about 9 s of screenshot time on Android (around 5% of the run), and about 5–6 s on iOS or web.

Pulling the frames from the video instead costs nothing while the run is going. After the run,
`trailblaze strings extract` saved all 51 frames from that session's 3-minute recording in 8 s,
including JVM start.

## Where it applies

- **Only when the run records video.** Video is off by default (`capture-video`). Recording has its
own cost, which this entry does not measure, so the saving is only net for runs that record
anyway, or once recording is shown to cost less than the screenshots it replaces. That cost is
measured in [Take Report Pictures from the Video](2026-09-30-video-instead-of-screenshots.md).
- **Only for captures whose picture nobody needs during the run.** A capture that goes to the model
needs its screenshot right away, since the model reads the image. The candidates are log
captures: replayed recorded steps, asserts, and the "before" picture of an action that the report
shows.
- **Android already moves part of this off the critical path.** For action logs, the scale and
encode are deferred until the bytes are first read. What still sits in the capture is the grab
itself: on-device `screencap` of a raw frame took 85–134 ms, including starting the process.

## What didn't work

Fresh timings on the laptop. Other sessions pushed the load average past 400, and iOS calls
measured then took seconds (`axe describe-ui` 5–6 s, `simctl io screenshot` 5–13 s). That is why
the iOS and web figures come from traces recorded in real runs, and why the Android A/B is given
with both of its passes.

## Open questions

- What does recording cost per platform? Answered in
[Take Report Pictures from the Video](2026-09-30-video-instead-of-screenshots.md): free during the run on
iOS, about 100 ms per tool call on web, and on an Android emulator free on a still screen but
costly while the screen animates.
- These are laptop and emulator numbers. Farm devices should be measured too before this is made
the default.
- iOS on AXe has too few traced screenshots to report (n=2).

## Future work

- A switch that skips the in-run screenshot for log captures when video is recording, relying on
the frames saved from the recording.
Loading
Loading