Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -226,7 +226,7 @@ All plugins follow semantic versioning (semver). Key versioning rules:
| claude-setup | 1.1.0 |
| development-codebase-tools | 2.7.0 |
| development-planning | 2.1.0 |
| development-pr-workflow | 2.4.0 |
| development-pr-workflow | 2.5.0 |
| development-productivity | 2.5.0 |
| skill-management | 1.3.0 |
| spec-workflow | 2.1.0 |
Expand Down
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "development-pr-workflow",
"version": "2.4.0",
"version": "2.5.0",
"description": "Pull request review, issue resolution, and Graphite stack management",
"author": {
"name": "Uniswap Labs",
Expand Down
Original file line number Diff line number Diff line change
@@ -1,19 +1,6 @@
---
name: backtest-change
description: >
Validate a data-driven change against LIVE historical data before it ships —
replay old-vs-new over a real window, report whether it achieves its goal, and
refuse to ship when the data disproves the premise. Fires whenever someone
proposes a measurable change and names a number: "add a monitor at 700MB",
"set the threshold to N", "warn at X / critical at Y", "alert when it exceeds
N", "raise the timeout to 5s", "change the sampling rate", "bump the cache
TTL", "tighten this alert", "loosen the threshold", "this should reduce the
noise", "that will fix the p95" — and before opening any PR for a monitor
threshold, alert routing or renotify cadence, metric/log/trace query, sampling
rate, rate limit, autoscaling parameter, or a perf change with a latency or
throughput target. A proposed number is a hypothesis, not a decision: backtest
it and let the data override it. Always report old N vs new M with the window
and data source. The /backtest-change command loads this same skill.
description: Validate a data-driven change against LIVE historical data before it ships — replay old-vs-new over a real window, report whether it achieves its goal, and refuse to ship when the data disproves the premise. Fires whenever someone proposes a measurable change and names a number — "add a monitor at 700MB", "set the threshold to N", "warn at X / critical at Y", "raise the timeout to 5s", "change the sampling rate", "tighten this alert", "this should reduce the noise" — and before opening any PR for a monitor threshold, alert routing or renotify cadence, metric/log/trace query, sampling rate, rate limit, autoscaling parameter, or a perf change with a latency or throughput target. Fires for a brand-new monitor, deploy gate, or alert just as it does for an edit to an existing one, including the joint replay of every condition in a composite (e.g. rate AND floor together). A proposed number is a hypothesis, not a decision — backtest it and let the data override it.
allowed-tools: Bash, Read, Grep, Glob, AskUserQuestion
model: opus
---
Expand Down Expand Up @@ -92,6 +79,10 @@ post-merge validation.
both the old and the new condition against the historical series and count
transitions. Identify *which groups/series* change, not just the aggregate.
For a brand-new monitor, "old" is 0 — say so explicitly rather than omitting it.
For a composite (a gate combining, say, a rate condition and a traffic floor),
replay the conditions *jointly* over the same window. Each one alone will fire
on windows the composite would have suppressed, so per-condition counts
overstate what the gate actually does.

5. **Separate the two populations.** The threshold's whole job is to divide
incident from healthy. Report the highest *legitimate* value observed and the
Expand Down
Loading