The independent verifier for trading strategies written by AI agents and humans.
Catch look-ahead bias, hidden trading costs and overfitting before a backtest reaches your money.
Docs · Quick start · Use from agents · Verifier API · Trap Suite · Changelog
A coding agent can turn a trading idea into a backtest in minutes. It will then tell you the strategy returns 40% a year with a Sharpe of 3. Most of the time that number is wrong, for the same few reasons:
- Look-ahead bias. The code reads future bars:
shift(-1), centred windows,bfill, statistics over the whole series,np.gradient, an FFT filter. - Missing costs. The edge is smaller than fees and slippage, or it disappears when the fill comes one bar later.
- Selection bias. The agent tried 300 variants and reports the best one as if it were the only one.
Backtest libraries run whatever code you give them. None of them tell you the backtest itself is broken.
Monte-Neo checks the backtest, not the idea. Give it the price data and the strategy code or its positions. It returns one of four verdicts, the checks behind the verdict, concrete next steps and a reproducible, optionally signed certificate.
The agent's strategy used shift(-1), so it knew the next close. Monte-Neo found the leak in four
independent ways and named the line. After the fix, no look-ahead is left, and the verifier tells
the truth: on a random walk, the strategy has no edge after costs.
The same run as a report for people (--html report.html: one file, no scripts, no network):
More reports, including an honest strategy and a parameter search, are in the gallery.
Text output of the first run
$ monte-neo verify --ohlcv prices.csv --strategy agent_strategy.py --n-trials 40
REJECT certificate 7f0603b0e604bea5
check category status summary
data_integrity integrity pass OHLCV is clean
lookahead_truncation lookahead fail truncation probe: LEAK DETECTED
lookahead_perturbation lookahead fail future-perturbation probe: LEAK DETECTED
lookahead_static_lint lookahead fail static lint: negative_shift
implausible_accuracy lookahead fail next-bar hit rate 1.000 over 2999 bars (z 54.8)
net_profitability economics fail net total return -64.97% after costs
deflated_sharpe statistics fail deflated Sharpe 0.000 over 40 trial(s)
... (16 more checks)
→ The signal at bar t changes when later bars are removed: compute features only from rows <= t
(no shift(-k), centered windows, bfill or full-sample stats).
→ Fix the flagged source lines (negative shift, center=True, backward fill) and re-run verify. Lines: 6.The run used a synthetic random walk; output shortened.
| Verdict | Meaning | CLI exit code |
|---|---|---|
PASS |
No problems found | 0 |
PASS_WITH_WARNINGS |
Usable; read the warnings | 0 |
NEEDS_MORE_EVIDENCE |
Too few trades, or the Sharpe does not survive the number of variants tried | 1 |
REJECT |
The backtest is broken or loses money after costs | 2 |
The problem. A quote has two times: when the exchange stamped it and when it reached you. Most backtests on intraday quotes bin them by the stamp, so a fast strategy trades on prices it could not have seen yet. No code lint finds this: the leak is in the clock, not in the code.
What Monte-Neo does (monte-neo verify --quotes quotes.csv --strategy my_strategy.py):
- Arrival look-ahead. The same strategy sees bars built from the arrival time and trades against the market. If the profit exists only with zero latency, it is a leak at millisecond scale.
- The delay where the profit vanishes. One number you can act on: "the profit disappears at 38 ms of extra delay".
- Latency Monte Carlo. Each quote's latency is redrawn from the latency you actually observed (inside its venue), 200 seeded draws: the distribution of the return and the probability of a loss. It shows when a profit was lucky latency.
- Order delay.
--order-latency-msdelays every fill, so an edge that lives in instant execution is rejected. - No server near the exchange?
--latency-model lognormal:8,25assumes the latency (median and p95 in ms) instead of reading it, and the certificate says plainly that it is assumed, not measured. - Several feeds in one strategy.
--symboland--feedslet a strategy trade one instrument on the prices of another (lead-lag, futures against spot); each feed has its own latency. A leader that arrives after the follower moved is caught. - Look-ahead probes too. The truncation and perturbation probes of
verify_strategyrun on the arrival bars, so a strategy that reads the next bar (shift(-1)) is rejected: that leak exists on every clock. - Quote quality. Crossed quotes, negative latency, out-of-order arrivals and bursty latency: a congested path (VPN, Wi-Fi) holds messages and releases them in bursts, and the check says so instead of reporting your network as a result.
- A certificate. Signed, and reproducible with
--rechecklike every other certificate.
$ monte-neo verify --demo-quotes # no files: a fast strategy and an honest one
$ python -m monte_neo.data.quote_recorder --symbol BTCUSDT --seconds 600 --out quotes.csv # your own latencyWhat sets it apart: it is an audit, not an environment. It takes a finished strategy from any source (any signal(df)
function, or an agent's code) and your own quote recording, and says whether the profit survives the time the data
arrived, with a certificate others can reproduce. The strategy does not have to be written for a particular engine.
Honest scope: experimental. The latency checks are context (warn at most) and were calibrated on synthetic data, so record
real latency on the machine that will trade, or state an assumed latency and read the verdict as conditional on it. Fills
are taker orders at the mid, with the data latency from the quotes (or assumed) plus a fixed order delay. Queue position and
partial fills are not modelled, on purpose (they need the order book and a passive-order model; a rough formula would give
false precision), so a market-making strategy cannot be validated here; the certificate lists these assumptions.
Read the guide.
monte-neo discover --ohlcv data.csv --out found/Tries thousands of causal formulas, counts the search, tests it against shuffled markets and a lockbox, and writes a certificate for the winner. On random walks it names nothing (0 of 60 searches in the calibration), although a plain search would report a Sharpe above 6. See the guide.
verify --ledgercounts every variant you try;verify --suggest-fixrewrites leaks as a diff (guide).- Signals must not repaint: a signal that changes after it was shown fails, and a signal that needs the bar's own prices is reported (guide).
history,register,oracle,portfolio,doctor: certificate history with a regression exit code, pre-registered hypotheses, a Thresholdout hold-out for agents, many strategies at once, and known problems of data exports (guide).- The honesty scorecard shows how often the verifier catches leaks and spares honest strategies (hypothesis zoo, with intervals and known gaps).
| Family | Checks |
|---|---|
| Look-ahead | Truncation probe (does bar t change when later bars are removed?), future-perturbation probe, outside-data watch (files or network read by the strategy), static AST lint (23 rules), implausible hit rate |
| Economics | Net return after commission and slippage (with optional funding, short-borrow fees, stops inside the bar and per-symbol costs for universes), break-even cost in bps, one- and two-bar execution delay, a spread estimate from high and low next to the modeled cost, capacity from volume (for universes with the tightest symbols named) |
| Statistics | Probabilistic and Deflated Sharpe priced by n_trials, Monte Carlo timing test (does the signal beat shifted copies of itself, or just ride the market?), sample size, holdout consistency, bootstrap confidence intervals, the minimum track record length and the result of each of six equal windows; for grid searches, walk-forward out-of-sample, parameter-plateau, probability-of-backtest-overfitting (PBO) and Reality Check / SPA checks |
| Integrity | Broken OHLCV (NaN, bad prices, bars out of time order), bad data that makes fake profit (one-bar spikes, frozen prices, unadjusted splits, gaps in time), non-deterministic signals, survivorship bias in a universe |
| Latency (quotes) | Arrival look-ahead (profit on exchange time against arrival time), the delay in ms where the profit vanishes, a Monte Carlo over the observed latency distribution, order delay, quote quality including bursty latency. Guide |
| Claims | --claim: the Sharpe, return, drawdown and trade count that were reported, checked against the verified ones (an overclaim fails) |
| Context | Buy-and-hold on the same data and costs, results by year / quarter / month and by market regime, a warning when one period makes all the profit |
Every rule is backed by the Trap Suite: 80 strategies that are known to lie and 35 honest controls. It runs on every build, so the verifier cannot silently stop catching a leak or start accusing honest code.
| Kind of tool | What it does | What Monte-Neo adds |
|---|---|---|
| Performance report libraries | Charts and ratios from a return series | Checks whether the backtest behind the returns is broken, prices the number of variants tried, issues a certificate |
| Backtesting frameworks | Run whatever strategy code they are given | An independent second opinion on the result; adapters for vectorbt, Backtrader, backtesting.py, bt, Nautilus Trader, Zipline, Freqtrade, Lean and plain fills (the first six are tested against the real frameworks) |
| Latency simulation inside a backtest | Models feed and order latency as part of the backtest run, for strategies written for that environment | Audits a finished strategy from any source on your own quote recording: where its profit vanishes, a Monte Carlo over the real latency distribution, a signed certificate |
| Look-ahead checkers inside one framework | Compare indicator values on cut data for that framework's strategies | Works with any strategy function or positions, adds costs, selection bias and claim checks |
| Statistics libraries (Deflated Sharpe, PBO) | Formulas you wire up yourself | The whole pipeline: probes, engine, statistics, report, certificate, agent tools |
| You are… | Use Monte-Neo to… |
|---|---|
| Building strategies with Claude Code, Codex, Gemini CLI or Cursor | Make the agent verify its own backtest before it reports results. The MCP server and the Claude Code plugin do this automatically. |
| Running a strategy repository | Add the GitHub Action. A pull request whose backtest leaks or loses money after costs fails CI, and the verdict is posted as a PR comment. |
| A quant, reviewer or allocator | Check a strategy someone else sends you, with their data and code, in one command. Re-check or verify the signature of the certificate they hand over. |
| A prop firm, strategy marketplace or trading course | Screen submissions before a human looks at them. Publish signed certificates next to listed strategies. |
| A researcher comparing agents | Run the Honesty Bench: the same tasks for every agent, scored by how often each one claims profit that is not there. |
- Independent. It checks code it did not write, with probes that do not trust the strategy's own numbers.
- Careful with accusations. 35 honest strategies (loops, windows, resampling, fits inside rolling windows, a real edge with a high hit rate) must never be flagged for look-ahead, on every build.
- Built for agents. An MCP server, a Claude Code plugin with a skill, a slash command and a reminder hook, plus rules for Codex, Gemini CLI and Cursor. Every failed check returns a
next_actionthe agent can act on. - Time-aware. It checks whether the profit survives the time the quotes really arrived, not only the time the exchange stamped them.
- Reproducible. The same data, code and
n_trialsalways give the samecertificate_id. Anyone can reproduce a certificate with--recheck. - Signed. Ed25519 signatures show who issued a certificate and that nobody edited it.
- Honest about selection bias. Declare how many variants you tried, or let
verify_gridcount them for you. The Deflated Sharpe prices them in. - Local and private. Your data and code never leave your machine. MIT licensed.
pip install monte-neo
monte-neo verify --demo # no files needed: a strategy that peeks at tomorrow's close is rejected
monte-neo init-ci # write a GitHub Actions workflow that verifies your strategy on every pull requestCommand line
monte-neo verify --ohlcv btc_1h.csv --strategy my_strategy.py --n-trials 12 --out verdict.jsonmy_strategy.py defines signal(df), which returns one position per bar: +1 long, 0 flat,
-1 short, or a weight such as 0.5 (half the equity long). The position decided on bar t is
filled at the open of bar t + 1.
def signal(df):
fast = df["close"].rolling(20).mean()
slow = df["close"].rolling(80).mean()
return (fast > slow).astype(int)Python
from monte_neo.verify import verify_strategy
report = verify_strategy("btc_1h.csv", strategy="my_strategy.py", n_trials=12)
print(report["verdict"], report["next_actions"])Parameter search with honest trial counting
monte-neo verify --ohlcv btc_1h.csv --strategy sma.py --grid '{"fast": [10, 20], "slow": [80, 120]}'Several symbols, and a report for people
# universe.csv: timestamp, symbol, open, high, low, close; signal(df) returns a weight per row
monte-neo verify --ohlcv universe.csv --strategy xs_momentum.py --html report.htmlThe HTML report is a tear sheet in one file with no scripts: the reason for the verdict in one sentence, equity on a log scale against buy-and-hold, drawdown, rolling Sharpe, a monthly heat map, the timing test, sensitivity to costs, trade statistics, the evidence for a leak (bars and source lines), every check, results by period and market regime, and the hashes to reproduce the run. It prints to PDF from the browser. See the example reports.
Claude Code (plugin with the MCP server, the verify-strategy skill and /verify):
/plugin marketplace add NeoZorK/Monte-Neo
/plugin install monte-neo@monte-neo
Any MCP client (Codex, Gemini CLI, Cursor, and others):
uvx monte-neo mcpIt is also listed in the official MCP Registry as io.github.NeoZorK/monte-neo. Setup for each
client: Use from agents.
Local agents (Qwen Code, or any terminal agent on a model served from your machine through Ollama, LM Studio or llama.cpp):
qwen mcp add monte-neo uvx --from "monte-neo[mcp]>=0.52.0" monte-neo-mcpqwen-code is also a built-in agent of monte-neo bench run, so you can measure how honest a local model is.
A small model often fails the bench tasks; that is a result of the model, not of the verifier.
Setup, --base-url and troubleshooting: Local agents.
MCP tools (17): verify_strategy, verify_grid, verify_quotes, verify_portfolio, probe_lookahead,
cost_stress, recheck_certificate, check_signature, render_report, suggest_fix, discover_indicator,
register_hypothesis, compare_certificates, holdout_query, diagnose_data, verdict_schema, verifier_manifest.
- uses: NeoZorK/Monte-Neo@v0.52.0
with:
ohlcv: data/btc_1h.csv
strategy: strategies/momentum.py
n-trials: "12"
comment: "true" # post the verdict on the pull request
signing-key: ${{ secrets.MONTE_NEO_SIGNING_KEY }} # optional: sign the certificate
upload-certificate: "true" # optional: keep it as a workflow artifactThe job fails on REJECT. The verdict and every check appear in the step summary.
Each run produces a strategy-verdict/1 JSON certificate. It contains the verdict, every check,
the metrics and the SHA-256 hashes of the data, signals and code.
monte-neo verify --recheck verdict.json --ohlcv btc_1h.csv --strategy my_strategy.py # reproduce it
pip install "monte-neo[sign]"
monte-neo verify --keygen issuer # issuer.key + issuer.pub
monte-neo verify --ohlcv btc_1h.csv --strategy my_strategy.py --sign issuer.key --out verdict.json
monte-neo verify --check-signature verdict.json --public-key issuer.pubShow that a strategy passed:
[](https://github.com/NeoZorK/Monte-Neo)Link the badge to the verification page with your certificate and public
key (/verify/?cert=<https URL>&key=ed25519:<key>). Readers then check the signature in their
browser with one click; nothing is uploaded.
pip install monte-neo # verifier, CLI and MCP server
pip install "monte-neo[sign]" # + Ed25519 certificate signing
pip install "monte-neo[parquet]" # + read .parquet tables
pip install "monte-neo[research]" # + interactive research CLI
pip install "monte-neo[plot]" # + charts
pip install "monte-neo[apple]" # + Metal / MLX research engine (Apple Silicon)
pip install "monte-neo[full]" # everythingPython 3.11+ on macOS or Linux.
Research engine (fee-aware bar backtests, sweeps, Monte Carlo)
Monte-Neo started as a fast local research engine for Apple Silicon, and the verifier runs on it. The engine is still available. It is in maintenance mode: bug fixes only.
from monte_neo.backtest import ExecutionModel, export_sma_sweep, synthetic_ohlcv
ohlc = synthetic_ohlcv(100_000, seed=42)
model = ExecutionModel(commission_bps=5.0, slippage_bps=5.0, warmup_bars=50)
out = export_sma_sweep(ohlc["open"], ohlc["high"], ohlc["low"], ohlc["close"], combos=16, model=model, device="auto")
print(out["device"], out["combos"])- Next-bar fills, costs in bps, SL/TP/trailing stops, funding, sessions
- Export API with golden vectors, holdout, walk-forward, CSCV/PBO, Monte Carlo helpers
- Metal / MLX / Numba device selection with a memory planner that falls back to CPU instead of hanging
- Paper OMS for event-level validation
Docs: quick start · export API · backtest engine
Monte-Neo is in active development (beta). The verifier API and the strategy-verdict/1
schema are stable across minor releases. See the
roadmap.
Not investment advice. Monte-Neo checks backtest methodology. It does not predict future profit.
Found a way a backtest fooled you or your agent? Submit it as a trap. Did the verifier accuse an honest strategy? Report a false accusation. Bug reports and pull requests are welcome; see the contributing guide. Report security issues privately: SECURITY.md.
The method is described in the technical note.
If Monte-Neo helps your research, please cite it. GitHub shows the citation under Cite this repository (CITATION.cff).


