Skip to content
NeoZorKPublic

About

Independent verifier for trading strategies written by AI agents and humans: catches look-ahead bias, hidden costs and overfitting. MCP server, CLI, GitHub Action, signed certificates.

Topics

Resources

Contributing

Security policy

Stars

8 stars

Watchers

1 watching

Forks

Latest commit

 

History

178 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Monte-Neo: the independent verifier for trading strategies written by AI agents and humans

The independent verifier for trading strategies written by AI agents and humans.
Catch look-ahead bias, hidden trading costs and overfitting before a backtest reaches your money.

PyPI version Total downloads Downloads per month Python versions CI MIT license
MCP Registry Works with coding agents GitHub stars

Docs · Quick start · Use from agents · Verifier API · Trap Suite · Changelog


The problem

A coding agent can turn a trading idea into a backtest in minutes. It will then tell you the strategy returns 40% a year with a Sharpe of 3. Most of the time that number is wrong, for the same few reasons:

  • Look-ahead bias. The code reads future bars: shift(-1), centred windows, bfill, statistics over the whole series, np.gradient, an FFT filter.
  • Missing costs. The edge is smaller than fees and slippage, or it disappears when the fill comes one bar later.
  • Selection bias. The agent tried 300 variants and reports the best one as if it were the only one.

Backtest libraries run whatever code you give them. None of them tell you the backtest itself is broken.

The solution

Monte-Neo checks the backtest, not the idea. Give it the price data and the strategy code or its positions. It returns one of four verdicts, the checks behind the verdict, concrete next steps and a reproducible, optionally signed certificate.

monte-neo verify rejects a leaky agent strategy, the agent fixes it, and the honest verdict follows

The agent's strategy used shift(-1), so it knew the next close. Monte-Neo found the leak in four independent ways and named the line. After the fix, no look-ahead is left, and the verifier tells the truth: on a random walk, the strategy has no edge after costs.

The same run as a report for people (--html report.html: one file, no scripts, no network):

HTML report of a rejected strategy: the verdict in one sentence, a card per check family, the summary, the equity curve against buy and hold, the drawdown and the return distribution

More reports, including an honest strategy and a parameter search, are in the gallery.

Text output of the first run
$ monte-neo verify --ohlcv prices.csv --strategy agent_strategy.py --n-trials 40
REJECT  certificate 7f0603b0e604bea5
  check                    category    status  summary
  data_integrity           integrity   pass    OHLCV is clean
  lookahead_truncation     lookahead   fail    truncation probe: LEAK DETECTED
  lookahead_perturbation   lookahead   fail    future-perturbation probe: LEAK DETECTED
  lookahead_static_lint    lookahead   fail    static lint: negative_shift
  implausible_accuracy     lookahead   fail    next-bar hit rate 1.000 over 2999 bars (z 54.8)
  net_profitability        economics   fail    net total return -64.97% after costs
  deflated_sharpe          statistics  fail    deflated Sharpe 0.000 over 40 trial(s)
  ...                                          (16 more checks)
→ The signal at bar t changes when later bars are removed: compute features only from rows <= t
  (no shift(-k), centered windows, bfill or full-sample stats).
→ Fix the flagged source lines (negative shift, center=True, backward fill) and re-run verify. Lines: 6.

The run used a synthetic random walk; output shortened.

Verdict Meaning CLI exit code
PASS No problems found 0
PASS_WITH_WARNINGS Usable; read the warnings 0
NEEDS_MORE_EVIDENCE Too few trades, or the Sharpe does not survive the number of variants tried 1
REJECT The backtest is broken or loses money after costs 2

New: catch look-ahead that hides in milliseconds

monte-neo verify --quotes: a fast strategy earns +191% on exchange time and loses 23% on arrival time; a slow honest strategy keeps its profit

The problem. A quote has two times: when the exchange stamped it and when it reached you. Most backtests on intraday quotes bin them by the stamp, so a fast strategy trades on prices it could not have seen yet. No code lint finds this: the leak is in the clock, not in the code.

What Monte-Neo does (monte-neo verify --quotes quotes.csv --strategy my_strategy.py):

  • Arrival look-ahead. The same strategy sees bars built from the arrival time and trades against the market. If the profit exists only with zero latency, it is a leak at millisecond scale.
  • The delay where the profit vanishes. One number you can act on: "the profit disappears at 38 ms of extra delay".
  • Latency Monte Carlo. Each quote's latency is redrawn from the latency you actually observed (inside its venue), 200 seeded draws: the distribution of the return and the probability of a loss. It shows when a profit was lucky latency.
  • Order delay. --order-latency-ms delays every fill, so an edge that lives in instant execution is rejected.
  • No server near the exchange? --latency-model lognormal:8,25 assumes the latency (median and p95 in ms) instead of reading it, and the certificate says plainly that it is assumed, not measured.
  • Several feeds in one strategy. --symbol and --feeds let a strategy trade one instrument on the prices of another (lead-lag, futures against spot); each feed has its own latency. A leader that arrives after the follower moved is caught.
  • Look-ahead probes too. The truncation and perturbation probes of verify_strategy run on the arrival bars, so a strategy that reads the next bar (shift(-1)) is rejected: that leak exists on every clock.
  • Quote quality. Crossed quotes, negative latency, out-of-order arrivals and bursty latency: a congested path (VPN, Wi-Fi) holds messages and releases them in bursts, and the check says so instead of reporting your network as a result.
  • A certificate. Signed, and reproducible with --recheck like every other certificate.
$ monte-neo verify --demo-quotes     # no files: a fast strategy and an honest one
$ python -m monte_neo.data.quote_recorder --symbol BTCUSDT --seconds 600 --out quotes.csv   # your own latency

What sets it apart: it is an audit, not an environment. It takes a finished strategy from any source (any signal(df) function, or an agent's code) and your own quote recording, and says whether the profit survives the time the data arrived, with a certificate others can reproduce. The strategy does not have to be written for a particular engine.

Honest scope: experimental. The latency checks are context (warn at most) and were calibrated on synthetic data, so record real latency on the machine that will trade, or state an assumed latency and read the verdict as conditional on it. Fills are taker orders at the mid, with the data latency from the quotes (or assumed) plus a fixed order delay. Queue position and partial fills are not modelled, on purpose (they need the order book and a passive-order model; a rough formula would give false precision), so a market-making strategy cannot be validated here; the certificate lists these assumptions. Read the guide.

Search for an indicator, honestly

monte-neo discover --ohlcv data.csv --out found/

Tries thousands of causal formulas, counts the search, tests it against shuffled markets and a lockbox, and writes a certificate for the winner. On random walks it names nothing (0 of 60 searches in the calibration), although a plain search would report a Sharpe above 6. See the guide.

The rest of the workflow

  • verify --ledger counts every variant you try; verify --suggest-fix rewrites leaks as a diff (guide).
  • Signals must not repaint: a signal that changes after it was shown fails, and a signal that needs the bar's own prices is reported (guide).
  • history, register, oracle, portfolio, doctor: certificate history with a regression exit code, pre-registered hypotheses, a Thresholdout hold-out for agents, many strategies at once, and known problems of data exports (guide).
  • The honesty scorecard shows how often the verifier catches leaks and spares honest strategies (hypothesis zoo, with intervals and known gaps).

What it checks

Family Checks
Look-ahead Truncation probe (does bar t change when later bars are removed?), future-perturbation probe, outside-data watch (files or network read by the strategy), static AST lint (23 rules), implausible hit rate
Economics Net return after commission and slippage (with optional funding, short-borrow fees, stops inside the bar and per-symbol costs for universes), break-even cost in bps, one- and two-bar execution delay, a spread estimate from high and low next to the modeled cost, capacity from volume (for universes with the tightest symbols named)
Statistics Probabilistic and Deflated Sharpe priced by n_trials, Monte Carlo timing test (does the signal beat shifted copies of itself, or just ride the market?), sample size, holdout consistency, bootstrap confidence intervals, the minimum track record length and the result of each of six equal windows; for grid searches, walk-forward out-of-sample, parameter-plateau, probability-of-backtest-overfitting (PBO) and Reality Check / SPA checks
Integrity Broken OHLCV (NaN, bad prices, bars out of time order), bad data that makes fake profit (one-bar spikes, frozen prices, unadjusted splits, gaps in time), non-deterministic signals, survivorship bias in a universe
Latency (quotes) Arrival look-ahead (profit on exchange time against arrival time), the delay in ms where the profit vanishes, a Monte Carlo over the observed latency distribution, order delay, quote quality including bursty latency. Guide
Claims --claim: the Sharpe, return, drawdown and trade count that were reported, checked against the verified ones (an overclaim fails)
Context Buy-and-hold on the same data and costs, results by year / quarter / month and by market regime, a warning when one period makes all the profit

Every rule is backed by the Trap Suite: 80 strategies that are known to lie and 35 honest controls. It runs on every build, so the verifier cannot silently stop catching a leak or start accusing honest code.

How it differs from other tools

Kind of tool What it does What Monte-Neo adds
Performance report libraries Charts and ratios from a return series Checks whether the backtest behind the returns is broken, prices the number of variants tried, issues a certificate
Backtesting frameworks Run whatever strategy code they are given An independent second opinion on the result; adapters for vectorbt, Backtrader, backtesting.py, bt, Nautilus Trader, Zipline, Freqtrade, Lean and plain fills (the first six are tested against the real frameworks)
Latency simulation inside a backtest Models feed and order latency as part of the backtest run, for strategies written for that environment Audits a finished strategy from any source on your own quote recording: where its profit vanishes, a Monte Carlo over the real latency distribution, a signed certificate
Look-ahead checkers inside one framework Compare indicator values on cut data for that framework's strategies Works with any strategy function or positions, adds costs, selection bias and claim checks
Statistics libraries (Deflated Sharpe, PBO) Formulas you wire up yourself The whole pipeline: probes, engine, statistics, report, certificate, agent tools

Where to use it

You are… Use Monte-Neo to…
Building strategies with Claude Code, Codex, Gemini CLI or Cursor Make the agent verify its own backtest before it reports results. The MCP server and the Claude Code plugin do this automatically.
Running a strategy repository Add the GitHub Action. A pull request whose backtest leaks or loses money after costs fails CI, and the verdict is posted as a PR comment.
A quant, reviewer or allocator Check a strategy someone else sends you, with their data and code, in one command. Re-check or verify the signature of the certificate they hand over.
A prop firm, strategy marketplace or trading course Screen submissions before a human looks at them. Publish signed certificates next to listed strategies.
A researcher comparing agents Run the Honesty Bench: the same tasks for every agent, scored by how often each one claims profit that is not there.

Why Monte-Neo

  • Independent. It checks code it did not write, with probes that do not trust the strategy's own numbers.
  • Careful with accusations. 35 honest strategies (loops, windows, resampling, fits inside rolling windows, a real edge with a high hit rate) must never be flagged for look-ahead, on every build.
  • Built for agents. An MCP server, a Claude Code plugin with a skill, a slash command and a reminder hook, plus rules for Codex, Gemini CLI and Cursor. Every failed check returns a next_action the agent can act on.
  • Time-aware. It checks whether the profit survives the time the quotes really arrived, not only the time the exchange stamped them.
  • Reproducible. The same data, code and n_trials always give the same certificate_id. Anyone can reproduce a certificate with --recheck.
  • Signed. Ed25519 signatures show who issued a certificate and that nobody edited it.
  • Honest about selection bias. Declare how many variants you tried, or let verify_grid count them for you. The Deflated Sharpe prices them in.
  • Local and private. Your data and code never leave your machine. MIT licensed.

Quick start

pip install monte-neo
monte-neo verify --demo      # no files needed: a strategy that peeks at tomorrow's close is rejected
monte-neo init-ci            # write a GitHub Actions workflow that verifies your strategy on every pull request

Command line

monte-neo verify --ohlcv btc_1h.csv --strategy my_strategy.py --n-trials 12 --out verdict.json

my_strategy.py defines signal(df), which returns one position per bar: +1 long, 0 flat, -1 short, or a weight such as 0.5 (half the equity long). The position decided on bar t is filled at the open of bar t + 1.

def signal(df):
    fast = df["close"].rolling(20).mean()
    slow = df["close"].rolling(80).mean()
    return (fast > slow).astype(int)

Python

from monte_neo.verify import verify_strategy

report = verify_strategy("btc_1h.csv", strategy="my_strategy.py", n_trials=12)
print(report["verdict"], report["next_actions"])

Parameter search with honest trial counting

monte-neo verify --ohlcv btc_1h.csv --strategy sma.py --grid '{"fast": [10, 20], "slow": [80, 120]}'

Several symbols, and a report for people

# universe.csv: timestamp, symbol, open, high, low, close; signal(df) returns a weight per row
monte-neo verify --ohlcv universe.csv --strategy xs_momentum.py --html report.html

The HTML report is a tear sheet in one file with no scripts: the reason for the verdict in one sentence, equity on a log scale against buy-and-hold, drawdown, rolling Sharpe, a monthly heat map, the timing test, sensitivity to costs, trade statistics, the evidence for a leak (bars and source lines), every check, results by period and market regime, and the hashes to reproduce the run. It prints to PDF from the browser. See the example reports.

Use it from your coding agent

Claude Code (plugin with the MCP server, the verify-strategy skill and /verify):

/plugin marketplace add NeoZorK/Monte-Neo
/plugin install monte-neo@monte-neo

Any MCP client (Codex, Gemini CLI, Cursor, and others):

uvx monte-neo mcp

It is also listed in the official MCP Registry as io.github.NeoZorK/monte-neo. Setup for each client: Use from agents.

Local agents (Qwen Code, or any terminal agent on a model served from your machine through Ollama, LM Studio or llama.cpp):

qwen mcp add monte-neo uvx --from "monte-neo[mcp]>=0.52.0" monte-neo-mcp

qwen-code is also a built-in agent of monte-neo bench run, so you can measure how honest a local model is. A small model often fails the bench tasks; that is a result of the model, not of the verifier. Setup, --base-url and troubleshooting: Local agents.

MCP tools (17): verify_strategy, verify_grid, verify_quotes, verify_portfolio, probe_lookahead, cost_stress, recheck_certificate, check_signature, render_report, suggest_fix, discover_indicator, register_hypothesis, compare_certificates, holdout_query, diagnose_data, verdict_schema, verifier_manifest.

GitHub Action

- uses: NeoZorK/Monte-Neo@v0.52.0
  with:
    ohlcv: data/btc_1h.csv
    strategy: strategies/momentum.py
    n-trials: "12"
    comment: "true"                                        # post the verdict on the pull request
    signing-key: ${{ secrets.MONTE_NEO_SIGNING_KEY }}      # optional: sign the certificate
    upload-certificate: "true"                             # optional: keep it as a workflow artifact

The job fails on REJECT. The verdict and every check appear in the step summary.

Certificates you can check

Each run produces a strategy-verdict/1 JSON certificate. It contains the verdict, every check, the metrics and the SHA-256 hashes of the data, signals and code.

monte-neo verify --recheck verdict.json --ohlcv btc_1h.csv --strategy my_strategy.py  # reproduce it
pip install "monte-neo[sign]"
monte-neo verify --keygen issuer                                  # issuer.key + issuer.pub
monte-neo verify --ohlcv btc_1h.csv --strategy my_strategy.py --sign issuer.key --out verdict.json
monte-neo verify --check-signature verdict.json --public-key issuer.pub

Show that a strategy passed:

Verified by Monte-Neo

[![Verified by Monte-Neo](https://img.shields.io/badge/verified%20by-Monte--Neo-2ea44f)](https://github.com/NeoZorK/Monte-Neo)

Link the badge to the verification page with your certificate and public key (/verify/?cert=<https URL>&key=ed25519:<key>). Readers then check the signature in their browser with one click; nothing is uploaded.

Install options

pip install monte-neo              # verifier, CLI and MCP server
pip install "monte-neo[sign]"      # + Ed25519 certificate signing
pip install "monte-neo[parquet]"   # + read .parquet tables
pip install "monte-neo[research]"  # + interactive research CLI
pip install "monte-neo[plot]"      # + charts
pip install "monte-neo[apple]"     # + Metal / MLX research engine (Apple Silicon)
pip install "monte-neo[full]"      # everything

Python 3.11+ on macOS or Linux.

Research engine (fee-aware bar backtests, sweeps, Monte Carlo)

Monte-Neo started as a fast local research engine for Apple Silicon, and the verifier runs on it. The engine is still available. It is in maintenance mode: bug fixes only.

from monte_neo.backtest import ExecutionModel, export_sma_sweep, synthetic_ohlcv

ohlc = synthetic_ohlcv(100_000, seed=42)
model = ExecutionModel(commission_bps=5.0, slippage_bps=5.0, warmup_bars=50)
out = export_sma_sweep(ohlc["open"], ohlc["high"], ohlc["low"], ohlc["close"], combos=16, model=model, device="auto")
print(out["device"], out["combos"])
  • Next-bar fills, costs in bps, SL/TP/trailing stops, funding, sessions
  • Export API with golden vectors, holdout, walk-forward, CSCV/PBO, Monte Carlo helpers
  • Metal / MLX / Numba device selection with a memory planner that falls back to CPU instead of hanging
  • Paper OMS for event-level validation

Docs: quick start · export API · backtest engine

Project status

Monte-Neo is in active development (beta). The verifier API and the strategy-verdict/1 schema are stable across minor releases. See the roadmap.

Not investment advice. Monte-Neo checks backtest methodology. It does not predict future profit.

Contributing

Found a way a backtest fooled you or your agent? Submit it as a trap. Did the verifier accuse an honest strategy? Report a false accusation. Bug reports and pull requests are welcome; see the contributing guide. Report security issues privately: SECURITY.md.

Citation

The method is described in the technical note.

If Monte-Neo helps your research, please cite it. GitHub shows the citation under Cite this repository (CITATION.cff).

License

MIT

About

Independent verifier for trading strategies written by AI agents and humans: catches look-ahead bias, hidden costs and overfitting. MCP server, CLI, GitHub Action, signed certificates.

Topics

Resources

Contributing

Security policy

Stars

8 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages