Falsification-Driven Biological Law Engine 2026 — Rejected 194 of 203 Candidates
-
Updated
Oct 3, 2026 - HTML
Falsification-Driven Biological Law Engine 2026 — Rejected 194 of 203 Candidates
Agentic skills for Claude Code and Codex, built from published social-science methods sources. Covers experimental design, computational text analysis, manuscript QA, and transparent reporting.
Evidence cards for narrow public AI, ML and data claims: frozen protocol, hashed sources, honest result status.
A single-file Python CLI that pre-registers AI/ML accuracy claims with SHA-256. Lock the threshold before the data, or it didn't happen.
Code and frozen data from a pre-registered LLM-judge audit, with exploratory analyses of how rating-scale censoring can manufacture difference-in-differences effects. arXiv:2608.27309.
🪞 Catch false positives/negatives in AI eval claims — pre-registration, fair-baseline, small-sample CI. Zero training, zero deps.
ESM-2 protein-retrieval failure-mode study with preregistered evidence, calibration analysis, and explicit curation limits.
KPM-1: pre-registered, hash-verified UK election predictions. SHA-256 committed before voting opened. 65,000-persona synthetic panel; 136 councils. Falsifiable forecast, not retro-explained.
An Implementation Guide for Researchers - DePaul University Open Science SLC, 2025-26
Does cache-aware MoE routing degrade generation, and would the usual metrics notice? 432 pre-registered generations. Every number traced to a result file.
Pre-registered forecast of which stars get a substellar companion in Gaia DR4 (frozen lists, backtest, open evaluation)
Scientific TDD for AI coding agents: pre-registration before data, claims audit before submission, external novelty gate, reproducibility. Claude Code + 4 platforms.
Deterministic, drawdown-controlled ETF allocation framework (Convex Core) — reproducibility companion to the research paper: model engine, report code, tests, and computed result artifacts. Hypothetical/backtested; not investment advice.
Self-supervised vision-only world model for C. elegans — Phase 0 software validation on public data, pre-registered.
Replayable agent trajectories reconstructed from the July 2026 frontier-lab evaluation-containment failures, plus benign controls for measuring false-positive cost.
Resample or reroute after a weak-verifier stop? Pre-registered measurements of recoverable stopping debt on MBPP+, a two-sided action-support gate on BigCodeBench that fails closed, a LiveCodeBench observability ladder, and the exchangeable-actions reference showing realized-maximum gaps carry no selector signal. Artifacts for arXiv:2607.08665v3.
Tools for rigorously testing — and honestly falsifying — claims to decipher ancient undeciphered scripts, starting with Linear A.
Causal effect of contractionary US monetary policy shocks on US employment.
Post-hoc causal attribution over LLM-agent trajectories. Two total effects that both fail, a closed form for coupling once contexts diverge, and a traceability specification read against Regulation (EU) 2024/1689. Pre-registered experiment published in full and not run.
Deterministic 700-evaluation pilot on multi-agent configuration spaces and search operators; real-LLM replication preregistered and pending.
To associate your repository with the pre-registration topic, visit your repo's landing page and select "manage topics."