Code, data and pre-registrations for a study of what cache-aware Mixture-of-Experts routing does to free-running text generation, and of whether the metrics normally used to check would notice.
Cache-aware routing biases an MoE router toward experts already resident in fast memory, and it works: Skliar et al. (TMLR 2025) report a >50% cut in cache-miss rate at a perplexity cost of 0.1% to 3%, with 2× on-device speedups. Every published evaluation of it we could find scores the model as a reader: teacher-forced perplexity, multiple-choice accuracy, final-answer accuracy. We found none that computes a quality metric on text the model actually wrote. This is that evaluation, on 432 pre-registered generations. Two of its three findings are corrections to ourselves.
| Finding | What the data show |
|---|---|
| Our pre-registered hypothesis is refuted. | We predicted perplexity would be blind to damage visible in writing. It is not: it responds at λ = 0.1, well inside the deployable band and far below the λ = 0.5 the original paper works at. |
| The standard degeneration metric is the blind one. | Across the band where grammar visibly fails, judge NLL has a standardised paired effect of +2.48 and repetition rate has −0.27. Repetition's interval does not exclude zero until λ = 0.4, after the model has begun truncating and dropping out of English. |
| distinct-n adds nothing. | Within a generation it is 1 − repetition-n by construction. An evaluation reporting both reports one number twice. |
| The natural replacement metric fails on loops. | Scoring output under the untouched model rates a looped phrase at 0.180 and real human prose at 2.556. This confirms Holtzman et al. (2020) §4.3 in a new setting, in a paper we cite and did not read closely enough. |
| Per-user expert storage: the obvious statistic is the wrong one. | A per-sequence top-25% subset absorbs 84.4% of that sequence's selections, but one shared subset drops to 62.3%. Locality is real and short-horizon: a cache serves it, a persisted per-user store does not. |
| Implementation dominates algorithm. | The routing bias costs +0.465 ms/token when fused on-device and +10.298 ms/token when implemented the obvious way, a factor of 22.2, caused by 2.25 host synchronisations per layer per token. |
Two claims are withdrawn in the paper from its own earlier drafts, with the numbers that forced each: see §4.1 and §5.1.
No number in the paper was typed by hand. Each carries a tag in a LaTeX comment on the line above
it, and paper/verify_numbers.py resolves every tag to the file and the key it came from.
python3 paper/verify_numbers.py # the full provenance table
python3 paper/verify_numbers.py --check-tex # audit paper.tex against it
python3 paper/verify_numbers.py --text 0.3,1,0 # decode one generation (see below)--check-tex requires that every tag resolves and that the value actually appears near its tag.
It should report zero unresolved tags and zero mismatches. During drafting it caught two
transcription slips, and it re-ran clean after the estimator correction in §3.6 changed every
interval in the paper. paper/provenance.txt is its committed output.
--text decodes the stored tokens of one generation, so the prose quoted in §4.2 can be read
from the same file as the numbers. It needs the transformers package and the tokenizer of the
Ling-mini-2.0-3bit checkpoint, which it loads from models/Ling-mini-2.0-3bit under the
repository root. The checkpoint is not distributed here.
Two classes of number are not regenerated from a result file and are marked as such in the
table: figures quoted from a cited paper (tags X*, each carrying the sentence it came from), and
the three audit-finding tallies in §6 (tags A*), which are counts recorded in the audit reports
and are read back from those files.
paper/
paper.tex LaTeX source, self-contained (no \input, no images, inline bibliography)
verify_numbers.py regenerates every number; --check-tex audits the paper against it
provenance.txt/.json its committed output, 342 entries
data/
results/ the 17 result files the paper cites, and nothing else
gen_prompts.json 12 instructions across 6 categories
corpora/ held-out text for the reader curve and for measuring rho
src/ the experiment scripts (mapped to the paper below)
notes/
gen/PREREG.md the design, fixed before any code was written
gen/RESULT.md the first write-up
exp12/AUDIT.md hostile audit of experiments 1 and 2
v2/PREREG.md pre-registration of the v2 reruns
v2/AUDIT.md hostile audit of the v2 reruns
speed/PREREG.md pre-registration of the wall-clock cost measurement
speed/AUDIT.md hostile audit of the wall-clock cost measurement
tf-collapse/PREREG.md pre-registration of the teacher-forced collapse experiment
PRIOR-ART.md the prior-art check that found the method was not novel
| Paper | Result file | Script |
|---|---|---|
| Table 2, health gate | gen_health.json |
src/gen_health.py |
| Tables 3, 4, 9 and 11, and the prose quoted in §4.2 | gen_ling3_c25.json, gen_ling3_c50.json |
src/gen_bench.py |
| Table 5, reader curve with intervals | gen_ling3_ppl_windows.json |
not in this repository (see below) |
| Table 8, judge gate | judge_gate.json |
src/judge_gate.py |
| Table 10, per-sequence locality | workingset_probe.json; support figures from ra_regimes_ling_unique.json and ra_regimes_olmoe_unique.json |
src/workingset_probe.py; src/ra_regimes.py |
| §5.2, cost of the technique | sp_bench.json; savings model from p3_summary_ALL.json |
src/sp_bench.py and src/sp_impl.py, summarised by src/sp_report.py; src/p3_analyze.py |
| §6, measurement pathologies | tf_collapse.json, tf_collapse_v1_supersededbasis.json, tf_collapse_analysis.json, v2_bench.json, v2_collapse.json |
src/tf_collapse.py, src/tf_analyze.py, src/v2_bench.py, src/v2_collapse.py, summarised by src/v2_report.py |
| Appendix A, reader side of the excluded model | v2_bench.json |
src/v2_bench.py |
| Appendix D, the withdrawn first attempt | gen_ling_c25.json, gen_olmoe_c25.json |
src/gen_bench.py |
Tables 6 and 7 are read off Tables 4 and 5, and Table 1 is a reading of the cited papers; none of
the three cites a file of its own. src/gen_metrics.py holds the degeneration metrics and the
gate that shows they fire; src/memguard.py is the memory ceiling every script runs under.
One result file, gen_ling3_ppl_windows.json, was produced by a script that is not in this
repository: the re-run of the reader arm that stored every window (§3.6, D3). It is included
because the paper cites it, and every number the paper takes from it is listed in the provenance
table with the key it was read from. src/tf_analyze.py, src/workingset_probe.py and
src/p3_analyze.py are the versions that produced their result files, taken from the project's
history before the working directory was renamed.
The notes are the project's working records and keep the vocabulary they were written in. The
"toll" and "the invention" are the routing bias the paper calls Cache-Prior; γ is that bias in the
units of our implementation, related to the paper's λ by γ = λρ with ρ measured per model
(§3.2); ~/tollgate was the working directory. The notes also cite scripts, reports and result
files from earlier phases of the project (p3_toll_ppl.py, dt_bench.py, rc_collapse.py,
p4_ling_ppl.py, notes/phase1234/, notes/dt-bench/, v2_items.json) that are not included
here; the repository holds what the paper cites. The audits are reproduced as written.
The provenance tooling resolves paths relative to its own location and runs anywhere:
git clone https://github.com/arjvnv/cache-aware-moe-generation
cd cache-aware-moe-generation
pip install numpy scipy
python3 paper/verify_numbers.py --check-texThe experiment scripts in src/ are a different matter, and deliberately so. They resolve inputs
against ~/tollgate, which is the directory they ran in, and they need inputs that are not in
this repository: the model checkpoints under models/; for src/hostile_workingset.py and
src/workingset_probe.py, routing traces under data/traces/; and for src/p3_analyze.py, the
per-corpus perplexity files written by p3_toll_ppl.py that it summarises. They are shipped
unmodified rather than rewritten to match this repository's name: the point of the provenance
discipline is that the code here is the code that produced the result files here, and silently
editing paths afterwards would break exactly the correspondence the paper argues for. Clone to
~/tollgate if you want to run them.
python3 src/gen_health.py # the health gate; run this first, two of three checkpoints fail it
python3 src/gen_bench.py --model Ling-mini-2.0-3bit --ntok 320 --nprompts 12 \
--cache-fracs 0.25 --out data/results/gen_ling3_c25.jsonsrc/memguard.py enforces a hard memory ceiling and refuses to start on a loaded machine. That is
deliberate; TOLLGATE_MEM_GB raises it, up to 9 GB.
One model, one quantisation, English only, n = 24 per cell. No control arm: we did not test whether the damage is specific to cache-conditioning or is what happens when top-k routing is perturbed at all. No human evaluation: the qualitative readings are the author's own and unblinded. No distributional metric: MAUVE was not run. The reader curve rests on six held-out windows. §7 of the paper states these at length; none of them is hidden.
notes/gen/PREREG.md, the design, including the amendment that added the health gate and withdrew the first attempt.paper/paper.tex§4, the result.notes/exp12/AUDIT.md,notes/v2/AUDIT.mdandnotes/speed/AUDIT.md: three hostile audits of earlier work in which every must-fix defect ran in the direction that flattered the method. §6 comes from these.
A concatenation of Unix manual pages from the machine the experiments ran on. Included because exact reproduction of the reader curve needs the exact bytes: it is measured on tokens 60000 to 64096. The individual manual pages remain under their own respective licences, which vary by source package.
A preprint is in preparation. Until it is posted, cite this repository:
@misc{cacheawaremoegeneration,
title = {Evaluating Generation Quality Under Cache-Aware MoE Routing},
author = {Vivek, Arjun},
year = {2026},
note = {Preprint in preparation},
url = {https://github.com/arjvnv/cache-aware-moe-generation}
}Apache License 2.0, except where noted.
data/corpora/man_pages.txt is third-party content, a concatenation of Unix manual pages each
remaining under its own licence, and is not covered by the Apache licence. It is
redistributed only so the reader curve can be reproduced byte-for-byte. See NOTICE. No
model weights are distributed here.