Skip to content

Latest commit

 

History

5 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Evaluating Generation Quality Under Cache-Aware MoE Routing

provenance generations pre-registered status licence

Code, data and pre-registrations for a study of what cache-aware Mixture-of-Experts routing does to free-running text generation, and of whether the metrics normally used to check would notice.

Cache-aware routing biases an MoE router toward experts already resident in fast memory, and it works: Skliar et al. (TMLR 2025) report a >50% cut in cache-miss rate at a perplexity cost of 0.1% to 3%, with 2× on-device speedups. Every published evaluation of it we could find scores the model as a reader: teacher-forced perplexity, multiple-choice accuracy, final-answer accuracy. We found none that computes a quality metric on text the model actually wrote. This is that evaluation, on 432 pre-registered generations. Two of its three findings are corrections to ourselves.

Findings

Finding What the data show
Our pre-registered hypothesis is refuted. We predicted perplexity would be blind to damage visible in writing. It is not: it responds at λ = 0.1, well inside the deployable band and far below the λ = 0.5 the original paper works at.
The standard degeneration metric is the blind one. Across the band where grammar visibly fails, judge NLL has a standardised paired effect of +2.48 and repetition rate has −0.27. Repetition's interval does not exclude zero until λ = 0.4, after the model has begun truncating and dropping out of English.
distinct-n adds nothing. Within a generation it is 1 − repetition-n by construction. An evaluation reporting both reports one number twice.
The natural replacement metric fails on loops. Scoring output under the untouched model rates a looped phrase at 0.180 and real human prose at 2.556. This confirms Holtzman et al. (2020) §4.3 in a new setting, in a paper we cite and did not read closely enough.
Per-user expert storage: the obvious statistic is the wrong one. A per-sequence top-25% subset absorbs 84.4% of that sequence's selections, but one shared subset drops to 62.3%. Locality is real and short-horizon: a cache serves it, a persisted per-user store does not.
Implementation dominates algorithm. The routing bias costs +0.465 ms/token when fused on-device and +10.298 ms/token when implemented the obvious way, a factor of 22.2, caused by 2.25 host synchronisations per layer per token.

Two claims are withdrawn in the paper from its own earlier drafts, with the numbers that forced each: see §4.1 and §5.1.

Everything here is traceable to a file

No number in the paper was typed by hand. Each carries a tag in a LaTeX comment on the line above it, and paper/verify_numbers.py resolves every tag to the file and the key it came from.

python3 paper/verify_numbers.py                 # the full provenance table
python3 paper/verify_numbers.py --check-tex     # audit paper.tex against it
python3 paper/verify_numbers.py --text 0.3,1,0  # decode one generation (see below)

--check-tex requires that every tag resolves and that the value actually appears near its tag. It should report zero unresolved tags and zero mismatches. During drafting it caught two transcription slips, and it re-ran clean after the estimator correction in §3.6 changed every interval in the paper. paper/provenance.txt is its committed output.

--text decodes the stored tokens of one generation, so the prose quoted in §4.2 can be read from the same file as the numbers. It needs the transformers package and the tokenizer of the Ling-mini-2.0-3bit checkpoint, which it loads from models/Ling-mini-2.0-3bit under the repository root. The checkpoint is not distributed here.

Two classes of number are not regenerated from a result file and are marked as such in the table: figures quoted from a cited paper (tags X*, each carrying the sentence it came from), and the three audit-finding tallies in §6 (tags A*), which are counts recorded in the audit reports and are read back from those files.

Repository structure

paper/
  paper.tex              LaTeX source, self-contained (no \input, no images, inline bibliography)
  verify_numbers.py      regenerates every number; --check-tex audits the paper against it
  provenance.txt/.json   its committed output, 342 entries
data/
  results/               the 17 result files the paper cites, and nothing else
  gen_prompts.json       12 instructions across 6 categories
  corpora/               held-out text for the reader curve and for measuring rho
src/                     the experiment scripts (mapped to the paper below)
notes/
  gen/PREREG.md          the design, fixed before any code was written
  gen/RESULT.md          the first write-up
  exp12/AUDIT.md         hostile audit of experiments 1 and 2
  v2/PREREG.md           pre-registration of the v2 reruns
  v2/AUDIT.md            hostile audit of the v2 reruns
  speed/PREREG.md        pre-registration of the wall-clock cost measurement
  speed/AUDIT.md         hostile audit of the wall-clock cost measurement
  tf-collapse/PREREG.md  pre-registration of the teacher-forced collapse experiment
  PRIOR-ART.md           the prior-art check that found the method was not novel

What produced each table

Paper Result file Script
Table 2, health gate gen_health.json src/gen_health.py
Tables 3, 4, 9 and 11, and the prose quoted in §4.2 gen_ling3_c25.json, gen_ling3_c50.json src/gen_bench.py
Table 5, reader curve with intervals gen_ling3_ppl_windows.json not in this repository (see below)
Table 8, judge gate judge_gate.json src/judge_gate.py
Table 10, per-sequence locality workingset_probe.json; support figures from ra_regimes_ling_unique.json and ra_regimes_olmoe_unique.json src/workingset_probe.py; src/ra_regimes.py
§5.2, cost of the technique sp_bench.json; savings model from p3_summary_ALL.json src/sp_bench.py and src/sp_impl.py, summarised by src/sp_report.py; src/p3_analyze.py
§6, measurement pathologies tf_collapse.json, tf_collapse_v1_supersededbasis.json, tf_collapse_analysis.json, v2_bench.json, v2_collapse.json src/tf_collapse.py, src/tf_analyze.py, src/v2_bench.py, src/v2_collapse.py, summarised by src/v2_report.py
Appendix A, reader side of the excluded model v2_bench.json src/v2_bench.py
Appendix D, the withdrawn first attempt gen_ling_c25.json, gen_olmoe_c25.json src/gen_bench.py

Tables 6 and 7 are read off Tables 4 and 5, and Table 1 is a reading of the cited papers; none of the three cites a file of its own. src/gen_metrics.py holds the degeneration metrics and the gate that shows they fire; src/memguard.py is the memory ceiling every script runs under.

One result file, gen_ling3_ppl_windows.json, was produced by a script that is not in this repository: the re-run of the reader arm that stored every window (§3.6, D3). It is included because the paper cites it, and every number the paper takes from it is listed in the provenance table with the key it was read from. src/tf_analyze.py, src/workingset_probe.py and src/p3_analyze.py are the versions that produced their result files, taken from the project's history before the working directory was renamed.

Reading the notes

The notes are the project's working records and keep the vocabulary they were written in. The "toll" and "the invention" are the routing bias the paper calls Cache-Prior; γ is that bias in the units of our implementation, related to the paper's λ by γ = λρ with ρ measured per model (§3.2); ~/tollgate was the working directory. The notes also cite scripts, reports and result files from earlier phases of the project (p3_toll_ppl.py, dt_bench.py, rc_collapse.py, p4_ling_ppl.py, notes/phase1234/, notes/dt-bench/, v2_items.json) that are not included here; the repository holds what the paper cites. The audits are reproduced as written.

Reproducing

The provenance tooling resolves paths relative to its own location and runs anywhere:

git clone https://github.com/arjvnv/cache-aware-moe-generation
cd cache-aware-moe-generation
pip install numpy scipy
python3 paper/verify_numbers.py --check-tex

The experiment scripts in src/ are a different matter, and deliberately so. They resolve inputs against ~/tollgate, which is the directory they ran in, and they need inputs that are not in this repository: the model checkpoints under models/; for src/hostile_workingset.py and src/workingset_probe.py, routing traces under data/traces/; and for src/p3_analyze.py, the per-corpus perplexity files written by p3_toll_ppl.py that it summarises. They are shipped unmodified rather than rewritten to match this repository's name: the point of the provenance discipline is that the code here is the code that produced the result files here, and silently editing paths afterwards would break exactly the correspondence the paper argues for. Clone to ~/tollgate if you want to run them.

python3 src/gen_health.py       # the health gate; run this first, two of three checkpoints fail it
python3 src/gen_bench.py --model Ling-mini-2.0-3bit --ntok 320 --nprompts 12 \
    --cache-fracs 0.25 --out data/results/gen_ling3_c25.json

src/memguard.py enforces a hard memory ceiling and refuses to start on a loaded machine. That is deliberate; TOLLGATE_MEM_GB raises it, up to 9 GB.

What this study cannot establish

One model, one quantisation, English only, n = 24 per cell. No control arm: we did not test whether the damage is specific to cache-conditioning or is what happens when top-k routing is perturbed at all. No human evaluation: the qualitative readings are the author's own and unblinded. No distributional metric: MAUVE was not run. The reader curve rests on six held-out windows. §7 of the paper states these at length; none of them is hidden.

Reading order, if you want the argument rather than the numbers

  1. notes/gen/PREREG.md, the design, including the amendment that added the health gate and withdrew the first attempt.
  2. paper/paper.tex §4, the result.
  3. notes/exp12/AUDIT.md, notes/v2/AUDIT.md and notes/speed/AUDIT.md: three hostile audits of earlier work in which every must-fix defect ran in the direction that flattered the method. §6 comes from these.

On data/corpora/man_pages.txt

A concatenation of Unix manual pages from the machine the experiments ran on. Included because exact reproduction of the reader curve needs the exact bytes: it is measured on tokens 60000 to 64096. The individual manual pages remain under their own respective licences, which vary by source package.

Citation

A preprint is in preparation. Until it is posted, cite this repository:

@misc{cacheawaremoegeneration,
  title  = {Evaluating Generation Quality Under Cache-Aware MoE Routing},
  author = {Vivek, Arjun},
  year   = {2026},
  note   = {Preprint in preparation},
  url    = {https://github.com/arjvnv/cache-aware-moe-generation}
}

Licence

Apache License 2.0, except where noted.

data/corpora/man_pages.txt is third-party content, a concatenation of Unix manual pages each remaining under its own licence, and is not covered by the Apache licence. It is redistributed only so the reader curve can be reproduced byte-for-byte. See NOTICE. No model weights are distributed here.

About

Does cache-aware MoE routing degrade generation, and would the usual metrics notice? 432 pre-registered generations. Every number traced to a result file.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages