JIT Context OS: Running autonomous coding loops on local 9B models with sub-3ms RYOW & zero context bloat (EXP-009) #1885
wojciechwiesner
started this conversation in
Show and tell
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Background & The Agent Context Bottleneck
When running autonomous multi-turn agents (especially on local LLMs like Qwen or DeepSeek), long-running sessions inevitably suffer from the "Haystack Tax":
What is JIT Context OS?
We built and open-sourced JIT-Context OS (CERN Zenodo DOI: 10.5281/zenodo.22649542 · GitHub: wojciechwiesner/jit-context).
Instead of accumulating unbounded conversation history or relying on slow external vector DBs, JIT Context introduces a 3-tier cascade governed by 10 Epistemic Invariants:
0.0(speculations cannot pollute ground truth). Only deterministic tool executions (exit_code: 0or verified readbacks) earn authority1.0.Empirical Proof (EXP-009): Local 9B on Apple Silicon vs Google Gemini 3.8 Flash
To test whether local models can handle real multi-module autonomous engineering when context is kept lean, we ran a head-to-head duel on a task fixing 4 distinct root causes across 3 interconnected Python modules (
event_pipeline.py,retry_policy.py,storage.py), evaluated by a blindpytestsuite:Turn Breakdown for Local Qwen 3.8:
read_filecalls in one shot.write_filecalls addressing all 4 root causes simultaneously.run_tests()→ 4/4 passed (100% PASS, exit code 0) on the first attempt.Meanwhile, Gemini 3.8 Flash without JIT got caught in a regression loop, modifying test files instead of source code.
Mapping to Agent Zero
This pattern maps directly into Agent Zero's extension architecture (
/a0/extensions/python/):tool_execute_after/hist_add_tool_result: Intercept heavy tool outputs, record deterministic proofs to SQLite WAL, and keep raw logs out of active memory.before_main_llm_call: Inject a compiled working-set capsule (<1.8k tokens) instead of the accumulated monologue.system_prompt: Enforce strict authority gating (I1–I10).We're exploring packaging this as a drop-in extension for Agent Zero. Would love to hear thoughts from the Agent Zero community on context pruning and local model orchestration!
Repo: https://github.com/wojciechwiesner/jit-context
CERN Zenodo: https://doi.org/10.5281/zenodo.22649542
All reactions