A compact Rust router that puts each OpenAI-compatible inference request on
the healthy GPU replica where it can do the least repeated work.
ramjet sits between your clients and replicated model servers. It keeps conversations and shared system prompts near warm cache state, then lets live load override affinity before one replica becomes a hotspot. Clients keep the same OpenAI API; engines need no ramjet-specific integration.
How ramjet and the engines behind it were tuned for each model, written up on the Helix blog. GLM-5.3 ran on an 8× H200 server; the others on our 8× RTX PRO 6000 server:
| GLM-5.3 (8× H200) | GLM-5.3-Flash | Qwen3.8-Flash-Next | Qwen3.8-27B | DeepSeek V4 |
|---|---|---|---|---|
| 48 coding agents from one server (29 Sep) | Part 1: getting day-zero serving to work (27 Aug) | On eight GPUs: what actually helped (27 Aug) | Chasing a 454 tok/s tweet (22 Aug) | SGLang vs DwarfStar vs vLLM+DSpark (14 Aug) |
| Ramjet vs NVIDIA Dynamo (30 Sep) | Running on 2, 4 or 8 GPUs (14 Sep) | One model, two speeds: smart routing (28 Aug) | Doubling throughput by reading a log line (23 Aug) | V4.1 Flash: encoder, Engram and KV cache (10 Sep) |
| Ran out of cache snapshots, not cache tokens (26 Sep) | Swift 1.5 cut thinking tokens (27 Sep) | A better lm_head, tested and shipped (25 Aug) |
| Reuse more | Queue less | Fail cleanly |
|---|---|---|
| Bounded prefix fingerprints find the replica most likely to reuse prior work. | Size-weighted reservations spread cold prefills and concurrent decodes. | Active health probes, retryable failover, and immediate disconnect cancellation keep capacity honest. |
The ordinary router is stateless, privacy-bounded, and deliberately useful without raw KV-cache events. Optional DSpark enforcement persists only opaque quarantine commitments so an LB restart cannot forget a bad EngineCore. The production path remains the proxy plus your existing OpenAI-compatible engines.
System overview to see how your node is doing:
And specific serving tab:
You can also just plug it into prometheus, /metrics API is available.
Each metric has its own column, so green always means better: throughput and cache hits up, latency down. Red marks a measured cost, grey a change within noise, and a dash a metric that run did not record.
These are workload results, not theoretical peaks. Reproduce the DeepSeek rows from RESULTS.md; inspect every accepted and rejected experiment in EXPERIMENTS.md.
All measured on node06 — 8× RTX PRO 6000 Blackwell; the engine topology is listed per row, and each model name links to its Compose stack. The full-box column reports the best qualified saturation point recorded for that stack, not a shared concurrency level.
| Model | Engine · measured shape | Decode @ c1 | Full-box peak |
|---|---|---|---|
DeepSeek-V4-Flashdeepseek-v4-flashsparse MoE |
vLLM + DSpark 2× TP4 · c24/max256 |
245.1 tok/s | 1,891.2 tok/s |
Qwen3.8-27B FP8qwen3.8-27bdense |
vLLM 2× TP4 · c256/max256 · MTP off |
77 → 121 tok/s |
7,890.9 tok/s |
Qwen3.8-27B NVFP4 + BF16 headqwen3.8-27bdense |
SGLang + DFlash2 8× TP1 · 208 slots · bf16 SSM |
153.3 tok/s greedy median |
not yet requalified Inferact target: 7,882.6 tok/s |
Qwen3.8-Flash-Next FP8qwen3.8-flash-nextsparse MoE |
vLLM 2× TP4+EP · c64 · MTP3 on both |
113 → 202 tok/s |
3,340.5 tok/s |
GLM-5.3-Flash W4A16, FP8 expertsglm-5.3-flashsparse MoE |
SGLang + EAGLE TP2 · c4/max256 |
164.8 tok/s | 388.2 tok/s per 2-GPU replica; box not yet saturated |
No model — and neither Qwen3.8-27B stack — is simply better. Single-stream
decode is what an interactive user feels; the full-box figure is a capacity
landmark for a saturated agent fleet. These maxima come from separate
model-specific workloads, so they are not a matched head-to-head benchmark.
The vLLM row's saturation result has MTP off because speculation improves
low-concurrency latency but wastes rejected drafts once the batch saturates
the GPU. The SGLang row uses RadixArk's immutable
BF16-lm_head checkpoint. Its matched one-engine canary measured 153.3 tok/s
against 142.6 for the former Inferact target (+7.5%), with the same 7/8
objective answers and 20/25 deterministic agent-protocol cases. The smaller
target exposes 26 running slots and 582,246 KV tokens per engine: 208 slots
across the fleet. Full-box saturation has not yet been requalified on these
weights; the former Inferact target reached 7,882.6 tok/s, within 0.1% of the
vLLM reference. On that earlier SGLang stack, bf16 SSM state reduced c128 TTFT
p95 from 3.99s to 0.221s. The same
3-app × 4-session × 2-turn locality run measured 87.3% cached prompt
tokens, and 12 concurrent same-app requests spread across 7 of 8 engines
at 714 tok/s. Its cost is cold long-context prefill: a 196K-token first
turn pays ~57s of TTFT on one GPU, with prefix-cached follow-ups at 2–4s.
Qwen3.8-Flash-Next shows the same speculation trade-off: on 256-token outputs MTP3 adds 79% at c1 but only 7.5% at c32. The qualified pair therefore runs MTP3 on one engine and standard decoding on the other, and ramjet uses the requested output length to pick between them only once cache and load tie. Its full-box figure predates that split, with MTP3 on both engines. GLM-5.3-Flash runs on two GPUs per replica; its prefix cache is bounded by saved linear-attention states rather than KV tokens. Keeping two states per path instead of four, plus a 4 GB host tier, took a probe of 12 cyclic 20k-token sessions from 0/12 to 12/12 cached. The Kev stack adds a 0.8B decision model to the same server, sharing one Qwen GPU behind a second API profile: 77 ms p50 per short three-question request at c1 and 20.3 requests/s at c4, measured beside live traffic rather than saturated. Model profiles covers the sizing, sharding, and speculative-decoding trade-offs behind these numbers.
deploy/glm53_h200 runs the 753B FP8
checkpoint as one SGLang engine with eight data-parallel attention ranks, and
ramjet lists each rank as its own upstream (RJ_UPSTREAM_DP_RANKS). On a
simulated team of continuously working coding agents:
The blog post walks through each step, and Ramjet vs NVIDIA Dynamo compares the router against Dynamo 1.5.0's KV router on the same engine; the raw cells are in EXPERIMENTS.md (2026-09-29 and 2026-09-30).
For existing engines, the upstream list is normally the only setting you need:
services:
ramjet:
image: ghcr.io/helixml/ramjet:v0.8.0@sha256:fe432bbca183d2a457a7713fb150ea5ee36aba7a13f92280ef3ec7195ec23673
restart: unless-stopped
ports:
- "8000:8000" # OpenAI API + /health
- "9090:9090" # Prometheus
environment:
RJ_UPSTREAM: http://model-server-1:8000,http://model-server-2:8000
# RJ_UPSTREAM_TOKEN: ${MODEL_SERVER_API_KEY} # if requireddocker compose up -d
curl --fail http://localhost:8000/healthThe example pins a released image by immutable digest; see
CHANGELOG.md for what each version contains.
Safe defaults enable locality/load routing and keep tokenizer, raw KV-event,
exact-placement, and snapshot paths off. See the complete
configuration table, or start from the
eight-replica Compose stack currently
running in production. The
two-replica DeepSeek-V4-Flash stack
is the previous deployment, kept as a reviewed alternative and rollback
record.
Backend compatibility:
model-server-1andmodel-server-2are example Docker DNS names—replace them with your backends. The default router is not tied to vLLM: it forwards OpenAI-compatible APIs and health-checks each server withGET /v1/models. The opt-in/tokenize, KV-event, exact-routing, and snapshot research paths are currently designed for vLLM/DSpark.
score(replica) = min(prefix overlap, affinity cap) − α × live load
ramjet fingerprints only a bounded prefix, scores every healthy replica, and reserves load before forwarding. Warm state wins when it is valuable; idle capacity wins when reuse no longer pays for the queue. Score ties prefer the deeper raw overlap.
- OpenAI-compatible chat/completions, streaming, reasoning, and tool calls.
ok,degraded, andunhealthyreadiness atGET /health.- Optional SHA-pinned model/template compatibility admission for engines that expose the atomic identity contract, with fail-closed per-replica recovery; the node06 guide includes an opt-in, no-extra-hop vLLM middleware candidate.
- Optional DSpark reliability observation and sticky per-replica quarantine when active K5 acceptance collapses to zero across multiple complete metric windows; enforcement fsyncs an opaque EngineCore commitment and only a different compatibility-attested EngineCore can durably rearm it. A precommitted dirty marker keeps unresolved replicas fenced after an unclean LB exit or failed state mutation.
- Stable
ramjet_*Prometheus metrics on port9090. - Opaque
X-Ramjet-Upstreamroute correlation without leaking hosts. - Bounded memory, request sanitization, model metadata rewriting, and upstream cancellation when the client disappears.
Exact tokenization, fenced KV indexes, authenticated snapshot companions, exact-placement canaries, and session-affinity shadow telemetry remain opt-in research surfaces. The session path cannot change placement. These paths fail closed and are not dependencies of ordinary serving.
Naming: the project was renamed from ramjet to ramjet. Settings now use the
RJ_*prefix and responses carryX-Ramjet-*headers; the retiredMD_*prefix is refused at startup rather than silently ignored, so a stale overlay fails loudly instead of running a differently tuned proxy. Theramjet_*metric names are deliberately unchanged so existing Grafana history keeps resolving.
| Task | Start here |
|---|---|
| Deploy or roll back | Docker Compose operator guide |
| Configure the router | Environment reference |
| Serve a different model | Model profiles |
| Understand the design | Architecture and routing model |
| Inspect current work | Roadmap |
Codex-compatible repo skills are included for repeatable node operations:
$deploy-ramjet,
$optimize-ramjet-node,
$load-test-ramjet-node,
and
$troubleshoot-ramjet-node.
cargo fmt --check
cargo test --locked
cargo clippy --locked --all-targets --all-features -- -D warningsPrivacy-safe production-shape replay
For privacy-safe production-shape validation, bench/agent_trace.py accepts
only numeric/enumerated trace shapes and synthesizes all request content. A
bounded /tokenize preflight adjusts for the active chat-template overhead;
authoritative response usage still enforces the token-density gate. See the
sovereign trace replay contract.
See AGENTS.md for the GPU-free inner loop, full release gate, and node06 benchmark contract.
Everything we have written about serving on the Helix blog, grouped by topic and newest first. The per-model table above picks from the same posts.
Routing with ramjet
- Ramjet vs NVIDIA Dynamo: Which Router for Coding-Agent Traffic? (30 Sep)
- Serving Full GLM-5.3 to 48 Coding Agents From One 8×H200 Server (29 Sep)
- Self-Hosting Kev on an RTX PRO 6000 With Ramjet (22 Sep)
- One Qwen Model, Two Speeds: What Smart Routing Bought Us (28 Aug)
GLM-5.3-Flash
- GLM-5.3-Flash Ran Out of Cache Snapshots, Not Cache Tokens (26 Sep)
- Running GLM-5.3-Flash on 2, 4 or 8 RTX PRO 6000 GPUs (14 Sep)
- GLM-5.3-Flash on RTX PRO 6000, Part 1: Getting Day-Zero Serving to Work (27 Aug)
Qwen3.8
- Swift 1.5 Flash-Next Cut Qwen3.8's Thinking Tokens in Our Pilot (27 Sep)
- Qwen3.8-Flash-Next on Eight GPUs: What Actually Helped (27 Aug)
- A Better lm_head for Qwen3.8-27B: How We Tested and Shipped It (25 Aug)
- We Doubled Our Inference Throughput by Reading a Log Line (23 Aug)
- Chasing a 454 tok/s tweet: a day of tuning Qwen3.8-27B on the RTX PRO 6000 (22 Aug)
DeepSeek
- DeepSeek V4.1 Flash: Why Its Encoder, Engram and KV Cache Matter (10 Sep)
- SGLang vs DwarfStar vs vLLM+DSpark: Running DeepSeek 4 on the RTX Pro 6000 (14 Aug)
Hardware