Skip to content
helixmlPublic

About

Local inference stack (DeepSeek-V4-flash, DeepSeek-V4.1-Flash:, qwen3.8-27b, GLM 5.3, qwen3.8-next, )

Resources

Stars

6 stars

Watchers

0 watching

Forks

Latest commit

 

History

696 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

ramjet

Warm intake. Balanced burn.

A compact Rust router that puts each OpenAI-compatible inference request on
the healthy GPU replica where it can do the least repeated work.

Rust 1.95 or newer

An incoming prompt is scored by ramjet and routed to the GPU replica with the best combination of reusable prefix and available capacity

ramjet sits between your clients and replicated model servers. It keeps conversations and shared system prompts near warm cache state, then lets live load override affinity before one replica becomes a hotspot. Clients keep the same OpenAI API; engines need no ramjet-specific integration.

What we wrote about each model

How ramjet and the engines behind it were tuned for each model, written up on the Helix blog. GLM-5.3 ran on an 8× H200 server; the others on our 8× RTX PRO 6000 server:

GLM-5.3 (8× H200) GLM-5.3-Flash Qwen3.8-Flash-Next Qwen3.8-27B DeepSeek V4
48 coding agents from one server (29 Sep) Part 1: getting day-zero serving to work (27 Aug) On eight GPUs: what actually helped (27 Aug) Chasing a 454 tok/s tweet (22 Aug) SGLang vs DwarfStar vs vLLM+DSpark (14 Aug)
Ramjet vs NVIDIA Dynamo (30 Sep) Running on 2, 4 or 8 GPUs (14 Sep) One model, two speeds: smart routing (28 Aug) Doubling throughput by reading a log line (23 Aug) V4.1 Flash: encoder, Engram and KV cache (10 Sep)
Ran out of cache snapshots, not cache tokens (26 Sep) Swift 1.5 cut thinking tokens (27 Sep) A better lm_head, tested and shipped (25 Aug)

Why it exists

Reuse more Queue less Fail cleanly
Bounded prefix fingerprints find the replica most likely to reuse prior work. Size-weighted reservations spread cold prefills and concurrent decodes. Active health probes, retryable failover, and immediate disconnect cancellation keep capacity honest.

The ordinary router is stateless, privacy-bounded, and deliberately useful without raw KV-cache events. Optional DSpark enforcement persists only opaque quarantine commitments so an LB restart cannot forget a bad EngineCore. The production path remains the proxy plus your existing OpenAI-compatible engines.

Built-in dashboard

System overview to see how your node is doing:

image

And specific serving tab:

image

You can also just plug it into prometheus, /metrics API is available.

Measured on real hardware

Each metric has its own column, so green always means better: throughput and cache hits up, latency down. Red marks a measured cost, grey a change within noise, and a dash a metric that run did not record.

Result Throughput Latency Cache hits
DeepSeek-V4-Flash  2× TP4 · 8× RTX PRO 6000
12 same-app sessions, load-blind hash router → ramjet 298 → 469 tok/s
▲ 57%
batch wall time
7.5 → 4.5 s
▼ 40%
no locality loss
– tie
2–3 apps × 4 sessions × 2–3 turns, hash router → ramjet — — 82.9% → 82.9%
– tie
Whole-box deterministic code, c24/max256 1,820–1,844 tok/s TTFT p50 948–960 ms
p95 1,270–1,319 ms
—
Qwen3.8-Flash-Next  2× TP4 · 8× RTX PRO 6000
Load-only → prefix routing, returning request beside a long one long request
100% → 93.6%
▼ 6.4%
TTFT p50 1,119 → 908 ms
▼ 19%
23.9% → 35.8%
▲ 12 pts
Prefix routing → phase-aware load release, same probe long request
100% → 99.1%
▼ 0.9%
TTFT 2,496 → 287 ms
▼ 88.5%
—
Phase-aware load release under c32 load 2,639.9 → 2,711.8 tok/s
▲ 2.7%
TTFT p95
▲ 7.3%
—
Direct vLLM → same engine through ramjet c1 −0.03%
c16 −0.24%
– ≈ 0
TTFT slightly lower
– ≈ 0
—
GLM-5.3-Flash  2× TP4 · 8× H200
Coding-agent swarm, prefix routing → marginal affinity basis 194–197 → 243–249
turns/min
▲ 23–28%
TTFT p90 5.3–5.5 →
3.2–3.3 s
▼ 38–42%
85.5–86.0% →
92.0–92.4%
▲ 6–7 pts
DeepSeek-V4.1-Flash  2× TP4 · 8× H200
Coding-agent swarm, 64 agents, default absolute → marginal affinity basis 222.7 → 263.8
turns/min
▲ 18%
turn e2e p90 36.6 → 28.3 s
▼ 23%
TTFT p90 ▲ 8%
88.5% → 91.6%
▲ 3 pts
Replica answers /health but cannot generate, RJ_UPSTREAM_RANK_PROBE on → all, 12 requests 6 → 12 of 12 served
▲ 100%
batch wall time
182 → 2 s
▼ 99%
—
GLM-5.3  DP8 attention · 8× H200
NVIDIA Dynamo 1.5.0 → ramjet, same engine, 32 and 48 agents 73.0 → 88.9
77.3 → 99.5
turns/min
▲ 22%, 29%
TTFT p50 2.06 → 1.53 s
2.92 → 2.88 s
▼ 26%, 1%
p90 ▲ 13%, 2.5%
85.0% → 92.4%
84.7% → 91.7%
▲ 7 pts

These are workload results, not theoretical peaks. Reproduce the DeepSeek rows from RESULTS.md; inspect every accepted and rejected experiment in EXPERIMENTS.md.

Models with a validated stack

All measured on node06 — 8× RTX PRO 6000 Blackwell; the engine topology is listed per row, and each model name links to its Compose stack. The full-box column reports the best qualified saturation point recorded for that stack, not a shared concurrency level.

Model Engine · measured shape Decode @ c1 Full-box peak
DeepSeek-V4-Flash
deepseek-v4-flash
sparse MoE
vLLM + DSpark
2× TP4 · c24/max256
245.1 tok/s 1,891.2 tok/s
Qwen3.8-27B FP8
qwen3.8-27b
dense
vLLM
2× TP4 · c256/max256 · MTP off
77 → 121 tok/s
▲ 57% with MTP
7,890.9 tok/s
Qwen3.8-27B NVFP4 + BF16 head
qwen3.8-27b
dense
SGLang + DFlash2
8× TP1 · 208 slots · bf16 SSM
153.3 tok/s
greedy median
▲ 7.5% vs Inferact
not yet requalified
Inferact target:
7,882.6 tok/s
Qwen3.8-Flash-Next FP8
qwen3.8-flash-next
sparse MoE
vLLM
2× TP4+EP · c64 · MTP3 on both
113 → 202 tok/s
▲ 79% with MTP3
3,340.5 tok/s
GLM-5.3-Flash W4A16, FP8 experts
glm-5.3-flash
sparse MoE
SGLang + EAGLE
TP2 · c4/max256
164.8 tok/s 388.2 tok/s
per 2-GPU replica;
box not yet saturated

No model — and neither Qwen3.8-27B stack — is simply better. Single-stream decode is what an interactive user feels; the full-box figure is a capacity landmark for a saturated agent fleet. These maxima come from separate model-specific workloads, so they are not a matched head-to-head benchmark. The vLLM row's saturation result has MTP off because speculation improves low-concurrency latency but wastes rejected drafts once the batch saturates the GPU. The SGLang row uses RadixArk's immutable BF16-lm_head checkpoint. Its matched one-engine canary measured 153.3 tok/s against 142.6 for the former Inferact target (+7.5%), with the same 7/8 objective answers and 20/25 deterministic agent-protocol cases. The smaller target exposes 26 running slots and 582,246 KV tokens per engine: 208 slots across the fleet. Full-box saturation has not yet been requalified on these weights; the former Inferact target reached 7,882.6 tok/s, within 0.1% of the vLLM reference. On that earlier SGLang stack, bf16 SSM state reduced c128 TTFT p95 from 3.99s to 0.221s. The same 3-app × 4-session × 2-turn locality run measured 87.3% cached prompt tokens, and 12 concurrent same-app requests spread across 7 of 8 engines at 714 tok/s. Its cost is cold long-context prefill: a 196K-token first turn pays ~57s of TTFT on one GPU, with prefix-cached follow-ups at 2–4s.

Qwen3.8-Flash-Next shows the same speculation trade-off: on 256-token outputs MTP3 adds 79% at c1 but only 7.5% at c32. The qualified pair therefore runs MTP3 on one engine and standard decoding on the other, and ramjet uses the requested output length to pick between them only once cache and load tie. Its full-box figure predates that split, with MTP3 on both engines. GLM-5.3-Flash runs on two GPUs per replica; its prefix cache is bounded by saved linear-attention states rather than KV tokens. Keeping two states per path instead of four, plus a 4 GB host tier, took a probe of 12 cyclic 20k-token sessions from 0/12 to 12/12 cached. The Kev stack adds a 0.8B decision model to the same server, sharing one Qwen GPU behind a second API profile: 77 ms p50 per short three-question request at c1 and 20.3 requests/s at c4, measured beside live traffic rather than saturated. Model profiles covers the sizing, sharding, and speculative-decoding trade-offs behind these numbers.

Full GLM-5.3 on one 8× H200 server

deploy/glm53_h200 runs the 753B FP8 checkpoint as one SGLang engine with eight data-parallel attention ranks, and ramjet lists each rank as its own upstream (RJ_UPSTREAM_DP_RANKS). On a simulated team of continuously working coding agents:

Result Throughput Latency Cache hits
16 agents: SGLang rank placement → ramjet per-rank routing 43.0 → 72.7
turns/min
▲ 69%
TTFT p50 2.6 → 0.76 s
p90 5.0 → 1.7 s
▼ 66–71%
65.6% → 92.8%
▲ 27 pts
64 agents on ramjet per-rank routing: adding a 32 GB host KV tier per rank 45.9 → 95.1
turns/min
▲ 107%
TTFT p90 67 → 46 s
▼ 31%
71.8% → 91.8%
▲ 20 pts
Capacity per server: 32–48 agents 98–109
turns/min
TTFT p50 1.3–2.8 s —

The blog post walks through each step, and Ramjet vs NVIDIA Dynamo compares the router against Dynamo 1.5.0's KV router on the same engine; the raw cells are in EXPERIMENTS.md (2026-09-29 and 2026-09-30).

Start in one minute

For existing engines, the upstream list is normally the only setting you need:

services:
  ramjet:
    image: ghcr.io/helixml/ramjet:v0.8.0@sha256:fe432bbca183d2a457a7713fb150ea5ee36aba7a13f92280ef3ec7195ec23673
    restart: unless-stopped
    ports:
      - "8000:8000" # OpenAI API + /health
      - "9090:9090" # Prometheus
    environment:
      RJ_UPSTREAM: http://model-server-1:8000,http://model-server-2:8000
      # RJ_UPSTREAM_TOKEN: ${MODEL_SERVER_API_KEY} # if required
docker compose up -d
curl --fail http://localhost:8000/health

The example pins a released image by immutable digest; see CHANGELOG.md for what each version contains. Safe defaults enable locality/load routing and keep tokenizer, raw KV-event, exact-placement, and snapshot paths off. See the complete configuration table, or start from the eight-replica Compose stack currently running in production. The two-replica DeepSeek-V4-Flash stack is the previous deployment, kept as a reviewed alternative and rollback record.

Backend compatibility: model-server-1 and model-server-2 are example Docker DNS names—replace them with your backends. The default router is not tied to vLLM: it forwards OpenAI-compatible APIs and health-checks each server with GET /v1/models. The opt-in /tokenize, KV-event, exact-routing, and snapshot research paths are currently designed for vLLM/DSpark.

The routing rule

score(replica) = min(prefix overlap, affinity cap) − α × live load

ramjet fingerprints only a bounded prefix, scores every healthy replica, and reserves load before forwarding. Warm state wins when it is valuable; idle capacity wins when reuse no longer pays for the queue. Score ties prefer the deeper raw overlap.

Production surface

  • OpenAI-compatible chat/completions, streaming, reasoning, and tool calls.
  • ok, degraded, and unhealthy readiness at GET /health.
  • Optional SHA-pinned model/template compatibility admission for engines that expose the atomic identity contract, with fail-closed per-replica recovery; the node06 guide includes an opt-in, no-extra-hop vLLM middleware candidate.
  • Optional DSpark reliability observation and sticky per-replica quarantine when active K5 acceptance collapses to zero across multiple complete metric windows; enforcement fsyncs an opaque EngineCore commitment and only a different compatibility-attested EngineCore can durably rearm it. A precommitted dirty marker keeps unresolved replicas fenced after an unclean LB exit or failed state mutation.
  • Stable ramjet_* Prometheus metrics on port 9090.
  • Opaque X-Ramjet-Upstream route correlation without leaking hosts.
  • Bounded memory, request sanitization, model metadata rewriting, and upstream cancellation when the client disappears.

Exact tokenization, fenced KV indexes, authenticated snapshot companions, exact-placement canaries, and session-affinity shadow telemetry remain opt-in research surfaces. The session path cannot change placement. These paths fail closed and are not dependencies of ordinary serving.

Naming: the project was renamed from ramjet to ramjet. Settings now use the RJ_* prefix and responses carry X-Ramjet-* headers; the retired MD_* prefix is refused at startup rather than silently ignored, so a stale overlay fails loudly instead of running a differently tuned proxy. The ramjet_* metric names are deliberately unchanged so existing Grafana history keeps resolving.

Operate it

Task Start here
Deploy or roll back Docker Compose operator guide
Configure the router Environment reference
Serve a different model Model profiles
Understand the design Architecture and routing model
Inspect current work Roadmap

Codex-compatible repo skills are included for repeatable node operations: $deploy-ramjet, $optimize-ramjet-node, $load-test-ramjet-node, and $troubleshoot-ramjet-node.

Develop

cargo fmt --check
cargo test --locked
cargo clippy --locked --all-targets --all-features -- -D warnings
Privacy-safe production-shape replay

For privacy-safe production-shape validation, bench/agent_trace.py accepts only numeric/enumerated trace shapes and synthesizes all request content. A bounded /tokenize preflight adjusts for the active chat-template overhead; authoritative response usage still enforces the token-density gate. See the sovereign trace replay contract.

See AGENTS.md for the GPU-free inner loop, full release gate, and node06 benchmark contract.

Resources

Everything we have written about serving on the Helix blog, grouped by topic and newest first. The per-model table above picks from the same posts.

Routing with ramjet

GLM-5.3-Flash

Qwen3.8

DeepSeek

Hardware

License

Apache-2.0.

About

Local inference stack (DeepSeek-V4-flash, DeepSeek-V4.1-Flash:, qwen3.8-27b, GLM 5.3, qwen3.8-next, )

Resources

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages