The problem
A 35B mixture-of-experts model at Q4 is about 22 GB of weights, and the card in this machine holds 16. Buying more card is not the interesting answer for an operator whose inference has to sit next to the data. The interesting answer is that routed experts are touched sparsely: a token uses 8 of 256 per layer, so the working set of a decode step is a fraction of the model. Everything hinges on what happens when the expert you need is not in VRAM.
Results against llama.cpp
Same machine, same model file (Qwen3.6-35B-A3B UD-Q4_K_XL), measured 2026-09-29. llama.cpp is build acecd56 in its best configuration from a recorded sweep (--n-cpu-moe 16, 19 with its MTP draft, -t 6 -ub 2048 -fa on). Prompts are cold: a nonce at the start defeats any prompt cache. Decoding is greedy, one request at a time, medians of repeated runs.
| titan-engine, MTP=2 | titan-engine, no MTP | llama.cpp, MTP draft | llama.cpp, no MTP | |
|---|---|---|---|---|
| decode tok/s | 133.0 | 103.4 | 83.7 | 75.7 |
| time to first token (short prompt) | 98 ms | 103 ms | 172 ms | 152 ms |
| prompt tok/s, 4k | 2262 | 2314 | 1539 | 1771 |
| prompt tok/s, 13k | 2576 | 2667 | 1564 | 1711 |
| prompt tok/s, 28k | 2391 | 2384 | 1614 | - |
| q100 sanity check | 95 | 95 | 96 | - |
Decode is 1.59x llama.cpp with MTP on both sides and 1.37x without; prompt processing is 1.47x at 4k, 1.65x at 13k and 1.48x at 28k. A prompt that repeats an earlier prefix (an agent resending its system prompt and tools each turn) resumes from the hybrid prefix cache and starts generating in 0.4-0.9 s. Larger models are preliminary: Qwen3-Next-80B decodes at 41.5 tok/s against llama.cpp's 15.2, with experts streamed from NVMe.
What we built
- Tiered experts. A profiled share of every layer's experts stays in VRAM; the rest live in host memory. Missed experts run on the CPU in one pass per miss (gate/up, SwiGLU, down), and the GPU hands work over through a mapped-memory doorbell instead of blocking copies. That cut the GPU's idle share of a decode step from about 55% to 33%.
- Prefill expert streaming. Prompt chunks of 2048 tokens stream every non-resident expert to the GPU through a ring of staging buffers while earlier layers compute, and run them through a grouped MMQ kernel. A 13k prompt went from 336 to 2576 tok/s.
- Our own attention kernels. A tensor-core flash-prefill kernel with pipelined loads (64 TFLOP/s on this card) and a split-K flash-decoding kernel that reads each KV head once for all its query heads and all MTP verify rows. Decode after a 28k prompt went from 58 to 119 tok/s.
- Multi-token prediction, a hybrid prefix cache and CUDA graphs for the decode segments, with MTP output byte-identical to decoding without it.
- Every kernel is ours: written as Rust that emits PTX via NVIDIA's cuda-oxide, gated against llama.cpp's outputs, so the build needs no nvcc.
How we got there
A standing benchmark suite measures the matmul kernels against llama.cpp on identical device buffers, one MoE block's forward from a real profile, and end-to-end decode and prompt speed. Profiling first, not guessing, is what moved the numbers: the slow long prompts turned out to be a sum kernel launching a million tiny blocks and an attention path whose cost grew with every chunk, not the expert copies we suspected. The same suite killed ideas that looked good on paper: AVX-VNNI for the CPU dots (bit-identical, worth nothing end to end) and a learned expert-eviction policy (within 1.6% of simple LFU).