ST4TS

Measured on hardware we own. Click a column header to sort.
Every row is dated, stamped, and ships with its raw data. Words/s ≈ tok/s × 0.75 (stated conversion, English prose average — code and structured output tokenize differently, so treat it as prose-only). Corrections get published, not edited away.
The gate changed. An independent cross-check — a different model family, reading this page against its own raw data — now runs before rows land or change here. The first pass has run. It caught things; those are fixed in place, with the prior values still visible. The pass that clears the current page is still in flight, and this line gets updated when it lands — not before.

// the matrix — the under-50B class, three cards, decode tok/s

The under-50B class — models an ordinary card can hold. One model per row, same-family GGUF on all three hardware classes, pure-generation decode. “Built on” is the flag-accuracy column: a model’s country tells you who published it, not whose weights it started from. Holo is French and built on a Chinese base; Selene is British and built on an American one. “own” means the publisher pretrained it; “undisclosed” means we could not receipt a base model — a GGUF architecture tag names the code family that runs a model, not whose weights it started from, so we never publish an arch tag as lineage. Active-param counts written with a “~” are inferred from bandwidth arithmetic, not vendor-published (receipted in the CSV). ※ = the one model where the $1,299 card beat the workstation card — flagged for re-measurement, published as measured. ¶ = arch refusal on that box’s build, published rather than hidden: the model was attempted there and the engine refused to load its architecture (build named in the CSV). ⚑ = measured on a transient llama.cpp build newer than our pinned comparability build — several new architectures refuse to load on the pinned build at all, so the choice was a flagged number or no number (build hashes in the CSV; compare ⚑ cells to each other, not to unflagged ones). ⁂ = the same weights under two names: Ornith-1.0-35B (US) and Agents-A1-35B (CN) ship GGUF files 128 bytes apart with identical headers — 733 tensors, 256 experts — and bench within 3.4% everywhere, as identical files must. Two publishers, one model; we caught it because the numbers were implausibly identical. ◊ = ships a malformed chat template — benched with template parsing off; it cannot do tool-calling as shipped. ᶜ = community GGUF quant (quantizer in the CSV) where the vendor publishes none — Muse-Glimmer-30B is the sharpest case: no first-party GGUF exists for it at all, so its rows run a community ~4.5 bpw k-quant rather than the field-standard Q4_K_M, which makes its speeds approximate against the Q4_K_M ladder, not like-for-like. Four further models — an Indian 12B, a Korean 32B, and two US models — were benched but are held off this page pending license review: their licenses carry non-commercial, non-compete or read-before-publish clauses we read before publishing, not after. Bare dashes fill in public as test windows reach them; a dash carrying ¶ was attempted and refused. The 70B+ class gets its own table below.
Why these three: what’s feasible at home. One is the speed ceiling, one is a quiet whole-box with a huge memory pool, one is an ordinary $1,299 card. Read a row across and the trade-off is the answer.
ModelOriginParamsActiveTypeRTX PRO 6000 96GBStrix Halo 128GBR9700 32GBBuilt onQuant
LFM2.5-1.2B-ThinkingUS1.2B1.2Bdense979.8⚑455.7ownQ4_K_M
LFM2.5-8B-A1BUS8.5B1BMoE637.3⚑162.4303.7ownQ4_K_M
LFM2.5-2.6BUS2.7B2.7Bdense482.3⚑111.7231.0ownQ4_K_M
LFM2-24B-A2BUS24B2BMoE380.5119.2220.7ownQ4_K_M
Sarvam 30BIN30B~3BMoE364.683.7198.4ownQ4_K_M
NVIDIA Nemotron-3.5-Lightning-30B-A3BUS32.9B3BMoE hybrid (mamba)298.4⚑59.2⚑ownQ4_K_Mᶜ
Nemotron-3-Nano-30B-A3BUS31.6B3.5BMoE hybrid (mamba)286.465.9137.4ownQ4_K_M
Agents-A1-4BCN4.2B4.2Bdense285.1⚑131.0undisclosedQ4_K_M
Qwen3-Coder-30B-A3BCN30.5B3.3BMoE273.391.6182.0ownQ4_K_M
Cohere North Mini Code 1.0CA30B3BMoE270.6⚑134.0undisclosedUD-Q4_K_M
Nemotron-Cascade-2-30B-A3BUS30B3BMoE hybrid (mamba)268.562.3131.9ownQ4_K_M
Agents-A1-35B-A3B ⁂CN34.7B3BMoE258.6⚑72.4Qwen3.5-35B (CN)Q4_K_M
KAT-Coder-V2.5-DevCN34.7B3BMoE253.8⚑undisclosedQ4_K_Mᶜ
Ornith-1.0-35B ⁂US34.7B3BMoE250.2⚑71.3Qwen3.5-35B (CN)Q4_K_M
Kimi-Linear-48B-A3BCN48B3BMoE linear-attn224.170.4131.6ownQ4_K_M
H Company Holo 3.1 35B-A3BFR35B3BMoE221.372.9127.3Qwen3.6-35B (CN)Q4_K_M
Nanbeige-4.2-3BCN4.2B4.2Bdense213.0⚑108.5⚑ownQ4_K_Mᶜ
Qwen-AgentWorld-35B-A3BCN35B3BMoE206.461.2100.2ownUD-Q4_K_M / Q4_K_M (R9700, 2026-07)
Qwen3.6-35B-A3BCN35B3BMoE hybrid (SSM)204.059.9107.0ownUD-Q4_K_M
GLM-4.7-FlashCNMoE196.470.2122.0ownQ4_K_M
Ornith-1.0-9BUS9B9Bdense193.4⚑86.2Qwen3.5-9B (CN)Q4_K_M
Gemma-4-26B-A4B-itUS26B4BMoE190.052.5100.7ownUD-Q4_K_M
Mamba-Codestral-7BFR7B7Bpure mamba184.237.3⚑85.5ownQ4_0
Atla Selene-1-Mini 8BUK8B8Bdense162.327.069.9Llama-3.1-8B (US)Q8_0
DeepSeek-Coder-V2-LiteCN16B2.4BMoE156.4107.8206.4※ownQ4_K_M
Phi-4 14BUS14B14Bdense133.223.262.0ownQ4_K_M
StarCoder2-15B-InstructUS15B15Bdense115.620.656.1ownQ4_K_M
IBM Granite 4.1 30BUS30B30Bdense69.111.631.2ownQ4_K_M
Muse-Glimmer-30BUS30B30Bdense70.2⚑12.5⚑32.7⚑ownkquant ~4.5bpwᶜ
Gemma 4 31B ITUS31B31Bdense58.910.427.3ownQ4_K_M
Seed-OSS 36BCN36B36Bdense57.410.126.8ownQ4_K_M
Falcon-H1 34BAE34B34Bhybrid (mamba)51.89.924.5ownQ4_K_M
Nemotron-Super-49B v1.5US49B49Bdense40.5⚑Llama-3.3-70B (US)Q4_K_M
Nemotron-Super-49B v1US49B49Bdense41.0⚑Llama-3.3-70B (US)Q4_K_M
gpt-oss-20bUS21B3.6BMoE168.4ownMXFP4
Nemotron-Nano-9B-v2US9B9Bhybrid (mamba)73.0ownQ4_K_Mᶜ
IBM Granite-4.0-H-SmallUS32B9BMoE hybrid (mamba)72.8ownQ4_K_M
Apriel-1.5-15B-ThinkerUS15B15Bdense62.3◊ownQ4_K_Mᶜ
Phi-4-ReasoningUS14B14Bdense61.3ownQ4_K_Mᶜ
Reka-Flash-3.1US21B21Bdense41.7ownQ4_K_Mᶜ
Magistral-SmallFR24B24Bdense40.6ownQ4_K_M
Devstral-Small-2507FR24B24Bdense40.6ownQ4_K_M
Mistral-Small-3.2-24BFR24B24Bdense39.5ownQ4_K_M
Qwen-SEA-LION-v4.5-27BSG27B27Bdense31.1Qwen (CN)Q4_K_M
Qwen3-32BCN32.8B32.8Bdense29.0ownQ4_K_M
Olmo-3-32B-ThinkUS32B32Bdense28.8ownQ4_K_Mᶜ
2026-08-12: a 5 GB file took the top of the mini-PC column. LFM2.5-8B-A1B — 8.5B parameters, ~1B active — posts 637⚑ / 162 / 304 tok/s across all three cards, and it takes the Strix Halo column outright at 162.4, a 36% jump in that box’s sustained record (the previous holder, LFM2-24B-A2B, sat at 119.2). It does not sweep the matrix — its own 1.2B sibling leads the other two columns, at 979.8⚑ on the workstation card and 455.7 on the $1,299 card, with no Strix Halo leg yet. At the other end, the same night’s head-to-head: Nemotron-Super-49B v1 vs v1.5 is a dead heat (41.0 vs 40.5, identical VRAM to the MiB) — the .5 is a quality release, and a throughput bench cannot see quality. Same lesson from the qwen35moe cohort: Ornith and Agents-A1 (⁂ — one file under two names) plus KAT-Coder, a third publisher on the same architecture, all land within 3.4% of each other on the card where all three ran — choosing between them is a license and quality question, not a speed question. Same-family GGUF quants only, so the cards are the variable — vLLM/NVFP4 configs live in the detail table below. The Granite row: one dense 30B, three cards, 69 / 12 / 31 tok/s. At 100B+ the pattern holds harder: active-param MoEs span 84–226 tok/s on the Blackwell card — 5.1B-active gpt-oss-120b at the ceiling, 17B-active Llama-4-Scout at the floor, the 105–122B middle clustered at 117–139 — where 111–123B dense runs 17–20, and dense-72B collapses to ~4 on unified memory. The gpt-oss row is now the identical file on both machines: the workstation card is 4.2× the throughput, at roughly 3.4× the power on the closest measured comparison (a 119B MoE at 225 W on that card against the mini-PC’s measured 67 W) — not the 6× we printed. The mini-PC’s sibling quant samples at 0.83 tok/s/W (within 0.7% of this file, so we cite it as the nearest measured point rather than as this file’s own number); the workstation card’s leg of this pair was never power-sampled, so we withdraw the “per watt the mini-PC ties or wins” claim rather than replace it with another unmeasured one. Wave-3 pattern: Strix Halo’s 30B class is a three-tier ladder — plain-attention A3B MoE at 61–92 tok/s, hybrids around 60–73, dense pinned at the ~10 tok/s bandwidth wall. And the two 80B “twins” (Coder-Next vs Next-thinking, same architecture, different finetune) benched identical to the decimal on Blackwell — architecture sets the row, not the finetune.

// the big memory pools — 70B+ where 96 GB addressable meets 128 GB unified

The 70–123B class needs more memory than an ordinary card carries, so this is a memory-pool fight: 96 GB of addressable VRAM (the workstation card) vs 128 GB of unified memory (the mini-PC, where CPU and GPU share one pool and a model can spill past what a discrete card could hold). Same GGUF discipline, same decode measure. Params/Active tell the story before the numbers do — active-param MoEs run 5–10× the dense speed at every size. And one thing to keep in mind reading the 30B matrix above against this table: nobody buys a 96 GB card to run one 30B. That card's real job is either this class — or several 30Bs at once: one recent bakeoff held five separate models on it at once, each on its own port. Co-residency economics is a queued comparison of its own.
ModelOriginParamsActiveTypeRTX PRO 6000 96GBStrix Halo 128GBQuant
gpt-oss-120bUS117B5.1BMoE (iSWA)225.954.2MXFP4 (same file)
Qwen3-Coder-Next (0129)CN80B3BMoE hybrid (linear)174.459.0Q4_K_M
Qwen3-Next-80B-A3B (thinking)CN80B3BMoE hybrid (linear)174.356.9Q4_K_M
Mistral Small 4 119B (2603)FR119BMoE138.926.0Q4-class
Hunyuan-A13B-80BCN80B13BMoE138.326.7Q4-class
Sarvam 105BIN105BMoE131.7Q4-class
Qwen3.5-122B-A10BCN122B10BMoE117.8Q4-class
Llama-4-ScoutUS109B17BMoE84.0Q4-class
KAT-Dev-72BCN72B72Bdense27.94.4Q4-class
GLM-4.5-AirCNMoE23.2Q4-class
Cohere Command A 111BCA111B111Bdense19.8Q4-class
Devstral 2 123BFR123B123Bdense17.8Q4-class
Mistral Large 2407FR123B123Bdense17.3Q4-class
Llama-3.3-70BUS70B70Bdense5.1Q4-class
Hermes 70BUS70B70Bdense5.1Q4-class
Athene-V2-72BUS72B72Bdense4.4Q4-class
Cohere Command A Reasoning 111BCA111B111Bdense (iSWA)3.3Q4-class
Dashes are queued legs, not refusals — every model here loads on at least one box. Multi-model co-residency (several engines sharing one card, the as-deployed reality) is a queued comparison of its own — the Blackwell “short” rows in the detail table were measured exactly that way.

// the spill test — when the model doesn’t fit the card

These models are 44–59 GiB; the R9700 has 32 GB. llama.cpp does not refuse — it loads anyway, spilling the overflow into host RAM, and serves degraded. Nobody publishes this table. It answers a different question than every other number on this page — not “what does this card do?” but “what happens when you exceed it?” — which is why these rows live here and never beside protocol rows.
ModelTypeWeights on diskSpilled to host RAMdecode tok/s
Qwen3-Coder-Next 80BMoE A3B45.09 GiB14.21 GiB21.61
Qwen3-Next-80B-A3BMoE A3B45.09 GiB14.21 GiB21.61
gpt-oss-120b (MXFP4)MoE A5B59.03 GiB28.09 GiB15.17
Hunyuan-A13B-80BMoE A13B45.43 GiB18.08 GiB9.41
KAT-Dev-72Bdense44.16 GiB22.24 GiB1.75
Method, disclosed: these are load-and-probe figures — a single 128-token generation, no warm-ups, no sustained run, no power sampling, no prefill sweep — a different protocol from every other row on this page, which is why they get their own table instead of a column. VRAM pinned at 99.4–99.8% of the card in every case. The lesson stands: “it loaded” is the trap — a card that technically runs a model at walking pace is not a card that runs it. The MoE pattern is real, though: active-3B experts spill 3× more gracefully (21.6 tok/s) than a 72B dense crawl (1.75).

// performance per watt — the electric-bill column

Decode tok/s divided by GPU-reported power during sustained generation. This is where the $8K workstation card and the quiet mini-PC stop being rivals: on efficiency they’re peers for MoE models — and dense models are power hogs on both.
ModelRTX PRO 6000: tok/s @ Wper wattStrix Halo: tok/s @ Wper wattR9700: tok/s @ Wper watt
LFM2.5-8B-A1B637.3 @ 292⚑2.18162.4 @ 792.05303.7 @ 2541.20
LFM2.5-1.2B-Thinking (thin power sample)979.8 @ 216⚑4.55⚠455.7 @ 2481.84⚠
LFM2.5-2.6B482.3 @ 302⚑1.60⚠111.7 @ 861.31231.0 @ 2950.78
LFM2-24B-A2B378.2 @ 2401.57119.2 @ 771.54220.7 @ 2560.86
Cohere North Mini Code 1.0270.6 @ 268⚑1.01134.0 @ —
GLM-4.7-Flash196.4 @ 2140.9268.6 @ 820.84122.0 @ 2670.46
DeepSeek-Coder-V2-Lite113.3 @ 1280.89104.0 @ 801.30206.4 @ 2630.78
Gemma-4-26B-A4B-it190.0 @ 2240.8550.6 @ 780.65100.7 @ 2630.38
Mamba-Codestral-7B184.2 @ 2750.6785.5 @ 2900.29
Atla Selene-1-Mini 8B162.3 @ 300 cap0.5426.5 @ 860.3169.9 @ 3000.23
Phi-4 14B131.6 @ 300 cap0.4423.2 @ 900.2662.0 @ 3000.21
StarCoder2-15B115.6 @ 300 cap0.3920.6 @ 900.2356.1 @ 3000.19
Sarvam 30B364.6 @ 2831.2983.7 @ ~900.93
Qwen3-Coder-30B-A3B273.3 @ 2301.1991.6 @ ~851.08
Nemotron-3-Nano-30B-A3B286.4 @ 2561.1265.9 @ ~890.74
H Company Holo 3.1 35B-A3B221.3 @ 2161.0372.9 @ ~790.93
Kimi-Linear-48B-A3B224.1 @ 2330.96131.6 @ 2170.61
Qwen-AgentWorld-35B-A3B206.4 @ 2210.9361.2 @ ~780.78
Qwen3-Next-80B-A3B174.3 @ 2100.83
Qwen3-Coder-Next (0129)174.4 @ 2140.82
Gemma 4 31B IT (dense)58.9 @ 300 cap0.2010.4 @ ~1010.10
Seed-OSS 36B (dense)57.4 @ 300 cap0.1910.1 @ ~1030.10
Falcon-H1 34B (hybrid)51.8 @ 300 cap0.179.9 @ ~980.10
Qwen3.6-35B-A3B107.0 @ —
Nemotron-Cascade-2-30B-A3B131.9 @ —
IBM Granite 4.1 30B (dense)69.1 @ 300 cap0.2331.2 @ —
NVIDIA Nemotron-3.5-Lightning-30B-A3B298.4 @ 277⚑1.0859.2 @ 89⚑0.66
Agents-A1-35B-A3B ⁂258.6 @ 254⚑1.0272.4 @ 800.91
Ornith-1.0-35B ⁂250.2 @ 246⚑1.0271.3 @ 830.86
KAT-Coder-V2.5-Dev253.8 @ 258⚑0.98
Agents-A1-4B285.1 @ 301⚑0.95131.0 @ 2810.47
Ornith-1.0-9B193.4 @ 300 cap⚑0.6486.2 @ 3000.29
gpt-oss-20b (MXFP4)168.4 @ 3000.56
Muse-Glimmer-30B (dense)70.2 @ 300 cap⚑0.2312.5 @ 86⚑0.1532.7 @ 300⚑0.11
Nemotron-Super-49B v1.5 (dense)40.5 @ 300 cap⚑0.135
Method, honestly: Blackwell watts = nvidia-smi median board power during a 512-token sustained gen; Strix Halo watts = sysfs/hwmon GPU range midpoint during a 384-token gen — GPU-reported, not wall power, so cross-box comparison is directional, not billing-grade. Basis, stated: the tok/s here is the 512-token sustained rate where one was sampled and the 128-token bench rate otherwise, because power is only meaningful over the window it was sampled on — so cells here can sit a few percent under the matrix above. Rows whose wattage was never sampled now read “@ —” with no per-watt figure, rather than borrowing a card-class assumption. From the 30B class up, every dense model slams the Blackwell 300 W cap for about a fifth of the MoE speed; small dense models are a real exception — LFM2.5-1.2B-Thinking runs 979.8 tok/s at 216 W, Mamba-Codestral-7B 184.2 at 275 W. In English: if you care about the power bill, pick MoE — and on the mini-PC the draw is nearly flat (77–103 W whatever you load), so there the only thing that moves efficiency is speed.
2026-08-12: the crown moved on the mini-PC and the floor dropped. LFM2.5-8B-A1B posts 2.18⚠ / 2.05 / 1.20 tok/s/W from a file one-third the size of the model it displaced. The Strix Halo figure is the receipted record — 2.05 on 79 W, 33% clear of the previous holder (LFM2-24B at 1.54) — and on the R9700 it beats that model 1.20 to 0.86. On the workstation card it is the best settled row, but its own 1.2B sibling reads higher on both the Blackwell (4.55⚠) and R9700 (1.84⚠) legs, and its Blackwell leg is a ⚑ transient build measured against LFM2-24B’s pinned one — so no all-cards crown is claimed here until the longer power window lands. The floor for this table is the 49B dense class: Nemotron-Super v1.5 pins the 300 W cap for 40.5 tok/s = 0.135 tok/s/W (v1 reads 41.0 = 0.137), 16× worse than the champion on the same card — and the 70B+ table below goes lower still. The MoE-vs-dense power split is a strong tendency, not a law: at 30B and up on the Blackwell card every cap-pinned row is dense while the A3B MoEs cruise 210–292 W, but small dense models stay well under the cap, and on the R9700 sparse models pin it too (gpt-oss-20b 300 W, Granite-4.0-H-Small 299 W). ⚠ = the power median rests on 3–8 samples (4.55 = 3, 2.18 = 5, 1.60 = 6) — indicative, not settled; rows carrying it are not counted as records on this page until a 2,048-token window re-measures them.
70B+ class on the workstation cardtok/s @ Wper watt
Mistral Small 4 119B (MoE)138.9 @ 2250.62
Hunyuan-A13B-80B (MoE)138.3 @ 3000.46
Llama-4-Scout (MoE A17B)84.0 @ 3050.28
Corrected 2026-08-12: this table used to carry ten rows. Seven of them — gpt-oss-120b, Sarvam 105B, Qwen3.5-122B-A10B, KAT-Dev-72B, Cohere Command A 111B, Devstral 2 123B and Mistral Large 2407 — had no power sample behind them; their wattage was assumed from the card’s 300 W cap, not measured. Our own audit caught it and removed them; the three rows left are the ones whose watts are in the receipts. Their throughput is unaffected and still published in the tables above. What the three surviving rows show: even inside one architecture family and one card, per-watt spans 0.62 down to 0.28 — and the top row earns it partly by drawing 225 W where the others pin the 300 W cap, so read the watts column, not just the ratio. The dense-vs-MoE throughput gap is real but lives in the tables above; the dense rows are not in this one, because their power was never sampled. Strix Halo watts for this class: one measured point so far — the mini-PC’s sibling quant of the gpt-oss pair sampled at 67 W and 0.83 tok/s/W. That is a measurement of the sibling file, not of the MXFP4 file in the pair — the two read within 0.7% of each other on that box, which is why we cite it, and it is still an inference across quants rather than a direct sample. It puts the mini-PC near one-third the workstation card’s draw on the nearest comparable row, not the one-sixth we printed. Multi-model co-residency economics: queued.

// the bench — hardware classes

Specs from vendor datasheets/whitepapers (NVIDIA RTX Blackwell PRO architecture whitepaper Table 4; AMD product pages), fetched and verified 2026-08-11. Plural nodes per class; per-row stamps carry the exact box config.

// the detail — every engine config, one row each [ + 128 rows — click to open ][ − collapse ]

128 rows · one row per model × hardware × quant · 3 hardware classes · origins US / CN / EU / FR / UK / CA / IN / AE / SG. One row per model × hardware × quant — a model appears once. Sort any column, filter by class. A “—” means not yet measured — never a guess. Depth curves and config experiments live in their own sections below.
ModelDecode liveDecode benchPrefill t/s~words/sOriginArchQuantEngineHardware KV/tok KiBDate
poolside Laguna S 2.1 + draft119.4294.7~90USMoE 118B-A8.5BNVFP4vLLMRTX PRO 6000 96GB38.22026-08
Mistral Small 4 119B (2603)137.0~103FRMoE 119BNVFP4vLLMRTX PRO 6000 96GB2026-07
Qwen3.6-35B-A3B 238.4~179CNMoE 35B-A3BNVFP4 (NVIDIA)vLLMRTX PRO 6000 96GB10.352026-08-11
Cohere North Mini Code 1.0195.9194.5≥15,200~147CAMoE 30B-A3BNVFP4ᶜvLLMRTX PRO 6000 96GB16.682026-08-11
GLM-4.7-Flash177.211,535~133CNMoENVFP4ᶜvLLMRTX PRO 6000 96GB2026-08-04
Qwen3.5-122B-A10B 96.2~72CNMoE 122B-A10BNVFP4vLLMRTX PRO 6000 96GB2026-07
gpt-oss-120b225.9~169USMoE 117B-A5BGGUF MXFP4llama.cppRTX PRO 6000 96GB36.02026-08-11
Nemotron-Cascade-2-30B-A3B268.5~201USMoE 30B-A3BGGUF Q4_K_Mllama.cppRTX PRO 6000 96GB2026-08-11
Qwen3.6-35B-A3B204.06,364~153CNMoE 35B-A3BGGUF UD-Q4_K_Mllama.cppRTX PRO 6000 96GB2026-08-11
IBM Granite 4.1 30B73.069.12,956~55USdense 30BNVFP4ᶜ / GGUFvLLM / llama.cppRTX PRO 6000 96GB1282026-08-11
gpt-oss-120b51.854.6‖508‖~39USMoE 117B-A5BGGUF UD-Q4_K_XLllama.cppStrix Halo 128GB36.02026-08-11
gpt-oss-120b (same-quant pair)54.2~41USMoE 117B-A5BGGUF MXFP4llama.cppStrix Halo 128GB2026-08-11
Qwen3-Coder-Next (0129)55.559.0715~42CNMoE 80B-A3BGGUF Q4_K_Mllama.cppStrix Halo 128GB2026-08-11
Nemotron-Cascade-2-30B-A3B53.962.31,026~40USMoE 30B-A3BGGUF Q4_K_Mllama.cppStrix Halo 128GB2026-08-11
Qwen3.6-35B-A3B59.9~45CNMoE 35B-A3BGGUF UD-Q4_K_Mllama.cppStrix Halo 128GB2026-08-11
Kimi-Linear-48B-A3B-Instruct64.470.4683~53CNMoE 48B-A3B (linear attn)GGUF Q4_K_Mllama.cppStrix Halo 128GB7.92026-08-11
SmolLM3-3B91.2~68EUdense 3BGGUF Q4_K_Mllama.cppStrix Halo 128GB2026-08-11
Phi-4-mini71.6~54USdense 3.8BGGUF Q4llama.cppStrix Halo 128GB2026-08-11
Flow-Judge v0.1 (3.8B)47.750.2‖1,502~36EUdense 3.8BGGUF Q8_0llama.cppStrix Halo 128GB2026-08-11
Qwen3-8B40.6~30CNdense 8BGGUF Q4_K_Mllama.cppStrix Halo 128GB2026-08-11
Atla Selene 1 Mini (8B)24.927.0871~19UKdense 8BGGUF Q8_0llama.cppStrix Halo 128GB2026-08-11
IBM Granite 4.1 30B11.6~9USdense 30BGGUF Q4_K_Mllama.cppStrix Halo 128GB2026-08-11
Qwen3.6-35B-A3B107.02,502~80CNMoE 35B-A3BGGUF UD-Q4_K_Mllama.cppR9700 32GB2026-08-11
Nemotron-Cascade-2-30B-A3B131.92,825~99USMoE 30B-A3BGGUF Q4_K_Mllama.cppR9700 32GB2026-08-11
LFM2-24B-A2B380.59,990~285USMoE 24B-A2BGGUF Q4_K_Mllama.cppRTX PRO 6000 96GB2026-08-11
Cohere North Mini Code 1.0270.6⚑7,759~203CAMoE 30B-A3BGGUF UD-Q4_K_Mllama.cppRTX PRO 6000 96GB2026-08-11
GLM-4.7-Flash196.45,972~147CNMoEGGUF Q4_K_Mllama.cppRTX PRO 6000 96GB2026-08-11
Gemma-4-26B-A4B-it190.01,330~143USMoE 26B-A4BGGUF UD-Q4_K_Mllama.cppRTX PRO 6000 96GB2026-08-11
Mamba-Codestral-7B184.25,772~138FRpure mamba 7BGGUF Q4_0llama.cppRTX PRO 6000 96GB2026-08-11
Atla Selene-1-Mini 8B162.311,996~122UKdense 8BGGUF Q8_0llama.cppRTX PRO 6000 96GB2026-08-11
DeepSeek-Coder-V2-Lite156.4~117CNMoE 16B-A2.4BGGUF Q4_K_Mllama.cppRTX PRO 6000 96GB2026-08-11
Phi-4 14B133.2~100USdense 14BGGUF Q4_K_Mllama.cppRTX PRO 6000 96GB2026-08-11
StarCoder2-15B-Instruct115.65,088~87USdense 15BGGUF Q4_K_Mllama.cppRTX PRO 6000 96GB2026-08-11
LFM2-24B-A2B220.74,015~166USMoE 24B-A2BGGUF Q4_K_Mllama.cppR9700 32GB2026-08-11
DeepSeek-Coder-V2-Lite206.43,837~155CNMoE 16B-A2.4BGGUF Q4_K_Mllama.cppR9700 32GB2026-08-11
GLM-4.7-Flash122.01,493~92CNMoEGGUF Q4_K_Mllama.cppR9700 32GB2026-08-11
Gemma-4-26B-A4B-it100.72,670~76USMoE 26B-A4BGGUF UD-Q4_K_Mllama.cppR9700 32GB2026-08-11
Mamba-Codestral-7B85.52,339~64FRpure mamba 7BGGUF Q4_0llama.cppR9700 32GB2026-08-11
Atla Selene-1-Mini 8B69.92,804~52UKdense 8BGGUF Q8_0llama.cppR9700 32GB2026-08-11
Phi-4 14B62.01,478~47USdense 14BGGUF Q4_K_Mllama.cppR9700 32GB2026-08-11
StarCoder2-15B-Instruct56.11,108~42USdense 15BGGUF Q4_K_Mllama.cppR9700 32GB2026-08-11
Kimi-Linear-48B-A3B131.62,039~99CNMoE 48B-A3B (linear attn)GGUF Q4_K_Mllama.cppR9700 32GB7.92026-08-11
Cohere North Mini Code 1.0134.02,362~101CAMoE 30B-A3BGGUF UD-Q4_K_Mllama.cppR9700 32GB2026-08-11
Mistral Small 4 119B (2603)138.9~295 (anomaly, re-measure queued)~104FRMoE 119BGGUF Q4-classllama.cppRTX PRO 6000 96GB2026-08-11
Hunyuan-A13B-80B138.33,229~104CNMoE 80B-A13BGGUF Q4-classllama.cppRTX PRO 6000 96GB2026-08-11
Sarvam 105B131.72,931~99INMoE 105BGGUF Q4-classllama.cppRTX PRO 6000 96GB2026-08-11
Qwen3.5-122B-A10B117.83,106~88CNMoE 122B-A10BGGUF Q4-classllama.cppRTX PRO 6000 96GB2026-08-11
Llama-4-Scout84.02,347~63USMoE A17BGGUF Q4-classllama.cppRTX PRO 6000 96GB2026-08-11
KAT-Dev-72B27.91,252~21CNdense 72BGGUF Q4-classllama.cppRTX PRO 6000 96GB2026-08-11
Cohere Command A 111B19.8778~15CAdense 111BGGUF Q4-classllama.cppRTX PRO 6000 96GB2026-08-11
Devstral 2 123B17.8702~13FRdense 123BGGUF Q4-classllama.cppRTX PRO 6000 96GB2026-08-11
Mistral Large 240717.3697~13FRdense 123BGGUF Q4-classllama.cppRTX PRO 6000 96GB2026-08-11
Hunyuan-A13B-80B26.7200~20CNMoE 80B-A13BGGUF Q4-classllama.cppStrix Halo 128GB2026-08-11
KAT-Dev-72B4.4~3CNdense 72BGGUF Q4-classllama.cppStrix Halo 128GB2026-08-11
Athene-V2-72B4.4~3USdense 72BGGUF Q4-classllama.cppStrix Halo 128GB2026-08-11
Cohere Command A Reasoning 111B3.3~2CAdense 111BGGUF Q4-classllama.cppStrix Halo 128GB2026-08-11
Qwen3-Next-80B-A3B56.9~43CNMoE 80B-A3BGGUF Q4-classllama.cppStrix Halo 128GB2026-08-11
Mistral Small 4 119B (2603)26.0~20FRMoE 119B (MLA)GGUF Q4-classllama.cppStrix Halo 128GB2026-08-11
GLM-4.5-Air23.2~17CNMoEGGUF Q4-classllama.cppStrix Halo 128GB2026-08-11
Llama-3.3-70B5.1~4USdense 70BGGUF Q4-classllama.cppStrix Halo 128GB2026-08-11
Hermes 70B5.1~4USdense 70BGGUF Q4-classllama.cppStrix Halo 128GB2026-08-11
Sarvam 30B198.43,613~149INMoE 32BGGUF Q4_K_Mllama-benchR9700 32GB2026-08-11
Qwen3-Coder-30B-A3B182.03,035~137CNMoE 30B-A3BGGUF Q4_K_Mllama-benchR9700 32GB2026-08-11
Nemotron-3-Nano-30B-A3B137.42,677~103USMoE 31B-A3.5BGGUF Q4_K_Mllama-benchR9700 32GB2026-08-11
H Company Holo 3.1 35B-A3B127.32,681~95FRMoE 35B-A3BGGUF Q4_K_Mllama-benchR9700 32GB2026-08-11
Qwen-AgentWorld-35B-A3B100.2415~75CNMoE 35B-A3BGGUF Q4_K_Mllama.cppR9700 32GB2026-07
IBM Granite 4.1 30B31.2848~23USdense 30BGGUF Q4_K_Mllama-benchR9700 32GB2026-08-11
Gemma 4 31B IT25.527.3756~19USdense 31BGGUF Q4_K_Mllama.cppR9700 32GB2026-08-11
Seed-OSS 36B26.8707~20CNdense 36BGGUF Q4_K_Mllama-benchR9700 32GB2026-08-11
Falcon-H1 34B24.5725~18AEhybrid 34BGGUF Q4_K_Mllama-benchR9700 32GB2026-08-11
Sarvam 30B364.611,021~273INMoE 30B (A3B-class)GGUF Q4_K_Mllama.cppRTX PRO 6000 96GB2026-08-11
Nemotron-3-Nano-30B-A3B286.48,754~215USMoE 31B-A3.5BGGUF Q4_K_Mllama.cppRTX PRO 6000 96GB2026-08-11
Qwen3-Coder-30B-A3B273.37,194~205CNMoE 30B-A3BGGUF Q4_K_Mllama.cppRTX PRO 6000 96GB2026-08-11
Kimi-Linear-48B-A3B224.16,874~168CNMoE 48B-A3B (linear attn)GGUF Q4_K_Mllama.cppRTX PRO 6000 96GB7.92026-08-11
H Company Holo 3.1 35B-A3B221.36,917~166FRMoE 35B-A3BGGUF Q4_K_Mllama.cppRTX PRO 6000 96GB2026-08-11
Qwen-AgentWorld-35B-A3B206.47,543~155CNMoE 35B-A3BGGUF UD-Q4_K_Mllama.cppRTX PRO 6000 96GB2026-08-11
Qwen3-Coder-Next (0129)174.44,258~131CNMoE 80B-A3BGGUF Q4_K_Mllama.cppRTX PRO 6000 96GB24.02026-08-11
Qwen3-Next-80B-A3B (thinking)174.34,430~131CNMoE 80B-A3BGGUF Q4_K_Mllama.cppRTX PRO 6000 96GB2026-08-11
Gemma 4 31B IT58.9927~44USdense 31BGGUF Q4_K_Mllama.cppRTX PRO 6000 96GB42.52026-08-11
Seed-OSS 36B57.42,296~43CNdense 36BGGUF Q4_K_Mllama.cppRTX PRO 6000 96GB2026-08-11
Falcon-H1 34B51.82,190~39AEhybrid 34BGGUF Q4_K_Mllama.cppRTX PRO 6000 96GB2026-08-11
Qwen3-Coder-30B-A3B91.6993~69CNMoE 30B-A3BGGUF Q4_K_Mllama.cppStrix Halo 128GB2026-08-11
Sarvam 30B83.71,121~63INMoE 30B (A3B-class)GGUF Q4_K_Mllama.cppStrix Halo 128GB2026-08-11
H Company Holo 3.1 35B-A3B72.9861~55FRMoE 35B-A3BGGUF Q4_K_Mllama.cppStrix Halo 128GB2026-08-11
Nemotron-3-Nano-30B-A3B65.9945~50USMoE 31B-A3.5BGGUF Q4_K_Mllama.cppStrix Halo 128GB2026-08-11
Qwen-AgentWorld-35B-A3B61.2872~46CNMoE 35B-A3BGGUF UD-Q4_K_Mllama.cppStrix Halo 128GB2026-08-11
Gemma 4 31B IT10.4207~8USdense 31BGGUF Q4_K_Mllama.cppStrix Halo 128GB2026-08-11
Falcon-H1 34B9.9211~7AEhybrid 34BGGUF Q4_K_Mllama.cppStrix Halo 128GB2026-08-11
Seed-OSS 36B10.1178~8CNdense 36BGGUF Q4_K_Mllama.cppStrix Halo 128GB2026-08-11
LFM2.5-1.2B-Thinking979.8⚑58,202~735USdense 1.2BGGUF Q4_K_Mllama.cppRTX PRO 6000 96GB2026-08-12
LFM2.5-8B-A1B637.3⚑22,298~478USMoE 8.5B-A1BGGUF Q4_K_Mllama.cppRTX PRO 6000 96GB2026-08-12
LFM2.5-2.6B482.3⚑26,626~362USdense 2.7BGGUF Q4_K_Mllama.cppRTX PRO 6000 96GB2026-08-12
NVIDIA Nemotron-3.5-Lightning-30B-A3B298.4⚑9,716~224USMoE 33B-A3B hybrid (mamba)GGUF Q4_K_Mᶜllama.cppRTX PRO 6000 96GB7.02026-08-12
Agents-A1-4B285.1⚑13,290~214CNdense 4.2BGGUF Q4_K_Mllama.cppRTX PRO 6000 96GB2026-08-12
Agents-A1-35B-A3B ⁂258.6⚑8,048~194CNMoE 35B-A3BGGUF Q4_K_Mllama.cppRTX PRO 6000 96GB2026-08-12
KAT-Coder-V2.5-Dev253.8⚑8,254~190CNMoE 35B-A3BGGUF Q4_K_Mᶜllama.cppRTX PRO 6000 96GB2026-08-12
Ornith-1.0-35B ⁂250.2⚑8,183~188USMoE 35B-A3BGGUF Q4_K_Mllama.cppRTX PRO 6000 96GB2026-08-12
Nanbeige-4.2-3B213⚑9,592~160CNdense 4.2BGGUF Q4_K_Mᶜllama.cppRTX PRO 6000 96GB2026-08-12
Ornith-1.0-9B193.4⚑9,004~145USdense 9BGGUF Q4_K_Mllama.cppRTX PRO 6000 96GB2026-08-12
Muse-Glimmer-30B70.2⚑3,301~53USdense 30B (thinking)GGUF kquant ~4.5bpwᶜllama.cppRTX PRO 6000 96GB2026-08-12
Nemotron-Super-49B v141⚑1,821~31USdense 49BGGUF Q4_K_Mllama.cppRTX PRO 6000 96GB2026-08-12
Nemotron-Super-49B v1.540.5⚑1,812~30USdense 49BGGUF Q4_K_Mllama.cppRTX PRO 6000 96GB2026-08-12
LFM2.5-8B-A1B162.42,662~122USMoE 8.5B-A1BGGUF Q4_K_Mllama.cppStrix Halo 128GB2026-08-12
LFM2.5-2.6B111.72,772~84USdense 2.7BGGUF Q4_K_Mllama.cppStrix Halo 128GB2026-08-12
Agents-A1-35B-A3B ⁂72.4844~54CNMoE 35B-A3BGGUF Q4_K_Mllama.cppStrix Halo 128GB2026-08-12
Ornith-1.0-35B ⁂71.3845~53USMoE 35B-A3BGGUF Q4_K_Mllama.cppStrix Halo 128GB2026-08-12
NVIDIA Nemotron-3.5-Lightning-30B-A3B59.2⚑722~44USMoE 33B-A3B hybrid (mamba)GGUF Q4_K_Mᶜllama.cppStrix Halo 128GB2026-08-12
Mamba-Codestral-7B37.3⚑478~28FRpure mamba 7BGGUF Q4_0llama.cppStrix Halo 128GB2026-08-12
Muse-Glimmer-30B12.5⚑240~9USdense 30B (thinking)GGUF kquant ~4.5bpwᶜllama.cppStrix Halo 128GB2026-08-12
LFM2.5-1.2B-Thinking455.715,849~342USdense 1.2BGGUF Q4_K_Mllama.cppR9700 32GB2026-08-12
LFM2.5-8B-A1B303.77,350~228USMoE 8.5B-A1BGGUF Q4_K_Mllama.cppR9700 32GB2026-08-12
LFM2.5-2.6B2317,591~173USdense 2.7BGGUF Q4_K_Mllama.cppR9700 32GB2026-08-12
gpt-oss-20b168.43,270~126USMoE 21B-A3.6BGGUF MXFP4llama.cppR9700 32GB2026-08-12
Agents-A1-4B1314,464~98CNdense 4.2BGGUF Q4_K_Mllama.cppR9700 32GB2026-08-12
Nanbeige-4.2-3B108.5⚑2,210~81CNdense 4.2BGGUF Q4_K_Mᶜllama.cppR9700 32GB2026-08-12
Ornith-1.0-9B86.22,780~65USdense 9BGGUF Q4_K_Mllama.cppR9700 32GB2026-08-12
Nemotron-Nano-9B-v2732,177~55UShybrid 9B (mamba)GGUF Q4_K_Mᶜllama.cppR9700 32GB2026-08-12
IBM Granite-4.0-H-Small72.81,408~55USMoE 32B-A9B hybrid (mamba)GGUF Q4_K_Mllama.cppR9700 32GB2026-08-12
Apriel-1.5-15B-Thinker ◊62.31,481~47USdense 15BGGUF Q4_K_Mᶜllama.cppR9700 32GB2026-08-12
Phi-4-Reasoning61.31,475~46USdense 14BGGUF Q4_K_Mᶜllama.cppR9700 32GB2026-08-12
Reka-Flash-3.141.7850~31USdense 21BGGUF Q4_K_Mᶜllama.cppR9700 32GB2026-08-12
Magistral-Small40.6939~30FRdense 24BGGUF Q4_K_Mllama.cppR9700 32GB2026-08-12
Devstral-Small-250740.6938~30FRdense 24BGGUF Q4_K_Mllama.cppR9700 32GB2026-08-12
Mistral-Small-3.2-24B39.5939~30FRdense 24BGGUF Q4_K_Mllama.cppR9700 32GB2026-08-12
Muse-Glimmer-30B32.7⚑859~25USdense 30B (thinking)GGUF kquant ~4.5bpwᶜllama.cppR9700 32GB2026-08-12
Qwen-SEA-LION-v4.5-27B31.1837~23SGdense 27BGGUF Q4_K_Mllama.cppR9700 32GB2026-08-12
Qwen3-32B29614~22CNdense 32.8BGGUF Q4_K_Mllama.cppR9700 32GB2026-08-12
Olmo-3-32B-Think28.8679~22USdense 32BGGUF Q4_K_Mᶜllama.cppR9700 32GB2026-08-12
Two decode columns, one honest difference: live = salted timed completion direct to the serving engine, short prompt, includes prefill + one LAN hop, n=5–6 — the floor you'd feel. bench = pure generation timing, prompt cost excluded (llama-bench tg128 r=5, or 128-token warm gens n=2–3) — the number most sites publish. Where both exist they sit side by side; bench usually reads higher, and now you can see by how much. Where a row’s Quant/Engine cells carry a slash, the two decode columns came from different configs — left of the slash feeds the live cell, right feeds the bench cell — so that pair is a config difference, not a live-vs-bench method delta. Per-cell n and per-run samples: CSV.
‖ = re-measured 2026-08-11 after our own audit found these two cells were short bursts (82 and 21 tokens) published as full 128-token runs. Corrected: gpt-oss-120b 59.1 → 54.6 with prefill 469 → 508 (we were 8% too fast) and Flow-Judge 36.2 → 50.2 with prefill 751 → 1,502 (we were 39% too slow — that run was truncated and cold-started). Independent check on gpt-oss: a different quant of the same model on the same box reads 54.2, agreeing to 0.7%. ¶ = arch refusal on that box’s llama.cpp build, published rather than hidden. Flags: † = July one-shot (n=1), current-build re-runs queued. ‡ = adversarially re-swept with per-run samples in the CSV after an outside review challenged the zero-width range — stability real, zero width was rounding. ᶜ = community quant, not vendor-official (quantizer named in the CSV) — official checkpoints exist for Laguna (poolside) and Qwen3.6 NVFP4 (NVIDIA ModelOpt); on Muse-Glimmer-30B it also marks a quant deviation, a ~4.5 bpw k-quant standing in for a Q4_K_M that does not exist, so those rows are not same-quant comparable with the Q4_K_M ladder. poolside re-released the Laguna NVFP4 repo in Aug 2026; our revision is pinned in the CSV. The 250 W power-cap experiment lives in the postmortem — buyer data here, ops archaeology there.
Two rows carry a ceiling worth knowing before you buy on speed alone: Phi-4 and StarCoder2-15B are native 16,384-context models — they benched fine at our 32K protocol but warned, and they can never serve long context no matter which card you put them on. Speed is not the only axis. KV/tok is engine-reported at boot, not a formula — parameter count does not predict memory cost, and the spread in that column is 18× (7.0 to 128 KiB/token; the full KV ladder below spans 120×). Gaps there are queued: the current test window records it for every model booted. Blackwell “short” rows were measured with all three engines co-resident on one node — the as-deployed floor, not the card's isolated ceiling; single-tenant re-runs land with the three-card matrix. Value stat, card MSRP only: Sarvam 30B decodes 198.4 tok/s on the $1,299 R9700 — 152.8 tok/s per $1,000 of card. Words/s ≈ tok/s × 0.75, stated conversion.

// long context — what depth costs

Model · hardware@32K prefill / decode@100K prefill / decode@200K prefill / decode
poolside Laguna S 2.1 · RTX PRO 600012,775 / 69.18,892 / 62.05,926 / —*
Qwen3-Coder-Next · Strix Halo616 / 48.7426 / 37.2255 / 27.1
Kimi-Linear-48B · R9700 32GB ($1,299 card)2,039 / 131.62,095 / 131.6 (@128K)
Qwen3.5-122B-A10B · RTX PRO 6000
Cells are prefill t/s / decode tok/s at that actual prompt depth (measured 27–28K / 85–87K / 170–175K against the nominal buckets; n=3 per cell except the Strix 200K point, n=1 at 688 s per rep — one honest sample, disclosed). * = the 2-call method can't resolve vLLM decode at 200K; empty until a method can. Corrected 2026-08-12: the Qwen3.5-122B row now carries no cells. The prefill and decode figures printed there were a July single-shot vLLM run at short prompt, not a depth sweep on this protocol — the row stays listed, and empty, until one is run. Decode slows as the KV grows; prefill slows as the prompt grows — pick the box by your bottleneck. The 128K fit test: a 48B linear-attention model at full 131,072 context on the $1,299 32 GB card — full GPU offload, f16 KV, zero tweaks, 2.66 GiB headroom left. Going 32K → 128K cost 852 MiB and nothing in speed (131.6 both). We then pushed 38,400 real tokens of prompt through it at 1,463 tok/s — allocatable and usable are different claims, and this is the second one.

// the depth curves — does it slow down at 100K? measured, not vibes

The question every long-context spec sheet dodges: what speed do you have left once the context is actually full? We prefilled real tokens to each depth (tokenizer-calibrated per model, actual prompt lengths 26 / 32,671 / 65,440 / 99,001) and timed full 128-token decodes at every point — every generation length-verified, no truncated runs. Four architectures on the 128 GB unified-memory mini-PC, launched at 131,072 context, f16 KV. Nobody publishes this column. Every model slows. The only question is the slope — and parameter count does not set it: the 48B here decays faster than both the 80B and the 120B.
Model · Strix Halo 128GBAttentionKV @ 131Kdecode @ 0@ 32K@ 64K@ 100Kretained @ 100K
Qwen3-Next-80B-A3B (thinking)hybrid linear3.46 GiB59.045.239.634.558.6%
gpt-oss-120bsliding window4.49 GiB53.434.432.126.850.3%
Kimi-Linear-48B-A3Blinear1.40 GiB69.047.738.732.146.4%
Hunyuan-A13B-80Bfull attention16.54 GiB26.715.212.9cut††48.2% @ 64K
Same runs — the prefill wallprefill @ 32K@ 64K@ 100Ktime to first token at 100K depth
Qwen3-Next-80B-A3B604.5502.5405.74m 04s
Kimi-Linear-48B-A3B480.0328.1246.26m 42s
gpt-oss-120b364.4257.4177.49m 18s
Hunyuan-A13B-80B163.756.519m 18s just to reach 64K
What the curves say: (1) A fat KV cache is a hard tax at the extreme — but cache size alone does not rank the slope — Hunyuan carries 12× Kimi’s KV and is the only model to lose half its speed by 64K, and on unified memory that cache competes with the weights for the same bandwidth. Below that extreme the ranking inverts: Kimi has the smallest cache of the four (1.40 GiB) and the worst retention (46.4% @100K), while Qwen3-Next carries 2.5× the cache and retains best (58.6%). Parameter count predicts nothing either — 48B, 80B and 120B interleave in both columns. (2) Most of the damage lands in the first 32K (−23% to −43%) — which is exactly why capping context at 32K does not buy the speed back. By 32K you have already paid 56–83% of the total decode loss, so dropping from 100K back to 32K returns only 17–44% of what depth cost you (Qwen3-Next 44%, Kimi 42%, gpt-oss 29%, Hunyuan 17% measured to 64K). Budget for the deep number, not the shallow one. (3) Prefill is the real wall, not decode — decode at 100K is merely ~2× slower, but getting there costs minutes: 9m 18s of silence before gpt-oss-120b’s first token. For interactive use the binding constraint is time-to-first-token. (4) Sliding-window buys a plateau, not immunity — gpt-oss drops hard to 32K, then nearly flattens for the next 32K. †† = the Hunyuan 100K cell was cut on purpose and disclosed: its own curve projected ~35–40 minutes of prefill for that single number — the missing cell is the finding. Depth-0 speeds at 131K allocation land within −2.2% to +3.7% of our published 32K-protocol numbers for the same models on the same box (Hunyuan 26.7 vs 26.7, Kimi 69.0 vs 70.4, gpt-oss 53.4 vs 54.6, Qwen3-Next 59.0 vs 56.9); only Hunyuan reproduces to ±0.3%. The widest gap, Qwen3-Next at +3.7%, sits on top of a +3.4% spread between our own two runs of that model on that box, and the errors are signed both ways rather than systematically slower — so allocating a big context costs essentially nothing until you fill it. Blackwell cross-check, same night — both rows ⚑, measured on a transient build rather than the pinned build the Strix curves above ran on, so read them against each other and not against the rows above: Nemotron-3.5-Lightning-30B-A3B⚑ at 131,072 context costs ~0.9 GiB of KV, decodes 321 tok/s shallow and 254 tok/s with 100K tokens in front of it (−21%), prefilling those 100K in 10 seconds — the big card moves the wall, it doesn’t remove it. Meta’s Muse-Glimmer-30B⚑ at 128K on the same card — on a ~4.5 bpw community k-quantᶜ, not a Q4_K_M — holds −12% decode at 100K depth with 2,831 tok/s deep prefill.

// kv cache — the memory bill nobody prints

Every token of context costs memory beyond the weights — the KV cache. Spec sheets never print it, and it decides whether long context is nearly free or costs you a second GPU. Engine-reported values, converted across context depths. The spread below is 120× — and it tracks architecture, not model size.
ModelKV dtypeKiB / token@8K GiB@32K GiB@128K GiB@model max
Nemotron-Cascade-2-30B (6 attn / 52L)q8_03.190.020.100.390.80 @262K
Nemotron-3.5-Lightning-30B (hybrid mamba)f167.0*0.050.220.87(1M declared)
Kimi-Linear-48B (linear attn)f167.90.060.250.98
Qwen3.6-35B-A3B (hybrid)fp810.350.080.321.262.59 @262K
Cohere North Mini Code 1.0fp816.680.130.522.082.08 @131K
Qwen3-Coder-Next 80B (12 attn / 48L)f1624.00.180.753.06.0 @262K
gpt-oss-120b (iSWA)f1636.0*0.281.134.54.53 @131K
Gemma 4 31B (10 global + 50 SWA)q8_042.5*0.331.335.311.05 @262K
Cohere Command A Reasoning 111B (iSWA)f16~91*0.712.84
IBM Granite 4.1 30B (dense, 64L)fp8128.01.04.016.08.0 @64K served
Hunyuan-A13B-80Bf16128.01.04.016.5
Atla Selene 1 Mini 8Bf16128.01.02.0 @16K
KAT-Dev / Athene 72B (dense)f16320.02.510.040.0
Flow-Judge 3.8B (dense, fat heads)f16384.03.012.03.0 @8K
* = SWA/iSWA models: the scaling part only — their sliding-window layers add a small fixed pool, so long-context totals run slightly above the pure product. The 3.8B judge model costs 120× more memory per token than the 30B at the top: attention-layer count × KV-head count × head-dim × dtype, not parameter count. Measured sibling: at f16/32K the hybrid 100B-class models hold 0.63–1.27 GiB of KV where the dense-72B pair holds 10.0 GiB — 13× on the same box, same context. In English: if you want long documents in a home rig, pick a model with hybrid, linear, or sliding-window attention — classic dense makes context a second GPU.

// config experiments — same model, different config

ModelVariantResultBaselineThe lesson
poolside Laguna S 2.1speculative draft OFF100.7 tok/s294.7 w/ draftthe draft model is +192% — but KV goes 24.0 → 38.2 KiB/tok, so not free
Qwen3.5-122B-A10B--enforce-eager23.1 tok/s96.2 configuredthe most-copied “OOM fix” costs 76%
GLM-4.7-FlashMLA disabled (fallback)~470 KiB/tok KV177.2 tok/s configuredone flag = a ~470 KiB/token memory bill
Published on purpose: every number here is a config people actually copy from forums. Field reports had the 122B class at 23–33 tok/s — the gap to ours was these traps, not the hardware.

// pick by workload — what our own numbers say

Every arrow above cites a measurement on this page, not a spec sheet. If your workload isn't here, ask for it — requests section below.

// test queue — ranked by what people actually ask for

// how we test — no secrets, run it yourself

// requests — tell us what to measure

// updates — dated, append-only

[ + ] earlier updates (10)[ − ] earlier updates (10)
// directive4.ai · don't rent your intelligence · full stamps: vLLM/llama.cpp versions per row where recorded; gaps say so rather than guess · build #056