Measured on hardware we own. Click a column header to sort.
Every row is dated, stamped, and ships with its raw data. Words/s ≈ tok/s × 0.75 (stated conversion, English prose average — code and structured output tokenize differently, so treat it as prose-only). Corrections get published, not edited away.
The gate changed. An independent cross-check — a different model family, reading this page against its own raw data — now runs before rows land or change here. The first pass has run. It caught things; those are fixed in place, with the prior values still visible. The pass that clears the current page is still in flight, and this line gets updated when it lands — not before.
// the matrix — the under-50B class, three cards, decode tok/s
The under-50B class — models an ordinary card can hold. One model per row, same-family GGUF on all three hardware classes, pure-generation decode. “Built on” is the flag-accuracy column: a model’s country tells you who published it, not whose weights it started from. Holo is French and built on a Chinese base; Selene is British and built on an American one. “own” means the publisher pretrained it; “undisclosed” means we could not receipt a base model — a GGUF architecture tag names the code family that runs a model, not whose weights it started from, so we never publish an arch tag as lineage. Active-param counts written with a “~” are inferred from bandwidth arithmetic, not vendor-published (receipted in the CSV). ※ = the one model where the $1,299 card beat the workstation card — flagged for re-measurement, published as measured. ¶ = arch refusal on that box’s build, published rather than hidden: the model was attempted there and the engine refused to load its architecture (build named in the CSV). ⚑ = measured on a transient llama.cpp build newer than our pinned comparability build — several new architectures refuse to load on the pinned build at all, so the choice was a flagged number or no number (build hashes in the CSV; compare ⚑ cells to each other, not to unflagged ones). ⁂ = the same weights under two names: Ornith-1.0-35B (US) and Agents-A1-35B (CN) ship GGUF files 128 bytes apart with identical headers — 733 tensors, 256 experts — and bench within 3.4% everywhere, as identical files must. Two publishers, one model; we caught it because the numbers were implausibly identical. ◊ = ships a malformed chat template — benched with template parsing off; it cannot do tool-calling as shipped. ᶜ = community GGUF quant (quantizer in the CSV) where the vendor publishes none — Muse-Glimmer-30B is the sharpest case: no first-party GGUF exists for it at all, so its rows run a community ~4.5 bpw k-quant rather than the field-standard Q4_K_M, which makes its speeds approximate against the Q4_K_M ladder, not like-for-like. Four further models — an Indian 12B, a Korean 32B, and two US models — were benched but are held off this page pending license review: their licenses carry non-commercial, non-compete or read-before-publish clauses we read before publishing, not after. Bare dashes fill in public as test windows reach them; a dash carrying ¶ was attempted and refused. The 70B+ class gets its own table below.
Why these three: what’s feasible at home. One is the speed ceiling, one is a quiet whole-box with a huge memory pool, one is an ordinary $1,299 card. Read a row across and the trade-off is the answer.
Model
Origin
Params
Active
Type
RTX PRO 6000 96GB
Strix Halo 128GB
R9700 32GB
Built on
Quant
LFM2.5-1.2B-Thinking
US
1.2B
1.2B
dense
979.8⚑
—
455.7
own
Q4_K_M
LFM2.5-8B-A1B
US
8.5B
1B
MoE
637.3⚑
162.4
303.7
own
Q4_K_M
LFM2.5-2.6B
US
2.7B
2.7B
dense
482.3⚑
111.7
231.0
own
Q4_K_M
LFM2-24B-A2B
US
24B
2B
MoE
380.5
119.2
220.7
own
Q4_K_M
Sarvam 30B
IN
30B
~3B
MoE
364.6
83.7
198.4
own
Q4_K_M
NVIDIA Nemotron-3.5-Lightning-30B-A3B
US
32.9B
3B
MoE hybrid (mamba)
298.4⚑
59.2⚑
—
own
Q4_K_Mᶜ
Nemotron-3-Nano-30B-A3B
US
31.6B
3.5B
MoE hybrid (mamba)
286.4
65.9
137.4
own
Q4_K_M
Agents-A1-4B
CN
4.2B
4.2B
dense
285.1⚑
—
131.0
undisclosed
Q4_K_M
Qwen3-Coder-30B-A3B
CN
30.5B
3.3B
MoE
273.3
91.6
182.0
own
Q4_K_M
Cohere North Mini Code 1.0
CA
30B
3B
MoE
270.6⚑
—¶
134.0
undisclosed
UD-Q4_K_M
Nemotron-Cascade-2-30B-A3B
US
30B
3B
MoE hybrid (mamba)
268.5
62.3
131.9
own
Q4_K_M
Agents-A1-35B-A3B ⁂
CN
34.7B
3B
MoE
258.6⚑
72.4
—
Qwen3.5-35B (CN)
Q4_K_M
KAT-Coder-V2.5-Dev
CN
34.7B
3B
MoE
253.8⚑
—
—
undisclosed
Q4_K_Mᶜ
Ornith-1.0-35B ⁂
US
34.7B
3B
MoE
250.2⚑
71.3
—
Qwen3.5-35B (CN)
Q4_K_M
Kimi-Linear-48B-A3B
CN
48B
3B
MoE linear-attn
224.1
70.4
131.6
own
Q4_K_M
H Company Holo 3.1 35B-A3B
FR
35B
3B
MoE
221.3
72.9
127.3
Qwen3.6-35B (CN)
Q4_K_M
Nanbeige-4.2-3B
CN
4.2B
4.2B
dense
213.0⚑
—
108.5⚑
own
Q4_K_Mᶜ
Qwen-AgentWorld-35B-A3B
CN
35B
3B
MoE
206.4
61.2
100.2
own
UD-Q4_K_M / Q4_K_M (R9700, 2026-07)
Qwen3.6-35B-A3B
CN
35B
3B
MoE hybrid (SSM)
204.0
59.9
107.0
own
UD-Q4_K_M
GLM-4.7-Flash
CN
—
—
MoE
196.4
70.2
122.0
own
Q4_K_M
Ornith-1.0-9B
US
9B
9B
dense
193.4⚑
—
86.2
Qwen3.5-9B (CN)
Q4_K_M
Gemma-4-26B-A4B-it
US
26B
4B
MoE
190.0
52.5
100.7
own
UD-Q4_K_M
Mamba-Codestral-7B
FR
7B
7B
pure mamba
184.2
37.3⚑
85.5
own
Q4_0
Atla Selene-1-Mini 8B
UK
8B
8B
dense
162.3
27.0
69.9
Llama-3.1-8B (US)
Q8_0
DeepSeek-Coder-V2-Lite
CN
16B
2.4B
MoE
156.4
107.8
206.4※
own
Q4_K_M
Phi-4 14B
US
14B
14B
dense
133.2
23.2
62.0
own
Q4_K_M
StarCoder2-15B-Instruct
US
15B
15B
dense
115.6
20.6
56.1
own
Q4_K_M
IBM Granite 4.1 30B
US
30B
30B
dense
69.1
11.6
31.2
own
Q4_K_M
Muse-Glimmer-30B
US
30B
30B
dense
70.2⚑
12.5⚑
32.7⚑
own
kquant ~4.5bpwᶜ
Gemma 4 31B IT
US
31B
31B
dense
58.9
10.4
27.3
own
Q4_K_M
Seed-OSS 36B
CN
36B
36B
dense
57.4
10.1
26.8
own
Q4_K_M
Falcon-H1 34B
AE
34B
34B
hybrid (mamba)
51.8
9.9
24.5
own
Q4_K_M
Nemotron-Super-49B v1.5
US
49B
49B
dense
40.5⚑
—
—
Llama-3.3-70B (US)
Q4_K_M
Nemotron-Super-49B v1
US
49B
49B
dense
41.0⚑
—
—
Llama-3.3-70B (US)
Q4_K_M
gpt-oss-20b
US
21B
3.6B
MoE
—
—
168.4
own
MXFP4
Nemotron-Nano-9B-v2
US
9B
9B
hybrid (mamba)
—
—
73.0
own
Q4_K_Mᶜ
IBM Granite-4.0-H-Small
US
32B
9B
MoE hybrid (mamba)
—
—
72.8
own
Q4_K_M
Apriel-1.5-15B-Thinker
US
15B
15B
dense
—
—
62.3◊
own
Q4_K_Mᶜ
Phi-4-Reasoning
US
14B
14B
dense
—
—
61.3
own
Q4_K_Mᶜ
Reka-Flash-3.1
US
21B
21B
dense
—
—
41.7
own
Q4_K_Mᶜ
Magistral-Small
FR
24B
24B
dense
—
—
40.6
own
Q4_K_M
Devstral-Small-2507
FR
24B
24B
dense
—
—
40.6
own
Q4_K_M
Mistral-Small-3.2-24B
FR
24B
24B
dense
—
—
39.5
own
Q4_K_M
Qwen-SEA-LION-v4.5-27B
SG
27B
27B
dense
—
—
31.1
Qwen (CN)
Q4_K_M
Qwen3-32B
CN
32.8B
32.8B
dense
—
—
29.0
own
Q4_K_M
Olmo-3-32B-Think
US
32B
32B
dense
—
—
28.8
own
Q4_K_Mᶜ
2026-08-12: a 5 GB file took the top of the mini-PC column. LFM2.5-8B-A1B — 8.5B parameters, ~1B active — posts 637⚑ / 162 / 304 tok/s across all three cards, and it takes the Strix Halo column outright at 162.4, a 36% jump in that box’s sustained record (the previous holder, LFM2-24B-A2B, sat at 119.2). It does not sweep the matrix — its own 1.2B sibling leads the other two columns, at 979.8⚑ on the workstation card and 455.7 on the $1,299 card, with no Strix Halo leg yet. At the other end, the same night’s head-to-head: Nemotron-Super-49B v1 vs v1.5 is a dead heat (41.0 vs 40.5, identical VRAM to the MiB) — the .5 is a quality release, and a throughput bench cannot see quality. Same lesson from the qwen35moe cohort: Ornith and Agents-A1 (⁂ — one file under two names) plus KAT-Coder, a third publisher on the same architecture, all land within 3.4% of each other on the card where all three ran — choosing between them is a license and quality question, not a speed question. Same-family GGUF quants only, so the cards are the variable — vLLM/NVFP4 configs live in the detail table below. The Granite row: one dense 30B, three cards, 69 / 12 / 31 tok/s. At 100B+ the pattern holds harder: active-param MoEs span 84–226 tok/s on the Blackwell card — 5.1B-active gpt-oss-120b at the ceiling, 17B-active Llama-4-Scout at the floor, the 105–122B middle clustered at 117–139 — where 111–123B dense runs 17–20, and dense-72B collapses to ~4 on unified memory. The gpt-oss row is now the identical file on both machines: the workstation card is 4.2× the throughput, at roughly 3.4× the power on the closest measured comparison (a 119B MoE at 225 W on that card against the mini-PC’s measured 67 W) — not the 6× we printed. The mini-PC’s sibling quant samples at 0.83 tok/s/W (within 0.7% of this file, so we cite it as the nearest measured point rather than as this file’s own number); the workstation card’s leg of this pair was never power-sampled, so we withdraw the “per watt the mini-PC ties or wins” claim rather than replace it with another unmeasured one. Wave-3 pattern: Strix Halo’s 30B class is a three-tier ladder — plain-attention A3B MoE at 61–92 tok/s, hybrids around 60–73, dense pinned at the ~10 tok/s bandwidth wall. And the two 80B “twins” (Coder-Next vs Next-thinking, same architecture, different finetune) benched identical to the decimal on Blackwell — architecture sets the row, not the finetune.
// the big memory pools — 70B+ where 96 GB addressable meets 128 GB unified
The 70–123B class needs more memory than an ordinary card carries, so this is a memory-pool fight: 96 GB of addressable VRAM (the workstation card) vs 128 GB of unified memory (the mini-PC, where CPU and GPU share one pool and a model can spill past what a discrete card could hold). Same GGUF discipline, same decode measure. Params/Active tell the story before the numbers do — active-param MoEs run 5–10× the dense speed at every size. And one thing to keep in mind reading the 30B matrix above against this table: nobody buys a 96 GB card to run one 30B. That card's real job is either this class — or several 30Bs at once: one recent bakeoff held five separate models on it at once, each on its own port. Co-residency economics is a queued comparison of its own.
Model
Origin
Params
Active
Type
RTX PRO 6000 96GB
Strix Halo 128GB
Quant
gpt-oss-120b
US
117B
5.1B
MoE (iSWA)
225.9
54.2
MXFP4 (same file)
Qwen3-Coder-Next (0129)
CN
80B
3B
MoE hybrid (linear)
174.4
59.0
Q4_K_M
Qwen3-Next-80B-A3B (thinking)
CN
80B
3B
MoE hybrid (linear)
174.3
56.9
Q4_K_M
Mistral Small 4 119B (2603)
FR
119B
—
MoE
138.9
26.0
Q4-class
Hunyuan-A13B-80B
CN
80B
13B
MoE
138.3
26.7
Q4-class
Sarvam 105B
IN
105B
—
MoE
131.7
—
Q4-class
Qwen3.5-122B-A10B
CN
122B
10B
MoE
117.8
—
Q4-class
Llama-4-Scout
US
109B
17B
MoE
84.0
—
Q4-class
KAT-Dev-72B
CN
72B
72B
dense
27.9
4.4
Q4-class
GLM-4.5-Air
CN
—
—
MoE
—
23.2
Q4-class
Cohere Command A 111B
CA
111B
111B
dense
19.8
—
Q4-class
Devstral 2 123B
FR
123B
123B
dense
17.8
—
Q4-class
Mistral Large 2407
FR
123B
123B
dense
17.3
—
Q4-class
Llama-3.3-70B
US
70B
70B
dense
—
5.1
Q4-class
Hermes 70B
US
70B
70B
dense
—
5.1
Q4-class
Athene-V2-72B
US
72B
72B
dense
—
4.4
Q4-class
Cohere Command A Reasoning 111B
CA
111B
111B
dense (iSWA)
—
3.3
Q4-class
Dashes are queued legs, not refusals — every model here loads on at least one box. Multi-model co-residency (several engines sharing one card, the as-deployed reality) is a queued comparison of its own — the Blackwell “short” rows in the detail table were measured exactly that way.
// the spill test — when the model doesn’t fit the card
These models are 44–59 GiB; the R9700 has 32 GB. llama.cpp does not refuse — it loads anyway, spilling the overflow into host RAM, and serves degraded. Nobody publishes this table. It answers a different question than every other number on this page — not “what does this card do?” but “what happens when you exceed it?” — which is why these rows live here and never beside protocol rows.
Model
Type
Weights on disk
Spilled to host RAM
decode tok/s
Qwen3-Coder-Next 80B
MoE A3B
45.09 GiB
14.21 GiB
21.61
Qwen3-Next-80B-A3B
MoE A3B
45.09 GiB
14.21 GiB
21.61
gpt-oss-120b (MXFP4)
MoE A5B
59.03 GiB
28.09 GiB
15.17
Hunyuan-A13B-80B
MoE A13B
45.43 GiB
18.08 GiB
9.41
KAT-Dev-72B
dense
44.16 GiB
22.24 GiB
1.75
Method, disclosed: these are load-and-probe figures — a single 128-token generation, no warm-ups, no sustained run, no power sampling, no prefill sweep — a different protocol from every other row on this page, which is why they get their own table instead of a column. VRAM pinned at 99.4–99.8% of the card in every case. The lesson stands: “it loaded” is the trap — a card that technically runs a model at walking pace is not a card that runs it. The MoE pattern is real, though: active-3B experts spill 3× more gracefully (21.6 tok/s) than a 72B dense crawl (1.75).
// performance per watt — the electric-bill column
Decode tok/s divided by GPU-reported power during sustained generation. This is where the $8K workstation card and the quiet mini-PC stop being rivals: on efficiency they’re peers for MoE models — and dense models are power hogs on both.
Model
RTX PRO 6000: tok/s @ W
per watt
Strix Halo: tok/s @ W
per watt
R9700: tok/s @ W
per watt
LFM2.5-8B-A1B
637.3 @ 292⚑
2.18⚠
162.4 @ 79
2.05
303.7 @ 254
1.20
LFM2.5-1.2B-Thinking (thin power sample)
979.8 @ 216⚑
4.55⚠
—
—
455.7 @ 248
1.84⚠
LFM2.5-2.6B
482.3 @ 302⚑
1.60⚠
111.7 @ 86
1.31
231.0 @ 295
0.78
LFM2-24B-A2B
378.2 @ 240
1.57
119.2 @ 77
1.54
220.7 @ 256
0.86
Cohere North Mini Code 1.0
270.6 @ 268⚑
1.01
—
—
134.0 @ —
—
GLM-4.7-Flash
196.4 @ 214
0.92
68.6 @ 82
0.84
122.0 @ 267
0.46
DeepSeek-Coder-V2-Lite
113.3 @ 128
0.89
104.0 @ 80
1.30
206.4 @ 263
0.78
Gemma-4-26B-A4B-it
190.0 @ 224
0.85
50.6 @ 78
0.65
100.7 @ 263
0.38
Mamba-Codestral-7B
184.2 @ 275
0.67
—
—
85.5 @ 290
0.29
Atla Selene-1-Mini 8B
162.3 @ 300 cap
0.54
26.5 @ 86
0.31
69.9 @ 300
0.23
Phi-4 14B
131.6 @ 300 cap
0.44
23.2 @ 90
0.26
62.0 @ 300
0.21
StarCoder2-15B
115.6 @ 300 cap
0.39
20.6 @ 90
0.23
56.1 @ 300
0.19
Sarvam 30B
364.6 @ 283
1.29
83.7 @ ~90
0.93
—
—
Qwen3-Coder-30B-A3B
273.3 @ 230
1.19
91.6 @ ~85
1.08
—
—
Nemotron-3-Nano-30B-A3B
286.4 @ 256
1.12
65.9 @ ~89
0.74
—
—
H Company Holo 3.1 35B-A3B
221.3 @ 216
1.03
72.9 @ ~79
0.93
—
—
Kimi-Linear-48B-A3B
224.1 @ 233
0.96
—
—
131.6 @ 217
0.61
Qwen-AgentWorld-35B-A3B
206.4 @ 221
0.93
61.2 @ ~78
0.78
—
—
Qwen3-Next-80B-A3B
174.3 @ 210
0.83
—
—
—
—
Qwen3-Coder-Next (0129)
174.4 @ 214
0.82
—
—
—
—
Gemma 4 31B IT (dense)
58.9 @ 300 cap
0.20
10.4 @ ~101
0.10
—
—
Seed-OSS 36B (dense)
57.4 @ 300 cap
0.19
10.1 @ ~103
0.10
—
—
Falcon-H1 34B (hybrid)
51.8 @ 300 cap
0.17
9.9 @ ~98
0.10
—
—
Qwen3.6-35B-A3B
—
—
—
—
107.0 @ —
—
Nemotron-Cascade-2-30B-A3B
—
—
—
—
131.9 @ —
—
IBM Granite 4.1 30B (dense)
69.1 @ 300 cap
0.23
—
—
31.2 @ —
—
NVIDIA Nemotron-3.5-Lightning-30B-A3B
298.4 @ 277⚑
1.08
59.2 @ 89⚑
0.66
—
—
Agents-A1-35B-A3B ⁂
258.6 @ 254⚑
1.02
72.4 @ 80
0.91
—
—
Ornith-1.0-35B ⁂
250.2 @ 246⚑
1.02
71.3 @ 83
0.86
—
—
KAT-Coder-V2.5-Dev
253.8 @ 258⚑
0.98
—
—
—
—
Agents-A1-4B
285.1 @ 301⚑
0.95
—
—
131.0 @ 281
0.47
Ornith-1.0-9B
193.4 @ 300 cap⚑
0.64
—
—
86.2 @ 300
0.29
gpt-oss-20b (MXFP4)
—
—
—
—
168.4 @ 300
0.56
Muse-Glimmer-30B (dense)
70.2 @ 300 cap⚑
0.23
12.5 @ 86⚑
0.15
32.7 @ 300⚑
0.11
Nemotron-Super-49B v1.5 (dense)
40.5 @ 300 cap⚑
0.135
—
—
—
—
Method, honestly: Blackwell watts = nvidia-smi median board power during a 512-token sustained gen; Strix Halo watts = sysfs/hwmon GPU range midpoint during a 384-token gen — GPU-reported, not wall power, so cross-box comparison is directional, not billing-grade. Basis, stated: the tok/s here is the 512-token sustained rate where one was sampled and the 128-token bench rate otherwise, because power is only meaningful over the window it was sampled on — so cells here can sit a few percent under the matrix above. Rows whose wattage was never sampled now read “@ —” with no per-watt figure, rather than borrowing a card-class assumption. From the 30B class up, every dense model slams the Blackwell 300 W cap for about a fifth of the MoE speed; small dense models are a real exception — LFM2.5-1.2B-Thinking runs 979.8 tok/s at 216 W, Mamba-Codestral-7B 184.2 at 275 W. In English: if you care about the power bill, pick MoE — and on the mini-PC the draw is nearly flat (77–103 W whatever you load), so there the only thing that moves efficiency is speed.
2026-08-12: the crown moved on the mini-PC and the floor dropped. LFM2.5-8B-A1B posts 2.18⚠ / 2.05 / 1.20 tok/s/W from a file one-third the size of the model it displaced. The Strix Halo figure is the receipted record — 2.05 on 79 W, 33% clear of the previous holder (LFM2-24B at 1.54) — and on the R9700 it beats that model 1.20 to 0.86. On the workstation card it is the best settled row, but its own 1.2B sibling reads higher on both the Blackwell (4.55⚠) and R9700 (1.84⚠) legs, and its Blackwell leg is a ⚑ transient build measured against LFM2-24B’s pinned one — so no all-cards crown is claimed here until the longer power window lands. The floor for this table is the 49B dense class: Nemotron-Super v1.5 pins the 300 W cap for 40.5 tok/s = 0.135 tok/s/W (v1 reads 41.0 = 0.137), 16× worse than the champion on the same card — and the 70B+ table below goes lower still. The MoE-vs-dense power split is a strong tendency, not a law: at 30B and up on the Blackwell card every cap-pinned row is dense while the A3B MoEs cruise 210–292 W, but small dense models stay well under the cap, and on the R9700 sparse models pin it too (gpt-oss-20b 300 W, Granite-4.0-H-Small 299 W). ⚠ = the power median rests on 3–8 samples (4.55 = 3, 2.18 = 5, 1.60 = 6) — indicative, not settled; rows carrying it are not counted as records on this page until a 2,048-token window re-measures them.
70B+ class on the workstation card
tok/s @ W
per watt
Mistral Small 4 119B (MoE)
138.9 @ 225
0.62
Hunyuan-A13B-80B (MoE)
138.3 @ 300
0.46
Llama-4-Scout (MoE A17B)
84.0 @ 305
0.28
Corrected 2026-08-12: this table used to carry ten rows. Seven of them — gpt-oss-120b, Sarvam 105B, Qwen3.5-122B-A10B, KAT-Dev-72B, Cohere Command A 111B, Devstral 2 123B and Mistral Large 2407 — had no power sample behind them; their wattage was assumed from the card’s 300 W cap, not measured. Our own audit caught it and removed them; the three rows left are the ones whose watts are in the receipts. Their throughput is unaffected and still published in the tables above. What the three surviving rows show: even inside one architecture family and one card, per-watt spans 0.62 down to 0.28 — and the top row earns it partly by drawing 225 W where the others pin the 300 W cap, so read the watts column, not just the ratio. The dense-vs-MoE throughput gap is real but lives in the tables above; the dense rows are not in this one, because their power was never sampled. Strix Halo watts for this class: one measured point so far — the mini-PC’s sibling quant of the gpt-oss pair sampled at 67 W and 0.83 tok/s/W. That is a measurement of the sibling file, not of the MXFP4 file in the pair — the two read within 0.7% of each other on that box, which is why we cite it, and it is still an inference across quants rather than a direct sample. It puts the mini-PC near one-third the workstation card’s draw on the nearest comparable row, not the one-sixth we printed. Multi-model co-residency economics: queued.
// the bench — hardware classes
RTX Blackwell nodes — NVIDIA RTX PRO 6000 Blackwell: 96 GB GDDR7 ECC · 512-bit · 1,792 GB/s · 24,064 CUDA cores · 752 5th-gen Tensor cores · 188 RT cores · PCIe 5.0 ×16. The watts shown in the perf/watt table are GPU-reported draw during the run, not a per-node limit, and node identity lives in the CSV rather than in a page column. The 250 W power-cap experiment lives in the L4B postmortem — no row on this page was measured under it. In English: the fastest workstation card — this page’s speed ceiling.
Strix Halo nodes — AMD Ryzen AI Max+ 395: 16× Zen 5 (32 threads, 5.1 GHz boost) · Radeon 8060S iGPU, 40 RDNA 3.5 CUs · XDNA 2 NPU · 128 GB unified LPDDR5X-8000 · 256-bit · 256 GB/s · 45–120 W cTDP. Slow lane, huge memory — MoE country. In English: a quiet mini-PC where CPU and GPU share one memory pool — holds the biggest models on little power.
RDNA4 nodes — AMD Radeon AI PRO R9700: RDNA 4 (Navi 48) · 32 GB GDDR6 (ECC on Linux) · 256-bit · 640 GB/s · 64 CUs / 4,096 stream processors · 128 AI accelerators · 64 MB Infinity Cache · PCIe 5.0 · 300 W TBP · $1,299 MSRP. In English: an ordinary graphics card, one slot, $1,299.
Specs from vendor datasheets/whitepapers (NVIDIA RTX Blackwell PRO architecture whitepaper Table 4; AMD product pages), fetched and verified 2026-08-11. Plural nodes per class; per-row stamps carry the exact box config.
// the detail — every engine config, one row each [ + 128 rows — click to open ][ − collapse ]
128 rows · one row per model × hardware × quant · 3 hardware classes · origins US / CN / EU / FR / UK / CA / IN / AE / SG. One row per model × hardware × quant — a model appears once. Sort any column, filter by class. A “—” means not yet measured — never a guess. Depth curves and config experiments live in their own sections below.
Model
Decode live
Decode bench
Prefill t/s
~words/s
Origin
Arch
Quant
Engine
Hardware
KV/tok KiB
Date
poolside Laguna S 2.1 + draft
119.4
294.7
—
~90
US
MoE 118B-A8.5B
NVFP4
vLLM
RTX PRO 6000 96GB
38.2
2026-08
Mistral Small 4 119B (2603)
—
137.0
—
~103
FR
MoE 119B
NVFP4
vLLM
RTX PRO 6000 96GB
—
2026-07
Qwen3.6-35B-A3B ‡
238.4
—
—
~179
CN
MoE 35B-A3B
NVFP4 (NVIDIA)
vLLM
RTX PRO 6000 96GB
10.35
2026-08-11
Cohere North Mini Code 1.0
195.9
194.5
≥15,200
~147
CA
MoE 30B-A3B
NVFP4ᶜ
vLLM
RTX PRO 6000 96GB
16.68
2026-08-11
GLM-4.7-Flash
—
177.2
11,535
~133
CN
MoE
NVFP4ᶜ
vLLM
RTX PRO 6000 96GB
—
2026-08-04
Qwen3.5-122B-A10B †
—
96.2
—
~72
CN
MoE 122B-A10B
NVFP4
vLLM
RTX PRO 6000 96GB
—
2026-07
gpt-oss-120b
—
225.9
—
~169
US
MoE 117B-A5B
GGUF MXFP4
llama.cpp
RTX PRO 6000 96GB
36.0
2026-08-11
Nemotron-Cascade-2-30B-A3B
—
268.5
—
~201
US
MoE 30B-A3B
GGUF Q4_K_M
llama.cpp
RTX PRO 6000 96GB
—
2026-08-11
Qwen3.6-35B-A3B
—
204.0
6,364
~153
CN
MoE 35B-A3B
GGUF UD-Q4_K_M
llama.cpp
RTX PRO 6000 96GB
—
2026-08-11
IBM Granite 4.1 30B
73.0
69.1
2,956
~55
US
dense 30B
NVFP4ᶜ / GGUF
vLLM / llama.cpp
RTX PRO 6000 96GB
128
2026-08-11
gpt-oss-120b
51.8
54.6‖
508‖
~39
US
MoE 117B-A5B
GGUF UD-Q4_K_XL
llama.cpp
Strix Halo 128GB
36.0
2026-08-11
gpt-oss-120b (same-quant pair)
—
54.2
—
~41
US
MoE 117B-A5B
GGUF MXFP4
llama.cpp
Strix Halo 128GB
—
2026-08-11
Qwen3-Coder-Next (0129)
55.5
59.0
715
~42
CN
MoE 80B-A3B
GGUF Q4_K_M
llama.cpp
Strix Halo 128GB
—
2026-08-11
Nemotron-Cascade-2-30B-A3B
53.9
62.3
1,026
~40
US
MoE 30B-A3B
GGUF Q4_K_M
llama.cpp
Strix Halo 128GB
—
2026-08-11
Qwen3.6-35B-A3B
—
59.9
—
~45
CN
MoE 35B-A3B
GGUF UD-Q4_K_M
llama.cpp
Strix Halo 128GB
—
2026-08-11
Kimi-Linear-48B-A3B-Instruct
64.4
70.4
683
~53
CN
MoE 48B-A3B (linear attn)
GGUF Q4_K_M
llama.cpp
Strix Halo 128GB
7.9
2026-08-11
SmolLM3-3B
91.2
—
—
~68
EU
dense 3B
GGUF Q4_K_M
llama.cpp
Strix Halo 128GB
—
2026-08-11
Phi-4-mini
71.6
—
—
~54
US
dense 3.8B
GGUF Q4
llama.cpp
Strix Halo 128GB
—
2026-08-11
Flow-Judge v0.1 (3.8B)
47.7
50.2‖
1,502
~36
EU
dense 3.8B
GGUF Q8_0
llama.cpp
Strix Halo 128GB
—
2026-08-11
Qwen3-8B
40.6
—
—
~30
CN
dense 8B
GGUF Q4_K_M
llama.cpp
Strix Halo 128GB
—
2026-08-11
Atla Selene 1 Mini (8B)
24.9
27.0
871
~19
UK
dense 8B
GGUF Q8_0
llama.cpp
Strix Halo 128GB
—
2026-08-11
IBM Granite 4.1 30B
—
11.6
—
~9
US
dense 30B
GGUF Q4_K_M
llama.cpp
Strix Halo 128GB
—
2026-08-11
Qwen3.6-35B-A3B
—
107.0
2,502
~80
CN
MoE 35B-A3B
GGUF UD-Q4_K_M
llama.cpp
R9700 32GB
—
2026-08-11
Nemotron-Cascade-2-30B-A3B
—
131.9
2,825
~99
US
MoE 30B-A3B
GGUF Q4_K_M
llama.cpp
R9700 32GB
—
2026-08-11
LFM2-24B-A2B
—
380.5
9,990
~285
US
MoE 24B-A2B
GGUF Q4_K_M
llama.cpp
RTX PRO 6000 96GB
—
2026-08-11
Cohere North Mini Code 1.0
—
270.6⚑
7,759
~203
CA
MoE 30B-A3B
GGUF UD-Q4_K_M
llama.cpp
RTX PRO 6000 96GB
—
2026-08-11
GLM-4.7-Flash
—
196.4
5,972
~147
CN
MoE
GGUF Q4_K_M
llama.cpp
RTX PRO 6000 96GB
—
2026-08-11
Gemma-4-26B-A4B-it
—
190.0
1,330
~143
US
MoE 26B-A4B
GGUF UD-Q4_K_M
llama.cpp
RTX PRO 6000 96GB
—
2026-08-11
Mamba-Codestral-7B
—
184.2
5,772
~138
FR
pure mamba 7B
GGUF Q4_0
llama.cpp
RTX PRO 6000 96GB
—
2026-08-11
Atla Selene-1-Mini 8B
—
162.3
11,996
~122
UK
dense 8B
GGUF Q8_0
llama.cpp
RTX PRO 6000 96GB
—
2026-08-11
DeepSeek-Coder-V2-Lite
—
156.4
—
~117
CN
MoE 16B-A2.4B
GGUF Q4_K_M
llama.cpp
RTX PRO 6000 96GB
—
2026-08-11
Phi-4 14B
—
133.2
—
~100
US
dense 14B
GGUF Q4_K_M
llama.cpp
RTX PRO 6000 96GB
—
2026-08-11
StarCoder2-15B-Instruct
—
115.6
5,088
~87
US
dense 15B
GGUF Q4_K_M
llama.cpp
RTX PRO 6000 96GB
—
2026-08-11
LFM2-24B-A2B
—
220.7
4,015
~166
US
MoE 24B-A2B
GGUF Q4_K_M
llama.cpp
R9700 32GB
—
2026-08-11
DeepSeek-Coder-V2-Lite
—
206.4
3,837
~155
CN
MoE 16B-A2.4B
GGUF Q4_K_M
llama.cpp
R9700 32GB
—
2026-08-11
GLM-4.7-Flash
—
122.0
1,493
~92
CN
MoE
GGUF Q4_K_M
llama.cpp
R9700 32GB
—
2026-08-11
Gemma-4-26B-A4B-it
—
100.7
2,670
~76
US
MoE 26B-A4B
GGUF UD-Q4_K_M
llama.cpp
R9700 32GB
—
2026-08-11
Mamba-Codestral-7B
—
85.5
2,339
~64
FR
pure mamba 7B
GGUF Q4_0
llama.cpp
R9700 32GB
—
2026-08-11
Atla Selene-1-Mini 8B
—
69.9
2,804
~52
UK
dense 8B
GGUF Q8_0
llama.cpp
R9700 32GB
—
2026-08-11
Phi-4 14B
—
62.0
1,478
~47
US
dense 14B
GGUF Q4_K_M
llama.cpp
R9700 32GB
—
2026-08-11
StarCoder2-15B-Instruct
—
56.1
1,108
~42
US
dense 15B
GGUF Q4_K_M
llama.cpp
R9700 32GB
—
2026-08-11
Kimi-Linear-48B-A3B
—
131.6
2,039
~99
CN
MoE 48B-A3B (linear attn)
GGUF Q4_K_M
llama.cpp
R9700 32GB
7.9
2026-08-11
Cohere North Mini Code 1.0
—
134.0
2,362
~101
CA
MoE 30B-A3B
GGUF UD-Q4_K_M
llama.cpp
R9700 32GB
—
2026-08-11
Mistral Small 4 119B (2603)
—
138.9
~295 (anomaly, re-measure queued)
~104
FR
MoE 119B
GGUF Q4-class
llama.cpp
RTX PRO 6000 96GB
—
2026-08-11
Hunyuan-A13B-80B
—
138.3
3,229
~104
CN
MoE 80B-A13B
GGUF Q4-class
llama.cpp
RTX PRO 6000 96GB
—
2026-08-11
Sarvam 105B
—
131.7
2,931
~99
IN
MoE 105B
GGUF Q4-class
llama.cpp
RTX PRO 6000 96GB
—
2026-08-11
Qwen3.5-122B-A10B
—
117.8
3,106
~88
CN
MoE 122B-A10B
GGUF Q4-class
llama.cpp
RTX PRO 6000 96GB
—
2026-08-11
Llama-4-Scout
—
84.0
2,347
~63
US
MoE A17B
GGUF Q4-class
llama.cpp
RTX PRO 6000 96GB
—
2026-08-11
KAT-Dev-72B
—
27.9
1,252
~21
CN
dense 72B
GGUF Q4-class
llama.cpp
RTX PRO 6000 96GB
—
2026-08-11
Cohere Command A 111B
—
19.8
778
~15
CA
dense 111B
GGUF Q4-class
llama.cpp
RTX PRO 6000 96GB
—
2026-08-11
Devstral 2 123B
—
17.8
702
~13
FR
dense 123B
GGUF Q4-class
llama.cpp
RTX PRO 6000 96GB
—
2026-08-11
Mistral Large 2407
—
17.3
697
~13
FR
dense 123B
GGUF Q4-class
llama.cpp
RTX PRO 6000 96GB
—
2026-08-11
Hunyuan-A13B-80B
—
26.7
200
~20
CN
MoE 80B-A13B
GGUF Q4-class
llama.cpp
Strix Halo 128GB
—
2026-08-11
KAT-Dev-72B
—
4.4
—
~3
CN
dense 72B
GGUF Q4-class
llama.cpp
Strix Halo 128GB
—
2026-08-11
Athene-V2-72B
—
4.4
—
~3
US
dense 72B
GGUF Q4-class
llama.cpp
Strix Halo 128GB
—
2026-08-11
Cohere Command A Reasoning 111B
—
3.3
—
~2
CA
dense 111B
GGUF Q4-class
llama.cpp
Strix Halo 128GB
—
2026-08-11
Qwen3-Next-80B-A3B
—
56.9
—
~43
CN
MoE 80B-A3B
GGUF Q4-class
llama.cpp
Strix Halo 128GB
—
2026-08-11
Mistral Small 4 119B (2603)
—
26.0
—
~20
FR
MoE 119B (MLA)
GGUF Q4-class
llama.cpp
Strix Halo 128GB
—
2026-08-11
GLM-4.5-Air
—
23.2
—
~17
CN
MoE
GGUF Q4-class
llama.cpp
Strix Halo 128GB
—
2026-08-11
Llama-3.3-70B
—
5.1
—
~4
US
dense 70B
GGUF Q4-class
llama.cpp
Strix Halo 128GB
—
2026-08-11
Hermes 70B
—
5.1
—
~4
US
dense 70B
GGUF Q4-class
llama.cpp
Strix Halo 128GB
—
2026-08-11
Sarvam 30B
—
198.4
3,613
~149
IN
MoE 32B
GGUF Q4_K_M
llama-bench
R9700 32GB
—
2026-08-11
Qwen3-Coder-30B-A3B
—
182.0
3,035
~137
CN
MoE 30B-A3B
GGUF Q4_K_M
llama-bench
R9700 32GB
—
2026-08-11
Nemotron-3-Nano-30B-A3B
—
137.4
2,677
~103
US
MoE 31B-A3.5B
GGUF Q4_K_M
llama-bench
R9700 32GB
—
2026-08-11
H Company Holo 3.1 35B-A3B
—
127.3
2,681
~95
FR
MoE 35B-A3B
GGUF Q4_K_M
llama-bench
R9700 32GB
—
2026-08-11
Qwen-AgentWorld-35B-A3B
—
100.2
415
~75
CN
MoE 35B-A3B
GGUF Q4_K_M
llama.cpp
R9700 32GB
—
2026-07
IBM Granite 4.1 30B
—
31.2
848
~23
US
dense 30B
GGUF Q4_K_M
llama-bench
R9700 32GB
—
2026-08-11
Gemma 4 31B IT
25.5
27.3
756
~19
US
dense 31B
GGUF Q4_K_M
llama.cpp
R9700 32GB
—
2026-08-11
Seed-OSS 36B
—
26.8
707
~20
CN
dense 36B
GGUF Q4_K_M
llama-bench
R9700 32GB
—
2026-08-11
Falcon-H1 34B
—
24.5
725
~18
AE
hybrid 34B
GGUF Q4_K_M
llama-bench
R9700 32GB
—
2026-08-11
Sarvam 30B
—
364.6
11,021
~273
IN
MoE 30B (A3B-class)
GGUF Q4_K_M
llama.cpp
RTX PRO 6000 96GB
—
2026-08-11
Nemotron-3-Nano-30B-A3B
—
286.4
8,754
~215
US
MoE 31B-A3.5B
GGUF Q4_K_M
llama.cpp
RTX PRO 6000 96GB
—
2026-08-11
Qwen3-Coder-30B-A3B
—
273.3
7,194
~205
CN
MoE 30B-A3B
GGUF Q4_K_M
llama.cpp
RTX PRO 6000 96GB
—
2026-08-11
Kimi-Linear-48B-A3B
—
224.1
6,874
~168
CN
MoE 48B-A3B (linear attn)
GGUF Q4_K_M
llama.cpp
RTX PRO 6000 96GB
7.9
2026-08-11
H Company Holo 3.1 35B-A3B
—
221.3
6,917
~166
FR
MoE 35B-A3B
GGUF Q4_K_M
llama.cpp
RTX PRO 6000 96GB
—
2026-08-11
Qwen-AgentWorld-35B-A3B
—
206.4
7,543
~155
CN
MoE 35B-A3B
GGUF UD-Q4_K_M
llama.cpp
RTX PRO 6000 96GB
—
2026-08-11
Qwen3-Coder-Next (0129)
—
174.4
4,258
~131
CN
MoE 80B-A3B
GGUF Q4_K_M
llama.cpp
RTX PRO 6000 96GB
24.0
2026-08-11
Qwen3-Next-80B-A3B (thinking)
—
174.3
4,430
~131
CN
MoE 80B-A3B
GGUF Q4_K_M
llama.cpp
RTX PRO 6000 96GB
—
2026-08-11
Gemma 4 31B IT
—
58.9
927
~44
US
dense 31B
GGUF Q4_K_M
llama.cpp
RTX PRO 6000 96GB
42.5
2026-08-11
Seed-OSS 36B
—
57.4
2,296
~43
CN
dense 36B
GGUF Q4_K_M
llama.cpp
RTX PRO 6000 96GB
—
2026-08-11
Falcon-H1 34B
—
51.8
2,190
~39
AE
hybrid 34B
GGUF Q4_K_M
llama.cpp
RTX PRO 6000 96GB
—
2026-08-11
Qwen3-Coder-30B-A3B
—
91.6
993
~69
CN
MoE 30B-A3B
GGUF Q4_K_M
llama.cpp
Strix Halo 128GB
—
2026-08-11
Sarvam 30B
—
83.7
1,121
~63
IN
MoE 30B (A3B-class)
GGUF Q4_K_M
llama.cpp
Strix Halo 128GB
—
2026-08-11
H Company Holo 3.1 35B-A3B
—
72.9
861
~55
FR
MoE 35B-A3B
GGUF Q4_K_M
llama.cpp
Strix Halo 128GB
—
2026-08-11
Nemotron-3-Nano-30B-A3B
—
65.9
945
~50
US
MoE 31B-A3.5B
GGUF Q4_K_M
llama.cpp
Strix Halo 128GB
—
2026-08-11
Qwen-AgentWorld-35B-A3B
—
61.2
872
~46
CN
MoE 35B-A3B
GGUF UD-Q4_K_M
llama.cpp
Strix Halo 128GB
—
2026-08-11
Gemma 4 31B IT
—
10.4
207
~8
US
dense 31B
GGUF Q4_K_M
llama.cpp
Strix Halo 128GB
—
2026-08-11
Falcon-H1 34B
—
9.9
211
~7
AE
hybrid 34B
GGUF Q4_K_M
llama.cpp
Strix Halo 128GB
—
2026-08-11
Seed-OSS 36B
—
10.1
178
~8
CN
dense 36B
GGUF Q4_K_M
llama.cpp
Strix Halo 128GB
—
2026-08-11
LFM2.5-1.2B-Thinking
—
979.8⚑
58,202
~735
US
dense 1.2B
GGUF Q4_K_M
llama.cpp
RTX PRO 6000 96GB
—
2026-08-12
LFM2.5-8B-A1B
—
637.3⚑
22,298
~478
US
MoE 8.5B-A1B
GGUF Q4_K_M
llama.cpp
RTX PRO 6000 96GB
—
2026-08-12
LFM2.5-2.6B
—
482.3⚑
26,626
~362
US
dense 2.7B
GGUF Q4_K_M
llama.cpp
RTX PRO 6000 96GB
—
2026-08-12
NVIDIA Nemotron-3.5-Lightning-30B-A3B
—
298.4⚑
9,716
~224
US
MoE 33B-A3B hybrid (mamba)
GGUF Q4_K_Mᶜ
llama.cpp
RTX PRO 6000 96GB
7.0
2026-08-12
Agents-A1-4B
—
285.1⚑
13,290
~214
CN
dense 4.2B
GGUF Q4_K_M
llama.cpp
RTX PRO 6000 96GB
—
2026-08-12
Agents-A1-35B-A3B ⁂
—
258.6⚑
8,048
~194
CN
MoE 35B-A3B
GGUF Q4_K_M
llama.cpp
RTX PRO 6000 96GB
—
2026-08-12
KAT-Coder-V2.5-Dev
—
253.8⚑
8,254
~190
CN
MoE 35B-A3B
GGUF Q4_K_Mᶜ
llama.cpp
RTX PRO 6000 96GB
—
2026-08-12
Ornith-1.0-35B ⁂
—
250.2⚑
8,183
~188
US
MoE 35B-A3B
GGUF Q4_K_M
llama.cpp
RTX PRO 6000 96GB
—
2026-08-12
Nanbeige-4.2-3B
—
213⚑
9,592
~160
CN
dense 4.2B
GGUF Q4_K_Mᶜ
llama.cpp
RTX PRO 6000 96GB
—
2026-08-12
Ornith-1.0-9B
—
193.4⚑
9,004
~145
US
dense 9B
GGUF Q4_K_M
llama.cpp
RTX PRO 6000 96GB
—
2026-08-12
Muse-Glimmer-30B
—
70.2⚑
3,301
~53
US
dense 30B (thinking)
GGUF kquant ~4.5bpwᶜ
llama.cpp
RTX PRO 6000 96GB
—
2026-08-12
Nemotron-Super-49B v1
—
41⚑
1,821
~31
US
dense 49B
GGUF Q4_K_M
llama.cpp
RTX PRO 6000 96GB
—
2026-08-12
Nemotron-Super-49B v1.5
—
40.5⚑
1,812
~30
US
dense 49B
GGUF Q4_K_M
llama.cpp
RTX PRO 6000 96GB
—
2026-08-12
LFM2.5-8B-A1B
—
162.4
2,662
~122
US
MoE 8.5B-A1B
GGUF Q4_K_M
llama.cpp
Strix Halo 128GB
—
2026-08-12
LFM2.5-2.6B
—
111.7
2,772
~84
US
dense 2.7B
GGUF Q4_K_M
llama.cpp
Strix Halo 128GB
—
2026-08-12
Agents-A1-35B-A3B ⁂
—
72.4
844
~54
CN
MoE 35B-A3B
GGUF Q4_K_M
llama.cpp
Strix Halo 128GB
—
2026-08-12
Ornith-1.0-35B ⁂
—
71.3
845
~53
US
MoE 35B-A3B
GGUF Q4_K_M
llama.cpp
Strix Halo 128GB
—
2026-08-12
NVIDIA Nemotron-3.5-Lightning-30B-A3B
—
59.2⚑
722
~44
US
MoE 33B-A3B hybrid (mamba)
GGUF Q4_K_Mᶜ
llama.cpp
Strix Halo 128GB
—
2026-08-12
Mamba-Codestral-7B
—
37.3⚑
478
~28
FR
pure mamba 7B
GGUF Q4_0
llama.cpp
Strix Halo 128GB
—
2026-08-12
Muse-Glimmer-30B
—
12.5⚑
240
~9
US
dense 30B (thinking)
GGUF kquant ~4.5bpwᶜ
llama.cpp
Strix Halo 128GB
—
2026-08-12
LFM2.5-1.2B-Thinking
—
455.7
15,849
~342
US
dense 1.2B
GGUF Q4_K_M
llama.cpp
R9700 32GB
—
2026-08-12
LFM2.5-8B-A1B
—
303.7
7,350
~228
US
MoE 8.5B-A1B
GGUF Q4_K_M
llama.cpp
R9700 32GB
—
2026-08-12
LFM2.5-2.6B
—
231
7,591
~173
US
dense 2.7B
GGUF Q4_K_M
llama.cpp
R9700 32GB
—
2026-08-12
gpt-oss-20b
—
168.4
3,270
~126
US
MoE 21B-A3.6B
GGUF MXFP4
llama.cpp
R9700 32GB
—
2026-08-12
Agents-A1-4B
—
131
4,464
~98
CN
dense 4.2B
GGUF Q4_K_M
llama.cpp
R9700 32GB
—
2026-08-12
Nanbeige-4.2-3B
—
108.5⚑
2,210
~81
CN
dense 4.2B
GGUF Q4_K_Mᶜ
llama.cpp
R9700 32GB
—
2026-08-12
Ornith-1.0-9B
—
86.2
2,780
~65
US
dense 9B
GGUF Q4_K_M
llama.cpp
R9700 32GB
—
2026-08-12
Nemotron-Nano-9B-v2
—
73
2,177
~55
US
hybrid 9B (mamba)
GGUF Q4_K_Mᶜ
llama.cpp
R9700 32GB
—
2026-08-12
IBM Granite-4.0-H-Small
—
72.8
1,408
~55
US
MoE 32B-A9B hybrid (mamba)
GGUF Q4_K_M
llama.cpp
R9700 32GB
—
2026-08-12
Apriel-1.5-15B-Thinker ◊
—
62.3
1,481
~47
US
dense 15B
GGUF Q4_K_Mᶜ
llama.cpp
R9700 32GB
—
2026-08-12
Phi-4-Reasoning
—
61.3
1,475
~46
US
dense 14B
GGUF Q4_K_Mᶜ
llama.cpp
R9700 32GB
—
2026-08-12
Reka-Flash-3.1
—
41.7
850
~31
US
dense 21B
GGUF Q4_K_Mᶜ
llama.cpp
R9700 32GB
—
2026-08-12
Magistral-Small
—
40.6
939
~30
FR
dense 24B
GGUF Q4_K_M
llama.cpp
R9700 32GB
—
2026-08-12
Devstral-Small-2507
—
40.6
938
~30
FR
dense 24B
GGUF Q4_K_M
llama.cpp
R9700 32GB
—
2026-08-12
Mistral-Small-3.2-24B
—
39.5
939
~30
FR
dense 24B
GGUF Q4_K_M
llama.cpp
R9700 32GB
—
2026-08-12
Muse-Glimmer-30B
—
32.7⚑
859
~25
US
dense 30B (thinking)
GGUF kquant ~4.5bpwᶜ
llama.cpp
R9700 32GB
—
2026-08-12
Qwen-SEA-LION-v4.5-27B
—
31.1
837
~23
SG
dense 27B
GGUF Q4_K_M
llama.cpp
R9700 32GB
—
2026-08-12
Qwen3-32B
—
29
614
~22
CN
dense 32.8B
GGUF Q4_K_M
llama.cpp
R9700 32GB
—
2026-08-12
Olmo-3-32B-Think
—
28.8
679
~22
US
dense 32B
GGUF Q4_K_Mᶜ
llama.cpp
R9700 32GB
—
2026-08-12
Two decode columns, one honest difference:live = salted timed completion direct to the serving engine, short prompt, includes prefill + one LAN hop, n=5–6 — the floor you'd feel. bench = pure generation timing, prompt cost excluded (llama-bench tg128 r=5, or 128-token warm gens n=2–3) — the number most sites publish. Where both exist they sit side by side; bench usually reads higher, and now you can see by how much. Where a row’s Quant/Engine cells carry a slash, the two decode columns came from different configs — left of the slash feeds the live cell, right feeds the bench cell — so that pair is a config difference, not a live-vs-bench method delta. Per-cell n and per-run samples: CSV.
‖ = re-measured 2026-08-11 after our own audit found these two cells were short bursts (82 and 21 tokens) published as full 128-token runs. Corrected: gpt-oss-120b 59.1 →54.6 with prefill 469 →508 (we were 8% too fast) and Flow-Judge 36.2 →50.2 with prefill 751 →1,502 (we were 39% too slow — that run was truncated and cold-started). Independent check on gpt-oss: a different quant of the same model on the same box reads 54.2, agreeing to 0.7%. ¶ = arch refusal on that box’s llama.cpp build, published rather than hidden. Flags: † = July one-shot (n=1), current-build re-runs queued. ‡ = adversarially re-swept with per-run samples in the CSV after an outside review challenged the zero-width range — stability real, zero width was rounding. ᶜ = community quant, not vendor-official (quantizer named in the CSV) — official checkpoints exist for Laguna (poolside) and Qwen3.6 NVFP4 (NVIDIA ModelOpt); on Muse-Glimmer-30B it also marks a quant deviation, a ~4.5 bpw k-quant standing in for a Q4_K_M that does not exist, so those rows are not same-quant comparable with the Q4_K_M ladder. poolside re-released the Laguna NVFP4 repo in Aug 2026; our revision is pinned in the CSV. The 250 W power-cap experiment lives in the postmortem — buyer data here, ops archaeology there.
Two rows carry a ceiling worth knowing before you buy on speed alone: Phi-4 and StarCoder2-15B are native 16,384-context models — they benched fine at our 32K protocol but warned, and they can never serve long context no matter which card you put them on. Speed is not the only axis. KV/tok is engine-reported at boot, not a formula — parameter count does not predict memory cost, and the spread in that column is 18× (7.0 to 128 KiB/token; the full KV ladder below spans 120×). Gaps there are queued: the current test window records it for every model booted. Blackwell “short” rows were measured with all three engines co-resident on one node — the as-deployed floor, not the card's isolated ceiling; single-tenant re-runs land with the three-card matrix. Value stat, card MSRP only: Sarvam 30B decodes 198.4 tok/s on the $1,299 R9700 — 152.8 tok/s per $1,000 of card. Words/s ≈ tok/s × 0.75, stated conversion.
// long context — what depth costs
Model · hardware
@32K prefill / decode
@100K prefill / decode
@200K prefill / decode
poolside Laguna S 2.1 · RTX PRO 6000
12,775 / 69.1
8,892 / 62.0
5,926 / —*
Qwen3-Coder-Next · Strix Halo
616 / 48.7
426 / 37.2
255 / 27.1
Kimi-Linear-48B · R9700 32GB ($1,299 card)
2,039 / 131.6
2,095 / 131.6 (@128K)
—
Qwen3.5-122B-A10B · RTX PRO 6000
—
—
—
Cells are prefill t/s / decode tok/s at that actual prompt depth (measured 27–28K / 85–87K / 170–175K against the nominal buckets; n=3 per cell except the Strix 200K point, n=1 at 688 s per rep — one honest sample, disclosed). * = the 2-call method can't resolve vLLM decode at 200K; empty until a method can. Corrected 2026-08-12: the Qwen3.5-122B row now carries no cells. The prefill and decode figures printed there were a July single-shot vLLM run at short prompt, not a depth sweep on this protocol — the row stays listed, and empty, until one is run. Decode slows as the KV grows; prefill slows as the prompt grows — pick the box by your bottleneck. The 128K fit test: a 48B linear-attention model at full 131,072 context on the $1,299 32 GB card — full GPU offload, f16 KV, zero tweaks, 2.66 GiB headroom left. Going 32K → 128K cost 852 MiB and nothing in speed (131.6 both). We then pushed 38,400 real tokens of prompt through it at 1,463 tok/s — allocatable and usable are different claims, and this is the second one.
// the depth curves — does it slow down at 100K? measured, not vibes
The question every long-context spec sheet dodges: what speed do you have left once the context is actually full? We prefilled real tokens to each depth (tokenizer-calibrated per model, actual prompt lengths 26 / 32,671 / 65,440 / 99,001) and timed full 128-token decodes at every point — every generation length-verified, no truncated runs. Four architectures on the 128 GB unified-memory mini-PC, launched at 131,072 context, f16 KV. Nobody publishes this column. Every model slows. The only question is the slope — and parameter count does not set it: the 48B here decays faster than both the 80B and the 120B.
Model · Strix Halo 128GB
Attention
KV @ 131K
decode @ 0
@ 32K
@ 64K
@ 100K
retained @ 100K
Qwen3-Next-80B-A3B (thinking)
hybrid linear
3.46 GiB
59.0
45.2
39.6
34.5
58.6%
gpt-oss-120b
sliding window
4.49 GiB
53.4
34.4
32.1
26.8
50.3%
Kimi-Linear-48B-A3B
linear
1.40 GiB
69.0
47.7
38.7
32.1
46.4%
Hunyuan-A13B-80B
full attention
16.54 GiB
26.7
15.2
12.9
cut††
48.2% @ 64K
Same runs — the prefill wall
prefill @ 32K
@ 64K
@ 100K
time to first token at 100K depth
Qwen3-Next-80B-A3B
604.5
502.5
405.7
4m 04s
Kimi-Linear-48B-A3B
480.0
328.1
246.2
6m 42s
gpt-oss-120b
364.4
257.4
177.4
9m 18s
Hunyuan-A13B-80B
163.7
56.5
—
19m 18s just to reach 64K
What the curves say: (1) A fat KV cache is a hard tax at the extreme — but cache size alone does not rank the slope — Hunyuan carries 12× Kimi’s KV and is the only model to lose half its speed by 64K, and on unified memory that cache competes with the weights for the same bandwidth. Below that extreme the ranking inverts: Kimi has the smallest cache of the four (1.40 GiB) and the worst retention (46.4% @100K), while Qwen3-Next carries 2.5× the cache and retains best (58.6%). Parameter count predicts nothing either — 48B, 80B and 120B interleave in both columns. (2) Most of the damage lands in the first 32K (−23% to −43%) — which is exactly why capping context at 32K does not buy the speed back. By 32K you have already paid 56–83% of the total decode loss, so dropping from 100K back to 32K returns only 17–44% of what depth cost you (Qwen3-Next 44%, Kimi 42%, gpt-oss 29%, Hunyuan 17% measured to 64K). Budget for the deep number, not the shallow one. (3) Prefill is the real wall, not decode — decode at 100K is merely ~2× slower, but getting there costs minutes: 9m 18s of silence before gpt-oss-120b’s first token. For interactive use the binding constraint is time-to-first-token. (4) Sliding-window buys a plateau, not immunity — gpt-oss drops hard to 32K, then nearly flattens for the next 32K. †† = the Hunyuan 100K cell was cut on purpose and disclosed: its own curve projected ~35–40 minutes of prefill for that single number — the missing cell is the finding. Depth-0 speeds at 131K allocation land within −2.2% to +3.7% of our published 32K-protocol numbers for the same models on the same box (Hunyuan 26.7 vs 26.7, Kimi 69.0 vs 70.4, gpt-oss 53.4 vs 54.6, Qwen3-Next 59.0 vs 56.9); only Hunyuan reproduces to ±0.3%. The widest gap, Qwen3-Next at +3.7%, sits on top of a +3.4% spread between our own two runs of that model on that box, and the errors are signed both ways rather than systematically slower — so allocating a big context costs essentially nothing until you fill it. Blackwell cross-check, same night — both rows ⚑, measured on a transient build rather than the pinned build the Strix curves above ran on, so read them against each other and not against the rows above: Nemotron-3.5-Lightning-30B-A3B⚑ at 131,072 context costs ~0.9 GiB of KV, decodes 321 tok/s shallow and 254 tok/s with 100K tokens in front of it (−21%), prefilling those 100K in 10 seconds — the big card moves the wall, it doesn’t remove it. Meta’s Muse-Glimmer-30B⚑ at 128K on the same card — on a ~4.5 bpw community k-quantᶜ, not a Q4_K_M — holds −12% decode at 100K depth with 2,831 tok/s deep prefill.
// kv cache — the memory bill nobody prints
Every token of context costs memory beyond the weights — the KV cache. Spec sheets never print it, and it decides whether long context is nearly free or costs you a second GPU. Engine-reported values, converted across context depths. The spread below is 120× — and it tracks architecture, not model size.
Model
KV dtype
KiB / token
@8K GiB
@32K GiB
@128K GiB
@model max
Nemotron-Cascade-2-30B (6 attn / 52L)
q8_0
3.19
0.02
0.10
0.39
0.80 @262K
Nemotron-3.5-Lightning-30B (hybrid mamba)
f16
7.0*
0.05
0.22
0.87
— (1M declared)
Kimi-Linear-48B (linear attn)
f16
7.9
0.06
0.25
0.98
—
Qwen3.6-35B-A3B (hybrid)
fp8
10.35
0.08
0.32
1.26
2.59 @262K
Cohere North Mini Code 1.0
fp8
16.68
0.13
0.52
2.08
2.08 @131K
Qwen3-Coder-Next 80B (12 attn / 48L)
f16
24.0
0.18
0.75
3.0
6.0 @262K
gpt-oss-120b (iSWA)
f16
36.0*
0.28
1.13
4.5
4.53 @131K
Gemma 4 31B (10 global + 50 SWA)
q8_0
42.5*
0.33
1.33
5.3
11.05 @262K
Cohere Command A Reasoning 111B (iSWA)
f16
~91*
0.71
2.84
—
—
IBM Granite 4.1 30B (dense, 64L)
fp8
128.0
1.0
4.0
16.0
8.0 @64K served
Hunyuan-A13B-80B
f16
128.0
1.0
4.0
16.5
—
Atla Selene 1 Mini 8B
f16
128.0
1.0
—
—
2.0 @16K
KAT-Dev / Athene 72B (dense)
f16
320.0
2.5
10.0
40.0
—
Flow-Judge 3.8B (dense, fat heads)
f16
384.0
3.0
12.0
—
3.0 @8K
* = SWA/iSWA models: the scaling part only — their sliding-window layers add a small fixed pool, so long-context totals run slightly above the pure product. The 3.8B judge model costs 120× more memory per token than the 30B at the top: attention-layer count × KV-head count × head-dim × dtype, not parameter count. Measured sibling: at f16/32K the hybrid 100B-class models hold 0.63–1.27 GiB of KV where the dense-72B pair holds 10.0 GiB — 13× on the same box, same context. In English: if you want long documents in a home rig, pick a model with hybrid, linear, or sliding-window attention — classic dense makes context a second GPU.
// config experiments — same model, different config
Model
Variant
Result
Baseline
The lesson
poolside Laguna S 2.1
speculative draft OFF
100.7 tok/s
294.7 w/ draft
the draft model is +192% — but KV goes 24.0 → 38.2 KiB/tok, so not free
Qwen3.5-122B-A10B
--enforce-eager
23.1 tok/s
96.2 configured
the most-copied “OOM fix” costs 76%
GLM-4.7-Flash
MLA disabled (fallback)
~470 KiB/tok KV
177.2 tok/s configured
one flag = a ~470 KiB/token memory bill
Published on purpose: every number here is a config people actually copy from forums. Field reports had the 122B class at 23–33 tok/s — the gap to ours was these traps, not the hardware.
// pick by workload — what our own numbers say
Chat & code at 8–32K on a quiet box, card budget ≤$1.5K → RDNA4 R9700 + a 30B MoE — 127–198 tok/s measured above.
Long-context RAG or agent runs at 100K+ → RTX Blackwell 96GB + vLLM — 5.9–12.8K t/s prefill measured above.
Biggest models on one low-power box → Strix Halo unified memory — a 120B-class MoE at ~52 tok/s in the live sweep. MoE country — the memory table above shows why.
Every arrow above cites a measurement on this page, not a spec sheet. If your workload isn't here, ask for it — requests section below.
// test queue — ranked by what people actually ask for
Under-50B class, three cards, same GGUF: fill every cell of the matrix — more American, more Indian, more origins; the NAS census holds the queue. [WAVE-4 DISPATCHED]
120B MoE battles: Blackwell 96 GB GDDR7 vs Strix Halo 128 GB unified — same models both ways, GGUF and NVFP4 on Blackwell. [IN PROGRESS]
Quant ladder with speed AND quality on one rig (Q4 vs Q5 vs Q8: tok/s + KLD) — the question every forum thread asks and no tool answers. [PLANNED]
Model load / swap time tables per size, storage, and runtime — chronic pain, near-zero published data. [PLANNED]
Long-context curves: TTFT + prefill t/s at 32K / 100K / 200K, cold vs warm. [FIRST DATA LIVE — table above]
Anchor rows: Llama-3.1-8B Q4_K_M on every box — the gut-check number everyone already knows, so our tables calibrate against yours. [PLANNED]
Raw vs harnessed: how these models perform inside open-source agent harnesses — tokens burned, wall-clock, task success, answer drift vs the bare endpoint, from a shop that sells no gateway. [NEXT — the next capture direction]
Co-residency economics: three 30Bs sharing the 96 GB card at once vs three boxes — throughput per model, per watt, per dollar. The as-deployed reality nobody benches. [PLANNED]
32 GB vs 32 GB: an NVIDIA 32 GB-class card beside the R9700 — same models, same GGUF, the $1,299 question answered side by side. [PLANNED — hardware not yet acquired]
YaRN at length: quality-vs-context ladders for extended-context configs — admitted-unsupplied even by the people who publish quant benchmarks. [PLANNED]
Tokens per joule + idle draw per box; unedited screen captures of real-time generation. [PLANNED]
Multi-model co-residency: N engines sharing one card vs single-tenant — the floors you actually get in deployment, both big boxes. [PLANNED]
// how we test — no secrets, run it yourself
llama-bench rows:llama-bench -m <model.gguf> -p 512 -p 8192 -n 128 -r 5 -o json — prompt-processing at 512 and 8,192 synthetic tokens, then 128 generated tokens, five repetitions, JSON out. Stock binary, stock flags (full GPU offload, mmap on). No prompt tricks — llama-bench feeds synthetic tokens, so there's nothing to cherry-pick.
Live-sweep rows: a real request, direct to the engine (no gateway): a randomly-salted prompt (“<fresh-hex>: Write a plain-English 200-word explanation of how a car alternator works.”), max_tokens=256, n=6, first run discarded as warmup, 1s spacing. tok/s = completion tokens ÷ wall-clock, so it includes prefill and one LAN hop — the number you'd actually feel, not the flattering one.
Every response's model identity is verified against the engine's own report — because gateways lie when a fallback kicks in, and we caught ours doing it.
Raw data: st4ts_data.csv ships alongside this page — mean/min/max and stddev per row, per-run samples where captured (the challenged rows carry them; all new sweeps will), and the llama.cpp build commit where recorded (current rows: bf2c86d). Gaps say so rather than guess. If your numbers disagree with ours, we want to see them.
// requests — tell us what to measure
The bench: NVIDIA RTX Blackwell nodes · AMD Strix Halo nodes · AMD RDNA4 R9700s · llama.cpp and vLLM.
Got a model, quant, context length, or head-to-head you want real numbers on? Send it: reed@directive4.ai — the benchmarks desk. Real-world requests get priority over synthetic ones.
We test models from everywhere — US, Europe, China — on the same hardware, same protocol. Balance is the point.
// updates — dated, append-only
2026-08-12 — the audit: a reader called one cell, so we re-checked every cell. The founder read the 70B+ table and called the R9700 column unbelievable. He was right to. The number was real — a receipted spill probe — but a table reorganization had dropped its context and left its flag pointing at the wrong legend. The full-page audit that followed checked every published number, flag, and claim against the raw-data receipts and confirmed 84 defects: 25 values unreceipted or contradicting their receipt (corrected to the receipted value or removed), 59 claims and flags missing load-bearing context (now carrying it). The big ones, printed rather than buried: the “leads all three cards” line for LFM2.5-8B-A1B was wrong — it leads one column, its 1.2B sibling leads another; a wave-3 “record” claim was wrong; seven perf/watt rows carried assumed wattage and are gone until measured; five big-table cells sat in the wrong hardware column; one long-context row had no probe behind it and is gone; spill numbers moved to their own section with their method disclosed. Corrections ran both ways again — Laguna’s draft-assisted decode was actually faster than we printed (283.1 → 294.7). Every glyph on the page now has exactly one meaning. 73 minor variances are queued. The raw CSV carried the true receipts all along — the page now matches it. Struck text in older entries below points here. Named, so you can check us: the “leads all three cards” claim for LFM2.5-8B-A1B — it leads one column, its 1.2B sibling leads another; a wave-3 “record” that wasn’t one; seven perf/watt rows whose wattage was assumed from the card’s cap rather than measured — removed until sampled; five 70B+ cells holding a mini-PC number in the workstation-card column; a long-context row with no probe behind it; a KV-cache row for a model with no receipt anywhere in the raw data; a spill-test number tabled beside protocol rows it was never comparable to — now its own section with its method stated; a July row that merged two engines and two quants into one line; and a “per watt the mini-PC ties or wins” claim we withdrew rather than replace with another unmeasured one. One correction ran in our favour: Laguna’s draft-assisted decode was faster than we printed (283.1 → 294.7).
2026-08-12 — wave-5: the depth curves land — the first measured answer to “does it slow down at 100K context?” on this class of hardware (yes: models keep 46–59% of their speed, the KV cache sets the slope, and prefill — minutes of it — is the real wall). 39 new rows, ~20 new models including the Mistral 24B trio, more US rows (Olmo-3, Granite-H, Reka, Phi-4-Reasoning, Nemotron-Nano-9B), Singapore’s SEA-LION, and Meta’s brand-new Muse-Glimmer-30B. Records: LFM2.5-8B-A1B takes the top of the matrix on all three cards at once and the perf/watt crown on each — it tops the Strix Halo column and takes that box’s perf/watt record outright; its 1.2B sibling leads the other two speed columns; the mini-PC sustained record jumps 36% to 162.4. Head-to-heads with verdicts: Nemotron-Super v1 vs v1.5 is a dead heat (the .5 is a quality release — throughput can’t see it), and the Qwen3.6 “MTP” GGUF is 529 MB of dead weight on llama.cpp — the engine discards those tensors at load, so the delta is zero by construction. Discovery: two “different” releases (Ornith-1.0-35B, US and Agents-A1-35B, CN) are the same weights under two names — GGUF files 128 bytes apart, identical headers. New ⚑ flag marks rows measured on a transient llama.cpp build (several new architectures refuse the pinned build); the wave-4 Mamba-Codestral refusal is closed the same way at 37.3. Two benched models are held off the page pending license review — four are — read first, publish second. — corrected 2026-08-12, see audit entry
2026-08-11 — the two withdrawn numbers came back, and the errors ran both ways: gpt-oss-120b was 8% too fast (59.1 → 54.6), Flow-Judge was 39% too slow (36.2 → 50.2, prefill 751 → 1,502 — that one was truncated and cold). A defect that only ever flattered us would be a thumb on the scale; this one cost us a number and gave one back. Strix Halo column completed for the wave-4 models — LFM2-24B sets a new record for that box at 119.2 and takes the efficiency crown on both cards (1.54 tok/s per watt, 30× the 72B dense).
2026-08-11 — wave-4: nine more models across two cards, and a new Built on column. A country flag says who published a model, not whose weights it started from — so Holo now reads “France, built on Qwen3.6 (CN)” and Selene reads “UK, built on Llama-3.1 (US),” from GGUF headers and publisher cards, 76 models classified. New leader: Liquid AI’s LFM2-24B at 380.5 tok/s. First pure-Mamba row (Codestral 7B). The 70B+ table gains an R9700 column showing what happens when a model is twice the size of the card: it loads by spilling into system RAM and crawls — 1.8 tok/s for a 72B dense. Published because “it loaded” is the trap.
2026-08-11 — correction, self-caught: an internal audit of every published row found four defects of one kind — a generation that stopped early being reported as a full-length run. Two were on this page as bench numbers and are now withdrawn pending re-measurement (gpt-oss-120b UD-Q4_K_XL, previously 59.1; Flow-Judge 3.8B, previously 36.2 — 82 and 21 actual tokens). Two were correctable from the logs and are now right: Falcon-H1 34B on Strix Halo 10.3 → 9.9, Nemotron-3-Nano 66.0 → 65.9. Old values stay printed here on purpose. The audit also cleared 45 other rows as verified-clean.
2026-08-11 — the 128K fit test on the cheap card: Kimi-Linear-48B runs full 131,072 context on the 32 GB R9700 with 2.66 GiB to spare and no speed penalty — 131.6 tok/s at both 32K and 128K. Its matrix cell fills at 131.6 (1.86× the mini-PC on the same file).
2026-08-11 — restructure by weight class: the matrix is now the 30B class (what an ordinary card holds), the 70B+ class gets its own two-box table, and both gained Params / Active / Type columns. Performance-per-watt adds the R9700 (hwmon-measured) and a 70B+ Blackwell table — the 10× MoE-vs-dense efficiency spread inside one weight class is the finding.
2026-08-11 — wave-3: 19 new rows in one day (11 Blackwell + 8 Strix Halo, zero failures, boxes restored and verified). New records: Sarvam 30B at 364.6 tok/s on the workstation card and 83.7 on the mini-PC — fastest GGUF we've ever measured on both Sarvam 30B at 364.6 tok/s on the workstation card, the fastest GGUF we've ever measured on that box; on the mini-PC the wave's record went to Qwen3-Coder-30B-A3B at 91.6, with Sarvam second at 83.7. Matrix gains an origin column; new performance-per-watt table (GPU-reported watts). Pattern finds: the two 80B "twins" benched identical to the decimal — architecture sets the row, not the finetune; the 30B dense class slams the 300 W cap at a fifth of MoE speed on every box. — corrected 2026-08-12, see audit entry
2026-08-11 — same-quant 120B row closed: the identical MXFP4 file on the workstation card and the mini-PC (225.9 vs 54.2 — 4.2× the speed at ~6× the power, so per-watt the mini-PC ties or wins roughly 3.4× the power on the nearest measured comparison; the mini-PC leg measures 0.83 tok/s/W on its sibling quant and the workstation card’s leg of this pair was never power-sampled, so no per-watt winner is claimed). Six more Strix Halo large-class rows, the linear-attention 48B leg, and a new KV-cache cost table: 14 models, a 120× per-token memory spread. — corrected 2026-08-12, see audit entry
2026-08-11 — tables unified into one master spreadsheet after founder review: one schema, origin on every row, per-class filters. The old layout scattered the same metric across four table shapes; this one shows its gaps instead of hiding them behind table boundaries.
2026-08-11 — first public cut: live fleet sweep (14 engines), R9700 30B ladder (8 models), long-context probes at 32K/100K/200K. Strix Halo + Blackwell ladder legs in progress. One correction already logged in the raw CSV: a mid-run cleanup truncated one benchmark's output; the run was discarded and re-run clean, and the incident is written down instead of erased.
// directive4.ai · don't rent your intelligence · full stamps: vLLM/llama.cpp versions per row where recorded; gaps say so rather than guess · build #056