← Steppe Integrations / Articles

Local AI · measured field report

What Consumer Hardware Actually Does

Twenty-nine indexed local-model labs plus the August rebuild and Qwen 3.8 campaigns, across an RTX 4070 Ti, an RTX 5070, and a pair of Intel Arc Pro B70s. Every number below is measured, dated, and attributed to the lab that produced it.

Sixty-four gigabytes of VRAM for near or less than the price of one flagship card is a trade almost nobody writes up. Here is what it actually returns, where it stops, and why the binding constraint depends on which phase of the workload you are in.

The article in seven findings

01 / 07 · auto

Contents

What used to be one crowded speed chart is now a walk through the lab with its cast: what the machines can hold, what the answers are worth, and where the verdict flips. Every panel is one idea, one chart, and the table behind it.

The cast

Every number below was pulled by somebody

The Blue Oxentwo Intel Arc Pro B70s Bought at MSRP — the whole reason this lab exists. They have pulled on Windows 10, on Linux, and now headless in the rebuilt workstation. Strong, stubborn, and they do not talk to each other directly.
The First LlamaRTX 4070 Ti How a GPU became an engine here at all: first local models, first whisper transcripts, the origin story in section 1.
AM4the Ryzen box, 32 GB DDR4 First host to the Oxen — 64 GB of VRAM hitched to a 32 GB wagon. Overload that memory and you did not reboot: you pulled cards to the workbench, retrained RAM, reflashed BIOS, and lost the long weekend. More than once. That tuition paid for finding 1.
OMENZ890, 285K, 128 GB DDR5 The rebuilt house where the Oxen now live, each on its own direct gen5 lane. Windows 11, Vulkan, and every “rebuilt” number in this report.
The FiverRTX 5070, 12 GB The CUDA control group: what a current 12 GB NVIDIA card does against the same workloads.

Same machines as the lineage table — now with the names they earned.

Interlude · the workbench years

The trail so far, paid for in weekends

The First Llama era: an RTX 4070 Ti, Ollama, whisper — how a GPU became an engine here The Oxen arrive at MSRP: 64 GB of VRAM hitched to a 32 GB wagon, Windows 10 The workbench loop: overload the memory and you pulled cards, retrained RAM, reflashed BIOS — days at a time The Linux passage: Ubuntu and the Battlematrix stack — the Oxen speak directly, 14.29 GB/s card to card OMEN, the rebuilt house: 128 GB DDR5, each Ox on its own direct gen5 lane, Windows 11 and Vulkan the First Llama the Oxen arrive the workbench loop the Linux passage OMEN
Nobody bought this as a lab. In the wagon years, pushing two Oxen against 32 GB of DDR4 meant that an overloaded run didn’t end with a reboot — maybe, if you waited fifteen minutes, it came back; more likely you were pulling B70s out onto the workbench, seating the 4070 Ti back in, retraining memory, reflashing BIOS. Hours, days, a long weekend and change — more than once, through every rebuild, in search of the ever-greater local AI lab. On the Linux passage the Oxen learned to speak directly to each other — 14.29 GB/s, measured — and in the rebuilt house they carry more than ever, though every word between them now goes through the walls. That tuition is why every number in this deck exists.

The same five steps as the lineage table, walked instead of tabulated.

Chapter one · Capacity

Where the bytes live, and what moving them costs

Seven charts about weight, in both senses: what the Oxen can hold on their backs, what stays in the house’s 128 GB, and the price of every trip between the two — the same limits that once ruled the workbench, now measured instead of survived.

  • Generation speed, one era at a time · the placement ladder · the MoE on tap
  • Load physics · parkable context · four models at once · one honest negative
Capacity · 01 The OxenThe Fiver

Generation speed, one stop of the trail at a time

OMEN, the rebuilt house · the Oxen · Win 11 / Vulkan
Qwen3-30B-A3B95.2
Flash-Next 125Ball 48 blocks on GPU27.7
Qwen3.8-27B dense23.8
The Linux passage · the Oxen · SYCL
Qwen3-30B-A3Bone card86.9
gpt-oss-120bboth cards27.7
The overloaded-wagon years · the Oxen on AM4 · Win 10
Qwen3-30B-A3B81.7
Qwen2.5-32Bone card23.7
Llama-3.3-70Bboth cards11.3
The Fiver’s house · RTX 5070 · CUDA
Qwen2.5-14B59.9
Mixtral 8x22B+ 128 GB system RAM4.2
Tokens per second, single stream, shallow context. Each panel keeps its own scale and its own harness — eras are never mixed on one axis, because the August sweep measured the cross-harness gap at about 19% on a fast model. The 70B appears at all because of one flag; the Mixtral’s 4.2 tok/s is a 140B model carried by system RAM — useless for chat, perfectly adequate for an overnight reviewer.
Data table
Q4-class quantization except the 120B (MXFP4), Mixtral (Q2_K), Flash-Next (IQ4_XS). May–August 2026.
ModelRig / eratok/s
Qwen3-30B-A3BOMEN rebuilt dual B70, Win 1195.2
Qwen3-30B-A3Bone B70, Linux86.9
Qwen3-30B-A3BAM4-era dual B70, Win 1081.7
Qwen2.5-14BRTX 5070, CUDA59.9
gpt-oss-120btwo B70, Linux27.7
Qwen3.8-Flash-Next 125B48 of 48 blocks on GPU, OMEN27.7
Qwen3.8-27B denseOMEN rebuilt dual B7023.8
Qwen2.5-32Bone B70, Win 1023.7
Llama-3.3-70Btwo B70, Win 1011.3
Mixtral 8x22B Q2_KRTX 5070 + system RAM4.2

Full tables in each rig’s section; harness skew in what the measurement costs.

Capacity · 02 The Oxenin OMEN

An 88 GiB model climbs onto the Oxen, one rung at a time

0 10 20 30 tok/s 0 of 48 blocks: 4.75 tok/s, 6/8 valid, 6.5 GB commit and no VRAM 8 blocks: 5.99 tok/s, 6/8 valid 16 blocks: 7.40 tok/s, 6/8 valid 24 blocks: 9.12 tok/s, 6/8 valid 32 blocks: 11.62 tok/s, 6/8 valid 40 blocks: 16.69 tok/s, 6/8 valid 48 of 48 blocks: 27.70 tok/s, 6/8 valid, 61.2 GB VRAM and 65.3 GB commit 27.70 4.75 0 16 32 48 blocks on GPU
Same model, same cards, same tasks — the only variable is how many of its 48 blocks ride on the Oxen. Speed climbs 5.8×; correctness never moves (six of eight at every rung, failing the same two tasks the same way). The memory bill runs the wrong way: full placement holds 61.2 GB of VRAM and 65.3 GB of commit.

Same measurement as the placement-ladder table, 27 August 2026.

Capacity · 03 The Oxenthe house’s DDR5

The MoE you can keep on tap

0 10 20 30 tok/s Full placement: 27.7 tok/s at 60.4 GB commit and 61.2 GB VRAM full placement · 27.7 Experts in system RAM, attention and routing on the GPUs: 10.6 tok/s at +6.3 GB commit experts in RAM · 10.6 Same placement loaded beside the live production model, no maintenance window: 4.1 tok/s beside live production · 4.1 0 committed memory, GB 70
Leave the 512 experts file-backed in the house’s DDR5 and route only attention through the Oxen: 10.6 tok/s at +6.3 GB of commit — a tenth the decode-per-GB price of full placement, and cheap enough to sit loaded beside the production model (4.1 tok/s, no window, no pagefile). Shallow-prompt decode is what was measured; deep-prompt prefill in this placement is not yet, and is the registered next probe.
Data table
Qwen3.8-Flash-Next IQ4_XS, single stream, short prompts. Rotation-physics probes, 28 August 2026.
PlacementDecode tok/sCommitNote
All 48 blocks on GPU27.760.4 GBplus 61.2 GB VRAM
Experts in system RAM10.6+6.3 GBexperts stay file-backed; commit charge never materializes
Same, beside live production4.1+6.3 GBproduction model kept serving and passed its health probe after

Rotation-physics probes, 28 August 2026. Deep-prompt prefill in this placement is unmeasured — see what is not measured.

Capacity · 04 The Oxenin OMEN

Hitching up: the warm page cache loses to a cold read

Cold load, read into memory17–19 GB models, steady state8.2 s
Warm reload from page cachememory-mapped, weights already cached12.7–13.3 s
How long it takes to get a model onto the Oxen’s backs and answering — the two 17–19 GB production models, seven trials each. The intuition says a warm page cache should win; the stopwatch says the plain cold read hitches up ~55% faster. Bonus honesty: the “direct I/O” flag involved turns out to be inert at the Windows file layer — reads are buffered either way at about 2.3 GB/s. The flag still matters, but only because it switches the loader off memory-mapping; the speed belongs to sequential reads, not to the flag’s name.
Data table
NVMe source, dual-B70 placement, 28 August 2026. Rotation cost on this rig is therefore a cold load plus KV restore — no warm pool to manage.
Load pathTrialsLoad to healthy
Read-into-memory, cold (30B-A3B)38.2–8.3 s
Read-into-memory, cold (27B dense)38.2 s
Memory-mapped, warm page cache (both)612.7–13.3 s
First load after a driver init219–27 s

Rotation-physics probes, 28 August 2026 · the flag’s sign also reverses with model size — see the reversal box.

Capacity · 05 The Oxenin OMEN

Setting a load down without dropping it

Recompute the context29K tokens, cold prefill102.8 s
Park itsave the KV state, 2.68 GB1.74 s
Bring it backrestore after a full server restart1.19 s
The Oxen learned to set a load down without dropping it. A 29,320-token context is 2.68 GB of KV state. Saving it takes 1.74 s; restoring it — across a complete server stop and start — takes 1.19 s, after which the identical prompt re-evaluates exactly one token and the greedy output matches byte-for-byte. The linear bars are the message: parking context costs about 3% of recomputing it. Deep context stops being a tax you re-pay per swap and becomes a thing you set down and pick back up.
Data table
Qwen3-30B-A3B on one B70, 28 August 2026. Same lever as the 306× prefix-miss measurement, driven from the other side.
OperationTimeDetail
Cold prefill, 29,313 tokens102.8 s286.9 tok/s prefill, single card
Save KV state1.74 s2.68 GB written, ~1.66 GB/s
Restore after full restart1.19 sidentical prompt then re-evaluates 1 token
Cross-model restore attemptrefusedHTTP 400 — the guard fails loudly, not silently

Rotation-physics probes, 28 August 2026 · companion to what a cache miss costs.

Capacity · 06 The Oxenin OMEN

Four models, two Oxen, one house

0 25 50 75 100 tok/s phi-4 · card 1 phi-4 alone: 48.7 tok/s phi-4 with all four answering: 22.5 tok/s Qwen2.5-14B · card 1 Qwen2.5-14B alone: 45.6 tok/s Qwen2.5-14B with all four answering: 28.6 tok/s gpt-oss-20b · card 2 gpt-oss-20b alone: 98.3 tok/s gpt-oss-20b with all four answering: 52.9 tok/s Mistral-24B · card 2 Mistral-Small-24B alone: 29.8 tok/s Mistral-Small-24B with all four answering: 23.9 tok/s
aloneall four answering at once
Each Ox learned to pull for two riders at once — four different models resident, two per card, 26 GB of weights per Ox. Under simultaneous fire each card fair-shares between its tenants — no crash, no eviction, about 128 tok/s aggregate. And with just one model per card, a heterogeneous pair holds full solo speed under concurrent load: 99.2 and 21.6 tok/s together, 95–100% of alone.
Data table
tok/s, single stream each, 28 August 2026. Two-per-card heterogeneous residency; unload and reload cycling stayed clean throughout.
ModelCardAloneAll four at once
phi-4 (8.3 GB)148.722.5
Qwen2.5-14B (8.4 GB)145.628.6
gpt-oss-20b (11.3 GB)298.352.9
Mistral-Small-24B (13.3 GB)229.823.9

Rotation-physics probes, 28 August 2026.

Capacity · 07 The Oxenin OMEN

In this house, the Oxen cannot hear each other

27B densetensor split vs layer split−30%
30B MoEtensor split vs layer split−72%
On the Linux passage the Oxen spoke directly — 14.29 GB/s ear to ear, measured. In the rebuilt house they cannot: peer access probes report the cards cannot reach each other on this board — the root complex drops the peer traffic at the silicon — so every word between them goes through the walls. That is why tensor parallelism, both Oxen pulling the same token, loads cleanly and loses: the dense model gives back 30%, the MoE 72%, all of it crossing-time. An honest negative, kept on the board for re-testing as the code path matures.
Data table
Single stream, shallow context, 28 August 2026. Peer-to-peer access between the two B70s measures unavailable on this platform; cross-card traffic is host-staged.
ModelLayer splitTensor splitDelta
Qwen3.8-27B dense22.7 tok/s15.9 tok/s−30%
Qwen3-30B-A3B MoE~104 tok/s28.8 tok/s−72%

Rotation-physics probes and Level-Zero reconnaissance, 28 August 2026.

Chapter two · Capability

What the answers are worth

Capacity says whether a load can be carried. Capability says whether it was worth hauling. Three charts where the download size tells you nothing.

  • Verbatim recall at depth · blind judging against frontier cloud · the speedup that changed the answers
Capability · 01 The Oxenin OMEN

Both get the gist; the little dense one remembers exactly

27B dense8K-token documents8 of 8
27B dense32K-token documents5 of 8
30B MoE8K-token documents2 of 8
30B MoE32K-token documents1 of 8
Eight probes per document, spread across its depth: find the line containing an exact phrase and quote it verbatim, at temperature zero. Both models locate the content almost every time — fragment recall is 6–8 of 8 everywhere. The difference is fidelity: the MoE strips list prefixes and reformats what it quotes; the dense model reproduces the line exactly. 5 of 8 at 32K against the MoE’s 1 of 8. If the job is quoting exact values back out of long documents, parameter-efficiency is not the metric that decides.
Data table
Organic engineering documents, probes at depth fractions 0.01–1.0, temperature 0, production dual-card shapes, 28 August 2026. One task family; eight probes per cell — a probe set, not a benchmark.
ModelDepthVerbatim lineFragment located
Qwen3.8-27B dense8K8/88/8
Qwen3.8-27B dense32K5/86/8
Qwen3-30B-A3B MoE8K2/88/8
Qwen3-30B-A3B MoE32K1/87/8

Depth-recall probes, 28 August 2026 · companion to the operating-point inversion.

Capability · 02 The Oxen, judged blind

Blind judging: the dense candidate vs everyone

vs the incumbent 30B MoE · 90 blind pairs
44 W42 T4
vs frontier cloud (Gemini 3.1 Pro reference) · 91 blind pairs
469 T18 L
wintieloss
Blind pairwise judging of full responses. Against the incumbent MoE the dense 27B wins or ties 95.6% of pairs. Against a pinned frontier reference it mostly ties — 69 of 91 — and loses 18: a fair sketch of where a consumer-hardware model actually sits in 2026, neither embarrassed nor victorious.

Qwen 3.8 promotion campaign, blind judgment files, 27 August 2026.

Capability · 03 The Oxenin OMEN

The 3× speedup that changed the answers

Speculative decoding off510 jobs/h
Speculative decoding ondraft acceptance 0.65 → 0.861591 jobs/h
The single largest configuration win in the campaign — and the one that is not deployed. At temperature zero the two paths should produce identical text. Across every compared cell there were zero identical responses, and all 126 long-prompt requests on the non-speculative path stopped early where the speculative path ran to length. A throughput chart that is really a correctness question. The campaign’s own gate had compared validity rates — 0.875 vs 0.875 — and never an output to an output.

Same measurement as speculative decoding tripled throughput, 27 August 2026.

Chapter three · Operating points

The verdict depends on where you measure

Every “which model is faster” answer in this report is true at one stop of the trail and false at another. Two charts about the flip.

  • The depth inversion · what a cache miss costs
Operating points · 01 The Oxenin OMEN

The inversion: slower at 512, five times faster at 32K

parity 512-token prompts: the dense model delivers 0.69 times the MoE's jobs per hour — the promotion gate's operating point 8K prompts: 2.63 times 32K prompts: 5.49 times the MoE's jobs per hour 0.69× 2.63× 5.49× 512 8K 32K prompt tokens
Dense-27B jobs per hour as a multiple of the MoE’s, by prompt depth. The promotion verdict was decided at 512 tokens, where the dense model loses. At the depths where long-document work actually happens, the same model is 2.6× to 5.5× the incumbent. A verdict is a property of an operating point, not of a model.

Same corpus rows as the operating-point table, 27 August 2026.

Operating points · 02 The Oxenin OMEN

What a cache miss costs, three requests in a row

0 30 60 s Unique prompt, request 1: 59.8 s to first token Unique prompt, request 2: 59.6 s Unique prompt, request 3: 55.4 s Shared prefix, request 1: 39.75 s — the one cold prefill Shared prefix, request 2: 0.13 s Shared prefix, request 3: 0.12 s unique every time shared prefix · 306× cheaper request 1 request 2 request 3
shared prefixunique prompts
Same model, same hardware, same 26K depth — the only difference is whether the prompt shares a prefix with the one before it. A prefix-cache miss costs about 306× the warm case. Between this chart and the parking chart, the theme of the whole report: on consumer hardware, context is the expensive thing — and it is entirely manageable.

Same accidental experiment as measured by accident, 27 August 2026.

The cast · 01 / 17

How the rig got here

None of this was bought as a lab. It accreted, one constraint at a time, and the sequence explains most of the numbers below — including the one that keeps breaking.

The hardware lineage. Each step was a purchase decision, not a plan.
StepThe machineWhat it unlocked
1A Ryzen daily driver on Windows 10 with an RTX 4070 Ti, 32 GB DDR4First local models, on Ollama. The planner-and-critic experiments in section 1 ran here.
2Same box, two Intel Arc Pro B70s go inBought at MSRP, which is the entire reason this configuration exists. 64 GB of VRAM on a machine with 32 GB of system RAM — the inversion behind finding 1.
3An open-box workstation, 128 GB DDR5 + RTX 5070, becomes the daily driverFrees the first box from having to be anyone's desktop. Also the rig for every 12 GB NVIDIA measurement in section 1.
4The B70 box converts to native LinuxOnly possible once it wasn't the daily driver. Unlocked Intel's Battlematrix multi-GPU stack — and everything in section 3.
5Completed 20 August: both B70s move into a rebuilt 128 GB DDR5 workstationASUS Pro WS Z890-ACE SE, Core Ultra 9 285K, Windows 11 and Vulkan. The RTX 5070 leaves; the Intel iGPU and ASPEED adapter drive the displays, so both B70s are headless.

Step 2 created the original experiment: two large GPUs bolted onto a machine whose system memory was specified for one. Step 5 did more than fix that mismatch. It held the cards constant while replacing the board, CPU, RAM, operating system, driver and display path. The result turned the planned upgrade into a controlled challenge to the article's own conclusions.

Three eras, the same cards. The B70s ran first on Windows 10 + Vulkan, then on native Ubuntu + SYCL / Level Zero, then in the rebuilt workstation on Windows 11 + Vulkan. Harnesses still define the comparison boundary: old and new numbers are comparable only where the campaign deliberately repeated the workload.

Seven findings, corrected by the rebuild

1. The binding constraint changes by phase

The original Windows box had 64 GB of VRAM and 32 GB of DDR4, and host memory broke repeatedly. A stack sweep found swap saturated with one 30B resident. A judging run OOM-killed until context came down from 32k to 8k. Three large models enabled together produced 81 crash-loop restarts and about 18 GiB of driver allocations recoverable only by rebooting. That dated conclusion remains true for that machine.

The 128 GB rebuild made the broader claim false. The eight-request knee remained with 78 GB of commit headroom; that wall was a Vulkan dispatch choice. Context-depth decay remained because KV lives in VRAM. And RAM still governed placement at load: a 70B loaded at 92% commit ran prompt processing at 11.8 instead of at least 150 tok/s, and freeing the co-tenant did not repair the process. Commit and placement bind at load, kernel dispatch binds under concurrency, and KV plus attention bind at depth.

2. MoE is the architecture that fits consumer hardware

At equal size — roughly 32B total parameters, Q4, one card — a 3B-active mixture-of-experts decoded 5.7× faster than a dense 32B (130 vs 23 tok/s) and prefilled 2.9× faster. Decode streams only the active parameters, so sparsity buys back exactly the memory bandwidth consumer cards don't have. Every practical result in this whole corpus rests on that.

3. Backend and kernel choice swing results more than the GPU does

Same model, same card, same prompt: CUDA processed the prompt at 2276 tok/s and Vulkan at 68.5. A 33× spread on identical silicon. The old Vulkan stack also made flash attention about 4× slower and q8 KV 2–4× slower. The rebuilt campaign narrowed both rules: flash-attention-off wins by 25–87% below a 16k–24k crossover, then hangs; q8 KV costs only 5–6% at 32k and 64k on the newer build, though its 128k hang remains. Configuration findings belong to a measured stack, not to the logo on the card.

The most extreme case is a single flag. Running a 70B across both B70s, turning memory-mapping off took prompt processing from 21.5 to 151.4 tok/s — a 7× win, and the difference between "unusable" and "usable". Nothing about the hardware changed. If you are benchmarking a local stack and the number looks absurd, suspect the configuration before the card.

4. Context depth is the real cliff, not model size

A 30B MoE that fits comfortably on one B70 decodes at 86.9 tok/s empty and 5.3 tok/s at 64k context. On the Windows/Vulkan side the same shape showed 132.8 tok/s empty falling to 11.5 at 64k — roughly linear to 32k, then a superlinear cliff. Meanwhile prompt ingest dominates the wall clock: a 57,000-token request spent 207 of its 213 seconds just reading the prompt.

Which means the thing worth engineering is not placement. It's caching. An identical 8k prompt sent twice went from 11.8 s to 0.17 s — 68× — on cache hit. And restoring a saved KV cache beat re-computing it by 29× on real bytes, with the advantage growing with context: 22× at 1k, 75× at 14k.

5. Local output held up against frontier cloud

On a scored documentation and decision-record task, the local 120B model scored 93.8 against Gemini 3.1 Pro's 90.6 and 3.5 Flash's 89.2. The cloud's measurable advantage was latency — about 4–5× faster — and that was the entire advantage. One task family, fifteen cells — not a claim that local wins everywhere. The narrower claim: "local is the cheap option you settle for" did not survive contact with a rubric.

6. And you can train on it, not just serve

The one most likely to be assumed impossible. A LoRA fine-tune of a 7B coder model, running on one Arc Pro B70 on Windows via PyTorch's XPU backend, took held-out loss from 2.96 to 1.13 in 108 seconds of training — then merged, quantized, and served on the same card in about six seconds, generating correct code. Full detail in section 6. Local inference is well-trodden ground by now. Local weight-changing, on a non-NVIDIA consumer card, is not.

7. A benchmark's operating point can decide the verdict

The newest correction, and the one aimed at this article's own method. A dense 27B failed a promotion gate at 0.69× the incumbent's throughput and beat the same incumbent by 5.49× on the same box the same day — the only difference was prompt length. The gate sampled 512-token prompts, which is an unstated claim about the workload. Where decode dominates, a sparse model with 3B active wins; where prefill dominates, that reverses. Detail in section 5.

1 · NVIDIA consumer — RTX 5070 and 4070 Ti, 12 GB each

Rig: HP OMEN 45L · Intel Core Ultra 9 285K (24 threads) · 128 GB DDR5 · RTX 5070 12 GB GDDR7 · Intel iGPU · Intel AI Boost NPU · Windows. The 5070 labs: 2 July 2026. A separate round on an RTX 4070 Ti 12 GB GDDR6X ran from June.

Backend lab — CUDA vs Vulkan vs the integrated GPU

The same model and prompt driven through four isolated servers on separate ports, to find out what the backend choice alone is worth.

Qwen2.5-14B, 4096 context, 957 prompt tokens, 120 output tokens. Measured 2 July 2026.
BackendDevice pathLoadPrompt evalOutput
CUDARTX 50703.11 s2276.17 tok/s59.88 tok/s
VulkanRTX 507011.11 s68.51 tok/s42.26 tok/s
Vulkan, mixedRTX + Intel visible9.13 s1921.88 tok/s45.32 tok/s
VulkanIntel integrated GPU12.35 s47.90 tok/s3.70 tok/s

Source: Ollama backend lab — results summary, raw CSV, per-profile JSON and runner logs. 2 July 2026.

The 52 GB sparse model, off system RAM

Mixtral 8x22B Instruct Q2_K (140.6B total params, 52 GB) on RTX 5070 + 128 GB DDR5. Measured 2 July 2026.
ContextPrompt tokensPrompt evalOutputResident
8,1926,875224.79 tok/s4.12 tok/s~53–56 GB
16,38414,425243.91 tok/s3.77 tok/s56 GB
32,76828,770205.41 tok/s2.27 tok/s60 GB

Cold load 39.6 s, warm load 0.03 s — residency is the whole cost. The machine stayed responsive throughout and never obviously paged.

Source: local MoE findings note. 2 July 2026.

The capability matrix and the baseline protocol

Two more documents from the same day turn those numbers into decisions rather than trivia. A capability matrix assigns each model class an operational role — resident controller, interactive specialist, batch specialist, fallback worker, overnight critic — and tags every cell measured, published, or inferred, so you can see which claims came off this rig and which came off a model card. Its operating conclusion: a 12 GB GPU strongly favours a control-plane and specialist split, and the constraint is residency, not active parameters. A 30B MoE behaves like a small model per token, but its total weights still drive placement.

The other is a dated baseline snapshot plus a written rerun protocol, built deliberately so a future Arc Pro B70 rerun would be comparable. Warm load for both test models collapses to about 0.10 s. Worth noting for anyone building on Ollama: same-service "concurrency" there behaved as queueing against one daemon, not true batched serving.

Sources: hardware capability matrix; RTX 5070 baseline snapshot and baseline protocol. 2 July 2026.

What the raw logs said that the summaries didn't

Going back through the raw per-call CSVs — which log GPU utilization, memory and power draw before and after every request — turned up two things no summary of mine had stated.

Per-call telemetry, CUDA on the RTX 5070. Measured 2 July 2026.
ScenarioModelLoadPrompt tok/sOutput tok/sGPU power
baseline, coldqwen2.5:14b4.52 s1015.6361.828 → 203.4 W
warm, back-to-backqwen2.5:14b0.13 s7154.0462.0030 → 203.8 W
baseline, coldqwen3-coder:30b13.10 s294.3757.4022 → 66.7 W
long prompt @ 8kqwen3-coder:30b10.27 s2154.1250.7921 → 82.7 W

Source: raw capability-bench and cross-service CSVs, RTX 5070. 2 July 2026.

The model matrix — on the 4070 Ti

Worth separating, because it's a different question and a different card. Where the 5070 work is two models deep and many scenarios wide, this one is model-wide: a planner-and-critic loop run across different model pairings to find where it breaks. It broke in three distinct places, each run isolating one.

That question is unresolved — and it is the single best argument for the dual-B70 box that follows: 12 GB forces a choice between a good planner and a good critic, and the choice is what breaks.

Source: local-inference research series and rehearsal run log, RTX 4070 Ti on Ollama/CUDA, from 5 June 2026.

2 · Intel Arc Pro B70 — the Windows / Vulkan era

Rig: 2× Intel Arc Pro B70 (Battlemage Xe2, 32 GB GDDR6 each) · AMD Ryzen 9 5900X · 32 GB DDR4 · Windows 10 Pro · PCIe 4.0 ×8/×8 (the cards are PCIe 5.0 capable; the host limits them). These three repos are public.

b70tools — is Task Manager telling you the truth?

It wasn't. That's why the tool exists. It reconciles four independent telemetry sources — kernel-mode adapter counters, per-process heap budgets, the Vulkan memory-budget extension, and Intel's own power/thermal API — and flags where they disagree in real time.

Measured on the dual-B70 rig, Windows 10, production Intel driver. May 2026.
WorkloadConfigGenerationPromptVRAM/card
Mistral-Small-3.2-24B Q4_K_Msingle card27.3 tok/s~400–443 tok/s14 GB
Qwen3-30B-A3B MoE Q4_K_Mdual split81.7 tok/s30.1 tok/s8.5 GB
Qwen2.5-32B-Instruct Q4_K_Mdual split20.7 tok/s242.2 tok/s9.3 GB

Look at the middle two rows: the 30B MoE generates four times faster than the 32B dense, while the dense model prompts eight times faster. Different bottlenecks, same box.

The other finding was about the cards themselves. They're physically identical and they are not identical in software: the top-slot card runs 10–15 °C hotter under identical load and its telemetry reports physically impossible values — 5.117 V and 8.55 GHz at idle. That's a driver bug, not a hardware one. It took several experiments to conclude the cards deliver identical inference throughput, which is the whole argument for building the observation tool before trusting the dashboard.

At the time, that conclusion was not airtight. Four days earlier, the overnight bench below measured the bottom card generating 18% faster than the top one on an identical 14B — 50.09 vs 42.37 tok/s. Both results remain here because that contradiction was real in May. The controlled August rebuild A/B later put the cards within ±1.6% and traced the old gap to ECC asymmetry; see the resolution.

Source: github.com/djcdevelopment/b70tools — Windows-native telemetry workbench, May 2026. Observation cost: 16.6 MiB RSS, under 130 ms init.

The overnight campaign — 14B, 32B, and a 70B across both cards

A phased ladder run overnight: each card alone on a 14B, then a 32B, then Llama-3.3-70B tensor-split across both, plus a context curve and a KV-precision comparison. About a hundred files of raw output. This is the run the rest of the Windows era was built on.

llama-bench, Q4_K_M throughout, all layers offloaded, dual Arc Pro B70 on Windows Vulkan. Measured 24 May 2026.
ModelPlacementPrompt tok/sGenerate tok/s
Qwen2.5-14Bsingle card (top slot)1261.0642.37
Qwen2.5-14Bsingle card (bottom slot)1335.6850.09
Qwen2.5-32Bsingle card591.8823.69
Llama-3.3-70Bdual, tensor-split, mmap on21.49
Llama-3.3-70Bdual, tensor-split, mmap off151.4111.28

Source: overnight bench campaign, dual Arc Pro B70 on Windows Vulkan, 24 May 2026. Roughly 100 raw JSONL and log files, retained.

denning — treating the KV cache as an OS-managed resource

The deepest lab in the set, and the one to hand a skeptic. The premise: on a box with no GPU fabric, less system RAM than VRAM, and an operating system arbitrating video memory, correct model-state management is bandwidth-roofline admission control — and you have to coexist with an OS memory manager that will act against you. Named for Peter J. Denning's working-set model. Predictions were committed to git before the data that confirmed them.

The May retraction. The two-card symmetric scaling result — a clean 1.96× by replication — was marked provisional and under correction. Driving both cards symmetrically, including the display card, reproducibly tripped a Windows display-driver timeout and reset: four resets across two sessions. The concurrency run collected across two of those resets was retracted outright rather than published. In August, the headless rebuild isolated the display path and resolved the retraction with clean 1.85× scaling; both stages of the record are retained.

Source: github.com/djcdevelopment/denning — on-rig validation 19–21 June 2026, dual Arc Pro B70, Windows / Vulkan.

The single-card battery — PCIe width turns out not to matter

After a thermal pull left the box on one card, a six-part battery ran at full PCIe ×16: prefill, model load, session rotation, restore-versus-re-prefill, an engine probe, and a 70B deliberately overcommitted. Predictions were locked to a git tag before the data was collected.

Source: pre-registered single-card ×16 battery, one Arc Pro B70 on Windows, 21 June 2026.

battlemage — where it started

The earliest of the three: single- and dual-card Windows 10 benchmarking, including the dual-card 70B-class run on Vulkan that first proved 64 GB of shared inference was real on these cards. Superseded by the two repos above for anything load-bearing, and worth knowing it exists.

Source: github.com/djcdevelopment/battlemage — from May 2026.

3 · Intel Arc Pro B70 — the Linux / SYCL era

Rig: the same box, migrated to native Ubuntu 26.04 on the modern xe kernel driver with a current Level Zero stack · llama.cpp SYCL build · both B70s · Ryzen 9 5900X · 32 GB DDR4.

Does splitting a model across two cards help? (29 June 2026)

The first Linux-native benchmark asked two things: does multi-card work on Ubuntu the way it didn't on Windows, and is splitting actually better than using one card?

Qwen3-30B-A3B Instruct Q4_K_M, KV cache q8_0, all layers offloaded. "single" = one card; "layer" = both cards, split 1:1. Measured 29 June 2026.
Context depthsingle — promptsingle — generatelayer — promptlayer — generate
0 (empty)1205.8586.921146.0184.30
8,192612.1229.12607.2328.64
16,384405.4417.65405.7317.48
32,768241.369.91240.649.86
65,536132.315.28130.714.89
131,072timed out — no throughput inside a 420-second budget

Multi-card layer split works, and buys nothing for a single stream. Parity through 32k, slightly worse at 64k. The second card is capacity, not speed. The more aggressive split modes weren't even candidates: row-split segfaulted on both variants, and tensor-split timed out and then segfaulted.

Source: SYCL context-ladder lab, dual Arc Pro B70 on Ubuntu, 29 June 2026.

The same question through a real server

Benchmarks measure the model. This measured what a client experiences — a resident server at 128k context behind an OpenAI-compatible facade, timing readiness, time-to-first-token, streaming and cache reuse.

Qwen3-30B-A3B Instruct Q4_K_M, KV cache q8_0. Single-card placement, one slot, 32-token generation cap. Measured 29 June 2026.
Prompt tokensTotal requestPrompt tok/sGenerate tok/sTime to first token
7,2089.8 s863.4922.441.03 s
14,34223.4 s671.7215.741.66 s
28,65166.0 s456.649.922.91 s
57,228213.3 s275.695.695.43 s

Then the concurrency ladder, at about 14,300 prompt tokens: 1 client 22.9 s, 2 clients 45.3 s, 4 clients 91.0 s, 8 clients 181.5 s. Almost perfectly linear — which is the bad news. Time-to-first-token is nearly identical to full completion time at 1 through 4 clients, meaning streaming stops being useful the moment concurrent prompt ingest dominates. The conclusion there is architectural, not hardware: what needs work is request scheduling, prompt-cache strategy and admission control.

Source: service-path placement matrix, depth ladder and concurrency ladder, 29 June 2026.

What the two cards can actually say to each other (18 July 2026)

Remember the 6.48 GB/s host-bounced number from Windows. On Linux, with the modern driver and Level Zero stack, the peer-transfer benchmark was built from source and run — 50 iterations, 8 bytes to 256 MB.

Peer-to-peer transfer between the two Arc Pro B70s, Ubuntu. Measured 18 July 2026.
ModeOperationBandwidth at 256 MBLatency floor at 8 B
UnidirectionalWrite14.29 GB/s6.17 µs
UnidirectionalRead13.06 GB/s7.75 µs
Bidirectional (aggregate)Write16.61 GB/s5.07 µs
Bidirectional (aggregate)Read20.92 GB/s6.67 µs
Dedicated copy engineWrite14.29 GB/s

14.29 GB/s is the practical ceiling of PCIe 4.0 ×8. Going through the root complex costs nothing. That is the same traffic that was host-bounced at 6.48 GB/s on Windows — and it is the single cleanest measurement of what the Linux migration was worth on this hardware. Transfers are perfectly symmetric in both directions, and the dedicated copy engine matches the compute engine exactly, so bulk weight and cache movement can ride the blitter without stealing compute.

Source: peer-transfer microbenchmark, dual Arc Pro B70 on Ubuntu, 18 July 2026.

A 63 GB model, resident, on two consumer cards (18 July 2026)

The question was whether a 120-billion-parameter MoE — 63.4 GB in a single file — can be an always-warm service on hardware like this.

gpt-oss-120b MXFP4 on dual Arc Pro B70, layer-split, four serving slots. Measured 18 July 2026.
MetricValue
Cold load to health-OK~3.2 min
Warm prefill221.6 tok/s
Decode, single stream26.6–28.7 tok/s
Concurrent decode, 4 slots~31 tok/s aggregate
Host RAM, steady state12 GiB used / 18 GiB available
Service restarts across the soak0

It fits only because about four layers of experts ride host memory rather than VRAM — that flag is required, not tuning, since 63.4 GB of weights exceed roughly 60 GiB of usable VRAM. Sparse architecture plus a little CPU offload is the whole trick.

Then the attempt to break it

The brief was blunt: find out if it breaks before we build things with it. Six sweeps — concurrency ladder, context depth, decode-heavy, sustained soak, overload burst, and memory pressure. 453 requests, 453 successes, zero restarts.

The one real trap: over-limit prompts truncate silently. A deliberate 17,500-token request against a 16,000-token slot ceiling did not error. It returned success, normal time-to-first-token, normal decode — the server had quietly truncated the prompt to fit. Context loss with no signal to the caller. If you build on a local inference server, enforce payload size yourself. It will not refuse for you.

Source: stress characterization of the resident 120B model, 18 July 2026. Instruments and raw per-request data retained so the campaign is rerunnable after any config change.

The standoff, 30 July 2026

Not a planned benchmark — an incident, and the most instructive record here for anyone about to run several local models at once. Three large models were all enabled as boot-persistent services on one box. They starved each other.

Fixed by making the services mutually exclusive at the init-system level, adding start-rate limits and alerting. The eventual topology decision cut the other two models entirely: one model serves every role.

Source: incident trace, dual Arc Pro B70 on Ubuntu, 30 July 2026.

4 · Rebuilt OMEN — the Windows 11 / Vulkan era

Rig: ASUS Pro WS Z890-ACE SE · Intel Core Ultra 9 285K · 128 GB DDR5-4400 · 2× Arc Pro B70 32 GB, ECC off · Windows 11 · Intel driver 32.0.101.8974. Neither B70 drives a display; the Intel iGPU and ASPEED adapter own the desktop. Campaign: 20–24 August 2026.

Step 5 became the experiment. The cards moved out of the 32 GB host and into a machine with four times the RAM, a new board and CPU, a newer driver, a fresh OS and no display attached to either accelerator. The campaign retained 261 result artifacts, 24 harness files, 42 server logs and a 13,577-sample board-level sensor capture. It was enough of a change to dissolve every old explanation that depended on the host. One of the old limits did not move at all.

The eight-request knee did not move

The rebuilt concurrency ladder climbed from 37 tok/s at one slot to 183 tok/s aggregate at eight. At ten it collapsed to about 80 tok/s, while p95 request latency jumped from roughly 6 to 22 seconds. The machine still had 78 GB of commit headroom. A one-card dense-model discriminator reproduced the same boundary, eliminating dual-card traffic, MoE expert routing and host-memory pressure as explanations.

The wall was one Vulkan dispatch constant. In llama.cpp, decode batches of eight or fewer columns selected a fast matrix-vector kernel. The ninth was accepted, but silently sent the entire batch through a much slower general matrix-multiply path. Turning that constant into the runtime knob GGML_VK_MMV_MAX_COLS moved the knee on command. On Mistral-24B, ninth-thread aggregate decode went from 19.7 to 131.5 tok/s, a 6.7× gain. Across four dense model families the ninth-thread gain was 6.6–6.9×. Ten real concurrent requests fell from 204 to 38 seconds median, with per-stream decode rising from 1.3 to 5.4 tok/s. MoE gains were smaller because expert multiplication uses a separate dispatch gate.

The complete seven-model study, correctness checks, raw data and patch live in the public evidence package; the narrative version is I Upgraded Everything Around the Bottleneck. It Did Not Move. The upstream llama.cpp pull request #27652 remained open as of 27 August 2026.

This did not make the current service 6.7× faster. Its production configuration exposes two deep slots. The knob changes dispatch only when parallel decode exceeds eight, so it is inert at -np 2. The gain is real and causal, but it belongs to a higher-concurrency operating point.

More RAM changed the failure shape, not the need for admission control

The extra memory converted some obvious failures into plausible-looking slow successes. A 70B loaded while commit sat at 92% placed badly and processed prompts at 11.8 tok/s instead of the campaign's 150 tok/s floor. Stopping the co-tenant afterward did not recover it; placement was fixed for the life of the process. Separately, one process configured for four 64k slots silently spilled 10.24 GB into shared memory and lost about 22% throughput. Two 64k slots remained resident.

The durable result of the rebuild: identify the phase before naming the bottleneck.
PhaseWhat boundWhat the campaign observed
LoadCommit headroom and placementA poisoned allocation stayed slow after pressure was removed.
ConcurrencyKernel dispatchThe knee stayed at eight until the Vulkan dispatch limit moved.
DepthKV capacity and attention workFour 64k slots spilled; decode still fell sharply with resident context.

A post-hoc depth replication supports, but does not upgrade, the older pre-registered result: short-prompt decode was 109.3 tok/s and fell to 6.3 at 57,279 prompt tokens, while prefill fell from 1,404 to 163 tok/s. One identical-prefix reuse observation was 284× cheaper than re-prefill. It is useful corroboration on a changed platform, not a new discovery or a distribution.

Four old configuration rules changed state

These updates supersede a blanket rule; they do not erase the older stack that produced it.
QuestionRebuilt resultEditorial state
Disable cooperative matrices?Re-enabling them took pp512 from 720 to 1451/1474 tok/s, with zero display resets over about 6.5 hours.Old defense retired.
Flash attention always loses?FA-off won by 25–87% below the measured 16k–24k crossover, then hard-hung at greater depth.Replace with crossover.
q8 KV is a 2–4× trap?Only a 5–6% tax at 32k and 64k on the newer build; the 128k empty result still reproduced.Stack-specific and superseded.
Can Windows serve symmetrically?Headless B70s delivered 1.85× scaling, 24/24 requests served and zero display resets or hardware errors.Retraction resolved.

The card-equality contradiction closed with it. A controlled A/B found the cards within ±1.6%; the old 18% gap tracked the earlier ECC asymmetry. The historical disagreement remains in section 2 because it was honestly unresolved at the time.

The largest model that fits was not the best service

Three clean server configurations on the rebuilt box. Aggregate figures use each tested configuration.
ModelSingle streamAggregatep95Operating result
Qwen3-30B-A3B95.2 tok/s116.3 tok/s4.1 sDefault service
gpt-oss-120b14.7 tok/s~21.6 tok/s25–32 sFits at the memory edge; banked, not resident
Llama-3.3-70B10.7 tok/s16.1 tok/s21.9 sBenchmark workload

The purchasing lesson is not "run the largest model that fits." Active architecture, latency target and memory margin decide whether a fitting model is a useful resident service.

The workstation held several jobs at once

A two-hour combined-load soak produced 847 serving waves at 103.5 tok/s mean and 5.1% coefficient of variation. Vulkan serving retained about 89% of solo speed beside CPU 70B benchmarking and hash-verified disk traffic. A one-hour finale added XPU LoRA training: serving, training, CPU inference and disk copying ran together with zero display resets, hardware errors or unexpected power events. A shorter serving-plus-training A/B measured a 0.99 throughput ratio.

The first thermal result was less reassuring: one card's VRAM reached 94 °C. A cardboard airflow duct cut the sustained peak by about 12 °C; the finale settled around 70 °C VRAM. That is consumer hardware in the literal sense: the software result depended on case airflow, and the fix was measured before it was trusted.

Sources: rebuilt-OMEN limit campaign, 20–24 August 2026; public Vulkan-cliff evidence package. Campaign claims are distilled from retained harness outputs, server logs and board telemetry.

5 · Two Qwen 3.8 models — when the operating point decides the verdict

Rig: the same rebuilt workstation as section 4 — ASUS Pro WS Z890-ACE SE · Core Ultra 9 285K · 128 GB DDR5 · 2× Arc Pro B70 32 GB, headless · Windows 11 and Vulkan. A pinned engine build was used for both candidates so the incumbent was never disturbed. Campaign: 27 August 2026.

Two new models arrived: a dense 27B and a 125B sparse model with 6B active per token whose weights are 88.1 GiB — larger than the box's entire 64 GB of VRAM. The first was run as a gated promotion campaign against the resident MoE. The second could not be a promotion candidate at all, so it was run as a placement study: how does a model that does not fit behave when you slide it across the CPU/GPU boundary one layer at a time?

The gate measured the one prompt length where the candidate loses

The dense 27B lost its promotion decisively and won the same comparison decisively, depending entirely on prompt length. The gate sampled 512-token prompts at sixteen saturated clients, and there the candidate returned 0.69× the incumbent's completed jobs per hour. Moving along the prompt axis, on the same hardware and the same afternoon, it returned 5.49×.

Completed jobs per hour at sixteen saturated clients, dense 27B against the resident 3B-active MoE. Measured 27 August 2026.
Prompt lengthIncumbent MoEDense 27BRatioWhich wins
512 tokens231615910.69×MoE — decode dominates
8K75519832.63×Dense
32K23813065.49×Dense — prefill dominates

The verdict was do not promote on twelve of fifteen gates, and that is the correct mechanical answer to the question the gate asked. It failed throughput by 31.3% where it needed to gain 25%, and p95 latency at 1.542× against a 1.5× ceiling. It is also an answer about 512-token prompts, which is a claim about the workload that nobody had written down. Where decode dominates, a 3B-active MoE beats a dense 27B doing nine times the arithmetic per token. Where prefill dominates, that reverses.

Quality ran the other way from throughput. Blind pairwise judging, order-reversed and with order-inconsistent pairs discarded, gave the candidate 44 wins, 42 ties and 4 losses — a 95.6% win-or-tie rate against a 60% bar. The deterministic assay reads 0.861 against 0.472, but that gap is inflated and worth correcting here: the incumbent has no vision at all and scores zero on all three vision families. On text-only families the honest comparison is 0.835 against 0.708. A sixty-minute soak at sixteen clients returned 1605 of 1605 valid requests with no system events.

The same crossover, measured the way the rest of this article measures

Completed jobs per hour under sixteen clients is not the number the rest of this article reports, so the two new models were re-measured single-stream through llama-bench — the harness behind most of the figures above. The control that makes those rows admissible ran first: the same incumbent model, at the same placement, on both the production binary and the campaign binary the new models require. They agreed to 0.40%, so the build is not doing any of the work below.

llama-bench, single stream, dual Arc Pro B70 layer-split, flash attention on. Prompt and generation measured separately at three context depths. Measured 27 August 2026.
Modelpp512pp @ 8Kpp @ 32Ktg128tg @ 8Ktg @ 32K
Qwen3-30B-A3B MoE, 3B active2394.7469.3139.8112.0432.4610.45
Qwen3.8-27B dense796.4442.3181.623.7517.359.65
Qwen3.8-Flash-Next 125B, 6B active, host-placed481.7282.3173.38.916.285.08

Read along the rows rather than down the columns. The dense 27B begins 4.7× behind the mixture-of-experts model on generation and 3× behind on prompt processing, and by 32K it has closed to within 8% on generation and overtaken it on prompt processing, 181.6 against 139.8. The 125B model overtakes there too, at 173.3, while remaining the slowest of the three at every generation depth. The ranking at 512 tokens is not the ranking at 32K.

That is the jobs-per-hour inversion again, arrived at from a different direction: a different harness, a different measurement shape, one client instead of sixteen. Two independent routes to the same conclusion is worth more than either alone, and it is the reason this section exists rather than resting on the gate result.

One row in that table is not the same kind of number as the others, and the table does not show it. Every figure above is the mean of three repetitions. For the two models that fit in VRAM those three agree to better than 1.5%, so the mean describes them. Flash-Next's generation repetitions are 7.93, 5.45 and 5.45 tok/s at 8K — a 45% spread, and 19% to 48% across the three depths. The mean is real arithmetic on a series that is not steady.

The direction rules out the obvious explanation. The first repetition is the fastest and the rate then settles lower and stays there, so this is not a cache warming up. Nor is it heat: thermal throttling would punish prompt processing hardest, because that is the compute-bound half, and prompt processing barely moves (4–8%) while generation loses up to 48%. Losing ten times more on the memory-bound half than the compute-bound half is what losing residency looks like — prompt processing amortises a weight fetch over 512 tokens, generation pays it once per token. Flash-Next is the only row here that does not fit in video memory (88.1 GiB of weights against 64 GB of card), and it was the only arm loaded through a memory mapping rather than read into allocated memory.

The placement ladder below prices exactly that, and it says this row is not the model's speed — it is the speed of a half-placed model. Decode on Flash-Next rises monotonically with the number of blocks actually resident on the cards: 7.41 tok/s at sixteen blocks, 9.24 at twenty-four, 27.70 at all forty-eight. The 8.91 above sits between the sixteen- and twenty-four-block rungs. The sweep asked for all forty-eight and did not get them, because a memory mapping leaves the weights on disk to be paged rather than committing them to memory, and the run then slid down that ladder as pages were reclaimed — 9.97 tok/s on the first repetition, 8.4 by the third, about four blocks' worth of residency lost between them.

So the honest reading of the bottom row is not "this model is slow." It is that the same model on the same two cards runs at 8.9 or at 27.7 depending on a load flag, and nothing in a tokens-per-second figure tells you which one you are looking at. That is the entire argument of this article compressed into one row, and it caught the author of the row.

What the measurement itself costs

Two numbers fell out of the sweep that are not about these models at all, and are probably the more useful half for anyone comparing hardware from published figures.

Splitting a model across two cards costs about 7% of generation speed when it would have fit on one. The incumbent measured 121.18 tok/s on a single card and 112.50 tok/s layer-split across both, same binary, same harness, same afternoon. The gap shrinks with depth — about 2% at 8K, about 1% at 32K — and prompt processing is indifferent to placement throughout, varying by under 1.3% in either direction. For a 17.7 GB model on a 32 GB card, the second card buys capacity, not speed. That single-card arm was an accident: llama-bench reads a comma-separated tensor split as two separate configurations rather than one even split, so the run landed every layer on one card. It is kept because, paired with the corrected run, it isolates exactly what placement costs.

The harness is worth about 19% on the fast model and about 3% on the slow one. At matched placement, llama-bench reports 112.04 tok/s where the server reports 94.4 for the mixture-of-experts model, but 23.75 against 23.0 for the dense one. That is what a roughly fixed per-request serving cost looks like: a large share of a fast generation, a small share of a slow one. The practical consequence is that the gap cannot be subtracted as a constant from someone else's published number — two benchmarks of the same card can differ by a fifth before any hardware difference enters, and by more or less depending on how fast the model was to begin with.

The thermal ceiling belongs to one card, not to the pair

Two configurations died at 96 °C, and both times the reading came from VRAM on the same single adapter while its neighbour sat at 86 °C and both GPU cores idled at 77–79 °C. Memory was healthy in both aborts — 56.7 GB of commit headroom free, 1.2 GB shared — so nothing about capacity was involved. The casualties were the placement the campaign had predicted would win, one full model replica per card, which tripped the line at only four concurrent 512-token requests; and the 128K context tier.

At the production operating point the same model then soaked for an hour at sixteen clients and settled at 82 °C. So the ceiling is not a property of running this model, or of these cards as a pair. It belongs to airflow over one specific card, which is the same lesson the cardboard duct taught in section 4, now localised to a single sensor.

An 88 GiB model on 64 GB of VRAM

The 125B sparse model has 48 blocks, 512 experts with ten active per token, and only two key-value heads against twenty-four query heads. It was walked up a placement ladder, moving blocks from CPU to GPU one rung at a time.

Qwen3.8-Flash-Next IQ4_XS, 88.1 GiB, single stream, text-only tasks. Both B70s, host placement for the remainder. Measured 27 August 2026.
Blocks on GPUDecodeValidMemory
0 of 484.75 tok/s6/8No VRAM; 6.5 GB private commit
85.99 tok/s6/8
167.40 tok/s6/8
249.12 tok/s6/8
3211.62 tok/s6/8
4016.69 tok/s6/8
48 of 4827.70 tok/s6/861.2 GB VRAM and 65.3 GB private commit

Two results matter more than the speed curve. The first is that correctness did not move: every rung returned the same six of eight, failing the same two tasks in the same way, with no empty output, no corrupt tokens and no repetition loops anywhere on the ladder. Splitting a model across CPU and GPU changed how fast it answered, not what it answered. Time to first token fell alongside decode, from 27.5 s at no offload to 9.8 s at forty blocks.

The second is the memory bill, and it runs the wrong way. At no offload the process held 6.5 GB of private commit and no VRAM, its weights sitting in evictable file-backed pages. At full offload it held 61.2 GB of VRAM and 65.3 GB of private commit — roughly 126 GB of live memory for a 94 GB model, with both cards at about 95% and one of them pushing 0.93 GB back across the bus into system memory. Offloading to the GPUs did not trade host memory for card memory; it raised both. The per-layer embedding tables never become blocks, so they stay in DDR5 no matter how high the layer count goes. For scale, the resident MoE serves at 94 tok/s on roughly 30 GB of VRAM: the larger model is about 3.4× slower for about four times the memory.

The same flag reverses sign when the model stops fitting. Section 2 records --no-mmap taking a 70B's prompt processing from 21.5 to 151.4 tok/s — a 7× win, for a model that fits. For one that does not, the identical flag is the difference between running and not running: it forces all 88.1 GiB into commit, and the first attempt died in 41 seconds with 0.76 GB of headroom left. Direct I/O, requested separately, bypasses the page cache and defeats the memory mapping even when mapping is switched back on, so both had to go. A host-placed model must page from its mapping. Configuration findings belong to a measured stack and to a memory regime.

Speculative decoding tripled throughput and changed the answers

Multi-token prediction took the dense candidate from 510 to 1591 completed jobs per hour, with draft acceptance climbing from 0.65 to 0.86. It is the single largest configuration win in this campaign, and it is also the one that will not be deployed yet. Speculative decoding is only output-safe when verification is exact: a drafted token is supposed to be accepted only where it matches what the full model would have produced, so at temperature zero the two paths should be byte-identical. They were not. Across every cell compared there were zero identical responses, and all 126 long-prompt requests on the non-speculative path stopped early where the speculative path ran to length. Batch-dependent arithmetic explains some divergence; it does not explain a clean split like that.

The campaign's own compatibility gate had passed this configuration by comparing validity rates between the two paths — 0.875 against 0.875 — and never comparing an output to an output. Equal scores on a rubric are not equivalence.

What a cache miss costs, measured by accident

The campaign contained an unplanned controlled experiment. Its production-shaped runs reuse one filler prompt, so every request after the first hits a warm prefix; its deep-context runs insert a unique retrieval needle, so no request ever does. Same model, same hardware, same 26K depth. With a shared prefix, time to first token went 39.75 s, then 0.13 s, then 0.12 s. With unique prompts it stayed cold at 59.8, 59.6 and 55.4 s. A prefix-cache miss therefore costs about 306× the warm case at that depth, which is the same lever section 3 measured at 68× and section 4 at 284×, now bounded from the other direction.

Sources: Qwen 3.8 promotion campaign and Flash-Next placement study, 27 August 2026. Verdict, scorecard, per-request rows, watchdog telemetry and quarantine records retained; measurements normalised into the cross-harness corpus.

6 · Training on it — LoRA fine-tuning on an Intel card

Rig: one Intel Arc Pro B70, 32 GB · native Windows · PyTorch 2.12.1 with the XPU backend · base model Qwen2.5-Coder-7B. 22 June 2026.

Everything above is about serving models. llama.cpp serves a model; it cannot change one. This is the other half — and it is widely assumed to be unavailable on this hardware.

It got found by accident. A probe meant to assess a different inference engine refuted its own premise: the environment already had PyTorch's XPU build working on native Windows. Not WSL, not a container — native. torch.xpu saw the B70, and matrix multiply ran at 150 TFLOPS fp16 and 158 TFLOPS bf16, with the matrix engines engaged and bf16 available — a numeric format the Vulkan serving path doesn't have. The engine under test was blocked on Windows by packaging, not by any absence of hardware support. That reframed the whole question and opened the door to training.

The memory ceiling, mapped before trusting it

Before any real run, peak VRAM was swept against sequence length, batch size and gradient checkpointing — to find the envelope deliberately rather than discover it by crashing.

bf16 LoRA, rank 16, all linear layers, on a 7B base. One Arc Pro B70, 32 GB. Measured 22 June 2026.
Sequence · batch · gradient checkpointingPeak VRAMResult
1024 · 1 · off30.3 GiBfits, barely
2048 · 1 · off~45 GiBout of memory
4096 · 1 · off~43 GiBout of memory
2048 · 4 · off~40 GiBout of memory
4096 · 1 · ON27.3 GiBthe working envelope
8192 · 1 · on~37 GiBout of memory

Gradient checkpointing isn't a tuning knob here, it's a requirement. Without it the 7B runs out of memory by a 2048-token sequence. With it, 4096 tokens at batch size 1 fits in 27.3 GiB with room to spare. Past that you need 4-bit quantized training. Note also that the over-budget runs didn't fail cleanly — they climbed past 32 GiB by spilling into shared system RAM first, the same OS demotion behaviour that haunts the serving side.

The actual fine-tune, and closing the loop

Then a real run inside that envelope: 40.4 million trainable parameters — 0.53% of the model — held-out loss 2.96 to 1.13, training loss 2.33 to 0.91, peak 17.0 GiB, sixty optimizer updates in 108 seconds, producing a 157 MB adapter.

Then the part that makes it a loop rather than a demo. Merge the adapter into the base weights, convert to the serving format (14.2 GB at half precision), quantize to 4-bit (4.36 GB), and hand it to the ordinary Vulkan serving path. It loaded in about six seconds and generated a correct Fibonacci function.

Train on one backend, serve on another, same card. The fine-tuned model re-enters the existing serving path unchanged — nothing downstream needed to know it had been modified. That's the whole proposition: a consumer Intel GPU that can both change a model's weights and then serve the result, on Windows, without a datacenter or a cloud bill.

Two honest limits. This is LoRA, not full fine-tuning — half a percent of the parameters, not all of them. And a 7B is a small model; 14B and up, or longer sequences, need 4-bit training that has not been run. What is proven is that the path exists end to end and the envelope is mapped.

Source: on-rig fine-tuning project, phases 0–4, one Arc Pro B70 on Windows with PyTorch XPU, 22 June 2026.

The serving substrate underneath

Worth a line because it is the connective tissue: partway through, Ollama came off that box in favour of something leaner — its memory overhead was the motivation. A Windows-native Vulkan dual-card substrate that owns config, the exact launch recipe, process lifecycle and structured status, behind a minimal OpenAI-compatible surface. It ran a 30B mixture-of-experts dual-split at 128k context alongside a 14B critic, from June 2026, and it's the direct ancestor of the Linux serving work in section 3.

Its most-earned design decision is one line: readiness means "can serve", not "process is up." Ignoring it produced a production incident on the Linux side five weeks later, when health checks that only probed the socket cheerfully reported success while a model was still loading — or already dying.

Source: dual-B70 Windows Vulkan serving substrate, June 2026. A dozen dated run directories with per-run manifests and telemetry, largely unmined.

7 · Does the output hold up?

Throughput is the easy half. These labs asked whether local models produce good enough output — and ran into a harder problem, which is that when a language model grades the work, you may be measuring the ruler.

The matrix wind tunnel — does self-refinement help?

Six rounds of overnight experiments on idle hardware, planner and critic deliberately running on different physical boxes: the 30B MoE on the dual-B70 machine, a 30B coder model on the RTX 5070. Each cell was a combination of planner, critic, prompt archetype, refinement laps and ordering, scored 0–100 by a held-out judge panel. 192 cells in the largest round.

Mean planning score by refinement laps. Round 4 used the original judge; round 6 re-scored with a neutral one. July 2026.
Refinement lapsRound 4 (original judge)Round 6 (neutral judge)
1 lap81.287.5
2 laps86.285.6
3 laps82.684.5
4 laps75.982.0

Read those two columns side by side. Round 4 tells a dramatic story — two laps is the sweet spot, four laps collapses. Round 6 says lap count barely matters, and the collapse was largely an artifact of the first judge's bias toward brevity. The same judge inflated the one real effect in the data — a concise author prompt, "shortest complete answer, lead with the decision" — from +2.8 points to +8.0.

The follow-on work quantified it: judges are near-deterministic on repeat (0.40 point sampling deviation) and wildly variable across rubrics (6.5 points average disagreement, up to 47). So resolving a genuine ~3-point effect takes roughly 19 lens-diverse votes but only about 1 repeat vote. Spend the evaluation budget on different rubrics, not on asking the same judge twice.

The value here wasn't the answer about laps — that's genuinely "meh". It was that the instrument flagged an effect, corrected it, explained it, and then caught its own measurement bias. Entirely on idle local hardware, overnight, at no marginal cost.

Source: github.com/djcdevelopment/windtunnel — dual Arc Pro B70 plus RTX 5070, 5–10 July 2026.

8 · Local vs cloud — quality, latency, money

The head-to-head (21 July 2026)

A scored documentation and decision-record task, run identically across the local model and two frontier cloud models. Fifteen cells, all successful.

Documentation and decision-record authoring task. Measured 21 July 2026.
Where it ranModelMean qualityMean latencyMarginal cost
Local, dual B70gpt-oss-120b93.8105.2 s$0.00
CloudGemini 3.1 Pro (preview)90.628.1 smetered
CloudGemini 3.5 Flash89.219.8 smetered

The local model won on quality, on two consumer GPUs, at zero marginal cost. The cloud won on latency by 4–5×, and that was the whole of its advantage on this task.

What cloud actually costs (23 July 2026)

A billing audit after standing up a cloud agent that reaches back into the local GPUs across an audited network boundary. The finding was not about tokens. Idle standing infrastructure produced an unexpected $36–38 bill — two managed agent runtimes at roughly $3.50/day each plus a small VM — while actual inference spend was trivial. If you are comparing local against cloud, compare the standing cost, not the per-token rate.

The ledger

Not a lab — the production record. Every model call routed through the local gateway is logged with backend, model, token counts and outcome, so the economics are measured rather than argued.

Offload ledger, watermark 5 August 2026. Totals across all backends: 1,199 calls, 2.80 M tokens in, 363 k tokens out.
WhereModelCost classCallsTokens inTokens outReal spend
Local, dual B70gpt-oss-120bsunk237326,012140,228$0.00
Local, dual B70Qwen3-30B-A3B Q4sunk33236,97015,431$0.00
CloudGemini 3.5 Flashtrial credit237411,17956,033$0.91
CloudGemini 3.1 Protrial credit1852,022,132150,362$5.35

618 of those calls landed on already-owned hardware. Every backend listed ran at a 100% success rate.

9 · What is not measured

Stated plainly so nobody cites this for something it doesn't cover.

If you're deciding whether to try this

The honest summary is that a 30B-class mixture-of-experts model ran at 95.2 tok/s on the rebuilt dual-card service, while the much larger 120B fit at the memory edge and delivered 14.7 tok/s single-stream. A 70B works too, once the placement flags are right. The output quality is competitive with frontier cloud on real work, and the same hardware can fine-tune as well as serve. The largest model that fits is not automatically the best rung.

Buy enough RAM, but identify the phase before naming the wall. Probe commit and placement before load, enforce admission control at the measured kernel boundary, cache prefixes, and budget KV at depth. Check your own payload sizes too: one server in this record silently truncated an over-limit prompt and reported success. Then build the observation tool before trusting the dashboard — because this platform also reported 5.117 volts and 8.55 gigahertz at idle, and cheerfully meant it.

Cite, inspect, reproduce

This report is meant to be challenged at the operating point where you intend to use it. The public package below separates the claims registry, causal Vulkan result, narrative explanation, and raw demonstration so a correction can name the exact layer it changes.

Article version: 2026-08-28 Published: 2026-08-05 Registry: 25 checked claims

Ciula, Derek. “What Consumer Hardware Actually Does.” Steppe Integrations, 5 August 2026, updated 28 August 2026, https://steppeintegrations.com/articles/what-consumer-hardware-actually-does/.

Ran one of these cells on another card, driver, backend, or prompt shape? Open an issue with the model, quantization, build fingerprint, placement, context, concurrency, and raw log. A result that disagrees is more useful here than a vote that agrees.

Upstream implementation: llama.cpp pull request #27652.