I Upgraded Everything Around the Bottleneck. It Did Not Move.

A local-AI multitasking ceiling of exactly 8 survived a complete machine rebuild, because it was never hardware. Finding one line of code, proving it guilty by dose-response across seven models everyone knows, and moving it: the 9th parallel request gets 6.7× faster on the same GPU.

Illustrated explainer panel: a sketched throughput curve climbs to 8 concurrent requests and falls off a drawn cliff edge, with callouts for the fast matrix-vector kernel below the limit and the slow general path above it
The cliff, as a notebook sketch. Illustration; the measured curves are further down.

In May I found a capacity cliff and explained it. In August I accidentally proved my explanation wrong without touching the thing I was measuring. Then I found the real explanation, sitting in my own repository, one directory away from the wrong one, the whole time.

The cliff, and the story I told about it

May 2026. Two Intel Arc Pro B70s in an aging AMD box: Ryzen 9 5900X, 32 GB of DDR4, Windows 10. I was measuring how many concurrent LLM sessions the pair could serve before quality collapsed.

Throughput climbed with every session I added, peaked at eight concurrent sessions, then fell off a cliff. Not a taper. By ten sessions, aggregate throughput had collapsed to a third and latency had blown through every service target I had.

I went looking for a reason and found one immediately. Windows commit charge, the system's virtual memory commitment, was at 92 percent right at the cliff. The box had 32 GB of RAM and models are hungry. Plausible mechanism, measured correlation, tidy conclusion. I wrote it down: the host RAM is the constraint. Eight is what 32 GB buys you.

I believed that for three months.

The accidental experiment

In August the two cards moved into a rebuilt machine. Intel Core Ultra 9 285K. 128 GB of DDR5. Windows 11. A GPU driver a full year newer. Different motherboard, different PSU, different case. The only things that survived the migration were the two GPUs and the models on disk.

If you wanted to design an experiment to test “the host RAM is the constraint,” you could not do much better than this. Change everything around the suspect. Keep the suspect. See if the symptom moves.

I re-ran the concurrency ladder with a different serving harness and a different model, a 30B mixture-of-experts split across both cards. One slot: 40 tokens per second. Two: 72. Four: 113. Six: 145. Eight: 183.

Ten: 80. Latency p95 went from six seconds to twenty-two.

The identical knee. The identical cliff. At the identical count. And this time the machine had 78 GB of commit headroom free at the moment of collapse. The host was not merely comfortable. It was barely awake.

Four-panel illustrated card: three concurrency ladders on different machines and model classes all break at eight, and a verdict panel showing the code line mul_mat_vec_max_cols = 8 behind a padlock
Three ladders, two machines, three model classes, one knee. Illustration of the discovery arc.

The explanation was already on my disk

I went hunting for the mechanism and found it in the cost model of my own public benchmarking repo, written in May, during the same campaign that produced the wrong attribution:

N*_bw = SLO_TBT x B_vram / (k x L)     (bandwidth knee)
N*    = min( N*_bw , capacity_limit )

In words: aggregate decode throughput on a GPU is a memory-bandwidth roofline. Adding sessions does not add aggregate tokens per second; it divides the same bandwidth into thinner slices. Now look at the terms. A latency target. VRAM bandwidth. Model bytes per token. The host appears nowhere in that equation.

The formula I derived in May predicts the host-invariance I paid a full rebuild to discover in August. I had derived the answer, published it, and then filed the question under a different explanation one directory over. The commit-charge reading was real. It redlined at the same moment. It was the capacity branch of my own min() expression looking like the binding term while the other branch did the binding.

The twenty-minute experiment

Two things still did not fit. The formula predicts a plateau past the knee, and I measured a collapse. And a mixture-of-experts model reads far fewer bytes per token than a dense model, so its bandwidth knee should have sat somewhere else entirely. Which whispered that the binding term at eight was not bandwidth at all: something in the serving stack.

There was a twenty-minute experiment that discriminates. Run the same ladder on a small dense model, on one card, nothing else resident. If the knee is the bandwidth roofline, it moves. If it is still exactly eight, the limit lives in software.

0501001502001 slots: 40 tok/s12 slots: 72 tok/s24 slots: 113 tok/s46 slots: 145 tok/s68 slots: 183 tok/s810 slots: 80 tok/s1030B MoE, two GPUs, 128 GB hostaggregate tok/sknee at 80501001501 slots: 27 tok/s18 slots: 134 tok/s810 slots: 27 tok/s10dense 14B, ONE GPUknee at 8concurrent sessions - different machine, model, and card count; identical knee
Measured. Left: the 30B mixture-of-experts across both GPUs on the rebuilt box. Right: a dense 14B on a single GPU. Different bytes per token, different card count, different machine generation than May. The knee does not care.

Dense 14B, single card: one slot, 27 tokens per second. Eight slots, 134. Ten slots, 27, with p95 latency going from nine seconds to seventy. Exactly eight. Third configuration, most violent cliff of the three.

The wall has a name and a line number

That is not physics behaving like a constant. That is a constant.

static constexpr uint32_t mul_mat_vec_max_cols = 8;

ggml/src/ggml-vulkan/ggml-vulkan.cpp, the Vulkan backend of llama.cpp.

Decode batches of eight or fewer sequences take the fast matrix-vector kernel path. Nine or more fall through to the general matrix-multiply path, which at decode shapes is catastrophic until the batch grows much larger. That explains the count, on every host, at any card count, for any model, because a compile-time constant is blind to all of them. It explains the cliff instead of a plateau, because a kernel-path switch is discontinuous. It even explains the small recovery from ten to twelve slots that I had no story for: the slow path scaling up, which is exactly what a path switch looks like and exactly what resource exhaustion does not.

The number moved

A constant is an accusation until you move it. I patched it into a runtime switch: one file, two dozen lines, default behavior byte-identical to stock. Then I ran the same benchmark three times on the same binary with the limit set to 4, 8, and 16.

The cliff landed at 4, 8, and 16.

0501001501246891012162432parallel sequences (threads in flight)aggregate decode tok/slimit 4, B=1: 30.5 tok/slimit 4, B=2: 47.0 tok/slimit 4, B=4: 78.3 tok/slimit 4, B=6: 23.3 tok/slimit 4, B=8: 29.9 tok/slimit 4, B=9: 19.7 tok/slimit 4, B=10: 22.4 tok/slimit 4, B=12: 29.4 tok/slimit 4, B=16: 45.5 tok/slimit 4, B=24: 92.3 tok/slimit 4, B=32: 140.1 tok/slimit 8 (stock), B=1: 30.5 tok/slimit 8 (stock), B=2: 46.9 tok/slimit 8 (stock), B=4: 78.2 tok/slimit 8 (stock), B=6: 106.4 tok/slimit 8 (stock), B=8: 125.0 tok/slimit 8 (stock), B=9: 19.7 tok/slimit 8 (stock), B=10: 22.4 tok/slimit 8 (stock), B=12: 29.4 tok/slimit 8 (stock), B=16: 45.4 tok/slimit 8 (stock), B=24: 92.1 tok/slimit 8 (stock), B=32: 139.5 tok/slimit 16, B=1: 30.5 tok/slimit 16, B=2: 47.0 tok/slimit 16, B=4: 78.4 tok/slimit 16, B=6: 107.0 tok/slimit 16, B=8: 125.2 tok/slimit 16, B=9: 131.5 tok/slimit 16, B=10: 137.9 tok/slimit 16, B=12: 147.4 tok/slimit 16, B=16: 158.5 tok/slimit 16, B=24: 92.3 tok/slimit 16, B=32: 139.5 tok/scliff >4cliff >8cliff >16move the number, the cliff moveslimit 4limit 8 (stock)limit 16
Measured dose-response, Mistral-Small-24B on one Arc Pro B70. Line widths are stepped because the three runs produce identical numbers wherever the curves share a path; each collapses at exactly the first batch past its own limit, and all three converge on the same slow-path floor beyond it.

Set the limit to four and the machine breaks at the fifth session. Leave it at eight and it breaks at the ninth. Raise it to sixteen and eight sessions’ worth of ceiling stops existing: the ninth session goes from 20 tokens per second aggregate to 131. That is a dose-response curve, and dose-response is what causality looks like when it signs its work.

Then I asked whether it was a quirk of one model. Seven models people actually recognize, same Windows 11 machine, same Intel Arc Pro B70, stock limit versus 16:

modelclass9th thread, stock9th thread, patchedgain
Phi-4 14B (Microsoft)dense28.7195.1×6.8
Gemma-3-27B (Google)dense16.1110.3×6.8
Qwen2.5-32B (Alibaba)dense12.888.1×6.9
Mistral-Small-24B (Mistral)dense19.7130.6×6.6
Llama-3.3-70B, 2 GPUs (Meta)dense7.325.4×3.5
gpt-oss-20b (OpenAI)MoE34.863.2×1.8
Qwen3-30B-A3B (Alibaba)MoE65.874.3×1.1

Aggregate decode tokens per second at batch 9, warm runs, llama-batched-bench thread protocol. Full ladders for every model are in the repo.

Every dense model gains six to seven times at the ninth thread. The two sparse models gain less, and even that is evidence: only part of their compute goes through this kernel, and they improve by exactly that part. My favorite detail is Gemma: on the stock build, Gemma-3-27B at thirty-two sessions still moves fewer total tokens than it did at eight. With the shipped constant, adding sessions past eight can never win.

What the cliff feels like

Throughput charts undersell it. I run parallel coding agents against this box, so my unit of value is: how many assistants can this machine keep in flow at once? Gamers measure their experience in frame pacing and one-percent lows, so I measured mine the same way. Two hundred requests, ten at a time, per-request latency logged like a frame-time trace.

050100150200250request 1: 211.58 srequest 2: 211.58 srequest 3: 211.58 srequest 4: 211.58 srequest 5: 211.58 srequest 6: 211.58 srequest 7: 211.58 srequest 8: 211.58 srequest 9: 211.58 srequest 10: 211.58 srequest 11: 210.42 srequest 12: 210.42 srequest 13: 210.42 srequest 14: 210.42 srequest 15: 210.42 srequest 16: 210.42 srequest 17: 210.42 srequest 18: 210.42 srequest 19: 210.42 srequest 20: 210.42 srequest 21: 220.90 srequest 22: 220.90 srequest 23: 220.90 srequest 24: 220.90 srequest 25: 220.91 srequest 26: 220.90 srequest 27: 220.90 srequest 28: 220.91 srequest 29: 220.91 srequest 30: 220.91 srequest 31: 190.07 srequest 32: 190.07 srequest 33: 190.07 srequest 34: 190.07 srequest 35: 190.07 srequest 36: 190.07 srequest 37: 190.07 srequest 38: 190.07 srequest 39: 190.07 srequest 40: 190.07 srequest 41: 196.72 srequest 42: 196.72 srequest 43: 196.72 srequest 44: 196.72 srequest 45: 196.72 srequest 46: 196.72 srequest 47: 196.72 srequest 48: 196.72 srequest 49: 196.72 srequest 50: 196.72 srequest 51: 223.23 srequest 52: 223.23 srequest 53: 223.23 srequest 54: 223.23 srequest 55: 223.23 srequest 56: 223.23 srequest 57: 223.24 srequest 58: 223.23 srequest 59: 223.23 srequest 60: 223.23 srequest 61: 256.48 srequest 62: 256.48 srequest 63: 256.48 srequest 64: 256.48 srequest 65: 256.48 srequest 66: 256.48 srequest 67: 256.48 srequest 68: 256.48 srequest 69: 256.48 srequest 70: 256.48 srequest 71: 249.23 srequest 72: 249.22 srequest 73: 249.23 srequest 74: 249.22 srequest 75: 248.12 srequest 76: 249.22 srequest 77: 249.22 srequest 78: 249.23 srequest 79: 249.23 srequest 80: 249.23 srequest 81: 241.95 srequest 82: 241.95 srequest 83: 241.95 srequest 84: 241.95 srequest 85: 241.95 srequest 86: 241.95 srequest 87: 241.95 srequest 88: 241.95 srequest 89: 241.95 srequest 90: 241.95 srequest 91: 247.70 srequest 92: 247.70 srequest 93: 247.70 srequest 94: 247.70 srequest 95: 247.70 srequest 96: 247.70 srequest 97: 247.70 srequest 98: 247.70 srequest 99: 247.70 srequest 100: 247.70 srequest 101: 241.77 srequest 102: 241.77 srequest 103: 241.77 srequest 104: 240.42 srequest 105: 241.77 srequest 106: 241.77 srequest 107: 241.77 srequest 108: 241.77 srequest 109: 241.77 srequest 110: 241.77 srequest 111: 197.82 srequest 112: 197.82 srequest 113: 197.82 srequest 114: 197.82 srequest 115: 197.82 srequest 116: 197.82 srequest 117: 197.82 srequest 118: 197.82 srequest 119: 197.82 srequest 120: 197.82 srequest 121: 255.23 srequest 122: 255.23 srequest 123: 255.23 srequest 124: 255.23 srequest 125: 255.23 srequest 126: 255.23 srequest 127: 255.23 srequest 128: 255.23 srequest 129: 255.23 srequest 130: 255.23 srequest 131: 147.93 srequest 132: 147.93 srequest 133: 147.93 srequest 134: 147.82 srequest 135: 147.82 srequest 136: 147.93 srequest 137: 147.82 srequest 138: 147.93 srequest 139: 147.93 srequest 140: 147.93 srequest 141: 88.70 srequest 142: 88.70 srequest 143: 88.69 srequest 144: 88.70 srequest 145: 88.70 srequest 146: 88.70 srequest 147: 88.70 srequest 148: 88.70 srequest 149: 88.70 srequest 150: 88.70 srequest 151: 88.77 srequest 152: 88.77 srequest 153: 88.77 srequest 154: 88.77 srequest 155: 88.78 srequest 156: 88.77 srequest 157: 88.78 srequest 158: 88.77 srequest 159: 88.78 srequest 160: 88.77 srequest 161: 88.85 srequest 162: 88.85 srequest 163: 88.85 srequest 164: 88.85 srequest 165: 88.85 srequest 166: 88.85 srequest 167: 88.85 srequest 168: 88.85 srequest 169: 88.85 srequest 170: 88.85 srequest 171: 88.90 srequest 172: 88.90 srequest 173: 88.90 srequest 174: 88.91 srequest 175: 88.90 srequest 176: 88.48 srequest 177: 88.90 srequest 178: 88.91 srequest 179: 88.91 srequest 180: 88.91 srequest 181: 90.59 srequest 182: 90.58 srequest 183: 90.58 srequest 184: 90.58 srequest 185: 90.59 srequest 186: 90.59 srequest 187: 90.59 srequest 188: 90.59 srequest 189: 90.59 srequest 190: 90.59 srequest 191: 164.32 srequest 192: 164.32 srequest 193: 164.32 srequest 194: 164.32 srequest 195: 164.32 srequest 196: 164.33 srequest 197: 164.33 srequest 198: 164.33 srequest 199: 164.33 srequest 200: 164.33 sstock (limit 8) median 204.1 s per request - 200 requests, 10 in flight050100150200250request 1: 38.25 srequest 2: 38.25 srequest 3: 38.25 srequest 4: 38.25 srequest 5: 38.25 srequest 6: 38.25 srequest 7: 38.25 srequest 8: 38.15 srequest 9: 38.25 srequest 10: 38.25 srequest 11: 34.96 srequest 12: 34.96 srequest 13: 34.96 srequest 14: 34.96 srequest 15: 34.96 srequest 16: 34.96 srequest 17: 34.96 srequest 18: 34.96 srequest 19: 34.96 srequest 20: 34.96 srequest 21: 40.89 srequest 22: 40.89 srequest 23: 40.89 srequest 24: 40.89 srequest 25: 40.89 srequest 26: 40.76 srequest 27: 40.89 srequest 28: 40.89 srequest 29: 40.89 srequest 30: 40.89 srequest 31: 67.21 srequest 32: 67.21 srequest 33: 67.22 srequest 34: 67.21 srequest 35: 67.21 srequest 36: 67.21 srequest 37: 67.22 srequest 38: 67.22 srequest 39: 67.22 srequest 40: 67.21 srequest 41: 42.07 srequest 42: 42.07 srequest 43: 42.07 srequest 44: 42.07 srequest 45: 42.07 srequest 46: 42.08 srequest 47: 42.08 srequest 48: 42.08 srequest 49: 42.07 srequest 50: 42.07 srequest 51: 43.77 srequest 52: 43.77 srequest 53: 43.77 srequest 54: 43.77 srequest 55: 43.77 srequest 56: 43.77 srequest 57: 43.77 srequest 58: 43.77 srequest 59: 43.77 srequest 60: 43.77 srequest 61: 32.61 srequest 62: 32.61 srequest 63: 32.61 srequest 64: 32.53 srequest 65: 32.61 srequest 66: 32.61 srequest 67: 32.53 srequest 68: 32.61 srequest 69: 32.61 srequest 70: 32.61 srequest 71: 29.19 srequest 72: 29.19 srequest 73: 29.20 srequest 74: 29.19 srequest 75: 29.20 srequest 76: 29.20 srequest 77: 29.19 srequest 78: 29.20 srequest 79: 29.20 srequest 80: 29.20 srequest 81: 34.80 srequest 82: 34.80 srequest 83: 34.80 srequest 84: 34.80 srequest 85: 34.80 srequest 86: 34.80 srequest 87: 34.80 srequest 88: 34.80 srequest 89: 34.80 srequest 90: 34.80 srequest 91: 27.15 srequest 92: 27.15 srequest 93: 27.15 srequest 94: 27.15 srequest 95: 27.15 srequest 96: 27.15 srequest 97: 27.15 srequest 98: 27.15 srequest 99: 27.15 srequest 100: 27.15 srequest 101: 40.94 srequest 102: 40.94 srequest 103: 40.94 srequest 104: 40.94 srequest 105: 40.95 srequest 106: 40.95 srequest 107: 40.95 srequest 108: 40.94 srequest 109: 40.95 srequest 110: 40.95 srequest 111: 38.40 srequest 112: 38.40 srequest 113: 38.40 srequest 114: 38.40 srequest 115: 38.40 srequest 116: 38.40 srequest 117: 38.40 srequest 118: 38.40 srequest 119: 38.40 srequest 120: 38.40 srequest 121: 31.44 srequest 122: 31.44 srequest 123: 31.44 srequest 124: 31.44 srequest 125: 31.44 srequest 126: 31.44 srequest 127: 31.44 srequest 128: 31.44 srequest 129: 31.44 srequest 130: 31.44 srequest 131: 41.29 srequest 132: 41.29 srequest 133: 41.29 srequest 134: 41.29 srequest 135: 41.29 srequest 136: 41.29 srequest 137: 41.29 srequest 138: 41.29 srequest 139: 41.29 srequest 140: 41.29 srequest 141: 38.13 srequest 142: 38.13 srequest 143: 38.13 srequest 144: 38.13 srequest 145: 38.13 srequest 146: 38.13 srequest 147: 38.14 srequest 148: 38.13 srequest 149: 38.13 srequest 150: 38.13 srequest 151: 42.47 srequest 152: 42.47 srequest 153: 42.47 srequest 154: 42.47 srequest 155: 42.47 srequest 156: 42.47 srequest 157: 42.47 srequest 158: 42.47 srequest 159: 42.47 srequest 160: 42.47 srequest 161: 32.78 srequest 162: 32.77 srequest 163: 32.78 srequest 164: 32.78 srequest 165: 32.78 srequest 166: 32.78 srequest 167: 32.78 srequest 168: 32.78 srequest 169: 32.78 srequest 170: 32.78 srequest 171: 34.42 srequest 172: 34.42 srequest 173: 34.42 srequest 174: 34.42 srequest 175: 34.42 srequest 176: 34.42 srequest 177: 34.42 srequest 178: 34.42 srequest 179: 34.42 srequest 180: 34.42 srequest 181: 35.06 srequest 182: 35.06 srequest 183: 35.06 srequest 184: 35.07 srequest 185: 35.07 srequest 186: 35.07 srequest 187: 35.07 srequest 188: 35.07 srequest 189: 35.07 srequest 190: 35.07 srequest 191: 37.73 srequest 192: 37.73 srequest 193: 37.73 srequest 194: 37.73 srequest 195: 37.73 srequest 196: 37.73 srequest 197: 37.73 srequest 198: 37.73 srequest 199: 37.73 srequest 200: 37.73 spatched (limit 16) median 37.9 s per request - 200 requests, 10 in flightrequest index (completion order) - seconds per request, lower and flatter is better
Measured. Each bar is one request. Stock at ten concurrent is not stuttery; it is uniformly broken, every stream crawling at 1.3 tokens per second, below human reading speed. Same load on the patched build: median 38 seconds and 5.4 tokens per second per stream.
204 s → 38 s
median request latency at ten concurrent, stock → patched. A ×5.4 improvement with identical hardware and watts.
1.3 → 5.4
tokens per second each stream still delivers at ten concurrent. Below reading speed, versus comfortably above it.
8 → 16
the dispatch ceiling, moved. Twice the parallel capacity from the same silicon, at zero hardware cost.

Did the math survive?

Three checks, because a speedup that changes answers is just a bug with good marketing. The backend's kernel tests pass against the CPU reference at the raised limit, exercising batch widths 9 through 16 that stock builds can never reach. The patched binary at its default is stock within noise.

The third check took a correction that made it stronger. My first perplexity A/B matched to every digit, for the wrong reason: the tool's default micro-batch is wider than the knob's whole window, so the new code path was never taken and the test could not fail. The original reporter of the upstream issue caught it while reviewing the patch. Re-run with the micro-batch forced inside the window, the two kernel paths differ by three hundredths of a percent, far inside the estimate's error bar, and the same-kernel control matches to every digit. Different accumulation order, same quality. A null test retracted and replaced is worth more than a green checkmark that cannot go red.

Illustrated results panel: before and after bars for the ninth-thread recovery, a checklist of verification steps, and a torn-paper motif
The results, as the notebook drew them. Illustration; exact numbers above and in the repo.

The lesson

My May conclusion was a correlation wearing a causation costume. Two gauges redlined at the same moment and I charged the one I understood.

Capacity limits deserve attribution experiments, not just correlated readings. Admission control should pin at the measured knee, and the trigger for re-measuring is a software update, not a hardware upgrade. I had that exactly backwards: I re-benchmarked when the metal changed, when the thing that actually owned my ceiling was a header constant that has survived every machine I ran it on.

And the lesson I did not expect: past rigor is an asset you have to actually consult. The right equation sat in my own repo for three months while I repeated a wrong explanation from memory. Instruments do not just need to be built. They need to be read.

The serving stack in this story runs my local inference full time, and until this week it was capped below the knee. Next week the same machine serves twice as many assistants. The hardware will not have changed at all.

The whole hunt in six and a half minutes: the fake hardware wall, the discriminating experiment, the line of code, and the number moving.

Receipts. Everything is public in the vulkancliff repo: the one-file patch, every raw benchmark output, the spill-guarded harnesses that produced them, the charts and their generator, and the video above (also mirrored on GitHub). The same cliff is reported on AMD hardware in llama.cpp issue #25356; the CUDA backend merged the equivalent knob in August. The patch is now upstream as PR #27652, with this repo as its evidence base.