Why memory bandwidth decides how fast local AI runs

Why memory bandwidth decides how fast local AI runs

Brian SanchezSoftware engineer
Published
Reading time4 min

A language model reads every weight for every token, so the speed limit is memory bandwidth, not compute. Tokens per second for nine graphics cards and three model sizes.

info

We may earn a commission if you make a purchase through our links, at no extra cost to you. This helps support our work.

Two graphics cards with the same amount of memory can generate text at very different speeds, and the specification that predicts the gap is not the one on the front of the box. It is memory bandwidth. This article explains why, puts real figures next to nine cards, and shows where the rule stops holding.

Why bandwidth and not compute

A language model writes one token at a time, and to produce each one it has to read every weight it is going to use. An 8-billion-parameter model quantized to 4 bits is about 4 GB of weights. Generating 100 tokens a second means moving 400 GB a second, every second, from the card's memory into its compute units.

That is an enormous amount of traffic and almost no arithmetic per byte moved — one multiply-accumulate for each weight read. The compute units finish their work and wait. So the ceiling is set by how fast memory can be read, and the estimate falls out of one division:

tokens per second ≈ memory bandwidth ÷ model size in memory, times an efficiency factor for the traffic that is not weights. Our inference speed estimator uses 0.75, and the methodology page explains where that number comes from.

The figures

Four-bit weights (Q4_K_M), so model size in gigabytes is roughly half the parameter count. A dash means the weights do not fit in the card's memory at all.

CardBandwidthMemoryLlama 3.1 8BQwen3 14BQwen3 32B
RTX 50901792 GB/s32 GB335 tok/s182 tok/s82 tok/s
RTX 40901008 GB/s24 GB188 tok/s102 tok/s46 tok/s
RTX 5080960 GB/s16 GB179 tok/s97 tok/s—
RTX 5070 Ti896 GB/s16 GB167 tok/s91 tok/s—
RTX 4080717 GB/s16 GB134 tok/s73 tok/s—
RTX 5070672 GB/s12 GB126 tok/s68 tok/s—
RX 7800 XT624 GB/s16 GB117 tok/s63 tok/s—
RTX 3070 Ti608 GB/s8 GB114 tok/s62 tok/s *—
Arc A770512 GB/s16 GB96 tok/s52 tok/s—
* The 14B weights are 7.4 GB on an 8 GB card. They technically fit and leave nothing for the cache, so in practice that row runs at a short context or not at all.

Three things worth reading off it

Memory size and memory speed are separate gates

The RX 7800 XT and the RTX 5070 illustrate it. The Radeon has more memory — 16 GB against 12 — so it can hold a model the GeForce cannot. The GeForce has more bandwidth, so on a model they both hold, it generates faster. Capacity decides whether a model runs; bandwidth decides how fast. Buying for one and hoping for the other is the most common mistake here.

The RTX 3070 Ti is the sharper version of the same lesson. At 608 GB/s it moves data faster than an Arc A770, but with 8 GB it is shut out of everything above about 13 billion parameters. Its speed is real and mostly unusable.

The gap between generations is a memory gap

The RTX 5090 generates about 1.8 times as many tokens per second as an RTX 4090 on the same model. Its bandwidth is about 1.8 times higher. That is not a coincidence, and it is why generation-over-generation gaming benchmarks are a poor guide to this particular job: they measure a workload that is limited by something else.

Big models are slow on every card

Qwen3 32B needs about 16 GB of weights. Even on an RTX 5090 with the highest bandwidth on this list, that is 82 tokens a second — faster than you read, but a quarter of what the same card does with an 8B model. There is no card on which a 32B model is quick in the way a small one is. If responsiveness matters more than depth, the model choice moves the number far more than the card does.

The case that breaks the ranking

Apple's unified memory is the interesting exception. An M3 Ultra offers 819 GB/s — between an RTX 4090 and an RTX 5080 — but it can be configured with up to 256 GB, shared between the processor and the graphics units. On a 70B model, which no consumer graphics card holds at all, that machine is the one that runs it, at a speed a mid-range GeForce would manage if only it had the memory.

The trade is visible in the table's logic: Apple wins on capacity per euro at these sizes and loses on raw bandwidth against a flagship card. Which matters depends entirely on whether the model you want fits.

What the estimate does not cover

Reading a long prompt is a different job. Prefill processes all the input tokens at once, which is arithmetic-heavy rather than memory-heavy, so it runs roughly an order of magnitude faster and is limited by compute instead. A card that looks slow here can still ingest a long document quickly.

Batching changes the picture too: serving several requests at once reads the weights once for all of them, so tokens per second across the batch rises well above these figures. The numbers above describe one person typing to one model, which is what a desktop actually does.

And the efficiency factor is a single number standing in for many things — cache traffic, attention over the context, kernel quality, whether the runtime is llama.cpp or vLLM, and how hot the card is. Treat the table as a ranking you can trust and an absolute value you should not, and check your own combination in the inference speed estimator. To find out whether a model fits before worrying about how fast it runs, the VRAM calculator answers that first question.