The reason spec sheets mislead in this category is that the marketing numbers and the limiting numbers are different numbers. Four mechanisms explain almost every surprising result, and knowing them makes the comparison table above readable without any benchmarks at all.
Quantisation: why a 70B model is not 140GB
Model weights ship at 16-bit precision, so a 7B-parameter model is around 14GB and a 70B model around 140GB — beyond any laptop. Quantisation stores each weight in fewer bits: the common Q4_K_M format averages roughly 4.5 bits, which works out to about 0.6GB per billion parameters. That is the arithmetic behind every figure in the comparison table: 8B becomes ~5GB, 14B ~9GB, 32B ~20GB, 70B ~42GB. Quality loss at Q4 is small and well documented; below Q3 it stops being small.
Why generation speed is a bandwidth problem
Producing one token requires reading every weight in the model once. A 20GB model on a GPU with 896 GB/s of bandwidth therefore has a theoretical ceiling near 45 tokens per second, and no amount of extra compute raises it. This is why an RTX 5090 laptop generates several times faster than a Strix Halo machine holding the same model, and why the base M5 is noticeably slower than an M5 Max at identical capacity. Prompt processing is the exception — it is compute-bound, so raw GPU power does show up when you paste in a long document.
The KV cache, and why context costs memory
Beyond the weights, a model keeps a key-value cache for every token in the conversation so it does not recompute the whole history each step. That cache grows linearly with context length and can add several gigabytes at 32k tokens on a mid-sized model. It is the most commonly forgotten line in the budget: a model that loads with 1GB to spare will run out of memory partway through a long session, which is why our estimates leave headroom rather than filling the pool.
NPU TOPS, and what the number does not mean
Every 2026 laptop advertises 40–50 TOPS of NPU throughput, and it has almost nothing to do with running a language model. NPUs are built for sustained low-power inference on small, fixed, heavily optimised models — background blur, live translation, Windows Recall. They have no large memory pool of their own and the toolchains to target them are narrow. When you run Ollama or LM Studio, the work goes to the GPU or the CPU. Treat TOPS as a battery-life feature, not a capability.
CUDA, MLX and ROCm — the stack tax
Hardware is only half the purchase. CUDA is the default target of essentially all research code, so a new technique usually works on NVIDIA on day one and elsewhere some weeks later. MLX, Apple’s framework, is fast and pleasant for inference and handles LoRA fine-tuning well, but heavier training remains awkward. ROCm covers mainstream inference on AMD and is improving rapidly, with Windows support still behind Linux. If you only serve models, all three are fine; if you train, the calculus is much narrower.