Reference

LLM tokens/sec Benchmark Reference

Real, single-stream text-generation speeds for running LLMs locally — across NVIDIA GPUs, Apple Silicon, and DGX Spark. Every number links to the public source it came from. Where no verified benchmark exists, we say so instead of guessing.

Last reviewed 2026-08-23 · Next review 2026-11 · Reviewed quarterly

How to read this: figures are text generation (tok/s) at batch size 1 — the speed you feel when a model replies. A range (e.g. 82–124) means the source measured it across context depths; a single value is a single reported figure. Prompt processing (pp) is how fast the model ingests your input. Numbers are only comparable within the same model × quantization × backend — a CUDA number and a Metal number are different machines running different code.

Llama 7B · Q4_0

HardwareGenerationPromptBackendConfidenceSource
Mac Studio (M3 Ultra, 80-core GPU)92.14 tok/s1471.24 tok/sllama.cpp (Metal)Maintainer-measuredllama.cpp Discussion #4167
Mac (M4 Max, 40-core GPU)Check price$1,999 – $5,99983.06 tok/s885.68 tok/sllama.cpp (Metal)Maintainer-measuredllama.cpp Discussion #4167
Strix Halo (Ryzen AI Max+ 395, 128GB)Range across 5 backends.Check price$2,199 – $3,64950.59–55.73 tok/s1545.36 tok/sllama.cpp (ROCm/Vulkan)Community-measuredkyuz0 Strix Halo benchmark grid
Mac (M4 Pro, 20-core GPU)Check price$1,599 — discontinued50.74 tok/s439.78 tok/sllama.cpp (Metal)Maintainer-measuredllama.cpp Discussion #4167

Llama 3 8B · Q4_K_M

HardwareGenerationPromptBackendConfidenceSource
RTX 4090 (24GB)Check price$1,599 – $1,999127.74 tok/s—llama.cpp (CUDA)Community-measuredXiongjieDai/GPU-Benchmarks-on-LLM-Inference
RTX 3090 (24GB)Check price$699 – $999111.74 tok/s—llama.cpp (CUDA)Community-measuredXiongjieDai/GPU-Benchmarks-on-LLM-Inference
Mac Studio (M2 Ultra, 76-core GPU)76.28 tok/s—llama.cpp (Metal)Community-measuredXiongjieDai/GPU-Benchmarks-on-LLM-Inference
Mac (M3 Max, 40-core GPU)50.74 tok/s—llama.cpp (Metal)Community-measuredXiongjieDai/GPU-Benchmarks-on-LLM-Inference
Mac (M1 Max, 32-core GPU)34.49 tok/s—llama.cpp (Metal)Community-measuredXiongjieDai/GPU-Benchmarks-on-LLM-Inference

Llama 3 70B · Q4_K_M

HardwareGenerationPromptBackendConfidenceSource
Mac Studio (M2 Ultra, 76-core GPU)70B does not fit in a 24GB consumer GPU; unified memory is the enabler here.12.13 tok/s—llama.cpp (Metal)Community-measuredXiongjieDai/GPU-Benchmarks-on-LLM-Inference
Mac (M3 Max, 40-core GPU)7.53 tok/s—llama.cpp (Metal)Community-measuredXiongjieDai/GPU-Benchmarks-on-LLM-Inference
Mac (M1 Max, 32-core GPU)4.09 tok/s—llama.cpp (Metal)Community-measuredXiongjieDai/GPU-Benchmarks-on-LLM-Inference

Qwen3 30B-A3B (MoE) · Q4_K_XL

HardwareGenerationPromptBackendConfidenceSource
RTX 5090 (32GB)234.3 tok/s at 4K context, falling to 110.7 at 32K.Check price$1,999 – $2,199110.65–234.3 tok/s—llama.cpp (CUDA)Independent-measuredhardware-corner.net RTX 5090 LLM benchmarks

Qwen3 8B · Q4_K_XL

HardwareGenerationPromptBackendConfidenceSource
RTX 5090 (32GB)185.9 tok/s at 4K context, falling to 111.9 at 32K.Check price$1,999 – $2,199111.91–185.91 tok/s—llama.cpp (CUDA)Independent-measuredhardware-corner.net RTX 5090 LLM benchmarks

Qwen3 14B · Q4_K_XL

HardwareGenerationPromptBackendConfidenceSource
RTX 5090 (32GB)123.8 tok/s at 4K context, falling to 82.4 at 32K.Check price$1,999 – $2,19982.35–123.79 tok/s—llama.cpp (CUDA)Independent-measuredhardware-corner.net RTX 5090 LLM benchmarks

Qwen3 32B · Q4_K_XL

HardwareGenerationPromptBackendConfidenceSource
RTX 5090 (32GB)61.4 tok/s at 4K context, falling to 43.8 at 32K.Check price$1,999 – $2,19943.82–61.38 tok/s—llama.cpp (CUDA)Independent-measuredhardware-corner.net RTX 5090 LLM benchmarks

gpt-oss 20B · MXFP4

HardwareGenerationPromptBackendConfidenceSource
Strix Halo (Ryzen AI Max+ 395, 128GB)Range across 5 backends.Check price$2,199 – $3,64972.68–79.78 tok/s1812.57 tok/sllama.cpp (ROCm/Vulkan)Community-measuredkyuz0 Strix Halo benchmark grid
DGX Spark (GB10)Check price$4,699 – $7,99960.85 tok/s2009 tok/sllama.cpp (CUDA)Maintainer-measuredllama.cpp Discussion #16578

gpt-oss 120B · MXFP4

HardwareGenerationPromptBackendConfidenceSource
DGX Spark (GB10)60.6 tok/s at short context, falling to 51.5 at 8K depth.Check price$4,699 – $7,99951.54–60.57 tok/s1956 tok/sllama.cpp (CUDA)Maintainer-measuredllama.cpp Discussion #16578
Strix Halo (Ryzen AI Max+ 395, 128GB)Framework Desktop; same chip as GMKtec EVO-X2 / Beelink GTR9 Pro. Range across 5 backends.Check price$2,199 – $3,64951.02–56.61 tok/s719.91 tok/sllama.cpp (ROCm/Vulkan)Community-measuredkyuz0 Strix Halo benchmark grid

Qwen3 Coder 30B-A3B · Q8_0

HardwareGenerationPromptBackendConfidenceSource
DGX Spark (GB10)Check price$4,699 – $7,99944.26 tok/s1654 tok/sllama.cpp (CUDA)Maintainer-measuredllama.cpp Discussion #16578

Llama 3.1 8B · Q4_K_M

HardwareGenerationPromptBackendConfidenceSource
RTX 4090 (24GB)Check price$1,599 – $1,99991–120 tok/s6697 tok/sLocalScore (llamafile)Community-measuredLocalScore leaderboard
RTX 3090 (24GB)Check price$699 – $99995.7 tok/s3536 tok/sLocalScore (llamafile)Community-measuredLocalScore leaderboard
RTX 5090 (32GB)llamafile trails llama.cpp CUDA on Blackwell — compare only within this table.Check price$1,999 – $2,19966.3 tok/s6297 tok/sLocalScore (llamafile)Community-measuredLocalScore leaderboard
RTX 5070 Ti (16GB)Check price$1,049 – $1,29934.7–65.5 tok/s—LocalScore (llamafile)Community-measuredLocalScore leaderboard
Mac Studio (M3 Ultra, 80-core GPU)62.7–63.3 tok/s1109 tok/sLocalScore (llamafile)Community-measuredLocalScore leaderboard
Mac (M4 Max, 40-core GPU)Check price$1,999 – $5,99947.9–55.1 tok/s663 tok/sLocalScore (llamafile)Community-measuredLocalScore leaderboard
DGX Spark (GB10)Check price$4,699 – $7,99933.6–34.7 tok/s2010 tok/sLocalScore (llamafile)Community-measuredLocalScore leaderboard
Mac mini (M4 Pro, 20-core GPU)Check price$1,599 — discontinued32.5–32.9 tok/s361 tok/sLocalScore (llamafile)Community-measuredLocalScore leaderboard

Known gaps — no number we'd stand behind

These segments have no verified single-stream benchmark we can cite. We list them explicitly rather than fill the cell with a guess.

SegmentStatusWhy
Intel Arc B580Data exists, not reproducible as a single figureCommunity Vulkan-backend runs exist but are sparse and not yet standardized enough to quote. See the llama.cpp Vulkan thread. Source
A100 / H100 / MI250X (datacenter)Only batched-throughput data — not comparableServing benchmarks (vLLM) measure batched throughput, not single-stream tok/s — a different metric than everything else on this page.
CPU-only mini PCs (Beelink, GMKtec, etc.)No verified public benchmarkLocalScore has scattered CPU entries, but no standardized single-stream tok/s coverage for these specific SKUs. Source
RTX Spark / N1X (expected fall 2026)No verified public benchmarkUnreleased. No public benchmarks exist — any number you see quoted elsewhere is a guess.

Need to know what will fit before you worry about speed? Check memory requirements. Not sure which of these fits your workload? Use the GPU Advisor to match a model and budget to hardware, or browse the full hardware catalog.

Disclosure: Some links on this page are affiliate links. We may earn a commission if you make a purchase — at no extra cost to you. Benchmark source links above are not affiliate links.