AI hardware score · local LLM inference

AI Score MetricsWhich GPU or Mac for local LLMs?

Which hardware is fastest for your use case, and does your model even fit in memory? Memory bandwidth and compute are turned into tokens per second, wait times and image times. Prices and value for money can be shown optionally.

What is AI Score Metrics?

AI Score Metrics is a free calculator for hardware that runs large language models (LLMs) locally. It compares NVIDIA and AMD graphics cards, Apple Macs with M-series chips and Ryzen AI Max mini PCs by whether a model fits in memory, how fast it generates text, how long the first token takes and how much performance you get per euro. GPU prices for Germany are updated automatically every day and can be shown optionally.

How is the score calculated?

For the selected model, the calculator first checks whether the model weights, the KV cache for the context and about 1 GB of runtime overhead fit in usable graphics memory. If the model fits, generation speed follows from memory bandwidth, because the active weights are read once for every token. Prompt processing depends on compute (TFLOPS). Both are weighted by use case into a performance figure and divided by the system price; the best device gets 100 points.

Perf = fits · Generation^α · Prompt^(1−α) Score = Perf / system price (best device = 100)

α is 0.5 for agentic coding and 0.8 for chat. For graphics cards, the system price includes an adjustable amount for the rest of the PC.

Frequently asked questions

Which GPU is best for local LLMs?

Mostly it depends on model size. Graphics cards with 16 GB of VRAM such as the RTX 5060 Ti 16 GB or RX 9070 handle models up to about 12 billion parameters and gpt-oss-20b. With 24–32 GB (RTX 4090, RTX 5090), models in the 27–32B class run fast. 70B models and large MoE models such as gpt-oss-120b need 48–96 GB: workstation cards, Macs with plenty of unified memory, or Ryzen AI Max+ 395 mini PCs with 128 GB.

How much VRAM do I need for an LLM?

At the common Q4_K_M quantization, a model needs about 0.6 GB per billion parameters, plus the KV cache for the context and about 1 GB of overhead. A 27B model takes about 17 GB plus context; a 70B model a little over 42 GB plus up to 10 GB of KV cache at 32k tokens. The calculator shows per-device usage in the memory chart.

What is memory bandwidth and why does it matter?

Memory bandwidth (GB/s) is how fast the GPU can read data from its memory. Every generated token requires reading all active model weights once, so bandwidth almost directly sets tokens per second. An RTX 4090 with 1,008 GB/s reaches about 120 tokens/s with an 8B model.

Mac or graphics card for local AI?

Macs with M-series chips share RAM between CPU and GPU, so they can load very large models that do not fit on any single consumer graphics card. Graphics cards usually offer more bandwidth and far more compute per euro, so they are often faster for models that fit in VRAM – especially when processing long prompts.

What are MoE models?

Mixture-of-experts models such as gpt-oss-120b or Qwen3-Coder 30B-A3B use only a small share of their parameters per token. That makes them generate text much faster than dense models of the same size, but they still have to fit in memory completely.

How accurate are the numbers?

The numbers are a model estimate, not a measurement. Deviations of ±20% from real benchmarks are normal, depending on backend (CUDA, MLX, llama.cpp), drivers and quantization. A check against the LocalScore database shows generation within about ±3%.