Which hardware is fastest for your use case, and does your model even fit in memory? Memory bandwidth and compute are turned into tokens per second, wait times and image times. Prices and value for money can be shown optionally.
AI Score Metrics ist ein kostenloser Rechner für Hardware, auf der große Sprachmodelle (LLMs) lokal laufen. Er vergleicht NVIDIA- und AMD-Grafikkarten, Apple-Macs mit M-Chips und Mini-PCs mit Ryzen AI Max danach, ob ein Modell in den Speicher passt, wie schnell es Text erzeugt, wie lange der erste Token dauert und wie viel Leistung es pro Euro gibt. GPU-Preise für Deutschland werden täglich automatisch aktualisiert und lassen sich optional einblenden.
Für das gewählte Modell prüft der Rechner zuerst, ob Modellgewichte, KV-Cache für den Kontext und rund 1 GB Laufzeit-Overhead in den nutzbaren Grafikspeicher passen. Passt das Modell, ergibt sich die Generierungsgeschwindigkeit aus der Speicherbandbreite, weil für jeden Token die aktiven Gewichte einmal gelesen werden. Die Prompt-Verarbeitung hängt von der Rechenleistung (TFLOPS) ab. Beides wird je nach Einsatzzweck gewichtet zur Leistung zusammengefasst und durch den Systempreis geteilt; das beste Gerät erhält 100 Punkte.
α ist 0,5 für Agentic Coding und 0,8 für Chat. Bei Grafikkarten zählt zum Systempreis ein einstellbarer Betrag für den restlichen PC.
Das hängt vor allem von der Modellgröße ab. Grafikkarten mit 16 GB VRAM wie die RTX 5060 Ti 16 GB oder RX 9070 reichen für Modelle bis etwa 12 Milliarden Parameter und für gpt-oss-20b. Mit 24–32 GB (RTX 4090, RTX 5090) laufen Modelle der 27–32B-Klasse schnell. Für 70B-Modelle und große MoE-Modelle wie gpt-oss-120b braucht es 48–96 GB: Workstation-Karten, Macs mit viel Unified Memory oder Mini-PCs mit Ryzen AI Max+ 395 und 128 GB.
Bei der üblichen Quantisierung Q4_K_M braucht ein Modell rund 0,6 GB pro Milliarde Parameter, dazu kommen der KV-Cache für den Kontext und etwa 1 GB Overhead. Ein 27B-Modell belegt so rund 17 GB plus Kontext, ein 70B-Modell gut 42 GB plus bis zu 10 GB KV-Cache bei 32k Tokens. Der Rechner zeigt die Belegung je Gerät in der Speicherauslastung.
Die Speicherbandbreite (GB/s) gibt an, wie schnell die GPU Daten aus ihrem Speicher lesen kann. Beim Erzeugen jedes Tokens müssen alle aktiven Modellgewichte einmal gelesen werden, deshalb bestimmt die Bandbreite fast direkt die Tokens pro Sekunde. Eine RTX 4090 mit 1.008 GB/s schafft so mit einem 8B-Modell rund 120 Tokens/s.
Macs mit M-Chips teilen sich den Arbeitsspeicher zwischen CPU und GPU und können dadurch sehr große Modelle laden, die auf keine einzelne Consumer-Grafikkarte passen. Grafikkarten haben dafür meist mehr Bandbreite und deutlich mehr Rechenleistung pro Euro, sie sind bei Modellen, die in den VRAM passen, oft schneller – vor allem bei der Verarbeitung langer Prompts.
Mixture-of-Experts-Modelle wie gpt-oss-120b oder Qwen3-Coder 30B-A3B nutzen pro Token nur einen kleinen Teil ihrer Parameter. Sie erzeugen Text dadurch viel schneller als gleich große dichte Modelle, müssen aber trotzdem vollständig in den Speicher passen.
Die Werte sind eine Modellrechnung, keine Messung. Abweichungen von ±20 % zu echten Benchmarks sind je nach Backend (CUDA, MLX, llama.cpp), Treiber und Quantisierung normal. Der Abgleich mit der LocalScore-Datenbank zeigt bei der Generierung eine Abweichung von etwa ±3 %.
AI Score Metrics is a free calculator for hardware that runs large language models (LLMs) locally. It compares NVIDIA and AMD graphics cards, Apple Macs with M-series chips and Ryzen AI Max mini PCs by whether a model fits in memory, how fast it generates text, how long the first token takes and how much performance you get per euro. GPU prices for Germany are updated automatically every day and can be shown optionally.
For the selected model, the calculator first checks whether the model weights, the KV cache for the context and about 1 GB of runtime overhead fit in usable graphics memory. If the model fits, generation speed follows from memory bandwidth, because the active weights are read once for every token. Prompt processing depends on compute (TFLOPS). Both are weighted by use case into a performance figure and divided by the system price; the best device gets 100 points.
α is 0.5 for agentic coding and 0.8 for chat. For graphics cards, the system price includes an adjustable amount for the rest of the PC.
Mostly it depends on model size. Graphics cards with 16 GB of VRAM such as the RTX 5060 Ti 16 GB or RX 9070 handle models up to about 12 billion parameters and gpt-oss-20b. With 24–32 GB (RTX 4090, RTX 5090), models in the 27–32B class run fast. 70B models and large MoE models such as gpt-oss-120b need 48–96 GB: workstation cards, Macs with plenty of unified memory, or Ryzen AI Max+ 395 mini PCs with 128 GB.
At the common Q4_K_M quantization, a model needs about 0.6 GB per billion parameters, plus the KV cache for the context and about 1 GB of overhead. A 27B model takes about 17 GB plus context; a 70B model a little over 42 GB plus up to 10 GB of KV cache at 32k tokens. The calculator shows per-device usage in the memory chart.
Memory bandwidth (GB/s) is how fast the GPU can read data from its memory. Every generated token requires reading all active model weights once, so bandwidth almost directly sets tokens per second. An RTX 4090 with 1,008 GB/s reaches about 120 tokens/s with an 8B model.
Macs with M-series chips share RAM between CPU and GPU, so they can load very large models that do not fit on any single consumer graphics card. Graphics cards usually offer more bandwidth and far more compute per euro, so they are often faster for models that fit in VRAM – especially when processing long prompts.
Mixture-of-experts models such as gpt-oss-120b or Qwen3-Coder 30B-A3B use only a small share of their parameters per token. That makes them generate text much faster than dense models of the same size, but they still have to fit in memory completely.
The numbers are a model estimate, not a measurement. Deviations of ±20% from real benchmarks are normal, depending on backend (CUDA, MLX, llama.cpp), drivers and quantization. A check against the LocalScore database shows generation within about ±3%.