GPUs suited for local LLMs
This list includes GPUs suited for running local LLMs (Gemma, Qwen, Llama, DeepSeek, and others). As a rule of thumb, (quantized) 8B models need 8 GB of VRAM, 22B models need 16 GB, and 32B models need 24 GB. At the same VRAM capacity, newer generations also run faster.
Listing criteria: models with 16 GB of VRAM or more (the mainstream tier), ordered by VRAM capacity.
What to Check Beyond the Score
- Unlike AI image generation, local LLMs run fairly well on AMD and Intel GPUs too (Ollama, LM Studio, and others support them).
- At the same VRAM capacity, newer generations with faster VRAM generate quicker.
- Tiers reflect the model landscape as of 2026 and may shift as the technology evolves.