Best LLM for 12 GB of VRAM
Author · AI Local Check
12 GB of VRAM — an RTX 3060, 4070 or similar — is a real step up from 8 GB: it's the point where 14B models run comfortably at a good quantization, and where you can start squeezing in a 20B+ model at lower bits. Every number below is measured from real GGUF files.
The best models for 12 GB of VRAM
| Model | Size | Best quant that fits | Memory |
|---|---|---|---|
| Qwen3 14B | 14.8B | Q5_K_M | ≈ 11.2 GB |
| Phi-4 (14B) | 14.7B | Q5_K_M | ≈ 11.4 GB |
| Gemma 3 12B | 11.8B | Q6_K | ≈ 11.2 GB |
| DeepSeek-Coder-V2-Lite | 15.7B MoE | Q4_K_L | ≈ 11.8 GB |
| gpt-oss 20B | 20.9B MoE | Q4_K_XL | ≈ 11.9 GB |
Memory is weights plus KV cache and system margin at a 4,096-token context.
What to pick
The natural home for 12 GB is a 14B model at Q5 — Qwen3 14B and Phi-4 are both strong all-rounders that fit with room for a decent context. If you'd rather have maximum quality at a slightly smaller size, Gemma 3 12B at Q6 is excellent.
12 GB also unlocks your first taste of larger models: gpt-oss 20B, a Mixture-of-Experts model, fits at a 4-bit quant (≈ 12 GB), and DeepSeek-Coder-V2-Lite (16B MoE) is a capable coder at Q4. Both are faster than their size suggests because only a fraction of their experts run per token.
Can 12 GB run a 24B or 30B?
A 24B like Mistral Small only fits 12 GB at an aggressive ~3-bit quant, where quality starts to suffer — usable, but a 16 GB card handles it far better. A 30B is out of reach without heavy RAM offloading.
Tips for 12 GB
- 14B at Q5 is the sweet spot — the best balance of capability and quality this tier offers.
- Try an MoE 20B. gpt-oss 20B gives you a bigger model's breadth at 12 GB thanks to sparse activation.
- Mind the context. A long context can push a Q5 14B over 12 GB; drop to Q4_K_M or shorten the context if it spills.
Check any model against your exact card with the calculator, or compare tiers: 8 GB · 16 GB · 24 GB.