Best LLM for 12 GB of VRAM

LA

By Lefi Abdelmonem

Author · AI Local Check

12 GB of VRAM — an RTX 3060, 4070 or similar — is a real step up from 8 GB: it's the point where 14B models run comfortably at a good quantization, and where you can start squeezing in a 20B+ model at lower bits. Every number below is measured from real GGUF files.

The best models for 12 GB of VRAM

ModelSizeBest quant that fitsMemory
Qwen3 14B14.8BQ5_K_M≈ 11.2 GB
Phi-4 (14B)14.7BQ5_K_M≈ 11.4 GB
Gemma 3 12B11.8BQ6_K≈ 11.2 GB
DeepSeek-Coder-V2-Lite15.7B MoEQ4_K_L≈ 11.8 GB
gpt-oss 20B20.9B MoEQ4_K_XL≈ 11.9 GB

Memory is weights plus KV cache and system margin at a 4,096-token context.

What to pick

The natural home for 12 GB is a 14B model at Q5 — Qwen3 14B and Phi-4 are both strong all-rounders that fit with room for a decent context. If you'd rather have maximum quality at a slightly smaller size, Gemma 3 12B at Q6 is excellent.

12 GB also unlocks your first taste of larger models: gpt-oss 20B, a Mixture-of-Experts model, fits at a 4-bit quant (≈ 12 GB), and DeepSeek-Coder-V2-Lite (16B MoE) is a capable coder at Q4. Both are faster than their size suggests because only a fraction of their experts run per token.

Can 12 GB run a 24B or 30B?

A 24B like Mistral Small only fits 12 GB at an aggressive ~3-bit quant, where quality starts to suffer — usable, but a 16 GB card handles it far better. A 30B is out of reach without heavy RAM offloading.

Tips for 12 GB

  • 14B at Q5 is the sweet spot — the best balance of capability and quality this tier offers.
  • Try an MoE 20B. gpt-oss 20B gives you a bigger model's breadth at 12 GB thanks to sparse activation.
  • Mind the context. A long context can push a Q5 14B over 12 GB; drop to Q4_K_M or shorten the context if it spills.

Check any model against your exact card with the calculator, or compare tiers: 8 GB · 16 GB · 24 GB.