Best Local AI Models for AMD Radeon RX 7900 XT (20 GB VRAM)
The AMD Radeon RX 7900 XT has 20 GB of VRAM. Here are the popular AI models it can run locally (4,096-token context, ~32.0 GB system RAM assumed), ranked by popularity.
See also: Best GPU for running local LLMs.
The AMD Radeon RX 7900 XT is built on AMD's RDNA 3 architecture featuring GDDR6 320-bit delivering 800 GB/s of raw memory bandwidth. Generous 20 GB VRAM buffer: comfortably runs 14B at uncompressed Q8 and 32B at Q4.
Equipped with 20 GB of dedicated VRAM, the AMD Radeon RX 7900 XT can run 36 popular open-source models completely in GPU memory without offloading. This includes full-speed execution for weights like Qwen3-Coder-30B-A3B-Instruct-GGUF, LFM2.5-2.6B-GGUF, LFM2.5-8B-A1B-GGUF. For a comprehensive breakdown of compatible model weights, see our guide to the best LLMs for 16 GB VRAM.
With a memory bandwidth of 800 GB/s, this card can generate tokens at an estimated peak rate of ~133.3 tokens/second on an 8B parameter model (Q4_K_M). Its rated power draw is 315W TDP, so ensure your system's power supply and case ventilation are adequate for sustained local inferencing.
36 fit fully in VRAM · 1 run with offload
| Model | Size | Quant. | Quality | Memory | Speed~ | Verdict |
|---|---|---|---|---|---|---|
| unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF | 30.53B | Q4_1 | Very good |
19.05 GB
|
22.4 t/s | Fits in VRAM |
| LiquidAI/LFM2.5-2.6B-GGUF | 2.7B | BF16 | Excellent |
5.89 GB
|
79.5 t/s | Fits in VRAM |
| LiquidAI/LFM2.5-8B-A1B-GGUF | 8.47B | BF16 | Excellent |
16.63 GB
|
25.3 t/s | Fits in VRAM |
| LiquidAI/LFM2.5-230M-GGUF | 0.23B | BF16 | Excellent |
1.28 GB
|
929.9 t/s | Fits in VRAM |
| unsloth/gpt-oss-20b-GGUF | 20.91B | F16 | Very good |
13.74 GB
|
31.1 t/s | Fits in VRAM |
| Qwen/Qwen3-8B-GGUF | 8.19B | Q8_0 | Excellent |
9.47 GB
|
49.3 t/s | Fits in VRAM |
| unsloth/Ornith-1.0-9B-GGUF | — | BF16 | Excellent |
17.61 GB
|
24.0 t/s | Fits in VRAM |
| bartowski/DeepSeek-Coder-V2-Lite-Instruct-GGUF | 15.71B | Q8_0_L | Excellent |
17.39 GB
|
25.1 t/s | Fits in VRAM |
| unsloth/Qwen3-4B-GGUF | 4.02B | BF16 | Excellent |
8.86 GB
|
53.3 t/s | Fits in VRAM |
| unsloth/Qwen-AgentWorld-35B-A3B-GGUF | 34.66B | IQ4_NL | Fair |
17.75 GB
|
23.7 t/s | Fits in VRAM |
| bartowski/Kwaipilot_KAT-Coder-V2.5-Dev-GGUF | 34.66B | Q4_0 | Good |
19.45 GB
|
21.5 t/s | Fits in VRAM |
| Qwen/Qwen3-0.6B-GGUF | 0.75B | Q8_0 | Excellent |
1.83 GB
|
671.7 t/s | Fits in VRAM |
| bartowski/Meta-Llama-3.1-8B-Instruct-GGUF | 8.03B | Q8_0 | Excellent |
9.23 GB
|
50.3 t/s | Fits in VRAM |
| LiquidAI/LFM2.5-1.2B-Instruct-GGUF | 1.17B | BF16 | Excellent |
3.03 GB
|
183.3 t/s | Fits in VRAM |
| Qwen/Qwen2.5-Coder-7B-Instruct-GGUF | 7.62B | GGUF | Excellent |
15.21 GB
|
28.2 t/s | Fits in VRAM |
| janhq/Jan-v3.5-4B-gguf | 4.41B | GGUF | Excellent |
9.59 GB
|
48.6 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-14B-GGUF | 14.77B | Q6_K | Excellent |
12.71 GB
|
35.4 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-1.7B-GGUF | 2.03B | GGUF | Excellent |
5.03 GB
|
105.5 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-30B-A3B-GGUF | 30.53B | Q4_K_M | Good |
18.46 GB
|
23.1 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-32B-GGUF | 32.76B | Q3_K_L | Good |
17.94 GB
|
24.8 t/s | Fits in VRAM |
| ggml-org/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF | 31.58B | Q4_0 | Good |
18.42 GB
|
22.7 t/s | Fits in VRAM |
| Qwen/Qwen2.5-1.5B-Instruct-GGUF | 1.54B | GGUF | Excellent |
4.23 GB
|
120.6 t/s | Fits in VRAM |
| MaziyarPanahi/Yi-Coder-9B-Chat-GGUF | 8.83B | GGUF | Excellent |
17.62 GB
|
24.3 t/s | Fits in VRAM |
| bartowski/Qwen2.5-32B-Instruct-GGUF | 32.76B | Q4_K_S | Good |
19.29 GB
|
22.9 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-4B-Instruct-2507-GGUF | 4.02B | GGUF | Excellent |
8.86 GB
|
53.3 t/s | Fits in VRAM |
| Qwen/Qwen2.5-0.5B-Instruct-GGUF | 0.49B | GGUF | Excellent |
2.03 GB
|
339.1 t/s | Fits in VRAM |
| Qwen/Qwen2.5-3B-Instruct-GGUF | 3.09B | GGUF | Excellent |
7.27 GB
|
63.2 t/s | Fits in VRAM |
| bartowski/Qwen2.5-7B-Instruct-GGUF | 7.62B | F16 | Excellent |
15.21 GB
|
28.2 t/s | Fits in VRAM |
| unsloth/Llama-3.2-3B-Instruct-GGUF | 3.21B | F16 | Excellent |
7.23 GB
|
66.8 t/s | Fits in VRAM |
| MaziyarPanahi/Meta-Llama-3-8B-Instruct-GGUF | 8.03B | GGUF | Excellent |
16.24 GB
|
26.7 t/s | Fits in VRAM |
| unsloth/Qwen3-Coder-Next-GGUF | 79.67B | TQ1_0 | Very low |
18.82 GB
|
22.7 t/s | Fits in VRAM |
| MaziyarPanahi/Yi-Coder-1.5B-Chat-GGUF | 1.48B | GGUF | Excellent |
4.3 GB
|
145.4 t/s | Fits in VRAM |
| MaziyarPanahi/Phi-3.5-mini-instruct-GGUF | 3.82B | Q8_0 | Excellent |
6.08 GB
|
105.8 t/s | Fits in VRAM |
| MaziyarPanahi/Mistral-7B-Instruct-v0.3-GGUF | 7.25B | GGUF | Excellent |
14.8 GB
|
29.6 t/s | Fits in VRAM |
| MaziyarPanahi/gemma-3-4b-it-GGUF | 4.3B | GGUF | Excellent |
8.38 GB
|
55.3 t/s | Fits in VRAM |
| MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF | 1.24B | GGUF | Excellent |
3.3 GB
|
173.2 t/s | Fits in VRAM |
| MaziyarPanahi/Mixtral-8x22B-v0.1-GGUF | 140.62B | Q2_K | Low |
50.2 GB
|
1.0 t/s | Offload |
"Fits in VRAM" = fast, fully on GPU. "Offload" = part on system RAM, slower. Speed is a rough estimate.
Frequently asked questions
What is the VRAM and memory bandwidth of the AMD Radeon RX 7900 XT?
The AMD Radeon RX 7900 XT features 20 GB of VRAM and a memory bandwidth of 800 GB/s (GDDR6 320-bit). In local language model inference, VRAM determines which model sizes fit on the card, while memory bandwidth dictates how many tokens per second the GPU generates.
What is the best local AI model to run on a AMD Radeon RX 7900 XT?
The best overall model for the AMD Radeon RX 7900 XT is unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF using the recommended Q4_1 quantization (19.05 GB total memory). With 20 GB of VRAM, this GPU typically runs a 14B at high quality, or a 24–27B at 4-bit at full GPU speed. Check our guide to the best LLMs for 16 GB VRAM for details.
Can the AMD Radeon RX 7900 XT run 14B and 32B models?
Yes. A 16 GB VRAM buffer allows you to run 14B models (like Qwen 2.5 14B or DeepSeek-R1 14B) in high-fidelity Q8_0 quantizations, or 32B models (like Qwen 2.5 32B) in Q3_K_M or Q4_K_M quantizations fully on the GPU.
Does the AMD Radeon RX 7900 XT work with Ollama and LM Studio?
Yes. Under Linux, the AMD Radeon RX 7900 XT supports native ROCm acceleration in Ollama and llama.cpp. On Windows, Ollama and LM Studio utilize the Vulkan or DirectML backends to leverage the card's full 20 GB VRAM pool.
AMD Radeon RX 7900 XT Head-to-Head Comparisons
Compare specs, memory bandwidth, and AI model capability against other graphics cards.
AMD Radeon RX 7900 XT vs NVIDIA RTX 4090
Side-by-side local AI performance and supported model comparison.
AMD Radeon RX 7900 XT vs NVIDIA RTX 3090 Ti
Side-by-side local AI performance and supported model comparison.
AMD Radeon RX 7900 XT vs NVIDIA RTX 5080
Side-by-side local AI performance and supported model comparison.