Which AI models run on a AMD Radeon RX 7800 XT?
The AMD Radeon RX 7800 XT has 16 GB of VRAM. Here are the popular AI models it can run locally (4,096-token context, ~32.0 GB system RAM assumed), ranked by popularity.
See also: Best GPU for running local LLMs.
The AMD Radeon RX 7800 XT comes with 16 GB of VRAM. Among the popular GGUF models we track, it can run 34 of them entirely in VRAM — including Qwen3-Coder-30B-A3B-Instruct-GGUF, gpt-oss-20b-GGUF, Qwen-AgentWorld-35B-A3B-GGUF.
With 16 GB you can typically run a 14B at high quality, or a 24–27B at 4-bit. Which quantization is best depends on the exact model and your context length. For a full shortlist, see the best LLM for 16 GB of VRAM.
Larger models such as Laguna-S-2.1-GGUF still run on a AMD Radeon RX 7800 XT but require offloading part of the model to system RAM, which lowers speed. Models that exceed both VRAM and RAM are not listed.
34 fit fully in VRAM · 4 run with offload
| Model | Size | Quant. | Quality | Memory | Speed~ | Verdict |
|---|---|---|---|---|---|---|
| unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF | 30.53B | Q3_K_M | Fair |
14.88 GB
|
29.2 t/s | Fits in VRAM |
| unsloth/gpt-oss-20b-GGUF | 20.91B | F16 | Very good |
13.74 GB
|
31.1 t/s | Fits in VRAM |
| unsloth/Qwen-AgentWorld-35B-A3B-GGUF | 34.66B | IQ3_S | Fair |
14.83 GB
|
28.7 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-4B-GGUF | 4.02B | GGUF | Excellent |
8.86 GB
|
53.3 t/s | Fits in VRAM |
| bartowski/Meta-Llama-3.1-8B-Instruct-GGUF | 8.03B | Q8_0 | Excellent |
9.23 GB
|
50.3 t/s | Fits in VRAM |
| janhq/Jan-v3.5-4B-gguf | 4.41B | GGUF | Excellent |
9.59 GB
|
48.6 t/s | Fits in VRAM |
| Qwen/Qwen3-8B-GGUF | 8.19B | Q8_0 | Excellent |
9.47 GB
|
49.3 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-0.6B-GGUF | 0.75B | GGUF | Excellent |
2.64 GB
|
284.6 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-14B-GGUF | 14.77B | Q6_K | Excellent |
12.71 GB
|
35.4 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-1.7B-GGUF | 2.03B | GGUF | Excellent |
5.03 GB
|
105.5 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-32B-GGUF | 32.76B | Q2_K | Low |
13.3 GB
|
34.8 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-30B-A3B-GGUF | 30.53B | Q3_K_L | Fair |
15.98 GB
|
27.0 t/s | Fits in VRAM |
| bartowski/Qwen2.5-7B-Instruct-GGUF | 7.62B | F16 | Excellent |
15.21 GB
|
28.2 t/s | Fits in VRAM |
| unsloth/Ornith-1.0-35B-GGUF | — | IQ3_S | Excellent |
14.83 GB
|
28.7 t/s | Fits in VRAM |
| Qwen/Qwen2.5-3B-Instruct-GGUF | 3.09B | GGUF | Excellent |
7.27 GB
|
63.2 t/s | Fits in VRAM |
| LiquidAI/LFM2.5-2.6B-GGUF | 2.7B | BF16 | Excellent |
5.89 GB
|
79.5 t/s | Fits in VRAM |
| LiquidAI/LFM2.5-1.2B-Instruct-GGUF | 1.17B | BF16 | Excellent |
3.03 GB
|
183.3 t/s | Fits in VRAM |
| bartowski/Kwaipilot_KAT-Coder-V2.5-Dev-GGUF | 34.66B | Q3_K_M | Fair |
15.99 GB
|
26.5 t/s | Fits in VRAM |
| Qwen/Qwen2.5-1.5B-Instruct-GGUF | 1.54B | GGUF | Excellent |
4.23 GB
|
120.6 t/s | Fits in VRAM |
| Qwen/Qwen2.5-Coder-7B-Instruct-GGUF | 7.62B | GGUF | Excellent |
15.21 GB
|
28.2 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-4B-Instruct-2507-GGUF | 4.02B | GGUF | Excellent |
8.86 GB
|
53.3 t/s | Fits in VRAM |
| MaziyarPanahi/Meta-Llama-3-8B-Instruct-GGUF | 8.03B | Q8_0 | Excellent |
9.23 GB
|
50.3 t/s | Fits in VRAM |
| MaziyarPanahi/Mistral-7B-Instruct-v0.3-GGUF | 7.25B | GGUF | Excellent |
14.8 GB
|
29.6 t/s | Fits in VRAM |
| Qwen/Qwen2.5-0.5B-Instruct-GGUF | 0.49B | GGUF | Excellent |
2.03 GB
|
339.1 t/s | Fits in VRAM |
| MaziyarPanahi/gemma-3-4b-it-GGUF | 4.3B | GGUF | Excellent |
8.38 GB
|
55.3 t/s | Fits in VRAM |
| MaziyarPanahi/Phi-3.5-mini-instruct-GGUF | 3.82B | Q8_0 | Excellent |
6.08 GB
|
105.8 t/s | Fits in VRAM |
| lmstudio-community/Llama-3.2-3B-Instruct-GGUF | 3.21B | Q8_0 | Excellent |
4.29 GB
|
125.5 t/s | Fits in VRAM |
| MaziyarPanahi/Llama-3-8B-Instruct-32k-v0.1-GGUF | 8.03B | Q8_0 | Excellent |
9.23 GB
|
50.3 t/s | Fits in VRAM |
| MaziyarPanahi/Mistral-Small-24B-Instruct-2501-GGUF | 23.57B | Q4_K_M | Good |
14.77 GB
|
30.0 t/s | Fits in VRAM |
| MaziyarPanahi/Yi-1.5-6B-Chat-GGUF | 6.06B | Q6_K | Excellent |
5.68 GB
|
86.3 t/s | Fits in VRAM |
| MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF | 1.24B | GGUF | Excellent |
3.3 GB
|
173.2 t/s | Fits in VRAM |
| MaziyarPanahi/Mistral-Nemo-Instruct-2407-GGUF | 12.25B | Q8_0 | Excellent |
13.55 GB
|
33.0 t/s | Fits in VRAM |
| unsloth/GLM-4.7-Flash-GGUF | 31.22B | Q3_K_M | Fair |
14.62 GB
|
29.4 t/s | Fits in VRAM |
| MaziyarPanahi/WizardLM-2-7B-GGUF | 7.24B | GGUF | Excellent |
14.74 GB
|
29.7 t/s | Fits in VRAM |
| unsloth/Laguna-S-2.1-GGUF | 117.56B | IQ3_S | Low |
46.09 GB
|
1.1 t/s | Offload |
| unsloth/Qwen3-Coder-Next-GGUF | 79.67B | Q4_1 | Very good |
47.79 GB
|
1.1 t/s | Offload |
| MaziyarPanahi/Mixtral-8x22B-v0.1-GGUF | 140.62B | IQ1_M | Very low |
32.16 GB
|
1.6 t/s | Offload |
| MaziyarPanahi/Llama-3.3-70B-Instruct-GGUF | 70.55B | Q5_K_S | Very good |
47.53 GB
|
1.1 t/s | Offload |
"Fits in VRAM" = fast, fully on GPU. "Offload" = part on system RAM, slower. Speed is a rough estimate.
Frequently asked questions
How much VRAM does the AMD Radeon RX 7800 XT have?
The AMD Radeon RX 7800 XT has 16 GB of VRAM, which determines how large a model it can run entirely on the GPU.
What is the best LLM to run on a AMD Radeon RX 7800 XT?
Among popular models, unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF runs well on a AMD Radeon RX 7800 XT using the Q3_K_M quantization (about 14.88 GB). With 16 GB you can generally run a 14B at high quality, or a 24–27B at 4-bit. Larger models trade speed for capability via RAM offloading. See the best LLM for 16 GB of VRAM.
Can a AMD Radeon RX 7800 XT run a 7–8B model?
Yes. A 7–8B model like Meta-Llama-3.1-8B-Instruct-GGUF fits entirely in the 16 GB of a AMD Radeon RX 7800 XT (Q8_0).
Can a AMD Radeon RX 7800 XT run a 13–14B model?
Yes. A 13–14B model like Qwen3-14B-GGUF fits entirely in the 16 GB of a AMD Radeon RX 7800 XT (Q6_K).
Can a AMD Radeon RX 7800 XT run a 70B model?
Only with offloading. A 70B model like Qwen3-Coder-Next-GGUF runs on a AMD Radeon RX 7800 XT by using system RAM in addition to its 16 GB, which is slower.