Which AI models run on a NVIDIA RTX 3080?
The NVIDIA RTX 3080 has 10 GB of VRAM. Here are the popular AI models it can run locally (4,096-token context, ~16.0 GB system RAM assumed), ranked by popularity.
See also: Best GPU for running local LLMs.
The NVIDIA RTX 3080 comes with 10 GB of VRAM. Among the popular GGUF models we track, it can run 29 of them entirely in VRAM — including Qwen3-Coder-30B-A3B-Instruct-GGUF, Qwen3-4B-GGUF, Meta-Llama-3.1-8B-Instruct-GGUF.
With 10 GB you can typically run a 7–8B model at Q6, entirely in VRAM. Which quantization is best depends on the exact model and your context length. For a full shortlist, see the best LLM for 8 GB of VRAM.
Larger models such as gpt-oss-20b-GGUF still run on a NVIDIA RTX 3080 but require offloading part of the model to system RAM, which lowers speed. Models that exceed both VRAM and RAM are not listed.
29 fit fully in VRAM · 6 run with offload
| Model | Size | Quant. | Quality | Memory | Speed~ | Verdict |
|---|---|---|---|---|---|---|
| unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF | 30.53B | IQ1_S | Very low |
9.48 GB
|
48.2 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-4B-GGUF | 4.02B | GGUF | Excellent |
8.86 GB
|
53.3 t/s | Fits in VRAM |
| bartowski/Meta-Llama-3.1-8B-Instruct-GGUF | 8.03B | Q8_0 | Excellent |
9.23 GB
|
50.3 t/s | Fits in VRAM |
| janhq/Jan-v3.5-4B-gguf | 4.41B | GGUF | Excellent |
9.59 GB
|
48.6 t/s | Fits in VRAM |
| Qwen/Qwen3-8B-GGUF | 8.19B | Q8_0 | Excellent |
9.47 GB
|
49.3 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-0.6B-GGUF | 0.75B | GGUF | Excellent |
2.64 GB
|
284.6 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-14B-GGUF | 14.77B | Q4_K_M | Good |
9.81 GB
|
47.7 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-1.7B-GGUF | 2.03B | GGUF | Excellent |
5.03 GB
|
105.5 t/s | Fits in VRAM |
| bartowski/Qwen2.5-7B-Instruct-GGUF | 7.62B | Q8_0 | Excellent |
8.56 GB
|
53.0 t/s | Fits in VRAM |
| Qwen/Qwen2.5-3B-Instruct-GGUF | 3.09B | GGUF | Excellent |
7.27 GB
|
63.2 t/s | Fits in VRAM |
| LiquidAI/LFM2.5-2.6B-GGUF | 2.7B | BF16 | Excellent |
5.89 GB
|
79.5 t/s | Fits in VRAM |
| LiquidAI/LFM2.5-1.2B-Instruct-GGUF | 1.17B | BF16 | Excellent |
3.03 GB
|
183.3 t/s | Fits in VRAM |
| bartowski/Kwaipilot_KAT-Coder-V2.5-Dev-GGUF | 34.66B | IQ2_XXS | Very low |
9.99 GB
|
43.9 t/s | Fits in VRAM |
| Qwen/Qwen2.5-1.5B-Instruct-GGUF | 1.54B | GGUF | Excellent |
4.23 GB
|
120.6 t/s | Fits in VRAM |
| Qwen/Qwen2.5-Coder-7B-Instruct-GGUF | 7.62B | Q8_0 | Excellent |
8.56 GB
|
53.0 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-4B-Instruct-2507-GGUF | 4.02B | GGUF | Excellent |
8.86 GB
|
53.3 t/s | Fits in VRAM |
| MaziyarPanahi/Meta-Llama-3-8B-Instruct-GGUF | 8.03B | Q8_0 | Excellent |
9.23 GB
|
50.3 t/s | Fits in VRAM |
| MaziyarPanahi/Mistral-7B-Instruct-v0.3-GGUF | 7.25B | Q8_0 | Excellent |
8.47 GB
|
55.8 t/s | Fits in VRAM |
| Qwen/Qwen2.5-0.5B-Instruct-GGUF | 0.49B | GGUF | Excellent |
2.03 GB
|
339.1 t/s | Fits in VRAM |
| MaziyarPanahi/gemma-3-4b-it-GGUF | 4.3B | GGUF | Excellent |
8.38 GB
|
55.3 t/s | Fits in VRAM |
| MaziyarPanahi/Phi-3.5-mini-instruct-GGUF | 3.82B | Q8_0 | Excellent |
6.08 GB
|
105.8 t/s | Fits in VRAM |
| lmstudio-community/Llama-3.2-3B-Instruct-GGUF | 3.21B | Q8_0 | Excellent |
4.29 GB
|
125.5 t/s | Fits in VRAM |
| MaziyarPanahi/Llama-3-8B-Instruct-32k-v0.1-GGUF | 8.03B | Q8_0 | Excellent |
9.23 GB
|
50.3 t/s | Fits in VRAM |
| MaziyarPanahi/Mistral-Small-24B-Instruct-2501-GGUF | 23.57B | Q2_K | Low |
9.7 GB
|
48.3 t/s | Fits in VRAM |
| MaziyarPanahi/Yi-1.5-6B-Chat-GGUF | 6.06B | Q6_K | Excellent |
5.68 GB
|
86.3 t/s | Fits in VRAM |
| MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF | 1.24B | GGUF | Excellent |
3.3 GB
|
173.2 t/s | Fits in VRAM |
| MaziyarPanahi/Mistral-Nemo-Instruct-2407-GGUF | 12.25B | Q5_K_M | Very good |
9.55 GB
|
49.2 t/s | Fits in VRAM |
| unsloth/GLM-4.7-Flash-GGUF | 31.22B | IQ1_S | Very low |
9.62 GB
|
46.4 t/s | Fits in VRAM |
| MaziyarPanahi/WizardLM-2-7B-GGUF | 7.24B | Q8_0 | Excellent |
8.42 GB
|
55.8 t/s | Fits in VRAM |
| unsloth/gpt-oss-20b-GGUF | 20.91B | F16 | Very good |
13.74 GB
|
3.9 t/s | Offload |
| unsloth/Qwen-AgentWorld-35B-A3B-GGUF | 34.66B | Q5_K_XL | Very good |
25.58 GB
|
2.0 t/s | Offload |
| MaziyarPanahi/Qwen3-32B-GGUF | 32.76B | Q5_K_M | Very good |
23.42 GB
|
2.3 t/s | Offload |
| MaziyarPanahi/Qwen3-30B-A3B-GGUF | 30.53B | Q6_K | Excellent |
24.54 GB
|
2.1 t/s | Offload |
| unsloth/Ornith-1.0-35B-GGUF | — | Q5_K_XL | Excellent |
25.58 GB
|
2.0 t/s | Offload |
| unsloth/Qwen3-Coder-Next-GGUF | 79.67B | IQ2_M | Very low |
24.42 GB
|
2.2 t/s | Offload |
"Fits in VRAM" = fast, fully on GPU. "Offload" = part on system RAM, slower. Speed is a rough estimate.
Frequently asked questions
How much VRAM does the NVIDIA RTX 3080 have?
The NVIDIA RTX 3080 has 10 GB of VRAM, which determines how large a model it can run entirely on the GPU.
What is the best LLM to run on a NVIDIA RTX 3080?
Among popular models, unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF runs well on a NVIDIA RTX 3080 using the IQ1_S quantization (about 9.48 GB). With 10 GB you can generally run a 7–8B model at Q6, entirely in VRAM. Larger models trade speed for capability via RAM offloading. See the best LLM for 8 GB of VRAM.
Can a NVIDIA RTX 3080 run a 7–8B model?
Yes. A 7–8B model like Meta-Llama-3.1-8B-Instruct-GGUF fits entirely in the 10 GB of a NVIDIA RTX 3080 (Q8_0).
Can a NVIDIA RTX 3080 run a 13–14B model?
Yes. A 13–14B model like Qwen3-14B-GGUF fits entirely in the 10 GB of a NVIDIA RTX 3080 (Q4_K_M).
Can a NVIDIA RTX 3080 run a 70B model?
Only with offloading. A 70B model like Qwen3-Coder-Next-GGUF runs on a NVIDIA RTX 3080 by using system RAM in addition to its 10 GB, which is slower.