Best Local AI Models for NVIDIA RTX 2080 Ti (11 GB VRAM)
The NVIDIA RTX 2080 Ti has 11 GB of VRAM. Here are the popular AI models it can run locally (4,096-token context, ~16.0 GB system RAM assumed), ranked by popularity.
See also: Best GPU for running local LLMs.
The NVIDIA RTX 2080 Ti is built on NVIDIA's Turing architecture featuring GDDR6 352-bit delivering 616 GB/s of raw memory bandwidth. Classic 11 GB titan: high memory bandwidth that still outpaces many modern budget cards for 8B models.
Equipped with 11 GB of dedicated VRAM, the NVIDIA RTX 2080 Ti can run 30 popular open-source models completely in GPU memory without offloading. This includes full-speed execution for weights like Qwen3-Coder-30B-A3B-Instruct-GGUF, LFM2.5-2.6B-GGUF, LFM2.5-8B-A1B-GGUF. For a comprehensive breakdown of compatible model weights, see our guide to the best LLMs for 8 GB VRAM.
With a memory bandwidth of 616 GB/s, this card can generate tokens at an estimated peak rate of ~102.7 tokens/second on an 8B parameter model (Q4_K_M). Its rated power draw is 250W TDP, so ensure your system's power supply and case ventilation are adequate for sustained local inferencing.
30 fit fully in VRAM · 6 run with offload
| Model | Size | Quant. | Quality | Memory | Speed~ | Verdict |
|---|---|---|---|---|---|---|
| unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF | 30.53B | IQ2_XXS | Low |
10.8 GB
|
41.6 t/s | Fits in VRAM |
| LiquidAI/LFM2.5-2.6B-GGUF | 2.7B | BF16 | Excellent |
5.89 GB
|
79.5 t/s | Fits in VRAM |
| LiquidAI/LFM2.5-8B-A1B-GGUF | 8.47B | Q8_0 | Excellent |
9.24 GB
|
47.7 t/s | Fits in VRAM |
| LiquidAI/LFM2.5-230M-GGUF | 0.23B | BF16 | Excellent |
1.28 GB
|
929.9 t/s | Fits in VRAM |
| Qwen/Qwen3-8B-GGUF | 8.19B | Q8_0 | Excellent |
9.47 GB
|
49.3 t/s | Fits in VRAM |
| unsloth/Ornith-1.0-9B-GGUF | — | Q8_0 | Excellent |
9.8 GB
|
45.1 t/s | Fits in VRAM |
| bartowski/DeepSeek-Coder-V2-Lite-Instruct-GGUF | 15.71B | Q4_K_S | Good |
10.35 GB
|
45.1 t/s | Fits in VRAM |
| unsloth/Qwen3-4B-GGUF | 4.02B | BF16 | Excellent |
8.86 GB
|
53.3 t/s | Fits in VRAM |
| bartowski/Kwaipilot_KAT-Coder-V2.5-Dev-GGUF | 34.66B | IQ2_XS | Very low |
10.93 GB
|
39.8 t/s | Fits in VRAM |
| Qwen/Qwen3-0.6B-GGUF | 0.75B | Q8_0 | Excellent |
1.83 GB
|
671.7 t/s | Fits in VRAM |
| bartowski/Meta-Llama-3.1-8B-Instruct-GGUF | 8.03B | Q8_0 | Excellent |
9.23 GB
|
50.3 t/s | Fits in VRAM |
| LiquidAI/LFM2.5-1.2B-Instruct-GGUF | 1.17B | BF16 | Excellent |
3.03 GB
|
183.3 t/s | Fits in VRAM |
| Qwen/Qwen2.5-Coder-7B-Instruct-GGUF | 7.62B | Q8_0 | Excellent |
8.56 GB
|
53.0 t/s | Fits in VRAM |
| janhq/Jan-v3.5-4B-gguf | 4.41B | GGUF | Excellent |
9.59 GB
|
48.6 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-14B-GGUF | 14.77B | Q4_K_M | Good |
9.81 GB
|
47.7 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-1.7B-GGUF | 2.03B | GGUF | Excellent |
5.03 GB
|
105.5 t/s | Fits in VRAM |
| Qwen/Qwen2.5-1.5B-Instruct-GGUF | 1.54B | GGUF | Excellent |
4.23 GB
|
120.6 t/s | Fits in VRAM |
| MaziyarPanahi/Yi-Coder-9B-Chat-GGUF | 8.83B | Q5_K_M | Very good |
7.0 GB
|
68.6 t/s | Fits in VRAM |
| bartowski/Qwen2.5-32B-Instruct-GGUF | 32.76B | IQ2_XXS | Very low |
10.21 GB
|
47.6 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-4B-Instruct-2507-GGUF | 4.02B | GGUF | Excellent |
8.86 GB
|
53.3 t/s | Fits in VRAM |
| Qwen/Qwen2.5-0.5B-Instruct-GGUF | 0.49B | GGUF | Excellent |
2.03 GB
|
339.1 t/s | Fits in VRAM |
| Qwen/Qwen2.5-3B-Instruct-GGUF | 3.09B | GGUF | Excellent |
7.27 GB
|
63.2 t/s | Fits in VRAM |
| bartowski/Qwen2.5-7B-Instruct-GGUF | 7.62B | Q8_0 | Excellent |
8.56 GB
|
53.0 t/s | Fits in VRAM |
| unsloth/Llama-3.2-3B-Instruct-GGUF | 3.21B | F16 | Excellent |
7.23 GB
|
66.8 t/s | Fits in VRAM |
| MaziyarPanahi/Meta-Llama-3-8B-Instruct-GGUF | 8.03B | Q8_0 | Excellent |
9.23 GB
|
50.3 t/s | Fits in VRAM |
| MaziyarPanahi/Yi-Coder-1.5B-Chat-GGUF | 1.48B | GGUF | Excellent |
4.3 GB
|
145.4 t/s | Fits in VRAM |
| MaziyarPanahi/Phi-3.5-mini-instruct-GGUF | 3.82B | Q8_0 | Excellent |
6.08 GB
|
105.8 t/s | Fits in VRAM |
| MaziyarPanahi/Mistral-7B-Instruct-v0.3-GGUF | 7.25B | Q8_0 | Excellent |
8.47 GB
|
55.8 t/s | Fits in VRAM |
| MaziyarPanahi/gemma-3-4b-it-GGUF | 4.3B | GGUF | Excellent |
8.38 GB
|
55.3 t/s | Fits in VRAM |
| MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF | 1.24B | GGUF | Excellent |
3.3 GB
|
173.2 t/s | Fits in VRAM |
| unsloth/gpt-oss-20b-GGUF | 20.91B | F16 | Very good |
13.74 GB
|
3.9 t/s | Offload |
| unsloth/Qwen-AgentWorld-35B-A3B-GGUF | 34.66B | Q5_K_XL | Very good |
25.58 GB
|
2.0 t/s | Offload |
| MaziyarPanahi/Qwen3-30B-A3B-GGUF | 30.53B | Q6_K | Excellent |
24.54 GB
|
2.1 t/s | Offload |
| MaziyarPanahi/Qwen3-32B-GGUF | 32.76B | Q6_K | Excellent |
26.84 GB
|
2.0 t/s | Offload |
| ggml-org/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF | 31.58B | Q4_0 | Good |
18.42 GB
|
2.8 t/s | Offload |
| unsloth/Qwen3-Coder-Next-GGUF | 79.67B | Q2_K_XL | Low |
26.1 GB
|
2.0 t/s | Offload |
"Fits in VRAM" = fast, fully on GPU. "Offload" = part on system RAM, slower. Speed is a rough estimate.
Frequently asked questions
What is the VRAM and memory bandwidth of the NVIDIA RTX 2080 Ti?
The NVIDIA RTX 2080 Ti features 11 GB of VRAM and a memory bandwidth of 616 GB/s (GDDR6 352-bit). In local language model inference, VRAM determines which model sizes fit on the card, while memory bandwidth dictates how many tokens per second the GPU generates.
What is the best local AI model to run on a NVIDIA RTX 2080 Ti?
The best overall model for the NVIDIA RTX 2080 Ti is unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF using the recommended IQ2_XXS quantization (10.8 GB total memory). With 11 GB of VRAM, this GPU typically runs a 7–8B model at Q6, entirely in VRAM at full GPU speed. Check our guide to the best LLMs for 8 GB VRAM for details.
What models can you run on an 8 GB GPU like the NVIDIA RTX 2080 Ti?
With 8 GB of VRAM, the NVIDIA RTX 2080 Ti comfortably runs 7B and 8B models (such as Llama 3.1 8B, Mistral 7B, or Qwen 2.5 7B) using Q4_K_M or Q5_K_M quantizations (~5.5 to 6.8 GB). Running 14B models requires offloading memory to system RAM, which lowers token generation speed.
NVIDIA RTX 2080 Ti Head-to-Head Comparisons
Compare specs, memory bandwidth, and AI model capability against other graphics cards.
NVIDIA RTX 2080 Ti vs NVIDIA RTX 5080
Side-by-side local AI performance and supported model comparison.
NVIDIA RTX 2080 Ti vs NVIDIA RTX 5070 Ti
Side-by-side local AI performance and supported model comparison.
NVIDIA RTX 2080 Ti vs NVIDIA RTX 4080 Super
Side-by-side local AI performance and supported model comparison.