Best Local AI Models for NVIDIA RTX 5090 (32 GB VRAM)
The NVIDIA RTX 5090 has 32 GB of VRAM. Here are the popular AI models it can run locally (4,096-token context, ~32.0 GB system RAM assumed), ranked by popularity.
See also: Best GPU for running local LLMs.
The NVIDIA RTX 5090 is built on NVIDIA's Blackwell architecture featuring GDDR7 512-bit delivering 1792 GB/s of raw memory bandwidth. The ultimate consumer local AI powerhouse: 32 GB of ultra-fast GDDR7 runs 70B models comfortably in 4-bit quantizations.
Equipped with 32 GB of dedicated VRAM, the NVIDIA RTX 5090 can run 37 popular open-source models completely in GPU memory without offloading. This includes full-speed execution for weights like Qwen3-Coder-30B-A3B-Instruct-GGUF, LFM2.5-2.6B-GGUF, LFM2.5-8B-A1B-GGUF. For a comprehensive breakdown of compatible model weights, see our guide to the best LLMs for 32 GB VRAM.
With a memory bandwidth of 1792 GB/s, this card can generate tokens at an estimated peak rate of ~298.7 tokens/second on an 8B parameter model (Q4_K_M). Its rated power draw is 575W TDP, so ensure your system's power supply and case ventilation are adequate for sustained local inferencing.
37 fit fully in VRAM
| Model | Size | Quant. | Quality | Memory | Speed~ | Verdict |
|---|---|---|---|---|---|---|
| unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF | 30.53B | Q8_0 | Excellent |
31.43 GB
|
13.2 t/s | Fits in VRAM |
| LiquidAI/LFM2.5-2.6B-GGUF | 2.7B | BF16 | Excellent |
5.89 GB
|
79.5 t/s | Fits in VRAM |
| LiquidAI/LFM2.5-8B-A1B-GGUF | 8.47B | BF16 | Excellent |
16.63 GB
|
25.3 t/s | Fits in VRAM |
| LiquidAI/LFM2.5-230M-GGUF | 0.23B | BF16 | Excellent |
1.28 GB
|
929.9 t/s | Fits in VRAM |
| unsloth/gpt-oss-20b-GGUF | 20.91B | F16 | Very good |
13.74 GB
|
31.1 t/s | Fits in VRAM |
| Qwen/Qwen3-8B-GGUF | 8.19B | Q8_0 | Excellent |
9.47 GB
|
49.3 t/s | Fits in VRAM |
| unsloth/Ornith-1.0-9B-GGUF | — | BF16 | Excellent |
17.61 GB
|
24.0 t/s | Fits in VRAM |
| bartowski/DeepSeek-Coder-V2-Lite-Instruct-GGUF | 15.71B | Q8_0_L | Excellent |
17.39 GB
|
25.1 t/s | Fits in VRAM |
| unsloth/Qwen3-4B-GGUF | 4.02B | BF16 | Excellent |
8.86 GB
|
53.3 t/s | Fits in VRAM |
| unsloth/Qwen-AgentWorld-35B-A3B-GGUF | 34.66B | Q6_K_XL | Excellent |
30.53 GB
|
13.5 t/s | Fits in VRAM |
| bartowski/Kwaipilot_KAT-Coder-V2.5-Dev-GGUF | 34.66B | Q6_K_L | Excellent |
29.1 GB
|
14.2 t/s | Fits in VRAM |
| Qwen/Qwen3-0.6B-GGUF | 0.75B | Q8_0 | Excellent |
1.83 GB
|
671.7 t/s | Fits in VRAM |
| bartowski/Meta-Llama-3.1-8B-Instruct-GGUF | 8.03B | F32 | Excellent |
31.2 GB
|
13.4 t/s | Fits in VRAM |
| LiquidAI/LFM2.5-1.2B-Instruct-GGUF | 1.17B | BF16 | Excellent |
3.03 GB
|
183.3 t/s | Fits in VRAM |
| Qwen/Qwen2.5-Coder-7B-Instruct-GGUF | 7.62B | GGUF | Excellent |
15.21 GB
|
28.2 t/s | Fits in VRAM |
| janhq/Jan-v3.5-4B-gguf | 4.41B | GGUF | Excellent |
9.59 GB
|
48.6 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-14B-GGUF | 14.77B | GGUF | Excellent |
28.94 GB
|
14.5 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-1.7B-GGUF | 2.03B | GGUF | Excellent |
5.03 GB
|
105.5 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-30B-A3B-GGUF | 30.53B | Q6_K | Excellent |
24.54 GB
|
17.1 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-32B-GGUF | 32.76B | Q6_K | Excellent |
26.84 GB
|
16.0 t/s | Fits in VRAM |
| ggml-org/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF | 31.58B | Q4_0 | Good |
18.42 GB
|
22.7 t/s | Fits in VRAM |
| Qwen/Qwen2.5-1.5B-Instruct-GGUF | 1.54B | GGUF | Excellent |
4.23 GB
|
120.6 t/s | Fits in VRAM |
| MaziyarPanahi/Yi-Coder-9B-Chat-GGUF | 8.83B | GGUF | Excellent |
17.62 GB
|
24.3 t/s | Fits in VRAM |
| bartowski/Qwen2.5-32B-Instruct-GGUF | 32.76B | Q6_K_L | Excellent |
27.19 GB
|
15.8 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-4B-Instruct-2507-GGUF | 4.02B | GGUF | Excellent |
8.86 GB
|
53.3 t/s | Fits in VRAM |
| Qwen/Qwen2.5-0.5B-Instruct-GGUF | 0.49B | GGUF | Excellent |
2.03 GB
|
339.1 t/s | Fits in VRAM |
| Qwen/Qwen2.5-3B-Instruct-GGUF | 3.09B | GGUF | Excellent |
7.27 GB
|
63.2 t/s | Fits in VRAM |
| bartowski/Qwen2.5-7B-Instruct-GGUF | 7.62B | F16 | Excellent |
15.21 GB
|
28.2 t/s | Fits in VRAM |
| unsloth/Llama-3.2-3B-Instruct-GGUF | 3.21B | F16 | Excellent |
7.23 GB
|
66.8 t/s | Fits in VRAM |
| MaziyarPanahi/Meta-Llama-3-8B-Instruct-GGUF | 8.03B | GGUF | Excellent |
16.24 GB
|
26.7 t/s | Fits in VRAM |
| unsloth/Qwen3-Coder-Next-GGUF | 79.67B | IQ3_S | Low |
28.83 GB
|
14.5 t/s | Fits in VRAM |
| MaziyarPanahi/Yi-Coder-1.5B-Chat-GGUF | 1.48B | GGUF | Excellent |
4.3 GB
|
145.4 t/s | Fits in VRAM |
| MaziyarPanahi/Phi-3.5-mini-instruct-GGUF | 3.82B | Q8_0 | Excellent |
6.08 GB
|
105.8 t/s | Fits in VRAM |
| MaziyarPanahi/Mistral-7B-Instruct-v0.3-GGUF | 7.25B | GGUF | Excellent |
14.8 GB
|
29.6 t/s | Fits in VRAM |
| MaziyarPanahi/Mixtral-8x22B-v0.1-GGUF | 140.62B | IQ1_S | Very low |
29.28 GB
|
14.5 t/s | Fits in VRAM |
| MaziyarPanahi/gemma-3-4b-it-GGUF | 4.3B | GGUF | Excellent |
8.38 GB
|
55.3 t/s | Fits in VRAM |
| MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF | 1.24B | GGUF | Excellent |
3.3 GB
|
173.2 t/s | Fits in VRAM |
"Fits in VRAM" = fast, fully on GPU. "Offload" = part on system RAM, slower. Speed is a rough estimate.
Frequently asked questions
What is the VRAM and memory bandwidth of the NVIDIA RTX 5090?
The NVIDIA RTX 5090 features 32 GB of VRAM and a memory bandwidth of 1792 GB/s (GDDR7 512-bit). In local language model inference, VRAM determines which model sizes fit on the card, while memory bandwidth dictates how many tokens per second the GPU generates.
What is the best local AI model to run on a NVIDIA RTX 5090?
The best overall model for the NVIDIA RTX 5090 is unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF using the recommended Q8_0 quantization (31.43 GB total memory). With 32 GB of VRAM, this GPU typically runs a 30–32B model at high quality, or a 70B at low bit rates at full GPU speed. Check our guide to the best LLMs for 32 GB VRAM for details.
Can the NVIDIA RTX 5090 run 70B parameter models like Llama 3.3 70B?
Yes. With 32 GB of VRAM, the NVIDIA RTX 5090 can run 70B parameter models quantized to 3-bit or 4-bit (such as IQ3_M or Q4_K_M) with moderate context windows (4k tokens). For larger context windows (16k+), pairing with a second GPU or slight CPU offloading may be required.
What power supply (PSU) do you need for the NVIDIA RTX 5090 when running local AI?
The NVIDIA RTX 5090 has a rated TDP of 575W. While local inference typically draws less power than full 3D rasterization or gaming, continuous token generation can sustain high loads. A quality power supply of at least 875W is strongly recommended.
NVIDIA RTX 5090 Head-to-Head Comparisons
Compare specs, memory bandwidth, and AI model capability against other graphics cards.
RTX 5090 32GB vs RTX 4090 24GB
32 GB GDDR7 vs 24 GB: what new models become possible.
NVIDIA RTX 5090 vs NVIDIA RTX 4090
Side-by-side local AI performance and supported model comparison.
NVIDIA RTX 5090 vs NVIDIA RTX 3090 Ti
Side-by-side local AI performance and supported model comparison.