Best Local AI Models for NVIDIA RTX 5090 (32 GB VRAM)

The NVIDIA RTX 5090 has 32 GB of VRAM. Here are the popular AI models it can run locally (4,096-token context, ~32.0 GB system RAM assumed), ranked by popularity.

Verdict: The ultimate consumer local AI powerhouse: 32 GB of ultra-fast GDDR7 runs 70B models comfortably in 4-bit quantizations.

See also: Best GPU for running local LLMs.

VRAM & Memory
32 GB
GDDR7 512-bit
Bandwidth
1792 GB/s
Blackwell
Fits in VRAM
37 models
Zero offload
TDP / Assumed RAM
575W
RAM: 32.0 GB

The NVIDIA RTX 5090 is built on NVIDIA's Blackwell architecture featuring GDDR7 512-bit delivering 1792 GB/s of raw memory bandwidth. The ultimate consumer local AI powerhouse: 32 GB of ultra-fast GDDR7 runs 70B models comfortably in 4-bit quantizations.

Equipped with 32 GB of dedicated VRAM, the NVIDIA RTX 5090 can run 37 popular open-source models completely in GPU memory without offloading. This includes full-speed execution for weights like Qwen3-Coder-30B-A3B-Instruct-GGUF, LFM2.5-2.6B-GGUF, LFM2.5-8B-A1B-GGUF. For a comprehensive breakdown of compatible model weights, see our guide to the best LLMs for 32 GB VRAM.

With a memory bandwidth of 1792 GB/s, this card can generate tokens at an estimated peak rate of ~298.7 tokens/second on an 8B parameter model (Q4_K_M). Its rated power draw is 575W TDP, so ensure your system's power supply and case ventilation are adequate for sustained local inferencing.

New to this? Read: How much VRAM do you need?

37 fit fully in VRAM

ModelSize Quant.Quality MemorySpeed~ Verdict
unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF 30.53B Q8_0 Excellent
31.43 GB
13.2 t/s Fits in VRAM
LiquidAI/LFM2.5-2.6B-GGUF 2.7B BF16 Excellent
5.89 GB
79.5 t/s Fits in VRAM
LiquidAI/LFM2.5-8B-A1B-GGUF 8.47B BF16 Excellent
16.63 GB
25.3 t/s Fits in VRAM
LiquidAI/LFM2.5-230M-GGUF 0.23B BF16 Excellent
1.28 GB
929.9 t/s Fits in VRAM
unsloth/gpt-oss-20b-GGUF 20.91B F16 Very good
13.74 GB
31.1 t/s Fits in VRAM
Qwen/Qwen3-8B-GGUF 8.19B Q8_0 Excellent
9.47 GB
49.3 t/s Fits in VRAM
unsloth/Ornith-1.0-9B-GGUF — BF16 Excellent
17.61 GB
24.0 t/s Fits in VRAM
bartowski/DeepSeek-Coder-V2-Lite-Instruct-GGUF 15.71B Q8_0_L Excellent
17.39 GB
25.1 t/s Fits in VRAM
unsloth/Qwen3-4B-GGUF 4.02B BF16 Excellent
8.86 GB
53.3 t/s Fits in VRAM
unsloth/Qwen-AgentWorld-35B-A3B-GGUF 34.66B Q6_K_XL Excellent
30.53 GB
13.5 t/s Fits in VRAM
bartowski/Kwaipilot_KAT-Coder-V2.5-Dev-GGUF 34.66B Q6_K_L Excellent
29.1 GB
14.2 t/s Fits in VRAM
Qwen/Qwen3-0.6B-GGUF 0.75B Q8_0 Excellent
1.83 GB
671.7 t/s Fits in VRAM
bartowski/Meta-Llama-3.1-8B-Instruct-GGUF 8.03B F32 Excellent
31.2 GB
13.4 t/s Fits in VRAM
LiquidAI/LFM2.5-1.2B-Instruct-GGUF 1.17B BF16 Excellent
3.03 GB
183.3 t/s Fits in VRAM
Qwen/Qwen2.5-Coder-7B-Instruct-GGUF 7.62B GGUF Excellent
15.21 GB
28.2 t/s Fits in VRAM
janhq/Jan-v3.5-4B-gguf 4.41B GGUF Excellent
9.59 GB
48.6 t/s Fits in VRAM
MaziyarPanahi/Qwen3-14B-GGUF 14.77B GGUF Excellent
28.94 GB
14.5 t/s Fits in VRAM
MaziyarPanahi/Qwen3-1.7B-GGUF 2.03B GGUF Excellent
5.03 GB
105.5 t/s Fits in VRAM
MaziyarPanahi/Qwen3-30B-A3B-GGUF 30.53B Q6_K Excellent
24.54 GB
17.1 t/s Fits in VRAM
MaziyarPanahi/Qwen3-32B-GGUF 32.76B Q6_K Excellent
26.84 GB
16.0 t/s Fits in VRAM
ggml-org/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF 31.58B Q4_0 Good
18.42 GB
22.7 t/s Fits in VRAM
Qwen/Qwen2.5-1.5B-Instruct-GGUF 1.54B GGUF Excellent
4.23 GB
120.6 t/s Fits in VRAM
MaziyarPanahi/Yi-Coder-9B-Chat-GGUF 8.83B GGUF Excellent
17.62 GB
24.3 t/s Fits in VRAM
bartowski/Qwen2.5-32B-Instruct-GGUF 32.76B Q6_K_L Excellent
27.19 GB
15.8 t/s Fits in VRAM
MaziyarPanahi/Qwen3-4B-Instruct-2507-GGUF 4.02B GGUF Excellent
8.86 GB
53.3 t/s Fits in VRAM
Qwen/Qwen2.5-0.5B-Instruct-GGUF 0.49B GGUF Excellent
2.03 GB
339.1 t/s Fits in VRAM
Qwen/Qwen2.5-3B-Instruct-GGUF 3.09B GGUF Excellent
7.27 GB
63.2 t/s Fits in VRAM
bartowski/Qwen2.5-7B-Instruct-GGUF 7.62B F16 Excellent
15.21 GB
28.2 t/s Fits in VRAM
unsloth/Llama-3.2-3B-Instruct-GGUF 3.21B F16 Excellent
7.23 GB
66.8 t/s Fits in VRAM
MaziyarPanahi/Meta-Llama-3-8B-Instruct-GGUF 8.03B GGUF Excellent
16.24 GB
26.7 t/s Fits in VRAM
unsloth/Qwen3-Coder-Next-GGUF 79.67B IQ3_S Low
28.83 GB
14.5 t/s Fits in VRAM
MaziyarPanahi/Yi-Coder-1.5B-Chat-GGUF 1.48B GGUF Excellent
4.3 GB
145.4 t/s Fits in VRAM
MaziyarPanahi/Phi-3.5-mini-instruct-GGUF 3.82B Q8_0 Excellent
6.08 GB
105.8 t/s Fits in VRAM
MaziyarPanahi/Mistral-7B-Instruct-v0.3-GGUF 7.25B GGUF Excellent
14.8 GB
29.6 t/s Fits in VRAM
MaziyarPanahi/Mixtral-8x22B-v0.1-GGUF 140.62B IQ1_S Very low
29.28 GB
14.5 t/s Fits in VRAM
MaziyarPanahi/gemma-3-4b-it-GGUF 4.3B GGUF Excellent
8.38 GB
55.3 t/s Fits in VRAM
MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF 1.24B GGUF Excellent
3.3 GB
173.2 t/s Fits in VRAM

"Fits in VRAM" = fast, fully on GPU. "Offload" = part on system RAM, slower. Speed is a rough estimate.

Frequently asked questions

What is the VRAM and memory bandwidth of the NVIDIA RTX 5090?

The NVIDIA RTX 5090 features 32 GB of VRAM and a memory bandwidth of 1792 GB/s (GDDR7 512-bit). In local language model inference, VRAM determines which model sizes fit on the card, while memory bandwidth dictates how many tokens per second the GPU generates.

What is the best local AI model to run on a NVIDIA RTX 5090?

The best overall model for the NVIDIA RTX 5090 is unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF using the recommended Q8_0 quantization (31.43 GB total memory). With 32 GB of VRAM, this GPU typically runs a 30–32B model at high quality, or a 70B at low bit rates at full GPU speed. Check our guide to the best LLMs for 32 GB VRAM for details.

Can the NVIDIA RTX 5090 run 70B parameter models like Llama 3.3 70B?

Yes. With 32 GB of VRAM, the NVIDIA RTX 5090 can run 70B parameter models quantized to 3-bit or 4-bit (such as IQ3_M or Q4_K_M) with moderate context windows (4k tokens). For larger context windows (16k+), pairing with a second GPU or slight CPU offloading may be required.

What power supply (PSU) do you need for the NVIDIA RTX 5090 when running local AI?

The NVIDIA RTX 5090 has a rated TDP of 575W. While local inference typically draws less power than full 3D rasterization or gaming, continuous token generation can sustain high loads. A quality power supply of at least 875W is strongly recommended.

NVIDIA RTX 5090 Head-to-Head Comparisons

Compare specs, memory bandwidth, and AI model capability against other graphics cards.

Another graphics card

NVIDIA RTX 4090 24 GBNVIDIA RTX 3090 Ti 24 GBNVIDIA RTX 3090 24 GBNVIDIA RTX 5080 16 GBNVIDIA RTX 5070 Ti 16 GBNVIDIA RTX 4080 Super 16 GBNVIDIA RTX 4080 16 GBNVIDIA RTX 4070 Ti Super 16 GBNVIDIA RTX 5060 Ti 16 GB 16 GBNVIDIA RTX 4060 Ti 16 GB 16 GBNVIDIA RTX 5070 12 GBNVIDIA RTX 4070 Ti 12 GBNVIDIA RTX 4070 Super 12 GBNVIDIA RTX 4070 12 GBNVIDIA RTX 3080 Ti 12 GBNVIDIA RTX 3060 12 GB 12 GBNVIDIA RTX 2080 Ti 11 GBNVIDIA RTX 3080 10 GBNVIDIA RTX 5060 8 GBNVIDIA RTX 4060 Ti 8 GB 8 GBNVIDIA RTX 4060 8 GBNVIDIA RTX 3070 Ti 8 GBNVIDIA RTX 3070 8 GBNVIDIA RTX 3060 Ti 8 GBNVIDIA RTX 2080 Super 8 GBNVIDIA RTX 2070 Super 8 GBNVIDIA RTX 2060 Super 8 GBNVIDIA RTX 3050 8 GBNVIDIA RTX 2060 6 GBNVIDIA GTX 1660 Ti 6 GBNVIDIA GTX 1660 Super 6 GBNVIDIA GTX 1660 6 GBNVIDIA GTX 1650 4 GBNVIDIA RTX 5090 Laptop 24 GBNVIDIA RTX 5080 Laptop 16 GBNVIDIA RTX 5070 Ti Laptop 12 GBNVIDIA RTX 5070 Laptop 8 GBNVIDIA RTX 5060 Laptop 8 GBNVIDIA RTX 5050 Laptop 8 GBNVIDIA RTX 4090 Laptop 16 GBNVIDIA RTX 4080 Laptop 12 GBNVIDIA RTX 4070 Laptop 8 GBNVIDIA RTX 4060 Laptop 8 GBNVIDIA RTX 4050 Laptop 6 GBNVIDIA RTX 3080 Ti Laptop 16 GBNVIDIA RTX 3070 Ti Laptop 8 GBNVIDIA RTX 3070 Laptop 8 GBNVIDIA RTX 3060 Laptop 6 GBNVIDIA RTX 3050 Ti Laptop 4 GBNVIDIA RTX 3050 Laptop 4 GBAMD Radeon RX 7900 XTX 24 GBAMD Radeon RX 7900 XT 20 GBAMD Radeon RX 7800 XT 16 GBAMD Radeon RX 7600 XT 16 GBAMD Radeon RX 6950 XT 16 GBAMD Radeon RX 6800 XT 16 GBAMD Radeon RX 6800 16 GBAMD Radeon RX 7700 XT 12 GBAMD Radeon RX 6750 XT 12 GBAMD Radeon RX 6700 XT 12 GBAMD Radeon RX 7600 8 GBAMD Radeon RX 6650 XT 8 GBAMD Radeon RX 6600 8 GBIntel Arc A770 16 GBIntel Arc B580 12 GBIntel Arc A750 8 GB