Which AI models run on a NVIDIA RTX 5090 Laptop?
With 24 GB of VRAM, here are the popular models you can run locally (4,096-token context, ~32.0 GB system RAM assumed), ranked by popularity.
See also: Best GPU for running local LLMs.
The NVIDIA RTX 5090 Laptop comes with 24 GB of VRAM. Among the popular GGUF models we track, it can run 34 of them entirely in VRAM — including Qwen3-Coder-30B-A3B-Instruct-GGUF, Qwen-AgentWorld-35B-A3B-GGUF, gpt-oss-20b-GGUF.
With 24 GB you can typically run a 30–35B model at Q5, fully on the GPU. Which quantization is best depends on the exact model and your context length. For a full shortlist, see the best LLM for 24 GB of VRAM.
Larger models such as Qwen2.5-Coder-32B-Instruct-GGUF still run on a NVIDIA RTX 5090 Laptop but require offloading part of the model to system RAM, which lowers speed. Models that exceed both VRAM and RAM are not listed.
34 fit fully in VRAM · 3 run with offload
| Model | Size | Quant. | Quality | Memory | Speed~ | Verdict |
|---|---|---|---|---|---|---|
| unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF | 30.53B | Q5_K_XL | Very good |
21.42 GB
|
19.8 t/s | Fits in VRAM |
| unsloth/Qwen-AgentWorld-35B-A3B-GGUF | 34.66B | Q4_K_XL | Very good |
21.67 GB
|
19.2 t/s | Fits in VRAM |
| unsloth/gpt-oss-20b-GGUF | 20.91B | F16 | Very good |
13.74 GB
|
31.1 t/s | Fits in VRAM |
| janhq/Jan-v3.5-4B-gguf | 4.41B | GGUF | Excellent |
9.59 GB
|
48.6 t/s | Fits in VRAM |
| unsloth/Qwen3-8B-GGUF | 8.19B | BF16 | Excellent |
16.63 GB
|
26.2 t/s | Fits in VRAM |
| hugging-quants/Llama-3.2-1B-Instruct-Q8_0-GGUF | 1.24B | Q8_0 | Excellent |
2.22 GB
|
325.1 t/s | Fits in VRAM |
| bartowski/Meta-Llama-3.1-8B-Instruct-GGUF | 8.03B | Q8_0 | Excellent |
9.23 GB
|
50.3 t/s | Fits in VRAM |
| Qwen/Qwen3-4B-GGUF | 4.02B | Q8_0 | Excellent |
5.35 GB
|
100.3 t/s | Fits in VRAM |
| LiquidAI/LFM2.5-1.2B-Instruct-GGUF | 1.17B | BF16 | Excellent |
3.03 GB
|
183.3 t/s | Fits in VRAM |
| Qwen/Qwen2.5-1.5B-Instruct-GGUF | 1.78B | GGUF | Excellent |
4.23 GB
|
120.6 t/s | Fits in VRAM |
| unsloth/Qwen3-Coder-Next-GGUF | 79.67B | IQ2_XXS | Very low |
22.89 GB
|
18.4 t/s | Fits in VRAM |
| Qwen/Qwen2.5-Coder-7B-Instruct-GGUF | 7.62B | Q8_0 | Excellent |
16.1 GB
|
26.5 t/s | Fits in VRAM |
| Qwen/Qwen2.5-3B-Instruct-GGUF | 3.4B | GGUF | Excellent |
7.27 GB
|
63.2 t/s | Fits in VRAM |
| Qwen/Qwen2.5-0.5B-Instruct-GGUF | 0.63B | GGUF | Excellent |
2.03 GB
|
339.1 t/s | Fits in VRAM |
| ibm-granite/granite-4.1-3b-GGUF | 3.4B | BF16 | Excellent |
7.45 GB
|
63.1 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-0.6B-GGUF | 0.75B | GGUF | Excellent |
2.64 GB
|
284.6 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-14B-GGUF | 14.77B | Q6_K | Excellent |
12.71 GB
|
35.4 t/s | Fits in VRAM |
| bartowski/Llama-3.2-3B-Instruct-GGUF | 3.21B | F16 | Excellent |
7.09 GB
|
66.8 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-1.7B-GGUF | 2.03B | GGUF | Excellent |
5.03 GB
|
105.5 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-32B-GGUF | 32.76B | Q5_K_M | Very good |
23.42 GB
|
18.5 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-30B-A3B-GGUF | 30.53B | Q5_K_M | Very good |
21.41 GB
|
19.8 t/s | Fits in VRAM |
| unsloth/Ornith-1.0-35B-GGUF | 34.66B | Q4_K_XL | Very good |
21.67 GB
|
19.2 t/s | Fits in VRAM |
| bartowski/Qwen2.5-7B-Instruct-GGUF | 7.62B | F16 | Excellent |
15.21 GB
|
28.2 t/s | Fits in VRAM |
| bartowski/gemma-2-2b-it-GGUF | 2.61B | F32 | Excellent |
10.82 GB
|
41.0 t/s | Fits in VRAM |
| bartowski/Phi-3.5-mini-instruct-GGUF | 3.82B | F32 | Excellent |
16.54 GB
|
28.1 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-4B-Instruct-2507-GGUF | 4.02B | GGUF | Excellent |
8.86 GB
|
53.3 t/s | Fits in VRAM |
| lmstudio-community/DeepSeek-R1-0528-Qwen3-8B-GGUF | 8.19B | Q8_0 | Excellent |
9.47 GB
|
49.3 t/s | Fits in VRAM |
| google/gemma-2b | 2.51B | GGUF | Excellent |
10.41 GB
|
42.8 t/s | Fits in VRAM |
| bartowski/Qwen2.5-14B-Instruct-GGUF | 14.77B | Q8_0 | Excellent |
16.17 GB
|
27.4 t/s | Fits in VRAM |
| bartowski/Qwen2.5-32B-Instruct-GGUF | 32.76B | Q5_K_L | Very good |
23.91 GB
|
18.1 t/s | Fits in VRAM |
| bartowski/Kwaipilot_KAT-Coder-V2.5-Dev-GGUF | 34.66B | Q5_K_S | Very good |
23.38 GB
|
17.8 t/s | Fits in VRAM |
| MaziyarPanahi/Meta-Llama-3-8B-Instruct-GGUF | 8.03B | GGUF | Excellent |
16.24 GB
|
26.7 t/s | Fits in VRAM |
| MaziyarPanahi/Mistral-7B-Instruct-v0.3-GGUF | 7.25B | GGUF | Excellent |
14.8 GB
|
29.6 t/s | Fits in VRAM |
| MaziyarPanahi/gemma-3-4b-it-GGUF | 3.88B | GGUF | Excellent |
8.37 GB
|
55.3 t/s | Fits in VRAM |
| Qwen/Qwen2.5-Coder-32B-Instruct-GGUF | 32.76B | Q6_K | Excellent |
51.88 GB
|
1.0 t/s | Offload |
| unsloth/Laguna-S-2.1-GGUF | 117.56B | IQ4_NL | Fair |
55.7 GB
|
0.9 t/s | Offload |
| MaziyarPanahi/Mixtral-8x22B-v0.1-GGUF | 140.62B | IQ3_XS | Fair |
55.9 GB
|
0.9 t/s | Offload |
"Fits in VRAM" = fast, fully on GPU. "Offload" = part on system RAM, slower. Speed is a rough estimate.
Frequently asked questions
How much VRAM does the NVIDIA RTX 5090 Laptop have?
The NVIDIA RTX 5090 Laptop has 24 GB of VRAM, which determines how large a model it can run entirely on the GPU.
What is the best LLM to run on a NVIDIA RTX 5090 Laptop?
Among popular models, unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF runs well on a NVIDIA RTX 5090 Laptop using the Q5_K_XL quantization (about 21.42 GB). With 24 GB you can generally run a 30–35B model at Q5, fully on the GPU. Larger models trade speed for capability via RAM offloading. See the best LLM for 24 GB of VRAM.
Can a NVIDIA RTX 5090 Laptop run a 7–8B model?
Yes. A 7–8B model like Qwen3-8B-GGUF fits entirely in the 24 GB of a NVIDIA RTX 5090 Laptop (BF16).
Can a NVIDIA RTX 5090 Laptop run a 13–14B model?
Yes. A 13–14B model like Qwen3-14B-GGUF fits entirely in the 24 GB of a NVIDIA RTX 5090 Laptop (Q6_K).
Can a NVIDIA RTX 5090 Laptop run a 70B model?
Yes. A 70B model like Qwen3-Coder-Next-GGUF fits entirely in the 24 GB of a NVIDIA RTX 5090 Laptop (IQ2_XXS).