Which AI models run on a NVIDIA GTX 1660 Super?
The NVIDIA GTX 1660 Super has 6 GB of VRAM. Here are the popular AI models it can run locally (4,096-token context, ~16.0 GB system RAM assumed), ranked by popularity.
See also: Best GPU for running local LLMs.
The NVIDIA GTX 1660 Super comes with 6 GB of VRAM. Among the popular GGUF models we track, it can run 23 of them entirely in VRAM — including Qwen3-4B-GGUF, Meta-Llama-3.1-8B-Instruct-GGUF, Jan-v3.5-4B-gguf.
With 6 GB you can typically run smaller models, typically up to about 3–4B. Which quantization is best depends on the exact model and your context length.
Larger models such as Qwen3-Coder-30B-A3B-Instruct-GGUF still run on a NVIDIA GTX 1660 Super but require offloading part of the model to system RAM, which lowers speed. Models that exceed both VRAM and RAM are not listed.
23 fit fully in VRAM · 12 run with offload
| Model | Size | Quant. | Quality | Memory | Speed~ | Verdict |
|---|---|---|---|---|---|---|
| MaziyarPanahi/Qwen3-4B-GGUF | 4.02B | Q6_K | Excellent |
4.44 GB
|
129.9 t/s | Fits in VRAM |
| bartowski/Meta-Llama-3.1-8B-Instruct-GGUF | 8.03B | Q4_K_M | Good |
5.86 GB
|
87.3 t/s | Fits in VRAM |
| janhq/Jan-v3.5-4B-gguf | 4.41B | Q8_0 | Excellent |
5.73 GB
|
91.5 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-0.6B-GGUF | 0.75B | GGUF | Excellent |
2.64 GB
|
284.6 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-1.7B-GGUF | 2.03B | GGUF | Excellent |
5.03 GB
|
105.5 t/s | Fits in VRAM |
| bartowski/Qwen2.5-7B-Instruct-GGUF | 7.62B | Q5_K_S | Very good |
5.97 GB
|
80.8 t/s | Fits in VRAM |
| Qwen/Qwen2.5-3B-Instruct-GGUF | 3.09B | Q8_0 | Excellent |
4.31 GB
|
118.8 t/s | Fits in VRAM |
| LiquidAI/LFM2.5-2.6B-GGUF | 2.7B | BF16 | Excellent |
5.89 GB
|
79.5 t/s | Fits in VRAM |
| LiquidAI/LFM2.5-1.2B-Instruct-GGUF | 1.17B | BF16 | Excellent |
3.03 GB
|
183.3 t/s | Fits in VRAM |
| Qwen/Qwen2.5-1.5B-Instruct-GGUF | 1.54B | GGUF | Excellent |
4.23 GB
|
120.6 t/s | Fits in VRAM |
| Qwen/Qwen2.5-Coder-7B-Instruct-GGUF | 7.62B | Q5_0 | Very good |
5.97 GB
|
80.8 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-4B-Instruct-2507-GGUF | 4.02B | Q6_K | Excellent |
4.44 GB
|
129.9 t/s | Fits in VRAM |
| MaziyarPanahi/Meta-Llama-3-8B-Instruct-GGUF | 8.03B | Q4_K_M | Good |
5.86 GB
|
87.3 t/s | Fits in VRAM |
| MaziyarPanahi/Mistral-7B-Instruct-v0.3-GGUF | 7.25B | Q5_K_S | Very good |
5.96 GB
|
85.9 t/s | Fits in VRAM |
| Qwen/Qwen2.5-0.5B-Instruct-GGUF | 0.49B | GGUF | Excellent |
2.03 GB
|
339.1 t/s | Fits in VRAM |
| MaziyarPanahi/gemma-3-4b-it-GGUF | 4.3B | Q8_0 | Excellent |
5.0 GB
|
104.0 t/s | Fits in VRAM |
| MaziyarPanahi/Phi-3.5-mini-instruct-GGUF | 3.82B | Q6_K | Excellent |
5.22 GB
|
137.0 t/s | Fits in VRAM |
| lmstudio-community/Llama-3.2-3B-Instruct-GGUF | 3.21B | Q8_0 | Excellent |
4.29 GB
|
125.5 t/s | Fits in VRAM |
| MaziyarPanahi/Llama-3-8B-Instruct-32k-v0.1-GGUF | 8.03B | Q4_K_M | Good |
5.86 GB
|
87.3 t/s | Fits in VRAM |
| MaziyarPanahi/Yi-1.5-6B-Chat-GGUF | 6.06B | Q6_K | Excellent |
5.68 GB
|
86.3 t/s | Fits in VRAM |
| MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF | 1.24B | GGUF | Excellent |
3.3 GB
|
173.2 t/s | Fits in VRAM |
| MaziyarPanahi/Mistral-Nemo-Instruct-2407-GGUF | 12.25B | Q2_K | Low |
5.89 GB
|
89.6 t/s | Fits in VRAM |
| MaziyarPanahi/WizardLM-2-7B-GGUF | 7.24B | Q5_K_S | Very good |
5.91 GB
|
85.9 t/s | Fits in VRAM |
| unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF | 30.53B | Q5_K_XL | Very good |
21.42 GB
|
2.5 t/s | Offload |
| unsloth/gpt-oss-20b-GGUF | 20.91B | F16 | Very good |
13.74 GB
|
3.9 t/s | Offload |
| unsloth/Qwen-AgentWorld-35B-A3B-GGUF | 34.66B | Q4_K_XL | Very good |
21.67 GB
|
2.4 t/s | Offload |
| Qwen/Qwen3-8B-GGUF | 8.19B | Q8_0 | Excellent |
9.47 GB
|
6.2 t/s | Offload |
| MaziyarPanahi/Qwen3-14B-GGUF | 14.77B | Q6_K | Excellent |
12.71 GB
|
4.4 t/s | Offload |
| MaziyarPanahi/Qwen3-32B-GGUF | 32.76B | Q4_K_M | Good |
20.2 GB
|
2.7 t/s | Offload |
| MaziyarPanahi/Qwen3-30B-A3B-GGUF | 30.53B | Q5_K_M | Very good |
21.41 GB
|
2.5 t/s | Offload |
| unsloth/Ornith-1.0-35B-GGUF | — | Q4_K_XL | Excellent |
21.67 GB
|
2.4 t/s | Offload |
| bartowski/Kwaipilot_KAT-Coder-V2.5-Dev-GGUF | 34.66B | Q4_1 | Very good |
21.34 GB
|
2.4 t/s | Offload |
| unsloth/Qwen3-Coder-Next-GGUF | 79.67B | IQ1_M | Very low |
21.39 GB
|
2.5 t/s | Offload |
| MaziyarPanahi/Mistral-Small-24B-Instruct-2501-GGUF | 23.57B | Q6_K | Excellent |
19.44 GB
|
2.8 t/s | Offload |
| unsloth/GLM-4.7-Flash-GGUF | 31.22B | Q5_K_XL | Very good |
21.21 GB
|
2.5 t/s | Offload |
"Fits in VRAM" = fast, fully on GPU. "Offload" = part on system RAM, slower. Speed is a rough estimate.
Frequently asked questions
How much VRAM does the NVIDIA GTX 1660 Super have?
The NVIDIA GTX 1660 Super has 6 GB of VRAM, which determines how large a model it can run entirely on the GPU.
What is the best LLM to run on a NVIDIA GTX 1660 Super?
Among popular models, MaziyarPanahi/Qwen3-4B-GGUF runs well on a NVIDIA GTX 1660 Super using the Q6_K quantization (about 4.44 GB). With 6 GB you can generally run smaller models, typically up to about 3–4B. Larger models trade speed for capability via RAM offloading.
Can a NVIDIA GTX 1660 Super run a 7–8B model?
Yes. A 7–8B model like Meta-Llama-3.1-8B-Instruct-GGUF fits entirely in the 6 GB of a NVIDIA GTX 1660 Super (Q4_K_M).
Can a NVIDIA GTX 1660 Super run a 13–14B model?
Yes. A 13–14B model like Mistral-Nemo-Instruct-2407-GGUF fits entirely in the 6 GB of a NVIDIA GTX 1660 Super (Q2_K).
Can a NVIDIA GTX 1660 Super run a 70B model?
Only with offloading. A 70B model like Qwen3-Coder-Next-GGUF runs on a NVIDIA GTX 1660 Super by using system RAM in addition to its 6 GB, which is slower.