Best Local AI Models for NVIDIA GTX 1660 Super (6 GB VRAM)
The NVIDIA GTX 1660 Super has 6 GB of VRAM. Here are the popular AI models it can run locally (4,096-token context, ~16.0 GB system RAM assumed), ranked by popularity.
See also: Best GPU for running local LLMs.
The NVIDIA GTX 1660 Super is built on NVIDIA's Desktop architecture featuring GDDR6 delivering 192 GB/s of raw memory bandwidth. NVIDIA GTX 1660 Super with 6 GB VRAM for local language model execution.
Equipped with 6 GB of dedicated VRAM, the NVIDIA GTX 1660 Super can run 24 popular open-source models completely in GPU memory without offloading. This includes full-speed execution for weights like LFM2.5-2.6B-GGUF, LFM2.5-8B-A1B-GGUF, LFM2.5-230M-GGUF.
With a memory bandwidth of 192 GB/s, this card can generate tokens at an estimated peak rate of ~32.0 tokens/second on an 8B parameter model (Q4_K_M). Its rated power draw is 150W TDP, so ensure your system's power supply and case ventilation are adequate for sustained local inferencing.
24 fit fully in VRAM · 12 run with offload
| Model | Size | Quant. | Quality | Memory | Speed~ | Verdict |
|---|---|---|---|---|---|---|
| LiquidAI/LFM2.5-2.6B-GGUF | 2.7B | BF16 | Excellent |
5.89 GB
|
79.5 t/s | Fits in VRAM |
| LiquidAI/LFM2.5-8B-A1B-GGUF | 8.47B | Q4_K_M | Good |
5.65 GB
|
83.3 t/s | Fits in VRAM |
| LiquidAI/LFM2.5-230M-GGUF | 0.23B | BF16 | Excellent |
1.28 GB
|
929.9 t/s | Fits in VRAM |
| unsloth/Ornith-1.0-9B-GGUF | — | Q4_K_S | Excellent |
5.98 GB
|
79.2 t/s | Fits in VRAM |
| unsloth/Qwen3-4B-GGUF | 4.02B | Q8_0 | Excellent |
5.35 GB
|
100.3 t/s | Fits in VRAM |
| Qwen/Qwen3-0.6B-GGUF | 0.75B | Q8_0 | Excellent |
1.83 GB
|
671.7 t/s | Fits in VRAM |
| bartowski/Meta-Llama-3.1-8B-Instruct-GGUF | 8.03B | Q4_K_M | Good |
5.86 GB
|
87.3 t/s | Fits in VRAM |
| LiquidAI/LFM2.5-1.2B-Instruct-GGUF | 1.17B | BF16 | Excellent |
3.03 GB
|
183.3 t/s | Fits in VRAM |
| Qwen/Qwen2.5-Coder-7B-Instruct-GGUF | 7.62B | Q5_0 | Very good |
5.97 GB
|
80.8 t/s | Fits in VRAM |
| janhq/Jan-v3.5-4B-gguf | 4.41B | Q8_0 | Excellent |
5.73 GB
|
91.5 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-1.7B-GGUF | 2.03B | GGUF | Excellent |
5.03 GB
|
105.5 t/s | Fits in VRAM |
| Qwen/Qwen2.5-1.5B-Instruct-GGUF | 1.54B | GGUF | Excellent |
4.23 GB
|
120.6 t/s | Fits in VRAM |
| MaziyarPanahi/Yi-Coder-9B-Chat-GGUF | 8.83B | Q4_K_S | Good |
5.9 GB
|
84.7 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-4B-Instruct-2507-GGUF | 4.02B | Q6_K | Excellent |
4.44 GB
|
129.9 t/s | Fits in VRAM |
| Qwen/Qwen2.5-0.5B-Instruct-GGUF | 0.49B | GGUF | Excellent |
2.03 GB
|
339.1 t/s | Fits in VRAM |
| Qwen/Qwen2.5-3B-Instruct-GGUF | 3.09B | Q8_0 | Excellent |
4.31 GB
|
118.8 t/s | Fits in VRAM |
| bartowski/Qwen2.5-7B-Instruct-GGUF | 7.62B | Q5_K_S | Very good |
5.97 GB
|
80.8 t/s | Fits in VRAM |
| unsloth/Llama-3.2-3B-Instruct-GGUF | 3.21B | Q8_K_XL | Excellent |
5.15 GB
|
102.2 t/s | Fits in VRAM |
| MaziyarPanahi/Meta-Llama-3-8B-Instruct-GGUF | 8.03B | Q4_K_M | Good |
5.86 GB
|
87.3 t/s | Fits in VRAM |
| MaziyarPanahi/Yi-Coder-1.5B-Chat-GGUF | 1.48B | GGUF | Excellent |
4.3 GB
|
145.4 t/s | Fits in VRAM |
| MaziyarPanahi/Phi-3.5-mini-instruct-GGUF | 3.82B | Q6_K | Excellent |
5.22 GB
|
137.0 t/s | Fits in VRAM |
| MaziyarPanahi/Mistral-7B-Instruct-v0.3-GGUF | 7.25B | Q5_K_S | Very good |
5.96 GB
|
85.9 t/s | Fits in VRAM |
| MaziyarPanahi/gemma-3-4b-it-GGUF | 4.3B | Q8_0 | Excellent |
5.0 GB
|
104.0 t/s | Fits in VRAM |
| MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF | 1.24B | GGUF | Excellent |
3.3 GB
|
173.2 t/s | Fits in VRAM |
| unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF | 30.53B | Q5_K_XL | Very good |
21.42 GB
|
2.5 t/s | Offload |
| unsloth/gpt-oss-20b-GGUF | 20.91B | F16 | Very good |
13.74 GB
|
3.9 t/s | Offload |
| Qwen/Qwen3-8B-GGUF | 8.19B | Q8_0 | Excellent |
9.47 GB
|
6.2 t/s | Offload |
| bartowski/DeepSeek-Coder-V2-Lite-Instruct-GGUF | 15.71B | Q8_0_L | Excellent |
17.39 GB
|
3.1 t/s | Offload |
| unsloth/Qwen-AgentWorld-35B-A3B-GGUF | 34.66B | Q4_K_XL | Very good |
21.67 GB
|
2.4 t/s | Offload |
| bartowski/Kwaipilot_KAT-Coder-V2.5-Dev-GGUF | 34.66B | Q4_1 | Very good |
21.34 GB
|
2.4 t/s | Offload |
| MaziyarPanahi/Qwen3-14B-GGUF | 14.77B | Q6_K | Excellent |
12.71 GB
|
4.4 t/s | Offload |
| MaziyarPanahi/Qwen3-30B-A3B-GGUF | 30.53B | Q5_K_M | Very good |
21.41 GB
|
2.5 t/s | Offload |
| MaziyarPanahi/Qwen3-32B-GGUF | 32.76B | Q4_K_M | Good |
20.2 GB
|
2.7 t/s | Offload |
| ggml-org/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF | 31.58B | Q4_0 | Good |
18.42 GB
|
2.8 t/s | Offload |
| bartowski/Qwen2.5-32B-Instruct-GGUF | 32.76B | Q4_K_L | Good |
20.83 GB
|
2.6 t/s | Offload |
| unsloth/Qwen3-Coder-Next-GGUF | 79.67B | IQ1_M | Very low |
21.39 GB
|
2.5 t/s | Offload |
"Fits in VRAM" = fast, fully on GPU. "Offload" = part on system RAM, slower. Speed is a rough estimate.
Frequently asked questions
What is the VRAM and memory bandwidth of the NVIDIA GTX 1660 Super?
The NVIDIA GTX 1660 Super features 6 GB of VRAM and a memory bandwidth of 192 GB/s (GDDR6). In local language model inference, VRAM determines which model sizes fit on the card, while memory bandwidth dictates how many tokens per second the GPU generates.
What is the best local AI model to run on a NVIDIA GTX 1660 Super?
The best overall model for the NVIDIA GTX 1660 Super is LiquidAI/LFM2.5-2.6B-GGUF using the recommended BF16 quantization (5.89 GB total memory). With 6 GB of VRAM, this GPU typically runs smaller models, typically up to about 3–4B at full GPU speed.
What models can you run on an 8 GB GPU like the NVIDIA GTX 1660 Super?
With 8 GB of VRAM, the NVIDIA GTX 1660 Super comfortably runs 7B and 8B models (such as Llama 3.1 8B, Mistral 7B, or Qwen 2.5 7B) using Q4_K_M or Q5_K_M quantizations (~5.5 to 6.8 GB). Running 14B models requires offloading memory to system RAM, which lowers token generation speed.
NVIDIA GTX 1660 Super Head-to-Head Comparisons
Compare specs, memory bandwidth, and AI model capability against other graphics cards.
NVIDIA GTX 1660 Super vs NVIDIA RTX 5070
Side-by-side local AI performance and supported model comparison.
NVIDIA GTX 1660 Super vs NVIDIA RTX 4070 Ti
Side-by-side local AI performance and supported model comparison.
NVIDIA GTX 1660 Super vs NVIDIA RTX 4070 Super
Side-by-side local AI performance and supported model comparison.