Which AI models run on a NVIDIA RTX 4060 Ti 8 GB?
The NVIDIA RTX 4060 Ti 8 GB has 8 GB of VRAM. Here are the popular AI models it can run locally (4,096-token context, ~16.0 GB system RAM assumed), ranked by popularity.
See also: Best GPU for running local LLMs.
The NVIDIA RTX 4060 Ti 8 GB comes with 8 GB of VRAM. Among the popular GGUF models we track, it can run 25 of them entirely in VRAM — including Qwen3-4B-GGUF, Meta-Llama-3.1-8B-Instruct-GGUF, Jan-v3.5-4B-gguf.
With 8 GB you can typically run a 7–8B model at Q6, entirely in VRAM. Which quantization is best depends on the exact model and your context length. For a full shortlist, see the best LLM for 8 GB of VRAM.
Larger models such as Qwen3-Coder-30B-A3B-Instruct-GGUF still run on a NVIDIA RTX 4060 Ti 8 GB but require offloading part of the model to system RAM, which lowers speed. Models that exceed both VRAM and RAM are not listed.
25 fit fully in VRAM · 10 run with offload
| Model | Size | Quant. | Quality | Memory | Speed~ | Verdict |
|---|---|---|---|---|---|---|
| MaziyarPanahi/Qwen3-4B-GGUF | 4.02B | Q6_K | Excellent |
4.44 GB
|
129.9 t/s | Fits in VRAM |
| bartowski/Meta-Llama-3.1-8B-Instruct-GGUF | 8.03B | Q6_K_L | Excellent |
7.66 GB
|
62.7 t/s | Fits in VRAM |
| janhq/Jan-v3.5-4B-gguf | 4.41B | Q8_0 | Excellent |
5.73 GB
|
91.5 t/s | Fits in VRAM |
| Qwen/Qwen3-8B-GGUF | 8.19B | Q6_K | Excellent |
7.63 GB
|
63.9 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-0.6B-GGUF | 0.75B | GGUF | Excellent |
2.64 GB
|
284.6 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-14B-GGUF | 14.77B | Q2_K | Low |
6.78 GB
|
74.6 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-1.7B-GGUF | 2.03B | GGUF | Excellent |
5.03 GB
|
105.5 t/s | Fits in VRAM |
| bartowski/Qwen2.5-7B-Instruct-GGUF | 7.62B | Q6_K_L | Excellent |
7.09 GB
|
65.9 t/s | Fits in VRAM |
| Qwen/Qwen2.5-3B-Instruct-GGUF | 3.09B | GGUF | Excellent |
7.27 GB
|
63.2 t/s | Fits in VRAM |
| LiquidAI/LFM2.5-2.6B-GGUF | 2.7B | BF16 | Excellent |
5.89 GB
|
79.5 t/s | Fits in VRAM |
| LiquidAI/LFM2.5-1.2B-Instruct-GGUF | 1.17B | BF16 | Excellent |
3.03 GB
|
183.3 t/s | Fits in VRAM |
| Qwen/Qwen2.5-1.5B-Instruct-GGUF | 1.54B | GGUF | Excellent |
4.23 GB
|
120.6 t/s | Fits in VRAM |
| Qwen/Qwen2.5-Coder-7B-Instruct-GGUF | 7.62B | Q6_K | Excellent |
6.84 GB
|
68.7 t/s | Fits in VRAM |
| MaziyarPanahi/Qwen3-4B-Instruct-2507-GGUF | 4.02B | Q6_K | Excellent |
4.44 GB
|
129.9 t/s | Fits in VRAM |
| MaziyarPanahi/Meta-Llama-3-8B-Instruct-GGUF | 8.03B | Q6_K | Excellent |
7.42 GB
|
65.1 t/s | Fits in VRAM |
| MaziyarPanahi/Mistral-7B-Instruct-v0.3-GGUF | 7.25B | Q6_K | Excellent |
6.84 GB
|
72.2 t/s | Fits in VRAM |
| Qwen/Qwen2.5-0.5B-Instruct-GGUF | 0.49B | GGUF | Excellent |
2.03 GB
|
339.1 t/s | Fits in VRAM |
| MaziyarPanahi/gemma-3-4b-it-GGUF | 4.3B | Q8_0 | Excellent |
5.0 GB
|
104.0 t/s | Fits in VRAM |
| MaziyarPanahi/Phi-3.5-mini-instruct-GGUF | 3.82B | Q8_0 | Excellent |
6.08 GB
|
105.8 t/s | Fits in VRAM |
| lmstudio-community/Llama-3.2-3B-Instruct-GGUF | 3.21B | Q8_0 | Excellent |
4.29 GB
|
125.5 t/s | Fits in VRAM |
| MaziyarPanahi/Llama-3-8B-Instruct-32k-v0.1-GGUF | 8.03B | Q6_K | Excellent |
7.42 GB
|
65.1 t/s | Fits in VRAM |
| MaziyarPanahi/Yi-1.5-6B-Chat-GGUF | 6.06B | Q6_K | Excellent |
5.68 GB
|
86.3 t/s | Fits in VRAM |
| MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF | 1.24B | GGUF | Excellent |
3.3 GB
|
173.2 t/s | Fits in VRAM |
| MaziyarPanahi/Mistral-Nemo-Instruct-2407-GGUF | 12.25B | Q3_K_L | Good |
7.54 GB
|
65.5 t/s | Fits in VRAM |
| MaziyarPanahi/WizardLM-2-7B-GGUF | 7.24B | Q6_K | Excellent |
6.79 GB
|
72.3 t/s | Fits in VRAM |
| unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF | 30.53B | Q5_K_XL | Very good |
21.42 GB
|
2.5 t/s | Offload |
| unsloth/gpt-oss-20b-GGUF | 20.91B | F16 | Very good |
13.74 GB
|
3.9 t/s | Offload |
| unsloth/Qwen-AgentWorld-35B-A3B-GGUF | 34.66B | Q4_K_XL | Very good |
21.67 GB
|
2.4 t/s | Offload |
| MaziyarPanahi/Qwen3-32B-GGUF | 32.76B | Q5_K_M | Very good |
23.42 GB
|
2.3 t/s | Offload |
| MaziyarPanahi/Qwen3-30B-A3B-GGUF | 30.53B | Q5_K_M | Very good |
21.41 GB
|
2.5 t/s | Offload |
| unsloth/Ornith-1.0-35B-GGUF | — | Q4_K_XL | Excellent |
21.67 GB
|
2.4 t/s | Offload |
| bartowski/Kwaipilot_KAT-Coder-V2.5-Dev-GGUF | 34.66B | Q5_K_S | Very good |
23.38 GB
|
2.2 t/s | Offload |
| unsloth/Qwen3-Coder-Next-GGUF | 79.67B | IQ2_XXS | Very low |
22.89 GB
|
2.3 t/s | Offload |
| MaziyarPanahi/Mistral-Small-24B-Instruct-2501-GGUF | 23.57B | Q6_K | Excellent |
19.44 GB
|
2.8 t/s | Offload |
| unsloth/GLM-4.7-Flash-GGUF | 31.22B | Q5_K_XL | Very good |
21.21 GB
|
2.5 t/s | Offload |
"Fits in VRAM" = fast, fully on GPU. "Offload" = part on system RAM, slower. Speed is a rough estimate.
Frequently asked questions
How much VRAM does the NVIDIA RTX 4060 Ti 8 GB have?
The NVIDIA RTX 4060 Ti 8 GB has 8 GB of VRAM, which determines how large a model it can run entirely on the GPU.
What is the best LLM to run on a NVIDIA RTX 4060 Ti 8 GB?
Among popular models, MaziyarPanahi/Qwen3-4B-GGUF runs well on a NVIDIA RTX 4060 Ti 8 GB using the Q6_K quantization (about 4.44 GB). With 8 GB you can generally run a 7–8B model at Q6, entirely in VRAM. Larger models trade speed for capability via RAM offloading. See the best LLM for 8 GB of VRAM.
Can a NVIDIA RTX 4060 Ti 8 GB run a 7–8B model?
Yes. A 7–8B model like Meta-Llama-3.1-8B-Instruct-GGUF fits entirely in the 8 GB of a NVIDIA RTX 4060 Ti 8 GB (Q6_K_L).
Can a NVIDIA RTX 4060 Ti 8 GB run a 13–14B model?
Yes. A 13–14B model like Qwen3-14B-GGUF fits entirely in the 8 GB of a NVIDIA RTX 4060 Ti 8 GB (Q2_K).
Can a NVIDIA RTX 4060 Ti 8 GB run a 70B model?
Only with offloading. A 70B model like Qwen3-Coder-Next-GGUF runs on a NVIDIA RTX 4060 Ti 8 GB by using system RAM in addition to its 8 GB, which is slower.