Supported Local AI Models
Browse verified GGUF text and reasoning models ready for local execution with Ollama, LM Studio, or llama.cpp. Click any model to calculate required VRAM, RAM and tokens per second for your setup.
Qwen3-Coder-30B-A3B-Instruct
unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF
LFM2.5-2.6B
LiquidAI/LFM2.5-2.6B-GGUF
LFM2.5-230M
LiquidAI/LFM2.5-230M-GGUF
gpt-oss-20b
unsloth/gpt-oss-20b-GGUF
Ornith-1.0-9B
unsloth/Ornith-1.0-9B-GGUF
LFM2.5-8B-A1B
LiquidAI/LFM2.5-8B-A1B-GGUF
Qwen3-4B
unsloth/Qwen3-4B-GGUF
GLM-5.3-Flash
unsloth/GLM-5.3-Flash-GGUF
DeepSeek-Coder-V2-Lite-Instruct
bartowski/DeepSeek-Coder-V2-Lite-Instruct-GGUF
Qwen-AgentWorld-35B-A3B
unsloth/Qwen-AgentWorld-35B-A3B-GGUF
GLM-5.3
unsloth/GLM-5.3-GGUF
Kwaipilot_KAT-Coder-V2.5-Dev
bartowski/Kwaipilot_KAT-Coder-V2.5-Dev-GGUF
GLM-5.2
unsloth/GLM-5.2-GGUF
Qwen3-8B
Qwen/Qwen3-8B-GGUF
Qwen2.5-3B-Instruct
Qwen/Qwen2.5-3B-Instruct-GGUF
LFM2.5-1.2B-Instruct
LiquidAI/LFM2.5-1.2B-Instruct-GGUF
Meta-Llama-3.1-8B-Instruct
bartowski/Meta-Llama-3.1-8B-Instruct-GGUF
NVIDIA-Nemotron-3.5-Lightning-30B-A3B
ggml-org/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF
Qwen2.5-Coder-7B-Instruct
Qwen/Qwen2.5-Coder-7B-Instruct-GGUF
Jan-v3.5-4B
janhq/Jan-v3.5-4B-gguf
Qwen3-0.6B
MaziyarPanahi/Qwen3-0.6B-GGUF
Qwen3-14B
MaziyarPanahi/Qwen3-14B-GGUF
Qwen3-1.7B
MaziyarPanahi/Qwen3-1.7B-GGUF
Qwen_Qwen3-Next-80B-A3B-Thinking
bartowski/Qwen_Qwen3-Next-80B-A3B-Thinking-GGUF
Qwen3-30B-A3B
MaziyarPanahi/Qwen3-30B-A3B-GGUF
Qwen3-32B
MaziyarPanahi/Qwen3-32B-GGUF
Llama-3.2-3B-Instruct
unsloth/Llama-3.2-3B-Instruct-GGUF
Yi-Coder-9B-Chat
MaziyarPanahi/Yi-Coder-9B-Chat-GGUF
Qwen2.5-7B-Instruct
bartowski/Qwen2.5-7B-Instruct-GGUF
DeepSeek-R1-Distill-Llama-70B
unsloth/DeepSeek-R1-Distill-Llama-70B-GGUF
Qwen2.5-1.5B-Instruct
Qwen/Qwen2.5-1.5B-Instruct-GGUF
Qwen3-Coder-Next
unsloth/Qwen3-Coder-Next-GGUF
Qwen2.5-0.5B-Instruct
Qwen/Qwen2.5-0.5B-Instruct-GGUF
Qwen3-4B-Instruct-2507
MaziyarPanahi/Qwen3-4B-Instruct-2507-GGUF
gpt-oss-120b
ggml-org/gpt-oss-120b-GGUF
Yi-Coder-1.5B-Chat
MaziyarPanahi/Yi-Coder-1.5B-Chat-GGUF
Meta-Llama-3-8B-Instruct
MaziyarPanahi/Meta-Llama-3-8B-Instruct-GGUF
Phi-3.5-mini-instruct
MaziyarPanahi/Phi-3.5-mini-instruct-GGUF
Mixtral-8x22B-v0.1
MaziyarPanahi/Mixtral-8x22B-v0.1-GGUF
Mistral-7B-Instruct-v0.3
MaziyarPanahi/Mistral-7B-Instruct-v0.3-GGUF
Hermes-3-Llama-3.1-70B
bartowski/Hermes-3-Llama-3.1-70B-GGUF
Llama-3.2-1B-Instruct
MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF
gemma-3-4b-it
MaziyarPanahi/gemma-3-4b-it-GGUF
Llama-3-8B-Instruct-32k-v0.1
MaziyarPanahi/Llama-3-8B-Instruct-32k-v0.1-GGUF
gemma-2b
google/gemma-2b
solar-pro-preview-instruct
MaziyarPanahi/solar-pro-preview-instruct-GGUF
DeepSeek-R1-0528-Qwen3-8B
MaziyarPanahi/DeepSeek-R1-0528-Qwen3-8B-GGUF
Mistral-Small-24B-Instruct-2501
MaziyarPanahi/Mistral-Small-24B-Instruct-2501-GGUF
Phi-4-mini-instruct
MaziyarPanahi/Phi-4-mini-instruct-GGUF
Llama-3.3-70B-Instruct
MaziyarPanahi/Llama-3.3-70B-Instruct-GGUF
phi-4
MaziyarPanahi/phi-4-GGUF
Mistral-Nemo-Instruct-2407
MaziyarPanahi/Mistral-Nemo-Instruct-2407-GGUF
gemma-3-1b-it
MaziyarPanahi/gemma-3-1b-it-GGUF
QwQ-32B
MaziyarPanahi/QwQ-32B-GGUF
Qwen2-7B-Instruct
MaziyarPanahi/Qwen2-7B-Instruct-GGUF
gemma-2-2b-it
MaziyarPanahi/gemma-2-2b-it-GGUF
Mistral-Small-Instruct-2409
MaziyarPanahi/Mistral-Small-Instruct-2409-GGUF
Ministral-3-3B-Reasoning-2512
MaziyarPanahi/Ministral-3-3B-Reasoning-2512-GGUF
gemma-3-12b-it
MaziyarPanahi/gemma-3-12b-it-GGUF
mistral-small-3.1-24b-instruct-2503-hf
MaziyarPanahi/mistral-small-3.1-24b-instruct-2503-hf-GGUF
DeepSeek-V3-0324
MaziyarPanahi/DeepSeek-V3-0324-GGUF
Meta-Llama-3.1-70B-Instruct
MaziyarPanahi/Meta-Llama-3.1-70B-Instruct-GGUF
Yi-1.5-6B-Chat
MaziyarPanahi/Yi-1.5-6B-Chat-GGUF
WizardLM-2-7B
MaziyarPanahi/WizardLM-2-7B-GGUF
gemma-3-27b-it
MaziyarPanahi/gemma-3-27b-it-GGUF
mathstral-7B-v0.1
MaziyarPanahi/mathstral-7B-v0.1-GGUF
Llama-3-8B-Instruct-64k
MaziyarPanahi/Llama-3-8B-Instruct-64k-GGUF
Mistral-Large-Instruct-2411
MaziyarPanahi/Mistral-Large-Instruct-2411-GGUF
INTELLECT-2
MaziyarPanahi/INTELLECT-2-GGUF
Meta-Llama-3.1-405B-Instruct
MaziyarPanahi/Meta-Llama-3.1-405B-Instruct-GGUF
firefunction-v2
MaziyarPanahi/firefunction-v2-GGUF
GLM-4.7-Flash
unsloth/GLM-4.7-Flash-GGUF
LFM2.5-2.6B-DSpark
LiquidAI/LFM2.5-2.6B-DSpark-GGUF
Qwen2.5-Coder-1.5B-Instruct
Qwen/Qwen2.5-Coder-1.5B-Instruct-GGUF
Qwen2.5-Coder-32B-Instruct
Qwen/Qwen2.5-Coder-32B-Instruct-GGUF
Qwen2.5-72B-Instruct
bartowski/Qwen2.5-72B-Instruct-GGUF
Qwen3-30B-A3B-Instruct-2507
MaziyarPanahi/Qwen3-30B-A3B-Instruct-2507-GGUF
Qwen2.5-Coder-14B-Instruct
Qwen/Qwen2.5-Coder-14B-Instruct-GGUF
Qwen_Qwen3-14B
bartowski/Qwen_Qwen3-14B-GGUF
DeepSeek-R1-Distill-Qwen-1.5B
unsloth/DeepSeek-R1-Distill-Qwen-1.5B-GGUF
stories15M_MOE
ggml-org/stories15M_MOE
GLM-4.6V-Flash
MaziyarPanahi/GLM-4.6V-Flash-GGUF
gpt-oss-20b-Derestricted
MaziyarPanahi/gpt-oss-20b-Derestricted-GGUF
LFM2.5-350M
LiquidAI/LFM2.5-350M-GGUF
Nemotron-Orchestrator-8B
MaziyarPanahi/Nemotron-Orchestrator-8B-GGUF
Trinity-Mini
MaziyarPanahi/Trinity-Mini-GGUF
Qwen_Qwen3-0.6B
bartowski/Qwen_Qwen3-0.6B-GGUF
Qwen2.5-Coder-3B-Instruct
Qwen/Qwen2.5-Coder-3B-Instruct-GGUF
Ling-3.0-tiny
bartowski/Ling-3.0-tiny-GGUF
LFM2.5-8B-A1B-DSpark
LiquidAI/LFM2.5-8B-A1B-DSpark-GGUF
gemma-3-270m-it
unsloth/gemma-3-270m-it-GGUF
Qwen2.5-14B-Instruct
bartowski/Qwen2.5-14B-Instruct-GGUF
DeepSeek-R1-Distill-Qwen-7B
bartowski/DeepSeek-R1-Distill-Qwen-7B-GGUF
Ministral-3-14B-Reasoning-2512
MaziyarPanahi/Ministral-3-14B-Reasoning-2512-GGUF
NVIDIA-Nemotron-Nano-12B-v2
MaziyarPanahi/NVIDIA-Nemotron-Nano-12B-v2-GGUF
Qwen3-4B-Thinking-2507
MaziyarPanahi/Qwen3-4B-Thinking-2507-GGUF
gemma-2b-it
google/gemma-2b-it
DeepSeek-R1-Distill-Qwen-32B
bartowski/DeepSeek-R1-Distill-Qwen-32B-GGUF
DeepSeek-R1-Distill-Qwen-14B
bartowski/DeepSeek-R1-Distill-Qwen-14B-GGUF
DeepSeek-R1-Distill-Llama-8B
unsloth/DeepSeek-R1-Distill-Llama-8B-GGUF
Kimi-K2-Instruct
unsloth/Kimi-K2-Instruct-GGUF
Llama-2-7B-Chat
TheBloke/Llama-2-7B-Chat-GGUF
Qwen_Qwen3-1.7B
bartowski/Qwen_Qwen3-1.7B-GGUF
microsoft_Phi-4-mini-instruct
bartowski/microsoft_Phi-4-mini-instruct-GGUF
Qwen3-Next-80B-A3B-Instruct
unsloth/Qwen3-Next-80B-A3B-Instruct-GGUF
Qwen_Qwen3-8B
bartowski/Qwen_Qwen3-8B-GGUF
Mistral-7B-Instruct-v0.2
TheBloke/Mistral-7B-Instruct-v0.2-GGUF
Qwen2.5-32B-Instruct
bartowski/Qwen2.5-32B-Instruct-GGUF
Ornith-1.0-35B
unsloth/Ornith-1.0-35B-GGUF
Qwen3-235B-A22B
Qwen/Qwen3-235B-A22B-GGUF
moonshotai_Kimi-Linear-48B-A3B-Instruct
bartowski/moonshotai_Kimi-Linear-48B-A3B-Instruct-GGUF
Mistral-7B-Instruct-v0.1
TheBloke/Mistral-7B-Instruct-v0.1-GGUF
HuggingFaceTB_SmolLM3-3B
bartowski/HuggingFaceTB_SmolLM3-3B-GGUF
gemma-2-9b-it
bartowski/gemma-2-9b-it-GGUF
SmolLM2-135M-Instruct
bartowski/SmolLM2-135M-Instruct-GGUF
gemma-7b-it
google/gemma-7b-it
cognitivecomputations_Dolphin-Mistral-24B-Venice-Edition
bartowski/cognitivecomputations_Dolphin-Mistral-24B-Venice-Edition-GGUF
MiniMax-M2.7
unsloth/MiniMax-M2.7-GGUF
DeepSeek-R1
unsloth/DeepSeek-R1-GGUF
L3-8B-Stheno-v3.2
bartowski/L3-8B-Stheno-v3.2-GGUF
gemma-7b
google/gemma-7b
GLM-4.7-Flash-REAP-23B-A3B
unsloth/GLM-4.7-Flash-REAP-23B-A3B-GGUF
WhiteRabbitNeo_WhiteRabbitNeo-V3-7B
bartowski/WhiteRabbitNeo_WhiteRabbitNeo-V3-7B-GGUF
Qwen_Qwen3-4B-Instruct-2507
bartowski/Qwen_Qwen3-4B-Instruct-2507-GGUF
TheDrummer_Cydonia-24B-v4.3
bartowski/TheDrummer_Cydonia-24B-v4.3-GGUF
Laguna-S-2.1
unsloth/Laguna-S-2.1-GGUF
Phi-3-mini-4k-instruct
microsoft/Phi-3-mini-4k-instruct-gguf
granite-4.1-8b-fp8
ibm-granite/granite-4.1-8b-fp8
openai_gpt-oss-20b
bartowski/openai_gpt-oss-20b-GGUF
Hermes-3-Llama-3.2-3B
bartowski/Hermes-3-Llama-3.2-3B-GGUF
dolphin-2.9-llama3-8b
bartowski/dolphin-2.9-llama3-8b-GGUF
NemoMix-Unleashed-12B
bartowski/NemoMix-Unleashed-12B-GGUF
NousResearch_Hermes-4-14B
bartowski/NousResearch_Hermes-4-14B-GGUF
Qwen2.5-Math-1.5B-Instruct
bartowski/Qwen2.5-Math-1.5B-Instruct-GGUF
CodeLlama-7B-Instruct
TheBloke/CodeLlama-7B-Instruct-GGUF
Qwen2.5-Math-7B-Instruct
lmstudio-community/Qwen2.5-Math-7B-Instruct-GGUF
phi-2
TheBloke/phi-2-GGUF
deepseek-ai_DeepSeek-R1-0528
bartowski/deepseek-ai_DeepSeek-R1-0528-GGUF
Llama-2-7B
TheBloke/Llama-2-7B-GGUF
Qwen2.5-Coder-0.5B-Instruct
Qwen/Qwen2.5-Coder-0.5B-Instruct-GGUF
Ling-3.0-flash
bartowski/Ling-3.0-flash-GGUF
Phi-4-mini-reasoning
lmstudio-community/Phi-4-mini-reasoning-GGUF
Qwen2.5-VL-32B-Instruct
lmstudio-community/Qwen2.5-VL-32B-Instruct-GGUF
Qwen3.8-2.4T-A95B
unsloth/Qwen3.8-2.4T-A95B-GGUF
Qwen2-1.5B-Instruct
Qwen/Qwen2-1.5B-Instruct-GGUF
NVIDIA-Nemotron-3-Nano-Omni-30B-A3B-Reasoning
unsloth/NVIDIA-Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF
Qwen3-Coder-30B-A3B-Instruct-1M
unsloth/Qwen3-Coder-30B-A3B-Instruct-1M-GGUF
Wanabi-Novelist-24B
mradermacher/Wanabi-Novelist-24B-GGUF
zai-org_GLM-4.7-Flash
bartowski/zai-org_GLM-4.7-Flash-GGUF
Qwen_Qwen3-Coder-Next
bartowski/Qwen_Qwen3-Coder-Next-GGUF
Dolphin3.0-Llama3.1-8B
bartowski/Dolphin3.0-Llama3.1-8B-GGUF
Qwen2.5-7B-Instruct-1M
lmstudio-community/Qwen2.5-7B-Instruct-1M-GGUF
Jan-code-4b
janhq/Jan-code-4b-gguf
cognitivecomputations_Dolphin3.0-R1-Mistral-24B
bartowski/cognitivecomputations_Dolphin3.0-R1-Mistral-24B-GGUF
granite-4.2-30b
bartowski/granite-4.2-30b-GGUF
Nanbeige_Nanbeige4.2-3B
bartowski/Nanbeige_Nanbeige4.2-3B-GGUF
Qwen3-Next-80B-A3B-Thinking
unsloth/Qwen3-Next-80B-A3B-Thinking-GGUF
Nemotron-3-Nano-30B-A3B
unsloth/Nemotron-3-Nano-30B-A3B-GGUF
Cotype-Nano
QuantFactory/Cotype-Nano-GGUF
GLM-4.5-Air
unsloth/GLM-4.5-Air-GGUF
Qwen_Qwen3-30B-A3B
bartowski/Qwen_Qwen3-30B-A3B-GGUF
Qwen2-0.5B-Instruct
Qwen/Qwen2-0.5B-Instruct-GGUF
qwen2.5-7b-ins-v3
bartowski/qwen2.5-7b-ins-v3-GGUF
allenai_Olmo-3.1-32B-Think
bartowski/allenai_Olmo-3.1-32B-Think-GGUF
Qwen_Qwen3-30B-A3B-Instruct-2507
bartowski/Qwen_Qwen3-30B-A3B-Instruct-2507-GGUF
Qwen2.5-Coder-32B
lmstudio-community/Qwen2.5-Coder-32B-GGUF
google_gemma-3-1b-it
bartowski/google_gemma-3-1b-it-GGUF
Hermes-3-Llama-3.1-8B
bartowski/Hermes-3-Llama-3.1-8B-GGUF
LFM2-2.6B
LiquidAI/LFM2-2.6B-GGUF
Qwen_Qwen3-32B
bartowski/Qwen_Qwen3-32B-GGUF
Qwen1.5-7B-Chat
Qwen/Qwen1.5-7B-Chat-GGUF
Hermes-4-70B
lmstudio-community/Hermes-4-70B-GGUF
LFM2-1.2B
LiquidAI/LFM2-1.2B-GGUF
GLM-4.7
unsloth/GLM-4.7-GGUF
openai_gpt-oss-120b
bartowski/openai_gpt-oss-120b-GGUF
OLMoE-1B-7B-0924-Instruct
bartowski/OLMoE-1B-7B-0924-Instruct-GGUF
gemma-3-1B-it-qat
lmstudio-community/gemma-3-1B-it-qat-GGUF
No local AI model matches your search. Try searching for "Llama", "Qwen", "Mistral" or "DeepSeek".
Frequently asked questions about local models
What is the GGUF model format?
GGUF (GPT-Generated Unified Format) is the standard format used by llama.cpp, Ollama, LM Studio, and Jan to run large language models on consumer hardware. It stores weights and quantization metadata in a single, fast-loading binary file.
How do I know which model fits on my computer?
Select any model above to open our interactive calculator. It reads the model's exact weight sizes and calculates required memory for KV cache and system overhead across various context windows (4K, 8K, 32K+).
Can't find the model you need?
You can search for any GGUF repository directly on the homepage search bar, or browse models categorized by graphics card.