Download README.md from fastgpu/vram-and-gpu-cost-calculator: direct link, hf CLI and curl.
- Browser
- Download file 56.6 kB
-
https://huggingface.co/spaces/fastgpu/vram-and-gpu-cost-calculator/resolve/main/README.md
- Command line
-
hf download hf://spaces/fastgpu/vram-and-gpu-cost-calculator/README.md
-
curl -L -o README.md https://huggingface.co/spaces/fastgpu/vram-and-gpu-cost-calculator/resolve/main/README.md
title: VRAM & GPU Cost Calculator
emoji: ⚡
colorFrom: blue
colorTo: indigo
sdk: static
pinned: true
license: mit
short_description: VRAM any HF model needs + live GPU rental prices
tags:
- vram
- gpu
- gpu-prices
- calculator
- inference
- fine-tuning
- llm
- cloud-gpu
models:
- Qwen/Qwen3-0.6B
- google/gemma-4-26B-A4B-it
- Qwen/Qwen3-8B
- BAAI/bge-large-en-v1.5
- Qwen/Qwen3-VL-8B-Instruct
- Qwen/Qwen3-4B
- google/gemma-4-31B-it
- Qwen/Qwen3-Embedding-0.6B
- Qwen/Qwen3.5-9B
- Qwen/Qwen2.5-7B-Instruct
- Qwen/Qwen3.5-4B
- meta-llama/Llama-3.2-1B-Instruct
- intfloat/multilingual-e5-large
- Qwen/Qwen3.8-27B
- Qwen/Qwen2.5-1.5B-Instruct
- openai/whisper-large-v3-turbo
- zai-org/GLM-5.3-Flash
- RadixArk/Kimi-K3-DSpark
- meta-llama/Llama-3.1-8B-Instruct
- openai/gpt-oss-20b
- unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF
- Qwen/Qwen2.5-VL-7B-Instruct
- nvidia/Qwen3.6-35B-A3B-NVFP4
- ornith-ai/Ornith-1.5-9B-GGUF
- Qwen/Qwen3-14B
- dphn/dolphin-2.9.1-yi-1.5-34b
- Qwen/Qwen3.8-27B-FP8
- deepseek-ai/DeepSeek-V4-Flash-0731
- Qwen/Qwen3.6-35B-A3B-FP8
- google/gemma-4-E4B-it
- stabilityai/stable-diffusion-xl-base-1.0
- deepseek-ai/DeepSeek-V3.2
- prism-ml/Ternary-Bonsai-2-27B-gguf
- Qwen/Qwen2.5-3B-Instruct
- openai/gpt-oss-120b
- Qwen/Qwen3.5-2B
- openai/whisper-large-v3
- cdiamond/Qwen3.8-27B-iMatrix-NVFP4-MTP-GGUF
- ornith-ai/Ornith-1.5-35B-A3B-GGUF
- Qwen/Qwen3-32B
- Qwen/Qwen3-4B-Instruct-2507
- google/embeddinggemma-300m
- Qwen/Qwen3.6-35B-A3B
- ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF
- openai/whisper-small
- Qwen/Qwen3-VL-4B-Instruct
- google/gemma-3-1b-it
- Qwen/Qwen3-1.7B
- Qwen/Qwen3-14B-AWQ
- google/gemma-4-E2B-it
- Qwen/Qwen3-VL-2B-Instruct
- Qwen/Qwen2.5-7B-Instruct-AWQ
- MahmoudAshraf/mms-300m-1130-forced-aligner
- Qwen/Qwen3.5-0.8B
- CMSManhattan/JiRackUltra_1b
- ornith-ai/Ornith-1.0-9B-GGUF
- Qwen/Qwen3-Embedding-8B
- Qwen/Qwen3.6-27B-FP8
- BAAI/bge-reranker-large
- Qwen/Qwen3-8B-AWQ
- Qwen/Qwen3.6-27B
- Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice
- Qwen/Qwen2.5-VL-3B-Instruct
- vikhyatk/moondream2
- Qwen/Qwen2.5-Coder-7B-Instruct
- deepseek-ai/DeepSeek-OCR
- mixedbread-ai/mxbai-embed-large-v1
- jinaai/jina-embeddings-v3
- nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16
- mistralai/Voxtral-Mini-4B-Realtime-2602
- Qwen/Qwen3.5-27B
- handy-computer/nemotron-3.5-asr-streaming-0.6b-gguf
- ggml-org/Qwen3.8-27B-GGUF
- datalab-to/chandra-ocr-2
- Qwen/Qwen2.5-32B-Instruct
- Qwen/Qwen3-Embedding-4B
- handy-computer/parakeet-unified-en-0.6b-gguf
- farbodtavakkoli/OTel-2.0-LLM-31B-IT
- Qwen/Qwen3-VL-8B-Instruct-FP8
- peculiar-ragdoll/Tiel-Coder-35B-A3B-GGUF-MTP
- Qwen/Qwen2.5-14B-Instruct-AWQ
- google/gemma-4-12B-it
- Qwen/Qwen2.5-VL-7B-Instruct-AWQ
- Qwen/Qwen3.8-Flash-Next
- ornith-ai/Ornith-1.0-35B-GGUF
- stable-diffusion-v1-5/stable-diffusion-v1-5
- unsloth/gemma-4-12b-it-GGUF
- Qwen/Qwen2-VL-7B-Instruct-AWQ
- Qwen/Qwen2.5-Coder-14B-Instruct
- mistralai/Mistral-7B-Instruct-v0.2
- zai-org/GLM-4.7-Flash
- meta-llama/Llama-3.2-3B-Instruct
- zai-org/GLM-5.3
- intfloat/multilingual-e5-large-instruct
- autotrust/JEV-27B-VL
- ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF
- nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4
- Alibaba-NLP/gte-multilingual-base
- Qwen/Qwen3-VL-Embedding-8B
- facebook/w2v-bert-2.0
- Qwen/Qwen3.5-35B-A3B
- peculiar-ragdoll/Dirk-Qwen3.8-27B-GGUF
- ornith-ai/Ornith-1.0-35B
- k2-fsa/OmniVoice
- datalab-to/surya-ocr-2
- Qwen/Qwen3-30B-A3B
- nvidia/Gemma-4-31B-IT-NVFP4
- deepseek-ai/DeepSeek-V3-0324
- nvidia/Qwen2.5-VL-7B-Instruct-NVFP4
- deepseek-ai/DeepSeek-V3
- llava-hf/llava-1.5-7b-hf
- Qwen/Qwen3.5-35B-A3B-FP8
- Qwen/Qwen2.5-14B-Instruct
- unsloth/Qwen3.8-Flash-Next-GGUF
- baidu/Unlimited-OCR
- Qwen/Qwen2-VL-2B-Instruct
- unsloth/Qwen3.6-35B-A3B-GGUF
- TinyLlama/TinyLlama-1.1B-Chat-v1.0
- nvidia/nemotron-3.5-asr-streaming-0.6b
- HuggingFaceTB/SmolVLM2-500M-Video-Instruct
- deepseek-ai/DeepSeek-V4.1-Flash
- Alibaba-NLP/gte-large-en-v1.5
- Qwen/Qwen3-ASR-1.7B
- openbmb/MiniCPM5-2B
- antirez/deepseek-v4-gguf
- moonshotai/Kimi-K3
- google/gemma-3-4b-it
- ornith-ai/Ornith-1.5-397B-GGUF
- unsloth/Qwen3.5-9B-GGUF
- bigscience/bloomz-560m
- gigant/romanian-wav2vec2
- unsloth/Qwen3.5-4B-GGUF
- Qwen/Qwen2.5-32B-Instruct-AWQ
- intfloat/e5-large-v2
- Qwen/Qwen3-VL-Embedding-2B
- nomic-ai/nomic-embed-text-v2-moe
- deepseek-ai/DeepSeek-R1
- openai-community/gpt2-large
- dots-studio/dots.ocr
- ornith-ai/Ornith-1.5-35B-A3B-NVFP4
- Qwen/Qwen3-4B-Instruct-2507-FP8
- sahilchachra/Unlimited-OCR-AWQ
- byteshape/Qwen3.8-27B-GGUF
- gaunernst/gemma-3-27b-it-int4-awq
- Qwen/Qwen3-Coder-Next-FP8
- LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct-AWQ
- Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice
- handy-computer/cohere-transcribe-03-2026-gguf
- KBLab/wav2vec2-large-voxrex-swedish
- RedHatAI/Qwen3.6-35B-A3B-NVFP4
- deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B
- unsloth/GLM-5.3-Flash-GGUF
- unsloth/Qwen3.6-35B-A3B-MTP-GGUF
- gabor-hosu/e5-mistral-7b-instruct-bnb-4bit
- nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-FP8
- Qwen/Qwen3-32B-AWQ
- nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16
- LiquidAI/LFM2.5-2.6B-GGUF
- nvidia/Qwen3.5-122B-A10B-NVFP4
- Qwen/Qwen2.5-Coder-32B-Instruct-AWQ
- RadixArk/Qwen3.8-27B-NVFP4
- cyankiwi/gemma-4-26B-A4B-it-AWQ-4bit
- google/gemma-2-9b-it
- FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree
- deepseek-ai/DeepSeek-V4-Flash
- openbmb/MiniCPM5-2B-DSpark
- Qwen/Qwen3-30B-A3B-Instruct-2507
- RedHatAI/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-FP8
- Inferact/Qwen3.8-27B-NVFP4
- thinkingmachines/Inkling
- kingabzpro/wav2vec2-large-xls-r-300m-Urdu
- MiniMaxAI/MiniMax-M2.7
- ornith-ai/Ornith-1.0-9B
- tencent/HunyuanOCR
- Qwen/Qwen-72B
- Qwen/Qwen3-Coder-30B-A3B-Instruct-FP8
- unsloth/Qwen3.6-27B-NVFP4
- stabilityai/sdxl-turbo
- nvidia/Gemma-4-26B-A4B-NVFP4
- deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct
- CMSManhattan/JiRackUltra_14b
- nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
- kresnik/wav2vec2-large-xlsr-korean
- unsloth/Z-Image-Turbo-GGUF
- unsloth/Qwen3.5-35B-A3B-GGUF
- Snowflake/snowflake-arctic-embed-l-v2.0
- RedHatAI/gemma-4-26B-A4B-it-FP8-dynamic
- Octen/Octen-Embedding-8B
- Lykon/dreamshaper-7
- cyankiwi/Qwen3.6-27B-AWQ-INT4
- unsloth/Inkling-Small-GGUF
- bartowski/endless-frontier_BigBang-v1-GGUF
- Qwen/Qwen2.5-Coder-32B-Instruct
- meta-llama/Meta-Llama-3-8B-Instruct
- cyankiwi/Qwen3-Coder-30B-A3B-Instruct-AWQ-4bit
- google/medgemma-4b-it
- nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4
- google/gemma-4-12B-it-qat-w4a16-ct
- unsloth/Qwen3.6-35B-A3B-NVFP4
- Snowflake/snowflake-arctic-embed-l
- unsloth/Qwen-Image-2.1-GGUF
- Qwen/Qwen2.5-VL-32B-Instruct
- casperhansen/llama-3.3-70b-instruct-awq
- Qwen/Qwen2.5-7B
- theainerd/Wav2Vec2-large-xlsr-hindi
- google/gemma-4-12B-it-qat-q4_0-gguf
- Qwen/Qwen3-8B-Base
- comodoro/wav2vec2-xls-r-300m-cs-250
- jinaai/jina-embeddings-v5-omni-small
- Orion-zhen/Qwen2.5-Coder-7B-Instruct-AWQ
- meta-llama/Llama-2-7b-hf
- deepseek-ai/DeepSeek-OCR-2
- Infomaniak-AI/vllm-translategemma-4b-it
- deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
- esatapedico/Qwen3.8-27B-NVFP4-MTP-GGUF
- ornith-ai/Ornith-1.5-9B-NVFP4
- Qwen/Qwen2.5-Coder-7B-Instruct-AWQ
- Yehor/w2v-xls-r-uk
- microsoft/Phi-3.5-vision-instruct
- AtomicChat/Qwen3.8-Flash-Next-GGUF
- Qwen/Qwen3-0.6B-Base
- meta-models/Muse-Glimmer-30B-GGUF
- google/gemma-4-E4B-it-qat-q4_0-gguf
- openbmb/MiniCPM-o-4_5
- microsoft/VibeVoice-ASR
- imvladikon/wav2vec2-xls-r-300m-hebrew
- datalab-to/surya-ocr-2-gguf
- unsloth/Qwen3.6-35B-A3B-NVFP4-Fast
- empero-ai/Qwen3.8-35B-A3B-Distill-GGUF
- ISTA-DASLab/gemma-3-27b-it-GPTQ-4b-128g
- meta-llama/Llama-3.2-1B
- microsoft/VibeVoice-1.5B
- QuantTrio/Qwen3.6-35B-A3B-AWQ
- black-forest-labs/FLUX.1-dev
- swiss-ai/Apertus-v1.5-8B
- ggml-org/gemma-4-E4B-it-GGUF
- Qwen/Qwen3-Omni-30B-A3B-Instruct
- Lightricks/LTX-Video
- nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
- NbAiLab/nb-wav2vec2-1b-bokmaal-v2
- nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4
- unsloth/Qwen3.6-27B-GGUF
- zai-org/GLM-5.2
- apple/OpenELM-1_1B-Instruct
- unsloth/Qwen3.6-27B-MTP-GGUF
- Qwen/Qwen2-VL-7B-Instruct
- deepseek-ai/DeepSeek-R1-0528-Qwen3-8B
- thinkingmachines/Inkling-Small
- NbAiLab/nb-wav2vec2-1b-nynorsk
- black-forest-labs/FLUX.1-schnell
- Qwen/Qwen3.5-4B-Base
- Qwen/Qwen2.5-1.5B
- WhereIsAI/UAE-Large-V1
- meta-llama/Llama-3.1-8B
- OpenGVLab/InternVL2-1B
- OpenGVLab/InternVL3_5-GPT-OSS-20B-A4B-Preview-HF
- cl-nagoya/ruri-v3-310m
- deepseek-ai/deepseek-coder-7b-instruct-v1.5
- nvidia/Cosmos-Reason2-2B
- HuggingFaceTB/SmolLM3-3B
- yuxinlu1/gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-GGUF
- stabilityai/sd-turbo
- google/diffusiongemma-26B-A4B-it
- bartowski/Qwen2.5-32B-Instruct-GGUF
- empero-ai/Qwen3.8-9B-Distill-GGUF
- Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8
- nvidia/Qwen3.8-27B-NVFP4
- google/gemma-4-E4B
- microsoft/phi-2
- Tongyi-MAI/Z-Image-Turbo
- ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF
- handy-computer/parakeet-tdt-0.6b-v3-gguf
- Qwen/Qwen3-1.7B-Base
- nvidia/Qwen3.6-27B-NVFP4
- classla/wav2vec2-xls-r-parlaspeech-hr
- cyankiwi/MiniCPM-SALA-AWQ-8bit
- prism-ml/Ternary-Bonsai-27B-gguf
- Salesforce/blip2-opt-2.7b
- saattrupdan/wav2vec2-xls-r-300m-ftspeech
- Qwen/Qwen3-4B-Thinking-2507
- Qwen/Qwen3-4B-Base
- unsloth/gemma-4-E4B-it-GGUF
- nvidia/parakeet-tdt-0.6b-v3
- ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF
- Qwen/Qwen3-Coder-Next
- nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
- EleutherAI/pythia-6.9b
- sentence-transformers/all-roberta-large-v1
- Qwen/Qwen3-TTS-12Hz-0.6B-Base
- CompVis/stable-diffusion-v1-4
- handy-computer/whisper-medium-gguf
- empero-ai/Qwen3.8-4B-Distill-GGUF
- unsloth/Llama-3.2-1B-Instruct
- FINAL-Bench/POCKET-35B-GGUF
- Qwen/Qwen3-8B-FP8
- deepseek-ai/DeepSeek-V4-Flash-DSpark
- cyankiwi/Qwen3.8-27B-AWQ-INT4
- sentence-transformers/LaBSE
- Qwen/Qwen3.5-122B-A10B
- RedHatAI/Qwen3.8-27B-INT4
- distil-whisper/distil-large-v3
- openbmb/VoxCPM2
- unsloth/gemma-4-26B-A4B-it-GGUF
- OpenGVLab/InternVL2-2B
- ibm-granite/granite-4.1-3b
- QuantTrio/Qwen3.5-9B-AWQ
- CMSManhattan/JiRackDeltaNet_27b
- google/gemma-4-31B
- Qwen/Qwen2.5-72B-Instruct-AWQ
- Qwen/Qwen3-4B-GGUF
- Qwen/Qwen3-8B-GGUF
- Qwen/Qwen2-1.5B-Instruct
- QuantTrio/Qwen3-VL-30B-A3B-Instruct-AWQ
- stepfun-ai/GOT-OCR2_0
- SG161222/RealVisXL_V5.0
- ai-sage/Giga-Embeddings-instruct
- MiniMaxAI/MiniMax-M2.5
- abenzerps/Apodex-1.1-mini-GGUF
- google/gemma-2-2b-it
- empero-ai/Qwen3.8-2B-Distill-GGUF
- SG161222/Realistic_Vision_V5.1_noVAE
- Qwen/Qwen2.5-Coder-3B-Instruct
- NousResearch/Hermes-3-Llama-3.1-8B
- setu4993/LaBSE
- RedHatAI/Llama-3.2-3B-Instruct-FP8-dynamic
- empero-ai/Qwythos-9B-Claude-Mythos-5-1M-GGUF
- boboliu/bge-reranker-v2.5-gemma2-lightweight-gptq
- Qwen/Qwen3.5-122B-A10B-FP8
- Qwen/Qwen3-235B-A22B
- ornith-ai/Ornith-1.5-35B-A3B-FP8
- Alibaba-NLP/gte-Qwen2-1.5B-instruct
- zai-org/GLM-5
- zai-org/GLM-4.5-Air
- Qwen/Qwen2.5-Coder-1.5B-Instruct
- LeaderboardModel1/zeta-2.1-autoround-W4A16
- unsloth/gemma-4-12B-it-qat-GGUF
- bartowski/MiMo-V2.6-Distill-Qwen-9B-GGUF
- XHToken/Spark-X2.5-4B-GGUF
- llava-hf/llava-onevision-qwen2-0.5b-ov-hf
- bartowski/Qwen3.8-27B-GGUF
- Qwen/Qwen3-ASR-0.6B
- unsloth/gpt-oss-20b-GGUF
- allenai/Olmo-3-7B-Think
- bartowski/Meta-Llama-3.1-8B-Instruct-GGUF
- unsloth/GLM-5.3-GGUF
- Qwen/Qwen2.5-32B-Instruct-GPTQ-Int4
- nlpai-lab/KURE-v1
- unsloth/Qwen-Image-Edit-2511-GGUF
- Qwen/Qwen3.5-0.8B-Base
- stelterlab/Qwen3-30B-A3B-Instruct-2507-AWQ
- meta-llama/Llama-3.1-405B-FP8
- openbmb/MiniCPM5-2B-MLX
- Qwen/Qwen2.5-Coder-7B-Instruct-GPTQ-Int4
- allenai/Olmo-3-7B-Instruct
- nm-testing/SmolLM-1.7B-Instruct-quantized.w4a16
- nvidia/Qwen3-14B-NVFP4
- Qwen/Qwen-Image-Edit-2509
- ByteDance-Seed/UI-TARS-1.5-7B
- unsloth/Qwen3-4B-GGUF
- nomic-ai/nomic-embed-code
- bigscience/bloom-560m
- ornith-ai/Ornith-1.5-9B
- unsloth/gemma-4-26B-A4B-it-qat-GGUF
- Qwen/Qwen3.5-397B-A17B
- swiss-ai/Apertus-8B-Instruct-2509
- nvidia/GLM-5.2-NVFP4
- moonshotai/Kimi-K2.6
- k-chirkunov/gemma4-e4b-claims-comparison
- gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090
- Qwen/Qwen2.5-Coder-7B
- google/gemma-3-12b-it
- thenlper/gte-large
- Qwen/Qwen2.5-3B-Instruct-AWQ
- google/gemma-4-E4B-it-qat-w4a16-ct
- unsloth/inkling-GGUF
- mistralai/Mistral-7B-v0.1
- RedHatAI/gemma-4-31B-it-FP8-block
- QuantStack/Wan2.2-I2V-A14B-GGUF
- sakamakismile/Qwen3.8-27B-MTP-NVFP4
- GSAI-ML/LLaDA-8B-Instruct
- Accio-Lab/occamy-1.0-GGUF
- z-lab/Qwen3.8-27B-DFlash2-GGUF
- PrimeIntellect/Qwen3-0.6B
- migtissera/Tess-4-27B-GGUF
- deepseek-ai/DeepSeek-V4-Pro
- deepseek-ai/DeepSeek-R1-Distill-Qwen-32B
- ornith-ai/Ornith-1.5-397B-NVFP4
- ukisai/Swift-Qwen3.8-27B-GGUF
- cyankiwi/Qwen3-VL-8B-Instruct-AWQ-4bit
- Qwen/Qwen3-VL-30B-A3B-Instruct
- codellama/CodeLlama-7b-hf
- google/gemma-3-27b-it
- empero-ai/Qwythos-9B-v2-GGUF
- handy-computer/Voxtral-Mini-4B-Realtime-2602-gguf
- prism-ml/Bonsai-27B-gguf
- tencent/Hy3
- eddiegulay/wav2vec2-large-xlsr-mvc-swahili
- ibm-research/PowerMoE-3b
- RedHatAI/Qwen3-32B-NVFP4
- meta-llama/Llama-2-7b-chat-hf
- Qwen/Qwen3.5-397B-A17B-FP8
- black-forest-labs/FLUX.2-dev
- Qwen/Qwen3-Coder-30B-A3B-Instruct
- Qwen/Qwen3-ForcedAligner-0.6B
- Abiray/MiniMax-H3-GGUF
- black-forest-labs/FLUX.2-klein-4B
- IlyaGusev/saiga_llama3_8b
- Qwen/Qwen2.5-Coder-1.5B
- openbmb/MiniCPM5-1B
- black-forest-labs/FLUX.1-Kontext-dev
- Jackrong/Qwopus3.8-27B-Flash-V2-GGUF
- unsloth/gemma-4-E4B-it-qat-GGUF
- Qwen/Qwen3-VL-8B-Instruct-GGUF
- microsoft/phi-4
- google/gemma-4-E2B-it-qat-q4_0-gguf
- cagliostrolab/animagine-xl-4.0
- microsoft/Florence-2-large
- LiquidAI/LFM2.5-8B-A1B-GGUF
- unsloth/gemma-4-E2B-it-GGUF
- DavidAU/Qwen3.8-27B-Cold-Fusion-GAIN-V1.1-NM-DAU-NEO-MAX-MTP-GGUF
- EleutherAI/gpt-neox-20b
- Qwen/Qwen2.5-Coder-14B-Instruct-AWQ
- peculiar-ragdoll/Cyber-Tiel-Coder-35B-A3B-GGUF-MTP
- Qwen/Qwen3-235B-A22B-Instruct-2507-FP8
- OBLITERATUS/Ornith-1.5-9B-OBLITERATED
- unsloth/LTX-2.3-GGUF
- nvidia/llama-nemotron-embed-1b-v2
- Qwen/Qwen3.5-122B-A10B-GPTQ-Int4
- deepvk/USER-bge-m3
- unsloth/mistral-7b-v0.3-bnb-4bit
- prism-ml/Ternary-Bonsai-8B-gguf
- QuantTrio/GLM-4.7-Flash-AWQ
- tiiuae/falcon-7b
- llava-hf/llava-v1.6-mistral-7b-hf
- AxionML/Qwen3.5-9B-NVFP4
- handy-computer/whisper-large-v3-turbo-gguf
- deepseek-ai/DeepSeek-V3.1
- unsloth/Ornith-1.0-9B-GGUF
- microsoft/Phi-3-mini-4k-instruct
- microsoft/Phi-3.5-mini-instruct
- Qwen/Qwen2.5-VL-3B-Instruct-AWQ
- nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-NVFP4
- bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF
- Jackrong/Qwopus3.6-35B-A3B-Coder-MTP-GGUF
- OrdalieTech/Solon-embeddings-large-0.1
- cyankiwi/Qwen3.5-4B-AWQ-4bit
- bartowski/thomsonreuters_Thomson-1.0-Small-GGUF
- unsloth/FLUX.2-klein-9B-GGUF
- bartowski/google_gemma-4-E2B-it-GGUF
- michaelw9999/Qwen3.6-35B-A3B-NVFP4-MTP-GGUF
- datalab-to/chandra
- peculiar-ragdoll/Tiel-Coder-35B-A3B-GGUF
- meta-llama/Llama-3.3-70B-Instruct
- microsoft/Phi-4-mini-instruct
- Qwen/Qwen3-0.6B-GGUF
- empero-ai/Qwen3.8-27B-Ridge-GGUF
- nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16
- typhoon-ai/typhoon-ocr1.5-2b
- tvall43/Qwen3.6-14B-A3B-FableVibes-GGUF
- AbelZimba/whisper-bemba-stt
- voyageai/voyage-4-nano
- Jackrong/Qwopus3.6-27B-Coder-Compat-MTP-GGUF
- Qwen/Qwen3-30B-A3B-Instruct-2507-FP8
- zai-org/GLM-5.1
- google/gemma-4-26B-A4B-it-qat-q4_0-gguf
- Viggle/Qwen-Image-2.1-viggle-turbo
- Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign
- nvidia/Qwen3.8-Flash-Next-NVFP4
- LiquidAI/LFM2.5-1.2B-Instruct-GGUF
- unsloth/Qwen3.5-2B-GGUF
- unsloth/FLUX.2-klein-4B-GGUF
- unsloth/Kimi-K3-GGUF
- handy-computer/whisper-large-v3-gguf
- moonshotai/Kimi-K2-Instruct
- zai-org/GLM-5.2-FP8
- ornith-ai/Ornith-1.5-397B-FP8
- sanskar003/Qwen3.5-4B-AWQ
- Qwen/Qwen3-ASR-1.7B-hf
- nvidia/NVIDIA-Nemotron-Nano-9B-v2
- moonshotai/Kimi-K2.5
- unsloth/Qwen3-VL-4B-Instruct-GGUF
- AnkitAI/Parable-Qwen3-8B-Claude-Fable-5-GGUF
- BAAI/bge-multilingual-gemma2
- deepseek-ai/DeepSeek-R1-Distill-Qwen-14B
- ibm-granite/granite-3.3-2b-instruct
- yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF
- unsloth/Qwen-Image-2512-GGUF
- alpindale/Llama-Guard-3-1B
- unsloth/gemma-4-E2B-it-qat-GGUF
- unsloth/Qwen-AgentWorld-35B-A3B-GGUF
- Qwen/Qwen3-VL-32B-Instruct
- llm-jp/llm-jp-4-33b-thinking-gguf
- unsloth/Kimi-K2.7-Code-GGUF
- bartowski/XYZAILab_XYZ-Aquila-mini-GGUF
- google/gemma-4-31B-it-qat-q4_0-unquantized
- RedHatAI/gemma-4-26B-A4B-it-NVFP4
- allenai/OLMo-2-0425-1B
- ggml-org/gemma-4-26B-A4B-it-GGUF
- Qwen/Qwen2-7B-Instruct
- Tencent-Hunyuan/HunyuanDiT-v1.1-Diffusers-Distilled
- Qwen/Qwen2.5-Coder-7B-Instruct-GGUF
- Qwen/Qwen2.5-Omni-7B
- google/gemma-4-E2B-it-qat-w4a16-ct
- abenzerps/Spark-X2.5-4B-GGUF
- QuantTrio/Qwen3-Coder-30B-A3B-Instruct-AWQ
- unsloth/gemma-4-31B-it-qat-GGUF
- Qwen/Qwen3.5-9B-Base
- peculiar-ragdoll/Sharp-Spark-X2.5-4B-GGUF
- XLabs-AI/xflux_text_encoders
- Qwen/Qwen3.5-2B-Base
- black-forest-labs/FLUX.2-klein-base-4B
- unsloth/GLM-5.2-GGUF
- agentionai/Qwen3.8-27B-AP-GGUF
- Qwen/Qwen3-VL-235B-A22B-Instruct
- bartowski/DeepSeek-Coder-V2-Lite-Instruct-GGUF
- RadixArk/Qwen3.8-Flash-Next-NVFP4
- Qwen/Qwen3.8-Flash-Next-FP8
- microsoft/harrier-oss-v1-0.6b
- poolside/Laguna-XS.2
- Qwen/Qwen2.5-1.5B-Instruct-GGUF
- meta-llama/Llama-3.2-3B
- utter-project/EuroLLM-22B-Instruct-2512
- drbaph/Higgs-Audio-v3-Studio
- janhq/Jan-v3.5-4B-gguf
- RedHatAI/gemma-3-27b-it-quantized.w4a16
- deepseek-ai/DeepSeek-R1-Distill-Qwen-7B
- unsloth/gemma-4-31B-it-GGUF
- philbert440/Qwen3.8-27B-W4A16-AWQ
- Qwen/Qwen2.5-Math-1.5B-Instruct
- RedHatAI/Qwen3-Coder-Next-FP8-dynamic
- Qwen/Qwen2.5-72B-Instruct
- incoai/Qwen3.8-27B-DFlash2
- Qwen/Qwen2.5-Omni-3B
- superwhisper/s1-mini-GGUF
- LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct
- RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead
- AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF
- bartowski/Kwaipilot_KAT-Coder-V2.5-Dev-GGUF
- mixedbread-ai/deepset-mxbai-embed-de-large-v1
- empero-ai/Qwythos-27B-v1-GGUF
- KyleHessling1/Qwopus3.6-27B-Fusion-GGUF
- poolside/Laguna-M.1
- playgroundai/playground-v2.5-1024px-aesthetic
- huggyllama/llama-7b
- nanonets/Nanonets-OCR2-3B
- thinkingmachines/Inkling-Small-NVFP4
- ornith-ai/Ornith-1.0-397B-FP8
- webAI-Official/TwIL-LM3
- google/gemma-4-31B-it-qat-q4_0-gguf
- Serveurperso/Qwen3-TTS-GGUF
- MiniMaxAI/MiniMax-M2
- Lightricks/LTX-2
- ornith-ai/Ornith-1.0-35B-FP8
- mesolitica/llama2-embedding-1b-8k
- bartowski/kai-os_Grug-12B-GGUF
- meta-models/Muse-Glimmer-30B
- openbmb/MiniCPM5-2B-GGUF
- Qwen/Qwen3-Next-80B-A3B-Instruct
- google/gemma-4-31B-it-assistant
- LilaRest/gemma-4-31B-it-NVFP4-turbo
- Qwen/Qwen3-0.6B-FP8
- Qwen/Qwen-Image
- BDRC/tibetan-ocr
- Qwen/Qwen2.5-3B
- ornith-ai/Ornith-1.5-35B-A3B
- FINAL-Bench/POCKET-26B-GGUF
- EleutherAI/pythia-410m
- AtomicChat/Ling-3.0-flash-GGUF
- ibm-granite/granite-4.1-30b
- lightonai/LightOnOCR-2-1B
- Qwen/Qwen-Image-Edit-2511
- GnLOLot/MiniCPM5-1B-Claude-Opus-Fable5-Thinking-GGUF
- bloomer010/Ling-3.0-tiny-GGUF
- MaziyarPanahi/Qwen3-0.6B-GGUF
- GnLOLot/MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-GGUF
- google/gemma-4-31B-it-qat-w4a16-ct
- poolside/Laguna-XS-2.1
- unsloth/Wan2.2-TI2V-5B-GGUF
- AngelSlim/Hy3-GGUF
- Qwen/Qwen1.5-MoE-A2.7B
- sdadas/mmlw-retrieval-roberta-large
- nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-FP8
- ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF
- stepfun-ai/Step-3.7-Flash-NVFP4
- unsloth/diffusiongemma-26B-A4B-it-GGUF
- unsloth/Qwen2.5-14B-bnb-4bit
- google/gemma-4-E2B
- ATH-MaaS/OvisOCR2
- stelterlab/Mistral-Small-24B-Instruct-2501-AWQ
- Wan-AI/Wan2.1-T2V-1.3B-Diffusers
- openbmb/MiniCPM-V-4.6
- ChristianAzinn/mxbai-embed-large-v1-gguf
- Qwen/Qwen3-4B-AWQ
- ai4bharat/indic-parler-tts
- google/paligemma-3b-pt-224
- QuantTrio/Qwen3.5-27B-AWQ
- ornith-ai/Ornith-1.5-397B
- Qwen/Qwen2.5-0.5B-Instruct-GGUF
- Jackrong/Qwopus3.6-27B-Coder-MTP-GGUF
- opendatalab/MinerU2.5-Pro-2605-1.2B
- cyankiwi/Qwen3.6-35B-A3B-AWQ-4bit
- utter-project/EuroLLM-1.7B-Instruct
- RedHatAI/Llama-3.2-1B-Instruct-FP8-dynamic
- XHToken/Spark-X2.5-1.7B-GGUF
- google/medgemma-1.5-4b-it
- InternScience/Agents-A1-4B
- unsloth/gemma-4-E2B-it-qat-mobile-GGUF
- MaziyarPanahi/Qwen3-14B-GGUF
- ai-forever/FRIDA
- unsloth/Llama-3.2-3B-Instruct
- RedHatAI/gemma-4-31B-it-FP8-dynamic
- InternScience/Agents-A1-4B-Q8_0-GGUF
- intfloat/e5-mistral-7b-instruct
- unsloth/Muse-Glimmer-30B-GGUF
- moonshotai/Kimi-VL-A3B-Instruct
- MaziyarPanahi/Qwen3-4B-GGUF
- Qwen/Qwen3-VL-32B-Instruct-FP8
- NousResearch/Meta-Llama-3.1-8B-Instruct
- zenosai/MonkeyOCRv2-B-Parsing
- QuantStack/Wan2.2-TI2V-5B-GGUF
- MaziyarPanahi/Qwen3-1.7B-GGUF
- Qwen/Qwen2.5-7B-Instruct-GPTQ-Int4
- moonshotai/Kimi-Linear-48B-A3B-Instruct
- MaziyarPanahi/Qwen3-8B-GGUF
- XiaomiMiMo/MiMo-V2.5
- jinaai/jina-embeddings-v5-text-small
- bartowski/Qwen_Qwen3-4B-GGUF
- HuggingFaceTB/SmolLM2-1.7B-Instruct
- stepfun-ai/Step-3.5-Flash
- MaziyarPanahi/Qwen3-30B-A3B-GGUF
- MaziyarPanahi/Qwen3-32B-GGUF
- skt/A.X-K2-NVFP4
- QuantTrio/Qwen3.5-4B-AWQ
- Hcompany/Holo-3.1-35B-A3B-GGUF
- inclusionAI/LLaDA2.0-mini
- Qwen/Qwen2.5-VL-32B-Instruct-AWQ
- Qwen/Qwen3-14B-GGUF
- dots-studio/dots.mocr
- AnkitAI/Parable-Granite-4.1-3B-Claude-Fable-5-GGUF
- AnkitAI/Parable-Granite-4.1-8B-Claude-Fable-5-GGUF
- nvidia/LocateAnything-3B
- unsloth/Qwen3.5-0.8B-GGUF
- Qwen/Qwen2.5-14B-Instruct-GPTQ-Int4
- poolside/Laguna-S-2.1
- nvidia/NVIDIA-Nemotron-Nano-9B-v2-FP8
- unsloth/Meta-Llama-3.1-8B-Instruct
- Zyphra/Zamba2-1.2B-instruct
- poolside/Laguna-XS-2.1-NVFP4
- MiniMaxAI/MiniMax-M3
- meta-llama/Meta-Llama-3-8B
- QCRI/Fanar-2-27B-Instruct
- NousResearch/Meta-Llama-3.1-8B
- ggml-org/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF
- openbmb/MiniCPM-o-2_6
- cyankiwi/gemma-4-12B-it-AWQ-INT4
- MaziyarPanahi/Yi-Coder-9B-Chat-GGUF
- Qwen/Qwen2-7B
- microsoft/VibeVoice-Realtime-0.5B
- cyankiwi/Qwen3-30B-A3B-Instruct-2507-AWQ-4bit
- Qwen/Qwen2.5-3B-Instruct-GGUF
- deepseek-ai/DeepSeek-R1-0528
- Qwen/Qwen3.5-35B-A3B-GPTQ-Int4
- opendatalab/MinerU2.5-Pro-2604-1.2B
- black-forest-labs/FLUX.2-klein-9B
- nvidia/Nemotron-3-Embed-1B-BF16
- poolside/Laguna-XS-2.1-GGUF
- meta-llama/Llama-2-13b-chat-hf
- bytkim/Qwen3.6-27B-MTP-pi-tune-GGUF
- Qwen/Qwen3-Omni-30B-A3B-Thinking
- poolside/Laguna-S-2.1-NVFP4
- RadixArk/Qwen3.8-27B-DSpark
- Qwen/Qwen2.5-Math-1.5B
- bartowski/Qwen2.5-7B-Instruct-GGUF
- city96/FLUX.1-dev-gguf
- deepseek-ai/DeepSeek-V2-Lite
- ReadyArt/gemma-4-31B-it-scotoma-2-GGUF
- cyankiwi/Qwen3.5-27B-AWQ-4bit
- deepseek-ai/DeepSeek-R1-Distill-Llama-8B
- antirez/qwen3.8-flash-next-gguf
- Laxhar/noobai-XL-1.1
- RedHatAI/Llama-3.2-3B-Instruct-FP8
- ukisai/Swift-1.5-Qwen3.8-27B-GGUF
- Myric/Laguna-S-2.1-APEX-GGUF
- ornith-ai/Ornith-1.0-397B
- unsloth/Qwen2.5-VL-7B-Instruct-GGUF
- bartowski/Qwen_Qwen3.6-35B-A3B-GGUF
- unsloth/Qwen2.5-Coder-7B-Instruct
- Lorbus/Qwen3.6-27B-int4-AutoRound
- chatdb/natural-sql-7b
- unsloth/Qwen2.5-7B-Instruct
- ibm-granite/granite-vision-4.1-4b
- LSX-UniWue/LLaMmlein_1B_prerelease
- Qwen/Qwen3.5-27B-FP8
- InternScience/Agents-A1-4B-Q4_K_M-GGUF
- Qwen/Qwen3Guard-Gen-0.6B
- Qwen/Qwen3-VL-4B-Instruct-FP8
- unsloth/gemma-4-E4B-it-unsloth-bnb-4bit
- leejet/MiniMax-H3-GGUF
- Qwen/Qwen3-235B-A22B-Instruct-2507
- Qwen/Qwen3-30B-A3B-FP8
- JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-Q8_0
- google/gemma-3n-E2B-it
- Qwen/Qwen3-VL-4B-Instruct-GGUF
- MaziyarPanahi/Qwen3-4B-Instruct-2507-GGUF
- meta-llama/Llama-4-Scout-17B-16E-Instruct
- omlab/VLX-Seek-1.5-10B
- google/gemma-2-2b
- nvidia/NVIDIA-Nemotron-Nano-12B-v2-VL-NVFP4-QAD
- Jackrong/Qwen3.5-9B-DeepSeek-V4-Flash-GGUF
- nvidia/Cosmos-Reason2-8B
- nvidia/Kimi-K2-Thinking-NVFP4
- bartowski/Qwen_Qwen3.5-27B-GGUF
- ibm-granite/granite-embedding-311m-multilingual-r2
- EleutherAI/pythia-1b
- ai-forever/sbert_large_nlu_ru
- zai-org/GLM-4.1V-9B-Thinking
- unsloth/GLM-4.7-Flash-GGUF
- mistral-experimental/pixtral-12b
- pfnet/plamo-embedding-1b
- hugging-quants/Meta-Llama-3.1-8B-Instruct-AWQ-INT4
- dreamgen/lucid-v1-nemo
- cyankiwi/Qwen3-VL-30B-A3B-Instruct-AWQ-4bit
- unsloth/Step-3.7-Flash-GGUF
- hugging-quants/Meta-Llama-3.1-70B-Instruct-AWQ-INT4
- z-lab/Qwen3.8-27B-DFlash2
- TIGER-Lab/VLM2Vec-Full
- boboliu/Qwen3-Embedding-4B-W4A16-G128
- QuantStack/Wan2.2-T2V-A14B-GGUF
- microsoft/Phi-3-mini-128k-instruct
- mixedbread-ai/mxbai-embed-2d-large-v1
- Jackrong/Qwopus3.5-9B-Coder-MTP-GGUF
- nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4
- datalab-to/lift
- EleutherAI/pythia-1.4b
- ai-forever/ru-en-RoSBERTa
- bartowski/gemma-2-27b-it-GGUF
- unsloth/gemma-4-26B-A4B-it-NVFP4
- batiai/Qwen3.6-27B-GGUF
- ibm-granite/granite-4.1-8b
- jinaai/jina-embeddings-v3-hf
- bartowski/Qwen_Qwen3.5-4B-GGUF
- ggml-org/gpt-oss-20b-GGUF
- bartowski/Llama-3.2-3B-Instruct-GGUF
- Wan-AI/Wan2.2-TI2V-5B-Diffusers
- wikeeyang/Flux2-Klein-9B-True-V2
- allenai/wildguard
- MaziyarPanahi/Meta-Llama-3-8B-Instruct-GGUF
- SimianLuo/LCM_Dreamshaper_v7
- MaziyarPanahi/Mistral-7B-v0.1-GGUF
- QuantTrio/Qwen3-Coder-30B-A3B-Instruct-GPTQ-Int8
- mistralai/Mistral-7B-Instruct-v0.1
- unsloth/ERNIE-Image-Turbo-GGUF
- cagliostrolab/animagine-xl-3.1
- Jackrong/Qwopus3.6-27B-v2-MTP-GGUF
- unsloth/Qwen3-VL-8B-Instruct-GGUF
- ReliquaryForge/qwen3-4b-base-dapo-v4
- nvidia/NVIDIA-Nemotron-Parse-v1.1
- MaziyarPanahi/Mistral-Small-Instruct-2409-GGUF
- MaziyarPanahi/Mistral-Nemo-Instruct-2407-GGUF
- Qwen/Qwen2.5-7B-Instruct-GGUF
- unsloth/Qwen3-14B-unsloth-bnb-4bit
- MaziyarPanahi/Yi-Coder-1.5B-Chat-GGUF
- Wan-AI/Wan2.2-I2V-A14B-Diffusers
- ggml-org/gpt-oss-120b-GGUF
- gaunernst/gemma-3-12b-it-int4-awq
- Qwen/Qwen3-1.7B-GGUF
- MaziyarPanahi/firefunction-v2-GGUF
- unsloth/Qwen3-Coder-Next-GGUF
- Qwen/Qwen1.5-1.8B-Chat
- QUASAR-QAT/Qwen3.8-27B-QUASAR-NVFP4
- unsloth/Qwen3.5-27B-GGUF
- abhishekchohan/gemma-3-12b-it-quantized-W4A16
- ukisai/Swift-1.5-Qwen3.8-Flash-Next-GGUF
- nvidia/Qwen3-8B-NVFP4
- Qwen/Qwen3-VL-8B-Thinking
- HuggingFaceTB/SmolVLM2-2.2B-Instruct
- MaziyarPanahi/Qwen2-7B-Instruct-GGUF
- unsloth/Qwen3-0.6B-GGUF
- xinsir/controlnet-openpose-sdxl-1.0
- QuantTrio/Qwen3-VL-32B-Instruct-AWQ
- leejet/Qwen-Image-2.1-GGUF
- unsloth/Llama-3.1-8B-Instruct
- microsoft/Phi-3.5-MoE-instruct
- RedHatAI/Meta-Llama-3.1-70B-Instruct-FP8
- NousResearch/Llama-2-7b-hf
- MaziyarPanahi/Phi-3.5-mini-instruct-GGUF
- bullerwins/Wan2.2-I2V-A14B-GGUF
- MaziyarPanahi/Mistral-7B-Instruct-v0.3-GGUF
- MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF
- John6666/nova-furry-xl-il-v120-sdxl
- MaziyarPanahi/Mixtral-8x22B-v0.1-GGUF
- MaziyarPanahi/phi-4-GGUF
- MaziyarPanahi/Meta-Llama-3.1-8B-Instruct-GGUF
- Wan-AI/Wan2.2-T2V-A14B-Diffusers
- intfloat/e5-large
- bartowski/Qwen_Qwen3.5-9B-GGUF
- MaziyarPanahi/gemma-3-4b-it-GGUF
- MaziyarPanahi/solar-pro-preview-instruct-GGUF
- Qwen/Qwen3-VL-2B-Instruct-GGUF
- google/gemma-2b
- allenai/OLMoE-1B-7B-0125-Instruct
- deepseek-ai/DeepSeek-V2-Lite-Chat
- MaziyarPanahi/Llama-3-8B-Instruct-32k-v0.1-GGUF
- Jackrong/Qwopus3.8-27B-Flash-GGUF
- nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark
- CohereLabs/North-Micro-Vision-Instruct
- cyankiwi/gemma-4-31B-it-AWQ-4bit
- ibm-granite/granite-4.2-8b
- city96/Wan2.1-I2V-14B-480P-gguf
- Qwen/Qwen2.5-Coder-1.5B-Instruct-GGUF
- HuggingFaceTB/SmolVLM-500M-Instruct
- LiquidAI/LFM2.5-2.6B-DSpark-GGUF
- sesame/csm-1b
- MaziyarPanahi/Qwen2.5-7B-Instruct-GGUF
- unsloth/FLUX.2-dev-GGUF
- RedHatAI/DeepSeek-Coder-V2-Lite-Instruct-FP8
- MaziyarPanahi/QwQ-32B-GGUF
- MaziyarPanahi/DeepSeek-R1-0528-Qwen3-8B-GGUF
- unsloth/gemma-3n-E4B-it
- MaziyarPanahi/Mistral-Small-24B-Instruct-2501-GGUF
- xinsir/controlnet-union-sdxl-1.0
- bartowski/Llama-3.2-1B-Instruct-GGUF
- Qwen/Qwen3-Next-80B-A3B-Instruct-FP8
- google/gemma-3-4b-pt
- MaziyarPanahi/Phi-4-mini-instruct-GGUF
- MaziyarPanahi/Llama-3.2-3B-Instruct-GGUF
- OpenMOSS-Team/MOSS-TTS-v1.5
- MaziyarPanahi/Ministral-3-3B-Reasoning-2512-GGUF
- allenai/Olmo-3-1025-7B
- MaziyarPanahi/gemma-2-2b-it-GGUF
- unsloth/Llama-3.2-1B
- MaziyarPanahi/Qwen2.5-1.5B-Instruct-GGUF
- MaziyarPanahi/gemma-3-1b-it-GGUF
- MaziyarPanahi/Llama-3.3-70B-Instruct-GGUF
- RedHatAI/Qwen2.5-1.5B-quantized.w8a8
- RedHatAI/Meta-Llama-3.1-8B-Instruct-FP8
- MaziyarPanahi/Yi-1.5-6B-Chat-GGUF
- stabilityai/stable-video-diffusion-img2vid-xt
- MaziyarPanahi/WizardLM-2-7B-GGUF
- MaziyarPanahi/gemma-3-12b-it-GGUF
- MaziyarPanahi/mistral-small-3.1-24b-instruct-2503-hf-GGUF
- moonshotai/Kimi-Linear-48B-A3B-Base
- MaziyarPanahi/Llama-3-8B-Instruct-64k-GGUF
- zai-org/GLM-4.5
- MaziyarPanahi/mathstral-7B-v0.1-GGUF
- unsloth/Z-Image-GGUF
- MaziyarPanahi/Meta-Llama-3.1-70B-Instruct-GGUF
- MaziyarPanahi/DeepSeek-V3-0324-GGUF
- MaziyarPanahi/gemma-3-27b-it-GGUF
- dragonkue/snowflake-arctic-embed-l-v2.0-ko
- Qwen/Qwen2-1.5B
- MaziyarPanahi/Mistral-Large-Instruct-2411-GGUF
- MaziyarPanahi/INTELLECT-2-GGUF
- MaziyarPanahi/Meta-Llama-3.1-405B-Instruct-GGUF
- LiquidAI/LFM2.5-2.6B
- Qwen/Qwen3-32B-FP8
- inclusionAI/LLaDA2.1-mini
- ThorOdinson246/nl2sh-1.5b-Q4_K_M
- latam-gpt/Llama-3.1-70B-LatamGPT-SFT-1.0
- SG161222/RealVisXL_V4.0
- meta-llama/Meta-Llama-3-70B-Instruct
- RedHatAI/Llama-3.3-70B-Instruct-FP8-dynamic
- deepseek-ai/deepseek-coder-6.7b-instruct
- aari1995/German_Semantic_STS_V2
- deepseek-ai/deepseek-coder-6.7b-base
- prism-ml/Bonsai-8B-gguf
- Qwen/Qwen2.5-Coder-14B-Instruct-GGUF
- unsloth/Phi-4-mini-instruct-GGUF
- Qwen/Qwen-Image-2.1
- LiquidAI/LFM2.5-1.2B-Instruct
- Snowflake/snowflake-arctic-embed-m-v2.0
- MuXodious/gpt-oss-20b-RichardErkhov-heresy
- city96/FLUX.1-schnell-gguf
- SebastianBodza/Kartoffel_Orpheus-3B_german_natural-v0.1
- Qwen/Qwen2.5-1.5B-Instruct-AWQ
- Qwen/Qwen-Image-Edit
- canopylabs/3b-de-ft-research_release
- allenai/OLMoE-1B-7B-0924
- audio-cpp/AuK-Base-and-Flash-GGUF
- QuantStack/Qwen-Image-Edit-2509-GGUF
- Alibaba-NLP/gte-Qwen2-7B-instruct
- hmellor/Ilama-3.2-1B
- tencent/Hy3-preview
- sarvamai/sarvam-30b
- unsloth/Qwen3-4B-Instruct-2507-unsloth-bnb-4bit
- KyleHessling1/Qwopus-GLM-18B-Merged-GGUF
- google/gemma-4-12B
- lightseekorg/kimi-k2.6-eagle3-mla
- Qwen/Qwen2.5-Coder-3B-Instruct-GGUF
- unsloth/Qwen3-1.7B-GGUF
- unsloth/Qwen3-Coder-Next-FP8
- Wan-AI/Wan2.1-I2V-14B-480P
- jinaai/jina-embeddings-v5-text-small-retrieval
- HuggingFaceTB/SmolLM3-3B-Base
- diffusers/stable-diffusion-xl-1.0-inpainting-0.1
- unsloth/Qwen2.5-7B-Instruct-bnb-4bit
- pentacoxian-dev/Qwen3.8-Flash-Next-IQ3E-Q8D-MTP-GGUF
- NousResearch/Meta-Llama-3-8B-Instruct
- bosonai/higgs-tts-3-4b
- Qwen/Qwen3-30B-A3B-GGUF
- SulphurAI/Sulphur-2-base
- PrimeIntellect/Qwen3-1.7B
- John6666/amanatsu-illustrious-v11-sdxl
- dealignai/GLM-5.3-CYBERSECURITY-FP8
- stabilityai/stable-diffusion-3.5-large
- stabilityai/stable-diffusion-xl-refiner-1.0
- unsloth/Llama-3.2-3B-Instruct-GGUF
- bartowski/Qwen2.5-14B-Instruct-GGUF
- DevQuasar/amd.Instella-MoE-16B-A3B-Think-GGUF
- Qwen/Qwen3-32B-GGUF
- h2oai/h2o-danube3-500m-chat
- dealignai/Bonsai-2-27B-Ternary-CRACK-GGUF
- Nanbeige/Nanbeige4.2-3B
- nvidia/NV-Embed-v2
- MaziyarPanahi/Qwen3-30B-A3B-Instruct-2507-GGUF
- littlejohn-ai/bge-m3-spa-law-qa
- meta-llama/Meta-Llama-3-70B
- JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-Q4_K_M
- Salesforce/Llama-xLAM-2-8b-fc-r
- bartowski/Meta-Llama-3.1-70B-Instruct-GGUF
- RamManavalan/Qwen3-VL-Embedding-8B-FP8
- XiaomiMiMo/MiMo-V2-Flash
- cais/HarmBench-Llama-2-13b-cls
- molbal/ideogram-4-gguf
- zai-org/GLM-4.7
- CohereLabs/c4ai-command-r-v01
- XiaomiMiMo/MiMo-V2.6-Pro-RL
- LLMSafety/Qwen2.5-Math-7B-4bit
- unsloth/DeepSeek-R1-0528-Qwen3-8B-GGUF
- black-forest-labs/FLUX.2-klein-base-9B
- unsloth/llama-3-8b-Instruct-bnb-4bit
- nvidia/diffusiongemma-26B-A4B-it-NVFP4
- HuggingFaceH4/zephyr-7b-beta
- Abiray/Qwen-Image-2.1-GGUF
- silveroxides/Chroma-GGUF
- badtheorylabs/BTL-4-Compact
- lmstudio-community/Qwen3-Coder-30B-A3B-Instruct-GGUF
- unsloth/Qwen3-8B-GGUF
- krea/Krea-2-Turbo
- hyrelabs/Homura-30B-GGUF
- stabilityai/stable-diffusion-3.5-medium
- zenmagnets/Nemotron-3-Embed-1B-Q4_K_M-GGUF
- NovaSearch/stella_en_400M_v5
- deepseek-ai/DeepSeek-V4-Pro-0813
- nvidia/GLM-5.3-Flash-NVFP4
- OnomaAIResearch/Illustrious-xl-early-release-v0
- ibm-granite/granite-speech-4.1-2b-nar
- Qwen/Qwen3-14B-Base
- unsloth/gpt-oss-20b-BF16
- baidu/ERNIE-4.5-21B-A3B-Thinking
- city96/FLUX.2-dev-gguf
- ibm-granite/granite-3.0-1b-a400m-instruct
- unsloth/Qwen2.5-3B-Instruct-unsloth-bnb-4bit
- internlm/internlm2-chat-7b
- meta-llama/Llama-3.1-70B-Instruct
- krea/Krea-2-Raw
- Qwen/Qwen2.5-14B
- unsloth/Meta-Llama-3.1-8B
- OpenMOSS-Team/MOSS-TTS-Local-Transformer-v1.5
- unsloth/Qwen2.5-7B-Instruct-unsloth-bnb-4bit
- unsloth/Qwen3-8B
- meta-llama/Llama-3.1-405B
- Aratako/MioTTS-2.6B
- ibm-granite/granite-3.3-8b-instruct
- bartowski/Qwen_Qwen3-0.6B-GGUF
- bartowski/DeepSeek-R1-Distill-Qwen-7B-GGUF
- Wan-AI/Wan2.1-T2V-14B-Diffusers
- Wan-AI/Wan2.1-I2V-14B-720P-Diffusers
- jinaai/jina-clip-v2
- Qwen/Qwen3-30B-A3B-Thinking-2507
- vantagewithai/Krea-2-Turbo-GGUF
- Qwen/Qwen3-4B-FP8
- openai-community/gpt2-xl
- Qwen/Qwen2.5-Math-7B-Instruct
- Jackrong/Qwen3.5-4B-Claude-4.6-Opus-Reasoning-Distilled-GGUF
- nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-Base-BF16
- bartowski/microsoft_Phi-4-mini-instruct-GGUF
- Dream-org/Dream-v0-Instruct-7B
- leejet/Z-Image-Turbo-GGUF
- Xenova/sweep-next-edit-1.5B
- unsloth/gpt-oss-20b-unsloth-bnb-4bit
- bartowski/Qwen2.5-1.5B-Instruct-GGUF
- unsloth/Qwen2.5-Coder-7B-Instruct-bnb-4bit
- nvidia/Llama-3_3-Nemotron-Super-49B-v1_5-FP8
- frankjoshua/novaAnimeXL_ilV140
- unsloth/Qwen-Image-GGUF
- nytopop/Qwen3-30B-A3B.w8a8
- RadixArk/GLM-5.3-NVFP4
- bartowski/gemma-2-2b-it-GGUF
- Wan-AI/Wan2.1-VACE-14B
- MaziyarPanahi/GLM-4.6V-Flash-GGUF
- lmstudio-community/Meta-Llama-3.1-8B-Instruct-GGUF
- OpenMOSS-Team/MOSS-Audio-Tokenizer
- MaziyarPanahi/gpt-oss-20b-Derestricted-GGUF
- laion/voiceclap-large-v2
- lmstudio-community/Llama-3.2-3B-Instruct-GGUF
- MaziyarPanahi/Qwen3-Coder-Next-GGUF
- Qwen/Qwen1.5-7B
- nvidia/Llama-3_3-Nemotron-Super-49B-v1_5
- unsloth/gpt-oss-120b-GGUF
- bartowski/Qwen2.5-3B-Instruct-GGUF
- 0xSero/deepseek-v4-flash-0731-spark
- MaziyarPanahi/Nemotron-Orchestrator-8B-GGUF
- fla-hub/transformer-1.3B-100B
- lmstudio-community/gemma-3-1b-it-GGUF
- MaziyarPanahi/Trinity-Mini-GGUF
- unsloth/Qwen3-32B-GGUF
- QuantStack/Wan2.1_14B_VACE-GGUF
- tencent/Hunyuan-A13B-Instruct
- FastVideo/FastWan2.2-TI2V-5B-FullAttn-Diffusers
- Wan-AI/Wan2.1-I2V-14B-480P-Diffusers
- nvidia/Nemotron-Labs-Diffusion-8B-Base
- Qwen/Qwen1.5-0.5B-Chat
- nvidia/Nemotron-Labs-Diffusion-3B
- Inferact/GLM-5.3-NVFP4
- internlm/internlm3-8b-instruct
- Lykon/dreamshaper-8
- unsloth/DeepSeek-R1-Distill-Qwen-1.5B-GGUF
- bartowski/Qwen2.5-72B-Instruct-GGUF
- LiquidAI/LFM2-1.2B
- prism-ml/Bonsai-1.7B-gguf
- unsloth/Meta-Llama-3.1-8B-Instruct-bnb-4bit
- unsloth/Qwen3-14B-GGUF
- TheBloke/Llama-2-7B-Chat-GGUF
- Qwen/Qwen-7B
- bartowski/meta-llama_Llama-4-Scout-17B-16E-Instruct-old-GGUF
- TinyLlama/TinyLlama-1.1B-intermediate-step-1431k-3T
- John6666/diving-illustrious-real-asian-v50-sdxl
- LiquidAI/LFM2.5-8B-A1B-DSpark-GGUF
- tobil/qmd-query-expansion-1.7B-gguf
- LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-GGUF
- sbintuitions/sarashina2.2-0.5b-instruct-v0.1
- stabilityai/stable-diffusion-3-medium-diffusers
- agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF
- moonshotai/Kimi-K2-Instruct-0905
- CMSManhattan/JiRackUltra_7b
- NousResearch/Meta-Llama-3-8B
- CMSManhattan/JiRackUltra_32b
- joeygambino/MiniMax-H3-encoder-GGUF
- nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8
- ibm-granite/granite-4.1-8b-fp8
- moonshotai/Kimi-K2-Thinking
- Qwen/Qwen3Guard-Gen-4B
- Alittlehammmer/Qwen3.6-35B-A3B-DFlash-GGUF-llama.cpp
- black-forest-labs/FLUX.2-klein-9b-kv
- unsloth/Qwen3-1.7B-unsloth-bnb-4bit
- Qwen/Qwen3-30B-A3B-GPTQ-Int4
- jdopensource/JoyAI-Image-Edit-Diffusers
- unsloth/Qwen3-Next-80B-A3B-Instruct-GGUF
- RedHatAI/Llama-3.2-1B-Instruct-FP8
- unsloth/Llama-3.2-1B-Instruct-bnb-4bit
- hesamation/Qwen3.6-35B-A3B-Claude-4.6-Opus-Reasoning-Distilled-GGUF
- deepseek-ai/DeepSeek-R1-Distill-Llama-70B
- arianraje/qwen3-4b-mamba3-hybrid-stage2b-kd-bias0
- AtomicChat/Qwen3-4B-DFlash-GGUF
- molbal/MiniMax-H3-GGUF
- John6666/one-obsession-17-red-sdxl
- lmstudio-community/DeepSeek-R1-0528-Qwen3-8B-GGUF
- AtomicChat/Qwen3.5-9B-DFlash-GGUF
- casperhansen/llama-3-8b-instruct-awq
- microsoft/Phi-tiny-MoE-instruct
- AtomicChat/Qwen3.5-4B-DFlash-GGUF
- allenai/OLMo-2-0425-1B-Instruct
- Alittlehammmer/Qwen3.6-27B-DFlash-GGUF-llama.cpp
- microsoft/phi-1_5
- martineux/dvine82-xl
- kmhf/hf-moshiko
- RadixArk/Inkling-Small-DSpark-Preview
- bullerwins/FLUX.1-Kontext-dev-GGUF
- Efficient-Large-Model/gemma-2-2b-it
- DevQuasar-13/THUDM.GLM-Z1-32B-0414-GGUF
- Wan-AI/Wan2.1-T2V-14B
- perplexity-ai/pplx-embed-v1-0.6b
- solidrust/Hermes-3-Llama-3.1-8B-AWQ
- bartowski/Ling-3.0-tiny-GGUF
- lodestones/Chroma1-HD
- bartowski/Mistral-7B-Instruct-v0.3-GGUF
- sahilchachra/gemma-4-12B-coder-fable5-composer2.5-AWQ
- openthaigpt/openthaigpt1.5-7b-instruct
- QuantFactory/Qwen2.5-Coder-7B-GGUF
- RavichandranJ/Dolphin3-Cyber-8B-GGUF
- ibm-granite/granite-4.0-h-tiny
- tiiuae/Falcon-H1-0.5B-Base
- hugging-quants/Llama-3.2-1B-Instruct-Q4_K_M-GGUF
- lmstudio-community/Qwen3-4B-Instruct-2507-GGUF
- John6666/obsession-illustriousxl-v10-sdxl
- HuggingFaceTB/SmolLM-1.7B
- OpenOneRec/OneRec-1.7B
- unsloth/gemma-3-1b-it
- nvidia/MiniMax-M3-NVFP4
- microsoft/Phi-4-mini-reasoning
- XCurOS/XCurOS0.1-8B-Instruct
- bartowski/Qwen2.5-Coder-7B-Instruct-GGUF
- unsloth/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF
- avatargrim/Qwen2.5-7B_pyuigpt
- XiaomiMiMo/MiMo-V2.6-Flash-RL
- Qwen/Qwen3-14B-FP8
- unsloth/Qwen3-4B-unsloth-bnb-4bit
- SG161222/RealVisXL_V5.0_Lightning
- zai-org/GLM-4.5-Air-FP8
- ibm-granite/granite-guardian-4.1-8b
- deepseek-ai/DeepSeek-V3.2-Exp
- ibm-granite/granite-3.1-8b-instruct
- google/gemma-3-1b-pt
- LGAI-EXAONE/EXAONE-3.5-2.4B-Instruct
- Tongyi-MAI/Z-Image
- Abiray/LTX-2.5-Distilled-GGUF
- HiDream-ai/HiDream-I1-Fast
- ibm-granite/granite-4.0-tiny-preview
- unsloth/DeepSeek-R1-Distill-Llama-70B-GGUF
- nvidia/gpt-oss-120b-Eagle3-short-context
- Qwen/Qwen3-30B-A3B-Base
- allenai/OLMo-1B-hf
- MaziyarPanahi/Qwen3-4B-Thinking-2507-GGUF
- MaziyarPanahi/Ministral-3-14B-Reasoning-2512-GGUF
- tencent/Hy3-FP8
- MaziyarPanahi/NVIDIA-Nemotron-Nano-12B-v2-GGUF
- ibm-granite/granite-3.2-8b-instruct
- leejet/FLUX.2-klein-4B-GGUF
- NousResearch/Meta-Llama-3-70B-Instruct
- Lykon/dreamshaper-xl-v2-turbo
- Qwen/Qwen-7B-Chat
- nvidia/gpt-oss-120b-Eagle3-v3
- entrick/Security-SLM-Gemma-4-E2B-it-GGUF
- stabilityai/stable-diffusion-3.5-large-controlnet-canny
- microsoft/Phi-3-vision-128k-instruct
- ideogram-ai/ideogram-4-fp8
- nvidia/Nemotron-H-8B-Base-8K
- TheBloke/Mistral-7B-Instruct-v0.2-GGUF
- bartowski/Phi-3.5-mini-instruct-GGUF
- XingChen-AGI/Xing4.0-29B-A4B
- abenzerps/Nex-N2.5-mini-GGUF
- tencent/Hunyuan-7B-Instruct
- google/codegemma-7b-it
- Abiray/Qwen-Image-2.1-viggle-4-steps-turbo-GGUF
- openbmb/MiniCPM5-1B-MLX
- stabilityai/stable-diffusion-3.5-large-controlnet-depth
- ByteDance-Seed/Seed-OSS-36B-Instruct
- unsloth/Qwen2.5-32B-Instruct-bnb-4bit
- ibnzterrell/Meta-Llama-3.3-70B-Instruct-AWQ-INT4
- Wan-AI/Wan2.1-T2V-1.3B
- unsloth/Llama-3.3-70B-Instruct
- Qwen/Qwen2.5-7B-Instruct-1M
- nvidia/Llama-3.1-8B-Instruct-FP8
- QuantStack/FLUX.1-Kontext-dev-GGUF
- vcruz305/DeepSeek-V4.1-Flash-GGUF
- unsloth/Llama-3.2-3B-Instruct-unsloth-bnb-4bit
- ibm-granite/granite-4.2-3b
- bartowski/Altworld_Hemmingway-1-GGUF
- NousResearch/Llama-3.2-1B
- bigcode/starcoder2-3b
- cagliostrolab/animagine-xl-3.0
- nvidia/Mistral-NeMo-Minitron-8B-Instruct
- Vikhrmodels/Vikhr-Nemo-12B-Instruct-R-21-09-24
- wikeeyang/Flux2-Klein-9B-True-V3
- bartowski/Qwen_Qwen3-4B-Instruct-2507-GGUF
- IPostYellow/TurboWan2.1-T2V-1.3B-Diffusers
- NousResearch/Hermes-2-Pro-Mistral-7B
- John6666/prefect-illustrious-xl-v3-sdxl
- EldanRing/Winnow-12B
- z-lab/LLaMA3.1-8B-Instruct-DFlash-UltraChat
- bartowski/Qwen_Qwen3-1.7B-GGUF
- google/gemma-1.1-2b-it
- bartowski/Qwen2.5-Coder-1.5B-Instruct-GGUF
- Lightricks/LTX-2.5-Diffusers
- EleutherAI/pythia-410m-deduped
- unsloth/gemma-3-1b-it-GGUF
- unsloth/Qwen2.5-1.5B-Instruct-unsloth-bnb-4bit
- KBlueLeaf/TIPO-500M-ft
- LGAI-EXAONE/EXAONE-3.5-32B-Instruct-AWQ
- bartowski/THUDM_GLM-4-32B-0414-GGUF
- incoai/GLM-5.3-Flash-DFlash2
- openbmb/MiniCPM5-1B-GGUF
- Qwen/Qwen-Image-2512
- HuggingFaceTB/SmolLM2-1.7B
- RedHatAI/Meta-Llama-3.1-8B-Instruct-quantized.w4a16
- speakleash/Bielik-11B-v3.0-Instruct-awq
- z-lab/Qwen3.6-35B-A3B-DFlash
- lmstudio-community/Qwen2.5-Coder-14B-Instruct-GGUF
- google/t5gemma-2b-2b-ul2-it
- duyntnet/Chroma-GGUF-All-Versions
- ibm-granite/granite-4.0-1b-base
- unsloth/Qwen3-0.6B-unsloth-bnb-4bit
- microsoft/Phi-3-mini-4k-instruct-gguf
- FINAL-Bench/POCKET-Darwin-180B-GGUF
- ibm-granite/granite-3b-code-base-2k
- unsloth/Qwen2.5-14B-Instruct
- allenai/Olmo-3-32B-Think-SFT
- TaoLiveAIGC/TLive-Omni-4B
- unsloth/Llama-3.2-1B-Instruct-GGUF
- Qwen/QwQ-32B
- meta-llama/Llama-Guard-3-8B
- bartowski/DeepSeek-R1-Distill-Qwen-14B-GGUF
- unsloth/Llama-3.3-70B-Instruct-GGUF
- bartowski/Qwen2.5-Coder-14B-Instruct-GGUF
- black-forest-labs/FLUX.1-Krea-dev
- jayn7/Z-Image-Turbo-GGUF
- QuantFactory/Meta-Llama-3-8B-Instruct-GGUF
- Qwen/Qwen3-Next-80B-A3B-Thinking
- empero-ai/Qwen3.8-4B-Distill
- peculiar-ragdoll/Cyber-Tiel-Coder-35B-A3B-GGUF
- mudler/Qwen3.5-35B-A3B-APEX-GGUF
- Qwen/Qwen2.5-Math-7B
- TokenRhythm/NeoHorse-1-4B-GGUF
- second-state/stable-diffusion-v1-5-GGUF
- meta-llama/Llama-3.1-70B
- FireRedTeam/FireRed-Image-Edit-1.0
- unsloth/DeepSeek-R1-Distill-Llama-8B-GGUF
- bartowski/Qwen_Qwen3-8B-GGUF
- ggml-org/Qwen3-0.6B-GGUF
- deepseek-ai/deepseek-coder-1.3b-instruct
- swiss-ai/Apertus-8B-2509
- EleutherAI/gpt-neo-1.3B
- dealignai/Bonsai-2-27B-1bit-CRACK-GGUF
- Edge0/Audio8-ASR-Infinite
- QuantStack/Qwen-Image-Edit-GGUF
- inclusionAI/Ling-3.0-tiny-GGUF
- neuralcrew/neutrino-instruct
- nvidia/DeepSeek-V4-Flash-0731-NVFP4
- useful-quants/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-W4A16
- MeiGen-AI/PosterOmni_v1
- shafire/Zero-Gemma4-E4B-OpenZero-GGUF
- unsloth/FLUX.2-klein-base-9B-GGUF
- Laxhar/noobai-XL-1.0
- Laxhar/noobai-XL-Vpred-1.0
- unsloth/Llama-3.2-3B-Instruct-bnb-4bit
- Qwen/Qwen1.5-0.5B
- zai-org/GLM-5.3-BF16
- ibm-granite/granite-4.2-30b
- unsloth/Qwen2.5-3B-Instruct-bnb-4bit
- bartowski/DeepSeek-R1-Distill-Qwen-32B-GGUF
- nvidia/Llama-3_3-Nemotron-Super-49B-v1
- xinsir/controlnet-canny-sdxl-1.0
- Phil2Sat/Qwen-Image-Edit-Rapid-AIO-GGUF
- skt/A.X-4.0-Light
- Qwen/Qwen3-235B-A22B-Thinking-2507-FP8
- Qwen/Qwen3-Coder-480B-A35B-Instruct
- hugging-quants/Llama-3.2-3B-Instruct-Q4_K_M-GGUF
- stabilityai/stablelm-3b-4e1t
- audio-cpp/LiveAvatar-GGUF
- pipecat-ai/phonellm-alpha-1
- Qwen/Qwen3Guard-Gen-8B
- legraphista/glm-4-9b-chat-IMat-GGUF
- diffusers/controlnet-depth-sdxl-1.0
- TheBloke/Mistral-7B-Instruct-v0.1-GGUF
- bartowski/NousResearch_Hermes-4-14B-GGUF
- prithivMLmods/Qwen-Image-2.1-PE-T2I-GGUF
- solidrust/Mistral-7B-Instruct-v0.3-AWQ
- Qwen/Qwen2.5-Coder-3B
- John6666/hassaku-xl-illustrious-v31-sdxl
- openai/gpt-oss-safeguard-20b
- osllmai-community/Llama-3.2-1B
- unsloth/FLUX.1-dev-GGUF
- stabilityai/stable-video-diffusion-img2vid
- fdtn-ai/antares-1b
- ai21labs/AI21-Jamba-Reasoning-3B
- allenai/Olmo-3-7B-Instruct-SFT
- Venastine-Research/Xing4.0-29B-A4B-GGUF
- unsloth/LFM2.5-1.2B-Instruct-GGUF
- tiiuae/falcon-7b-instruct
- allenai/OLMo-2-1124-7B-Instruct
- Lykon/dreamshaper-xl-lightning
- OBLITERATUS/gemma-4-E4B-it-OBLITERATED
- lmstudio-community/Qwen3-14B-GGUF
- Agnuxo/CAJAL-4B
- outsourc-e/Qwen3.8-27B-Unleashed-GGUF
- PixArt-alpha/PixArt-Sigma-XL-2-1024-MS
- sakamakismile/Qwen3.6-27B-Text-NVFP4-MTP
- upstage/SOLAR-10.7B-Instruct-v1.0
- unsloth/llama-3-8b-Instruct
- bg-digitalservices/Gemma-4-26B-A4B-it-NVFP4
- bartowski/L3-8B-Stheno-v3.2-GGUF
- google/gemma-2-9b
- zai-org/glm-edge-4b-chat-gguf
- Manojb/stable-diffusion-2-1-base
- meta-llama/Llama-Guard-3-1B
- nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Base-BF16
- realrebelai/LTX-2.5_GGUFs
- magespace/Wan2.2-I2V-A14B-Lightning-Diffusers
- TinyLlama/TinyLlama-1.1B-step-50K-105b
- SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16
- stelterlab/Qwen3-Coder-30B-A3B-Instruct-AWQ
- GritLM/GritLM-7B
- diffusers/controlnet-canny-sdxl-1.0
- city96/Qwen-Image-gguf
- QuantStack/LTX-2.3-GGUF
- nvidia/GLM-5.3-NVFP4
- autotrust/GLM-5.3-Flash-GGUF-DGX-Spark
- chfm/Qwen-Image-2.1-GGUF
- IFM/K2-Horizon-MoVA-36B-A4B-GGUF
- Qwen/Qwen3.8-2.4T-A95B
- nvidia/llama-nemotron-embed-vl-1b-v2
- fishaudio/s2-pro
- sudoingx/Ternary-Bonsai-2-27B-PTQ1_0-MTP-GGUF
- joeygambino/MiniMax-H3-curve-GGUF
- unsloth/FLUX.1-schnell-GGUF
- rectangleworm/ideogram-4-gguf
- XHToken/Spark-X2.5-4B
- ai4bharat/IndicF5
- perplexity-ai/pplx-embed-v1-4b
- city96/Wan2.1-I2V-14B-720P-gguf
- realrebelai/SCAIL-2_GGUF
- bartowski/Cloudflare_clef-GGUF
- jayn7/WAN2.2-I2V_A14B-DISTILL-LIGHTX2V-4STEP-GGUF
- IFM/K2-Horizon-7B
- DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF
- QuantStack/Wan2.2-S2V-14B-GGUF
- bartowski/TheDrummer_Artemis-31B-v1.2-GGUF
- Hob-forge/Kolibri-1-GGUF
- ggml-org/Clef-Flash-GGUF
- bartowski/Cloudflare_clef-flash-GGUF
- city96/Wan2.1-FLF2V-14B-720P-gguf
- nvidia/Cosmos3-Super-Image2Video
- LiquidAI/LFM2.5-8B-A1B
- city96/Wan2.1-T2V-14B-gguf
- realrebelai/Wan-Animate-2_GGUFs
- >-
deadbydawn101/RavenX-CyberAgent-Qwen3.6-35B-A3B-Opus-4.7-OpenMythos-Pentester-BugHunter-RATH-GGUF
- jialinyyzz/humanizer
- IFM/K2-Horizon-375B-A23B
- hum-ma/Wan2.2-TI2V-5B-Turbo-GGUF
- Wan-AI/Wan2.1-VACE-1.3B
- IFM/K2-Horizon-3.7B
- unsloth/LTX-2-GGUF
- vantagewithai/SCAIL-2-GGUF-ComfyUI
- bartowski/FrogNano-4B-2609-GGUF
- google/embeddinggemma-2
- netease-youdao/Confucius4-R2T2
- ggml-org/Qwen3.8-Flash-Next-GGUF
- Abiray/MiniMax-H3-Singularity-GGUF
- Wan-AI/Wan2.1-I2V-14B-720P
- m-a-p/MERT-v2-FullSong
- XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
- robbyant/lingbot-world-fast
- calcuis/wan-1.3b-gguf
- bytkim/Qwen3.8-27B-pi-GGUF
- IFM/K2-Horizon-MoVA-36B-A4B
- tencent/WeMM-Embedding-2B
- Cloudflare/clef-flash
- ukisai/Swift-Bonsai-2-GGUF
- paradigma-inc/limite-1b-violetto
- zai-org/CogVideoX-2b
- zai-org/CogVideoX-5b
- bullerwins/Wan2.2-T2V-A14B-GGUF
- Rabinovich/LongLive-2.0-5B-Diffusers
- inclusionAI/Ling-3.0-tiny
- TaichuAI/ZDTaichu5.0-9B
- ggml-org/Clef-GGUF
- BreezeBlue/Breeze-TTS-2
- inclusionAI/Ling-3.0-flash
- Altworld/Hemmingway-1
- Cloudflare/clef
- moondream/parakeet-ultra
- Kwaipilot/KAT-Coder-V2.5-Dev
- nex-agi/Nex-N2.5-mini
- nvidia/NVIDIA-Nemotron-3-Nano-4B-GGUF
- CohereLabs/North-Mini-Code-1.0
- jinaai/jina-embeddings-v5-omni-nano
- LiquidAI/d1-omni-600M
- terrorswift/REDCELL-26B-A4B-OSINT-Cyber-APEX-GGUF
- dphn/Dolphin-Mistral-24B-Venice-Edition
- DavidAU/LFM2.5-8B-A1B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF
- peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF
- autotrust/GLM-5.3-GGUF-DGX-Spark
- empero-ai/Qwen3.8-35B-A3B-Distill
- Aleph-Alpha/Kolibri-1
- empero-ai/Qwythos-9B-Claude-Mythos-5-1M
- XiaomiMiMo/MiMo-V2.6-Flash-MOPD
- elix3r/gemma4-12b-with-proj-ltx-2.5-GGUF
- albucino/Qwen3.8-Flash-Next-W4A16-FP8PLE
- autotrust/GLM5.3-Flash-E224-DGX-Spark
- LiquidAI/d1-3B
- SkyIsNotGreen/Scion-35B-A3B
- microsoft/harrier-oss-v1-27b
- ukisai/Swift-1.5-Qwen3.8-27b
- orcarouter/OrcaSAQ-2-27B
- yandex/AliceAI-Foundation-80B-A3B-Base
- ggml-org/OpenJev-GGUF
- DeepHat/DeepHat-V1-7B
- pfeifferj/Ornith-1.5-35B-A3B-GSQ-RCO-GGUF
- SeerRay-Lab/Xiaomi-OCR-0
- Qwen/Qwen3Guard-Stream-0.6B
- KittenML/kitten-tts-2
- llm-jp/llm-jp-4.1-8b-thinking-gguf
- alesha-pro/Qwen3.8-27B-S-mirai-GGUF
- webAI-Official/TwIL-LM3-Pro
- XiaomiMiMo/MiMo-V2.6-Pro-MOPD
- apodex/Apodex-1.1-mini
- tencent/HunyuanImage-3.0
- logic65/Whittle-Qwen-3.8-35B-A3B
- rmonsurate/Victoria
- Lythri/Lythri-4B-A2B
- Mia-AiLab/GLM-5.3-Flash-EXL3-4bpw-TensorFold
- stabilityai/stable-video-diffusion-img2vid-xt-1-1
- emperorofrome/Gmcoder
- InternScience/Agents-A1
- aj9o9/Qwen3.8-27B-Escha-W2-GGUF
- OrionLLM/OxCoder-9B
- tsinghua-sigs-robot-lab/VeriLoop-E2
- jialinyyzz/humanizer-GGUF
- ArliAI/GLM-4.6-Derestricted-v3
- MaziyarPanahi/Llama-3-Groq-8B-Tool-Use-GGUF
- LiquidAI/d1-3B-GGUF
- AikidoSec/altar-1
- Mia-AiLab/GLM-5.3-Flash-EXL3-4bpw-TensorFold-Ablit
- bottlecapai/ThinkingCap-Qwen3.8-27B
- diffbot/DeepSeek-V4.1-Flash-EXL3-3bpw-2x-RTX-PRO-6000
- bytkim/Qwen3.8-27B-pi
- hiwaifu-research/WaifuGemma4-26b-a4b-v1
- FINAL-Bench/Darwin-27B-RSI-GGUF
- Blackfrost-AI/CYBER-FROST-3.8-BF16
- IFM/K2-Type-0.9B
- BAAI/AREX-2
- RedHatAI/Qwen3.8-Flash-Next-NVFP4
- unsloth/embeddinggemma-2
- IQuestLab/IQuest-Q1
- topk-io/topk-embed-v1-xsmall
- vcruz305/CYBER-FROST-3.8-EXL3-SAGE-3.87bpw
- Aleph-Alpha/Kolibri-1-BF16
- Eliasfpv28/Kolibri-1-Q3_K_S-GGUF
- CaryPalmer/Ternary-Bonsai-2-27B-262k-GGUF
- llm-jp/llm-jp-4.1-33b-thinking
- BuzzASR/persian
- cantina-security/apex-flash-1
- Blackfrost-Research/GLM-5.3-F.U-AnthraClaud-Edition-BF16
- troed/Qwen3.8-27B-ASCII-Condensed
- Abiray/Qwen-Image-2.1-viggle-turbo-v0.3-6step-GGUF
- >-
nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2
- perplexity-ai/pplx-embed-v2-context-9b-preview
- FINAL-Bench/Darwin-27B-RSI
- lmstudio-community/Llama-3-Groq-70B-Tool-Use-GGUF
- mehdi-hf/nemotron-asr-streaming-farsi
- Gryphe/Gemma-4-26B-A4B-StyleTune-V2
- EschaLabs/Qwen3.8-27B-Escha-W2
- nvidia/NVIDIA-NemotronLabs-AI-for-Media-Sports-Tennis
- FINAL-Bench/Darwin-180B-RSI
- Hcompany/Holo4-27B
- Blackfrost-AI/CYBER-FROST-3.8-NVFP4-V2
- ISTA-DASLab/Qwen3.8-Flash-Next-P48NVFP4-MoESQ
- LiquidAI/LFM2-1.2B-RAG
- vcruz305/RED-SNOW-5.3-FLASH-EXL3-SAGE-4.91bpw
- RedHatAI/DeepSeek-V4-Flash-0731-NVFP4
- kova-ai/kova-tts-1
- PleIAs/baguettotron-600m
- elyza/ELYZA-Thinking-1.0-llm-jp-4-33b
- elyza/ELYZA-Thinking-1.0-llm-jp-4-32b-a3b
- yamz-labs/GLM-5.3-Flash-EXL3-Yamz
- bitlamas/Qwen3.8-Flash-Next-Q4_K_XL-DN4
- kyutai/glm-4-voice-of-reason-9b
- YOON1v/Apex-2
- perplexity-ai/pplx-embed-v2-late-0.6b
- cklxx/laya-browser
- amazon/ALoDLM-8B
- Ateron/Gemma-4-Dark-Thoughts-V2-31B
- briaai/fibo-scene-analyzer
- Blockway/Agens-Volundr-32B-Preview
- WonseokJayJung/Connect-C1-0.8B-v0.1-GGUF
- kandinskylab/Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers
- kandinskylab/Kandinsky-6.0-Pro-distill-5s-Diffusers
- Remek/basal-1.5-4.5B
- perplexity-ai/pplx-embed-v2-late-9b
- kandinskylab/Kandinsky-6.0-VSR-5s-Diffusers
- Maincode/matilda-jev-v1
- FINAL-Bench/Darwin-180B-RSI-R3
- vcruz305/RED-SNOW-5.3-FLASH-EXL3-SAGE-2.49bpw
- h2oai/h2o-lightning-4b
- Blackfrost-AI/RED-SNOW-5.3-FLASH-BF16
- kandinskylab/Kandinsky-6.0-Lite-distill-5s-Diffusers
- kandinskylab/Kandinsky-6.0-Pro-5s-Diffusers
- Ateron/Gemma-4-Writers-31B-V2
- audreyt/Kolibri-1-NVFP4-W4A16
- lighteternal/biodecision-v2-4b
- malteos/most-embed-de
- kandinskylab/Kandinsky-6.0-Lite-5s-Diffusers
- ConwayResearch/Underdog-Saluki-27B-1.0
- PleIAs/Baguettotron-MoE
- JetBrains/Mellum2.1-12B-A2.5B-Thinking
- kandinskylab/Kandinsky-6.0-Pro-pretrain-5s-Diffusers
- kandinskylab/Kandinsky-6.0-Lite-pretrain-5s-Diffusers
- aina-tech/Anima-Lightning
- amazon/ALoDLM-1.7B
- facebook/meta-encoder
- speridlabs/iris-3b
- Modularcomputing/Native-Bird
- lightonai/LightOnOCR-3-0.8B
- lightonai/LightOnOCR-3-4B
- ApolloRaines/LTX-2.5-22b-OmniGen-v12
- JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF
- Kujira/Underdog-Saluki-27B-1.0-MTP-GGUF
- DavidAU/Qwen3.8-27B-Turbo-Brilliance-Power-35X-Reasoning-Instruct-modes-GGUF
- inclusionAI/Ming-Image-0.1-Design
- Aratako/Irodori-TTS-v4-Large
VRAM & GPU Cost Calculator
Pick any Hugging Face model and see:
- how much GPU memory it needs to run, fine-tune with LoRA, fully fine-tune, generate, transcribe or embed, at its published precision or at 16, 8 or 4 bits;
- the cheapest GPU setups that hold it (one H100, two A100s, four RTX 4090s...), each priced live across the GPU clouds FastGPU tracks, with a link to every offer.
It reads each model's own metadata from the Hub (the parameter count in its safetensors or GGUF header, its quantization config) and asks FastGPU's free match API for the sizing and the live prices. Nothing to install, no sign-in. Type a size such as 70B for a model that is not on the Hub.
How the memory is sized
The same way as on fastgpu.co, by FastGPU's matcher:
- Run / serve: the weights (parameters × bytes per parameter: 2 at 16-bit, 1 at 8-bit, 0.5 at 4-bit), plus the KV cache for one request at a 4K-token context and runtime headroom.
- Fine-tune with LoRA: the frozen weights plus the adapters, their optimizer state and activations.
- Full fine-tune: about 16 bytes per parameter (weights, gradients and Adam states) plus activations.
- Image, video, speech and embedding models FastGPU knows (FLUX, SDXL, Wan, Whisper and more) use a figure for the whole pipeline, text encoders and VAE included.
A quantized repo (AWQ, GPTQ, FP8, NVFP4, bitsandbytes, GGUF) is sized at the bits it is stored in unless you pick another precision. The full method: fastgpu.co/methodology.
Where the prices come from
FastGPU collects the GPU rental prices providers publish (marketplaces, neoclouds and hyperscalers) and refreshes them through the day. Each price here is the provider's own per-GPU rate times the number of GPUs in the setup. "Lowest price" (the default) sorts by price alone. "Best match" orders by FastGPU's score, led by price and weighing reliability and availability; providers that pay FastGPU a referral fee are marked ★ and can only win a near-tie there, within a few percent of the cheapest.
The price history is open: fastgpu/cloud-gpu-prices (CC BY 4.0, updated daily).
Use it from code
curl "https://fastgpu.co/api/v1/match?model=Qwen/Qwen3-32B¶ms_b=32.8&task=inference"
No key needed. Docs: fastgpu.co/docs.