--- title: VRAM & GPU Cost Calculator emoji: ⚡ colorFrom: blue colorTo: indigo sdk: static pinned: true license: mit short_description: VRAM any HF model needs + live GPU rental prices tags: - vram - gpu - gpu-prices - calculator - inference - fine-tuning - llm - cloud-gpu models: - Qwen/Qwen3-0.6B - google/gemma-4-26B-A4B-it - Qwen/Qwen3-8B - BAAI/bge-large-en-v1.5 - Qwen/Qwen3-VL-8B-Instruct - Qwen/Qwen3-4B - google/gemma-4-31B-it - Qwen/Qwen3-Embedding-0.6B - Qwen/Qwen3.5-9B - Qwen/Qwen2.5-7B-Instruct - Qwen/Qwen3.5-4B - meta-llama/Llama-3.2-1B-Instruct - intfloat/multilingual-e5-large - Qwen/Qwen3.8-27B - Qwen/Qwen2.5-1.5B-Instruct - openai/whisper-large-v3-turbo - zai-org/GLM-5.3-Flash - RadixArk/Kimi-K3-DSpark - meta-llama/Llama-3.1-8B-Instruct - openai/gpt-oss-20b - unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF - Qwen/Qwen2.5-VL-7B-Instruct - nvidia/Qwen3.6-35B-A3B-NVFP4 - ornith-ai/Ornith-1.5-9B-GGUF - Qwen/Qwen3-14B - dphn/dolphin-2.9.1-yi-1.5-34b - Qwen/Qwen3.8-27B-FP8 - deepseek-ai/DeepSeek-V4-Flash-0731 - Qwen/Qwen3.6-35B-A3B-FP8 - google/gemma-4-E4B-it - stabilityai/stable-diffusion-xl-base-1.0 - deepseek-ai/DeepSeek-V3.2 - prism-ml/Ternary-Bonsai-2-27B-gguf - Qwen/Qwen2.5-3B-Instruct - openai/gpt-oss-120b - Qwen/Qwen3.5-2B - openai/whisper-large-v3 - cdiamond/Qwen3.8-27B-iMatrix-NVFP4-MTP-GGUF - ornith-ai/Ornith-1.5-35B-A3B-GGUF - Qwen/Qwen3-32B - Qwen/Qwen3-4B-Instruct-2507 - google/embeddinggemma-300m - Qwen/Qwen3.6-35B-A3B - ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF - openai/whisper-small - Qwen/Qwen3-VL-4B-Instruct - google/gemma-3-1b-it - Qwen/Qwen3-1.7B - Qwen/Qwen3-14B-AWQ - google/gemma-4-E2B-it - Qwen/Qwen3-VL-2B-Instruct - Qwen/Qwen2.5-7B-Instruct-AWQ - MahmoudAshraf/mms-300m-1130-forced-aligner - Qwen/Qwen3.5-0.8B - CMSManhattan/JiRackUltra_1b - ornith-ai/Ornith-1.0-9B-GGUF - Qwen/Qwen3-Embedding-8B - Qwen/Qwen3.6-27B-FP8 - BAAI/bge-reranker-large - Qwen/Qwen3-8B-AWQ - Qwen/Qwen3.6-27B - Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice - Qwen/Qwen2.5-VL-3B-Instruct - vikhyatk/moondream2 - Qwen/Qwen2.5-Coder-7B-Instruct - deepseek-ai/DeepSeek-OCR - mixedbread-ai/mxbai-embed-large-v1 - jinaai/jina-embeddings-v3 - nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16 - mistralai/Voxtral-Mini-4B-Realtime-2602 - Qwen/Qwen3.5-27B - handy-computer/nemotron-3.5-asr-streaming-0.6b-gguf - ggml-org/Qwen3.8-27B-GGUF - datalab-to/chandra-ocr-2 - Qwen/Qwen2.5-32B-Instruct - Qwen/Qwen3-Embedding-4B - handy-computer/parakeet-unified-en-0.6b-gguf - farbodtavakkoli/OTel-2.0-LLM-31B-IT - Qwen/Qwen3-VL-8B-Instruct-FP8 - peculiar-ragdoll/Tiel-Coder-35B-A3B-GGUF-MTP - Qwen/Qwen2.5-14B-Instruct-AWQ - google/gemma-4-12B-it - Qwen/Qwen2.5-VL-7B-Instruct-AWQ - Qwen/Qwen3.8-Flash-Next - ornith-ai/Ornith-1.0-35B-GGUF - stable-diffusion-v1-5/stable-diffusion-v1-5 - unsloth/gemma-4-12b-it-GGUF - Qwen/Qwen2-VL-7B-Instruct-AWQ - Qwen/Qwen2.5-Coder-14B-Instruct - mistralai/Mistral-7B-Instruct-v0.2 - zai-org/GLM-4.7-Flash - meta-llama/Llama-3.2-3B-Instruct - zai-org/GLM-5.3 - intfloat/multilingual-e5-large-instruct - autotrust/JEV-27B-VL - ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF - nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4 - Alibaba-NLP/gte-multilingual-base - Qwen/Qwen3-VL-Embedding-8B - facebook/w2v-bert-2.0 - Qwen/Qwen3.5-35B-A3B - peculiar-ragdoll/Dirk-Qwen3.8-27B-GGUF - ornith-ai/Ornith-1.0-35B - k2-fsa/OmniVoice - datalab-to/surya-ocr-2 - Qwen/Qwen3-30B-A3B - nvidia/Gemma-4-31B-IT-NVFP4 - deepseek-ai/DeepSeek-V3-0324 - nvidia/Qwen2.5-VL-7B-Instruct-NVFP4 - deepseek-ai/DeepSeek-V3 - llava-hf/llava-1.5-7b-hf - Qwen/Qwen3.5-35B-A3B-FP8 - Qwen/Qwen2.5-14B-Instruct - unsloth/Qwen3.8-Flash-Next-GGUF - baidu/Unlimited-OCR - Qwen/Qwen2-VL-2B-Instruct - unsloth/Qwen3.6-35B-A3B-GGUF - TinyLlama/TinyLlama-1.1B-Chat-v1.0 - nvidia/nemotron-3.5-asr-streaming-0.6b - HuggingFaceTB/SmolVLM2-500M-Video-Instruct - deepseek-ai/DeepSeek-V4.1-Flash - Alibaba-NLP/gte-large-en-v1.5 - Qwen/Qwen3-ASR-1.7B - openbmb/MiniCPM5-2B - antirez/deepseek-v4-gguf - moonshotai/Kimi-K3 - google/gemma-3-4b-it - ornith-ai/Ornith-1.5-397B-GGUF - unsloth/Qwen3.5-9B-GGUF - bigscience/bloomz-560m - gigant/romanian-wav2vec2 - unsloth/Qwen3.5-4B-GGUF - Qwen/Qwen2.5-32B-Instruct-AWQ - intfloat/e5-large-v2 - Qwen/Qwen3-VL-Embedding-2B - nomic-ai/nomic-embed-text-v2-moe - deepseek-ai/DeepSeek-R1 - openai-community/gpt2-large - dots-studio/dots.ocr - ornith-ai/Ornith-1.5-35B-A3B-NVFP4 - Qwen/Qwen3-4B-Instruct-2507-FP8 - sahilchachra/Unlimited-OCR-AWQ - byteshape/Qwen3.8-27B-GGUF - gaunernst/gemma-3-27b-it-int4-awq - Qwen/Qwen3-Coder-Next-FP8 - LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct-AWQ - Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice - handy-computer/cohere-transcribe-03-2026-gguf - KBLab/wav2vec2-large-voxrex-swedish - RedHatAI/Qwen3.6-35B-A3B-NVFP4 - deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B - unsloth/GLM-5.3-Flash-GGUF - unsloth/Qwen3.6-35B-A3B-MTP-GGUF - gabor-hosu/e5-mistral-7b-instruct-bnb-4bit - nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-FP8 - Qwen/Qwen3-32B-AWQ - nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 - LiquidAI/LFM2.5-2.6B-GGUF - nvidia/Qwen3.5-122B-A10B-NVFP4 - Qwen/Qwen2.5-Coder-32B-Instruct-AWQ - RadixArk/Qwen3.8-27B-NVFP4 - cyankiwi/gemma-4-26B-A4B-it-AWQ-4bit - google/gemma-2-9b-it - FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree - deepseek-ai/DeepSeek-V4-Flash - openbmb/MiniCPM5-2B-DSpark - Qwen/Qwen3-30B-A3B-Instruct-2507 - RedHatAI/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-FP8 - Inferact/Qwen3.8-27B-NVFP4 - thinkingmachines/Inkling - kingabzpro/wav2vec2-large-xls-r-300m-Urdu - MiniMaxAI/MiniMax-M2.7 - ornith-ai/Ornith-1.0-9B - tencent/HunyuanOCR - Qwen/Qwen-72B - Qwen/Qwen3-Coder-30B-A3B-Instruct-FP8 - unsloth/Qwen3.6-27B-NVFP4 - stabilityai/sdxl-turbo - nvidia/Gemma-4-26B-A4B-NVFP4 - deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct - CMSManhattan/JiRackUltra_14b - nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 - kresnik/wav2vec2-large-xlsr-korean - unsloth/Z-Image-Turbo-GGUF - unsloth/Qwen3.5-35B-A3B-GGUF - Snowflake/snowflake-arctic-embed-l-v2.0 - RedHatAI/gemma-4-26B-A4B-it-FP8-dynamic - Octen/Octen-Embedding-8B - Lykon/dreamshaper-7 - cyankiwi/Qwen3.6-27B-AWQ-INT4 - unsloth/Inkling-Small-GGUF - bartowski/endless-frontier_BigBang-v1-GGUF - Qwen/Qwen2.5-Coder-32B-Instruct - meta-llama/Meta-Llama-3-8B-Instruct - cyankiwi/Qwen3-Coder-30B-A3B-Instruct-AWQ-4bit - google/medgemma-4b-it - nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 - google/gemma-4-12B-it-qat-w4a16-ct - unsloth/Qwen3.6-35B-A3B-NVFP4 - Snowflake/snowflake-arctic-embed-l - unsloth/Qwen-Image-2.1-GGUF - Qwen/Qwen2.5-VL-32B-Instruct - casperhansen/llama-3.3-70b-instruct-awq - Qwen/Qwen2.5-7B - theainerd/Wav2Vec2-large-xlsr-hindi - google/gemma-4-12B-it-qat-q4_0-gguf - Qwen/Qwen3-8B-Base - comodoro/wav2vec2-xls-r-300m-cs-250 - jinaai/jina-embeddings-v5-omni-small - Orion-zhen/Qwen2.5-Coder-7B-Instruct-AWQ - meta-llama/Llama-2-7b-hf - deepseek-ai/DeepSeek-OCR-2 - Infomaniak-AI/vllm-translategemma-4b-it - deepseek-ai/DeepSeek-V4-Flash-Vision-Exp - esatapedico/Qwen3.8-27B-NVFP4-MTP-GGUF - ornith-ai/Ornith-1.5-9B-NVFP4 - Qwen/Qwen2.5-Coder-7B-Instruct-AWQ - Yehor/w2v-xls-r-uk - microsoft/Phi-3.5-vision-instruct - AtomicChat/Qwen3.8-Flash-Next-GGUF - Qwen/Qwen3-0.6B-Base - meta-models/Muse-Glimmer-30B-GGUF - google/gemma-4-E4B-it-qat-q4_0-gguf - openbmb/MiniCPM-o-4_5 - microsoft/VibeVoice-ASR - imvladikon/wav2vec2-xls-r-300m-hebrew - datalab-to/surya-ocr-2-gguf - unsloth/Qwen3.6-35B-A3B-NVFP4-Fast - empero-ai/Qwen3.8-35B-A3B-Distill-GGUF - ISTA-DASLab/gemma-3-27b-it-GPTQ-4b-128g - meta-llama/Llama-3.2-1B - microsoft/VibeVoice-1.5B - QuantTrio/Qwen3.6-35B-A3B-AWQ - black-forest-labs/FLUX.1-dev - swiss-ai/Apertus-v1.5-8B - ggml-org/gemma-4-E4B-it-GGUF - Qwen/Qwen3-Omni-30B-A3B-Instruct - Lightricks/LTX-Video - nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 - NbAiLab/nb-wav2vec2-1b-bokmaal-v2 - nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 - unsloth/Qwen3.6-27B-GGUF - zai-org/GLM-5.2 - apple/OpenELM-1_1B-Instruct - unsloth/Qwen3.6-27B-MTP-GGUF - Qwen/Qwen2-VL-7B-Instruct - deepseek-ai/DeepSeek-R1-0528-Qwen3-8B - thinkingmachines/Inkling-Small - NbAiLab/nb-wav2vec2-1b-nynorsk - black-forest-labs/FLUX.1-schnell - Qwen/Qwen3.5-4B-Base - Qwen/Qwen2.5-1.5B - WhereIsAI/UAE-Large-V1 - meta-llama/Llama-3.1-8B - OpenGVLab/InternVL2-1B - OpenGVLab/InternVL3_5-GPT-OSS-20B-A4B-Preview-HF - cl-nagoya/ruri-v3-310m - deepseek-ai/deepseek-coder-7b-instruct-v1.5 - nvidia/Cosmos-Reason2-2B - HuggingFaceTB/SmolLM3-3B - yuxinlu1/gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-GGUF - stabilityai/sd-turbo - google/diffusiongemma-26B-A4B-it - bartowski/Qwen2.5-32B-Instruct-GGUF - empero-ai/Qwen3.8-9B-Distill-GGUF - Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8 - nvidia/Qwen3.8-27B-NVFP4 - google/gemma-4-E4B - microsoft/phi-2 - Tongyi-MAI/Z-Image-Turbo - ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF - handy-computer/parakeet-tdt-0.6b-v3-gguf - Qwen/Qwen3-1.7B-Base - nvidia/Qwen3.6-27B-NVFP4 - classla/wav2vec2-xls-r-parlaspeech-hr - cyankiwi/MiniCPM-SALA-AWQ-8bit - prism-ml/Ternary-Bonsai-27B-gguf - Salesforce/blip2-opt-2.7b - saattrupdan/wav2vec2-xls-r-300m-ftspeech - Qwen/Qwen3-4B-Thinking-2507 - Qwen/Qwen3-4B-Base - unsloth/gemma-4-E4B-it-GGUF - nvidia/parakeet-tdt-0.6b-v3 - ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF - Qwen/Qwen3-Coder-Next - nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 - EleutherAI/pythia-6.9b - sentence-transformers/all-roberta-large-v1 - Qwen/Qwen3-TTS-12Hz-0.6B-Base - CompVis/stable-diffusion-v1-4 - handy-computer/whisper-medium-gguf - empero-ai/Qwen3.8-4B-Distill-GGUF - unsloth/Llama-3.2-1B-Instruct - FINAL-Bench/POCKET-35B-GGUF - Qwen/Qwen3-8B-FP8 - deepseek-ai/DeepSeek-V4-Flash-DSpark - cyankiwi/Qwen3.8-27B-AWQ-INT4 - sentence-transformers/LaBSE - Qwen/Qwen3.5-122B-A10B - RedHatAI/Qwen3.8-27B-INT4 - distil-whisper/distil-large-v3 - openbmb/VoxCPM2 - unsloth/gemma-4-26B-A4B-it-GGUF - OpenGVLab/InternVL2-2B - ibm-granite/granite-4.1-3b - QuantTrio/Qwen3.5-9B-AWQ - CMSManhattan/JiRackDeltaNet_27b - google/gemma-4-31B - Qwen/Qwen2.5-72B-Instruct-AWQ - Qwen/Qwen3-4B-GGUF - Qwen/Qwen3-8B-GGUF - Qwen/Qwen2-1.5B-Instruct - QuantTrio/Qwen3-VL-30B-A3B-Instruct-AWQ - stepfun-ai/GOT-OCR2_0 - SG161222/RealVisXL_V5.0 - ai-sage/Giga-Embeddings-instruct - MiniMaxAI/MiniMax-M2.5 - abenzerps/Apodex-1.1-mini-GGUF - google/gemma-2-2b-it - empero-ai/Qwen3.8-2B-Distill-GGUF - SG161222/Realistic_Vision_V5.1_noVAE - Qwen/Qwen2.5-Coder-3B-Instruct - NousResearch/Hermes-3-Llama-3.1-8B - setu4993/LaBSE - RedHatAI/Llama-3.2-3B-Instruct-FP8-dynamic - empero-ai/Qwythos-9B-Claude-Mythos-5-1M-GGUF - boboliu/bge-reranker-v2.5-gemma2-lightweight-gptq - Qwen/Qwen3.5-122B-A10B-FP8 - Qwen/Qwen3-235B-A22B - ornith-ai/Ornith-1.5-35B-A3B-FP8 - Alibaba-NLP/gte-Qwen2-1.5B-instruct - zai-org/GLM-5 - zai-org/GLM-4.5-Air - Qwen/Qwen2.5-Coder-1.5B-Instruct - LeaderboardModel1/zeta-2.1-autoround-W4A16 - unsloth/gemma-4-12B-it-qat-GGUF - bartowski/MiMo-V2.6-Distill-Qwen-9B-GGUF - XHToken/Spark-X2.5-4B-GGUF - llava-hf/llava-onevision-qwen2-0.5b-ov-hf - bartowski/Qwen3.8-27B-GGUF - Qwen/Qwen3-ASR-0.6B - unsloth/gpt-oss-20b-GGUF - allenai/Olmo-3-7B-Think - bartowski/Meta-Llama-3.1-8B-Instruct-GGUF - unsloth/GLM-5.3-GGUF - Qwen/Qwen2.5-32B-Instruct-GPTQ-Int4 - nlpai-lab/KURE-v1 - unsloth/Qwen-Image-Edit-2511-GGUF - Qwen/Qwen3.5-0.8B-Base - stelterlab/Qwen3-30B-A3B-Instruct-2507-AWQ - meta-llama/Llama-3.1-405B-FP8 - openbmb/MiniCPM5-2B-MLX - Qwen/Qwen2.5-Coder-7B-Instruct-GPTQ-Int4 - allenai/Olmo-3-7B-Instruct - nm-testing/SmolLM-1.7B-Instruct-quantized.w4a16 - nvidia/Qwen3-14B-NVFP4 - Qwen/Qwen-Image-Edit-2509 - ByteDance-Seed/UI-TARS-1.5-7B - unsloth/Qwen3-4B-GGUF - nomic-ai/nomic-embed-code - bigscience/bloom-560m - ornith-ai/Ornith-1.5-9B - unsloth/gemma-4-26B-A4B-it-qat-GGUF - Qwen/Qwen3.5-397B-A17B - swiss-ai/Apertus-8B-Instruct-2509 - nvidia/GLM-5.2-NVFP4 - moonshotai/Kimi-K2.6 - k-chirkunov/gemma4-e4b-claims-comparison - gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 - Qwen/Qwen2.5-Coder-7B - google/gemma-3-12b-it - thenlper/gte-large - Qwen/Qwen2.5-3B-Instruct-AWQ - google/gemma-4-E4B-it-qat-w4a16-ct - unsloth/inkling-GGUF - mistralai/Mistral-7B-v0.1 - RedHatAI/gemma-4-31B-it-FP8-block - QuantStack/Wan2.2-I2V-A14B-GGUF - sakamakismile/Qwen3.8-27B-MTP-NVFP4 - GSAI-ML/LLaDA-8B-Instruct - Accio-Lab/occamy-1.0-GGUF - z-lab/Qwen3.8-27B-DFlash2-GGUF - PrimeIntellect/Qwen3-0.6B - migtissera/Tess-4-27B-GGUF - deepseek-ai/DeepSeek-V4-Pro - deepseek-ai/DeepSeek-R1-Distill-Qwen-32B - ornith-ai/Ornith-1.5-397B-NVFP4 - ukisai/Swift-Qwen3.8-27B-GGUF - cyankiwi/Qwen3-VL-8B-Instruct-AWQ-4bit - Qwen/Qwen3-VL-30B-A3B-Instruct - codellama/CodeLlama-7b-hf - google/gemma-3-27b-it - empero-ai/Qwythos-9B-v2-GGUF - handy-computer/Voxtral-Mini-4B-Realtime-2602-gguf - prism-ml/Bonsai-27B-gguf - tencent/Hy3 - eddiegulay/wav2vec2-large-xlsr-mvc-swahili - ibm-research/PowerMoE-3b - RedHatAI/Qwen3-32B-NVFP4 - meta-llama/Llama-2-7b-chat-hf - Qwen/Qwen3.5-397B-A17B-FP8 - black-forest-labs/FLUX.2-dev - Qwen/Qwen3-Coder-30B-A3B-Instruct - Qwen/Qwen3-ForcedAligner-0.6B - Abiray/MiniMax-H3-GGUF - black-forest-labs/FLUX.2-klein-4B - IlyaGusev/saiga_llama3_8b - Qwen/Qwen2.5-Coder-1.5B - openbmb/MiniCPM5-1B - black-forest-labs/FLUX.1-Kontext-dev - Jackrong/Qwopus3.8-27B-Flash-V2-GGUF - unsloth/gemma-4-E4B-it-qat-GGUF - Qwen/Qwen3-VL-8B-Instruct-GGUF - microsoft/phi-4 - google/gemma-4-E2B-it-qat-q4_0-gguf - cagliostrolab/animagine-xl-4.0 - microsoft/Florence-2-large - LiquidAI/LFM2.5-8B-A1B-GGUF - unsloth/gemma-4-E2B-it-GGUF - DavidAU/Qwen3.8-27B-Cold-Fusion-GAIN-V1.1-NM-DAU-NEO-MAX-MTP-GGUF - EleutherAI/gpt-neox-20b - Qwen/Qwen2.5-Coder-14B-Instruct-AWQ - peculiar-ragdoll/Cyber-Tiel-Coder-35B-A3B-GGUF-MTP - Qwen/Qwen3-235B-A22B-Instruct-2507-FP8 - OBLITERATUS/Ornith-1.5-9B-OBLITERATED - unsloth/LTX-2.3-GGUF - nvidia/llama-nemotron-embed-1b-v2 - Qwen/Qwen3.5-122B-A10B-GPTQ-Int4 - deepvk/USER-bge-m3 - unsloth/mistral-7b-v0.3-bnb-4bit - prism-ml/Ternary-Bonsai-8B-gguf - QuantTrio/GLM-4.7-Flash-AWQ - tiiuae/falcon-7b - llava-hf/llava-v1.6-mistral-7b-hf - AxionML/Qwen3.5-9B-NVFP4 - handy-computer/whisper-large-v3-turbo-gguf - deepseek-ai/DeepSeek-V3.1 - unsloth/Ornith-1.0-9B-GGUF - microsoft/Phi-3-mini-4k-instruct - microsoft/Phi-3.5-mini-instruct - Qwen/Qwen2.5-VL-3B-Instruct-AWQ - nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-NVFP4 - bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF - Jackrong/Qwopus3.6-35B-A3B-Coder-MTP-GGUF - OrdalieTech/Solon-embeddings-large-0.1 - cyankiwi/Qwen3.5-4B-AWQ-4bit - bartowski/thomsonreuters_Thomson-1.0-Small-GGUF - unsloth/FLUX.2-klein-9B-GGUF - bartowski/google_gemma-4-E2B-it-GGUF - michaelw9999/Qwen3.6-35B-A3B-NVFP4-MTP-GGUF - datalab-to/chandra - peculiar-ragdoll/Tiel-Coder-35B-A3B-GGUF - meta-llama/Llama-3.3-70B-Instruct - microsoft/Phi-4-mini-instruct - Qwen/Qwen3-0.6B-GGUF - empero-ai/Qwen3.8-27B-Ridge-GGUF - nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16 - typhoon-ai/typhoon-ocr1.5-2b - tvall43/Qwen3.6-14B-A3B-FableVibes-GGUF - AbelZimba/whisper-bemba-stt - voyageai/voyage-4-nano - Jackrong/Qwopus3.6-27B-Coder-Compat-MTP-GGUF - Qwen/Qwen3-30B-A3B-Instruct-2507-FP8 - zai-org/GLM-5.1 - google/gemma-4-26B-A4B-it-qat-q4_0-gguf - Viggle/Qwen-Image-2.1-viggle-turbo - Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign - nvidia/Qwen3.8-Flash-Next-NVFP4 - LiquidAI/LFM2.5-1.2B-Instruct-GGUF - unsloth/Qwen3.5-2B-GGUF - unsloth/FLUX.2-klein-4B-GGUF - unsloth/Kimi-K3-GGUF - handy-computer/whisper-large-v3-gguf - moonshotai/Kimi-K2-Instruct - zai-org/GLM-5.2-FP8 - ornith-ai/Ornith-1.5-397B-FP8 - sanskar003/Qwen3.5-4B-AWQ - Qwen/Qwen3-ASR-1.7B-hf - nvidia/NVIDIA-Nemotron-Nano-9B-v2 - moonshotai/Kimi-K2.5 - unsloth/Qwen3-VL-4B-Instruct-GGUF - AnkitAI/Parable-Qwen3-8B-Claude-Fable-5-GGUF - BAAI/bge-multilingual-gemma2 - deepseek-ai/DeepSeek-R1-Distill-Qwen-14B - ibm-granite/granite-3.3-2b-instruct - yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF - unsloth/Qwen-Image-2512-GGUF - alpindale/Llama-Guard-3-1B - unsloth/gemma-4-E2B-it-qat-GGUF - unsloth/Qwen-AgentWorld-35B-A3B-GGUF - Qwen/Qwen3-VL-32B-Instruct - llm-jp/llm-jp-4-33b-thinking-gguf - unsloth/Kimi-K2.7-Code-GGUF - bartowski/XYZAILab_XYZ-Aquila-mini-GGUF - google/gemma-4-31B-it-qat-q4_0-unquantized - RedHatAI/gemma-4-26B-A4B-it-NVFP4 - allenai/OLMo-2-0425-1B - ggml-org/gemma-4-26B-A4B-it-GGUF - Qwen/Qwen2-7B-Instruct - Tencent-Hunyuan/HunyuanDiT-v1.1-Diffusers-Distilled - Qwen/Qwen2.5-Coder-7B-Instruct-GGUF - Qwen/Qwen2.5-Omni-7B - google/gemma-4-E2B-it-qat-w4a16-ct - abenzerps/Spark-X2.5-4B-GGUF - QuantTrio/Qwen3-Coder-30B-A3B-Instruct-AWQ - unsloth/gemma-4-31B-it-qat-GGUF - Qwen/Qwen3.5-9B-Base - peculiar-ragdoll/Sharp-Spark-X2.5-4B-GGUF - XLabs-AI/xflux_text_encoders - Qwen/Qwen3.5-2B-Base - black-forest-labs/FLUX.2-klein-base-4B - unsloth/GLM-5.2-GGUF - agentionai/Qwen3.8-27B-AP-GGUF - Qwen/Qwen3-VL-235B-A22B-Instruct - bartowski/DeepSeek-Coder-V2-Lite-Instruct-GGUF - RadixArk/Qwen3.8-Flash-Next-NVFP4 - Qwen/Qwen3.8-Flash-Next-FP8 - microsoft/harrier-oss-v1-0.6b - poolside/Laguna-XS.2 - Qwen/Qwen2.5-1.5B-Instruct-GGUF - meta-llama/Llama-3.2-3B - utter-project/EuroLLM-22B-Instruct-2512 - drbaph/Higgs-Audio-v3-Studio - janhq/Jan-v3.5-4B-gguf - RedHatAI/gemma-3-27b-it-quantized.w4a16 - deepseek-ai/DeepSeek-R1-Distill-Qwen-7B - unsloth/gemma-4-31B-it-GGUF - philbert440/Qwen3.8-27B-W4A16-AWQ - Qwen/Qwen2.5-Math-1.5B-Instruct - RedHatAI/Qwen3-Coder-Next-FP8-dynamic - Qwen/Qwen2.5-72B-Instruct - incoai/Qwen3.8-27B-DFlash2 - Qwen/Qwen2.5-Omni-3B - superwhisper/s1-mini-GGUF - LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct - RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead - AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF - bartowski/Kwaipilot_KAT-Coder-V2.5-Dev-GGUF - mixedbread-ai/deepset-mxbai-embed-de-large-v1 - empero-ai/Qwythos-27B-v1-GGUF - KyleHessling1/Qwopus3.6-27B-Fusion-GGUF - poolside/Laguna-M.1 - playgroundai/playground-v2.5-1024px-aesthetic - huggyllama/llama-7b - nanonets/Nanonets-OCR2-3B - thinkingmachines/Inkling-Small-NVFP4 - ornith-ai/Ornith-1.0-397B-FP8 - webAI-Official/TwIL-LM3 - google/gemma-4-31B-it-qat-q4_0-gguf - Serveurperso/Qwen3-TTS-GGUF - MiniMaxAI/MiniMax-M2 - Lightricks/LTX-2 - ornith-ai/Ornith-1.0-35B-FP8 - mesolitica/llama2-embedding-1b-8k - bartowski/kai-os_Grug-12B-GGUF - meta-models/Muse-Glimmer-30B - openbmb/MiniCPM5-2B-GGUF - Qwen/Qwen3-Next-80B-A3B-Instruct - google/gemma-4-31B-it-assistant - LilaRest/gemma-4-31B-it-NVFP4-turbo - Qwen/Qwen3-0.6B-FP8 - Qwen/Qwen-Image - BDRC/tibetan-ocr - Qwen/Qwen2.5-3B - ornith-ai/Ornith-1.5-35B-A3B - FINAL-Bench/POCKET-26B-GGUF - EleutherAI/pythia-410m - AtomicChat/Ling-3.0-flash-GGUF - ibm-granite/granite-4.1-30b - lightonai/LightOnOCR-2-1B - Qwen/Qwen-Image-Edit-2511 - GnLOLot/MiniCPM5-1B-Claude-Opus-Fable5-Thinking-GGUF - bloomer010/Ling-3.0-tiny-GGUF - MaziyarPanahi/Qwen3-0.6B-GGUF - GnLOLot/MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-GGUF - google/gemma-4-31B-it-qat-w4a16-ct - poolside/Laguna-XS-2.1 - unsloth/Wan2.2-TI2V-5B-GGUF - AngelSlim/Hy3-GGUF - Qwen/Qwen1.5-MoE-A2.7B - sdadas/mmlw-retrieval-roberta-large - nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-FP8 - ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF - stepfun-ai/Step-3.7-Flash-NVFP4 - unsloth/diffusiongemma-26B-A4B-it-GGUF - unsloth/Qwen2.5-14B-bnb-4bit - google/gemma-4-E2B - ATH-MaaS/OvisOCR2 - stelterlab/Mistral-Small-24B-Instruct-2501-AWQ - Wan-AI/Wan2.1-T2V-1.3B-Diffusers - openbmb/MiniCPM-V-4.6 - ChristianAzinn/mxbai-embed-large-v1-gguf - Qwen/Qwen3-4B-AWQ - ai4bharat/indic-parler-tts - google/paligemma-3b-pt-224 - QuantTrio/Qwen3.5-27B-AWQ - ornith-ai/Ornith-1.5-397B - Qwen/Qwen2.5-0.5B-Instruct-GGUF - Jackrong/Qwopus3.6-27B-Coder-MTP-GGUF - opendatalab/MinerU2.5-Pro-2605-1.2B - cyankiwi/Qwen3.6-35B-A3B-AWQ-4bit - utter-project/EuroLLM-1.7B-Instruct - RedHatAI/Llama-3.2-1B-Instruct-FP8-dynamic - XHToken/Spark-X2.5-1.7B-GGUF - google/medgemma-1.5-4b-it - InternScience/Agents-A1-4B - unsloth/gemma-4-E2B-it-qat-mobile-GGUF - MaziyarPanahi/Qwen3-14B-GGUF - ai-forever/FRIDA - unsloth/Llama-3.2-3B-Instruct - RedHatAI/gemma-4-31B-it-FP8-dynamic - InternScience/Agents-A1-4B-Q8_0-GGUF - intfloat/e5-mistral-7b-instruct - unsloth/Muse-Glimmer-30B-GGUF - moonshotai/Kimi-VL-A3B-Instruct - MaziyarPanahi/Qwen3-4B-GGUF - Qwen/Qwen3-VL-32B-Instruct-FP8 - NousResearch/Meta-Llama-3.1-8B-Instruct - zenosai/MonkeyOCRv2-B-Parsing - QuantStack/Wan2.2-TI2V-5B-GGUF - MaziyarPanahi/Qwen3-1.7B-GGUF - Qwen/Qwen2.5-7B-Instruct-GPTQ-Int4 - moonshotai/Kimi-Linear-48B-A3B-Instruct - MaziyarPanahi/Qwen3-8B-GGUF - XiaomiMiMo/MiMo-V2.5 - jinaai/jina-embeddings-v5-text-small - bartowski/Qwen_Qwen3-4B-GGUF - HuggingFaceTB/SmolLM2-1.7B-Instruct - stepfun-ai/Step-3.5-Flash - MaziyarPanahi/Qwen3-30B-A3B-GGUF - MaziyarPanahi/Qwen3-32B-GGUF - skt/A.X-K2-NVFP4 - QuantTrio/Qwen3.5-4B-AWQ - Hcompany/Holo-3.1-35B-A3B-GGUF - inclusionAI/LLaDA2.0-mini - Qwen/Qwen2.5-VL-32B-Instruct-AWQ - Qwen/Qwen3-14B-GGUF - dots-studio/dots.mocr - AnkitAI/Parable-Granite-4.1-3B-Claude-Fable-5-GGUF - AnkitAI/Parable-Granite-4.1-8B-Claude-Fable-5-GGUF - nvidia/LocateAnything-3B - unsloth/Qwen3.5-0.8B-GGUF - Qwen/Qwen2.5-14B-Instruct-GPTQ-Int4 - poolside/Laguna-S-2.1 - nvidia/NVIDIA-Nemotron-Nano-9B-v2-FP8 - unsloth/Meta-Llama-3.1-8B-Instruct - Zyphra/Zamba2-1.2B-instruct - poolside/Laguna-XS-2.1-NVFP4 - MiniMaxAI/MiniMax-M3 - meta-llama/Meta-Llama-3-8B - QCRI/Fanar-2-27B-Instruct - NousResearch/Meta-Llama-3.1-8B - ggml-org/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF - openbmb/MiniCPM-o-2_6 - cyankiwi/gemma-4-12B-it-AWQ-INT4 - MaziyarPanahi/Yi-Coder-9B-Chat-GGUF - Qwen/Qwen2-7B - microsoft/VibeVoice-Realtime-0.5B - cyankiwi/Qwen3-30B-A3B-Instruct-2507-AWQ-4bit - Qwen/Qwen2.5-3B-Instruct-GGUF - deepseek-ai/DeepSeek-R1-0528 - Qwen/Qwen3.5-35B-A3B-GPTQ-Int4 - opendatalab/MinerU2.5-Pro-2604-1.2B - black-forest-labs/FLUX.2-klein-9B - nvidia/Nemotron-3-Embed-1B-BF16 - poolside/Laguna-XS-2.1-GGUF - meta-llama/Llama-2-13b-chat-hf - bytkim/Qwen3.6-27B-MTP-pi-tune-GGUF - Qwen/Qwen3-Omni-30B-A3B-Thinking - poolside/Laguna-S-2.1-NVFP4 - RadixArk/Qwen3.8-27B-DSpark - Qwen/Qwen2.5-Math-1.5B - bartowski/Qwen2.5-7B-Instruct-GGUF - city96/FLUX.1-dev-gguf - deepseek-ai/DeepSeek-V2-Lite - ReadyArt/gemma-4-31B-it-scotoma-2-GGUF - cyankiwi/Qwen3.5-27B-AWQ-4bit - deepseek-ai/DeepSeek-R1-Distill-Llama-8B - antirez/qwen3.8-flash-next-gguf - Laxhar/noobai-XL-1.1 - RedHatAI/Llama-3.2-3B-Instruct-FP8 - ukisai/Swift-1.5-Qwen3.8-27B-GGUF - Myric/Laguna-S-2.1-APEX-GGUF - ornith-ai/Ornith-1.0-397B - unsloth/Qwen2.5-VL-7B-Instruct-GGUF - bartowski/Qwen_Qwen3.6-35B-A3B-GGUF - unsloth/Qwen2.5-Coder-7B-Instruct - Lorbus/Qwen3.6-27B-int4-AutoRound - chatdb/natural-sql-7b - unsloth/Qwen2.5-7B-Instruct - ibm-granite/granite-vision-4.1-4b - LSX-UniWue/LLaMmlein_1B_prerelease - Qwen/Qwen3.5-27B-FP8 - InternScience/Agents-A1-4B-Q4_K_M-GGUF - Qwen/Qwen3Guard-Gen-0.6B - Qwen/Qwen3-VL-4B-Instruct-FP8 - unsloth/gemma-4-E4B-it-unsloth-bnb-4bit - leejet/MiniMax-H3-GGUF - Qwen/Qwen3-235B-A22B-Instruct-2507 - Qwen/Qwen3-30B-A3B-FP8 - JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-Q8_0 - google/gemma-3n-E2B-it - Qwen/Qwen3-VL-4B-Instruct-GGUF - MaziyarPanahi/Qwen3-4B-Instruct-2507-GGUF - meta-llama/Llama-4-Scout-17B-16E-Instruct - omlab/VLX-Seek-1.5-10B - google/gemma-2-2b - nvidia/NVIDIA-Nemotron-Nano-12B-v2-VL-NVFP4-QAD - Jackrong/Qwen3.5-9B-DeepSeek-V4-Flash-GGUF - nvidia/Cosmos-Reason2-8B - nvidia/Kimi-K2-Thinking-NVFP4 - bartowski/Qwen_Qwen3.5-27B-GGUF - ibm-granite/granite-embedding-311m-multilingual-r2 - EleutherAI/pythia-1b - ai-forever/sbert_large_nlu_ru - zai-org/GLM-4.1V-9B-Thinking - unsloth/GLM-4.7-Flash-GGUF - mistral-experimental/pixtral-12b - pfnet/plamo-embedding-1b - hugging-quants/Meta-Llama-3.1-8B-Instruct-AWQ-INT4 - dreamgen/lucid-v1-nemo - cyankiwi/Qwen3-VL-30B-A3B-Instruct-AWQ-4bit - unsloth/Step-3.7-Flash-GGUF - hugging-quants/Meta-Llama-3.1-70B-Instruct-AWQ-INT4 - z-lab/Qwen3.8-27B-DFlash2 - TIGER-Lab/VLM2Vec-Full - boboliu/Qwen3-Embedding-4B-W4A16-G128 - QuantStack/Wan2.2-T2V-A14B-GGUF - microsoft/Phi-3-mini-128k-instruct - mixedbread-ai/mxbai-embed-2d-large-v1 - Jackrong/Qwopus3.5-9B-Coder-MTP-GGUF - nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4 - datalab-to/lift - EleutherAI/pythia-1.4b - ai-forever/ru-en-RoSBERTa - bartowski/gemma-2-27b-it-GGUF - unsloth/gemma-4-26B-A4B-it-NVFP4 - batiai/Qwen3.6-27B-GGUF - ibm-granite/granite-4.1-8b - jinaai/jina-embeddings-v3-hf - bartowski/Qwen_Qwen3.5-4B-GGUF - ggml-org/gpt-oss-20b-GGUF - bartowski/Llama-3.2-3B-Instruct-GGUF - Wan-AI/Wan2.2-TI2V-5B-Diffusers - wikeeyang/Flux2-Klein-9B-True-V2 - allenai/wildguard - MaziyarPanahi/Meta-Llama-3-8B-Instruct-GGUF - SimianLuo/LCM_Dreamshaper_v7 - MaziyarPanahi/Mistral-7B-v0.1-GGUF - QuantTrio/Qwen3-Coder-30B-A3B-Instruct-GPTQ-Int8 - mistralai/Mistral-7B-Instruct-v0.1 - unsloth/ERNIE-Image-Turbo-GGUF - cagliostrolab/animagine-xl-3.1 - Jackrong/Qwopus3.6-27B-v2-MTP-GGUF - unsloth/Qwen3-VL-8B-Instruct-GGUF - ReliquaryForge/qwen3-4b-base-dapo-v4 - nvidia/NVIDIA-Nemotron-Parse-v1.1 - MaziyarPanahi/Mistral-Small-Instruct-2409-GGUF - MaziyarPanahi/Mistral-Nemo-Instruct-2407-GGUF - Qwen/Qwen2.5-7B-Instruct-GGUF - unsloth/Qwen3-14B-unsloth-bnb-4bit - MaziyarPanahi/Yi-Coder-1.5B-Chat-GGUF - Wan-AI/Wan2.2-I2V-A14B-Diffusers - ggml-org/gpt-oss-120b-GGUF - gaunernst/gemma-3-12b-it-int4-awq - Qwen/Qwen3-1.7B-GGUF - MaziyarPanahi/firefunction-v2-GGUF - unsloth/Qwen3-Coder-Next-GGUF - Qwen/Qwen1.5-1.8B-Chat - QUASAR-QAT/Qwen3.8-27B-QUASAR-NVFP4 - unsloth/Qwen3.5-27B-GGUF - abhishekchohan/gemma-3-12b-it-quantized-W4A16 - ukisai/Swift-1.5-Qwen3.8-Flash-Next-GGUF - nvidia/Qwen3-8B-NVFP4 - Qwen/Qwen3-VL-8B-Thinking - HuggingFaceTB/SmolVLM2-2.2B-Instruct - MaziyarPanahi/Qwen2-7B-Instruct-GGUF - unsloth/Qwen3-0.6B-GGUF - xinsir/controlnet-openpose-sdxl-1.0 - QuantTrio/Qwen3-VL-32B-Instruct-AWQ - leejet/Qwen-Image-2.1-GGUF - unsloth/Llama-3.1-8B-Instruct - microsoft/Phi-3.5-MoE-instruct - RedHatAI/Meta-Llama-3.1-70B-Instruct-FP8 - NousResearch/Llama-2-7b-hf - MaziyarPanahi/Phi-3.5-mini-instruct-GGUF - bullerwins/Wan2.2-I2V-A14B-GGUF - MaziyarPanahi/Mistral-7B-Instruct-v0.3-GGUF - MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF - John6666/nova-furry-xl-il-v120-sdxl - MaziyarPanahi/Mixtral-8x22B-v0.1-GGUF - MaziyarPanahi/phi-4-GGUF - MaziyarPanahi/Meta-Llama-3.1-8B-Instruct-GGUF - Wan-AI/Wan2.2-T2V-A14B-Diffusers - intfloat/e5-large - bartowski/Qwen_Qwen3.5-9B-GGUF - MaziyarPanahi/gemma-3-4b-it-GGUF - MaziyarPanahi/solar-pro-preview-instruct-GGUF - Qwen/Qwen3-VL-2B-Instruct-GGUF - google/gemma-2b - allenai/OLMoE-1B-7B-0125-Instruct - deepseek-ai/DeepSeek-V2-Lite-Chat - MaziyarPanahi/Llama-3-8B-Instruct-32k-v0.1-GGUF - Jackrong/Qwopus3.8-27B-Flash-GGUF - nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark - CohereLabs/North-Micro-Vision-Instruct - cyankiwi/gemma-4-31B-it-AWQ-4bit - ibm-granite/granite-4.2-8b - city96/Wan2.1-I2V-14B-480P-gguf - Qwen/Qwen2.5-Coder-1.5B-Instruct-GGUF - HuggingFaceTB/SmolVLM-500M-Instruct - LiquidAI/LFM2.5-2.6B-DSpark-GGUF - sesame/csm-1b - MaziyarPanahi/Qwen2.5-7B-Instruct-GGUF - unsloth/FLUX.2-dev-GGUF - RedHatAI/DeepSeek-Coder-V2-Lite-Instruct-FP8 - MaziyarPanahi/QwQ-32B-GGUF - MaziyarPanahi/DeepSeek-R1-0528-Qwen3-8B-GGUF - unsloth/gemma-3n-E4B-it - MaziyarPanahi/Mistral-Small-24B-Instruct-2501-GGUF - xinsir/controlnet-union-sdxl-1.0 - bartowski/Llama-3.2-1B-Instruct-GGUF - Qwen/Qwen3-Next-80B-A3B-Instruct-FP8 - google/gemma-3-4b-pt - MaziyarPanahi/Phi-4-mini-instruct-GGUF - MaziyarPanahi/Llama-3.2-3B-Instruct-GGUF - OpenMOSS-Team/MOSS-TTS-v1.5 - MaziyarPanahi/Ministral-3-3B-Reasoning-2512-GGUF - allenai/Olmo-3-1025-7B - MaziyarPanahi/gemma-2-2b-it-GGUF - unsloth/Llama-3.2-1B - MaziyarPanahi/Qwen2.5-1.5B-Instruct-GGUF - MaziyarPanahi/gemma-3-1b-it-GGUF - MaziyarPanahi/Llama-3.3-70B-Instruct-GGUF - RedHatAI/Qwen2.5-1.5B-quantized.w8a8 - RedHatAI/Meta-Llama-3.1-8B-Instruct-FP8 - MaziyarPanahi/Yi-1.5-6B-Chat-GGUF - stabilityai/stable-video-diffusion-img2vid-xt - MaziyarPanahi/WizardLM-2-7B-GGUF - MaziyarPanahi/gemma-3-12b-it-GGUF - MaziyarPanahi/mistral-small-3.1-24b-instruct-2503-hf-GGUF - moonshotai/Kimi-Linear-48B-A3B-Base - MaziyarPanahi/Llama-3-8B-Instruct-64k-GGUF - zai-org/GLM-4.5 - MaziyarPanahi/mathstral-7B-v0.1-GGUF - unsloth/Z-Image-GGUF - MaziyarPanahi/Meta-Llama-3.1-70B-Instruct-GGUF - MaziyarPanahi/DeepSeek-V3-0324-GGUF - MaziyarPanahi/gemma-3-27b-it-GGUF - dragonkue/snowflake-arctic-embed-l-v2.0-ko - Qwen/Qwen2-1.5B - MaziyarPanahi/Mistral-Large-Instruct-2411-GGUF - MaziyarPanahi/INTELLECT-2-GGUF - MaziyarPanahi/Meta-Llama-3.1-405B-Instruct-GGUF - LiquidAI/LFM2.5-2.6B - Qwen/Qwen3-32B-FP8 - inclusionAI/LLaDA2.1-mini - ThorOdinson246/nl2sh-1.5b-Q4_K_M - latam-gpt/Llama-3.1-70B-LatamGPT-SFT-1.0 - SG161222/RealVisXL_V4.0 - meta-llama/Meta-Llama-3-70B-Instruct - RedHatAI/Llama-3.3-70B-Instruct-FP8-dynamic - deepseek-ai/deepseek-coder-6.7b-instruct - aari1995/German_Semantic_STS_V2 - deepseek-ai/deepseek-coder-6.7b-base - prism-ml/Bonsai-8B-gguf - Qwen/Qwen2.5-Coder-14B-Instruct-GGUF - unsloth/Phi-4-mini-instruct-GGUF - Qwen/Qwen-Image-2.1 - LiquidAI/LFM2.5-1.2B-Instruct - Snowflake/snowflake-arctic-embed-m-v2.0 - MuXodious/gpt-oss-20b-RichardErkhov-heresy - city96/FLUX.1-schnell-gguf - SebastianBodza/Kartoffel_Orpheus-3B_german_natural-v0.1 - Qwen/Qwen2.5-1.5B-Instruct-AWQ - Qwen/Qwen-Image-Edit - canopylabs/3b-de-ft-research_release - allenai/OLMoE-1B-7B-0924 - audio-cpp/AuK-Base-and-Flash-GGUF - QuantStack/Qwen-Image-Edit-2509-GGUF - Alibaba-NLP/gte-Qwen2-7B-instruct - hmellor/Ilama-3.2-1B - tencent/Hy3-preview - sarvamai/sarvam-30b - unsloth/Qwen3-4B-Instruct-2507-unsloth-bnb-4bit - KyleHessling1/Qwopus-GLM-18B-Merged-GGUF - google/gemma-4-12B - lightseekorg/kimi-k2.6-eagle3-mla - Qwen/Qwen2.5-Coder-3B-Instruct-GGUF - unsloth/Qwen3-1.7B-GGUF - unsloth/Qwen3-Coder-Next-FP8 - Wan-AI/Wan2.1-I2V-14B-480P - jinaai/jina-embeddings-v5-text-small-retrieval - HuggingFaceTB/SmolLM3-3B-Base - diffusers/stable-diffusion-xl-1.0-inpainting-0.1 - unsloth/Qwen2.5-7B-Instruct-bnb-4bit - pentacoxian-dev/Qwen3.8-Flash-Next-IQ3E-Q8D-MTP-GGUF - NousResearch/Meta-Llama-3-8B-Instruct - bosonai/higgs-tts-3-4b - Qwen/Qwen3-30B-A3B-GGUF - SulphurAI/Sulphur-2-base - PrimeIntellect/Qwen3-1.7B - John6666/amanatsu-illustrious-v11-sdxl - dealignai/GLM-5.3-CYBERSECURITY-FP8 - stabilityai/stable-diffusion-3.5-large - stabilityai/stable-diffusion-xl-refiner-1.0 - unsloth/Llama-3.2-3B-Instruct-GGUF - bartowski/Qwen2.5-14B-Instruct-GGUF - DevQuasar/amd.Instella-MoE-16B-A3B-Think-GGUF - Qwen/Qwen3-32B-GGUF - h2oai/h2o-danube3-500m-chat - dealignai/Bonsai-2-27B-Ternary-CRACK-GGUF - Nanbeige/Nanbeige4.2-3B - nvidia/NV-Embed-v2 - MaziyarPanahi/Qwen3-30B-A3B-Instruct-2507-GGUF - littlejohn-ai/bge-m3-spa-law-qa - meta-llama/Meta-Llama-3-70B - JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-Q4_K_M - Salesforce/Llama-xLAM-2-8b-fc-r - bartowski/Meta-Llama-3.1-70B-Instruct-GGUF - RamManavalan/Qwen3-VL-Embedding-8B-FP8 - XiaomiMiMo/MiMo-V2-Flash - cais/HarmBench-Llama-2-13b-cls - molbal/ideogram-4-gguf - zai-org/GLM-4.7 - CohereLabs/c4ai-command-r-v01 - XiaomiMiMo/MiMo-V2.6-Pro-RL - LLMSafety/Qwen2.5-Math-7B-4bit - unsloth/DeepSeek-R1-0528-Qwen3-8B-GGUF - black-forest-labs/FLUX.2-klein-base-9B - unsloth/llama-3-8b-Instruct-bnb-4bit - nvidia/diffusiongemma-26B-A4B-it-NVFP4 - HuggingFaceH4/zephyr-7b-beta - Abiray/Qwen-Image-2.1-GGUF - silveroxides/Chroma-GGUF - badtheorylabs/BTL-4-Compact - lmstudio-community/Qwen3-Coder-30B-A3B-Instruct-GGUF - unsloth/Qwen3-8B-GGUF - krea/Krea-2-Turbo - hyrelabs/Homura-30B-GGUF - stabilityai/stable-diffusion-3.5-medium - zenmagnets/Nemotron-3-Embed-1B-Q4_K_M-GGUF - NovaSearch/stella_en_400M_v5 - deepseek-ai/DeepSeek-V4-Pro-0813 - nvidia/GLM-5.3-Flash-NVFP4 - OnomaAIResearch/Illustrious-xl-early-release-v0 - ibm-granite/granite-speech-4.1-2b-nar - Qwen/Qwen3-14B-Base - unsloth/gpt-oss-20b-BF16 - baidu/ERNIE-4.5-21B-A3B-Thinking - city96/FLUX.2-dev-gguf - ibm-granite/granite-3.0-1b-a400m-instruct - unsloth/Qwen2.5-3B-Instruct-unsloth-bnb-4bit - internlm/internlm2-chat-7b - meta-llama/Llama-3.1-70B-Instruct - krea/Krea-2-Raw - Qwen/Qwen2.5-14B - unsloth/Meta-Llama-3.1-8B - OpenMOSS-Team/MOSS-TTS-Local-Transformer-v1.5 - unsloth/Qwen2.5-7B-Instruct-unsloth-bnb-4bit - unsloth/Qwen3-8B - meta-llama/Llama-3.1-405B - Aratako/MioTTS-2.6B - ibm-granite/granite-3.3-8b-instruct - bartowski/Qwen_Qwen3-0.6B-GGUF - bartowski/DeepSeek-R1-Distill-Qwen-7B-GGUF - Wan-AI/Wan2.1-T2V-14B-Diffusers - Wan-AI/Wan2.1-I2V-14B-720P-Diffusers - jinaai/jina-clip-v2 - Qwen/Qwen3-30B-A3B-Thinking-2507 - vantagewithai/Krea-2-Turbo-GGUF - Qwen/Qwen3-4B-FP8 - openai-community/gpt2-xl - Qwen/Qwen2.5-Math-7B-Instruct - Jackrong/Qwen3.5-4B-Claude-4.6-Opus-Reasoning-Distilled-GGUF - nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-Base-BF16 - bartowski/microsoft_Phi-4-mini-instruct-GGUF - Dream-org/Dream-v0-Instruct-7B - leejet/Z-Image-Turbo-GGUF - Xenova/sweep-next-edit-1.5B - unsloth/gpt-oss-20b-unsloth-bnb-4bit - bartowski/Qwen2.5-1.5B-Instruct-GGUF - unsloth/Qwen2.5-Coder-7B-Instruct-bnb-4bit - nvidia/Llama-3_3-Nemotron-Super-49B-v1_5-FP8 - frankjoshua/novaAnimeXL_ilV140 - unsloth/Qwen-Image-GGUF - nytopop/Qwen3-30B-A3B.w8a8 - RadixArk/GLM-5.3-NVFP4 - bartowski/gemma-2-2b-it-GGUF - Wan-AI/Wan2.1-VACE-14B - MaziyarPanahi/GLM-4.6V-Flash-GGUF - lmstudio-community/Meta-Llama-3.1-8B-Instruct-GGUF - OpenMOSS-Team/MOSS-Audio-Tokenizer - MaziyarPanahi/gpt-oss-20b-Derestricted-GGUF - laion/voiceclap-large-v2 - lmstudio-community/Llama-3.2-3B-Instruct-GGUF - MaziyarPanahi/Qwen3-Coder-Next-GGUF - Qwen/Qwen1.5-7B - nvidia/Llama-3_3-Nemotron-Super-49B-v1_5 - unsloth/gpt-oss-120b-GGUF - bartowski/Qwen2.5-3B-Instruct-GGUF - 0xSero/deepseek-v4-flash-0731-spark - MaziyarPanahi/Nemotron-Orchestrator-8B-GGUF - fla-hub/transformer-1.3B-100B - lmstudio-community/gemma-3-1b-it-GGUF - MaziyarPanahi/Trinity-Mini-GGUF - unsloth/Qwen3-32B-GGUF - QuantStack/Wan2.1_14B_VACE-GGUF - tencent/Hunyuan-A13B-Instruct - FastVideo/FastWan2.2-TI2V-5B-FullAttn-Diffusers - Wan-AI/Wan2.1-I2V-14B-480P-Diffusers - nvidia/Nemotron-Labs-Diffusion-8B-Base - Qwen/Qwen1.5-0.5B-Chat - nvidia/Nemotron-Labs-Diffusion-3B - Inferact/GLM-5.3-NVFP4 - internlm/internlm3-8b-instruct - Lykon/dreamshaper-8 - unsloth/DeepSeek-R1-Distill-Qwen-1.5B-GGUF - bartowski/Qwen2.5-72B-Instruct-GGUF - LiquidAI/LFM2-1.2B - prism-ml/Bonsai-1.7B-gguf - unsloth/Meta-Llama-3.1-8B-Instruct-bnb-4bit - unsloth/Qwen3-14B-GGUF - TheBloke/Llama-2-7B-Chat-GGUF - Qwen/Qwen-7B - bartowski/meta-llama_Llama-4-Scout-17B-16E-Instruct-old-GGUF - TinyLlama/TinyLlama-1.1B-intermediate-step-1431k-3T - John6666/diving-illustrious-real-asian-v50-sdxl - LiquidAI/LFM2.5-8B-A1B-DSpark-GGUF - tobil/qmd-query-expansion-1.7B-gguf - LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-GGUF - sbintuitions/sarashina2.2-0.5b-instruct-v0.1 - stabilityai/stable-diffusion-3-medium-diffusers - agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF - moonshotai/Kimi-K2-Instruct-0905 - CMSManhattan/JiRackUltra_7b - NousResearch/Meta-Llama-3-8B - CMSManhattan/JiRackUltra_32b - joeygambino/MiniMax-H3-encoder-GGUF - nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8 - ibm-granite/granite-4.1-8b-fp8 - moonshotai/Kimi-K2-Thinking - Qwen/Qwen3Guard-Gen-4B - Alittlehammmer/Qwen3.6-35B-A3B-DFlash-GGUF-llama.cpp - black-forest-labs/FLUX.2-klein-9b-kv - unsloth/Qwen3-1.7B-unsloth-bnb-4bit - Qwen/Qwen3-30B-A3B-GPTQ-Int4 - jdopensource/JoyAI-Image-Edit-Diffusers - unsloth/Qwen3-Next-80B-A3B-Instruct-GGUF - RedHatAI/Llama-3.2-1B-Instruct-FP8 - unsloth/Llama-3.2-1B-Instruct-bnb-4bit - hesamation/Qwen3.6-35B-A3B-Claude-4.6-Opus-Reasoning-Distilled-GGUF - deepseek-ai/DeepSeek-R1-Distill-Llama-70B - arianraje/qwen3-4b-mamba3-hybrid-stage2b-kd-bias0 - AtomicChat/Qwen3-4B-DFlash-GGUF - molbal/MiniMax-H3-GGUF - John6666/one-obsession-17-red-sdxl - lmstudio-community/DeepSeek-R1-0528-Qwen3-8B-GGUF - AtomicChat/Qwen3.5-9B-DFlash-GGUF - casperhansen/llama-3-8b-instruct-awq - microsoft/Phi-tiny-MoE-instruct - AtomicChat/Qwen3.5-4B-DFlash-GGUF - allenai/OLMo-2-0425-1B-Instruct - Alittlehammmer/Qwen3.6-27B-DFlash-GGUF-llama.cpp - microsoft/phi-1_5 - martineux/dvine82-xl - kmhf/hf-moshiko - RadixArk/Inkling-Small-DSpark-Preview - bullerwins/FLUX.1-Kontext-dev-GGUF - Efficient-Large-Model/gemma-2-2b-it - DevQuasar-13/THUDM.GLM-Z1-32B-0414-GGUF - Wan-AI/Wan2.1-T2V-14B - perplexity-ai/pplx-embed-v1-0.6b - solidrust/Hermes-3-Llama-3.1-8B-AWQ - bartowski/Ling-3.0-tiny-GGUF - lodestones/Chroma1-HD - bartowski/Mistral-7B-Instruct-v0.3-GGUF - sahilchachra/gemma-4-12B-coder-fable5-composer2.5-AWQ - openthaigpt/openthaigpt1.5-7b-instruct - QuantFactory/Qwen2.5-Coder-7B-GGUF - RavichandranJ/Dolphin3-Cyber-8B-GGUF - ibm-granite/granite-4.0-h-tiny - tiiuae/Falcon-H1-0.5B-Base - hugging-quants/Llama-3.2-1B-Instruct-Q4_K_M-GGUF - lmstudio-community/Qwen3-4B-Instruct-2507-GGUF - John6666/obsession-illustriousxl-v10-sdxl - HuggingFaceTB/SmolLM-1.7B - OpenOneRec/OneRec-1.7B - unsloth/gemma-3-1b-it - nvidia/MiniMax-M3-NVFP4 - microsoft/Phi-4-mini-reasoning - XCurOS/XCurOS0.1-8B-Instruct - bartowski/Qwen2.5-Coder-7B-Instruct-GGUF - unsloth/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF - avatargrim/Qwen2.5-7B_pyuigpt - XiaomiMiMo/MiMo-V2.6-Flash-RL - Qwen/Qwen3-14B-FP8 - unsloth/Qwen3-4B-unsloth-bnb-4bit - SG161222/RealVisXL_V5.0_Lightning - zai-org/GLM-4.5-Air-FP8 - ibm-granite/granite-guardian-4.1-8b - deepseek-ai/DeepSeek-V3.2-Exp - ibm-granite/granite-3.1-8b-instruct - google/gemma-3-1b-pt - LGAI-EXAONE/EXAONE-3.5-2.4B-Instruct - Tongyi-MAI/Z-Image - Abiray/LTX-2.5-Distilled-GGUF - HiDream-ai/HiDream-I1-Fast - ibm-granite/granite-4.0-tiny-preview - unsloth/DeepSeek-R1-Distill-Llama-70B-GGUF - nvidia/gpt-oss-120b-Eagle3-short-context - Qwen/Qwen3-30B-A3B-Base - allenai/OLMo-1B-hf - MaziyarPanahi/Qwen3-4B-Thinking-2507-GGUF - MaziyarPanahi/Ministral-3-14B-Reasoning-2512-GGUF - tencent/Hy3-FP8 - MaziyarPanahi/NVIDIA-Nemotron-Nano-12B-v2-GGUF - ibm-granite/granite-3.2-8b-instruct - leejet/FLUX.2-klein-4B-GGUF - NousResearch/Meta-Llama-3-70B-Instruct - Lykon/dreamshaper-xl-v2-turbo - Qwen/Qwen-7B-Chat - nvidia/gpt-oss-120b-Eagle3-v3 - entrick/Security-SLM-Gemma-4-E2B-it-GGUF - stabilityai/stable-diffusion-3.5-large-controlnet-canny - microsoft/Phi-3-vision-128k-instruct - ideogram-ai/ideogram-4-fp8 - nvidia/Nemotron-H-8B-Base-8K - TheBloke/Mistral-7B-Instruct-v0.2-GGUF - bartowski/Phi-3.5-mini-instruct-GGUF - XingChen-AGI/Xing4.0-29B-A4B - abenzerps/Nex-N2.5-mini-GGUF - tencent/Hunyuan-7B-Instruct - google/codegemma-7b-it - Abiray/Qwen-Image-2.1-viggle-4-steps-turbo-GGUF - openbmb/MiniCPM5-1B-MLX - stabilityai/stable-diffusion-3.5-large-controlnet-depth - ByteDance-Seed/Seed-OSS-36B-Instruct - unsloth/Qwen2.5-32B-Instruct-bnb-4bit - ibnzterrell/Meta-Llama-3.3-70B-Instruct-AWQ-INT4 - Wan-AI/Wan2.1-T2V-1.3B - unsloth/Llama-3.3-70B-Instruct - Qwen/Qwen2.5-7B-Instruct-1M - nvidia/Llama-3.1-8B-Instruct-FP8 - QuantStack/FLUX.1-Kontext-dev-GGUF - vcruz305/DeepSeek-V4.1-Flash-GGUF - unsloth/Llama-3.2-3B-Instruct-unsloth-bnb-4bit - ibm-granite/granite-4.2-3b - bartowski/Altworld_Hemmingway-1-GGUF - NousResearch/Llama-3.2-1B - bigcode/starcoder2-3b - cagliostrolab/animagine-xl-3.0 - nvidia/Mistral-NeMo-Minitron-8B-Instruct - Vikhrmodels/Vikhr-Nemo-12B-Instruct-R-21-09-24 - wikeeyang/Flux2-Klein-9B-True-V3 - bartowski/Qwen_Qwen3-4B-Instruct-2507-GGUF - IPostYellow/TurboWan2.1-T2V-1.3B-Diffusers - NousResearch/Hermes-2-Pro-Mistral-7B - John6666/prefect-illustrious-xl-v3-sdxl - EldanRing/Winnow-12B - z-lab/LLaMA3.1-8B-Instruct-DFlash-UltraChat - bartowski/Qwen_Qwen3-1.7B-GGUF - google/gemma-1.1-2b-it - bartowski/Qwen2.5-Coder-1.5B-Instruct-GGUF - Lightricks/LTX-2.5-Diffusers - EleutherAI/pythia-410m-deduped - unsloth/gemma-3-1b-it-GGUF - unsloth/Qwen2.5-1.5B-Instruct-unsloth-bnb-4bit - KBlueLeaf/TIPO-500M-ft - LGAI-EXAONE/EXAONE-3.5-32B-Instruct-AWQ - bartowski/THUDM_GLM-4-32B-0414-GGUF - incoai/GLM-5.3-Flash-DFlash2 - openbmb/MiniCPM5-1B-GGUF - Qwen/Qwen-Image-2512 - HuggingFaceTB/SmolLM2-1.7B - RedHatAI/Meta-Llama-3.1-8B-Instruct-quantized.w4a16 - speakleash/Bielik-11B-v3.0-Instruct-awq - z-lab/Qwen3.6-35B-A3B-DFlash - lmstudio-community/Qwen2.5-Coder-14B-Instruct-GGUF - google/t5gemma-2b-2b-ul2-it - duyntnet/Chroma-GGUF-All-Versions - ibm-granite/granite-4.0-1b-base - unsloth/Qwen3-0.6B-unsloth-bnb-4bit - microsoft/Phi-3-mini-4k-instruct-gguf - FINAL-Bench/POCKET-Darwin-180B-GGUF - ibm-granite/granite-3b-code-base-2k - unsloth/Qwen2.5-14B-Instruct - allenai/Olmo-3-32B-Think-SFT - TaoLiveAIGC/TLive-Omni-4B - unsloth/Llama-3.2-1B-Instruct-GGUF - Qwen/QwQ-32B - meta-llama/Llama-Guard-3-8B - bartowski/DeepSeek-R1-Distill-Qwen-14B-GGUF - unsloth/Llama-3.3-70B-Instruct-GGUF - bartowski/Qwen2.5-Coder-14B-Instruct-GGUF - black-forest-labs/FLUX.1-Krea-dev - jayn7/Z-Image-Turbo-GGUF - QuantFactory/Meta-Llama-3-8B-Instruct-GGUF - Qwen/Qwen3-Next-80B-A3B-Thinking - empero-ai/Qwen3.8-4B-Distill - peculiar-ragdoll/Cyber-Tiel-Coder-35B-A3B-GGUF - mudler/Qwen3.5-35B-A3B-APEX-GGUF - Qwen/Qwen2.5-Math-7B - TokenRhythm/NeoHorse-1-4B-GGUF - second-state/stable-diffusion-v1-5-GGUF - meta-llama/Llama-3.1-70B - FireRedTeam/FireRed-Image-Edit-1.0 - unsloth/DeepSeek-R1-Distill-Llama-8B-GGUF - bartowski/Qwen_Qwen3-8B-GGUF - ggml-org/Qwen3-0.6B-GGUF - deepseek-ai/deepseek-coder-1.3b-instruct - swiss-ai/Apertus-8B-2509 - EleutherAI/gpt-neo-1.3B - dealignai/Bonsai-2-27B-1bit-CRACK-GGUF - Edge0/Audio8-ASR-Infinite - QuantStack/Qwen-Image-Edit-GGUF - inclusionAI/Ling-3.0-tiny-GGUF - neuralcrew/neutrino-instruct - nvidia/DeepSeek-V4-Flash-0731-NVFP4 - useful-quants/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-W4A16 - MeiGen-AI/PosterOmni_v1 - shafire/Zero-Gemma4-E4B-OpenZero-GGUF - unsloth/FLUX.2-klein-base-9B-GGUF - Laxhar/noobai-XL-1.0 - Laxhar/noobai-XL-Vpred-1.0 - unsloth/Llama-3.2-3B-Instruct-bnb-4bit - Qwen/Qwen1.5-0.5B - zai-org/GLM-5.3-BF16 - ibm-granite/granite-4.2-30b - unsloth/Qwen2.5-3B-Instruct-bnb-4bit - bartowski/DeepSeek-R1-Distill-Qwen-32B-GGUF - nvidia/Llama-3_3-Nemotron-Super-49B-v1 - xinsir/controlnet-canny-sdxl-1.0 - Phil2Sat/Qwen-Image-Edit-Rapid-AIO-GGUF - skt/A.X-4.0-Light - Qwen/Qwen3-235B-A22B-Thinking-2507-FP8 - Qwen/Qwen3-Coder-480B-A35B-Instruct - hugging-quants/Llama-3.2-3B-Instruct-Q4_K_M-GGUF - stabilityai/stablelm-3b-4e1t - audio-cpp/LiveAvatar-GGUF - pipecat-ai/phonellm-alpha-1 - Qwen/Qwen3Guard-Gen-8B - legraphista/glm-4-9b-chat-IMat-GGUF - diffusers/controlnet-depth-sdxl-1.0 - TheBloke/Mistral-7B-Instruct-v0.1-GGUF - bartowski/NousResearch_Hermes-4-14B-GGUF - prithivMLmods/Qwen-Image-2.1-PE-T2I-GGUF - solidrust/Mistral-7B-Instruct-v0.3-AWQ - Qwen/Qwen2.5-Coder-3B - John6666/hassaku-xl-illustrious-v31-sdxl - openai/gpt-oss-safeguard-20b - osllmai-community/Llama-3.2-1B - unsloth/FLUX.1-dev-GGUF - stabilityai/stable-video-diffusion-img2vid - fdtn-ai/antares-1b - ai21labs/AI21-Jamba-Reasoning-3B - allenai/Olmo-3-7B-Instruct-SFT - Venastine-Research/Xing4.0-29B-A4B-GGUF - unsloth/LFM2.5-1.2B-Instruct-GGUF - tiiuae/falcon-7b-instruct - allenai/OLMo-2-1124-7B-Instruct - Lykon/dreamshaper-xl-lightning - OBLITERATUS/gemma-4-E4B-it-OBLITERATED - lmstudio-community/Qwen3-14B-GGUF - Agnuxo/CAJAL-4B - outsourc-e/Qwen3.8-27B-Unleashed-GGUF - PixArt-alpha/PixArt-Sigma-XL-2-1024-MS - sakamakismile/Qwen3.6-27B-Text-NVFP4-MTP - upstage/SOLAR-10.7B-Instruct-v1.0 - unsloth/llama-3-8b-Instruct - bg-digitalservices/Gemma-4-26B-A4B-it-NVFP4 - bartowski/L3-8B-Stheno-v3.2-GGUF - google/gemma-2-9b - zai-org/glm-edge-4b-chat-gguf - Manojb/stable-diffusion-2-1-base - meta-llama/Llama-Guard-3-1B - nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Base-BF16 - realrebelai/LTX-2.5_GGUFs - magespace/Wan2.2-I2V-A14B-Lightning-Diffusers - TinyLlama/TinyLlama-1.1B-step-50K-105b - SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16 - stelterlab/Qwen3-Coder-30B-A3B-Instruct-AWQ - GritLM/GritLM-7B - diffusers/controlnet-canny-sdxl-1.0 - city96/Qwen-Image-gguf - QuantStack/LTX-2.3-GGUF - nvidia/GLM-5.3-NVFP4 - autotrust/GLM-5.3-Flash-GGUF-DGX-Spark - chfm/Qwen-Image-2.1-GGUF - IFM/K2-Horizon-MoVA-36B-A4B-GGUF - Qwen/Qwen3.8-2.4T-A95B - nvidia/llama-nemotron-embed-vl-1b-v2 - fishaudio/s2-pro - sudoingx/Ternary-Bonsai-2-27B-PTQ1_0-MTP-GGUF - joeygambino/MiniMax-H3-curve-GGUF - unsloth/FLUX.1-schnell-GGUF - rectangleworm/ideogram-4-gguf - XHToken/Spark-X2.5-4B - ai4bharat/IndicF5 - perplexity-ai/pplx-embed-v1-4b - city96/Wan2.1-I2V-14B-720P-gguf - realrebelai/SCAIL-2_GGUF - bartowski/Cloudflare_clef-GGUF - jayn7/WAN2.2-I2V_A14B-DISTILL-LIGHTX2V-4STEP-GGUF - IFM/K2-Horizon-7B - DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF - QuantStack/Wan2.2-S2V-14B-GGUF - bartowski/TheDrummer_Artemis-31B-v1.2-GGUF - Hob-forge/Kolibri-1-GGUF - ggml-org/Clef-Flash-GGUF - bartowski/Cloudflare_clef-flash-GGUF - city96/Wan2.1-FLF2V-14B-720P-gguf - nvidia/Cosmos3-Super-Image2Video - LiquidAI/LFM2.5-8B-A1B - city96/Wan2.1-T2V-14B-gguf - realrebelai/Wan-Animate-2_GGUFs - deadbydawn101/RavenX-CyberAgent-Qwen3.6-35B-A3B-Opus-4.7-OpenMythos-Pentester-BugHunter-RATH-GGUF - jialinyyzz/humanizer - IFM/K2-Horizon-375B-A23B - hum-ma/Wan2.2-TI2V-5B-Turbo-GGUF - Wan-AI/Wan2.1-VACE-1.3B - IFM/K2-Horizon-3.7B - unsloth/LTX-2-GGUF - vantagewithai/SCAIL-2-GGUF-ComfyUI - bartowski/FrogNano-4B-2609-GGUF - google/embeddinggemma-2 - netease-youdao/Confucius4-R2T2 - ggml-org/Qwen3.8-Flash-Next-GGUF - Abiray/MiniMax-H3-Singularity-GGUF - Wan-AI/Wan2.1-I2V-14B-720P - m-a-p/MERT-v2-FullSong - XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B - robbyant/lingbot-world-fast - calcuis/wan-1.3b-gguf - bytkim/Qwen3.8-27B-pi-GGUF - IFM/K2-Horizon-MoVA-36B-A4B - tencent/WeMM-Embedding-2B - Cloudflare/clef-flash - ukisai/Swift-Bonsai-2-GGUF - paradigma-inc/limite-1b-violetto - zai-org/CogVideoX-2b - zai-org/CogVideoX-5b - bullerwins/Wan2.2-T2V-A14B-GGUF - Rabinovich/LongLive-2.0-5B-Diffusers - inclusionAI/Ling-3.0-tiny - TaichuAI/ZDTaichu5.0-9B - ggml-org/Clef-GGUF - BreezeBlue/Breeze-TTS-2 - inclusionAI/Ling-3.0-flash - Altworld/Hemmingway-1 - Cloudflare/clef - moondream/parakeet-ultra - Kwaipilot/KAT-Coder-V2.5-Dev - nex-agi/Nex-N2.5-mini - nvidia/NVIDIA-Nemotron-3-Nano-4B-GGUF - CohereLabs/North-Mini-Code-1.0 - jinaai/jina-embeddings-v5-omni-nano - LiquidAI/d1-omni-600M - terrorswift/REDCELL-26B-A4B-OSINT-Cyber-APEX-GGUF - dphn/Dolphin-Mistral-24B-Venice-Edition - DavidAU/LFM2.5-8B-A1B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF - peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF - autotrust/GLM-5.3-GGUF-DGX-Spark - empero-ai/Qwen3.8-35B-A3B-Distill - Aleph-Alpha/Kolibri-1 - empero-ai/Qwythos-9B-Claude-Mythos-5-1M - XiaomiMiMo/MiMo-V2.6-Flash-MOPD - elix3r/gemma4-12b-with-proj-ltx-2.5-GGUF - albucino/Qwen3.8-Flash-Next-W4A16-FP8PLE - autotrust/GLM5.3-Flash-E224-DGX-Spark - LiquidAI/d1-3B - SkyIsNotGreen/Scion-35B-A3B - microsoft/harrier-oss-v1-27b - ukisai/Swift-1.5-Qwen3.8-27b - orcarouter/OrcaSAQ-2-27B - yandex/AliceAI-Foundation-80B-A3B-Base - ggml-org/OpenJev-GGUF - DeepHat/DeepHat-V1-7B - pfeifferj/Ornith-1.5-35B-A3B-GSQ-RCO-GGUF - SeerRay-Lab/Xiaomi-OCR-0 - Qwen/Qwen3Guard-Stream-0.6B - KittenML/kitten-tts-2 - llm-jp/llm-jp-4.1-8b-thinking-gguf - alesha-pro/Qwen3.8-27B-S-mirai-GGUF - webAI-Official/TwIL-LM3-Pro - XiaomiMiMo/MiMo-V2.6-Pro-MOPD - apodex/Apodex-1.1-mini - tencent/HunyuanImage-3.0 - logic65/Whittle-Qwen-3.8-35B-A3B - rmonsurate/Victoria - Lythri/Lythri-4B-A2B - Mia-AiLab/GLM-5.3-Flash-EXL3-4bpw-TensorFold - stabilityai/stable-video-diffusion-img2vid-xt-1-1 - emperorofrome/Gmcoder - InternScience/Agents-A1 - aj9o9/Qwen3.8-27B-Escha-W2-GGUF - OrionLLM/OxCoder-9B - tsinghua-sigs-robot-lab/VeriLoop-E2 - jialinyyzz/humanizer-GGUF - ArliAI/GLM-4.6-Derestricted-v3 - MaziyarPanahi/Llama-3-Groq-8B-Tool-Use-GGUF - LiquidAI/d1-3B-GGUF - AikidoSec/altar-1 - Mia-AiLab/GLM-5.3-Flash-EXL3-4bpw-TensorFold-Ablit - bottlecapai/ThinkingCap-Qwen3.8-27B - diffbot/DeepSeek-V4.1-Flash-EXL3-3bpw-2x-RTX-PRO-6000 - bytkim/Qwen3.8-27B-pi - hiwaifu-research/WaifuGemma4-26b-a4b-v1 - FINAL-Bench/Darwin-27B-RSI-GGUF - Blackfrost-AI/CYBER-FROST-3.8-BF16 - IFM/K2-Type-0.9B - BAAI/AREX-2 - RedHatAI/Qwen3.8-Flash-Next-NVFP4 - unsloth/embeddinggemma-2 - IQuestLab/IQuest-Q1 - topk-io/topk-embed-v1-xsmall - vcruz305/CYBER-FROST-3.8-EXL3-SAGE-3.87bpw - Aleph-Alpha/Kolibri-1-BF16 - Eliasfpv28/Kolibri-1-Q3_K_S-GGUF - CaryPalmer/Ternary-Bonsai-2-27B-262k-GGUF - llm-jp/llm-jp-4.1-33b-thinking - BuzzASR/persian - cantina-security/apex-flash-1 - Blackfrost-Research/GLM-5.3-F.U-AnthraClaud-Edition-BF16 - troed/Qwen3.8-27B-ASCII-Condensed - Abiray/Qwen-Image-2.1-viggle-turbo-v0.3-6step-GGUF - nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2 - perplexity-ai/pplx-embed-v2-context-9b-preview - FINAL-Bench/Darwin-27B-RSI - lmstudio-community/Llama-3-Groq-70B-Tool-Use-GGUF - mehdi-hf/nemotron-asr-streaming-farsi - Gryphe/Gemma-4-26B-A4B-StyleTune-V2 - EschaLabs/Qwen3.8-27B-Escha-W2 - nvidia/NVIDIA-NemotronLabs-AI-for-Media-Sports-Tennis - FINAL-Bench/Darwin-180B-RSI - Hcompany/Holo4-27B - Blackfrost-AI/CYBER-FROST-3.8-NVFP4-V2 - ISTA-DASLab/Qwen3.8-Flash-Next-P48NVFP4-MoESQ - LiquidAI/LFM2-1.2B-RAG - vcruz305/RED-SNOW-5.3-FLASH-EXL3-SAGE-4.91bpw - RedHatAI/DeepSeek-V4-Flash-0731-NVFP4 - kova-ai/kova-tts-1 - PleIAs/baguettotron-600m - elyza/ELYZA-Thinking-1.0-llm-jp-4-33b - elyza/ELYZA-Thinking-1.0-llm-jp-4-32b-a3b - yamz-labs/GLM-5.3-Flash-EXL3-Yamz - bitlamas/Qwen3.8-Flash-Next-Q4_K_XL-DN4 - kyutai/glm-4-voice-of-reason-9b - YOON1v/Apex-2 - perplexity-ai/pplx-embed-v2-late-0.6b - cklxx/laya-browser - amazon/ALoDLM-8B - Ateron/Gemma-4-Dark-Thoughts-V2-31B - briaai/fibo-scene-analyzer - Blockway/Agens-Volundr-32B-Preview - WonseokJayJung/Connect-C1-0.8B-v0.1-GGUF - kandinskylab/Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers - kandinskylab/Kandinsky-6.0-Pro-distill-5s-Diffusers - Remek/basal-1.5-4.5B - perplexity-ai/pplx-embed-v2-late-9b - kandinskylab/Kandinsky-6.0-VSR-5s-Diffusers - Maincode/matilda-jev-v1 - FINAL-Bench/Darwin-180B-RSI-R3 - vcruz305/RED-SNOW-5.3-FLASH-EXL3-SAGE-2.49bpw - h2oai/h2o-lightning-4b - Blackfrost-AI/RED-SNOW-5.3-FLASH-BF16 - kandinskylab/Kandinsky-6.0-Lite-distill-5s-Diffusers - kandinskylab/Kandinsky-6.0-Pro-5s-Diffusers - Ateron/Gemma-4-Writers-31B-V2 - audreyt/Kolibri-1-NVFP4-W4A16 - lighteternal/biodecision-v2-4b - malteos/most-embed-de - kandinskylab/Kandinsky-6.0-Lite-5s-Diffusers - ConwayResearch/Underdog-Saluki-27B-1.0 - PleIAs/Baguettotron-MoE - JetBrains/Mellum2.1-12B-A2.5B-Thinking - kandinskylab/Kandinsky-6.0-Pro-pretrain-5s-Diffusers - kandinskylab/Kandinsky-6.0-Lite-pretrain-5s-Diffusers - aina-tech/Anima-Lightning - amazon/ALoDLM-1.7B - facebook/meta-encoder - speridlabs/iris-3b - Modularcomputing/Native-Bird - lightonai/LightOnOCR-3-0.8B - lightonai/LightOnOCR-3-4B - ApolloRaines/LTX-2.5-22b-OmniGen-v12 - JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF - Kujira/Underdog-Saluki-27B-1.0-MTP-GGUF - DavidAU/Qwen3.8-27B-Turbo-Brilliance-Power-35X-Reasoning-Instruct-modes-GGUF - inclusionAI/Ming-Image-0.1-Design - Aratako/Irodori-TTS-v4-Large --- # VRAM & GPU Cost Calculator Pick any Hugging Face model and see: - how much GPU memory it needs to run, fine-tune with LoRA, fully fine-tune, generate, transcribe or embed, at its published precision or at 16, 8 or 4 bits; - the cheapest GPU setups that hold it (one H100, two A100s, four RTX 4090s...), each priced live across the GPU clouds [FastGPU](https://fastgpu.co/?utm_source=huggingface&utm_medium=space&utm_campaign=gpu-cost-calculator) tracks, with a link to every offer. It reads each model's own metadata from the Hub (the parameter count in its safetensors or GGUF header, its quantization config) and asks FastGPU's free match API for the sizing and the live prices. Nothing to install, no sign-in. Type a size such as `70B` for a model that is not on the Hub. ## How the memory is sized The same way as on fastgpu.co, by FastGPU's matcher: - **Run / serve:** the weights (parameters × bytes per parameter: 2 at 16-bit, 1 at 8-bit, 0.5 at 4-bit), plus the KV cache for one request at a 4K-token context and runtime headroom. - **Fine-tune with LoRA:** the frozen weights plus the adapters, their optimizer state and activations. - **Full fine-tune:** about 16 bytes per parameter (weights, gradients and Adam states) plus activations. - **Image, video, speech and embedding models** FastGPU knows (FLUX, SDXL, Wan, Whisper and more) use a figure for the whole pipeline, text encoders and VAE included. A quantized repo (AWQ, GPTQ, FP8, NVFP4, bitsandbytes, GGUF) is sized at the bits it is stored in unless you pick another precision. The full method: [fastgpu.co/methodology](https://fastgpu.co/methodology?utm_source=huggingface&utm_medium=space&utm_campaign=gpu-cost-calculator). ## Where the prices come from FastGPU collects the GPU rental prices providers publish (marketplaces, neoclouds and hyperscalers) and refreshes them through the day. Each price here is the provider's own per-GPU rate times the number of GPUs in the setup. "Lowest price" (the default) sorts by price alone. "Best match" orders by FastGPU's score, led by price and weighing reliability and availability; providers that pay FastGPU a referral fee are marked ★ and can only win a near-tie there, within a few percent of the cheapest. The price history is open: [fastgpu/cloud-gpu-prices](https://huggingface.co/datasets/fastgpu/cloud-gpu-prices) (CC BY 4.0, updated daily). ## Use it from code ```bash curl "https://fastgpu.co/api/v1/match?model=Qwen/Qwen3-32B¶ms_b=32.8&task=inference" ``` No key needed. Docs: [fastgpu.co/docs](https://fastgpu.co/docs?utm_source=huggingface&utm_medium=space&utm_campaign=gpu-cost-calculator).