fastgpu's picture
VRAM & GPU cost calculator: live GPU rental prices for any Hugging Face model
8bbafb9 verified
|
Raw History Blame Contribute Delete
56.6 kB
metadata
title: VRAM & GPU Cost Calculator
emoji: ⚡
colorFrom: blue
colorTo: indigo
sdk: static
pinned: true
license: mit
short_description: VRAM any HF model needs + live GPU rental prices
tags:
  - vram
  - gpu
  - gpu-prices
  - calculator
  - inference
  - fine-tuning
  - llm
  - cloud-gpu
models:
  - Qwen/Qwen3-0.6B
  - google/gemma-4-26B-A4B-it
  - Qwen/Qwen3-8B
  - BAAI/bge-large-en-v1.5
  - Qwen/Qwen3-VL-8B-Instruct
  - Qwen/Qwen3-4B
  - google/gemma-4-31B-it
  - Qwen/Qwen3-Embedding-0.6B
  - Qwen/Qwen3.5-9B
  - Qwen/Qwen2.5-7B-Instruct
  - Qwen/Qwen3.5-4B
  - meta-llama/Llama-3.2-1B-Instruct
  - intfloat/multilingual-e5-large
  - Qwen/Qwen3.8-27B
  - Qwen/Qwen2.5-1.5B-Instruct
  - openai/whisper-large-v3-turbo
  - zai-org/GLM-5.3-Flash
  - RadixArk/Kimi-K3-DSpark
  - meta-llama/Llama-3.1-8B-Instruct
  - openai/gpt-oss-20b
  - unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF
  - Qwen/Qwen2.5-VL-7B-Instruct
  - nvidia/Qwen3.6-35B-A3B-NVFP4
  - ornith-ai/Ornith-1.5-9B-GGUF
  - Qwen/Qwen3-14B
  - dphn/dolphin-2.9.1-yi-1.5-34b
  - Qwen/Qwen3.8-27B-FP8
  - deepseek-ai/DeepSeek-V4-Flash-0731
  - Qwen/Qwen3.6-35B-A3B-FP8
  - google/gemma-4-E4B-it
  - stabilityai/stable-diffusion-xl-base-1.0
  - deepseek-ai/DeepSeek-V3.2
  - prism-ml/Ternary-Bonsai-2-27B-gguf
  - Qwen/Qwen2.5-3B-Instruct
  - openai/gpt-oss-120b
  - Qwen/Qwen3.5-2B
  - openai/whisper-large-v3
  - cdiamond/Qwen3.8-27B-iMatrix-NVFP4-MTP-GGUF
  - ornith-ai/Ornith-1.5-35B-A3B-GGUF
  - Qwen/Qwen3-32B
  - Qwen/Qwen3-4B-Instruct-2507
  - google/embeddinggemma-300m
  - Qwen/Qwen3.6-35B-A3B
  - ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF
  - openai/whisper-small
  - Qwen/Qwen3-VL-4B-Instruct
  - google/gemma-3-1b-it
  - Qwen/Qwen3-1.7B
  - Qwen/Qwen3-14B-AWQ
  - google/gemma-4-E2B-it
  - Qwen/Qwen3-VL-2B-Instruct
  - Qwen/Qwen2.5-7B-Instruct-AWQ
  - MahmoudAshraf/mms-300m-1130-forced-aligner
  - Qwen/Qwen3.5-0.8B
  - CMSManhattan/JiRackUltra_1b
  - ornith-ai/Ornith-1.0-9B-GGUF
  - Qwen/Qwen3-Embedding-8B
  - Qwen/Qwen3.6-27B-FP8
  - BAAI/bge-reranker-large
  - Qwen/Qwen3-8B-AWQ
  - Qwen/Qwen3.6-27B
  - Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice
  - Qwen/Qwen2.5-VL-3B-Instruct
  - vikhyatk/moondream2
  - Qwen/Qwen2.5-Coder-7B-Instruct
  - deepseek-ai/DeepSeek-OCR
  - mixedbread-ai/mxbai-embed-large-v1
  - jinaai/jina-embeddings-v3
  - nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16
  - mistralai/Voxtral-Mini-4B-Realtime-2602
  - Qwen/Qwen3.5-27B
  - handy-computer/nemotron-3.5-asr-streaming-0.6b-gguf
  - ggml-org/Qwen3.8-27B-GGUF
  - datalab-to/chandra-ocr-2
  - Qwen/Qwen2.5-32B-Instruct
  - Qwen/Qwen3-Embedding-4B
  - handy-computer/parakeet-unified-en-0.6b-gguf
  - farbodtavakkoli/OTel-2.0-LLM-31B-IT
  - Qwen/Qwen3-VL-8B-Instruct-FP8
  - peculiar-ragdoll/Tiel-Coder-35B-A3B-GGUF-MTP
  - Qwen/Qwen2.5-14B-Instruct-AWQ
  - google/gemma-4-12B-it
  - Qwen/Qwen2.5-VL-7B-Instruct-AWQ
  - Qwen/Qwen3.8-Flash-Next
  - ornith-ai/Ornith-1.0-35B-GGUF
  - stable-diffusion-v1-5/stable-diffusion-v1-5
  - unsloth/gemma-4-12b-it-GGUF
  - Qwen/Qwen2-VL-7B-Instruct-AWQ
  - Qwen/Qwen2.5-Coder-14B-Instruct
  - mistralai/Mistral-7B-Instruct-v0.2
  - zai-org/GLM-4.7-Flash
  - meta-llama/Llama-3.2-3B-Instruct
  - zai-org/GLM-5.3
  - intfloat/multilingual-e5-large-instruct
  - autotrust/JEV-27B-VL
  - ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF
  - nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4
  - Alibaba-NLP/gte-multilingual-base
  - Qwen/Qwen3-VL-Embedding-8B
  - facebook/w2v-bert-2.0
  - Qwen/Qwen3.5-35B-A3B
  - peculiar-ragdoll/Dirk-Qwen3.8-27B-GGUF
  - ornith-ai/Ornith-1.0-35B
  - k2-fsa/OmniVoice
  - datalab-to/surya-ocr-2
  - Qwen/Qwen3-30B-A3B
  - nvidia/Gemma-4-31B-IT-NVFP4
  - deepseek-ai/DeepSeek-V3-0324
  - nvidia/Qwen2.5-VL-7B-Instruct-NVFP4
  - deepseek-ai/DeepSeek-V3
  - llava-hf/llava-1.5-7b-hf
  - Qwen/Qwen3.5-35B-A3B-FP8
  - Qwen/Qwen2.5-14B-Instruct
  - unsloth/Qwen3.8-Flash-Next-GGUF
  - baidu/Unlimited-OCR
  - Qwen/Qwen2-VL-2B-Instruct
  - unsloth/Qwen3.6-35B-A3B-GGUF
  - TinyLlama/TinyLlama-1.1B-Chat-v1.0
  - nvidia/nemotron-3.5-asr-streaming-0.6b
  - HuggingFaceTB/SmolVLM2-500M-Video-Instruct
  - deepseek-ai/DeepSeek-V4.1-Flash
  - Alibaba-NLP/gte-large-en-v1.5
  - Qwen/Qwen3-ASR-1.7B
  - openbmb/MiniCPM5-2B
  - antirez/deepseek-v4-gguf
  - moonshotai/Kimi-K3
  - google/gemma-3-4b-it
  - ornith-ai/Ornith-1.5-397B-GGUF
  - unsloth/Qwen3.5-9B-GGUF
  - bigscience/bloomz-560m
  - gigant/romanian-wav2vec2
  - unsloth/Qwen3.5-4B-GGUF
  - Qwen/Qwen2.5-32B-Instruct-AWQ
  - intfloat/e5-large-v2
  - Qwen/Qwen3-VL-Embedding-2B
  - nomic-ai/nomic-embed-text-v2-moe
  - deepseek-ai/DeepSeek-R1
  - openai-community/gpt2-large
  - dots-studio/dots.ocr
  - ornith-ai/Ornith-1.5-35B-A3B-NVFP4
  - Qwen/Qwen3-4B-Instruct-2507-FP8
  - sahilchachra/Unlimited-OCR-AWQ
  - byteshape/Qwen3.8-27B-GGUF
  - gaunernst/gemma-3-27b-it-int4-awq
  - Qwen/Qwen3-Coder-Next-FP8
  - LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct-AWQ
  - Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice
  - handy-computer/cohere-transcribe-03-2026-gguf
  - KBLab/wav2vec2-large-voxrex-swedish
  - RedHatAI/Qwen3.6-35B-A3B-NVFP4
  - deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B
  - unsloth/GLM-5.3-Flash-GGUF
  - unsloth/Qwen3.6-35B-A3B-MTP-GGUF
  - gabor-hosu/e5-mistral-7b-instruct-bnb-4bit
  - nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-FP8
  - Qwen/Qwen3-32B-AWQ
  - nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16
  - LiquidAI/LFM2.5-2.6B-GGUF
  - nvidia/Qwen3.5-122B-A10B-NVFP4
  - Qwen/Qwen2.5-Coder-32B-Instruct-AWQ
  - RadixArk/Qwen3.8-27B-NVFP4
  - cyankiwi/gemma-4-26B-A4B-it-AWQ-4bit
  - google/gemma-2-9b-it
  - FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree
  - deepseek-ai/DeepSeek-V4-Flash
  - openbmb/MiniCPM5-2B-DSpark
  - Qwen/Qwen3-30B-A3B-Instruct-2507
  - RedHatAI/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-FP8
  - Inferact/Qwen3.8-27B-NVFP4
  - thinkingmachines/Inkling
  - kingabzpro/wav2vec2-large-xls-r-300m-Urdu
  - MiniMaxAI/MiniMax-M2.7
  - ornith-ai/Ornith-1.0-9B
  - tencent/HunyuanOCR
  - Qwen/Qwen-72B
  - Qwen/Qwen3-Coder-30B-A3B-Instruct-FP8
  - unsloth/Qwen3.6-27B-NVFP4
  - stabilityai/sdxl-turbo
  - nvidia/Gemma-4-26B-A4B-NVFP4
  - deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct
  - CMSManhattan/JiRackUltra_14b
  - nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
  - kresnik/wav2vec2-large-xlsr-korean
  - unsloth/Z-Image-Turbo-GGUF
  - unsloth/Qwen3.5-35B-A3B-GGUF
  - Snowflake/snowflake-arctic-embed-l-v2.0
  - RedHatAI/gemma-4-26B-A4B-it-FP8-dynamic
  - Octen/Octen-Embedding-8B
  - Lykon/dreamshaper-7
  - cyankiwi/Qwen3.6-27B-AWQ-INT4
  - unsloth/Inkling-Small-GGUF
  - bartowski/endless-frontier_BigBang-v1-GGUF
  - Qwen/Qwen2.5-Coder-32B-Instruct
  - meta-llama/Meta-Llama-3-8B-Instruct
  - cyankiwi/Qwen3-Coder-30B-A3B-Instruct-AWQ-4bit
  - google/medgemma-4b-it
  - nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4
  - google/gemma-4-12B-it-qat-w4a16-ct
  - unsloth/Qwen3.6-35B-A3B-NVFP4
  - Snowflake/snowflake-arctic-embed-l
  - unsloth/Qwen-Image-2.1-GGUF
  - Qwen/Qwen2.5-VL-32B-Instruct
  - casperhansen/llama-3.3-70b-instruct-awq
  - Qwen/Qwen2.5-7B
  - theainerd/Wav2Vec2-large-xlsr-hindi
  - google/gemma-4-12B-it-qat-q4_0-gguf
  - Qwen/Qwen3-8B-Base
  - comodoro/wav2vec2-xls-r-300m-cs-250
  - jinaai/jina-embeddings-v5-omni-small
  - Orion-zhen/Qwen2.5-Coder-7B-Instruct-AWQ
  - meta-llama/Llama-2-7b-hf
  - deepseek-ai/DeepSeek-OCR-2
  - Infomaniak-AI/vllm-translategemma-4b-it
  - deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
  - esatapedico/Qwen3.8-27B-NVFP4-MTP-GGUF
  - ornith-ai/Ornith-1.5-9B-NVFP4
  - Qwen/Qwen2.5-Coder-7B-Instruct-AWQ
  - Yehor/w2v-xls-r-uk
  - microsoft/Phi-3.5-vision-instruct
  - AtomicChat/Qwen3.8-Flash-Next-GGUF
  - Qwen/Qwen3-0.6B-Base
  - meta-models/Muse-Glimmer-30B-GGUF
  - google/gemma-4-E4B-it-qat-q4_0-gguf
  - openbmb/MiniCPM-o-4_5
  - microsoft/VibeVoice-ASR
  - imvladikon/wav2vec2-xls-r-300m-hebrew
  - datalab-to/surya-ocr-2-gguf
  - unsloth/Qwen3.6-35B-A3B-NVFP4-Fast
  - empero-ai/Qwen3.8-35B-A3B-Distill-GGUF
  - ISTA-DASLab/gemma-3-27b-it-GPTQ-4b-128g
  - meta-llama/Llama-3.2-1B
  - microsoft/VibeVoice-1.5B
  - QuantTrio/Qwen3.6-35B-A3B-AWQ
  - black-forest-labs/FLUX.1-dev
  - swiss-ai/Apertus-v1.5-8B
  - ggml-org/gemma-4-E4B-it-GGUF
  - Qwen/Qwen3-Omni-30B-A3B-Instruct
  - Lightricks/LTX-Video
  - nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
  - NbAiLab/nb-wav2vec2-1b-bokmaal-v2
  - nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4
  - unsloth/Qwen3.6-27B-GGUF
  - zai-org/GLM-5.2
  - apple/OpenELM-1_1B-Instruct
  - unsloth/Qwen3.6-27B-MTP-GGUF
  - Qwen/Qwen2-VL-7B-Instruct
  - deepseek-ai/DeepSeek-R1-0528-Qwen3-8B
  - thinkingmachines/Inkling-Small
  - NbAiLab/nb-wav2vec2-1b-nynorsk
  - black-forest-labs/FLUX.1-schnell
  - Qwen/Qwen3.5-4B-Base
  - Qwen/Qwen2.5-1.5B
  - WhereIsAI/UAE-Large-V1
  - meta-llama/Llama-3.1-8B
  - OpenGVLab/InternVL2-1B
  - OpenGVLab/InternVL3_5-GPT-OSS-20B-A4B-Preview-HF
  - cl-nagoya/ruri-v3-310m
  - deepseek-ai/deepseek-coder-7b-instruct-v1.5
  - nvidia/Cosmos-Reason2-2B
  - HuggingFaceTB/SmolLM3-3B
  - yuxinlu1/gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-GGUF
  - stabilityai/sd-turbo
  - google/diffusiongemma-26B-A4B-it
  - bartowski/Qwen2.5-32B-Instruct-GGUF
  - empero-ai/Qwen3.8-9B-Distill-GGUF
  - Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8
  - nvidia/Qwen3.8-27B-NVFP4
  - google/gemma-4-E4B
  - microsoft/phi-2
  - Tongyi-MAI/Z-Image-Turbo
  - ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF
  - handy-computer/parakeet-tdt-0.6b-v3-gguf
  - Qwen/Qwen3-1.7B-Base
  - nvidia/Qwen3.6-27B-NVFP4
  - classla/wav2vec2-xls-r-parlaspeech-hr
  - cyankiwi/MiniCPM-SALA-AWQ-8bit
  - prism-ml/Ternary-Bonsai-27B-gguf
  - Salesforce/blip2-opt-2.7b
  - saattrupdan/wav2vec2-xls-r-300m-ftspeech
  - Qwen/Qwen3-4B-Thinking-2507
  - Qwen/Qwen3-4B-Base
  - unsloth/gemma-4-E4B-it-GGUF
  - nvidia/parakeet-tdt-0.6b-v3
  - ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF
  - Qwen/Qwen3-Coder-Next
  - nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
  - EleutherAI/pythia-6.9b
  - sentence-transformers/all-roberta-large-v1
  - Qwen/Qwen3-TTS-12Hz-0.6B-Base
  - CompVis/stable-diffusion-v1-4
  - handy-computer/whisper-medium-gguf
  - empero-ai/Qwen3.8-4B-Distill-GGUF
  - unsloth/Llama-3.2-1B-Instruct
  - FINAL-Bench/POCKET-35B-GGUF
  - Qwen/Qwen3-8B-FP8
  - deepseek-ai/DeepSeek-V4-Flash-DSpark
  - cyankiwi/Qwen3.8-27B-AWQ-INT4
  - sentence-transformers/LaBSE
  - Qwen/Qwen3.5-122B-A10B
  - RedHatAI/Qwen3.8-27B-INT4
  - distil-whisper/distil-large-v3
  - openbmb/VoxCPM2
  - unsloth/gemma-4-26B-A4B-it-GGUF
  - OpenGVLab/InternVL2-2B
  - ibm-granite/granite-4.1-3b
  - QuantTrio/Qwen3.5-9B-AWQ
  - CMSManhattan/JiRackDeltaNet_27b
  - google/gemma-4-31B
  - Qwen/Qwen2.5-72B-Instruct-AWQ
  - Qwen/Qwen3-4B-GGUF
  - Qwen/Qwen3-8B-GGUF
  - Qwen/Qwen2-1.5B-Instruct
  - QuantTrio/Qwen3-VL-30B-A3B-Instruct-AWQ
  - stepfun-ai/GOT-OCR2_0
  - SG161222/RealVisXL_V5.0
  - ai-sage/Giga-Embeddings-instruct
  - MiniMaxAI/MiniMax-M2.5
  - abenzerps/Apodex-1.1-mini-GGUF
  - google/gemma-2-2b-it
  - empero-ai/Qwen3.8-2B-Distill-GGUF
  - SG161222/Realistic_Vision_V5.1_noVAE
  - Qwen/Qwen2.5-Coder-3B-Instruct
  - NousResearch/Hermes-3-Llama-3.1-8B
  - setu4993/LaBSE
  - RedHatAI/Llama-3.2-3B-Instruct-FP8-dynamic
  - empero-ai/Qwythos-9B-Claude-Mythos-5-1M-GGUF
  - boboliu/bge-reranker-v2.5-gemma2-lightweight-gptq
  - Qwen/Qwen3.5-122B-A10B-FP8
  - Qwen/Qwen3-235B-A22B
  - ornith-ai/Ornith-1.5-35B-A3B-FP8
  - Alibaba-NLP/gte-Qwen2-1.5B-instruct
  - zai-org/GLM-5
  - zai-org/GLM-4.5-Air
  - Qwen/Qwen2.5-Coder-1.5B-Instruct
  - LeaderboardModel1/zeta-2.1-autoround-W4A16
  - unsloth/gemma-4-12B-it-qat-GGUF
  - bartowski/MiMo-V2.6-Distill-Qwen-9B-GGUF
  - XHToken/Spark-X2.5-4B-GGUF
  - llava-hf/llava-onevision-qwen2-0.5b-ov-hf
  - bartowski/Qwen3.8-27B-GGUF
  - Qwen/Qwen3-ASR-0.6B
  - unsloth/gpt-oss-20b-GGUF
  - allenai/Olmo-3-7B-Think
  - bartowski/Meta-Llama-3.1-8B-Instruct-GGUF
  - unsloth/GLM-5.3-GGUF
  - Qwen/Qwen2.5-32B-Instruct-GPTQ-Int4
  - nlpai-lab/KURE-v1
  - unsloth/Qwen-Image-Edit-2511-GGUF
  - Qwen/Qwen3.5-0.8B-Base
  - stelterlab/Qwen3-30B-A3B-Instruct-2507-AWQ
  - meta-llama/Llama-3.1-405B-FP8
  - openbmb/MiniCPM5-2B-MLX
  - Qwen/Qwen2.5-Coder-7B-Instruct-GPTQ-Int4
  - allenai/Olmo-3-7B-Instruct
  - nm-testing/SmolLM-1.7B-Instruct-quantized.w4a16
  - nvidia/Qwen3-14B-NVFP4
  - Qwen/Qwen-Image-Edit-2509
  - ByteDance-Seed/UI-TARS-1.5-7B
  - unsloth/Qwen3-4B-GGUF
  - nomic-ai/nomic-embed-code
  - bigscience/bloom-560m
  - ornith-ai/Ornith-1.5-9B
  - unsloth/gemma-4-26B-A4B-it-qat-GGUF
  - Qwen/Qwen3.5-397B-A17B
  - swiss-ai/Apertus-8B-Instruct-2509
  - nvidia/GLM-5.2-NVFP4
  - moonshotai/Kimi-K2.6
  - k-chirkunov/gemma4-e4b-claims-comparison
  - gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090
  - Qwen/Qwen2.5-Coder-7B
  - google/gemma-3-12b-it
  - thenlper/gte-large
  - Qwen/Qwen2.5-3B-Instruct-AWQ
  - google/gemma-4-E4B-it-qat-w4a16-ct
  - unsloth/inkling-GGUF
  - mistralai/Mistral-7B-v0.1
  - RedHatAI/gemma-4-31B-it-FP8-block
  - QuantStack/Wan2.2-I2V-A14B-GGUF
  - sakamakismile/Qwen3.8-27B-MTP-NVFP4
  - GSAI-ML/LLaDA-8B-Instruct
  - Accio-Lab/occamy-1.0-GGUF
  - z-lab/Qwen3.8-27B-DFlash2-GGUF
  - PrimeIntellect/Qwen3-0.6B
  - migtissera/Tess-4-27B-GGUF
  - deepseek-ai/DeepSeek-V4-Pro
  - deepseek-ai/DeepSeek-R1-Distill-Qwen-32B
  - ornith-ai/Ornith-1.5-397B-NVFP4
  - ukisai/Swift-Qwen3.8-27B-GGUF
  - cyankiwi/Qwen3-VL-8B-Instruct-AWQ-4bit
  - Qwen/Qwen3-VL-30B-A3B-Instruct
  - codellama/CodeLlama-7b-hf
  - google/gemma-3-27b-it
  - empero-ai/Qwythos-9B-v2-GGUF
  - handy-computer/Voxtral-Mini-4B-Realtime-2602-gguf
  - prism-ml/Bonsai-27B-gguf
  - tencent/Hy3
  - eddiegulay/wav2vec2-large-xlsr-mvc-swahili
  - ibm-research/PowerMoE-3b
  - RedHatAI/Qwen3-32B-NVFP4
  - meta-llama/Llama-2-7b-chat-hf
  - Qwen/Qwen3.5-397B-A17B-FP8
  - black-forest-labs/FLUX.2-dev
  - Qwen/Qwen3-Coder-30B-A3B-Instruct
  - Qwen/Qwen3-ForcedAligner-0.6B
  - Abiray/MiniMax-H3-GGUF
  - black-forest-labs/FLUX.2-klein-4B
  - IlyaGusev/saiga_llama3_8b
  - Qwen/Qwen2.5-Coder-1.5B
  - openbmb/MiniCPM5-1B
  - black-forest-labs/FLUX.1-Kontext-dev
  - Jackrong/Qwopus3.8-27B-Flash-V2-GGUF
  - unsloth/gemma-4-E4B-it-qat-GGUF
  - Qwen/Qwen3-VL-8B-Instruct-GGUF
  - microsoft/phi-4
  - google/gemma-4-E2B-it-qat-q4_0-gguf
  - cagliostrolab/animagine-xl-4.0
  - microsoft/Florence-2-large
  - LiquidAI/LFM2.5-8B-A1B-GGUF
  - unsloth/gemma-4-E2B-it-GGUF
  - DavidAU/Qwen3.8-27B-Cold-Fusion-GAIN-V1.1-NM-DAU-NEO-MAX-MTP-GGUF
  - EleutherAI/gpt-neox-20b
  - Qwen/Qwen2.5-Coder-14B-Instruct-AWQ
  - peculiar-ragdoll/Cyber-Tiel-Coder-35B-A3B-GGUF-MTP
  - Qwen/Qwen3-235B-A22B-Instruct-2507-FP8
  - OBLITERATUS/Ornith-1.5-9B-OBLITERATED
  - unsloth/LTX-2.3-GGUF
  - nvidia/llama-nemotron-embed-1b-v2
  - Qwen/Qwen3.5-122B-A10B-GPTQ-Int4
  - deepvk/USER-bge-m3
  - unsloth/mistral-7b-v0.3-bnb-4bit
  - prism-ml/Ternary-Bonsai-8B-gguf
  - QuantTrio/GLM-4.7-Flash-AWQ
  - tiiuae/falcon-7b
  - llava-hf/llava-v1.6-mistral-7b-hf
  - AxionML/Qwen3.5-9B-NVFP4
  - handy-computer/whisper-large-v3-turbo-gguf
  - deepseek-ai/DeepSeek-V3.1
  - unsloth/Ornith-1.0-9B-GGUF
  - microsoft/Phi-3-mini-4k-instruct
  - microsoft/Phi-3.5-mini-instruct
  - Qwen/Qwen2.5-VL-3B-Instruct-AWQ
  - nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-NVFP4
  - bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF
  - Jackrong/Qwopus3.6-35B-A3B-Coder-MTP-GGUF
  - OrdalieTech/Solon-embeddings-large-0.1
  - cyankiwi/Qwen3.5-4B-AWQ-4bit
  - bartowski/thomsonreuters_Thomson-1.0-Small-GGUF
  - unsloth/FLUX.2-klein-9B-GGUF
  - bartowski/google_gemma-4-E2B-it-GGUF
  - michaelw9999/Qwen3.6-35B-A3B-NVFP4-MTP-GGUF
  - datalab-to/chandra
  - peculiar-ragdoll/Tiel-Coder-35B-A3B-GGUF
  - meta-llama/Llama-3.3-70B-Instruct
  - microsoft/Phi-4-mini-instruct
  - Qwen/Qwen3-0.6B-GGUF
  - empero-ai/Qwen3.8-27B-Ridge-GGUF
  - nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16
  - typhoon-ai/typhoon-ocr1.5-2b
  - tvall43/Qwen3.6-14B-A3B-FableVibes-GGUF
  - AbelZimba/whisper-bemba-stt
  - voyageai/voyage-4-nano
  - Jackrong/Qwopus3.6-27B-Coder-Compat-MTP-GGUF
  - Qwen/Qwen3-30B-A3B-Instruct-2507-FP8
  - zai-org/GLM-5.1
  - google/gemma-4-26B-A4B-it-qat-q4_0-gguf
  - Viggle/Qwen-Image-2.1-viggle-turbo
  - Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign
  - nvidia/Qwen3.8-Flash-Next-NVFP4
  - LiquidAI/LFM2.5-1.2B-Instruct-GGUF
  - unsloth/Qwen3.5-2B-GGUF
  - unsloth/FLUX.2-klein-4B-GGUF
  - unsloth/Kimi-K3-GGUF
  - handy-computer/whisper-large-v3-gguf
  - moonshotai/Kimi-K2-Instruct
  - zai-org/GLM-5.2-FP8
  - ornith-ai/Ornith-1.5-397B-FP8
  - sanskar003/Qwen3.5-4B-AWQ
  - Qwen/Qwen3-ASR-1.7B-hf
  - nvidia/NVIDIA-Nemotron-Nano-9B-v2
  - moonshotai/Kimi-K2.5
  - unsloth/Qwen3-VL-4B-Instruct-GGUF
  - AnkitAI/Parable-Qwen3-8B-Claude-Fable-5-GGUF
  - BAAI/bge-multilingual-gemma2
  - deepseek-ai/DeepSeek-R1-Distill-Qwen-14B
  - ibm-granite/granite-3.3-2b-instruct
  - yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF
  - unsloth/Qwen-Image-2512-GGUF
  - alpindale/Llama-Guard-3-1B
  - unsloth/gemma-4-E2B-it-qat-GGUF
  - unsloth/Qwen-AgentWorld-35B-A3B-GGUF
  - Qwen/Qwen3-VL-32B-Instruct
  - llm-jp/llm-jp-4-33b-thinking-gguf
  - unsloth/Kimi-K2.7-Code-GGUF
  - bartowski/XYZAILab_XYZ-Aquila-mini-GGUF
  - google/gemma-4-31B-it-qat-q4_0-unquantized
  - RedHatAI/gemma-4-26B-A4B-it-NVFP4
  - allenai/OLMo-2-0425-1B
  - ggml-org/gemma-4-26B-A4B-it-GGUF
  - Qwen/Qwen2-7B-Instruct
  - Tencent-Hunyuan/HunyuanDiT-v1.1-Diffusers-Distilled
  - Qwen/Qwen2.5-Coder-7B-Instruct-GGUF
  - Qwen/Qwen2.5-Omni-7B
  - google/gemma-4-E2B-it-qat-w4a16-ct
  - abenzerps/Spark-X2.5-4B-GGUF
  - QuantTrio/Qwen3-Coder-30B-A3B-Instruct-AWQ
  - unsloth/gemma-4-31B-it-qat-GGUF
  - Qwen/Qwen3.5-9B-Base
  - peculiar-ragdoll/Sharp-Spark-X2.5-4B-GGUF
  - XLabs-AI/xflux_text_encoders
  - Qwen/Qwen3.5-2B-Base
  - black-forest-labs/FLUX.2-klein-base-4B
  - unsloth/GLM-5.2-GGUF
  - agentionai/Qwen3.8-27B-AP-GGUF
  - Qwen/Qwen3-VL-235B-A22B-Instruct
  - bartowski/DeepSeek-Coder-V2-Lite-Instruct-GGUF
  - RadixArk/Qwen3.8-Flash-Next-NVFP4
  - Qwen/Qwen3.8-Flash-Next-FP8
  - microsoft/harrier-oss-v1-0.6b
  - poolside/Laguna-XS.2
  - Qwen/Qwen2.5-1.5B-Instruct-GGUF
  - meta-llama/Llama-3.2-3B
  - utter-project/EuroLLM-22B-Instruct-2512
  - drbaph/Higgs-Audio-v3-Studio
  - janhq/Jan-v3.5-4B-gguf
  - RedHatAI/gemma-3-27b-it-quantized.w4a16
  - deepseek-ai/DeepSeek-R1-Distill-Qwen-7B
  - unsloth/gemma-4-31B-it-GGUF
  - philbert440/Qwen3.8-27B-W4A16-AWQ
  - Qwen/Qwen2.5-Math-1.5B-Instruct
  - RedHatAI/Qwen3-Coder-Next-FP8-dynamic
  - Qwen/Qwen2.5-72B-Instruct
  - incoai/Qwen3.8-27B-DFlash2
  - Qwen/Qwen2.5-Omni-3B
  - superwhisper/s1-mini-GGUF
  - LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct
  - RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead
  - AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF
  - bartowski/Kwaipilot_KAT-Coder-V2.5-Dev-GGUF
  - mixedbread-ai/deepset-mxbai-embed-de-large-v1
  - empero-ai/Qwythos-27B-v1-GGUF
  - KyleHessling1/Qwopus3.6-27B-Fusion-GGUF
  - poolside/Laguna-M.1
  - playgroundai/playground-v2.5-1024px-aesthetic
  - huggyllama/llama-7b
  - nanonets/Nanonets-OCR2-3B
  - thinkingmachines/Inkling-Small-NVFP4
  - ornith-ai/Ornith-1.0-397B-FP8
  - webAI-Official/TwIL-LM3
  - google/gemma-4-31B-it-qat-q4_0-gguf
  - Serveurperso/Qwen3-TTS-GGUF
  - MiniMaxAI/MiniMax-M2
  - Lightricks/LTX-2
  - ornith-ai/Ornith-1.0-35B-FP8
  - mesolitica/llama2-embedding-1b-8k
  - bartowski/kai-os_Grug-12B-GGUF
  - meta-models/Muse-Glimmer-30B
  - openbmb/MiniCPM5-2B-GGUF
  - Qwen/Qwen3-Next-80B-A3B-Instruct
  - google/gemma-4-31B-it-assistant
  - LilaRest/gemma-4-31B-it-NVFP4-turbo
  - Qwen/Qwen3-0.6B-FP8
  - Qwen/Qwen-Image
  - BDRC/tibetan-ocr
  - Qwen/Qwen2.5-3B
  - ornith-ai/Ornith-1.5-35B-A3B
  - FINAL-Bench/POCKET-26B-GGUF
  - EleutherAI/pythia-410m
  - AtomicChat/Ling-3.0-flash-GGUF
  - ibm-granite/granite-4.1-30b
  - lightonai/LightOnOCR-2-1B
  - Qwen/Qwen-Image-Edit-2511
  - GnLOLot/MiniCPM5-1B-Claude-Opus-Fable5-Thinking-GGUF
  - bloomer010/Ling-3.0-tiny-GGUF
  - MaziyarPanahi/Qwen3-0.6B-GGUF
  - GnLOLot/MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-GGUF
  - google/gemma-4-31B-it-qat-w4a16-ct
  - poolside/Laguna-XS-2.1
  - unsloth/Wan2.2-TI2V-5B-GGUF
  - AngelSlim/Hy3-GGUF
  - Qwen/Qwen1.5-MoE-A2.7B
  - sdadas/mmlw-retrieval-roberta-large
  - nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-FP8
  - ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF
  - stepfun-ai/Step-3.7-Flash-NVFP4
  - unsloth/diffusiongemma-26B-A4B-it-GGUF
  - unsloth/Qwen2.5-14B-bnb-4bit
  - google/gemma-4-E2B
  - ATH-MaaS/OvisOCR2
  - stelterlab/Mistral-Small-24B-Instruct-2501-AWQ
  - Wan-AI/Wan2.1-T2V-1.3B-Diffusers
  - openbmb/MiniCPM-V-4.6
  - ChristianAzinn/mxbai-embed-large-v1-gguf
  - Qwen/Qwen3-4B-AWQ
  - ai4bharat/indic-parler-tts
  - google/paligemma-3b-pt-224
  - QuantTrio/Qwen3.5-27B-AWQ
  - ornith-ai/Ornith-1.5-397B
  - Qwen/Qwen2.5-0.5B-Instruct-GGUF
  - Jackrong/Qwopus3.6-27B-Coder-MTP-GGUF
  - opendatalab/MinerU2.5-Pro-2605-1.2B
  - cyankiwi/Qwen3.6-35B-A3B-AWQ-4bit
  - utter-project/EuroLLM-1.7B-Instruct
  - RedHatAI/Llama-3.2-1B-Instruct-FP8-dynamic
  - XHToken/Spark-X2.5-1.7B-GGUF
  - google/medgemma-1.5-4b-it
  - InternScience/Agents-A1-4B
  - unsloth/gemma-4-E2B-it-qat-mobile-GGUF
  - MaziyarPanahi/Qwen3-14B-GGUF
  - ai-forever/FRIDA
  - unsloth/Llama-3.2-3B-Instruct
  - RedHatAI/gemma-4-31B-it-FP8-dynamic
  - InternScience/Agents-A1-4B-Q8_0-GGUF
  - intfloat/e5-mistral-7b-instruct
  - unsloth/Muse-Glimmer-30B-GGUF
  - moonshotai/Kimi-VL-A3B-Instruct
  - MaziyarPanahi/Qwen3-4B-GGUF
  - Qwen/Qwen3-VL-32B-Instruct-FP8
  - NousResearch/Meta-Llama-3.1-8B-Instruct
  - zenosai/MonkeyOCRv2-B-Parsing
  - QuantStack/Wan2.2-TI2V-5B-GGUF
  - MaziyarPanahi/Qwen3-1.7B-GGUF
  - Qwen/Qwen2.5-7B-Instruct-GPTQ-Int4
  - moonshotai/Kimi-Linear-48B-A3B-Instruct
  - MaziyarPanahi/Qwen3-8B-GGUF
  - XiaomiMiMo/MiMo-V2.5
  - jinaai/jina-embeddings-v5-text-small
  - bartowski/Qwen_Qwen3-4B-GGUF
  - HuggingFaceTB/SmolLM2-1.7B-Instruct
  - stepfun-ai/Step-3.5-Flash
  - MaziyarPanahi/Qwen3-30B-A3B-GGUF
  - MaziyarPanahi/Qwen3-32B-GGUF
  - skt/A.X-K2-NVFP4
  - QuantTrio/Qwen3.5-4B-AWQ
  - Hcompany/Holo-3.1-35B-A3B-GGUF
  - inclusionAI/LLaDA2.0-mini
  - Qwen/Qwen2.5-VL-32B-Instruct-AWQ
  - Qwen/Qwen3-14B-GGUF
  - dots-studio/dots.mocr
  - AnkitAI/Parable-Granite-4.1-3B-Claude-Fable-5-GGUF
  - AnkitAI/Parable-Granite-4.1-8B-Claude-Fable-5-GGUF
  - nvidia/LocateAnything-3B
  - unsloth/Qwen3.5-0.8B-GGUF
  - Qwen/Qwen2.5-14B-Instruct-GPTQ-Int4
  - poolside/Laguna-S-2.1
  - nvidia/NVIDIA-Nemotron-Nano-9B-v2-FP8
  - unsloth/Meta-Llama-3.1-8B-Instruct
  - Zyphra/Zamba2-1.2B-instruct
  - poolside/Laguna-XS-2.1-NVFP4
  - MiniMaxAI/MiniMax-M3
  - meta-llama/Meta-Llama-3-8B
  - QCRI/Fanar-2-27B-Instruct
  - NousResearch/Meta-Llama-3.1-8B
  - ggml-org/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF
  - openbmb/MiniCPM-o-2_6
  - cyankiwi/gemma-4-12B-it-AWQ-INT4
  - MaziyarPanahi/Yi-Coder-9B-Chat-GGUF
  - Qwen/Qwen2-7B
  - microsoft/VibeVoice-Realtime-0.5B
  - cyankiwi/Qwen3-30B-A3B-Instruct-2507-AWQ-4bit
  - Qwen/Qwen2.5-3B-Instruct-GGUF
  - deepseek-ai/DeepSeek-R1-0528
  - Qwen/Qwen3.5-35B-A3B-GPTQ-Int4
  - opendatalab/MinerU2.5-Pro-2604-1.2B
  - black-forest-labs/FLUX.2-klein-9B
  - nvidia/Nemotron-3-Embed-1B-BF16
  - poolside/Laguna-XS-2.1-GGUF
  - meta-llama/Llama-2-13b-chat-hf
  - bytkim/Qwen3.6-27B-MTP-pi-tune-GGUF
  - Qwen/Qwen3-Omni-30B-A3B-Thinking
  - poolside/Laguna-S-2.1-NVFP4
  - RadixArk/Qwen3.8-27B-DSpark
  - Qwen/Qwen2.5-Math-1.5B
  - bartowski/Qwen2.5-7B-Instruct-GGUF
  - city96/FLUX.1-dev-gguf
  - deepseek-ai/DeepSeek-V2-Lite
  - ReadyArt/gemma-4-31B-it-scotoma-2-GGUF
  - cyankiwi/Qwen3.5-27B-AWQ-4bit
  - deepseek-ai/DeepSeek-R1-Distill-Llama-8B
  - antirez/qwen3.8-flash-next-gguf
  - Laxhar/noobai-XL-1.1
  - RedHatAI/Llama-3.2-3B-Instruct-FP8
  - ukisai/Swift-1.5-Qwen3.8-27B-GGUF
  - Myric/Laguna-S-2.1-APEX-GGUF
  - ornith-ai/Ornith-1.0-397B
  - unsloth/Qwen2.5-VL-7B-Instruct-GGUF
  - bartowski/Qwen_Qwen3.6-35B-A3B-GGUF
  - unsloth/Qwen2.5-Coder-7B-Instruct
  - Lorbus/Qwen3.6-27B-int4-AutoRound
  - chatdb/natural-sql-7b
  - unsloth/Qwen2.5-7B-Instruct
  - ibm-granite/granite-vision-4.1-4b
  - LSX-UniWue/LLaMmlein_1B_prerelease
  - Qwen/Qwen3.5-27B-FP8
  - InternScience/Agents-A1-4B-Q4_K_M-GGUF
  - Qwen/Qwen3Guard-Gen-0.6B
  - Qwen/Qwen3-VL-4B-Instruct-FP8
  - unsloth/gemma-4-E4B-it-unsloth-bnb-4bit
  - leejet/MiniMax-H3-GGUF
  - Qwen/Qwen3-235B-A22B-Instruct-2507
  - Qwen/Qwen3-30B-A3B-FP8
  - JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-Q8_0
  - google/gemma-3n-E2B-it
  - Qwen/Qwen3-VL-4B-Instruct-GGUF
  - MaziyarPanahi/Qwen3-4B-Instruct-2507-GGUF
  - meta-llama/Llama-4-Scout-17B-16E-Instruct
  - omlab/VLX-Seek-1.5-10B
  - google/gemma-2-2b
  - nvidia/NVIDIA-Nemotron-Nano-12B-v2-VL-NVFP4-QAD
  - Jackrong/Qwen3.5-9B-DeepSeek-V4-Flash-GGUF
  - nvidia/Cosmos-Reason2-8B
  - nvidia/Kimi-K2-Thinking-NVFP4
  - bartowski/Qwen_Qwen3.5-27B-GGUF
  - ibm-granite/granite-embedding-311m-multilingual-r2
  - EleutherAI/pythia-1b
  - ai-forever/sbert_large_nlu_ru
  - zai-org/GLM-4.1V-9B-Thinking
  - unsloth/GLM-4.7-Flash-GGUF
  - mistral-experimental/pixtral-12b
  - pfnet/plamo-embedding-1b
  - hugging-quants/Meta-Llama-3.1-8B-Instruct-AWQ-INT4
  - dreamgen/lucid-v1-nemo
  - cyankiwi/Qwen3-VL-30B-A3B-Instruct-AWQ-4bit
  - unsloth/Step-3.7-Flash-GGUF
  - hugging-quants/Meta-Llama-3.1-70B-Instruct-AWQ-INT4
  - z-lab/Qwen3.8-27B-DFlash2
  - TIGER-Lab/VLM2Vec-Full
  - boboliu/Qwen3-Embedding-4B-W4A16-G128
  - QuantStack/Wan2.2-T2V-A14B-GGUF
  - microsoft/Phi-3-mini-128k-instruct
  - mixedbread-ai/mxbai-embed-2d-large-v1
  - Jackrong/Qwopus3.5-9B-Coder-MTP-GGUF
  - nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4
  - datalab-to/lift
  - EleutherAI/pythia-1.4b
  - ai-forever/ru-en-RoSBERTa
  - bartowski/gemma-2-27b-it-GGUF
  - unsloth/gemma-4-26B-A4B-it-NVFP4
  - batiai/Qwen3.6-27B-GGUF
  - ibm-granite/granite-4.1-8b
  - jinaai/jina-embeddings-v3-hf
  - bartowski/Qwen_Qwen3.5-4B-GGUF
  - ggml-org/gpt-oss-20b-GGUF
  - bartowski/Llama-3.2-3B-Instruct-GGUF
  - Wan-AI/Wan2.2-TI2V-5B-Diffusers
  - wikeeyang/Flux2-Klein-9B-True-V2
  - allenai/wildguard
  - MaziyarPanahi/Meta-Llama-3-8B-Instruct-GGUF
  - SimianLuo/LCM_Dreamshaper_v7
  - MaziyarPanahi/Mistral-7B-v0.1-GGUF
  - QuantTrio/Qwen3-Coder-30B-A3B-Instruct-GPTQ-Int8
  - mistralai/Mistral-7B-Instruct-v0.1
  - unsloth/ERNIE-Image-Turbo-GGUF
  - cagliostrolab/animagine-xl-3.1
  - Jackrong/Qwopus3.6-27B-v2-MTP-GGUF
  - unsloth/Qwen3-VL-8B-Instruct-GGUF
  - ReliquaryForge/qwen3-4b-base-dapo-v4
  - nvidia/NVIDIA-Nemotron-Parse-v1.1
  - MaziyarPanahi/Mistral-Small-Instruct-2409-GGUF
  - MaziyarPanahi/Mistral-Nemo-Instruct-2407-GGUF
  - Qwen/Qwen2.5-7B-Instruct-GGUF
  - unsloth/Qwen3-14B-unsloth-bnb-4bit
  - MaziyarPanahi/Yi-Coder-1.5B-Chat-GGUF
  - Wan-AI/Wan2.2-I2V-A14B-Diffusers
  - ggml-org/gpt-oss-120b-GGUF
  - gaunernst/gemma-3-12b-it-int4-awq
  - Qwen/Qwen3-1.7B-GGUF
  - MaziyarPanahi/firefunction-v2-GGUF
  - unsloth/Qwen3-Coder-Next-GGUF
  - Qwen/Qwen1.5-1.8B-Chat
  - QUASAR-QAT/Qwen3.8-27B-QUASAR-NVFP4
  - unsloth/Qwen3.5-27B-GGUF
  - abhishekchohan/gemma-3-12b-it-quantized-W4A16
  - ukisai/Swift-1.5-Qwen3.8-Flash-Next-GGUF
  - nvidia/Qwen3-8B-NVFP4
  - Qwen/Qwen3-VL-8B-Thinking
  - HuggingFaceTB/SmolVLM2-2.2B-Instruct
  - MaziyarPanahi/Qwen2-7B-Instruct-GGUF
  - unsloth/Qwen3-0.6B-GGUF
  - xinsir/controlnet-openpose-sdxl-1.0
  - QuantTrio/Qwen3-VL-32B-Instruct-AWQ
  - leejet/Qwen-Image-2.1-GGUF
  - unsloth/Llama-3.1-8B-Instruct
  - microsoft/Phi-3.5-MoE-instruct
  - RedHatAI/Meta-Llama-3.1-70B-Instruct-FP8
  - NousResearch/Llama-2-7b-hf
  - MaziyarPanahi/Phi-3.5-mini-instruct-GGUF
  - bullerwins/Wan2.2-I2V-A14B-GGUF
  - MaziyarPanahi/Mistral-7B-Instruct-v0.3-GGUF
  - MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF
  - John6666/nova-furry-xl-il-v120-sdxl
  - MaziyarPanahi/Mixtral-8x22B-v0.1-GGUF
  - MaziyarPanahi/phi-4-GGUF
  - MaziyarPanahi/Meta-Llama-3.1-8B-Instruct-GGUF
  - Wan-AI/Wan2.2-T2V-A14B-Diffusers
  - intfloat/e5-large
  - bartowski/Qwen_Qwen3.5-9B-GGUF
  - MaziyarPanahi/gemma-3-4b-it-GGUF
  - MaziyarPanahi/solar-pro-preview-instruct-GGUF
  - Qwen/Qwen3-VL-2B-Instruct-GGUF
  - google/gemma-2b
  - allenai/OLMoE-1B-7B-0125-Instruct
  - deepseek-ai/DeepSeek-V2-Lite-Chat
  - MaziyarPanahi/Llama-3-8B-Instruct-32k-v0.1-GGUF
  - Jackrong/Qwopus3.8-27B-Flash-GGUF
  - nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark
  - CohereLabs/North-Micro-Vision-Instruct
  - cyankiwi/gemma-4-31B-it-AWQ-4bit
  - ibm-granite/granite-4.2-8b
  - city96/Wan2.1-I2V-14B-480P-gguf
  - Qwen/Qwen2.5-Coder-1.5B-Instruct-GGUF
  - HuggingFaceTB/SmolVLM-500M-Instruct
  - LiquidAI/LFM2.5-2.6B-DSpark-GGUF
  - sesame/csm-1b
  - MaziyarPanahi/Qwen2.5-7B-Instruct-GGUF
  - unsloth/FLUX.2-dev-GGUF
  - RedHatAI/DeepSeek-Coder-V2-Lite-Instruct-FP8
  - MaziyarPanahi/QwQ-32B-GGUF
  - MaziyarPanahi/DeepSeek-R1-0528-Qwen3-8B-GGUF
  - unsloth/gemma-3n-E4B-it
  - MaziyarPanahi/Mistral-Small-24B-Instruct-2501-GGUF
  - xinsir/controlnet-union-sdxl-1.0
  - bartowski/Llama-3.2-1B-Instruct-GGUF
  - Qwen/Qwen3-Next-80B-A3B-Instruct-FP8
  - google/gemma-3-4b-pt
  - MaziyarPanahi/Phi-4-mini-instruct-GGUF
  - MaziyarPanahi/Llama-3.2-3B-Instruct-GGUF
  - OpenMOSS-Team/MOSS-TTS-v1.5
  - MaziyarPanahi/Ministral-3-3B-Reasoning-2512-GGUF
  - allenai/Olmo-3-1025-7B
  - MaziyarPanahi/gemma-2-2b-it-GGUF
  - unsloth/Llama-3.2-1B
  - MaziyarPanahi/Qwen2.5-1.5B-Instruct-GGUF
  - MaziyarPanahi/gemma-3-1b-it-GGUF
  - MaziyarPanahi/Llama-3.3-70B-Instruct-GGUF
  - RedHatAI/Qwen2.5-1.5B-quantized.w8a8
  - RedHatAI/Meta-Llama-3.1-8B-Instruct-FP8
  - MaziyarPanahi/Yi-1.5-6B-Chat-GGUF
  - stabilityai/stable-video-diffusion-img2vid-xt
  - MaziyarPanahi/WizardLM-2-7B-GGUF
  - MaziyarPanahi/gemma-3-12b-it-GGUF
  - MaziyarPanahi/mistral-small-3.1-24b-instruct-2503-hf-GGUF
  - moonshotai/Kimi-Linear-48B-A3B-Base
  - MaziyarPanahi/Llama-3-8B-Instruct-64k-GGUF
  - zai-org/GLM-4.5
  - MaziyarPanahi/mathstral-7B-v0.1-GGUF
  - unsloth/Z-Image-GGUF
  - MaziyarPanahi/Meta-Llama-3.1-70B-Instruct-GGUF
  - MaziyarPanahi/DeepSeek-V3-0324-GGUF
  - MaziyarPanahi/gemma-3-27b-it-GGUF
  - dragonkue/snowflake-arctic-embed-l-v2.0-ko
  - Qwen/Qwen2-1.5B
  - MaziyarPanahi/Mistral-Large-Instruct-2411-GGUF
  - MaziyarPanahi/INTELLECT-2-GGUF
  - MaziyarPanahi/Meta-Llama-3.1-405B-Instruct-GGUF
  - LiquidAI/LFM2.5-2.6B
  - Qwen/Qwen3-32B-FP8
  - inclusionAI/LLaDA2.1-mini
  - ThorOdinson246/nl2sh-1.5b-Q4_K_M
  - latam-gpt/Llama-3.1-70B-LatamGPT-SFT-1.0
  - SG161222/RealVisXL_V4.0
  - meta-llama/Meta-Llama-3-70B-Instruct
  - RedHatAI/Llama-3.3-70B-Instruct-FP8-dynamic
  - deepseek-ai/deepseek-coder-6.7b-instruct
  - aari1995/German_Semantic_STS_V2
  - deepseek-ai/deepseek-coder-6.7b-base
  - prism-ml/Bonsai-8B-gguf
  - Qwen/Qwen2.5-Coder-14B-Instruct-GGUF
  - unsloth/Phi-4-mini-instruct-GGUF
  - Qwen/Qwen-Image-2.1
  - LiquidAI/LFM2.5-1.2B-Instruct
  - Snowflake/snowflake-arctic-embed-m-v2.0
  - MuXodious/gpt-oss-20b-RichardErkhov-heresy
  - city96/FLUX.1-schnell-gguf
  - SebastianBodza/Kartoffel_Orpheus-3B_german_natural-v0.1
  - Qwen/Qwen2.5-1.5B-Instruct-AWQ
  - Qwen/Qwen-Image-Edit
  - canopylabs/3b-de-ft-research_release
  - allenai/OLMoE-1B-7B-0924
  - audio-cpp/AuK-Base-and-Flash-GGUF
  - QuantStack/Qwen-Image-Edit-2509-GGUF
  - Alibaba-NLP/gte-Qwen2-7B-instruct
  - hmellor/Ilama-3.2-1B
  - tencent/Hy3-preview
  - sarvamai/sarvam-30b
  - unsloth/Qwen3-4B-Instruct-2507-unsloth-bnb-4bit
  - KyleHessling1/Qwopus-GLM-18B-Merged-GGUF
  - google/gemma-4-12B
  - lightseekorg/kimi-k2.6-eagle3-mla
  - Qwen/Qwen2.5-Coder-3B-Instruct-GGUF
  - unsloth/Qwen3-1.7B-GGUF
  - unsloth/Qwen3-Coder-Next-FP8
  - Wan-AI/Wan2.1-I2V-14B-480P
  - jinaai/jina-embeddings-v5-text-small-retrieval
  - HuggingFaceTB/SmolLM3-3B-Base
  - diffusers/stable-diffusion-xl-1.0-inpainting-0.1
  - unsloth/Qwen2.5-7B-Instruct-bnb-4bit
  - pentacoxian-dev/Qwen3.8-Flash-Next-IQ3E-Q8D-MTP-GGUF
  - NousResearch/Meta-Llama-3-8B-Instruct
  - bosonai/higgs-tts-3-4b
  - Qwen/Qwen3-30B-A3B-GGUF
  - SulphurAI/Sulphur-2-base
  - PrimeIntellect/Qwen3-1.7B
  - John6666/amanatsu-illustrious-v11-sdxl
  - dealignai/GLM-5.3-CYBERSECURITY-FP8
  - stabilityai/stable-diffusion-3.5-large
  - stabilityai/stable-diffusion-xl-refiner-1.0
  - unsloth/Llama-3.2-3B-Instruct-GGUF
  - bartowski/Qwen2.5-14B-Instruct-GGUF
  - DevQuasar/amd.Instella-MoE-16B-A3B-Think-GGUF
  - Qwen/Qwen3-32B-GGUF
  - h2oai/h2o-danube3-500m-chat
  - dealignai/Bonsai-2-27B-Ternary-CRACK-GGUF
  - Nanbeige/Nanbeige4.2-3B
  - nvidia/NV-Embed-v2
  - MaziyarPanahi/Qwen3-30B-A3B-Instruct-2507-GGUF
  - littlejohn-ai/bge-m3-spa-law-qa
  - meta-llama/Meta-Llama-3-70B
  - JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-Q4_K_M
  - Salesforce/Llama-xLAM-2-8b-fc-r
  - bartowski/Meta-Llama-3.1-70B-Instruct-GGUF
  - RamManavalan/Qwen3-VL-Embedding-8B-FP8
  - XiaomiMiMo/MiMo-V2-Flash
  - cais/HarmBench-Llama-2-13b-cls
  - molbal/ideogram-4-gguf
  - zai-org/GLM-4.7
  - CohereLabs/c4ai-command-r-v01
  - XiaomiMiMo/MiMo-V2.6-Pro-RL
  - LLMSafety/Qwen2.5-Math-7B-4bit
  - unsloth/DeepSeek-R1-0528-Qwen3-8B-GGUF
  - black-forest-labs/FLUX.2-klein-base-9B
  - unsloth/llama-3-8b-Instruct-bnb-4bit
  - nvidia/diffusiongemma-26B-A4B-it-NVFP4
  - HuggingFaceH4/zephyr-7b-beta
  - Abiray/Qwen-Image-2.1-GGUF
  - silveroxides/Chroma-GGUF
  - badtheorylabs/BTL-4-Compact
  - lmstudio-community/Qwen3-Coder-30B-A3B-Instruct-GGUF
  - unsloth/Qwen3-8B-GGUF
  - krea/Krea-2-Turbo
  - hyrelabs/Homura-30B-GGUF
  - stabilityai/stable-diffusion-3.5-medium
  - zenmagnets/Nemotron-3-Embed-1B-Q4_K_M-GGUF
  - NovaSearch/stella_en_400M_v5
  - deepseek-ai/DeepSeek-V4-Pro-0813
  - nvidia/GLM-5.3-Flash-NVFP4
  - OnomaAIResearch/Illustrious-xl-early-release-v0
  - ibm-granite/granite-speech-4.1-2b-nar
  - Qwen/Qwen3-14B-Base
  - unsloth/gpt-oss-20b-BF16
  - baidu/ERNIE-4.5-21B-A3B-Thinking
  - city96/FLUX.2-dev-gguf
  - ibm-granite/granite-3.0-1b-a400m-instruct
  - unsloth/Qwen2.5-3B-Instruct-unsloth-bnb-4bit
  - internlm/internlm2-chat-7b
  - meta-llama/Llama-3.1-70B-Instruct
  - krea/Krea-2-Raw
  - Qwen/Qwen2.5-14B
  - unsloth/Meta-Llama-3.1-8B
  - OpenMOSS-Team/MOSS-TTS-Local-Transformer-v1.5
  - unsloth/Qwen2.5-7B-Instruct-unsloth-bnb-4bit
  - unsloth/Qwen3-8B
  - meta-llama/Llama-3.1-405B
  - Aratako/MioTTS-2.6B
  - ibm-granite/granite-3.3-8b-instruct
  - bartowski/Qwen_Qwen3-0.6B-GGUF
  - bartowski/DeepSeek-R1-Distill-Qwen-7B-GGUF
  - Wan-AI/Wan2.1-T2V-14B-Diffusers
  - Wan-AI/Wan2.1-I2V-14B-720P-Diffusers
  - jinaai/jina-clip-v2
  - Qwen/Qwen3-30B-A3B-Thinking-2507
  - vantagewithai/Krea-2-Turbo-GGUF
  - Qwen/Qwen3-4B-FP8
  - openai-community/gpt2-xl
  - Qwen/Qwen2.5-Math-7B-Instruct
  - Jackrong/Qwen3.5-4B-Claude-4.6-Opus-Reasoning-Distilled-GGUF
  - nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-Base-BF16
  - bartowski/microsoft_Phi-4-mini-instruct-GGUF
  - Dream-org/Dream-v0-Instruct-7B
  - leejet/Z-Image-Turbo-GGUF
  - Xenova/sweep-next-edit-1.5B
  - unsloth/gpt-oss-20b-unsloth-bnb-4bit
  - bartowski/Qwen2.5-1.5B-Instruct-GGUF
  - unsloth/Qwen2.5-Coder-7B-Instruct-bnb-4bit
  - nvidia/Llama-3_3-Nemotron-Super-49B-v1_5-FP8
  - frankjoshua/novaAnimeXL_ilV140
  - unsloth/Qwen-Image-GGUF
  - nytopop/Qwen3-30B-A3B.w8a8
  - RadixArk/GLM-5.3-NVFP4
  - bartowski/gemma-2-2b-it-GGUF
  - Wan-AI/Wan2.1-VACE-14B
  - MaziyarPanahi/GLM-4.6V-Flash-GGUF
  - lmstudio-community/Meta-Llama-3.1-8B-Instruct-GGUF
  - OpenMOSS-Team/MOSS-Audio-Tokenizer
  - MaziyarPanahi/gpt-oss-20b-Derestricted-GGUF
  - laion/voiceclap-large-v2
  - lmstudio-community/Llama-3.2-3B-Instruct-GGUF
  - MaziyarPanahi/Qwen3-Coder-Next-GGUF
  - Qwen/Qwen1.5-7B
  - nvidia/Llama-3_3-Nemotron-Super-49B-v1_5
  - unsloth/gpt-oss-120b-GGUF
  - bartowski/Qwen2.5-3B-Instruct-GGUF
  - 0xSero/deepseek-v4-flash-0731-spark
  - MaziyarPanahi/Nemotron-Orchestrator-8B-GGUF
  - fla-hub/transformer-1.3B-100B
  - lmstudio-community/gemma-3-1b-it-GGUF
  - MaziyarPanahi/Trinity-Mini-GGUF
  - unsloth/Qwen3-32B-GGUF
  - QuantStack/Wan2.1_14B_VACE-GGUF
  - tencent/Hunyuan-A13B-Instruct
  - FastVideo/FastWan2.2-TI2V-5B-FullAttn-Diffusers
  - Wan-AI/Wan2.1-I2V-14B-480P-Diffusers
  - nvidia/Nemotron-Labs-Diffusion-8B-Base
  - Qwen/Qwen1.5-0.5B-Chat
  - nvidia/Nemotron-Labs-Diffusion-3B
  - Inferact/GLM-5.3-NVFP4
  - internlm/internlm3-8b-instruct
  - Lykon/dreamshaper-8
  - unsloth/DeepSeek-R1-Distill-Qwen-1.5B-GGUF
  - bartowski/Qwen2.5-72B-Instruct-GGUF
  - LiquidAI/LFM2-1.2B
  - prism-ml/Bonsai-1.7B-gguf
  - unsloth/Meta-Llama-3.1-8B-Instruct-bnb-4bit
  - unsloth/Qwen3-14B-GGUF
  - TheBloke/Llama-2-7B-Chat-GGUF
  - Qwen/Qwen-7B
  - bartowski/meta-llama_Llama-4-Scout-17B-16E-Instruct-old-GGUF
  - TinyLlama/TinyLlama-1.1B-intermediate-step-1431k-3T
  - John6666/diving-illustrious-real-asian-v50-sdxl
  - LiquidAI/LFM2.5-8B-A1B-DSpark-GGUF
  - tobil/qmd-query-expansion-1.7B-gguf
  - LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-GGUF
  - sbintuitions/sarashina2.2-0.5b-instruct-v0.1
  - stabilityai/stable-diffusion-3-medium-diffusers
  - agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF
  - moonshotai/Kimi-K2-Instruct-0905
  - CMSManhattan/JiRackUltra_7b
  - NousResearch/Meta-Llama-3-8B
  - CMSManhattan/JiRackUltra_32b
  - joeygambino/MiniMax-H3-encoder-GGUF
  - nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8
  - ibm-granite/granite-4.1-8b-fp8
  - moonshotai/Kimi-K2-Thinking
  - Qwen/Qwen3Guard-Gen-4B
  - Alittlehammmer/Qwen3.6-35B-A3B-DFlash-GGUF-llama.cpp
  - black-forest-labs/FLUX.2-klein-9b-kv
  - unsloth/Qwen3-1.7B-unsloth-bnb-4bit
  - Qwen/Qwen3-30B-A3B-GPTQ-Int4
  - jdopensource/JoyAI-Image-Edit-Diffusers
  - unsloth/Qwen3-Next-80B-A3B-Instruct-GGUF
  - RedHatAI/Llama-3.2-1B-Instruct-FP8
  - unsloth/Llama-3.2-1B-Instruct-bnb-4bit
  - hesamation/Qwen3.6-35B-A3B-Claude-4.6-Opus-Reasoning-Distilled-GGUF
  - deepseek-ai/DeepSeek-R1-Distill-Llama-70B
  - arianraje/qwen3-4b-mamba3-hybrid-stage2b-kd-bias0
  - AtomicChat/Qwen3-4B-DFlash-GGUF
  - molbal/MiniMax-H3-GGUF
  - John6666/one-obsession-17-red-sdxl
  - lmstudio-community/DeepSeek-R1-0528-Qwen3-8B-GGUF
  - AtomicChat/Qwen3.5-9B-DFlash-GGUF
  - casperhansen/llama-3-8b-instruct-awq
  - microsoft/Phi-tiny-MoE-instruct
  - AtomicChat/Qwen3.5-4B-DFlash-GGUF
  - allenai/OLMo-2-0425-1B-Instruct
  - Alittlehammmer/Qwen3.6-27B-DFlash-GGUF-llama.cpp
  - microsoft/phi-1_5
  - martineux/dvine82-xl
  - kmhf/hf-moshiko
  - RadixArk/Inkling-Small-DSpark-Preview
  - bullerwins/FLUX.1-Kontext-dev-GGUF
  - Efficient-Large-Model/gemma-2-2b-it
  - DevQuasar-13/THUDM.GLM-Z1-32B-0414-GGUF
  - Wan-AI/Wan2.1-T2V-14B
  - perplexity-ai/pplx-embed-v1-0.6b
  - solidrust/Hermes-3-Llama-3.1-8B-AWQ
  - bartowski/Ling-3.0-tiny-GGUF
  - lodestones/Chroma1-HD
  - bartowski/Mistral-7B-Instruct-v0.3-GGUF
  - sahilchachra/gemma-4-12B-coder-fable5-composer2.5-AWQ
  - openthaigpt/openthaigpt1.5-7b-instruct
  - QuantFactory/Qwen2.5-Coder-7B-GGUF
  - RavichandranJ/Dolphin3-Cyber-8B-GGUF
  - ibm-granite/granite-4.0-h-tiny
  - tiiuae/Falcon-H1-0.5B-Base
  - hugging-quants/Llama-3.2-1B-Instruct-Q4_K_M-GGUF
  - lmstudio-community/Qwen3-4B-Instruct-2507-GGUF
  - John6666/obsession-illustriousxl-v10-sdxl
  - HuggingFaceTB/SmolLM-1.7B
  - OpenOneRec/OneRec-1.7B
  - unsloth/gemma-3-1b-it
  - nvidia/MiniMax-M3-NVFP4
  - microsoft/Phi-4-mini-reasoning
  - XCurOS/XCurOS0.1-8B-Instruct
  - bartowski/Qwen2.5-Coder-7B-Instruct-GGUF
  - unsloth/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF
  - avatargrim/Qwen2.5-7B_pyuigpt
  - XiaomiMiMo/MiMo-V2.6-Flash-RL
  - Qwen/Qwen3-14B-FP8
  - unsloth/Qwen3-4B-unsloth-bnb-4bit
  - SG161222/RealVisXL_V5.0_Lightning
  - zai-org/GLM-4.5-Air-FP8
  - ibm-granite/granite-guardian-4.1-8b
  - deepseek-ai/DeepSeek-V3.2-Exp
  - ibm-granite/granite-3.1-8b-instruct
  - google/gemma-3-1b-pt
  - LGAI-EXAONE/EXAONE-3.5-2.4B-Instruct
  - Tongyi-MAI/Z-Image
  - Abiray/LTX-2.5-Distilled-GGUF
  - HiDream-ai/HiDream-I1-Fast
  - ibm-granite/granite-4.0-tiny-preview
  - unsloth/DeepSeek-R1-Distill-Llama-70B-GGUF
  - nvidia/gpt-oss-120b-Eagle3-short-context
  - Qwen/Qwen3-30B-A3B-Base
  - allenai/OLMo-1B-hf
  - MaziyarPanahi/Qwen3-4B-Thinking-2507-GGUF
  - MaziyarPanahi/Ministral-3-14B-Reasoning-2512-GGUF
  - tencent/Hy3-FP8
  - MaziyarPanahi/NVIDIA-Nemotron-Nano-12B-v2-GGUF
  - ibm-granite/granite-3.2-8b-instruct
  - leejet/FLUX.2-klein-4B-GGUF
  - NousResearch/Meta-Llama-3-70B-Instruct
  - Lykon/dreamshaper-xl-v2-turbo
  - Qwen/Qwen-7B-Chat
  - nvidia/gpt-oss-120b-Eagle3-v3
  - entrick/Security-SLM-Gemma-4-E2B-it-GGUF
  - stabilityai/stable-diffusion-3.5-large-controlnet-canny
  - microsoft/Phi-3-vision-128k-instruct
  - ideogram-ai/ideogram-4-fp8
  - nvidia/Nemotron-H-8B-Base-8K
  - TheBloke/Mistral-7B-Instruct-v0.2-GGUF
  - bartowski/Phi-3.5-mini-instruct-GGUF
  - XingChen-AGI/Xing4.0-29B-A4B
  - abenzerps/Nex-N2.5-mini-GGUF
  - tencent/Hunyuan-7B-Instruct
  - google/codegemma-7b-it
  - Abiray/Qwen-Image-2.1-viggle-4-steps-turbo-GGUF
  - openbmb/MiniCPM5-1B-MLX
  - stabilityai/stable-diffusion-3.5-large-controlnet-depth
  - ByteDance-Seed/Seed-OSS-36B-Instruct
  - unsloth/Qwen2.5-32B-Instruct-bnb-4bit
  - ibnzterrell/Meta-Llama-3.3-70B-Instruct-AWQ-INT4
  - Wan-AI/Wan2.1-T2V-1.3B
  - unsloth/Llama-3.3-70B-Instruct
  - Qwen/Qwen2.5-7B-Instruct-1M
  - nvidia/Llama-3.1-8B-Instruct-FP8
  - QuantStack/FLUX.1-Kontext-dev-GGUF
  - vcruz305/DeepSeek-V4.1-Flash-GGUF
  - unsloth/Llama-3.2-3B-Instruct-unsloth-bnb-4bit
  - ibm-granite/granite-4.2-3b
  - bartowski/Altworld_Hemmingway-1-GGUF
  - NousResearch/Llama-3.2-1B
  - bigcode/starcoder2-3b
  - cagliostrolab/animagine-xl-3.0
  - nvidia/Mistral-NeMo-Minitron-8B-Instruct
  - Vikhrmodels/Vikhr-Nemo-12B-Instruct-R-21-09-24
  - wikeeyang/Flux2-Klein-9B-True-V3
  - bartowski/Qwen_Qwen3-4B-Instruct-2507-GGUF
  - IPostYellow/TurboWan2.1-T2V-1.3B-Diffusers
  - NousResearch/Hermes-2-Pro-Mistral-7B
  - John6666/prefect-illustrious-xl-v3-sdxl
  - EldanRing/Winnow-12B
  - z-lab/LLaMA3.1-8B-Instruct-DFlash-UltraChat
  - bartowski/Qwen_Qwen3-1.7B-GGUF
  - google/gemma-1.1-2b-it
  - bartowski/Qwen2.5-Coder-1.5B-Instruct-GGUF
  - Lightricks/LTX-2.5-Diffusers
  - EleutherAI/pythia-410m-deduped
  - unsloth/gemma-3-1b-it-GGUF
  - unsloth/Qwen2.5-1.5B-Instruct-unsloth-bnb-4bit
  - KBlueLeaf/TIPO-500M-ft
  - LGAI-EXAONE/EXAONE-3.5-32B-Instruct-AWQ
  - bartowski/THUDM_GLM-4-32B-0414-GGUF
  - incoai/GLM-5.3-Flash-DFlash2
  - openbmb/MiniCPM5-1B-GGUF
  - Qwen/Qwen-Image-2512
  - HuggingFaceTB/SmolLM2-1.7B
  - RedHatAI/Meta-Llama-3.1-8B-Instruct-quantized.w4a16
  - speakleash/Bielik-11B-v3.0-Instruct-awq
  - z-lab/Qwen3.6-35B-A3B-DFlash
  - lmstudio-community/Qwen2.5-Coder-14B-Instruct-GGUF
  - google/t5gemma-2b-2b-ul2-it
  - duyntnet/Chroma-GGUF-All-Versions
  - ibm-granite/granite-4.0-1b-base
  - unsloth/Qwen3-0.6B-unsloth-bnb-4bit
  - microsoft/Phi-3-mini-4k-instruct-gguf
  - FINAL-Bench/POCKET-Darwin-180B-GGUF
  - ibm-granite/granite-3b-code-base-2k
  - unsloth/Qwen2.5-14B-Instruct
  - allenai/Olmo-3-32B-Think-SFT
  - TaoLiveAIGC/TLive-Omni-4B
  - unsloth/Llama-3.2-1B-Instruct-GGUF
  - Qwen/QwQ-32B
  - meta-llama/Llama-Guard-3-8B
  - bartowski/DeepSeek-R1-Distill-Qwen-14B-GGUF
  - unsloth/Llama-3.3-70B-Instruct-GGUF
  - bartowski/Qwen2.5-Coder-14B-Instruct-GGUF
  - black-forest-labs/FLUX.1-Krea-dev
  - jayn7/Z-Image-Turbo-GGUF
  - QuantFactory/Meta-Llama-3-8B-Instruct-GGUF
  - Qwen/Qwen3-Next-80B-A3B-Thinking
  - empero-ai/Qwen3.8-4B-Distill
  - peculiar-ragdoll/Cyber-Tiel-Coder-35B-A3B-GGUF
  - mudler/Qwen3.5-35B-A3B-APEX-GGUF
  - Qwen/Qwen2.5-Math-7B
  - TokenRhythm/NeoHorse-1-4B-GGUF
  - second-state/stable-diffusion-v1-5-GGUF
  - meta-llama/Llama-3.1-70B
  - FireRedTeam/FireRed-Image-Edit-1.0
  - unsloth/DeepSeek-R1-Distill-Llama-8B-GGUF
  - bartowski/Qwen_Qwen3-8B-GGUF
  - ggml-org/Qwen3-0.6B-GGUF
  - deepseek-ai/deepseek-coder-1.3b-instruct
  - swiss-ai/Apertus-8B-2509
  - EleutherAI/gpt-neo-1.3B
  - dealignai/Bonsai-2-27B-1bit-CRACK-GGUF
  - Edge0/Audio8-ASR-Infinite
  - QuantStack/Qwen-Image-Edit-GGUF
  - inclusionAI/Ling-3.0-tiny-GGUF
  - neuralcrew/neutrino-instruct
  - nvidia/DeepSeek-V4-Flash-0731-NVFP4
  - useful-quants/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-W4A16
  - MeiGen-AI/PosterOmni_v1
  - shafire/Zero-Gemma4-E4B-OpenZero-GGUF
  - unsloth/FLUX.2-klein-base-9B-GGUF
  - Laxhar/noobai-XL-1.0
  - Laxhar/noobai-XL-Vpred-1.0
  - unsloth/Llama-3.2-3B-Instruct-bnb-4bit
  - Qwen/Qwen1.5-0.5B
  - zai-org/GLM-5.3-BF16
  - ibm-granite/granite-4.2-30b
  - unsloth/Qwen2.5-3B-Instruct-bnb-4bit
  - bartowski/DeepSeek-R1-Distill-Qwen-32B-GGUF
  - nvidia/Llama-3_3-Nemotron-Super-49B-v1
  - xinsir/controlnet-canny-sdxl-1.0
  - Phil2Sat/Qwen-Image-Edit-Rapid-AIO-GGUF
  - skt/A.X-4.0-Light
  - Qwen/Qwen3-235B-A22B-Thinking-2507-FP8
  - Qwen/Qwen3-Coder-480B-A35B-Instruct
  - hugging-quants/Llama-3.2-3B-Instruct-Q4_K_M-GGUF
  - stabilityai/stablelm-3b-4e1t
  - audio-cpp/LiveAvatar-GGUF
  - pipecat-ai/phonellm-alpha-1
  - Qwen/Qwen3Guard-Gen-8B
  - legraphista/glm-4-9b-chat-IMat-GGUF
  - diffusers/controlnet-depth-sdxl-1.0
  - TheBloke/Mistral-7B-Instruct-v0.1-GGUF
  - bartowski/NousResearch_Hermes-4-14B-GGUF
  - prithivMLmods/Qwen-Image-2.1-PE-T2I-GGUF
  - solidrust/Mistral-7B-Instruct-v0.3-AWQ
  - Qwen/Qwen2.5-Coder-3B
  - John6666/hassaku-xl-illustrious-v31-sdxl
  - openai/gpt-oss-safeguard-20b
  - osllmai-community/Llama-3.2-1B
  - unsloth/FLUX.1-dev-GGUF
  - stabilityai/stable-video-diffusion-img2vid
  - fdtn-ai/antares-1b
  - ai21labs/AI21-Jamba-Reasoning-3B
  - allenai/Olmo-3-7B-Instruct-SFT
  - Venastine-Research/Xing4.0-29B-A4B-GGUF
  - unsloth/LFM2.5-1.2B-Instruct-GGUF
  - tiiuae/falcon-7b-instruct
  - allenai/OLMo-2-1124-7B-Instruct
  - Lykon/dreamshaper-xl-lightning
  - OBLITERATUS/gemma-4-E4B-it-OBLITERATED
  - lmstudio-community/Qwen3-14B-GGUF
  - Agnuxo/CAJAL-4B
  - outsourc-e/Qwen3.8-27B-Unleashed-GGUF
  - PixArt-alpha/PixArt-Sigma-XL-2-1024-MS
  - sakamakismile/Qwen3.6-27B-Text-NVFP4-MTP
  - upstage/SOLAR-10.7B-Instruct-v1.0
  - unsloth/llama-3-8b-Instruct
  - bg-digitalservices/Gemma-4-26B-A4B-it-NVFP4
  - bartowski/L3-8B-Stheno-v3.2-GGUF
  - google/gemma-2-9b
  - zai-org/glm-edge-4b-chat-gguf
  - Manojb/stable-diffusion-2-1-base
  - meta-llama/Llama-Guard-3-1B
  - nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Base-BF16
  - realrebelai/LTX-2.5_GGUFs
  - magespace/Wan2.2-I2V-A14B-Lightning-Diffusers
  - TinyLlama/TinyLlama-1.1B-step-50K-105b
  - SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16
  - stelterlab/Qwen3-Coder-30B-A3B-Instruct-AWQ
  - GritLM/GritLM-7B
  - diffusers/controlnet-canny-sdxl-1.0
  - city96/Qwen-Image-gguf
  - QuantStack/LTX-2.3-GGUF
  - nvidia/GLM-5.3-NVFP4
  - autotrust/GLM-5.3-Flash-GGUF-DGX-Spark
  - chfm/Qwen-Image-2.1-GGUF
  - IFM/K2-Horizon-MoVA-36B-A4B-GGUF
  - Qwen/Qwen3.8-2.4T-A95B
  - nvidia/llama-nemotron-embed-vl-1b-v2
  - fishaudio/s2-pro
  - sudoingx/Ternary-Bonsai-2-27B-PTQ1_0-MTP-GGUF
  - joeygambino/MiniMax-H3-curve-GGUF
  - unsloth/FLUX.1-schnell-GGUF
  - rectangleworm/ideogram-4-gguf
  - XHToken/Spark-X2.5-4B
  - ai4bharat/IndicF5
  - perplexity-ai/pplx-embed-v1-4b
  - city96/Wan2.1-I2V-14B-720P-gguf
  - realrebelai/SCAIL-2_GGUF
  - bartowski/Cloudflare_clef-GGUF
  - jayn7/WAN2.2-I2V_A14B-DISTILL-LIGHTX2V-4STEP-GGUF
  - IFM/K2-Horizon-7B
  - DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF
  - QuantStack/Wan2.2-S2V-14B-GGUF
  - bartowski/TheDrummer_Artemis-31B-v1.2-GGUF
  - Hob-forge/Kolibri-1-GGUF
  - ggml-org/Clef-Flash-GGUF
  - bartowski/Cloudflare_clef-flash-GGUF
  - city96/Wan2.1-FLF2V-14B-720P-gguf
  - nvidia/Cosmos3-Super-Image2Video
  - LiquidAI/LFM2.5-8B-A1B
  - city96/Wan2.1-T2V-14B-gguf
  - realrebelai/Wan-Animate-2_GGUFs
  - >-
    deadbydawn101/RavenX-CyberAgent-Qwen3.6-35B-A3B-Opus-4.7-OpenMythos-Pentester-BugHunter-RATH-GGUF
  - jialinyyzz/humanizer
  - IFM/K2-Horizon-375B-A23B
  - hum-ma/Wan2.2-TI2V-5B-Turbo-GGUF
  - Wan-AI/Wan2.1-VACE-1.3B
  - IFM/K2-Horizon-3.7B
  - unsloth/LTX-2-GGUF
  - vantagewithai/SCAIL-2-GGUF-ComfyUI
  - bartowski/FrogNano-4B-2609-GGUF
  - google/embeddinggemma-2
  - netease-youdao/Confucius4-R2T2
  - ggml-org/Qwen3.8-Flash-Next-GGUF
  - Abiray/MiniMax-H3-Singularity-GGUF
  - Wan-AI/Wan2.1-I2V-14B-720P
  - m-a-p/MERT-v2-FullSong
  - XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
  - robbyant/lingbot-world-fast
  - calcuis/wan-1.3b-gguf
  - bytkim/Qwen3.8-27B-pi-GGUF
  - IFM/K2-Horizon-MoVA-36B-A4B
  - tencent/WeMM-Embedding-2B
  - Cloudflare/clef-flash
  - ukisai/Swift-Bonsai-2-GGUF
  - paradigma-inc/limite-1b-violetto
  - zai-org/CogVideoX-2b
  - zai-org/CogVideoX-5b
  - bullerwins/Wan2.2-T2V-A14B-GGUF
  - Rabinovich/LongLive-2.0-5B-Diffusers
  - inclusionAI/Ling-3.0-tiny
  - TaichuAI/ZDTaichu5.0-9B
  - ggml-org/Clef-GGUF
  - BreezeBlue/Breeze-TTS-2
  - inclusionAI/Ling-3.0-flash
  - Altworld/Hemmingway-1
  - Cloudflare/clef
  - moondream/parakeet-ultra
  - Kwaipilot/KAT-Coder-V2.5-Dev
  - nex-agi/Nex-N2.5-mini
  - nvidia/NVIDIA-Nemotron-3-Nano-4B-GGUF
  - CohereLabs/North-Mini-Code-1.0
  - jinaai/jina-embeddings-v5-omni-nano
  - LiquidAI/d1-omni-600M
  - terrorswift/REDCELL-26B-A4B-OSINT-Cyber-APEX-GGUF
  - dphn/Dolphin-Mistral-24B-Venice-Edition
  - DavidAU/LFM2.5-8B-A1B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF
  - peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF
  - autotrust/GLM-5.3-GGUF-DGX-Spark
  - empero-ai/Qwen3.8-35B-A3B-Distill
  - Aleph-Alpha/Kolibri-1
  - empero-ai/Qwythos-9B-Claude-Mythos-5-1M
  - XiaomiMiMo/MiMo-V2.6-Flash-MOPD
  - elix3r/gemma4-12b-with-proj-ltx-2.5-GGUF
  - albucino/Qwen3.8-Flash-Next-W4A16-FP8PLE
  - autotrust/GLM5.3-Flash-E224-DGX-Spark
  - LiquidAI/d1-3B
  - SkyIsNotGreen/Scion-35B-A3B
  - microsoft/harrier-oss-v1-27b
  - ukisai/Swift-1.5-Qwen3.8-27b
  - orcarouter/OrcaSAQ-2-27B
  - yandex/AliceAI-Foundation-80B-A3B-Base
  - ggml-org/OpenJev-GGUF
  - DeepHat/DeepHat-V1-7B
  - pfeifferj/Ornith-1.5-35B-A3B-GSQ-RCO-GGUF
  - SeerRay-Lab/Xiaomi-OCR-0
  - Qwen/Qwen3Guard-Stream-0.6B
  - KittenML/kitten-tts-2
  - llm-jp/llm-jp-4.1-8b-thinking-gguf
  - alesha-pro/Qwen3.8-27B-S-mirai-GGUF
  - webAI-Official/TwIL-LM3-Pro
  - XiaomiMiMo/MiMo-V2.6-Pro-MOPD
  - apodex/Apodex-1.1-mini
  - tencent/HunyuanImage-3.0
  - logic65/Whittle-Qwen-3.8-35B-A3B
  - rmonsurate/Victoria
  - Lythri/Lythri-4B-A2B
  - Mia-AiLab/GLM-5.3-Flash-EXL3-4bpw-TensorFold
  - stabilityai/stable-video-diffusion-img2vid-xt-1-1
  - emperorofrome/Gmcoder
  - InternScience/Agents-A1
  - aj9o9/Qwen3.8-27B-Escha-W2-GGUF
  - OrionLLM/OxCoder-9B
  - tsinghua-sigs-robot-lab/VeriLoop-E2
  - jialinyyzz/humanizer-GGUF
  - ArliAI/GLM-4.6-Derestricted-v3
  - MaziyarPanahi/Llama-3-Groq-8B-Tool-Use-GGUF
  - LiquidAI/d1-3B-GGUF
  - AikidoSec/altar-1
  - Mia-AiLab/GLM-5.3-Flash-EXL3-4bpw-TensorFold-Ablit
  - bottlecapai/ThinkingCap-Qwen3.8-27B
  - diffbot/DeepSeek-V4.1-Flash-EXL3-3bpw-2x-RTX-PRO-6000
  - bytkim/Qwen3.8-27B-pi
  - hiwaifu-research/WaifuGemma4-26b-a4b-v1
  - FINAL-Bench/Darwin-27B-RSI-GGUF
  - Blackfrost-AI/CYBER-FROST-3.8-BF16
  - IFM/K2-Type-0.9B
  - BAAI/AREX-2
  - RedHatAI/Qwen3.8-Flash-Next-NVFP4
  - unsloth/embeddinggemma-2
  - IQuestLab/IQuest-Q1
  - topk-io/topk-embed-v1-xsmall
  - vcruz305/CYBER-FROST-3.8-EXL3-SAGE-3.87bpw
  - Aleph-Alpha/Kolibri-1-BF16
  - Eliasfpv28/Kolibri-1-Q3_K_S-GGUF
  - CaryPalmer/Ternary-Bonsai-2-27B-262k-GGUF
  - llm-jp/llm-jp-4.1-33b-thinking
  - BuzzASR/persian
  - cantina-security/apex-flash-1
  - Blackfrost-Research/GLM-5.3-F.U-AnthraClaud-Edition-BF16
  - troed/Qwen3.8-27B-ASCII-Condensed
  - Abiray/Qwen-Image-2.1-viggle-turbo-v0.3-6step-GGUF
  - >-
    nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2
  - perplexity-ai/pplx-embed-v2-context-9b-preview
  - FINAL-Bench/Darwin-27B-RSI
  - lmstudio-community/Llama-3-Groq-70B-Tool-Use-GGUF
  - mehdi-hf/nemotron-asr-streaming-farsi
  - Gryphe/Gemma-4-26B-A4B-StyleTune-V2
  - EschaLabs/Qwen3.8-27B-Escha-W2
  - nvidia/NVIDIA-NemotronLabs-AI-for-Media-Sports-Tennis
  - FINAL-Bench/Darwin-180B-RSI
  - Hcompany/Holo4-27B
  - Blackfrost-AI/CYBER-FROST-3.8-NVFP4-V2
  - ISTA-DASLab/Qwen3.8-Flash-Next-P48NVFP4-MoESQ
  - LiquidAI/LFM2-1.2B-RAG
  - vcruz305/RED-SNOW-5.3-FLASH-EXL3-SAGE-4.91bpw
  - RedHatAI/DeepSeek-V4-Flash-0731-NVFP4
  - kova-ai/kova-tts-1
  - PleIAs/baguettotron-600m
  - elyza/ELYZA-Thinking-1.0-llm-jp-4-33b
  - elyza/ELYZA-Thinking-1.0-llm-jp-4-32b-a3b
  - yamz-labs/GLM-5.3-Flash-EXL3-Yamz
  - bitlamas/Qwen3.8-Flash-Next-Q4_K_XL-DN4
  - kyutai/glm-4-voice-of-reason-9b
  - YOON1v/Apex-2
  - perplexity-ai/pplx-embed-v2-late-0.6b
  - cklxx/laya-browser
  - amazon/ALoDLM-8B
  - Ateron/Gemma-4-Dark-Thoughts-V2-31B
  - briaai/fibo-scene-analyzer
  - Blockway/Agens-Volundr-32B-Preview
  - WonseokJayJung/Connect-C1-0.8B-v0.1-GGUF
  - kandinskylab/Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers
  - kandinskylab/Kandinsky-6.0-Pro-distill-5s-Diffusers
  - Remek/basal-1.5-4.5B
  - perplexity-ai/pplx-embed-v2-late-9b
  - kandinskylab/Kandinsky-6.0-VSR-5s-Diffusers
  - Maincode/matilda-jev-v1
  - FINAL-Bench/Darwin-180B-RSI-R3
  - vcruz305/RED-SNOW-5.3-FLASH-EXL3-SAGE-2.49bpw
  - h2oai/h2o-lightning-4b
  - Blackfrost-AI/RED-SNOW-5.3-FLASH-BF16
  - kandinskylab/Kandinsky-6.0-Lite-distill-5s-Diffusers
  - kandinskylab/Kandinsky-6.0-Pro-5s-Diffusers
  - Ateron/Gemma-4-Writers-31B-V2
  - audreyt/Kolibri-1-NVFP4-W4A16
  - lighteternal/biodecision-v2-4b
  - malteos/most-embed-de
  - kandinskylab/Kandinsky-6.0-Lite-5s-Diffusers
  - ConwayResearch/Underdog-Saluki-27B-1.0
  - PleIAs/Baguettotron-MoE
  - JetBrains/Mellum2.1-12B-A2.5B-Thinking
  - kandinskylab/Kandinsky-6.0-Pro-pretrain-5s-Diffusers
  - kandinskylab/Kandinsky-6.0-Lite-pretrain-5s-Diffusers
  - aina-tech/Anima-Lightning
  - amazon/ALoDLM-1.7B
  - facebook/meta-encoder
  - speridlabs/iris-3b
  - Modularcomputing/Native-Bird
  - lightonai/LightOnOCR-3-0.8B
  - lightonai/LightOnOCR-3-4B
  - ApolloRaines/LTX-2.5-22b-OmniGen-v12
  - JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF
  - Kujira/Underdog-Saluki-27B-1.0-MTP-GGUF
  - DavidAU/Qwen3.8-27B-Turbo-Brilliance-Power-35X-Reasoning-Instruct-modes-GGUF
  - inclusionAI/Ming-Image-0.1-Design
  - Aratako/Irodori-TTS-v4-Large

VRAM & GPU Cost Calculator

Pick any Hugging Face model and see:

  • how much GPU memory it needs to run, fine-tune with LoRA, fully fine-tune, generate, transcribe or embed, at its published precision or at 16, 8 or 4 bits;
  • the cheapest GPU setups that hold it (one H100, two A100s, four RTX 4090s...), each priced live across the GPU clouds FastGPU tracks, with a link to every offer.

It reads each model's own metadata from the Hub (the parameter count in its safetensors or GGUF header, its quantization config) and asks FastGPU's free match API for the sizing and the live prices. Nothing to install, no sign-in. Type a size such as 70B for a model that is not on the Hub.

How the memory is sized

The same way as on fastgpu.co, by FastGPU's matcher:

  • Run / serve: the weights (parameters × bytes per parameter: 2 at 16-bit, 1 at 8-bit, 0.5 at 4-bit), plus the KV cache for one request at a 4K-token context and runtime headroom.
  • Fine-tune with LoRA: the frozen weights plus the adapters, their optimizer state and activations.
  • Full fine-tune: about 16 bytes per parameter (weights, gradients and Adam states) plus activations.
  • Image, video, speech and embedding models FastGPU knows (FLUX, SDXL, Wan, Whisper and more) use a figure for the whole pipeline, text encoders and VAE included.

A quantized repo (AWQ, GPTQ, FP8, NVFP4, bitsandbytes, GGUF) is sized at the bits it is stored in unless you pick another precision. The full method: fastgpu.co/methodology.

Where the prices come from

FastGPU collects the GPU rental prices providers publish (marketplaces, neoclouds and hyperscalers) and refreshes them through the day. Each price here is the provider's own per-GPU rate times the number of GPUs in the setup. "Lowest price" (the default) sorts by price alone. "Best match" orders by FastGPU's score, led by price and weighing reliability and availability; providers that pay FastGPU a referral fee are marked ★ and can only win a near-tie there, within a few percent of the cheapest.

The price history is open: fastgpu/cloud-gpu-prices (CC BY 4.0, updated daily).

Use it from code

curl "https://fastgpu.co/api/v1/match?model=Qwen/Qwen3-32B&params_b=32.8&task=inference"

No key needed. Docs: fastgpu.co/docs.