FastVLM Qwen2 (BF16, GGUF)

A BF16 Qwen2 GGUF model for higher-quality FastVLM language decoding. This model provides improved numerical fidelity compared to aggressive quantization and is suitable for BF16-capable inference.

Model file

  • fastvlm_qwen2_bf16.gguf

Usage

Use this model for higher-quality multimodal reasoning when BF16 performance is available.

./llama-cli -m fastvlm_qwen2_bf16.gguf -p "Your prompt here"

Base models

Derived from:

Downloads last month
21
GGUF
Model size
0.6B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for musk12/fastvlm-qwen2-bf16

Quantized
(116)
this model