How to use from
OpenClaw
Start the llama.cpp server
# Install llama.cpp:
brew install llama.cpp
# Start a local OpenAI-compatible server:
llama serve -hf robertzty/Cosmos-Reason2-8B-GGUF:BF16
Configure OpenClaw
# Install OpenClaw:
npm install -g openclaw@latest
# Register the local server and set it as the default model:
openclaw onboard --non-interactive --mode local \
  --auth-choice custom-api-key \
  --custom-base-url http://127.0.0.1:8080/v1 \
  --custom-model-id "robertzty/Cosmos-Reason2-8B-GGUF:BF16" \
  --custom-provider-id llama-cpp \
  --custom-compatibility openai \
  --custom-text-input \
  --accept-risk \
  --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Quick Links

Cosmos-Reason2-8B-GGUF

GGUF conversion of nvidia/Cosmos-Reason2-8B.

Files

  • Cosmos-Reason2-8B-BF16-split-00001-of-00005.gguf to ...-00005-of-00005.gguf: Main model (BF16, split into 5 parts)
  • mmproj-Cosmos-Reason2-8B-BF16.gguf: Vision encoder (multimodal projector)

Download

huggingface-cli download robertzty/Cosmos-Reason2-8B-GGUF --local-dir ./Cosmos-Reason2-8B-GGUF

Usage with llama.cpp

No need to merge split files - llama.cpp loads them automatically:

llama-cli -m ./Cosmos-Reason2-8B-GGUF/Cosmos-Reason2-8B-BF16-split-00001-of-00005.gguf \
  --mmproj ./Cosmos-Reason2-8B-GGUF/mmproj-Cosmos-Reason2-8B-BF16.gguf -cnv

Optional: merge into single file:

llama-gguf-split --merge Cosmos-Reason2-8B-BF16-split-00001-of-00005.gguf Cosmos-Reason2-8B-BF16.gguf

Notes

  • Requires llama.cpp b7480+ for qwen3vl architecture support
  • Ollama and LM Studio may have compatibility issues with qwen3vl models
Downloads last month
7
GGUF
Model size
8B params
Architecture
qwen3vl
Hardware compatibility
Log In to add your hardware

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for robertzty/Cosmos-Reason2-8B-GGUF

Quantized
(12)
this model