How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf SebastianAldrin/Qwen3.5-9B-minecraft-distill-v1-GGUF:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf SebastianAldrin/Qwen3.5-9B-minecraft-distill-v1-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf SebastianAldrin/Qwen3.5-9B-minecraft-distill-v1-GGUF:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf SebastianAldrin/Qwen3.5-9B-minecraft-distill-v1-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf SebastianAldrin/Qwen3.5-9B-minecraft-distill-v1-GGUF:Q4_K_M
# Run inference directly in the terminal:
./llama-cli -hf SebastianAldrin/Qwen3.5-9B-minecraft-distill-v1-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf SebastianAldrin/Qwen3.5-9B-minecraft-distill-v1-GGUF:Q4_K_M
# Run inference directly in the terminal:
./build/bin/llama-cli -hf SebastianAldrin/Qwen3.5-9B-minecraft-distill-v1-GGUF:Q4_K_M
Use Docker
docker model run hf.co/SebastianAldrin/Qwen3.5-9B-minecraft-distill-v1-GGUF:Q4_K_M
Quick Links

Qwen3.5-9B — Minecraft Agent Distill v1 (GGUF)

A LoRA fine-tune of Qwen/Qwen3.5-9B that distils Claude Sonnet's per-tick decisions in the Agent Society Minecraft sandbox — so the agent's fast decision tier runs locally and offline instead of calling a frontier model every tick.

Base Qwen/Qwen3.5-9B
Method LoRA SFT — rank 16, alpha 32, dropout 0.05, all-linear, 3 epochs
Teacher Claude Sonnet
Data SebastianAldrin/agent-society-distill-v1 — 1,355 examples
Code Agent Society
Format GGUF, Q4_K_M (~5.4 GB) — runs in llama.cpp / Ollama

Given the agent's situation as a prompt (felt needs, a local block-map, bearings, the current plan step, recent memory, who else is nearby), it returns one in-character decision as JSON: {"thought": "...", "action": "...", "args": {...}}. It is the fast per-tick tier of a three-tier agent mind; planning and reflection stay on a stronger model.

Run

llama-server -m Qwen3.5-9B-minecraft-distill-v1-Q4_K_M.gguf -c 8192 --jinja

Send the decide prompt with enable_thinking: false and a JSON-schema response format (the model is trained to answer with the decision JSON only). Full system prompt and prompt format live in the Agent Society repository.

On CUDA, add --flash-attn off — llama.cpp's auto default silently corrupts this hybrid-SSM architecture: the JSON shape survives but the words inside turn to noise.

Limits

  • Trained on a small, mostly-social set (1,355 examples), so it talks a lot and is weak at rare actions.
  • Needs a recent llama.cpp that knows the qwen3_5 arch — older builds won't load it.
  • The prompt isn't cacheable on this arch, so it re-reads the whole prompt every tick. Slow on a weak GPU or CPU.
  • Measured: on 250 held-out teacher decisions it picks the teacher's action 78.0% of the time; the untuned base scores 56.8%. Method in the evaluation doc.

License

MIT.

Downloads last month
58
GGUF
Model size
9B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for SebastianAldrin/Qwen3.5-9B-minecraft-distill-v1-GGUF

Finetuned
Qwen/Qwen3.5-9B
Quantized
(544)
this model