Instructions to use ArizonaZZZ/ralph-v2-qwen3-8b-binary-p2A2-step500-05c56816 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ArizonaZZZ/ralph-v2-qwen3-8b-binary-p2A2-step500-05c56816 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ArizonaZZZ/ralph-v2-qwen3-8b-binary-p2A2-step500-05c56816 # Run inference directly in the terminal: llama cli -hf ArizonaZZZ/ralph-v2-qwen3-8b-binary-p2A2-step500-05c56816
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ArizonaZZZ/ralph-v2-qwen3-8b-binary-p2A2-step500-05c56816 # Run inference directly in the terminal: llama cli -hf ArizonaZZZ/ralph-v2-qwen3-8b-binary-p2A2-step500-05c56816
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ArizonaZZZ/ralph-v2-qwen3-8b-binary-p2A2-step500-05c56816 # Run inference directly in the terminal: ./llama-cli -hf ArizonaZZZ/ralph-v2-qwen3-8b-binary-p2A2-step500-05c56816
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ArizonaZZZ/ralph-v2-qwen3-8b-binary-p2A2-step500-05c56816 # Run inference directly in the terminal: ./build/bin/llama-cli -hf ArizonaZZZ/ralph-v2-qwen3-8b-binary-p2A2-step500-05c56816
Use Docker
docker model run hf.co/ArizonaZZZ/ralph-v2-qwen3-8b-binary-p2A2-step500-05c56816
- LM Studio
- Jan
- Ollama
How to use ArizonaZZZ/ralph-v2-qwen3-8b-binary-p2A2-step500-05c56816 with Ollama:
ollama run hf.co/ArizonaZZZ/ralph-v2-qwen3-8b-binary-p2A2-step500-05c56816
- Unsloth Desktop
- Docker Model Runner
How to use ArizonaZZZ/ralph-v2-qwen3-8b-binary-p2A2-step500-05c56816 with Docker Model Runner:
docker model run hf.co/ArizonaZZZ/ralph-v2-qwen3-8b-binary-p2A2-step500-05c56816
- Lemonade
How to use ArizonaZZZ/ralph-v2-qwen3-8b-binary-p2A2-step500-05c56816 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ArizonaZZZ/ralph-v2-qwen3-8b-binary-p2A2-step500-05c56816
Run and chat with the model
lemonade run user.ralph-v2-qwen3-8b-binary-p2A2-step500-05c56816-{{QUANT_TAG}}List all available models
lemonade list
- Atomic Chat
ralph-v2-qwen3-8b-binary-p2A2-step500
A 1-bit (Q1_0 GGUF, 1.0 code bits / 1.125 container bits per weight) Qwen3-8B-architecture model, the SN40 (Ralph v2) binary-tier crown from round 7.
License: Apache-2.0 — for the exact file model.gguf in this repository (sha256
05c568169fc180067172cbdd38c1a9f5c249a556ffdef1dacd7172bad40fab58, chain-pinned revision
29c525c5c3166df2d62d5e180d09255199f778fe).
Lineage and upstream terms to preserve
| component | source | license | what we preserve |
|---|---|---|---|
| binary weights (signs) | prism-ml/Bonsai-8B-unpacked (PrismML) | Apache-2.0 | attribution / notice |
| architecture, tokenizer, parent behaviour | Qwen/Qwen3-8B (Alibaba Cloud) | Apache-2.0 (LICENSE) | attribution / notice |
| container / kernels | llama.cpp GGUF Q1_0 | MIT (tooling; not embedded) | — |
Training data used to fit the per-block scales and norms (no dataset text is embedded in the weights): prompts drawn from public datasets, with targets generated by Qwen/Qwen3-8B.
| dataset | license | note |
|---|---|---|
| nvidia/OpenMathReasoning | CC-BY-4.0 | attribution required |
| zake7749/OpenScience-Chinese-Reasoning-SFT | CC-BY-4.0 | attribution required |
| glaiveai/reasoning-v1-20m | Apache-2.0 | |
| sarvamai/samvaad-hi-v1 | Apache-2.0 | |
| ricdomolm/mini-coder-trajs-400k | MIT |
Method (brief)
Bonsai-8B's 1-bit signs are kept bit-exact; only the 128-element group scales and the RMSNorm weights were re-fitted by distillation from Qwen3-8B's own generated continuations (KL(parent‖student) on parent-step tokens, with extra weight on the first tokens of each step and an unlikelihood penalty on chat-template leak tokens), then exported losslessly to Q1_0. Selection used a 720-item private pool scored under three observers. Contains no auto_map, no code, and no dataset text.
Attribution
- Qwen3-8B © Alibaba Cloud, Apache-2.0.
- Bonsai-8B © PrismML, Apache-2.0.
- OpenMathReasoning © NVIDIA, CC-BY-4.0. OpenScience-Chinese-Reasoning-SFT © zake7749, CC-BY-4.0.
- Downloads last month
- 136
We're not able to determine the quantization variants.