--- license: apache-2.0 library_name: gguf pipeline_tag: text-generation base_model: Qwen/Qwen3-8B tags: [gguf, conversational, qwen3, quantized, compression, bittensor] ---

Ralph — open model compression on Bittensor SN40

# The same model. Small enough for your phone. ### Ralph crowns — Qwen3-8B compressed in public **Four downloadable GGUFs, from 1.16 GB to 4.61 GB. One admitted Qwen3-8B architecture. Fewer bits. Each file holds a current bit-tier crown in Ralph, the open model-compression tournament on Bittensor SN40.** “Same model” means the same admitted architecture and parameter structure—not identical weights, guaranteed identical behavior, or universal phone compatibility. The files are published together so users have one stable place to download the current crowns, while the source builders and exact scored revisions remain visible. ## Start here: the 2.94 GB sub2 crown `ralph-qwen3-8b-sub2.gguf` is the first file we recommend for a phone test. | | | |---|---| | Exact size | 2,937,263,168 bytes (2.94 GB decimal) | | Code bits / weight | 2.2626 | | Crowned | Round 7 | | Embedded chat-template key | Yes | | SHA-256 | `9c45bf0ac486e793bd761a99ce4b9d965ec12e70f2d420e02af16797c71f6b50` | ```bash hf download RalphLabsAI/ralph-crowns ralph-qwen3-8b-sub2.gguf --local-dir ./ralph-crowns llama-cli -m ./ralph-crowns/ralph-qwen3-8b-sub2.gguf -cnv --jinja -p "Explain one practical use of low-bit quantization." ``` ### Find it in PocketPal 1. Open **Models → Add from Hugging Face**. 2. Search for **RalphLabsAI**, then open **ralph-crowns**. 3. Choose **ralph-qwen3-8b-sub2.gguf**. ## Physical-iPhone device lab **[Join Device Lab #1 — share a success or failure](https://huggingface.co/RalphLabsAI/ralph-crowns/discussions/3).** These are Ralph team-authored tests. Add your own device, app version, settings, and result to the pinned discussion; a failed load is useful evidence too. One configured 2.94 GB Round 7 sub2 crown completed the same normal PocketPal chat on four physical cloud iPhones through the local llama.cpp Metal runtime. | Physical device | OS | Normal PocketPal chat | TTFT | |---|---|---:|---:| | iPhone 15 | iOS 26.6 | 10.58 tok/s | 838 ms | | iPhone 16 | iOS 18.6 | 12.36 tok/s | 439 ms | | iPhone 17 | iOS 26.6 | 16.77 tok/s | 427 ms | | iPhone 17 Pro Max | iOS 26.6 | 18.54 tok/s | 400 ms | [Watch the 30-second matrix](https://huggingface.co/datasets/RalphLabsAI/ralph-device-lab/resolve/7f5f1cfbf7882726a6c55a1d9b2f7c0ff06b3da8/social/ralph-sub2-four-physical-iphones-horizontal.mp4) · [Inspect all four raw receipts](https://huggingface.co/datasets/RalphLabsAI/ralph-device-lab/tree/7f5f1cfbf7882726a6c55a1d9b2f7c0ff06b3da8) · [Verify the public payload](https://huggingface.co/datasets/RalphLabsAI/ralph-device-lab/blob/7f5f1cfbf7882726a6c55a1d9b2f7c0ff06b3da8/SHA256SUMS) Test configuration: PocketPal 1.17.3 (147), context 2,048, fixed prompt, one connected session per device. The expected model digest was verified against Ralph's local mirror, not computed on-device. Every device displayed PocketPal's **Memory tight** warning and nevertheless loaded and completed the prompt. > Physical BrowserStack devices; connected test. BrowserStack re-signing strips > PocketPal's increased-memory and extended-virtual-addressing entitlements, so > these are local-Metal backend/throughput observations—not App Store > memory-fit, offline, universal-compatibility, or broad-quality certification. > The dataset includes full settings, exact responses, screenshots, checksums, > and untouched recordings. ## Current crowns — through Round 7 | File | Tier | Code bits / weight | Exact size | Crown status | Retention at crown | Builder | Chat-template key | |---|---|---:|---:|---|---:|---|---| | `ralph-qwen3-8b-binary.gguf` | binary | 1.0 | 1,158,964,160 B | New in Round 7 | 0.263616 | [ArizonaZZZ](https://huggingface.co/ArizonaZZZ) | No | | `ralph-qwen3-8b-ternary.gguf` | ternary | 1.714 | 2,477,212,736 B | Crowned in Round 5; held in Round 7 | 0.253732 | [tensor-tailor](https://huggingface.co/tensor-tailor) | Yes | | `ralph-qwen3-8b-sub2.gguf` | sub2 | 2.2626 | 2,937,263,168 B | New in Round 7 | 0.282457 | [boweizh1204](https://huggingface.co/boweizh1204) | Yes | | `ralph-qwen3-8b-sub4.gguf` | sub4 | 4.0 | 4,614,304,896 B | New in Round 7 | 0.320923 | [boweizh1204](https://huggingface.co/boweizh1204) | Yes | ### Exact checksums ```text 05c568169fc180067172cbdd38c1a9f5c249a556ffdef1dacd7172bad40fab58 ralph-qwen3-8b-binary.gguf c643cb42575a7a85b7518fd935973202c5686db270ef8b593afa8616a986154a ralph-qwen3-8b-ternary.gguf 9c45bf0ac486e793bd761a99ce4b9d965ec12e70f2d420e02af16797c71f6b50 ralph-qwen3-8b-sub2.gguf b2ec80dce90258bbe9555a67983dea595d91f35cae48ca0d54c714eae8100edf ralph-qwen3-8b-sub4.gguf ``` The same filenames, byte counts, hashes, pinned sources, and chat-template flags are available in [`crowns.json`](./crowns.json). ## What the crown means Ralph retention is a within-protocol compression-fidelity measurement: it measures how closely an artifact reproduces the pinned comparison parent's effect on configured judge models, including the protocol's live language/depth floors. It is **not** a general capability benchmark, an intelligence percentage, or evidence of device speed. Round 7 uses three recorded judges and a fresh public exam. A challenger takes a crown only through the protocol's recorded point-estimate margin or persistence path and must satisfy the per-judge floor rule. A positive paired-bootstrap lower bound can affect the candidate weight vector without changing the crown. - [Round 7 signed record](https://huggingface.co/datasets/RalphLabsAI/ralph-v2-rounds/resolve/main/rounds/round-00000007-f42ae20c6ed2a796.json) - [Round 5 record that crowned the current ternary](https://huggingface.co/datasets/RalphLabsAI/ralph-v2-rounds/resolve/main/rounds/round-00000005-25e6fd0528a93e74.json) - [Ralph protocol and verifier](https://github.com/RalphLabsAI/ralph) ## Provenance The crown publisher fetches the miner's full behavior-affecting package at the **pinned revision** in the round record and recomputes its content hash. It mirrors the GGUF only when that hash matches the record's `model_id`; the GGUF is uploaded without conversion or repacking. For this release, the source owners subsequently added artifact-level Apache-2.0 declarations and notices. Hugging Face LFS metadata was checked again on 2026-09-14: every current source GGUF, scored revision, and mirrored file has the same SHA-256 shown above. The documentation commits did not change the models. Credit for the weights belongs to the builders named above and the miner hotkeys recorded in `crowns.json`. Ralph provides a content-verified mirror and tournament record; it is not the origin of the weights. Admission checks architecture and tensor shapes against `Qwen/Qwen3-8B`. They do not cryptographically prove lineage from the parent's weights; `base_model` identifies the pinned comparison and admission target. ## License Each exact GGUF listed by checksum above is released under the **Apache License 2.0** by its source owner. This repository includes the license text in [`LICENSE`](./LICENSE) and preserves the four source notices in [`NOTICE`](./NOTICE). All four are compressed derivatives of [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B), Copyright 2024 Alibaba Cloud, Apache-2.0. The binary crown also incorporates bit-exact signs from [prism-ml/Bonsai-8B-unpacked](https://huggingface.co/prism-ml/Bonsai-8B-unpacked), Apache-2.0. The source notices describe each modification, data provenance, and MIT-licensed `llama.cpp` tooling. Preserve `LICENSE`, `NOTICE`, source attribution, and modification statements when redistributing these files. The grant is tied here to the exact listed files and hashes. A future crown change requires a new artifact-level license check; this card does not grant terms on behalf of a future builder. ## Running the other crowns For the ternary and sub4 files, replace the filename in the sub2 commands above. Runtime support depends on the exact quantization type, runtime build, backend, available memory, and context settings. Test the exact file on the target hardware before making compatibility or performance claims. The binary file has no embedded chat-template key, so this card does not prescribe a canonical chat command for it. Establish and report a tested external template/runtime combination without repacking or replacing the canonical bytes. ## Help build the device evidence If you test a crown, share both successes and failures in [the pinned Device Lab discussion](https://huggingface.co/RalphLabsAI/ralph-crowns/discussions/3). A useful receipt includes device model, OS and app version, exact GGUF filename, checksum status, context size, runtime settings, load result, prompt-processing speed, generation speed, and the exact failure if it did not run. Remove notifications, account identifiers, serial numbers, and other private device information from screenshots. ## Audit levels - **L0** verifies signatures and recomputes arithmetic, crown decisions, and candidate vectors from published measurements. - **L1** re-derives item selection and checks the pinned pool. - **L2** re-runs a recorded judge over frozen text; complete multi-judge coverage requires every recorded judge. - **L3** loads artifacts and regenerates their steps, binding frozen text to model execution. L2/L3 require the recorded inputs and compatible runtimes. The records and verifier distinguish what can be checked from data alone from what requires inference hardware.