
Qwen3-0.6B-ahana
Original model name with AhanaAI pack: Qwen3-0.6B-ahana.
0.6B tiny serve face · AhanaAI Company · AhanaAI Company.
Shop this SKU: www.ahanazoo.com
Size
Finished file (model.aarm) |
0.825 GiB |
| Dense F16-class | 1.20 GiB |
| AhanaZip / AhanaAI | ~1.45× smaller |
Tiny ladder face. Live on AhanaBot. Inference from the packed file — no dense restore.
AhanaAI Company
AhanaAI is a Delaware company. We build density: smaller files, same truth, served compressed.
Most of the industry shrinks LLMs by throwing information away — aggressive quant, then hoping the model still talks. We took the other path.
- Lossless delivery. AhanaZip / AhanaAI compression packs the weights you actually ship. Expand matches the packed source. You download the finished file size. You do not unpack a giant GGUF to run it.
- Near-lossless performance at ~0.8-quant class size. When the product goal is a tiny footprint, we size the runtime class near 0.8-quant (about 1-bit class, often 1/16–1/2 of dense F16). Quality stays with the source class we packed — not a reconstructed guess.
- Serve compressed. The
.aarmis a self-contained serve unit. Inference isserve_aarm. Working-set stream only. Never llama. Never full dense restore.
Company: www.AhanaAi.com · Shop: www.ahanazoo.com · Assemble: www.ahanabot.com · Files: ahanazip.com.
Inference from the compressed container
Serving inference from compression is not a prototype. It is how these models run.
AhanaAI wrote a custom codec and custom kernels so the same packed .aarm is both the ship format and the runtime format. The stack compresses and decompresses, and it also executes against packed weights without restoring the full dense tensor.
Concretely:
- Parameters remain inside the compressed container for the life of the process.
- A forward pass streams only the active working set (layer / tile window) into a small device buffer.
- Matmul and attention run on that window; the window is released. The rest of the model stays packed.
- Peak GPU memory and load follow the hot window, not the dense parameter footprint. You do not pay full-size VRAM to read the weights.
There is a runtime inside the compression: decode is fused with the compute path so the GPU never needs a full-size restored copy of the model. That is serve_aarm (--mega layer-stream). Not llama. Not “unpack GGUF then infer.”
This is how a 4B-class helm ships at 0.893 GiB and still chats, and how smaller faces stay on CPU or mixed devices without a dense expand.
What this file is
One AARMv1 container, ready to serve. The product is the finished .aarm byte count.
python -m ahanatool.v5.inference.serve_aarm model.aarm --mega --prefetch 2
Licensing
If you ship devices, clouds, or a model catalog and you want this stack under contract — talk to us about a licensing deal.
We license:
- lossless pack of LLM weights for delivery
- inference from the packed container (no full-size restore)
- near-0.8-quant class runtimes with near-lossless performance
- the
.aarm/serve_aarmpath for your own faces
Shop: www.ahanazoo.com
Contact: jeremiah@AhanaAi.com
Subject: AhanaAI licensing
Company: AhanaAI · Delaware
We do not publish the codec or kernel sources on this card. Deals are written.