AhanaAI / AhanaZip

Qwen3-0.6B-ahana

Original model name with AhanaAI pack: Qwen3-0.6B-ahana.

0.6B tiny serve face · AhanaAI Company · AhanaAI Company.

Shop this SKU: www.ahanazoo.com

Size

Finished file (model.aarm) 0.825 GiB
Dense F16-class 1.20 GiB
AhanaZip / AhanaAI ~1.45× smaller

Tiny ladder face. Live on AhanaBot. Inference from the packed file — no dense restore.

AhanaAI Company

AhanaAI is a Delaware company. We build density: smaller files, same truth, served compressed.

Most of the industry shrinks LLMs by throwing information away — aggressive quant, then hoping the model still talks. We took the other path.

  1. Lossless delivery. AhanaZip / AhanaAI compression packs the weights you actually ship. Expand matches the packed source. You download the finished file size. You do not unpack a giant GGUF to run it.
  2. Near-lossless performance at ~0.8-quant class size. When the product goal is a tiny footprint, we size the runtime class near 0.8-quant (about 1-bit class, often 1/16–1/2 of dense F16). Quality stays with the source class we packed — not a reconstructed guess.
  3. Serve compressed. The .aarm is a self-contained serve unit. Inference is serve_aarm. Working-set stream only. Never llama. Never full dense restore.

Company: www.AhanaAi.com · Shop: www.ahanazoo.com · Assemble: www.ahanabot.com · Files: ahanazip.com.

Inference from the compressed container

Serving inference from compression is not a prototype. It is how these models run.

AhanaAI wrote a custom codec and custom kernels so the same packed .aarm is both the ship format and the runtime format. The stack compresses and decompresses, and it also executes against packed weights without restoring the full dense tensor.

Concretely:

  • Parameters remain inside the compressed container for the life of the process.
  • A forward pass streams only the active working set (layer / tile window) into a small device buffer.
  • Matmul and attention run on that window; the window is released. The rest of the model stays packed.
  • Peak GPU memory and load follow the hot window, not the dense parameter footprint. You do not pay full-size VRAM to read the weights.

There is a runtime inside the compression: decode is fused with the compute path so the GPU never needs a full-size restored copy of the model. That is serve_aarm (--mega layer-stream). Not llama. Not “unpack GGUF then infer.”

This is how a 4B-class helm ships at 0.893 GiB and still chats, and how smaller faces stay on CPU or mixed devices without a dense expand.

What this file is

One AARMv1 container, ready to serve. The product is the finished .aarm byte count.

python -m ahanatool.v5.inference.serve_aarm model.aarm --mega --prefetch 2

Licensing

If you ship devices, clouds, or a model catalog and you want this stack under contract — talk to us about a licensing deal.

We license:

  • lossless pack of LLM weights for delivery
  • inference from the packed container (no full-size restore)
  • near-0.8-quant class runtimes with near-lossless performance
  • the .aarm / serve_aarm path for your own faces

Shop: www.ahanazoo.com
Contact: jeremiah@AhanaAi.com
Subject: AhanaAI licensing
Company: AhanaAI · Delaware

We do not publish the codec or kernel sources on this card. Deals are written.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support