gpt2-nano

This is a 44M-parameter decoder-only transformer implemented in PyTorch. The published checkpoint predates the mixed-precision extension.

Measured extension

The extension assigns BF16, INT8, or packed INT4 KV-cache storage to each of the model's 96 attention heads. A greedy allocator uses isolated single-head output-KL measurements as a calibration proxy.

On an RTX A5000, the selected policy uses 70 INT4 heads, 19 INT8 heads, and 7 BF16 heads. It stores 579,840 bytes for a batch-one, 64-token cache, compared with 1,572,864 bytes for BF16. Held-out teacher-forced output KL is 0.001528, and next-token agreement is 97.85%.

The fixed cache-noise adaptation does not establish a robustness gain.

Intended use

The model and code are intended for education and controlled cache experiments. They are not intended for production text generation or factual question answering.

Limits

  • Agreement is measured on shared teacher-forced prefixes, not generated text.
  • The additive calibration score is not a bound on joint held-out KL.
  • Persistent cache bytes exclude model weights and temporary dequantized tensors.
  • The implementation dequantizes before attention and does not provide a speedup.
  • The study uses one checkpoint, one held-out window, and one GPU.

The measured evidence is in artifacts/kv-cache-publication/results.json.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Articles mentioning kotlarmilos/gpt2-nano