CapSel-2 routing artifact

This card accompanies the article "Reducing dropped tokens in a small MoE model" and the measured artifact in artifacts/publication-routing-study.json.

Result

CapSel-2 uses one shared per-expert capacity. It admits primary routes first, tries to rescue rejected primaries through their second choice, then uses remaining capacity for uncertain tokens that already have one accepted route.

Across three seeds on an NVIDIA RTX A5000, mean zero-expert token drop is 8.09% for CapSel-2 and 13.84% for fixed top-2 at the same capacity factor. CapSel-2 requests 15.4% fewer slots and rescues 58.27% of rejected primaries.

Intended use

Use this repository to inspect or reproduce the measured routing experiment. It does not provide a selected pretrained checkpoint.

Training data

The study uses the verified FineWeb-Edu token shard recorded in artifacts/manifest.json. Each condition trains for 500 steps and uses the same initialization and batches within each seed.

Limitations

The held-out cross-entropy difference is 0.0089 in favor of CapSel-2 on the three-seed mean, but it is smaller than the seed spread and changes sign for one seed. The study does not establish a quality winner. Requested slots are not latency measurements. The uncertainty threshold is calibrated once at initialization and stays fixed.

The implementation is available at https://github.com/kotlarmilos/gpt2-nano-moe.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Article mentioning kotlarmilos/gpt2-nano-moe