This page presents the average per-position byte compression rate, demonstrating the long-context capabilities of the models. Calculated at the byte level, these metrics allow for direct comparison across different tokenizers. Benchmarks utilize the [UncheatableEval-Long dataset series](https://huggingface.co/collections/Jellyfish042/uncheatableeval), with sequence lengths extending up to 32k characters.