What datasets did you use for calibration?

#5
by qwaen - opened

What datasets did you use for calibration?

I can't share the corpus itself β€” it's built from private traffic, and the imatrix is basically a fingerprint of it. But the corpus was never the secret. The shape was, and that part reproduces from your own data:

~525K tokens, 30% agentic / 30% code / 15% conversation / 12.5% math / 12.5% writing. The ratio is the product, not the provenance.

Three things that actually moved the needle:

  • Put your system prompt and tool definitions in it. In agentic use those tokens are in every single forward pass, and they're completely absent from a wikitext-style imatrix. It's the activation pattern the model always fires.
  • Go bigger than the usual ~41K tokens. At 10x that, mixed, the expert ranking gets much more stable.
  • Dedupe before you weight. Agentic logs repeat the same system prompt thousands of times and the imatrix collapses onto it. And keep ~10% general text as insurance β€” a corpus overfitted to your own traffic quietly costs general capability.

Why it's worth the trouble: same K, same bits, same disk, changing only the corpus took HumanEval from 89.6% to 95.1% and dropped fabrication on a knowledge probe from 58% to 17%. Which experts you keep matters more than how many.

One warning, because it cost me: don't rank two selections by retained routing mass. The metric is circular β€” each selection wins when scored against the corpus that produced it. Mine confidently predicted the better build would lose. Budget for a real benchmark.

Sign up or log in to comment