Correct GoLLeM-v5 rows (64M, 32M, 16M)
#81
by Maggio33 - opened
This corrects our own three rows. We found an error in our earlier evaluation script, so some numbers we submitted were too high.
GoLLeM-v5 64M: replaced with the final 64M model and its measurement.
- Checkpoint:
final_64m_14x576/ckpt_760k.ptin SlayerLab/gollem-v5-ckpts, revision3fa6f3ba(LFS sha256d144c26f…), 760,000 steps / 24.9B tokens. - BLiMP 76.16, ARC-Easy 48.19, WikiText-2 byte perplexity 2.3474. These are our measurements with the board's
benchmark.pyprotocol (BLiMP 67,000 pairs, ARC-Easy test 2,376 questions, WikiText-2 test in 256-token windows), not an official score. - Parameters: 62.9M. The input and output embeddings are tied (one shared matrix), so it is counted once. A plain sum over the state_dict counts it twice (69.9M).
- The previous row (BLiMP 77.84, wiki 2.016, from PR #78) came from our faulty script. Compactbot's re-measurement in #78 was right on ARC-Easy (47.94) and WikiText-2 (byte ppl 2.370). Its BLiMP (75.99) used the 73,000 pairs of our old public script; on the board's 67,000 pairs that checkpoint scores 75.83.
GoLLeM-v5 32M / 16M: BLiMP only. The old values (73.77 / 70.53) were computed on 73,000 pairs. With the board's 67,000 pairs they are 73.48 / 70.08. ARC-Easy and WikiText-2 are unchanged.
We would be grateful if Compactbot could re-benchmark the 64M checkpoint above. Evaluation details and all checkpoints are on the model card: https://huggingface.co/SlayerLab/gollem-v5-ckpts
CompactAI changed pull request status to merged