MTP or ngram head?

#7
by TheREZOR - opened

Standard gemma heads does not give any visible performance improvement. Would you consider releasing MTP or draft head?

Owner
•
This comment has been hidden

Yeah, I'll probably do it later.

Better to have an MTP (EAGLE-ish) head trained on this model's own outputs, I think. A stock Gemma draft doesn't match the fine-tuned distribution, and acceptance stays low because human prose is less predictable than code or math, plus sampling runs at temperature 1.0. Token speed isn't that important for this use case, so it's not a priority yet, but it'd be a good improvement.

Sign up or log in to comment