Mielikki/Erebus-87k
Viewer • Updated • 87.1k • 34 • 12
This is a base model that has had an experimental reward model RL training done over it for a subset of the Erebus dataset (creative writing).
Reward function files can be found here: verifiers
This model was trained using my chunked pref reward model baseline: pretrain-rm-baseline-7b