Palette: New Series of Small Model Creative Distills With Only The Best Teachers!

5258078962007613676

Quick Overview: What is Palette-RP-4B-2609-v0.1?

Palette-RP-4B-2609-v0.1 is Qwen 3.5 4B model trained on roleplay data generated by Hy4, why v0.1? Because it's my first try with it, I will be still adjusting training parameters.

It is recommended to keep thinking OFF as all the training data had it turned off, Hy4 scored anomalously high among models with thinking disabled, hence why I choose it to distill, I may give it a try again next month, the data I was able to get is already massive, but I also contemplate going for three times more data next, which will massively hit my wallet.

Still needs a lot of testing in my opinion, I am also experimenting with LFM 2.5 2.6B again, and of course, I will release the next finetune of it with thinking fully removed.

Quants are as always in the repo, safetensors are there too.

Now, on to the advantages this model holds(according to my plan, I dunno if its a success yet lol):

  • Very strong RP and EPR: The teacher model was very unaligned, which is not surprising, I have not found a single refusal across over 25 thousand assisstant completions.
  • Vivid prose: The teacher model is a massive 770B A49B MoE, its RP feels incredibly mature.
  • Improved in-roleplay intelligence: For the reason spoken above, the teacher model is extremely intelligent, which the student also(partially) inherits.
Downloads last month
515
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Emerald7664/Palette-RP-4B-2609-v0.1

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(840)
this model