Not bad, but still not Jev.

#3
by Cypherfox - opened

Greetings,
Running a test on human-calibrated data (my own email, souped and specific decision-relevant fields extracted) this doesn't match up to Jev. It does the same thing as a base Qwen 3.5-4B does (my own first test comparison as well): it over marks emails as 'Important', especially for emphatic lying, e.g. "Important news about your account, Roger" which is an AT&T marketing email.

Same input, same labels, lower accuracy on both my triage sets, half the 'Important' precision, worse calibration, and twice the false positives. At first I thought it was because I'd tuned the length of what I was sending to get better accuracy from Jev, so I shortened it for this model, and it got much worse. I'm running some more tests, this time with few-shot examples. (Fwiw, Jev got much worse with few-shot examples, focusing on the examples, not the 'things that are considered important' prompt. But an LLM-based model may behave differently. I'll comment if it changes anything.)

Not dismissing the project, just sharing my own testing results. I'd love to find something that competes with Jev, to run locally.

Good luck!

Intern Large Models org

Thanks for your feedback!

In the email classification scenario, it seems the model follows some superficial patterns from the text itself (like Important news). We actually didn't mix too much classic classification task data into our training pipeline, and I think that's why it underperforms in your test.

This is a good chance for us to improve Intern-Decision. We'll collect relevant dataset to train a new version of the model. Hope future version would help!

Sign up or log in to comment