How good this model is at agentic coding?

#1
by DedeProGames - opened

So in teory this model is better than Opus 5.5 huh?

idk if we should trust the SeaWolf 🌊🐺

Yeah
Furthermore, I’ve already benchmarked Darwin-28B-Opus; it claims to score 89 on GPQA, but actually scores less than 70.

Well, at the base is Qwen3.8 Flash Next, so is not a stupid model. But is can compete the cloud models, i don't think so, even if i would like that.
Is the same as DavidAU with Qwen3.8-27B, they try make it better then it's is, and we end with some ridiculous affirmations results in benchmarks.
In the end, a small model will remain a small model. Is like comparing a new born baby with an adult. :D

Well, at the base is Qwen3.8 Flash Next, so is not a stupid model. But is can compete the cloud models, i don't think so, even if i would like that.
Is the same as DavidAU with Qwen3.8-27B, they try make it better then it's is, and we end with some ridiculous affirmations results in benchmarks.
In the end, a small model will remain a small model. Is like comparing a new born baby with an adult. :D

The problem is that the results don't actually outperform others, yet they claim it's the best model in the world—check this out => https://huggingface.co/blog/DedeProGames/when-benchmark-numbers-become-marketing

Well, at the base is Qwen3.8 Flash Next, so is not a stupid model. But is can compete the cloud models, i don't think so, even if i would like that.
Is the same as DavidAU with Qwen3.8-27B, they try make it better then it's is, and we end with some ridiculous affirmations results in benchmarks.
In the end, a small model will remain a small model. Is like comparing a new born baby with an adult. :D

The problem is that the results don't actually outperform others, yet they claim it's the best model in the world—check this out => https://huggingface.co/blog/DedeProGames/when-benchmark-numbers-become-marketing

Yeah, i know. That's why i have stopped since long ago running any of so said these "super intelligent" models. And most of the times i do myself the quants from the base model.
I rather stay with a "stupid model then loose my time with these so called "claims". :))

FINAL_Bench org

So in teory this model is better than Opus 5.5 huh?

Ha, we'd love to say yes, but we never claimed that. 🙂 The card doesn't compare against closed models at all. We report what we measured, under the exact conditions written in the model-index, and nothing more.

FINAL_Bench org

idk if we should trust the SeaWolf 🌊🐺

Totally fair. Don't trust the wolf, trust the script. 🐺 Every number on the card comes with its protocol: sample count, voting rule, thinking budget. Set up the same environment and you should get the same result. That's exactly why we write it all down.

FINAL_Bench org

Yeah
Furthermore, I’ve already benchmarked Darwin-28B-Opus; it claims to score 89 on GPQA, but actually scores less than 70.

Thanks for actually running it. That's the best kind of feedback. Your number matches ours: under a single-sample, fixed-budget protocol, Darwin-27B-Opus scores 72.85 (it's on the Darwin-27B-RSI card). The higher figure comes from a different protocol (multiple samples, longer thinking budget). Same model, different ruler. 📏 We'll put both protocols side by side on the card so nobody has to guess which ruler was used. If you share your setup, we're happy to compare line by line.

FINAL_Bench org

Well, at the base is Qwen3.8 Flash Next, so is not a stupid model. But is can compete the cloud models, i don't think so, even if i would like that.
Is the same as DavidAU with Qwen3.8-27B, they try make it better then it's is, and we end with some ridiculous affirmations results in benchmarks.
In the end, a small model will remain a small model. Is like comparing a new born baby with an adult. :D

Agreed on the base: Qwen3.8 Flash-Next is a strong starting point, and we say so on the card. As for "a small model stays small": a 180B MoE with 512 experts is quite a large baby. 👶 But we share the spirit: benchmarks should do the talking, not adjectives. That's why we publish the protocol with every number.

FINAL_Bench org

Well, at the base is Qwen3.8 Flash Next, so is not a stupid model. But is can compete the cloud models, i don't think so, even if i would like that.
Is the same as DavidAU with Qwen3.8-27B, they try make it better then it's is, and we end with some ridiculous affirmations results in benchmarks.
In the end, a small model will remain a small model. Is like comparing a new born baby with an adult. :D

The problem is that the results don't actually outperform others, yet they claim it's the best model in the world—check this out => https://huggingface.co/blog/DedeProGames/when-benchmark-numbers-become-marketing

Good article. Separating "the number is real" from "the conclusion is supported" is exactly the right line to draw. We try to stay on the right side of it: every score comes with its protocol, and our Decision Index run is fully public with engine code, so anyone can re-score it instead of taking our word for it. If you ever audit our cards, we'll gladly help with setups. We'd rather be checked than believed.

FINAL_Bench org

Well, at the base is Qwen3.8 Flash Next, so is not a stupid model. But is can compete the cloud models, i don't think so, even if i would like that.
Is the same as DavidAU with Qwen3.8-27B, they try make it better then it's is, and we end with some ridiculous affirmations results in benchmarks.
In the end, a small model will remain a small model. Is like comparing a new born baby with an adult. :D

The problem is that the results don't actually outperform others, yet they claim it's the best model in the world—check this out => https://huggingface.co/blog/DedeProGames/when-benchmark-numbers-become-marketing

Yeah, i know. That's why i have stopped since long ago running any of so said these "super intelligent" models. And most of the times i do myself the quants from the base model.
I rather stay with a "stupid model then loose my time with these so called "claims". :))

Honestly, that's a healthy habit. Quantizing from the base yourself is the best way to know what you're running. 🙂 If you ever want to give this one a spin, the protocols are on the card, so you can check our claims on your own hardware before believing any of them.

•
This comment has been hidden (marked as Abuse)
FINAL_Bench org

Let me correct the record.

Opus 5.5: We have never claimed this model beats Opus 5.5 or any other closed model. What we report is our rank on five Hugging Face official leaderboards, and nothing beyond that.
Protocol: The protocol behind every score (number of samples, majority vote or not, thinking budget) has been public on the model card and in the model-index from day one. Reporting majority vote (maj@k) over multiple samples is a standard practice widely used by major labs and leaderboards. A single-sample score differing from a maj@k score is not fabrication. It is a different protocol.
"Hardcoded": This claim has no basis. The full weights are public. If you reproduce our setup exactly as written on the card and get a different result, please share your config and outputs, and we will gladly review them.

@SeaWolf-AI "We never claimed it beats Opus 5.5." Fine. But your card says "#1 on three Hugging Face official leaderboards," calls this the "flagship," and includes zero agentic or coding evals. You let the implication do the work, then hide behind the literal wording when someone calls it out.

Your headline number is 94.44 on GPQA Diamond. That's majority vote over up to 16 samples, with a 131K thinking budget, on a 198-question set. 94.44% is 187/198, so one question is worth 0.5 points, and there are no confidence intervals and no seeds. Putting that next to other models' single-sample scores and calling it "#1" isn't "a different ruler." It's an unfair comparison presented as a ranking. And those leaderboards are self-reported. Your own card says "all numbers are self-measured."

Where is the parent, Qwen3.8-Flash-Next, on GPQA under the same maj@16 protocol? It's not on the card. Without it, nobody can tell whether 94.44 comes from your "RSI" or just from the base model plus voting. The only RSI evidence you show is MMLU-Pro 88.04 → 88.12. That's +0.08, which is noise, with no ablation. "Self-improving" in the headline isn't backed by that.

About the GPQA gap: you confirmed the single-sample score is 72.85, which is in the ballpark of what I measured. Your own family table lists 27B-Opus at 86.9, so the headline number is 14+ points above what the model does under a plain protocol.

ZTC is in the headline badges, yet the card says its readout "is being fitted and will ship." You're advertising a feature that doesn't exist for this model yet.

"Anyone can reproduce it": open weights don't make the scores reproducible. Where are the eval harness, the per-question outputs and the logs for this model? The "Decision Index run" you mention is a different benchmark and isn't linked on the card. Publish the outputs and the harness, or stop saying "measured, not claimed."

And please stop replying with AI-generated text. Answer the questions above yourself.

FINAL_Bench org

Answering your points directly.

On protocol: how many samples to draw and whether to vote is the publisher's call. Plenty of labs report majority-vote numbers without ever saying so. We state samples, voting rule and thinking budget next to every score. Hiding the protocol would be the problem. Disclosing it isn't.

On RSI: Darwin-180B-RSI scores above the published numbers of its parent, Qwen3.8-Flash-Next, on several of these benchmarks. We changed about 0.02% of the parameters and used no human-written labels to get there.

On ZTC: the probe for this model will be added to the repo.

•
This comment has been hidden (marked as Abuse)
FINAL_Bench org

A note on this thread.

We welcome criticism of our models, our numbers and our methods. That includes questions about evaluation protocols, reproducibility and comparisons. If you disagree, tell us why, and we'll answer on the substance.

What we won't host is abuse. Insults, name-calling and repeated hostile comments aimed at people rather than the work don't help anyone reading this page, so accounts that post them will be blocked.

Thanks to everyone who keeps the discussion technical.

@SeaWolf-AI That's a statement about tone, not an answer. Three questions are still unanswered, and the whole thread is here for anyone to read:

  1. What does Qwen3.8-Flash-Next score on GPQA Diamond under your exact protocol (maj@16, 131K thinking budget)? Without it, nobody can tell whether 94.44 comes from "RSI" or from the base model plus voting.
  2. What is Darwin-180B-RSI's GPQA Diamond score single-sample, with seed count and confidence interval?
  3. Where are the eval harness and per-question outputs for this model?

I don't need to prove anything is hardcoded. When a publisher claims "#1" and "measured, not claimed," the burden is on the publisher to release what's needed to re-score it. Right now that's missing.

Also still open: was the "89" from Darwin-27B-Opus or 28B-Opus, and who wrote the "verifiable answer keys" if there were "no human-written labels"?

You said you'd answer on the substance. These are the substance questions.

FINAL_Bench org

Closing this thread. Thanks to everyone who kept the discussion technical.

SeaWolf-AI changed discussion status to closed

Sign up or log in to comment