benchmarks, comparable to the other ones I've given for your qwen fusion models

#8
by pkp24 - opened

I stopped the benchmarking because I didn't want to dedicate a huge amount more time to it, testing every single thinking mode was a ton of testing.

tldr nothing set for reasoning effort was by far the best compared to anything else.

Reasoning-mode benchmark: LFM2.5-2.6B Turbo-Brilliance NEO-MAX Q8_0

Tested model: DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF, file LFM2.5-2.6B-Q3.8-TBrilliance-NEO-MAX-Q8_0.gguf.

Snapshot taken 2026-09-26, 09:23. Run 1 is complete. Run 2 was still in progress: easy, hard and ultrahard were done for every mode, and abyssal was still running.

Setup

  • What changes between modes: only the reasoning mode. Every mode uses the same weights and the same llama.cpp server layout (4 slots Γ— 128K context, f16 KV cache). The mode is set server-side with --reasoning-effort <mode>, which chooses the system block the chat template injects.
  • Baselines: off is stock LFM2.5 thinking with nothing injected. high is the template's default mode.
  • Sampling: thinking on for every test, with the model card's tester settings: temp 1.0, top-k 64, min-p 0.05, top-p 0.95, repeat penalty 1.0.
  • Limits: no token cap, so a response can run to the full 128K context. The only other stops are a 1800 s timeout per request and a repetition (loop) detector.
  • Test suites: four difficulty tiers (easy, hard, ultrahard, abyssal). The categories are coding (the program is run against tests), logic, JSON, distill, constraint, simulation, cx and flawstep. Planning tests are recorded but not scored.

How to read the scores:

  • Each test is a single sample at temp 1.0. A single suite can swing 10–20 points from one run to the next, so small gaps are not meaningful yet.
  • trunc, loop and blank each score 0 on that test:
    • trunc: the response ran out of context.
    • loop: the repetition detector abandoned it.
    • blank: it finished without an answer.
  • wall min is not a speed comparison. Two modes ran at the same time on two servers sharing one GPU, so a mode's time depends on what the other server was doing.

Run 1 (complete: easy, hard, ultrahard, abyssal)

Scores (%)

mode easy hard ultrahard abyssal mean vs off vs high tok/test trunc loop blank err wall min
off 95.0 75.0 15.8 1.2 46.8 +0.0 +9.1 11,594 0 0 5 0 37
deeptree 91.0 78.3 15.1 0.6 46.3 -0.5 +8.6 12,017 0 0 3 0 38
omni 93.8 64.6 14.8 3.1 44.1 -2.7 +6.4 11,287 0 0 3 0 33
einstein 93.8 65.0 16.7 0.5 44.0 -2.8 +6.4 13,393 0 0 0 0 43
spoon 68.8 65.0 23.8 0.5 39.5 -7.2 +1.9 11,801 0 2 0 0 28
low 83.8 55.4 12.5 4.1 39.0 -7.8 +1.3 12,454 0 0 3 0 39
socrates 70.0 73.8 8.3 3.7 38.9 -7.8 +1.3 12,350 0 0 2 0 44
high 73.8 60.8 12.4 3.5 37.6 -9.1 +0.0 13,835 0 0 0 0 43
hyper 67.5 57.8 16.4 0.6 35.6 -11.2 -2.0 12,438 0 0 1 0 38
logic 68.8 42.1 29.6 0.2 35.1 -11.6 -2.5 11,909 0 0 3 0 38
ultra 73.8 47.9 15.8 0.2 34.4 -12.3 -3.2 14,049 0 0 3 0 52
medium-low 78.8 41.7 13.9 3.2 34.4 -12.4 -3.2 11,652 0 1 0 0 42
medium 68.8 42.5 18.9 3.8 33.5 -13.3 -4.1 12,344 0 0 1 0 37

By category (all four suites pooled, %)

mode coding constraint cx distill flawstep json logic sim
off 57.1 4.2 0.0 25.0 0.0 50.0 62.5 12.5
deeptree 57.9 0.0 0.0 32.5 0.0 50.0 56.2 4.2
omni 62.5 0.0 0.0 29.2 0.0 50.0 43.8 12.5
einstein 46.6 4.2 0.0 27.5 0.0 50.0 56.2 12.5
spoon 24.0 0.0 0.0 27.5 0.0 58.3 50.0 25.0
low 27.3 12.5 0.0 32.5 0.0 50.0 50.0 1.8
socrates 49.2 0.0 0.0 30.0 0.0 41.7 43.8 12.5
high 20.2 3.1 0.0 27.5 0.0 50.0 56.2 12.5
hyper 30.7 8.3 0.0 25.0 0.0 50.0 37.5 12.5
logic 12.2 12.5 0.0 27.5 0.0 58.3 37.5 25.0
ultra 18.6 0.0 0.0 25.0 0.0 58.3 43.8 0.0
medium-low 21.1 4.2 0.0 25.0 0.0 50.0 43.8 12.5
medium 9.2 0.0 0.0 30.0 0.0 50.0 50.0 12.5

Where the modes disagree

Of 72 graded tests, 8 passed in every mode and 24 failed in every mode. 31 split, meaning some mode scored at least 50 points above another on that test.

Number of split tests on which each mode had the top score (ties count for every tied mode): off 20, omni 18, deeptree 17, einstein 17, socrates 15, spoon 14, low 13, high 13, medium-low 11, hyper 11, ultra 10, logic 10, medium 9.

Per-test scores on the 31 split tests (score Γ—100, blank = 0)
test off low medium-low medium high ultra omni deeptree hyper socrates logic einstein spoon
abyssal:ab_chain3 100
abyssal:ab_chain4 100
abyssal:ab_chain5 100 100 100
easy:code_eval 100 100 100 100 100 100
easy:code_lvp 100 100 100 100 88 100 100 100
easy:code_minwin 100 100 100 100 17 100
easy:code_sliding 100 100 100 100 100
easy:code_wordbreak 100 100 100 100 100 100
easy:json_primes 100 100 100 100 100 100 100 100 100 100 100 100
easy:json_tool 100 100 100 100 100 100 100 100 100 100 100 100
easy:logic_div 100 100
easy:logic_race 100 100 100 100 100 100 100 100 100 100 100
hard:h_coin 100 100 100 100 100
hard:h_countsmaller 100 100 100 100 100 100 100
hard:h_domino 100 100 100 100 100 100 100 100 100 100 100 100
hard:h_edit 100 100 100 100 100 100 100 100
hard:h_lis 100 100 100 100 100
hard:h_lockers 100 100 100 100 100 100 100 100 100 100 100
hard:h_modexp 100 100 100 100 100 100 100 100 100 100
hard:h_regex 100 100 38 100 50 100
hard:h_sieve 100 100 100 100 100 100 100 100
hard:h_trap 100 100 100 100 100 100 100
ultrahard:uh_cover 100 100 100 100 100 59 100
ultrahard:uh_cs1 33 100 33 67 100 33
ultrahard:uh_j2 100 100 100 100
ultrahard:uh_mul 100 100 100
ultrahard:uh_parse 39 72 39 50 33 11 11 39
ultrahard:uh_sim15 100 100 100 100 33 100 100
ultrahard:uh_sim250 100
ultrahard:uh_sim60 100 100 100 100 100
ultrahard:uh_window 100 54 8 62 15 100 62

Run 2 (partial: easy, hard and ultrahard done for every mode; abyssal still running)

The mean below covers easy, hard and ultrahard only, so it is not comparable to run 1's four-suite mean.

Scores (%)

mode easy hard ultrahard mean vs off vs high tok/test trunc loop blank err wall min
off 93.8 66.7 28.3 62.9 +0.0 +14.0 6,766 0 0 1 0 15
deeptree 95.0 75.0 12.8 60.9 -2.0 +12.1 9,083 0 0 0 0 19
omni 88.8 70.8 18.4 59.3 -3.6 +10.5 8,614 0 0 2 0 18
einstein 88.8 52.5 16.6 52.6 -10.3 +3.8 8,718 0 0 0 0 16
socrates 88.0 51.0 14.1 51.1 -11.8 +2.2 7,524 0 0 1 0 17
hyper 72.5 66.7 13.4 50.9 -12.0 +2.0 9,289 0 0 0 0 19
low 68.8 61.7 19.5 50.0 -12.9 +1.1 6,887 0 0 0 0 13
logic 73.8 53.5 19.4 48.9 -14.0 +0.1 8,195 0 0 1 0 15
high 75.0 61.3 10.3 48.8 -14.0 +0.0 8,769 0 0 1 0 19
medium-low 67.5 61.7 17.1 48.8 -14.1 -0.1 8,495 0 0 0 0 21
ultra 73.8 53.3 10.4 45.8 -17.1 -3.0 7,646 0 0 0 0 16
medium 61.3 54.6 18.1 44.6 -18.3 -4.2 8,989 0 0 3 0 20
spoon 78.8 35.4 12.3 42.2 -20.7 -6.7 8,320 0 0 0 0 11

By category (easy + hard + ultrahard pooled, %)

mode coding constraint distill json logic sim
off 77.5 8.3 33.3 77.8 66.7 50.0
deeptree 73.8 0.0 33.3 66.7 75.0 25.0
omni 69.5 25.0 33.3 66.7 66.7 25.0
einstein 59.3 6.2 36.7 55.6 66.7 8.3
socrates 64.5 0.0 33.3 55.6 58.3 8.3
hyper 54.8 0.0 33.3 66.7 58.3 0.0
low 15.6 0.0 43.3 77.8 66.7 25.0
logic 25.0 8.3 36.7 77.8 58.3 25.0
high 29.8 0.0 40.0 66.7 66.7 0.0
medium-low 18.5 8.3 43.3 66.7 58.3 50.0
ultra 21.2 0.0 43.3 77.8 50.0 0.0
medium 26.7 16.7 36.7 77.8 41.7 8.3
spoon 17.5 8.3 33.3 55.6 58.3 25.0

Where the modes disagree

Of 47 graded tests, 7 passed in every mode, 9 failed in every mode and 28 split.

Number of split tests on which each mode had the top score: off 22, deeptree 21, omni 20, hyper 14, socrates 14, einstein 14, low 12, high 12, logic 12, medium-low 11, medium 10, ultra 10, spoon 9.

Per-test scores on the 28 split tests (score Γ—100, blank = 0)
test off low medium-low medium high ultra omni deeptree hyper socrates logic einstein spoon
easy:code_eval 100 100 100 100 100 100 86 100
easy:code_lvp 100 100 100 100 100
easy:code_minwin 100 100 100 100
easy:code_sliding 100 100 100 100 100 100 100
easy:code_wordbreak 100 100 100 100 100
easy:logic_div 100 100
easy:logic_fact 100 100 100 100 100 100 100 100 100 100 100
easy:logic_hh 100 100 100 100 100 100 100 100 100 100 100
hard:h_coin 100 100 100 100 100 100
hard:h_countsmaller 100 100 100 100 100 100 100 100
hard:h_domino 100 100 100 100 100 100 100 100 100 100 100
hard:h_edit 100 100 100 100 100
hard:h_graph 100 100 100 100 100 100 100 100 100 100 100
hard:h_lis 100 100 100
hard:h_modexp 100 100 100 100 100 100 100 100 100 100 100 100
hard:h_regex 50 100 100 100 100 100 100 75 75 100
hard:h_sieve 100 100 100 100 100 100 100
hard:h_toolevent 100 100 100 100 100 100 100 100 100 100 100 100
hard:h_trap 100 100 100 100 100 100 100 50
ultrahard:uh_cover 100 100 100 47 100 100 100 24 24
ultrahard:uh_crt 100 100 100 100 100 100 100 100 100 100 100 100
ultrahard:uh_cs1 33 33 67 100 33 33
ultrahard:uh_j2 100 100 100 100 100
ultrahard:uh_mul 100 100 100
ultrahard:uh_parse 6 28 11 50 39 11 39
ultrahard:uh_sim15 100 100 100 33 100 33 100 33 100
ultrahard:uh_sim60 100 100 100
ultrahard:uh_window 77 100 92 77 100 100 62

Observations so far

  • No mode beat stock thinking. off (no injection) led both runs, with deeptree and omni close behind. The template default high sat mid-pack, 9–14 points behind off.
  • Coding drives most of the differences. The coding scores span roughly 9–78% across modes, while distill and JSON barely move.
  • Ultrahard and abyssal are beyond this model. It scores about 8–30% on ultrahard and 0–4% on abyssal in every mode, so those suites say little about which mode is better.

Thank you for detailed testing and notes.
These are very helpful.

A few notes:

  • 2.6B is a limiting factor ; that being said LiquidAI did an exceptional job here however.
  • In run #1, on very hard/abyssal some of the reasoning modes clearly pulled ahead of "off" / "standard" - some by a long shot. ; I don't know if I would "mean" these kinds of numbers.
  • Without "gating" the specialized reasoning modes will attempt to solve problems they were not designed for. (this matches our own internal testing).
  • Gating would select the best mode automatically during testing, rather than "manually" as is the case now; IE you could run one test, for all modes VS 12.
  • The specialized reasoning is use case specific ; this is part explains the low scores. Larger models will disallow the "reasoning mode" automatically based on context ; if it doesn't match -> standard.
  • The reasoning modes were NOT calibrated for this specific model, they are "general", especially the advanced modes. IE min: 9B/27B sized models. (or larger)
  • An odd issue: sometimes the model will override/disagree with the reasoning mode, and cause issues itself.
  • Parameters for testing (ours) differ greatly from LiquidAi's published ones. (temp is especially critical, rep pen too -> especially as LFM suggests 1.1 vs 1 [off] from us)
  • Testing smaller para models is always more difficult // less range (our experience here) // likewise RAISING benchmarks (thru tuning) is also a lot harder too.
  • It appears your tests are primarily coding, coding related, and logic ; the alignment for all modes was "general" - with aligned reasoning modes the scores will likely reflect a change.

Again ; thank for publishing - this will help with refinements to reasoning, reasoning modes and other planned improvements.

Yeah I generally use it for a lot of logic chains, tool usage, and coding so that's where most of the benchmarks are at.

"In run #1, on very hard/abyssal some of the reasoning modes clearly pulled ahead of "off" / "standard" - some by a long shot. ; I don't know if I would "mean" these kinds of numbers."
Due to how low the passing scores were with some of the much harder tests I don't know how well it really would do when run over and over again. sometimes models are on the cusp of failing or passing a test so it may flip between passing/failing multiple times if I run it multiple times. Since I only ran the full suite 1.2 times idk that the barely passing numbers mean too much.

"The specialized reasoning is use case specific ; this is part explains the low scores. Larger models will disallow the "reasoning mode" automatically based on context ; if it doesn't match -> standard."
If you can get auto gating somehow that'd certainly be far far easier to test with, and I could more reasonably run a full 10 runs as it'd be ~10 hours vs 72+ hours.

Note:
Another user alerted us to some jinja template bugs.
These have been corrected, and model is being retested.

Likely this may have affected some results/benches ; performance changes were noted after fixes -> ie net # of tokens as one direct result.

Revised ggufs now up.

Sign up or log in to comment