Instructions to use DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF:IQ4_XS # Run inference directly in the terminal: llama cli -hf DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF:IQ4_XS
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF:IQ4_XS # Run inference directly in the terminal: llama cli -hf DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF:IQ4_XS
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF:IQ4_XS # Run inference directly in the terminal: ./llama-cli -hf DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF:IQ4_XS
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF:IQ4_XS # Run inference directly in the terminal: ./build/bin/llama-cli -hf DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF:IQ4_XS
Use Docker
docker model run hf.co/DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF:IQ4_XS
- LM Studio
- Jan
- vLLM
How to use DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF:IQ4_XS
- Ollama
How to use DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF with Ollama:
ollama run hf.co/DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF:IQ4_XS
- Unsloth Desktop
- Pi
How to use DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF:IQ4_XS
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF:IQ4_XS" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF with Docker Model Runner:
docker model run hf.co/DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF:IQ4_XS
- Lemonade
How to use DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF:IQ4_XS
Run and chat with the model
lemonade run user.LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF-IQ4_XS
List all available models
lemonade list
- Hermes Agent
How to use DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF:IQ4_XS
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF:IQ4_XS
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF:IQ4_XS
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF:IQ4_XS" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
benchmarks, comparable to the other ones I've given for your qwen fusion models
I stopped the benchmarking because I didn't want to dedicate a huge amount more time to it, testing every single thinking mode was a ton of testing.
tldr nothing set for reasoning effort was by far the best compared to anything else.
Reasoning-mode benchmark: LFM2.5-2.6B Turbo-Brilliance NEO-MAX Q8_0
Tested model: DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF, file LFM2.5-2.6B-Q3.8-TBrilliance-NEO-MAX-Q8_0.gguf.
Snapshot taken 2026-09-26, 09:23. Run 1 is complete. Run 2 was still in progress: easy, hard and ultrahard were done for every mode, and abyssal was still running.
Setup
- What changes between modes: only the reasoning mode. Every mode uses the same weights and the same llama.cpp server layout (4 slots Γ 128K context, f16 KV cache). The mode is set server-side with
--reasoning-effort <mode>, which chooses the system block the chat template injects. - Baselines:
offis stock LFM2.5 thinking with nothing injected.highis the template's default mode. - Sampling: thinking on for every test, with the model card's tester settings: temp 1.0, top-k 64, min-p 0.05, top-p 0.95, repeat penalty 1.0.
- Limits: no token cap, so a response can run to the full 128K context. The only other stops are a 1800 s timeout per request and a repetition (loop) detector.
- Test suites: four difficulty tiers (easy, hard, ultrahard, abyssal). The categories are coding (the program is run against tests), logic, JSON, distill, constraint, simulation, cx and flawstep. Planning tests are recorded but not scored.
How to read the scores:
- Each test is a single sample at temp 1.0. A single suite can swing 10β20 points from one run to the next, so small gaps are not meaningful yet.
trunc,loopandblankeach score 0 on that test:trunc: the response ran out of context.loop: the repetition detector abandoned it.blank: it finished without an answer.
wall minis not a speed comparison. Two modes ran at the same time on two servers sharing one GPU, so a mode's time depends on what the other server was doing.
Run 1 (complete: easy, hard, ultrahard, abyssal)
Scores (%)
| mode | easy | hard | ultrahard | abyssal | mean | vs off | vs high | tok/test | trunc | loop | blank | err | wall min |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| off | 95.0 | 75.0 | 15.8 | 1.2 | 46.8 | +0.0 | +9.1 | 11,594 | 0 | 0 | 5 | 0 | 37 |
| deeptree | 91.0 | 78.3 | 15.1 | 0.6 | 46.3 | -0.5 | +8.6 | 12,017 | 0 | 0 | 3 | 0 | 38 |
| omni | 93.8 | 64.6 | 14.8 | 3.1 | 44.1 | -2.7 | +6.4 | 11,287 | 0 | 0 | 3 | 0 | 33 |
| einstein | 93.8 | 65.0 | 16.7 | 0.5 | 44.0 | -2.8 | +6.4 | 13,393 | 0 | 0 | 0 | 0 | 43 |
| spoon | 68.8 | 65.0 | 23.8 | 0.5 | 39.5 | -7.2 | +1.9 | 11,801 | 0 | 2 | 0 | 0 | 28 |
| low | 83.8 | 55.4 | 12.5 | 4.1 | 39.0 | -7.8 | +1.3 | 12,454 | 0 | 0 | 3 | 0 | 39 |
| socrates | 70.0 | 73.8 | 8.3 | 3.7 | 38.9 | -7.8 | +1.3 | 12,350 | 0 | 0 | 2 | 0 | 44 |
| high | 73.8 | 60.8 | 12.4 | 3.5 | 37.6 | -9.1 | +0.0 | 13,835 | 0 | 0 | 0 | 0 | 43 |
| hyper | 67.5 | 57.8 | 16.4 | 0.6 | 35.6 | -11.2 | -2.0 | 12,438 | 0 | 0 | 1 | 0 | 38 |
| logic | 68.8 | 42.1 | 29.6 | 0.2 | 35.1 | -11.6 | -2.5 | 11,909 | 0 | 0 | 3 | 0 | 38 |
| ultra | 73.8 | 47.9 | 15.8 | 0.2 | 34.4 | -12.3 | -3.2 | 14,049 | 0 | 0 | 3 | 0 | 52 |
| medium-low | 78.8 | 41.7 | 13.9 | 3.2 | 34.4 | -12.4 | -3.2 | 11,652 | 0 | 1 | 0 | 0 | 42 |
| medium | 68.8 | 42.5 | 18.9 | 3.8 | 33.5 | -13.3 | -4.1 | 12,344 | 0 | 0 | 1 | 0 | 37 |
By category (all four suites pooled, %)
| mode | coding | constraint | cx | distill | flawstep | json | logic | sim |
|---|---|---|---|---|---|---|---|---|
| off | 57.1 | 4.2 | 0.0 | 25.0 | 0.0 | 50.0 | 62.5 | 12.5 |
| deeptree | 57.9 | 0.0 | 0.0 | 32.5 | 0.0 | 50.0 | 56.2 | 4.2 |
| omni | 62.5 | 0.0 | 0.0 | 29.2 | 0.0 | 50.0 | 43.8 | 12.5 |
| einstein | 46.6 | 4.2 | 0.0 | 27.5 | 0.0 | 50.0 | 56.2 | 12.5 |
| spoon | 24.0 | 0.0 | 0.0 | 27.5 | 0.0 | 58.3 | 50.0 | 25.0 |
| low | 27.3 | 12.5 | 0.0 | 32.5 | 0.0 | 50.0 | 50.0 | 1.8 |
| socrates | 49.2 | 0.0 | 0.0 | 30.0 | 0.0 | 41.7 | 43.8 | 12.5 |
| high | 20.2 | 3.1 | 0.0 | 27.5 | 0.0 | 50.0 | 56.2 | 12.5 |
| hyper | 30.7 | 8.3 | 0.0 | 25.0 | 0.0 | 50.0 | 37.5 | 12.5 |
| logic | 12.2 | 12.5 | 0.0 | 27.5 | 0.0 | 58.3 | 37.5 | 25.0 |
| ultra | 18.6 | 0.0 | 0.0 | 25.0 | 0.0 | 58.3 | 43.8 | 0.0 |
| medium-low | 21.1 | 4.2 | 0.0 | 25.0 | 0.0 | 50.0 | 43.8 | 12.5 |
| medium | 9.2 | 0.0 | 0.0 | 30.0 | 0.0 | 50.0 | 50.0 | 12.5 |
Where the modes disagree
Of 72 graded tests, 8 passed in every mode and 24 failed in every mode. 31 split, meaning some mode scored at least 50 points above another on that test.
Number of split tests on which each mode had the top score (ties count for every tied mode): off 20, omni 18, deeptree 17, einstein 17, socrates 15, spoon 14, low 13, high 13, medium-low 11, hyper 11, ultra 10, logic 10, medium 9.
Per-test scores on the 31 split tests (score Γ100, blank = 0)
| test | off | low | medium-low | medium | high | ultra | omni | deeptree | hyper | socrates | logic | einstein | spoon |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| abyssal:ab_chain3 | 100 | ||||||||||||
| abyssal:ab_chain4 | 100 | ||||||||||||
| abyssal:ab_chain5 | 100 | 100 | 100 | ||||||||||
| easy:code_eval | 100 | 100 | 100 | 100 | 100 | 100 | |||||||
| easy:code_lvp | 100 | 100 | 100 | 100 | 88 | 100 | 100 | 100 | |||||
| easy:code_minwin | 100 | 100 | 100 | 100 | 17 | 100 | |||||||
| easy:code_sliding | 100 | 100 | 100 | 100 | 100 | ||||||||
| easy:code_wordbreak | 100 | 100 | 100 | 100 | 100 | 100 | |||||||
| easy:json_primes | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | |
| easy:json_tool | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | |
| easy:logic_div | 100 | 100 | |||||||||||
| easy:logic_race | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | ||
| hard:h_coin | 100 | 100 | 100 | 100 | 100 | ||||||||
| hard:h_countsmaller | 100 | 100 | 100 | 100 | 100 | 100 | 100 | ||||||
| hard:h_domino | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | |
| hard:h_edit | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | |||||
| hard:h_lis | 100 | 100 | 100 | 100 | 100 | ||||||||
| hard:h_lockers | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | ||
| hard:h_modexp | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | |||
| hard:h_regex | 100 | 100 | 38 | 100 | 50 | 100 | |||||||
| hard:h_sieve | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | |||||
| hard:h_trap | 100 | 100 | 100 | 100 | 100 | 100 | 100 | ||||||
| ultrahard:uh_cover | 100 | 100 | 100 | 100 | 100 | 59 | 100 | ||||||
| ultrahard:uh_cs1 | 33 | 100 | 33 | 67 | 100 | 33 | |||||||
| ultrahard:uh_j2 | 100 | 100 | 100 | 100 | |||||||||
| ultrahard:uh_mul | 100 | 100 | 100 | ||||||||||
| ultrahard:uh_parse | 39 | 72 | 39 | 50 | 33 | 11 | 11 | 39 | |||||
| ultrahard:uh_sim15 | 100 | 100 | 100 | 100 | 33 | 100 | 100 | ||||||
| ultrahard:uh_sim250 | 100 | ||||||||||||
| ultrahard:uh_sim60 | 100 | 100 | 100 | 100 | 100 | ||||||||
| ultrahard:uh_window | 100 | 54 | 8 | 62 | 15 | 100 | 62 |
Run 2 (partial: easy, hard and ultrahard done for every mode; abyssal still running)
The mean below covers easy, hard and ultrahard only, so it is not comparable to run 1's four-suite mean.
Scores (%)
| mode | easy | hard | ultrahard | mean | vs off | vs high | tok/test | trunc | loop | blank | err | wall min |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| off | 93.8 | 66.7 | 28.3 | 62.9 | +0.0 | +14.0 | 6,766 | 0 | 0 | 1 | 0 | 15 |
| deeptree | 95.0 | 75.0 | 12.8 | 60.9 | -2.0 | +12.1 | 9,083 | 0 | 0 | 0 | 0 | 19 |
| omni | 88.8 | 70.8 | 18.4 | 59.3 | -3.6 | +10.5 | 8,614 | 0 | 0 | 2 | 0 | 18 |
| einstein | 88.8 | 52.5 | 16.6 | 52.6 | -10.3 | +3.8 | 8,718 | 0 | 0 | 0 | 0 | 16 |
| socrates | 88.0 | 51.0 | 14.1 | 51.1 | -11.8 | +2.2 | 7,524 | 0 | 0 | 1 | 0 | 17 |
| hyper | 72.5 | 66.7 | 13.4 | 50.9 | -12.0 | +2.0 | 9,289 | 0 | 0 | 0 | 0 | 19 |
| low | 68.8 | 61.7 | 19.5 | 50.0 | -12.9 | +1.1 | 6,887 | 0 | 0 | 0 | 0 | 13 |
| logic | 73.8 | 53.5 | 19.4 | 48.9 | -14.0 | +0.1 | 8,195 | 0 | 0 | 1 | 0 | 15 |
| high | 75.0 | 61.3 | 10.3 | 48.8 | -14.0 | +0.0 | 8,769 | 0 | 0 | 1 | 0 | 19 |
| medium-low | 67.5 | 61.7 | 17.1 | 48.8 | -14.1 | -0.1 | 8,495 | 0 | 0 | 0 | 0 | 21 |
| ultra | 73.8 | 53.3 | 10.4 | 45.8 | -17.1 | -3.0 | 7,646 | 0 | 0 | 0 | 0 | 16 |
| medium | 61.3 | 54.6 | 18.1 | 44.6 | -18.3 | -4.2 | 8,989 | 0 | 0 | 3 | 0 | 20 |
| spoon | 78.8 | 35.4 | 12.3 | 42.2 | -20.7 | -6.7 | 8,320 | 0 | 0 | 0 | 0 | 11 |
By category (easy + hard + ultrahard pooled, %)
| mode | coding | constraint | distill | json | logic | sim |
|---|---|---|---|---|---|---|
| off | 77.5 | 8.3 | 33.3 | 77.8 | 66.7 | 50.0 |
| deeptree | 73.8 | 0.0 | 33.3 | 66.7 | 75.0 | 25.0 |
| omni | 69.5 | 25.0 | 33.3 | 66.7 | 66.7 | 25.0 |
| einstein | 59.3 | 6.2 | 36.7 | 55.6 | 66.7 | 8.3 |
| socrates | 64.5 | 0.0 | 33.3 | 55.6 | 58.3 | 8.3 |
| hyper | 54.8 | 0.0 | 33.3 | 66.7 | 58.3 | 0.0 |
| low | 15.6 | 0.0 | 43.3 | 77.8 | 66.7 | 25.0 |
| logic | 25.0 | 8.3 | 36.7 | 77.8 | 58.3 | 25.0 |
| high | 29.8 | 0.0 | 40.0 | 66.7 | 66.7 | 0.0 |
| medium-low | 18.5 | 8.3 | 43.3 | 66.7 | 58.3 | 50.0 |
| ultra | 21.2 | 0.0 | 43.3 | 77.8 | 50.0 | 0.0 |
| medium | 26.7 | 16.7 | 36.7 | 77.8 | 41.7 | 8.3 |
| spoon | 17.5 | 8.3 | 33.3 | 55.6 | 58.3 | 25.0 |
Where the modes disagree
Of 47 graded tests, 7 passed in every mode, 9 failed in every mode and 28 split.
Number of split tests on which each mode had the top score: off 22, deeptree 21, omni 20, hyper 14, socrates 14, einstein 14, low 12, high 12, logic 12, medium-low 11, medium 10, ultra 10, spoon 9.
Per-test scores on the 28 split tests (score Γ100, blank = 0)
| test | off | low | medium-low | medium | high | ultra | omni | deeptree | hyper | socrates | logic | einstein | spoon |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| easy:code_eval | 100 | 100 | 100 | 100 | 100 | 100 | 86 | 100 | |||||
| easy:code_lvp | 100 | 100 | 100 | 100 | 100 | ||||||||
| easy:code_minwin | 100 | 100 | 100 | 100 | |||||||||
| easy:code_sliding | 100 | 100 | 100 | 100 | 100 | 100 | 100 | ||||||
| easy:code_wordbreak | 100 | 100 | 100 | 100 | 100 | ||||||||
| easy:logic_div | 100 | 100 | |||||||||||
| easy:logic_fact | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | ||
| easy:logic_hh | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | ||
| hard:h_coin | 100 | 100 | 100 | 100 | 100 | 100 | |||||||
| hard:h_countsmaller | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | |||||
| hard:h_domino | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | ||
| hard:h_edit | 100 | 100 | 100 | 100 | 100 | ||||||||
| hard:h_graph | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | ||
| hard:h_lis | 100 | 100 | 100 | ||||||||||
| hard:h_modexp | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | |
| hard:h_regex | 50 | 100 | 100 | 100 | 100 | 100 | 100 | 75 | 75 | 100 | |||
| hard:h_sieve | 100 | 100 | 100 | 100 | 100 | 100 | 100 | ||||||
| hard:h_toolevent | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | |
| hard:h_trap | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 50 | |||||
| ultrahard:uh_cover | 100 | 100 | 100 | 47 | 100 | 100 | 100 | 24 | 24 | ||||
| ultrahard:uh_crt | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | |
| ultrahard:uh_cs1 | 33 | 33 | 67 | 100 | 33 | 33 | |||||||
| ultrahard:uh_j2 | 100 | 100 | 100 | 100 | 100 | ||||||||
| ultrahard:uh_mul | 100 | 100 | 100 | ||||||||||
| ultrahard:uh_parse | 6 | 28 | 11 | 50 | 39 | 11 | 39 | ||||||
| ultrahard:uh_sim15 | 100 | 100 | 100 | 33 | 100 | 33 | 100 | 33 | 100 | ||||
| ultrahard:uh_sim60 | 100 | 100 | 100 | ||||||||||
| ultrahard:uh_window | 77 | 100 | 92 | 77 | 100 | 100 | 62 |
Observations so far
- No mode beat stock thinking.
off(no injection) led both runs, withdeeptreeandomniclose behind. The template defaulthighsat mid-pack, 9β14 points behindoff. - Coding drives most of the differences. The coding scores span roughly 9β78% across modes, while distill and JSON barely move.
- Ultrahard and abyssal are beyond this model. It scores about 8β30% on ultrahard and 0β4% on abyssal in every mode, so those suites say little about which mode is better.
Thank you for detailed testing and notes.
These are very helpful.
A few notes:
- 2.6B is a limiting factor ; that being said LiquidAI did an exceptional job here however.
- In run #1, on very hard/abyssal some of the reasoning modes clearly pulled ahead of "off" / "standard" - some by a long shot. ; I don't know if I would "mean" these kinds of numbers.
- Without "gating" the specialized reasoning modes will attempt to solve problems they were not designed for. (this matches our own internal testing).
- Gating would select the best mode automatically during testing, rather than "manually" as is the case now; IE you could run one test, for all modes VS 12.
- The specialized reasoning is use case specific ; this is part explains the low scores. Larger models will disallow the "reasoning mode" automatically based on context ; if it doesn't match -> standard.
- The reasoning modes were NOT calibrated for this specific model, they are "general", especially the advanced modes. IE min: 9B/27B sized models. (or larger)
- An odd issue: sometimes the model will override/disagree with the reasoning mode, and cause issues itself.
- Parameters for testing (ours) differ greatly from LiquidAi's published ones. (temp is especially critical, rep pen too -> especially as LFM suggests 1.1 vs 1 [off] from us)
- Testing smaller para models is always more difficult // less range (our experience here) // likewise RAISING benchmarks (thru tuning) is also a lot harder too.
- It appears your tests are primarily coding, coding related, and logic ; the alignment for all modes was "general" - with aligned reasoning modes the scores will likely reflect a change.
Again ; thank for publishing - this will help with refinements to reasoning, reasoning modes and other planned improvements.
Yeah I generally use it for a lot of logic chains, tool usage, and coding so that's where most of the benchmarks are at.
"In run #1, on very hard/abyssal some of the reasoning modes clearly pulled ahead of "off" / "standard" - some by a long shot. ; I don't know if I would "mean" these kinds of numbers."
Due to how low the passing scores were with some of the much harder tests I don't know how well it really would do when run over and over again. sometimes models are on the cusp of failing or passing a test so it may flip between passing/failing multiple times if I run it multiple times. Since I only ran the full suite 1.2 times idk that the barely passing numbers mean too much.
"The specialized reasoning is use case specific ; this is part explains the low scores. Larger models will disallow the "reasoning mode" automatically based on context ; if it doesn't match -> standard."
If you can get auto gating somehow that'd certainly be far far easier to test with, and I could more reasonably run a full 10 runs as it'd be ~10 hours vs 72+ hours.
Note:
Another user alerted us to some jinja template bugs.
These have been corrected, and model is being retested.
Likely this may have affected some results/benches ; performance changes were noted after fixes -> ie net # of tokens as one direct result.
Revised ggufs now up.