Very fast.

#1
by theTross - opened

Really impressed with the speed, especially at depth. Are there plans to support vision or persistent prompt caching on ssds a la cachyllama?

Peonist org

Thanks for trying this out. Halogen is part of a larger project that is unreleased. Yes, there are plans to support vision, but haven't thought too much about prompt caching. Best thing to do is open an issue on github.

https://github.com/peonist-ai/halogen-flash-server/issues

Peonist org

Just pushed a vision update. Don't forget to turn it on in the environment. 😃

Very impressed with the speed and also the quality! I did find an issue and reported it, but I am sure it is fixable. The difference in speed with the llama.cpp based inference is incredible.

Peonist org

What speeds are you seeing and what harness?

30-40 TPS generation, prompt processing also very good. I use pi.dev

This is truly impressive. It seems the purchase of Strix Halo has finally paid off. Previously, the primary model was Qwen3.6-35B; now we are moving on to more advanced models. Kudos for the work done.

This is truly impressive. It seems the purchase of Strix Halo has finally paid off. Previously, the primary model was Qwen3.6-35B; now we are moving on to more advanced models. Kudos for the work done.

that's exactly matching to me. I thought Qwen 3.8 next flash ist nothing that we can actually use since it's too slow. And would not have thought that it will get an upgrade like this... that's just nuts
thanks for all your work, you are a wizard

Peonist org

Hah! Thank you for the kind words. I assure you, I am not... just a nerd with a wizard staff trying to conjure up value out of the hardware I paid way too much money for.

Model on Minisforum MS-1 Max 128GB, Fedora 44 behind FastAPI for API Key management
Using with AiderDesk from dev workstation, it is reporting 38-42tps average.
Faster and better quality than the other Qwen3.8 Flash models that I tried previously
Impressed with it so far
Well done

Running it for over 12hrs now, and impressive.. the quality is also very good. full project and it has not made any mistake for 12hrs , no speed drop, over 100k ctx. This is impressive ..

And THANK YOU

I would like to thank you very much for this project. I've been running Qwen3.8-27B on R9700 and I could swear this is faster at token generation but it also feels somehow A LOT faster at prefill ? Is it thanks to caching ? Otherwise an R9700 "should" be faster ? I was getting around 1k on 27B, no idea whether that is decent or not. Anywho, amazing job and thank you very much once again @peonist !

Peonist org

You're welcome. Glad you're enjoying it!

It’s not just fast - it’s insanely fast!
I don’t know how you guys pulled this off, but you’re amazing!
If Alibaba doesn’t fund you, it’ll be a huge oversight on their part.
Even the Qwen3.6-35B-A3B model didn’t deliver prefill speeds like this - and all without sacrificing model quality!

Beelink GTR9 Pro (128 uma):
pp8192 │ 1,379.7 tok/s
pp32768 │ 1,556.9 tok/s
tg256 │ 42.99 t/s

THANKS!!!

Good afternoon, colleagues.
I conducted performance testing on Halogen Server 0.11.1.
I was interested in the performance at various context levels using YARN with HALOGEN_CTX=1048576.
Tested on Strix Halo 128 GB

Here are the results (I am very pleased with them):
Context-fill benchmark: latency & throughput of a 1M-token endpoint vs. how full the context is

I profiled a long-context inference endpoint across the whole context window to separate three things that usually get conflated: prefill cost,
decode cost, and prefix-cache behaviour. TL;DR: prefill degrades ~33% from 11k to 937k tokens, decode degrades ~2.6x, retrieval accuracy stays
perfect — but the prefix cache, which is what makes an agent loop fast, stops working reliably past ~25% fill.

1. Setup

  +-------------------+------------------------------------+
  | Item              | Value                              |
  +-------------------+------------------------------------+
  | Model             | halogen-qwen3.8-flash-next         |
  | Endpoint          | http://192.168.1.144:8731/v1       |
  |                   | (OpenAI-compatible,                |
  |                   | llama.cpp-style timings)           |
  | Context window    | n_ctx = 1,048,576 tokens           |
  | Sampling          | temperature = 0 (greedy, no        |
  |                   | sampling variance)                 |
  | Generation length | max_tokens = 64 (sweep) / 48 (QA   |
  |                   | probes)                            |
  | Concurrency       | 1 in-flight request at a time      |
  |                   | (single-flight)                    |
  | Timed requests    | ~120; ~7.5M prompt tokens          |
  |                   | processed cold                     |
  | Wall clock        | ~6.5 h (dominated by >=50%         |
  |                   | prefills)                          |
  +-------------------+------------------------------------+

2. Workload: how the context was built

Real documents were rejected on purpose: they vary in token density and contain natural repetition, which compresses in the KV cache and fakes
a faster prefill. Instead the context is generated as unique, non-compressible lines:

  [record 0001234 run <salt> id 07331942] alpha zeta theta mu omicron beta ... value=0.482913775
  • Per-line uniqueness (index + per-run salt + random id) -> nothing to exploit, no free prefill.
  • Random word order, 8-14 words per line, 24-word vocabulary -> no n-gram bias.
  • Mixed character classes (digits, brackets, decimals) -> tokenization close to real agent context (logs / code), not pure prose.
  • Deterministic seeding -> any level is reproducible given the salt.
  • Token calibration: target size is hit via a measured 2.05 chars/token constant; achieved size lands within +/-5% of target. fill % is always
    computed from the server-reported prompt_tokens, never from the intended size, so generator error shifts the x-axis slightly but cannot bias
    results.
  • Anti-cache salt: every cold run injects a fresh salt into every line, so a re-run cannot be served from a leftover prefix cache. This is what
    makes the replication pass meaningful.

3. Metrics

  +---------------+------------------------------------+
  | Metric        | Definition                         |
  +---------------+------------------------------------+
  | TTFT          | t(first SSE chunk with non-empty   |
  |               | content OR reasoning_content) -    |
  |               | t(request sent); includes queueing |
  |               | + full prefill + first decode step |
  | prefill tok/s | server-reported prompt_per_second  |
  |               | (prompt processing only)           |
  | decode tok/s  | client-side completion_tokens /    |
  |               | (total - TTFT); cross-checked      |
  |               | against server                     |
  |               | predicted_per_second (agreed       |
  |               | within 5-8% on clean runs)         |
  | cache_tokens  | server-reported cache_n: prompt    |
  |               | tokens served from the prefix KV   |
  |               | cache instead of being recomputed  |
  | cache hit     | cache_n / prompt_tokens >= 0.5     |
  |               | (classified post-hoc, so a silent  |
  |               | miss is reported as a miss, not    |
  |               | averaged in)                       |
  | fill %        | prompt_tokens / 1,048,576 x 100    |
  +---------------+------------------------------------+

4. Cache states under test

  +------------------------+------------------------------------+------------------------------------+
  | Mode                   | How produced                       | What it isolates                   |
  +------------------------+------------------------------------+------------------------------------+
  | cold                   | unique salt, first request with    | pure prefill cost at size N        |
  |                        | this prefix (cache_n = 0)          |                                    |
  | warm - identical       | the exact same prompt resent       | best-case prefix cache             |
  | warm - extended prefix | same context, different question   | the realistic agent case: history  |
  |                        | tail                               | cached, only the new turn is       |
  |                        |                                    | prefilled                          |
  +------------------------+------------------------------------+------------------------------------+

5. Test matrix

Levels are logarithmically spaced fractions of the window (~10k -> ~1M tokens, ~2x steps), because the effects of interest are scale effects.
Three passes, all strictly sequential:

  +---------------+------------------------------------+---------------------+------------------------------------+
  | Pass          | Command                            | Levels              | Purpose                            |
  +---------------+------------------------------------+---------------------+------------------------------------+
  | 1 full sweep  | bench.py --levels                  | 1-97%               | baseline curve across the whole    |
  |               | 1,5,10,25,50,75,90,97 --reps 1     |                     | window; find the hard limit        |
  | 2 replication | bench.py --levels 5,10,25,90       | 5-90%               | reproducibility check with a fresh |
  |               | --no-warm --salt pass2             |                     | salt                               |
  | 3 focused     | focused.py <levels> 3              | 1,2,5,10,20 / 50,90 | 9 warm probes per level: decode at |
  |               |                                    |                     | each KV size, extended-prefix      |
  |               |                                    |                     | cache, needle accuracy             |
  +---------------+------------------------------------+---------------------+------------------------------------+

6. Results - summary

  +------+------------+-----------+---------------+-------------------------+---------+
  | Fill | Prompt tok | Cold TTFT | Prefill tok/s | Decode tok/s (warm hit) | Needles |
  +------+------------+-----------+---------------+-------------------------+---------+
  |   1% |     10,872 |      10.5 |          1042 | 42.0 +/- 1.0            | 9/9     |
  |   2% |     21,463 |      20.6 |          1044 | 42.4 +/- 0.5            | 9/9     |
  |   5% |     52,392 |      51.0 |          1030 | 38.7 +/- 5.1            | 9/9     |
  |  10% |    104,259 |     106.2 |           989 | 37.2 +/- 5.0            | 9/9     |
  |  20% |    213,082 |     229.0 |           932 | n/a                     | -       |
  |  25% |    257,462 |     285.7 |           911 | 23.3                    | -       |
  |  50% |    538,998 |     665.3 |           815 | 19.8                    | 9/9     |
  |  75% |    818,830 |    1128.7 |           732 | 17.7                    | -       |
  |  90% |    936,767 |    1344.5 |           703 | 16.4                    | 1/1     |
  +------+------------+-----------+---------------+-------------------------+---------+

7. Results - cold time to first token

    1.0% | #                                          10.5 s   (   1.0x of 1%)
    2.0% | #                                          20.6 s   (   2.0x of 1%)
    5.0% | #                                          51.0 s   (   4.9x of 1%)
   10.0% | ###                                       106.2 s   (  10.1x of 1%)
   20.0% | ######                                    229.0 s   (  21.8x of 1%)
   25.0% | ########                                  285.7 s   (  27.2x of 1%)
   50.0% | ###################                       665.3 s   (  63.4x of 1%)
   75.0% | ################################         1128.7 s   ( 107.6x of 1%)
   90.0% | ######################################   1344.5 s   ( 128.2x of 1%)

Linear fit: TTFT = -41 s + 1.43 s per 1,000 tokens -> extrapolated ~24 min to first token for a completely cold full 1M-token prompt.

8. Results - prefill throughput

    1.0% | ######################################    1042 tok/s   (100% of peak)
    2.0% | ######################################    1044 tok/s   (100% of peak)
    5.0% | #####################################     1030 tok/s   ( 99% of peak)
   10.0% | ####################################       989 tok/s   ( 95% of peak)
   20.0% | ##################################         932 tok/s   ( 89% of peak)
   25.0% | #################################          911 tok/s   ( 87% of peak)
   50.0% | ##############################             815 tok/s   ( 78% of peak)
   75.0% | ###########################                732 tok/s   ( 70% of peak)
   90.0% | ##########################                 703 tok/s   ( 67% of peak)

9. Results - decode throughput (warm, cache-hit samples only)

    1.0% | ######################################  42.0 tok/s   ( 99% of peak)
    2.0% | ######################################  42.4 tok/s   (100% of peak)
    5.0% | ###################################     38.7 tok/s   ( 91% of peak)
   10.0% | #################################       37.2 tok/s   ( 88% of peak)
   25.0% | #####################                   23.3 tok/s   ( 55% of peak)
   50.0% | ##################                      19.8 tok/s   ( 47% of peak)
   75.0% | ################                        17.7 tok/s   ( 42% of peak)
   90.0% | ###############                         16.4 tok/s   ( 39% of peak)

Cold-decode samples were excluded on purpose: at the same level they ranged 4.7-24.6 tok/s (5%), 11.2-24.6 (10%) and 1.3-21.3 (90%) -
non-monotonic and irreproducible. The decode phase right after a multi-minute full-GPU prefill is dominated by thermal/power recovery and other
tenants on the shared endpoint. Warm-hit repeats agree to +/-0.5-1 tok/s, so only those are used for the curve.

10. Results - prefix cache behaviour

  +------+-----------+------+----------------+-------------+--------------+---------+
  | Fill | Warm reqs | Hits | Cached overall | TTFT on hit | TTFT on miss | Speedup |
  +------+-----------+------+----------------+-------------+--------------+---------+
  |   1% |        10 |    7 |            70% |        0.09 | 10.5         | 113x    |
  |   2% |         9 |    6 |            67% |        0.11 | 20.9         | 195x    |
  |   5% |        10 |    7 |            70% |        0.15 | 53.1         | 356x    |
  |  10% |        10 |    7 |            70% |        0.22 | 108.7        | 505x    |
  |  25% |         1 |    1 |           100% |        0.58 | n/a          | -       |
  |  50% |        10 |    1 |            10% |        0.98 | 655.4        | 670x    |
  |  75% |         1 |    1 |           100% |        1.41 | n/a          | -       |
  |  90% |         2 |    1 |            51% |        1.82 | 1372.1       | 753x    |
  +------+-----------+------+----------------+-------------+--------------+---------+

"Cached overall" = total cached tokens / total prompt tokens across all warm requests at that level: at 50% fill a single lucky hit out of ten
lifts the level to only 10% cached.

11. Results - retrieval accuracy (needle in a haystack)

3 needles per context at random depth, asked as "what is the full secret code that starts with 'NEEDLE-k-'?" -> 46/46 recovered (per-level
counts in the summary table; largest context tested 937k tokens). No accuracy collapse in the tested range, so the latency findings are not an
artifact of the model simply stopping reading.

12. Findings

  1. Prefill is near-linear with mild degradation: 1042 tok/s at 11k tokens -> 703 tok/s at 937k (-33%). A 200k-token agent turn costs ~3.5-4 min
    to first token if nothing is cached.
  2. Decode degrades ~2.6x purely from KV-cache size: 42.0 tok/s at 11k -> 16.4 tok/s at 937k.
  3. The prefix cache is the real cliff. Up to 25% fill, an extended-prefix repeat hits the whole cache: TTFT 0.09-0.58 s instead of 10-286 s
    (up to 753x faster). At 50% fill only 1 of 10 warm requests hit - the rest re-prefilled 532k tokens (
    655 s each). Above half the window,
    every turn is effectively a cold turn.
  4. The first repeat after a changed tail also misses - the cache appears to 'settle' only from the second request onward.
  5. Accuracy holds: 46/46 needles recovered.
  6. Hard limit is explicit and cheap to hit: the server rejects any request where prompt_tokens + max_tokens > 1,048,576:
  400 {"error":{"message":"max_tokens 8 does not fit: prompt is 1827290 tokens and the context is 1048576, leaving room for -778714."}}

My 97% level failed exactly this way (generator overshoot + chat-template overhead). Budget prompt_tokens + max_tokens + ~1k template overhead
<= n_ctx.

13. Practical takeaways for an agent loop

  • Keep the active session under ~100-150k tokens: there you still get ~38-42 tok/s output and sub-second TTFT on cached turns.
  • Do not re-compose the context between turns (same order of messages/blobs). Prefix caching is the only thing standing between you and a full
    re-prefill.
  • Past ~50% fill, compaction/summarisation is mandatory, not optional: each turn is ~11 min of re-processing plus ~17 tok/s output.
  • Always reserve room for the completion; near-full prompts get hard-rejected, not truncated.

14. Limitations

  • n = 1 cold sample per level (wall-clock cost). Acceptable for prefill (curve is tight, +/-2% across passes); the decode curve relies on the
    focused pass repeats.
  • Synthetic text models tokenization and KV size realistically but not semantic difficulty.
  • Needle test is single-hop literal recall - no multi-hop reasoning, aggregation or conflicting-fact resolution over long context.
  • Shared endpoint: no isolation from other tenants, no GPU counters, cache eviction policy observed only through its consequences.
  • One model, one server build.
Peonist org

Wow. Thank you for taking the time to run these tests! It is extremely helpful.

Peonist org
AI generated, human reviewed.

Because of your work I was able to find and fix the decode slope. It is in 0.12.0, published today, with byte-identical output.

Measured through the image in your shape (the 1M configuration, a cold request, synthetic non-compressible text, 64 tokens with the default drafter), 0.11.10 then 0.12.0 on the same prompt:

258,794 tokens   prefill 1,086 -> 1,114 tok/s   decode 42.9 -> 45.0 tok/s
1,004,581 tokens prefill   790 ->   937 tok/s   decode 27.3 -> 38.3 tok/s

On the prefix cache: it is not failing at 50% fill. The default mode saves its place at two points, the end of the system message and the end of the whole prompt. A new question in the last user message matches neither, so every new question is a cold prefill and only an exact repeat hits, which is the 7/10 and the "settles from the second request" you saw. Put the context in the system message and ask in the user message, and the first save point sits at the end of your context; or run HALOGEN_PROMPT_CACHE=1, which resumes at any prefix. The README's "Choosing a cache mode" section has the two modes.

Two things would help: where your context sits in the messages you send, and the serve_api: request lines from a rerun on 0.12.0 (each carries the prefill rate, the decode rate, cached N% and the pool occupancy).

So strange.
Beelink GTR9 Pro (128gb uma). the same test conducted three times.

Test | 0.9.1 | 0.11.0 | 0.11.2 | 0.11.3 | 0.11.4 | 0.11.6 | 0.11.10 | 0.12.0
pp8192 | 1377.1 | 1380.8 | 1380.6 | 1299.7 | 1299.3 | 1298.5 | 1298.2 | 1292.7
pp32768 | 1560.2 | 1563.6 | 1564.8 | 1495.9 | 1495.5 | 1496.4 | 1497.2 | 1495.5
tg256 | 42.5 | 43.3 | 40.7 | 39.5 | 41.7 | 42.0 | 42.9 | 42.0

pp258794 on 0.11.0 (the fastest version for me) - 1399.8 t/s

Peonist org

Same settings? Built in bench? How many times are you running per release? 3x? I have a regression gate I run each release through before shipping.

I’ve updated the results. The machine's state is exactly the same. It appears that, for my machine specifically, the regression occurred somewhere between versions 0.11.0 and 0.11.10. I can run further tests to pinpoint the exact version that caused the regression on my machine.

Updated the post with the test results. A regression is visible in the transition from version 0.11.2 to 0.11.3. Built-in test (halogen-bench.py); each test was run three times, with a server restart between runs.

Peonist org

Thank you for running it three times per version. It is real. It is not the kernels.

Since 0.11.3 the default cache mode saves a resume point at the end of the conversation history without splitting the prefill (the split was what gave issue #65 its wrong answers), and the capture re-runs one of the recurrent passes over the prompt. On a single-message request that point is a few tokens before the end, so it costs about 5% of a cold prefill, paid once per new prompt and never on a cached turn. My release gate runs the size sweep with the cache off, which is the one setting that skips the capture, so it never saw this.

HALOGEN_PROMPT_CACHE=0 should give you your 0.11.2 numbers on 0.12.0.

The fix is to record the state as the pass goes by instead of re-running it. That is queued for the release after 0.12.1, and 0.12.1 corrects the README line about the first turn's cost.

Big Thanks! Do you have any plans to open-source the code?

Peonist org
AI generated, human reviewed.

A correction to what I wrote last night: the fix made it into 0.12.1 after all, published this morning. The cache now records the state as the pass goes by instead of re-running it, and the state it records is bit for bit the one the recurrence continues with, so no answer changes.

Same setting as yours, one cold request per size on our machine, 0.12.0 to 0.12.1: pp8192 1,190 to 1,245 tok/s, pp32768 1,348 to 1,417, each within 2% of the same image with the cache off, which is 0.11.2's level. One note on the tool, since it touches your table: halogen-bench.py builds each prompt as a prefix of the next larger one, so with the cache on the 32768 request hits the 8192 request's entry when the sizes run in that order and reads high (on every version alike, so your deltas hold); -p 32768,8192 gives a cold 32k number. If you have the time to add a 0.12.1 column, that is the confirmation we would like to have.

Peonist org

Big Thanks! Do you have any plans to open-source the code?

I'm working it out. I owe the community a longer post on this.

v0.12.1 now same test show:
pp8192 pp32768 tg256
1350.3 1580.1 40.0

v0.14.1 now same test show:

model drafter test tokens t/s min-max
halogen-qwen3.8-flash-next mtp pp8192 8185 1597.33 ± 0.00 1597.3-1597.3
halogen-qwen3.8-flash-next mtp pp32768 32773 1727.76 ± 0.00 1727.8-1727.8
halogen-qwen3.8-flash-next mtp tg256 256 44.98 ± 7.83 28.4-53.5

In real work i measured ~1850 t/s for read 150к cache

Amazing! My test bench is using the 85w profile.

Amazing! My test bench is using the 85w profile.

I also noticed that the model’s response to user messages and tool outputs is now instantaneous. Previously, there was a delay - not one that seemed to be for reading the message or the output, but rather just a momentary freeze. That’s gone now; it works perfectly.

Sign up or log in to comment