File size: 28,875 Bytes
27a5002
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
0.00.114.439 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
0.00.114.444 I device_info:
0.00.114.555 I   - ROCm0   : AMD Radeon Graphics (131072 MiB, 123723 MiB free)
0.00.114.708 I   - Vulkan0 : AMD Radeon Graphics (RADV GFX1151) (132096 MiB, 131922 MiB free)
0.00.114.715 I   - CPU     : AMD RYZEN AI MAX+ 395 w/ Radeon 8060S (127438 MiB, 127438 MiB free)
0.00.114.796 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 | 
0.00.114.847 I srv          init: running without SSL
0.00.114.871 I srv          init: using 31 threads for HTTP server
0.00.114.873 I srv          init: the WebUI is disabled
0.00.114.942 I srv         start: binding port with default address family
0.00.116.175 I srv          main: loading model
0.00.116.182 I srv    load_model: loading model '/mnt/models/nex-n2.5-mini/out/Nex-N2.5-mini-Q4_0_ROCMFP4_STRIX_LEAN.gguf'
0.00.160.730 W llama_model_loader: direct I/O is enabled, disabling mmap
0.21.513.834 W llama_context: n_ctx_seq (65536) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
0.21.793.651 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
0.22.217.020 I srv    load_model: initializing slots, n_slots = 1
0.22.487.377 W srv    load_model: speculative decoding will use checkpoints
0.22.487.387 W common_speculative_init: no implementations specified for speculative decoding
0.22.487.390 I slot   load_model: id  0 | task -1 | new slot, n_ctx = 65536
0.22.487.471 I srv    load_model: prompt cache RAM enabled: limit_mib=8192
0.22.487.473 I srv    load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
0.22.487.495 I srv          init: idle slots will be saved to prompt cache upon starting a new task
0.22.540.918 I init: chat template, example_format: '<|im_start|>system
You are a helpful assistant<|im_end|>
<|im_start|>user
Hello<|im_end|>
<|im_start|>assistant
<think>

</think>

Hi there<|im_end|>
<|im_start|>user
How are you?<|im_end|>
<|im_start|>assistant
<think>

</think>

'
0.22.582.904 I srv          init: init: chat template, thinking = 0
0.22.582.984 I srv          main: model loaded
0.22.582.995 I srv          main: server is listening on http://127.0.0.1:18600
0.22.583.003 I srv  update_slots: all slots are idle
0.23.994.344 I srv  params_from_: Chat format: peg-native
0.23.995.890 I slot get_availabl: id  0 | task -1 | selected slot by LRU, t_last = -1
0.23.995.893 I srv  get_availabl: updating prompt cache
0.23.995.900 I srv          load:  - looking for better prompt, base f_keep = -1.000, sim = 0.000
0.23.995.907 I srv        update:  - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 65536 tokens, 8589934592 est)
0.23.995.909 I srv  get_availabl: prompt cache update took 0.01 ms
0.23.996.222 I reasoning-budget: activated, budget=2147483647 tokens
0.23.996.243 I slot launch_slot_: id  0 | task 0 | processing task, is_child = 0
0.24.656.206 I slot create_check: id  0 | task 0 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
0.25.028.892 I reasoning-budget: deactivated (natural end)
0.25.816.313 I slot print_timing: id  0 | task 0 | 
prompt eval time =     698.88 ms /   422 tokens (    1.66 ms per token,   603.83 tokens per second)
       eval time =    1121.16 ms /    54 tokens (   20.76 ms per token,    48.16 tokens per second)
      total time =    1820.04 ms /   476 tokens
0.25.816.410 I slot      release: id  0 | task 0 | stop processing: n_tokens = 475, truncated = 0
0.25.816.424 I srv  update_slots: all slots are idle
0.25.830.282 I srv  params_from_: Chat format: peg-native
0.25.830.661 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.906 (> 0.100 thold), f_keep = 0.853
0.25.830.794 I reasoning-budget: activated, budget=2147483647 tokens
0.25.830.832 I slot launch_slot_: id  0 | task 56 | processing task, is_child = 0
0.25.830.841 W slot update_slots: id  0 | task 56 | n_past = 405, slot.prompt.tokens.size() = 475, seq_id = 0, pos_min = 474, n_swa = 0
0.25.830.842 I slot update_slots: id  0 | task 56 | Checking checkpoint with [417, 417] against 405...
0.25.830.844 W slot update_slots: id  0 | task 56 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.25.830.846 W slot update_slots: id  0 | task 56 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.26.482.806 I slot create_check: id  0 | task 56 | created context checkpoint 1 of 32 (pos_min = 442, pos_max = 442, n_tokens = 443, size = 62.813 MiB)
0.26.903.330 I reasoning-budget: deactivated (natural end)
0.28.669.262 I slot print_timing: id  0 | task 56 | n_decoded =    100, tg =  47.24 t/s
0.28.729.150 I slot print_timing: id  0 | task 56 | 
prompt eval time =     721.52 ms /   447 tokens (    1.61 ms per token,   619.52 tokens per second)
       eval time =    2176.78 ms /   103 tokens (   21.13 ms per token,    47.32 tokens per second)
      total time =    2898.30 ms /   550 tokens
0.28.729.226 I slot      release: id  0 | task 56 | stop processing: n_tokens = 549, truncated = 0
0.28.729.266 I srv  update_slots: all slots are idle
0.28.742.730 I srv  params_from_: Chat format: peg-native
0.28.743.168 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.953 (> 0.100 thold), f_keep = 0.738
0.28.743.356 I reasoning-budget: activated, budget=2147483647 tokens
0.28.743.392 I slot launch_slot_: id  0 | task 161 | processing task, is_child = 0
0.28.743.406 W slot update_slots: id  0 | task 161 | n_past = 405, slot.prompt.tokens.size() = 549, seq_id = 0, pos_min = 548, n_swa = 0
0.28.743.407 I slot update_slots: id  0 | task 161 | Checking checkpoint with [442, 442] against 405...
0.28.743.407 W slot update_slots: id  0 | task 161 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.28.743.411 W slot update_slots: id  0 | task 161 | erased invalidated context checkpoint (pos_min = 442, pos_max = 442, n_tokens = 443, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.29.274.427 I slot create_check: id  0 | task 161 | created context checkpoint 1 of 32 (pos_min = 420, pos_max = 420, n_tokens = 421, size = 62.813 MiB)
0.29.681.788 I reasoning-budget: deactivated (natural end)
0.30.465.485 I slot print_timing: id  0 | task 161 | 
prompt eval time =     568.14 ms /   425 tokens (    1.34 ms per token,   748.05 tokens per second)
       eval time =    1153.92 ms /    57 tokens (   20.24 ms per token,    49.40 tokens per second)
      total time =    1722.06 ms /   482 tokens
0.30.465.571 I slot      release: id  0 | task 161 | stop processing: n_tokens = 481, truncated = 0
0.30.465.599 I srv  update_slots: all slots are idle
0.30.480.783 I srv  params_from_: Chat format: peg-native
0.30.481.142 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.953 (> 0.100 thold), f_keep = 0.842
0.30.481.346 I reasoning-budget: activated, budget=2147483647 tokens
0.30.481.386 I slot launch_slot_: id  0 | task 220 | processing task, is_child = 0
0.30.481.397 W slot update_slots: id  0 | task 220 | n_past = 405, slot.prompt.tokens.size() = 481, seq_id = 0, pos_min = 480, n_swa = 0
0.30.481.397 I slot update_slots: id  0 | task 220 | Checking checkpoint with [420, 420] against 405...
0.30.481.399 W slot update_slots: id  0 | task 220 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.30.481.401 W slot update_slots: id  0 | task 220 | erased invalidated context checkpoint (pos_min = 420, pos_max = 420, n_tokens = 421, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.31.008.867 I slot create_check: id  0 | task 220 | created context checkpoint 1 of 32 (pos_min = 420, pos_max = 420, n_tokens = 421, size = 62.813 MiB)
0.31.355.463 I reasoning-budget: deactivated (natural end)
0.31.458.677 I slot print_timing: id  0 | task 220 | 
prompt eval time =     568.77 ms /   425 tokens (    1.34 ms per token,   747.23 tokens per second)
       eval time =     408.50 ms /    20 tokens (   20.42 ms per token,    48.96 tokens per second)
      total time =     977.27 ms /   445 tokens
0.31.458.774 I slot      release: id  0 | task 220 | stop processing: n_tokens = 444, truncated = 0
0.31.458.805 I srv  update_slots: all slots are idle
0.31.510.659 I srv  params_from_: Chat format: peg-native
0.31.511.238 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.914
0.31.511.571 I reasoning-budget: activated, budget=2147483647 tokens
0.31.511.663 I slot launch_slot_: id  0 | task 242 | processing task, is_child = 0
0.31.511.677 W slot update_slots: id  0 | task 242 | n_past = 406, slot.prompt.tokens.size() = 444, seq_id = 0, pos_min = 443, n_swa = 0
0.31.511.684 I slot update_slots: id  0 | task 242 | Checking checkpoint with [420, 420] against 406...
0.31.511.685 W slot update_slots: id  0 | task 242 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.31.511.690 W slot update_slots: id  0 | task 242 | erased invalidated context checkpoint (pos_min = 420, pos_max = 420, n_tokens = 421, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.32.149.342 I slot create_check: id  0 | task 242 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
0.32.405.933 I reasoning-budget: deactivated (natural end)
0.33.283.206 I slot print_timing: id  0 | task 242 | 
prompt eval time =     677.52 ms /   422 tokens (    1.61 ms per token,   622.86 tokens per second)
       eval time =    1093.99 ms /    52 tokens (   21.04 ms per token,    47.53 tokens per second)
      total time =    1771.51 ms /   474 tokens
0.33.283.283 I slot      release: id  0 | task 242 | stop processing: n_tokens = 473, truncated = 0
0.33.283.309 I srv  update_slots: all slots are idle
0.33.311.488 I srv  params_from_: Chat format: peg-native
0.33.312.015 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.854 (> 0.100 thold), f_keep = 0.890
0.33.312.441 I reasoning-budget: activated, budget=2147483647 tokens
0.33.312.556 I slot launch_slot_: id  0 | task 296 | processing task, is_child = 0
0.33.312.580 W slot update_slots: id  0 | task 296 | n_past = 421, slot.prompt.tokens.size() = 473, seq_id = 0, pos_min = 472, n_swa = 0
0.33.312.583 I slot update_slots: id  0 | task 296 | Checking checkpoint with [417, 417] against 421...
0.33.321.205 W slot update_slots: id  0 | task 296 | restored context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_past = 418, size = 62.813 MiB)
0.33.558.612 I slot create_check: id  0 | task 296 | created context checkpoint 2 of 32 (pos_min = 488, pos_max = 488, n_tokens = 489, size = 62.813 MiB)
0.34.218.104 I reasoning-budget: deactivated (natural end)
0.34.556.664 I slot print_timing: id  0 | task 296 | 
prompt eval time =     286.32 ms /    75 tokens (    3.82 ms per token,   261.95 tokens per second)
       eval time =     957.75 ms /    41 tokens (   23.36 ms per token,    42.81 tokens per second)
      total time =    1244.07 ms /   116 tokens
0.34.556.763 I slot      release: id  0 | task 296 | stop processing: n_tokens = 533, truncated = 0
0.34.556.805 I srv  update_slots: all slots are idle
0.34.571.001 I srv  params_from_: Chat format: peg-native
0.34.571.465 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.972 (> 0.100 thold), f_keep = 0.769
0.34.571.696 I reasoning-budget: activated, budget=2147483647 tokens
0.34.571.740 I slot launch_slot_: id  0 | task 339 | processing task, is_child = 0
0.34.571.755 W slot update_slots: id  0 | task 339 | n_past = 410, slot.prompt.tokens.size() = 533, seq_id = 0, pos_min = 532, n_swa = 0
0.34.571.756 I slot update_slots: id  0 | task 339 | Checking checkpoint with [488, 488] against 410...
0.34.571.758 I slot update_slots: id  0 | task 339 | Checking checkpoint with [417, 417] against 410...
0.34.571.759 W slot update_slots: id  0 | task 339 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.34.571.763 W slot update_slots: id  0 | task 339 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.34.573.311 W slot update_slots: id  0 | task 339 | erased invalidated context checkpoint (pos_min = 488, pos_max = 488, n_tokens = 489, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.35.110.746 I slot create_check: id  0 | task 339 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
0.35.584.142 I reasoning-budget: deactivated (natural end)
0.36.392.408 I slot print_timing: id  0 | task 339 | 
prompt eval time =     576.37 ms /   422 tokens (    1.37 ms per token,   732.16 tokens per second)
       eval time =    1244.26 ms /    62 tokens (   20.07 ms per token,    49.83 tokens per second)
      total time =    1820.64 ms /   484 tokens
0.36.392.474 I slot      release: id  0 | task 339 | stop processing: n_tokens = 483, truncated = 0
0.36.392.500 I srv  update_slots: all slots are idle
0.36.404.680 I srv  params_from_: Chat format: peg-native
0.36.405.094 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.935 (> 0.100 thold), f_keep = 0.839
0.36.405.518 I reasoning-budget: activated, budget=2147483647 tokens
0.36.405.599 I slot launch_slot_: id  0 | task 403 | processing task, is_child = 0
0.36.405.620 W slot update_slots: id  0 | task 403 | n_past = 405, slot.prompt.tokens.size() = 483, seq_id = 0, pos_min = 482, n_swa = 0
0.36.405.621 I slot update_slots: id  0 | task 403 | Checking checkpoint with [417, 417] against 405...
0.36.405.625 W slot update_slots: id  0 | task 403 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.36.405.630 W slot update_slots: id  0 | task 403 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.36.974.524 I slot create_check: id  0 | task 403 | created context checkpoint 1 of 32 (pos_min = 428, pos_max = 428, n_tokens = 429, size = 62.813 MiB)
0.39.070.997 I slot print_timing: id  0 | task 403 | n_decoded =    100, tg =  48.88 t/s
0.39.448.467 I reasoning-budget: deactivated (natural end)
0.41.042.559 I slot print_timing: id  0 | task 403 | 
prompt eval time =     619.66 ms /   433 tokens (    1.43 ms per token,   698.77 tokens per second)
       eval time =    4017.27 ms /   200 tokens (   20.09 ms per token,    49.79 tokens per second)
      total time =    4636.93 ms /   633 tokens
0.41.042.634 I slot      release: id  0 | task 403 | stop processing: n_tokens = 632, truncated = 0
0.41.042.665 I srv  update_slots: all slots are idle
0.41.054.733 I srv  params_from_: Chat format: peg-native
0.41.055.120 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.641
0.41.055.469 I reasoning-budget: activated, budget=2147483647 tokens
0.41.055.471 I reasoning-budget: deactivated (natural end)
0.41.055.514 I slot launch_slot_: id  0 | task 605 | processing task, is_child = 0
0.41.055.527 W slot update_slots: id  0 | task 605 | n_past = 405, slot.prompt.tokens.size() = 632, seq_id = 0, pos_min = 631, n_swa = 0
0.41.055.529 I slot update_slots: id  0 | task 605 | Checking checkpoint with [428, 428] against 405...
0.41.055.530 W slot update_slots: id  0 | task 605 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.41.055.533 W slot update_slots: id  0 | task 605 | erased invalidated context checkpoint (pos_min = 428, pos_max = 428, n_tokens = 429, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.41.597.061 I slot create_check: id  0 | task 605 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
0.42.459.956 I slot print_timing: id  0 | task 605 | 
prompt eval time =     582.62 ms /   424 tokens (    1.37 ms per token,   727.74 tokens per second)
       eval time =     821.78 ms /    39 tokens (   21.07 ms per token,    47.46 tokens per second)
      total time =    1404.41 ms /   463 tokens
0.42.460.047 I slot      release: id  0 | task 605 | stop processing: n_tokens = 462, truncated = 0
0.42.460.077 I srv  update_slots: all slots are idle
0.42.475.720 I srv  params_from_: Chat format: peg-native
0.42.476.233 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.902 (> 0.100 thold), f_keep = 0.877
0.42.476.486 I reasoning-budget: activated, budget=2147483647 tokens
0.42.476.491 I reasoning-budget: deactivated (natural end)
0.42.476.536 I slot launch_slot_: id  0 | task 646 | processing task, is_child = 0
0.42.476.550 W slot update_slots: id  0 | task 646 | n_past = 405, slot.prompt.tokens.size() = 462, seq_id = 0, pos_min = 461, n_swa = 0
0.42.476.551 I slot update_slots: id  0 | task 646 | Checking checkpoint with [419, 419] against 405...
0.42.476.553 W slot update_slots: id  0 | task 646 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.42.476.571 W slot update_slots: id  0 | task 646 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.43.126.158 I slot create_check: id  0 | task 646 | created context checkpoint 1 of 32 (pos_min = 444, pos_max = 444, n_tokens = 445, size = 62.813 MiB)
0.44.988.511 I slot print_timing: id  0 | task 646 | 
prompt eval time =     693.47 ms /   449 tokens (    1.54 ms per token,   647.47 tokens per second)
       eval time =    1818.46 ms /    86 tokens (   21.14 ms per token,    47.29 tokens per second)
      total time =    2511.93 ms /   535 tokens
0.44.988.703 I slot      release: id  0 | task 646 | stop processing: n_tokens = 534, truncated = 0
0.44.988.760 I srv  update_slots: all slots are idle
0.45.002.229 I srv  params_from_: Chat format: peg-native
0.45.002.631 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.948 (> 0.100 thold), f_keep = 0.758
0.45.003.110 I reasoning-budget: activated, budget=2147483647 tokens
0.45.003.114 I reasoning-budget: deactivated (natural end)
0.45.003.190 I slot launch_slot_: id  0 | task 734 | processing task, is_child = 0
0.45.003.216 W slot update_slots: id  0 | task 734 | n_past = 405, slot.prompt.tokens.size() = 534, seq_id = 0, pos_min = 533, n_swa = 0
0.45.003.220 I slot update_slots: id  0 | task 734 | Checking checkpoint with [444, 444] against 405...
0.45.003.221 W slot update_slots: id  0 | task 734 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.45.003.227 W slot update_slots: id  0 | task 734 | erased invalidated context checkpoint (pos_min = 444, pos_max = 444, n_tokens = 445, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.45.561.942 I slot create_check: id  0 | task 734 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
0.46.436.833 I slot print_timing: id  0 | task 734 | 
prompt eval time =     595.84 ms /   427 tokens (    1.40 ms per token,   716.64 tokens per second)
       eval time =     837.77 ms /    39 tokens (   21.48 ms per token,    46.55 tokens per second)
      total time =    1433.60 ms /   466 tokens
0.46.436.923 I slot      release: id  0 | task 734 | stop processing: n_tokens = 465, truncated = 0
0.46.436.955 I srv  update_slots: all slots are idle
0.46.450.626 I srv  params_from_: Chat format: peg-native
0.46.451.043 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.948 (> 0.100 thold), f_keep = 0.871
0.46.451.527 I reasoning-budget: activated, budget=2147483647 tokens
0.46.451.532 I reasoning-budget: deactivated (natural end)
0.46.451.617 I slot launch_slot_: id  0 | task 775 | processing task, is_child = 0
0.46.451.637 W slot update_slots: id  0 | task 775 | n_past = 405, slot.prompt.tokens.size() = 465, seq_id = 0, pos_min = 464, n_swa = 0
0.46.451.641 I slot update_slots: id  0 | task 775 | Checking checkpoint with [422, 422] against 405...
0.46.451.643 W slot update_slots: id  0 | task 775 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.46.451.649 W slot update_slots: id  0 | task 775 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.46.982.640 I slot create_check: id  0 | task 775 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
0.47.142.299 I slot print_timing: id  0 | task 775 | 
prompt eval time =     585.65 ms /   427 tokens (    1.37 ms per token,   729.10 tokens per second)
       eval time =     104.99 ms /     4 tokens (   26.25 ms per token,    38.10 tokens per second)
      total time =     690.65 ms /   431 tokens
0.47.142.414 I slot      release: id  0 | task 775 | stop processing: n_tokens = 430, truncated = 0
0.47.142.449 I srv  update_slots: all slots are idle
0.47.193.526 I srv  params_from_: Chat format: peg-native
0.47.195.446 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.958 (> 0.100 thold), f_keep = 0.944
0.47.195.987 I reasoning-budget: activated, budget=2147483647 tokens
0.47.195.992 I reasoning-budget: deactivated (natural end)
0.47.196.101 I slot launch_slot_: id  0 | task 781 | processing task, is_child = 0
0.47.196.122 W slot update_slots: id  0 | task 781 | n_past = 406, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
0.47.196.128 I slot update_slots: id  0 | task 781 | Checking checkpoint with [422, 422] against 406...
0.47.196.129 W slot update_slots: id  0 | task 781 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.47.196.139 W slot update_slots: id  0 | task 781 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.47.774.566 I slot create_check: id  0 | task 781 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
0.48.760.957 I slot print_timing: id  0 | task 781 | 
prompt eval time =     635.18 ms /   424 tokens (    1.50 ms per token,   667.53 tokens per second)
       eval time =     929.64 ms /    40 tokens (   23.24 ms per token,    43.03 tokens per second)
      total time =    1564.83 ms /   464 tokens
0.48.761.052 I slot      release: id  0 | task 781 | stop processing: n_tokens = 463, truncated = 0
0.48.761.080 I srv  update_slots: all slots are idle
0.48.783.508 I srv  params_from_: Chat format: peg-native
0.48.783.860 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.935 (> 0.100 thold), f_keep = 1.000
0.48.784.082 I reasoning-budget: activated, budget=2147483647 tokens
0.48.784.084 I reasoning-budget: deactivated (natural end)
0.48.784.127 I slot launch_slot_: id  0 | task 823 | processing task, is_child = 0
0.48.935.107 I slot create_check: id  0 | task 823 | created context checkpoint 2 of 32 (pos_min = 490, pos_max = 490, n_tokens = 491, size = 62.813 MiB)
0.49.277.790 I slot print_timing: id  0 | task 823 | 
prompt eval time =     217.60 ms /    32 tokens (    6.80 ms per token,   147.06 tokens per second)
       eval time =     276.04 ms /    13 tokens (   21.23 ms per token,    47.09 tokens per second)
      total time =     493.64 ms /    45 tokens
0.49.277.883 I slot      release: id  0 | task 823 | stop processing: n_tokens = 507, truncated = 0
0.49.277.910 I srv  update_slots: all slots are idle
0.49.324.008 I srv  params_from_: Chat format: peg-native
0.49.324.548 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.967 (> 0.100 thold), f_keep = 0.809
0.49.324.947 I reasoning-budget: activated, budget=2147483647 tokens
0.49.324.952 I reasoning-budget: deactivated (natural end)
0.49.325.029 I slot launch_slot_: id  0 | task 838 | processing task, is_child = 0
0.49.325.041 W slot update_slots: id  0 | task 838 | n_past = 410, slot.prompt.tokens.size() = 507, seq_id = 0, pos_min = 506, n_swa = 0
0.49.325.042 I slot update_slots: id  0 | task 838 | Checking checkpoint with [490, 490] against 410...
0.49.325.043 I slot update_slots: id  0 | task 838 | Checking checkpoint with [419, 419] against 410...
0.49.325.048 W slot update_slots: id  0 | task 838 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.49.325.051 W slot update_slots: id  0 | task 838 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.49.328.874 W slot update_slots: id  0 | task 838 | erased invalidated context checkpoint (pos_min = 490, pos_max = 490, n_tokens = 491, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.49.907.729 I slot create_check: id  0 | task 838 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
0.50.729.586 I slot print_timing: id  0 | task 838 | 
prompt eval time =     624.90 ms /   424 tokens (    1.47 ms per token,   678.50 tokens per second)
       eval time =     779.62 ms /    40 tokens (   19.49 ms per token,    51.31 tokens per second)
      total time =    1404.52 ms /   464 tokens
0.50.729.654 I slot      release: id  0 | task 838 | stop processing: n_tokens = 463, truncated = 0
0.50.729.686 I srv  update_slots: all slots are idle
0.50.775.784 I srv  params_from_: Chat format: peg-native
0.50.776.334 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.931 (> 0.100 thold), f_keep = 0.875
0.50.776.559 I reasoning-budget: activated, budget=2147483647 tokens
0.50.776.562 I reasoning-budget: deactivated (natural end)
0.50.776.609 I slot launch_slot_: id  0 | task 880 | processing task, is_child = 0
0.50.776.620 W slot update_slots: id  0 | task 880 | n_past = 405, slot.prompt.tokens.size() = 463, seq_id = 0, pos_min = 462, n_swa = 0
0.50.776.622 I slot update_slots: id  0 | task 880 | Checking checkpoint with [419, 419] against 405...
0.50.776.623 W slot update_slots: id  0 | task 880 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.50.776.627 W slot update_slots: id  0 | task 880 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.51.328.909 I slot create_check: id  0 | task 880 | created context checkpoint 1 of 32 (pos_min = 430, pos_max = 430, n_tokens = 431, size = 62.813 MiB)
0.52.962.266 I slot print_timing: id  0 | task 880 | 
prompt eval time =     596.05 ms /   435 tokens (    1.37 ms per token,   729.81 tokens per second)
       eval time =    1589.57 ms /    80 tokens (   19.87 ms per token,    50.33 tokens per second)
      total time =    2185.62 ms /   515 tokens
0.52.962.487 I slot      release: id  0 | task 880 | stop processing: n_tokens = 514, truncated = 0
0.52.962.523 I srv  update_slots: all slots are idle
0.52.963.931 I srv    operator(): operator(): cleaning up before exit...