File size: 28,916 Bytes
27a5002
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
0.00.120.044 W Setting 'enable_thinking' via --chat-template-kwargs is deprecated. Use --reasoning on / --reasoning off instead.
0.00.163.630 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
0.00.163.643 I device_info:
0.00.163.785 I   - ROCm0   : AMD Radeon Graphics (131072 MiB, 123858 MiB free)
0.00.163.974 I   - Vulkan0 : AMD Radeon Graphics (RADV GFX1151) (132096 MiB, 131922 MiB free)
0.00.163.982 I   - CPU     : AMD RYZEN AI MAX+ 395 w/ Radeon 8060S (127438 MiB, 127438 MiB free)
0.00.164.078 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 | 
0.00.164.142 I srv          init: running without SSL
0.00.164.166 I srv          init: using 31 threads for HTTP server
0.00.164.167 I srv          init: the WebUI is disabled
0.00.164.248 I srv         start: binding port with default address family
0.00.165.500 I srv          main: loading model
0.00.165.507 I srv    load_model: loading model '/mnt/models/nex-n2.5-mini/out/Nex-N2.5-mini-Q4_0_ROCMFP4_STRIX_LEAN.gguf'
0.00.216.012 W llama_model_loader: direct I/O is enabled, disabling mmap
0.22.544.117 W llama_context: n_ctx_seq (65536) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
0.22.849.861 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
0.23.278.480 I srv    load_model: initializing slots, n_slots = 1
0.23.567.593 W srv    load_model: speculative decoding will use checkpoints
0.23.567.602 W common_speculative_init: no implementations specified for speculative decoding
0.23.567.605 I slot   load_model: id  0 | task -1 | new slot, n_ctx = 65536
0.23.567.670 I srv    load_model: prompt cache RAM enabled: limit_mib=8192
0.23.567.688 I srv    load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
0.23.567.709 I srv          init: idle slots will be saved to prompt cache upon starting a new task
0.23.579.583 I init: chat template, example_format: '<|im_start|>system
You are a helpful assistant<|im_end|>
<|im_start|>user
Hello<|im_end|>
<|im_start|>assistant
<think>

</think>

Hi there<|im_end|>
<|im_start|>user
How are you?<|im_end|>
<|im_start|>assistant
<think>

</think>

'
0.23.587.209 I srv          init: init: chat template, thinking = 1
0.23.587.253 I srv          main: model loaded
0.23.587.256 I srv          main: server is listening on http://127.0.0.1:18600
0.23.587.260 I srv  update_slots: all slots are idle
0.24.570.949 I srv  params_from_: Chat format: peg-native
0.24.571.406 I slot get_availabl: id  0 | task -1 | selected slot by LRU, t_last = -1
0.24.571.411 I srv  get_availabl: updating prompt cache
0.24.571.419 I srv          load:  - looking for better prompt, base f_keep = -1.000, sim = 0.000
0.24.571.426 I srv        update:  - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 65536 tokens, 8589934592 est)
0.24.571.428 I srv  get_availabl: prompt cache update took 0.01 ms
0.24.571.805 I reasoning-budget: activated, budget=2147483647 tokens
0.24.571.827 I slot launch_slot_: id  0 | task 0 | processing task, is_child = 0
0.25.235.195 I slot create_check: id  0 | task 0 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
0.25.806.187 I reasoning-budget: deactivated (natural end)
0.26.650.056 I slot print_timing: id  0 | task 0 | 
prompt eval time =     715.70 ms /   422 tokens (    1.70 ms per token,   589.63 tokens per second)
       eval time =    1362.48 ms /    60 tokens (   22.71 ms per token,    44.04 tokens per second)
      total time =    2078.18 ms /   482 tokens
0.26.650.247 I slot      release: id  0 | task 0 | stop processing: n_tokens = 481, truncated = 0
0.26.650.264 I srv  update_slots: all slots are idle
0.26.675.130 I srv  params_from_: Chat format: peg-native
0.26.675.629 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.906 (> 0.100 thold), f_keep = 0.842
0.26.676.106 I reasoning-budget: activated, budget=2147483647 tokens
0.26.676.196 I slot launch_slot_: id  0 | task 62 | processing task, is_child = 0
0.26.676.216 W slot update_slots: id  0 | task 62 | n_past = 405, slot.prompt.tokens.size() = 481, seq_id = 0, pos_min = 480, n_swa = 0
0.26.676.222 I slot update_slots: id  0 | task 62 | Checking checkpoint with [417, 417] against 405...
0.26.676.224 W slot update_slots: id  0 | task 62 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.26.676.229 W slot update_slots: id  0 | task 62 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.27.231.790 I slot create_check: id  0 | task 62 | created context checkpoint 1 of 32 (pos_min = 442, pos_max = 442, n_tokens = 443, size = 62.813 MiB)
0.27.733.232 I reasoning-budget: deactivated (natural end)
0.28.618.139 I slot print_timing: id  0 | task 62 | 
prompt eval time =     593.45 ms /   447 tokens (    1.33 ms per token,   753.22 tokens per second)
       eval time =    1348.46 ms /    65 tokens (   20.75 ms per token,    48.20 tokens per second)
      total time =    1941.91 ms /   512 tokens
0.28.618.225 I slot      release: id  0 | task 62 | stop processing: n_tokens = 511, truncated = 0
0.28.618.260 I srv  update_slots: all slots are idle
0.28.632.505 I srv  params_from_: Chat format: peg-native
0.28.632.871 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.953 (> 0.100 thold), f_keep = 0.793
0.28.633.063 I reasoning-budget: activated, budget=2147483647 tokens
0.28.633.092 I slot launch_slot_: id  0 | task 129 | processing task, is_child = 0
0.28.633.101 W slot update_slots: id  0 | task 129 | n_past = 405, slot.prompt.tokens.size() = 511, seq_id = 0, pos_min = 510, n_swa = 0
0.28.633.102 I slot update_slots: id  0 | task 129 | Checking checkpoint with [442, 442] against 405...
0.28.633.103 W slot update_slots: id  0 | task 129 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.28.633.104 W slot update_slots: id  0 | task 129 | erased invalidated context checkpoint (pos_min = 442, pos_max = 442, n_tokens = 443, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.29.161.759 I slot create_check: id  0 | task 129 | created context checkpoint 1 of 32 (pos_min = 420, pos_max = 420, n_tokens = 421, size = 62.813 MiB)
0.29.478.654 I reasoning-budget: deactivated (natural end)
0.30.331.330 I slot print_timing: id  0 | task 129 | 
prompt eval time =     580.01 ms /   425 tokens (    1.36 ms per token,   732.75 tokens per second)
       eval time =    1118.18 ms /    52 tokens (   21.50 ms per token,    46.50 tokens per second)
      total time =    1698.19 ms /   477 tokens
0.30.331.521 I slot      release: id  0 | task 129 | stop processing: n_tokens = 476, truncated = 0
0.30.331.586 I srv  update_slots: all slots are idle
0.30.386.522 I srv  params_from_: Chat format: peg-native
0.30.388.735 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.953 (> 0.100 thold), f_keep = 0.851
0.30.389.319 I reasoning-budget: activated, budget=2147483647 tokens
0.30.389.437 I slot launch_slot_: id  0 | task 183 | processing task, is_child = 0
0.30.389.465 W slot update_slots: id  0 | task 183 | n_past = 405, slot.prompt.tokens.size() = 476, seq_id = 0, pos_min = 475, n_swa = 0
0.30.389.469 I slot update_slots: id  0 | task 183 | Checking checkpoint with [420, 420] against 405...
0.30.389.471 W slot update_slots: id  0 | task 183 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.30.389.478 W slot update_slots: id  0 | task 183 | erased invalidated context checkpoint (pos_min = 420, pos_max = 420, n_tokens = 421, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.30.972.513 I slot create_check: id  0 | task 183 | created context checkpoint 1 of 32 (pos_min = 420, pos_max = 420, n_tokens = 421, size = 62.813 MiB)
0.31.338.635 I reasoning-budget: deactivated (natural end)
0.31.461.621 I slot print_timing: id  0 | task 183 | 
prompt eval time =     630.63 ms /   425 tokens (    1.48 ms per token,   673.92 tokens per second)
       eval time =     441.50 ms /    17 tokens (   25.97 ms per token,    38.51 tokens per second)
      total time =    1072.13 ms /   442 tokens
0.31.461.844 I slot      release: id  0 | task 183 | stop processing: n_tokens = 441, truncated = 0
0.31.461.907 I srv  update_slots: all slots are idle
0.31.524.318 I srv  params_from_: Chat format: peg-native
0.31.526.203 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.921
0.31.526.820 I reasoning-budget: activated, budget=2147483647 tokens
0.31.526.906 I slot launch_slot_: id  0 | task 202 | processing task, is_child = 0
0.31.526.929 W slot update_slots: id  0 | task 202 | n_past = 406, slot.prompt.tokens.size() = 441, seq_id = 0, pos_min = 440, n_swa = 0
0.31.526.933 I slot update_slots: id  0 | task 202 | Checking checkpoint with [420, 420] against 406...
0.31.526.935 W slot update_slots: id  0 | task 202 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.31.526.940 W slot update_slots: id  0 | task 202 | erased invalidated context checkpoint (pos_min = 420, pos_max = 420, n_tokens = 421, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.32.129.825 I slot create_check: id  0 | task 202 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
0.32.519.469 I reasoning-budget: deactivated (natural end)
0.33.361.317 I slot print_timing: id  0 | task 202 | 
prompt eval time =     645.36 ms /   422 tokens (    1.53 ms per token,   653.90 tokens per second)
       eval time =    1189.02 ms /    55 tokens (   21.62 ms per token,    46.26 tokens per second)
      total time =    1834.38 ms /   477 tokens
0.33.361.399 I slot      release: id  0 | task 202 | stop processing: n_tokens = 476, truncated = 0
0.33.361.431 I srv  update_slots: all slots are idle
0.33.406.059 I srv  params_from_: Chat format: peg-native
0.33.406.536 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.854 (> 0.100 thold), f_keep = 0.884
0.33.406.842 I reasoning-budget: activated, budget=2147483647 tokens
0.33.406.887 I slot launch_slot_: id  0 | task 259 | processing task, is_child = 0
0.33.406.903 W slot update_slots: id  0 | task 259 | n_past = 421, slot.prompt.tokens.size() = 476, seq_id = 0, pos_min = 475, n_swa = 0
0.33.406.904 I slot update_slots: id  0 | task 259 | Checking checkpoint with [417, 417] against 421...
0.33.410.865 W slot update_slots: id  0 | task 259 | restored context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_past = 418, size = 62.813 MiB)
0.33.668.081 I slot create_check: id  0 | task 259 | created context checkpoint 2 of 32 (pos_min = 488, pos_max = 488, n_tokens = 489, size = 62.813 MiB)
0.34.301.983 I reasoning-budget: deactivated (natural end)
0.34.590.965 I slot print_timing: id  0 | task 259 | 
prompt eval time =     311.59 ms /    75 tokens (    4.15 ms per token,   240.70 tokens per second)
       eval time =     872.45 ms /    38 tokens (   22.96 ms per token,    43.56 tokens per second)
      total time =    1184.04 ms /   113 tokens
0.34.591.058 I slot      release: id  0 | task 259 | stop processing: n_tokens = 530, truncated = 0
0.34.591.093 I srv  update_slots: all slots are idle
0.34.626.053 I srv  params_from_: Chat format: peg-native
0.34.626.459 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.972 (> 0.100 thold), f_keep = 0.774
0.34.626.655 I reasoning-budget: activated, budget=2147483647 tokens
0.34.626.694 I slot launch_slot_: id  0 | task 299 | processing task, is_child = 0
0.34.626.705 W slot update_slots: id  0 | task 299 | n_past = 410, slot.prompt.tokens.size() = 530, seq_id = 0, pos_min = 529, n_swa = 0
0.34.626.705 I slot update_slots: id  0 | task 299 | Checking checkpoint with [488, 488] against 410...
0.34.626.706 I slot update_slots: id  0 | task 299 | Checking checkpoint with [417, 417] against 410...
0.34.626.706 W slot update_slots: id  0 | task 299 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.34.626.709 W slot update_slots: id  0 | task 299 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.34.627.674 W slot update_slots: id  0 | task 299 | erased invalidated context checkpoint (pos_min = 488, pos_max = 488, n_tokens = 489, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.35.160.501 I slot create_check: id  0 | task 299 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
0.35.462.626 I reasoning-budget: deactivated (natural end)
0.36.334.370 I slot print_timing: id  0 | task 299 | 
prompt eval time =     570.77 ms /   422 tokens (    1.35 ms per token,   739.35 tokens per second)
       eval time =    1136.85 ms /    53 tokens (   21.45 ms per token,    46.62 tokens per second)
      total time =    1707.63 ms /   475 tokens
0.36.334.535 I slot      release: id  0 | task 299 | stop processing: n_tokens = 474, truncated = 0
0.36.334.587 I srv  update_slots: all slots are idle
0.36.346.401 I srv  params_from_: Chat format: peg-native
0.36.346.871 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.935 (> 0.100 thold), f_keep = 0.854
0.36.347.275 I reasoning-budget: activated, budget=2147483647 tokens
0.36.347.355 I slot launch_slot_: id  0 | task 354 | processing task, is_child = 0
0.36.347.371 W slot update_slots: id  0 | task 354 | n_past = 405, slot.prompt.tokens.size() = 474, seq_id = 0, pos_min = 473, n_swa = 0
0.36.347.374 I slot update_slots: id  0 | task 354 | Checking checkpoint with [417, 417] against 405...
0.36.347.376 W slot update_slots: id  0 | task 354 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.36.347.381 W slot update_slots: id  0 | task 354 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.36.950.505 I slot create_check: id  0 | task 354 | created context checkpoint 1 of 32 (pos_min = 428, pos_max = 428, n_tokens = 429, size = 62.813 MiB)
0.37.901.136 I reasoning-budget: deactivated (natural end)
0.39.133.735 I slot print_timing: id  0 | task 354 | n_decoded =    100, tg =  46.90 t/s
0.39.607.166 I slot print_timing: id  0 | task 354 | 
prompt eval time =     653.96 ms /   433 tokens (    1.51 ms per token,   662.12 tokens per second)
       eval time =    2605.80 ms /   123 tokens (   21.19 ms per token,    47.20 tokens per second)
      total time =    3259.76 ms /   556 tokens
0.39.607.372 I slot      release: id  0 | task 354 | stop processing: n_tokens = 555, truncated = 0
0.39.607.435 I srv  update_slots: all slots are idle
0.39.655.799 I srv  params_from_: Chat format: peg-native
0.39.656.354 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.730
0.39.656.713 I reasoning-budget: activated, budget=2147483647 tokens
0.39.656.717 I reasoning-budget: deactivated (natural end)
0.39.656.803 I slot launch_slot_: id  0 | task 479 | processing task, is_child = 0
0.39.656.826 W slot update_slots: id  0 | task 479 | n_past = 405, slot.prompt.tokens.size() = 555, seq_id = 0, pos_min = 554, n_swa = 0
0.39.656.828 I slot update_slots: id  0 | task 479 | Checking checkpoint with [428, 428] against 405...
0.39.656.830 W slot update_slots: id  0 | task 479 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.39.656.835 W slot update_slots: id  0 | task 479 | erased invalidated context checkpoint (pos_min = 428, pos_max = 428, n_tokens = 429, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.40.205.076 I slot create_check: id  0 | task 479 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
0.41.065.168 I slot print_timing: id  0 | task 479 | 
prompt eval time =     582.67 ms /   424 tokens (    1.37 ms per token,   727.68 tokens per second)
       eval time =     825.66 ms /    39 tokens (   21.17 ms per token,    47.23 tokens per second)
      total time =    1408.33 ms /   463 tokens
0.41.065.250 I slot      release: id  0 | task 479 | stop processing: n_tokens = 462, truncated = 0
0.41.065.277 I srv  update_slots: all slots are idle
0.41.077.853 I srv  params_from_: Chat format: peg-native
0.41.078.228 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.902 (> 0.100 thold), f_keep = 0.877
0.41.078.482 I reasoning-budget: activated, budget=2147483647 tokens
0.41.078.484 I reasoning-budget: deactivated (natural end)
0.41.078.524 I slot launch_slot_: id  0 | task 520 | processing task, is_child = 0
0.41.078.534 W slot update_slots: id  0 | task 520 | n_past = 405, slot.prompt.tokens.size() = 462, seq_id = 0, pos_min = 461, n_swa = 0
0.41.078.535 I slot update_slots: id  0 | task 520 | Checking checkpoint with [419, 419] against 405...
0.41.078.536 W slot update_slots: id  0 | task 520 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.41.078.537 W slot update_slots: id  0 | task 520 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.41.589.485 I slot create_check: id  0 | task 520 | created context checkpoint 1 of 32 (pos_min = 444, pos_max = 444, n_tokens = 445, size = 62.813 MiB)
0.43.355.441 I slot print_timing: id  0 | task 520 | 
prompt eval time =     549.58 ms /   449 tokens (    1.22 ms per token,   816.99 tokens per second)
       eval time =    1727.31 ms /    86 tokens (   20.09 ms per token,    49.79 tokens per second)
      total time =    2276.89 ms /   535 tokens
0.43.355.528 I slot      release: id  0 | task 520 | stop processing: n_tokens = 534, truncated = 0
0.43.355.559 I srv  update_slots: all slots are idle
0.43.407.484 I srv  params_from_: Chat format: peg-native
0.43.408.099 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.948 (> 0.100 thold), f_keep = 0.758
0.43.408.372 I reasoning-budget: activated, budget=2147483647 tokens
0.43.408.374 I reasoning-budget: deactivated (natural end)
0.43.408.422 I slot launch_slot_: id  0 | task 608 | processing task, is_child = 0
0.43.408.434 W slot update_slots: id  0 | task 608 | n_past = 405, slot.prompt.tokens.size() = 534, seq_id = 0, pos_min = 533, n_swa = 0
0.43.408.435 I slot update_slots: id  0 | task 608 | Checking checkpoint with [444, 444] against 405...
0.43.408.437 W slot update_slots: id  0 | task 608 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.43.408.440 W slot update_slots: id  0 | task 608 | erased invalidated context checkpoint (pos_min = 444, pos_max = 444, n_tokens = 445, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.43.970.219 I slot create_check: id  0 | task 608 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
0.44.791.545 I slot print_timing: id  0 | task 608 | 
prompt eval time =     601.11 ms /   427 tokens (    1.41 ms per token,   710.35 tokens per second)
       eval time =     781.99 ms /    39 tokens (   20.05 ms per token,    49.87 tokens per second)
      total time =    1383.10 ms /   466 tokens
0.44.791.624 I slot      release: id  0 | task 608 | stop processing: n_tokens = 465, truncated = 0
0.44.791.654 I srv  update_slots: all slots are idle
0.44.833.558 I srv  params_from_: Chat format: peg-native
0.44.834.049 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.948 (> 0.100 thold), f_keep = 0.871
0.44.834.271 I reasoning-budget: activated, budget=2147483647 tokens
0.44.834.272 I reasoning-budget: deactivated (natural end)
0.44.834.315 I slot launch_slot_: id  0 | task 649 | processing task, is_child = 0
0.44.834.326 W slot update_slots: id  0 | task 649 | n_past = 405, slot.prompt.tokens.size() = 465, seq_id = 0, pos_min = 464, n_swa = 0
0.44.834.327 I slot update_slots: id  0 | task 649 | Checking checkpoint with [422, 422] against 405...
0.44.834.328 W slot update_slots: id  0 | task 649 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.44.834.332 W slot update_slots: id  0 | task 649 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.45.316.040 I slot create_check: id  0 | task 649 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
0.45.454.550 I slot print_timing: id  0 | task 649 | 
prompt eval time =     522.74 ms /   427 tokens (    1.22 ms per token,   816.85 tokens per second)
       eval time =      97.47 ms /     4 tokens (   24.37 ms per token,    41.04 tokens per second)
      total time =     620.21 ms /   431 tokens
0.45.454.641 I slot      release: id  0 | task 649 | stop processing: n_tokens = 430, truncated = 0
0.45.454.670 I srv  update_slots: all slots are idle
0.45.491.620 I srv  params_from_: Chat format: peg-native
0.45.492.052 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.958 (> 0.100 thold), f_keep = 0.944
0.45.492.301 I reasoning-budget: activated, budget=2147483647 tokens
0.45.492.303 I reasoning-budget: deactivated (natural end)
0.45.492.343 I slot launch_slot_: id  0 | task 655 | processing task, is_child = 0
0.45.492.355 W slot update_slots: id  0 | task 655 | n_past = 406, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
0.45.492.357 I slot update_slots: id  0 | task 655 | Checking checkpoint with [422, 422] against 406...
0.45.492.358 W slot update_slots: id  0 | task 655 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.45.492.362 W slot update_slots: id  0 | task 655 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.46.038.260 I slot create_check: id  0 | task 655 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
0.46.950.011 I slot print_timing: id  0 | task 655 | 
prompt eval time =     592.07 ms /   424 tokens (    1.40 ms per token,   716.14 tokens per second)
       eval time =     865.56 ms /    40 tokens (   21.64 ms per token,    46.21 tokens per second)
      total time =    1457.62 ms /   464 tokens
0.46.950.215 I slot      release: id  0 | task 655 | stop processing: n_tokens = 463, truncated = 0
0.46.950.270 I srv  update_slots: all slots are idle
0.46.968.851 I srv  params_from_: Chat format: peg-native
0.46.969.311 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.935 (> 0.100 thold), f_keep = 1.000
0.46.969.946 I reasoning-budget: activated, budget=2147483647 tokens
0.46.969.951 I reasoning-budget: deactivated (natural end)
0.46.970.047 I slot launch_slot_: id  0 | task 697 | processing task, is_child = 0
0.47.121.064 I slot create_check: id  0 | task 697 | created context checkpoint 2 of 32 (pos_min = 490, pos_max = 490, n_tokens = 491, size = 62.813 MiB)
0.47.570.159 I slot print_timing: id  0 | task 697 | 
prompt eval time =     211.63 ms /    32 tokens (    6.61 ms per token,   151.21 tokens per second)
       eval time =     388.43 ms /    15 tokens (   25.90 ms per token,    38.62 tokens per second)
      total time =     600.06 ms /    47 tokens
0.47.570.388 I slot      release: id  0 | task 697 | stop processing: n_tokens = 509, truncated = 0
0.47.570.448 I srv  update_slots: all slots are idle
0.47.607.526 I srv  params_from_: Chat format: peg-native
0.47.608.040 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.967 (> 0.100 thold), f_keep = 0.806
0.47.608.345 I reasoning-budget: activated, budget=2147483647 tokens
0.47.608.347 I reasoning-budget: deactivated (natural end)
0.47.608.398 I slot launch_slot_: id  0 | task 714 | processing task, is_child = 0
0.47.608.410 W slot update_slots: id  0 | task 714 | n_past = 410, slot.prompt.tokens.size() = 509, seq_id = 0, pos_min = 508, n_swa = 0
0.47.608.412 I slot update_slots: id  0 | task 714 | Checking checkpoint with [490, 490] against 410...
0.47.608.412 I slot update_slots: id  0 | task 714 | Checking checkpoint with [419, 419] against 410...
0.47.608.413 W slot update_slots: id  0 | task 714 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.47.608.416 W slot update_slots: id  0 | task 714 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.47.609.303 W slot update_slots: id  0 | task 714 | erased invalidated context checkpoint (pos_min = 490, pos_max = 490, n_tokens = 491, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.48.271.002 I slot create_check: id  0 | task 714 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
0.49.108.254 I slot print_timing: id  0 | task 714 | 
prompt eval time =     731.88 ms /   424 tokens (    1.73 ms per token,   579.33 tokens per second)
       eval time =     767.92 ms /    40 tokens (   19.20 ms per token,    52.09 tokens per second)
      total time =    1499.80 ms /   464 tokens
0.49.108.437 I slot      release: id  0 | task 714 | stop processing: n_tokens = 463, truncated = 0
0.49.108.493 I srv  update_slots: all slots are idle
0.49.161.163 I srv  params_from_: Chat format: peg-native
0.49.163.186 I slot get_availabl: id  0 | task -1 | selected slot by LCP similarity, sim_best = 0.931 (> 0.100 thold), f_keep = 0.875
0.49.163.878 I reasoning-budget: activated, budget=2147483647 tokens
0.49.163.884 I reasoning-budget: deactivated (natural end)
0.49.163.990 I slot launch_slot_: id  0 | task 756 | processing task, is_child = 0
0.49.164.016 W slot update_slots: id  0 | task 756 | n_past = 405, slot.prompt.tokens.size() = 463, seq_id = 0, pos_min = 462, n_swa = 0
0.49.164.018 I slot update_slots: id  0 | task 756 | Checking checkpoint with [419, 419] against 405...
0.49.164.020 W slot update_slots: id  0 | task 756 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.49.164.027 W slot update_slots: id  0 | task 756 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.49.744.671 I slot create_check: id  0 | task 756 | created context checkpoint 1 of 32 (pos_min = 430, pos_max = 430, n_tokens = 431, size = 62.813 MiB)
0.51.381.802 I slot print_timing: id  0 | task 756 | 
prompt eval time =     632.00 ms /   435 tokens (    1.45 ms per token,   688.30 tokens per second)
       eval time =    1585.77 ms /    80 tokens (   19.82 ms per token,    50.45 tokens per second)
      total time =    2217.77 ms /   515 tokens
0.51.382.017 I slot      release: id  0 | task 756 | stop processing: n_tokens = 514, truncated = 0
0.51.382.050 I srv  update_slots: all slots are idle
0.51.383.517 I srv    operator(): operator(): cleaning up before exit...