File size: 28,875 Bytes
27a5002 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 | 0.00.114.439 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
0.00.114.444 I device_info:
0.00.114.555 I - ROCm0 : AMD Radeon Graphics (131072 MiB, 123723 MiB free)
0.00.114.708 I - Vulkan0 : AMD Radeon Graphics (RADV GFX1151) (132096 MiB, 131922 MiB free)
0.00.114.715 I - CPU : AMD RYZEN AI MAX+ 395 w/ Radeon 8060S (127438 MiB, 127438 MiB free)
0.00.114.796 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
0.00.114.847 I srv init: running without SSL
0.00.114.871 I srv init: using 31 threads for HTTP server
0.00.114.873 I srv init: the WebUI is disabled
0.00.114.942 I srv start: binding port with default address family
0.00.116.175 I srv main: loading model
0.00.116.182 I srv load_model: loading model '/mnt/models/nex-n2.5-mini/out/Nex-N2.5-mini-Q4_0_ROCMFP4_STRIX_LEAN.gguf'
0.00.160.730 W llama_model_loader: direct I/O is enabled, disabling mmap
0.21.513.834 W llama_context: n_ctx_seq (65536) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
0.21.793.651 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
0.22.217.020 I srv load_model: initializing slots, n_slots = 1
0.22.487.377 W srv load_model: speculative decoding will use checkpoints
0.22.487.387 W common_speculative_init: no implementations specified for speculative decoding
0.22.487.390 I slot load_model: id 0 | task -1 | new slot, n_ctx = 65536
0.22.487.471 I srv load_model: prompt cache RAM enabled: limit_mib=8192
0.22.487.473 I srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
0.22.487.495 I srv init: idle slots will be saved to prompt cache upon starting a new task
0.22.540.918 I init: chat template, example_format: '<|im_start|>system
You are a helpful assistant<|im_end|>
<|im_start|>user
Hello<|im_end|>
<|im_start|>assistant
<think>
</think>
Hi there<|im_end|>
<|im_start|>user
How are you?<|im_end|>
<|im_start|>assistant
<think>
</think>
'
0.22.582.904 I srv init: init: chat template, thinking = 0
0.22.582.984 I srv main: model loaded
0.22.582.995 I srv main: server is listening on http://127.0.0.1:18600
0.22.583.003 I srv update_slots: all slots are idle
0.23.994.344 I srv params_from_: Chat format: peg-native
0.23.995.890 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1
0.23.995.893 I srv get_availabl: updating prompt cache
0.23.995.900 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000
0.23.995.907 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 65536 tokens, 8589934592 est)
0.23.995.909 I srv get_availabl: prompt cache update took 0.01 ms
0.23.996.222 I reasoning-budget: activated, budget=2147483647 tokens
0.23.996.243 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0
0.24.656.206 I slot create_check: id 0 | task 0 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
0.25.028.892 I reasoning-budget: deactivated (natural end)
0.25.816.313 I slot print_timing: id 0 | task 0 |
prompt eval time = 698.88 ms / 422 tokens ( 1.66 ms per token, 603.83 tokens per second)
eval time = 1121.16 ms / 54 tokens ( 20.76 ms per token, 48.16 tokens per second)
total time = 1820.04 ms / 476 tokens
0.25.816.410 I slot release: id 0 | task 0 | stop processing: n_tokens = 475, truncated = 0
0.25.816.424 I srv update_slots: all slots are idle
0.25.830.282 I srv params_from_: Chat format: peg-native
0.25.830.661 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.906 (> 0.100 thold), f_keep = 0.853
0.25.830.794 I reasoning-budget: activated, budget=2147483647 tokens
0.25.830.832 I slot launch_slot_: id 0 | task 56 | processing task, is_child = 0
0.25.830.841 W slot update_slots: id 0 | task 56 | n_past = 405, slot.prompt.tokens.size() = 475, seq_id = 0, pos_min = 474, n_swa = 0
0.25.830.842 I slot update_slots: id 0 | task 56 | Checking checkpoint with [417, 417] against 405...
0.25.830.844 W slot update_slots: id 0 | task 56 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.25.830.846 W slot update_slots: id 0 | task 56 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.26.482.806 I slot create_check: id 0 | task 56 | created context checkpoint 1 of 32 (pos_min = 442, pos_max = 442, n_tokens = 443, size = 62.813 MiB)
0.26.903.330 I reasoning-budget: deactivated (natural end)
0.28.669.262 I slot print_timing: id 0 | task 56 | n_decoded = 100, tg = 47.24 t/s
0.28.729.150 I slot print_timing: id 0 | task 56 |
prompt eval time = 721.52 ms / 447 tokens ( 1.61 ms per token, 619.52 tokens per second)
eval time = 2176.78 ms / 103 tokens ( 21.13 ms per token, 47.32 tokens per second)
total time = 2898.30 ms / 550 tokens
0.28.729.226 I slot release: id 0 | task 56 | stop processing: n_tokens = 549, truncated = 0
0.28.729.266 I srv update_slots: all slots are idle
0.28.742.730 I srv params_from_: Chat format: peg-native
0.28.743.168 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.953 (> 0.100 thold), f_keep = 0.738
0.28.743.356 I reasoning-budget: activated, budget=2147483647 tokens
0.28.743.392 I slot launch_slot_: id 0 | task 161 | processing task, is_child = 0
0.28.743.406 W slot update_slots: id 0 | task 161 | n_past = 405, slot.prompt.tokens.size() = 549, seq_id = 0, pos_min = 548, n_swa = 0
0.28.743.407 I slot update_slots: id 0 | task 161 | Checking checkpoint with [442, 442] against 405...
0.28.743.407 W slot update_slots: id 0 | task 161 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.28.743.411 W slot update_slots: id 0 | task 161 | erased invalidated context checkpoint (pos_min = 442, pos_max = 442, n_tokens = 443, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.29.274.427 I slot create_check: id 0 | task 161 | created context checkpoint 1 of 32 (pos_min = 420, pos_max = 420, n_tokens = 421, size = 62.813 MiB)
0.29.681.788 I reasoning-budget: deactivated (natural end)
0.30.465.485 I slot print_timing: id 0 | task 161 |
prompt eval time = 568.14 ms / 425 tokens ( 1.34 ms per token, 748.05 tokens per second)
eval time = 1153.92 ms / 57 tokens ( 20.24 ms per token, 49.40 tokens per second)
total time = 1722.06 ms / 482 tokens
0.30.465.571 I slot release: id 0 | task 161 | stop processing: n_tokens = 481, truncated = 0
0.30.465.599 I srv update_slots: all slots are idle
0.30.480.783 I srv params_from_: Chat format: peg-native
0.30.481.142 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.953 (> 0.100 thold), f_keep = 0.842
0.30.481.346 I reasoning-budget: activated, budget=2147483647 tokens
0.30.481.386 I slot launch_slot_: id 0 | task 220 | processing task, is_child = 0
0.30.481.397 W slot update_slots: id 0 | task 220 | n_past = 405, slot.prompt.tokens.size() = 481, seq_id = 0, pos_min = 480, n_swa = 0
0.30.481.397 I slot update_slots: id 0 | task 220 | Checking checkpoint with [420, 420] against 405...
0.30.481.399 W slot update_slots: id 0 | task 220 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.30.481.401 W slot update_slots: id 0 | task 220 | erased invalidated context checkpoint (pos_min = 420, pos_max = 420, n_tokens = 421, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.31.008.867 I slot create_check: id 0 | task 220 | created context checkpoint 1 of 32 (pos_min = 420, pos_max = 420, n_tokens = 421, size = 62.813 MiB)
0.31.355.463 I reasoning-budget: deactivated (natural end)
0.31.458.677 I slot print_timing: id 0 | task 220 |
prompt eval time = 568.77 ms / 425 tokens ( 1.34 ms per token, 747.23 tokens per second)
eval time = 408.50 ms / 20 tokens ( 20.42 ms per token, 48.96 tokens per second)
total time = 977.27 ms / 445 tokens
0.31.458.774 I slot release: id 0 | task 220 | stop processing: n_tokens = 444, truncated = 0
0.31.458.805 I srv update_slots: all slots are idle
0.31.510.659 I srv params_from_: Chat format: peg-native
0.31.511.238 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.914
0.31.511.571 I reasoning-budget: activated, budget=2147483647 tokens
0.31.511.663 I slot launch_slot_: id 0 | task 242 | processing task, is_child = 0
0.31.511.677 W slot update_slots: id 0 | task 242 | n_past = 406, slot.prompt.tokens.size() = 444, seq_id = 0, pos_min = 443, n_swa = 0
0.31.511.684 I slot update_slots: id 0 | task 242 | Checking checkpoint with [420, 420] against 406...
0.31.511.685 W slot update_slots: id 0 | task 242 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.31.511.690 W slot update_slots: id 0 | task 242 | erased invalidated context checkpoint (pos_min = 420, pos_max = 420, n_tokens = 421, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.32.149.342 I slot create_check: id 0 | task 242 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
0.32.405.933 I reasoning-budget: deactivated (natural end)
0.33.283.206 I slot print_timing: id 0 | task 242 |
prompt eval time = 677.52 ms / 422 tokens ( 1.61 ms per token, 622.86 tokens per second)
eval time = 1093.99 ms / 52 tokens ( 21.04 ms per token, 47.53 tokens per second)
total time = 1771.51 ms / 474 tokens
0.33.283.283 I slot release: id 0 | task 242 | stop processing: n_tokens = 473, truncated = 0
0.33.283.309 I srv update_slots: all slots are idle
0.33.311.488 I srv params_from_: Chat format: peg-native
0.33.312.015 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.854 (> 0.100 thold), f_keep = 0.890
0.33.312.441 I reasoning-budget: activated, budget=2147483647 tokens
0.33.312.556 I slot launch_slot_: id 0 | task 296 | processing task, is_child = 0
0.33.312.580 W slot update_slots: id 0 | task 296 | n_past = 421, slot.prompt.tokens.size() = 473, seq_id = 0, pos_min = 472, n_swa = 0
0.33.312.583 I slot update_slots: id 0 | task 296 | Checking checkpoint with [417, 417] against 421...
0.33.321.205 W slot update_slots: id 0 | task 296 | restored context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_past = 418, size = 62.813 MiB)
0.33.558.612 I slot create_check: id 0 | task 296 | created context checkpoint 2 of 32 (pos_min = 488, pos_max = 488, n_tokens = 489, size = 62.813 MiB)
0.34.218.104 I reasoning-budget: deactivated (natural end)
0.34.556.664 I slot print_timing: id 0 | task 296 |
prompt eval time = 286.32 ms / 75 tokens ( 3.82 ms per token, 261.95 tokens per second)
eval time = 957.75 ms / 41 tokens ( 23.36 ms per token, 42.81 tokens per second)
total time = 1244.07 ms / 116 tokens
0.34.556.763 I slot release: id 0 | task 296 | stop processing: n_tokens = 533, truncated = 0
0.34.556.805 I srv update_slots: all slots are idle
0.34.571.001 I srv params_from_: Chat format: peg-native
0.34.571.465 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.972 (> 0.100 thold), f_keep = 0.769
0.34.571.696 I reasoning-budget: activated, budget=2147483647 tokens
0.34.571.740 I slot launch_slot_: id 0 | task 339 | processing task, is_child = 0
0.34.571.755 W slot update_slots: id 0 | task 339 | n_past = 410, slot.prompt.tokens.size() = 533, seq_id = 0, pos_min = 532, n_swa = 0
0.34.571.756 I slot update_slots: id 0 | task 339 | Checking checkpoint with [488, 488] against 410...
0.34.571.758 I slot update_slots: id 0 | task 339 | Checking checkpoint with [417, 417] against 410...
0.34.571.759 W slot update_slots: id 0 | task 339 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.34.571.763 W slot update_slots: id 0 | task 339 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.34.573.311 W slot update_slots: id 0 | task 339 | erased invalidated context checkpoint (pos_min = 488, pos_max = 488, n_tokens = 489, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.35.110.746 I slot create_check: id 0 | task 339 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
0.35.584.142 I reasoning-budget: deactivated (natural end)
0.36.392.408 I slot print_timing: id 0 | task 339 |
prompt eval time = 576.37 ms / 422 tokens ( 1.37 ms per token, 732.16 tokens per second)
eval time = 1244.26 ms / 62 tokens ( 20.07 ms per token, 49.83 tokens per second)
total time = 1820.64 ms / 484 tokens
0.36.392.474 I slot release: id 0 | task 339 | stop processing: n_tokens = 483, truncated = 0
0.36.392.500 I srv update_slots: all slots are idle
0.36.404.680 I srv params_from_: Chat format: peg-native
0.36.405.094 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.935 (> 0.100 thold), f_keep = 0.839
0.36.405.518 I reasoning-budget: activated, budget=2147483647 tokens
0.36.405.599 I slot launch_slot_: id 0 | task 403 | processing task, is_child = 0
0.36.405.620 W slot update_slots: id 0 | task 403 | n_past = 405, slot.prompt.tokens.size() = 483, seq_id = 0, pos_min = 482, n_swa = 0
0.36.405.621 I slot update_slots: id 0 | task 403 | Checking checkpoint with [417, 417] against 405...
0.36.405.625 W slot update_slots: id 0 | task 403 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.36.405.630 W slot update_slots: id 0 | task 403 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.36.974.524 I slot create_check: id 0 | task 403 | created context checkpoint 1 of 32 (pos_min = 428, pos_max = 428, n_tokens = 429, size = 62.813 MiB)
0.39.070.997 I slot print_timing: id 0 | task 403 | n_decoded = 100, tg = 48.88 t/s
0.39.448.467 I reasoning-budget: deactivated (natural end)
0.41.042.559 I slot print_timing: id 0 | task 403 |
prompt eval time = 619.66 ms / 433 tokens ( 1.43 ms per token, 698.77 tokens per second)
eval time = 4017.27 ms / 200 tokens ( 20.09 ms per token, 49.79 tokens per second)
total time = 4636.93 ms / 633 tokens
0.41.042.634 I slot release: id 0 | task 403 | stop processing: n_tokens = 632, truncated = 0
0.41.042.665 I srv update_slots: all slots are idle
0.41.054.733 I srv params_from_: Chat format: peg-native
0.41.055.120 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.641
0.41.055.469 I reasoning-budget: activated, budget=2147483647 tokens
0.41.055.471 I reasoning-budget: deactivated (natural end)
0.41.055.514 I slot launch_slot_: id 0 | task 605 | processing task, is_child = 0
0.41.055.527 W slot update_slots: id 0 | task 605 | n_past = 405, slot.prompt.tokens.size() = 632, seq_id = 0, pos_min = 631, n_swa = 0
0.41.055.529 I slot update_slots: id 0 | task 605 | Checking checkpoint with [428, 428] against 405...
0.41.055.530 W slot update_slots: id 0 | task 605 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.41.055.533 W slot update_slots: id 0 | task 605 | erased invalidated context checkpoint (pos_min = 428, pos_max = 428, n_tokens = 429, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.41.597.061 I slot create_check: id 0 | task 605 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
0.42.459.956 I slot print_timing: id 0 | task 605 |
prompt eval time = 582.62 ms / 424 tokens ( 1.37 ms per token, 727.74 tokens per second)
eval time = 821.78 ms / 39 tokens ( 21.07 ms per token, 47.46 tokens per second)
total time = 1404.41 ms / 463 tokens
0.42.460.047 I slot release: id 0 | task 605 | stop processing: n_tokens = 462, truncated = 0
0.42.460.077 I srv update_slots: all slots are idle
0.42.475.720 I srv params_from_: Chat format: peg-native
0.42.476.233 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.902 (> 0.100 thold), f_keep = 0.877
0.42.476.486 I reasoning-budget: activated, budget=2147483647 tokens
0.42.476.491 I reasoning-budget: deactivated (natural end)
0.42.476.536 I slot launch_slot_: id 0 | task 646 | processing task, is_child = 0
0.42.476.550 W slot update_slots: id 0 | task 646 | n_past = 405, slot.prompt.tokens.size() = 462, seq_id = 0, pos_min = 461, n_swa = 0
0.42.476.551 I slot update_slots: id 0 | task 646 | Checking checkpoint with [419, 419] against 405...
0.42.476.553 W slot update_slots: id 0 | task 646 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.42.476.571 W slot update_slots: id 0 | task 646 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.43.126.158 I slot create_check: id 0 | task 646 | created context checkpoint 1 of 32 (pos_min = 444, pos_max = 444, n_tokens = 445, size = 62.813 MiB)
0.44.988.511 I slot print_timing: id 0 | task 646 |
prompt eval time = 693.47 ms / 449 tokens ( 1.54 ms per token, 647.47 tokens per second)
eval time = 1818.46 ms / 86 tokens ( 21.14 ms per token, 47.29 tokens per second)
total time = 2511.93 ms / 535 tokens
0.44.988.703 I slot release: id 0 | task 646 | stop processing: n_tokens = 534, truncated = 0
0.44.988.760 I srv update_slots: all slots are idle
0.45.002.229 I srv params_from_: Chat format: peg-native
0.45.002.631 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.948 (> 0.100 thold), f_keep = 0.758
0.45.003.110 I reasoning-budget: activated, budget=2147483647 tokens
0.45.003.114 I reasoning-budget: deactivated (natural end)
0.45.003.190 I slot launch_slot_: id 0 | task 734 | processing task, is_child = 0
0.45.003.216 W slot update_slots: id 0 | task 734 | n_past = 405, slot.prompt.tokens.size() = 534, seq_id = 0, pos_min = 533, n_swa = 0
0.45.003.220 I slot update_slots: id 0 | task 734 | Checking checkpoint with [444, 444] against 405...
0.45.003.221 W slot update_slots: id 0 | task 734 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.45.003.227 W slot update_slots: id 0 | task 734 | erased invalidated context checkpoint (pos_min = 444, pos_max = 444, n_tokens = 445, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.45.561.942 I slot create_check: id 0 | task 734 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
0.46.436.833 I slot print_timing: id 0 | task 734 |
prompt eval time = 595.84 ms / 427 tokens ( 1.40 ms per token, 716.64 tokens per second)
eval time = 837.77 ms / 39 tokens ( 21.48 ms per token, 46.55 tokens per second)
total time = 1433.60 ms / 466 tokens
0.46.436.923 I slot release: id 0 | task 734 | stop processing: n_tokens = 465, truncated = 0
0.46.436.955 I srv update_slots: all slots are idle
0.46.450.626 I srv params_from_: Chat format: peg-native
0.46.451.043 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.948 (> 0.100 thold), f_keep = 0.871
0.46.451.527 I reasoning-budget: activated, budget=2147483647 tokens
0.46.451.532 I reasoning-budget: deactivated (natural end)
0.46.451.617 I slot launch_slot_: id 0 | task 775 | processing task, is_child = 0
0.46.451.637 W slot update_slots: id 0 | task 775 | n_past = 405, slot.prompt.tokens.size() = 465, seq_id = 0, pos_min = 464, n_swa = 0
0.46.451.641 I slot update_slots: id 0 | task 775 | Checking checkpoint with [422, 422] against 405...
0.46.451.643 W slot update_slots: id 0 | task 775 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.46.451.649 W slot update_slots: id 0 | task 775 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.46.982.640 I slot create_check: id 0 | task 775 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
0.47.142.299 I slot print_timing: id 0 | task 775 |
prompt eval time = 585.65 ms / 427 tokens ( 1.37 ms per token, 729.10 tokens per second)
eval time = 104.99 ms / 4 tokens ( 26.25 ms per token, 38.10 tokens per second)
total time = 690.65 ms / 431 tokens
0.47.142.414 I slot release: id 0 | task 775 | stop processing: n_tokens = 430, truncated = 0
0.47.142.449 I srv update_slots: all slots are idle
0.47.193.526 I srv params_from_: Chat format: peg-native
0.47.195.446 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.958 (> 0.100 thold), f_keep = 0.944
0.47.195.987 I reasoning-budget: activated, budget=2147483647 tokens
0.47.195.992 I reasoning-budget: deactivated (natural end)
0.47.196.101 I slot launch_slot_: id 0 | task 781 | processing task, is_child = 0
0.47.196.122 W slot update_slots: id 0 | task 781 | n_past = 406, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
0.47.196.128 I slot update_slots: id 0 | task 781 | Checking checkpoint with [422, 422] against 406...
0.47.196.129 W slot update_slots: id 0 | task 781 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.47.196.139 W slot update_slots: id 0 | task 781 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.47.774.566 I slot create_check: id 0 | task 781 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
0.48.760.957 I slot print_timing: id 0 | task 781 |
prompt eval time = 635.18 ms / 424 tokens ( 1.50 ms per token, 667.53 tokens per second)
eval time = 929.64 ms / 40 tokens ( 23.24 ms per token, 43.03 tokens per second)
total time = 1564.83 ms / 464 tokens
0.48.761.052 I slot release: id 0 | task 781 | stop processing: n_tokens = 463, truncated = 0
0.48.761.080 I srv update_slots: all slots are idle
0.48.783.508 I srv params_from_: Chat format: peg-native
0.48.783.860 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.935 (> 0.100 thold), f_keep = 1.000
0.48.784.082 I reasoning-budget: activated, budget=2147483647 tokens
0.48.784.084 I reasoning-budget: deactivated (natural end)
0.48.784.127 I slot launch_slot_: id 0 | task 823 | processing task, is_child = 0
0.48.935.107 I slot create_check: id 0 | task 823 | created context checkpoint 2 of 32 (pos_min = 490, pos_max = 490, n_tokens = 491, size = 62.813 MiB)
0.49.277.790 I slot print_timing: id 0 | task 823 |
prompt eval time = 217.60 ms / 32 tokens ( 6.80 ms per token, 147.06 tokens per second)
eval time = 276.04 ms / 13 tokens ( 21.23 ms per token, 47.09 tokens per second)
total time = 493.64 ms / 45 tokens
0.49.277.883 I slot release: id 0 | task 823 | stop processing: n_tokens = 507, truncated = 0
0.49.277.910 I srv update_slots: all slots are idle
0.49.324.008 I srv params_from_: Chat format: peg-native
0.49.324.548 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.967 (> 0.100 thold), f_keep = 0.809
0.49.324.947 I reasoning-budget: activated, budget=2147483647 tokens
0.49.324.952 I reasoning-budget: deactivated (natural end)
0.49.325.029 I slot launch_slot_: id 0 | task 838 | processing task, is_child = 0
0.49.325.041 W slot update_slots: id 0 | task 838 | n_past = 410, slot.prompt.tokens.size() = 507, seq_id = 0, pos_min = 506, n_swa = 0
0.49.325.042 I slot update_slots: id 0 | task 838 | Checking checkpoint with [490, 490] against 410...
0.49.325.043 I slot update_slots: id 0 | task 838 | Checking checkpoint with [419, 419] against 410...
0.49.325.048 W slot update_slots: id 0 | task 838 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.49.325.051 W slot update_slots: id 0 | task 838 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.49.328.874 W slot update_slots: id 0 | task 838 | erased invalidated context checkpoint (pos_min = 490, pos_max = 490, n_tokens = 491, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.49.907.729 I slot create_check: id 0 | task 838 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
0.50.729.586 I slot print_timing: id 0 | task 838 |
prompt eval time = 624.90 ms / 424 tokens ( 1.47 ms per token, 678.50 tokens per second)
eval time = 779.62 ms / 40 tokens ( 19.49 ms per token, 51.31 tokens per second)
total time = 1404.52 ms / 464 tokens
0.50.729.654 I slot release: id 0 | task 838 | stop processing: n_tokens = 463, truncated = 0
0.50.729.686 I srv update_slots: all slots are idle
0.50.775.784 I srv params_from_: Chat format: peg-native
0.50.776.334 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.931 (> 0.100 thold), f_keep = 0.875
0.50.776.559 I reasoning-budget: activated, budget=2147483647 tokens
0.50.776.562 I reasoning-budget: deactivated (natural end)
0.50.776.609 I slot launch_slot_: id 0 | task 880 | processing task, is_child = 0
0.50.776.620 W slot update_slots: id 0 | task 880 | n_past = 405, slot.prompt.tokens.size() = 463, seq_id = 0, pos_min = 462, n_swa = 0
0.50.776.622 I slot update_slots: id 0 | task 880 | Checking checkpoint with [419, 419] against 405...
0.50.776.623 W slot update_slots: id 0 | task 880 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
0.50.776.627 W slot update_slots: id 0 | task 880 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
0.51.328.909 I slot create_check: id 0 | task 880 | created context checkpoint 1 of 32 (pos_min = 430, pos_max = 430, n_tokens = 431, size = 62.813 MiB)
0.52.962.266 I slot print_timing: id 0 | task 880 |
prompt eval time = 596.05 ms / 435 tokens ( 1.37 ms per token, 729.81 tokens per second)
eval time = 1589.57 ms / 80 tokens ( 19.87 ms per token, 50.33 tokens per second)
total time = 2185.62 ms / 515 tokens
0.52.962.487 I slot release: id 0 | task 880 | stop processing: n_tokens = 514, truncated = 0
0.52.962.523 I srv update_slots: all slots are idle
0.52.963.931 I srv operator(): operator(): cleaning up before exit...
|