| 0.00.113.666 W Setting 'enable_thinking' via --chat-template-kwargs is deprecated. Use --reasoning on / --reasoning off instead. |
| 0.00.124.179 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg) |
| 0.00.124.183 I device_info: |
| 0.00.124.292 I - ROCm0 : AMD Radeon Graphics (131072 MiB, 123913 MiB free) |
| 0.00.124.424 I - Vulkan0 : AMD Radeon Graphics (RADV GFX1151) (132096 MiB, 131922 MiB free) |
| 0.00.124.432 I - CPU : AMD RYZEN AI MAX+ 395 w/ Radeon 8060S (127438 MiB, 127438 MiB free) |
| 0.00.124.514 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 | |
| 0.00.124.572 I srv init: running without SSL |
| 0.00.124.596 I srv init: using 31 threads for HTTP server |
| 0.00.124.597 I srv init: the WebUI is disabled |
| 0.00.124.672 I srv start: binding port with default address family |
| 0.00.125.865 I srv main: loading model |
| 0.00.125.873 I srv load_model: loading model '/mnt/models/nex-n2.5-mini/out/Nex-N2.5-mini-Q4_0_ROCMFP4_STRIX_LEAN.gguf' |
| 0.00.172.940 W llama_model_loader: direct I/O is enabled, disabling mmap |
| 0.22.203.780 W llama_context: n_ctx_seq (65536) < n_ctx_train (262144) -- the full capacity of the model will not be utilized |
| 0.22.400.933 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable) |
| 0.22.718.577 I srv load_model: initializing slots, n_slots = 1 |
| 0.22.944.184 W srv load_model: speculative decoding will use checkpoints |
| 0.22.944.193 W common_speculative_init: no implementations specified for speculative decoding |
| 0.22.944.194 I slot load_model: id 0 | task -1 | new slot, n_ctx = 65536 |
| 0.22.944.259 I srv load_model: prompt cache RAM enabled: limit_mib=8192 |
| 0.22.944.261 I srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391 |
| 0.22.944.283 I srv init: idle slots will be saved to prompt cache upon starting a new task |
| 0.22.956.362 I init: chat template, example_format: '<|im_start|>system |
| You are a helpful assistant<|im_end|> |
| <|im_start|>user |
| Hello<|im_end|> |
| <|im_start|>assistant |
| <think> |
|
|
| </think> |
|
|
| Hi there<|im_end|> |
| <|im_start|>user |
| How are you?<|im_end|> |
| <|im_start|>assistant |
| <think> |
|
|
| </think> |
|
|
| ' |
| 0.22.964.154 I srv init: init: chat template, thinking = 1 |
| 0.22.964.188 I srv main: model loaded |
| 0.22.964.191 I srv main: server is listening on http://127.0.0.1:18652 |
| 0.22.964.207 I srv update_slots: all slots are idle |
| 0.24.203.095 I srv params_from_: Chat format: peg-native |
| 0.24.203.661 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1 |
| 0.24.203.663 I srv get_availabl: updating prompt cache |
| 0.24.203.670 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000 |
| 0.24.203.676 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 65536 tokens, 8589934592 est) |
| 0.24.203.678 I srv get_availabl: prompt cache update took 0.01 ms |
| 0.24.203.955 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.24.203.957 I reasoning-budget: deactivated (natural end) |
| 0.24.203.972 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0 |
| 0.24.920.604 I slot create_check: id 0 | task 0 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB) |
| 0.25.129.936 I slot print_timing: id 0 | task 0 | |
| prompt eval time = 758.65 ms / 427 tokens ( 1.78 ms per token, 562.84 tokens per second) |
| eval time = 167.29 ms / 4 tokens ( 41.82 ms per token, 23.91 tokens per second) |
| total time = 925.94 ms / 431 tokens |
| 0.25.130.011 I slot release: id 0 | task 0 | stop processing: n_tokens = 430, truncated = 0 |
| 0.25.130.019 I srv update_slots: all slots are idle |
| 0.25.185.780 I srv params_from_: Chat format: peg-native |
| 0.25.186.376 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.942 |
| 0.25.186.613 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.25.186.618 I reasoning-budget: deactivated (natural end) |
| 0.25.186.779 I slot launch_slot_: id 0 | task 6 | processing task, is_child = 0 |
| 0.25.186.816 W slot update_slots: id 0 | task 6 | n_past = 405, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0 |
| 0.25.186.819 I slot update_slots: id 0 | task 6 | Checking checkpoint with [422, 422] against 405... |
| 0.25.186.822 W slot update_slots: id 0 | task 6 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) |
| 0.25.186.829 W slot update_slots: id 0 | task 6 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB) |
| 0.25.897.246 I slot create_check: id 0 | task 6 | created context checkpoint 1 of 32 (pos_min = 416, pos_max = 416, n_tokens = 417, size = 62.813 MiB) |
| 0.25.983.622 I slot print_timing: id 0 | task 6 | |
| prompt eval time = 763.27 ms / 421 tokens ( 1.81 ms per token, 551.57 tokens per second) |
| eval time = 33.53 ms / 2 tokens ( 16.77 ms per token, 59.64 tokens per second) |
| total time = 796.80 ms / 423 tokens |
| 0.25.983.694 I slot release: id 0 | task 6 | stop processing: n_tokens = 422, truncated = 0 |
| 0.25.983.724 I srv update_slots: all slots are idle |
| 0.26.010.893 I srv params_from_: Chat format: peg-native |
| 0.26.011.386 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.960 |
| 0.26.011.582 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.26.011.584 I reasoning-budget: deactivated (natural end) |
| 0.26.011.615 I slot launch_slot_: id 0 | task 10 | processing task, is_child = 0 |
| 0.26.011.627 W slot update_slots: id 0 | task 10 | n_past = 405, slot.prompt.tokens.size() = 422, seq_id = 0, pos_min = 421, n_swa = 0 |
| 0.26.011.628 I slot update_slots: id 0 | task 10 | Checking checkpoint with [416, 416] against 405... |
| 0.26.011.629 W slot update_slots: id 0 | task 10 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) |
| 0.26.011.634 W slot update_slots: id 0 | task 10 | erased invalidated context checkpoint (pos_min = 416, pos_max = 416, n_tokens = 417, n_swa = 0, pos_next = 0, size = 62.813 MiB) |
| 0.26.506.256 I slot create_check: id 0 | task 10 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB) |
| 0.27.421.771 I slot print_timing: id 0 | task 10 | |
| prompt eval time = 545.66 ms / 424 tokens ( 1.29 ms per token, 777.04 tokens per second) |
| eval time = 864.45 ms / 39 tokens ( 22.17 ms per token, 45.12 tokens per second) |
| total time = 1410.11 ms / 463 tokens |
| 0.27.421.958 I slot release: id 0 | task 10 | stop processing: n_tokens = 462, truncated = 0 |
| 0.27.422.017 I srv update_slots: all slots are idle |
| 0.27.483.797 I srv params_from_: Chat format: peg-native |
| 0.27.486.309 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.951 (> 0.100 thold), f_keep = 0.879 |
| 0.27.486.919 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.27.486.923 I reasoning-budget: deactivated (natural end) |
| 0.27.487.006 I slot launch_slot_: id 0 | task 51 | processing task, is_child = 0 |
| 0.27.487.029 W slot update_slots: id 0 | task 51 | n_past = 406, slot.prompt.tokens.size() = 462, seq_id = 0, pos_min = 461, n_swa = 0 |
| 0.27.487.032 I slot update_slots: id 0 | task 51 | Checking checkpoint with [419, 419] against 406... |
| 0.27.487.034 W slot update_slots: id 0 | task 51 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) |
| 0.27.487.041 W slot update_slots: id 0 | task 51 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB) |
| 0.28.076.337 I slot create_check: id 0 | task 51 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB) |
| 0.28.262.498 I slot print_timing: id 0 | task 51 | |
| prompt eval time = 638.39 ms / 427 tokens ( 1.50 ms per token, 668.86 tokens per second) |
| eval time = 137.05 ms / 4 tokens ( 34.26 ms per token, 29.19 tokens per second) |
| total time = 775.45 ms / 431 tokens |
| 0.28.262.688 I slot release: id 0 | task 51 | stop processing: n_tokens = 430, truncated = 0 |
| 0.28.262.742 I srv update_slots: all slots are idle |
| 0.28.315.880 I srv params_from_: Chat format: peg-native |
| 0.28.318.140 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.942 |
| 0.28.318.686 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.28.318.690 I reasoning-budget: deactivated (natural end) |
| 0.28.318.777 I slot launch_slot_: id 0 | task 57 | processing task, is_child = 0 |
| 0.28.318.800 W slot update_slots: id 0 | task 57 | n_past = 405, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0 |
| 0.28.318.803 I slot update_slots: id 0 | task 57 | Checking checkpoint with [422, 422] against 405... |
| 0.28.318.805 W slot update_slots: id 0 | task 57 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) |
| 0.28.318.810 W slot update_slots: id 0 | task 57 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB) |
| 0.28.907.669 I slot create_check: id 0 | task 57 | created context checkpoint 1 of 32 (pos_min = 416, pos_max = 416, n_tokens = 417, size = 62.813 MiB) |
| 0.28.982.494 I slot print_timing: id 0 | task 57 | |
| prompt eval time = 636.88 ms / 421 tokens ( 1.51 ms per token, 661.04 tokens per second) |
| eval time = 26.81 ms / 2 tokens ( 13.41 ms per token, 74.60 tokens per second) |
| total time = 663.69 ms / 423 tokens |
| 0.28.982.572 I slot release: id 0 | task 57 | stop processing: n_tokens = 422, truncated = 0 |
| 0.28.982.598 I srv update_slots: all slots are idle |
| 0.29.009.222 I srv params_from_: Chat format: peg-native |
| 0.29.011.294 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.960 |
| 0.29.011.848 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.29.011.851 I reasoning-budget: deactivated (natural end) |
| 0.29.011.922 I slot launch_slot_: id 0 | task 61 | processing task, is_child = 0 |
| 0.29.011.941 W slot update_slots: id 0 | task 61 | n_past = 405, slot.prompt.tokens.size() = 422, seq_id = 0, pos_min = 421, n_swa = 0 |
| 0.29.011.943 I slot update_slots: id 0 | task 61 | Checking checkpoint with [416, 416] against 405... |
| 0.29.011.945 W slot update_slots: id 0 | task 61 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) |
| 0.29.011.949 W slot update_slots: id 0 | task 61 | erased invalidated context checkpoint (pos_min = 416, pos_max = 416, n_tokens = 417, n_swa = 0, pos_next = 0, size = 62.813 MiB) |
| 0.29.554.822 I slot create_check: id 0 | task 61 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB) |
| 0.30.389.784 I slot print_timing: id 0 | task 61 | |
| prompt eval time = 589.42 ms / 424 tokens ( 1.39 ms per token, 719.35 tokens per second) |
| eval time = 788.41 ms / 39 tokens ( 20.22 ms per token, 49.47 tokens per second) |
| total time = 1377.83 ms / 463 tokens |
| 0.30.389.858 I slot release: id 0 | task 61 | stop processing: n_tokens = 462, truncated = 0 |
| 0.30.389.883 I srv update_slots: all slots are idle |
| 0.30.401.499 I srv params_from_: Chat format: peg-native |
| 0.30.401.853 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.879 |
| 0.30.402.127 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.30.402.129 I reasoning-budget: deactivated (natural end) |
| 0.30.402.192 I slot launch_slot_: id 0 | task 102 | processing task, is_child = 0 |
| 0.30.402.206 W slot update_slots: id 0 | task 102 | n_past = 406, slot.prompt.tokens.size() = 462, seq_id = 0, pos_min = 461, n_swa = 0 |
| 0.30.402.208 I slot update_slots: id 0 | task 102 | Checking checkpoint with [419, 419] against 406... |
| 0.30.402.209 W slot update_slots: id 0 | task 102 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) |
| 0.30.402.212 W slot update_slots: id 0 | task 102 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB) |
| 0.30.923.747 I slot create_check: id 0 | task 102 | created context checkpoint 1 of 32 (pos_min = 420, pos_max = 420, n_tokens = 421, size = 62.813 MiB) |
| 0.31.379.211 I slot print_timing: id 0 | task 102 | |
| prompt eval time = 572.31 ms / 425 tokens ( 1.35 ms per token, 742.61 tokens per second) |
| eval time = 404.68 ms / 17 tokens ( 23.80 ms per token, 42.01 tokens per second) |
| total time = 976.98 ms / 442 tokens |
| 0.31.379.332 I slot release: id 0 | task 102 | stop processing: n_tokens = 441, truncated = 0 |
| 0.31.379.373 I srv update_slots: all slots are idle |
| 0.31.405.914 I srv params_from_: Chat format: peg-native |
| 0.31.406.236 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.967 (> 0.100 thold), f_keep = 0.918 |
| 0.31.406.397 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.31.406.398 I reasoning-budget: deactivated (natural end) |
| 0.31.406.427 I slot launch_slot_: id 0 | task 121 | processing task, is_child = 0 |
| 0.31.406.435 W slot update_slots: id 0 | task 121 | n_past = 405, slot.prompt.tokens.size() = 441, seq_id = 0, pos_min = 440, n_swa = 0 |
| 0.31.406.436 I slot update_slots: id 0 | task 121 | Checking checkpoint with [420, 420] against 405... |
| 0.31.406.437 W slot update_slots: id 0 | task 121 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) |
| 0.31.406.439 W slot update_slots: id 0 | task 121 | erased invalidated context checkpoint (pos_min = 420, pos_max = 420, n_tokens = 421, n_swa = 0, pos_next = 0, size = 62.813 MiB) |
| 0.31.861.565 I slot create_check: id 0 | task 121 | created context checkpoint 1 of 32 (pos_min = 414, pos_max = 414, n_tokens = 415, size = 62.813 MiB) |
| 0.32.160.128 I slot print_timing: id 0 | task 121 | |
| prompt eval time = 491.31 ms / 419 tokens ( 1.17 ms per token, 852.82 tokens per second) |
| eval time = 262.37 ms / 12 tokens ( 21.86 ms per token, 45.74 tokens per second) |
| total time = 753.68 ms / 431 tokens |
| 0.32.160.211 I slot release: id 0 | task 121 | stop processing: n_tokens = 430, truncated = 0 |
| 0.32.160.241 I srv update_slots: all slots are idle |
| 0.32.209.194 I srv params_from_: Chat format: peg-native |
| 0.32.209.712 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.960 (> 0.100 thold), f_keep = 0.942 |
| 0.32.209.953 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.32.209.954 I reasoning-budget: deactivated (natural end) |
| 0.32.209.994 I slot launch_slot_: id 0 | task 135 | processing task, is_child = 0 |
| 0.32.210.006 W slot update_slots: id 0 | task 135 | n_past = 405, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0 |
| 0.32.210.007 I slot update_slots: id 0 | task 135 | Checking checkpoint with [414, 414] against 405... |
| 0.32.210.008 W slot update_slots: id 0 | task 135 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) |
| 0.32.210.011 W slot update_slots: id 0 | task 135 | erased invalidated context checkpoint (pos_min = 414, pos_max = 414, n_tokens = 415, n_swa = 0, pos_next = 0, size = 62.813 MiB) |
| 0.32.786.978 I slot create_check: id 0 | task 135 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB) |
| 0.33.839.717 I slot print_timing: id 0 | task 135 | |
| prompt eval time = 614.39 ms / 422 tokens ( 1.46 ms per token, 686.86 tokens per second) |
| eval time = 1015.30 ms / 53 tokens ( 19.16 ms per token, 52.20 tokens per second) |
| total time = 1629.69 ms / 475 tokens |
| 0.33.839.794 I slot release: id 0 | task 135 | stop processing: n_tokens = 474, truncated = 0 |
| 0.33.839.825 I srv update_slots: all slots are idle |
| 0.33.879.962 I srv params_from_: Chat format: peg-native |
| 0.33.880.516 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.958 (> 0.100 thold), f_keep = 0.857 |
| 0.33.880.743 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.33.880.744 I reasoning-budget: deactivated (natural end) |
| 0.33.880.792 I slot launch_slot_: id 0 | task 190 | processing task, is_child = 0 |
| 0.33.880.805 W slot update_slots: id 0 | task 190 | n_past = 406, slot.prompt.tokens.size() = 474, seq_id = 0, pos_min = 473, n_swa = 0 |
| 0.33.880.806 I slot update_slots: id 0 | task 190 | Checking checkpoint with [417, 417] against 406... |
| 0.33.880.807 W slot update_slots: id 0 | task 190 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) |
| 0.33.880.810 W slot update_slots: id 0 | task 190 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB) |
| 0.34.423.705 I slot create_check: id 0 | task 190 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB) |
| 0.34.628.920 I slot print_timing: id 0 | task 190 | |
| prompt eval time = 583.31 ms / 424 tokens ( 1.38 ms per token, 726.89 tokens per second) |
| eval time = 164.80 ms / 7 tokens ( 23.54 ms per token, 42.48 tokens per second) |
| total time = 748.11 ms / 431 tokens |
| 0.34.628.991 I slot release: id 0 | task 190 | stop processing: n_tokens = 430, truncated = 0 |
| 0.34.629.018 I srv update_slots: all slots are idle |
| 0.34.642.421 I srv params_from_: Chat format: peg-native |
| 0.34.642.945 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.969 (> 0.100 thold), f_keep = 0.942 |
| 0.34.643.148 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.34.643.150 I reasoning-budget: deactivated (natural end) |
| 0.34.643.185 I slot launch_slot_: id 0 | task 199 | processing task, is_child = 0 |
| 0.34.643.195 W slot update_slots: id 0 | task 199 | n_past = 405, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0 |
| 0.34.643.196 I slot update_slots: id 0 | task 199 | Checking checkpoint with [419, 419] against 405... |
| 0.34.643.197 W slot update_slots: id 0 | task 199 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) |
| 0.34.643.199 W slot update_slots: id 0 | task 199 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB) |
| 0.35.154.839 I slot create_check: id 0 | task 199 | created context checkpoint 1 of 32 (pos_min = 413, pos_max = 413, n_tokens = 414, size = 62.813 MiB) |
| 0.35.329.011 I slot print_timing: id 0 | task 199 | |
| prompt eval time = 563.78 ms / 418 tokens ( 1.35 ms per token, 741.42 tokens per second) |
| eval time = 122.02 ms / 5 tokens ( 24.40 ms per token, 40.98 tokens per second) |
| total time = 685.80 ms / 423 tokens |
| 0.35.329.095 I slot release: id 0 | task 199 | stop processing: n_tokens = 422, truncated = 0 |
| 0.35.329.136 I srv update_slots: all slots are idle |
| 0.35.355.115 I srv params_from_: Chat format: peg-native |
| 0.35.355.649 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.960 |
| 0.35.356.326 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.35.356.340 I reasoning-budget: deactivated (natural end) |
| 0.35.356.552 I slot launch_slot_: id 0 | task 206 | processing task, is_child = 0 |
| 0.35.356.603 W slot update_slots: id 0 | task 206 | n_past = 405, slot.prompt.tokens.size() = 422, seq_id = 0, pos_min = 421, n_swa = 0 |
| 0.35.356.611 I slot update_slots: id 0 | task 206 | Checking checkpoint with [413, 413] against 405... |
| 0.35.356.616 W slot update_slots: id 0 | task 206 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) |
| 0.35.356.634 W slot update_slots: id 0 | task 206 | erased invalidated context checkpoint (pos_min = 413, pos_max = 413, n_tokens = 414, n_swa = 0, pos_next = 0, size = 62.813 MiB) |
| 0.35.935.879 I slot create_check: id 0 | task 206 | created context checkpoint 1 of 32 (pos_min = 416, pos_max = 416, n_tokens = 417, size = 62.813 MiB) |
| 0.36.896.774 I slot print_timing: id 0 | task 206 | |
| prompt eval time = 625.99 ms / 421 tokens ( 1.49 ms per token, 672.53 tokens per second) |
| eval time = 914.17 ms / 42 tokens ( 21.77 ms per token, 45.94 tokens per second) |
| total time = 1540.17 ms / 463 tokens |
| 0.36.896.862 I slot release: id 0 | task 206 | stop processing: n_tokens = 462, truncated = 0 |
| 0.36.896.896 I srv update_slots: all slots are idle |
| 0.36.936.572 I srv params_from_: Chat format: peg-native |
| 0.36.937.088 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.951 (> 0.100 thold), f_keep = 0.879 |
| 0.36.937.669 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.36.937.673 I reasoning-budget: deactivated (natural end) |
| 0.36.937.771 I slot launch_slot_: id 0 | task 250 | processing task, is_child = 0 |
| 0.36.937.796 W slot update_slots: id 0 | task 250 | n_past = 406, slot.prompt.tokens.size() = 462, seq_id = 0, pos_min = 461, n_swa = 0 |
| 0.36.937.799 I slot update_slots: id 0 | task 250 | Checking checkpoint with [416, 416] against 406... |
| 0.36.937.801 W slot update_slots: id 0 | task 250 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) |
| 0.36.937.807 W slot update_slots: id 0 | task 250 | erased invalidated context checkpoint (pos_min = 416, pos_max = 416, n_tokens = 417, n_swa = 0, pos_next = 0, size = 62.813 MiB) |
| 0.37.457.001 I slot create_check: id 0 | task 250 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB) |
| 0.37.596.606 I slot print_timing: id 0 | task 250 | |
| prompt eval time = 561.66 ms / 427 tokens ( 1.32 ms per token, 760.24 tokens per second) |
| eval time = 97.14 ms / 4 tokens ( 24.28 ms per token, 41.18 tokens per second) |
| total time = 658.80 ms / 431 tokens |
| 0.37.596.716 I slot release: id 0 | task 250 | stop processing: n_tokens = 430, truncated = 0 |
| 0.37.596.750 I srv update_slots: all slots are idle |
| 0.37.626.974 I srv params_from_: Chat format: peg-native |
| 0.37.627.483 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.942 |
| 0.37.627.753 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.37.627.755 I reasoning-budget: deactivated (natural end) |
| 0.37.627.791 I slot launch_slot_: id 0 | task 256 | processing task, is_child = 0 |
| 0.37.627.804 W slot update_slots: id 0 | task 256 | n_past = 405, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0 |
| 0.37.627.806 I slot update_slots: id 0 | task 256 | Checking checkpoint with [422, 422] against 405... |
| 0.37.627.807 W slot update_slots: id 0 | task 256 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) |
| 0.37.627.811 W slot update_slots: id 0 | task 256 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB) |
| 0.38.154.752 I slot create_check: id 0 | task 256 | created context checkpoint 1 of 32 (pos_min = 416, pos_max = 416, n_tokens = 417, size = 62.813 MiB) |
| 0.38.242.158 I slot print_timing: id 0 | task 256 | |
| prompt eval time = 577.92 ms / 421 tokens ( 1.37 ms per token, 728.48 tokens per second) |
| eval time = 36.41 ms / 2 tokens ( 18.21 ms per token, 54.92 tokens per second) |
| total time = 614.33 ms / 423 tokens |
| 0.38.242.323 I slot release: id 0 | task 256 | stop processing: n_tokens = 422, truncated = 0 |
| 0.38.242.358 I srv update_slots: all slots are idle |
| 0.38.257.074 I srv params_from_: Chat format: peg-native |
| 0.38.257.513 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.960 |
| 0.38.257.760 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.38.257.761 I reasoning-budget: deactivated (natural end) |
| 0.38.257.798 I slot launch_slot_: id 0 | task 260 | processing task, is_child = 0 |
| 0.38.257.809 W slot update_slots: id 0 | task 260 | n_past = 405, slot.prompt.tokens.size() = 422, seq_id = 0, pos_min = 421, n_swa = 0 |
| 0.38.257.809 I slot update_slots: id 0 | task 260 | Checking checkpoint with [416, 416] against 405... |
| 0.38.257.810 W slot update_slots: id 0 | task 260 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) |
| 0.38.257.813 W slot update_slots: id 0 | task 260 | erased invalidated context checkpoint (pos_min = 416, pos_max = 416, n_tokens = 417, n_swa = 0, pos_next = 0, size = 62.813 MiB) |
| 0.38.777.417 I slot create_check: id 0 | task 260 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB) |
| 0.39.603.048 I slot print_timing: id 0 | task 260 | |
| prompt eval time = 562.75 ms / 424 tokens ( 1.33 ms per token, 753.44 tokens per second) |
| eval time = 782.47 ms / 39 tokens ( 20.06 ms per token, 49.84 tokens per second) |
| total time = 1345.22 ms / 463 tokens |
| 0.39.603.254 I slot release: id 0 | task 260 | stop processing: n_tokens = 462, truncated = 0 |
| 0.39.603.336 I srv update_slots: all slots are idle |
| 0.39.604.550 I srv operator(): operator(): cleaning up before exit... |
|
|