| 0.00.114.439 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg) |
| 0.00.114.444 I device_info: |
| 0.00.114.555 I - ROCm0 : AMD Radeon Graphics (131072 MiB, 123723 MiB free) |
| 0.00.114.708 I - Vulkan0 : AMD Radeon Graphics (RADV GFX1151) (132096 MiB, 131922 MiB free) |
| 0.00.114.715 I - CPU : AMD RYZEN AI MAX+ 395 w/ Radeon 8060S (127438 MiB, 127438 MiB free) |
| 0.00.114.796 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 | |
| 0.00.114.847 I srv init: running without SSL |
| 0.00.114.871 I srv init: using 31 threads for HTTP server |
| 0.00.114.873 I srv init: the WebUI is disabled |
| 0.00.114.942 I srv start: binding port with default address family |
| 0.00.116.175 I srv main: loading model |
| 0.00.116.182 I srv load_model: loading model '/mnt/models/nex-n2.5-mini/out/Nex-N2.5-mini-Q4_0_ROCMFP4_STRIX_LEAN.gguf' |
| 0.00.160.730 W llama_model_loader: direct I/O is enabled, disabling mmap |
| 0.21.513.834 W llama_context: n_ctx_seq (65536) < n_ctx_train (262144) -- the full capacity of the model will not be utilized |
| 0.21.793.651 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable) |
| 0.22.217.020 I srv load_model: initializing slots, n_slots = 1 |
| 0.22.487.377 W srv load_model: speculative decoding will use checkpoints |
| 0.22.487.387 W common_speculative_init: no implementations specified for speculative decoding |
| 0.22.487.390 I slot load_model: id 0 | task -1 | new slot, n_ctx = 65536 |
| 0.22.487.471 I srv load_model: prompt cache RAM enabled: limit_mib=8192 |
| 0.22.487.473 I srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391 |
| 0.22.487.495 I srv init: idle slots will be saved to prompt cache upon starting a new task |
| 0.22.540.918 I init: chat template, example_format: '<|im_start|>system |
| You are a helpful assistant<|im_end|> |
| <|im_start|>user |
| Hello<|im_end|> |
| <|im_start|>assistant |
| <think> |
|
|
| </think> |
|
|
| Hi there<|im_end|> |
| <|im_start|>user |
| How are you?<|im_end|> |
| <|im_start|>assistant |
| <think> |
|
|
| </think> |
|
|
| ' |
| 0.22.582.904 I srv init: init: chat template, thinking = 0 |
| 0.22.582.984 I srv main: model loaded |
| 0.22.582.995 I srv main: server is listening on http://127.0.0.1:18600 |
| 0.22.583.003 I srv update_slots: all slots are idle |
| 0.23.994.344 I srv params_from_: Chat format: peg-native |
| 0.23.995.890 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1 |
| 0.23.995.893 I srv get_availabl: updating prompt cache |
| 0.23.995.900 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000 |
| 0.23.995.907 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 65536 tokens, 8589934592 est) |
| 0.23.995.909 I srv get_availabl: prompt cache update took 0.01 ms |
| 0.23.996.222 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.23.996.243 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0 |
| 0.24.656.206 I slot create_check: id 0 | task 0 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB) |
| 0.25.028.892 I reasoning-budget: deactivated (natural end) |
| 0.25.816.313 I slot print_timing: id 0 | task 0 | |
| prompt eval time = 698.88 ms / 422 tokens ( 1.66 ms per token, 603.83 tokens per second) |
| eval time = 1121.16 ms / 54 tokens ( 20.76 ms per token, 48.16 tokens per second) |
| total time = 1820.04 ms / 476 tokens |
| 0.25.816.410 I slot release: id 0 | task 0 | stop processing: n_tokens = 475, truncated = 0 |
| 0.25.816.424 I srv update_slots: all slots are idle |
| 0.25.830.282 I srv params_from_: Chat format: peg-native |
| 0.25.830.661 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.906 (> 0.100 thold), f_keep = 0.853 |
| 0.25.830.794 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.25.830.832 I slot launch_slot_: id 0 | task 56 | processing task, is_child = 0 |
| 0.25.830.841 W slot update_slots: id 0 | task 56 | n_past = 405, slot.prompt.tokens.size() = 475, seq_id = 0, pos_min = 474, n_swa = 0 |
| 0.25.830.842 I slot update_slots: id 0 | task 56 | Checking checkpoint with [417, 417] against 405... |
| 0.25.830.844 W slot update_slots: id 0 | task 56 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) |
| 0.25.830.846 W slot update_slots: id 0 | task 56 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB) |
| 0.26.482.806 I slot create_check: id 0 | task 56 | created context checkpoint 1 of 32 (pos_min = 442, pos_max = 442, n_tokens = 443, size = 62.813 MiB) |
| 0.26.903.330 I reasoning-budget: deactivated (natural end) |
| 0.28.669.262 I slot print_timing: id 0 | task 56 | n_decoded = 100, tg = 47.24 t/s |
| 0.28.729.150 I slot print_timing: id 0 | task 56 | |
| prompt eval time = 721.52 ms / 447 tokens ( 1.61 ms per token, 619.52 tokens per second) |
| eval time = 2176.78 ms / 103 tokens ( 21.13 ms per token, 47.32 tokens per second) |
| total time = 2898.30 ms / 550 tokens |
| 0.28.729.226 I slot release: id 0 | task 56 | stop processing: n_tokens = 549, truncated = 0 |
| 0.28.729.266 I srv update_slots: all slots are idle |
| 0.28.742.730 I srv params_from_: Chat format: peg-native |
| 0.28.743.168 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.953 (> 0.100 thold), f_keep = 0.738 |
| 0.28.743.356 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.28.743.392 I slot launch_slot_: id 0 | task 161 | processing task, is_child = 0 |
| 0.28.743.406 W slot update_slots: id 0 | task 161 | n_past = 405, slot.prompt.tokens.size() = 549, seq_id = 0, pos_min = 548, n_swa = 0 |
| 0.28.743.407 I slot update_slots: id 0 | task 161 | Checking checkpoint with [442, 442] against 405... |
| 0.28.743.407 W slot update_slots: id 0 | task 161 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) |
| 0.28.743.411 W slot update_slots: id 0 | task 161 | erased invalidated context checkpoint (pos_min = 442, pos_max = 442, n_tokens = 443, n_swa = 0, pos_next = 0, size = 62.813 MiB) |
| 0.29.274.427 I slot create_check: id 0 | task 161 | created context checkpoint 1 of 32 (pos_min = 420, pos_max = 420, n_tokens = 421, size = 62.813 MiB) |
| 0.29.681.788 I reasoning-budget: deactivated (natural end) |
| 0.30.465.485 I slot print_timing: id 0 | task 161 | |
| prompt eval time = 568.14 ms / 425 tokens ( 1.34 ms per token, 748.05 tokens per second) |
| eval time = 1153.92 ms / 57 tokens ( 20.24 ms per token, 49.40 tokens per second) |
| total time = 1722.06 ms / 482 tokens |
| 0.30.465.571 I slot release: id 0 | task 161 | stop processing: n_tokens = 481, truncated = 0 |
| 0.30.465.599 I srv update_slots: all slots are idle |
| 0.30.480.783 I srv params_from_: Chat format: peg-native |
| 0.30.481.142 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.953 (> 0.100 thold), f_keep = 0.842 |
| 0.30.481.346 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.30.481.386 I slot launch_slot_: id 0 | task 220 | processing task, is_child = 0 |
| 0.30.481.397 W slot update_slots: id 0 | task 220 | n_past = 405, slot.prompt.tokens.size() = 481, seq_id = 0, pos_min = 480, n_swa = 0 |
| 0.30.481.397 I slot update_slots: id 0 | task 220 | Checking checkpoint with [420, 420] against 405... |
| 0.30.481.399 W slot update_slots: id 0 | task 220 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) |
| 0.30.481.401 W slot update_slots: id 0 | task 220 | erased invalidated context checkpoint (pos_min = 420, pos_max = 420, n_tokens = 421, n_swa = 0, pos_next = 0, size = 62.813 MiB) |
| 0.31.008.867 I slot create_check: id 0 | task 220 | created context checkpoint 1 of 32 (pos_min = 420, pos_max = 420, n_tokens = 421, size = 62.813 MiB) |
| 0.31.355.463 I reasoning-budget: deactivated (natural end) |
| 0.31.458.677 I slot print_timing: id 0 | task 220 | |
| prompt eval time = 568.77 ms / 425 tokens ( 1.34 ms per token, 747.23 tokens per second) |
| eval time = 408.50 ms / 20 tokens ( 20.42 ms per token, 48.96 tokens per second) |
| total time = 977.27 ms / 445 tokens |
| 0.31.458.774 I slot release: id 0 | task 220 | stop processing: n_tokens = 444, truncated = 0 |
| 0.31.458.805 I srv update_slots: all slots are idle |
| 0.31.510.659 I srv params_from_: Chat format: peg-native |
| 0.31.511.238 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.914 |
| 0.31.511.571 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.31.511.663 I slot launch_slot_: id 0 | task 242 | processing task, is_child = 0 |
| 0.31.511.677 W slot update_slots: id 0 | task 242 | n_past = 406, slot.prompt.tokens.size() = 444, seq_id = 0, pos_min = 443, n_swa = 0 |
| 0.31.511.684 I slot update_slots: id 0 | task 242 | Checking checkpoint with [420, 420] against 406... |
| 0.31.511.685 W slot update_slots: id 0 | task 242 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) |
| 0.31.511.690 W slot update_slots: id 0 | task 242 | erased invalidated context checkpoint (pos_min = 420, pos_max = 420, n_tokens = 421, n_swa = 0, pos_next = 0, size = 62.813 MiB) |
| 0.32.149.342 I slot create_check: id 0 | task 242 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB) |
| 0.32.405.933 I reasoning-budget: deactivated (natural end) |
| 0.33.283.206 I slot print_timing: id 0 | task 242 | |
| prompt eval time = 677.52 ms / 422 tokens ( 1.61 ms per token, 622.86 tokens per second) |
| eval time = 1093.99 ms / 52 tokens ( 21.04 ms per token, 47.53 tokens per second) |
| total time = 1771.51 ms / 474 tokens |
| 0.33.283.283 I slot release: id 0 | task 242 | stop processing: n_tokens = 473, truncated = 0 |
| 0.33.283.309 I srv update_slots: all slots are idle |
| 0.33.311.488 I srv params_from_: Chat format: peg-native |
| 0.33.312.015 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.854 (> 0.100 thold), f_keep = 0.890 |
| 0.33.312.441 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.33.312.556 I slot launch_slot_: id 0 | task 296 | processing task, is_child = 0 |
| 0.33.312.580 W slot update_slots: id 0 | task 296 | n_past = 421, slot.prompt.tokens.size() = 473, seq_id = 0, pos_min = 472, n_swa = 0 |
| 0.33.312.583 I slot update_slots: id 0 | task 296 | Checking checkpoint with [417, 417] against 421... |
| 0.33.321.205 W slot update_slots: id 0 | task 296 | restored context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_past = 418, size = 62.813 MiB) |
| 0.33.558.612 I slot create_check: id 0 | task 296 | created context checkpoint 2 of 32 (pos_min = 488, pos_max = 488, n_tokens = 489, size = 62.813 MiB) |
| 0.34.218.104 I reasoning-budget: deactivated (natural end) |
| 0.34.556.664 I slot print_timing: id 0 | task 296 | |
| prompt eval time = 286.32 ms / 75 tokens ( 3.82 ms per token, 261.95 tokens per second) |
| eval time = 957.75 ms / 41 tokens ( 23.36 ms per token, 42.81 tokens per second) |
| total time = 1244.07 ms / 116 tokens |
| 0.34.556.763 I slot release: id 0 | task 296 | stop processing: n_tokens = 533, truncated = 0 |
| 0.34.556.805 I srv update_slots: all slots are idle |
| 0.34.571.001 I srv params_from_: Chat format: peg-native |
| 0.34.571.465 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.972 (> 0.100 thold), f_keep = 0.769 |
| 0.34.571.696 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.34.571.740 I slot launch_slot_: id 0 | task 339 | processing task, is_child = 0 |
| 0.34.571.755 W slot update_slots: id 0 | task 339 | n_past = 410, slot.prompt.tokens.size() = 533, seq_id = 0, pos_min = 532, n_swa = 0 |
| 0.34.571.756 I slot update_slots: id 0 | task 339 | Checking checkpoint with [488, 488] against 410... |
| 0.34.571.758 I slot update_slots: id 0 | task 339 | Checking checkpoint with [417, 417] against 410... |
| 0.34.571.759 W slot update_slots: id 0 | task 339 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) |
| 0.34.571.763 W slot update_slots: id 0 | task 339 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB) |
| 0.34.573.311 W slot update_slots: id 0 | task 339 | erased invalidated context checkpoint (pos_min = 488, pos_max = 488, n_tokens = 489, n_swa = 0, pos_next = 0, size = 62.813 MiB) |
| 0.35.110.746 I slot create_check: id 0 | task 339 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB) |
| 0.35.584.142 I reasoning-budget: deactivated (natural end) |
| 0.36.392.408 I slot print_timing: id 0 | task 339 | |
| prompt eval time = 576.37 ms / 422 tokens ( 1.37 ms per token, 732.16 tokens per second) |
| eval time = 1244.26 ms / 62 tokens ( 20.07 ms per token, 49.83 tokens per second) |
| total time = 1820.64 ms / 484 tokens |
| 0.36.392.474 I slot release: id 0 | task 339 | stop processing: n_tokens = 483, truncated = 0 |
| 0.36.392.500 I srv update_slots: all slots are idle |
| 0.36.404.680 I srv params_from_: Chat format: peg-native |
| 0.36.405.094 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.935 (> 0.100 thold), f_keep = 0.839 |
| 0.36.405.518 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.36.405.599 I slot launch_slot_: id 0 | task 403 | processing task, is_child = 0 |
| 0.36.405.620 W slot update_slots: id 0 | task 403 | n_past = 405, slot.prompt.tokens.size() = 483, seq_id = 0, pos_min = 482, n_swa = 0 |
| 0.36.405.621 I slot update_slots: id 0 | task 403 | Checking checkpoint with [417, 417] against 405... |
| 0.36.405.625 W slot update_slots: id 0 | task 403 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) |
| 0.36.405.630 W slot update_slots: id 0 | task 403 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB) |
| 0.36.974.524 I slot create_check: id 0 | task 403 | created context checkpoint 1 of 32 (pos_min = 428, pos_max = 428, n_tokens = 429, size = 62.813 MiB) |
| 0.39.070.997 I slot print_timing: id 0 | task 403 | n_decoded = 100, tg = 48.88 t/s |
| 0.39.448.467 I reasoning-budget: deactivated (natural end) |
| 0.41.042.559 I slot print_timing: id 0 | task 403 | |
| prompt eval time = 619.66 ms / 433 tokens ( 1.43 ms per token, 698.77 tokens per second) |
| eval time = 4017.27 ms / 200 tokens ( 20.09 ms per token, 49.79 tokens per second) |
| total time = 4636.93 ms / 633 tokens |
| 0.41.042.634 I slot release: id 0 | task 403 | stop processing: n_tokens = 632, truncated = 0 |
| 0.41.042.665 I srv update_slots: all slots are idle |
| 0.41.054.733 I srv params_from_: Chat format: peg-native |
| 0.41.055.120 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.641 |
| 0.41.055.469 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.41.055.471 I reasoning-budget: deactivated (natural end) |
| 0.41.055.514 I slot launch_slot_: id 0 | task 605 | processing task, is_child = 0 |
| 0.41.055.527 W slot update_slots: id 0 | task 605 | n_past = 405, slot.prompt.tokens.size() = 632, seq_id = 0, pos_min = 631, n_swa = 0 |
| 0.41.055.529 I slot update_slots: id 0 | task 605 | Checking checkpoint with [428, 428] against 405... |
| 0.41.055.530 W slot update_slots: id 0 | task 605 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) |
| 0.41.055.533 W slot update_slots: id 0 | task 605 | erased invalidated context checkpoint (pos_min = 428, pos_max = 428, n_tokens = 429, n_swa = 0, pos_next = 0, size = 62.813 MiB) |
| 0.41.597.061 I slot create_check: id 0 | task 605 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB) |
| 0.42.459.956 I slot print_timing: id 0 | task 605 | |
| prompt eval time = 582.62 ms / 424 tokens ( 1.37 ms per token, 727.74 tokens per second) |
| eval time = 821.78 ms / 39 tokens ( 21.07 ms per token, 47.46 tokens per second) |
| total time = 1404.41 ms / 463 tokens |
| 0.42.460.047 I slot release: id 0 | task 605 | stop processing: n_tokens = 462, truncated = 0 |
| 0.42.460.077 I srv update_slots: all slots are idle |
| 0.42.475.720 I srv params_from_: Chat format: peg-native |
| 0.42.476.233 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.902 (> 0.100 thold), f_keep = 0.877 |
| 0.42.476.486 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.42.476.491 I reasoning-budget: deactivated (natural end) |
| 0.42.476.536 I slot launch_slot_: id 0 | task 646 | processing task, is_child = 0 |
| 0.42.476.550 W slot update_slots: id 0 | task 646 | n_past = 405, slot.prompt.tokens.size() = 462, seq_id = 0, pos_min = 461, n_swa = 0 |
| 0.42.476.551 I slot update_slots: id 0 | task 646 | Checking checkpoint with [419, 419] against 405... |
| 0.42.476.553 W slot update_slots: id 0 | task 646 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) |
| 0.42.476.571 W slot update_slots: id 0 | task 646 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB) |
| 0.43.126.158 I slot create_check: id 0 | task 646 | created context checkpoint 1 of 32 (pos_min = 444, pos_max = 444, n_tokens = 445, size = 62.813 MiB) |
| 0.44.988.511 I slot print_timing: id 0 | task 646 | |
| prompt eval time = 693.47 ms / 449 tokens ( 1.54 ms per token, 647.47 tokens per second) |
| eval time = 1818.46 ms / 86 tokens ( 21.14 ms per token, 47.29 tokens per second) |
| total time = 2511.93 ms / 535 tokens |
| 0.44.988.703 I slot release: id 0 | task 646 | stop processing: n_tokens = 534, truncated = 0 |
| 0.44.988.760 I srv update_slots: all slots are idle |
| 0.45.002.229 I srv params_from_: Chat format: peg-native |
| 0.45.002.631 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.948 (> 0.100 thold), f_keep = 0.758 |
| 0.45.003.110 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.45.003.114 I reasoning-budget: deactivated (natural end) |
| 0.45.003.190 I slot launch_slot_: id 0 | task 734 | processing task, is_child = 0 |
| 0.45.003.216 W slot update_slots: id 0 | task 734 | n_past = 405, slot.prompt.tokens.size() = 534, seq_id = 0, pos_min = 533, n_swa = 0 |
| 0.45.003.220 I slot update_slots: id 0 | task 734 | Checking checkpoint with [444, 444] against 405... |
| 0.45.003.221 W slot update_slots: id 0 | task 734 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) |
| 0.45.003.227 W slot update_slots: id 0 | task 734 | erased invalidated context checkpoint (pos_min = 444, pos_max = 444, n_tokens = 445, n_swa = 0, pos_next = 0, size = 62.813 MiB) |
| 0.45.561.942 I slot create_check: id 0 | task 734 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB) |
| 0.46.436.833 I slot print_timing: id 0 | task 734 | |
| prompt eval time = 595.84 ms / 427 tokens ( 1.40 ms per token, 716.64 tokens per second) |
| eval time = 837.77 ms / 39 tokens ( 21.48 ms per token, 46.55 tokens per second) |
| total time = 1433.60 ms / 466 tokens |
| 0.46.436.923 I slot release: id 0 | task 734 | stop processing: n_tokens = 465, truncated = 0 |
| 0.46.436.955 I srv update_slots: all slots are idle |
| 0.46.450.626 I srv params_from_: Chat format: peg-native |
| 0.46.451.043 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.948 (> 0.100 thold), f_keep = 0.871 |
| 0.46.451.527 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.46.451.532 I reasoning-budget: deactivated (natural end) |
| 0.46.451.617 I slot launch_slot_: id 0 | task 775 | processing task, is_child = 0 |
| 0.46.451.637 W slot update_slots: id 0 | task 775 | n_past = 405, slot.prompt.tokens.size() = 465, seq_id = 0, pos_min = 464, n_swa = 0 |
| 0.46.451.641 I slot update_slots: id 0 | task 775 | Checking checkpoint with [422, 422] against 405... |
| 0.46.451.643 W slot update_slots: id 0 | task 775 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) |
| 0.46.451.649 W slot update_slots: id 0 | task 775 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB) |
| 0.46.982.640 I slot create_check: id 0 | task 775 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB) |
| 0.47.142.299 I slot print_timing: id 0 | task 775 | |
| prompt eval time = 585.65 ms / 427 tokens ( 1.37 ms per token, 729.10 tokens per second) |
| eval time = 104.99 ms / 4 tokens ( 26.25 ms per token, 38.10 tokens per second) |
| total time = 690.65 ms / 431 tokens |
| 0.47.142.414 I slot release: id 0 | task 775 | stop processing: n_tokens = 430, truncated = 0 |
| 0.47.142.449 I srv update_slots: all slots are idle |
| 0.47.193.526 I srv params_from_: Chat format: peg-native |
| 0.47.195.446 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.958 (> 0.100 thold), f_keep = 0.944 |
| 0.47.195.987 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.47.195.992 I reasoning-budget: deactivated (natural end) |
| 0.47.196.101 I slot launch_slot_: id 0 | task 781 | processing task, is_child = 0 |
| 0.47.196.122 W slot update_slots: id 0 | task 781 | n_past = 406, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0 |
| 0.47.196.128 I slot update_slots: id 0 | task 781 | Checking checkpoint with [422, 422] against 406... |
| 0.47.196.129 W slot update_slots: id 0 | task 781 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) |
| 0.47.196.139 W slot update_slots: id 0 | task 781 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB) |
| 0.47.774.566 I slot create_check: id 0 | task 781 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB) |
| 0.48.760.957 I slot print_timing: id 0 | task 781 | |
| prompt eval time = 635.18 ms / 424 tokens ( 1.50 ms per token, 667.53 tokens per second) |
| eval time = 929.64 ms / 40 tokens ( 23.24 ms per token, 43.03 tokens per second) |
| total time = 1564.83 ms / 464 tokens |
| 0.48.761.052 I slot release: id 0 | task 781 | stop processing: n_tokens = 463, truncated = 0 |
| 0.48.761.080 I srv update_slots: all slots are idle |
| 0.48.783.508 I srv params_from_: Chat format: peg-native |
| 0.48.783.860 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.935 (> 0.100 thold), f_keep = 1.000 |
| 0.48.784.082 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.48.784.084 I reasoning-budget: deactivated (natural end) |
| 0.48.784.127 I slot launch_slot_: id 0 | task 823 | processing task, is_child = 0 |
| 0.48.935.107 I slot create_check: id 0 | task 823 | created context checkpoint 2 of 32 (pos_min = 490, pos_max = 490, n_tokens = 491, size = 62.813 MiB) |
| 0.49.277.790 I slot print_timing: id 0 | task 823 | |
| prompt eval time = 217.60 ms / 32 tokens ( 6.80 ms per token, 147.06 tokens per second) |
| eval time = 276.04 ms / 13 tokens ( 21.23 ms per token, 47.09 tokens per second) |
| total time = 493.64 ms / 45 tokens |
| 0.49.277.883 I slot release: id 0 | task 823 | stop processing: n_tokens = 507, truncated = 0 |
| 0.49.277.910 I srv update_slots: all slots are idle |
| 0.49.324.008 I srv params_from_: Chat format: peg-native |
| 0.49.324.548 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.967 (> 0.100 thold), f_keep = 0.809 |
| 0.49.324.947 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.49.324.952 I reasoning-budget: deactivated (natural end) |
| 0.49.325.029 I slot launch_slot_: id 0 | task 838 | processing task, is_child = 0 |
| 0.49.325.041 W slot update_slots: id 0 | task 838 | n_past = 410, slot.prompt.tokens.size() = 507, seq_id = 0, pos_min = 506, n_swa = 0 |
| 0.49.325.042 I slot update_slots: id 0 | task 838 | Checking checkpoint with [490, 490] against 410... |
| 0.49.325.043 I slot update_slots: id 0 | task 838 | Checking checkpoint with [419, 419] against 410... |
| 0.49.325.048 W slot update_slots: id 0 | task 838 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) |
| 0.49.325.051 W slot update_slots: id 0 | task 838 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB) |
| 0.49.328.874 W slot update_slots: id 0 | task 838 | erased invalidated context checkpoint (pos_min = 490, pos_max = 490, n_tokens = 491, n_swa = 0, pos_next = 0, size = 62.813 MiB) |
| 0.49.907.729 I slot create_check: id 0 | task 838 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB) |
| 0.50.729.586 I slot print_timing: id 0 | task 838 | |
| prompt eval time = 624.90 ms / 424 tokens ( 1.47 ms per token, 678.50 tokens per second) |
| eval time = 779.62 ms / 40 tokens ( 19.49 ms per token, 51.31 tokens per second) |
| total time = 1404.52 ms / 464 tokens |
| 0.50.729.654 I slot release: id 0 | task 838 | stop processing: n_tokens = 463, truncated = 0 |
| 0.50.729.686 I srv update_slots: all slots are idle |
| 0.50.775.784 I srv params_from_: Chat format: peg-native |
| 0.50.776.334 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.931 (> 0.100 thold), f_keep = 0.875 |
| 0.50.776.559 I reasoning-budget: activated, budget=2147483647 tokens |
| 0.50.776.562 I reasoning-budget: deactivated (natural end) |
| 0.50.776.609 I slot launch_slot_: id 0 | task 880 | processing task, is_child = 0 |
| 0.50.776.620 W slot update_slots: id 0 | task 880 | n_past = 405, slot.prompt.tokens.size() = 463, seq_id = 0, pos_min = 462, n_swa = 0 |
| 0.50.776.622 I slot update_slots: id 0 | task 880 | Checking checkpoint with [419, 419] against 405... |
| 0.50.776.623 W slot update_slots: id 0 | task 880 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) |
| 0.50.776.627 W slot update_slots: id 0 | task 880 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB) |
| 0.51.328.909 I slot create_check: id 0 | task 880 | created context checkpoint 1 of 32 (pos_min = 430, pos_max = 430, n_tokens = 431, size = 62.813 MiB) |
| 0.52.962.266 I slot print_timing: id 0 | task 880 | |
| prompt eval time = 596.05 ms / 435 tokens ( 1.37 ms per token, 729.81 tokens per second) |
| eval time = 1589.57 ms / 80 tokens ( 19.87 ms per token, 50.33 tokens per second) |
| total time = 2185.62 ms / 515 tokens |
| 0.52.962.487 I slot release: id 0 | task 880 | stop processing: n_tokens = 514, truncated = 0 |
| 0.52.962.523 I srv update_slots: all slots are idle |
| 0.52.963.931 I srv operator(): operator(): cleaning up before exit... |
|
|