Add files using upload-large-folder tool
Browse filesThis view is limited to 50 files because it contains too many changes. See raw diff
- recipe/logs/C1_convert.log +934 -0
- recipe/logs/C2_mmproj.log +362 -0
- recipe/logs/C_readback.log +2 -0
- recipe/logs/D2_verify_download.log +2 -0
- recipe/logs/N1_ppl_bf16.log +15 -0
- recipe/logs/N1c_ppl_bf16_cpu.log +15 -0
- recipe/logs/N2c_imatrix_cpu.log +114 -0
- recipe/logs/N3_q102i.log +0 -0
- recipe/logs/N3_q103i.log +0 -0
- recipe/logs/N3_q106i.log +0 -0
- recipe/logs/N3_readback.log +4 -0
- recipe/logs/N4_kld_q102i.log +92 -0
- recipe/logs/N4_kld_q103i.log +92 -0
- recipe/logs/N4_kld_q106.log +92 -0
- recipe/logs/N4_kld_q106i.log +92 -0
- recipe/logs/N4v_kld_q102i.log +93 -0
- recipe/logs/N4v_kld_q103.log +93 -0
- recipe/logs/N4v_kld_q103i.log +93 -0
- recipe/logs/N4v_kld_q106.log +93 -0
- recipe/logs/N4v_kld_q106i.log +93 -0
- recipe/logs/N5_kld_q106_repeat.log +92 -0
- recipe/logs/N5v_kld_q106_repeat.log +93 -0
- recipe/logs/N6t_tools_c1.log +33 -0
- recipe/logs/N6t_tools_tpl.log +22 -0
- recipe/logs/N6t_tools_tpl_medium.log +31 -0
- recipe/logs/N7_sizing.log +5 -0
- recipe/logs/N8_unice.log +3 -0
- recipe/logs/N8a_seats.log +8 -0
- recipe/logs/N8d_seats.log +8 -0
- recipe/logs/Q1_q102.log +0 -0
- recipe/logs/Q1_q103.log +0 -0
- recipe/logs/Q_readback.log +3 -0
- recipe/logs/Q_sizes.log +6 -0
- recipe/logs/b_n-c3-q106.log +284 -0
- recipe/logs/b_n-tools-q106-c1-probe.log +286 -0
- recipe/logs/b_n-tools-q106-c1.log +302 -0
- recipe/logs/b_n-tools-q106-roff-r3.log +302 -0
- recipe/logs/b_n-tools-q106-roff.log +301 -0
- recipe/logs/b_n-tools-q106-tpl-medium-probe.log +278 -0
- recipe/logs/b_n-tools-q106-tpl-medium.log +294 -0
- recipe/logs/b_n-tools-q106.log +302 -0
- recipe/logs/b_n-vision-q106-faoff.log +66 -0
- recipe/logs/b_n-vision-q106-faon.log +66 -0
- recipe/logs/b_n-vision-q106-roff-faon.log +70 -0
- recipe/logs/diag_bf16.log +5 -0
- recipe/logs/diag_bf16_cpu.log +13 -0
- recipe/logs/diag_bf16_purecpu_c1.log +14 -0
- recipe/logs/diag_bf16_rocm_faoff.log +13 -0
- recipe/logs/diag_bf16_vk_faon.log +13 -0
- recipe/logs/diag_ppl_q106_rocm_c4.log +13 -0
recipe/logs/C1_convert.log
ADDED
|
@@ -0,0 +1,934 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
INFO:hf-to-gguf:Loading model: hf
|
| 2 |
+
INFO:hf-to-gguf:Model architecture: Qwen3_5MoeForConditionalGeneration
|
| 3 |
+
INFO:hf-to-gguf:gguf: loading model weight map from 'model.safetensors.index.json'
|
| 4 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00001-of-00016.safetensors'
|
| 5 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00002-of-00016.safetensors'
|
| 6 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00003-of-00016.safetensors'
|
| 7 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00004-of-00016.safetensors'
|
| 8 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00005-of-00016.safetensors'
|
| 9 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00006-of-00016.safetensors'
|
| 10 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00007-of-00016.safetensors'
|
| 11 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00008-of-00016.safetensors'
|
| 12 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00009-of-00016.safetensors'
|
| 13 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00010-of-00016.safetensors'
|
| 14 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00011-of-00016.safetensors'
|
| 15 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00012-of-00016.safetensors'
|
| 16 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00013-of-00016.safetensors'
|
| 17 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00014-of-00016.safetensors'
|
| 18 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00015-of-00016.safetensors'
|
| 19 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00016-of-00016.safetensors'
|
| 20 |
+
INFO:gguf.gguf_writer:gguf: This GGUF file is for Little Endian only
|
| 21 |
+
INFO:hf-to-gguf:Exporting model...
|
| 22 |
+
INFO:hf-to-gguf:token_embd.weight, torch.bfloat16 --> BF16, shape = {2048, 248320}
|
| 23 |
+
INFO:hf-to-gguf:blk.0.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 24 |
+
INFO:hf-to-gguf:blk.0.ssm_a, torch.bfloat16 --> F32, shape = {32}
|
| 25 |
+
INFO:hf-to-gguf:blk.0.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
|
| 26 |
+
INFO:hf-to-gguf:blk.0.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
|
| 27 |
+
INFO:hf-to-gguf:blk.0.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 28 |
+
INFO:hf-to-gguf:blk.0.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 29 |
+
INFO:hf-to-gguf:blk.0.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 30 |
+
INFO:hf-to-gguf:blk.0.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
|
| 31 |
+
INFO:hf-to-gguf:blk.0.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
|
| 32 |
+
INFO:hf-to-gguf:blk.0.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 33 |
+
INFO:hf-to-gguf:blk.0.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 34 |
+
INFO:hf-to-gguf:blk.0.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 35 |
+
INFO:hf-to-gguf:blk.0.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 36 |
+
INFO:hf-to-gguf:blk.0.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 37 |
+
INFO:hf-to-gguf:blk.0.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 38 |
+
INFO:hf-to-gguf:blk.0.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 39 |
+
INFO:hf-to-gguf:blk.0.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 40 |
+
INFO:hf-to-gguf:blk.0.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 41 |
+
INFO:hf-to-gguf:blk.0.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 42 |
+
INFO:hf-to-gguf:blk.1.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 43 |
+
INFO:hf-to-gguf:blk.1.ssm_a, torch.bfloat16 --> F32, shape = {32}
|
| 44 |
+
INFO:hf-to-gguf:blk.1.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
|
| 45 |
+
INFO:hf-to-gguf:blk.1.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
|
| 46 |
+
INFO:hf-to-gguf:blk.1.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 47 |
+
INFO:hf-to-gguf:blk.1.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 48 |
+
INFO:hf-to-gguf:blk.1.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 49 |
+
INFO:hf-to-gguf:blk.1.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
|
| 50 |
+
INFO:hf-to-gguf:blk.1.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
|
| 51 |
+
INFO:hf-to-gguf:blk.1.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 52 |
+
INFO:hf-to-gguf:blk.1.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 53 |
+
INFO:hf-to-gguf:blk.1.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 54 |
+
INFO:hf-to-gguf:blk.1.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 55 |
+
INFO:hf-to-gguf:blk.1.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 56 |
+
INFO:hf-to-gguf:blk.1.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 57 |
+
INFO:hf-to-gguf:blk.1.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 58 |
+
INFO:hf-to-gguf:blk.1.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 59 |
+
INFO:hf-to-gguf:blk.1.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 60 |
+
INFO:hf-to-gguf:blk.1.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 61 |
+
INFO:hf-to-gguf:blk.2.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 62 |
+
INFO:hf-to-gguf:blk.2.ssm_a, torch.bfloat16 --> F32, shape = {32}
|
| 63 |
+
INFO:hf-to-gguf:blk.2.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
|
| 64 |
+
INFO:hf-to-gguf:blk.2.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
|
| 65 |
+
INFO:hf-to-gguf:blk.2.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 66 |
+
INFO:hf-to-gguf:blk.2.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 67 |
+
INFO:hf-to-gguf:blk.2.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 68 |
+
INFO:hf-to-gguf:blk.2.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
|
| 69 |
+
INFO:hf-to-gguf:blk.2.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
|
| 70 |
+
INFO:hf-to-gguf:blk.2.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 71 |
+
INFO:hf-to-gguf:blk.2.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 72 |
+
INFO:hf-to-gguf:blk.2.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 73 |
+
INFO:hf-to-gguf:blk.2.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 74 |
+
INFO:hf-to-gguf:blk.2.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 75 |
+
INFO:hf-to-gguf:blk.2.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 76 |
+
INFO:hf-to-gguf:blk.2.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 77 |
+
INFO:hf-to-gguf:blk.2.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 78 |
+
INFO:hf-to-gguf:blk.2.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 79 |
+
INFO:hf-to-gguf:blk.2.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 80 |
+
INFO:hf-to-gguf:blk.3.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 81 |
+
INFO:hf-to-gguf:blk.3.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 82 |
+
INFO:hf-to-gguf:blk.3.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 83 |
+
INFO:hf-to-gguf:blk.3.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 84 |
+
INFO:hf-to-gguf:blk.3.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 85 |
+
INFO:hf-to-gguf:blk.3.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 86 |
+
INFO:hf-to-gguf:blk.3.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 87 |
+
INFO:hf-to-gguf:blk.3.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 88 |
+
INFO:hf-to-gguf:blk.3.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 89 |
+
INFO:hf-to-gguf:blk.3.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 90 |
+
INFO:hf-to-gguf:blk.3.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
|
| 91 |
+
INFO:hf-to-gguf:blk.3.attn_k.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 92 |
+
INFO:hf-to-gguf:blk.3.attn_output.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 93 |
+
INFO:hf-to-gguf:blk.3.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
|
| 94 |
+
INFO:hf-to-gguf:blk.3.attn_q.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 95 |
+
INFO:hf-to-gguf:blk.3.attn_v.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 96 |
+
INFO:hf-to-gguf:blk.4.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 97 |
+
INFO:hf-to-gguf:blk.4.ssm_a, torch.bfloat16 --> F32, shape = {32}
|
| 98 |
+
INFO:hf-to-gguf:blk.4.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
|
| 99 |
+
INFO:hf-to-gguf:blk.4.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
|
| 100 |
+
INFO:hf-to-gguf:blk.4.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 101 |
+
INFO:hf-to-gguf:blk.4.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 102 |
+
INFO:hf-to-gguf:blk.4.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 103 |
+
INFO:hf-to-gguf:blk.4.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
|
| 104 |
+
INFO:hf-to-gguf:blk.4.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
|
| 105 |
+
INFO:hf-to-gguf:blk.4.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 106 |
+
INFO:hf-to-gguf:blk.4.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 107 |
+
INFO:hf-to-gguf:blk.4.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 108 |
+
INFO:hf-to-gguf:blk.4.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 109 |
+
INFO:hf-to-gguf:blk.4.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 110 |
+
INFO:hf-to-gguf:blk.4.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 111 |
+
INFO:hf-to-gguf:blk.4.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 112 |
+
INFO:hf-to-gguf:blk.4.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 113 |
+
INFO:hf-to-gguf:blk.4.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 114 |
+
INFO:hf-to-gguf:blk.4.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 115 |
+
INFO:hf-to-gguf:blk.5.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 116 |
+
INFO:hf-to-gguf:blk.5.ssm_a, torch.bfloat16 --> F32, shape = {32}
|
| 117 |
+
INFO:hf-to-gguf:blk.5.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
|
| 118 |
+
INFO:hf-to-gguf:blk.5.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
|
| 119 |
+
INFO:hf-to-gguf:blk.5.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 120 |
+
INFO:hf-to-gguf:blk.5.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 121 |
+
INFO:hf-to-gguf:blk.5.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 122 |
+
INFO:hf-to-gguf:blk.5.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
|
| 123 |
+
INFO:hf-to-gguf:blk.5.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
|
| 124 |
+
INFO:hf-to-gguf:blk.5.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 125 |
+
INFO:hf-to-gguf:blk.5.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 126 |
+
INFO:hf-to-gguf:blk.5.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 127 |
+
INFO:hf-to-gguf:blk.5.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 128 |
+
INFO:hf-to-gguf:blk.5.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 129 |
+
INFO:hf-to-gguf:blk.5.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 130 |
+
INFO:hf-to-gguf:blk.5.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 131 |
+
INFO:hf-to-gguf:blk.5.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 132 |
+
INFO:hf-to-gguf:blk.5.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 133 |
+
INFO:hf-to-gguf:blk.5.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 134 |
+
INFO:hf-to-gguf:blk.6.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 135 |
+
INFO:hf-to-gguf:blk.6.ssm_a, torch.bfloat16 --> F32, shape = {32}
|
| 136 |
+
INFO:hf-to-gguf:blk.6.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
|
| 137 |
+
INFO:hf-to-gguf:blk.6.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
|
| 138 |
+
INFO:hf-to-gguf:blk.6.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 139 |
+
INFO:hf-to-gguf:blk.6.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 140 |
+
INFO:hf-to-gguf:blk.6.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 141 |
+
INFO:hf-to-gguf:blk.6.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
|
| 142 |
+
INFO:hf-to-gguf:blk.6.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
|
| 143 |
+
INFO:hf-to-gguf:blk.6.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 144 |
+
INFO:hf-to-gguf:blk.6.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 145 |
+
INFO:hf-to-gguf:blk.6.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 146 |
+
INFO:hf-to-gguf:blk.6.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 147 |
+
INFO:hf-to-gguf:blk.6.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 148 |
+
INFO:hf-to-gguf:blk.6.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 149 |
+
INFO:hf-to-gguf:blk.6.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 150 |
+
INFO:hf-to-gguf:blk.6.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 151 |
+
INFO:hf-to-gguf:blk.6.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 152 |
+
INFO:hf-to-gguf:blk.6.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 153 |
+
INFO:hf-to-gguf:blk.7.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 154 |
+
INFO:hf-to-gguf:blk.7.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 155 |
+
INFO:hf-to-gguf:blk.7.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 156 |
+
INFO:hf-to-gguf:blk.7.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 157 |
+
INFO:hf-to-gguf:blk.7.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 158 |
+
INFO:hf-to-gguf:blk.7.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 159 |
+
INFO:hf-to-gguf:blk.7.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 160 |
+
INFO:hf-to-gguf:blk.7.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 161 |
+
INFO:hf-to-gguf:blk.7.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 162 |
+
INFO:hf-to-gguf:blk.7.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 163 |
+
INFO:hf-to-gguf:blk.7.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
|
| 164 |
+
INFO:hf-to-gguf:blk.7.attn_k.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 165 |
+
INFO:hf-to-gguf:blk.7.attn_output.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 166 |
+
INFO:hf-to-gguf:blk.7.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
|
| 167 |
+
INFO:hf-to-gguf:blk.7.attn_q.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 168 |
+
INFO:hf-to-gguf:blk.7.attn_v.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 169 |
+
INFO:hf-to-gguf:blk.8.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 170 |
+
INFO:hf-to-gguf:blk.8.ssm_a, torch.bfloat16 --> F32, shape = {32}
|
| 171 |
+
INFO:hf-to-gguf:blk.8.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
|
| 172 |
+
INFO:hf-to-gguf:blk.8.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
|
| 173 |
+
INFO:hf-to-gguf:blk.8.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 174 |
+
INFO:hf-to-gguf:blk.8.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 175 |
+
INFO:hf-to-gguf:blk.8.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 176 |
+
INFO:hf-to-gguf:blk.8.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
|
| 177 |
+
INFO:hf-to-gguf:blk.8.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
|
| 178 |
+
INFO:hf-to-gguf:blk.8.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 179 |
+
INFO:hf-to-gguf:blk.8.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 180 |
+
INFO:hf-to-gguf:blk.8.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 181 |
+
INFO:hf-to-gguf:blk.8.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 182 |
+
INFO:hf-to-gguf:blk.8.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 183 |
+
INFO:hf-to-gguf:blk.8.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 184 |
+
INFO:hf-to-gguf:blk.8.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 185 |
+
INFO:hf-to-gguf:blk.8.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 186 |
+
INFO:hf-to-gguf:blk.8.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 187 |
+
INFO:hf-to-gguf:blk.8.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 188 |
+
INFO:hf-to-gguf:blk.9.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 189 |
+
INFO:hf-to-gguf:blk.9.ssm_a, torch.bfloat16 --> F32, shape = {32}
|
| 190 |
+
INFO:hf-to-gguf:blk.9.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
|
| 191 |
+
INFO:hf-to-gguf:blk.9.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
|
| 192 |
+
INFO:hf-to-gguf:blk.9.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 193 |
+
INFO:hf-to-gguf:blk.9.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 194 |
+
INFO:hf-to-gguf:blk.9.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 195 |
+
INFO:hf-to-gguf:blk.9.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
|
| 196 |
+
INFO:hf-to-gguf:blk.9.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
|
| 197 |
+
INFO:hf-to-gguf:blk.9.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 198 |
+
INFO:hf-to-gguf:blk.9.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 199 |
+
INFO:hf-to-gguf:blk.9.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 200 |
+
INFO:hf-to-gguf:blk.9.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 201 |
+
INFO:hf-to-gguf:blk.9.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 202 |
+
INFO:hf-to-gguf:blk.9.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 203 |
+
INFO:hf-to-gguf:blk.9.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 204 |
+
INFO:hf-to-gguf:blk.9.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 205 |
+
INFO:hf-to-gguf:blk.10.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 206 |
+
INFO:hf-to-gguf:blk.10.ssm_a, torch.bfloat16 --> F32, shape = {32}
|
| 207 |
+
INFO:hf-to-gguf:blk.10.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
|
| 208 |
+
INFO:hf-to-gguf:blk.10.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
|
| 209 |
+
INFO:hf-to-gguf:blk.10.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 210 |
+
INFO:hf-to-gguf:blk.10.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 211 |
+
INFO:hf-to-gguf:blk.10.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 212 |
+
INFO:hf-to-gguf:blk.10.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
|
| 213 |
+
INFO:hf-to-gguf:blk.10.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
|
| 214 |
+
INFO:hf-to-gguf:blk.10.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 215 |
+
INFO:hf-to-gguf:blk.10.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 216 |
+
INFO:hf-to-gguf:blk.10.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 217 |
+
INFO:hf-to-gguf:blk.10.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 218 |
+
INFO:hf-to-gguf:blk.10.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 219 |
+
INFO:hf-to-gguf:blk.10.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 220 |
+
INFO:hf-to-gguf:blk.10.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 221 |
+
INFO:hf-to-gguf:blk.10.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 222 |
+
INFO:hf-to-gguf:blk.10.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 223 |
+
INFO:hf-to-gguf:blk.10.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 224 |
+
INFO:hf-to-gguf:blk.11.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 225 |
+
INFO:hf-to-gguf:blk.11.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 226 |
+
INFO:hf-to-gguf:blk.11.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 227 |
+
INFO:hf-to-gguf:blk.11.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 228 |
+
INFO:hf-to-gguf:blk.11.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 229 |
+
INFO:hf-to-gguf:blk.11.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 230 |
+
INFO:hf-to-gguf:blk.11.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 231 |
+
INFO:hf-to-gguf:blk.11.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 232 |
+
INFO:hf-to-gguf:blk.11.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 233 |
+
INFO:hf-to-gguf:blk.11.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 234 |
+
INFO:hf-to-gguf:blk.11.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
|
| 235 |
+
INFO:hf-to-gguf:blk.11.attn_k.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 236 |
+
INFO:hf-to-gguf:blk.11.attn_output.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 237 |
+
INFO:hf-to-gguf:blk.11.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
|
| 238 |
+
INFO:hf-to-gguf:blk.11.attn_q.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 239 |
+
INFO:hf-to-gguf:blk.11.attn_v.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 240 |
+
INFO:hf-to-gguf:blk.12.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 241 |
+
INFO:hf-to-gguf:blk.12.ssm_a, torch.bfloat16 --> F32, shape = {32}
|
| 242 |
+
INFO:hf-to-gguf:blk.12.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
|
| 243 |
+
INFO:hf-to-gguf:blk.12.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
|
| 244 |
+
INFO:hf-to-gguf:blk.12.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 245 |
+
INFO:hf-to-gguf:blk.12.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 246 |
+
INFO:hf-to-gguf:blk.12.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 247 |
+
INFO:hf-to-gguf:blk.12.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
|
| 248 |
+
INFO:hf-to-gguf:blk.12.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
|
| 249 |
+
INFO:hf-to-gguf:blk.12.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 250 |
+
INFO:hf-to-gguf:blk.12.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 251 |
+
INFO:hf-to-gguf:blk.12.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 252 |
+
INFO:hf-to-gguf:blk.12.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 253 |
+
INFO:hf-to-gguf:blk.12.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 254 |
+
INFO:hf-to-gguf:blk.12.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 255 |
+
INFO:hf-to-gguf:blk.9.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 256 |
+
INFO:hf-to-gguf:blk.9.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 257 |
+
INFO:hf-to-gguf:blk.12.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 258 |
+
INFO:hf-to-gguf:blk.12.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 259 |
+
INFO:hf-to-gguf:blk.12.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 260 |
+
INFO:hf-to-gguf:blk.12.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 261 |
+
INFO:hf-to-gguf:blk.13.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 262 |
+
INFO:hf-to-gguf:blk.13.ssm_a, torch.bfloat16 --> F32, shape = {32}
|
| 263 |
+
INFO:hf-to-gguf:blk.13.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
|
| 264 |
+
INFO:hf-to-gguf:blk.13.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
|
| 265 |
+
INFO:hf-to-gguf:blk.13.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 266 |
+
INFO:hf-to-gguf:blk.13.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 267 |
+
INFO:hf-to-gguf:blk.13.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 268 |
+
INFO:hf-to-gguf:blk.13.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
|
| 269 |
+
INFO:hf-to-gguf:blk.13.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
|
| 270 |
+
INFO:hf-to-gguf:blk.13.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 271 |
+
INFO:hf-to-gguf:blk.13.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 272 |
+
INFO:hf-to-gguf:blk.13.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 273 |
+
INFO:hf-to-gguf:blk.13.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 274 |
+
INFO:hf-to-gguf:blk.13.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 275 |
+
INFO:hf-to-gguf:blk.13.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 276 |
+
INFO:hf-to-gguf:blk.13.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 277 |
+
INFO:hf-to-gguf:blk.13.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 278 |
+
INFO:hf-to-gguf:blk.13.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 279 |
+
INFO:hf-to-gguf:blk.13.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 280 |
+
INFO:hf-to-gguf:blk.14.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 281 |
+
INFO:hf-to-gguf:blk.14.ssm_a, torch.bfloat16 --> F32, shape = {32}
|
| 282 |
+
INFO:hf-to-gguf:blk.14.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
|
| 283 |
+
INFO:hf-to-gguf:blk.14.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
|
| 284 |
+
INFO:hf-to-gguf:blk.14.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 285 |
+
INFO:hf-to-gguf:blk.14.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 286 |
+
INFO:hf-to-gguf:blk.14.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 287 |
+
INFO:hf-to-gguf:blk.14.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
|
| 288 |
+
INFO:hf-to-gguf:blk.14.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
|
| 289 |
+
INFO:hf-to-gguf:blk.14.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 290 |
+
INFO:hf-to-gguf:blk.14.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 291 |
+
INFO:hf-to-gguf:blk.14.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 292 |
+
INFO:hf-to-gguf:blk.14.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 293 |
+
INFO:hf-to-gguf:blk.14.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 294 |
+
INFO:hf-to-gguf:blk.14.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 295 |
+
INFO:hf-to-gguf:blk.14.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 296 |
+
INFO:hf-to-gguf:blk.14.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 297 |
+
INFO:hf-to-gguf:blk.14.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 298 |
+
INFO:hf-to-gguf:blk.14.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 299 |
+
INFO:hf-to-gguf:blk.15.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 300 |
+
INFO:hf-to-gguf:blk.15.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 301 |
+
INFO:hf-to-gguf:blk.15.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 302 |
+
INFO:hf-to-gguf:blk.15.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 303 |
+
INFO:hf-to-gguf:blk.15.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 304 |
+
INFO:hf-to-gguf:blk.15.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 305 |
+
INFO:hf-to-gguf:blk.15.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 306 |
+
INFO:hf-to-gguf:blk.15.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 307 |
+
INFO:hf-to-gguf:blk.15.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 308 |
+
INFO:hf-to-gguf:blk.15.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 309 |
+
INFO:hf-to-gguf:blk.15.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
|
| 310 |
+
INFO:hf-to-gguf:blk.15.attn_k.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 311 |
+
INFO:hf-to-gguf:blk.15.attn_output.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 312 |
+
INFO:hf-to-gguf:blk.15.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
|
| 313 |
+
INFO:hf-to-gguf:blk.15.attn_q.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 314 |
+
INFO:hf-to-gguf:blk.15.attn_v.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 315 |
+
INFO:hf-to-gguf:blk.16.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 316 |
+
INFO:hf-to-gguf:blk.16.ssm_a, torch.bfloat16 --> F32, shape = {32}
|
| 317 |
+
INFO:hf-to-gguf:blk.16.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
|
| 318 |
+
INFO:hf-to-gguf:blk.16.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
|
| 319 |
+
INFO:hf-to-gguf:blk.16.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 320 |
+
INFO:hf-to-gguf:blk.16.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 321 |
+
INFO:hf-to-gguf:blk.16.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 322 |
+
INFO:hf-to-gguf:blk.16.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
|
| 323 |
+
INFO:hf-to-gguf:blk.16.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
|
| 324 |
+
INFO:hf-to-gguf:blk.16.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 325 |
+
INFO:hf-to-gguf:blk.16.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 326 |
+
INFO:hf-to-gguf:blk.16.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 327 |
+
INFO:hf-to-gguf:blk.16.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 328 |
+
INFO:hf-to-gguf:blk.16.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 329 |
+
INFO:hf-to-gguf:blk.16.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 330 |
+
INFO:hf-to-gguf:blk.16.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 331 |
+
INFO:hf-to-gguf:blk.16.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 332 |
+
INFO:hf-to-gguf:blk.16.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 333 |
+
INFO:hf-to-gguf:blk.16.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 334 |
+
INFO:hf-to-gguf:blk.17.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 335 |
+
INFO:hf-to-gguf:blk.17.ssm_a, torch.bfloat16 --> F32, shape = {32}
|
| 336 |
+
INFO:hf-to-gguf:blk.17.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
|
| 337 |
+
INFO:hf-to-gguf:blk.17.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
|
| 338 |
+
INFO:hf-to-gguf:blk.17.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 339 |
+
INFO:hf-to-gguf:blk.17.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 340 |
+
INFO:hf-to-gguf:blk.17.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 341 |
+
INFO:hf-to-gguf:blk.17.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
|
| 342 |
+
INFO:hf-to-gguf:blk.17.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
|
| 343 |
+
INFO:hf-to-gguf:blk.17.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 344 |
+
INFO:hf-to-gguf:blk.17.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 345 |
+
INFO:hf-to-gguf:blk.17.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 346 |
+
INFO:hf-to-gguf:blk.17.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 347 |
+
INFO:hf-to-gguf:blk.17.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 348 |
+
INFO:hf-to-gguf:blk.17.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 349 |
+
INFO:hf-to-gguf:blk.17.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 350 |
+
INFO:hf-to-gguf:blk.17.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 351 |
+
INFO:hf-to-gguf:blk.17.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 352 |
+
INFO:hf-to-gguf:blk.17.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 353 |
+
INFO:hf-to-gguf:blk.18.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 354 |
+
INFO:hf-to-gguf:blk.18.ssm_a, torch.bfloat16 --> F32, shape = {32}
|
| 355 |
+
INFO:hf-to-gguf:blk.18.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
|
| 356 |
+
INFO:hf-to-gguf:blk.18.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
|
| 357 |
+
INFO:hf-to-gguf:blk.18.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 358 |
+
INFO:hf-to-gguf:blk.18.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 359 |
+
INFO:hf-to-gguf:blk.18.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 360 |
+
INFO:hf-to-gguf:blk.18.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
|
| 361 |
+
INFO:hf-to-gguf:blk.18.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
|
| 362 |
+
INFO:hf-to-gguf:blk.18.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 363 |
+
INFO:hf-to-gguf:blk.18.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 364 |
+
INFO:hf-to-gguf:blk.18.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 365 |
+
INFO:hf-to-gguf:blk.18.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 366 |
+
INFO:hf-to-gguf:blk.18.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 367 |
+
INFO:hf-to-gguf:blk.18.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 368 |
+
INFO:hf-to-gguf:blk.18.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 369 |
+
INFO:hf-to-gguf:blk.18.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 370 |
+
INFO:hf-to-gguf:blk.18.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 371 |
+
INFO:hf-to-gguf:blk.18.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 372 |
+
INFO:hf-to-gguf:blk.19.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 373 |
+
INFO:hf-to-gguf:blk.19.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 374 |
+
INFO:hf-to-gguf:blk.19.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 375 |
+
INFO:hf-to-gguf:blk.19.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 376 |
+
INFO:hf-to-gguf:blk.19.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 377 |
+
INFO:hf-to-gguf:blk.19.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 378 |
+
INFO:hf-to-gguf:blk.19.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 379 |
+
INFO:hf-to-gguf:blk.19.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 380 |
+
INFO:hf-to-gguf:blk.19.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 381 |
+
INFO:hf-to-gguf:blk.19.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 382 |
+
INFO:hf-to-gguf:blk.19.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
|
| 383 |
+
INFO:hf-to-gguf:blk.19.attn_k.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 384 |
+
INFO:hf-to-gguf:blk.19.attn_output.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 385 |
+
INFO:hf-to-gguf:blk.19.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
|
| 386 |
+
INFO:hf-to-gguf:blk.19.attn_q.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 387 |
+
INFO:hf-to-gguf:blk.19.attn_v.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 388 |
+
INFO:hf-to-gguf:blk.20.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 389 |
+
INFO:hf-to-gguf:blk.20.ssm_a, torch.bfloat16 --> F32, shape = {32}
|
| 390 |
+
INFO:hf-to-gguf:blk.20.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
|
| 391 |
+
INFO:hf-to-gguf:blk.20.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
|
| 392 |
+
INFO:hf-to-gguf:blk.20.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 393 |
+
INFO:hf-to-gguf:blk.20.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 394 |
+
INFO:hf-to-gguf:blk.20.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 395 |
+
INFO:hf-to-gguf:blk.20.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
|
| 396 |
+
INFO:hf-to-gguf:blk.20.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
|
| 397 |
+
INFO:hf-to-gguf:blk.20.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 398 |
+
INFO:hf-to-gguf:blk.20.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 399 |
+
INFO:hf-to-gguf:blk.20.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 400 |
+
INFO:hf-to-gguf:blk.20.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 401 |
+
INFO:hf-to-gguf:blk.20.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 402 |
+
INFO:hf-to-gguf:blk.20.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 403 |
+
INFO:hf-to-gguf:blk.20.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 404 |
+
INFO:hf-to-gguf:blk.20.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 405 |
+
INFO:hf-to-gguf:blk.20.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 406 |
+
INFO:hf-to-gguf:blk.20.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 407 |
+
INFO:hf-to-gguf:blk.21.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 408 |
+
INFO:hf-to-gguf:blk.21.ssm_a, torch.bfloat16 --> F32, shape = {32}
|
| 409 |
+
INFO:hf-to-gguf:blk.21.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
|
| 410 |
+
INFO:hf-to-gguf:blk.21.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
|
| 411 |
+
INFO:hf-to-gguf:blk.21.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 412 |
+
INFO:hf-to-gguf:blk.21.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 413 |
+
INFO:hf-to-gguf:blk.21.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 414 |
+
INFO:hf-to-gguf:blk.21.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
|
| 415 |
+
INFO:hf-to-gguf:blk.21.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
|
| 416 |
+
INFO:hf-to-gguf:blk.21.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 417 |
+
INFO:hf-to-gguf:blk.21.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 418 |
+
INFO:hf-to-gguf:blk.21.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 419 |
+
INFO:hf-to-gguf:blk.21.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 420 |
+
INFO:hf-to-gguf:blk.21.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 421 |
+
INFO:hf-to-gguf:blk.21.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 422 |
+
INFO:hf-to-gguf:blk.21.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 423 |
+
INFO:hf-to-gguf:blk.21.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 424 |
+
INFO:hf-to-gguf:blk.21.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 425 |
+
INFO:hf-to-gguf:blk.21.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 426 |
+
INFO:hf-to-gguf:blk.22.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 427 |
+
INFO:hf-to-gguf:blk.22.ssm_a, torch.bfloat16 --> F32, shape = {32}
|
| 428 |
+
INFO:hf-to-gguf:blk.22.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
|
| 429 |
+
INFO:hf-to-gguf:blk.22.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
|
| 430 |
+
INFO:hf-to-gguf:blk.22.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 431 |
+
INFO:hf-to-gguf:blk.22.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 432 |
+
INFO:hf-to-gguf:blk.22.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 433 |
+
INFO:hf-to-gguf:blk.22.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
|
| 434 |
+
INFO:hf-to-gguf:blk.22.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
|
| 435 |
+
INFO:hf-to-gguf:blk.22.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 436 |
+
INFO:hf-to-gguf:blk.22.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 437 |
+
INFO:hf-to-gguf:blk.22.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 438 |
+
INFO:hf-to-gguf:blk.22.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 439 |
+
INFO:hf-to-gguf:blk.22.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 440 |
+
INFO:hf-to-gguf:blk.22.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 441 |
+
INFO:hf-to-gguf:blk.22.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 442 |
+
INFO:hf-to-gguf:blk.22.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 443 |
+
INFO:hf-to-gguf:blk.22.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 444 |
+
INFO:hf-to-gguf:blk.22.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 445 |
+
INFO:hf-to-gguf:blk.23.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 446 |
+
INFO:hf-to-gguf:blk.23.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 447 |
+
INFO:hf-to-gguf:blk.23.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 448 |
+
INFO:hf-to-gguf:blk.23.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 449 |
+
INFO:hf-to-gguf:blk.23.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 450 |
+
INFO:hf-to-gguf:blk.23.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 451 |
+
INFO:hf-to-gguf:blk.23.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 452 |
+
INFO:hf-to-gguf:blk.23.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 453 |
+
INFO:hf-to-gguf:blk.23.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 454 |
+
INFO:hf-to-gguf:blk.23.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 455 |
+
INFO:hf-to-gguf:blk.23.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
|
| 456 |
+
INFO:hf-to-gguf:blk.23.attn_k.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 457 |
+
INFO:hf-to-gguf:blk.23.attn_output.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 458 |
+
INFO:hf-to-gguf:blk.23.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
|
| 459 |
+
INFO:hf-to-gguf:blk.23.attn_q.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 460 |
+
INFO:hf-to-gguf:blk.23.attn_v.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 461 |
+
INFO:hf-to-gguf:blk.24.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 462 |
+
INFO:hf-to-gguf:blk.24.ssm_a, torch.bfloat16 --> F32, shape = {32}
|
| 463 |
+
INFO:hf-to-gguf:blk.24.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
|
| 464 |
+
INFO:hf-to-gguf:blk.24.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
|
| 465 |
+
INFO:hf-to-gguf:blk.24.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 466 |
+
INFO:hf-to-gguf:blk.24.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 467 |
+
INFO:hf-to-gguf:blk.24.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 468 |
+
INFO:hf-to-gguf:blk.24.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
|
| 469 |
+
INFO:hf-to-gguf:blk.24.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
|
| 470 |
+
INFO:hf-to-gguf:blk.24.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 471 |
+
INFO:hf-to-gguf:blk.24.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 472 |
+
INFO:hf-to-gguf:blk.24.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 473 |
+
INFO:hf-to-gguf:blk.24.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 474 |
+
INFO:hf-to-gguf:blk.24.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 475 |
+
INFO:hf-to-gguf:blk.24.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 476 |
+
INFO:hf-to-gguf:blk.24.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 477 |
+
INFO:hf-to-gguf:blk.24.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 478 |
+
INFO:hf-to-gguf:blk.24.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 479 |
+
INFO:hf-to-gguf:blk.24.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 480 |
+
INFO:hf-to-gguf:blk.25.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 481 |
+
INFO:hf-to-gguf:blk.25.ssm_a, torch.bfloat16 --> F32, shape = {32}
|
| 482 |
+
INFO:hf-to-gguf:blk.25.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
|
| 483 |
+
INFO:hf-to-gguf:blk.25.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
|
| 484 |
+
INFO:hf-to-gguf:blk.25.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 485 |
+
INFO:hf-to-gguf:blk.25.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 486 |
+
INFO:hf-to-gguf:blk.25.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 487 |
+
INFO:hf-to-gguf:blk.25.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
|
| 488 |
+
INFO:hf-to-gguf:blk.25.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
|
| 489 |
+
INFO:hf-to-gguf:blk.25.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 490 |
+
INFO:hf-to-gguf:blk.25.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 491 |
+
INFO:hf-to-gguf:blk.25.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 492 |
+
INFO:hf-to-gguf:blk.25.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 493 |
+
INFO:hf-to-gguf:blk.25.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 494 |
+
INFO:hf-to-gguf:blk.25.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 495 |
+
INFO:hf-to-gguf:blk.25.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 496 |
+
INFO:hf-to-gguf:blk.25.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 497 |
+
INFO:hf-to-gguf:blk.25.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 498 |
+
INFO:hf-to-gguf:blk.25.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 499 |
+
INFO:hf-to-gguf:blk.26.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 500 |
+
INFO:hf-to-gguf:blk.26.ssm_a, torch.bfloat16 --> F32, shape = {32}
|
| 501 |
+
INFO:hf-to-gguf:blk.26.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
|
| 502 |
+
INFO:hf-to-gguf:blk.26.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
|
| 503 |
+
INFO:hf-to-gguf:blk.26.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 504 |
+
INFO:hf-to-gguf:blk.26.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 505 |
+
INFO:hf-to-gguf:blk.26.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 506 |
+
INFO:hf-to-gguf:blk.26.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
|
| 507 |
+
INFO:hf-to-gguf:blk.26.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
|
| 508 |
+
INFO:hf-to-gguf:blk.26.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 509 |
+
INFO:hf-to-gguf:blk.26.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 510 |
+
INFO:hf-to-gguf:blk.26.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 511 |
+
INFO:hf-to-gguf:blk.26.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 512 |
+
INFO:hf-to-gguf:blk.26.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 513 |
+
INFO:hf-to-gguf:blk.26.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 514 |
+
INFO:hf-to-gguf:blk.26.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 515 |
+
INFO:hf-to-gguf:blk.26.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 516 |
+
INFO:hf-to-gguf:blk.26.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 517 |
+
INFO:hf-to-gguf:blk.26.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 518 |
+
INFO:hf-to-gguf:blk.27.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 519 |
+
INFO:hf-to-gguf:blk.27.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 520 |
+
INFO:hf-to-gguf:blk.27.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 521 |
+
INFO:hf-to-gguf:blk.27.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 522 |
+
INFO:hf-to-gguf:blk.27.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 523 |
+
INFO:hf-to-gguf:blk.27.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 524 |
+
INFO:hf-to-gguf:blk.27.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 525 |
+
INFO:hf-to-gguf:blk.27.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 526 |
+
INFO:hf-to-gguf:blk.27.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 527 |
+
INFO:hf-to-gguf:blk.27.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 528 |
+
INFO:hf-to-gguf:blk.27.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
|
| 529 |
+
INFO:hf-to-gguf:blk.27.attn_k.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 530 |
+
INFO:hf-to-gguf:blk.27.attn_output.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 531 |
+
INFO:hf-to-gguf:blk.27.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
|
| 532 |
+
INFO:hf-to-gguf:blk.27.attn_q.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 533 |
+
INFO:hf-to-gguf:blk.27.attn_v.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 534 |
+
INFO:hf-to-gguf:blk.28.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 535 |
+
INFO:hf-to-gguf:blk.28.ssm_a, torch.bfloat16 --> F32, shape = {32}
|
| 536 |
+
INFO:hf-to-gguf:blk.28.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
|
| 537 |
+
INFO:hf-to-gguf:blk.28.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
|
| 538 |
+
INFO:hf-to-gguf:blk.28.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 539 |
+
INFO:hf-to-gguf:blk.28.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 540 |
+
INFO:hf-to-gguf:blk.28.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 541 |
+
INFO:hf-to-gguf:blk.28.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
|
| 542 |
+
INFO:hf-to-gguf:blk.28.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
|
| 543 |
+
INFO:hf-to-gguf:blk.28.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 544 |
+
INFO:hf-to-gguf:blk.28.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 545 |
+
INFO:hf-to-gguf:blk.28.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 546 |
+
INFO:hf-to-gguf:blk.28.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 547 |
+
INFO:hf-to-gguf:blk.28.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 548 |
+
INFO:hf-to-gguf:blk.28.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 549 |
+
INFO:hf-to-gguf:blk.28.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 550 |
+
INFO:hf-to-gguf:blk.28.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 551 |
+
INFO:hf-to-gguf:blk.28.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 552 |
+
INFO:hf-to-gguf:blk.28.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 553 |
+
INFO:hf-to-gguf:blk.29.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 554 |
+
INFO:hf-to-gguf:blk.29.ssm_a, torch.bfloat16 --> F32, shape = {32}
|
| 555 |
+
INFO:hf-to-gguf:blk.29.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
|
| 556 |
+
INFO:hf-to-gguf:blk.29.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
|
| 557 |
+
INFO:hf-to-gguf:blk.29.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 558 |
+
INFO:hf-to-gguf:blk.29.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 559 |
+
INFO:hf-to-gguf:blk.29.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 560 |
+
INFO:hf-to-gguf:blk.29.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
|
| 561 |
+
INFO:hf-to-gguf:blk.29.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
|
| 562 |
+
INFO:hf-to-gguf:blk.29.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 563 |
+
INFO:hf-to-gguf:blk.29.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 564 |
+
INFO:hf-to-gguf:blk.29.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 565 |
+
INFO:hf-to-gguf:blk.29.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 566 |
+
INFO:hf-to-gguf:blk.29.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 567 |
+
INFO:hf-to-gguf:blk.29.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 568 |
+
INFO:hf-to-gguf:blk.29.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 569 |
+
INFO:hf-to-gguf:blk.29.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 570 |
+
INFO:hf-to-gguf:blk.29.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 571 |
+
INFO:hf-to-gguf:blk.29.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 572 |
+
INFO:hf-to-gguf:blk.30.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 573 |
+
INFO:hf-to-gguf:blk.30.ssm_a, torch.bfloat16 --> F32, shape = {32}
|
| 574 |
+
INFO:hf-to-gguf:blk.30.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
|
| 575 |
+
INFO:hf-to-gguf:blk.30.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
|
| 576 |
+
INFO:hf-to-gguf:blk.30.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 577 |
+
INFO:hf-to-gguf:blk.30.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 578 |
+
INFO:hf-to-gguf:blk.30.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 579 |
+
INFO:hf-to-gguf:blk.30.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
|
| 580 |
+
INFO:hf-to-gguf:blk.30.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
|
| 581 |
+
INFO:hf-to-gguf:blk.30.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 582 |
+
INFO:hf-to-gguf:blk.30.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 583 |
+
INFO:hf-to-gguf:blk.30.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 584 |
+
INFO:hf-to-gguf:blk.30.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 585 |
+
INFO:hf-to-gguf:blk.30.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 586 |
+
INFO:hf-to-gguf:blk.30.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 587 |
+
INFO:hf-to-gguf:blk.30.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 588 |
+
INFO:hf-to-gguf:blk.30.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 589 |
+
INFO:hf-to-gguf:blk.30.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 590 |
+
INFO:hf-to-gguf:blk.30.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 591 |
+
INFO:hf-to-gguf:blk.31.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 592 |
+
INFO:hf-to-gguf:blk.31.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 593 |
+
INFO:hf-to-gguf:blk.31.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 594 |
+
INFO:hf-to-gguf:blk.31.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 595 |
+
INFO:hf-to-gguf:blk.31.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 596 |
+
INFO:hf-to-gguf:blk.31.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 597 |
+
INFO:hf-to-gguf:blk.31.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 598 |
+
INFO:hf-to-gguf:blk.31.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 599 |
+
INFO:hf-to-gguf:blk.31.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 600 |
+
INFO:hf-to-gguf:blk.31.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 601 |
+
INFO:hf-to-gguf:blk.31.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
|
| 602 |
+
INFO:hf-to-gguf:blk.31.attn_k.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 603 |
+
INFO:hf-to-gguf:blk.31.attn_output.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 604 |
+
INFO:hf-to-gguf:blk.31.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
|
| 605 |
+
INFO:hf-to-gguf:blk.31.attn_q.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 606 |
+
INFO:hf-to-gguf:blk.31.attn_v.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 607 |
+
INFO:hf-to-gguf:blk.32.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 608 |
+
INFO:hf-to-gguf:blk.32.ssm_a, torch.bfloat16 --> F32, shape = {32}
|
| 609 |
+
INFO:hf-to-gguf:blk.32.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
|
| 610 |
+
INFO:hf-to-gguf:blk.32.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
|
| 611 |
+
INFO:hf-to-gguf:blk.32.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 612 |
+
INFO:hf-to-gguf:blk.32.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 613 |
+
INFO:hf-to-gguf:blk.32.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 614 |
+
INFO:hf-to-gguf:blk.32.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
|
| 615 |
+
INFO:hf-to-gguf:blk.32.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
|
| 616 |
+
INFO:hf-to-gguf:blk.32.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 617 |
+
INFO:hf-to-gguf:blk.32.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 618 |
+
INFO:hf-to-gguf:blk.32.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 619 |
+
INFO:hf-to-gguf:blk.32.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 620 |
+
INFO:hf-to-gguf:blk.32.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 621 |
+
INFO:hf-to-gguf:blk.32.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 622 |
+
INFO:hf-to-gguf:blk.32.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 623 |
+
INFO:hf-to-gguf:blk.32.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 624 |
+
INFO:hf-to-gguf:blk.32.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 625 |
+
INFO:hf-to-gguf:blk.32.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 626 |
+
INFO:hf-to-gguf:blk.33.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 627 |
+
INFO:hf-to-gguf:blk.33.ssm_a, torch.bfloat16 --> F32, shape = {32}
|
| 628 |
+
INFO:hf-to-gguf:blk.33.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
|
| 629 |
+
INFO:hf-to-gguf:blk.33.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
|
| 630 |
+
INFO:hf-to-gguf:blk.33.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 631 |
+
INFO:hf-to-gguf:blk.33.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 632 |
+
INFO:hf-to-gguf:blk.33.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 633 |
+
INFO:hf-to-gguf:blk.33.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
|
| 634 |
+
INFO:hf-to-gguf:blk.33.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
|
| 635 |
+
INFO:hf-to-gguf:blk.33.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 636 |
+
INFO:hf-to-gguf:blk.33.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 637 |
+
INFO:hf-to-gguf:blk.33.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 638 |
+
INFO:hf-to-gguf:blk.33.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 639 |
+
INFO:hf-to-gguf:blk.33.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 640 |
+
INFO:hf-to-gguf:blk.33.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 641 |
+
INFO:hf-to-gguf:blk.33.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 642 |
+
INFO:hf-to-gguf:blk.33.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 643 |
+
INFO:hf-to-gguf:blk.33.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 644 |
+
INFO:hf-to-gguf:blk.33.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 645 |
+
INFO:hf-to-gguf:blk.34.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 646 |
+
INFO:hf-to-gguf:blk.34.ssm_a, torch.bfloat16 --> F32, shape = {32}
|
| 647 |
+
INFO:hf-to-gguf:blk.34.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
|
| 648 |
+
INFO:hf-to-gguf:blk.34.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
|
| 649 |
+
INFO:hf-to-gguf:blk.34.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 650 |
+
INFO:hf-to-gguf:blk.34.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 651 |
+
INFO:hf-to-gguf:blk.34.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 652 |
+
INFO:hf-to-gguf:blk.34.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
|
| 653 |
+
INFO:hf-to-gguf:blk.34.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
|
| 654 |
+
INFO:hf-to-gguf:blk.34.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 655 |
+
INFO:hf-to-gguf:blk.34.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 656 |
+
INFO:hf-to-gguf:blk.34.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 657 |
+
INFO:hf-to-gguf:blk.34.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 658 |
+
INFO:hf-to-gguf:blk.34.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 659 |
+
INFO:hf-to-gguf:blk.34.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 660 |
+
INFO:hf-to-gguf:blk.34.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 661 |
+
INFO:hf-to-gguf:blk.34.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 662 |
+
INFO:hf-to-gguf:blk.34.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 663 |
+
INFO:hf-to-gguf:blk.34.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 664 |
+
INFO:hf-to-gguf:blk.35.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 665 |
+
INFO:hf-to-gguf:blk.35.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 666 |
+
INFO:hf-to-gguf:blk.35.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 667 |
+
INFO:hf-to-gguf:blk.35.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 668 |
+
INFO:hf-to-gguf:blk.35.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 669 |
+
INFO:hf-to-gguf:blk.35.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 670 |
+
INFO:hf-to-gguf:blk.35.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 671 |
+
INFO:hf-to-gguf:blk.35.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 672 |
+
INFO:hf-to-gguf:blk.35.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 673 |
+
INFO:hf-to-gguf:blk.35.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 674 |
+
INFO:hf-to-gguf:blk.35.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
|
| 675 |
+
INFO:hf-to-gguf:blk.35.attn_k.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 676 |
+
INFO:hf-to-gguf:blk.35.attn_output.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 677 |
+
INFO:hf-to-gguf:blk.35.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
|
| 678 |
+
INFO:hf-to-gguf:blk.35.attn_q.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 679 |
+
INFO:hf-to-gguf:blk.35.attn_v.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 680 |
+
INFO:hf-to-gguf:blk.36.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 681 |
+
INFO:hf-to-gguf:blk.36.ssm_a, torch.bfloat16 --> F32, shape = {32}
|
| 682 |
+
INFO:hf-to-gguf:blk.36.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
|
| 683 |
+
INFO:hf-to-gguf:blk.36.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
|
| 684 |
+
INFO:hf-to-gguf:blk.36.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 685 |
+
INFO:hf-to-gguf:blk.36.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 686 |
+
INFO:hf-to-gguf:blk.36.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 687 |
+
INFO:hf-to-gguf:blk.36.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
|
| 688 |
+
INFO:hf-to-gguf:blk.36.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
|
| 689 |
+
INFO:hf-to-gguf:blk.36.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 690 |
+
INFO:hf-to-gguf:blk.36.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 691 |
+
INFO:hf-to-gguf:blk.36.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 692 |
+
INFO:hf-to-gguf:blk.36.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 693 |
+
INFO:hf-to-gguf:blk.36.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 694 |
+
INFO:hf-to-gguf:blk.36.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 695 |
+
INFO:hf-to-gguf:blk.36.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 696 |
+
INFO:hf-to-gguf:blk.36.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 697 |
+
INFO:hf-to-gguf:blk.36.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 698 |
+
INFO:hf-to-gguf:blk.36.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 699 |
+
INFO:hf-to-gguf:blk.37.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 700 |
+
INFO:hf-to-gguf:blk.37.ssm_a, torch.bfloat16 --> F32, shape = {32}
|
| 701 |
+
INFO:hf-to-gguf:blk.37.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
|
| 702 |
+
INFO:hf-to-gguf:blk.37.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
|
| 703 |
+
INFO:hf-to-gguf:blk.37.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 704 |
+
INFO:hf-to-gguf:blk.37.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 705 |
+
INFO:hf-to-gguf:blk.37.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 706 |
+
INFO:hf-to-gguf:blk.37.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
|
| 707 |
+
INFO:hf-to-gguf:blk.37.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
|
| 708 |
+
INFO:hf-to-gguf:blk.37.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 709 |
+
INFO:hf-to-gguf:blk.37.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 710 |
+
INFO:hf-to-gguf:blk.37.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 711 |
+
INFO:hf-to-gguf:blk.37.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 712 |
+
INFO:hf-to-gguf:blk.37.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 713 |
+
INFO:hf-to-gguf:blk.37.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 714 |
+
INFO:hf-to-gguf:blk.37.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 715 |
+
INFO:hf-to-gguf:blk.37.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 716 |
+
INFO:hf-to-gguf:blk.37.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 717 |
+
INFO:hf-to-gguf:blk.37.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 718 |
+
INFO:hf-to-gguf:blk.38.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 719 |
+
INFO:hf-to-gguf:blk.38.ssm_a, torch.bfloat16 --> F32, shape = {32}
|
| 720 |
+
INFO:hf-to-gguf:blk.38.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
|
| 721 |
+
INFO:hf-to-gguf:blk.38.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
|
| 722 |
+
INFO:hf-to-gguf:blk.38.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 723 |
+
INFO:hf-to-gguf:blk.38.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
|
| 724 |
+
INFO:hf-to-gguf:blk.38.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 725 |
+
INFO:hf-to-gguf:blk.38.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
|
| 726 |
+
INFO:hf-to-gguf:blk.38.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
|
| 727 |
+
INFO:hf-to-gguf:blk.38.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 728 |
+
INFO:hf-to-gguf:blk.38.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 729 |
+
INFO:hf-to-gguf:blk.38.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 730 |
+
INFO:hf-to-gguf:blk.38.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 731 |
+
INFO:hf-to-gguf:blk.38.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 732 |
+
INFO:hf-to-gguf:blk.38.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 733 |
+
INFO:hf-to-gguf:blk.38.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 734 |
+
INFO:hf-to-gguf:blk.38.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 735 |
+
INFO:hf-to-gguf:blk.38.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 736 |
+
INFO:hf-to-gguf:blk.38.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 737 |
+
INFO:hf-to-gguf:output.weight, torch.bfloat16 --> BF16, shape = {2048, 248320}
|
| 738 |
+
INFO:hf-to-gguf:blk.39.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 739 |
+
INFO:hf-to-gguf:blk.39.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
|
| 740 |
+
INFO:hf-to-gguf:blk.39.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 741 |
+
INFO:hf-to-gguf:blk.39.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
|
| 742 |
+
INFO:hf-to-gguf:blk.39.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
|
| 743 |
+
INFO:hf-to-gguf:blk.39.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
|
| 744 |
+
INFO:hf-to-gguf:blk.39.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 745 |
+
INFO:hf-to-gguf:blk.39.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 746 |
+
INFO:hf-to-gguf:blk.39.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
|
| 747 |
+
INFO:hf-to-gguf:blk.39.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 748 |
+
INFO:hf-to-gguf:blk.39.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
|
| 749 |
+
INFO:hf-to-gguf:blk.39.attn_k.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 750 |
+
INFO:hf-to-gguf:blk.39.attn_output.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
|
| 751 |
+
INFO:hf-to-gguf:blk.39.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
|
| 752 |
+
INFO:hf-to-gguf:blk.39.attn_q.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
|
| 753 |
+
INFO:hf-to-gguf:blk.39.attn_v.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
|
| 754 |
+
INFO:hf-to-gguf:output_norm.weight, torch.bfloat16 --> F32, shape = {2048}
|
| 755 |
+
INFO:hf-to-gguf:Set meta model
|
| 756 |
+
INFO:hf-to-gguf:Set model parameters
|
| 757 |
+
INFO:hf-to-gguf:gguf: context length = 262144
|
| 758 |
+
INFO:hf-to-gguf:gguf: embedding length = 2048
|
| 759 |
+
INFO:hf-to-gguf:gguf: head count = 16
|
| 760 |
+
INFO:hf-to-gguf:gguf: key-value head count = 2
|
| 761 |
+
WARNING:hf-to-gguf:Unknown RoPE type: default
|
| 762 |
+
INFO:hf-to-gguf:gguf: rope scaling type = NONE
|
| 763 |
+
INFO:hf-to-gguf:gguf: mrope sections: [11, 11, 10, 0]
|
| 764 |
+
INFO:hf-to-gguf:gguf: rope theta = 10000000
|
| 765 |
+
INFO:hf-to-gguf:gguf: rms norm epsilon = 1e-06
|
| 766 |
+
INFO:hf-to-gguf:gguf: expert count = 256
|
| 767 |
+
INFO:hf-to-gguf:gguf: experts used count = 8
|
| 768 |
+
INFO:hf-to-gguf:gguf: file type = 32
|
| 769 |
+
INFO:hf-to-gguf:gguf: expert feed forward length = 512
|
| 770 |
+
INFO:hf-to-gguf:gguf: expert shared feed forward length = 512
|
| 771 |
+
INFO:hf-to-gguf:Set model quantization version
|
| 772 |
+
INFO:hf-to-gguf:Set model tokenizer
|
| 773 |
+
INFO:gguf.vocab:Adding 247587 merge(s).
|
| 774 |
+
INFO:gguf.vocab:Setting special token type eos to 248046
|
| 775 |
+
INFO:gguf.vocab:Setting special token type pad to 248044
|
| 776 |
+
INFO:gguf.vocab:Setting chat_template to {%- set image_count = namespace(value=0) %}
|
| 777 |
+
{%- set video_count = namespace(value=0) %}
|
| 778 |
+
{%- macro render_content(content, do_vision_count, is_system_content=false) %}
|
| 779 |
+
{%- if content is string %}
|
| 780 |
+
{{- content }}
|
| 781 |
+
{%- elif content is iterable and content is not mapping %}
|
| 782 |
+
{%- for item in content %}
|
| 783 |
+
{%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
|
| 784 |
+
{%- if is_system_content %}
|
| 785 |
+
{{- raise_exception('System message cannot contain images.') }}
|
| 786 |
+
{%- endif %}
|
| 787 |
+
{%- if do_vision_count %}
|
| 788 |
+
{%- set image_count.value = image_count.value + 1 %}
|
| 789 |
+
{%- endif %}
|
| 790 |
+
{%- if add_vision_id %}
|
| 791 |
+
{{- 'Picture ' ~ image_count.value ~ ': ' }}
|
| 792 |
+
{%- endif %}
|
| 793 |
+
{{- '<|vision_start|><|image_pad|><|vision_end|>' }}
|
| 794 |
+
{%- elif 'video' in item or item.type == 'video' %}
|
| 795 |
+
{%- if is_system_content %}
|
| 796 |
+
{{- raise_exception('System message cannot contain videos.') }}
|
| 797 |
+
{%- endif %}
|
| 798 |
+
{%- if do_vision_count %}
|
| 799 |
+
{%- set video_count.value = video_count.value + 1 %}
|
| 800 |
+
{%- endif %}
|
| 801 |
+
{%- if add_vision_id %}
|
| 802 |
+
{{- 'Video ' ~ video_count.value ~ ': ' }}
|
| 803 |
+
{%- endif %}
|
| 804 |
+
{{- '<|vision_start|><|video_pad|><|vision_end|>' }}
|
| 805 |
+
{%- elif 'text' in item %}
|
| 806 |
+
{{- item.text }}
|
| 807 |
+
{%- else %}
|
| 808 |
+
{{- raise_exception('Unexpected item type in content.') }}
|
| 809 |
+
{%- endif %}
|
| 810 |
+
{%- endfor %}
|
| 811 |
+
{%- elif content is none or content is undefined %}
|
| 812 |
+
{{- '' }}
|
| 813 |
+
{%- else %}
|
| 814 |
+
{{- raise_exception('Unexpected content type.') }}
|
| 815 |
+
{%- endif %}
|
| 816 |
+
{%- endmacro %}
|
| 817 |
+
{%- if not messages %}
|
| 818 |
+
{{- raise_exception('No messages provided.') }}
|
| 819 |
+
{%- endif %}
|
| 820 |
+
{%- if tools and tools is iterable and tools is not mapping %}
|
| 821 |
+
{{- '<|im_start|>system\n' }}
|
| 822 |
+
{{- "# Tools\n\nYou have access to the following functions:\n\n<tools>" }}
|
| 823 |
+
{%- for tool in tools %}
|
| 824 |
+
{{- "\n" }}
|
| 825 |
+
{{- tool | tojson }}
|
| 826 |
+
{%- endfor %}
|
| 827 |
+
{{- "\n</tools>" }}
|
| 828 |
+
{{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n<tool_call>\n<function=example_function_name>\n<parameter=example_parameter_1>\nvalue_1\n</parameter>\n<parameter=example_parameter_2>\nThis is the value for the second parameter\nthat can span\nmultiple lines\n</parameter>\n</function>\n</tool_call>\n\n<IMPORTANT>\nReminder:\n- Function calls MUST follow the specified format: an inner <function=...></function> block must be nested within <tool_call></tool_call> XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n</IMPORTANT>' }}
|
| 829 |
+
{%- if messages[0].role == 'system' %}
|
| 830 |
+
{%- set content = render_content(messages[0].content, false, true)|trim %}
|
| 831 |
+
{%- if content %}
|
| 832 |
+
{{- '\n\n' + content }}
|
| 833 |
+
{%- endif %}
|
| 834 |
+
{%- endif %}
|
| 835 |
+
{{- '<|im_end|>\n' }}
|
| 836 |
+
{%- else %}
|
| 837 |
+
{%- if messages[0].role == 'system' %}
|
| 838 |
+
{%- set content = render_content(messages[0].content, false, true)|trim %}
|
| 839 |
+
{{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
|
| 840 |
+
{%- endif %}
|
| 841 |
+
{%- endif %}
|
| 842 |
+
{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
|
| 843 |
+
{%- for message in messages[::-1] %}
|
| 844 |
+
{%- set index = (messages|length - 1) - loop.index0 %}
|
| 845 |
+
{%- if ns.multi_step_tool and message.role == "user" %}
|
| 846 |
+
{%- set content = render_content(message.content, false)|trim %}
|
| 847 |
+
{%- if not(content.startswith('<tool_response>') and content.endswith('</tool_response>')) %}
|
| 848 |
+
{%- set ns.multi_step_tool = false %}
|
| 849 |
+
{%- set ns.last_query_index = index %}
|
| 850 |
+
{%- endif %}
|
| 851 |
+
{%- endif %}
|
| 852 |
+
{%- endfor %}
|
| 853 |
+
{%- if ns.multi_step_tool %}
|
| 854 |
+
{{- raise_exception('No user query found in messages.') }}
|
| 855 |
+
{%- endif %}
|
| 856 |
+
{%- for message in messages %}
|
| 857 |
+
{%- set content = render_content(message.content, true)|trim %}
|
| 858 |
+
{%- if message.role == "system" %}
|
| 859 |
+
{%- if not loop.first %}
|
| 860 |
+
{{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
|
| 861 |
+
{%- endif %}
|
| 862 |
+
{%- elif message.role == "user" %}
|
| 863 |
+
{{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
|
| 864 |
+
{%- elif message.role == "assistant" %}
|
| 865 |
+
{%- set reasoning_content = '' %}
|
| 866 |
+
{%- if message.reasoning_content is string %}
|
| 867 |
+
{%- set reasoning_content = message.reasoning_content %}
|
| 868 |
+
{%- else %}
|
| 869 |
+
{%- if '</think>' in content %}
|
| 870 |
+
{%- set reasoning_content = content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
|
| 871 |
+
{%- set content = content.split('</think>')[-1].lstrip('\n') %}
|
| 872 |
+
{%- endif %}
|
| 873 |
+
{%- endif %}
|
| 874 |
+
{%- set reasoning_content = reasoning_content|trim %}
|
| 875 |
+
{{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content + '\n</think>\n\n' + content }}
|
| 876 |
+
{%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
|
| 877 |
+
{%- for tool_call in message.tool_calls %}
|
| 878 |
+
{%- if tool_call.function is defined %}
|
| 879 |
+
{%- set tool_call = tool_call.function %}
|
| 880 |
+
{%- endif %}
|
| 881 |
+
{%- if loop.first %}
|
| 882 |
+
{%- if content|trim %}
|
| 883 |
+
{{- '\n\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
|
| 884 |
+
{%- else %}
|
| 885 |
+
{{- '<tool_call>\n<function=' + tool_call.name + '>\n' }}
|
| 886 |
+
{%- endif %}
|
| 887 |
+
{%- else %}
|
| 888 |
+
{{- '\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
|
| 889 |
+
{%- endif %}
|
| 890 |
+
{%- if tool_call.arguments is defined %}
|
| 891 |
+
{%- for args_name, args_value in tool_call.arguments|items %}
|
| 892 |
+
{{- '<parameter=' + args_name + '>\n' }}
|
| 893 |
+
{%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
|
| 894 |
+
{{- args_value }}
|
| 895 |
+
{{- '\n</parameter>\n' }}
|
| 896 |
+
{%- endfor %}
|
| 897 |
+
{%- endif %}
|
| 898 |
+
{{- '</function>\n</tool_call>' }}
|
| 899 |
+
{%- endfor %}
|
| 900 |
+
{%- endif %}
|
| 901 |
+
{{- '<|im_end|>\n' }}
|
| 902 |
+
{%- elif message.role == "tool" %}
|
| 903 |
+
{%- if loop.previtem and loop.previtem.role != "tool" %}
|
| 904 |
+
{{- '<|im_start|>user' }}
|
| 905 |
+
{%- endif %}
|
| 906 |
+
{{- '\n<tool_response>\n' }}
|
| 907 |
+
{{- content }}
|
| 908 |
+
{{- '\n</tool_response>' }}
|
| 909 |
+
{%- if not loop.last and loop.nextitem.role != "tool" %}
|
| 910 |
+
{{- '<|im_end|>\n' }}
|
| 911 |
+
{%- elif loop.last %}
|
| 912 |
+
{{- '<|im_end|>\n' }}
|
| 913 |
+
{%- endif %}
|
| 914 |
+
{%- else %}
|
| 915 |
+
{{- raise_exception('Unexpected message role.') }}
|
| 916 |
+
{%- endif %}
|
| 917 |
+
{%- endfor %}
|
| 918 |
+
{%- if add_generation_prompt %}
|
| 919 |
+
{{- '<|im_start|>assistant\n' }}
|
| 920 |
+
{%- if reasoning_effort is not defined or reasoning_effort is none %}
|
| 921 |
+
{{- '<think>' }}
|
| 922 |
+
{%- elif reasoning_effort == 'none' %}
|
| 923 |
+
{{- '<think>\n\n</think>\n\n' }}
|
| 924 |
+
{%- elif reasoning_effort == 'high' %}
|
| 925 |
+
{{- '<think>\n' }}
|
| 926 |
+
{%- else %}
|
| 927 |
+
{{- '<think>' }}
|
| 928 |
+
{%- endif %}
|
| 929 |
+
{%- endif %}
|
| 930 |
+
|
| 931 |
+
INFO:gguf.gguf_writer:Writing the following files:
|
| 932 |
+
INFO:gguf.gguf_writer:gguf/Nex-N2.5-mini-BF16.gguf: n_tensors = 733, total_size = 69.4G
|
| 933 |
+
|
| 934 |
+
INFO:hf-to-gguf:Model successfully exported to gguf/Nex-N2.5-mini-BF16.gguf
|
recipe/logs/C2_mmproj.log
ADDED
|
@@ -0,0 +1,362 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
INFO:hf-to-gguf:Loading model: hf
|
| 2 |
+
INFO:hf-to-gguf:Model architecture: Qwen3_5MoeForConditionalGeneration
|
| 3 |
+
INFO:hf-to-gguf:gguf: loading model weight map from 'model.safetensors.index.json'
|
| 4 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00001-of-00016.safetensors'
|
| 5 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00002-of-00016.safetensors'
|
| 6 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00003-of-00016.safetensors'
|
| 7 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00004-of-00016.safetensors'
|
| 8 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00005-of-00016.safetensors'
|
| 9 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00006-of-00016.safetensors'
|
| 10 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00007-of-00016.safetensors'
|
| 11 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00008-of-00016.safetensors'
|
| 12 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00009-of-00016.safetensors'
|
| 13 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00010-of-00016.safetensors'
|
| 14 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00011-of-00016.safetensors'
|
| 15 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00012-of-00016.safetensors'
|
| 16 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00013-of-00016.safetensors'
|
| 17 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00014-of-00016.safetensors'
|
| 18 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00015-of-00016.safetensors'
|
| 19 |
+
INFO:hf-to-gguf:gguf: indexing model part 'model-00016-of-00016.safetensors'
|
| 20 |
+
INFO:gguf.gguf_writer:gguf: This GGUF file is for Little Endian only
|
| 21 |
+
INFO:hf-to-gguf:Exporting model...
|
| 22 |
+
INFO:hf-to-gguf:v.blk.0.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 23 |
+
INFO:hf-to-gguf:v.blk.0.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
|
| 24 |
+
INFO:hf-to-gguf:v.blk.0.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
|
| 25 |
+
INFO:hf-to-gguf:v.blk.0.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
|
| 26 |
+
INFO:hf-to-gguf:v.blk.0.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
|
| 27 |
+
INFO:hf-to-gguf:v.blk.0.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
|
| 28 |
+
INFO:hf-to-gguf:v.blk.0.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 29 |
+
INFO:hf-to-gguf:v.blk.0.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
|
| 30 |
+
INFO:hf-to-gguf:v.blk.0.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 31 |
+
INFO:hf-to-gguf:v.blk.0.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 32 |
+
INFO:hf-to-gguf:v.blk.0.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 33 |
+
INFO:hf-to-gguf:v.blk.0.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 34 |
+
INFO:hf-to-gguf:v.blk.1.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 35 |
+
INFO:hf-to-gguf:v.blk.1.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
|
| 36 |
+
INFO:hf-to-gguf:v.blk.1.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
|
| 37 |
+
INFO:hf-to-gguf:v.blk.1.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
|
| 38 |
+
INFO:hf-to-gguf:v.blk.1.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
|
| 39 |
+
INFO:hf-to-gguf:v.blk.1.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
|
| 40 |
+
INFO:hf-to-gguf:v.blk.1.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 41 |
+
INFO:hf-to-gguf:v.blk.1.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
|
| 42 |
+
INFO:hf-to-gguf:v.blk.1.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 43 |
+
INFO:hf-to-gguf:v.blk.1.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 44 |
+
INFO:hf-to-gguf:v.blk.1.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 45 |
+
INFO:hf-to-gguf:v.blk.1.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 46 |
+
INFO:hf-to-gguf:v.blk.10.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 47 |
+
INFO:hf-to-gguf:v.blk.10.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
|
| 48 |
+
INFO:hf-to-gguf:v.blk.10.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
|
| 49 |
+
INFO:hf-to-gguf:v.blk.10.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
|
| 50 |
+
INFO:hf-to-gguf:v.blk.10.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
|
| 51 |
+
INFO:hf-to-gguf:v.blk.10.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
|
| 52 |
+
INFO:hf-to-gguf:v.blk.10.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 53 |
+
INFO:hf-to-gguf:v.blk.10.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
|
| 54 |
+
INFO:hf-to-gguf:v.blk.10.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 55 |
+
INFO:hf-to-gguf:v.blk.10.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 56 |
+
INFO:hf-to-gguf:v.blk.10.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 57 |
+
INFO:hf-to-gguf:v.blk.10.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 58 |
+
INFO:hf-to-gguf:v.blk.11.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 59 |
+
INFO:hf-to-gguf:v.blk.11.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
|
| 60 |
+
INFO:hf-to-gguf:v.blk.11.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
|
| 61 |
+
INFO:hf-to-gguf:v.blk.11.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
|
| 62 |
+
INFO:hf-to-gguf:v.blk.11.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
|
| 63 |
+
INFO:hf-to-gguf:v.blk.11.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
|
| 64 |
+
INFO:hf-to-gguf:v.blk.11.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 65 |
+
INFO:hf-to-gguf:v.blk.11.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
|
| 66 |
+
INFO:hf-to-gguf:v.blk.11.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 67 |
+
INFO:hf-to-gguf:v.blk.11.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 68 |
+
INFO:hf-to-gguf:v.blk.11.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 69 |
+
INFO:hf-to-gguf:v.blk.11.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 70 |
+
INFO:hf-to-gguf:v.blk.12.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 71 |
+
INFO:hf-to-gguf:v.blk.12.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
|
| 72 |
+
INFO:hf-to-gguf:v.blk.12.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
|
| 73 |
+
INFO:hf-to-gguf:v.blk.12.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
|
| 74 |
+
INFO:hf-to-gguf:v.blk.12.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
|
| 75 |
+
INFO:hf-to-gguf:v.blk.12.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
|
| 76 |
+
INFO:hf-to-gguf:v.blk.12.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 77 |
+
INFO:hf-to-gguf:v.blk.12.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
|
| 78 |
+
INFO:hf-to-gguf:v.blk.12.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 79 |
+
INFO:hf-to-gguf:v.blk.12.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 80 |
+
INFO:hf-to-gguf:v.blk.12.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 81 |
+
INFO:hf-to-gguf:v.blk.12.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 82 |
+
INFO:hf-to-gguf:v.blk.13.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 83 |
+
INFO:hf-to-gguf:v.blk.13.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
|
| 84 |
+
INFO:hf-to-gguf:v.blk.13.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
|
| 85 |
+
INFO:hf-to-gguf:v.blk.13.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
|
| 86 |
+
INFO:hf-to-gguf:v.blk.13.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
|
| 87 |
+
INFO:hf-to-gguf:v.blk.13.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
|
| 88 |
+
INFO:hf-to-gguf:v.blk.13.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 89 |
+
INFO:hf-to-gguf:v.blk.13.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
|
| 90 |
+
INFO:hf-to-gguf:v.blk.13.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 91 |
+
INFO:hf-to-gguf:v.blk.13.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 92 |
+
INFO:hf-to-gguf:v.blk.13.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 93 |
+
INFO:hf-to-gguf:v.blk.13.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 94 |
+
INFO:hf-to-gguf:v.blk.14.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 95 |
+
INFO:hf-to-gguf:v.blk.14.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
|
| 96 |
+
INFO:hf-to-gguf:v.blk.14.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
|
| 97 |
+
INFO:hf-to-gguf:v.blk.14.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
|
| 98 |
+
INFO:hf-to-gguf:v.blk.14.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
|
| 99 |
+
INFO:hf-to-gguf:v.blk.14.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
|
| 100 |
+
INFO:hf-to-gguf:v.blk.14.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 101 |
+
INFO:hf-to-gguf:v.blk.14.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
|
| 102 |
+
INFO:hf-to-gguf:v.blk.14.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 103 |
+
INFO:hf-to-gguf:v.blk.14.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 104 |
+
INFO:hf-to-gguf:v.blk.14.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 105 |
+
INFO:hf-to-gguf:v.blk.14.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 106 |
+
INFO:hf-to-gguf:v.blk.15.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 107 |
+
INFO:hf-to-gguf:v.blk.15.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
|
| 108 |
+
INFO:hf-to-gguf:v.blk.15.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
|
| 109 |
+
INFO:hf-to-gguf:v.blk.15.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
|
| 110 |
+
INFO:hf-to-gguf:v.blk.15.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
|
| 111 |
+
INFO:hf-to-gguf:v.blk.15.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
|
| 112 |
+
INFO:hf-to-gguf:v.blk.15.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 113 |
+
INFO:hf-to-gguf:v.blk.15.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
|
| 114 |
+
INFO:hf-to-gguf:v.blk.15.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 115 |
+
INFO:hf-to-gguf:v.blk.15.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 116 |
+
INFO:hf-to-gguf:v.blk.15.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 117 |
+
INFO:hf-to-gguf:v.blk.15.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 118 |
+
INFO:hf-to-gguf:v.blk.16.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 119 |
+
INFO:hf-to-gguf:v.blk.16.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
|
| 120 |
+
INFO:hf-to-gguf:v.blk.16.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
|
| 121 |
+
INFO:hf-to-gguf:v.blk.16.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
|
| 122 |
+
INFO:hf-to-gguf:v.blk.16.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
|
| 123 |
+
INFO:hf-to-gguf:v.blk.16.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
|
| 124 |
+
INFO:hf-to-gguf:v.blk.16.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 125 |
+
INFO:hf-to-gguf:v.blk.16.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
|
| 126 |
+
INFO:hf-to-gguf:v.blk.16.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 127 |
+
INFO:hf-to-gguf:v.blk.16.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 128 |
+
INFO:hf-to-gguf:v.blk.16.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 129 |
+
INFO:hf-to-gguf:v.blk.16.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 130 |
+
INFO:hf-to-gguf:v.blk.17.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 131 |
+
INFO:hf-to-gguf:v.blk.17.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
|
| 132 |
+
INFO:hf-to-gguf:v.blk.17.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
|
| 133 |
+
INFO:hf-to-gguf:v.blk.17.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
|
| 134 |
+
INFO:hf-to-gguf:v.blk.17.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
|
| 135 |
+
INFO:hf-to-gguf:v.blk.17.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
|
| 136 |
+
INFO:hf-to-gguf:v.blk.17.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 137 |
+
INFO:hf-to-gguf:v.blk.17.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
|
| 138 |
+
INFO:hf-to-gguf:v.blk.17.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 139 |
+
INFO:hf-to-gguf:v.blk.17.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 140 |
+
INFO:hf-to-gguf:v.blk.17.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 141 |
+
INFO:hf-to-gguf:v.blk.17.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 142 |
+
INFO:hf-to-gguf:v.blk.18.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 143 |
+
INFO:hf-to-gguf:v.blk.18.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
|
| 144 |
+
INFO:hf-to-gguf:v.blk.18.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
|
| 145 |
+
INFO:hf-to-gguf:v.blk.18.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
|
| 146 |
+
INFO:hf-to-gguf:v.blk.18.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
|
| 147 |
+
INFO:hf-to-gguf:v.blk.18.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
|
| 148 |
+
INFO:hf-to-gguf:v.blk.18.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 149 |
+
INFO:hf-to-gguf:v.blk.18.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
|
| 150 |
+
INFO:hf-to-gguf:v.blk.18.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 151 |
+
INFO:hf-to-gguf:v.blk.18.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 152 |
+
INFO:hf-to-gguf:v.blk.18.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 153 |
+
INFO:hf-to-gguf:v.blk.18.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 154 |
+
INFO:hf-to-gguf:v.blk.19.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 155 |
+
INFO:hf-to-gguf:v.blk.19.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
|
| 156 |
+
INFO:hf-to-gguf:v.blk.19.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
|
| 157 |
+
INFO:hf-to-gguf:v.blk.19.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
|
| 158 |
+
INFO:hf-to-gguf:v.blk.19.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
|
| 159 |
+
INFO:hf-to-gguf:v.blk.19.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
|
| 160 |
+
INFO:hf-to-gguf:v.blk.19.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 161 |
+
INFO:hf-to-gguf:v.blk.19.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
|
| 162 |
+
INFO:hf-to-gguf:v.blk.19.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 163 |
+
INFO:hf-to-gguf:v.blk.19.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 164 |
+
INFO:hf-to-gguf:v.blk.19.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 165 |
+
INFO:hf-to-gguf:v.blk.19.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 166 |
+
INFO:hf-to-gguf:v.blk.2.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 167 |
+
INFO:hf-to-gguf:v.blk.2.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
|
| 168 |
+
INFO:hf-to-gguf:v.blk.2.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
|
| 169 |
+
INFO:hf-to-gguf:v.blk.2.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
|
| 170 |
+
INFO:hf-to-gguf:v.blk.2.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
|
| 171 |
+
INFO:hf-to-gguf:v.blk.2.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
|
| 172 |
+
INFO:hf-to-gguf:v.blk.2.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 173 |
+
INFO:hf-to-gguf:v.blk.2.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
|
| 174 |
+
INFO:hf-to-gguf:v.blk.2.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 175 |
+
INFO:hf-to-gguf:v.blk.2.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 176 |
+
INFO:hf-to-gguf:v.blk.2.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 177 |
+
INFO:hf-to-gguf:v.blk.2.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 178 |
+
INFO:hf-to-gguf:v.blk.20.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 179 |
+
INFO:hf-to-gguf:v.blk.20.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
|
| 180 |
+
INFO:hf-to-gguf:v.blk.20.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
|
| 181 |
+
INFO:hf-to-gguf:v.blk.20.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
|
| 182 |
+
INFO:hf-to-gguf:v.blk.20.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
|
| 183 |
+
INFO:hf-to-gguf:v.blk.20.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
|
| 184 |
+
INFO:hf-to-gguf:v.blk.20.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 185 |
+
INFO:hf-to-gguf:v.blk.20.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
|
| 186 |
+
INFO:hf-to-gguf:v.blk.20.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 187 |
+
INFO:hf-to-gguf:v.blk.20.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 188 |
+
INFO:hf-to-gguf:v.blk.20.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 189 |
+
INFO:hf-to-gguf:v.blk.20.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 190 |
+
INFO:hf-to-gguf:v.blk.21.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 191 |
+
INFO:hf-to-gguf:v.blk.21.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
|
| 192 |
+
INFO:hf-to-gguf:v.blk.21.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
|
| 193 |
+
INFO:hf-to-gguf:v.blk.21.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
|
| 194 |
+
INFO:hf-to-gguf:v.blk.21.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
|
| 195 |
+
INFO:hf-to-gguf:v.blk.21.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
|
| 196 |
+
INFO:hf-to-gguf:v.blk.21.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 197 |
+
INFO:hf-to-gguf:v.blk.21.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
|
| 198 |
+
INFO:hf-to-gguf:v.blk.21.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 199 |
+
INFO:hf-to-gguf:v.blk.21.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 200 |
+
INFO:hf-to-gguf:v.blk.21.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 201 |
+
INFO:hf-to-gguf:v.blk.21.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 202 |
+
INFO:hf-to-gguf:v.blk.22.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 203 |
+
INFO:hf-to-gguf:v.blk.22.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
|
| 204 |
+
INFO:hf-to-gguf:v.blk.22.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
|
| 205 |
+
INFO:hf-to-gguf:v.blk.22.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
|
| 206 |
+
INFO:hf-to-gguf:v.blk.22.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
|
| 207 |
+
INFO:hf-to-gguf:v.blk.22.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
|
| 208 |
+
INFO:hf-to-gguf:v.blk.22.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 209 |
+
INFO:hf-to-gguf:v.blk.22.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
|
| 210 |
+
INFO:hf-to-gguf:v.blk.22.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 211 |
+
INFO:hf-to-gguf:v.blk.22.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 212 |
+
INFO:hf-to-gguf:v.blk.22.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 213 |
+
INFO:hf-to-gguf:v.blk.22.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 214 |
+
INFO:hf-to-gguf:v.blk.23.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 215 |
+
INFO:hf-to-gguf:v.blk.23.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
|
| 216 |
+
INFO:hf-to-gguf:v.blk.23.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
|
| 217 |
+
INFO:hf-to-gguf:v.blk.23.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
|
| 218 |
+
INFO:hf-to-gguf:v.blk.23.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
|
| 219 |
+
INFO:hf-to-gguf:v.blk.23.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
|
| 220 |
+
INFO:hf-to-gguf:v.blk.23.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 221 |
+
INFO:hf-to-gguf:v.blk.23.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
|
| 222 |
+
INFO:hf-to-gguf:v.blk.23.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 223 |
+
INFO:hf-to-gguf:v.blk.23.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 224 |
+
INFO:hf-to-gguf:v.blk.23.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 225 |
+
INFO:hf-to-gguf:v.blk.23.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 226 |
+
INFO:hf-to-gguf:v.blk.24.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 227 |
+
INFO:hf-to-gguf:v.blk.24.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
|
| 228 |
+
INFO:hf-to-gguf:v.blk.24.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
|
| 229 |
+
INFO:hf-to-gguf:v.blk.24.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
|
| 230 |
+
INFO:hf-to-gguf:v.blk.24.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
|
| 231 |
+
INFO:hf-to-gguf:v.blk.24.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
|
| 232 |
+
INFO:hf-to-gguf:v.blk.24.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 233 |
+
INFO:hf-to-gguf:v.blk.24.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
|
| 234 |
+
INFO:hf-to-gguf:v.blk.24.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 235 |
+
INFO:hf-to-gguf:v.blk.24.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 236 |
+
INFO:hf-to-gguf:v.blk.24.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 237 |
+
INFO:hf-to-gguf:v.blk.24.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 238 |
+
INFO:hf-to-gguf:v.blk.25.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 239 |
+
INFO:hf-to-gguf:v.blk.25.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
|
| 240 |
+
INFO:hf-to-gguf:v.blk.25.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
|
| 241 |
+
INFO:hf-to-gguf:v.blk.25.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
|
| 242 |
+
INFO:hf-to-gguf:v.blk.25.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
|
| 243 |
+
INFO:hf-to-gguf:v.blk.25.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
|
| 244 |
+
INFO:hf-to-gguf:v.blk.25.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 245 |
+
INFO:hf-to-gguf:v.blk.25.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
|
| 246 |
+
INFO:hf-to-gguf:v.blk.25.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 247 |
+
INFO:hf-to-gguf:v.blk.25.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 248 |
+
INFO:hf-to-gguf:v.blk.25.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 249 |
+
INFO:hf-to-gguf:v.blk.25.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 250 |
+
INFO:hf-to-gguf:v.blk.26.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 251 |
+
INFO:hf-to-gguf:v.blk.26.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
|
| 252 |
+
INFO:hf-to-gguf:v.blk.26.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
|
| 253 |
+
INFO:hf-to-gguf:v.blk.26.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
|
| 254 |
+
INFO:hf-to-gguf:v.blk.26.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
|
| 255 |
+
INFO:hf-to-gguf:v.blk.26.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
|
| 256 |
+
INFO:hf-to-gguf:v.blk.26.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 257 |
+
INFO:hf-to-gguf:v.blk.26.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
|
| 258 |
+
INFO:hf-to-gguf:v.blk.26.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 259 |
+
INFO:hf-to-gguf:v.blk.26.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 260 |
+
INFO:hf-to-gguf:v.blk.26.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 261 |
+
INFO:hf-to-gguf:v.blk.26.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 262 |
+
INFO:hf-to-gguf:v.blk.3.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 263 |
+
INFO:hf-to-gguf:v.blk.3.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
|
| 264 |
+
INFO:hf-to-gguf:v.blk.3.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
|
| 265 |
+
INFO:hf-to-gguf:v.blk.3.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
|
| 266 |
+
INFO:hf-to-gguf:v.blk.3.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
|
| 267 |
+
INFO:hf-to-gguf:v.blk.3.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
|
| 268 |
+
INFO:hf-to-gguf:v.blk.3.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 269 |
+
INFO:hf-to-gguf:v.blk.3.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
|
| 270 |
+
INFO:hf-to-gguf:v.blk.3.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 271 |
+
INFO:hf-to-gguf:v.blk.3.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 272 |
+
INFO:hf-to-gguf:v.blk.3.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 273 |
+
INFO:hf-to-gguf:v.blk.3.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 274 |
+
INFO:hf-to-gguf:v.blk.4.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 275 |
+
INFO:hf-to-gguf:v.blk.4.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
|
| 276 |
+
INFO:hf-to-gguf:v.blk.4.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
|
| 277 |
+
INFO:hf-to-gguf:v.blk.4.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
|
| 278 |
+
INFO:hf-to-gguf:v.blk.4.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
|
| 279 |
+
INFO:hf-to-gguf:v.blk.4.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
|
| 280 |
+
INFO:hf-to-gguf:v.blk.4.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 281 |
+
INFO:hf-to-gguf:v.blk.4.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
|
| 282 |
+
INFO:hf-to-gguf:v.blk.4.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 283 |
+
INFO:hf-to-gguf:v.blk.4.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 284 |
+
INFO:hf-to-gguf:v.blk.4.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 285 |
+
INFO:hf-to-gguf:v.blk.4.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 286 |
+
INFO:hf-to-gguf:v.blk.5.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 287 |
+
INFO:hf-to-gguf:v.blk.5.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
|
| 288 |
+
INFO:hf-to-gguf:v.blk.5.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
|
| 289 |
+
INFO:hf-to-gguf:v.blk.5.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
|
| 290 |
+
INFO:hf-to-gguf:v.blk.5.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
|
| 291 |
+
INFO:hf-to-gguf:v.blk.5.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
|
| 292 |
+
INFO:hf-to-gguf:v.blk.5.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 293 |
+
INFO:hf-to-gguf:v.blk.5.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
|
| 294 |
+
INFO:hf-to-gguf:v.blk.5.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 295 |
+
INFO:hf-to-gguf:v.blk.5.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 296 |
+
INFO:hf-to-gguf:v.blk.5.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 297 |
+
INFO:hf-to-gguf:v.blk.5.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 298 |
+
INFO:hf-to-gguf:v.blk.6.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 299 |
+
INFO:hf-to-gguf:v.blk.6.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
|
| 300 |
+
INFO:hf-to-gguf:v.blk.6.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
|
| 301 |
+
INFO:hf-to-gguf:v.blk.6.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
|
| 302 |
+
INFO:hf-to-gguf:v.blk.6.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
|
| 303 |
+
INFO:hf-to-gguf:v.blk.6.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
|
| 304 |
+
INFO:hf-to-gguf:v.blk.6.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 305 |
+
INFO:hf-to-gguf:v.blk.6.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
|
| 306 |
+
INFO:hf-to-gguf:v.blk.6.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 307 |
+
INFO:hf-to-gguf:v.blk.6.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 308 |
+
INFO:hf-to-gguf:v.blk.6.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 309 |
+
INFO:hf-to-gguf:v.blk.6.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 310 |
+
INFO:hf-to-gguf:v.blk.7.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 311 |
+
INFO:hf-to-gguf:v.blk.7.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
|
| 312 |
+
INFO:hf-to-gguf:v.blk.7.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
|
| 313 |
+
INFO:hf-to-gguf:v.blk.7.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
|
| 314 |
+
INFO:hf-to-gguf:v.blk.7.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
|
| 315 |
+
INFO:hf-to-gguf:v.blk.7.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
|
| 316 |
+
INFO:hf-to-gguf:v.blk.7.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 317 |
+
INFO:hf-to-gguf:v.blk.7.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
|
| 318 |
+
INFO:hf-to-gguf:v.blk.7.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 319 |
+
INFO:hf-to-gguf:v.blk.7.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 320 |
+
INFO:hf-to-gguf:v.blk.7.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 321 |
+
INFO:hf-to-gguf:v.blk.7.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 322 |
+
INFO:hf-to-gguf:v.blk.8.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 323 |
+
INFO:hf-to-gguf:v.blk.8.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
|
| 324 |
+
INFO:hf-to-gguf:v.blk.8.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
|
| 325 |
+
INFO:hf-to-gguf:v.blk.8.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
|
| 326 |
+
INFO:hf-to-gguf:v.blk.8.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
|
| 327 |
+
INFO:hf-to-gguf:v.blk.8.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
|
| 328 |
+
INFO:hf-to-gguf:v.blk.8.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 329 |
+
INFO:hf-to-gguf:v.blk.8.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
|
| 330 |
+
INFO:hf-to-gguf:v.blk.8.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 331 |
+
INFO:hf-to-gguf:v.blk.8.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 332 |
+
INFO:hf-to-gguf:v.blk.8.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 333 |
+
INFO:hf-to-gguf:v.blk.8.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 334 |
+
INFO:hf-to-gguf:v.blk.9.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 335 |
+
INFO:hf-to-gguf:v.blk.9.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
|
| 336 |
+
INFO:hf-to-gguf:v.blk.9.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
|
| 337 |
+
INFO:hf-to-gguf:v.blk.9.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
|
| 338 |
+
INFO:hf-to-gguf:v.blk.9.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
|
| 339 |
+
INFO:hf-to-gguf:v.blk.9.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
|
| 340 |
+
INFO:hf-to-gguf:v.blk.9.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 341 |
+
INFO:hf-to-gguf:v.blk.9.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
|
| 342 |
+
INFO:hf-to-gguf:v.blk.9.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 343 |
+
INFO:hf-to-gguf:v.blk.9.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 344 |
+
INFO:hf-to-gguf:v.blk.9.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 345 |
+
INFO:hf-to-gguf:v.blk.9.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 346 |
+
INFO:hf-to-gguf:mm.0.bias, torch.bfloat16 --> F32, shape = {4608}
|
| 347 |
+
INFO:hf-to-gguf:mm.0.weight, torch.bfloat16 --> BF16, shape = {4608, 4608}
|
| 348 |
+
INFO:hf-to-gguf:mm.2.bias, torch.bfloat16 --> F32, shape = {2048}
|
| 349 |
+
INFO:hf-to-gguf:mm.2.weight, torch.bfloat16 --> BF16, shape = {4608, 2048}
|
| 350 |
+
INFO:hf-to-gguf:v.post_ln.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 351 |
+
INFO:hf-to-gguf:v.post_ln.weight, torch.bfloat16 --> F32, shape = {1152}
|
| 352 |
+
INFO:hf-to-gguf:v.patch_embd.bias, torch.bfloat16 --> F32, shape = {1152}
|
| 353 |
+
INFO:hf-to-gguf:v.patch_embd.weight, torch.bfloat16 --> F32, shape = {16, 16, 3, 1152}
|
| 354 |
+
INFO:hf-to-gguf:v.patch_embd.weight.1, torch.bfloat16 --> F32, shape = {16, 16, 3, 1152}
|
| 355 |
+
INFO:hf-to-gguf:v.position_embd.weight, torch.bfloat16 --> F32, shape = {1152, 2304}
|
| 356 |
+
INFO:hf-to-gguf:Set meta model
|
| 357 |
+
INFO:hf-to-gguf:Set model parameters
|
| 358 |
+
INFO:hf-to-gguf:Set model quantization version
|
| 359 |
+
INFO:gguf.gguf_writer:Writing the following files:
|
| 360 |
+
INFO:gguf.gguf_writer:out/mmproj-Nex-N2.5-mini-BF16.gguf: n_tensors = 334, total_size = 902.8M
|
| 361 |
+
|
| 362 |
+
INFO:hf-to-gguf:Model successfully exported to out/mmproj-Nex-N2.5-mini-BF16.gguf
|
recipe/logs/C_readback.log
ADDED
|
@@ -0,0 +1,2 @@
|
|
|
|
|
|
|
|
|
|
| 1 |
+
PASS Nex-N2.5-mini-BF16.gguf arch=qwen35moe ftype=32 tensors=733 nextn=0 output.weight=BF16 token_embd.weight=BF16
|
| 2 |
+
PASS mmproj-Nex-N2.5-mini-BF16.gguf arch=clip ftype=32 tensors=334 nextn=0 output.weight=None token_embd.weight=None
|
recipe/logs/D2_verify_download.log
ADDED
|
@@ -0,0 +1,2 @@
|
|
|
|
|
|
|
|
|
|
| 1 |
+
files=28 sha-verified=19 size-only=9 bad=0
|
| 2 |
+
RESULT: PASS
|
recipe/logs/N1_ppl_bf16.log
ADDED
|
@@ -0,0 +1,15 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
0.00.044.003 I common_init_result: fitting params to device memory ...
|
| 2 |
+
0.00.044.006 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
|
| 3 |
+
0.00.378.699 W llama_model_loader: direct I/O is enabled, disabling mmap
|
| 4 |
+
1.17.123.477 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
|
| 5 |
+
1.17.377.358 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|
| 6 |
+
1.18.283.894 I
|
| 7 |
+
1.18.284.021 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
|
| 8 |
+
1.18.284.129 I perplexity: saving all logits to kld/bf16.kld
|
| 9 |
+
1.18.284.134 I perplexity: tokenizing the input ..
|
| 10 |
+
1.18.617.280 I perplexity: tokenization took 333.135 ms
|
| 11 |
+
1.18.617.370 I perplexity: calculating perplexity over 40 chunks, n_ctx=2048, batch_size=2048, n_seq=1
|
| 12 |
+
1.23.386.958 I perplexity: 4.58 seconds per pass - ETA 3.05 minutes
|
| 13 |
+
[1]139.1466,[2]166.0748,[3]174.5814,[4]170.5873,[5]171.7312,[6]127.5767,[7]99.6563,[8]94.8236,[9]101.8520,[10]104.1594,[11]108.4984,[12]113.4569,[13]109.8859,[14]107.0003,[15]104.7124,[16]107.6891,[17]105.9352,[18]106.7826,[19]104.4693,[20]101.7712,[21]102.7759,[22]104.9236,[23]107.8566,[24]107.9859,[25]109.0207,[26]109.5885,[27]111.5203,[28]113.7558,[29]116.8379,[30]117.2211,[31]115.3159,[32]113.7423,[33]112.7430,[34]112.9143,[35]113.3873,[36]113.5581,[37]111.6761,[38]111.1729,[39]108.9264,[40]105.9103,
|
| 14 |
+
4.26.432.497 I Final estimate: PPL = 105.9103 +/- 1.77389
|
| 15 |
+
|
recipe/logs/N1c_ppl_bf16_cpu.log
ADDED
|
@@ -0,0 +1,15 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
0.00.040.030 I common_init_result: fitting params to device memory ...
|
| 2 |
+
0.00.040.033 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
|
| 3 |
+
0.00.314.545 I common_params_fit_impl: projected to use 66756 MiB of host memory vs. 127438 MiB of total host memory
|
| 4 |
+
0.00.652.877 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
|
| 5 |
+
0.00.672.925 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|
| 6 |
+
0.01.325.340 I
|
| 7 |
+
0.01.325.484 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
|
| 8 |
+
0.01.326.996 I perplexity: saving all logits to kld/bf16.kld
|
| 9 |
+
0.01.327.002 I perplexity: tokenizing the input ..
|
| 10 |
+
0.01.631.676 I perplexity: tokenization took 304.664 ms
|
| 11 |
+
0.01.631.766 I perplexity: calculating perplexity over 40 chunks, n_ctx=2048, batch_size=2048, n_seq=1
|
| 12 |
+
0.13.030.135 I perplexity: 11.26 seconds per pass - ETA 7.50 minutes
|
| 13 |
+
[1]5.6964,[2]6.6666,[3]7.0328,[4]7.2903,[5]7.1445,[6]6.1493,[7]5.7412,[8]5.6706,[9]5.9764,[10]6.0886,[11]6.1422,[12]6.3827,[13]6.4266,[14]6.4831,[15]6.5233,[16]6.6999,[17]6.7466,[18]6.8361,[19]6.7889,[20]6.5275,[21]6.5437,[22]6.5571,[23]6.6058,[24]6.6060,[25]6.6371,[26]6.6074,[27]6.7700,[28]6.8576,[29]6.8518,[30]6.7971,[31]6.6991,[32]6.6042,[33]6.5393,[34]6.5251,[35]6.5369,[36]6.5530,[37]6.4632,[38]6.3955,[39]6.3171,[40]6.2303,
|
| 14 |
+
7.32.185.106 I Final estimate: PPL = 6.2303 +/- 0.07538
|
| 15 |
+
|
recipe/logs/N2c_imatrix_cpu.log
ADDED
|
@@ -0,0 +1,114 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
0.00.047.060 I common_init_result: fitting params to device memory ...
|
| 2 |
+
0.00.047.066 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
|
| 3 |
+
0.00.538.240 I common_params_fit_impl: projected to use 66727 MiB of host memory vs. 127438 MiB of total host memory
|
| 4 |
+
0.01.009.512 W llama_context: n_ctx_seq (512) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
|
| 5 |
+
0.01.026.428 I
|
| 6 |
+
0.01.026.542 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
|
| 7 |
+
0.01.026.548 I compute_imatrix: tokenizing the input ..
|
| 8 |
+
0.01.115.784 I compute_imatrix: tokenization took 89.228 ms
|
| 9 |
+
0.01.115.820 I compute_imatrix: computing over 129 chunks, n_ctx=512, batch_size=512, n_seq=1
|
| 10 |
+
0.05.010.249 I compute_imatrix: 3.89 seconds per pass - ETA 8.37 minutes
|
| 11 |
+
[1]4.8280,[2]3.5310,[3]3.3565,[4]3.5543,[5]3.5306,[6]3.2674,[7]3.7535,[8]3.7744,[9]4.1909,0.35.054.817 W
|
| 12 |
+
save_imatrix: saving imatrix using GGUF format with a different suffix than .gguf
|
| 13 |
+
0.35.054.819 W save_imatrix: if you want the previous imatrix format, use --output-format dat
|
| 14 |
+
0.35.054.834 I
|
| 15 |
+
0.35.054.839 W save_imatrix: entry ' blk.36.ffn_up_exps.weight' has partial data (99.61%)
|
| 16 |
+
0.35.054.845 W save_imatrix: entry ' blk.34.ffn_down_exps.weight' has partial data (98.83%)
|
| 17 |
+
0.35.054.846 W save_imatrix: entry ' blk.34.ffn_up_exps.weight' has partial data (98.83%)
|
| 18 |
+
0.35.054.849 W save_imatrix: entry ' blk.33.ffn_gate_exps.weight' has partial data (99.61%)
|
| 19 |
+
0.35.054.852 W save_imatrix: entry ' blk.32.ffn_down_exps.weight' has partial data (99.61%)
|
| 20 |
+
0.35.054.856 W save_imatrix: entry ' blk.31.ffn_down_exps.weight' has partial data (99.61%)
|
| 21 |
+
0.35.054.857 W save_imatrix: entry ' blk.31.ffn_gate_exps.weight' has partial data (99.61%)
|
| 22 |
+
0.35.054.860 W save_imatrix: entry ' blk.30.ffn_down_exps.weight' has partial data (99.61%)
|
| 23 |
+
0.35.054.864 W save_imatrix: entry ' blk.29.ffn_up_exps.weight' has partial data (98.83%)
|
| 24 |
+
0.35.054.865 W save_imatrix: entry ' blk.29.ffn_gate_exps.weight' has partial data (98.83%)
|
| 25 |
+
0.35.054.874 W save_imatrix: entry ' blk.27.ffn_gate_exps.weight' has partial data (99.61%)
|
| 26 |
+
0.35.054.877 W save_imatrix: entry ' blk.26.ffn_up_exps.weight' has partial data (99.61%)
|
| 27 |
+
0.35.054.883 W save_imatrix: entry ' blk.24.ffn_gate_exps.weight' has partial data (98.83%)
|
| 28 |
+
0.35.054.893 W save_imatrix: entry ' blk.32.ffn_gate_exps.weight' has partial data (99.61%)
|
| 29 |
+
0.35.054.895 W save_imatrix: entry ' blk.33.ffn_down_exps.weight' has partial data (99.61%)
|
| 30 |
+
0.35.054.897 W save_imatrix: entry ' blk.30.ffn_gate_exps.weight' has partial data (99.61%)
|
| 31 |
+
0.35.054.903 W save_imatrix: entry ' blk.26.ffn_gate_exps.weight' has partial data (99.61%)
|
| 32 |
+
0.35.054.907 W save_imatrix: entry ' blk.20.ffn_down_exps.weight' has partial data (98.83%)
|
| 33 |
+
0.35.054.910 W save_imatrix: entry ' blk.22.ffn_down_exps.weight' has partial data (99.22%)
|
| 34 |
+
0.35.054.914 W save_imatrix: entry ' blk.27.ffn_down_exps.weight' has partial data (99.61%)
|
| 35 |
+
0.35.054.916 W save_imatrix: entry ' blk.26.ffn_down_exps.weight' has partial data (99.61%)
|
| 36 |
+
0.35.054.919 W save_imatrix: entry ' blk.20.ffn_gate_exps.weight' has partial data (98.83%)
|
| 37 |
+
0.35.054.926 W save_imatrix: entry ' blk.22.ffn_up_exps.weight' has partial data (99.22%)
|
| 38 |
+
0.35.054.928 W save_imatrix: entry ' blk.7.ffn_down_exps.weight' has partial data (99.22%)
|
| 39 |
+
0.35.054.929 W save_imatrix: entry ' blk.36.ffn_gate_exps.weight' has partial data (99.61%)
|
| 40 |
+
0.35.054.933 W save_imatrix: entry ' blk.31.ffn_up_exps.weight' has partial data (99.61%)
|
| 41 |
+
0.35.054.936 W save_imatrix: entry ' blk.20.ffn_up_exps.weight' has partial data (98.83%)
|
| 42 |
+
0.35.054.949 W save_imatrix: entry ' blk.15.ffn_down_exps.weight' has partial data (99.22%)
|
| 43 |
+
0.35.054.968 W save_imatrix: entry ' blk.34.ffn_gate_exps.weight' has partial data (98.83%)
|
| 44 |
+
0.35.054.969 W save_imatrix: entry ' blk.7.ffn_gate_exps.weight' has partial data (99.22%)
|
| 45 |
+
0.35.054.971 W save_imatrix: entry ' blk.33.ffn_up_exps.weight' has partial data (99.61%)
|
| 46 |
+
0.35.054.974 W save_imatrix: entry ' blk.24.ffn_down_exps.weight' has partial data (98.83%)
|
| 47 |
+
0.35.054.983 W save_imatrix: entry ' blk.32.ffn_up_exps.weight' has partial data (99.61%)
|
| 48 |
+
0.35.054.989 W save_imatrix: entry ' blk.18.ffn_down_exps.weight' has partial data (98.83%)
|
| 49 |
+
0.35.054.997 W save_imatrix: entry ' blk.29.ffn_down_exps.weight' has partial data (98.83%)
|
| 50 |
+
0.35.055.001 W save_imatrix: entry ' blk.7.ffn_up_exps.weight' has partial data (99.22%)
|
| 51 |
+
0.35.055.004 W save_imatrix: entry ' blk.18.ffn_up_exps.weight' has partial data (98.83%)
|
| 52 |
+
0.35.055.006 W save_imatrix: entry ' blk.27.ffn_up_exps.weight' has partial data (99.61%)
|
| 53 |
+
0.35.055.013 W save_imatrix: entry ' blk.16.ffn_gate_exps.weight' has partial data (98.83%)
|
| 54 |
+
0.35.055.016 W save_imatrix: entry ' blk.30.ffn_up_exps.weight' has partial data (99.61%)
|
| 55 |
+
0.35.055.019 W save_imatrix: entry ' blk.36.ffn_down_exps.weight' has partial data (99.61%)
|
| 56 |
+
0.35.055.022 W save_imatrix: entry ' blk.16.ffn_up_exps.weight' has partial data (98.83%)
|
| 57 |
+
0.35.055.026 W save_imatrix: entry ' blk.18.ffn_gate_exps.weight' has partial data (98.83%)
|
| 58 |
+
0.35.055.036 W save_imatrix: entry ' blk.24.ffn_up_exps.weight' has partial data (98.83%)
|
| 59 |
+
0.35.055.040 W save_imatrix: entry ' blk.15.ffn_gate_exps.weight' has partial data (99.22%)
|
| 60 |
+
0.35.055.041 W save_imatrix: entry ' blk.15.ffn_up_exps.weight' has partial data (99.22%)
|
| 61 |
+
0.35.055.042 W save_imatrix: entry ' blk.22.ffn_gate_exps.weight' has partial data (99.22%)
|
| 62 |
+
0.35.055.045 W save_imatrix: entry ' blk.16.ffn_down_exps.weight' has partial data (98.83%)
|
| 63 |
+
|
| 64 |
+
[10]4.3240,[11]4.0211,[12]4.4049,[13]4.9418,[14]5.1664,[15]5.6128,[16]5.8282,[17]6.1242,[18]6.4430,[19]6.1720,1.16.116.747 W
|
| 65 |
+
save_imatrix: saving imatrix using GGUF format with a different suffix than .gguf
|
| 66 |
+
1.16.116.747 W save_imatrix: if you want the previous imatrix format, use --output-format dat
|
| 67 |
+
|
| 68 |
+
[20]6.1999,[21]6.2244,[22]6.2384,[23]6.2298,[24]6.4258,[25]6.5635,[26]6.5954,[27]6.6758,[28]6.7712,[29]7.0466,1.59.082.255 W
|
| 69 |
+
save_imatrix: saving imatrix using GGUF format with a different suffix than .gguf
|
| 70 |
+
1.59.082.256 W save_imatrix: if you want the previous imatrix format, use --output-format dat
|
| 71 |
+
|
| 72 |
+
[30]7.0971,[31]6.8928,[32]6.5976,[33]6.3884,[34]6.2660,[35]6.2001,[36]6.3146,[37]6.4046,[38]6.4913,[39]6.6513,2.41.868.430 W
|
| 73 |
+
save_imatrix: saving imatrix using GGUF format with a different suffix than .gguf
|
| 74 |
+
2.41.868.431 W save_imatrix: if you want the previous imatrix format, use --output-format dat
|
| 75 |
+
|
| 76 |
+
[40]6.8336,[41]6.9385,[42]7.2103,[43]7.3229,[44]7.4938,[45]7.5505,[46]7.5098,[47]7.4647,[48]7.6200,[49]7.6994,3.24.706.480 W
|
| 77 |
+
save_imatrix: saving imatrix using GGUF format with a different suffix than .gguf
|
| 78 |
+
3.24.706.520 W save_imatrix: if you want the previous imatrix format, use --output-format dat
|
| 79 |
+
|
| 80 |
+
[50]7.6500,[51]7.5877,[52]7.6542,[53]7.7868,[54]7.8872,[55]7.9814,[56]8.0100,[57]8.0181,[58]8.0214,[59]8.0351,4.08.161.216 W
|
| 81 |
+
save_imatrix: saving imatrix using GGUF format with a different suffix than .gguf
|
| 82 |
+
4.08.161.218 W save_imatrix: if you want the previous imatrix format, use --output-format dat
|
| 83 |
+
|
| 84 |
+
[60]8.0326,[61]7.9840,[62]7.9372,[63]7.9683,[64]8.0068,[65]7.9439,[66]7.9358,[67]7.9336,[68]7.8425,[69]7.8031,5.53.139.568 W
|
| 85 |
+
save_imatrix: saving imatrix using GGUF format with a different suffix than .gguf
|
| 86 |
+
5.53.139.570 W save_imatrix: if you want the previous imatrix format, use --output-format dat
|
| 87 |
+
|
| 88 |
+
[70]7.8080,[71]7.7716,[72]7.7330,[73]7.7239,[74]7.6582,[75]7.5914,[76]7.5604,[77]7.5344,[78]7.5000,[79]7.4563,6.36.748.709 W
|
| 89 |
+
save_imatrix: saving imatrix using GGUF format with a different suffix than .gguf
|
| 90 |
+
6.36.748.710 W save_imatrix: if you want the previous imatrix format, use --output-format dat
|
| 91 |
+
|
| 92 |
+
[80]7.3732,[81]7.3905,[82]7.3907,[83]7.3636,[84]7.3800,[85]7.4067,[86]7.3364,[87]7.3262,[88]7.3183,[89]7.3359,7.14.145.562 W
|
| 93 |
+
save_imatrix: saving imatrix using GGUF format with a different suffix than .gguf
|
| 94 |
+
7.14.145.563 W save_imatrix: if you want the previous imatrix format, use --output-format dat
|
| 95 |
+
|
| 96 |
+
[90]7.3358,[91]7.3166,[92]7.2408,[93]7.1671,[94]7.0855,[95]7.0144,[96]6.9469,[97]6.8749,[98]6.8120,[99]6.7501,7.49.421.037 W
|
| 97 |
+
save_imatrix: saving imatrix using GGUF format with a different suffix than .gguf
|
| 98 |
+
7.49.421.039 W save_imatrix: if you want the previous imatrix format, use --output-format dat
|
| 99 |
+
|
| 100 |
+
[100]6.7749,[101]6.7937,[102]6.8753,[103]6.9573,[104]7.0285,[105]7.1404,[106]7.2385,[107]7.2672,[108]7.2873,[109]7.2996,8.25.152.163 W
|
| 101 |
+
save_imatrix: saving imatrix using GGUF format with a different suffix than .gguf
|
| 102 |
+
8.25.152.164 W save_imatrix: if you want the previous imatrix format, use --output-format dat
|
| 103 |
+
|
| 104 |
+
[110]7.3026,[111]7.2653,[112]7.1830,[113]7.1010,[114]7.1408,[115]7.1761,[116]7.1992,[117]7.2053,[118]7.2436,[119]7.2755,9.00.696.257 W
|
| 105 |
+
save_imatrix: saving imatrix using GGUF format with a different suffix than .gguf
|
| 106 |
+
9.00.696.259 W save_imatrix: if you want the previous imatrix format, use --output-format dat
|
| 107 |
+
|
| 108 |
+
[120]7.2818,[121]7.2985,[122]7.3085,[123]7.2788,[124]7.3333,[125]7.3907,[126]7.4360,[127]7.5037,[128]7.5475,[129]7.6012,
|
| 109 |
+
Final estimate: PPL = 7.6012 +/- 0.10821
|
| 110 |
+
9.36.441.617 W
|
| 111 |
+
save_imatrix: saving imatrix using GGUF format with a different suffix than .gguf
|
| 112 |
+
9.36.441.617 W save_imatrix: if you want the previous imatrix format, use --output-format dat
|
| 113 |
+
|
| 114 |
+
|
recipe/logs/N3_q102i.log
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
recipe/logs/N3_q103i.log
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
recipe/logs/N3_q106i.log
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
recipe/logs/N3_readback.log
ADDED
|
@@ -0,0 +1,4 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
PASS Nex-N2.5-mini-imatrix-Q4_0_ROCMFP4_STRIX_LEAN.gguf arch=qwen35moe ftype=106 tensors=733 nextn=0 output.weight=Q6_K token_embd.weight=Q5_K
|
| 2 |
+
PASS Nex-N2.5-mini-imatrix-Q4_0_ROCMFP4_COHERENT.gguf arch=qwen35moe ftype=102 tensors=733 nextn=0 output.weight=Q6_K token_embd.weight=Q6_K
|
| 3 |
+
PASS Nex-N2.5-mini-imatrix-Q4_0_ROCMFP4_FAST.gguf arch=qwen35moe ftype=103 tensors=733 nextn=0 output.weight=Q6_K token_embd.weight=Q4_0_ROCMFP4_FAST
|
| 4 |
+
N3_DONE
|
recipe/logs/N4_kld_q102i.log
ADDED
|
@@ -0,0 +1,92 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
0.00.037.317 I common_init_result: fitting params to device memory ...
|
| 2 |
+
0.00.037.321 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
|
| 3 |
+
0.00.390.383 W llama_model_loader: direct I/O is enabled, disabling mmap
|
| 4 |
+
0.21.870.791 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
|
| 5 |
+
0.21.936.318 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|
| 6 |
+
0.22.143.467 I
|
| 7 |
+
0.22.143.580 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
|
| 8 |
+
0.22.258.917 I kl_divergence: computing over 40 chunks, n_ctx=2048, batch_size=2048, n_seq=1
|
| 9 |
+
0.24.621.838 I kl_divergence: 2.36 seconds per pass - ETA 1.57 minutes
|
| 10 |
+
|
| 11 |
+
chunk PPL ln(PPL(Q)/PPL(base)) KL Divergence Δp RMS Same top p
|
| 12 |
+
1 5.6299 ± 0.4284 -0.01123 ± 0.01762 0.10095 ± 0.00634 9.176 ± 0.605 % 88.270 ± 1.007 %
|
| 13 |
+
2 6.6587 ± 0.3601 -0.00093 ± 0.01117 0.08359 ± 0.00360 7.978 ± 0.395 % 88.025 ± 0.718 %
|
| 14 |
+
3 7.1385 ± 0.3226 0.01508 ± 0.00875 0.08242 ± 0.00314 7.696 ± 0.308 % 87.911 ± 0.589 %
|
| 15 |
+
4 7.4207 ± 0.2990 0.01815 ± 0.00760 0.08463 ± 0.00438 7.777 ± 0.308 % 88.465 ± 0.499 %
|
| 16 |
+
5 7.2795 ± 0.2642 0.01905 ± 0.00665 0.08239 ± 0.00359 7.768 ± 0.267 % 88.426 ± 0.447 %
|
| 17 |
+
6 6.2745 ± 0.2014 0.02044 ± 0.00615 0.08353 ± 0.00339 8.576 ± 0.294 % 88.987 ± 0.400 %
|
| 18 |
+
7 5.8589 ± 0.1719 0.02053 ± 0.00613 0.09196 ± 0.00373 9.221 ± 0.305 % 88.926 ± 0.371 %
|
| 19 |
+
8 5.7865 ± 0.1577 0.02044 ± 0.00558 0.09026 ± 0.00333 9.081 ± 0.275 % 88.734 ± 0.350 %
|
| 20 |
+
9 6.0902 ± 0.1574 0.01904 ± 0.00522 0.09017 ± 0.00300 8.885 ± 0.253 % 88.530 ± 0.332 %
|
| 21 |
+
10 6.2036 ± 0.1532 0.01905 ± 0.00484 0.08664 ± 0.00271 8.652 ± 0.235 % 88.592 ± 0.314 %
|
| 22 |
+
11 6.2550 ± 0.1467 0.01852 ± 0.00454 0.08389 ± 0.00248 8.471 ± 0.220 % 88.661 ± 0.299 %
|
| 23 |
+
12 6.5002 ± 0.1472 0.01864 ± 0.00426 0.08096 ± 0.00228 8.271 ± 0.208 % 88.702 ± 0.286 %
|
| 24 |
+
13 6.5547 ± 0.1424 0.02010 ± 0.00407 0.07947 ± 0.00212 8.179 ± 0.196 % 88.668 ± 0.275 %
|
| 25 |
+
14 6.6139 ± 0.1383 0.02031 ± 0.00387 0.07794 ± 0.00199 8.054 ± 0.186 % 88.773 ± 0.264 %
|
| 26 |
+
15 6.6603 ± 0.1347 0.02109 ± 0.00372 0.07724 ± 0.00186 7.982 ± 0.177 % 88.739 ± 0.255 %
|
| 27 |
+
16 6.8133 ± 0.1336 0.01708 ± 0.00360 0.07665 ± 0.00178 7.924 ± 0.170 % 88.612 ± 0.248 %
|
| 28 |
+
17 6.8557 ± 0.1300 0.01631 ± 0.00347 0.07576 ± 0.00169 7.851 ± 0.164 % 88.615 ± 0.241 %
|
| 29 |
+
18 6.9493 ± 0.1282 0.01668 ± 0.00337 0.07548 ± 0.00160 7.811 ± 0.158 % 88.541 ± 0.235 %
|
| 30 |
+
19 6.9014 ± 0.1244 0.01668 ± 0.00325 0.07445 ± 0.00154 7.725 ± 0.153 % 88.645 ± 0.228 %
|
| 31 |
+
20 6.6526 ± 0.1161 0.01930 ± 0.00323 0.07848 ± 0.00152 7.994 ± 0.146 % 88.514 ± 0.223 %
|
| 32 |
+
21 6.6703 ± 0.1135 0.01945 ± 0.00316 0.07876 ± 0.00146 7.982 ± 0.141 % 88.423 ± 0.218 %
|
| 33 |
+
22 6.6837 ± 0.1111 0.01946 ± 0.00308 0.07952 ± 0.00143 8.019 ± 0.138 % 88.376 ± 0.214 %
|
| 34 |
+
23 6.7415 ± 0.1096 0.02066 ± 0.00300 0.07952 ± 0.00138 7.991 ± 0.134 % 88.278 ± 0.210 %
|
| 35 |
+
24 6.7373 ± 0.1070 0.02000 ± 0.00295 0.07956 ± 0.00138 7.986 ± 0.133 % 88.270 ± 0.205 %
|
| 36 |
+
25 6.7719 ± 0.1054 0.02041 ± 0.00289 0.07925 ± 0.00134 7.953 ± 0.130 % 88.258 ± 0.201 %
|
| 37 |
+
26 6.7420 ± 0.1028 0.02046 ± 0.00283 0.07951 ± 0.00131 7.991 ± 0.129 % 88.274 ± 0.197 %
|
| 38 |
+
27 6.9113 ± 0.1041 0.02093 ± 0.00283 0.07927 ± 0.00136 7.933 ± 0.126 % 88.248 ± 0.194 %
|
| 39 |
+
28 7.0022 ± 0.1039 0.02114 ± 0.00276 0.07854 ± 0.00132 7.873 ± 0.123 % 88.256 ± 0.190 %
|
| 40 |
+
29 6.9981 ± 0.1020 0.02141 ± 0.00271 0.07833 ± 0.00128 7.882 ± 0.121 % 88.250 ± 0.187 %
|
| 41 |
+
30 6.9536 ± 0.0995 0.02303 ± 0.00267 0.07848 ± 0.00125 7.892 ± 0.119 % 88.244 ± 0.184 %
|
| 42 |
+
31 6.8548 ± 0.0962 0.02323 ± 0.00261 0.07786 ± 0.00122 7.868 ± 0.117 % 88.320 ± 0.180 %
|
| 43 |
+
32 6.7489 ± 0.0931 0.02193 ± 0.00261 0.08043 ± 0.00142 8.006 ± 0.117 % 88.273 ± 0.178 %
|
| 44 |
+
33 6.6865 ± 0.0906 0.02252 ± 0.00256 0.07990 ± 0.00139 7.989 ± 0.115 % 88.299 ± 0.175 %
|
| 45 |
+
34 6.6723 ± 0.0889 0.02256 ± 0.00251 0.07907 ± 0.00135 7.928 ± 0.113 % 88.365 ± 0.172 %
|
| 46 |
+
35 6.6827 ± 0.0878 0.02231 ± 0.00246 0.07849 ± 0.00131 7.908 ± 0.111 % 88.404 ± 0.169 %
|
| 47 |
+
36 6.6954 ± 0.0868 0.02172 ± 0.00242 0.07826 ± 0.00128 7.892 ± 0.110 % 88.414 ± 0.167 %
|
| 48 |
+
37 6.6026 ± 0.0841 0.02157 ± 0.00238 0.07763 ± 0.00125 7.875 ± 0.108 % 88.407 ± 0.165 %
|
| 49 |
+
38 6.5309 ± 0.0818 0.02118 ± 0.00235 0.07760 ± 0.00123 7.895 ± 0.106 % 88.432 ± 0.162 %
|
| 50 |
+
39 6.4522 ± 0.0795 0.02138 ± 0.00232 0.07746 ± 0.00121 7.905 ± 0.105 % 88.423 ± 0.160 %
|
| 51 |
+
40 6.3601 ± 0.0770 0.02082 ± 0.00229 0.07695 ± 0.00118 7.909 ± 0.104 % 88.463 ± 0.158 %
|
| 52 |
+
|
| 53 |
+
====== Perplexity statistics ======
|
| 54 |
+
Mean PPL(Q) : 6.360050 ± 0.077045
|
| 55 |
+
Mean PPL(base) : 6.228979 ± 0.075322
|
| 56 |
+
Cor(ln(PPL(Q)), ln(PPL(base))): 98.21%
|
| 57 |
+
Mean ln(PPL(Q)/PPL(base)) : 0.020824 ± 0.002287
|
| 58 |
+
Mean PPL(Q)/PPL(base) : 1.021042 ± 0.002336
|
| 59 |
+
Mean PPL(Q)-PPL(base) : 0.131071 ± 0.014499
|
| 60 |
+
|
| 61 |
+
====== KL divergence statistics ======
|
| 62 |
+
Mean KLD: 0.076947 ± 0.001181
|
| 63 |
+
Maximum KLD: 16.293211
|
| 64 |
+
99.9% KLD: 2.499913
|
| 65 |
+
99.0% KLD: 0.685225
|
| 66 |
+
95.0% KLD: 0.259541
|
| 67 |
+
90.0% KLD: 0.162843
|
| 68 |
+
Median KLD: 0.034241
|
| 69 |
+
10.0% KLD: 0.000500
|
| 70 |
+
5.0% KLD: 0.000134
|
| 71 |
+
1.0% KLD: -0.000030
|
| 72 |
+
0.1% KLD: -0.000274
|
| 73 |
+
Minimum KLD: -0.000654
|
| 74 |
+
|
| 75 |
+
====== Token probability statistics ======
|
| 76 |
+
Mean Δp: -0.505 ± 0.039 %
|
| 77 |
+
Maximum Δp: 98.233%
|
| 78 |
+
99.9% Δp: 50.439%
|
| 79 |
+
99.0% Δp: 21.309%
|
| 80 |
+
95.0% Δp: 9.386%
|
| 81 |
+
90.0% Δp: 5.261%
|
| 82 |
+
75.0% Δp: 0.918%
|
| 83 |
+
Median Δp: -0.011%
|
| 84 |
+
25.0% Δp: -1.620%
|
| 85 |
+
10.0% Δp: -6.708%
|
| 86 |
+
5.0% Δp: -11.584%
|
| 87 |
+
1.0% Δp: -26.637%
|
| 88 |
+
0.1% Δp: -64.007%
|
| 89 |
+
Minimum Δp: -99.856%
|
| 90 |
+
RMS Δp : 7.909 ± 0.104 %
|
| 91 |
+
Same top p: 88.463 ± 0.158 %
|
| 92 |
+
|
recipe/logs/N4_kld_q103i.log
ADDED
|
@@ -0,0 +1,92 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
0.00.039.723 I common_init_result: fitting params to device memory ...
|
| 2 |
+
0.00.039.726 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
|
| 3 |
+
0.00.396.902 W llama_model_loader: direct I/O is enabled, disabling mmap
|
| 4 |
+
0.20.542.121 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
|
| 5 |
+
0.20.588.569 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|
| 6 |
+
0.20.783.826 I
|
| 7 |
+
0.20.783.939 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
|
| 8 |
+
0.20.898.218 I kl_divergence: computing over 40 chunks, n_ctx=2048, batch_size=2048, n_seq=1
|
| 9 |
+
0.22.619.669 I kl_divergence: 1.72 seconds per pass - ETA 1.13 minutes
|
| 10 |
+
|
| 11 |
+
chunk PPL ln(PPL(Q)/PPL(base)) KL Divergence Δp RMS Same top p
|
| 12 |
+
1 5.7545 ± 0.4426 0.01066 ± 0.01514 0.10310 ± 0.00565 9.252 ± 0.547 % 86.901 ± 1.055 %
|
| 13 |
+
2 6.7692 ± 0.3702 0.01553 ± 0.01057 0.08979 ± 0.00354 8.313 ± 0.392 % 87.097 ± 0.741 %
|
| 14 |
+
3 7.2118 ± 0.3286 0.02530 ± 0.00859 0.09142 ± 0.00371 7.924 ± 0.306 % 86.901 ± 0.609 %
|
| 15 |
+
4 7.4361 ± 0.3008 0.02022 ± 0.00807 0.09891 ± 0.00569 8.314 ± 0.317 % 87.414 ± 0.519 %
|
| 16 |
+
5 7.2981 ± 0.2661 0.02160 ± 0.00714 0.09587 ± 0.00463 8.276 ± 0.280 % 87.527 ± 0.462 %
|
| 17 |
+
6 6.3222 ± 0.2036 0.02801 ± 0.00679 0.09876 ± 0.00434 9.164 ± 0.306 % 87.814 ± 0.418 %
|
| 18 |
+
7 5.8837 ± 0.1729 0.02475 ± 0.00651 0.10562 ± 0.00468 9.504 ± 0.293 % 87.837 ± 0.386 %
|
| 19 |
+
8 5.8219 ± 0.1591 0.02653 ± 0.00598 0.10395 ± 0.00416 9.429 ± 0.268 % 87.732 ± 0.363 %
|
| 20 |
+
9 6.1195 ± 0.1584 0.02383 ± 0.00560 0.10437 ± 0.00375 9.307 ± 0.250 % 87.281 ± 0.347 %
|
| 21 |
+
10 6.2384 ± 0.1544 0.02466 ± 0.00519 0.10032 ± 0.00340 9.080 ± 0.233 % 87.331 ± 0.329 %
|
| 22 |
+
11 6.3092 ± 0.1486 0.02715 ± 0.00488 0.09769 ± 0.00312 8.996 ± 0.220 % 87.292 ± 0.314 %
|
| 23 |
+
12 6.5586 ± 0.1493 0.02759 ± 0.00459 0.09446 ± 0.00286 8.773 ± 0.208 % 87.235 ± 0.301 %
|
| 24 |
+
13 6.6005 ± 0.1440 0.02706 ± 0.00438 0.09304 ± 0.00266 8.721 ± 0.198 % 87.292 ± 0.289 %
|
| 25 |
+
14 6.6554 ± 0.1398 0.02657 ± 0.00417 0.09137 ± 0.00249 8.562 ± 0.187 % 87.264 ± 0.279 %
|
| 26 |
+
15 6.6958 ± 0.1359 0.02640 ± 0.00402 0.09068 ± 0.00234 8.515 ± 0.178 % 87.175 ± 0.270 %
|
| 27 |
+
16 6.8581 ± 0.1350 0.02363 ± 0.00387 0.08962 ± 0.00221 8.427 ± 0.172 % 87.060 ± 0.262 %
|
| 28 |
+
17 6.9067 ± 0.1315 0.02373 ± 0.00373 0.08844 ± 0.00209 8.335 ± 0.164 % 87.120 ± 0.254 %
|
| 29 |
+
18 6.9978 ± 0.1297 0.02365 ± 0.00361 0.08829 ± 0.00199 8.313 ± 0.158 % 87.075 ± 0.247 %
|
| 30 |
+
19 6.9441 ± 0.1257 0.02285 ± 0.00349 0.08712 ± 0.00190 8.208 ± 0.152 % 87.143 ± 0.240 %
|
| 31 |
+
20 6.6745 ± 0.1169 0.02258 ± 0.00348 0.09166 ± 0.00187 8.508 ± 0.146 % 87.082 ± 0.234 %
|
| 32 |
+
21 6.6839 ± 0.1140 0.02149 ± 0.00341 0.09200 ± 0.00180 8.487 ± 0.142 % 87.041 ± 0.229 %
|
| 33 |
+
22 6.6958 ± 0.1115 0.02127 ± 0.00336 0.09301 ± 0.00182 8.516 ± 0.141 % 87.079 ± 0.224 %
|
| 34 |
+
23 6.7561 ± 0.1101 0.02283 ± 0.00329 0.09310 ± 0.00178 8.515 ± 0.138 % 86.974 ± 0.219 %
|
| 35 |
+
24 6.7528 ± 0.1075 0.02230 ± 0.00321 0.09247 ± 0.00172 8.468 ± 0.134 % 87.036 ± 0.214 %
|
| 36 |
+
25 6.7907 ± 0.1060 0.02318 ± 0.00313 0.09180 ± 0.00165 8.416 ± 0.131 % 87.003 ± 0.210 %
|
| 37 |
+
26 6.7642 ± 0.1034 0.02375 ± 0.00307 0.09231 ± 0.00162 8.481 ± 0.129 % 87.067 ± 0.206 %
|
| 38 |
+
27 6.9325 ± 0.1047 0.02399 ± 0.00301 0.09164 ± 0.00159 8.400 ± 0.127 % 87.068 ± 0.202 %
|
| 39 |
+
28 7.0252 ± 0.1045 0.02441 ± 0.00295 0.09094 ± 0.00155 8.369 ± 0.125 % 87.125 ± 0.198 %
|
| 40 |
+
29 7.0215 ± 0.1026 0.02475 ± 0.00290 0.09104 ± 0.00151 8.421 ± 0.124 % 87.144 ± 0.194 %
|
| 41 |
+
30 6.9731 ± 0.1000 0.02583 ± 0.00286 0.09121 ± 0.00147 8.431 ± 0.121 % 87.120 ± 0.191 %
|
| 42 |
+
31 6.8743 ± 0.0967 0.02607 ± 0.00280 0.09057 ± 0.00143 8.408 ± 0.119 % 87.207 ± 0.188 %
|
| 43 |
+
32 6.7690 ± 0.0936 0.02490 ± 0.00278 0.09337 ± 0.00160 8.561 ± 0.119 % 87.201 ± 0.185 %
|
| 44 |
+
33 6.7036 ± 0.0910 0.02507 ± 0.00273 0.09280 ± 0.00156 8.553 ± 0.117 % 87.263 ± 0.181 %
|
| 45 |
+
34 6.6897 ± 0.0893 0.02516 ± 0.00267 0.09179 ± 0.00151 8.490 ± 0.115 % 87.315 ± 0.178 %
|
| 46 |
+
35 6.6976 ± 0.0881 0.02453 ± 0.00261 0.09113 ± 0.00148 8.437 ± 0.113 % 87.393 ± 0.175 %
|
| 47 |
+
36 6.7103 ± 0.0872 0.02395 ± 0.00257 0.09082 ± 0.00144 8.395 ± 0.110 % 87.404 ± 0.173 %
|
| 48 |
+
37 6.6143 ± 0.0845 0.02333 ± 0.00253 0.08990 ± 0.00141 8.356 ± 0.108 % 87.451 ± 0.170 %
|
| 49 |
+
38 6.5394 ± 0.0821 0.02247 ± 0.00250 0.08996 ± 0.00138 8.378 ± 0.107 % 87.452 ± 0.168 %
|
| 50 |
+
39 6.4598 ± 0.0798 0.02255 ± 0.00247 0.08964 ± 0.00135 8.375 ± 0.105 % 87.430 ± 0.166 %
|
| 51 |
+
40 6.3681 ± 0.0773 0.02208 ± 0.00243 0.08897 ± 0.00132 8.370 ± 0.103 % 87.454 ± 0.164 %
|
| 52 |
+
|
| 53 |
+
====== Perplexity statistics ======
|
| 54 |
+
Mean PPL(Q) : 6.368075 ± 0.077317
|
| 55 |
+
Mean PPL(base) : 6.228979 ± 0.075322
|
| 56 |
+
Cor(ln(PPL(Q)), ln(PPL(base))): 97.99%
|
| 57 |
+
Mean ln(PPL(Q)/PPL(base)) : 0.022085 ± 0.002430
|
| 58 |
+
Mean PPL(Q)/PPL(base) : 1.022330 ± 0.002484
|
| 59 |
+
Mean PPL(Q)-PPL(base) : 0.139096 ± 0.015431
|
| 60 |
+
|
| 61 |
+
====== KL divergence statistics ======
|
| 62 |
+
Mean KLD: 0.088974 ± 0.001318
|
| 63 |
+
Maximum KLD: 14.587295
|
| 64 |
+
99.9% KLD: 2.825675
|
| 65 |
+
99.0% KLD: 0.784672
|
| 66 |
+
95.0% KLD: 0.301152
|
| 67 |
+
90.0% KLD: 0.188552
|
| 68 |
+
Median KLD: 0.039738
|
| 69 |
+
10.0% KLD: 0.000526
|
| 70 |
+
5.0% KLD: 0.000124
|
| 71 |
+
1.0% KLD: -0.000056
|
| 72 |
+
0.1% KLD: -0.000307
|
| 73 |
+
Minimum KLD: -0.000906
|
| 74 |
+
|
| 75 |
+
====== Token probability statistics ======
|
| 76 |
+
Mean Δp: -0.439 ± 0.041 %
|
| 77 |
+
Maximum Δp: 97.378%
|
| 78 |
+
99.9% Δp: 51.015%
|
| 79 |
+
99.0% Δp: 22.861%
|
| 80 |
+
95.0% Δp: 10.138%
|
| 81 |
+
90.0% Δp: 5.884%
|
| 82 |
+
75.0% Δp: 1.147%
|
| 83 |
+
Median Δp: -0.001%
|
| 84 |
+
25.0% Δp: -1.525%
|
| 85 |
+
10.0% Δp: -6.868%
|
| 86 |
+
5.0% Δp: -12.274%
|
| 87 |
+
1.0% Δp: -29.206%
|
| 88 |
+
0.1% Δp: -67.613%
|
| 89 |
+
Minimum Δp: -98.545%
|
| 90 |
+
RMS Δp : 8.370 ± 0.103 %
|
| 91 |
+
Same top p: 87.454 ± 0.164 %
|
| 92 |
+
|
recipe/logs/N4_kld_q106.log
ADDED
|
@@ -0,0 +1,92 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
0.00.042.882 I common_init_result: fitting params to device memory ...
|
| 2 |
+
0.00.042.887 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
|
| 3 |
+
0.00.541.611 W llama_model_loader: direct I/O is enabled, disabling mmap
|
| 4 |
+
0.22.255.822 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
|
| 5 |
+
0.22.359.058 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|
| 6 |
+
0.22.643.059 I
|
| 7 |
+
0.22.643.196 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
|
| 8 |
+
0.22.776.319 I kl_divergence: computing over 40 chunks, n_ctx=2048, batch_size=2048, n_seq=1
|
| 9 |
+
0.25.307.079 I kl_divergence: 2.53 seconds per pass - ETA 1.68 minutes
|
| 10 |
+
|
| 11 |
+
chunk PPL ln(PPL(Q)/PPL(base)) KL Divergence Δp RMS Same top p
|
| 12 |
+
1 6.2617 ± 0.5005 0.09513 ± 0.01958 0.13331 ± 0.00723 11.273 ± 0.670 % 85.728 ± 1.094 %
|
| 13 |
+
2 7.1013 ± 0.3976 0.06343 ± 0.01246 0.11005 ± 0.00427 9.422 ± 0.440 % 86.413 ± 0.758 %
|
| 14 |
+
3 7.5058 ± 0.3468 0.06527 ± 0.01057 0.11313 ± 0.00423 9.250 ± 0.346 % 86.022 ± 0.626 %
|
| 15 |
+
4 7.7849 ± 0.3204 0.06605 ± 0.00908 0.11249 ± 0.00383 9.264 ± 0.306 % 86.388 ± 0.536 %
|
| 16 |
+
5 7.5801 ± 0.2810 0.05952 ± 0.00798 0.11043 ± 0.00322 9.087 ± 0.266 % 86.725 ± 0.474 %
|
| 17 |
+
6 6.5014 ± 0.2131 0.05596 ± 0.00724 0.11026 ± 0.00320 9.619 ± 0.279 % 87.178 ± 0.427 %
|
| 18 |
+
7 6.0313 ± 0.1806 0.04953 ± 0.00696 0.12172 ± 0.00436 10.069 ± 0.275 % 87.236 ± 0.394 %
|
| 19 |
+
8 5.9472 ± 0.1655 0.04783 ± 0.00647 0.12058 ± 0.00391 9.949 ± 0.252 % 87.121 ± 0.370 %
|
| 20 |
+
9 6.2446 ± 0.1643 0.04408 ± 0.00605 0.12012 ± 0.00353 9.791 ± 0.233 % 86.597 ± 0.355 %
|
| 21 |
+
10 6.3764 ± 0.1603 0.04654 ± 0.00565 0.11624 ± 0.00321 9.609 ± 0.217 % 86.588 ± 0.337 %
|
| 22 |
+
11 6.4374 ± 0.1538 0.04726 ± 0.00530 0.11268 ± 0.00294 9.442 ± 0.204 % 86.484 ± 0.322 %
|
| 23 |
+
12 6.6950 ± 0.1547 0.04816 ± 0.00499 0.10924 ± 0.00271 9.238 ± 0.193 % 86.388 ± 0.310 %
|
| 24 |
+
13 6.7262 ± 0.1490 0.04592 ± 0.00474 0.10773 ± 0.00253 9.174 ± 0.183 % 86.413 ± 0.297 %
|
| 25 |
+
14 6.7857 ± 0.1447 0.04595 ± 0.00454 0.10583 ± 0.00237 9.050 ± 0.174 % 86.426 ± 0.286 %
|
| 26 |
+
15 6.8279 ± 0.1408 0.04594 ± 0.00437 0.10522 ± 0.00224 9.021 ± 0.166 % 86.419 ± 0.277 %
|
| 27 |
+
16 6.9800 ± 0.1394 0.04125 ± 0.00424 0.10488 ± 0.00213 8.960 ± 0.160 % 86.339 ± 0.268 %
|
| 28 |
+
17 7.0246 ± 0.1356 0.04064 ± 0.00408 0.10348 ± 0.00202 8.863 ± 0.153 % 86.424 ± 0.260 %
|
| 29 |
+
18 7.1128 ± 0.1335 0.03994 ± 0.00395 0.10297 ± 0.00192 8.819 ± 0.147 % 86.385 ± 0.253 %
|
| 30 |
+
19 7.0597 ± 0.1294 0.03935 ± 0.00383 0.10135 ± 0.00184 8.726 ± 0.143 % 86.474 ± 0.245 %
|
| 31 |
+
20 6.7990 ± 0.1207 0.04105 ± 0.00383 0.10693 ± 0.00185 9.139 ± 0.143 % 86.393 ± 0.240 %
|
| 32 |
+
21 6.8267 ± 0.1182 0.04262 ± 0.00375 0.10765 ± 0.00179 9.144 ± 0.139 % 86.375 ± 0.234 %
|
| 33 |
+
22 6.8456 ± 0.1159 0.04339 ± 0.00369 0.10893 ± 0.00181 9.184 ± 0.137 % 86.417 ± 0.228 %
|
| 34 |
+
23 6.9000 ± 0.1143 0.04390 ± 0.00360 0.10839 ± 0.00175 9.147 ± 0.133 % 86.383 ± 0.224 %
|
| 35 |
+
24 6.8942 ± 0.1115 0.04302 ± 0.00351 0.10851 ± 0.00171 9.165 ± 0.132 % 86.376 ± 0.219 %
|
| 36 |
+
25 6.9297 ± 0.1099 0.04345 ± 0.00344 0.10819 ± 0.00165 9.128 ± 0.128 % 86.334 ± 0.215 %
|
| 37 |
+
26 6.8989 ± 0.1072 0.04347 ± 0.00338 0.10899 ± 0.00166 9.185 ± 0.127 % 86.371 ± 0.210 %
|
| 38 |
+
27 7.0668 ± 0.1084 0.04318 ± 0.00330 0.10805 ± 0.00160 9.103 ± 0.124 % 86.384 ± 0.206 %
|
| 39 |
+
28 7.1505 ± 0.1080 0.04210 ± 0.00321 0.10685 ± 0.00155 9.028 ± 0.122 % 86.395 ± 0.203 %
|
| 40 |
+
29 7.1560 ± 0.1063 0.04372 ± 0.00317 0.10671 ± 0.00152 9.048 ± 0.120 % 86.392 ± 0.199 %
|
| 41 |
+
30 7.0985 ± 0.1034 0.04365 ± 0.00312 0.10649 ± 0.00148 9.044 ± 0.118 % 86.341 ± 0.196 %
|
| 42 |
+
31 6.9900 ± 0.0998 0.04276 ± 0.00305 0.10563 ± 0.00144 9.018 ± 0.116 % 86.444 ± 0.192 %
|
| 43 |
+
32 6.8747 ± 0.0965 0.04039 ± 0.00305 0.10917 ± 0.00170 9.201 ± 0.119 % 86.416 ± 0.189 %
|
| 44 |
+
33 6.8130 ± 0.0939 0.04125 ± 0.00301 0.10876 ± 0.00166 9.199 ± 0.117 % 86.457 ± 0.186 %
|
| 45 |
+
34 6.7927 ± 0.0920 0.04044 ± 0.00294 0.10768 ± 0.00161 9.128 ± 0.115 % 86.441 ± 0.184 %
|
| 46 |
+
35 6.8056 ± 0.0909 0.04054 ± 0.00288 0.10692 ± 0.00157 9.083 ± 0.113 % 86.494 ± 0.181 %
|
| 47 |
+
36 6.8222 ± 0.0899 0.04049 ± 0.00284 0.10659 ± 0.00154 9.060 ± 0.111 % 86.546 ± 0.178 %
|
| 48 |
+
37 6.7263 ± 0.0872 0.04013 ± 0.00279 0.10571 ± 0.00150 9.033 ± 0.109 % 86.566 ± 0.175 %
|
| 49 |
+
38 6.6472 ± 0.0847 0.03882 ± 0.00275 0.10525 ± 0.00147 9.020 ± 0.107 % 86.621 ± 0.173 %
|
| 50 |
+
39 6.5675 ± 0.0823 0.03909 ± 0.00271 0.10501 ± 0.00144 9.022 ± 0.106 % 86.638 ± 0.170 %
|
| 51 |
+
40 6.4740 ± 0.0798 0.03858 ± 0.00267 0.10444 ± 0.00141 9.028 ± 0.105 % 86.659 ± 0.168 %
|
| 52 |
+
|
| 53 |
+
====== Perplexity statistics ======
|
| 54 |
+
Mean PPL(Q) : 6.473964 ± 0.079799
|
| 55 |
+
Mean PPL(base) : 6.228979 ± 0.075322
|
| 56 |
+
Cor(ln(PPL(Q)), ln(PPL(base))): 97.63%
|
| 57 |
+
Mean ln(PPL(Q)/PPL(base)) : 0.038576 ± 0.002667
|
| 58 |
+
Mean PPL(Q)/PPL(base) : 1.039330 ± 0.002772
|
| 59 |
+
Mean PPL(Q)-PPL(base) : 0.244985 ± 0.017453
|
| 60 |
+
|
| 61 |
+
====== KL divergence statistics ======
|
| 62 |
+
Mean KLD: 0.104436 ± 0.001413
|
| 63 |
+
Maximum KLD: 13.909879
|
| 64 |
+
99.9% KLD: 3.189192
|
| 65 |
+
99.0% KLD: 0.926687
|
| 66 |
+
95.0% KLD: 0.354656
|
| 67 |
+
90.0% KLD: 0.225739
|
| 68 |
+
Median KLD: 0.047702
|
| 69 |
+
10.0% KLD: 0.000698
|
| 70 |
+
5.0% KLD: 0.000183
|
| 71 |
+
1.0% KLD: -0.000033
|
| 72 |
+
0.1% KLD: -0.000321
|
| 73 |
+
Minimum KLD: -0.000718
|
| 74 |
+
|
| 75 |
+
====== Token probability statistics ======
|
| 76 |
+
Mean Δp: -0.244 ± 0.045 %
|
| 77 |
+
Maximum Δp: 99.706%
|
| 78 |
+
99.9% Δp: 59.493%
|
| 79 |
+
99.0% Δp: 25.100%
|
| 80 |
+
95.0% Δp: 11.707%
|
| 81 |
+
90.0% Δp: 6.963%
|
| 82 |
+
75.0% Δp: 1.453%
|
| 83 |
+
Median Δp: -0.001%
|
| 84 |
+
25.0% Δp: -1.479%
|
| 85 |
+
10.0% Δp: -7.240%
|
| 86 |
+
5.0% Δp: -12.992%
|
| 87 |
+
1.0% Δp: -31.268%
|
| 88 |
+
0.1% Δp: -68.467%
|
| 89 |
+
Minimum Δp: -98.412%
|
| 90 |
+
RMS Δp : 9.028 ± 0.105 %
|
| 91 |
+
Same top p: 86.659 ± 0.168 %
|
| 92 |
+
|
recipe/logs/N4_kld_q106i.log
ADDED
|
@@ -0,0 +1,92 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
0.00.061.730 I common_init_result: fitting params to device memory ...
|
| 2 |
+
0.00.061.734 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
|
| 3 |
+
0.00.418.389 W llama_model_loader: direct I/O is enabled, disabling mmap
|
| 4 |
+
0.21.286.189 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
|
| 5 |
+
0.21.331.409 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|
| 6 |
+
0.21.529.206 I
|
| 7 |
+
0.21.529.313 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
|
| 8 |
+
0.21.655.658 I kl_divergence: computing over 40 chunks, n_ctx=2048, batch_size=2048, n_seq=1
|
| 9 |
+
0.23.558.651 I kl_divergence: 1.90 seconds per pass - ETA 1.27 minutes
|
| 10 |
+
|
| 11 |
+
chunk PPL ln(PPL(Q)/PPL(base)) KL Divergence Δp RMS Same top p
|
| 12 |
+
1 5.7279 ± 0.4388 0.00602 ± 0.01614 0.10315 ± 0.00583 9.735 ± 0.571 % 85.924 ± 1.088 %
|
| 13 |
+
2 6.7923 ± 0.3712 0.01894 ± 0.01065 0.08743 ± 0.00335 8.343 ± 0.375 % 86.413 ± 0.758 %
|
| 14 |
+
3 7.2151 ± 0.3282 0.02577 ± 0.00843 0.08904 ± 0.00386 8.063 ± 0.325 % 86.999 ± 0.607 %
|
| 15 |
+
4 7.4647 ± 0.3021 0.02405 ± 0.00738 0.09108 ± 0.00535 8.270 ± 0.318 % 87.634 ± 0.515 %
|
| 16 |
+
5 7.3108 ± 0.2664 0.02334 ± 0.00659 0.08904 ± 0.00437 8.208 ± 0.278 % 87.801 ± 0.458 %
|
| 17 |
+
6 6.2917 ± 0.2025 0.02318 ± 0.00607 0.08932 ± 0.00383 8.727 ± 0.274 % 88.237 ± 0.411 %
|
| 18 |
+
7 5.8298 ± 0.1710 0.01556 ± 0.00599 0.09705 ± 0.00425 9.211 ± 0.278 % 88.242 ± 0.381 %
|
| 19 |
+
8 5.7710 ± 0.1572 0.01775 ± 0.00555 0.09614 ± 0.00379 9.187 ± 0.255 % 88.038 ± 0.359 %
|
| 20 |
+
9 6.0701 ± 0.1567 0.01574 ± 0.00523 0.09667 ± 0.00342 9.059 ± 0.236 % 87.553 ± 0.344 %
|
| 21 |
+
10 6.1910 ± 0.1528 0.01702 ± 0.00487 0.09332 ± 0.00310 8.829 ± 0.220 % 87.625 ± 0.326 %
|
| 22 |
+
11 6.2600 ± 0.1469 0.01931 ± 0.00460 0.09110 ± 0.00284 8.756 ± 0.210 % 87.630 ± 0.310 %
|
| 23 |
+
12 6.5141 ± 0.1478 0.02077 ± 0.00433 0.08827 ± 0.00261 8.573 ± 0.198 % 87.569 ± 0.298 %
|
| 24 |
+
13 6.5510 ± 0.1424 0.01954 ± 0.00414 0.08704 ± 0.00243 8.533 ± 0.188 % 87.533 ± 0.286 %
|
| 25 |
+
14 6.6159 ± 0.1386 0.02062 ± 0.00397 0.08562 ± 0.00227 8.395 ± 0.178 % 87.565 ± 0.276 %
|
| 26 |
+
15 6.6597 ± 0.1349 0.02101 ± 0.00382 0.08498 ± 0.00213 8.353 ± 0.170 % 87.481 ± 0.267 %
|
| 27 |
+
16 6.8126 ± 0.1337 0.01697 ± 0.00370 0.08434 ± 0.00203 8.278 ± 0.164 % 87.390 ± 0.259 %
|
| 28 |
+
17 6.8618 ± 0.1303 0.01721 ± 0.00357 0.08322 ± 0.00192 8.183 ± 0.157 % 87.511 ± 0.251 %
|
| 29 |
+
18 6.9534 ± 0.1285 0.01728 ± 0.00345 0.08273 ± 0.00182 8.130 ± 0.150 % 87.471 ± 0.244 %
|
| 30 |
+
19 6.9010 ± 0.1245 0.01662 ± 0.00334 0.08204 ± 0.00175 8.059 ± 0.147 % 87.508 ± 0.237 %
|
| 31 |
+
20 6.6434 ± 0.1160 0.01790 ± 0.00333 0.08651 ± 0.00173 8.357 ± 0.142 % 87.454 ± 0.232 %
|
| 32 |
+
21 6.6543 ± 0.1132 0.01704 ± 0.00325 0.08675 ± 0.00166 8.313 ± 0.137 % 87.399 ± 0.226 %
|
| 33 |
+
22 6.6714 ± 0.1109 0.01762 ± 0.00322 0.08786 ± 0.00171 8.337 ± 0.136 % 87.412 ± 0.221 %
|
| 34 |
+
23 6.7286 ± 0.1095 0.01874 ± 0.00316 0.08790 ± 0.00167 8.321 ± 0.134 % 87.390 ± 0.216 %
|
| 35 |
+
24 6.7323 ± 0.1070 0.01926 ± 0.00310 0.08781 ± 0.00164 8.305 ± 0.133 % 87.447 ± 0.211 %
|
| 36 |
+
25 6.7692 ± 0.1055 0.02001 ± 0.00302 0.08721 ± 0.00158 8.253 ± 0.129 % 87.457 ± 0.207 %
|
| 37 |
+
26 6.7420 ± 0.1029 0.02046 ± 0.00296 0.08758 ± 0.00155 8.276 ± 0.128 % 87.522 ± 0.203 %
|
| 38 |
+
27 6.9087 ± 0.1041 0.02056 ± 0.00290 0.08699 ± 0.00151 8.214 ± 0.125 % 87.520 ± 0.199 %
|
| 39 |
+
28 6.9975 ± 0.1039 0.02047 ± 0.00285 0.08640 ± 0.00147 8.197 ± 0.124 % 87.558 ± 0.195 %
|
| 40 |
+
29 6.9967 ± 0.1020 0.02121 ± 0.00280 0.08643 ± 0.00144 8.249 ± 0.123 % 87.538 ± 0.192 %
|
| 41 |
+
30 6.9524 ± 0.0996 0.02286 ± 0.00276 0.08671 ± 0.00140 8.260 ± 0.120 % 87.511 ± 0.189 %
|
| 42 |
+
31 6.8530 ± 0.0962 0.02296 ± 0.00270 0.08616 ± 0.00136 8.244 ± 0.118 % 87.623 ± 0.185 %
|
| 43 |
+
32 6.7437 ± 0.0931 0.02116 ± 0.00271 0.08891 ± 0.00158 8.411 ± 0.120 % 87.628 ± 0.182 %
|
| 44 |
+
33 6.6801 ± 0.0905 0.02156 ± 0.00266 0.08849 ± 0.00154 8.418 ± 0.119 % 87.666 ± 0.179 %
|
| 45 |
+
34 6.6705 ± 0.0889 0.02230 ± 0.00261 0.08778 ± 0.00150 8.378 ± 0.117 % 87.715 ± 0.176 %
|
| 46 |
+
35 6.6811 ± 0.0878 0.02207 ± 0.00255 0.08711 ± 0.00147 8.332 ± 0.115 % 87.792 ± 0.173 %
|
| 47 |
+
36 6.6976 ± 0.0869 0.02206 ± 0.00252 0.08700 ± 0.00143 8.296 ± 0.113 % 87.814 ± 0.170 %
|
| 48 |
+
37 6.6017 ± 0.0842 0.02143 ± 0.00248 0.08616 ± 0.00140 8.262 ± 0.111 % 87.823 ± 0.168 %
|
| 49 |
+
38 6.5268 ± 0.0818 0.02054 ± 0.00245 0.08605 ± 0.00137 8.265 ± 0.109 % 87.853 ± 0.166 %
|
| 50 |
+
39 6.4448 ± 0.0795 0.02022 ± 0.00241 0.08583 ± 0.00134 8.292 ± 0.107 % 87.829 ± 0.164 %
|
| 51 |
+
40 6.3536 ± 0.0770 0.01981 ± 0.00238 0.08516 ± 0.00131 8.281 ± 0.106 % 87.849 ± 0.162 %
|
| 52 |
+
|
| 53 |
+
====== Perplexity statistics ======
|
| 54 |
+
Mean PPL(Q) : 6.353616 ± 0.077001
|
| 55 |
+
Mean PPL(base) : 6.228979 ± 0.075322
|
| 56 |
+
Cor(ln(PPL(Q)), ln(PPL(base))): 98.08%
|
| 57 |
+
Mean ln(PPL(Q)/PPL(base)) : 0.019812 ± 0.002375
|
| 58 |
+
Mean PPL(Q)/PPL(base) : 1.020009 ± 0.002423
|
| 59 |
+
Mean PPL(Q)-PPL(base) : 0.124637 ± 0.015034
|
| 60 |
+
|
| 61 |
+
====== KL divergence statistics ======
|
| 62 |
+
Mean KLD: 0.085158 ± 0.001309
|
| 63 |
+
Maximum KLD: 16.350994
|
| 64 |
+
99.9% KLD: 2.569785
|
| 65 |
+
99.0% KLD: 0.737186
|
| 66 |
+
95.0% KLD: 0.287234
|
| 67 |
+
90.0% KLD: 0.182120
|
| 68 |
+
Median KLD: 0.038436
|
| 69 |
+
10.0% KLD: 0.000529
|
| 70 |
+
5.0% KLD: 0.000132
|
| 71 |
+
1.0% KLD: -0.000040
|
| 72 |
+
0.1% KLD: -0.000280
|
| 73 |
+
Minimum KLD: -0.000917
|
| 74 |
+
|
| 75 |
+
====== Token probability statistics ======
|
| 76 |
+
Mean Δp: -0.453 ± 0.041 %
|
| 77 |
+
Maximum Δp: 99.545%
|
| 78 |
+
99.9% Δp: 55.136%
|
| 79 |
+
99.0% Δp: 23.249%
|
| 80 |
+
95.0% Δp: 9.945%
|
| 81 |
+
90.0% Δp: 5.745%
|
| 82 |
+
75.0% Δp: 0.987%
|
| 83 |
+
Median Δp: -0.006%
|
| 84 |
+
25.0% Δp: -1.591%
|
| 85 |
+
10.0% Δp: -7.009%
|
| 86 |
+
5.0% Δp: -11.986%
|
| 87 |
+
1.0% Δp: -27.836%
|
| 88 |
+
0.1% Δp: -65.986%
|
| 89 |
+
Minimum Δp: -99.892%
|
| 90 |
+
RMS Δp : 8.281 ± 0.106 %
|
| 91 |
+
Same top p: 87.849 ± 0.162 %
|
| 92 |
+
|
recipe/logs/N4v_kld_q102i.log
ADDED
|
@@ -0,0 +1,93 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
0.00.035.061 I common_init_result: fitting params to device memory ...
|
| 2 |
+
0.00.035.064 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
|
| 3 |
+
0.00.364.958 W llama_model_loader: direct I/O is enabled, disabling mmap
|
| 4 |
+
0.01.365.200 W read_raw_unsafe: Falling back to buffered IO due to Bad address
|
| 5 |
+
0.20.724.587 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
|
| 6 |
+
0.20.776.655 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|
| 7 |
+
0.20.958.449 I
|
| 8 |
+
0.20.958.535 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
|
| 9 |
+
0.21.072.467 I kl_divergence: computing over 40 chunks, n_ctx=2048, batch_size=2048, n_seq=1
|
| 10 |
+
0.23.478.059 I kl_divergence: 2.41 seconds per pass - ETA 1.60 minutes
|
| 11 |
+
|
| 12 |
+
chunk PPL ln(PPL(Q)/PPL(base)) KL Divergence Δp RMS Same top p
|
| 13 |
+
1 5.5867 ± 0.4228 -0.01894 ± 0.01754 0.10256 ± 0.00648 9.538 ± 0.610 % 88.368 ± 1.003 %
|
| 14 |
+
2 6.6450 ± 0.3596 -0.00300 ± 0.01113 0.08488 ± 0.00372 8.059 ± 0.404 % 88.270 ± 0.712 %
|
| 15 |
+
3 7.0875 ± 0.3197 0.00792 ± 0.00873 0.08299 ± 0.00332 7.716 ± 0.310 % 88.172 ± 0.583 %
|
| 16 |
+
4 7.3455 ± 0.2950 0.00796 ± 0.00746 0.08559 ± 0.00479 7.862 ± 0.310 % 88.759 ± 0.494 %
|
| 17 |
+
5 7.2194 ± 0.2615 0.01075 ± 0.00655 0.08256 ± 0.00391 7.727 ± 0.267 % 88.524 ± 0.446 %
|
| 18 |
+
6 6.2395 ± 0.1998 0.01485 ± 0.00603 0.08240 ± 0.00357 8.402 ± 0.286 % 88.970 ± 0.400 %
|
| 19 |
+
7 5.8467 ± 0.1714 0.01845 ± 0.00617 0.09250 ± 0.00438 9.020 ± 0.301 % 88.884 ± 0.371 %
|
| 20 |
+
8 5.7766 ± 0.1573 0.01872 ± 0.00560 0.09000 ± 0.00387 8.841 ± 0.272 % 88.869 ± 0.348 %
|
| 21 |
+
9 6.0753 ± 0.1568 0.01659 ± 0.00524 0.08968 ± 0.00347 8.651 ± 0.250 % 88.617 ± 0.331 %
|
| 22 |
+
10 6.1938 ± 0.1528 0.01748 ± 0.00486 0.08628 ± 0.00314 8.443 ± 0.232 % 88.778 ± 0.312 %
|
| 23 |
+
11 6.2507 ± 0.1465 0.01783 ± 0.00456 0.08369 ± 0.00287 8.284 ± 0.217 % 88.812 ± 0.297 %
|
| 24 |
+
12 6.4984 ± 0.1471 0.01836 ± 0.00428 0.08087 ± 0.00264 8.100 ± 0.205 % 88.840 ± 0.284 %
|
| 25 |
+
13 6.5527 ± 0.1424 0.01980 ± 0.00408 0.07955 ± 0.00245 8.022 ± 0.193 % 88.646 ± 0.275 %
|
| 26 |
+
14 6.6129 ± 0.1383 0.02016 ± 0.00389 0.07798 ± 0.00228 7.902 ± 0.183 % 88.752 ± 0.264 %
|
| 27 |
+
15 6.6619 ± 0.1348 0.02133 ± 0.00373 0.07730 ± 0.00214 7.842 ± 0.174 % 88.693 ± 0.256 %
|
| 28 |
+
16 6.8209 ± 0.1338 0.01819 ± 0.00361 0.07652 ± 0.00203 7.773 ± 0.167 % 88.618 ± 0.248 %
|
| 29 |
+
17 6.8676 ± 0.1304 0.01805 ± 0.00348 0.07551 ± 0.00192 7.696 ± 0.162 % 88.609 ± 0.241 %
|
| 30 |
+
18 6.9618 ± 0.1286 0.01849 ± 0.00338 0.07510 ± 0.00182 7.653 ± 0.155 % 88.552 ± 0.235 %
|
| 31 |
+
19 6.9170 ± 0.1249 0.01894 ± 0.00325 0.07404 ± 0.00175 7.561 ± 0.151 % 88.661 ± 0.227 %
|
| 32 |
+
20 6.6681 ± 0.1166 0.02161 ± 0.00323 0.07796 ± 0.00171 7.836 ± 0.145 % 88.548 ± 0.223 %
|
| 33 |
+
21 6.6815 ± 0.1138 0.02112 ± 0.00316 0.07831 ± 0.00165 7.820 ± 0.139 % 88.428 ± 0.218 %
|
| 34 |
+
22 6.6992 ± 0.1116 0.02177 ± 0.00313 0.07958 ± 0.00169 7.862 ± 0.139 % 88.394 ± 0.214 %
|
| 35 |
+
23 6.7561 ± 0.1101 0.02283 ± 0.00306 0.07988 ± 0.00164 7.864 ± 0.136 % 88.270 ± 0.210 %
|
| 36 |
+
24 6.7528 ± 0.1074 0.02230 ± 0.00301 0.07975 ± 0.00161 7.848 ± 0.134 % 88.262 ± 0.205 %
|
| 37 |
+
25 6.7834 ± 0.1058 0.02210 ± 0.00294 0.07935 ± 0.00155 7.808 ± 0.131 % 88.223 ± 0.202 %
|
| 38 |
+
26 6.7551 ± 0.1032 0.02241 ± 0.00288 0.07966 ± 0.00152 7.846 ± 0.129 % 88.270 ± 0.197 %
|
| 39 |
+
27 6.9239 ± 0.1045 0.02275 ± 0.00288 0.07942 ± 0.00156 7.792 ± 0.127 % 88.244 ± 0.194 %
|
| 40 |
+
28 7.0172 ± 0.1043 0.02329 ± 0.00282 0.07891 ± 0.00153 7.749 ± 0.124 % 88.259 ± 0.190 %
|
| 41 |
+
29 7.0127 ± 0.1024 0.02349 ± 0.00277 0.07872 ± 0.00148 7.761 ± 0.122 % 88.263 ± 0.187 %
|
| 42 |
+
30 6.9664 ± 0.0999 0.02487 ± 0.00272 0.07871 ± 0.00144 7.761 ± 0.119 % 88.263 ± 0.184 %
|
| 43 |
+
31 6.8666 ± 0.0966 0.02496 ± 0.00265 0.07794 ± 0.00140 7.729 ± 0.117 % 88.377 ± 0.180 %
|
| 44 |
+
32 6.7575 ± 0.0934 0.02320 ± 0.00266 0.08099 ± 0.00163 7.910 ± 0.120 % 88.313 ± 0.178 %
|
| 45 |
+
33 6.6952 ± 0.0909 0.02382 ± 0.00260 0.08035 ± 0.00159 7.884 ± 0.118 % 88.371 ± 0.174 %
|
| 46 |
+
34 6.6814 ± 0.0892 0.02392 ± 0.00254 0.07951 ± 0.00154 7.831 ± 0.115 % 88.416 ± 0.172 %
|
| 47 |
+
35 6.6908 ± 0.0880 0.02351 ± 0.00249 0.07890 ± 0.00150 7.807 ± 0.113 % 88.465 ± 0.169 %
|
| 48 |
+
36 6.7037 ± 0.0870 0.02296 ± 0.00245 0.07859 ± 0.00146 7.786 ± 0.111 % 88.498 ± 0.166 %
|
| 49 |
+
37 6.6096 ± 0.0844 0.02263 ± 0.00241 0.07791 ± 0.00143 7.765 ± 0.109 % 88.492 ± 0.164 %
|
| 50 |
+
38 6.5383 ± 0.0821 0.02230 ± 0.00238 0.07778 ± 0.00140 7.790 ± 0.108 % 88.514 ± 0.162 %
|
| 51 |
+
39 6.4567 ± 0.0797 0.02207 ± 0.00234 0.07744 ± 0.00136 7.798 ± 0.106 % 88.508 ± 0.160 %
|
| 52 |
+
40 6.3632 ± 0.0772 0.02132 ± 0.00231 0.07684 ± 0.00133 7.783 ± 0.104 % 88.556 ± 0.157 %
|
| 53 |
+
|
| 54 |
+
====== Perplexity statistics ======
|
| 55 |
+
Mean PPL(Q) : 6.363210 ± 0.077202
|
| 56 |
+
Mean PPL(base) : 6.228979 ± 0.075322
|
| 57 |
+
Cor(ln(PPL(Q)), ln(PPL(base))): 98.19%
|
| 58 |
+
Mean ln(PPL(Q)/PPL(base)) : 0.021321 ± 0.002307
|
| 59 |
+
Mean PPL(Q)/PPL(base) : 1.021549 ± 0.002357
|
| 60 |
+
Mean PPL(Q)-PPL(base) : 0.134231 ± 0.014643
|
| 61 |
+
|
| 62 |
+
====== KL divergence statistics ======
|
| 63 |
+
Mean KLD: 0.076844 ± 0.001332
|
| 64 |
+
Maximum KLD: 16.142012
|
| 65 |
+
99.9% KLD: 2.816024
|
| 66 |
+
99.0% KLD: 0.669352
|
| 67 |
+
95.0% KLD: 0.255460
|
| 68 |
+
90.0% KLD: 0.161866
|
| 69 |
+
Median KLD: 0.033844
|
| 70 |
+
10.0% KLD: 0.000495
|
| 71 |
+
5.0% KLD: 0.000131
|
| 72 |
+
1.0% KLD: -0.000030
|
| 73 |
+
0.1% KLD: -0.000287
|
| 74 |
+
Minimum KLD: -0.000637
|
| 75 |
+
|
| 76 |
+
====== Token probability statistics ======
|
| 77 |
+
Mean Δp: -0.458 ± 0.038 %
|
| 78 |
+
Maximum Δp: 99.702%
|
| 79 |
+
99.9% Δp: 52.654%
|
| 80 |
+
99.0% Δp: 20.917%
|
| 81 |
+
95.0% Δp: 9.373%
|
| 82 |
+
90.0% Δp: 5.248%
|
| 83 |
+
75.0% Δp: 0.940%
|
| 84 |
+
Median Δp: -0.010%
|
| 85 |
+
25.0% Δp: -1.540%
|
| 86 |
+
10.0% Δp: -6.629%
|
| 87 |
+
5.0% Δp: -11.288%
|
| 88 |
+
1.0% Δp: -25.973%
|
| 89 |
+
0.1% Δp: -60.301%
|
| 90 |
+
Minimum Δp: -99.838%
|
| 91 |
+
RMS Δp : 7.783 ± 0.104 %
|
| 92 |
+
Same top p: 88.556 ± 0.157 %
|
| 93 |
+
|
recipe/logs/N4v_kld_q103.log
ADDED
|
@@ -0,0 +1,93 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
0.00.042.837 I common_init_result: fitting params to device memory ...
|
| 2 |
+
0.00.042.842 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
|
| 3 |
+
0.00.444.710 W llama_model_loader: direct I/O is enabled, disabling mmap
|
| 4 |
+
0.01.459.450 W read_raw_unsafe: Falling back to buffered IO due to Bad address
|
| 5 |
+
0.21.377.631 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
|
| 6 |
+
0.21.428.701 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|
| 7 |
+
0.21.707.133 I
|
| 8 |
+
0.21.707.313 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
|
| 9 |
+
0.21.843.964 I kl_divergence: computing over 40 chunks, n_ctx=2048, batch_size=2048, n_seq=1
|
| 10 |
+
0.24.857.762 I kl_divergence: 3.01 seconds per pass - ETA 2.00 minutes
|
| 11 |
+
|
| 12 |
+
chunk PPL ln(PPL(Q)/PPL(base)) KL Divergence Δp RMS Same top p
|
| 13 |
+
1 6.1749 ± 0.4909 0.08117 ± 0.01859 0.12763 ± 0.00703 10.392 ± 0.610 % 85.337 ± 1.107 %
|
| 14 |
+
2 7.1137 ± 0.3991 0.06516 ± 0.01199 0.10730 ± 0.00416 9.008 ± 0.411 % 86.266 ± 0.761 %
|
| 15 |
+
3 7.5580 ± 0.3515 0.07219 ± 0.00958 0.10875 ± 0.00396 8.609 ± 0.316 % 85.989 ± 0.627 %
|
| 16 |
+
4 7.8280 ± 0.3239 0.07158 ± 0.00898 0.12059 ± 0.00615 9.140 ± 0.324 % 86.193 ± 0.539 %
|
| 17 |
+
5 7.6146 ± 0.2833 0.06405 ± 0.00801 0.11743 ± 0.00502 9.043 ± 0.279 % 86.549 ± 0.477 %
|
| 18 |
+
6 6.5667 ± 0.2161 0.06595 ± 0.00747 0.11888 ± 0.00464 9.905 ± 0.299 % 86.804 ± 0.432 %
|
| 19 |
+
7 6.0824 ± 0.1827 0.05798 ± 0.00709 0.12603 ± 0.00484 10.284 ± 0.293 % 86.929 ± 0.398 %
|
| 20 |
+
8 5.9850 ± 0.1670 0.05416 ± 0.00651 0.12399 ± 0.00430 10.146 ± 0.267 % 86.877 ± 0.373 %
|
| 21 |
+
9 6.2982 ± 0.1664 0.05262 ± 0.00612 0.12384 ± 0.00387 9.948 ± 0.246 % 86.445 ± 0.357 %
|
| 22 |
+
10 6.4302 ± 0.1624 0.05493 ± 0.00573 0.12014 ± 0.00351 9.781 ± 0.229 % 86.452 ± 0.338 %
|
| 23 |
+
11 6.4862 ± 0.1556 0.05482 ± 0.00537 0.11666 ± 0.00323 9.576 ± 0.215 % 86.564 ± 0.322 %
|
| 24 |
+
12 6.7484 ± 0.1564 0.05611 ± 0.00505 0.11329 ± 0.00297 9.389 ± 0.203 % 86.470 ± 0.309 %
|
| 25 |
+
13 6.7729 ± 0.1504 0.05285 ± 0.00482 0.11177 ± 0.00277 9.295 ± 0.192 % 86.525 ± 0.296 %
|
| 26 |
+
14 6.8361 ± 0.1461 0.05335 ± 0.00462 0.11058 ± 0.00260 9.192 ± 0.183 % 86.385 ± 0.287 %
|
| 27 |
+
15 6.8710 ± 0.1419 0.05223 ± 0.00443 0.11010 ± 0.00248 9.182 ± 0.175 % 86.373 ± 0.277 %
|
| 28 |
+
16 7.0369 ± 0.1408 0.04937 ± 0.00428 0.10941 ± 0.00236 9.094 ± 0.169 % 86.278 ± 0.269 %
|
| 29 |
+
17 7.0798 ± 0.1370 0.04848 ± 0.00412 0.10796 ± 0.00224 9.005 ± 0.162 % 86.361 ± 0.260 %
|
| 30 |
+
18 7.1674 ± 0.1348 0.04759 ± 0.00399 0.10751 ± 0.00213 8.990 ± 0.157 % 86.201 ± 0.254 %
|
| 31 |
+
19 7.1262 ± 0.1310 0.04873 ± 0.00389 0.10651 ± 0.00204 8.894 ± 0.152 % 86.227 ± 0.247 %
|
| 32 |
+
20 6.8634 ± 0.1222 0.05049 ± 0.00387 0.11186 ± 0.00203 9.240 ± 0.148 % 86.202 ± 0.241 %
|
| 33 |
+
21 6.8908 ± 0.1196 0.05197 ± 0.00379 0.11250 ± 0.00196 9.228 ± 0.144 % 86.175 ± 0.235 %
|
| 34 |
+
22 6.9164 ± 0.1174 0.05368 ± 0.00373 0.11404 ± 0.00194 9.255 ± 0.141 % 86.164 ± 0.230 %
|
| 35 |
+
23 6.9778 ± 0.1159 0.05511 ± 0.00365 0.11379 ± 0.00188 9.234 ± 0.139 % 86.128 ± 0.225 %
|
| 36 |
+
24 6.9773 ± 0.1133 0.05500 ± 0.00356 0.11399 ± 0.00184 9.243 ± 0.137 % 86.038 ± 0.221 %
|
| 37 |
+
25 7.0132 ± 0.1116 0.05542 ± 0.00348 0.11365 ± 0.00178 9.211 ± 0.133 % 86.037 ± 0.217 %
|
| 38 |
+
26 6.9824 ± 0.1089 0.05549 ± 0.00342 0.11396 ± 0.00176 9.257 ± 0.131 % 86.070 ± 0.212 %
|
| 39 |
+
27 7.1525 ± 0.1102 0.05524 ± 0.00337 0.11346 ± 0.00174 9.179 ± 0.129 % 86.072 ± 0.208 %
|
| 40 |
+
28 7.2389 ± 0.1098 0.05439 ± 0.00329 0.11219 ± 0.00168 9.107 ± 0.126 % 86.084 ± 0.205 %
|
| 41 |
+
29 7.2404 ± 0.1079 0.05545 ± 0.00323 0.11206 ± 0.00164 9.137 ± 0.125 % 86.048 ± 0.201 %
|
| 42 |
+
30 7.1812 ± 0.1050 0.05523 ± 0.00318 0.11167 ± 0.00159 9.127 ± 0.122 % 86.022 ± 0.198 %
|
| 43 |
+
31 7.0772 ± 0.1015 0.05516 ± 0.00311 0.11097 ± 0.00155 9.104 ± 0.119 % 86.151 ± 0.194 %
|
| 44 |
+
32 6.9648 ± 0.0981 0.05342 ± 0.00309 0.11357 ± 0.00169 9.252 ± 0.121 % 86.092 ± 0.191 %
|
| 45 |
+
33 6.9013 ± 0.0955 0.05413 ± 0.00306 0.11324 ± 0.00166 9.238 ± 0.119 % 86.087 ± 0.188 %
|
| 46 |
+
34 6.8839 ± 0.0937 0.05378 ± 0.00300 0.11241 ± 0.00162 9.209 ± 0.119 % 86.093 ± 0.186 %
|
| 47 |
+
35 6.8905 ± 0.0924 0.05293 ± 0.00293 0.11139 ± 0.00158 9.152 ± 0.116 % 86.178 ± 0.182 %
|
| 48 |
+
36 6.9081 ± 0.0914 0.05301 ± 0.00290 0.11111 ± 0.00155 9.127 ± 0.114 % 86.190 ± 0.180 %
|
| 49 |
+
37 6.8115 ± 0.0887 0.05270 ± 0.00284 0.11014 ± 0.00151 9.096 ± 0.112 % 86.238 ± 0.177 %
|
| 50 |
+
38 6.7312 ± 0.0861 0.05138 ± 0.00280 0.10982 ± 0.00148 9.100 ± 0.110 % 86.299 ± 0.174 %
|
| 51 |
+
39 6.6496 ± 0.0837 0.05151 ± 0.00276 0.10945 ± 0.00145 9.096 ± 0.108 % 86.312 ± 0.172 %
|
| 52 |
+
40 6.5513 ± 0.0811 0.05045 ± 0.00272 0.10876 ± 0.00142 9.101 ± 0.107 % 86.356 ± 0.170 %
|
| 53 |
+
|
| 54 |
+
====== Perplexity statistics ======
|
| 55 |
+
Mean PPL(Q) : 6.551305 ± 0.081055
|
| 56 |
+
Mean PPL(base) : 6.228979 ± 0.075322
|
| 57 |
+
Cor(ln(PPL(Q)), ln(PPL(base))): 97.55%
|
| 58 |
+
Mean ln(PPL(Q)/PPL(base)) : 0.050452 ± 0.002720
|
| 59 |
+
Mean PPL(Q)/PPL(base) : 1.051746 ± 0.002861
|
| 60 |
+
Mean PPL(Q)-PPL(base) : 0.322326 ± 0.018212
|
| 61 |
+
|
| 62 |
+
====== KL divergence statistics ======
|
| 63 |
+
Mean KLD: 0.108756 ± 0.001422
|
| 64 |
+
Maximum KLD: 12.410495
|
| 65 |
+
99.9% KLD: 3.419987
|
| 66 |
+
99.0% KLD: 0.930024
|
| 67 |
+
95.0% KLD: 0.366912
|
| 68 |
+
90.0% KLD: 0.235489
|
| 69 |
+
Median KLD: 0.050197
|
| 70 |
+
10.0% KLD: 0.000729
|
| 71 |
+
5.0% KLD: 0.000189
|
| 72 |
+
1.0% KLD: -0.000051
|
| 73 |
+
0.1% KLD: -0.000341
|
| 74 |
+
Minimum KLD: -0.000842
|
| 75 |
+
|
| 76 |
+
====== Token probability statistics ======
|
| 77 |
+
Mean Δp: -0.471 ± 0.045 %
|
| 78 |
+
Maximum Δp: 96.576%
|
| 79 |
+
99.9% Δp: 54.353%
|
| 80 |
+
99.0% Δp: 24.394%
|
| 81 |
+
95.0% Δp: 11.471%
|
| 82 |
+
90.0% Δp: 6.631%
|
| 83 |
+
75.0% Δp: 1.298%
|
| 84 |
+
Median Δp: -0.006%
|
| 85 |
+
25.0% Δp: -1.652%
|
| 86 |
+
10.0% Δp: -7.598%
|
| 87 |
+
5.0% Δp: -13.545%
|
| 88 |
+
1.0% Δp: -32.141%
|
| 89 |
+
0.1% Δp: -70.328%
|
| 90 |
+
Minimum Δp: -99.714%
|
| 91 |
+
RMS Δp : 9.101 ± 0.107 %
|
| 92 |
+
Same top p: 86.356 ± 0.170 %
|
| 93 |
+
|
recipe/logs/N4v_kld_q103i.log
ADDED
|
@@ -0,0 +1,93 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
0.00.033.949 I common_init_result: fitting params to device memory ...
|
| 2 |
+
0.00.033.953 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
|
| 3 |
+
0.00.371.029 W llama_model_loader: direct I/O is enabled, disabling mmap
|
| 4 |
+
0.01.312.258 W read_raw_unsafe: Falling back to buffered IO due to Bad address
|
| 5 |
+
0.19.504.583 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
|
| 6 |
+
0.19.553.410 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|
| 7 |
+
0.19.731.487 I
|
| 8 |
+
0.19.731.587 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
|
| 9 |
+
0.19.847.243 I kl_divergence: computing over 40 chunks, n_ctx=2048, batch_size=2048, n_seq=1
|
| 10 |
+
0.22.241.158 I kl_divergence: 2.39 seconds per pass - ETA 1.58 minutes
|
| 11 |
+
|
| 12 |
+
chunk PPL ln(PPL(Q)/PPL(base)) KL Divergence Δp RMS Same top p
|
| 13 |
+
1 5.7501 ± 0.4413 0.00989 ± 0.01572 0.10954 ± 0.00645 9.805 ± 0.583 % 85.630 ± 1.097 %
|
| 14 |
+
2 6.7714 ± 0.3709 0.01585 ± 0.01070 0.09308 ± 0.00377 8.559 ± 0.392 % 86.706 ± 0.751 %
|
| 15 |
+
3 7.1773 ± 0.3266 0.02051 ± 0.00847 0.09216 ± 0.00368 7.997 ± 0.296 % 86.934 ± 0.608 %
|
| 16 |
+
4 7.4782 ± 0.3040 0.02587 ± 0.00786 0.09738 ± 0.00472 8.531 ± 0.336 % 87.341 ± 0.520 %
|
| 17 |
+
5 7.3506 ± 0.2694 0.02877 ± 0.00704 0.09557 ± 0.00398 8.489 ± 0.308 % 87.468 ± 0.463 %
|
| 18 |
+
6 6.3429 ± 0.2051 0.03128 ± 0.00664 0.09662 ± 0.00373 9.145 ± 0.308 % 87.781 ± 0.418 %
|
| 19 |
+
7 5.9472 ± 0.1761 0.03550 ± 0.00668 0.10791 ± 0.00462 9.889 ± 0.320 % 87.795 ± 0.387 %
|
| 20 |
+
8 5.8786 ± 0.1618 0.03622 ± 0.00612 0.10626 ± 0.00411 9.788 ± 0.292 % 87.683 ± 0.363 %
|
| 21 |
+
9 6.1674 ± 0.1605 0.03164 ± 0.00571 0.10586 ± 0.00370 9.627 ± 0.271 % 87.227 ± 0.348 %
|
| 22 |
+
10 6.2930 ± 0.1566 0.03336 ± 0.00529 0.10181 ± 0.00335 9.388 ± 0.252 % 87.390 ± 0.328 %
|
| 23 |
+
11 6.3574 ± 0.1505 0.03475 ± 0.00499 0.09913 ± 0.00308 9.301 ± 0.238 % 87.292 ± 0.314 %
|
| 24 |
+
12 6.6087 ± 0.1511 0.03519 ± 0.00469 0.09595 ± 0.00284 9.076 ± 0.225 % 87.317 ± 0.300 %
|
| 25 |
+
13 6.6444 ± 0.1456 0.03370 ± 0.00446 0.09418 ± 0.00263 8.998 ± 0.213 % 87.367 ± 0.288 %
|
| 26 |
+
14 6.6988 ± 0.1413 0.03307 ± 0.00425 0.09228 ± 0.00246 8.826 ± 0.202 % 87.467 ± 0.277 %
|
| 27 |
+
15 6.7357 ± 0.1373 0.03235 ± 0.00408 0.09142 ± 0.00231 8.745 ± 0.192 % 87.488 ± 0.267 %
|
| 28 |
+
16 6.8972 ± 0.1363 0.02932 ± 0.00393 0.09032 ± 0.00219 8.621 ± 0.183 % 87.390 ± 0.259 %
|
| 29 |
+
17 6.9396 ± 0.1326 0.02848 ± 0.00379 0.08911 ± 0.00207 8.521 ± 0.175 % 87.436 ± 0.251 %
|
| 30 |
+
18 7.0315 ± 0.1308 0.02845 ± 0.00366 0.08879 ± 0.00197 8.486 ± 0.168 % 87.374 ± 0.245 %
|
| 31 |
+
19 6.9777 ± 0.1268 0.02767 ± 0.00353 0.08759 ± 0.00188 8.384 ± 0.163 % 87.385 ± 0.238 %
|
| 32 |
+
20 6.7038 ± 0.1178 0.02695 ± 0.00352 0.09228 ± 0.00186 8.685 ± 0.156 % 87.273 ± 0.233 %
|
| 33 |
+
21 6.7149 ± 0.1149 0.02612 ± 0.00345 0.09269 ± 0.00179 8.668 ± 0.151 % 87.162 ± 0.228 %
|
| 34 |
+
22 6.7275 ± 0.1124 0.02599 ± 0.00339 0.09343 ± 0.00179 8.678 ± 0.150 % 87.186 ± 0.223 %
|
| 35 |
+
23 6.7910 ± 0.1111 0.02798 ± 0.00333 0.09354 ± 0.00176 8.670 ± 0.147 % 87.080 ± 0.219 %
|
| 36 |
+
24 6.7874 ± 0.1085 0.02741 ± 0.00325 0.09330 ± 0.00171 8.652 ± 0.144 % 87.097 ± 0.214 %
|
| 37 |
+
25 6.8244 ± 0.1069 0.02813 ± 0.00316 0.09247 ± 0.00164 8.585 ± 0.140 % 87.077 ± 0.210 %
|
| 38 |
+
26 6.7976 ± 0.1043 0.02868 ± 0.00310 0.09292 ± 0.00161 8.642 ± 0.138 % 87.142 ± 0.205 %
|
| 39 |
+
27 6.9623 ± 0.1055 0.02829 ± 0.00303 0.09202 ± 0.00157 8.538 ± 0.135 % 87.129 ± 0.201 %
|
| 40 |
+
28 7.0526 ± 0.1053 0.02832 ± 0.00296 0.09127 ± 0.00153 8.496 ± 0.133 % 87.149 ± 0.198 %
|
| 41 |
+
29 7.0500 ± 0.1034 0.02880 ± 0.00292 0.09137 ± 0.00149 8.557 ± 0.131 % 87.191 ± 0.194 %
|
| 42 |
+
30 7.0006 ± 0.1008 0.02977 ± 0.00287 0.09152 ± 0.00145 8.567 ± 0.129 % 87.165 ± 0.191 %
|
| 43 |
+
31 6.9004 ± 0.0974 0.02987 ± 0.00281 0.09083 ± 0.00141 8.530 ± 0.126 % 87.251 ± 0.187 %
|
| 44 |
+
32 6.7986 ± 0.0943 0.02926 ± 0.00280 0.09330 ± 0.00152 8.688 ± 0.125 % 87.225 ± 0.184 %
|
| 45 |
+
33 6.7336 ± 0.0917 0.02953 ± 0.00275 0.09274 ± 0.00149 8.677 ± 0.123 % 87.266 ± 0.181 %
|
| 46 |
+
34 6.7179 ± 0.0899 0.02937 ± 0.00268 0.09177 ± 0.00145 8.617 ± 0.121 % 87.333 ± 0.178 %
|
| 47 |
+
35 6.7270 ± 0.0888 0.02891 ± 0.00263 0.09111 ± 0.00141 8.570 ± 0.118 % 87.398 ± 0.175 %
|
| 48 |
+
36 6.7405 ± 0.0878 0.02844 ± 0.00259 0.09075 ± 0.00138 8.522 ± 0.116 % 87.414 ± 0.173 %
|
| 49 |
+
37 6.6432 ± 0.0851 0.02769 ± 0.00255 0.08997 ± 0.00135 8.495 ± 0.114 % 87.409 ± 0.171 %
|
| 50 |
+
38 6.5666 ± 0.0827 0.02662 ± 0.00252 0.09001 ± 0.00132 8.514 ± 0.112 % 87.434 ± 0.168 %
|
| 51 |
+
39 6.4867 ± 0.0803 0.02671 ± 0.00249 0.08987 ± 0.00130 8.528 ± 0.111 % 87.408 ± 0.166 %
|
| 52 |
+
40 6.3929 ± 0.0779 0.02598 ± 0.00245 0.08915 ± 0.00127 8.509 ± 0.109 % 87.424 ± 0.164 %
|
| 53 |
+
|
| 54 |
+
====== Perplexity statistics ======
|
| 55 |
+
Mean PPL(Q) : 6.392929 ± 0.077851
|
| 56 |
+
Mean PPL(base) : 6.228979 ± 0.075322
|
| 57 |
+
Cor(ln(PPL(Q)), ln(PPL(base))): 97.97%
|
| 58 |
+
Mean ln(PPL(Q)/PPL(base)) : 0.025980 ± 0.002447
|
| 59 |
+
Mean PPL(Q)/PPL(base) : 1.026321 ± 0.002511
|
| 60 |
+
Mean PPL(Q)-PPL(base) : 0.163950 ± 0.015638
|
| 61 |
+
|
| 62 |
+
====== KL divergence statistics ======
|
| 63 |
+
Mean KLD: 0.089150 ± 0.001266
|
| 64 |
+
Maximum KLD: 13.245127
|
| 65 |
+
99.9% KLD: 3.279078
|
| 66 |
+
99.0% KLD: 0.784911
|
| 67 |
+
95.0% KLD: 0.302743
|
| 68 |
+
90.0% KLD: 0.190118
|
| 69 |
+
Median KLD: 0.039583
|
| 70 |
+
10.0% KLD: 0.000539
|
| 71 |
+
5.0% KLD: 0.000127
|
| 72 |
+
1.0% KLD: -0.000056
|
| 73 |
+
0.1% KLD: -0.000311
|
| 74 |
+
Minimum KLD: -0.000916
|
| 75 |
+
|
| 76 |
+
====== Token probability statistics ======
|
| 77 |
+
Mean Δp: -0.437 ± 0.042 %
|
| 78 |
+
Maximum Δp: 97.638%
|
| 79 |
+
99.9% Δp: 52.324%
|
| 80 |
+
99.0% Δp: 22.989%
|
| 81 |
+
95.0% Δp: 10.252%
|
| 82 |
+
90.0% Δp: 5.909%
|
| 83 |
+
75.0% Δp: 1.135%
|
| 84 |
+
Median Δp: -0.001%
|
| 85 |
+
25.0% Δp: -1.497%
|
| 86 |
+
10.0% Δp: -6.935%
|
| 87 |
+
5.0% Δp: -12.278%
|
| 88 |
+
1.0% Δp: -29.036%
|
| 89 |
+
0.1% Δp: -71.577%
|
| 90 |
+
Minimum Δp: -99.874%
|
| 91 |
+
RMS Δp : 8.509 ± 0.109 %
|
| 92 |
+
Same top p: 87.424 ± 0.164 %
|
| 93 |
+
|
recipe/logs/N4v_kld_q106.log
ADDED
|
@@ -0,0 +1,93 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
0.00.073.166 I common_init_result: fitting params to device memory ...
|
| 2 |
+
0.00.073.169 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
|
| 3 |
+
0.00.618.000 W llama_model_loader: direct I/O is enabled, disabling mmap
|
| 4 |
+
0.01.632.328 W read_raw_unsafe: Falling back to buffered IO due to Bad address
|
| 5 |
+
0.21.360.980 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
|
| 6 |
+
0.21.411.524 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|
| 7 |
+
0.21.696.824 I
|
| 8 |
+
0.21.696.976 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
|
| 9 |
+
0.21.834.422 I kl_divergence: computing over 40 chunks, n_ctx=2048, batch_size=2048, n_seq=1
|
| 10 |
+
0.25.227.075 I kl_divergence: 3.39 seconds per pass - ETA 2.25 minutes
|
| 11 |
+
|
| 12 |
+
chunk PPL ln(PPL(Q)/PPL(base)) KL Divergence Δp RMS Same top p
|
| 13 |
+
1 6.1880 ± 0.4939 0.08329 ± 0.01978 0.13233 ± 0.00748 10.697 ± 0.585 % 86.022 ± 1.085 %
|
| 14 |
+
2 7.0357 ± 0.3926 0.05414 ± 0.01246 0.10961 ± 0.00433 9.075 ± 0.394 % 86.559 ± 0.754 %
|
| 15 |
+
3 7.4624 ± 0.3452 0.05947 ± 0.01021 0.11204 ± 0.00434 9.072 ± 0.327 % 86.250 ± 0.622 %
|
| 16 |
+
4 7.7345 ± 0.3193 0.05957 ± 0.00891 0.11467 ± 0.00464 9.213 ± 0.314 % 86.486 ± 0.535 %
|
| 17 |
+
5 7.5351 ± 0.2801 0.05356 ± 0.00785 0.11152 ± 0.00383 8.995 ± 0.270 % 86.667 ± 0.475 %
|
| 18 |
+
6 6.4595 ± 0.2123 0.04949 ± 0.00714 0.11027 ± 0.00353 9.495 ± 0.276 % 87.113 ± 0.428 %
|
| 19 |
+
7 5.9967 ± 0.1800 0.04378 ± 0.00691 0.12283 ± 0.00457 10.169 ± 0.285 % 87.250 ± 0.394 %
|
| 20 |
+
8 5.9141 ± 0.1649 0.04224 ± 0.00637 0.12112 ± 0.00408 10.053 ± 0.261 % 87.170 ± 0.370 %
|
| 21 |
+
9 6.2093 ± 0.1636 0.03840 ± 0.00600 0.12020 ± 0.00367 9.847 ± 0.240 % 86.825 ± 0.352 %
|
| 22 |
+
10 6.3404 ± 0.1596 0.04087 ± 0.00561 0.11641 ± 0.00333 9.682 ± 0.223 % 86.823 ± 0.334 %
|
| 23 |
+
11 6.4032 ± 0.1532 0.04194 ± 0.00528 0.11306 ± 0.00305 9.498 ± 0.210 % 86.795 ± 0.319 %
|
| 24 |
+
12 6.6610 ± 0.1540 0.04308 ± 0.00496 0.10959 ± 0.00281 9.281 ± 0.198 % 86.763 ± 0.306 %
|
| 25 |
+
13 6.6960 ± 0.1485 0.04143 ± 0.00472 0.10783 ± 0.00262 9.199 ± 0.188 % 86.773 ± 0.294 %
|
| 26 |
+
14 6.7584 ± 0.1443 0.04192 ± 0.00452 0.10607 ± 0.00245 9.074 ± 0.179 % 86.657 ± 0.284 %
|
| 27 |
+
15 6.8026 ± 0.1405 0.04223 ± 0.00435 0.10576 ± 0.00235 9.049 ± 0.171 % 86.660 ± 0.274 %
|
| 28 |
+
16 6.9572 ± 0.1392 0.03797 ± 0.00421 0.10524 ± 0.00223 8.990 ± 0.164 % 86.571 ± 0.267 %
|
| 29 |
+
17 7.0033 ± 0.1354 0.03762 ± 0.00406 0.10400 ± 0.00212 8.906 ± 0.157 % 86.522 ± 0.259 %
|
| 30 |
+
18 7.0947 ± 0.1334 0.03740 ± 0.00393 0.10350 ± 0.00202 8.881 ± 0.152 % 86.532 ± 0.252 %
|
| 31 |
+
19 7.0450 ± 0.1294 0.03728 ± 0.00382 0.10187 ± 0.00193 8.789 ± 0.147 % 86.639 ± 0.244 %
|
| 32 |
+
20 6.7808 ± 0.1206 0.03838 ± 0.00381 0.10695 ± 0.00193 9.167 ± 0.147 % 86.569 ± 0.238 %
|
| 33 |
+
21 6.8071 ± 0.1180 0.03975 ± 0.00374 0.10769 ± 0.00187 9.179 ± 0.142 % 86.506 ± 0.233 %
|
| 34 |
+
22 6.8328 ± 0.1159 0.04152 ± 0.00369 0.10920 ± 0.00187 9.217 ± 0.139 % 86.501 ± 0.228 %
|
| 35 |
+
23 6.8923 ± 0.1144 0.04279 ± 0.00360 0.10882 ± 0.00181 9.197 ± 0.137 % 86.498 ± 0.223 %
|
| 36 |
+
24 6.8910 ± 0.1117 0.04255 ± 0.00351 0.10896 ± 0.00176 9.191 ± 0.134 % 86.490 ± 0.218 %
|
| 37 |
+
25 6.9237 ± 0.1100 0.04257 ± 0.00344 0.10863 ± 0.00170 9.157 ± 0.130 % 86.424 ± 0.214 %
|
| 38 |
+
26 6.8938 ± 0.1073 0.04273 ± 0.00338 0.10952 ± 0.00170 9.229 ± 0.130 % 86.435 ± 0.210 %
|
| 39 |
+
27 7.0651 ± 0.1086 0.04294 ± 0.00331 0.10883 ± 0.00167 9.149 ± 0.127 % 86.452 ± 0.206 %
|
| 40 |
+
28 7.1473 ± 0.1082 0.04166 ± 0.00323 0.10759 ± 0.00162 9.083 ± 0.125 % 86.514 ± 0.202 %
|
| 41 |
+
29 7.1505 ± 0.1064 0.04295 ± 0.00318 0.10739 ± 0.00158 9.106 ± 0.123 % 86.517 ± 0.198 %
|
| 42 |
+
30 7.0929 ± 0.1035 0.04287 ± 0.00312 0.10706 ± 0.00153 9.094 ± 0.121 % 86.452 ± 0.195 %
|
| 43 |
+
31 6.9865 ± 0.1000 0.04226 ± 0.00306 0.10620 ± 0.00149 9.057 ± 0.118 % 86.573 ± 0.191 %
|
| 44 |
+
32 6.8731 ± 0.0966 0.04016 ± 0.00304 0.10897 ± 0.00167 9.204 ± 0.119 % 86.550 ± 0.189 %
|
| 45 |
+
33 6.8112 ± 0.0941 0.04100 ± 0.00300 0.10867 ± 0.00163 9.201 ± 0.117 % 86.561 ± 0.186 %
|
| 46 |
+
34 6.7925 ± 0.0922 0.04042 ± 0.00293 0.10760 ± 0.00159 9.133 ± 0.115 % 86.594 ± 0.183 %
|
| 47 |
+
35 6.8061 ± 0.0911 0.04061 ± 0.00287 0.10675 ± 0.00155 9.093 ± 0.113 % 86.655 ± 0.180 %
|
| 48 |
+
36 6.8188 ± 0.0901 0.03999 ± 0.00283 0.10633 ± 0.00151 9.056 ± 0.111 % 86.717 ± 0.177 %
|
| 49 |
+
37 6.7235 ± 0.0873 0.03971 ± 0.00278 0.10555 ± 0.00148 9.021 ± 0.109 % 86.753 ± 0.174 %
|
| 50 |
+
38 6.6468 ± 0.0849 0.03876 ± 0.00274 0.10510 ± 0.00145 9.019 ± 0.107 % 86.762 ± 0.172 %
|
| 51 |
+
39 6.5673 ± 0.0825 0.03906 ± 0.00270 0.10487 ± 0.00142 9.014 ± 0.105 % 86.773 ± 0.170 %
|
| 52 |
+
40 6.4745 ± 0.0799 0.03865 ± 0.00266 0.10441 ± 0.00139 9.030 ± 0.104 % 86.794 ± 0.167 %
|
| 53 |
+
|
| 54 |
+
====== Perplexity statistics ======
|
| 55 |
+
Mean PPL(Q) : 6.474466 ± 0.079947
|
| 56 |
+
Mean PPL(base) : 6.228979 ± 0.075322
|
| 57 |
+
Cor(ln(PPL(Q)), ln(PPL(base))): 97.64%
|
| 58 |
+
Mean ln(PPL(Q)/PPL(base)) : 0.038654 ± 0.002665
|
| 59 |
+
Mean PPL(Q)/PPL(base) : 1.039410 ± 0.002770
|
| 60 |
+
Mean PPL(Q)-PPL(base) : 0.245487 ± 0.017468
|
| 61 |
+
|
| 62 |
+
====== KL divergence statistics ======
|
| 63 |
+
Mean KLD: 0.104408 ± 0.001395
|
| 64 |
+
Maximum KLD: 16.437456
|
| 65 |
+
99.9% KLD: 3.401226
|
| 66 |
+
99.0% KLD: 0.899358
|
| 67 |
+
95.0% KLD: 0.354356
|
| 68 |
+
90.0% KLD: 0.226252
|
| 69 |
+
Median KLD: 0.048009
|
| 70 |
+
10.0% KLD: 0.000697
|
| 71 |
+
5.0% KLD: 0.000180
|
| 72 |
+
1.0% KLD: -0.000044
|
| 73 |
+
0.1% KLD: -0.000336
|
| 74 |
+
Minimum KLD: -0.000813
|
| 75 |
+
|
| 76 |
+
====== Token probability statistics ======
|
| 77 |
+
Mean Δp: -0.202 ± 0.045 %
|
| 78 |
+
Maximum Δp: 99.758%
|
| 79 |
+
99.9% Δp: 56.495%
|
| 80 |
+
99.0% Δp: 25.703%
|
| 81 |
+
95.0% Δp: 11.848%
|
| 82 |
+
90.0% Δp: 6.930%
|
| 83 |
+
75.0% Δp: 1.471%
|
| 84 |
+
Median Δp: -0.001%
|
| 85 |
+
25.0% Δp: -1.409%
|
| 86 |
+
10.0% Δp: -7.175%
|
| 87 |
+
5.0% Δp: -13.010%
|
| 88 |
+
1.0% Δp: -31.209%
|
| 89 |
+
0.1% Δp: -69.269%
|
| 90 |
+
Minimum Δp: -98.138%
|
| 91 |
+
RMS Δp : 9.030 ± 0.104 %
|
| 92 |
+
Same top p: 86.794 ± 0.167 %
|
| 93 |
+
|
recipe/logs/N4v_kld_q106i.log
ADDED
|
@@ -0,0 +1,93 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
0.00.033.687 I common_init_result: fitting params to device memory ...
|
| 2 |
+
0.00.033.690 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
|
| 3 |
+
0.00.365.970 W llama_model_loader: direct I/O is enabled, disabling mmap
|
| 4 |
+
0.01.319.923 W read_raw_unsafe: Falling back to buffered IO due to Bad address
|
| 5 |
+
0.19.602.693 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
|
| 6 |
+
0.19.656.222 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|
| 7 |
+
0.19.834.989 I
|
| 8 |
+
0.19.835.081 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
|
| 9 |
+
0.19.951.235 I kl_divergence: computing over 40 chunks, n_ctx=2048, batch_size=2048, n_seq=1
|
| 10 |
+
0.22.342.997 I kl_divergence: 2.39 seconds per pass - ETA 1.58 minutes
|
| 11 |
+
|
| 12 |
+
chunk PPL ln(PPL(Q)/PPL(base)) KL Divergence Δp RMS Same top p
|
| 13 |
+
1 5.6983 ± 0.4331 0.00085 ± 0.01708 0.10779 ± 0.00685 10.030 ± 0.635 % 86.217 ± 1.078 %
|
| 14 |
+
2 6.7307 ± 0.3660 0.00982 ± 0.01106 0.08982 ± 0.00389 8.657 ± 0.431 % 86.950 ± 0.745 %
|
| 15 |
+
3 7.1804 ± 0.3254 0.02094 ± 0.00857 0.08982 ± 0.00369 8.264 ± 0.335 % 87.260 ± 0.602 %
|
| 16 |
+
4 7.4637 ± 0.3024 0.02393 ± 0.00766 0.09566 ± 0.00581 8.559 ± 0.352 % 87.732 ± 0.513 %
|
| 17 |
+
5 7.3233 ± 0.2674 0.02505 ± 0.00682 0.09342 ± 0.00473 8.497 ± 0.302 % 87.977 ± 0.455 %
|
| 18 |
+
6 6.3148 ± 0.2036 0.02684 ± 0.00631 0.09237 ± 0.00410 8.998 ± 0.290 % 88.368 ± 0.409 %
|
| 19 |
+
7 5.8753 ± 0.1731 0.02333 ± 0.00607 0.09805 ± 0.00414 9.342 ± 0.282 % 88.479 ± 0.377 %
|
| 20 |
+
8 5.8033 ± 0.1588 0.02333 ± 0.00556 0.09655 ± 0.00369 9.256 ± 0.258 % 88.331 ± 0.355 %
|
| 21 |
+
9 6.1048 ± 0.1583 0.02143 ± 0.00523 0.09704 ± 0.00332 9.085 ± 0.238 % 87.846 ± 0.341 %
|
| 22 |
+
10 6.2216 ± 0.1541 0.02196 ± 0.00487 0.09358 ± 0.00301 8.857 ± 0.221 % 87.869 ± 0.323 %
|
| 23 |
+
11 6.2877 ± 0.1481 0.02373 ± 0.00459 0.09124 ± 0.00277 8.771 ± 0.212 % 87.914 ± 0.307 %
|
| 24 |
+
12 6.5381 ± 0.1489 0.02446 ± 0.00432 0.08825 ± 0.00255 8.571 ± 0.200 % 87.854 ± 0.295 %
|
| 25 |
+
13 6.5779 ± 0.1436 0.02363 ± 0.00412 0.08706 ± 0.00237 8.553 ± 0.189 % 87.796 ± 0.284 %
|
| 26 |
+
14 6.6420 ± 0.1396 0.02456 ± 0.00395 0.08566 ± 0.00221 8.419 ± 0.180 % 87.767 ± 0.274 %
|
| 27 |
+
15 6.6847 ± 0.1359 0.02475 ± 0.00380 0.08492 ± 0.00208 8.367 ± 0.171 % 87.683 ± 0.265 %
|
| 28 |
+
16 6.8394 ± 0.1348 0.02090 ± 0.00368 0.08419 ± 0.00198 8.280 ± 0.165 % 87.616 ± 0.257 %
|
| 29 |
+
17 6.8859 ± 0.1312 0.02070 ± 0.00355 0.08311 ± 0.00187 8.185 ± 0.158 % 87.695 ± 0.249 %
|
| 30 |
+
18 6.9788 ± 0.1294 0.02093 ± 0.00343 0.08257 ± 0.00178 8.136 ± 0.151 % 87.640 ± 0.243 %
|
| 31 |
+
19 6.9229 ± 0.1253 0.01979 ± 0.00332 0.08171 ± 0.00170 8.036 ± 0.146 % 87.647 ± 0.236 %
|
| 32 |
+
20 6.6565 ± 0.1166 0.01988 ± 0.00330 0.08616 ± 0.00169 8.356 ± 0.141 % 87.542 ± 0.231 %
|
| 33 |
+
21 6.6723 ± 0.1139 0.01975 ± 0.00323 0.08644 ± 0.00162 8.308 ± 0.137 % 87.488 ± 0.226 %
|
| 34 |
+
22 6.6896 ± 0.1116 0.02034 ± 0.00316 0.08711 ± 0.00160 8.355 ± 0.137 % 87.514 ± 0.220 %
|
| 35 |
+
23 6.7506 ± 0.1102 0.02201 ± 0.00309 0.08713 ± 0.00156 8.332 ± 0.134 % 87.479 ± 0.216 %
|
| 36 |
+
24 6.7515 ± 0.1077 0.02210 ± 0.00303 0.08701 ± 0.00153 8.313 ± 0.132 % 87.553 ± 0.211 %
|
| 37 |
+
25 6.7898 ± 0.1062 0.02305 ± 0.00295 0.08650 ± 0.00148 8.258 ± 0.129 % 87.464 ± 0.207 %
|
| 38 |
+
26 6.7594 ± 0.1035 0.02304 ± 0.00290 0.08694 ± 0.00146 8.287 ± 0.127 % 87.540 ± 0.203 %
|
| 39 |
+
27 6.9231 ± 0.1046 0.02264 ± 0.00283 0.08601 ± 0.00141 8.197 ± 0.124 % 87.553 ± 0.199 %
|
| 40 |
+
28 7.0130 ± 0.1044 0.02268 ± 0.00277 0.08539 ± 0.00138 8.177 ± 0.123 % 87.586 ± 0.195 %
|
| 41 |
+
29 7.0144 ± 0.1026 0.02374 ± 0.00273 0.08551 ± 0.00135 8.241 ± 0.122 % 87.569 ± 0.192 %
|
| 42 |
+
30 6.9699 ± 0.1001 0.02537 ± 0.00270 0.08585 ± 0.00132 8.254 ± 0.120 % 87.563 ± 0.188 %
|
| 43 |
+
31 6.8694 ± 0.0967 0.02536 ± 0.00264 0.08520 ± 0.00128 8.219 ± 0.117 % 87.668 ± 0.185 %
|
| 44 |
+
32 6.7608 ± 0.0936 0.02369 ± 0.00262 0.08761 ± 0.00146 8.350 ± 0.118 % 87.662 ± 0.182 %
|
| 45 |
+
33 6.6975 ± 0.0910 0.02416 ± 0.00258 0.08714 ± 0.00143 8.347 ± 0.117 % 87.704 ± 0.179 %
|
| 46 |
+
34 6.6852 ± 0.0894 0.02449 ± 0.00252 0.08623 ± 0.00139 8.290 ± 0.114 % 87.775 ± 0.176 %
|
| 47 |
+
35 6.6941 ± 0.0882 0.02401 ± 0.00247 0.08558 ± 0.00135 8.241 ± 0.112 % 87.845 ± 0.173 %
|
| 48 |
+
36 6.7094 ± 0.0873 0.02381 ± 0.00243 0.08528 ± 0.00132 8.207 ± 0.110 % 87.846 ± 0.170 %
|
| 49 |
+
37 6.6134 ± 0.0846 0.02319 ± 0.00240 0.08448 ± 0.00129 8.182 ± 0.108 % 87.852 ± 0.168 %
|
| 50 |
+
38 6.5369 ± 0.0822 0.02208 ± 0.00237 0.08434 ± 0.00126 8.180 ± 0.106 % 87.897 ± 0.165 %
|
| 51 |
+
39 6.4545 ± 0.0798 0.02173 ± 0.00234 0.08418 ± 0.00123 8.209 ± 0.105 % 87.896 ± 0.163 %
|
| 52 |
+
40 6.3642 ± 0.0773 0.02147 ± 0.00231 0.08356 ± 0.00121 8.201 ± 0.103 % 87.903 ± 0.161 %
|
| 53 |
+
|
| 54 |
+
====== Perplexity statistics ======
|
| 55 |
+
Mean PPL(Q) : 6.364168 ± 0.077334
|
| 56 |
+
Mean PPL(base) : 6.228979 ± 0.075322
|
| 57 |
+
Cor(ln(PPL(Q)), ln(PPL(base))): 98.19%
|
| 58 |
+
Mean ln(PPL(Q)/PPL(base)) : 0.021471 ± 0.002306
|
| 59 |
+
Mean PPL(Q)/PPL(base) : 1.021703 ± 0.002356
|
| 60 |
+
Mean PPL(Q)-PPL(base) : 0.135189 ± 0.014650
|
| 61 |
+
|
| 62 |
+
====== KL divergence statistics ======
|
| 63 |
+
Mean KLD: 0.083563 ± 0.001206
|
| 64 |
+
Maximum KLD: 15.989422
|
| 65 |
+
99.9% KLD: 2.686481
|
| 66 |
+
99.0% KLD: 0.713036
|
| 67 |
+
95.0% KLD: 0.286778
|
| 68 |
+
90.0% KLD: 0.180925
|
| 69 |
+
Median KLD: 0.038000
|
| 70 |
+
10.0% KLD: 0.000531
|
| 71 |
+
5.0% KLD: 0.000125
|
| 72 |
+
1.0% KLD: -0.000041
|
| 73 |
+
0.1% KLD: -0.000270
|
| 74 |
+
Minimum KLD: -0.000789
|
| 75 |
+
|
| 76 |
+
====== Token probability statistics ======
|
| 77 |
+
Mean Δp: -0.435 ± 0.040 %
|
| 78 |
+
Maximum Δp: 99.469%
|
| 79 |
+
99.9% Δp: 52.136%
|
| 80 |
+
99.0% Δp: 22.860%
|
| 81 |
+
95.0% Δp: 9.882%
|
| 82 |
+
90.0% Δp: 5.626%
|
| 83 |
+
75.0% Δp: 1.034%
|
| 84 |
+
Median Δp: -0.004%
|
| 85 |
+
25.0% Δp: -1.525%
|
| 86 |
+
10.0% Δp: -6.894%
|
| 87 |
+
5.0% Δp: -11.769%
|
| 88 |
+
1.0% Δp: -27.478%
|
| 89 |
+
0.1% Δp: -66.247%
|
| 90 |
+
Minimum Δp: -99.814%
|
| 91 |
+
RMS Δp : 8.201 ± 0.103 %
|
| 92 |
+
Same top p: 87.903 ± 0.161 %
|
| 93 |
+
|
recipe/logs/N5_kld_q106_repeat.log
ADDED
|
@@ -0,0 +1,92 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
0.00.038.287 I common_init_result: fitting params to device memory ...
|
| 2 |
+
0.00.038.289 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
|
| 3 |
+
0.00.399.780 W llama_model_loader: direct I/O is enabled, disabling mmap
|
| 4 |
+
0.20.682.112 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
|
| 5 |
+
0.20.729.253 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|
| 6 |
+
0.20.921.065 I
|
| 7 |
+
0.20.921.189 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
|
| 8 |
+
0.21.035.122 I kl_divergence: computing over 40 chunks, n_ctx=2048, batch_size=2048, n_seq=1
|
| 9 |
+
0.22.991.105 I kl_divergence: 1.96 seconds per pass - ETA 1.30 minutes
|
| 10 |
+
|
| 11 |
+
chunk PPL ln(PPL(Q)/PPL(base)) KL Divergence Δp RMS Same top p
|
| 12 |
+
1 6.2617 ± 0.5005 0.09513 ± 0.01958 0.13331 ± 0.00723 11.273 ± 0.670 % 85.728 ± 1.094 %
|
| 13 |
+
2 7.1013 ± 0.3976 0.06343 ± 0.01246 0.11005 ± 0.00427 9.422 ± 0.440 % 86.413 ± 0.758 %
|
| 14 |
+
3 7.5058 ± 0.3468 0.06527 ± 0.01057 0.11313 ± 0.00423 9.250 ± 0.346 % 86.022 ± 0.626 %
|
| 15 |
+
4 7.7849 ± 0.3204 0.06605 ± 0.00908 0.11249 ± 0.00383 9.264 ± 0.306 % 86.388 ± 0.536 %
|
| 16 |
+
5 7.5801 ± 0.2810 0.05952 ± 0.00798 0.11043 ± 0.00322 9.087 ± 0.266 % 86.725 ± 0.474 %
|
| 17 |
+
6 6.5014 ± 0.2131 0.05596 ± 0.00724 0.11026 ± 0.00320 9.619 ± 0.279 % 87.178 ± 0.427 %
|
| 18 |
+
7 6.0313 ± 0.1806 0.04953 ± 0.00696 0.12172 ± 0.00436 10.069 ± 0.275 % 87.236 ± 0.394 %
|
| 19 |
+
8 5.9472 ± 0.1655 0.04783 ± 0.00647 0.12058 ± 0.00391 9.949 ± 0.252 % 87.121 ± 0.370 %
|
| 20 |
+
9 6.2446 ± 0.1643 0.04408 ± 0.00605 0.12012 ± 0.00353 9.791 ± 0.233 % 86.597 ± 0.355 %
|
| 21 |
+
10 6.3764 ± 0.1603 0.04654 ± 0.00565 0.11624 ± 0.00321 9.609 ± 0.217 % 86.588 ± 0.337 %
|
| 22 |
+
11 6.4374 ± 0.1538 0.04726 ± 0.00530 0.11268 ± 0.00294 9.442 ± 0.204 % 86.484 ± 0.322 %
|
| 23 |
+
12 6.6950 ± 0.1547 0.04816 ± 0.00499 0.10924 ± 0.00271 9.238 ± 0.193 % 86.388 ± 0.310 %
|
| 24 |
+
13 6.7262 ± 0.1490 0.04592 ± 0.00474 0.10773 ± 0.00253 9.174 ± 0.183 % 86.413 ± 0.297 %
|
| 25 |
+
14 6.7857 ± 0.1447 0.04595 ± 0.00454 0.10583 ± 0.00237 9.050 ± 0.174 % 86.426 ± 0.286 %
|
| 26 |
+
15 6.8279 ± 0.1408 0.04594 ± 0.00437 0.10522 ± 0.00224 9.021 ± 0.166 % 86.419 ± 0.277 %
|
| 27 |
+
16 6.9800 ± 0.1394 0.04125 ± 0.00424 0.10488 ± 0.00213 8.960 ± 0.160 % 86.339 ± 0.268 %
|
| 28 |
+
17 7.0246 ± 0.1356 0.04064 ± 0.00408 0.10348 ± 0.00202 8.863 ± 0.153 % 86.424 ± 0.260 %
|
| 29 |
+
18 7.1128 ± 0.1335 0.03994 ± 0.00395 0.10297 ± 0.00192 8.819 ± 0.147 % 86.385 ± 0.253 %
|
| 30 |
+
19 7.0597 ± 0.1294 0.03935 ± 0.00383 0.10135 ± 0.00184 8.726 ± 0.143 % 86.474 ± 0.245 %
|
| 31 |
+
20 6.7990 ± 0.1207 0.04105 ± 0.00383 0.10693 ± 0.00185 9.139 ± 0.143 % 86.393 ± 0.240 %
|
| 32 |
+
21 6.8267 ± 0.1182 0.04262 ± 0.00375 0.10765 ± 0.00179 9.144 ± 0.139 % 86.375 ± 0.234 %
|
| 33 |
+
22 6.8456 ± 0.1159 0.04339 ± 0.00369 0.10893 ± 0.00181 9.184 ± 0.137 % 86.417 ± 0.228 %
|
| 34 |
+
23 6.9000 ± 0.1143 0.04390 ± 0.00360 0.10839 ± 0.00175 9.147 ± 0.133 % 86.383 ± 0.224 %
|
| 35 |
+
24 6.8942 ± 0.1115 0.04302 ± 0.00351 0.10851 ± 0.00171 9.165 ± 0.132 % 86.376 ± 0.219 %
|
| 36 |
+
25 6.9297 ± 0.1099 0.04345 ± 0.00344 0.10819 ± 0.00165 9.128 ± 0.128 % 86.334 ± 0.215 %
|
| 37 |
+
26 6.8989 ± 0.1072 0.04347 ± 0.00338 0.10899 ± 0.00166 9.185 ± 0.127 % 86.371 ± 0.210 %
|
| 38 |
+
27 7.0668 ± 0.1084 0.04318 ± 0.00330 0.10805 ± 0.00160 9.103 ± 0.124 % 86.384 ± 0.206 %
|
| 39 |
+
28 7.1505 ± 0.1080 0.04210 ± 0.00321 0.10685 ± 0.00155 9.028 ± 0.122 % 86.395 ± 0.203 %
|
| 40 |
+
29 7.1560 ± 0.1063 0.04372 ± 0.00317 0.10671 ± 0.00152 9.048 ± 0.120 % 86.392 ± 0.199 %
|
| 41 |
+
30 7.0985 ± 0.1034 0.04365 ± 0.00312 0.10649 ± 0.00148 9.044 ± 0.118 % 86.341 ± 0.196 %
|
| 42 |
+
31 6.9900 ± 0.0998 0.04276 ± 0.00305 0.10563 ± 0.00144 9.018 ± 0.116 % 86.444 ± 0.192 %
|
| 43 |
+
32 6.8747 ± 0.0965 0.04039 ± 0.00305 0.10917 ± 0.00170 9.201 ± 0.119 % 86.416 ± 0.189 %
|
| 44 |
+
33 6.8130 ± 0.0939 0.04125 ± 0.00301 0.10876 ± 0.00166 9.199 ± 0.117 % 86.457 ± 0.186 %
|
| 45 |
+
34 6.7927 ± 0.0920 0.04044 ± 0.00294 0.10768 ± 0.00161 9.128 ± 0.115 % 86.441 ± 0.184 %
|
| 46 |
+
35 6.8056 ± 0.0909 0.04054 ± 0.00288 0.10692 ± 0.00157 9.083 ± 0.113 % 86.494 ± 0.181 %
|
| 47 |
+
36 6.8222 ± 0.0899 0.04049 ± 0.00284 0.10659 ± 0.00154 9.060 ± 0.111 % 86.546 ± 0.178 %
|
| 48 |
+
37 6.7263 ± 0.0872 0.04013 ± 0.00279 0.10571 ± 0.00150 9.033 ± 0.109 % 86.566 ± 0.175 %
|
| 49 |
+
38 6.6472 ± 0.0847 0.03882 ± 0.00275 0.10525 ± 0.00147 9.020 ± 0.107 % 86.621 ± 0.173 %
|
| 50 |
+
39 6.5675 ± 0.0823 0.03909 ± 0.00271 0.10501 ± 0.00144 9.022 ± 0.106 % 86.638 ± 0.170 %
|
| 51 |
+
40 6.4740 ± 0.0798 0.03858 ± 0.00267 0.10444 ± 0.00141 9.028 ± 0.105 % 86.659 ± 0.168 %
|
| 52 |
+
|
| 53 |
+
====== Perplexity statistics ======
|
| 54 |
+
Mean PPL(Q) : 6.473964 ± 0.079799
|
| 55 |
+
Mean PPL(base) : 6.228979 ± 0.075322
|
| 56 |
+
Cor(ln(PPL(Q)), ln(PPL(base))): 97.63%
|
| 57 |
+
Mean ln(PPL(Q)/PPL(base)) : 0.038576 ± 0.002667
|
| 58 |
+
Mean PPL(Q)/PPL(base) : 1.039330 ± 0.002772
|
| 59 |
+
Mean PPL(Q)-PPL(base) : 0.244985 ± 0.017453
|
| 60 |
+
|
| 61 |
+
====== KL divergence statistics ======
|
| 62 |
+
Mean KLD: 0.104436 ± 0.001413
|
| 63 |
+
Maximum KLD: 13.909879
|
| 64 |
+
99.9% KLD: 3.189192
|
| 65 |
+
99.0% KLD: 0.926687
|
| 66 |
+
95.0% KLD: 0.354656
|
| 67 |
+
90.0% KLD: 0.225739
|
| 68 |
+
Median KLD: 0.047702
|
| 69 |
+
10.0% KLD: 0.000698
|
| 70 |
+
5.0% KLD: 0.000183
|
| 71 |
+
1.0% KLD: -0.000033
|
| 72 |
+
0.1% KLD: -0.000321
|
| 73 |
+
Minimum KLD: -0.000718
|
| 74 |
+
|
| 75 |
+
====== Token probability statistics ======
|
| 76 |
+
Mean Δp: -0.244 ± 0.045 %
|
| 77 |
+
Maximum Δp: 99.706%
|
| 78 |
+
99.9% Δp: 59.493%
|
| 79 |
+
99.0% Δp: 25.100%
|
| 80 |
+
95.0% Δp: 11.707%
|
| 81 |
+
90.0% Δp: 6.963%
|
| 82 |
+
75.0% Δp: 1.453%
|
| 83 |
+
Median Δp: -0.001%
|
| 84 |
+
25.0% Δp: -1.479%
|
| 85 |
+
10.0% Δp: -7.240%
|
| 86 |
+
5.0% Δp: -12.992%
|
| 87 |
+
1.0% Δp: -31.268%
|
| 88 |
+
0.1% Δp: -68.467%
|
| 89 |
+
Minimum Δp: -98.412%
|
| 90 |
+
RMS Δp : 9.028 ± 0.105 %
|
| 91 |
+
Same top p: 86.659 ± 0.168 %
|
| 92 |
+
|
recipe/logs/N5v_kld_q106_repeat.log
ADDED
|
@@ -0,0 +1,93 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
0.00.034.429 I common_init_result: fitting params to device memory ...
|
| 2 |
+
0.00.034.432 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
|
| 3 |
+
0.00.372.660 W llama_model_loader: direct I/O is enabled, disabling mmap
|
| 4 |
+
0.01.317.954 W read_raw_unsafe: Falling back to buffered IO due to Bad address
|
| 5 |
+
0.19.648.031 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
|
| 6 |
+
0.19.696.562 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|
| 7 |
+
0.19.866.239 I
|
| 8 |
+
0.19.866.320 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
|
| 9 |
+
0.19.981.669 I kl_divergence: computing over 40 chunks, n_ctx=2048, batch_size=2048, n_seq=1
|
| 10 |
+
0.22.378.306 I kl_divergence: 2.40 seconds per pass - ETA 1.58 minutes
|
| 11 |
+
|
| 12 |
+
chunk PPL ln(PPL(Q)/PPL(base)) KL Divergence Δp RMS Same top p
|
| 13 |
+
1 6.1880 ± 0.4939 0.08329 ± 0.01978 0.13233 ± 0.00748 10.697 ± 0.585 % 86.022 ± 1.085 %
|
| 14 |
+
2 7.0357 ± 0.3926 0.05414 ± 0.01246 0.10961 ± 0.00433 9.075 ± 0.394 % 86.559 ± 0.754 %
|
| 15 |
+
3 7.4624 ± 0.3452 0.05947 ± 0.01021 0.11204 ± 0.00434 9.072 ± 0.327 % 86.250 ± 0.622 %
|
| 16 |
+
4 7.7345 ± 0.3193 0.05957 ± 0.00891 0.11467 ± 0.00464 9.213 ± 0.314 % 86.486 ± 0.535 %
|
| 17 |
+
5 7.5351 ± 0.2801 0.05356 ± 0.00785 0.11152 ± 0.00383 8.995 ± 0.270 % 86.667 ± 0.475 %
|
| 18 |
+
6 6.4595 ± 0.2123 0.04949 ± 0.00714 0.11027 ± 0.00353 9.495 ± 0.276 % 87.113 ± 0.428 %
|
| 19 |
+
7 5.9967 ± 0.1800 0.04378 ± 0.00691 0.12283 ± 0.00457 10.169 ± 0.285 % 87.250 ± 0.394 %
|
| 20 |
+
8 5.9141 ± 0.1649 0.04224 ± 0.00637 0.12112 ± 0.00408 10.053 ± 0.261 % 87.170 ± 0.370 %
|
| 21 |
+
9 6.2093 ± 0.1636 0.03840 ± 0.00600 0.12020 ± 0.00367 9.847 ± 0.240 % 86.825 ± 0.352 %
|
| 22 |
+
10 6.3404 ± 0.1596 0.04087 ± 0.00561 0.11641 ± 0.00333 9.682 ± 0.223 % 86.823 ± 0.334 %
|
| 23 |
+
11 6.4032 ± 0.1532 0.04194 ± 0.00528 0.11306 ± 0.00305 9.498 ± 0.210 % 86.795 ± 0.319 %
|
| 24 |
+
12 6.6610 ± 0.1540 0.04308 ± 0.00496 0.10959 ± 0.00281 9.281 ± 0.198 % 86.763 ± 0.306 %
|
| 25 |
+
13 6.6960 ± 0.1485 0.04143 ± 0.00472 0.10783 ± 0.00262 9.199 ± 0.188 % 86.773 ± 0.294 %
|
| 26 |
+
14 6.7584 ± 0.1443 0.04192 ± 0.00452 0.10607 ± 0.00245 9.074 ± 0.179 % 86.657 ± 0.284 %
|
| 27 |
+
15 6.8026 ± 0.1405 0.04223 ± 0.00435 0.10576 ± 0.00235 9.049 ± 0.171 % 86.660 ± 0.274 %
|
| 28 |
+
16 6.9572 ± 0.1392 0.03797 ± 0.00421 0.10524 ± 0.00223 8.990 ± 0.164 % 86.571 ± 0.267 %
|
| 29 |
+
17 7.0033 ± 0.1354 0.03762 ± 0.00406 0.10400 ± 0.00212 8.906 ± 0.157 % 86.522 ± 0.259 %
|
| 30 |
+
18 7.0947 ± 0.1334 0.03740 ± 0.00393 0.10350 ± 0.00202 8.881 ± 0.152 % 86.532 ± 0.252 %
|
| 31 |
+
19 7.0450 ± 0.1294 0.03728 ± 0.00382 0.10187 ± 0.00193 8.789 ± 0.147 % 86.639 ± 0.244 %
|
| 32 |
+
20 6.7808 ± 0.1206 0.03838 ± 0.00381 0.10695 ± 0.00193 9.167 ± 0.147 % 86.569 ± 0.238 %
|
| 33 |
+
21 6.8071 ± 0.1180 0.03975 ± 0.00374 0.10769 ± 0.00187 9.179 ± 0.142 % 86.506 ± 0.233 %
|
| 34 |
+
22 6.8328 ± 0.1159 0.04152 ± 0.00369 0.10920 ± 0.00187 9.217 ± 0.139 % 86.501 ± 0.228 %
|
| 35 |
+
23 6.8923 ± 0.1144 0.04279 ± 0.00360 0.10882 ± 0.00181 9.197 ± 0.137 % 86.498 ± 0.223 %
|
| 36 |
+
24 6.8910 ± 0.1117 0.04255 ± 0.00351 0.10896 ± 0.00176 9.191 ± 0.134 % 86.490 ± 0.218 %
|
| 37 |
+
25 6.9237 ± 0.1100 0.04257 ± 0.00344 0.10863 ± 0.00170 9.157 ± 0.130 % 86.424 ± 0.214 %
|
| 38 |
+
26 6.8938 ± 0.1073 0.04273 ± 0.00338 0.10952 ± 0.00170 9.229 ± 0.130 % 86.435 ± 0.210 %
|
| 39 |
+
27 7.0651 ± 0.1086 0.04294 ± 0.00331 0.10883 ± 0.00167 9.149 ± 0.127 % 86.452 ± 0.206 %
|
| 40 |
+
28 7.1473 ± 0.1082 0.04166 ± 0.00323 0.10759 ± 0.00162 9.083 ± 0.125 % 86.514 ± 0.202 %
|
| 41 |
+
29 7.1505 ± 0.1064 0.04295 ± 0.00318 0.10739 ± 0.00158 9.106 ± 0.123 % 86.517 ± 0.198 %
|
| 42 |
+
30 7.0929 ± 0.1035 0.04287 ± 0.00312 0.10706 ± 0.00153 9.094 ± 0.121 % 86.452 ± 0.195 %
|
| 43 |
+
31 6.9865 ± 0.1000 0.04226 ± 0.00306 0.10620 ± 0.00149 9.057 ± 0.118 % 86.573 ± 0.191 %
|
| 44 |
+
32 6.8731 ± 0.0966 0.04016 ± 0.00304 0.10897 ± 0.00167 9.204 ± 0.119 % 86.550 ± 0.189 %
|
| 45 |
+
33 6.8112 ± 0.0941 0.04100 ± 0.00300 0.10867 ± 0.00163 9.201 ± 0.117 % 86.561 ± 0.186 %
|
| 46 |
+
34 6.7925 ± 0.0922 0.04042 ± 0.00293 0.10760 ± 0.00159 9.133 ± 0.115 % 86.594 ± 0.183 %
|
| 47 |
+
35 6.8061 ± 0.0911 0.04061 ± 0.00287 0.10675 ± 0.00155 9.093 ± 0.113 % 86.655 ± 0.180 %
|
| 48 |
+
36 6.8188 ± 0.0901 0.03999 ± 0.00283 0.10633 ± 0.00151 9.056 ± 0.111 % 86.717 ± 0.177 %
|
| 49 |
+
37 6.7235 ± 0.0873 0.03971 ± 0.00278 0.10555 ± 0.00148 9.021 ± 0.109 % 86.753 ± 0.174 %
|
| 50 |
+
38 6.6468 ± 0.0849 0.03876 ± 0.00274 0.10510 ± 0.00145 9.019 ± 0.107 % 86.762 ± 0.172 %
|
| 51 |
+
39 6.5673 ± 0.0825 0.03906 ± 0.00270 0.10487 ± 0.00142 9.014 ± 0.105 % 86.773 ± 0.170 %
|
| 52 |
+
40 6.4745 ± 0.0799 0.03865 ± 0.00266 0.10441 ± 0.00139 9.030 ± 0.104 % 86.794 ± 0.167 %
|
| 53 |
+
|
| 54 |
+
====== Perplexity statistics ======
|
| 55 |
+
Mean PPL(Q) : 6.474466 ± 0.079947
|
| 56 |
+
Mean PPL(base) : 6.228979 ± 0.075322
|
| 57 |
+
Cor(ln(PPL(Q)), ln(PPL(base))): 97.64%
|
| 58 |
+
Mean ln(PPL(Q)/PPL(base)) : 0.038654 ± 0.002665
|
| 59 |
+
Mean PPL(Q)/PPL(base) : 1.039410 ± 0.002770
|
| 60 |
+
Mean PPL(Q)-PPL(base) : 0.245487 ± 0.017468
|
| 61 |
+
|
| 62 |
+
====== KL divergence statistics ======
|
| 63 |
+
Mean KLD: 0.104408 ± 0.001395
|
| 64 |
+
Maximum KLD: 16.437456
|
| 65 |
+
99.9% KLD: 3.401226
|
| 66 |
+
99.0% KLD: 0.899358
|
| 67 |
+
95.0% KLD: 0.354356
|
| 68 |
+
90.0% KLD: 0.226252
|
| 69 |
+
Median KLD: 0.048009
|
| 70 |
+
10.0% KLD: 0.000697
|
| 71 |
+
5.0% KLD: 0.000180
|
| 72 |
+
1.0% KLD: -0.000044
|
| 73 |
+
0.1% KLD: -0.000336
|
| 74 |
+
Minimum KLD: -0.000813
|
| 75 |
+
|
| 76 |
+
====== Token probability statistics ======
|
| 77 |
+
Mean Δp: -0.202 ± 0.045 %
|
| 78 |
+
Maximum Δp: 99.758%
|
| 79 |
+
99.9% Δp: 56.495%
|
| 80 |
+
99.0% Δp: 25.703%
|
| 81 |
+
95.0% Δp: 11.848%
|
| 82 |
+
90.0% Δp: 6.930%
|
| 83 |
+
75.0% Δp: 1.471%
|
| 84 |
+
Median Δp: -0.001%
|
| 85 |
+
25.0% Δp: -1.409%
|
| 86 |
+
10.0% Δp: -7.175%
|
| 87 |
+
5.0% Δp: -13.010%
|
| 88 |
+
1.0% Δp: -31.209%
|
| 89 |
+
0.1% Δp: -69.269%
|
| 90 |
+
Minimum Δp: -98.138%
|
| 91 |
+
RMS Δp : 9.030 ± 0.104 %
|
| 92 |
+
Same top p: 86.794 ± 0.167 %
|
| 93 |
+
|
recipe/logs/N6t_tools_c1.log
ADDED
|
@@ -0,0 +1,33 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
PASS think=True multi-arg: args={'city': 'Paris', 'unit': 'celsius'}
|
| 2 |
+
FAIL think=True nested-object: exception KeyError('tool_calls')
|
| 3 |
+
PASS think=True enum: unit=fahrenheit
|
| 4 |
+
PASS think=True correct-decline: content='391'
|
| 5 |
+
PASS think=True multi-turn: final='Tokyo is **21°C** with **clear skies**.'
|
| 6 |
+
PASS think=True streaming: stream args={'city': 'Rome', 'unit': 'celsius'}
|
| 7 |
+
PASS think=True parallel: calls=['lima', 'oslo']
|
| 8 |
+
PASS think=False multi-arg: args={'city': 'Paris', 'unit': 'celsius'}
|
| 9 |
+
PASS think=False nested-object: args={'title': 'Design review', 'when': {'date': '2026-10-02', 'time': '14:00'}, 'attendees': ['ana@x.io', 'bo@x.io']}
|
| 10 |
+
PASS think=False enum: unit=fahrenheit
|
| 11 |
+
PASS think=False correct-decline: content='391'
|
| 12 |
+
PASS think=False multi-turn: final='Tokyo is currently **21°C** with **clear skies**.'
|
| 13 |
+
PASS think=False streaming: stream args={'city': 'Rome', 'unit': 'celsius'}
|
| 14 |
+
PASS think=False parallel: calls=['lima', 'oslo']
|
| 15 |
+
{"label": "n-tools-q106-c1", "passed": 13, "total": 14, "detail": {"multi-arg|think=True": true, "nested-object|think=True": false, "enum|think=True": true, "correct-decline|think=True": true, "multi-turn|think=True": true, "streaming|think=True": true, "parallel|think=True": true, "multi-arg|think=False": true, "nested-object|think=False": true, "enum|think=False": true, "correct-decline|think=False": true, "multi-turn|think=False": true, "streaming|think=False": true, "parallel|think=False": true}}
|
| 16 |
+
{"label": "n-vision-q106-c1-faon", "fa": "on", "mtp": false, "expected": "red,blue,circle,square", "answer": "The image shows two shapes: a red circle on the left and a blue square on the right.", "hits": ["red", "blue", "circle", "square"], "error": null, "server_died": false, "server_log_errors": [], "result": "PASS"}
|
| 17 |
+
probe no-kwargs correct-decline {"content": "391", "reasoning_len": 0, "tool_calls": [], "leaks": []}
|
| 18 |
+
probe no-kwargs single-word {"content": "ready", "reasoning_len": 0, "tool_calls": [], "leaks": []}
|
| 19 |
+
probe no-kwargs multi-arg {"content": "", "reasoning_len": 0, "tool_calls": ["get_weather"], "leaks": []}
|
| 20 |
+
probe enable_thinking=false correct-decline {"content": "391", "reasoning_len": 0, "tool_calls": [], "leaks": []}
|
| 21 |
+
probe enable_thinking=false single-word {"content": "ready", "reasoning_len": 0, "tool_calls": [], "leaks": []}
|
| 22 |
+
probe enable_thinking=false multi-arg {"content": "", "reasoning_len": 0, "tool_calls": ["get_weather"], "leaks": []}
|
| 23 |
+
probe reasoning_effort=high correct-decline {"content": "We need answer directly. 391.\n</think>\n\n391", "reasoning_len": 0, "tool_calls": [], "leaks": ["</think>"]}
|
| 24 |
+
probe reasoning_effort=high single-word {"content": "We need need output exactly ready.\n</think>\n\nready", "reasoning_len": 0, "tool_calls": [], "leaks": ["</think>"]}
|
| 25 |
+
probe reasoning_effort=high multi-arg {"content": "We need need tool. Current weather Paris celsius.\n</think>\n\n", "reasoning_len": 0, "tool_calls": ["get_weather"], "leaks": ["</think>"]}
|
| 26 |
+
probe reasoning_effort=medium correct-decline {"content": "\n\n</think>\n\n391", "reasoning_len": 0, "tool_calls": [], "leaks": ["</think>"]}
|
| 27 |
+
probe reasoning_effort=medium single-word {"content": "\n\n</think>\n\nready", "reasoning_len": 0, "tool_calls": [], "leaks": ["</think>"]}
|
| 28 |
+
probe reasoning_effort=medium multi-arg {"content": "\n\n</think>\n\n", "reasoning_len": 0, "tool_calls": ["get_weather"], "leaks": ["</think>"]}
|
| 29 |
+
probe reasoning_effort=none correct-decline {"content": "391", "reasoning_len": 0, "tool_calls": [], "leaks": []}
|
| 30 |
+
probe reasoning_effort=none single-word {"content": "ready", "reasoning_len": 0, "tool_calls": [], "leaks": []}
|
| 31 |
+
probe reasoning_effort=none multi-arg {"content": "", "reasoning_len": 0, "tool_calls": ["get_weather"], "leaks": []}
|
| 32 |
+
NEX_TOOLS_TPL_DONE
|
| 33 |
+
rc=0
|
recipe/logs/N6t_tools_tpl.log
ADDED
|
@@ -0,0 +1,22 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
PASS think=True multi-arg: args={'city': 'Paris', 'unit': 'celsius'}
|
| 2 |
+
PASS think=True nested-object: args={'title': 'Design review', 'when': {'date': '2026-10-02', 'time': '14:00'}, 'attendees': ['ana@x.io', 'bo@x.io']}
|
| 3 |
+
PASS think=True enum: unit=fahrenheit
|
| 4 |
+
PASS think=True correct-decline: content='391'
|
| 5 |
+
PASS think=True multi-turn: final='Tokyo is currently **21°C and clear**.'
|
| 6 |
+
PASS think=True streaming: stream args={'city': 'Rome', 'unit': 'celsius'}
|
| 7 |
+
PASS think=True parallel: calls=['lima', 'oslo']
|
| 8 |
+
PASS think=False multi-arg: args={'city': 'Paris', 'unit': 'celsius'}
|
| 9 |
+
PASS think=False nested-object: args={'title': 'Design review', 'when': {'date': '2026-10-02', 'time': '14:00'}, 'attendees': ['ana@x.io', 'bo@x.io']}
|
| 10 |
+
PASS think=False enum: unit=fahrenheit
|
| 11 |
+
PASS think=False correct-decline: content='391'
|
| 12 |
+
PASS think=False multi-turn: final='Tokyo is currently **21°C**, with **clear skies**.'
|
| 13 |
+
PASS think=False streaming: stream args={'city': 'Rome', 'unit': 'celsius'}
|
| 14 |
+
PASS think=False parallel: calls=['lima', 'oslo']
|
| 15 |
+
{"label": "n-tools-q106-tpl", "passed": 14, "total": 14, "detail": {"multi-arg|think=True": true, "nested-object|think=True": true, "enum|think=True": true, "correct-decline|think=True": true, "multi-turn|think=True": true, "streaming|think=True": true, "parallel|think=True": true, "multi-arg|think=False": true, "nested-object|think=False": true, "enum|think=False": true, "correct-decline|think=False": true, "multi-turn|think=False": true, "streaming|think=False": true, "parallel|think=False": true}}
|
| 16 |
+
probe no-kwargs correct-decline {"content": "391", "reasoning_len": 30, "tool_calls": [], "leaks": []}
|
| 17 |
+
probe no-kwargs multi-arg {"content": "", "reasoning_len": 50, "tool_calls": ["get_weather"], "leaks": []}
|
| 18 |
+
probe reasoning_effort=medium correct-decline {"content": "391", "reasoning_len": 0, "tool_calls": [], "leaks": []}
|
| 19 |
+
probe reasoning_effort=medium multi-arg {"content": "", "reasoning_len": 0, "tool_calls": ["get_weather"], "leaks": []}
|
| 20 |
+
probe reasoning_effort=none correct-decline {"content": "", "reasoning_len": 3, "tool_calls": [], "leaks": []}
|
| 21 |
+
probe reasoning_effort=none multi-arg {"content": "", "reasoning_len": 133, "tool_calls": [], "leaks": []}
|
| 22 |
+
NEX_TOOLS_TPL_DONE
|
recipe/logs/N6t_tools_tpl_medium.log
ADDED
|
@@ -0,0 +1,31 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
FAIL think=True multi-arg: args={'city': 'Paris', 'unit': 'celsius'}
|
| 2 |
+
FAIL think=True nested-object: exception KeyError('tool_calls')
|
| 3 |
+
FAIL think=True enum: unit=fahrenheit
|
| 4 |
+
FAIL think=True correct-decline: content='</think>\n\n391'
|
| 5 |
+
FAIL think=True multi-turn: final='</think>\n\nTokyo is currently **21°C** and **clear**.'
|
| 6 |
+
FAIL think=True streaming: stream args={'city': 'Rome', 'unit': 'celsius'}
|
| 7 |
+
FAIL think=True parallel: calls=['lima', 'oslo']
|
| 8 |
+
PASS think=False multi-arg: args={'city': 'Paris', 'unit': 'celsius'}
|
| 9 |
+
PASS think=False nested-object: args={'title': 'Design review', 'when': {'date': '2026-10-02', 'time': '14:00'}, 'attendees': ['ana@x.io', 'bo@x.io']}
|
| 10 |
+
PASS think=False enum: unit=fahrenheit
|
| 11 |
+
PASS think=False correct-decline: content='391'
|
| 12 |
+
PASS think=False multi-turn: final='Tokyo is currently **21°C** with **clear skies**.'
|
| 13 |
+
PASS think=False streaming: stream args={'city': 'Rome', 'unit': 'celsius'}
|
| 14 |
+
PASS think=False parallel: calls=['lima', 'oslo']
|
| 15 |
+
{"label": "n-tools-q106-tpl-medium", "passed": 7, "total": 14, "detail": {"multi-arg|think=True": false, "nested-object|think=True": false, "enum|think=True": false, "correct-decline|think=True": false, "multi-turn|think=True": false, "streaming|think=True": false, "parallel|think=True": false, "multi-arg|think=False": true, "nested-object|think=False": true, "enum|think=False": true, "correct-decline|think=False": true, "multi-turn|think=False": true, "streaming|think=False": true, "parallel|think=False": true}}
|
| 16 |
+
probe no-kwargs correct-decline {"content": "</think>\n\n391", "reasoning_len": 0, "tool_calls": [], "leaks": ["</think>"]}
|
| 17 |
+
probe no-kwargs single-word {"content": "</think>\n\nready", "reasoning_len": 0, "tool_calls": [], "leaks": ["</think>"]}
|
| 18 |
+
probe no-kwargs multi-arg {"content": "</think>\n\n", "reasoning_len": 0, "tool_calls": ["get_weather"], "leaks": ["</think>"]}
|
| 19 |
+
probe enable_thinking=false correct-decline {"content": "391", "reasoning_len": 0, "tool_calls": [], "leaks": []}
|
| 20 |
+
probe enable_thinking=false single-word {"content": "ready", "reasoning_len": 0, "tool_calls": [], "leaks": []}
|
| 21 |
+
probe enable_thinking=false multi-arg {"content": "", "reasoning_len": 0, "tool_calls": ["get_weather"], "leaks": []}
|
| 22 |
+
probe reasoning_effort=high correct-decline {"content": "We need answer directly. 391.\n</think>\n\n391", "reasoning_len": 0, "tool_calls": [], "leaks": ["</think>"]}
|
| 23 |
+
probe reasoning_effort=high single-word {"content": "We need need output exactly ready.\n</think>\n\nready", "reasoning_len": 0, "tool_calls": [], "leaks": ["</think>"]}
|
| 24 |
+
probe reasoning_effort=high multi-arg {"content": "We need need tool. Current weather Paris celsius.\n</think>\n\n", "reasoning_len": 0, "tool_calls": ["get_weather"], "leaks": ["</think>"]}
|
| 25 |
+
probe reasoning_effort=medium correct-decline {"content": "</think>\n\n391", "reasoning_len": 0, "tool_calls": [], "leaks": ["</think>"]}
|
| 26 |
+
probe reasoning_effort=medium single-word {"content": "</think>\n\nready", "reasoning_len": 0, "tool_calls": [], "leaks": ["</think>"]}
|
| 27 |
+
probe reasoning_effort=medium multi-arg {"content": "</think>\n\n", "reasoning_len": 0, "tool_calls": ["get_weather"], "leaks": ["</think>"]}
|
| 28 |
+
probe reasoning_effort=none correct-decline {"content": "391", "reasoning_len": 0, "tool_calls": [], "leaks": []}
|
| 29 |
+
probe reasoning_effort=none single-word {"content": "ready", "reasoning_len": 0, "tool_calls": [], "leaks": []}
|
| 30 |
+
probe reasoning_effort=none multi-arg {"content": "", "reasoning_len": 0, "tool_calls": ["get_weather"], "leaks": []}
|
| 31 |
+
NEX_TOOLS_TPL_DONE
|
recipe/logs/N7_sizing.log
ADDED
|
@@ -0,0 +1,5 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
[2026-09-17T00:50:08Z] waiting for the quiet-box lock
|
| 2 |
+
[2026-09-17T00:50:08Z] quiet-box lock held
|
| 3 |
+
{"label":"strix-lean","ctx":65536,"avail_before":122.34,"footprint_loaded_gib":21.11,"footprint_after_8k_gib":21.29}
|
| 4 |
+
{"label":"strix-lean","ctx":262144,"avail_before":122.15,"footprint_loaded_gib":24.36,"footprint_after_8k_gib":24.52}
|
| 5 |
+
NEX_SIZING_DONE
|
recipe/logs/N8_unice.log
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
[2026-09-17T00:50:56Z] Nex FAST test seats: create + smoke test (box still iced)
|
| 2 |
+
[2026-09-17T00:51:38Z] nex_seats exit=0 ([2026-09-17T00:51:38Z] NEX_SEATS_DONE fail=0)
|
| 3 |
+
[2026-09-17T00:51:38Z] UNICE_DEFERRED: OxCoder-9B build (King 2026-09-16 22:33Z: "when that is done do this one"); its final step un-ices
|
recipe/logs/N8a_seats.log
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
[2026-09-17T00:50:56Z] plan: [('max1-nex-fast', 'ROCm0', 'fa on', 262144, '31000M', True), ('max1-nex-fast-imat', 'ROCm0', 'fa on', 262144, '31000M', True)]
|
| 2 |
+
[2026-09-17T00:50:56Z] max1-nex-fast written (ROCm0, -fa on, ctx 262144, MemoryMax 31000M) -> smoke test
|
| 3 |
+
{"unit": "max1-nex-fast", "port": 8097, "load_s": 25, "time": "2026-09-17T00:51:21Z", "direct_reply": "ready", "direct_tg": 92.93680297397769, "gateway_model": "nex-n2.5-mini-fast@max1", "gateway_reply": "ready", "result": "PASS"}
|
| 4 |
+
[2026-09-17T00:51:22Z] max1-nex-fast stopped (enabled: disabled)
|
| 5 |
+
[2026-09-17T00:51:27Z] max1-nex-fast-imat written (ROCm0, -fa on, ctx 262144, MemoryMax 31000M) -> smoke test
|
| 6 |
+
{"unit": "max1-nex-fast-imat", "port": 8098, "load_s": 5, "time": "2026-09-17T00:51:32Z", "direct_reply": "ready", "direct_tg": 94.37078280564337, "gateway_model": "nex-n2.5-mini-fast-imatrix@max1", "gateway_reply": "ready", "result": "PASS"}
|
| 7 |
+
[2026-09-17T00:51:33Z] max1-nex-fast-imat stopped (enabled: disabled)
|
| 8 |
+
[2026-09-17T00:51:38Z] NEX_SEATS_DONE fail=0
|
recipe/logs/N8d_seats.log
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
[2026-09-17T01:20:44Z] plan: [('max1-nex-fast', 'ROCm0', 'fa on', 262144, '31000M', True), ('max1-nex-fast-imat', 'ROCm0', 'fa on', 262144, '31000M', True)]
|
| 2 |
+
[2026-09-17T01:20:45Z] max1-nex-fast written (ROCm0, -fa on, ctx 262144, MemoryMax 31000M) -> smoke test
|
| 3 |
+
{"unit": "max1-nex-fast", "port": 8097, "load_s": 25, "time": "2026-09-17T01:21:10Z", "direct_reply": "ready", "direct_tg": 41.41386950489719, "default_reply": "ready", "default_reasoning_len": 0, "default_leak": false, "thinking_reply": "", "thinking_reasoning_len": 5, "thinking_leak": false, "gateway_model": "nex-n2.5-mini-fast@max1", "gateway_reply": "ready", "result": "PASS"}
|
| 4 |
+
[2026-09-17T01:21:12Z] max1-nex-fast stopped (enabled: disabled)
|
| 5 |
+
[2026-09-17T01:21:17Z] max1-nex-fast-imat written (ROCm0, -fa on, ctx 262144, MemoryMax 31000M) -> smoke test
|
| 6 |
+
{"unit": "max1-nex-fast-imat", "port": 8098, "load_s": 25, "time": "2026-09-17T01:21:43Z", "direct_reply": "ready", "direct_tg": 41.54290343352097, "default_reply": "ready", "default_reasoning_len": 0, "default_leak": false, "thinking_reply": "", "thinking_reasoning_len": 5, "thinking_leak": false, "gateway_model": "nex-n2.5-mini-fast-imatrix@max1", "gateway_reply": "ready", "result": "PASS"}
|
| 7 |
+
[2026-09-17T01:21:44Z] max1-nex-fast-imat stopped (enabled: disabled)
|
| 8 |
+
[2026-09-17T01:21:49Z] NEX_SEATS_DONE fail=0
|
recipe/logs/Q1_q102.log
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
recipe/logs/Q1_q103.log
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
recipe/logs/Q_readback.log
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
PASS Nex-N2.5-mini-Q4_0_ROCMFP4_STRIX_LEAN.gguf arch=qwen35moe ftype=106 tensors=733 nextn=0 output.weight=Q6_K token_embd.weight=Q5_K
|
| 2 |
+
PASS Nex-N2.5-mini-Q4_0_ROCMFP4_COHERENT.gguf arch=qwen35moe ftype=102 tensors=733 nextn=0 output.weight=Q6_K token_embd.weight=Q6_K
|
| 3 |
+
PASS Nex-N2.5-mini-Q4_0_ROCMFP4_FAST.gguf arch=qwen35moe ftype=103 tensors=733 nextn=0 output.weight=Q6_K token_embd.weight=Q4_0_ROCMFP4_FAST
|
recipe/logs/Q_sizes.log
ADDED
|
@@ -0,0 +1,6 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Q1_q106 quant size = 17865.52 MiB (4.32 BPW)
|
| 2 |
+
Q1_q102 quant size = 18916.30 MiB (4.58 BPW)
|
| 3 |
+
Q1_q103 quant size = 17774.11 MiB (4.30 BPW)
|
| 4 |
+
N3_q106i quant size = 17865.52 MiB (4.32 BPW)
|
| 5 |
+
N3_q102i quant size = 18916.30 MiB (4.58 BPW)
|
| 6 |
+
N3_q103i quant size = 17774.11 MiB (4.30 BPW)
|
recipe/logs/b_n-c3-q106.log
ADDED
|
@@ -0,0 +1,284 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
0.00.041.768 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
|
| 2 |
+
0.00.041.771 I device_info:
|
| 3 |
+
0.00.041.819 I - ROCm0 : AMD Radeon Graphics (131072 MiB, 125259 MiB free)
|
| 4 |
+
0.00.041.885 I - Vulkan0 : AMD Radeon Graphics (RADV GFX1151) (132096 MiB, 131926 MiB free)
|
| 5 |
+
0.00.041.888 I - CPU : AMD RYZEN AI MAX+ 395 w/ Radeon 8060S (127438 MiB, 127438 MiB free)
|
| 6 |
+
0.00.041.930 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
|
| 7 |
+
0.00.041.949 I srv init: running without SSL
|
| 8 |
+
0.00.041.972 I srv init: using 31 threads for HTTP server
|
| 9 |
+
0.00.041.973 I srv init: the WebUI is disabled
|
| 10 |
+
0.00.042.029 I srv start: binding port with default address family
|
| 11 |
+
0.00.043.171 I srv main: loading model
|
| 12 |
+
0.00.043.173 I srv load_model: loading model '/mnt/models/nex-n2.5-mini/out/Nex-N2.5-mini-Q4_0_ROCMFP4_STRIX_LEAN.gguf'
|
| 13 |
+
0.00.076.857 W llama_model_loader: direct I/O is enabled, disabling mmap
|
| 14 |
+
0.20.443.494 W llama_context: n_ctx_seq (65536) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
|
| 15 |
+
0.20.605.862 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|
| 16 |
+
0.20.804.349 I srv load_model: initializing slots, n_slots = 1
|
| 17 |
+
0.20.970.363 W srv load_model: speculative decoding will use checkpoints
|
| 18 |
+
0.20.970.368 W common_speculative_init: no implementations specified for speculative decoding
|
| 19 |
+
0.20.970.369 I slot load_model: id 0 | task -1 | new slot, n_ctx = 65536
|
| 20 |
+
0.20.970.402 I srv load_model: prompt cache RAM enabled: limit_mib=8192
|
| 21 |
+
0.20.970.414 I srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
|
| 22 |
+
0.20.970.429 I srv init: idle slots will be saved to prompt cache upon starting a new task
|
| 23 |
+
0.20.979.135 I init: chat template, example_format: '<|im_start|>system
|
| 24 |
+
You are a helpful assistant<|im_end|>
|
| 25 |
+
<|im_start|>user
|
| 26 |
+
Hello<|im_end|>
|
| 27 |
+
<|im_start|>assistant
|
| 28 |
+
<think>
|
| 29 |
+
|
| 30 |
+
</think>
|
| 31 |
+
|
| 32 |
+
Hi there<|im_end|>
|
| 33 |
+
<|im_start|>user
|
| 34 |
+
How are you?<|im_end|>
|
| 35 |
+
<|im_start|>assistant
|
| 36 |
+
<think>'
|
| 37 |
+
0.20.985.480 I srv init: init: chat template, thinking = 1
|
| 38 |
+
0.20.985.501 I srv main: model loaded
|
| 39 |
+
0.20.985.504 I srv main: server is listening on http://127.0.0.1:18600
|
| 40 |
+
0.20.985.506 I srv update_slots: all slots are idle
|
| 41 |
+
0.23.012.828 I srv params_from_: Chat format: peg-native
|
| 42 |
+
0.23.012.941 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1
|
| 43 |
+
0.23.012.943 I srv get_availabl: updating prompt cache
|
| 44 |
+
0.23.012.949 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000
|
| 45 |
+
0.23.012.953 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 65536 tokens, 8589934592 est)
|
| 46 |
+
0.23.012.954 I srv get_availabl: prompt cache update took 0.01 ms
|
| 47 |
+
0.23.013.012 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0
|
| 48 |
+
0.26.234.628 I slot print_timing: id 0 | task 0 | prompt processing, n_tokens = 4096, progress = 0.58, t = 3.22 s / 1271.42 tokens per second
|
| 49 |
+
0.27.896.790 I slot print_timing: id 0 | task 0 | prompt processing, n_tokens = 6011, progress = 0.85, t = 4.88 s / 1230.81 tokens per second
|
| 50 |
+
0.27.944.722 I slot create_check: id 0 | task 0 | created context checkpoint 1 of 32 (pos_min = 6010, pos_max = 6010, n_tokens = 6011, size = 62.813 MiB)
|
| 51 |
+
0.28.834.529 I slot print_timing: id 0 | task 0 | prompt processing, n_tokens = 7035, progress = 1.00, t = 5.82 s / 1208.45 tokens per second
|
| 52 |
+
0.28.892.962 I slot create_check: id 0 | task 0 | created context checkpoint 2 of 32 (pos_min = 7034, pos_max = 7034, n_tokens = 7035, size = 62.813 MiB)
|
| 53 |
+
0.29.188.213 I slot print_timing: id 0 | task 0 |
|
| 54 |
+
prompt eval time = 5911.23 ms / 7039 tokens ( 0.84 ms per token, 1190.78 tokens per second)
|
| 55 |
+
eval time = 263.95 ms / 16 tokens ( 16.50 ms per token, 60.62 tokens per second)
|
| 56 |
+
total time = 6175.18 ms / 7055 tokens
|
| 57 |
+
0.29.188.522 I slot release: id 0 | task 0 | stop processing: n_tokens = 7054, truncated = 0
|
| 58 |
+
0.29.188.528 I srv update_slots: all slots are idle
|
| 59 |
+
0.29.206.128 I srv params_from_: Chat format: peg-native
|
| 60 |
+
0.29.206.239 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.996 (> 0.100 thold), f_keep = 0.994
|
| 61 |
+
0.29.206.302 I slot launch_slot_: id 0 | task 21 | processing task, is_child = 0
|
| 62 |
+
0.29.206.311 W slot update_slots: id 0 | task 21 | n_past = 7014, slot.prompt.tokens.size() = 7054, seq_id = 0, pos_min = 7053, n_swa = 0
|
| 63 |
+
0.29.206.312 I slot update_slots: id 0 | task 21 | Checking checkpoint with [7034, 7034] against 7014...
|
| 64 |
+
0.29.206.312 I slot update_slots: id 0 | task 21 | Checking checkpoint with [6010, 6010] against 7014...
|
| 65 |
+
0.29.210.086 W slot update_slots: id 0 | task 21 | restored context checkpoint (pos_min = 6010, pos_max = 6010, n_tokens = 6011, n_past = 6011, size = 62.813 MiB)
|
| 66 |
+
0.29.210.090 W slot update_slots: id 0 | task 21 | erased invalidated context checkpoint (pos_min = 7034, pos_max = 7034, n_tokens = 7035, n_swa = 0, pos_next = 6011, size = 62.813 MiB)
|
| 67 |
+
0.30.175.583 I slot create_check: id 0 | task 21 | created context checkpoint 2 of 32 (pos_min = 7034, pos_max = 7034, n_tokens = 7035, size = 62.813 MiB)
|
| 68 |
+
0.31.806.656 I slot print_timing: id 0 | task 21 | n_decoded = 100, tg = 62.48 t/s
|
| 69 |
+
0.33.255.966 I slot print_timing: id 0 | task 21 |
|
| 70 |
+
prompt eval time = 999.89 ms / 1028 tokens ( 0.97 ms per token, 1028.11 tokens per second)
|
| 71 |
+
eval time = 3049.75 ms / 192 tokens ( 15.88 ms per token, 62.96 tokens per second)
|
| 72 |
+
total time = 4049.65 ms / 1220 tokens
|
| 73 |
+
0.33.256.219 I slot release: id 0 | task 21 | stop processing: n_tokens = 7230, truncated = 0
|
| 74 |
+
0.33.256.240 I srv update_slots: all slots are idle
|
| 75 |
+
0.33.273.500 I srv params_from_: Chat format: peg-native
|
| 76 |
+
0.33.273.615 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.974
|
| 77 |
+
0.33.273.668 I slot launch_slot_: id 0 | task 215 | processing task, is_child = 0
|
| 78 |
+
0.33.273.676 W slot update_slots: id 0 | task 215 | erased invalidated context checkpoint (pos_min = 6010, pos_max = 6010, n_tokens = 6011, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 79 |
+
0.33.275.258 W slot update_slots: id 0 | task 215 | erased invalidated context checkpoint (pos_min = 7034, pos_max = 7034, n_tokens = 7035, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 80 |
+
0.36.625.268 I slot print_timing: id 0 | task 215 | prompt processing, n_tokens = 4096, progress = 0.58, t = 3.35 s / 1222.11 tokens per second
|
| 81 |
+
0.38.375.032 I slot print_timing: id 0 | task 215 | prompt processing, n_tokens = 6011, progress = 0.85, t = 5.10 s / 1178.32 tokens per second
|
| 82 |
+
0.38.426.513 I slot create_check: id 0 | task 215 | created context checkpoint 1 of 32 (pos_min = 6010, pos_max = 6010, n_tokens = 6011, size = 62.813 MiB)
|
| 83 |
+
0.39.343.585 I slot print_timing: id 0 | task 215 | prompt processing, n_tokens = 7035, progress = 1.00, t = 6.07 s / 1159.00 tokens per second
|
| 84 |
+
0.39.401.811 I slot create_check: id 0 | task 215 | created context checkpoint 2 of 32 (pos_min = 7034, pos_max = 7034, n_tokens = 7035, size = 62.813 MiB)
|
| 85 |
+
0.41.031.975 I slot print_timing: id 0 | task 215 | n_decoded = 100, tg = 62.51 t/s
|
| 86 |
+
0.42.481.538 I slot print_timing: id 0 | task 215 |
|
| 87 |
+
prompt eval time = 6158.66 ms / 7039 tokens ( 0.87 ms per token, 1142.94 tokens per second)
|
| 88 |
+
eval time = 3049.19 ms / 192 tokens ( 15.88 ms per token, 62.97 tokens per second)
|
| 89 |
+
total time = 9207.85 ms / 7231 tokens
|
| 90 |
+
0.42.481.799 I slot release: id 0 | task 215 | stop processing: n_tokens = 7230, truncated = 0
|
| 91 |
+
0.42.481.812 I srv update_slots: all slots are idle
|
| 92 |
+
0.42.499.185 I srv params_from_: Chat format: peg-native
|
| 93 |
+
0.42.499.292 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.996 (> 0.100 thold), f_keep = 0.970
|
| 94 |
+
0.42.499.320 I slot launch_slot_: id 0 | task 412 | processing task, is_child = 0
|
| 95 |
+
0.42.499.327 W slot update_slots: id 0 | task 412 | n_past = 7014, slot.prompt.tokens.size() = 7230, seq_id = 0, pos_min = 7229, n_swa = 0
|
| 96 |
+
0.42.499.328 I slot update_slots: id 0 | task 412 | Checking checkpoint with [7034, 7034] against 7014...
|
| 97 |
+
0.42.499.328 I slot update_slots: id 0 | task 412 | Checking checkpoint with [6010, 6010] against 7014...
|
| 98 |
+
0.42.502.995 W slot update_slots: id 0 | task 412 | restored context checkpoint (pos_min = 6010, pos_max = 6010, n_tokens = 6011, n_past = 6011, size = 62.813 MiB)
|
| 99 |
+
0.42.502.999 W slot update_slots: id 0 | task 412 | erased invalidated context checkpoint (pos_min = 7034, pos_max = 7034, n_tokens = 7035, n_swa = 0, pos_next = 6011, size = 62.813 MiB)
|
| 100 |
+
0.43.471.709 I slot create_check: id 0 | task 412 | created context checkpoint 2 of 32 (pos_min = 7034, pos_max = 7034, n_tokens = 7035, size = 62.813 MiB)
|
| 101 |
+
0.43.764.311 I slot print_timing: id 0 | task 412 |
|
| 102 |
+
prompt eval time = 1002.90 ms / 1028 tokens ( 0.98 ms per token, 1025.03 tokens per second)
|
| 103 |
+
eval time = 262.06 ms / 16 tokens ( 16.38 ms per token, 61.05 tokens per second)
|
| 104 |
+
total time = 1264.96 ms / 1044 tokens
|
| 105 |
+
0.43.764.770 I slot release: id 0 | task 412 | stop processing: n_tokens = 7054, truncated = 0
|
| 106 |
+
0.43.764.786 I srv update_slots: all slots are idle
|
| 107 |
+
0.43.792.285 I srv params_from_: Chat format: peg-native
|
| 108 |
+
0.43.792.395 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.996 (> 0.100 thold), f_keep = 0.994
|
| 109 |
+
0.43.792.452 I slot launch_slot_: id 0 | task 430 | processing task, is_child = 0
|
| 110 |
+
0.43.792.461 W slot update_slots: id 0 | task 430 | n_past = 7014, slot.prompt.tokens.size() = 7054, seq_id = 0, pos_min = 7053, n_swa = 0
|
| 111 |
+
0.43.792.462 I slot update_slots: id 0 | task 430 | Checking checkpoint with [7034, 7034] against 7014...
|
| 112 |
+
0.43.792.462 I slot update_slots: id 0 | task 430 | Checking checkpoint with [6010, 6010] against 7014...
|
| 113 |
+
0.43.796.220 W slot update_slots: id 0 | task 430 | restored context checkpoint (pos_min = 6010, pos_max = 6010, n_tokens = 6011, n_past = 6011, size = 62.813 MiB)
|
| 114 |
+
0.43.796.224 W slot update_slots: id 0 | task 430 | erased invalidated context checkpoint (pos_min = 7034, pos_max = 7034, n_tokens = 7035, n_swa = 0, pos_next = 6011, size = 62.813 MiB)
|
| 115 |
+
0.44.764.733 I slot create_check: id 0 | task 430 | created context checkpoint 2 of 32 (pos_min = 7034, pos_max = 7034, n_tokens = 7035, size = 62.813 MiB)
|
| 116 |
+
0.46.401.785 I slot print_timing: id 0 | task 430 | n_decoded = 100, tg = 62.25 t/s
|
| 117 |
+
0.47.852.134 I slot print_timing: id 0 | task 430 |
|
| 118 |
+
prompt eval time = 1002.97 ms / 1028 tokens ( 0.98 ms per token, 1024.96 tokens per second)
|
| 119 |
+
eval time = 3056.70 ms / 192 tokens ( 15.92 ms per token, 62.81 tokens per second)
|
| 120 |
+
total time = 4059.66 ms / 1220 tokens
|
| 121 |
+
0.47.852.392 I slot release: id 0 | task 430 | stop processing: n_tokens = 7230, truncated = 0
|
| 122 |
+
0.47.852.409 I srv update_slots: all slots are idle
|
| 123 |
+
0.47.869.577 I srv params_from_: Chat format: peg-native
|
| 124 |
+
0.47.869.673 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.974
|
| 125 |
+
0.47.869.703 I slot launch_slot_: id 0 | task 624 | processing task, is_child = 0
|
| 126 |
+
0.47.869.708 W slot update_slots: id 0 | task 624 | erased invalidated context checkpoint (pos_min = 6010, pos_max = 6010, n_tokens = 6011, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 127 |
+
0.47.871.268 W slot update_slots: id 0 | task 624 | erased invalidated context checkpoint (pos_min = 7034, pos_max = 7034, n_tokens = 7035, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 128 |
+
0.51.222.915 I slot print_timing: id 0 | task 624 | prompt processing, n_tokens = 4096, progress = 0.58, t = 3.35 s / 1221.52 tokens per second
|
| 129 |
+
0.52.975.285 I slot print_timing: id 0 | task 624 | prompt processing, n_tokens = 6011, progress = 0.85, t = 5.11 s / 1177.34 tokens per second
|
| 130 |
+
0.53.025.961 I slot create_check: id 0 | task 624 | created context checkpoint 1 of 32 (pos_min = 6010, pos_max = 6010, n_tokens = 6011, size = 62.813 MiB)
|
| 131 |
+
0.53.946.619 I slot print_timing: id 0 | task 624 | prompt processing, n_tokens = 7035, progress = 1.00, t = 6.08 s / 1157.66 tokens per second
|
| 132 |
+
0.54.004.891 I slot create_check: id 0 | task 624 | created context checkpoint 2 of 32 (pos_min = 7034, pos_max = 7034, n_tokens = 7035, size = 62.813 MiB)
|
| 133 |
+
0.55.641.849 I slot print_timing: id 0 | task 624 | n_decoded = 100, tg = 62.25 t/s
|
| 134 |
+
0.57.095.298 I slot print_timing: id 0 | task 624 |
|
| 135 |
+
prompt eval time = 6165.72 ms / 7039 tokens ( 0.88 ms per token, 1141.64 tokens per second)
|
| 136 |
+
eval time = 3059.86 ms / 192 tokens ( 15.94 ms per token, 62.75 tokens per second)
|
| 137 |
+
total time = 9225.58 ms / 7231 tokens
|
| 138 |
+
0.57.095.560 I slot release: id 0 | task 624 | stop processing: n_tokens = 7230, truncated = 0
|
| 139 |
+
0.57.095.573 I srv update_slots: all slots are idle
|
| 140 |
+
0.57.112.963 I srv params_from_: Chat format: peg-native
|
| 141 |
+
0.57.113.061 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.996 (> 0.100 thold), f_keep = 0.970
|
| 142 |
+
0.57.113.099 I slot launch_slot_: id 0 | task 821 | processing task, is_child = 0
|
| 143 |
+
0.57.113.105 W slot update_slots: id 0 | task 821 | n_past = 7014, slot.prompt.tokens.size() = 7230, seq_id = 0, pos_min = 7229, n_swa = 0
|
| 144 |
+
0.57.113.106 I slot update_slots: id 0 | task 821 | Checking checkpoint with [7034, 7034] against 7014...
|
| 145 |
+
0.57.113.106 I slot update_slots: id 0 | task 821 | Checking checkpoint with [6010, 6010] against 7014...
|
| 146 |
+
0.57.116.771 W slot update_slots: id 0 | task 821 | restored context checkpoint (pos_min = 6010, pos_max = 6010, n_tokens = 6011, n_past = 6011, size = 62.813 MiB)
|
| 147 |
+
0.57.116.773 W slot update_slots: id 0 | task 821 | erased invalidated context checkpoint (pos_min = 7034, pos_max = 7034, n_tokens = 7035, n_swa = 0, pos_next = 6011, size = 62.813 MiB)
|
| 148 |
+
0.58.086.071 I slot create_check: id 0 | task 821 | created context checkpoint 2 of 32 (pos_min = 7034, pos_max = 7034, n_tokens = 7035, size = 62.813 MiB)
|
| 149 |
+
0.58.378.533 I slot print_timing: id 0 | task 821 |
|
| 150 |
+
prompt eval time = 1003.26 ms / 1028 tokens ( 0.98 ms per token, 1024.66 tokens per second)
|
| 151 |
+
eval time = 262.14 ms / 16 tokens ( 16.38 ms per token, 61.04 tokens per second)
|
| 152 |
+
total time = 1265.40 ms / 1044 tokens
|
| 153 |
+
0.58.378.987 I slot release: id 0 | task 821 | stop processing: n_tokens = 7054, truncated = 0
|
| 154 |
+
0.58.379.007 I srv update_slots: all slots are idle
|
| 155 |
+
0.58.404.936 I srv params_from_: Chat format: peg-native
|
| 156 |
+
0.58.405.048 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.996 (> 0.100 thold), f_keep = 0.994
|
| 157 |
+
0.58.405.100 I slot launch_slot_: id 0 | task 839 | processing task, is_child = 0
|
| 158 |
+
0.58.405.109 W slot update_slots: id 0 | task 839 | n_past = 7014, slot.prompt.tokens.size() = 7054, seq_id = 0, pos_min = 7053, n_swa = 0
|
| 159 |
+
0.58.405.109 I slot update_slots: id 0 | task 839 | Checking checkpoint with [7034, 7034] against 7014...
|
| 160 |
+
0.58.405.110 I slot update_slots: id 0 | task 839 | Checking checkpoint with [6010, 6010] against 7014...
|
| 161 |
+
0.58.408.850 W slot update_slots: id 0 | task 839 | restored context checkpoint (pos_min = 6010, pos_max = 6010, n_tokens = 6011, n_past = 6011, size = 62.813 MiB)
|
| 162 |
+
0.58.408.855 W slot update_slots: id 0 | task 839 | erased invalidated context checkpoint (pos_min = 7034, pos_max = 7034, n_tokens = 7035, n_swa = 0, pos_next = 6011, size = 62.813 MiB)
|
| 163 |
+
0.59.378.102 I slot create_check: id 0 | task 839 | created context checkpoint 2 of 32 (pos_min = 7034, pos_max = 7034, n_tokens = 7035, size = 62.813 MiB)
|
| 164 |
+
1.01.014.891 I slot print_timing: id 0 | task 839 | n_decoded = 100, tg = 62.26 t/s
|
| 165 |
+
1.02.468.788 I slot print_timing: id 0 | task 839 |
|
| 166 |
+
prompt eval time = 1003.48 ms / 1028 tokens ( 0.98 ms per token, 1024.43 tokens per second)
|
| 167 |
+
eval time = 3060.19 ms / 192 tokens ( 15.94 ms per token, 62.74 tokens per second)
|
| 168 |
+
total time = 4063.67 ms / 1220 tokens
|
| 169 |
+
1.02.469.046 I slot release: id 0 | task 839 | stop processing: n_tokens = 7230, truncated = 0
|
| 170 |
+
1.02.469.057 I srv update_slots: all slots are idle
|
| 171 |
+
1.02.486.085 I srv params_from_: Chat format: peg-native
|
| 172 |
+
1.02.486.197 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.974
|
| 173 |
+
1.02.486.250 I slot launch_slot_: id 0 | task 1033 | processing task, is_child = 0
|
| 174 |
+
1.02.486.266 W slot update_slots: id 0 | task 1033 | erased invalidated context checkpoint (pos_min = 6010, pos_max = 6010, n_tokens = 6011, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 175 |
+
1.02.488.866 W slot update_slots: id 0 | task 1033 | erased invalidated context checkpoint (pos_min = 7034, pos_max = 7034, n_tokens = 7035, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 176 |
+
1.05.842.274 I slot print_timing: id 0 | task 1033 | prompt processing, n_tokens = 4096, progress = 0.58, t = 3.36 s / 1220.51 tokens per second
|
| 177 |
+
1.07.594.394 I slot print_timing: id 0 | task 1033 | prompt processing, n_tokens = 6011, progress = 0.85, t = 5.11 s / 1176.75 tokens per second
|
| 178 |
+
1.07.645.015 I slot create_check: id 0 | task 1033 | created context checkpoint 1 of 32 (pos_min = 6010, pos_max = 6010, n_tokens = 6011, size = 62.813 MiB)
|
| 179 |
+
1.08.566.063 I slot print_timing: id 0 | task 1033 | prompt processing, n_tokens = 7035, progress = 1.00, t = 6.08 s / 1157.11 tokens per second
|
| 180 |
+
1.08.624.878 I slot create_check: id 0 | task 1033 | created context checkpoint 2 of 32 (pos_min = 7034, pos_max = 7034, n_tokens = 7035, size = 62.813 MiB)
|
| 181 |
+
1.10.262.102 I slot print_timing: id 0 | task 1033 | n_decoded = 100, tg = 62.24 t/s
|
| 182 |
+
1.11.717.464 I slot print_timing: id 0 | task 1033 |
|
| 183 |
+
prompt eval time = 6169.04 ms / 7039 tokens ( 0.88 ms per token, 1141.02 tokens per second)
|
| 184 |
+
eval time = 3062.15 ms / 192 tokens ( 15.95 ms per token, 62.70 tokens per second)
|
| 185 |
+
total time = 9231.19 ms / 7231 tokens
|
| 186 |
+
1.11.717.716 I slot release: id 0 | task 1033 | stop processing: n_tokens = 7230, truncated = 0
|
| 187 |
+
1.11.717.731 I srv update_slots: all slots are idle
|
| 188 |
+
1.11.734.739 I srv params_from_: Chat format: peg-native
|
| 189 |
+
1.11.734.856 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.996 (> 0.100 thold), f_keep = 0.970
|
| 190 |
+
1.11.734.912 I slot launch_slot_: id 0 | task 1230 | processing task, is_child = 0
|
| 191 |
+
1.11.734.923 W slot update_slots: id 0 | task 1230 | n_past = 7014, slot.prompt.tokens.size() = 7230, seq_id = 0, pos_min = 7229, n_swa = 0
|
| 192 |
+
1.11.734.923 I slot update_slots: id 0 | task 1230 | Checking checkpoint with [7034, 7034] against 7014...
|
| 193 |
+
1.11.734.924 I slot update_slots: id 0 | task 1230 | Checking checkpoint with [6010, 6010] against 7014...
|
| 194 |
+
1.11.738.681 W slot update_slots: id 0 | task 1230 | restored context checkpoint (pos_min = 6010, pos_max = 6010, n_tokens = 6011, n_past = 6011, size = 62.813 MiB)
|
| 195 |
+
1.11.738.683 W slot update_slots: id 0 | task 1230 | erased invalidated context checkpoint (pos_min = 7034, pos_max = 7034, n_tokens = 7035, n_swa = 0, pos_next = 6011, size = 62.813 MiB)
|
| 196 |
+
1.12.708.145 I slot create_check: id 0 | task 1230 | created context checkpoint 2 of 32 (pos_min = 7034, pos_max = 7034, n_tokens = 7035, size = 62.813 MiB)
|
| 197 |
+
1.13.000.920 I slot print_timing: id 0 | task 1230 |
|
| 198 |
+
prompt eval time = 1003.98 ms / 1028 tokens ( 0.98 ms per token, 1023.92 tokens per second)
|
| 199 |
+
eval time = 262.00 ms / 16 tokens ( 16.37 ms per token, 61.07 tokens per second)
|
| 200 |
+
total time = 1265.98 ms / 1044 tokens
|
| 201 |
+
1.13.001.384 I slot release: id 0 | task 1230 | stop processing: n_tokens = 7054, truncated = 0
|
| 202 |
+
1.13.001.412 I srv update_slots: all slots are idle
|
| 203 |
+
1.13.029.139 I srv params_from_: Chat format: peg-native
|
| 204 |
+
1.13.029.237 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.996 (> 0.100 thold), f_keep = 0.994
|
| 205 |
+
1.13.029.274 I slot launch_slot_: id 0 | task 1248 | processing task, is_child = 0
|
| 206 |
+
1.13.029.280 W slot update_slots: id 0 | task 1248 | n_past = 7014, slot.prompt.tokens.size() = 7054, seq_id = 0, pos_min = 7053, n_swa = 0
|
| 207 |
+
1.13.029.280 I slot update_slots: id 0 | task 1248 | Checking checkpoint with [7034, 7034] against 7014...
|
| 208 |
+
1.13.029.281 I slot update_slots: id 0 | task 1248 | Checking checkpoint with [6010, 6010] against 7014...
|
| 209 |
+
1.13.032.947 W slot update_slots: id 0 | task 1248 | restored context checkpoint (pos_min = 6010, pos_max = 6010, n_tokens = 6011, n_past = 6011, size = 62.813 MiB)
|
| 210 |
+
1.13.032.950 W slot update_slots: id 0 | task 1248 | erased invalidated context checkpoint (pos_min = 7034, pos_max = 7034, n_tokens = 7035, n_swa = 0, pos_next = 6011, size = 62.813 MiB)
|
| 211 |
+
1.14.002.308 I slot create_check: id 0 | task 1248 | created context checkpoint 2 of 32 (pos_min = 7034, pos_max = 7034, n_tokens = 7035, size = 62.813 MiB)
|
| 212 |
+
1.15.640.869 I slot print_timing: id 0 | task 1248 | n_decoded = 100, tg = 62.18 t/s
|
| 213 |
+
1.17.096.125 I slot print_timing: id 0 | task 1248 |
|
| 214 |
+
prompt eval time = 1003.44 ms / 1028 tokens ( 0.98 ms per token, 1024.48 tokens per second)
|
| 215 |
+
eval time = 3063.39 ms / 192 tokens ( 15.96 ms per token, 62.68 tokens per second)
|
| 216 |
+
total time = 4066.83 ms / 1220 tokens
|
| 217 |
+
1.17.096.385 I slot release: id 0 | task 1248 | stop processing: n_tokens = 7230, truncated = 0
|
| 218 |
+
1.17.096.398 I srv update_slots: all slots are idle
|
| 219 |
+
1.17.113.652 I srv params_from_: Chat format: peg-native
|
| 220 |
+
1.17.113.773 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.974
|
| 221 |
+
1.17.113.827 I slot launch_slot_: id 0 | task 1442 | processing task, is_child = 0
|
| 222 |
+
1.17.113.835 W slot update_slots: id 0 | task 1442 | erased invalidated context checkpoint (pos_min = 6010, pos_max = 6010, n_tokens = 6011, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 223 |
+
1.17.115.551 W slot update_slots: id 0 | task 1442 | erased invalidated context checkpoint (pos_min = 7034, pos_max = 7034, n_tokens = 7035, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 224 |
+
1.20.471.772 I slot print_timing: id 0 | task 1442 | prompt processing, n_tokens = 4096, progress = 0.58, t = 3.36 s / 1219.80 tokens per second
|
| 225 |
+
1.22.222.821 I slot print_timing: id 0 | task 1442 | prompt processing, n_tokens = 6011, progress = 0.85, t = 5.11 s / 1176.56 tokens per second
|
| 226 |
+
1.22.273.999 I slot create_check: id 0 | task 1442 | created context checkpoint 1 of 32 (pos_min = 6010, pos_max = 6010, n_tokens = 6011, size = 62.813 MiB)
|
| 227 |
+
1.23.196.105 I slot print_timing: id 0 | task 1442 | prompt processing, n_tokens = 7035, progress = 1.00, t = 6.08 s / 1156.64 tokens per second
|
| 228 |
+
1.23.254.243 I slot create_check: id 0 | task 1442 | created context checkpoint 2 of 32 (pos_min = 7034, pos_max = 7034, n_tokens = 7035, size = 62.813 MiB)
|
| 229 |
+
1.24.893.156 I slot print_timing: id 0 | task 1442 | n_decoded = 100, tg = 62.17 t/s
|
| 230 |
+
1.26.350.431 I slot print_timing: id 0 | task 1442 |
|
| 231 |
+
prompt eval time = 6170.80 ms / 7039 tokens ( 0.88 ms per token, 1140.70 tokens per second)
|
| 232 |
+
eval time = 3065.79 ms / 192 tokens ( 15.97 ms per token, 62.63 tokens per second)
|
| 233 |
+
total time = 9236.59 ms / 7231 tokens
|
| 234 |
+
1.26.350.694 I slot release: id 0 | task 1442 | stop processing: n_tokens = 7230, truncated = 0
|
| 235 |
+
1.26.350.710 I srv update_slots: all slots are idle
|
| 236 |
+
1.26.367.931 I srv params_from_: Chat format: peg-native
|
| 237 |
+
1.26.368.047 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.996 (> 0.100 thold), f_keep = 0.970
|
| 238 |
+
1.26.368.104 I slot launch_slot_: id 0 | task 1639 | processing task, is_child = 0
|
| 239 |
+
1.26.368.113 W slot update_slots: id 0 | task 1639 | n_past = 7014, slot.prompt.tokens.size() = 7230, seq_id = 0, pos_min = 7229, n_swa = 0
|
| 240 |
+
1.26.368.113 I slot update_slots: id 0 | task 1639 | Checking checkpoint with [7034, 7034] against 7014...
|
| 241 |
+
1.26.368.114 I slot update_slots: id 0 | task 1639 | Checking checkpoint with [6010, 6010] against 7014...
|
| 242 |
+
1.26.371.873 W slot update_slots: id 0 | task 1639 | restored context checkpoint (pos_min = 6010, pos_max = 6010, n_tokens = 6011, n_past = 6011, size = 62.813 MiB)
|
| 243 |
+
1.26.371.876 W slot update_slots: id 0 | task 1639 | erased invalidated context checkpoint (pos_min = 7034, pos_max = 7034, n_tokens = 7035, n_swa = 0, pos_next = 6011, size = 62.813 MiB)
|
| 244 |
+
1.27.342.456 I slot create_check: id 0 | task 1639 | created context checkpoint 2 of 32 (pos_min = 7034, pos_max = 7034, n_tokens = 7035, size = 62.813 MiB)
|
| 245 |
+
1.27.635.050 I slot print_timing: id 0 | task 1639 |
|
| 246 |
+
prompt eval time = 1004.71 ms / 1028 tokens ( 0.98 ms per token, 1023.18 tokens per second)
|
| 247 |
+
eval time = 262.20 ms / 16 tokens ( 16.39 ms per token, 61.02 tokens per second)
|
| 248 |
+
total time = 1266.92 ms / 1044 tokens
|
| 249 |
+
1.27.635.512 I slot release: id 0 | task 1639 | stop processing: n_tokens = 7054, truncated = 0
|
| 250 |
+
1.27.635.532 I srv update_slots: all slots are idle
|
| 251 |
+
1.27.666.084 I srv params_from_: Chat format: peg-native
|
| 252 |
+
1.27.666.197 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.996 (> 0.100 thold), f_keep = 0.994
|
| 253 |
+
1.27.666.254 I slot launch_slot_: id 0 | task 1657 | processing task, is_child = 0
|
| 254 |
+
1.27.666.272 W slot update_slots: id 0 | task 1657 | n_past = 7014, slot.prompt.tokens.size() = 7054, seq_id = 0, pos_min = 7053, n_swa = 0
|
| 255 |
+
1.27.666.273 I slot update_slots: id 0 | task 1657 | Checking checkpoint with [7034, 7034] against 7014...
|
| 256 |
+
1.27.666.273 I slot update_slots: id 0 | task 1657 | Checking checkpoint with [6010, 6010] against 7014...
|
| 257 |
+
1.27.670.025 W slot update_slots: id 0 | task 1657 | restored context checkpoint (pos_min = 6010, pos_max = 6010, n_tokens = 6011, n_past = 6011, size = 62.813 MiB)
|
| 258 |
+
1.27.670.030 W slot update_slots: id 0 | task 1657 | erased invalidated context checkpoint (pos_min = 7034, pos_max = 7034, n_tokens = 7035, n_swa = 0, pos_next = 6011, size = 62.813 MiB)
|
| 259 |
+
1.28.641.307 I slot create_check: id 0 | task 1657 | created context checkpoint 2 of 32 (pos_min = 7034, pos_max = 7034, n_tokens = 7035, size = 62.813 MiB)
|
| 260 |
+
1.30.279.245 I slot print_timing: id 0 | task 1657 | n_decoded = 100, tg = 62.21 t/s
|
| 261 |
+
1.31.732.686 I slot print_timing: id 0 | task 1657 |
|
| 262 |
+
prompt eval time = 1005.61 ms / 1028 tokens ( 0.98 ms per token, 1022.27 tokens per second)
|
| 263 |
+
eval time = 3060.81 ms / 192 tokens ( 15.94 ms per token, 62.73 tokens per second)
|
| 264 |
+
total time = 4066.42 ms / 1220 tokens
|
| 265 |
+
1.31.732.942 I slot release: id 0 | task 1657 | stop processing: n_tokens = 7230, truncated = 0
|
| 266 |
+
1.31.732.953 I srv update_slots: all slots are idle
|
| 267 |
+
1.31.750.071 I srv params_from_: Chat format: peg-native
|
| 268 |
+
1.31.750.187 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.974
|
| 269 |
+
1.31.750.241 I slot launch_slot_: id 0 | task 1851 | processing task, is_child = 0
|
| 270 |
+
1.31.750.248 W slot update_slots: id 0 | task 1851 | erased invalidated context checkpoint (pos_min = 6010, pos_max = 6010, n_tokens = 6011, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 271 |
+
1.31.751.779 W slot update_slots: id 0 | task 1851 | erased invalidated context checkpoint (pos_min = 7034, pos_max = 7034, n_tokens = 7035, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 272 |
+
1.35.110.825 I slot print_timing: id 0 | task 1851 | prompt processing, n_tokens = 4096, progress = 0.58, t = 3.36 s / 1218.84 tokens per second
|
| 273 |
+
1.36.865.814 I slot print_timing: id 0 | task 1851 | prompt processing, n_tokens = 6011, progress = 0.85, t = 5.12 s / 1175.04 tokens per second
|
| 274 |
+
1.36.917.039 I slot create_check: id 0 | task 1851 | created context checkpoint 1 of 32 (pos_min = 6010, pos_max = 6010, n_tokens = 6011, size = 62.813 MiB)
|
| 275 |
+
1.37.839.294 I slot print_timing: id 0 | task 1851 | prompt processing, n_tokens = 7035, progress = 1.00, t = 6.09 s / 1155.36 tokens per second
|
| 276 |
+
1.37.897.427 I slot create_check: id 0 | task 1851 | created context checkpoint 2 of 32 (pos_min = 7034, pos_max = 7034, n_tokens = 7035, size = 62.813 MiB)
|
| 277 |
+
1.39.536.081 I slot print_timing: id 0 | task 1851 | n_decoded = 100, tg = 62.18 t/s
|
| 278 |
+
1.40.995.612 I slot print_timing: id 0 | task 1851 |
|
| 279 |
+
prompt eval time = 6177.63 ms / 7039 tokens ( 0.88 ms per token, 1139.43 tokens per second)
|
| 280 |
+
eval time = 3067.72 ms / 192 tokens ( 15.98 ms per token, 62.59 tokens per second)
|
| 281 |
+
total time = 9245.35 ms / 7231 tokens
|
| 282 |
+
1.40.995.868 I slot release: id 0 | task 1851 | stop processing: n_tokens = 7230, truncated = 0
|
| 283 |
+
1.40.995.883 I srv update_slots: all slots are idle
|
| 284 |
+
1.40.996.543 I srv operator(): operator(): cleaning up before exit...
|
recipe/logs/b_n-tools-q106-c1-probe.log
ADDED
|
@@ -0,0 +1,286 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
0.00.113.666 W Setting 'enable_thinking' via --chat-template-kwargs is deprecated. Use --reasoning on / --reasoning off instead.
|
| 2 |
+
0.00.124.179 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
|
| 3 |
+
0.00.124.183 I device_info:
|
| 4 |
+
0.00.124.292 I - ROCm0 : AMD Radeon Graphics (131072 MiB, 123913 MiB free)
|
| 5 |
+
0.00.124.424 I - Vulkan0 : AMD Radeon Graphics (RADV GFX1151) (132096 MiB, 131922 MiB free)
|
| 6 |
+
0.00.124.432 I - CPU : AMD RYZEN AI MAX+ 395 w/ Radeon 8060S (127438 MiB, 127438 MiB free)
|
| 7 |
+
0.00.124.514 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
|
| 8 |
+
0.00.124.572 I srv init: running without SSL
|
| 9 |
+
0.00.124.596 I srv init: using 31 threads for HTTP server
|
| 10 |
+
0.00.124.597 I srv init: the WebUI is disabled
|
| 11 |
+
0.00.124.672 I srv start: binding port with default address family
|
| 12 |
+
0.00.125.865 I srv main: loading model
|
| 13 |
+
0.00.125.873 I srv load_model: loading model '/mnt/models/nex-n2.5-mini/out/Nex-N2.5-mini-Q4_0_ROCMFP4_STRIX_LEAN.gguf'
|
| 14 |
+
0.00.172.940 W llama_model_loader: direct I/O is enabled, disabling mmap
|
| 15 |
+
0.22.203.780 W llama_context: n_ctx_seq (65536) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
|
| 16 |
+
0.22.400.933 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|
| 17 |
+
0.22.718.577 I srv load_model: initializing slots, n_slots = 1
|
| 18 |
+
0.22.944.184 W srv load_model: speculative decoding will use checkpoints
|
| 19 |
+
0.22.944.193 W common_speculative_init: no implementations specified for speculative decoding
|
| 20 |
+
0.22.944.194 I slot load_model: id 0 | task -1 | new slot, n_ctx = 65536
|
| 21 |
+
0.22.944.259 I srv load_model: prompt cache RAM enabled: limit_mib=8192
|
| 22 |
+
0.22.944.261 I srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
|
| 23 |
+
0.22.944.283 I srv init: idle slots will be saved to prompt cache upon starting a new task
|
| 24 |
+
0.22.956.362 I init: chat template, example_format: '<|im_start|>system
|
| 25 |
+
You are a helpful assistant<|im_end|>
|
| 26 |
+
<|im_start|>user
|
| 27 |
+
Hello<|im_end|>
|
| 28 |
+
<|im_start|>assistant
|
| 29 |
+
<think>
|
| 30 |
+
|
| 31 |
+
</think>
|
| 32 |
+
|
| 33 |
+
Hi there<|im_end|>
|
| 34 |
+
<|im_start|>user
|
| 35 |
+
How are you?<|im_end|>
|
| 36 |
+
<|im_start|>assistant
|
| 37 |
+
<think>
|
| 38 |
+
|
| 39 |
+
</think>
|
| 40 |
+
|
| 41 |
+
'
|
| 42 |
+
0.22.964.154 I srv init: init: chat template, thinking = 1
|
| 43 |
+
0.22.964.188 I srv main: model loaded
|
| 44 |
+
0.22.964.191 I srv main: server is listening on http://127.0.0.1:18652
|
| 45 |
+
0.22.964.207 I srv update_slots: all slots are idle
|
| 46 |
+
0.24.203.095 I srv params_from_: Chat format: peg-native
|
| 47 |
+
0.24.203.661 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1
|
| 48 |
+
0.24.203.663 I srv get_availabl: updating prompt cache
|
| 49 |
+
0.24.203.670 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000
|
| 50 |
+
0.24.203.676 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 65536 tokens, 8589934592 est)
|
| 51 |
+
0.24.203.678 I srv get_availabl: prompt cache update took 0.01 ms
|
| 52 |
+
0.24.203.955 I reasoning-budget: activated, budget=2147483647 tokens
|
| 53 |
+
0.24.203.957 I reasoning-budget: deactivated (natural end)
|
| 54 |
+
0.24.203.972 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0
|
| 55 |
+
0.24.920.604 I slot create_check: id 0 | task 0 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
|
| 56 |
+
0.25.129.936 I slot print_timing: id 0 | task 0 |
|
| 57 |
+
prompt eval time = 758.65 ms / 427 tokens ( 1.78 ms per token, 562.84 tokens per second)
|
| 58 |
+
eval time = 167.29 ms / 4 tokens ( 41.82 ms per token, 23.91 tokens per second)
|
| 59 |
+
total time = 925.94 ms / 431 tokens
|
| 60 |
+
0.25.130.011 I slot release: id 0 | task 0 | stop processing: n_tokens = 430, truncated = 0
|
| 61 |
+
0.25.130.019 I srv update_slots: all slots are idle
|
| 62 |
+
0.25.185.780 I srv params_from_: Chat format: peg-native
|
| 63 |
+
0.25.186.376 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.942
|
| 64 |
+
0.25.186.613 I reasoning-budget: activated, budget=2147483647 tokens
|
| 65 |
+
0.25.186.618 I reasoning-budget: deactivated (natural end)
|
| 66 |
+
0.25.186.779 I slot launch_slot_: id 0 | task 6 | processing task, is_child = 0
|
| 67 |
+
0.25.186.816 W slot update_slots: id 0 | task 6 | n_past = 405, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
|
| 68 |
+
0.25.186.819 I slot update_slots: id 0 | task 6 | Checking checkpoint with [422, 422] against 405...
|
| 69 |
+
0.25.186.822 W slot update_slots: id 0 | task 6 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 70 |
+
0.25.186.829 W slot update_slots: id 0 | task 6 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 71 |
+
0.25.897.246 I slot create_check: id 0 | task 6 | created context checkpoint 1 of 32 (pos_min = 416, pos_max = 416, n_tokens = 417, size = 62.813 MiB)
|
| 72 |
+
0.25.983.622 I slot print_timing: id 0 | task 6 |
|
| 73 |
+
prompt eval time = 763.27 ms / 421 tokens ( 1.81 ms per token, 551.57 tokens per second)
|
| 74 |
+
eval time = 33.53 ms / 2 tokens ( 16.77 ms per token, 59.64 tokens per second)
|
| 75 |
+
total time = 796.80 ms / 423 tokens
|
| 76 |
+
0.25.983.694 I slot release: id 0 | task 6 | stop processing: n_tokens = 422, truncated = 0
|
| 77 |
+
0.25.983.724 I srv update_slots: all slots are idle
|
| 78 |
+
0.26.010.893 I srv params_from_: Chat format: peg-native
|
| 79 |
+
0.26.011.386 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.960
|
| 80 |
+
0.26.011.582 I reasoning-budget: activated, budget=2147483647 tokens
|
| 81 |
+
0.26.011.584 I reasoning-budget: deactivated (natural end)
|
| 82 |
+
0.26.011.615 I slot launch_slot_: id 0 | task 10 | processing task, is_child = 0
|
| 83 |
+
0.26.011.627 W slot update_slots: id 0 | task 10 | n_past = 405, slot.prompt.tokens.size() = 422, seq_id = 0, pos_min = 421, n_swa = 0
|
| 84 |
+
0.26.011.628 I slot update_slots: id 0 | task 10 | Checking checkpoint with [416, 416] against 405...
|
| 85 |
+
0.26.011.629 W slot update_slots: id 0 | task 10 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 86 |
+
0.26.011.634 W slot update_slots: id 0 | task 10 | erased invalidated context checkpoint (pos_min = 416, pos_max = 416, n_tokens = 417, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 87 |
+
0.26.506.256 I slot create_check: id 0 | task 10 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
|
| 88 |
+
0.27.421.771 I slot print_timing: id 0 | task 10 |
|
| 89 |
+
prompt eval time = 545.66 ms / 424 tokens ( 1.29 ms per token, 777.04 tokens per second)
|
| 90 |
+
eval time = 864.45 ms / 39 tokens ( 22.17 ms per token, 45.12 tokens per second)
|
| 91 |
+
total time = 1410.11 ms / 463 tokens
|
| 92 |
+
0.27.421.958 I slot release: id 0 | task 10 | stop processing: n_tokens = 462, truncated = 0
|
| 93 |
+
0.27.422.017 I srv update_slots: all slots are idle
|
| 94 |
+
0.27.483.797 I srv params_from_: Chat format: peg-native
|
| 95 |
+
0.27.486.309 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.951 (> 0.100 thold), f_keep = 0.879
|
| 96 |
+
0.27.486.919 I reasoning-budget: activated, budget=2147483647 tokens
|
| 97 |
+
0.27.486.923 I reasoning-budget: deactivated (natural end)
|
| 98 |
+
0.27.487.006 I slot launch_slot_: id 0 | task 51 | processing task, is_child = 0
|
| 99 |
+
0.27.487.029 W slot update_slots: id 0 | task 51 | n_past = 406, slot.prompt.tokens.size() = 462, seq_id = 0, pos_min = 461, n_swa = 0
|
| 100 |
+
0.27.487.032 I slot update_slots: id 0 | task 51 | Checking checkpoint with [419, 419] against 406...
|
| 101 |
+
0.27.487.034 W slot update_slots: id 0 | task 51 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 102 |
+
0.27.487.041 W slot update_slots: id 0 | task 51 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 103 |
+
0.28.076.337 I slot create_check: id 0 | task 51 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
|
| 104 |
+
0.28.262.498 I slot print_timing: id 0 | task 51 |
|
| 105 |
+
prompt eval time = 638.39 ms / 427 tokens ( 1.50 ms per token, 668.86 tokens per second)
|
| 106 |
+
eval time = 137.05 ms / 4 tokens ( 34.26 ms per token, 29.19 tokens per second)
|
| 107 |
+
total time = 775.45 ms / 431 tokens
|
| 108 |
+
0.28.262.688 I slot release: id 0 | task 51 | stop processing: n_tokens = 430, truncated = 0
|
| 109 |
+
0.28.262.742 I srv update_slots: all slots are idle
|
| 110 |
+
0.28.315.880 I srv params_from_: Chat format: peg-native
|
| 111 |
+
0.28.318.140 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.942
|
| 112 |
+
0.28.318.686 I reasoning-budget: activated, budget=2147483647 tokens
|
| 113 |
+
0.28.318.690 I reasoning-budget: deactivated (natural end)
|
| 114 |
+
0.28.318.777 I slot launch_slot_: id 0 | task 57 | processing task, is_child = 0
|
| 115 |
+
0.28.318.800 W slot update_slots: id 0 | task 57 | n_past = 405, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
|
| 116 |
+
0.28.318.803 I slot update_slots: id 0 | task 57 | Checking checkpoint with [422, 422] against 405...
|
| 117 |
+
0.28.318.805 W slot update_slots: id 0 | task 57 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 118 |
+
0.28.318.810 W slot update_slots: id 0 | task 57 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 119 |
+
0.28.907.669 I slot create_check: id 0 | task 57 | created context checkpoint 1 of 32 (pos_min = 416, pos_max = 416, n_tokens = 417, size = 62.813 MiB)
|
| 120 |
+
0.28.982.494 I slot print_timing: id 0 | task 57 |
|
| 121 |
+
prompt eval time = 636.88 ms / 421 tokens ( 1.51 ms per token, 661.04 tokens per second)
|
| 122 |
+
eval time = 26.81 ms / 2 tokens ( 13.41 ms per token, 74.60 tokens per second)
|
| 123 |
+
total time = 663.69 ms / 423 tokens
|
| 124 |
+
0.28.982.572 I slot release: id 0 | task 57 | stop processing: n_tokens = 422, truncated = 0
|
| 125 |
+
0.28.982.598 I srv update_slots: all slots are idle
|
| 126 |
+
0.29.009.222 I srv params_from_: Chat format: peg-native
|
| 127 |
+
0.29.011.294 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.960
|
| 128 |
+
0.29.011.848 I reasoning-budget: activated, budget=2147483647 tokens
|
| 129 |
+
0.29.011.851 I reasoning-budget: deactivated (natural end)
|
| 130 |
+
0.29.011.922 I slot launch_slot_: id 0 | task 61 | processing task, is_child = 0
|
| 131 |
+
0.29.011.941 W slot update_slots: id 0 | task 61 | n_past = 405, slot.prompt.tokens.size() = 422, seq_id = 0, pos_min = 421, n_swa = 0
|
| 132 |
+
0.29.011.943 I slot update_slots: id 0 | task 61 | Checking checkpoint with [416, 416] against 405...
|
| 133 |
+
0.29.011.945 W slot update_slots: id 0 | task 61 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 134 |
+
0.29.011.949 W slot update_slots: id 0 | task 61 | erased invalidated context checkpoint (pos_min = 416, pos_max = 416, n_tokens = 417, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 135 |
+
0.29.554.822 I slot create_check: id 0 | task 61 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
|
| 136 |
+
0.30.389.784 I slot print_timing: id 0 | task 61 |
|
| 137 |
+
prompt eval time = 589.42 ms / 424 tokens ( 1.39 ms per token, 719.35 tokens per second)
|
| 138 |
+
eval time = 788.41 ms / 39 tokens ( 20.22 ms per token, 49.47 tokens per second)
|
| 139 |
+
total time = 1377.83 ms / 463 tokens
|
| 140 |
+
0.30.389.858 I slot release: id 0 | task 61 | stop processing: n_tokens = 462, truncated = 0
|
| 141 |
+
0.30.389.883 I srv update_slots: all slots are idle
|
| 142 |
+
0.30.401.499 I srv params_from_: Chat format: peg-native
|
| 143 |
+
0.30.401.853 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.879
|
| 144 |
+
0.30.402.127 I reasoning-budget: activated, budget=2147483647 tokens
|
| 145 |
+
0.30.402.129 I reasoning-budget: deactivated (natural end)
|
| 146 |
+
0.30.402.192 I slot launch_slot_: id 0 | task 102 | processing task, is_child = 0
|
| 147 |
+
0.30.402.206 W slot update_slots: id 0 | task 102 | n_past = 406, slot.prompt.tokens.size() = 462, seq_id = 0, pos_min = 461, n_swa = 0
|
| 148 |
+
0.30.402.208 I slot update_slots: id 0 | task 102 | Checking checkpoint with [419, 419] against 406...
|
| 149 |
+
0.30.402.209 W slot update_slots: id 0 | task 102 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 150 |
+
0.30.402.212 W slot update_slots: id 0 | task 102 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 151 |
+
0.30.923.747 I slot create_check: id 0 | task 102 | created context checkpoint 1 of 32 (pos_min = 420, pos_max = 420, n_tokens = 421, size = 62.813 MiB)
|
| 152 |
+
0.31.379.211 I slot print_timing: id 0 | task 102 |
|
| 153 |
+
prompt eval time = 572.31 ms / 425 tokens ( 1.35 ms per token, 742.61 tokens per second)
|
| 154 |
+
eval time = 404.68 ms / 17 tokens ( 23.80 ms per token, 42.01 tokens per second)
|
| 155 |
+
total time = 976.98 ms / 442 tokens
|
| 156 |
+
0.31.379.332 I slot release: id 0 | task 102 | stop processing: n_tokens = 441, truncated = 0
|
| 157 |
+
0.31.379.373 I srv update_slots: all slots are idle
|
| 158 |
+
0.31.405.914 I srv params_from_: Chat format: peg-native
|
| 159 |
+
0.31.406.236 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.967 (> 0.100 thold), f_keep = 0.918
|
| 160 |
+
0.31.406.397 I reasoning-budget: activated, budget=2147483647 tokens
|
| 161 |
+
0.31.406.398 I reasoning-budget: deactivated (natural end)
|
| 162 |
+
0.31.406.427 I slot launch_slot_: id 0 | task 121 | processing task, is_child = 0
|
| 163 |
+
0.31.406.435 W slot update_slots: id 0 | task 121 | n_past = 405, slot.prompt.tokens.size() = 441, seq_id = 0, pos_min = 440, n_swa = 0
|
| 164 |
+
0.31.406.436 I slot update_slots: id 0 | task 121 | Checking checkpoint with [420, 420] against 405...
|
| 165 |
+
0.31.406.437 W slot update_slots: id 0 | task 121 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 166 |
+
0.31.406.439 W slot update_slots: id 0 | task 121 | erased invalidated context checkpoint (pos_min = 420, pos_max = 420, n_tokens = 421, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 167 |
+
0.31.861.565 I slot create_check: id 0 | task 121 | created context checkpoint 1 of 32 (pos_min = 414, pos_max = 414, n_tokens = 415, size = 62.813 MiB)
|
| 168 |
+
0.32.160.128 I slot print_timing: id 0 | task 121 |
|
| 169 |
+
prompt eval time = 491.31 ms / 419 tokens ( 1.17 ms per token, 852.82 tokens per second)
|
| 170 |
+
eval time = 262.37 ms / 12 tokens ( 21.86 ms per token, 45.74 tokens per second)
|
| 171 |
+
total time = 753.68 ms / 431 tokens
|
| 172 |
+
0.32.160.211 I slot release: id 0 | task 121 | stop processing: n_tokens = 430, truncated = 0
|
| 173 |
+
0.32.160.241 I srv update_slots: all slots are idle
|
| 174 |
+
0.32.209.194 I srv params_from_: Chat format: peg-native
|
| 175 |
+
0.32.209.712 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.960 (> 0.100 thold), f_keep = 0.942
|
| 176 |
+
0.32.209.953 I reasoning-budget: activated, budget=2147483647 tokens
|
| 177 |
+
0.32.209.954 I reasoning-budget: deactivated (natural end)
|
| 178 |
+
0.32.209.994 I slot launch_slot_: id 0 | task 135 | processing task, is_child = 0
|
| 179 |
+
0.32.210.006 W slot update_slots: id 0 | task 135 | n_past = 405, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
|
| 180 |
+
0.32.210.007 I slot update_slots: id 0 | task 135 | Checking checkpoint with [414, 414] against 405...
|
| 181 |
+
0.32.210.008 W slot update_slots: id 0 | task 135 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 182 |
+
0.32.210.011 W slot update_slots: id 0 | task 135 | erased invalidated context checkpoint (pos_min = 414, pos_max = 414, n_tokens = 415, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 183 |
+
0.32.786.978 I slot create_check: id 0 | task 135 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
|
| 184 |
+
0.33.839.717 I slot print_timing: id 0 | task 135 |
|
| 185 |
+
prompt eval time = 614.39 ms / 422 tokens ( 1.46 ms per token, 686.86 tokens per second)
|
| 186 |
+
eval time = 1015.30 ms / 53 tokens ( 19.16 ms per token, 52.20 tokens per second)
|
| 187 |
+
total time = 1629.69 ms / 475 tokens
|
| 188 |
+
0.33.839.794 I slot release: id 0 | task 135 | stop processing: n_tokens = 474, truncated = 0
|
| 189 |
+
0.33.839.825 I srv update_slots: all slots are idle
|
| 190 |
+
0.33.879.962 I srv params_from_: Chat format: peg-native
|
| 191 |
+
0.33.880.516 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.958 (> 0.100 thold), f_keep = 0.857
|
| 192 |
+
0.33.880.743 I reasoning-budget: activated, budget=2147483647 tokens
|
| 193 |
+
0.33.880.744 I reasoning-budget: deactivated (natural end)
|
| 194 |
+
0.33.880.792 I slot launch_slot_: id 0 | task 190 | processing task, is_child = 0
|
| 195 |
+
0.33.880.805 W slot update_slots: id 0 | task 190 | n_past = 406, slot.prompt.tokens.size() = 474, seq_id = 0, pos_min = 473, n_swa = 0
|
| 196 |
+
0.33.880.806 I slot update_slots: id 0 | task 190 | Checking checkpoint with [417, 417] against 406...
|
| 197 |
+
0.33.880.807 W slot update_slots: id 0 | task 190 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 198 |
+
0.33.880.810 W slot update_slots: id 0 | task 190 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 199 |
+
0.34.423.705 I slot create_check: id 0 | task 190 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
|
| 200 |
+
0.34.628.920 I slot print_timing: id 0 | task 190 |
|
| 201 |
+
prompt eval time = 583.31 ms / 424 tokens ( 1.38 ms per token, 726.89 tokens per second)
|
| 202 |
+
eval time = 164.80 ms / 7 tokens ( 23.54 ms per token, 42.48 tokens per second)
|
| 203 |
+
total time = 748.11 ms / 431 tokens
|
| 204 |
+
0.34.628.991 I slot release: id 0 | task 190 | stop processing: n_tokens = 430, truncated = 0
|
| 205 |
+
0.34.629.018 I srv update_slots: all slots are idle
|
| 206 |
+
0.34.642.421 I srv params_from_: Chat format: peg-native
|
| 207 |
+
0.34.642.945 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.969 (> 0.100 thold), f_keep = 0.942
|
| 208 |
+
0.34.643.148 I reasoning-budget: activated, budget=2147483647 tokens
|
| 209 |
+
0.34.643.150 I reasoning-budget: deactivated (natural end)
|
| 210 |
+
0.34.643.185 I slot launch_slot_: id 0 | task 199 | processing task, is_child = 0
|
| 211 |
+
0.34.643.195 W slot update_slots: id 0 | task 199 | n_past = 405, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
|
| 212 |
+
0.34.643.196 I slot update_slots: id 0 | task 199 | Checking checkpoint with [419, 419] against 405...
|
| 213 |
+
0.34.643.197 W slot update_slots: id 0 | task 199 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 214 |
+
0.34.643.199 W slot update_slots: id 0 | task 199 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 215 |
+
0.35.154.839 I slot create_check: id 0 | task 199 | created context checkpoint 1 of 32 (pos_min = 413, pos_max = 413, n_tokens = 414, size = 62.813 MiB)
|
| 216 |
+
0.35.329.011 I slot print_timing: id 0 | task 199 |
|
| 217 |
+
prompt eval time = 563.78 ms / 418 tokens ( 1.35 ms per token, 741.42 tokens per second)
|
| 218 |
+
eval time = 122.02 ms / 5 tokens ( 24.40 ms per token, 40.98 tokens per second)
|
| 219 |
+
total time = 685.80 ms / 423 tokens
|
| 220 |
+
0.35.329.095 I slot release: id 0 | task 199 | stop processing: n_tokens = 422, truncated = 0
|
| 221 |
+
0.35.329.136 I srv update_slots: all slots are idle
|
| 222 |
+
0.35.355.115 I srv params_from_: Chat format: peg-native
|
| 223 |
+
0.35.355.649 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.960
|
| 224 |
+
0.35.356.326 I reasoning-budget: activated, budget=2147483647 tokens
|
| 225 |
+
0.35.356.340 I reasoning-budget: deactivated (natural end)
|
| 226 |
+
0.35.356.552 I slot launch_slot_: id 0 | task 206 | processing task, is_child = 0
|
| 227 |
+
0.35.356.603 W slot update_slots: id 0 | task 206 | n_past = 405, slot.prompt.tokens.size() = 422, seq_id = 0, pos_min = 421, n_swa = 0
|
| 228 |
+
0.35.356.611 I slot update_slots: id 0 | task 206 | Checking checkpoint with [413, 413] against 405...
|
| 229 |
+
0.35.356.616 W slot update_slots: id 0 | task 206 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 230 |
+
0.35.356.634 W slot update_slots: id 0 | task 206 | erased invalidated context checkpoint (pos_min = 413, pos_max = 413, n_tokens = 414, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 231 |
+
0.35.935.879 I slot create_check: id 0 | task 206 | created context checkpoint 1 of 32 (pos_min = 416, pos_max = 416, n_tokens = 417, size = 62.813 MiB)
|
| 232 |
+
0.36.896.774 I slot print_timing: id 0 | task 206 |
|
| 233 |
+
prompt eval time = 625.99 ms / 421 tokens ( 1.49 ms per token, 672.53 tokens per second)
|
| 234 |
+
eval time = 914.17 ms / 42 tokens ( 21.77 ms per token, 45.94 tokens per second)
|
| 235 |
+
total time = 1540.17 ms / 463 tokens
|
| 236 |
+
0.36.896.862 I slot release: id 0 | task 206 | stop processing: n_tokens = 462, truncated = 0
|
| 237 |
+
0.36.896.896 I srv update_slots: all slots are idle
|
| 238 |
+
0.36.936.572 I srv params_from_: Chat format: peg-native
|
| 239 |
+
0.36.937.088 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.951 (> 0.100 thold), f_keep = 0.879
|
| 240 |
+
0.36.937.669 I reasoning-budget: activated, budget=2147483647 tokens
|
| 241 |
+
0.36.937.673 I reasoning-budget: deactivated (natural end)
|
| 242 |
+
0.36.937.771 I slot launch_slot_: id 0 | task 250 | processing task, is_child = 0
|
| 243 |
+
0.36.937.796 W slot update_slots: id 0 | task 250 | n_past = 406, slot.prompt.tokens.size() = 462, seq_id = 0, pos_min = 461, n_swa = 0
|
| 244 |
+
0.36.937.799 I slot update_slots: id 0 | task 250 | Checking checkpoint with [416, 416] against 406...
|
| 245 |
+
0.36.937.801 W slot update_slots: id 0 | task 250 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 246 |
+
0.36.937.807 W slot update_slots: id 0 | task 250 | erased invalidated context checkpoint (pos_min = 416, pos_max = 416, n_tokens = 417, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 247 |
+
0.37.457.001 I slot create_check: id 0 | task 250 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
|
| 248 |
+
0.37.596.606 I slot print_timing: id 0 | task 250 |
|
| 249 |
+
prompt eval time = 561.66 ms / 427 tokens ( 1.32 ms per token, 760.24 tokens per second)
|
| 250 |
+
eval time = 97.14 ms / 4 tokens ( 24.28 ms per token, 41.18 tokens per second)
|
| 251 |
+
total time = 658.80 ms / 431 tokens
|
| 252 |
+
0.37.596.716 I slot release: id 0 | task 250 | stop processing: n_tokens = 430, truncated = 0
|
| 253 |
+
0.37.596.750 I srv update_slots: all slots are idle
|
| 254 |
+
0.37.626.974 I srv params_from_: Chat format: peg-native
|
| 255 |
+
0.37.627.483 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.942
|
| 256 |
+
0.37.627.753 I reasoning-budget: activated, budget=2147483647 tokens
|
| 257 |
+
0.37.627.755 I reasoning-budget: deactivated (natural end)
|
| 258 |
+
0.37.627.791 I slot launch_slot_: id 0 | task 256 | processing task, is_child = 0
|
| 259 |
+
0.37.627.804 W slot update_slots: id 0 | task 256 | n_past = 405, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
|
| 260 |
+
0.37.627.806 I slot update_slots: id 0 | task 256 | Checking checkpoint with [422, 422] against 405...
|
| 261 |
+
0.37.627.807 W slot update_slots: id 0 | task 256 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 262 |
+
0.37.627.811 W slot update_slots: id 0 | task 256 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 263 |
+
0.38.154.752 I slot create_check: id 0 | task 256 | created context checkpoint 1 of 32 (pos_min = 416, pos_max = 416, n_tokens = 417, size = 62.813 MiB)
|
| 264 |
+
0.38.242.158 I slot print_timing: id 0 | task 256 |
|
| 265 |
+
prompt eval time = 577.92 ms / 421 tokens ( 1.37 ms per token, 728.48 tokens per second)
|
| 266 |
+
eval time = 36.41 ms / 2 tokens ( 18.21 ms per token, 54.92 tokens per second)
|
| 267 |
+
total time = 614.33 ms / 423 tokens
|
| 268 |
+
0.38.242.323 I slot release: id 0 | task 256 | stop processing: n_tokens = 422, truncated = 0
|
| 269 |
+
0.38.242.358 I srv update_slots: all slots are idle
|
| 270 |
+
0.38.257.074 I srv params_from_: Chat format: peg-native
|
| 271 |
+
0.38.257.513 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.960
|
| 272 |
+
0.38.257.760 I reasoning-budget: activated, budget=2147483647 tokens
|
| 273 |
+
0.38.257.761 I reasoning-budget: deactivated (natural end)
|
| 274 |
+
0.38.257.798 I slot launch_slot_: id 0 | task 260 | processing task, is_child = 0
|
| 275 |
+
0.38.257.809 W slot update_slots: id 0 | task 260 | n_past = 405, slot.prompt.tokens.size() = 422, seq_id = 0, pos_min = 421, n_swa = 0
|
| 276 |
+
0.38.257.809 I slot update_slots: id 0 | task 260 | Checking checkpoint with [416, 416] against 405...
|
| 277 |
+
0.38.257.810 W slot update_slots: id 0 | task 260 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 278 |
+
0.38.257.813 W slot update_slots: id 0 | task 260 | erased invalidated context checkpoint (pos_min = 416, pos_max = 416, n_tokens = 417, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 279 |
+
0.38.777.417 I slot create_check: id 0 | task 260 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
|
| 280 |
+
0.39.603.048 I slot print_timing: id 0 | task 260 |
|
| 281 |
+
prompt eval time = 562.75 ms / 424 tokens ( 1.33 ms per token, 753.44 tokens per second)
|
| 282 |
+
eval time = 782.47 ms / 39 tokens ( 20.06 ms per token, 49.84 tokens per second)
|
| 283 |
+
total time = 1345.22 ms / 463 tokens
|
| 284 |
+
0.39.603.254 I slot release: id 0 | task 260 | stop processing: n_tokens = 462, truncated = 0
|
| 285 |
+
0.39.603.336 I srv update_slots: all slots are idle
|
| 286 |
+
0.39.604.550 I srv operator(): operator(): cleaning up before exit...
|
recipe/logs/b_n-tools-q106-c1.log
ADDED
|
@@ -0,0 +1,302 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
0.00.120.044 W Setting 'enable_thinking' via --chat-template-kwargs is deprecated. Use --reasoning on / --reasoning off instead.
|
| 2 |
+
0.00.163.630 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
|
| 3 |
+
0.00.163.643 I device_info:
|
| 4 |
+
0.00.163.785 I - ROCm0 : AMD Radeon Graphics (131072 MiB, 123858 MiB free)
|
| 5 |
+
0.00.163.974 I - Vulkan0 : AMD Radeon Graphics (RADV GFX1151) (132096 MiB, 131922 MiB free)
|
| 6 |
+
0.00.163.982 I - CPU : AMD RYZEN AI MAX+ 395 w/ Radeon 8060S (127438 MiB, 127438 MiB free)
|
| 7 |
+
0.00.164.078 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
|
| 8 |
+
0.00.164.142 I srv init: running without SSL
|
| 9 |
+
0.00.164.166 I srv init: using 31 threads for HTTP server
|
| 10 |
+
0.00.164.167 I srv init: the WebUI is disabled
|
| 11 |
+
0.00.164.248 I srv start: binding port with default address family
|
| 12 |
+
0.00.165.500 I srv main: loading model
|
| 13 |
+
0.00.165.507 I srv load_model: loading model '/mnt/models/nex-n2.5-mini/out/Nex-N2.5-mini-Q4_0_ROCMFP4_STRIX_LEAN.gguf'
|
| 14 |
+
0.00.216.012 W llama_model_loader: direct I/O is enabled, disabling mmap
|
| 15 |
+
0.22.544.117 W llama_context: n_ctx_seq (65536) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
|
| 16 |
+
0.22.849.861 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|
| 17 |
+
0.23.278.480 I srv load_model: initializing slots, n_slots = 1
|
| 18 |
+
0.23.567.593 W srv load_model: speculative decoding will use checkpoints
|
| 19 |
+
0.23.567.602 W common_speculative_init: no implementations specified for speculative decoding
|
| 20 |
+
0.23.567.605 I slot load_model: id 0 | task -1 | new slot, n_ctx = 65536
|
| 21 |
+
0.23.567.670 I srv load_model: prompt cache RAM enabled: limit_mib=8192
|
| 22 |
+
0.23.567.688 I srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
|
| 23 |
+
0.23.567.709 I srv init: idle slots will be saved to prompt cache upon starting a new task
|
| 24 |
+
0.23.579.583 I init: chat template, example_format: '<|im_start|>system
|
| 25 |
+
You are a helpful assistant<|im_end|>
|
| 26 |
+
<|im_start|>user
|
| 27 |
+
Hello<|im_end|>
|
| 28 |
+
<|im_start|>assistant
|
| 29 |
+
<think>
|
| 30 |
+
|
| 31 |
+
</think>
|
| 32 |
+
|
| 33 |
+
Hi there<|im_end|>
|
| 34 |
+
<|im_start|>user
|
| 35 |
+
How are you?<|im_end|>
|
| 36 |
+
<|im_start|>assistant
|
| 37 |
+
<think>
|
| 38 |
+
|
| 39 |
+
</think>
|
| 40 |
+
|
| 41 |
+
'
|
| 42 |
+
0.23.587.209 I srv init: init: chat template, thinking = 1
|
| 43 |
+
0.23.587.253 I srv main: model loaded
|
| 44 |
+
0.23.587.256 I srv main: server is listening on http://127.0.0.1:18600
|
| 45 |
+
0.23.587.260 I srv update_slots: all slots are idle
|
| 46 |
+
0.24.570.949 I srv params_from_: Chat format: peg-native
|
| 47 |
+
0.24.571.406 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1
|
| 48 |
+
0.24.571.411 I srv get_availabl: updating prompt cache
|
| 49 |
+
0.24.571.419 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000
|
| 50 |
+
0.24.571.426 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 65536 tokens, 8589934592 est)
|
| 51 |
+
0.24.571.428 I srv get_availabl: prompt cache update took 0.01 ms
|
| 52 |
+
0.24.571.805 I reasoning-budget: activated, budget=2147483647 tokens
|
| 53 |
+
0.24.571.827 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0
|
| 54 |
+
0.25.235.195 I slot create_check: id 0 | task 0 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
|
| 55 |
+
0.25.806.187 I reasoning-budget: deactivated (natural end)
|
| 56 |
+
0.26.650.056 I slot print_timing: id 0 | task 0 |
|
| 57 |
+
prompt eval time = 715.70 ms / 422 tokens ( 1.70 ms per token, 589.63 tokens per second)
|
| 58 |
+
eval time = 1362.48 ms / 60 tokens ( 22.71 ms per token, 44.04 tokens per second)
|
| 59 |
+
total time = 2078.18 ms / 482 tokens
|
| 60 |
+
0.26.650.247 I slot release: id 0 | task 0 | stop processing: n_tokens = 481, truncated = 0
|
| 61 |
+
0.26.650.264 I srv update_slots: all slots are idle
|
| 62 |
+
0.26.675.130 I srv params_from_: Chat format: peg-native
|
| 63 |
+
0.26.675.629 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.906 (> 0.100 thold), f_keep = 0.842
|
| 64 |
+
0.26.676.106 I reasoning-budget: activated, budget=2147483647 tokens
|
| 65 |
+
0.26.676.196 I slot launch_slot_: id 0 | task 62 | processing task, is_child = 0
|
| 66 |
+
0.26.676.216 W slot update_slots: id 0 | task 62 | n_past = 405, slot.prompt.tokens.size() = 481, seq_id = 0, pos_min = 480, n_swa = 0
|
| 67 |
+
0.26.676.222 I slot update_slots: id 0 | task 62 | Checking checkpoint with [417, 417] against 405...
|
| 68 |
+
0.26.676.224 W slot update_slots: id 0 | task 62 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 69 |
+
0.26.676.229 W slot update_slots: id 0 | task 62 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 70 |
+
0.27.231.790 I slot create_check: id 0 | task 62 | created context checkpoint 1 of 32 (pos_min = 442, pos_max = 442, n_tokens = 443, size = 62.813 MiB)
|
| 71 |
+
0.27.733.232 I reasoning-budget: deactivated (natural end)
|
| 72 |
+
0.28.618.139 I slot print_timing: id 0 | task 62 |
|
| 73 |
+
prompt eval time = 593.45 ms / 447 tokens ( 1.33 ms per token, 753.22 tokens per second)
|
| 74 |
+
eval time = 1348.46 ms / 65 tokens ( 20.75 ms per token, 48.20 tokens per second)
|
| 75 |
+
total time = 1941.91 ms / 512 tokens
|
| 76 |
+
0.28.618.225 I slot release: id 0 | task 62 | stop processing: n_tokens = 511, truncated = 0
|
| 77 |
+
0.28.618.260 I srv update_slots: all slots are idle
|
| 78 |
+
0.28.632.505 I srv params_from_: Chat format: peg-native
|
| 79 |
+
0.28.632.871 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.953 (> 0.100 thold), f_keep = 0.793
|
| 80 |
+
0.28.633.063 I reasoning-budget: activated, budget=2147483647 tokens
|
| 81 |
+
0.28.633.092 I slot launch_slot_: id 0 | task 129 | processing task, is_child = 0
|
| 82 |
+
0.28.633.101 W slot update_slots: id 0 | task 129 | n_past = 405, slot.prompt.tokens.size() = 511, seq_id = 0, pos_min = 510, n_swa = 0
|
| 83 |
+
0.28.633.102 I slot update_slots: id 0 | task 129 | Checking checkpoint with [442, 442] against 405...
|
| 84 |
+
0.28.633.103 W slot update_slots: id 0 | task 129 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 85 |
+
0.28.633.104 W slot update_slots: id 0 | task 129 | erased invalidated context checkpoint (pos_min = 442, pos_max = 442, n_tokens = 443, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 86 |
+
0.29.161.759 I slot create_check: id 0 | task 129 | created context checkpoint 1 of 32 (pos_min = 420, pos_max = 420, n_tokens = 421, size = 62.813 MiB)
|
| 87 |
+
0.29.478.654 I reasoning-budget: deactivated (natural end)
|
| 88 |
+
0.30.331.330 I slot print_timing: id 0 | task 129 |
|
| 89 |
+
prompt eval time = 580.01 ms / 425 tokens ( 1.36 ms per token, 732.75 tokens per second)
|
| 90 |
+
eval time = 1118.18 ms / 52 tokens ( 21.50 ms per token, 46.50 tokens per second)
|
| 91 |
+
total time = 1698.19 ms / 477 tokens
|
| 92 |
+
0.30.331.521 I slot release: id 0 | task 129 | stop processing: n_tokens = 476, truncated = 0
|
| 93 |
+
0.30.331.586 I srv update_slots: all slots are idle
|
| 94 |
+
0.30.386.522 I srv params_from_: Chat format: peg-native
|
| 95 |
+
0.30.388.735 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.953 (> 0.100 thold), f_keep = 0.851
|
| 96 |
+
0.30.389.319 I reasoning-budget: activated, budget=2147483647 tokens
|
| 97 |
+
0.30.389.437 I slot launch_slot_: id 0 | task 183 | processing task, is_child = 0
|
| 98 |
+
0.30.389.465 W slot update_slots: id 0 | task 183 | n_past = 405, slot.prompt.tokens.size() = 476, seq_id = 0, pos_min = 475, n_swa = 0
|
| 99 |
+
0.30.389.469 I slot update_slots: id 0 | task 183 | Checking checkpoint with [420, 420] against 405...
|
| 100 |
+
0.30.389.471 W slot update_slots: id 0 | task 183 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 101 |
+
0.30.389.478 W slot update_slots: id 0 | task 183 | erased invalidated context checkpoint (pos_min = 420, pos_max = 420, n_tokens = 421, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 102 |
+
0.30.972.513 I slot create_check: id 0 | task 183 | created context checkpoint 1 of 32 (pos_min = 420, pos_max = 420, n_tokens = 421, size = 62.813 MiB)
|
| 103 |
+
0.31.338.635 I reasoning-budget: deactivated (natural end)
|
| 104 |
+
0.31.461.621 I slot print_timing: id 0 | task 183 |
|
| 105 |
+
prompt eval time = 630.63 ms / 425 tokens ( 1.48 ms per token, 673.92 tokens per second)
|
| 106 |
+
eval time = 441.50 ms / 17 tokens ( 25.97 ms per token, 38.51 tokens per second)
|
| 107 |
+
total time = 1072.13 ms / 442 tokens
|
| 108 |
+
0.31.461.844 I slot release: id 0 | task 183 | stop processing: n_tokens = 441, truncated = 0
|
| 109 |
+
0.31.461.907 I srv update_slots: all slots are idle
|
| 110 |
+
0.31.524.318 I srv params_from_: Chat format: peg-native
|
| 111 |
+
0.31.526.203 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.921
|
| 112 |
+
0.31.526.820 I reasoning-budget: activated, budget=2147483647 tokens
|
| 113 |
+
0.31.526.906 I slot launch_slot_: id 0 | task 202 | processing task, is_child = 0
|
| 114 |
+
0.31.526.929 W slot update_slots: id 0 | task 202 | n_past = 406, slot.prompt.tokens.size() = 441, seq_id = 0, pos_min = 440, n_swa = 0
|
| 115 |
+
0.31.526.933 I slot update_slots: id 0 | task 202 | Checking checkpoint with [420, 420] against 406...
|
| 116 |
+
0.31.526.935 W slot update_slots: id 0 | task 202 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 117 |
+
0.31.526.940 W slot update_slots: id 0 | task 202 | erased invalidated context checkpoint (pos_min = 420, pos_max = 420, n_tokens = 421, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 118 |
+
0.32.129.825 I slot create_check: id 0 | task 202 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
|
| 119 |
+
0.32.519.469 I reasoning-budget: deactivated (natural end)
|
| 120 |
+
0.33.361.317 I slot print_timing: id 0 | task 202 |
|
| 121 |
+
prompt eval time = 645.36 ms / 422 tokens ( 1.53 ms per token, 653.90 tokens per second)
|
| 122 |
+
eval time = 1189.02 ms / 55 tokens ( 21.62 ms per token, 46.26 tokens per second)
|
| 123 |
+
total time = 1834.38 ms / 477 tokens
|
| 124 |
+
0.33.361.399 I slot release: id 0 | task 202 | stop processing: n_tokens = 476, truncated = 0
|
| 125 |
+
0.33.361.431 I srv update_slots: all slots are idle
|
| 126 |
+
0.33.406.059 I srv params_from_: Chat format: peg-native
|
| 127 |
+
0.33.406.536 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.854 (> 0.100 thold), f_keep = 0.884
|
| 128 |
+
0.33.406.842 I reasoning-budget: activated, budget=2147483647 tokens
|
| 129 |
+
0.33.406.887 I slot launch_slot_: id 0 | task 259 | processing task, is_child = 0
|
| 130 |
+
0.33.406.903 W slot update_slots: id 0 | task 259 | n_past = 421, slot.prompt.tokens.size() = 476, seq_id = 0, pos_min = 475, n_swa = 0
|
| 131 |
+
0.33.406.904 I slot update_slots: id 0 | task 259 | Checking checkpoint with [417, 417] against 421...
|
| 132 |
+
0.33.410.865 W slot update_slots: id 0 | task 259 | restored context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_past = 418, size = 62.813 MiB)
|
| 133 |
+
0.33.668.081 I slot create_check: id 0 | task 259 | created context checkpoint 2 of 32 (pos_min = 488, pos_max = 488, n_tokens = 489, size = 62.813 MiB)
|
| 134 |
+
0.34.301.983 I reasoning-budget: deactivated (natural end)
|
| 135 |
+
0.34.590.965 I slot print_timing: id 0 | task 259 |
|
| 136 |
+
prompt eval time = 311.59 ms / 75 tokens ( 4.15 ms per token, 240.70 tokens per second)
|
| 137 |
+
eval time = 872.45 ms / 38 tokens ( 22.96 ms per token, 43.56 tokens per second)
|
| 138 |
+
total time = 1184.04 ms / 113 tokens
|
| 139 |
+
0.34.591.058 I slot release: id 0 | task 259 | stop processing: n_tokens = 530, truncated = 0
|
| 140 |
+
0.34.591.093 I srv update_slots: all slots are idle
|
| 141 |
+
0.34.626.053 I srv params_from_: Chat format: peg-native
|
| 142 |
+
0.34.626.459 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.972 (> 0.100 thold), f_keep = 0.774
|
| 143 |
+
0.34.626.655 I reasoning-budget: activated, budget=2147483647 tokens
|
| 144 |
+
0.34.626.694 I slot launch_slot_: id 0 | task 299 | processing task, is_child = 0
|
| 145 |
+
0.34.626.705 W slot update_slots: id 0 | task 299 | n_past = 410, slot.prompt.tokens.size() = 530, seq_id = 0, pos_min = 529, n_swa = 0
|
| 146 |
+
0.34.626.705 I slot update_slots: id 0 | task 299 | Checking checkpoint with [488, 488] against 410...
|
| 147 |
+
0.34.626.706 I slot update_slots: id 0 | task 299 | Checking checkpoint with [417, 417] against 410...
|
| 148 |
+
0.34.626.706 W slot update_slots: id 0 | task 299 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 149 |
+
0.34.626.709 W slot update_slots: id 0 | task 299 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 150 |
+
0.34.627.674 W slot update_slots: id 0 | task 299 | erased invalidated context checkpoint (pos_min = 488, pos_max = 488, n_tokens = 489, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 151 |
+
0.35.160.501 I slot create_check: id 0 | task 299 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
|
| 152 |
+
0.35.462.626 I reasoning-budget: deactivated (natural end)
|
| 153 |
+
0.36.334.370 I slot print_timing: id 0 | task 299 |
|
| 154 |
+
prompt eval time = 570.77 ms / 422 tokens ( 1.35 ms per token, 739.35 tokens per second)
|
| 155 |
+
eval time = 1136.85 ms / 53 tokens ( 21.45 ms per token, 46.62 tokens per second)
|
| 156 |
+
total time = 1707.63 ms / 475 tokens
|
| 157 |
+
0.36.334.535 I slot release: id 0 | task 299 | stop processing: n_tokens = 474, truncated = 0
|
| 158 |
+
0.36.334.587 I srv update_slots: all slots are idle
|
| 159 |
+
0.36.346.401 I srv params_from_: Chat format: peg-native
|
| 160 |
+
0.36.346.871 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.935 (> 0.100 thold), f_keep = 0.854
|
| 161 |
+
0.36.347.275 I reasoning-budget: activated, budget=2147483647 tokens
|
| 162 |
+
0.36.347.355 I slot launch_slot_: id 0 | task 354 | processing task, is_child = 0
|
| 163 |
+
0.36.347.371 W slot update_slots: id 0 | task 354 | n_past = 405, slot.prompt.tokens.size() = 474, seq_id = 0, pos_min = 473, n_swa = 0
|
| 164 |
+
0.36.347.374 I slot update_slots: id 0 | task 354 | Checking checkpoint with [417, 417] against 405...
|
| 165 |
+
0.36.347.376 W slot update_slots: id 0 | task 354 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 166 |
+
0.36.347.381 W slot update_slots: id 0 | task 354 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 167 |
+
0.36.950.505 I slot create_check: id 0 | task 354 | created context checkpoint 1 of 32 (pos_min = 428, pos_max = 428, n_tokens = 429, size = 62.813 MiB)
|
| 168 |
+
0.37.901.136 I reasoning-budget: deactivated (natural end)
|
| 169 |
+
0.39.133.735 I slot print_timing: id 0 | task 354 | n_decoded = 100, tg = 46.90 t/s
|
| 170 |
+
0.39.607.166 I slot print_timing: id 0 | task 354 |
|
| 171 |
+
prompt eval time = 653.96 ms / 433 tokens ( 1.51 ms per token, 662.12 tokens per second)
|
| 172 |
+
eval time = 2605.80 ms / 123 tokens ( 21.19 ms per token, 47.20 tokens per second)
|
| 173 |
+
total time = 3259.76 ms / 556 tokens
|
| 174 |
+
0.39.607.372 I slot release: id 0 | task 354 | stop processing: n_tokens = 555, truncated = 0
|
| 175 |
+
0.39.607.435 I srv update_slots: all slots are idle
|
| 176 |
+
0.39.655.799 I srv params_from_: Chat format: peg-native
|
| 177 |
+
0.39.656.354 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.730
|
| 178 |
+
0.39.656.713 I reasoning-budget: activated, budget=2147483647 tokens
|
| 179 |
+
0.39.656.717 I reasoning-budget: deactivated (natural end)
|
| 180 |
+
0.39.656.803 I slot launch_slot_: id 0 | task 479 | processing task, is_child = 0
|
| 181 |
+
0.39.656.826 W slot update_slots: id 0 | task 479 | n_past = 405, slot.prompt.tokens.size() = 555, seq_id = 0, pos_min = 554, n_swa = 0
|
| 182 |
+
0.39.656.828 I slot update_slots: id 0 | task 479 | Checking checkpoint with [428, 428] against 405...
|
| 183 |
+
0.39.656.830 W slot update_slots: id 0 | task 479 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 184 |
+
0.39.656.835 W slot update_slots: id 0 | task 479 | erased invalidated context checkpoint (pos_min = 428, pos_max = 428, n_tokens = 429, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 185 |
+
0.40.205.076 I slot create_check: id 0 | task 479 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
|
| 186 |
+
0.41.065.168 I slot print_timing: id 0 | task 479 |
|
| 187 |
+
prompt eval time = 582.67 ms / 424 tokens ( 1.37 ms per token, 727.68 tokens per second)
|
| 188 |
+
eval time = 825.66 ms / 39 tokens ( 21.17 ms per token, 47.23 tokens per second)
|
| 189 |
+
total time = 1408.33 ms / 463 tokens
|
| 190 |
+
0.41.065.250 I slot release: id 0 | task 479 | stop processing: n_tokens = 462, truncated = 0
|
| 191 |
+
0.41.065.277 I srv update_slots: all slots are idle
|
| 192 |
+
0.41.077.853 I srv params_from_: Chat format: peg-native
|
| 193 |
+
0.41.078.228 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.902 (> 0.100 thold), f_keep = 0.877
|
| 194 |
+
0.41.078.482 I reasoning-budget: activated, budget=2147483647 tokens
|
| 195 |
+
0.41.078.484 I reasoning-budget: deactivated (natural end)
|
| 196 |
+
0.41.078.524 I slot launch_slot_: id 0 | task 520 | processing task, is_child = 0
|
| 197 |
+
0.41.078.534 W slot update_slots: id 0 | task 520 | n_past = 405, slot.prompt.tokens.size() = 462, seq_id = 0, pos_min = 461, n_swa = 0
|
| 198 |
+
0.41.078.535 I slot update_slots: id 0 | task 520 | Checking checkpoint with [419, 419] against 405...
|
| 199 |
+
0.41.078.536 W slot update_slots: id 0 | task 520 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 200 |
+
0.41.078.537 W slot update_slots: id 0 | task 520 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 201 |
+
0.41.589.485 I slot create_check: id 0 | task 520 | created context checkpoint 1 of 32 (pos_min = 444, pos_max = 444, n_tokens = 445, size = 62.813 MiB)
|
| 202 |
+
0.43.355.441 I slot print_timing: id 0 | task 520 |
|
| 203 |
+
prompt eval time = 549.58 ms / 449 tokens ( 1.22 ms per token, 816.99 tokens per second)
|
| 204 |
+
eval time = 1727.31 ms / 86 tokens ( 20.09 ms per token, 49.79 tokens per second)
|
| 205 |
+
total time = 2276.89 ms / 535 tokens
|
| 206 |
+
0.43.355.528 I slot release: id 0 | task 520 | stop processing: n_tokens = 534, truncated = 0
|
| 207 |
+
0.43.355.559 I srv update_slots: all slots are idle
|
| 208 |
+
0.43.407.484 I srv params_from_: Chat format: peg-native
|
| 209 |
+
0.43.408.099 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.948 (> 0.100 thold), f_keep = 0.758
|
| 210 |
+
0.43.408.372 I reasoning-budget: activated, budget=2147483647 tokens
|
| 211 |
+
0.43.408.374 I reasoning-budget: deactivated (natural end)
|
| 212 |
+
0.43.408.422 I slot launch_slot_: id 0 | task 608 | processing task, is_child = 0
|
| 213 |
+
0.43.408.434 W slot update_slots: id 0 | task 608 | n_past = 405, slot.prompt.tokens.size() = 534, seq_id = 0, pos_min = 533, n_swa = 0
|
| 214 |
+
0.43.408.435 I slot update_slots: id 0 | task 608 | Checking checkpoint with [444, 444] against 405...
|
| 215 |
+
0.43.408.437 W slot update_slots: id 0 | task 608 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 216 |
+
0.43.408.440 W slot update_slots: id 0 | task 608 | erased invalidated context checkpoint (pos_min = 444, pos_max = 444, n_tokens = 445, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 217 |
+
0.43.970.219 I slot create_check: id 0 | task 608 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
|
| 218 |
+
0.44.791.545 I slot print_timing: id 0 | task 608 |
|
| 219 |
+
prompt eval time = 601.11 ms / 427 tokens ( 1.41 ms per token, 710.35 tokens per second)
|
| 220 |
+
eval time = 781.99 ms / 39 tokens ( 20.05 ms per token, 49.87 tokens per second)
|
| 221 |
+
total time = 1383.10 ms / 466 tokens
|
| 222 |
+
0.44.791.624 I slot release: id 0 | task 608 | stop processing: n_tokens = 465, truncated = 0
|
| 223 |
+
0.44.791.654 I srv update_slots: all slots are idle
|
| 224 |
+
0.44.833.558 I srv params_from_: Chat format: peg-native
|
| 225 |
+
0.44.834.049 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.948 (> 0.100 thold), f_keep = 0.871
|
| 226 |
+
0.44.834.271 I reasoning-budget: activated, budget=2147483647 tokens
|
| 227 |
+
0.44.834.272 I reasoning-budget: deactivated (natural end)
|
| 228 |
+
0.44.834.315 I slot launch_slot_: id 0 | task 649 | processing task, is_child = 0
|
| 229 |
+
0.44.834.326 W slot update_slots: id 0 | task 649 | n_past = 405, slot.prompt.tokens.size() = 465, seq_id = 0, pos_min = 464, n_swa = 0
|
| 230 |
+
0.44.834.327 I slot update_slots: id 0 | task 649 | Checking checkpoint with [422, 422] against 405...
|
| 231 |
+
0.44.834.328 W slot update_slots: id 0 | task 649 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 232 |
+
0.44.834.332 W slot update_slots: id 0 | task 649 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 233 |
+
0.45.316.040 I slot create_check: id 0 | task 649 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
|
| 234 |
+
0.45.454.550 I slot print_timing: id 0 | task 649 |
|
| 235 |
+
prompt eval time = 522.74 ms / 427 tokens ( 1.22 ms per token, 816.85 tokens per second)
|
| 236 |
+
eval time = 97.47 ms / 4 tokens ( 24.37 ms per token, 41.04 tokens per second)
|
| 237 |
+
total time = 620.21 ms / 431 tokens
|
| 238 |
+
0.45.454.641 I slot release: id 0 | task 649 | stop processing: n_tokens = 430, truncated = 0
|
| 239 |
+
0.45.454.670 I srv update_slots: all slots are idle
|
| 240 |
+
0.45.491.620 I srv params_from_: Chat format: peg-native
|
| 241 |
+
0.45.492.052 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.958 (> 0.100 thold), f_keep = 0.944
|
| 242 |
+
0.45.492.301 I reasoning-budget: activated, budget=2147483647 tokens
|
| 243 |
+
0.45.492.303 I reasoning-budget: deactivated (natural end)
|
| 244 |
+
0.45.492.343 I slot launch_slot_: id 0 | task 655 | processing task, is_child = 0
|
| 245 |
+
0.45.492.355 W slot update_slots: id 0 | task 655 | n_past = 406, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
|
| 246 |
+
0.45.492.357 I slot update_slots: id 0 | task 655 | Checking checkpoint with [422, 422] against 406...
|
| 247 |
+
0.45.492.358 W slot update_slots: id 0 | task 655 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 248 |
+
0.45.492.362 W slot update_slots: id 0 | task 655 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 249 |
+
0.46.038.260 I slot create_check: id 0 | task 655 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
|
| 250 |
+
0.46.950.011 I slot print_timing: id 0 | task 655 |
|
| 251 |
+
prompt eval time = 592.07 ms / 424 tokens ( 1.40 ms per token, 716.14 tokens per second)
|
| 252 |
+
eval time = 865.56 ms / 40 tokens ( 21.64 ms per token, 46.21 tokens per second)
|
| 253 |
+
total time = 1457.62 ms / 464 tokens
|
| 254 |
+
0.46.950.215 I slot release: id 0 | task 655 | stop processing: n_tokens = 463, truncated = 0
|
| 255 |
+
0.46.950.270 I srv update_slots: all slots are idle
|
| 256 |
+
0.46.968.851 I srv params_from_: Chat format: peg-native
|
| 257 |
+
0.46.969.311 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.935 (> 0.100 thold), f_keep = 1.000
|
| 258 |
+
0.46.969.946 I reasoning-budget: activated, budget=2147483647 tokens
|
| 259 |
+
0.46.969.951 I reasoning-budget: deactivated (natural end)
|
| 260 |
+
0.46.970.047 I slot launch_slot_: id 0 | task 697 | processing task, is_child = 0
|
| 261 |
+
0.47.121.064 I slot create_check: id 0 | task 697 | created context checkpoint 2 of 32 (pos_min = 490, pos_max = 490, n_tokens = 491, size = 62.813 MiB)
|
| 262 |
+
0.47.570.159 I slot print_timing: id 0 | task 697 |
|
| 263 |
+
prompt eval time = 211.63 ms / 32 tokens ( 6.61 ms per token, 151.21 tokens per second)
|
| 264 |
+
eval time = 388.43 ms / 15 tokens ( 25.90 ms per token, 38.62 tokens per second)
|
| 265 |
+
total time = 600.06 ms / 47 tokens
|
| 266 |
+
0.47.570.388 I slot release: id 0 | task 697 | stop processing: n_tokens = 509, truncated = 0
|
| 267 |
+
0.47.570.448 I srv update_slots: all slots are idle
|
| 268 |
+
0.47.607.526 I srv params_from_: Chat format: peg-native
|
| 269 |
+
0.47.608.040 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.967 (> 0.100 thold), f_keep = 0.806
|
| 270 |
+
0.47.608.345 I reasoning-budget: activated, budget=2147483647 tokens
|
| 271 |
+
0.47.608.347 I reasoning-budget: deactivated (natural end)
|
| 272 |
+
0.47.608.398 I slot launch_slot_: id 0 | task 714 | processing task, is_child = 0
|
| 273 |
+
0.47.608.410 W slot update_slots: id 0 | task 714 | n_past = 410, slot.prompt.tokens.size() = 509, seq_id = 0, pos_min = 508, n_swa = 0
|
| 274 |
+
0.47.608.412 I slot update_slots: id 0 | task 714 | Checking checkpoint with [490, 490] against 410...
|
| 275 |
+
0.47.608.412 I slot update_slots: id 0 | task 714 | Checking checkpoint with [419, 419] against 410...
|
| 276 |
+
0.47.608.413 W slot update_slots: id 0 | task 714 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 277 |
+
0.47.608.416 W slot update_slots: id 0 | task 714 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 278 |
+
0.47.609.303 W slot update_slots: id 0 | task 714 | erased invalidated context checkpoint (pos_min = 490, pos_max = 490, n_tokens = 491, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 279 |
+
0.48.271.002 I slot create_check: id 0 | task 714 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
|
| 280 |
+
0.49.108.254 I slot print_timing: id 0 | task 714 |
|
| 281 |
+
prompt eval time = 731.88 ms / 424 tokens ( 1.73 ms per token, 579.33 tokens per second)
|
| 282 |
+
eval time = 767.92 ms / 40 tokens ( 19.20 ms per token, 52.09 tokens per second)
|
| 283 |
+
total time = 1499.80 ms / 464 tokens
|
| 284 |
+
0.49.108.437 I slot release: id 0 | task 714 | stop processing: n_tokens = 463, truncated = 0
|
| 285 |
+
0.49.108.493 I srv update_slots: all slots are idle
|
| 286 |
+
0.49.161.163 I srv params_from_: Chat format: peg-native
|
| 287 |
+
0.49.163.186 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.931 (> 0.100 thold), f_keep = 0.875
|
| 288 |
+
0.49.163.878 I reasoning-budget: activated, budget=2147483647 tokens
|
| 289 |
+
0.49.163.884 I reasoning-budget: deactivated (natural end)
|
| 290 |
+
0.49.163.990 I slot launch_slot_: id 0 | task 756 | processing task, is_child = 0
|
| 291 |
+
0.49.164.016 W slot update_slots: id 0 | task 756 | n_past = 405, slot.prompt.tokens.size() = 463, seq_id = 0, pos_min = 462, n_swa = 0
|
| 292 |
+
0.49.164.018 I slot update_slots: id 0 | task 756 | Checking checkpoint with [419, 419] against 405...
|
| 293 |
+
0.49.164.020 W slot update_slots: id 0 | task 756 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 294 |
+
0.49.164.027 W slot update_slots: id 0 | task 756 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 295 |
+
0.49.744.671 I slot create_check: id 0 | task 756 | created context checkpoint 1 of 32 (pos_min = 430, pos_max = 430, n_tokens = 431, size = 62.813 MiB)
|
| 296 |
+
0.51.381.802 I slot print_timing: id 0 | task 756 |
|
| 297 |
+
prompt eval time = 632.00 ms / 435 tokens ( 1.45 ms per token, 688.30 tokens per second)
|
| 298 |
+
eval time = 1585.77 ms / 80 tokens ( 19.82 ms per token, 50.45 tokens per second)
|
| 299 |
+
total time = 2217.77 ms / 515 tokens
|
| 300 |
+
0.51.382.017 I slot release: id 0 | task 756 | stop processing: n_tokens = 514, truncated = 0
|
| 301 |
+
0.51.382.050 I srv update_slots: all slots are idle
|
| 302 |
+
0.51.383.517 I srv operator(): operator(): cleaning up before exit...
|
recipe/logs/b_n-tools-q106-roff-r3.log
ADDED
|
@@ -0,0 +1,302 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
0.00.114.439 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
|
| 2 |
+
0.00.114.444 I device_info:
|
| 3 |
+
0.00.114.555 I - ROCm0 : AMD Radeon Graphics (131072 MiB, 123723 MiB free)
|
| 4 |
+
0.00.114.708 I - Vulkan0 : AMD Radeon Graphics (RADV GFX1151) (132096 MiB, 131922 MiB free)
|
| 5 |
+
0.00.114.715 I - CPU : AMD RYZEN AI MAX+ 395 w/ Radeon 8060S (127438 MiB, 127438 MiB free)
|
| 6 |
+
0.00.114.796 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
|
| 7 |
+
0.00.114.847 I srv init: running without SSL
|
| 8 |
+
0.00.114.871 I srv init: using 31 threads for HTTP server
|
| 9 |
+
0.00.114.873 I srv init: the WebUI is disabled
|
| 10 |
+
0.00.114.942 I srv start: binding port with default address family
|
| 11 |
+
0.00.116.175 I srv main: loading model
|
| 12 |
+
0.00.116.182 I srv load_model: loading model '/mnt/models/nex-n2.5-mini/out/Nex-N2.5-mini-Q4_0_ROCMFP4_STRIX_LEAN.gguf'
|
| 13 |
+
0.00.160.730 W llama_model_loader: direct I/O is enabled, disabling mmap
|
| 14 |
+
0.21.513.834 W llama_context: n_ctx_seq (65536) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
|
| 15 |
+
0.21.793.651 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|
| 16 |
+
0.22.217.020 I srv load_model: initializing slots, n_slots = 1
|
| 17 |
+
0.22.487.377 W srv load_model: speculative decoding will use checkpoints
|
| 18 |
+
0.22.487.387 W common_speculative_init: no implementations specified for speculative decoding
|
| 19 |
+
0.22.487.390 I slot load_model: id 0 | task -1 | new slot, n_ctx = 65536
|
| 20 |
+
0.22.487.471 I srv load_model: prompt cache RAM enabled: limit_mib=8192
|
| 21 |
+
0.22.487.473 I srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
|
| 22 |
+
0.22.487.495 I srv init: idle slots will be saved to prompt cache upon starting a new task
|
| 23 |
+
0.22.540.918 I init: chat template, example_format: '<|im_start|>system
|
| 24 |
+
You are a helpful assistant<|im_end|>
|
| 25 |
+
<|im_start|>user
|
| 26 |
+
Hello<|im_end|>
|
| 27 |
+
<|im_start|>assistant
|
| 28 |
+
<think>
|
| 29 |
+
|
| 30 |
+
</think>
|
| 31 |
+
|
| 32 |
+
Hi there<|im_end|>
|
| 33 |
+
<|im_start|>user
|
| 34 |
+
How are you?<|im_end|>
|
| 35 |
+
<|im_start|>assistant
|
| 36 |
+
<think>
|
| 37 |
+
|
| 38 |
+
</think>
|
| 39 |
+
|
| 40 |
+
'
|
| 41 |
+
0.22.582.904 I srv init: init: chat template, thinking = 0
|
| 42 |
+
0.22.582.984 I srv main: model loaded
|
| 43 |
+
0.22.582.995 I srv main: server is listening on http://127.0.0.1:18600
|
| 44 |
+
0.22.583.003 I srv update_slots: all slots are idle
|
| 45 |
+
0.23.994.344 I srv params_from_: Chat format: peg-native
|
| 46 |
+
0.23.995.890 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1
|
| 47 |
+
0.23.995.893 I srv get_availabl: updating prompt cache
|
| 48 |
+
0.23.995.900 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000
|
| 49 |
+
0.23.995.907 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 65536 tokens, 8589934592 est)
|
| 50 |
+
0.23.995.909 I srv get_availabl: prompt cache update took 0.01 ms
|
| 51 |
+
0.23.996.222 I reasoning-budget: activated, budget=2147483647 tokens
|
| 52 |
+
0.23.996.243 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0
|
| 53 |
+
0.24.656.206 I slot create_check: id 0 | task 0 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
|
| 54 |
+
0.25.028.892 I reasoning-budget: deactivated (natural end)
|
| 55 |
+
0.25.816.313 I slot print_timing: id 0 | task 0 |
|
| 56 |
+
prompt eval time = 698.88 ms / 422 tokens ( 1.66 ms per token, 603.83 tokens per second)
|
| 57 |
+
eval time = 1121.16 ms / 54 tokens ( 20.76 ms per token, 48.16 tokens per second)
|
| 58 |
+
total time = 1820.04 ms / 476 tokens
|
| 59 |
+
0.25.816.410 I slot release: id 0 | task 0 | stop processing: n_tokens = 475, truncated = 0
|
| 60 |
+
0.25.816.424 I srv update_slots: all slots are idle
|
| 61 |
+
0.25.830.282 I srv params_from_: Chat format: peg-native
|
| 62 |
+
0.25.830.661 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.906 (> 0.100 thold), f_keep = 0.853
|
| 63 |
+
0.25.830.794 I reasoning-budget: activated, budget=2147483647 tokens
|
| 64 |
+
0.25.830.832 I slot launch_slot_: id 0 | task 56 | processing task, is_child = 0
|
| 65 |
+
0.25.830.841 W slot update_slots: id 0 | task 56 | n_past = 405, slot.prompt.tokens.size() = 475, seq_id = 0, pos_min = 474, n_swa = 0
|
| 66 |
+
0.25.830.842 I slot update_slots: id 0 | task 56 | Checking checkpoint with [417, 417] against 405...
|
| 67 |
+
0.25.830.844 W slot update_slots: id 0 | task 56 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 68 |
+
0.25.830.846 W slot update_slots: id 0 | task 56 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 69 |
+
0.26.482.806 I slot create_check: id 0 | task 56 | created context checkpoint 1 of 32 (pos_min = 442, pos_max = 442, n_tokens = 443, size = 62.813 MiB)
|
| 70 |
+
0.26.903.330 I reasoning-budget: deactivated (natural end)
|
| 71 |
+
0.28.669.262 I slot print_timing: id 0 | task 56 | n_decoded = 100, tg = 47.24 t/s
|
| 72 |
+
0.28.729.150 I slot print_timing: id 0 | task 56 |
|
| 73 |
+
prompt eval time = 721.52 ms / 447 tokens ( 1.61 ms per token, 619.52 tokens per second)
|
| 74 |
+
eval time = 2176.78 ms / 103 tokens ( 21.13 ms per token, 47.32 tokens per second)
|
| 75 |
+
total time = 2898.30 ms / 550 tokens
|
| 76 |
+
0.28.729.226 I slot release: id 0 | task 56 | stop processing: n_tokens = 549, truncated = 0
|
| 77 |
+
0.28.729.266 I srv update_slots: all slots are idle
|
| 78 |
+
0.28.742.730 I srv params_from_: Chat format: peg-native
|
| 79 |
+
0.28.743.168 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.953 (> 0.100 thold), f_keep = 0.738
|
| 80 |
+
0.28.743.356 I reasoning-budget: activated, budget=2147483647 tokens
|
| 81 |
+
0.28.743.392 I slot launch_slot_: id 0 | task 161 | processing task, is_child = 0
|
| 82 |
+
0.28.743.406 W slot update_slots: id 0 | task 161 | n_past = 405, slot.prompt.tokens.size() = 549, seq_id = 0, pos_min = 548, n_swa = 0
|
| 83 |
+
0.28.743.407 I slot update_slots: id 0 | task 161 | Checking checkpoint with [442, 442] against 405...
|
| 84 |
+
0.28.743.407 W slot update_slots: id 0 | task 161 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 85 |
+
0.28.743.411 W slot update_slots: id 0 | task 161 | erased invalidated context checkpoint (pos_min = 442, pos_max = 442, n_tokens = 443, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 86 |
+
0.29.274.427 I slot create_check: id 0 | task 161 | created context checkpoint 1 of 32 (pos_min = 420, pos_max = 420, n_tokens = 421, size = 62.813 MiB)
|
| 87 |
+
0.29.681.788 I reasoning-budget: deactivated (natural end)
|
| 88 |
+
0.30.465.485 I slot print_timing: id 0 | task 161 |
|
| 89 |
+
prompt eval time = 568.14 ms / 425 tokens ( 1.34 ms per token, 748.05 tokens per second)
|
| 90 |
+
eval time = 1153.92 ms / 57 tokens ( 20.24 ms per token, 49.40 tokens per second)
|
| 91 |
+
total time = 1722.06 ms / 482 tokens
|
| 92 |
+
0.30.465.571 I slot release: id 0 | task 161 | stop processing: n_tokens = 481, truncated = 0
|
| 93 |
+
0.30.465.599 I srv update_slots: all slots are idle
|
| 94 |
+
0.30.480.783 I srv params_from_: Chat format: peg-native
|
| 95 |
+
0.30.481.142 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.953 (> 0.100 thold), f_keep = 0.842
|
| 96 |
+
0.30.481.346 I reasoning-budget: activated, budget=2147483647 tokens
|
| 97 |
+
0.30.481.386 I slot launch_slot_: id 0 | task 220 | processing task, is_child = 0
|
| 98 |
+
0.30.481.397 W slot update_slots: id 0 | task 220 | n_past = 405, slot.prompt.tokens.size() = 481, seq_id = 0, pos_min = 480, n_swa = 0
|
| 99 |
+
0.30.481.397 I slot update_slots: id 0 | task 220 | Checking checkpoint with [420, 420] against 405...
|
| 100 |
+
0.30.481.399 W slot update_slots: id 0 | task 220 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 101 |
+
0.30.481.401 W slot update_slots: id 0 | task 220 | erased invalidated context checkpoint (pos_min = 420, pos_max = 420, n_tokens = 421, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 102 |
+
0.31.008.867 I slot create_check: id 0 | task 220 | created context checkpoint 1 of 32 (pos_min = 420, pos_max = 420, n_tokens = 421, size = 62.813 MiB)
|
| 103 |
+
0.31.355.463 I reasoning-budget: deactivated (natural end)
|
| 104 |
+
0.31.458.677 I slot print_timing: id 0 | task 220 |
|
| 105 |
+
prompt eval time = 568.77 ms / 425 tokens ( 1.34 ms per token, 747.23 tokens per second)
|
| 106 |
+
eval time = 408.50 ms / 20 tokens ( 20.42 ms per token, 48.96 tokens per second)
|
| 107 |
+
total time = 977.27 ms / 445 tokens
|
| 108 |
+
0.31.458.774 I slot release: id 0 | task 220 | stop processing: n_tokens = 444, truncated = 0
|
| 109 |
+
0.31.458.805 I srv update_slots: all slots are idle
|
| 110 |
+
0.31.510.659 I srv params_from_: Chat format: peg-native
|
| 111 |
+
0.31.511.238 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.914
|
| 112 |
+
0.31.511.571 I reasoning-budget: activated, budget=2147483647 tokens
|
| 113 |
+
0.31.511.663 I slot launch_slot_: id 0 | task 242 | processing task, is_child = 0
|
| 114 |
+
0.31.511.677 W slot update_slots: id 0 | task 242 | n_past = 406, slot.prompt.tokens.size() = 444, seq_id = 0, pos_min = 443, n_swa = 0
|
| 115 |
+
0.31.511.684 I slot update_slots: id 0 | task 242 | Checking checkpoint with [420, 420] against 406...
|
| 116 |
+
0.31.511.685 W slot update_slots: id 0 | task 242 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 117 |
+
0.31.511.690 W slot update_slots: id 0 | task 242 | erased invalidated context checkpoint (pos_min = 420, pos_max = 420, n_tokens = 421, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 118 |
+
0.32.149.342 I slot create_check: id 0 | task 242 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
|
| 119 |
+
0.32.405.933 I reasoning-budget: deactivated (natural end)
|
| 120 |
+
0.33.283.206 I slot print_timing: id 0 | task 242 |
|
| 121 |
+
prompt eval time = 677.52 ms / 422 tokens ( 1.61 ms per token, 622.86 tokens per second)
|
| 122 |
+
eval time = 1093.99 ms / 52 tokens ( 21.04 ms per token, 47.53 tokens per second)
|
| 123 |
+
total time = 1771.51 ms / 474 tokens
|
| 124 |
+
0.33.283.283 I slot release: id 0 | task 242 | stop processing: n_tokens = 473, truncated = 0
|
| 125 |
+
0.33.283.309 I srv update_slots: all slots are idle
|
| 126 |
+
0.33.311.488 I srv params_from_: Chat format: peg-native
|
| 127 |
+
0.33.312.015 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.854 (> 0.100 thold), f_keep = 0.890
|
| 128 |
+
0.33.312.441 I reasoning-budget: activated, budget=2147483647 tokens
|
| 129 |
+
0.33.312.556 I slot launch_slot_: id 0 | task 296 | processing task, is_child = 0
|
| 130 |
+
0.33.312.580 W slot update_slots: id 0 | task 296 | n_past = 421, slot.prompt.tokens.size() = 473, seq_id = 0, pos_min = 472, n_swa = 0
|
| 131 |
+
0.33.312.583 I slot update_slots: id 0 | task 296 | Checking checkpoint with [417, 417] against 421...
|
| 132 |
+
0.33.321.205 W slot update_slots: id 0 | task 296 | restored context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_past = 418, size = 62.813 MiB)
|
| 133 |
+
0.33.558.612 I slot create_check: id 0 | task 296 | created context checkpoint 2 of 32 (pos_min = 488, pos_max = 488, n_tokens = 489, size = 62.813 MiB)
|
| 134 |
+
0.34.218.104 I reasoning-budget: deactivated (natural end)
|
| 135 |
+
0.34.556.664 I slot print_timing: id 0 | task 296 |
|
| 136 |
+
prompt eval time = 286.32 ms / 75 tokens ( 3.82 ms per token, 261.95 tokens per second)
|
| 137 |
+
eval time = 957.75 ms / 41 tokens ( 23.36 ms per token, 42.81 tokens per second)
|
| 138 |
+
total time = 1244.07 ms / 116 tokens
|
| 139 |
+
0.34.556.763 I slot release: id 0 | task 296 | stop processing: n_tokens = 533, truncated = 0
|
| 140 |
+
0.34.556.805 I srv update_slots: all slots are idle
|
| 141 |
+
0.34.571.001 I srv params_from_: Chat format: peg-native
|
| 142 |
+
0.34.571.465 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.972 (> 0.100 thold), f_keep = 0.769
|
| 143 |
+
0.34.571.696 I reasoning-budget: activated, budget=2147483647 tokens
|
| 144 |
+
0.34.571.740 I slot launch_slot_: id 0 | task 339 | processing task, is_child = 0
|
| 145 |
+
0.34.571.755 W slot update_slots: id 0 | task 339 | n_past = 410, slot.prompt.tokens.size() = 533, seq_id = 0, pos_min = 532, n_swa = 0
|
| 146 |
+
0.34.571.756 I slot update_slots: id 0 | task 339 | Checking checkpoint with [488, 488] against 410...
|
| 147 |
+
0.34.571.758 I slot update_slots: id 0 | task 339 | Checking checkpoint with [417, 417] against 410...
|
| 148 |
+
0.34.571.759 W slot update_slots: id 0 | task 339 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 149 |
+
0.34.571.763 W slot update_slots: id 0 | task 339 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 150 |
+
0.34.573.311 W slot update_slots: id 0 | task 339 | erased invalidated context checkpoint (pos_min = 488, pos_max = 488, n_tokens = 489, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 151 |
+
0.35.110.746 I slot create_check: id 0 | task 339 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
|
| 152 |
+
0.35.584.142 I reasoning-budget: deactivated (natural end)
|
| 153 |
+
0.36.392.408 I slot print_timing: id 0 | task 339 |
|
| 154 |
+
prompt eval time = 576.37 ms / 422 tokens ( 1.37 ms per token, 732.16 tokens per second)
|
| 155 |
+
eval time = 1244.26 ms / 62 tokens ( 20.07 ms per token, 49.83 tokens per second)
|
| 156 |
+
total time = 1820.64 ms / 484 tokens
|
| 157 |
+
0.36.392.474 I slot release: id 0 | task 339 | stop processing: n_tokens = 483, truncated = 0
|
| 158 |
+
0.36.392.500 I srv update_slots: all slots are idle
|
| 159 |
+
0.36.404.680 I srv params_from_: Chat format: peg-native
|
| 160 |
+
0.36.405.094 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.935 (> 0.100 thold), f_keep = 0.839
|
| 161 |
+
0.36.405.518 I reasoning-budget: activated, budget=2147483647 tokens
|
| 162 |
+
0.36.405.599 I slot launch_slot_: id 0 | task 403 | processing task, is_child = 0
|
| 163 |
+
0.36.405.620 W slot update_slots: id 0 | task 403 | n_past = 405, slot.prompt.tokens.size() = 483, seq_id = 0, pos_min = 482, n_swa = 0
|
| 164 |
+
0.36.405.621 I slot update_slots: id 0 | task 403 | Checking checkpoint with [417, 417] against 405...
|
| 165 |
+
0.36.405.625 W slot update_slots: id 0 | task 403 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 166 |
+
0.36.405.630 W slot update_slots: id 0 | task 403 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 167 |
+
0.36.974.524 I slot create_check: id 0 | task 403 | created context checkpoint 1 of 32 (pos_min = 428, pos_max = 428, n_tokens = 429, size = 62.813 MiB)
|
| 168 |
+
0.39.070.997 I slot print_timing: id 0 | task 403 | n_decoded = 100, tg = 48.88 t/s
|
| 169 |
+
0.39.448.467 I reasoning-budget: deactivated (natural end)
|
| 170 |
+
0.41.042.559 I slot print_timing: id 0 | task 403 |
|
| 171 |
+
prompt eval time = 619.66 ms / 433 tokens ( 1.43 ms per token, 698.77 tokens per second)
|
| 172 |
+
eval time = 4017.27 ms / 200 tokens ( 20.09 ms per token, 49.79 tokens per second)
|
| 173 |
+
total time = 4636.93 ms / 633 tokens
|
| 174 |
+
0.41.042.634 I slot release: id 0 | task 403 | stop processing: n_tokens = 632, truncated = 0
|
| 175 |
+
0.41.042.665 I srv update_slots: all slots are idle
|
| 176 |
+
0.41.054.733 I srv params_from_: Chat format: peg-native
|
| 177 |
+
0.41.055.120 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.641
|
| 178 |
+
0.41.055.469 I reasoning-budget: activated, budget=2147483647 tokens
|
| 179 |
+
0.41.055.471 I reasoning-budget: deactivated (natural end)
|
| 180 |
+
0.41.055.514 I slot launch_slot_: id 0 | task 605 | processing task, is_child = 0
|
| 181 |
+
0.41.055.527 W slot update_slots: id 0 | task 605 | n_past = 405, slot.prompt.tokens.size() = 632, seq_id = 0, pos_min = 631, n_swa = 0
|
| 182 |
+
0.41.055.529 I slot update_slots: id 0 | task 605 | Checking checkpoint with [428, 428] against 405...
|
| 183 |
+
0.41.055.530 W slot update_slots: id 0 | task 605 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 184 |
+
0.41.055.533 W slot update_slots: id 0 | task 605 | erased invalidated context checkpoint (pos_min = 428, pos_max = 428, n_tokens = 429, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 185 |
+
0.41.597.061 I slot create_check: id 0 | task 605 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
|
| 186 |
+
0.42.459.956 I slot print_timing: id 0 | task 605 |
|
| 187 |
+
prompt eval time = 582.62 ms / 424 tokens ( 1.37 ms per token, 727.74 tokens per second)
|
| 188 |
+
eval time = 821.78 ms / 39 tokens ( 21.07 ms per token, 47.46 tokens per second)
|
| 189 |
+
total time = 1404.41 ms / 463 tokens
|
| 190 |
+
0.42.460.047 I slot release: id 0 | task 605 | stop processing: n_tokens = 462, truncated = 0
|
| 191 |
+
0.42.460.077 I srv update_slots: all slots are idle
|
| 192 |
+
0.42.475.720 I srv params_from_: Chat format: peg-native
|
| 193 |
+
0.42.476.233 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.902 (> 0.100 thold), f_keep = 0.877
|
| 194 |
+
0.42.476.486 I reasoning-budget: activated, budget=2147483647 tokens
|
| 195 |
+
0.42.476.491 I reasoning-budget: deactivated (natural end)
|
| 196 |
+
0.42.476.536 I slot launch_slot_: id 0 | task 646 | processing task, is_child = 0
|
| 197 |
+
0.42.476.550 W slot update_slots: id 0 | task 646 | n_past = 405, slot.prompt.tokens.size() = 462, seq_id = 0, pos_min = 461, n_swa = 0
|
| 198 |
+
0.42.476.551 I slot update_slots: id 0 | task 646 | Checking checkpoint with [419, 419] against 405...
|
| 199 |
+
0.42.476.553 W slot update_slots: id 0 | task 646 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 200 |
+
0.42.476.571 W slot update_slots: id 0 | task 646 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 201 |
+
0.43.126.158 I slot create_check: id 0 | task 646 | created context checkpoint 1 of 32 (pos_min = 444, pos_max = 444, n_tokens = 445, size = 62.813 MiB)
|
| 202 |
+
0.44.988.511 I slot print_timing: id 0 | task 646 |
|
| 203 |
+
prompt eval time = 693.47 ms / 449 tokens ( 1.54 ms per token, 647.47 tokens per second)
|
| 204 |
+
eval time = 1818.46 ms / 86 tokens ( 21.14 ms per token, 47.29 tokens per second)
|
| 205 |
+
total time = 2511.93 ms / 535 tokens
|
| 206 |
+
0.44.988.703 I slot release: id 0 | task 646 | stop processing: n_tokens = 534, truncated = 0
|
| 207 |
+
0.44.988.760 I srv update_slots: all slots are idle
|
| 208 |
+
0.45.002.229 I srv params_from_: Chat format: peg-native
|
| 209 |
+
0.45.002.631 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.948 (> 0.100 thold), f_keep = 0.758
|
| 210 |
+
0.45.003.110 I reasoning-budget: activated, budget=2147483647 tokens
|
| 211 |
+
0.45.003.114 I reasoning-budget: deactivated (natural end)
|
| 212 |
+
0.45.003.190 I slot launch_slot_: id 0 | task 734 | processing task, is_child = 0
|
| 213 |
+
0.45.003.216 W slot update_slots: id 0 | task 734 | n_past = 405, slot.prompt.tokens.size() = 534, seq_id = 0, pos_min = 533, n_swa = 0
|
| 214 |
+
0.45.003.220 I slot update_slots: id 0 | task 734 | Checking checkpoint with [444, 444] against 405...
|
| 215 |
+
0.45.003.221 W slot update_slots: id 0 | task 734 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 216 |
+
0.45.003.227 W slot update_slots: id 0 | task 734 | erased invalidated context checkpoint (pos_min = 444, pos_max = 444, n_tokens = 445, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 217 |
+
0.45.561.942 I slot create_check: id 0 | task 734 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
|
| 218 |
+
0.46.436.833 I slot print_timing: id 0 | task 734 |
|
| 219 |
+
prompt eval time = 595.84 ms / 427 tokens ( 1.40 ms per token, 716.64 tokens per second)
|
| 220 |
+
eval time = 837.77 ms / 39 tokens ( 21.48 ms per token, 46.55 tokens per second)
|
| 221 |
+
total time = 1433.60 ms / 466 tokens
|
| 222 |
+
0.46.436.923 I slot release: id 0 | task 734 | stop processing: n_tokens = 465, truncated = 0
|
| 223 |
+
0.46.436.955 I srv update_slots: all slots are idle
|
| 224 |
+
0.46.450.626 I srv params_from_: Chat format: peg-native
|
| 225 |
+
0.46.451.043 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.948 (> 0.100 thold), f_keep = 0.871
|
| 226 |
+
0.46.451.527 I reasoning-budget: activated, budget=2147483647 tokens
|
| 227 |
+
0.46.451.532 I reasoning-budget: deactivated (natural end)
|
| 228 |
+
0.46.451.617 I slot launch_slot_: id 0 | task 775 | processing task, is_child = 0
|
| 229 |
+
0.46.451.637 W slot update_slots: id 0 | task 775 | n_past = 405, slot.prompt.tokens.size() = 465, seq_id = 0, pos_min = 464, n_swa = 0
|
| 230 |
+
0.46.451.641 I slot update_slots: id 0 | task 775 | Checking checkpoint with [422, 422] against 405...
|
| 231 |
+
0.46.451.643 W slot update_slots: id 0 | task 775 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 232 |
+
0.46.451.649 W slot update_slots: id 0 | task 775 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 233 |
+
0.46.982.640 I slot create_check: id 0 | task 775 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
|
| 234 |
+
0.47.142.299 I slot print_timing: id 0 | task 775 |
|
| 235 |
+
prompt eval time = 585.65 ms / 427 tokens ( 1.37 ms per token, 729.10 tokens per second)
|
| 236 |
+
eval time = 104.99 ms / 4 tokens ( 26.25 ms per token, 38.10 tokens per second)
|
| 237 |
+
total time = 690.65 ms / 431 tokens
|
| 238 |
+
0.47.142.414 I slot release: id 0 | task 775 | stop processing: n_tokens = 430, truncated = 0
|
| 239 |
+
0.47.142.449 I srv update_slots: all slots are idle
|
| 240 |
+
0.47.193.526 I srv params_from_: Chat format: peg-native
|
| 241 |
+
0.47.195.446 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.958 (> 0.100 thold), f_keep = 0.944
|
| 242 |
+
0.47.195.987 I reasoning-budget: activated, budget=2147483647 tokens
|
| 243 |
+
0.47.195.992 I reasoning-budget: deactivated (natural end)
|
| 244 |
+
0.47.196.101 I slot launch_slot_: id 0 | task 781 | processing task, is_child = 0
|
| 245 |
+
0.47.196.122 W slot update_slots: id 0 | task 781 | n_past = 406, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
|
| 246 |
+
0.47.196.128 I slot update_slots: id 0 | task 781 | Checking checkpoint with [422, 422] against 406...
|
| 247 |
+
0.47.196.129 W slot update_slots: id 0 | task 781 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 248 |
+
0.47.196.139 W slot update_slots: id 0 | task 781 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 249 |
+
0.47.774.566 I slot create_check: id 0 | task 781 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
|
| 250 |
+
0.48.760.957 I slot print_timing: id 0 | task 781 |
|
| 251 |
+
prompt eval time = 635.18 ms / 424 tokens ( 1.50 ms per token, 667.53 tokens per second)
|
| 252 |
+
eval time = 929.64 ms / 40 tokens ( 23.24 ms per token, 43.03 tokens per second)
|
| 253 |
+
total time = 1564.83 ms / 464 tokens
|
| 254 |
+
0.48.761.052 I slot release: id 0 | task 781 | stop processing: n_tokens = 463, truncated = 0
|
| 255 |
+
0.48.761.080 I srv update_slots: all slots are idle
|
| 256 |
+
0.48.783.508 I srv params_from_: Chat format: peg-native
|
| 257 |
+
0.48.783.860 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.935 (> 0.100 thold), f_keep = 1.000
|
| 258 |
+
0.48.784.082 I reasoning-budget: activated, budget=2147483647 tokens
|
| 259 |
+
0.48.784.084 I reasoning-budget: deactivated (natural end)
|
| 260 |
+
0.48.784.127 I slot launch_slot_: id 0 | task 823 | processing task, is_child = 0
|
| 261 |
+
0.48.935.107 I slot create_check: id 0 | task 823 | created context checkpoint 2 of 32 (pos_min = 490, pos_max = 490, n_tokens = 491, size = 62.813 MiB)
|
| 262 |
+
0.49.277.790 I slot print_timing: id 0 | task 823 |
|
| 263 |
+
prompt eval time = 217.60 ms / 32 tokens ( 6.80 ms per token, 147.06 tokens per second)
|
| 264 |
+
eval time = 276.04 ms / 13 tokens ( 21.23 ms per token, 47.09 tokens per second)
|
| 265 |
+
total time = 493.64 ms / 45 tokens
|
| 266 |
+
0.49.277.883 I slot release: id 0 | task 823 | stop processing: n_tokens = 507, truncated = 0
|
| 267 |
+
0.49.277.910 I srv update_slots: all slots are idle
|
| 268 |
+
0.49.324.008 I srv params_from_: Chat format: peg-native
|
| 269 |
+
0.49.324.548 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.967 (> 0.100 thold), f_keep = 0.809
|
| 270 |
+
0.49.324.947 I reasoning-budget: activated, budget=2147483647 tokens
|
| 271 |
+
0.49.324.952 I reasoning-budget: deactivated (natural end)
|
| 272 |
+
0.49.325.029 I slot launch_slot_: id 0 | task 838 | processing task, is_child = 0
|
| 273 |
+
0.49.325.041 W slot update_slots: id 0 | task 838 | n_past = 410, slot.prompt.tokens.size() = 507, seq_id = 0, pos_min = 506, n_swa = 0
|
| 274 |
+
0.49.325.042 I slot update_slots: id 0 | task 838 | Checking checkpoint with [490, 490] against 410...
|
| 275 |
+
0.49.325.043 I slot update_slots: id 0 | task 838 | Checking checkpoint with [419, 419] against 410...
|
| 276 |
+
0.49.325.048 W slot update_slots: id 0 | task 838 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 277 |
+
0.49.325.051 W slot update_slots: id 0 | task 838 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 278 |
+
0.49.328.874 W slot update_slots: id 0 | task 838 | erased invalidated context checkpoint (pos_min = 490, pos_max = 490, n_tokens = 491, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 279 |
+
0.49.907.729 I slot create_check: id 0 | task 838 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
|
| 280 |
+
0.50.729.586 I slot print_timing: id 0 | task 838 |
|
| 281 |
+
prompt eval time = 624.90 ms / 424 tokens ( 1.47 ms per token, 678.50 tokens per second)
|
| 282 |
+
eval time = 779.62 ms / 40 tokens ( 19.49 ms per token, 51.31 tokens per second)
|
| 283 |
+
total time = 1404.52 ms / 464 tokens
|
| 284 |
+
0.50.729.654 I slot release: id 0 | task 838 | stop processing: n_tokens = 463, truncated = 0
|
| 285 |
+
0.50.729.686 I srv update_slots: all slots are idle
|
| 286 |
+
0.50.775.784 I srv params_from_: Chat format: peg-native
|
| 287 |
+
0.50.776.334 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.931 (> 0.100 thold), f_keep = 0.875
|
| 288 |
+
0.50.776.559 I reasoning-budget: activated, budget=2147483647 tokens
|
| 289 |
+
0.50.776.562 I reasoning-budget: deactivated (natural end)
|
| 290 |
+
0.50.776.609 I slot launch_slot_: id 0 | task 880 | processing task, is_child = 0
|
| 291 |
+
0.50.776.620 W slot update_slots: id 0 | task 880 | n_past = 405, slot.prompt.tokens.size() = 463, seq_id = 0, pos_min = 462, n_swa = 0
|
| 292 |
+
0.50.776.622 I slot update_slots: id 0 | task 880 | Checking checkpoint with [419, 419] against 405...
|
| 293 |
+
0.50.776.623 W slot update_slots: id 0 | task 880 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 294 |
+
0.50.776.627 W slot update_slots: id 0 | task 880 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 295 |
+
0.51.328.909 I slot create_check: id 0 | task 880 | created context checkpoint 1 of 32 (pos_min = 430, pos_max = 430, n_tokens = 431, size = 62.813 MiB)
|
| 296 |
+
0.52.962.266 I slot print_timing: id 0 | task 880 |
|
| 297 |
+
prompt eval time = 596.05 ms / 435 tokens ( 1.37 ms per token, 729.81 tokens per second)
|
| 298 |
+
eval time = 1589.57 ms / 80 tokens ( 19.87 ms per token, 50.33 tokens per second)
|
| 299 |
+
total time = 2185.62 ms / 515 tokens
|
| 300 |
+
0.52.962.487 I slot release: id 0 | task 880 | stop processing: n_tokens = 514, truncated = 0
|
| 301 |
+
0.52.962.523 I srv update_slots: all slots are idle
|
| 302 |
+
0.52.963.931 I srv operator(): operator(): cleaning up before exit...
|
recipe/logs/b_n-tools-q106-roff.log
ADDED
|
@@ -0,0 +1,301 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
0.00.077.092 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
|
| 2 |
+
0.00.077.102 I device_info:
|
| 3 |
+
0.00.077.242 I - ROCm0 : AMD Radeon Graphics (131072 MiB, 123819 MiB free)
|
| 4 |
+
0.00.077.530 I - Vulkan0 : AMD Radeon Graphics (RADV GFX1151) (132096 MiB, 131922 MiB free)
|
| 5 |
+
0.00.077.538 I - CPU : AMD RYZEN AI MAX+ 395 w/ Radeon 8060S (127438 MiB, 127438 MiB free)
|
| 6 |
+
0.00.077.648 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
|
| 7 |
+
0.00.077.722 I srv init: running without SSL
|
| 8 |
+
0.00.077.768 I srv init: using 31 threads for HTTP server
|
| 9 |
+
0.00.077.770 I srv init: the WebUI is disabled
|
| 10 |
+
0.00.077.863 I srv start: binding port with default address family
|
| 11 |
+
0.00.079.093 I srv main: loading model
|
| 12 |
+
0.00.079.096 I srv load_model: loading model '/mnt/models/nex-n2.5-mini/out/Nex-N2.5-mini-Q4_0_ROCMFP4_STRIX_LEAN.gguf'
|
| 13 |
+
0.00.120.315 W llama_model_loader: direct I/O is enabled, disabling mmap
|
| 14 |
+
0.21.810.861 W llama_context: n_ctx_seq (65536) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
|
| 15 |
+
0.22.004.505 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|
| 16 |
+
0.22.307.874 I srv load_model: initializing slots, n_slots = 1
|
| 17 |
+
0.22.526.550 W srv load_model: speculative decoding will use checkpoints
|
| 18 |
+
0.22.526.576 W common_speculative_init: no implementations specified for speculative decoding
|
| 19 |
+
0.22.526.584 I slot load_model: id 0 | task -1 | new slot, n_ctx = 65536
|
| 20 |
+
0.22.526.740 I srv load_model: prompt cache RAM enabled: limit_mib=8192
|
| 21 |
+
0.22.526.745 I srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
|
| 22 |
+
0.22.526.855 I srv init: idle slots will be saved to prompt cache upon starting a new task
|
| 23 |
+
0.22.576.014 I init: chat template, example_format: '<|im_start|>system
|
| 24 |
+
You are a helpful assistant<|im_end|>
|
| 25 |
+
<|im_start|>user
|
| 26 |
+
Hello<|im_end|>
|
| 27 |
+
<|im_start|>assistant
|
| 28 |
+
<think>
|
| 29 |
+
|
| 30 |
+
</think>
|
| 31 |
+
|
| 32 |
+
Hi there<|im_end|>
|
| 33 |
+
<|im_start|>user
|
| 34 |
+
How are you?<|im_end|>
|
| 35 |
+
<|im_start|>assistant
|
| 36 |
+
<think>
|
| 37 |
+
|
| 38 |
+
</think>
|
| 39 |
+
|
| 40 |
+
'
|
| 41 |
+
0.22.604.308 I srv init: init: chat template, thinking = 0
|
| 42 |
+
0.22.604.360 I srv main: model loaded
|
| 43 |
+
0.22.604.364 I srv main: server is listening on http://127.0.0.1:18600
|
| 44 |
+
0.22.604.370 I srv update_slots: all slots are idle
|
| 45 |
+
0.24.142.693 I srv params_from_: Chat format: peg-native
|
| 46 |
+
0.24.143.568 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1
|
| 47 |
+
0.24.143.571 I srv get_availabl: updating prompt cache
|
| 48 |
+
0.24.143.578 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000
|
| 49 |
+
0.24.143.583 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 65536 tokens, 8589934592 est)
|
| 50 |
+
0.24.143.585 I srv get_availabl: prompt cache update took 0.01 ms
|
| 51 |
+
0.24.143.943 I reasoning-budget: activated, budget=2147483647 tokens
|
| 52 |
+
0.24.143.962 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0
|
| 53 |
+
0.24.903.283 I slot create_check: id 0 | task 0 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
|
| 54 |
+
0.25.273.393 I reasoning-budget: deactivated (natural end)
|
| 55 |
+
0.26.099.925 I slot print_timing: id 0 | task 0 |
|
| 56 |
+
prompt eval time = 810.27 ms / 422 tokens ( 1.92 ms per token, 520.81 tokens per second)
|
| 57 |
+
eval time = 1145.63 ms / 51 tokens ( 22.46 ms per token, 44.52 tokens per second)
|
| 58 |
+
total time = 1955.91 ms / 473 tokens
|
| 59 |
+
0.26.100.146 I slot release: id 0 | task 0 | stop processing: n_tokens = 472, truncated = 0
|
| 60 |
+
0.26.100.163 I srv update_slots: all slots are idle
|
| 61 |
+
0.26.148.514 I srv params_from_: Chat format: peg-native
|
| 62 |
+
0.26.149.069 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.906 (> 0.100 thold), f_keep = 0.858
|
| 63 |
+
0.26.149.246 I reasoning-budget: activated, budget=2147483647 tokens
|
| 64 |
+
0.26.149.294 I slot launch_slot_: id 0 | task 53 | processing task, is_child = 0
|
| 65 |
+
0.26.149.305 W slot update_slots: id 0 | task 53 | n_past = 405, slot.prompt.tokens.size() = 472, seq_id = 0, pos_min = 471, n_swa = 0
|
| 66 |
+
0.26.149.306 I slot update_slots: id 0 | task 53 | Checking checkpoint with [417, 417] against 405...
|
| 67 |
+
0.26.149.307 W slot update_slots: id 0 | task 53 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 68 |
+
0.26.149.310 W slot update_slots: id 0 | task 53 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 69 |
+
0.26.723.793 I slot create_check: id 0 | task 53 | created context checkpoint 1 of 32 (pos_min = 442, pos_max = 442, n_tokens = 443, size = 62.813 MiB)
|
| 70 |
+
0.27.315.116 I reasoning-budget: deactivated (natural end)
|
| 71 |
+
0.28.126.542 I slot print_timing: id 0 | task 53 |
|
| 72 |
+
prompt eval time = 626.21 ms / 447 tokens ( 1.40 ms per token, 713.82 tokens per second)
|
| 73 |
+
eval time = 1350.95 ms / 68 tokens ( 19.87 ms per token, 50.34 tokens per second)
|
| 74 |
+
total time = 1977.16 ms / 515 tokens
|
| 75 |
+
0.28.126.747 I slot release: id 0 | task 53 | stop processing: n_tokens = 514, truncated = 0
|
| 76 |
+
0.28.126.781 I srv update_slots: all slots are idle
|
| 77 |
+
0.28.164.510 I srv params_from_: Chat format: peg-native
|
| 78 |
+
0.28.166.615 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.953 (> 0.100 thold), f_keep = 0.788
|
| 79 |
+
0.28.167.130 I reasoning-budget: activated, budget=2147483647 tokens
|
| 80 |
+
0.28.167.205 I slot launch_slot_: id 0 | task 123 | processing task, is_child = 0
|
| 81 |
+
0.28.167.226 W slot update_slots: id 0 | task 123 | n_past = 405, slot.prompt.tokens.size() = 514, seq_id = 0, pos_min = 513, n_swa = 0
|
| 82 |
+
0.28.167.229 I slot update_slots: id 0 | task 123 | Checking checkpoint with [442, 442] against 405...
|
| 83 |
+
0.28.167.232 W slot update_slots: id 0 | task 123 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 84 |
+
0.28.167.239 W slot update_slots: id 0 | task 123 | erased invalidated context checkpoint (pos_min = 442, pos_max = 442, n_tokens = 443, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 85 |
+
0.28.745.049 I slot create_check: id 0 | task 123 | created context checkpoint 1 of 32 (pos_min = 420, pos_max = 420, n_tokens = 421, size = 62.813 MiB)
|
| 86 |
+
0.28.952.307 I reasoning-budget: deactivated (natural end)
|
| 87 |
+
0.29.756.441 I slot print_timing: id 0 | task 123 |
|
| 88 |
+
prompt eval time = 613.73 ms / 425 tokens ( 1.44 ms per token, 692.49 tokens per second)
|
| 89 |
+
eval time = 975.48 ms / 48 tokens ( 20.32 ms per token, 49.21 tokens per second)
|
| 90 |
+
total time = 1589.20 ms / 473 tokens
|
| 91 |
+
0.29.756.519 I slot release: id 0 | task 123 | stop processing: n_tokens = 472, truncated = 0
|
| 92 |
+
0.29.756.549 I srv update_slots: all slots are idle
|
| 93 |
+
0.29.771.836 I srv params_from_: Chat format: peg-native
|
| 94 |
+
0.29.772.265 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.953 (> 0.100 thold), f_keep = 0.858
|
| 95 |
+
0.29.772.460 I reasoning-budget: activated, budget=2147483647 tokens
|
| 96 |
+
0.29.772.502 I slot launch_slot_: id 0 | task 173 | processing task, is_child = 0
|
| 97 |
+
0.29.772.513 W slot update_slots: id 0 | task 173 | n_past = 405, slot.prompt.tokens.size() = 472, seq_id = 0, pos_min = 471, n_swa = 0
|
| 98 |
+
0.29.772.513 I slot update_slots: id 0 | task 173 | Checking checkpoint with [420, 420] against 405...
|
| 99 |
+
0.29.772.514 W slot update_slots: id 0 | task 173 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 100 |
+
0.29.772.518 W slot update_slots: id 0 | task 173 | erased invalidated context checkpoint (pos_min = 420, pos_max = 420, n_tokens = 421, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 101 |
+
0.30.386.059 I slot create_check: id 0 | task 173 | created context checkpoint 1 of 32 (pos_min = 420, pos_max = 420, n_tokens = 421, size = 62.813 MiB)
|
| 102 |
+
0.30.672.836 I reasoning-budget: deactivated (natural end)
|
| 103 |
+
0.30.764.932 I slot print_timing: id 0 | task 173 |
|
| 104 |
+
prompt eval time = 647.59 ms / 425 tokens ( 1.52 ms per token, 656.28 tokens per second)
|
| 105 |
+
eval time = 344.81 ms / 17 tokens ( 20.28 ms per token, 49.30 tokens per second)
|
| 106 |
+
total time = 992.40 ms / 442 tokens
|
| 107 |
+
0.30.765.023 I slot release: id 0 | task 173 | stop processing: n_tokens = 441, truncated = 0
|
| 108 |
+
0.30.765.050 I srv update_slots: all slots are idle
|
| 109 |
+
0.30.804.618 I srv params_from_: Chat format: peg-native
|
| 110 |
+
0.30.806.300 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.921
|
| 111 |
+
0.30.806.744 I reasoning-budget: activated, budget=2147483647 tokens
|
| 112 |
+
0.30.806.810 I slot launch_slot_: id 0 | task 192 | processing task, is_child = 0
|
| 113 |
+
0.30.806.826 W slot update_slots: id 0 | task 192 | n_past = 406, slot.prompt.tokens.size() = 441, seq_id = 0, pos_min = 440, n_swa = 0
|
| 114 |
+
0.30.806.829 I slot update_slots: id 0 | task 192 | Checking checkpoint with [420, 420] against 406...
|
| 115 |
+
0.30.806.830 W slot update_slots: id 0 | task 192 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 116 |
+
0.30.806.835 W slot update_slots: id 0 | task 192 | erased invalidated context checkpoint (pos_min = 420, pos_max = 420, n_tokens = 421, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 117 |
+
0.31.351.459 I slot create_check: id 0 | task 192 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
|
| 118 |
+
0.31.884.045 I reasoning-budget: deactivated (natural end)
|
| 119 |
+
0.32.727.839 I slot print_timing: id 0 | task 192 |
|
| 120 |
+
prompt eval time = 584.38 ms / 422 tokens ( 1.38 ms per token, 722.13 tokens per second)
|
| 121 |
+
eval time = 1336.62 ms / 62 tokens ( 21.56 ms per token, 46.39 tokens per second)
|
| 122 |
+
total time = 1921.00 ms / 484 tokens
|
| 123 |
+
0.32.727.928 I slot release: id 0 | task 192 | stop processing: n_tokens = 483, truncated = 0
|
| 124 |
+
0.32.727.959 I srv update_slots: all slots are idle
|
| 125 |
+
0.32.742.237 I srv params_from_: Chat format: peg-native
|
| 126 |
+
0.32.742.652 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.854 (> 0.100 thold), f_keep = 0.872
|
| 127 |
+
0.32.743.158 I reasoning-budget: activated, budget=2147483647 tokens
|
| 128 |
+
0.32.743.231 I slot launch_slot_: id 0 | task 256 | processing task, is_child = 0
|
| 129 |
+
0.32.743.251 W slot update_slots: id 0 | task 256 | n_past = 421, slot.prompt.tokens.size() = 483, seq_id = 0, pos_min = 482, n_swa = 0
|
| 130 |
+
0.32.743.254 I slot update_slots: id 0 | task 256 | Checking checkpoint with [417, 417] against 421...
|
| 131 |
+
0.32.748.604 W slot update_slots: id 0 | task 256 | restored context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_past = 418, size = 62.813 MiB)
|
| 132 |
+
0.32.958.098 I slot create_check: id 0 | task 256 | created context checkpoint 2 of 32 (pos_min = 488, pos_max = 488, n_tokens = 489, size = 62.813 MiB)
|
| 133 |
+
0.33.556.876 I reasoning-budget: deactivated (natural end)
|
| 134 |
+
0.33.886.761 I slot print_timing: id 0 | task 256 |
|
| 135 |
+
prompt eval time = 258.37 ms / 75 tokens ( 3.44 ms per token, 290.28 tokens per second)
|
| 136 |
+
eval time = 885.13 ms / 42 tokens ( 21.07 ms per token, 47.45 tokens per second)
|
| 137 |
+
total time = 1143.50 ms / 117 tokens
|
| 138 |
+
0.33.886.848 I slot release: id 0 | task 256 | stop processing: n_tokens = 534, truncated = 0
|
| 139 |
+
0.33.886.880 I srv update_slots: all slots are idle
|
| 140 |
+
0.33.924.683 I srv params_from_: Chat format: peg-native
|
| 141 |
+
0.33.925.111 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.972 (> 0.100 thold), f_keep = 0.768
|
| 142 |
+
0.33.925.551 I reasoning-budget: activated, budget=2147483647 tokens
|
| 143 |
+
0.33.925.629 I slot launch_slot_: id 0 | task 300 | processing task, is_child = 0
|
| 144 |
+
0.33.925.649 W slot update_slots: id 0 | task 300 | n_past = 410, slot.prompt.tokens.size() = 534, seq_id = 0, pos_min = 533, n_swa = 0
|
| 145 |
+
0.33.925.650 I slot update_slots: id 0 | task 300 | Checking checkpoint with [488, 488] against 410...
|
| 146 |
+
0.33.925.652 I slot update_slots: id 0 | task 300 | Checking checkpoint with [417, 417] against 410...
|
| 147 |
+
0.33.925.653 W slot update_slots: id 0 | task 300 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 148 |
+
0.33.925.658 W slot update_slots: id 0 | task 300 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 149 |
+
0.33.927.212 W slot update_slots: id 0 | task 300 | erased invalidated context checkpoint (pos_min = 488, pos_max = 488, n_tokens = 489, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 150 |
+
0.34.478.068 I slot create_check: id 0 | task 300 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
|
| 151 |
+
0.34.936.115 I reasoning-budget: deactivated (natural end)
|
| 152 |
+
0.35.820.698 I slot print_timing: id 0 | task 300 |
|
| 153 |
+
prompt eval time = 590.50 ms / 422 tokens ( 1.40 ms per token, 714.65 tokens per second)
|
| 154 |
+
eval time = 1304.53 ms / 61 tokens ( 21.39 ms per token, 46.76 tokens per second)
|
| 155 |
+
total time = 1895.03 ms / 483 tokens
|
| 156 |
+
0.35.820.771 I slot release: id 0 | task 300 | stop processing: n_tokens = 482, truncated = 0
|
| 157 |
+
0.35.820.802 I srv update_slots: all slots are idle
|
| 158 |
+
0.35.833.343 I srv params_from_: Chat format: peg-native
|
| 159 |
+
0.35.833.721 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.935 (> 0.100 thold), f_keep = 0.840
|
| 160 |
+
0.35.833.997 I reasoning-budget: activated, budget=2147483647 tokens
|
| 161 |
+
0.35.834.089 I slot launch_slot_: id 0 | task 363 | processing task, is_child = 0
|
| 162 |
+
0.35.834.109 W slot update_slots: id 0 | task 363 | n_past = 405, slot.prompt.tokens.size() = 482, seq_id = 0, pos_min = 481, n_swa = 0
|
| 163 |
+
0.35.834.109 I slot update_slots: id 0 | task 363 | Checking checkpoint with [417, 417] against 405...
|
| 164 |
+
0.35.834.111 W slot update_slots: id 0 | task 363 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 165 |
+
0.35.834.118 W slot update_slots: id 0 | task 363 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 166 |
+
0.36.300.366 I slot create_check: id 0 | task 363 | created context checkpoint 1 of 32 (pos_min = 428, pos_max = 428, n_tokens = 429, size = 62.813 MiB)
|
| 167 |
+
0.36.903.862 I reasoning-budget: deactivated (natural end)
|
| 168 |
+
0.38.352.136 I slot print_timing: id 0 | task 363 | n_decoded = 100, tg = 49.65 t/s
|
| 169 |
+
0.38.552.887 I slot print_timing: id 0 | task 363 |
|
| 170 |
+
prompt eval time = 504.10 ms / 433 tokens ( 1.16 ms per token, 858.96 tokens per second)
|
| 171 |
+
eval time = 2214.64 ms / 109 tokens ( 20.32 ms per token, 49.22 tokens per second)
|
| 172 |
+
total time = 2718.74 ms / 542 tokens
|
| 173 |
+
0.38.553.137 I slot release: id 0 | task 363 | stop processing: n_tokens = 541, truncated = 0
|
| 174 |
+
0.38.553.221 I srv update_slots: all slots are idle
|
| 175 |
+
0.38.576.788 I srv params_from_: Chat format: peg-native
|
| 176 |
+
0.38.577.342 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.749
|
| 177 |
+
0.38.577.906 I reasoning-budget: activated, budget=2147483647 tokens
|
| 178 |
+
0.38.577.907 I reasoning-budget: deactivated (natural end)
|
| 179 |
+
0.38.577.985 I slot launch_slot_: id 0 | task 474 | processing task, is_child = 0
|
| 180 |
+
0.38.578.015 W slot update_slots: id 0 | task 474 | n_past = 405, slot.prompt.tokens.size() = 541, seq_id = 0, pos_min = 540, n_swa = 0
|
| 181 |
+
0.38.578.016 I slot update_slots: id 0 | task 474 | Checking checkpoint with [428, 428] against 405...
|
| 182 |
+
0.38.578.018 W slot update_slots: id 0 | task 474 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 183 |
+
0.38.578.022 W slot update_slots: id 0 | task 474 | erased invalidated context checkpoint (pos_min = 428, pos_max = 428, n_tokens = 429, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 184 |
+
0.39.129.005 I slot create_check: id 0 | task 474 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
|
| 185 |
+
0.40.014.664 I slot print_timing: id 0 | task 474 |
|
| 186 |
+
prompt eval time = 601.68 ms / 424 tokens ( 1.42 ms per token, 704.69 tokens per second)
|
| 187 |
+
eval time = 834.95 ms / 39 tokens ( 21.41 ms per token, 46.71 tokens per second)
|
| 188 |
+
total time = 1436.64 ms / 463 tokens
|
| 189 |
+
0.40.014.755 I slot release: id 0 | task 474 | stop processing: n_tokens = 462, truncated = 0
|
| 190 |
+
0.40.014.785 I srv update_slots: all slots are idle
|
| 191 |
+
0.40.063.657 I srv params_from_: Chat format: peg-native
|
| 192 |
+
0.40.065.988 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.902 (> 0.100 thold), f_keep = 0.877
|
| 193 |
+
0.40.066.550 I reasoning-budget: activated, budget=2147483647 tokens
|
| 194 |
+
0.40.066.555 I reasoning-budget: deactivated (natural end)
|
| 195 |
+
0.40.066.646 I slot launch_slot_: id 0 | task 515 | processing task, is_child = 0
|
| 196 |
+
0.40.066.666 W slot update_slots: id 0 | task 515 | n_past = 405, slot.prompt.tokens.size() = 462, seq_id = 0, pos_min = 461, n_swa = 0
|
| 197 |
+
0.40.066.671 I slot update_slots: id 0 | task 515 | Checking checkpoint with [419, 419] against 405...
|
| 198 |
+
0.40.066.673 W slot update_slots: id 0 | task 515 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 199 |
+
0.40.066.679 W slot update_slots: id 0 | task 515 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 200 |
+
0.40.651.450 I slot create_check: id 0 | task 515 | created context checkpoint 1 of 32 (pos_min = 444, pos_max = 444, n_tokens = 445, size = 62.813 MiB)
|
| 201 |
+
0.42.596.129 I slot print_timing: id 0 | task 515 |
|
| 202 |
+
prompt eval time = 621.59 ms / 449 tokens ( 1.38 ms per token, 722.34 tokens per second)
|
| 203 |
+
eval time = 1907.84 ms / 86 tokens ( 22.18 ms per token, 45.08 tokens per second)
|
| 204 |
+
total time = 2529.43 ms / 535 tokens
|
| 205 |
+
0.42.596.351 I slot release: id 0 | task 515 | stop processing: n_tokens = 534, truncated = 0
|
| 206 |
+
0.42.596.411 I srv update_slots: all slots are idle
|
| 207 |
+
0.42.613.151 I srv params_from_: Chat format: peg-native
|
| 208 |
+
0.42.613.661 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.948 (> 0.100 thold), f_keep = 0.758
|
| 209 |
+
0.42.613.938 I reasoning-budget: activated, budget=2147483647 tokens
|
| 210 |
+
0.42.613.941 I reasoning-budget: deactivated (natural end)
|
| 211 |
+
0.42.613.986 I slot launch_slot_: id 0 | task 603 | processing task, is_child = 0
|
| 212 |
+
0.42.614.003 W slot update_slots: id 0 | task 603 | n_past = 405, slot.prompt.tokens.size() = 534, seq_id = 0, pos_min = 533, n_swa = 0
|
| 213 |
+
0.42.614.004 I slot update_slots: id 0 | task 603 | Checking checkpoint with [444, 444] against 405...
|
| 214 |
+
0.42.614.006 W slot update_slots: id 0 | task 603 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 215 |
+
0.42.614.008 W slot update_slots: id 0 | task 603 | erased invalidated context checkpoint (pos_min = 444, pos_max = 444, n_tokens = 445, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 216 |
+
0.43.137.655 I slot create_check: id 0 | task 603 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
|
| 217 |
+
0.44.068.310 I slot print_timing: id 0 | task 603 |
|
| 218 |
+
prompt eval time = 589.04 ms / 427 tokens ( 1.38 ms per token, 724.90 tokens per second)
|
| 219 |
+
eval time = 865.23 ms / 39 tokens ( 22.19 ms per token, 45.07 tokens per second)
|
| 220 |
+
total time = 1454.27 ms / 466 tokens
|
| 221 |
+
0.44.068.519 I slot release: id 0 | task 603 | stop processing: n_tokens = 465, truncated = 0
|
| 222 |
+
0.44.068.586 I srv update_slots: all slots are idle
|
| 223 |
+
0.44.108.932 I srv params_from_: Chat format: peg-native
|
| 224 |
+
0.44.109.519 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.948 (> 0.100 thold), f_keep = 0.871
|
| 225 |
+
0.44.109.808 I reasoning-budget: activated, budget=2147483647 tokens
|
| 226 |
+
0.44.109.810 I reasoning-budget: deactivated (natural end)
|
| 227 |
+
0.44.109.863 I slot launch_slot_: id 0 | task 644 | processing task, is_child = 0
|
| 228 |
+
0.44.109.876 W slot update_slots: id 0 | task 644 | n_past = 405, slot.prompt.tokens.size() = 465, seq_id = 0, pos_min = 464, n_swa = 0
|
| 229 |
+
0.44.109.878 I slot update_slots: id 0 | task 644 | Checking checkpoint with [422, 422] against 405...
|
| 230 |
+
0.44.109.879 W slot update_slots: id 0 | task 644 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 231 |
+
0.44.109.881 W slot update_slots: id 0 | task 644 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 232 |
+
0.44.678.393 I slot create_check: id 0 | task 644 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
|
| 233 |
+
0.44.867.520 I slot print_timing: id 0 | task 644 |
|
| 234 |
+
prompt eval time = 621.79 ms / 427 tokens ( 1.46 ms per token, 686.72 tokens per second)
|
| 235 |
+
eval time = 135.82 ms / 4 tokens ( 33.96 ms per token, 29.45 tokens per second)
|
| 236 |
+
total time = 757.61 ms / 431 tokens
|
| 237 |
+
0.44.867.722 I slot release: id 0 | task 644 | stop processing: n_tokens = 430, truncated = 0
|
| 238 |
+
0.44.867.777 I srv update_slots: all slots are idle
|
| 239 |
+
0.44.920.558 I srv params_from_: Chat format: peg-native
|
| 240 |
+
0.44.922.823 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.958 (> 0.100 thold), f_keep = 0.944
|
| 241 |
+
0.44.923.382 I reasoning-budget: activated, budget=2147483647 tokens
|
| 242 |
+
0.44.923.386 I reasoning-budget: deactivated (natural end)
|
| 243 |
+
0.44.923.473 I slot launch_slot_: id 0 | task 650 | processing task, is_child = 0
|
| 244 |
+
0.44.923.490 W slot update_slots: id 0 | task 650 | n_past = 406, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
|
| 245 |
+
0.44.923.492 I slot update_slots: id 0 | task 650 | Checking checkpoint with [422, 422] against 406...
|
| 246 |
+
0.44.923.496 W slot update_slots: id 0 | task 650 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 247 |
+
0.44.923.503 W slot update_slots: id 0 | task 650 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 248 |
+
0.45.494.241 I slot create_check: id 0 | task 650 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
|
| 249 |
+
0.46.360.864 I slot print_timing: id 0 | task 650 |
|
| 250 |
+
prompt eval time = 619.67 ms / 424 tokens ( 1.46 ms per token, 684.24 tokens per second)
|
| 251 |
+
eval time = 817.69 ms / 40 tokens ( 20.44 ms per token, 48.92 tokens per second)
|
| 252 |
+
total time = 1437.36 ms / 464 tokens
|
| 253 |
+
0.46.360.962 I slot release: id 0 | task 650 | stop processing: n_tokens = 463, truncated = 0
|
| 254 |
+
0.46.360.994 I srv update_slots: all slots are idle
|
| 255 |
+
0.46.376.404 I srv params_from_: Chat format: peg-native
|
| 256 |
+
0.46.376.859 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.935 (> 0.100 thold), f_keep = 1.000
|
| 257 |
+
0.46.377.382 I reasoning-budget: activated, budget=2147483647 tokens
|
| 258 |
+
0.46.377.385 I reasoning-budget: deactivated (natural end)
|
| 259 |
+
0.46.377.478 I slot launch_slot_: id 0 | task 692 | processing task, is_child = 0
|
| 260 |
+
0.46.580.028 I slot create_check: id 0 | task 692 | created context checkpoint 2 of 32 (pos_min = 490, pos_max = 490, n_tokens = 491, size = 62.813 MiB)
|
| 261 |
+
0.46.886.342 I slot print_timing: id 0 | task 692 |
|
| 262 |
+
prompt eval time = 260.43 ms / 32 tokens ( 8.14 ms per token, 122.87 tokens per second)
|
| 263 |
+
eval time = 248.41 ms / 13 tokens ( 19.11 ms per token, 52.33 tokens per second)
|
| 264 |
+
total time = 508.84 ms / 45 tokens
|
| 265 |
+
0.46.886.415 I slot release: id 0 | task 692 | stop processing: n_tokens = 507, truncated = 0
|
| 266 |
+
0.46.886.437 I srv update_slots: all slots are idle
|
| 267 |
+
0.46.902.435 I srv params_from_: Chat format: peg-native
|
| 268 |
+
0.46.903.190 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.967 (> 0.100 thold), f_keep = 0.809
|
| 269 |
+
0.46.903.420 I reasoning-budget: activated, budget=2147483647 tokens
|
| 270 |
+
0.46.903.422 I reasoning-budget: deactivated (natural end)
|
| 271 |
+
0.46.903.471 I slot launch_slot_: id 0 | task 707 | processing task, is_child = 0
|
| 272 |
+
0.46.903.481 W slot update_slots: id 0 | task 707 | n_past = 410, slot.prompt.tokens.size() = 507, seq_id = 0, pos_min = 506, n_swa = 0
|
| 273 |
+
0.46.903.483 I slot update_slots: id 0 | task 707 | Checking checkpoint with [490, 490] against 410...
|
| 274 |
+
0.46.903.484 I slot update_slots: id 0 | task 707 | Checking checkpoint with [419, 419] against 410...
|
| 275 |
+
0.46.903.485 W slot update_slots: id 0 | task 707 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 276 |
+
0.46.903.488 W slot update_slots: id 0 | task 707 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 277 |
+
0.46.904.537 W slot update_slots: id 0 | task 707 | erased invalidated context checkpoint (pos_min = 490, pos_max = 490, n_tokens = 491, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 278 |
+
0.47.441.027 I slot create_check: id 0 | task 707 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
|
| 279 |
+
0.48.289.348 I slot print_timing: id 0 | task 707 |
|
| 280 |
+
prompt eval time = 577.77 ms / 424 tokens ( 1.36 ms per token, 733.85 tokens per second)
|
| 281 |
+
eval time = 808.09 ms / 40 tokens ( 20.20 ms per token, 49.50 tokens per second)
|
| 282 |
+
total time = 1385.86 ms / 464 tokens
|
| 283 |
+
0.48.289.404 I slot release: id 0 | task 707 | stop processing: n_tokens = 463, truncated = 0
|
| 284 |
+
0.48.289.426 I srv update_slots: all slots are idle
|
| 285 |
+
0.48.302.317 I srv params_from_: Chat format: peg-native
|
| 286 |
+
0.48.302.639 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.931 (> 0.100 thold), f_keep = 0.875
|
| 287 |
+
0.48.302.821 I reasoning-budget: activated, budget=2147483647 tokens
|
| 288 |
+
0.48.302.822 I reasoning-budget: deactivated (natural end)
|
| 289 |
+
0.48.302.871 I slot launch_slot_: id 0 | task 749 | processing task, is_child = 0
|
| 290 |
+
0.48.302.880 W slot update_slots: id 0 | task 749 | n_past = 405, slot.prompt.tokens.size() = 463, seq_id = 0, pos_min = 462, n_swa = 0
|
| 291 |
+
0.48.302.881 I slot update_slots: id 0 | task 749 | Checking checkpoint with [419, 419] against 405...
|
| 292 |
+
0.48.302.882 W slot update_slots: id 0 | task 749 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 293 |
+
0.48.302.884 W slot update_slots: id 0 | task 749 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 294 |
+
0.48.829.428 I slot create_check: id 0 | task 749 | created context checkpoint 1 of 32 (pos_min = 430, pos_max = 430, n_tokens = 431, size = 62.813 MiB)
|
| 295 |
+
0.50.509.107 I slot print_timing: id 0 | task 749 |
|
| 296 |
+
prompt eval time = 572.46 ms / 435 tokens ( 1.32 ms per token, 759.87 tokens per second)
|
| 297 |
+
eval time = 1633.74 ms / 80 tokens ( 20.42 ms per token, 48.97 tokens per second)
|
| 298 |
+
total time = 2206.20 ms / 515 tokens
|
| 299 |
+
0.50.509.286 I slot release: id 0 | task 749 | stop processing: n_tokens = 514, truncated = 0
|
| 300 |
+
0.50.509.314 I srv update_slots: all slots are idle
|
| 301 |
+
0.50.510.178 I srv operator(): operator(): cleaning up before exit...
|
recipe/logs/b_n-tools-q106-tpl-medium-probe.log
ADDED
|
@@ -0,0 +1,278 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
0.00.129.990 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
|
| 2 |
+
0.00.129.993 I device_info:
|
| 3 |
+
0.00.130.113 I - ROCm0 : AMD Radeon Graphics (131072 MiB, 122319 MiB free)
|
| 4 |
+
0.00.130.278 I - Vulkan0 : AMD Radeon Graphics (RADV GFX1151) (132096 MiB, 131922 MiB free)
|
| 5 |
+
0.00.130.286 I - CPU : AMD RYZEN AI MAX+ 395 w/ Radeon 8060S (127438 MiB, 127438 MiB free)
|
| 6 |
+
0.00.130.376 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
|
| 7 |
+
0.00.130.453 I srv init: running without SSL
|
| 8 |
+
0.00.130.481 I srv init: using 31 threads for HTTP server
|
| 9 |
+
0.00.130.482 I srv init: the WebUI is disabled
|
| 10 |
+
0.00.130.551 I srv start: binding port with default address family
|
| 11 |
+
0.00.131.742 I srv main: loading model
|
| 12 |
+
0.00.131.749 I srv load_model: loading model '/mnt/models/nex-n2.5-mini/out/Nex-N2.5-mini-Q4_0_ROCMFP4_STRIX_LEAN.gguf'
|
| 13 |
+
0.00.190.887 W llama_model_loader: direct I/O is enabled, disabling mmap
|
| 14 |
+
0.23.864.154 W llama_context: n_ctx_seq (65536) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
|
| 15 |
+
0.24.195.940 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|
| 16 |
+
0.24.556.150 I srv load_model: initializing slots, n_slots = 1
|
| 17 |
+
0.24.849.854 W srv load_model: speculative decoding will use checkpoints
|
| 18 |
+
0.24.849.866 W common_speculative_init: no implementations specified for speculative decoding
|
| 19 |
+
0.24.849.870 I slot load_model: id 0 | task -1 | new slot, n_ctx = 65536
|
| 20 |
+
0.24.849.977 I srv load_model: prompt cache RAM enabled: limit_mib=8192
|
| 21 |
+
0.24.849.981 I srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
|
| 22 |
+
0.24.850.006 I srv init: idle slots will be saved to prompt cache upon starting a new task
|
| 23 |
+
0.24.907.565 I init: chat template, example_format: '<|im_start|>system
|
| 24 |
+
You are a helpful assistant<|im_end|>
|
| 25 |
+
<|im_start|>user
|
| 26 |
+
Hello<|im_end|>
|
| 27 |
+
<|im_start|>assistant
|
| 28 |
+
<think>
|
| 29 |
+
|
| 30 |
+
</think>
|
| 31 |
+
|
| 32 |
+
Hi there<|im_end|>
|
| 33 |
+
<|im_start|>user
|
| 34 |
+
How are you?<|im_end|>
|
| 35 |
+
<|im_start|>assistant
|
| 36 |
+
<think>'
|
| 37 |
+
0.24.952.737 I srv init: init: chat template, thinking = 1
|
| 38 |
+
0.24.952.792 I srv main: model loaded
|
| 39 |
+
0.24.952.798 I srv main: server is listening on http://127.0.0.1:18652
|
| 40 |
+
0.24.952.803 I srv update_slots: all slots are idle
|
| 41 |
+
0.26.377.211 I srv params_from_: Chat format: peg-native
|
| 42 |
+
0.26.379.299 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1
|
| 43 |
+
0.26.379.305 I srv get_availabl: updating prompt cache
|
| 44 |
+
0.26.379.315 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000
|
| 45 |
+
0.26.379.323 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 65536 tokens, 8589934592 est)
|
| 46 |
+
0.26.379.326 I srv get_availabl: prompt cache update took 0.02 ms
|
| 47 |
+
0.26.380.071 I reasoning-budget: activated, budget=2147483647 tokens
|
| 48 |
+
0.26.380.099 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0
|
| 49 |
+
0.27.073.027 I slot create_check: id 0 | task 0 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
|
| 50 |
+
0.27.138.087 I reasoning-budget: deactivated (natural end)
|
| 51 |
+
0.27.269.343 I slot print_timing: id 0 | task 0 |
|
| 52 |
+
prompt eval time = 728.42 ms / 424 tokens ( 1.72 ms per token, 582.08 tokens per second)
|
| 53 |
+
eval time = 160.80 ms / 7 tokens ( 22.97 ms per token, 43.53 tokens per second)
|
| 54 |
+
total time = 889.21 ms / 431 tokens
|
| 55 |
+
0.27.269.422 I slot release: id 0 | task 0 | stop processing: n_tokens = 430, truncated = 0
|
| 56 |
+
0.27.269.432 I srv update_slots: all slots are idle
|
| 57 |
+
0.27.299.600 I srv params_from_: Chat format: peg-native
|
| 58 |
+
0.27.300.148 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.969 (> 0.100 thold), f_keep = 0.942
|
| 59 |
+
0.27.300.612 I reasoning-budget: activated, budget=2147483647 tokens
|
| 60 |
+
0.27.300.707 I slot launch_slot_: id 0 | task 9 | processing task, is_child = 0
|
| 61 |
+
0.27.300.733 W slot update_slots: id 0 | task 9 | n_past = 405, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
|
| 62 |
+
0.27.300.737 I slot update_slots: id 0 | task 9 | Checking checkpoint with [419, 419] against 405...
|
| 63 |
+
0.27.300.739 W slot update_slots: id 0 | task 9 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 64 |
+
0.27.300.746 W slot update_slots: id 0 | task 9 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 65 |
+
0.27.937.713 I slot create_check: id 0 | task 9 | created context checkpoint 1 of 32 (pos_min = 413, pos_max = 413, n_tokens = 414, size = 62.813 MiB)
|
| 66 |
+
0.28.043.202 I reasoning-budget: deactivated (natural end)
|
| 67 |
+
0.28.142.801 I slot print_timing: id 0 | task 9 |
|
| 68 |
+
prompt eval time = 701.16 ms / 418 tokens ( 1.68 ms per token, 596.15 tokens per second)
|
| 69 |
+
eval time = 140.89 ms / 5 tokens ( 28.18 ms per token, 35.49 tokens per second)
|
| 70 |
+
total time = 842.06 ms / 423 tokens
|
| 71 |
+
0.28.142.910 I slot release: id 0 | task 9 | stop processing: n_tokens = 422, truncated = 0
|
| 72 |
+
0.28.142.959 I srv update_slots: all slots are idle
|
| 73 |
+
0.28.158.005 I srv params_from_: Chat format: peg-native
|
| 74 |
+
0.28.158.364 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.960
|
| 75 |
+
0.28.158.578 I reasoning-budget: activated, budget=2147483647 tokens
|
| 76 |
+
0.28.158.608 I slot launch_slot_: id 0 | task 16 | processing task, is_child = 0
|
| 77 |
+
0.28.158.616 W slot update_slots: id 0 | task 16 | n_past = 405, slot.prompt.tokens.size() = 422, seq_id = 0, pos_min = 421, n_swa = 0
|
| 78 |
+
0.28.158.617 I slot update_slots: id 0 | task 16 | Checking checkpoint with [413, 413] against 405...
|
| 79 |
+
0.28.158.618 W slot update_slots: id 0 | task 16 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 80 |
+
0.28.158.620 W slot update_slots: id 0 | task 16 | erased invalidated context checkpoint (pos_min = 413, pos_max = 413, n_tokens = 414, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 81 |
+
0.28.549.814 I slot create_check: id 0 | task 16 | created context checkpoint 1 of 32 (pos_min = 416, pos_max = 416, n_tokens = 417, size = 62.813 MiB)
|
| 82 |
+
0.28.603.918 I reasoning-budget: deactivated (natural end)
|
| 83 |
+
0.29.232.552 I slot print_timing: id 0 | task 16 |
|
| 84 |
+
prompt eval time = 422.03 ms / 421 tokens ( 1.00 ms per token, 997.56 tokens per second)
|
| 85 |
+
eval time = 651.88 ms / 42 tokens ( 15.52 ms per token, 64.43 tokens per second)
|
| 86 |
+
total time = 1073.91 ms / 463 tokens
|
| 87 |
+
0.29.232.642 I slot release: id 0 | task 16 | stop processing: n_tokens = 462, truncated = 0
|
| 88 |
+
0.29.232.679 I srv update_slots: all slots are idle
|
| 89 |
+
0.29.259.355 I srv params_from_: Chat format: peg-native
|
| 90 |
+
0.29.259.776 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.951 (> 0.100 thold), f_keep = 0.879
|
| 91 |
+
0.29.259.946 I reasoning-budget: activated, budget=2147483647 tokens
|
| 92 |
+
0.29.259.949 I reasoning-budget: deactivated (natural end)
|
| 93 |
+
0.29.259.987 I slot launch_slot_: id 0 | task 60 | processing task, is_child = 0
|
| 94 |
+
0.29.259.997 W slot update_slots: id 0 | task 60 | n_past = 406, slot.prompt.tokens.size() = 462, seq_id = 0, pos_min = 461, n_swa = 0
|
| 95 |
+
0.29.259.998 I slot update_slots: id 0 | task 60 | Checking checkpoint with [416, 416] against 406...
|
| 96 |
+
0.29.260.000 W slot update_slots: id 0 | task 60 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 97 |
+
0.29.260.002 W slot update_slots: id 0 | task 60 | erased invalidated context checkpoint (pos_min = 416, pos_max = 416, n_tokens = 417, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 98 |
+
0.29.608.781 I slot create_check: id 0 | task 60 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
|
| 99 |
+
0.29.705.934 I slot print_timing: id 0 | task 60 |
|
| 100 |
+
prompt eval time = 378.91 ms / 427 tokens ( 0.89 ms per token, 1126.93 tokens per second)
|
| 101 |
+
eval time = 67.01 ms / 4 tokens ( 16.75 ms per token, 59.69 tokens per second)
|
| 102 |
+
total time = 445.92 ms / 431 tokens
|
| 103 |
+
0.29.706.025 I slot release: id 0 | task 60 | stop processing: n_tokens = 430, truncated = 0
|
| 104 |
+
0.29.706.063 I srv update_slots: all slots are idle
|
| 105 |
+
0.29.723.346 I srv params_from_: Chat format: peg-native
|
| 106 |
+
0.29.723.698 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.942
|
| 107 |
+
0.29.723.878 I reasoning-budget: activated, budget=2147483647 tokens
|
| 108 |
+
0.29.723.880 I reasoning-budget: deactivated (natural end)
|
| 109 |
+
0.29.723.911 I slot launch_slot_: id 0 | task 66 | processing task, is_child = 0
|
| 110 |
+
0.29.723.921 W slot update_slots: id 0 | task 66 | n_past = 405, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
|
| 111 |
+
0.29.723.922 I slot update_slots: id 0 | task 66 | Checking checkpoint with [422, 422] against 405...
|
| 112 |
+
0.29.723.924 W slot update_slots: id 0 | task 66 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 113 |
+
0.29.723.926 W slot update_slots: id 0 | task 66 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 114 |
+
0.30.074.713 I slot create_check: id 0 | task 66 | created context checkpoint 1 of 32 (pos_min = 416, pos_max = 416, n_tokens = 417, size = 62.813 MiB)
|
| 115 |
+
0.30.128.419 I slot print_timing: id 0 | task 66 |
|
| 116 |
+
prompt eval time = 380.89 ms / 421 tokens ( 0.90 ms per token, 1105.32 tokens per second)
|
| 117 |
+
eval time = 23.59 ms / 2 tokens ( 11.79 ms per token, 84.79 tokens per second)
|
| 118 |
+
total time = 404.48 ms / 423 tokens
|
| 119 |
+
0.30.128.524 I slot release: id 0 | task 66 | stop processing: n_tokens = 422, truncated = 0
|
| 120 |
+
0.30.128.572 I srv update_slots: all slots are idle
|
| 121 |
+
0.30.150.200 I srv params_from_: Chat format: peg-native
|
| 122 |
+
0.30.150.557 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.960
|
| 123 |
+
0.30.150.711 I reasoning-budget: activated, budget=2147483647 tokens
|
| 124 |
+
0.30.150.713 I reasoning-budget: deactivated (natural end)
|
| 125 |
+
0.30.150.739 I slot launch_slot_: id 0 | task 70 | processing task, is_child = 0
|
| 126 |
+
0.30.150.744 W slot update_slots: id 0 | task 70 | n_past = 405, slot.prompt.tokens.size() = 422, seq_id = 0, pos_min = 421, n_swa = 0
|
| 127 |
+
0.30.150.745 I slot update_slots: id 0 | task 70 | Checking checkpoint with [416, 416] against 405...
|
| 128 |
+
0.30.150.746 W slot update_slots: id 0 | task 70 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 129 |
+
0.30.150.752 W slot update_slots: id 0 | task 70 | erased invalidated context checkpoint (pos_min = 416, pos_max = 416, n_tokens = 417, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 130 |
+
0.30.504.785 I slot create_check: id 0 | task 70 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
|
| 131 |
+
0.31.192.911 I slot print_timing: id 0 | task 70 |
|
| 132 |
+
prompt eval time = 383.87 ms / 424 tokens ( 0.91 ms per token, 1104.53 tokens per second)
|
| 133 |
+
eval time = 658.28 ms / 39 tokens ( 16.88 ms per token, 59.25 tokens per second)
|
| 134 |
+
total time = 1042.15 ms / 463 tokens
|
| 135 |
+
0.31.192.995 I slot release: id 0 | task 70 | stop processing: n_tokens = 462, truncated = 0
|
| 136 |
+
0.31.193.032 I srv update_slots: all slots are idle
|
| 137 |
+
0.31.213.646 I srv params_from_: Chat format: peg-native
|
| 138 |
+
0.31.213.997 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.879
|
| 139 |
+
0.31.214.182 I reasoning-budget: activated, budget=2147483647 tokens
|
| 140 |
+
0.31.214.218 I slot launch_slot_: id 0 | task 111 | processing task, is_child = 0
|
| 141 |
+
0.31.214.228 W slot update_slots: id 0 | task 111 | n_past = 406, slot.prompt.tokens.size() = 462, seq_id = 0, pos_min = 461, n_swa = 0
|
| 142 |
+
0.31.214.230 I slot update_slots: id 0 | task 111 | Checking checkpoint with [419, 419] against 406...
|
| 143 |
+
0.31.214.231 W slot update_slots: id 0 | task 111 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 144 |
+
0.31.214.234 W slot update_slots: id 0 | task 111 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 145 |
+
0.31.741.413 I slot create_check: id 0 | task 111 | created context checkpoint 1 of 32 (pos_min = 420, pos_max = 420, n_tokens = 421, size = 62.813 MiB)
|
| 146 |
+
0.32.002.855 I reasoning-budget: deactivated (natural end)
|
| 147 |
+
0.32.103.990 I slot print_timing: id 0 | task 111 |
|
| 148 |
+
prompt eval time = 558.28 ms / 425 tokens ( 1.31 ms per token, 761.27 tokens per second)
|
| 149 |
+
eval time = 331.47 ms / 17 tokens ( 19.50 ms per token, 51.29 tokens per second)
|
| 150 |
+
total time = 889.75 ms / 442 tokens
|
| 151 |
+
0.32.104.074 I slot release: id 0 | task 111 | stop processing: n_tokens = 441, truncated = 0
|
| 152 |
+
0.32.104.109 I srv update_slots: all slots are idle
|
| 153 |
+
0.32.155.224 I srv params_from_: Chat format: peg-native
|
| 154 |
+
0.32.157.476 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.967 (> 0.100 thold), f_keep = 0.918
|
| 155 |
+
0.32.157.999 I reasoning-budget: activated, budget=2147483647 tokens
|
| 156 |
+
0.32.158.095 I slot launch_slot_: id 0 | task 130 | processing task, is_child = 0
|
| 157 |
+
0.32.158.121 W slot update_slots: id 0 | task 130 | n_past = 405, slot.prompt.tokens.size() = 441, seq_id = 0, pos_min = 440, n_swa = 0
|
| 158 |
+
0.32.158.123 I slot update_slots: id 0 | task 130 | Checking checkpoint with [420, 420] against 405...
|
| 159 |
+
0.32.158.124 W slot update_slots: id 0 | task 130 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 160 |
+
0.32.158.130 W slot update_slots: id 0 | task 130 | erased invalidated context checkpoint (pos_min = 420, pos_max = 420, n_tokens = 421, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 161 |
+
0.32.714.547 I slot create_check: id 0 | task 130 | created context checkpoint 1 of 32 (pos_min = 414, pos_max = 414, n_tokens = 415, size = 62.813 MiB)
|
| 162 |
+
0.32.960.511 I reasoning-budget: deactivated (natural end)
|
| 163 |
+
0.33.017.861 I slot print_timing: id 0 | task 130 |
|
| 164 |
+
prompt eval time = 601.42 ms / 419 tokens ( 1.44 ms per token, 696.69 tokens per second)
|
| 165 |
+
eval time = 258.30 ms / 12 tokens ( 21.53 ms per token, 46.46 tokens per second)
|
| 166 |
+
total time = 859.72 ms / 431 tokens
|
| 167 |
+
0.33.018.052 I slot release: id 0 | task 130 | stop processing: n_tokens = 430, truncated = 0
|
| 168 |
+
0.33.018.108 I srv update_slots: all slots are idle
|
| 169 |
+
0.33.050.367 I srv params_from_: Chat format: peg-native
|
| 170 |
+
0.33.050.835 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.960 (> 0.100 thold), f_keep = 0.942
|
| 171 |
+
0.33.051.413 I reasoning-budget: activated, budget=2147483647 tokens
|
| 172 |
+
0.33.051.498 I slot launch_slot_: id 0 | task 144 | processing task, is_child = 0
|
| 173 |
+
0.33.051.517 W slot update_slots: id 0 | task 144 | n_past = 405, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
|
| 174 |
+
0.33.051.521 I slot update_slots: id 0 | task 144 | Checking checkpoint with [414, 414] against 405...
|
| 175 |
+
0.33.051.522 W slot update_slots: id 0 | task 144 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 176 |
+
0.33.051.528 W slot update_slots: id 0 | task 144 | erased invalidated context checkpoint (pos_min = 414, pos_max = 414, n_tokens = 415, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 177 |
+
0.33.615.892 I slot create_check: id 0 | task 144 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
|
| 178 |
+
0.33.916.408 I reasoning-budget: deactivated (natural end)
|
| 179 |
+
0.34.765.727 I slot print_timing: id 0 | task 144 |
|
| 180 |
+
prompt eval time = 597.52 ms / 422 tokens ( 1.42 ms per token, 706.25 tokens per second)
|
| 181 |
+
eval time = 1116.68 ms / 53 tokens ( 21.07 ms per token, 47.46 tokens per second)
|
| 182 |
+
total time = 1714.20 ms / 475 tokens
|
| 183 |
+
0.34.765.808 I slot release: id 0 | task 144 | stop processing: n_tokens = 474, truncated = 0
|
| 184 |
+
0.34.765.835 I srv update_slots: all slots are idle
|
| 185 |
+
0.34.781.504 I srv params_from_: Chat format: peg-native
|
| 186 |
+
0.34.781.937 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.958 (> 0.100 thold), f_keep = 0.857
|
| 187 |
+
0.34.782.155 I reasoning-budget: activated, budget=2147483647 tokens
|
| 188 |
+
0.34.782.201 I slot launch_slot_: id 0 | task 199 | processing task, is_child = 0
|
| 189 |
+
0.34.782.213 W slot update_slots: id 0 | task 199 | n_past = 406, slot.prompt.tokens.size() = 474, seq_id = 0, pos_min = 473, n_swa = 0
|
| 190 |
+
0.34.782.213 I slot update_slots: id 0 | task 199 | Checking checkpoint with [417, 417] against 406...
|
| 191 |
+
0.34.782.214 W slot update_slots: id 0 | task 199 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 192 |
+
0.34.782.218 W slot update_slots: id 0 | task 199 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 193 |
+
0.35.312.464 I slot create_check: id 0 | task 199 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
|
| 194 |
+
0.35.382.494 I reasoning-budget: deactivated (natural end)
|
| 195 |
+
0.35.500.672 I slot print_timing: id 0 | task 199 |
|
| 196 |
+
prompt eval time = 569.17 ms / 424 tokens ( 1.34 ms per token, 744.95 tokens per second)
|
| 197 |
+
eval time = 149.28 ms / 7 tokens ( 21.33 ms per token, 46.89 tokens per second)
|
| 198 |
+
total time = 718.45 ms / 431 tokens
|
| 199 |
+
0.35.500.758 I slot release: id 0 | task 199 | stop processing: n_tokens = 430, truncated = 0
|
| 200 |
+
0.35.500.792 I srv update_slots: all slots are idle
|
| 201 |
+
0.35.523.836 I srv params_from_: Chat format: peg-native
|
| 202 |
+
0.35.524.275 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.969 (> 0.100 thold), f_keep = 0.942
|
| 203 |
+
0.35.524.794 I reasoning-budget: activated, budget=2147483647 tokens
|
| 204 |
+
0.35.524.873 I slot launch_slot_: id 0 | task 208 | processing task, is_child = 0
|
| 205 |
+
0.35.524.891 W slot update_slots: id 0 | task 208 | n_past = 405, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
|
| 206 |
+
0.35.524.892 I slot update_slots: id 0 | task 208 | Checking checkpoint with [419, 419] against 405...
|
| 207 |
+
0.35.524.894 W slot update_slots: id 0 | task 208 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 208 |
+
0.35.524.899 W slot update_slots: id 0 | task 208 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 209 |
+
0.36.142.960 I slot create_check: id 0 | task 208 | created context checkpoint 1 of 32 (pos_min = 413, pos_max = 413, n_tokens = 414, size = 62.813 MiB)
|
| 210 |
+
0.36.260.514 I reasoning-budget: deactivated (natural end)
|
| 211 |
+
0.36.391.695 I slot print_timing: id 0 | task 208 |
|
| 212 |
+
prompt eval time = 684.84 ms / 418 tokens ( 1.64 ms per token, 610.36 tokens per second)
|
| 213 |
+
eval time = 181.96 ms / 5 tokens ( 36.39 ms per token, 27.48 tokens per second)
|
| 214 |
+
total time = 866.79 ms / 423 tokens
|
| 215 |
+
0.36.391.789 I slot release: id 0 | task 208 | stop processing: n_tokens = 422, truncated = 0
|
| 216 |
+
0.36.391.820 I srv update_slots: all slots are idle
|
| 217 |
+
0.36.404.024 I srv params_from_: Chat format: peg-native
|
| 218 |
+
0.36.404.587 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.960
|
| 219 |
+
0.36.404.848 I reasoning-budget: activated, budget=2147483647 tokens
|
| 220 |
+
0.36.404.897 I slot launch_slot_: id 0 | task 215 | processing task, is_child = 0
|
| 221 |
+
0.36.404.910 W slot update_slots: id 0 | task 215 | n_past = 405, slot.prompt.tokens.size() = 422, seq_id = 0, pos_min = 421, n_swa = 0
|
| 222 |
+
0.36.404.912 I slot update_slots: id 0 | task 215 | Checking checkpoint with [413, 413] against 405...
|
| 223 |
+
0.36.404.913 W slot update_slots: id 0 | task 215 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 224 |
+
0.36.404.916 W slot update_slots: id 0 | task 215 | erased invalidated context checkpoint (pos_min = 413, pos_max = 413, n_tokens = 414, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 225 |
+
0.36.899.390 I slot create_check: id 0 | task 215 | created context checkpoint 1 of 32 (pos_min = 416, pos_max = 416, n_tokens = 417, size = 62.813 MiB)
|
| 226 |
+
0.37.012.108 I reasoning-budget: deactivated (natural end)
|
| 227 |
+
0.37.909.549 I slot print_timing: id 0 | task 215 |
|
| 228 |
+
prompt eval time = 555.94 ms / 421 tokens ( 1.32 ms per token, 757.28 tokens per second)
|
| 229 |
+
eval time = 948.68 ms / 42 tokens ( 22.59 ms per token, 44.27 tokens per second)
|
| 230 |
+
total time = 1504.62 ms / 463 tokens
|
| 231 |
+
0.37.909.654 I slot release: id 0 | task 215 | stop processing: n_tokens = 462, truncated = 0
|
| 232 |
+
0.37.909.691 I srv update_slots: all slots are idle
|
| 233 |
+
0.37.937.856 I srv params_from_: Chat format: peg-native
|
| 234 |
+
0.37.938.240 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.951 (> 0.100 thold), f_keep = 0.879
|
| 235 |
+
0.37.938.471 I reasoning-budget: activated, budget=2147483647 tokens
|
| 236 |
+
0.37.938.517 I slot launch_slot_: id 0 | task 259 | processing task, is_child = 0
|
| 237 |
+
0.37.938.529 W slot update_slots: id 0 | task 259 | n_past = 406, slot.prompt.tokens.size() = 462, seq_id = 0, pos_min = 461, n_swa = 0
|
| 238 |
+
0.37.938.531 I slot update_slots: id 0 | task 259 | Checking checkpoint with [416, 416] against 406...
|
| 239 |
+
0.37.938.532 W slot update_slots: id 0 | task 259 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 240 |
+
0.37.938.535 W slot update_slots: id 0 | task 259 | erased invalidated context checkpoint (pos_min = 416, pos_max = 416, n_tokens = 417, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 241 |
+
0.38.478.969 I slot create_check: id 0 | task 259 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
|
| 242 |
+
0.38.599.473 I slot print_timing: id 0 | task 259 |
|
| 243 |
+
prompt eval time = 580.20 ms / 427 tokens ( 1.36 ms per token, 735.95 tokens per second)
|
| 244 |
+
eval time = 80.65 ms / 4 tokens ( 20.16 ms per token, 49.60 tokens per second)
|
| 245 |
+
total time = 660.85 ms / 431 tokens
|
| 246 |
+
0.38.599.762 I slot release: id 0 | task 259 | stop processing: n_tokens = 430, truncated = 0
|
| 247 |
+
0.38.599.853 I srv update_slots: all slots are idle
|
| 248 |
+
0.38.629.562 I srv params_from_: Chat format: peg-native
|
| 249 |
+
0.38.630.012 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.942
|
| 250 |
+
0.38.630.609 I reasoning-budget: activated, budget=2147483647 tokens
|
| 251 |
+
0.38.630.693 I slot launch_slot_: id 0 | task 265 | processing task, is_child = 0
|
| 252 |
+
0.38.630.711 W slot update_slots: id 0 | task 265 | n_past = 405, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
|
| 253 |
+
0.38.630.716 I slot update_slots: id 0 | task 265 | Checking checkpoint with [422, 422] against 405...
|
| 254 |
+
0.38.630.719 W slot update_slots: id 0 | task 265 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 255 |
+
0.38.630.724 W slot update_slots: id 0 | task 265 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 256 |
+
0.39.188.051 I slot create_check: id 0 | task 265 | created context checkpoint 1 of 32 (pos_min = 416, pos_max = 416, n_tokens = 417, size = 62.813 MiB)
|
| 257 |
+
0.39.260.024 I slot print_timing: id 0 | task 265 |
|
| 258 |
+
prompt eval time = 591.28 ms / 421 tokens ( 1.40 ms per token, 712.01 tokens per second)
|
| 259 |
+
eval time = 37.99 ms / 2 tokens ( 19.00 ms per token, 52.64 tokens per second)
|
| 260 |
+
total time = 629.28 ms / 423 tokens
|
| 261 |
+
0.39.260.276 I slot release: id 0 | task 265 | stop processing: n_tokens = 422, truncated = 0
|
| 262 |
+
0.39.260.338 I srv update_slots: all slots are idle
|
| 263 |
+
0.39.312.544 I srv params_from_: Chat format: peg-native
|
| 264 |
+
0.39.314.833 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.960
|
| 265 |
+
0.39.315.431 I reasoning-budget: activated, budget=2147483647 tokens
|
| 266 |
+
0.39.315.516 I slot launch_slot_: id 0 | task 269 | processing task, is_child = 0
|
| 267 |
+
0.39.315.539 W slot update_slots: id 0 | task 269 | n_past = 405, slot.prompt.tokens.size() = 422, seq_id = 0, pos_min = 421, n_swa = 0
|
| 268 |
+
0.39.315.543 I slot update_slots: id 0 | task 269 | Checking checkpoint with [416, 416] against 405...
|
| 269 |
+
0.39.315.545 W slot update_slots: id 0 | task 269 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 270 |
+
0.39.315.552 W slot update_slots: id 0 | task 269 | erased invalidated context checkpoint (pos_min = 416, pos_max = 416, n_tokens = 417, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 271 |
+
0.39.893.482 I slot create_check: id 0 | task 269 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
|
| 272 |
+
0.40.793.357 I slot print_timing: id 0 | task 269 |
|
| 273 |
+
prompt eval time = 633.79 ms / 424 tokens ( 1.49 ms per token, 668.99 tokens per second)
|
| 274 |
+
eval time = 844.02 ms / 39 tokens ( 21.64 ms per token, 46.21 tokens per second)
|
| 275 |
+
total time = 1477.81 ms / 463 tokens
|
| 276 |
+
0.40.793.457 I slot release: id 0 | task 269 | stop processing: n_tokens = 462, truncated = 0
|
| 277 |
+
0.40.793.487 I srv update_slots: all slots are idle
|
| 278 |
+
0.40.794.731 I srv operator(): operator(): cleaning up before exit...
|
recipe/logs/b_n-tools-q106-tpl-medium.log
ADDED
|
@@ -0,0 +1,294 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
0.00.105.454 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
|
| 2 |
+
0.00.105.464 I device_info:
|
| 3 |
+
0.00.105.601 I - ROCm0 : AMD Radeon Graphics (131072 MiB, 122408 MiB free)
|
| 4 |
+
0.00.105.795 I - Vulkan0 : AMD Radeon Graphics (RADV GFX1151) (132096 MiB, 131922 MiB free)
|
| 5 |
+
0.00.105.803 I - CPU : AMD RYZEN AI MAX+ 395 w/ Radeon 8060S (127438 MiB, 127438 MiB free)
|
| 6 |
+
0.00.105.902 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
|
| 7 |
+
0.00.105.988 I srv init: running without SSL
|
| 8 |
+
0.00.106.040 I srv init: using 31 threads for HTTP server
|
| 9 |
+
0.00.106.042 I srv init: the WebUI is disabled
|
| 10 |
+
0.00.106.129 I srv start: binding port with default address family
|
| 11 |
+
0.00.107.354 I srv main: loading model
|
| 12 |
+
0.00.107.360 I srv load_model: loading model '/mnt/models/nex-n2.5-mini/out/Nex-N2.5-mini-Q4_0_ROCMFP4_STRIX_LEAN.gguf'
|
| 13 |
+
0.00.146.109 W llama_model_loader: direct I/O is enabled, disabling mmap
|
| 14 |
+
0.23.588.699 W llama_context: n_ctx_seq (65536) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
|
| 15 |
+
0.23.912.034 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|
| 16 |
+
0.24.349.867 I srv load_model: initializing slots, n_slots = 1
|
| 17 |
+
0.24.567.364 W srv load_model: speculative decoding will use checkpoints
|
| 18 |
+
0.24.567.384 W common_speculative_init: no implementations specified for speculative decoding
|
| 19 |
+
0.24.567.388 I slot load_model: id 0 | task -1 | new slot, n_ctx = 65536
|
| 20 |
+
0.24.567.468 I srv load_model: prompt cache RAM enabled: limit_mib=8192
|
| 21 |
+
0.24.567.470 I srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
|
| 22 |
+
0.24.567.551 I srv init: idle slots will be saved to prompt cache upon starting a new task
|
| 23 |
+
0.24.579.260 I init: chat template, example_format: '<|im_start|>system
|
| 24 |
+
You are a helpful assistant<|im_end|>
|
| 25 |
+
<|im_start|>user
|
| 26 |
+
Hello<|im_end|>
|
| 27 |
+
<|im_start|>assistant
|
| 28 |
+
<think>
|
| 29 |
+
|
| 30 |
+
</think>
|
| 31 |
+
|
| 32 |
+
Hi there<|im_end|>
|
| 33 |
+
<|im_start|>user
|
| 34 |
+
How are you?<|im_end|>
|
| 35 |
+
<|im_start|>assistant
|
| 36 |
+
<think>'
|
| 37 |
+
0.24.587.661 I srv init: init: chat template, thinking = 1
|
| 38 |
+
0.24.587.712 I srv main: model loaded
|
| 39 |
+
0.24.587.715 I srv main: server is listening on http://127.0.0.1:18600
|
| 40 |
+
0.24.587.720 I srv update_slots: all slots are idle
|
| 41 |
+
0.26.148.362 I srv params_from_: Chat format: peg-native
|
| 42 |
+
0.26.149.988 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1
|
| 43 |
+
0.26.149.996 I srv get_availabl: updating prompt cache
|
| 44 |
+
0.26.150.007 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000
|
| 45 |
+
0.26.150.015 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 65536 tokens, 8589934592 est)
|
| 46 |
+
0.26.150.018 I srv get_availabl: prompt cache update took 0.02 ms
|
| 47 |
+
0.26.150.786 I reasoning-budget: activated, budget=2147483647 tokens
|
| 48 |
+
0.26.150.816 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0
|
| 49 |
+
0.26.871.165 I slot create_check: id 0 | task 0 | created context checkpoint 1 of 32 (pos_min = 416, pos_max = 416, n_tokens = 417, size = 62.813 MiB)
|
| 50 |
+
0.27.133.131 I reasoning-budget: deactivated (natural end)
|
| 51 |
+
0.28.001.263 I slot print_timing: id 0 | task 0 |
|
| 52 |
+
prompt eval time = 773.86 ms / 421 tokens ( 1.84 ms per token, 544.03 tokens per second)
|
| 53 |
+
eval time = 1076.54 ms / 48 tokens ( 22.43 ms per token, 44.59 tokens per second)
|
| 54 |
+
total time = 1850.39 ms / 469 tokens
|
| 55 |
+
0.28.001.472 I slot release: id 0 | task 0 | stop processing: n_tokens = 468, truncated = 0
|
| 56 |
+
0.28.001.492 I srv update_slots: all slots are idle
|
| 57 |
+
0.28.065.185 I srv params_from_: Chat format: peg-native
|
| 58 |
+
0.28.065.968 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.908 (> 0.100 thold), f_keep = 0.865
|
| 59 |
+
0.28.066.222 I reasoning-budget: activated, budget=2147483647 tokens
|
| 60 |
+
0.28.066.277 I slot launch_slot_: id 0 | task 50 | processing task, is_child = 0
|
| 61 |
+
0.28.066.290 W slot update_slots: id 0 | task 50 | n_past = 405, slot.prompt.tokens.size() = 468, seq_id = 0, pos_min = 467, n_swa = 0
|
| 62 |
+
0.28.066.292 I slot update_slots: id 0 | task 50 | Checking checkpoint with [416, 416] against 405...
|
| 63 |
+
0.28.066.293 W slot update_slots: id 0 | task 50 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 64 |
+
0.28.066.297 W slot update_slots: id 0 | task 50 | erased invalidated context checkpoint (pos_min = 416, pos_max = 416, n_tokens = 417, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 65 |
+
0.28.663.976 I slot create_check: id 0 | task 50 | created context checkpoint 1 of 32 (pos_min = 441, pos_max = 441, n_tokens = 442, size = 62.813 MiB)
|
| 66 |
+
0.29.408.376 I reasoning-budget: deactivated (natural end)
|
| 67 |
+
0.30.568.909 I slot print_timing: id 0 | task 50 |
|
| 68 |
+
prompt eval time = 632.88 ms / 446 tokens ( 1.42 ms per token, 704.71 tokens per second)
|
| 69 |
+
eval time = 1869.72 ms / 81 tokens ( 23.08 ms per token, 43.32 tokens per second)
|
| 70 |
+
total time = 2502.60 ms / 527 tokens
|
| 71 |
+
0.30.568.986 I slot release: id 0 | task 50 | stop processing: n_tokens = 526, truncated = 0
|
| 72 |
+
0.30.569.017 I srv update_slots: all slots are idle
|
| 73 |
+
0.30.584.452 I srv params_from_: Chat format: peg-native
|
| 74 |
+
0.30.584.982 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.770
|
| 75 |
+
0.30.585.566 I reasoning-budget: activated, budget=2147483647 tokens
|
| 76 |
+
0.30.585.633 I slot launch_slot_: id 0 | task 133 | processing task, is_child = 0
|
| 77 |
+
0.30.585.660 W slot update_slots: id 0 | task 133 | n_past = 405, slot.prompt.tokens.size() = 526, seq_id = 0, pos_min = 525, n_swa = 0
|
| 78 |
+
0.30.585.663 I slot update_slots: id 0 | task 133 | Checking checkpoint with [441, 441] against 405...
|
| 79 |
+
0.30.585.666 W slot update_slots: id 0 | task 133 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 80 |
+
0.30.585.671 W slot update_slots: id 0 | task 133 | erased invalidated context checkpoint (pos_min = 441, pos_max = 441, n_tokens = 442, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 81 |
+
0.31.126.980 I slot create_check: id 0 | task 133 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
|
| 82 |
+
0.31.358.672 I reasoning-budget: deactivated (natural end)
|
| 83 |
+
0.32.155.583 I slot print_timing: id 0 | task 133 |
|
| 84 |
+
prompt eval time = 578.11 ms / 424 tokens ( 1.36 ms per token, 733.42 tokens per second)
|
| 85 |
+
eval time = 991.80 ms / 48 tokens ( 20.66 ms per token, 48.40 tokens per second)
|
| 86 |
+
total time = 1569.91 ms / 472 tokens
|
| 87 |
+
0.32.155.659 I slot release: id 0 | task 133 | stop processing: n_tokens = 471, truncated = 0
|
| 88 |
+
0.32.155.694 I srv update_slots: all slots are idle
|
| 89 |
+
0.32.187.576 I srv params_from_: Chat format: peg-native
|
| 90 |
+
0.32.188.058 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.860
|
| 91 |
+
0.32.188.654 I reasoning-budget: activated, budget=2147483647 tokens
|
| 92 |
+
0.32.188.751 I slot launch_slot_: id 0 | task 183 | processing task, is_child = 0
|
| 93 |
+
0.32.188.776 W slot update_slots: id 0 | task 183 | n_past = 405, slot.prompt.tokens.size() = 471, seq_id = 0, pos_min = 470, n_swa = 0
|
| 94 |
+
0.32.188.780 I slot update_slots: id 0 | task 183 | Checking checkpoint with [419, 419] against 405...
|
| 95 |
+
0.32.188.782 W slot update_slots: id 0 | task 183 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 96 |
+
0.32.188.790 W slot update_slots: id 0 | task 183 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 97 |
+
0.32.763.234 I slot create_check: id 0 | task 183 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
|
| 98 |
+
0.32.827.843 I reasoning-budget: deactivated (natural end)
|
| 99 |
+
0.32.942.756 I slot print_timing: id 0 | task 183 |
|
| 100 |
+
prompt eval time = 610.58 ms / 424 tokens ( 1.44 ms per token, 694.43 tokens per second)
|
| 101 |
+
eval time = 143.38 ms / 7 tokens ( 20.48 ms per token, 48.82 tokens per second)
|
| 102 |
+
total time = 753.96 ms / 431 tokens
|
| 103 |
+
0.32.942.940 I slot release: id 0 | task 183 | stop processing: n_tokens = 430, truncated = 0
|
| 104 |
+
0.32.942.992 I srv update_slots: all slots are idle
|
| 105 |
+
0.32.959.301 I srv params_from_: Chat format: peg-native
|
| 106 |
+
0.32.959.816 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.964 (> 0.100 thold), f_keep = 0.944
|
| 107 |
+
0.32.960.404 I reasoning-budget: activated, budget=2147483647 tokens
|
| 108 |
+
0.32.960.503 I slot launch_slot_: id 0 | task 192 | processing task, is_child = 0
|
| 109 |
+
0.32.960.525 W slot update_slots: id 0 | task 192 | n_past = 406, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
|
| 110 |
+
0.32.960.529 I slot update_slots: id 0 | task 192 | Checking checkpoint with [419, 419] against 406...
|
| 111 |
+
0.32.960.530 W slot update_slots: id 0 | task 192 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 112 |
+
0.32.960.538 W slot update_slots: id 0 | task 192 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 113 |
+
0.33.460.168 I slot create_check: id 0 | task 192 | created context checkpoint 1 of 32 (pos_min = 416, pos_max = 416, n_tokens = 417, size = 62.813 MiB)
|
| 114 |
+
0.33.682.213 I reasoning-budget: deactivated (natural end)
|
| 115 |
+
0.34.478.330 I slot print_timing: id 0 | task 192 |
|
| 116 |
+
prompt eval time = 535.13 ms / 421 tokens ( 1.27 ms per token, 786.73 tokens per second)
|
| 117 |
+
eval time = 982.66 ms / 50 tokens ( 19.65 ms per token, 50.88 tokens per second)
|
| 118 |
+
total time = 1517.79 ms / 471 tokens
|
| 119 |
+
0.34.478.426 I slot release: id 0 | task 192 | stop processing: n_tokens = 470, truncated = 0
|
| 120 |
+
0.34.478.463 I srv update_slots: all slots are idle
|
| 121 |
+
0.34.496.222 I srv params_from_: Chat format: peg-native
|
| 122 |
+
0.34.496.734 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.942 (> 0.100 thold), f_keep = 1.000
|
| 123 |
+
0.34.496.948 I reasoning-budget: activated, budget=2147483647 tokens
|
| 124 |
+
0.34.496.989 I slot launch_slot_: id 0 | task 244 | processing task, is_child = 0
|
| 125 |
+
0.34.639.263 I slot create_check: id 0 | task 244 | created context checkpoint 2 of 32 (pos_min = 494, pos_max = 494, n_tokens = 495, size = 62.813 MiB)
|
| 126 |
+
0.34.705.972 I reasoning-budget: deactivated (natural end)
|
| 127 |
+
0.35.076.072 I slot print_timing: id 0 | task 244 |
|
| 128 |
+
prompt eval time = 177.70 ms / 29 tokens ( 6.13 ms per token, 163.19 tokens per second)
|
| 129 |
+
eval time = 401.35 ms / 17 tokens ( 23.61 ms per token, 42.36 tokens per second)
|
| 130 |
+
total time = 579.05 ms / 46 tokens
|
| 131 |
+
0.35.076.179 I slot release: id 0 | task 244 | stop processing: n_tokens = 515, truncated = 0
|
| 132 |
+
0.35.076.214 I srv update_slots: all slots are idle
|
| 133 |
+
0.35.091.806 I srv params_from_: Chat format: peg-native
|
| 134 |
+
0.35.092.529 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.974 (> 0.100 thold), f_keep = 0.796
|
| 135 |
+
0.35.093.006 I reasoning-budget: activated, budget=2147483647 tokens
|
| 136 |
+
0.35.093.087 I slot launch_slot_: id 0 | task 263 | processing task, is_child = 0
|
| 137 |
+
0.35.093.108 W slot update_slots: id 0 | task 263 | n_past = 410, slot.prompt.tokens.size() = 515, seq_id = 0, pos_min = 514, n_swa = 0
|
| 138 |
+
0.35.093.110 I slot update_slots: id 0 | task 263 | Checking checkpoint with [494, 494] against 410...
|
| 139 |
+
0.35.093.111 I slot update_slots: id 0 | task 263 | Checking checkpoint with [416, 416] against 410...
|
| 140 |
+
0.35.093.113 W slot update_slots: id 0 | task 263 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 141 |
+
0.35.093.119 W slot update_slots: id 0 | task 263 | erased invalidated context checkpoint (pos_min = 416, pos_max = 416, n_tokens = 417, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 142 |
+
0.35.094.217 W slot update_slots: id 0 | task 263 | erased invalidated context checkpoint (pos_min = 494, pos_max = 494, n_tokens = 495, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 143 |
+
0.35.626.770 I slot create_check: id 0 | task 263 | created context checkpoint 1 of 32 (pos_min = 416, pos_max = 416, n_tokens = 417, size = 62.813 MiB)
|
| 144 |
+
0.36.142.008 I reasoning-budget: deactivated (natural end)
|
| 145 |
+
0.36.985.537 I slot print_timing: id 0 | task 263 |
|
| 146 |
+
prompt eval time = 581.67 ms / 421 tokens ( 1.38 ms per token, 723.78 tokens per second)
|
| 147 |
+
eval time = 1310.72 ms / 64 tokens ( 20.48 ms per token, 48.83 tokens per second)
|
| 148 |
+
total time = 1892.39 ms / 485 tokens
|
| 149 |
+
0.36.985.731 I slot release: id 0 | task 263 | stop processing: n_tokens = 484, truncated = 0
|
| 150 |
+
0.36.985.791 I srv update_slots: all slots are idle
|
| 151 |
+
0.37.040.565 I srv params_from_: Chat format: peg-native
|
| 152 |
+
0.37.041.116 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.938 (> 0.100 thold), f_keep = 0.837
|
| 153 |
+
0.37.041.402 I reasoning-budget: activated, budget=2147483647 tokens
|
| 154 |
+
0.37.041.512 I slot launch_slot_: id 0 | task 329 | processing task, is_child = 0
|
| 155 |
+
0.37.041.525 W slot update_slots: id 0 | task 329 | n_past = 405, slot.prompt.tokens.size() = 484, seq_id = 0, pos_min = 483, n_swa = 0
|
| 156 |
+
0.37.041.528 I slot update_slots: id 0 | task 329 | Checking checkpoint with [416, 416] against 405...
|
| 157 |
+
0.37.041.528 W slot update_slots: id 0 | task 329 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 158 |
+
0.37.041.531 W slot update_slots: id 0 | task 329 | erased invalidated context checkpoint (pos_min = 416, pos_max = 416, n_tokens = 417, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 159 |
+
0.37.608.593 I slot create_check: id 0 | task 329 | created context checkpoint 1 of 32 (pos_min = 427, pos_max = 427, n_tokens = 428, size = 62.813 MiB)
|
| 160 |
+
0.38.415.584 I reasoning-budget: deactivated (natural end)
|
| 161 |
+
0.40.014.106 I slot print_timing: id 0 | task 329 | n_decoded = 100, tg = 42.63 t/s
|
| 162 |
+
0.40.267.618 I slot print_timing: id 0 | task 329 |
|
| 163 |
+
prompt eval time = 626.94 ms / 432 tokens ( 1.45 ms per token, 689.06 tokens per second)
|
| 164 |
+
eval time = 2599.14 ms / 113 tokens ( 23.00 ms per token, 43.48 tokens per second)
|
| 165 |
+
total time = 3226.08 ms / 545 tokens
|
| 166 |
+
0.40.267.691 I slot release: id 0 | task 329 | stop processing: n_tokens = 544, truncated = 0
|
| 167 |
+
0.40.267.739 I srv update_slots: all slots are idle
|
| 168 |
+
0.40.291.843 I srv params_from_: Chat format: peg-native
|
| 169 |
+
0.40.292.295 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.744
|
| 170 |
+
0.40.292.805 I reasoning-budget: activated, budget=2147483647 tokens
|
| 171 |
+
0.40.292.809 I reasoning-budget: deactivated (natural end)
|
| 172 |
+
0.40.292.882 I slot launch_slot_: id 0 | task 444 | processing task, is_child = 0
|
| 173 |
+
0.40.292.903 W slot update_slots: id 0 | task 444 | n_past = 405, slot.prompt.tokens.size() = 544, seq_id = 0, pos_min = 543, n_swa = 0
|
| 174 |
+
0.40.292.906 I slot update_slots: id 0 | task 444 | Checking checkpoint with [427, 427] against 405...
|
| 175 |
+
0.40.292.907 W slot update_slots: id 0 | task 444 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 176 |
+
0.40.292.912 W slot update_slots: id 0 | task 444 | erased invalidated context checkpoint (pos_min = 427, pos_max = 427, n_tokens = 428, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 177 |
+
0.40.808.790 I slot create_check: id 0 | task 444 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
|
| 178 |
+
0.41.645.500 I slot print_timing: id 0 | task 444 |
|
| 179 |
+
prompt eval time = 551.51 ms / 424 tokens ( 1.30 ms per token, 768.79 tokens per second)
|
| 180 |
+
eval time = 801.07 ms / 39 tokens ( 20.54 ms per token, 48.68 tokens per second)
|
| 181 |
+
total time = 1352.58 ms / 463 tokens
|
| 182 |
+
0.41.645.589 I slot release: id 0 | task 444 | stop processing: n_tokens = 462, truncated = 0
|
| 183 |
+
0.41.645.620 I srv update_slots: all slots are idle
|
| 184 |
+
0.41.661.261 I srv params_from_: Chat format: peg-native
|
| 185 |
+
0.41.661.789 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.902 (> 0.100 thold), f_keep = 0.877
|
| 186 |
+
0.41.662.003 I reasoning-budget: activated, budget=2147483647 tokens
|
| 187 |
+
0.41.662.005 I reasoning-budget: deactivated (natural end)
|
| 188 |
+
0.41.662.045 I slot launch_slot_: id 0 | task 485 | processing task, is_child = 0
|
| 189 |
+
0.41.662.055 W slot update_slots: id 0 | task 485 | n_past = 405, slot.prompt.tokens.size() = 462, seq_id = 0, pos_min = 461, n_swa = 0
|
| 190 |
+
0.41.662.056 I slot update_slots: id 0 | task 485 | Checking checkpoint with [419, 419] against 405...
|
| 191 |
+
0.41.662.057 W slot update_slots: id 0 | task 485 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 192 |
+
0.41.662.060 W slot update_slots: id 0 | task 485 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 193 |
+
0.42.169.075 I slot create_check: id 0 | task 485 | created context checkpoint 1 of 32 (pos_min = 444, pos_max = 444, n_tokens = 445, size = 62.813 MiB)
|
| 194 |
+
0.44.038.103 I slot print_timing: id 0 | task 485 |
|
| 195 |
+
prompt eval time = 548.61 ms / 449 tokens ( 1.22 ms per token, 818.43 tokens per second)
|
| 196 |
+
eval time = 1827.42 ms / 87 tokens ( 21.00 ms per token, 47.61 tokens per second)
|
| 197 |
+
total time = 2376.03 ms / 536 tokens
|
| 198 |
+
0.44.038.196 I slot release: id 0 | task 485 | stop processing: n_tokens = 535, truncated = 0
|
| 199 |
+
0.44.038.227 I srv update_slots: all slots are idle
|
| 200 |
+
0.44.070.792 I srv params_from_: Chat format: peg-native
|
| 201 |
+
0.44.071.263 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.948 (> 0.100 thold), f_keep = 0.757
|
| 202 |
+
0.44.071.950 I reasoning-budget: activated, budget=2147483647 tokens
|
| 203 |
+
0.44.071.955 I reasoning-budget: deactivated (natural end)
|
| 204 |
+
0.44.072.053 I slot launch_slot_: id 0 | task 574 | processing task, is_child = 0
|
| 205 |
+
0.44.072.077 W slot update_slots: id 0 | task 574 | n_past = 405, slot.prompt.tokens.size() = 535, seq_id = 0, pos_min = 534, n_swa = 0
|
| 206 |
+
0.44.072.081 I slot update_slots: id 0 | task 574 | Checking checkpoint with [444, 444] against 405...
|
| 207 |
+
0.44.072.083 W slot update_slots: id 0 | task 574 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 208 |
+
0.44.072.088 W slot update_slots: id 0 | task 574 | erased invalidated context checkpoint (pos_min = 444, pos_max = 444, n_tokens = 445, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 209 |
+
0.44.616.066 I slot create_check: id 0 | task 574 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
|
| 210 |
+
0.45.562.940 I slot print_timing: id 0 | task 574 |
|
| 211 |
+
prompt eval time = 603.10 ms / 427 tokens ( 1.41 ms per token, 708.01 tokens per second)
|
| 212 |
+
eval time = 887.75 ms / 39 tokens ( 22.76 ms per token, 43.93 tokens per second)
|
| 213 |
+
total time = 1490.84 ms / 466 tokens
|
| 214 |
+
0.45.563.042 I slot release: id 0 | task 574 | stop processing: n_tokens = 465, truncated = 0
|
| 215 |
+
0.45.563.082 I srv update_slots: all slots are idle
|
| 216 |
+
0.45.579.534 I srv params_from_: Chat format: peg-native
|
| 217 |
+
0.45.580.094 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.948 (> 0.100 thold), f_keep = 0.871
|
| 218 |
+
0.45.580.642 I reasoning-budget: activated, budget=2147483647 tokens
|
| 219 |
+
0.45.580.646 I reasoning-budget: deactivated (natural end)
|
| 220 |
+
0.45.582.811 I slot launch_slot_: id 0 | task 615 | processing task, is_child = 0
|
| 221 |
+
0.45.582.848 W slot update_slots: id 0 | task 615 | n_past = 405, slot.prompt.tokens.size() = 465, seq_id = 0, pos_min = 464, n_swa = 0
|
| 222 |
+
0.45.582.849 I slot update_slots: id 0 | task 615 | Checking checkpoint with [422, 422] against 405...
|
| 223 |
+
0.45.582.851 W slot update_slots: id 0 | task 615 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 224 |
+
0.45.582.860 W slot update_slots: id 0 | task 615 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 225 |
+
0.46.041.817 I slot create_check: id 0 | task 615 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
|
| 226 |
+
0.46.171.798 I slot print_timing: id 0 | task 615 |
|
| 227 |
+
prompt eval time = 507.02 ms / 427 tokens ( 1.19 ms per token, 842.17 tokens per second)
|
| 228 |
+
eval time = 81.92 ms / 4 tokens ( 20.48 ms per token, 48.83 tokens per second)
|
| 229 |
+
total time = 588.94 ms / 431 tokens
|
| 230 |
+
0.46.171.890 I slot release: id 0 | task 615 | stop processing: n_tokens = 430, truncated = 0
|
| 231 |
+
0.46.171.920 I srv update_slots: all slots are idle
|
| 232 |
+
0.46.188.295 I srv params_from_: Chat format: peg-native
|
| 233 |
+
0.46.188.729 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.958 (> 0.100 thold), f_keep = 0.944
|
| 234 |
+
0.46.189.261 I reasoning-budget: activated, budget=2147483647 tokens
|
| 235 |
+
0.46.189.264 I reasoning-budget: deactivated (natural end)
|
| 236 |
+
0.46.189.348 I slot launch_slot_: id 0 | task 621 | processing task, is_child = 0
|
| 237 |
+
0.46.189.367 W slot update_slots: id 0 | task 621 | n_past = 406, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
|
| 238 |
+
0.46.189.370 I slot update_slots: id 0 | task 621 | Checking checkpoint with [422, 422] against 406...
|
| 239 |
+
0.46.189.372 W slot update_slots: id 0 | task 621 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 240 |
+
0.46.189.377 W slot update_slots: id 0 | task 621 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 241 |
+
0.46.716.804 I slot create_check: id 0 | task 621 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
|
| 242 |
+
0.47.566.188 I slot print_timing: id 0 | task 621 |
|
| 243 |
+
prompt eval time = 563.24 ms / 424 tokens ( 1.33 ms per token, 752.78 tokens per second)
|
| 244 |
+
eval time = 813.55 ms / 40 tokens ( 20.34 ms per token, 49.17 tokens per second)
|
| 245 |
+
total time = 1376.79 ms / 464 tokens
|
| 246 |
+
0.47.566.402 I slot release: id 0 | task 621 | stop processing: n_tokens = 463, truncated = 0
|
| 247 |
+
0.47.566.467 I srv update_slots: all slots are idle
|
| 248 |
+
0.47.589.246 I srv params_from_: Chat format: peg-native
|
| 249 |
+
0.47.589.845 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.935 (> 0.100 thold), f_keep = 1.000
|
| 250 |
+
0.47.590.510 I reasoning-budget: activated, budget=2147483647 tokens
|
| 251 |
+
0.47.590.517 I reasoning-budget: deactivated (natural end)
|
| 252 |
+
0.47.590.637 I slot launch_slot_: id 0 | task 663 | processing task, is_child = 0
|
| 253 |
+
0.47.740.622 I slot create_check: id 0 | task 663 | created context checkpoint 2 of 32 (pos_min = 490, pos_max = 490, n_tokens = 491, size = 62.813 MiB)
|
| 254 |
+
0.48.115.507 I slot print_timing: id 0 | task 663 |
|
| 255 |
+
prompt eval time = 197.73 ms / 32 tokens ( 6.18 ms per token, 161.84 tokens per second)
|
| 256 |
+
eval time = 327.11 ms / 15 tokens ( 21.81 ms per token, 45.86 tokens per second)
|
| 257 |
+
total time = 524.83 ms / 47 tokens
|
| 258 |
+
0.48.115.614 I slot release: id 0 | task 663 | stop processing: n_tokens = 509, truncated = 0
|
| 259 |
+
0.48.115.648 I srv update_slots: all slots are idle
|
| 260 |
+
0.48.182.251 I srv params_from_: Chat format: peg-native
|
| 261 |
+
0.48.182.770 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.967 (> 0.100 thold), f_keep = 0.806
|
| 262 |
+
0.48.183.002 I reasoning-budget: activated, budget=2147483647 tokens
|
| 263 |
+
0.48.183.003 I reasoning-budget: deactivated (natural end)
|
| 264 |
+
0.48.183.055 I slot launch_slot_: id 0 | task 680 | processing task, is_child = 0
|
| 265 |
+
0.48.183.066 W slot update_slots: id 0 | task 680 | n_past = 410, slot.prompt.tokens.size() = 509, seq_id = 0, pos_min = 508, n_swa = 0
|
| 266 |
+
0.48.183.067 I slot update_slots: id 0 | task 680 | Checking checkpoint with [490, 490] against 410...
|
| 267 |
+
0.48.183.067 I slot update_slots: id 0 | task 680 | Checking checkpoint with [419, 419] against 410...
|
| 268 |
+
0.48.183.069 W slot update_slots: id 0 | task 680 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 269 |
+
0.48.183.072 W slot update_slots: id 0 | task 680 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 270 |
+
0.48.184.972 W slot update_slots: id 0 | task 680 | erased invalidated context checkpoint (pos_min = 490, pos_max = 490, n_tokens = 491, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 271 |
+
0.48.742.624 I slot create_check: id 0 | task 680 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
|
| 272 |
+
0.49.641.904 I slot print_timing: id 0 | task 680 |
|
| 273 |
+
prompt eval time = 599.46 ms / 424 tokens ( 1.41 ms per token, 707.31 tokens per second)
|
| 274 |
+
eval time = 859.36 ms / 40 tokens ( 21.48 ms per token, 46.55 tokens per second)
|
| 275 |
+
total time = 1458.81 ms / 464 tokens
|
| 276 |
+
0.49.641.977 I slot release: id 0 | task 680 | stop processing: n_tokens = 463, truncated = 0
|
| 277 |
+
0.49.642.005 I srv update_slots: all slots are idle
|
| 278 |
+
0.49.654.625 I srv params_from_: Chat format: peg-native
|
| 279 |
+
0.49.655.158 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.931 (> 0.100 thold), f_keep = 0.875
|
| 280 |
+
0.49.655.408 I reasoning-budget: activated, budget=2147483647 tokens
|
| 281 |
+
0.49.655.411 I reasoning-budget: deactivated (natural end)
|
| 282 |
+
0.49.655.459 I slot launch_slot_: id 0 | task 722 | processing task, is_child = 0
|
| 283 |
+
0.49.655.469 W slot update_slots: id 0 | task 722 | n_past = 405, slot.prompt.tokens.size() = 463, seq_id = 0, pos_min = 462, n_swa = 0
|
| 284 |
+
0.49.655.470 I slot update_slots: id 0 | task 722 | Checking checkpoint with [419, 419] against 405...
|
| 285 |
+
0.49.655.470 W slot update_slots: id 0 | task 722 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 286 |
+
0.49.655.472 W slot update_slots: id 0 | task 722 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 287 |
+
0.50.147.360 I slot create_check: id 0 | task 722 | created context checkpoint 1 of 32 (pos_min = 430, pos_max = 430, n_tokens = 431, size = 62.813 MiB)
|
| 288 |
+
0.51.806.641 I slot print_timing: id 0 | task 722 |
|
| 289 |
+
prompt eval time = 536.58 ms / 435 tokens ( 1.23 ms per token, 810.69 tokens per second)
|
| 290 |
+
eval time = 1614.55 ms / 80 tokens ( 20.18 ms per token, 49.55 tokens per second)
|
| 291 |
+
total time = 2151.13 ms / 515 tokens
|
| 292 |
+
0.51.807.033 I slot release: id 0 | task 722 | stop processing: n_tokens = 514, truncated = 0
|
| 293 |
+
0.51.807.128 I srv update_slots: all slots are idle
|
| 294 |
+
0.51.808.793 I srv operator(): operator(): cleaning up before exit...
|
recipe/logs/b_n-tools-q106.log
ADDED
|
@@ -0,0 +1,302 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
0.00.039.657 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
|
| 2 |
+
0.00.039.660 I device_info:
|
| 3 |
+
0.00.039.707 I - ROCm0 : AMD Radeon Graphics (131072 MiB, 125232 MiB free)
|
| 4 |
+
0.00.039.793 I - Vulkan0 : AMD Radeon Graphics (RADV GFX1151) (132096 MiB, 131926 MiB free)
|
| 5 |
+
0.00.039.797 I - CPU : AMD RYZEN AI MAX+ 395 w/ Radeon 8060S (127438 MiB, 127438 MiB free)
|
| 6 |
+
0.00.039.840 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
|
| 7 |
+
0.00.039.855 I srv init: running without SSL
|
| 8 |
+
0.00.039.876 I srv init: using 31 threads for HTTP server
|
| 9 |
+
0.00.039.877 I srv init: the WebUI is disabled
|
| 10 |
+
0.00.039.934 I srv start: binding port with default address family
|
| 11 |
+
0.00.041.085 I srv main: loading model
|
| 12 |
+
0.00.041.087 I srv load_model: loading model '/mnt/models/nex-n2.5-mini/out/Nex-N2.5-mini-Q4_0_ROCMFP4_STRIX_LEAN.gguf'
|
| 13 |
+
0.00.075.504 W llama_model_loader: direct I/O is enabled, disabling mmap
|
| 14 |
+
0.20.409.072 W llama_context: n_ctx_seq (65536) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
|
| 15 |
+
0.20.574.487 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|
| 16 |
+
0.20.774.364 I srv load_model: initializing slots, n_slots = 1
|
| 17 |
+
0.20.940.846 W srv load_model: speculative decoding will use checkpoints
|
| 18 |
+
0.20.940.851 W common_speculative_init: no implementations specified for speculative decoding
|
| 19 |
+
0.20.940.852 I slot load_model: id 0 | task -1 | new slot, n_ctx = 65536
|
| 20 |
+
0.20.940.886 I srv load_model: prompt cache RAM enabled: limit_mib=8192
|
| 21 |
+
0.20.940.894 I srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
|
| 22 |
+
0.20.940.907 I srv init: idle slots will be saved to prompt cache upon starting a new task
|
| 23 |
+
0.20.950.043 I init: chat template, example_format: '<|im_start|>system
|
| 24 |
+
You are a helpful assistant<|im_end|>
|
| 25 |
+
<|im_start|>user
|
| 26 |
+
Hello<|im_end|>
|
| 27 |
+
<|im_start|>assistant
|
| 28 |
+
<think>
|
| 29 |
+
|
| 30 |
+
</think>
|
| 31 |
+
|
| 32 |
+
Hi there<|im_end|>
|
| 33 |
+
<|im_start|>user
|
| 34 |
+
How are you?<|im_end|>
|
| 35 |
+
<|im_start|>assistant
|
| 36 |
+
<think>'
|
| 37 |
+
0.20.957.344 I srv init: init: chat template, thinking = 1
|
| 38 |
+
0.20.957.368 I srv main: model loaded
|
| 39 |
+
0.20.957.372 I srv main: server is listening on http://127.0.0.1:18600
|
| 40 |
+
0.20.957.374 I srv update_slots: all slots are idle
|
| 41 |
+
0.22.001.528 I srv params_from_: Chat format: peg-native
|
| 42 |
+
0.22.001.842 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1
|
| 43 |
+
0.22.001.845 I srv get_availabl: updating prompt cache
|
| 44 |
+
0.22.001.851 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000
|
| 45 |
+
0.22.001.856 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 65536 tokens, 8589934592 est)
|
| 46 |
+
0.22.001.858 I srv get_availabl: prompt cache update took 0.01 ms
|
| 47 |
+
0.22.010.813 I reasoning-budget: activated, budget=2147483647 tokens
|
| 48 |
+
0.22.010.830 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0
|
| 49 |
+
0.22.354.924 I slot create_check: id 0 | task 0 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
|
| 50 |
+
0.22.608.502 I reasoning-budget: deactivated (natural end)
|
| 51 |
+
0.23.202.843 I slot print_timing: id 0 | task 0 |
|
| 52 |
+
prompt eval time = 372.12 ms / 422 tokens ( 0.88 ms per token, 1134.04 tokens per second)
|
| 53 |
+
eval time = 819.87 ms / 55 tokens ( 14.91 ms per token, 67.08 tokens per second)
|
| 54 |
+
total time = 1191.99 ms / 477 tokens
|
| 55 |
+
0.23.202.893 I slot release: id 0 | task 0 | stop processing: n_tokens = 476, truncated = 0
|
| 56 |
+
0.23.202.897 I srv update_slots: all slots are idle
|
| 57 |
+
0.23.213.207 I srv params_from_: Chat format: peg-native
|
| 58 |
+
0.23.213.731 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.906 (> 0.100 thold), f_keep = 0.851
|
| 59 |
+
0.23.213.930 I reasoning-budget: activated, budget=2147483647 tokens
|
| 60 |
+
0.23.213.965 I slot launch_slot_: id 0 | task 57 | processing task, is_child = 0
|
| 61 |
+
0.23.213.971 W slot update_slots: id 0 | task 57 | n_past = 405, slot.prompt.tokens.size() = 476, seq_id = 0, pos_min = 475, n_swa = 0
|
| 62 |
+
0.23.213.971 I slot update_slots: id 0 | task 57 | Checking checkpoint with [417, 417] against 405...
|
| 63 |
+
0.23.213.973 W slot update_slots: id 0 | task 57 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 64 |
+
0.23.213.975 W slot update_slots: id 0 | task 57 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 65 |
+
0.23.551.876 I slot create_check: id 0 | task 57 | created context checkpoint 1 of 32 (pos_min = 442, pos_max = 442, n_tokens = 443, size = 62.813 MiB)
|
| 66 |
+
0.24.213.049 I reasoning-budget: deactivated (natural end)
|
| 67 |
+
0.24.949.926 I slot print_timing: id 0 | task 57 |
|
| 68 |
+
prompt eval time = 365.53 ms / 447 tokens ( 0.82 ms per token, 1222.88 tokens per second)
|
| 69 |
+
eval time = 1370.41 ms / 92 tokens ( 14.90 ms per token, 67.13 tokens per second)
|
| 70 |
+
total time = 1735.95 ms / 539 tokens
|
| 71 |
+
0.24.949.971 I slot release: id 0 | task 57 | stop processing: n_tokens = 538, truncated = 0
|
| 72 |
+
0.24.949.992 I srv update_slots: all slots are idle
|
| 73 |
+
0.24.959.419 I srv params_from_: Chat format: peg-native
|
| 74 |
+
0.24.959.740 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.953 (> 0.100 thold), f_keep = 0.753
|
| 75 |
+
0.24.959.952 I reasoning-budget: activated, budget=2147483647 tokens
|
| 76 |
+
0.24.959.981 I slot launch_slot_: id 0 | task 151 | processing task, is_child = 0
|
| 77 |
+
0.24.959.987 W slot update_slots: id 0 | task 151 | n_past = 405, slot.prompt.tokens.size() = 538, seq_id = 0, pos_min = 537, n_swa = 0
|
| 78 |
+
0.24.959.987 I slot update_slots: id 0 | task 151 | Checking checkpoint with [442, 442] against 405...
|
| 79 |
+
0.24.959.988 W slot update_slots: id 0 | task 151 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 80 |
+
0.24.959.991 W slot update_slots: id 0 | task 151 | erased invalidated context checkpoint (pos_min = 442, pos_max = 442, n_tokens = 443, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 81 |
+
0.25.284.086 I slot create_check: id 0 | task 151 | created context checkpoint 1 of 32 (pos_min = 420, pos_max = 420, n_tokens = 421, size = 62.813 MiB)
|
| 82 |
+
0.25.471.508 I reasoning-budget: deactivated (natural end)
|
| 83 |
+
0.26.064.107 I slot print_timing: id 0 | task 151 |
|
| 84 |
+
prompt eval time = 351.76 ms / 425 tokens ( 0.83 ms per token, 1208.20 tokens per second)
|
| 85 |
+
eval time = 752.35 ms / 51 tokens ( 14.75 ms per token, 67.79 tokens per second)
|
| 86 |
+
total time = 1104.11 ms / 476 tokens
|
| 87 |
+
0.26.064.151 I slot release: id 0 | task 151 | stop processing: n_tokens = 475, truncated = 0
|
| 88 |
+
0.26.064.171 I srv update_slots: all slots are idle
|
| 89 |
+
0.26.074.471 I srv params_from_: Chat format: peg-native
|
| 90 |
+
0.26.074.789 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.953 (> 0.100 thold), f_keep = 0.853
|
| 91 |
+
0.26.074.999 I reasoning-budget: activated, budget=2147483647 tokens
|
| 92 |
+
0.26.075.029 I slot launch_slot_: id 0 | task 204 | processing task, is_child = 0
|
| 93 |
+
0.26.075.034 W slot update_slots: id 0 | task 204 | n_past = 405, slot.prompt.tokens.size() = 475, seq_id = 0, pos_min = 474, n_swa = 0
|
| 94 |
+
0.26.075.045 I slot update_slots: id 0 | task 204 | Checking checkpoint with [420, 420] against 405...
|
| 95 |
+
0.26.075.048 W slot update_slots: id 0 | task 204 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 96 |
+
0.26.075.050 W slot update_slots: id 0 | task 204 | erased invalidated context checkpoint (pos_min = 420, pos_max = 420, n_tokens = 421, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 97 |
+
0.26.399.124 I slot create_check: id 0 | task 204 | created context checkpoint 1 of 32 (pos_min = 420, pos_max = 420, n_tokens = 421, size = 62.813 MiB)
|
| 98 |
+
0.26.601.525 I reasoning-budget: deactivated (natural end)
|
| 99 |
+
0.26.675.559 I slot print_timing: id 0 | task 204 |
|
| 100 |
+
prompt eval time = 351.60 ms / 425 tokens ( 0.83 ms per token, 1208.76 tokens per second)
|
| 101 |
+
eval time = 248.92 ms / 17 tokens ( 14.64 ms per token, 68.30 tokens per second)
|
| 102 |
+
total time = 600.52 ms / 442 tokens
|
| 103 |
+
0.26.675.604 I slot release: id 0 | task 204 | stop processing: n_tokens = 441, truncated = 0
|
| 104 |
+
0.26.675.625 I srv update_slots: all slots are idle
|
| 105 |
+
0.26.685.901 I srv params_from_: Chat format: peg-native
|
| 106 |
+
0.26.686.221 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.921
|
| 107 |
+
0.26.686.435 I reasoning-budget: activated, budget=2147483647 tokens
|
| 108 |
+
0.26.686.468 I slot launch_slot_: id 0 | task 223 | processing task, is_child = 0
|
| 109 |
+
0.26.686.473 W slot update_slots: id 0 | task 223 | n_past = 406, slot.prompt.tokens.size() = 441, seq_id = 0, pos_min = 440, n_swa = 0
|
| 110 |
+
0.26.686.474 I slot update_slots: id 0 | task 223 | Checking checkpoint with [420, 420] against 406...
|
| 111 |
+
0.26.686.475 W slot update_slots: id 0 | task 223 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 112 |
+
0.26.686.476 W slot update_slots: id 0 | task 223 | erased invalidated context checkpoint (pos_min = 420, pos_max = 420, n_tokens = 421, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 113 |
+
0.27.011.819 I slot create_check: id 0 | task 223 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
|
| 114 |
+
0.27.481.528 I reasoning-budget: deactivated (natural end)
|
| 115 |
+
0.28.088.823 I slot print_timing: id 0 | task 223 |
|
| 116 |
+
prompt eval time = 353.00 ms / 422 tokens ( 0.84 ms per token, 1195.46 tokens per second)
|
| 117 |
+
eval time = 1049.34 ms / 71 tokens ( 14.78 ms per token, 67.66 tokens per second)
|
| 118 |
+
total time = 1402.34 ms / 493 tokens
|
| 119 |
+
0.28.088.866 I slot release: id 0 | task 223 | stop processing: n_tokens = 492, truncated = 0
|
| 120 |
+
0.28.088.892 I srv update_slots: all slots are idle
|
| 121 |
+
0.28.099.105 I srv params_from_: Chat format: peg-native
|
| 122 |
+
0.28.099.413 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.943 (> 0.100 thold), f_keep = 1.000
|
| 123 |
+
0.28.099.624 I reasoning-budget: activated, budget=2147483647 tokens
|
| 124 |
+
0.28.099.652 I slot launch_slot_: id 0 | task 296 | processing task, is_child = 0
|
| 125 |
+
0.28.109.016 I slot create_check: id 0 | task 296 | created context checkpoint 2 of 32 (pos_min = 491, pos_max = 491, n_tokens = 492, size = 62.813 MiB)
|
| 126 |
+
0.28.638.379 I reasoning-budget: deactivated (natural end)
|
| 127 |
+
0.28.905.201 I slot print_timing: id 0 | task 296 |
|
| 128 |
+
prompt eval time = 111.40 ms / 30 tokens ( 3.71 ms per token, 269.30 tokens per second)
|
| 129 |
+
eval time = 694.13 ms / 47 tokens ( 14.77 ms per token, 67.71 tokens per second)
|
| 130 |
+
total time = 805.53 ms / 77 tokens
|
| 131 |
+
0.28.905.251 I slot release: id 0 | task 296 | stop processing: n_tokens = 568, truncated = 0
|
| 132 |
+
0.28.905.277 I srv update_slots: all slots are idle
|
| 133 |
+
0.28.915.951 I srv params_from_: Chat format: peg-native
|
| 134 |
+
0.28.916.241 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.972 (> 0.100 thold), f_keep = 0.722
|
| 135 |
+
0.28.916.410 I reasoning-budget: activated, budget=2147483647 tokens
|
| 136 |
+
0.28.916.437 I slot launch_slot_: id 0 | task 345 | processing task, is_child = 0
|
| 137 |
+
0.28.916.443 W slot update_slots: id 0 | task 345 | n_past = 410, slot.prompt.tokens.size() = 568, seq_id = 0, pos_min = 567, n_swa = 0
|
| 138 |
+
0.28.916.454 I slot update_slots: id 0 | task 345 | Checking checkpoint with [491, 491] against 410...
|
| 139 |
+
0.28.916.456 I slot update_slots: id 0 | task 345 | Checking checkpoint with [417, 417] against 410...
|
| 140 |
+
0.28.916.457 W slot update_slots: id 0 | task 345 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 141 |
+
0.28.916.459 W slot update_slots: id 0 | task 345 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 142 |
+
0.28.918.102 W slot update_slots: id 0 | task 345 | erased invalidated context checkpoint (pos_min = 491, pos_max = 491, n_tokens = 492, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 143 |
+
0.29.241.970 I slot create_check: id 0 | task 345 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
|
| 144 |
+
0.29.578.948 I reasoning-budget: deactivated (natural end)
|
| 145 |
+
0.30.187.657 I slot print_timing: id 0 | task 345 |
|
| 146 |
+
prompt eval time = 354.33 ms / 422 tokens ( 0.84 ms per token, 1190.98 tokens per second)
|
| 147 |
+
eval time = 916.87 ms / 62 tokens ( 14.79 ms per token, 67.62 tokens per second)
|
| 148 |
+
total time = 1271.20 ms / 484 tokens
|
| 149 |
+
0.30.187.698 I slot release: id 0 | task 345 | stop processing: n_tokens = 483, truncated = 0
|
| 150 |
+
0.30.187.716 I srv update_slots: all slots are idle
|
| 151 |
+
0.30.197.642 I srv params_from_: Chat format: peg-native
|
| 152 |
+
0.30.197.942 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.935 (> 0.100 thold), f_keep = 0.839
|
| 153 |
+
0.30.198.086 I reasoning-budget: activated, budget=2147483647 tokens
|
| 154 |
+
0.30.198.110 I slot launch_slot_: id 0 | task 409 | processing task, is_child = 0
|
| 155 |
+
0.30.198.115 W slot update_slots: id 0 | task 409 | n_past = 405, slot.prompt.tokens.size() = 483, seq_id = 0, pos_min = 482, n_swa = 0
|
| 156 |
+
0.30.198.116 I slot update_slots: id 0 | task 409 | Checking checkpoint with [417, 417] against 405...
|
| 157 |
+
0.30.198.117 W slot update_slots: id 0 | task 409 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 158 |
+
0.30.198.120 W slot update_slots: id 0 | task 409 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 159 |
+
0.30.540.883 I slot create_check: id 0 | task 409 | created context checkpoint 1 of 32 (pos_min = 428, pos_max = 428, n_tokens = 429, size = 62.813 MiB)
|
| 160 |
+
0.32.074.181 I reasoning-budget: deactivated (natural end)
|
| 161 |
+
0.32.119.135 I slot print_timing: id 0 | task 409 | n_decoded = 100, tg = 64.60 t/s
|
| 162 |
+
0.33.280.892 I slot print_timing: id 0 | task 409 |
|
| 163 |
+
prompt eval time = 373.02 ms / 433 tokens ( 0.86 ms per token, 1160.78 tokens per second)
|
| 164 |
+
eval time = 2709.74 ms / 178 tokens ( 15.22 ms per token, 65.69 tokens per second)
|
| 165 |
+
total time = 3082.77 ms / 611 tokens
|
| 166 |
+
0.33.280.942 I slot release: id 0 | task 409 | stop processing: n_tokens = 610, truncated = 0
|
| 167 |
+
0.33.280.983 I srv update_slots: all slots are idle
|
| 168 |
+
0.33.291.660 I srv params_from_: Chat format: peg-native
|
| 169 |
+
0.33.291.953 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.664
|
| 170 |
+
0.33.292.103 I reasoning-budget: activated, budget=2147483647 tokens
|
| 171 |
+
0.33.292.127 I slot launch_slot_: id 0 | task 589 | processing task, is_child = 0
|
| 172 |
+
0.33.292.133 W slot update_slots: id 0 | task 589 | n_past = 405, slot.prompt.tokens.size() = 610, seq_id = 0, pos_min = 609, n_swa = 0
|
| 173 |
+
0.33.292.133 I slot update_slots: id 0 | task 589 | Checking checkpoint with [428, 428] against 405...
|
| 174 |
+
0.33.292.134 W slot update_slots: id 0 | task 589 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 175 |
+
0.33.292.136 W slot update_slots: id 0 | task 589 | erased invalidated context checkpoint (pos_min = 428, pos_max = 428, n_tokens = 429, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 176 |
+
0.33.632.176 I slot create_check: id 0 | task 589 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
|
| 177 |
+
0.34.264.101 I slot print_timing: id 0 | task 589 |
|
| 178 |
+
prompt eval time = 370.43 ms / 424 tokens ( 0.87 ms per token, 1144.62 tokens per second)
|
| 179 |
+
eval time = 601.52 ms / 39 tokens ( 15.42 ms per token, 64.84 tokens per second)
|
| 180 |
+
total time = 971.94 ms / 463 tokens
|
| 181 |
+
0.34.264.186 I slot release: id 0 | task 589 | stop processing: n_tokens = 462, truncated = 0
|
| 182 |
+
0.34.264.222 I srv update_slots: all slots are idle
|
| 183 |
+
0.34.291.838 I srv params_from_: Chat format: peg-native
|
| 184 |
+
0.34.292.413 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.902 (> 0.100 thold), f_keep = 0.877
|
| 185 |
+
0.34.292.609 I reasoning-budget: activated, budget=2147483647 tokens
|
| 186 |
+
0.34.292.640 I slot launch_slot_: id 0 | task 630 | processing task, is_child = 0
|
| 187 |
+
0.34.292.647 W slot update_slots: id 0 | task 630 | n_past = 405, slot.prompt.tokens.size() = 462, seq_id = 0, pos_min = 461, n_swa = 0
|
| 188 |
+
0.34.292.648 I slot update_slots: id 0 | task 630 | Checking checkpoint with [419, 419] against 405...
|
| 189 |
+
0.34.292.649 W slot update_slots: id 0 | task 630 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 190 |
+
0.34.292.653 W slot update_slots: id 0 | task 630 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 191 |
+
0.34.652.933 I slot create_check: id 0 | task 630 | created context checkpoint 1 of 32 (pos_min = 444, pos_max = 444, n_tokens = 445, size = 62.813 MiB)
|
| 192 |
+
0.36.023.509 I slot print_timing: id 0 | task 630 |
|
| 193 |
+
prompt eval time = 390.02 ms / 449 tokens ( 0.87 ms per token, 1151.23 tokens per second)
|
| 194 |
+
eval time = 1340.82 ms / 86 tokens ( 15.59 ms per token, 64.14 tokens per second)
|
| 195 |
+
total time = 1730.84 ms / 535 tokens
|
| 196 |
+
0.36.023.602 I slot release: id 0 | task 630 | stop processing: n_tokens = 534, truncated = 0
|
| 197 |
+
0.36.023.637 I srv update_slots: all slots are idle
|
| 198 |
+
0.36.023.710 W common_chat_peg_parse: unparsed peg-native output: <tool_call>
|
| 199 |
+
<function=create_event>
|
| 200 |
+
<parameter=attendees>
|
| 201 |
+
["ana@x.io", "bo@x.io"]
|
| 202 |
+
</parameter>
|
| 203 |
+
<parameter=when>
|
| 204 |
+
{"date": "2026-10-02", "time": "14:00"}
|
| 205 |
+
</parameter>
|
| 206 |
+
<parameter=title>
|
| 207 |
+
Design review
|
| 208 |
+
</parameter>
|
| 209 |
+
</function>
|
| 210 |
+
</tool_call>
|
| 211 |
+
0.36.025.846 W srv stop: cancel task, id_task = 630
|
| 212 |
+
0.36.025.894 I srv update_slots: all slots are idle
|
| 213 |
+
0.36.025.957 W srv operator(): got exception: {"error":{"code":500,"message":"The model produced output that does not match the expected peg-native format","type":"server_error"}}
|
| 214 |
+
0.36.045.619 I srv params_from_: Chat format: peg-native
|
| 215 |
+
0.36.045.957 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.948 (> 0.100 thold), f_keep = 0.758
|
| 216 |
+
0.36.046.107 I reasoning-budget: activated, budget=2147483647 tokens
|
| 217 |
+
0.36.046.133 I slot launch_slot_: id 0 | task 719 | processing task, is_child = 0
|
| 218 |
+
0.36.046.138 W slot update_slots: id 0 | task 719 | n_past = 405, slot.prompt.tokens.size() = 534, seq_id = 0, pos_min = 533, n_swa = 0
|
| 219 |
+
0.36.046.139 I slot update_slots: id 0 | task 719 | Checking checkpoint with [444, 444] against 405...
|
| 220 |
+
0.36.046.142 W slot update_slots: id 0 | task 719 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 221 |
+
0.36.046.144 W slot update_slots: id 0 | task 719 | erased invalidated context checkpoint (pos_min = 444, pos_max = 444, n_tokens = 445, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 222 |
+
0.36.387.995 I slot create_check: id 0 | task 719 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
|
| 223 |
+
0.37.019.525 I slot print_timing: id 0 | task 719 |
|
| 224 |
+
prompt eval time = 371.74 ms / 427 tokens ( 0.87 ms per token, 1148.64 tokens per second)
|
| 225 |
+
eval time = 601.62 ms / 39 tokens ( 15.43 ms per token, 64.82 tokens per second)
|
| 226 |
+
total time = 973.36 ms / 466 tokens
|
| 227 |
+
0.37.019.612 I slot release: id 0 | task 719 | stop processing: n_tokens = 465, truncated = 0
|
| 228 |
+
0.37.019.648 I srv update_slots: all slots are idle
|
| 229 |
+
0.37.047.661 I srv params_from_: Chat format: peg-native
|
| 230 |
+
0.37.048.192 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.948 (> 0.100 thold), f_keep = 0.871
|
| 231 |
+
0.37.048.354 I reasoning-budget: activated, budget=2147483647 tokens
|
| 232 |
+
0.37.048.382 I slot launch_slot_: id 0 | task 760 | processing task, is_child = 0
|
| 233 |
+
0.37.048.388 W slot update_slots: id 0 | task 760 | n_past = 405, slot.prompt.tokens.size() = 465, seq_id = 0, pos_min = 464, n_swa = 0
|
| 234 |
+
0.37.048.390 I slot update_slots: id 0 | task 760 | Checking checkpoint with [422, 422] against 405...
|
| 235 |
+
0.37.048.391 W slot update_slots: id 0 | task 760 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 236 |
+
0.37.048.393 W slot update_slots: id 0 | task 760 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 237 |
+
0.37.396.973 I slot create_check: id 0 | task 760 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
|
| 238 |
+
0.37.495.139 I slot print_timing: id 0 | task 760 |
|
| 239 |
+
prompt eval time = 378.68 ms / 427 tokens ( 0.89 ms per token, 1127.60 tokens per second)
|
| 240 |
+
eval time = 68.05 ms / 4 tokens ( 17.01 ms per token, 58.78 tokens per second)
|
| 241 |
+
total time = 446.73 ms / 431 tokens
|
| 242 |
+
0.37.495.225 I slot release: id 0 | task 760 | stop processing: n_tokens = 430, truncated = 0
|
| 243 |
+
0.37.495.261 I srv update_slots: all slots are idle
|
| 244 |
+
0.37.511.484 I srv params_from_: Chat format: peg-native
|
| 245 |
+
0.37.511.815 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.958 (> 0.100 thold), f_keep = 0.944
|
| 246 |
+
0.37.511.971 I reasoning-budget: activated, budget=2147483647 tokens
|
| 247 |
+
0.37.512.000 I slot launch_slot_: id 0 | task 766 | processing task, is_child = 0
|
| 248 |
+
0.37.512.007 W slot update_slots: id 0 | task 766 | n_past = 406, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
|
| 249 |
+
0.37.512.008 I slot update_slots: id 0 | task 766 | Checking checkpoint with [422, 422] against 406...
|
| 250 |
+
0.37.512.009 W slot update_slots: id 0 | task 766 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 251 |
+
0.37.512.012 W slot update_slots: id 0 | task 766 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 252 |
+
0.37.860.025 I slot create_check: id 0 | task 766 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
|
| 253 |
+
0.38.509.296 I slot print_timing: id 0 | task 766 |
|
| 254 |
+
prompt eval time = 377.88 ms / 424 tokens ( 0.89 ms per token, 1122.06 tokens per second)
|
| 255 |
+
eval time = 619.39 ms / 40 tokens ( 15.48 ms per token, 64.58 tokens per second)
|
| 256 |
+
total time = 997.27 ms / 464 tokens
|
| 257 |
+
0.38.509.383 I slot release: id 0 | task 766 | stop processing: n_tokens = 463, truncated = 0
|
| 258 |
+
0.38.509.426 I srv update_slots: all slots are idle
|
| 259 |
+
0.38.537.775 I srv params_from_: Chat format: peg-native
|
| 260 |
+
0.38.538.231 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.935 (> 0.100 thold), f_keep = 1.000
|
| 261 |
+
0.38.538.382 I reasoning-budget: activated, budget=2147483647 tokens
|
| 262 |
+
0.38.538.411 I slot launch_slot_: id 0 | task 808 | processing task, is_child = 0
|
| 263 |
+
0.38.628.133 I slot create_check: id 0 | task 808 | created context checkpoint 2 of 32 (pos_min = 490, pos_max = 490, n_tokens = 491, size = 62.813 MiB)
|
| 264 |
+
0.38.869.113 I slot print_timing: id 0 | task 808 |
|
| 265 |
+
prompt eval time = 118.73 ms / 32 tokens ( 3.71 ms per token, 269.51 tokens per second)
|
| 266 |
+
eval time = 211.94 ms / 14 tokens ( 15.14 ms per token, 66.06 tokens per second)
|
| 267 |
+
total time = 330.67 ms / 46 tokens
|
| 268 |
+
0.38.869.220 I slot release: id 0 | task 808 | stop processing: n_tokens = 508, truncated = 0
|
| 269 |
+
0.38.869.257 I srv update_slots: all slots are idle
|
| 270 |
+
0.38.890.188 I srv params_from_: Chat format: peg-native
|
| 271 |
+
0.38.890.530 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.967 (> 0.100 thold), f_keep = 0.807
|
| 272 |
+
0.38.890.683 I reasoning-budget: activated, budget=2147483647 tokens
|
| 273 |
+
0.38.890.717 I slot launch_slot_: id 0 | task 824 | processing task, is_child = 0
|
| 274 |
+
0.38.890.724 W slot update_slots: id 0 | task 824 | n_past = 410, slot.prompt.tokens.size() = 508, seq_id = 0, pos_min = 507, n_swa = 0
|
| 275 |
+
0.38.890.724 I slot update_slots: id 0 | task 824 | Checking checkpoint with [490, 490] against 410...
|
| 276 |
+
0.38.890.725 I slot update_slots: id 0 | task 824 | Checking checkpoint with [419, 419] against 410...
|
| 277 |
+
0.38.890.726 W slot update_slots: id 0 | task 824 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 278 |
+
0.38.890.729 W slot update_slots: id 0 | task 824 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 279 |
+
0.38.892.293 W slot update_slots: id 0 | task 824 | erased invalidated context checkpoint (pos_min = 490, pos_max = 490, n_tokens = 491, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 280 |
+
0.39.240.897 I slot create_check: id 0 | task 824 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
|
| 281 |
+
0.39.889.179 I slot print_timing: id 0 | task 824 |
|
| 282 |
+
prompt eval time = 380.09 ms / 424 tokens ( 0.90 ms per token, 1115.53 tokens per second)
|
| 283 |
+
eval time = 618.35 ms / 40 tokens ( 15.46 ms per token, 64.69 tokens per second)
|
| 284 |
+
total time = 998.43 ms / 464 tokens
|
| 285 |
+
0.39.889.258 I slot release: id 0 | task 824 | stop processing: n_tokens = 463, truncated = 0
|
| 286 |
+
0.39.889.291 I srv update_slots: all slots are idle
|
| 287 |
+
0.39.916.715 I srv params_from_: Chat format: peg-native
|
| 288 |
+
0.39.917.343 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.931 (> 0.100 thold), f_keep = 0.875
|
| 289 |
+
0.39.917.545 I reasoning-budget: activated, budget=2147483647 tokens
|
| 290 |
+
0.39.917.577 I slot launch_slot_: id 0 | task 866 | processing task, is_child = 0
|
| 291 |
+
0.39.917.585 W slot update_slots: id 0 | task 866 | n_past = 405, slot.prompt.tokens.size() = 463, seq_id = 0, pos_min = 462, n_swa = 0
|
| 292 |
+
0.39.917.586 I slot update_slots: id 0 | task 866 | Checking checkpoint with [419, 419] against 405...
|
| 293 |
+
0.39.917.587 W slot update_slots: id 0 | task 866 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
|
| 294 |
+
0.39.917.590 W slot update_slots: id 0 | task 866 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
|
| 295 |
+
0.40.271.964 I slot create_check: id 0 | task 866 | created context checkpoint 1 of 32 (pos_min = 430, pos_max = 430, n_tokens = 431, size = 62.813 MiB)
|
| 296 |
+
0.41.551.306 I slot print_timing: id 0 | task 866 |
|
| 297 |
+
prompt eval time = 384.12 ms / 435 tokens ( 0.88 ms per token, 1132.46 tokens per second)
|
| 298 |
+
eval time = 1249.57 ms / 80 tokens ( 15.62 ms per token, 64.02 tokens per second)
|
| 299 |
+
total time = 1633.69 ms / 515 tokens
|
| 300 |
+
0.41.551.447 I slot release: id 0 | task 866 | stop processing: n_tokens = 514, truncated = 0
|
| 301 |
+
0.41.551.496 I srv update_slots: all slots are idle
|
| 302 |
+
0.41.552.857 I srv operator(): operator(): cleaning up before exit...
|
recipe/logs/b_n-vision-q106-faoff.log
ADDED
|
@@ -0,0 +1,66 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
0.00.044.023 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
|
| 2 |
+
0.00.044.026 I device_info:
|
| 3 |
+
0.00.044.069 I - ROCm0 : AMD Radeon Graphics (131072 MiB, 125177 MiB free)
|
| 4 |
+
0.00.044.147 I - Vulkan0 : AMD Radeon Graphics (RADV GFX1151) (132096 MiB, 131926 MiB free)
|
| 5 |
+
0.00.044.151 I - CPU : AMD RYZEN AI MAX+ 395 w/ Radeon 8060S (127438 MiB, 127438 MiB free)
|
| 6 |
+
0.00.044.190 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
|
| 7 |
+
0.00.044.225 I srv init: running without SSL
|
| 8 |
+
0.00.044.247 I srv init: using 31 threads for HTTP server
|
| 9 |
+
0.00.044.248 I srv init: the WebUI is disabled
|
| 10 |
+
0.00.044.304 I srv start: binding port with default address family
|
| 11 |
+
0.00.045.454 I srv main: loading model
|
| 12 |
+
0.00.045.457 I srv load_model: loading model '/mnt/models/nex-n2.5-mini/out/Nex-N2.5-mini-Q4_0_ROCMFP4_STRIX_LEAN.gguf'
|
| 13 |
+
0.00.079.109 W llama_model_loader: direct I/O is enabled, disabling mmap
|
| 14 |
+
0.20.570.878 W llama_context: n_ctx_seq (65536) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
|
| 15 |
+
0.21.071.707 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|
| 16 |
+
0.21.277.411 W load_hparams: Qwen-VL models require at minimum 1024 image tokens to function correctly on grounding tasks
|
| 17 |
+
0.21.277.414 W load_hparams: if you encounter problems with accuracy, try adding --image-min-tokens 1024
|
| 18 |
+
0.21.277.414 W load_hparams: more info: https://github.com/ggml-org/llama.cpp/issues/16842
|
| 19 |
+
|
| 20 |
+
0.22.242.916 I srv load_model: loaded multimodal model, '/mnt/models/nex-n2.5-mini/out/mmproj-Nex-N2.5-mini-BF16.gguf'
|
| 21 |
+
0.22.242.925 I srv load_model: initializing slots, n_slots = 1
|
| 22 |
+
0.22.421.079 W srv load_model: speculative decoding will use checkpoints
|
| 23 |
+
0.22.421.086 W common_speculative_init: no implementations specified for speculative decoding
|
| 24 |
+
0.22.421.088 I slot load_model: id 0 | task -1 | new slot, n_ctx = 65536
|
| 25 |
+
0.22.421.122 I srv load_model: prompt cache RAM enabled: limit_mib=8192
|
| 26 |
+
0.22.421.129 I srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
|
| 27 |
+
0.22.421.144 I srv init: idle slots will be saved to prompt cache upon starting a new task
|
| 28 |
+
0.22.429.869 I init: chat template, example_format: '<|im_start|>system
|
| 29 |
+
You are a helpful assistant<|im_end|>
|
| 30 |
+
<|im_start|>user
|
| 31 |
+
Hello<|im_end|>
|
| 32 |
+
<|im_start|>assistant
|
| 33 |
+
<think>
|
| 34 |
+
|
| 35 |
+
</think>
|
| 36 |
+
|
| 37 |
+
Hi there<|im_end|>
|
| 38 |
+
<|im_start|>user
|
| 39 |
+
How are you?<|im_end|>
|
| 40 |
+
<|im_start|>assistant
|
| 41 |
+
<think>'
|
| 42 |
+
0.22.436.195 I srv init: init: chat template, thinking = 1
|
| 43 |
+
0.22.436.220 I srv main: model loaded
|
| 44 |
+
0.22.436.223 I srv main: server is listening on http://127.0.0.1:18600
|
| 45 |
+
0.22.436.225 I srv update_slots: all slots are idle
|
| 46 |
+
0.24.206.034 I srv params_from_: Chat format: peg-native
|
| 47 |
+
0.24.206.169 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1
|
| 48 |
+
0.24.206.172 I srv get_availabl: updating prompt cache
|
| 49 |
+
0.24.206.177 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000
|
| 50 |
+
0.24.206.182 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 65536 tokens, 8589934592 est)
|
| 51 |
+
0.24.206.183 I srv get_availabl: prompt cache update took 0.01 ms
|
| 52 |
+
0.24.206.233 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0
|
| 53 |
+
0.24.246.538 I srv process_chun: processing image...
|
| 54 |
+
0.24.375.032 W find_slot: non-consecutive token position 4 after 3 for sequence 0 with 196 new tokens
|
| 55 |
+
0.24.375.099 W find_slot: non-consecutive token position 4 after 3 for sequence 0 with 196 new tokens
|
| 56 |
+
0.24.571.922 I srv process_chun: image processed in 326 ms
|
| 57 |
+
0.24.572.141 W find_slot: non-consecutive token position 34 after 4 for sequence 0 with 17 new tokens
|
| 58 |
+
0.24.572.170 W find_slot: non-consecutive token position 34 after 4 for sequence 0 with 17 new tokens
|
| 59 |
+
0.24.650.730 I slot create_check: id 0 | task 0 | created context checkpoint 1 of 32 (pos_min = 34, pos_max = 34, n_tokens = 217, size = 62.813 MiB)
|
| 60 |
+
0.25.135.589 I slot print_timing: id 0 | task 0 |
|
| 61 |
+
prompt eval time = 471.32 ms / 221 tokens ( 2.13 ms per token, 468.89 tokens per second)
|
| 62 |
+
eval time = 458.01 ms / 30 tokens ( 15.27 ms per token, 65.50 tokens per second)
|
| 63 |
+
total time = 929.34 ms / 251 tokens
|
| 64 |
+
0.25.135.616 I slot release: id 0 | task 0 | stop processing: n_tokens = 250, truncated = 0
|
| 65 |
+
0.25.135.619 I srv update_slots: all slots are idle
|
| 66 |
+
0.26.136.415 I srv operator(): operator(): cleaning up before exit...
|
recipe/logs/b_n-vision-q106-faon.log
ADDED
|
@@ -0,0 +1,66 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
0.00.039.566 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
|
| 2 |
+
0.00.039.569 I device_info:
|
| 3 |
+
0.00.039.614 I - ROCm0 : AMD Radeon Graphics (131072 MiB, 125113 MiB free)
|
| 4 |
+
0.00.039.689 I - Vulkan0 : AMD Radeon Graphics (RADV GFX1151) (132096 MiB, 131926 MiB free)
|
| 5 |
+
0.00.039.693 I - CPU : AMD RYZEN AI MAX+ 395 w/ Radeon 8060S (127438 MiB, 127438 MiB free)
|
| 6 |
+
0.00.039.733 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
|
| 7 |
+
0.00.039.777 I srv init: running without SSL
|
| 8 |
+
0.00.039.798 I srv init: using 31 threads for HTTP server
|
| 9 |
+
0.00.039.799 I srv init: the WebUI is disabled
|
| 10 |
+
0.00.039.854 I srv start: binding port with default address family
|
| 11 |
+
0.00.041.005 I srv main: loading model
|
| 12 |
+
0.00.041.007 I srv load_model: loading model '/mnt/models/nex-n2.5-mini/out/Nex-N2.5-mini-Q4_0_ROCMFP4_STRIX_LEAN.gguf'
|
| 13 |
+
0.00.072.658 W llama_model_loader: direct I/O is enabled, disabling mmap
|
| 14 |
+
0.20.385.335 W llama_context: n_ctx_seq (65536) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
|
| 15 |
+
0.20.547.529 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|
| 16 |
+
0.20.746.298 W load_hparams: Qwen-VL models require at minimum 1024 image tokens to function correctly on grounding tasks
|
| 17 |
+
0.20.746.300 W load_hparams: if you encounter problems with accuracy, try adding --image-min-tokens 1024
|
| 18 |
+
0.20.746.300 W load_hparams: more info: https://github.com/ggml-org/llama.cpp/issues/16842
|
| 19 |
+
|
| 20 |
+
0.21.800.919 I srv load_model: loaded multimodal model, '/mnt/models/nex-n2.5-mini/out/mmproj-Nex-N2.5-mini-BF16.gguf'
|
| 21 |
+
0.21.800.929 I srv load_model: initializing slots, n_slots = 1
|
| 22 |
+
0.21.998.188 W srv load_model: speculative decoding will use checkpoints
|
| 23 |
+
0.21.998.194 W common_speculative_init: no implementations specified for speculative decoding
|
| 24 |
+
0.21.998.195 I slot load_model: id 0 | task -1 | new slot, n_ctx = 65536
|
| 25 |
+
0.21.998.237 I srv load_model: prompt cache RAM enabled: limit_mib=8192
|
| 26 |
+
0.21.998.246 I srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
|
| 27 |
+
0.21.998.259 I srv init: idle slots will be saved to prompt cache upon starting a new task
|
| 28 |
+
0.22.007.086 I init: chat template, example_format: '<|im_start|>system
|
| 29 |
+
You are a helpful assistant<|im_end|>
|
| 30 |
+
<|im_start|>user
|
| 31 |
+
Hello<|im_end|>
|
| 32 |
+
<|im_start|>assistant
|
| 33 |
+
<think>
|
| 34 |
+
|
| 35 |
+
</think>
|
| 36 |
+
|
| 37 |
+
Hi there<|im_end|>
|
| 38 |
+
<|im_start|>user
|
| 39 |
+
How are you?<|im_end|>
|
| 40 |
+
<|im_start|>assistant
|
| 41 |
+
<think>'
|
| 42 |
+
0.22.013.608 I srv init: init: chat template, thinking = 1
|
| 43 |
+
0.22.013.635 I srv main: model loaded
|
| 44 |
+
0.22.013.638 I srv main: server is listening on http://127.0.0.1:18600
|
| 45 |
+
0.22.013.640 I srv update_slots: all slots are idle
|
| 46 |
+
0.24.005.244 I srv params_from_: Chat format: peg-native
|
| 47 |
+
0.24.005.366 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1
|
| 48 |
+
0.24.005.368 I srv get_availabl: updating prompt cache
|
| 49 |
+
0.24.005.373 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000
|
| 50 |
+
0.24.005.376 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 65536 tokens, 8589934592 est)
|
| 51 |
+
0.24.005.377 I srv get_availabl: prompt cache update took 0.01 ms
|
| 52 |
+
0.24.005.401 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0
|
| 53 |
+
0.24.032.329 I srv process_chun: processing image...
|
| 54 |
+
0.24.141.298 W find_slot: non-consecutive token position 4 after 3 for sequence 0 with 196 new tokens
|
| 55 |
+
0.24.141.364 W find_slot: non-consecutive token position 4 after 3 for sequence 0 with 196 new tokens
|
| 56 |
+
0.24.326.150 I srv process_chun: image processed in 294 ms
|
| 57 |
+
0.24.326.369 W find_slot: non-consecutive token position 34 after 4 for sequence 0 with 17 new tokens
|
| 58 |
+
0.24.326.399 W find_slot: non-consecutive token position 34 after 4 for sequence 0 with 17 new tokens
|
| 59 |
+
0.24.404.937 I slot create_check: id 0 | task 0 | created context checkpoint 1 of 32 (pos_min = 34, pos_max = 34, n_tokens = 217, size = 62.813 MiB)
|
| 60 |
+
0.24.744.679 I slot print_timing: id 0 | task 0 |
|
| 61 |
+
prompt eval time = 426.45 ms / 221 tokens ( 1.93 ms per token, 518.23 tokens per second)
|
| 62 |
+
eval time = 312.81 ms / 21 tokens ( 14.90 ms per token, 67.13 tokens per second)
|
| 63 |
+
total time = 739.26 ms / 242 tokens
|
| 64 |
+
0.24.744.708 I slot release: id 0 | task 0 | stop processing: n_tokens = 241, truncated = 0
|
| 65 |
+
0.24.744.712 I srv update_slots: all slots are idle
|
| 66 |
+
0.25.745.484 I srv operator(): operator(): cleaning up before exit...
|
recipe/logs/b_n-vision-q106-roff-faon.log
ADDED
|
@@ -0,0 +1,70 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
0.00.208.975 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
|
| 2 |
+
0.00.208.993 I device_info:
|
| 3 |
+
0.00.209.271 I - ROCm0 : AMD Radeon Graphics (131072 MiB, 123872 MiB free)
|
| 4 |
+
0.00.209.724 I - Vulkan0 : AMD Radeon Graphics (RADV GFX1151) (132096 MiB, 131922 MiB free)
|
| 5 |
+
0.00.209.744 I - CPU : AMD RYZEN AI MAX+ 395 w/ Radeon 8060S (127438 MiB, 127438 MiB free)
|
| 6 |
+
0.00.209.942 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
|
| 7 |
+
0.00.210.045 I srv init: running without SSL
|
| 8 |
+
0.00.210.124 I srv init: using 31 threads for HTTP server
|
| 9 |
+
0.00.210.128 I srv init: the WebUI is disabled
|
| 10 |
+
0.00.210.347 I srv start: binding port with default address family
|
| 11 |
+
0.00.211.826 I srv main: loading model
|
| 12 |
+
0.00.211.839 I srv load_model: loading model '/mnt/models/nex-n2.5-mini/out/Nex-N2.5-mini-Q4_0_ROCMFP4_STRIX_LEAN.gguf'
|
| 13 |
+
0.00.341.319 W llama_model_loader: direct I/O is enabled, disabling mmap
|
| 14 |
+
0.21.841.225 W llama_context: n_ctx_seq (65536) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
|
| 15 |
+
0.22.147.423 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|
| 16 |
+
0.22.493.441 W load_hparams: Qwen-VL models require at minimum 1024 image tokens to function correctly on grounding tasks
|
| 17 |
+
0.22.493.444 W load_hparams: if you encounter problems with accuracy, try adding --image-min-tokens 1024
|
| 18 |
+
0.22.493.445 W load_hparams: more info: https://github.com/ggml-org/llama.cpp/issues/16842
|
| 19 |
+
|
| 20 |
+
0.22.758.348 I srv load_model: loaded multimodal model, '/mnt/models/nex-n2.5-mini/out/mmproj-Nex-N2.5-mini-BF16.gguf'
|
| 21 |
+
0.22.758.359 I srv load_model: initializing slots, n_slots = 1
|
| 22 |
+
0.23.027.404 W srv load_model: speculative decoding will use checkpoints
|
| 23 |
+
0.23.027.422 W common_speculative_init: no implementations specified for speculative decoding
|
| 24 |
+
0.23.027.426 I slot load_model: id 0 | task -1 | new slot, n_ctx = 65536
|
| 25 |
+
0.23.027.506 I srv load_model: prompt cache RAM enabled: limit_mib=8192
|
| 26 |
+
0.23.027.508 I srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
|
| 27 |
+
0.23.027.604 I srv init: idle slots will be saved to prompt cache upon starting a new task
|
| 28 |
+
0.23.040.255 I init: chat template, example_format: '<|im_start|>system
|
| 29 |
+
You are a helpful assistant<|im_end|>
|
| 30 |
+
<|im_start|>user
|
| 31 |
+
Hello<|im_end|>
|
| 32 |
+
<|im_start|>assistant
|
| 33 |
+
<think>
|
| 34 |
+
|
| 35 |
+
</think>
|
| 36 |
+
|
| 37 |
+
Hi there<|im_end|>
|
| 38 |
+
<|im_start|>user
|
| 39 |
+
How are you?<|im_end|>
|
| 40 |
+
<|im_start|>assistant
|
| 41 |
+
<think>
|
| 42 |
+
|
| 43 |
+
</think>
|
| 44 |
+
|
| 45 |
+
'
|
| 46 |
+
0.23.049.208 I srv init: init: chat template, thinking = 0
|
| 47 |
+
0.23.049.258 I srv main: model loaded
|
| 48 |
+
0.23.049.261 I srv main: server is listening on http://127.0.0.1:18600
|
| 49 |
+
0.23.049.265 I srv update_slots: all slots are idle
|
| 50 |
+
0.24.037.221 I srv params_from_: Chat format: peg-native
|
| 51 |
+
0.24.037.451 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1
|
| 52 |
+
0.24.037.454 I srv get_availabl: updating prompt cache
|
| 53 |
+
0.24.037.461 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000
|
| 54 |
+
0.24.037.467 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 65536 tokens, 8589934592 est)
|
| 55 |
+
0.24.037.470 I srv get_availabl: prompt cache update took 0.01 ms
|
| 56 |
+
0.24.037.553 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0
|
| 57 |
+
0.24.096.831 I srv process_chun: processing image...
|
| 58 |
+
0.24.348.055 W find_slot: non-consecutive token position 4 after 3 for sequence 0 with 196 new tokens
|
| 59 |
+
0.24.348.208 W find_slot: non-consecutive token position 4 after 3 for sequence 0 with 196 new tokens
|
| 60 |
+
0.24.670.293 I srv process_chun: image processed in 573 ms
|
| 61 |
+
0.24.670.967 W find_slot: non-consecutive token position 34 after 4 for sequence 0 with 17 new tokens
|
| 62 |
+
0.24.671.060 W find_slot: non-consecutive token position 34 after 4 for sequence 0 with 17 new tokens
|
| 63 |
+
0.24.787.337 I slot create_check: id 0 | task 0 | created context checkpoint 1 of 32 (pos_min = 34, pos_max = 34, n_tokens = 217, size = 62.813 MiB)
|
| 64 |
+
0.25.422.382 I slot print_timing: id 0 | task 0 |
|
| 65 |
+
prompt eval time = 788.44 ms / 221 tokens ( 3.57 ms per token, 280.30 tokens per second)
|
| 66 |
+
eval time = 596.35 ms / 21 tokens ( 28.40 ms per token, 35.21 tokens per second)
|
| 67 |
+
total time = 1384.78 ms / 242 tokens
|
| 68 |
+
0.25.422.466 I slot release: id 0 | task 0 | stop processing: n_tokens = 241, truncated = 0
|
| 69 |
+
0.25.422.481 I srv update_slots: all slots are idle
|
| 70 |
+
0.26.423.773 I srv operator(): operator(): cleaning up before exit...
|
recipe/logs/diag_bf16.log
ADDED
|
@@ -0,0 +1,5 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
[2026-09-16T22:27:17Z] bf16_rocm_faoff rc=0 [1]136.6879 [2]166.5229 [3]174.2693 [4]169.2609
|
| 2 |
+
[2026-09-16T22:29:15Z] bf16_vk_faon rc=0 [1]5.6953 [2]6.6594 [3]7.0323 [4]7.2857
|
| 3 |
+
[2026-09-16T22:29:46Z] q106_vk_faon rc=0 [1]6.1880 [2]7.0357 [3]7.4624 [4]7.7345
|
| 4 |
+
[2026-09-16T22:30:26Z] bf16_cpu rc=0 [1]141.8336 [2]168.6793
|
| 5 |
+
[2026-09-16T22:30:26Z] DIAG_BF16_DONE
|
recipe/logs/diag_bf16_cpu.log
ADDED
|
@@ -0,0 +1,13 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
0.00.032.646 I common_init_result: fitting params to device memory ...
|
| 2 |
+
0.00.032.649 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
|
| 3 |
+
0.08.593.776 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
|
| 4 |
+
0.08.699.313 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|
| 5 |
+
0.09.377.813 I
|
| 6 |
+
0.09.377.939 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
|
| 7 |
+
0.09.378.018 I perplexity: tokenizing the input ..
|
| 8 |
+
0.09.680.733 I perplexity: tokenization took 302.776 ms
|
| 9 |
+
0.09.680.810 I perplexity: calculating perplexity over 2 chunks, n_ctx=2048, batch_size=2048, n_seq=1
|
| 10 |
+
0.24.718.172 I perplexity: 15.04 seconds per pass - ETA 0.50 minutes
|
| 11 |
+
[1]141.8336,[2]168.6793,
|
| 12 |
+
0.39.305.358 I Final estimate: PPL = 168.6793 +/- 12.64504
|
| 13 |
+
|
recipe/logs/diag_bf16_purecpu_c1.log
ADDED
|
@@ -0,0 +1,14 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
0.00.039.976 I common_init_result: fitting params to device memory ...
|
| 2 |
+
0.00.039.980 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
|
| 3 |
+
0.00.307.048 I common_params_fit_impl: projected to use 66756 MiB of host memory vs. 127438 MiB of total host memory
|
| 4 |
+
0.00.633.915 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
|
| 5 |
+
0.00.653.775 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|
| 6 |
+
0.01.312.044 I
|
| 7 |
+
0.01.312.177 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
|
| 8 |
+
0.01.312.255 I perplexity: tokenizing the input ..
|
| 9 |
+
0.01.611.737 I perplexity: tokenization took 299.542 ms
|
| 10 |
+
0.01.611.820 I perplexity: calculating perplexity over 1 chunks, n_ctx=2048, batch_size=2048, n_seq=1
|
| 11 |
+
0.11.553.629 I perplexity: 9.94 seconds per pass - ETA 0.15 minutes
|
| 12 |
+
[1]5.6964,
|
| 13 |
+
0.11.592.689 I Final estimate: PPL = 5.6964 +/- 0.43869
|
| 14 |
+
|
recipe/logs/diag_bf16_rocm_faoff.log
ADDED
|
@@ -0,0 +1,13 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
0.00.041.213 I common_init_result: fitting params to device memory ...
|
| 2 |
+
0.00.041.216 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
|
| 3 |
+
2.25.152.472 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
|
| 4 |
+
2.25.329.776 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|
| 5 |
+
2.26.102.739 I
|
| 6 |
+
2.26.102.951 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
|
| 7 |
+
2.26.102.958 I perplexity: tokenizing the input ..
|
| 8 |
+
2.26.847.963 I perplexity: tokenization took 744.994 ms
|
| 9 |
+
2.26.848.073 I perplexity: calculating perplexity over 4 chunks, n_ctx=2048, batch_size=2048, n_seq=1
|
| 10 |
+
2.32.003.043 I perplexity: 5.15 seconds per pass - ETA 0.33 minutes
|
| 11 |
+
[1]136.6879,[2]166.5229,[3]174.2693,[4]169.2609,
|
| 12 |
+
2.45.605.000 I Final estimate: PPL = 169.2609 +/- 8.97119
|
| 13 |
+
|
recipe/logs/diag_bf16_vk_faon.log
ADDED
|
@@ -0,0 +1,13 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
0.00.081.666 I common_init_result: fitting params to device memory ...
|
| 2 |
+
0.00.081.669 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
|
| 3 |
+
1.04.951.000 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
|
| 4 |
+
1.05.009.711 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|
| 5 |
+
1.06.643.947 I
|
| 6 |
+
1.06.646.719 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
|
| 7 |
+
1.06.646.727 I perplexity: tokenizing the input ..
|
| 8 |
+
1.07.413.989 I perplexity: tokenization took 767.253 ms
|
| 9 |
+
1.07.414.080 I perplexity: calculating perplexity over 4 chunks, n_ctx=2048, batch_size=2048, n_seq=1
|
| 10 |
+
1.19.697.964 I perplexity: 12.28 seconds per pass - ETA 0.82 minutes
|
| 11 |
+
[1]5.6953,[2]6.6594,[3]7.0323,[4]7.2857,
|
| 12 |
+
1.57.415.169 I Final estimate: PPL = 7.2857 +/- 0.29353
|
| 13 |
+
|
recipe/logs/diag_ppl_q106_rocm_c4.log
ADDED
|
@@ -0,0 +1,13 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
0.00.044.069 I common_init_result: fitting params to device memory ...
|
| 2 |
+
0.00.044.071 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
|
| 3 |
+
0.03.928.272 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
|
| 4 |
+
0.03.976.550 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|
| 5 |
+
0.04.333.276 I
|
| 6 |
+
0.04.333.426 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
|
| 7 |
+
0.04.333.434 I perplexity: tokenizing the input ..
|
| 8 |
+
0.04.683.706 I perplexity: tokenization took 350.264 ms
|
| 9 |
+
0.04.683.803 I perplexity: calculating perplexity over 4 chunks, n_ctx=2048, batch_size=2048, n_seq=1
|
| 10 |
+
0.07.474.485 I perplexity: 2.79 seconds per pass - ETA 0.18 minutes
|
| 11 |
+
[1]6.2788,[2]7.1110,[3]7.5138,[4]7.8414,
|
| 12 |
+
0.14.413.643 I Final estimate: PPL = 7.8414 +/- 0.32420
|
| 13 |
+
|