kingjones777 commited on
Commit
27a5002
·
verified ·
1 Parent(s): 7f30bda

Add files using upload-large-folder tool

Browse files
This view is limited to 50 files because it contains too many changes.   See raw diff
Files changed (50) hide show
  1. recipe/logs/C1_convert.log +934 -0
  2. recipe/logs/C2_mmproj.log +362 -0
  3. recipe/logs/C_readback.log +2 -0
  4. recipe/logs/D2_verify_download.log +2 -0
  5. recipe/logs/N1_ppl_bf16.log +15 -0
  6. recipe/logs/N1c_ppl_bf16_cpu.log +15 -0
  7. recipe/logs/N2c_imatrix_cpu.log +114 -0
  8. recipe/logs/N3_q102i.log +0 -0
  9. recipe/logs/N3_q103i.log +0 -0
  10. recipe/logs/N3_q106i.log +0 -0
  11. recipe/logs/N3_readback.log +4 -0
  12. recipe/logs/N4_kld_q102i.log +92 -0
  13. recipe/logs/N4_kld_q103i.log +92 -0
  14. recipe/logs/N4_kld_q106.log +92 -0
  15. recipe/logs/N4_kld_q106i.log +92 -0
  16. recipe/logs/N4v_kld_q102i.log +93 -0
  17. recipe/logs/N4v_kld_q103.log +93 -0
  18. recipe/logs/N4v_kld_q103i.log +93 -0
  19. recipe/logs/N4v_kld_q106.log +93 -0
  20. recipe/logs/N4v_kld_q106i.log +93 -0
  21. recipe/logs/N5_kld_q106_repeat.log +92 -0
  22. recipe/logs/N5v_kld_q106_repeat.log +93 -0
  23. recipe/logs/N6t_tools_c1.log +33 -0
  24. recipe/logs/N6t_tools_tpl.log +22 -0
  25. recipe/logs/N6t_tools_tpl_medium.log +31 -0
  26. recipe/logs/N7_sizing.log +5 -0
  27. recipe/logs/N8_unice.log +3 -0
  28. recipe/logs/N8a_seats.log +8 -0
  29. recipe/logs/N8d_seats.log +8 -0
  30. recipe/logs/Q1_q102.log +0 -0
  31. recipe/logs/Q1_q103.log +0 -0
  32. recipe/logs/Q_readback.log +3 -0
  33. recipe/logs/Q_sizes.log +6 -0
  34. recipe/logs/b_n-c3-q106.log +284 -0
  35. recipe/logs/b_n-tools-q106-c1-probe.log +286 -0
  36. recipe/logs/b_n-tools-q106-c1.log +302 -0
  37. recipe/logs/b_n-tools-q106-roff-r3.log +302 -0
  38. recipe/logs/b_n-tools-q106-roff.log +301 -0
  39. recipe/logs/b_n-tools-q106-tpl-medium-probe.log +278 -0
  40. recipe/logs/b_n-tools-q106-tpl-medium.log +294 -0
  41. recipe/logs/b_n-tools-q106.log +302 -0
  42. recipe/logs/b_n-vision-q106-faoff.log +66 -0
  43. recipe/logs/b_n-vision-q106-faon.log +66 -0
  44. recipe/logs/b_n-vision-q106-roff-faon.log +70 -0
  45. recipe/logs/diag_bf16.log +5 -0
  46. recipe/logs/diag_bf16_cpu.log +13 -0
  47. recipe/logs/diag_bf16_purecpu_c1.log +14 -0
  48. recipe/logs/diag_bf16_rocm_faoff.log +13 -0
  49. recipe/logs/diag_bf16_vk_faon.log +13 -0
  50. recipe/logs/diag_ppl_q106_rocm_c4.log +13 -0
recipe/logs/C1_convert.log ADDED
@@ -0,0 +1,934 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ INFO:hf-to-gguf:Loading model: hf
2
+ INFO:hf-to-gguf:Model architecture: Qwen3_5MoeForConditionalGeneration
3
+ INFO:hf-to-gguf:gguf: loading model weight map from 'model.safetensors.index.json'
4
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00001-of-00016.safetensors'
5
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00002-of-00016.safetensors'
6
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00003-of-00016.safetensors'
7
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00004-of-00016.safetensors'
8
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00005-of-00016.safetensors'
9
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00006-of-00016.safetensors'
10
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00007-of-00016.safetensors'
11
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00008-of-00016.safetensors'
12
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00009-of-00016.safetensors'
13
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00010-of-00016.safetensors'
14
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00011-of-00016.safetensors'
15
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00012-of-00016.safetensors'
16
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00013-of-00016.safetensors'
17
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00014-of-00016.safetensors'
18
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00015-of-00016.safetensors'
19
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00016-of-00016.safetensors'
20
+ INFO:gguf.gguf_writer:gguf: This GGUF file is for Little Endian only
21
+ INFO:hf-to-gguf:Exporting model...
22
+ INFO:hf-to-gguf:token_embd.weight, torch.bfloat16 --> BF16, shape = {2048, 248320}
23
+ INFO:hf-to-gguf:blk.0.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
24
+ INFO:hf-to-gguf:blk.0.ssm_a, torch.bfloat16 --> F32, shape = {32}
25
+ INFO:hf-to-gguf:blk.0.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
26
+ INFO:hf-to-gguf:blk.0.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
27
+ INFO:hf-to-gguf:blk.0.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
28
+ INFO:hf-to-gguf:blk.0.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
29
+ INFO:hf-to-gguf:blk.0.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
30
+ INFO:hf-to-gguf:blk.0.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
31
+ INFO:hf-to-gguf:blk.0.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
32
+ INFO:hf-to-gguf:blk.0.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
33
+ INFO:hf-to-gguf:blk.0.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
34
+ INFO:hf-to-gguf:blk.0.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
35
+ INFO:hf-to-gguf:blk.0.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
36
+ INFO:hf-to-gguf:blk.0.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
37
+ INFO:hf-to-gguf:blk.0.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
38
+ INFO:hf-to-gguf:blk.0.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
39
+ INFO:hf-to-gguf:blk.0.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
40
+ INFO:hf-to-gguf:blk.0.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
41
+ INFO:hf-to-gguf:blk.0.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
42
+ INFO:hf-to-gguf:blk.1.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
43
+ INFO:hf-to-gguf:blk.1.ssm_a, torch.bfloat16 --> F32, shape = {32}
44
+ INFO:hf-to-gguf:blk.1.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
45
+ INFO:hf-to-gguf:blk.1.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
46
+ INFO:hf-to-gguf:blk.1.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
47
+ INFO:hf-to-gguf:blk.1.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
48
+ INFO:hf-to-gguf:blk.1.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
49
+ INFO:hf-to-gguf:blk.1.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
50
+ INFO:hf-to-gguf:blk.1.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
51
+ INFO:hf-to-gguf:blk.1.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
52
+ INFO:hf-to-gguf:blk.1.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
53
+ INFO:hf-to-gguf:blk.1.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
54
+ INFO:hf-to-gguf:blk.1.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
55
+ INFO:hf-to-gguf:blk.1.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
56
+ INFO:hf-to-gguf:blk.1.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
57
+ INFO:hf-to-gguf:blk.1.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
58
+ INFO:hf-to-gguf:blk.1.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
59
+ INFO:hf-to-gguf:blk.1.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
60
+ INFO:hf-to-gguf:blk.1.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
61
+ INFO:hf-to-gguf:blk.2.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
62
+ INFO:hf-to-gguf:blk.2.ssm_a, torch.bfloat16 --> F32, shape = {32}
63
+ INFO:hf-to-gguf:blk.2.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
64
+ INFO:hf-to-gguf:blk.2.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
65
+ INFO:hf-to-gguf:blk.2.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
66
+ INFO:hf-to-gguf:blk.2.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
67
+ INFO:hf-to-gguf:blk.2.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
68
+ INFO:hf-to-gguf:blk.2.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
69
+ INFO:hf-to-gguf:blk.2.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
70
+ INFO:hf-to-gguf:blk.2.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
71
+ INFO:hf-to-gguf:blk.2.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
72
+ INFO:hf-to-gguf:blk.2.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
73
+ INFO:hf-to-gguf:blk.2.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
74
+ INFO:hf-to-gguf:blk.2.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
75
+ INFO:hf-to-gguf:blk.2.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
76
+ INFO:hf-to-gguf:blk.2.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
77
+ INFO:hf-to-gguf:blk.2.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
78
+ INFO:hf-to-gguf:blk.2.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
79
+ INFO:hf-to-gguf:blk.2.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
80
+ INFO:hf-to-gguf:blk.3.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
81
+ INFO:hf-to-gguf:blk.3.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
82
+ INFO:hf-to-gguf:blk.3.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
83
+ INFO:hf-to-gguf:blk.3.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
84
+ INFO:hf-to-gguf:blk.3.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
85
+ INFO:hf-to-gguf:blk.3.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
86
+ INFO:hf-to-gguf:blk.3.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
87
+ INFO:hf-to-gguf:blk.3.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
88
+ INFO:hf-to-gguf:blk.3.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
89
+ INFO:hf-to-gguf:blk.3.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
90
+ INFO:hf-to-gguf:blk.3.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
91
+ INFO:hf-to-gguf:blk.3.attn_k.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
92
+ INFO:hf-to-gguf:blk.3.attn_output.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
93
+ INFO:hf-to-gguf:blk.3.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
94
+ INFO:hf-to-gguf:blk.3.attn_q.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
95
+ INFO:hf-to-gguf:blk.3.attn_v.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
96
+ INFO:hf-to-gguf:blk.4.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
97
+ INFO:hf-to-gguf:blk.4.ssm_a, torch.bfloat16 --> F32, shape = {32}
98
+ INFO:hf-to-gguf:blk.4.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
99
+ INFO:hf-to-gguf:blk.4.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
100
+ INFO:hf-to-gguf:blk.4.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
101
+ INFO:hf-to-gguf:blk.4.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
102
+ INFO:hf-to-gguf:blk.4.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
103
+ INFO:hf-to-gguf:blk.4.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
104
+ INFO:hf-to-gguf:blk.4.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
105
+ INFO:hf-to-gguf:blk.4.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
106
+ INFO:hf-to-gguf:blk.4.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
107
+ INFO:hf-to-gguf:blk.4.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
108
+ INFO:hf-to-gguf:blk.4.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
109
+ INFO:hf-to-gguf:blk.4.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
110
+ INFO:hf-to-gguf:blk.4.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
111
+ INFO:hf-to-gguf:blk.4.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
112
+ INFO:hf-to-gguf:blk.4.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
113
+ INFO:hf-to-gguf:blk.4.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
114
+ INFO:hf-to-gguf:blk.4.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
115
+ INFO:hf-to-gguf:blk.5.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
116
+ INFO:hf-to-gguf:blk.5.ssm_a, torch.bfloat16 --> F32, shape = {32}
117
+ INFO:hf-to-gguf:blk.5.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
118
+ INFO:hf-to-gguf:blk.5.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
119
+ INFO:hf-to-gguf:blk.5.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
120
+ INFO:hf-to-gguf:blk.5.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
121
+ INFO:hf-to-gguf:blk.5.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
122
+ INFO:hf-to-gguf:blk.5.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
123
+ INFO:hf-to-gguf:blk.5.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
124
+ INFO:hf-to-gguf:blk.5.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
125
+ INFO:hf-to-gguf:blk.5.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
126
+ INFO:hf-to-gguf:blk.5.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
127
+ INFO:hf-to-gguf:blk.5.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
128
+ INFO:hf-to-gguf:blk.5.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
129
+ INFO:hf-to-gguf:blk.5.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
130
+ INFO:hf-to-gguf:blk.5.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
131
+ INFO:hf-to-gguf:blk.5.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
132
+ INFO:hf-to-gguf:blk.5.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
133
+ INFO:hf-to-gguf:blk.5.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
134
+ INFO:hf-to-gguf:blk.6.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
135
+ INFO:hf-to-gguf:blk.6.ssm_a, torch.bfloat16 --> F32, shape = {32}
136
+ INFO:hf-to-gguf:blk.6.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
137
+ INFO:hf-to-gguf:blk.6.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
138
+ INFO:hf-to-gguf:blk.6.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
139
+ INFO:hf-to-gguf:blk.6.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
140
+ INFO:hf-to-gguf:blk.6.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
141
+ INFO:hf-to-gguf:blk.6.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
142
+ INFO:hf-to-gguf:blk.6.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
143
+ INFO:hf-to-gguf:blk.6.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
144
+ INFO:hf-to-gguf:blk.6.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
145
+ INFO:hf-to-gguf:blk.6.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
146
+ INFO:hf-to-gguf:blk.6.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
147
+ INFO:hf-to-gguf:blk.6.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
148
+ INFO:hf-to-gguf:blk.6.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
149
+ INFO:hf-to-gguf:blk.6.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
150
+ INFO:hf-to-gguf:blk.6.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
151
+ INFO:hf-to-gguf:blk.6.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
152
+ INFO:hf-to-gguf:blk.6.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
153
+ INFO:hf-to-gguf:blk.7.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
154
+ INFO:hf-to-gguf:blk.7.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
155
+ INFO:hf-to-gguf:blk.7.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
156
+ INFO:hf-to-gguf:blk.7.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
157
+ INFO:hf-to-gguf:blk.7.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
158
+ INFO:hf-to-gguf:blk.7.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
159
+ INFO:hf-to-gguf:blk.7.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
160
+ INFO:hf-to-gguf:blk.7.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
161
+ INFO:hf-to-gguf:blk.7.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
162
+ INFO:hf-to-gguf:blk.7.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
163
+ INFO:hf-to-gguf:blk.7.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
164
+ INFO:hf-to-gguf:blk.7.attn_k.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
165
+ INFO:hf-to-gguf:blk.7.attn_output.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
166
+ INFO:hf-to-gguf:blk.7.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
167
+ INFO:hf-to-gguf:blk.7.attn_q.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
168
+ INFO:hf-to-gguf:blk.7.attn_v.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
169
+ INFO:hf-to-gguf:blk.8.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
170
+ INFO:hf-to-gguf:blk.8.ssm_a, torch.bfloat16 --> F32, shape = {32}
171
+ INFO:hf-to-gguf:blk.8.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
172
+ INFO:hf-to-gguf:blk.8.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
173
+ INFO:hf-to-gguf:blk.8.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
174
+ INFO:hf-to-gguf:blk.8.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
175
+ INFO:hf-to-gguf:blk.8.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
176
+ INFO:hf-to-gguf:blk.8.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
177
+ INFO:hf-to-gguf:blk.8.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
178
+ INFO:hf-to-gguf:blk.8.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
179
+ INFO:hf-to-gguf:blk.8.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
180
+ INFO:hf-to-gguf:blk.8.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
181
+ INFO:hf-to-gguf:blk.8.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
182
+ INFO:hf-to-gguf:blk.8.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
183
+ INFO:hf-to-gguf:blk.8.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
184
+ INFO:hf-to-gguf:blk.8.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
185
+ INFO:hf-to-gguf:blk.8.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
186
+ INFO:hf-to-gguf:blk.8.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
187
+ INFO:hf-to-gguf:blk.8.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
188
+ INFO:hf-to-gguf:blk.9.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
189
+ INFO:hf-to-gguf:blk.9.ssm_a, torch.bfloat16 --> F32, shape = {32}
190
+ INFO:hf-to-gguf:blk.9.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
191
+ INFO:hf-to-gguf:blk.9.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
192
+ INFO:hf-to-gguf:blk.9.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
193
+ INFO:hf-to-gguf:blk.9.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
194
+ INFO:hf-to-gguf:blk.9.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
195
+ INFO:hf-to-gguf:blk.9.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
196
+ INFO:hf-to-gguf:blk.9.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
197
+ INFO:hf-to-gguf:blk.9.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
198
+ INFO:hf-to-gguf:blk.9.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
199
+ INFO:hf-to-gguf:blk.9.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
200
+ INFO:hf-to-gguf:blk.9.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
201
+ INFO:hf-to-gguf:blk.9.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
202
+ INFO:hf-to-gguf:blk.9.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
203
+ INFO:hf-to-gguf:blk.9.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
204
+ INFO:hf-to-gguf:blk.9.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
205
+ INFO:hf-to-gguf:blk.10.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
206
+ INFO:hf-to-gguf:blk.10.ssm_a, torch.bfloat16 --> F32, shape = {32}
207
+ INFO:hf-to-gguf:blk.10.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
208
+ INFO:hf-to-gguf:blk.10.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
209
+ INFO:hf-to-gguf:blk.10.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
210
+ INFO:hf-to-gguf:blk.10.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
211
+ INFO:hf-to-gguf:blk.10.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
212
+ INFO:hf-to-gguf:blk.10.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
213
+ INFO:hf-to-gguf:blk.10.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
214
+ INFO:hf-to-gguf:blk.10.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
215
+ INFO:hf-to-gguf:blk.10.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
216
+ INFO:hf-to-gguf:blk.10.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
217
+ INFO:hf-to-gguf:blk.10.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
218
+ INFO:hf-to-gguf:blk.10.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
219
+ INFO:hf-to-gguf:blk.10.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
220
+ INFO:hf-to-gguf:blk.10.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
221
+ INFO:hf-to-gguf:blk.10.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
222
+ INFO:hf-to-gguf:blk.10.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
223
+ INFO:hf-to-gguf:blk.10.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
224
+ INFO:hf-to-gguf:blk.11.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
225
+ INFO:hf-to-gguf:blk.11.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
226
+ INFO:hf-to-gguf:blk.11.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
227
+ INFO:hf-to-gguf:blk.11.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
228
+ INFO:hf-to-gguf:blk.11.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
229
+ INFO:hf-to-gguf:blk.11.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
230
+ INFO:hf-to-gguf:blk.11.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
231
+ INFO:hf-to-gguf:blk.11.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
232
+ INFO:hf-to-gguf:blk.11.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
233
+ INFO:hf-to-gguf:blk.11.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
234
+ INFO:hf-to-gguf:blk.11.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
235
+ INFO:hf-to-gguf:blk.11.attn_k.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
236
+ INFO:hf-to-gguf:blk.11.attn_output.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
237
+ INFO:hf-to-gguf:blk.11.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
238
+ INFO:hf-to-gguf:blk.11.attn_q.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
239
+ INFO:hf-to-gguf:blk.11.attn_v.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
240
+ INFO:hf-to-gguf:blk.12.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
241
+ INFO:hf-to-gguf:blk.12.ssm_a, torch.bfloat16 --> F32, shape = {32}
242
+ INFO:hf-to-gguf:blk.12.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
243
+ INFO:hf-to-gguf:blk.12.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
244
+ INFO:hf-to-gguf:blk.12.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
245
+ INFO:hf-to-gguf:blk.12.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
246
+ INFO:hf-to-gguf:blk.12.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
247
+ INFO:hf-to-gguf:blk.12.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
248
+ INFO:hf-to-gguf:blk.12.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
249
+ INFO:hf-to-gguf:blk.12.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
250
+ INFO:hf-to-gguf:blk.12.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
251
+ INFO:hf-to-gguf:blk.12.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
252
+ INFO:hf-to-gguf:blk.12.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
253
+ INFO:hf-to-gguf:blk.12.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
254
+ INFO:hf-to-gguf:blk.12.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
255
+ INFO:hf-to-gguf:blk.9.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
256
+ INFO:hf-to-gguf:blk.9.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
257
+ INFO:hf-to-gguf:blk.12.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
258
+ INFO:hf-to-gguf:blk.12.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
259
+ INFO:hf-to-gguf:blk.12.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
260
+ INFO:hf-to-gguf:blk.12.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
261
+ INFO:hf-to-gguf:blk.13.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
262
+ INFO:hf-to-gguf:blk.13.ssm_a, torch.bfloat16 --> F32, shape = {32}
263
+ INFO:hf-to-gguf:blk.13.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
264
+ INFO:hf-to-gguf:blk.13.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
265
+ INFO:hf-to-gguf:blk.13.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
266
+ INFO:hf-to-gguf:blk.13.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
267
+ INFO:hf-to-gguf:blk.13.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
268
+ INFO:hf-to-gguf:blk.13.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
269
+ INFO:hf-to-gguf:blk.13.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
270
+ INFO:hf-to-gguf:blk.13.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
271
+ INFO:hf-to-gguf:blk.13.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
272
+ INFO:hf-to-gguf:blk.13.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
273
+ INFO:hf-to-gguf:blk.13.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
274
+ INFO:hf-to-gguf:blk.13.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
275
+ INFO:hf-to-gguf:blk.13.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
276
+ INFO:hf-to-gguf:blk.13.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
277
+ INFO:hf-to-gguf:blk.13.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
278
+ INFO:hf-to-gguf:blk.13.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
279
+ INFO:hf-to-gguf:blk.13.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
280
+ INFO:hf-to-gguf:blk.14.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
281
+ INFO:hf-to-gguf:blk.14.ssm_a, torch.bfloat16 --> F32, shape = {32}
282
+ INFO:hf-to-gguf:blk.14.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
283
+ INFO:hf-to-gguf:blk.14.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
284
+ INFO:hf-to-gguf:blk.14.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
285
+ INFO:hf-to-gguf:blk.14.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
286
+ INFO:hf-to-gguf:blk.14.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
287
+ INFO:hf-to-gguf:blk.14.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
288
+ INFO:hf-to-gguf:blk.14.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
289
+ INFO:hf-to-gguf:blk.14.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
290
+ INFO:hf-to-gguf:blk.14.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
291
+ INFO:hf-to-gguf:blk.14.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
292
+ INFO:hf-to-gguf:blk.14.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
293
+ INFO:hf-to-gguf:blk.14.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
294
+ INFO:hf-to-gguf:blk.14.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
295
+ INFO:hf-to-gguf:blk.14.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
296
+ INFO:hf-to-gguf:blk.14.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
297
+ INFO:hf-to-gguf:blk.14.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
298
+ INFO:hf-to-gguf:blk.14.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
299
+ INFO:hf-to-gguf:blk.15.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
300
+ INFO:hf-to-gguf:blk.15.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
301
+ INFO:hf-to-gguf:blk.15.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
302
+ INFO:hf-to-gguf:blk.15.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
303
+ INFO:hf-to-gguf:blk.15.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
304
+ INFO:hf-to-gguf:blk.15.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
305
+ INFO:hf-to-gguf:blk.15.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
306
+ INFO:hf-to-gguf:blk.15.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
307
+ INFO:hf-to-gguf:blk.15.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
308
+ INFO:hf-to-gguf:blk.15.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
309
+ INFO:hf-to-gguf:blk.15.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
310
+ INFO:hf-to-gguf:blk.15.attn_k.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
311
+ INFO:hf-to-gguf:blk.15.attn_output.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
312
+ INFO:hf-to-gguf:blk.15.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
313
+ INFO:hf-to-gguf:blk.15.attn_q.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
314
+ INFO:hf-to-gguf:blk.15.attn_v.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
315
+ INFO:hf-to-gguf:blk.16.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
316
+ INFO:hf-to-gguf:blk.16.ssm_a, torch.bfloat16 --> F32, shape = {32}
317
+ INFO:hf-to-gguf:blk.16.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
318
+ INFO:hf-to-gguf:blk.16.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
319
+ INFO:hf-to-gguf:blk.16.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
320
+ INFO:hf-to-gguf:blk.16.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
321
+ INFO:hf-to-gguf:blk.16.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
322
+ INFO:hf-to-gguf:blk.16.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
323
+ INFO:hf-to-gguf:blk.16.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
324
+ INFO:hf-to-gguf:blk.16.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
325
+ INFO:hf-to-gguf:blk.16.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
326
+ INFO:hf-to-gguf:blk.16.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
327
+ INFO:hf-to-gguf:blk.16.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
328
+ INFO:hf-to-gguf:blk.16.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
329
+ INFO:hf-to-gguf:blk.16.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
330
+ INFO:hf-to-gguf:blk.16.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
331
+ INFO:hf-to-gguf:blk.16.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
332
+ INFO:hf-to-gguf:blk.16.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
333
+ INFO:hf-to-gguf:blk.16.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
334
+ INFO:hf-to-gguf:blk.17.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
335
+ INFO:hf-to-gguf:blk.17.ssm_a, torch.bfloat16 --> F32, shape = {32}
336
+ INFO:hf-to-gguf:blk.17.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
337
+ INFO:hf-to-gguf:blk.17.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
338
+ INFO:hf-to-gguf:blk.17.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
339
+ INFO:hf-to-gguf:blk.17.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
340
+ INFO:hf-to-gguf:blk.17.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
341
+ INFO:hf-to-gguf:blk.17.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
342
+ INFO:hf-to-gguf:blk.17.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
343
+ INFO:hf-to-gguf:blk.17.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
344
+ INFO:hf-to-gguf:blk.17.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
345
+ INFO:hf-to-gguf:blk.17.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
346
+ INFO:hf-to-gguf:blk.17.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
347
+ INFO:hf-to-gguf:blk.17.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
348
+ INFO:hf-to-gguf:blk.17.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
349
+ INFO:hf-to-gguf:blk.17.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
350
+ INFO:hf-to-gguf:blk.17.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
351
+ INFO:hf-to-gguf:blk.17.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
352
+ INFO:hf-to-gguf:blk.17.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
353
+ INFO:hf-to-gguf:blk.18.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
354
+ INFO:hf-to-gguf:blk.18.ssm_a, torch.bfloat16 --> F32, shape = {32}
355
+ INFO:hf-to-gguf:blk.18.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
356
+ INFO:hf-to-gguf:blk.18.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
357
+ INFO:hf-to-gguf:blk.18.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
358
+ INFO:hf-to-gguf:blk.18.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
359
+ INFO:hf-to-gguf:blk.18.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
360
+ INFO:hf-to-gguf:blk.18.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
361
+ INFO:hf-to-gguf:blk.18.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
362
+ INFO:hf-to-gguf:blk.18.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
363
+ INFO:hf-to-gguf:blk.18.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
364
+ INFO:hf-to-gguf:blk.18.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
365
+ INFO:hf-to-gguf:blk.18.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
366
+ INFO:hf-to-gguf:blk.18.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
367
+ INFO:hf-to-gguf:blk.18.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
368
+ INFO:hf-to-gguf:blk.18.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
369
+ INFO:hf-to-gguf:blk.18.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
370
+ INFO:hf-to-gguf:blk.18.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
371
+ INFO:hf-to-gguf:blk.18.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
372
+ INFO:hf-to-gguf:blk.19.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
373
+ INFO:hf-to-gguf:blk.19.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
374
+ INFO:hf-to-gguf:blk.19.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
375
+ INFO:hf-to-gguf:blk.19.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
376
+ INFO:hf-to-gguf:blk.19.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
377
+ INFO:hf-to-gguf:blk.19.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
378
+ INFO:hf-to-gguf:blk.19.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
379
+ INFO:hf-to-gguf:blk.19.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
380
+ INFO:hf-to-gguf:blk.19.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
381
+ INFO:hf-to-gguf:blk.19.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
382
+ INFO:hf-to-gguf:blk.19.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
383
+ INFO:hf-to-gguf:blk.19.attn_k.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
384
+ INFO:hf-to-gguf:blk.19.attn_output.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
385
+ INFO:hf-to-gguf:blk.19.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
386
+ INFO:hf-to-gguf:blk.19.attn_q.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
387
+ INFO:hf-to-gguf:blk.19.attn_v.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
388
+ INFO:hf-to-gguf:blk.20.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
389
+ INFO:hf-to-gguf:blk.20.ssm_a, torch.bfloat16 --> F32, shape = {32}
390
+ INFO:hf-to-gguf:blk.20.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
391
+ INFO:hf-to-gguf:blk.20.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
392
+ INFO:hf-to-gguf:blk.20.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
393
+ INFO:hf-to-gguf:blk.20.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
394
+ INFO:hf-to-gguf:blk.20.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
395
+ INFO:hf-to-gguf:blk.20.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
396
+ INFO:hf-to-gguf:blk.20.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
397
+ INFO:hf-to-gguf:blk.20.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
398
+ INFO:hf-to-gguf:blk.20.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
399
+ INFO:hf-to-gguf:blk.20.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
400
+ INFO:hf-to-gguf:blk.20.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
401
+ INFO:hf-to-gguf:blk.20.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
402
+ INFO:hf-to-gguf:blk.20.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
403
+ INFO:hf-to-gguf:blk.20.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
404
+ INFO:hf-to-gguf:blk.20.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
405
+ INFO:hf-to-gguf:blk.20.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
406
+ INFO:hf-to-gguf:blk.20.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
407
+ INFO:hf-to-gguf:blk.21.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
408
+ INFO:hf-to-gguf:blk.21.ssm_a, torch.bfloat16 --> F32, shape = {32}
409
+ INFO:hf-to-gguf:blk.21.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
410
+ INFO:hf-to-gguf:blk.21.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
411
+ INFO:hf-to-gguf:blk.21.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
412
+ INFO:hf-to-gguf:blk.21.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
413
+ INFO:hf-to-gguf:blk.21.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
414
+ INFO:hf-to-gguf:blk.21.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
415
+ INFO:hf-to-gguf:blk.21.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
416
+ INFO:hf-to-gguf:blk.21.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
417
+ INFO:hf-to-gguf:blk.21.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
418
+ INFO:hf-to-gguf:blk.21.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
419
+ INFO:hf-to-gguf:blk.21.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
420
+ INFO:hf-to-gguf:blk.21.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
421
+ INFO:hf-to-gguf:blk.21.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
422
+ INFO:hf-to-gguf:blk.21.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
423
+ INFO:hf-to-gguf:blk.21.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
424
+ INFO:hf-to-gguf:blk.21.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
425
+ INFO:hf-to-gguf:blk.21.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
426
+ INFO:hf-to-gguf:blk.22.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
427
+ INFO:hf-to-gguf:blk.22.ssm_a, torch.bfloat16 --> F32, shape = {32}
428
+ INFO:hf-to-gguf:blk.22.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
429
+ INFO:hf-to-gguf:blk.22.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
430
+ INFO:hf-to-gguf:blk.22.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
431
+ INFO:hf-to-gguf:blk.22.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
432
+ INFO:hf-to-gguf:blk.22.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
433
+ INFO:hf-to-gguf:blk.22.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
434
+ INFO:hf-to-gguf:blk.22.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
435
+ INFO:hf-to-gguf:blk.22.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
436
+ INFO:hf-to-gguf:blk.22.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
437
+ INFO:hf-to-gguf:blk.22.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
438
+ INFO:hf-to-gguf:blk.22.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
439
+ INFO:hf-to-gguf:blk.22.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
440
+ INFO:hf-to-gguf:blk.22.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
441
+ INFO:hf-to-gguf:blk.22.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
442
+ INFO:hf-to-gguf:blk.22.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
443
+ INFO:hf-to-gguf:blk.22.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
444
+ INFO:hf-to-gguf:blk.22.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
445
+ INFO:hf-to-gguf:blk.23.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
446
+ INFO:hf-to-gguf:blk.23.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
447
+ INFO:hf-to-gguf:blk.23.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
448
+ INFO:hf-to-gguf:blk.23.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
449
+ INFO:hf-to-gguf:blk.23.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
450
+ INFO:hf-to-gguf:blk.23.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
451
+ INFO:hf-to-gguf:blk.23.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
452
+ INFO:hf-to-gguf:blk.23.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
453
+ INFO:hf-to-gguf:blk.23.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
454
+ INFO:hf-to-gguf:blk.23.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
455
+ INFO:hf-to-gguf:blk.23.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
456
+ INFO:hf-to-gguf:blk.23.attn_k.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
457
+ INFO:hf-to-gguf:blk.23.attn_output.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
458
+ INFO:hf-to-gguf:blk.23.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
459
+ INFO:hf-to-gguf:blk.23.attn_q.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
460
+ INFO:hf-to-gguf:blk.23.attn_v.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
461
+ INFO:hf-to-gguf:blk.24.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
462
+ INFO:hf-to-gguf:blk.24.ssm_a, torch.bfloat16 --> F32, shape = {32}
463
+ INFO:hf-to-gguf:blk.24.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
464
+ INFO:hf-to-gguf:blk.24.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
465
+ INFO:hf-to-gguf:blk.24.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
466
+ INFO:hf-to-gguf:blk.24.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
467
+ INFO:hf-to-gguf:blk.24.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
468
+ INFO:hf-to-gguf:blk.24.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
469
+ INFO:hf-to-gguf:blk.24.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
470
+ INFO:hf-to-gguf:blk.24.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
471
+ INFO:hf-to-gguf:blk.24.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
472
+ INFO:hf-to-gguf:blk.24.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
473
+ INFO:hf-to-gguf:blk.24.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
474
+ INFO:hf-to-gguf:blk.24.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
475
+ INFO:hf-to-gguf:blk.24.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
476
+ INFO:hf-to-gguf:blk.24.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
477
+ INFO:hf-to-gguf:blk.24.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
478
+ INFO:hf-to-gguf:blk.24.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
479
+ INFO:hf-to-gguf:blk.24.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
480
+ INFO:hf-to-gguf:blk.25.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
481
+ INFO:hf-to-gguf:blk.25.ssm_a, torch.bfloat16 --> F32, shape = {32}
482
+ INFO:hf-to-gguf:blk.25.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
483
+ INFO:hf-to-gguf:blk.25.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
484
+ INFO:hf-to-gguf:blk.25.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
485
+ INFO:hf-to-gguf:blk.25.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
486
+ INFO:hf-to-gguf:blk.25.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
487
+ INFO:hf-to-gguf:blk.25.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
488
+ INFO:hf-to-gguf:blk.25.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
489
+ INFO:hf-to-gguf:blk.25.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
490
+ INFO:hf-to-gguf:blk.25.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
491
+ INFO:hf-to-gguf:blk.25.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
492
+ INFO:hf-to-gguf:blk.25.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
493
+ INFO:hf-to-gguf:blk.25.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
494
+ INFO:hf-to-gguf:blk.25.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
495
+ INFO:hf-to-gguf:blk.25.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
496
+ INFO:hf-to-gguf:blk.25.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
497
+ INFO:hf-to-gguf:blk.25.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
498
+ INFO:hf-to-gguf:blk.25.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
499
+ INFO:hf-to-gguf:blk.26.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
500
+ INFO:hf-to-gguf:blk.26.ssm_a, torch.bfloat16 --> F32, shape = {32}
501
+ INFO:hf-to-gguf:blk.26.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
502
+ INFO:hf-to-gguf:blk.26.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
503
+ INFO:hf-to-gguf:blk.26.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
504
+ INFO:hf-to-gguf:blk.26.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
505
+ INFO:hf-to-gguf:blk.26.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
506
+ INFO:hf-to-gguf:blk.26.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
507
+ INFO:hf-to-gguf:blk.26.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
508
+ INFO:hf-to-gguf:blk.26.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
509
+ INFO:hf-to-gguf:blk.26.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
510
+ INFO:hf-to-gguf:blk.26.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
511
+ INFO:hf-to-gguf:blk.26.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
512
+ INFO:hf-to-gguf:blk.26.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
513
+ INFO:hf-to-gguf:blk.26.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
514
+ INFO:hf-to-gguf:blk.26.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
515
+ INFO:hf-to-gguf:blk.26.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
516
+ INFO:hf-to-gguf:blk.26.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
517
+ INFO:hf-to-gguf:blk.26.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
518
+ INFO:hf-to-gguf:blk.27.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
519
+ INFO:hf-to-gguf:blk.27.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
520
+ INFO:hf-to-gguf:blk.27.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
521
+ INFO:hf-to-gguf:blk.27.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
522
+ INFO:hf-to-gguf:blk.27.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
523
+ INFO:hf-to-gguf:blk.27.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
524
+ INFO:hf-to-gguf:blk.27.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
525
+ INFO:hf-to-gguf:blk.27.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
526
+ INFO:hf-to-gguf:blk.27.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
527
+ INFO:hf-to-gguf:blk.27.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
528
+ INFO:hf-to-gguf:blk.27.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
529
+ INFO:hf-to-gguf:blk.27.attn_k.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
530
+ INFO:hf-to-gguf:blk.27.attn_output.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
531
+ INFO:hf-to-gguf:blk.27.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
532
+ INFO:hf-to-gguf:blk.27.attn_q.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
533
+ INFO:hf-to-gguf:blk.27.attn_v.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
534
+ INFO:hf-to-gguf:blk.28.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
535
+ INFO:hf-to-gguf:blk.28.ssm_a, torch.bfloat16 --> F32, shape = {32}
536
+ INFO:hf-to-gguf:blk.28.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
537
+ INFO:hf-to-gguf:blk.28.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
538
+ INFO:hf-to-gguf:blk.28.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
539
+ INFO:hf-to-gguf:blk.28.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
540
+ INFO:hf-to-gguf:blk.28.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
541
+ INFO:hf-to-gguf:blk.28.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
542
+ INFO:hf-to-gguf:blk.28.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
543
+ INFO:hf-to-gguf:blk.28.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
544
+ INFO:hf-to-gguf:blk.28.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
545
+ INFO:hf-to-gguf:blk.28.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
546
+ INFO:hf-to-gguf:blk.28.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
547
+ INFO:hf-to-gguf:blk.28.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
548
+ INFO:hf-to-gguf:blk.28.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
549
+ INFO:hf-to-gguf:blk.28.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
550
+ INFO:hf-to-gguf:blk.28.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
551
+ INFO:hf-to-gguf:blk.28.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
552
+ INFO:hf-to-gguf:blk.28.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
553
+ INFO:hf-to-gguf:blk.29.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
554
+ INFO:hf-to-gguf:blk.29.ssm_a, torch.bfloat16 --> F32, shape = {32}
555
+ INFO:hf-to-gguf:blk.29.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
556
+ INFO:hf-to-gguf:blk.29.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
557
+ INFO:hf-to-gguf:blk.29.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
558
+ INFO:hf-to-gguf:blk.29.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
559
+ INFO:hf-to-gguf:blk.29.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
560
+ INFO:hf-to-gguf:blk.29.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
561
+ INFO:hf-to-gguf:blk.29.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
562
+ INFO:hf-to-gguf:blk.29.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
563
+ INFO:hf-to-gguf:blk.29.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
564
+ INFO:hf-to-gguf:blk.29.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
565
+ INFO:hf-to-gguf:blk.29.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
566
+ INFO:hf-to-gguf:blk.29.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
567
+ INFO:hf-to-gguf:blk.29.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
568
+ INFO:hf-to-gguf:blk.29.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
569
+ INFO:hf-to-gguf:blk.29.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
570
+ INFO:hf-to-gguf:blk.29.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
571
+ INFO:hf-to-gguf:blk.29.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
572
+ INFO:hf-to-gguf:blk.30.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
573
+ INFO:hf-to-gguf:blk.30.ssm_a, torch.bfloat16 --> F32, shape = {32}
574
+ INFO:hf-to-gguf:blk.30.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
575
+ INFO:hf-to-gguf:blk.30.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
576
+ INFO:hf-to-gguf:blk.30.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
577
+ INFO:hf-to-gguf:blk.30.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
578
+ INFO:hf-to-gguf:blk.30.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
579
+ INFO:hf-to-gguf:blk.30.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
580
+ INFO:hf-to-gguf:blk.30.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
581
+ INFO:hf-to-gguf:blk.30.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
582
+ INFO:hf-to-gguf:blk.30.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
583
+ INFO:hf-to-gguf:blk.30.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
584
+ INFO:hf-to-gguf:blk.30.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
585
+ INFO:hf-to-gguf:blk.30.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
586
+ INFO:hf-to-gguf:blk.30.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
587
+ INFO:hf-to-gguf:blk.30.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
588
+ INFO:hf-to-gguf:blk.30.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
589
+ INFO:hf-to-gguf:blk.30.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
590
+ INFO:hf-to-gguf:blk.30.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
591
+ INFO:hf-to-gguf:blk.31.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
592
+ INFO:hf-to-gguf:blk.31.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
593
+ INFO:hf-to-gguf:blk.31.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
594
+ INFO:hf-to-gguf:blk.31.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
595
+ INFO:hf-to-gguf:blk.31.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
596
+ INFO:hf-to-gguf:blk.31.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
597
+ INFO:hf-to-gguf:blk.31.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
598
+ INFO:hf-to-gguf:blk.31.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
599
+ INFO:hf-to-gguf:blk.31.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
600
+ INFO:hf-to-gguf:blk.31.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
601
+ INFO:hf-to-gguf:blk.31.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
602
+ INFO:hf-to-gguf:blk.31.attn_k.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
603
+ INFO:hf-to-gguf:blk.31.attn_output.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
604
+ INFO:hf-to-gguf:blk.31.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
605
+ INFO:hf-to-gguf:blk.31.attn_q.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
606
+ INFO:hf-to-gguf:blk.31.attn_v.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
607
+ INFO:hf-to-gguf:blk.32.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
608
+ INFO:hf-to-gguf:blk.32.ssm_a, torch.bfloat16 --> F32, shape = {32}
609
+ INFO:hf-to-gguf:blk.32.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
610
+ INFO:hf-to-gguf:blk.32.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
611
+ INFO:hf-to-gguf:blk.32.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
612
+ INFO:hf-to-gguf:blk.32.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
613
+ INFO:hf-to-gguf:blk.32.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
614
+ INFO:hf-to-gguf:blk.32.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
615
+ INFO:hf-to-gguf:blk.32.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
616
+ INFO:hf-to-gguf:blk.32.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
617
+ INFO:hf-to-gguf:blk.32.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
618
+ INFO:hf-to-gguf:blk.32.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
619
+ INFO:hf-to-gguf:blk.32.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
620
+ INFO:hf-to-gguf:blk.32.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
621
+ INFO:hf-to-gguf:blk.32.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
622
+ INFO:hf-to-gguf:blk.32.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
623
+ INFO:hf-to-gguf:blk.32.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
624
+ INFO:hf-to-gguf:blk.32.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
625
+ INFO:hf-to-gguf:blk.32.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
626
+ INFO:hf-to-gguf:blk.33.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
627
+ INFO:hf-to-gguf:blk.33.ssm_a, torch.bfloat16 --> F32, shape = {32}
628
+ INFO:hf-to-gguf:blk.33.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
629
+ INFO:hf-to-gguf:blk.33.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
630
+ INFO:hf-to-gguf:blk.33.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
631
+ INFO:hf-to-gguf:blk.33.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
632
+ INFO:hf-to-gguf:blk.33.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
633
+ INFO:hf-to-gguf:blk.33.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
634
+ INFO:hf-to-gguf:blk.33.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
635
+ INFO:hf-to-gguf:blk.33.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
636
+ INFO:hf-to-gguf:blk.33.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
637
+ INFO:hf-to-gguf:blk.33.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
638
+ INFO:hf-to-gguf:blk.33.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
639
+ INFO:hf-to-gguf:blk.33.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
640
+ INFO:hf-to-gguf:blk.33.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
641
+ INFO:hf-to-gguf:blk.33.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
642
+ INFO:hf-to-gguf:blk.33.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
643
+ INFO:hf-to-gguf:blk.33.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
644
+ INFO:hf-to-gguf:blk.33.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
645
+ INFO:hf-to-gguf:blk.34.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
646
+ INFO:hf-to-gguf:blk.34.ssm_a, torch.bfloat16 --> F32, shape = {32}
647
+ INFO:hf-to-gguf:blk.34.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
648
+ INFO:hf-to-gguf:blk.34.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
649
+ INFO:hf-to-gguf:blk.34.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
650
+ INFO:hf-to-gguf:blk.34.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
651
+ INFO:hf-to-gguf:blk.34.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
652
+ INFO:hf-to-gguf:blk.34.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
653
+ INFO:hf-to-gguf:blk.34.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
654
+ INFO:hf-to-gguf:blk.34.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
655
+ INFO:hf-to-gguf:blk.34.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
656
+ INFO:hf-to-gguf:blk.34.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
657
+ INFO:hf-to-gguf:blk.34.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
658
+ INFO:hf-to-gguf:blk.34.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
659
+ INFO:hf-to-gguf:blk.34.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
660
+ INFO:hf-to-gguf:blk.34.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
661
+ INFO:hf-to-gguf:blk.34.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
662
+ INFO:hf-to-gguf:blk.34.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
663
+ INFO:hf-to-gguf:blk.34.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
664
+ INFO:hf-to-gguf:blk.35.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
665
+ INFO:hf-to-gguf:blk.35.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
666
+ INFO:hf-to-gguf:blk.35.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
667
+ INFO:hf-to-gguf:blk.35.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
668
+ INFO:hf-to-gguf:blk.35.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
669
+ INFO:hf-to-gguf:blk.35.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
670
+ INFO:hf-to-gguf:blk.35.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
671
+ INFO:hf-to-gguf:blk.35.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
672
+ INFO:hf-to-gguf:blk.35.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
673
+ INFO:hf-to-gguf:blk.35.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
674
+ INFO:hf-to-gguf:blk.35.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
675
+ INFO:hf-to-gguf:blk.35.attn_k.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
676
+ INFO:hf-to-gguf:blk.35.attn_output.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
677
+ INFO:hf-to-gguf:blk.35.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
678
+ INFO:hf-to-gguf:blk.35.attn_q.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
679
+ INFO:hf-to-gguf:blk.35.attn_v.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
680
+ INFO:hf-to-gguf:blk.36.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
681
+ INFO:hf-to-gguf:blk.36.ssm_a, torch.bfloat16 --> F32, shape = {32}
682
+ INFO:hf-to-gguf:blk.36.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
683
+ INFO:hf-to-gguf:blk.36.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
684
+ INFO:hf-to-gguf:blk.36.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
685
+ INFO:hf-to-gguf:blk.36.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
686
+ INFO:hf-to-gguf:blk.36.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
687
+ INFO:hf-to-gguf:blk.36.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
688
+ INFO:hf-to-gguf:blk.36.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
689
+ INFO:hf-to-gguf:blk.36.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
690
+ INFO:hf-to-gguf:blk.36.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
691
+ INFO:hf-to-gguf:blk.36.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
692
+ INFO:hf-to-gguf:blk.36.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
693
+ INFO:hf-to-gguf:blk.36.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
694
+ INFO:hf-to-gguf:blk.36.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
695
+ INFO:hf-to-gguf:blk.36.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
696
+ INFO:hf-to-gguf:blk.36.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
697
+ INFO:hf-to-gguf:blk.36.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
698
+ INFO:hf-to-gguf:blk.36.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
699
+ INFO:hf-to-gguf:blk.37.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
700
+ INFO:hf-to-gguf:blk.37.ssm_a, torch.bfloat16 --> F32, shape = {32}
701
+ INFO:hf-to-gguf:blk.37.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
702
+ INFO:hf-to-gguf:blk.37.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
703
+ INFO:hf-to-gguf:blk.37.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
704
+ INFO:hf-to-gguf:blk.37.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
705
+ INFO:hf-to-gguf:blk.37.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
706
+ INFO:hf-to-gguf:blk.37.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
707
+ INFO:hf-to-gguf:blk.37.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
708
+ INFO:hf-to-gguf:blk.37.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
709
+ INFO:hf-to-gguf:blk.37.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
710
+ INFO:hf-to-gguf:blk.37.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
711
+ INFO:hf-to-gguf:blk.37.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
712
+ INFO:hf-to-gguf:blk.37.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
713
+ INFO:hf-to-gguf:blk.37.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
714
+ INFO:hf-to-gguf:blk.37.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
715
+ INFO:hf-to-gguf:blk.37.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
716
+ INFO:hf-to-gguf:blk.37.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
717
+ INFO:hf-to-gguf:blk.37.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
718
+ INFO:hf-to-gguf:blk.38.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
719
+ INFO:hf-to-gguf:blk.38.ssm_a, torch.bfloat16 --> F32, shape = {32}
720
+ INFO:hf-to-gguf:blk.38.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 8192}
721
+ INFO:hf-to-gguf:blk.38.ssm_dt.bias, torch.bfloat16 --> F32, shape = {32}
722
+ INFO:hf-to-gguf:blk.38.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
723
+ INFO:hf-to-gguf:blk.38.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {2048, 32}
724
+ INFO:hf-to-gguf:blk.38.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
725
+ INFO:hf-to-gguf:blk.38.attn_gate.weight, torch.bfloat16 --> BF16, shape = {2048, 4096}
726
+ INFO:hf-to-gguf:blk.38.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
727
+ INFO:hf-to-gguf:blk.38.ssm_out.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
728
+ INFO:hf-to-gguf:blk.38.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
729
+ INFO:hf-to-gguf:blk.38.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
730
+ INFO:hf-to-gguf:blk.38.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
731
+ INFO:hf-to-gguf:blk.38.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
732
+ INFO:hf-to-gguf:blk.38.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
733
+ INFO:hf-to-gguf:blk.38.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
734
+ INFO:hf-to-gguf:blk.38.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
735
+ INFO:hf-to-gguf:blk.38.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
736
+ INFO:hf-to-gguf:blk.38.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
737
+ INFO:hf-to-gguf:output.weight, torch.bfloat16 --> BF16, shape = {2048, 248320}
738
+ INFO:hf-to-gguf:blk.39.attn_norm.weight, torch.bfloat16 --> F32, shape = {2048}
739
+ INFO:hf-to-gguf:blk.39.ffn_down_exps.weight, torch.bfloat16 --> BF16, shape = {512, 2048, 256}
740
+ INFO:hf-to-gguf:blk.39.ffn_gate_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
741
+ INFO:hf-to-gguf:blk.39.ffn_up_exps.weight, torch.bfloat16 --> BF16, shape = {2048, 512, 256}
742
+ INFO:hf-to-gguf:blk.39.ffn_gate_inp.weight, torch.bfloat16 --> F32, shape = {2048, 256}
743
+ INFO:hf-to-gguf:blk.39.ffn_down_shexp.weight, torch.bfloat16 --> BF16, shape = {512, 2048}
744
+ INFO:hf-to-gguf:blk.39.ffn_gate_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
745
+ INFO:hf-to-gguf:blk.39.ffn_up_shexp.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
746
+ INFO:hf-to-gguf:blk.39.ffn_gate_inp_shexp.weight, torch.bfloat16 --> F32, shape = {2048, 1}
747
+ INFO:hf-to-gguf:blk.39.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {2048}
748
+ INFO:hf-to-gguf:blk.39.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
749
+ INFO:hf-to-gguf:blk.39.attn_k.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
750
+ INFO:hf-to-gguf:blk.39.attn_output.weight, torch.bfloat16 --> BF16, shape = {4096, 2048}
751
+ INFO:hf-to-gguf:blk.39.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
752
+ INFO:hf-to-gguf:blk.39.attn_q.weight, torch.bfloat16 --> BF16, shape = {2048, 8192}
753
+ INFO:hf-to-gguf:blk.39.attn_v.weight, torch.bfloat16 --> BF16, shape = {2048, 512}
754
+ INFO:hf-to-gguf:output_norm.weight, torch.bfloat16 --> F32, shape = {2048}
755
+ INFO:hf-to-gguf:Set meta model
756
+ INFO:hf-to-gguf:Set model parameters
757
+ INFO:hf-to-gguf:gguf: context length = 262144
758
+ INFO:hf-to-gguf:gguf: embedding length = 2048
759
+ INFO:hf-to-gguf:gguf: head count = 16
760
+ INFO:hf-to-gguf:gguf: key-value head count = 2
761
+ WARNING:hf-to-gguf:Unknown RoPE type: default
762
+ INFO:hf-to-gguf:gguf: rope scaling type = NONE
763
+ INFO:hf-to-gguf:gguf: mrope sections: [11, 11, 10, 0]
764
+ INFO:hf-to-gguf:gguf: rope theta = 10000000
765
+ INFO:hf-to-gguf:gguf: rms norm epsilon = 1e-06
766
+ INFO:hf-to-gguf:gguf: expert count = 256
767
+ INFO:hf-to-gguf:gguf: experts used count = 8
768
+ INFO:hf-to-gguf:gguf: file type = 32
769
+ INFO:hf-to-gguf:gguf: expert feed forward length = 512
770
+ INFO:hf-to-gguf:gguf: expert shared feed forward length = 512
771
+ INFO:hf-to-gguf:Set model quantization version
772
+ INFO:hf-to-gguf:Set model tokenizer
773
+ INFO:gguf.vocab:Adding 247587 merge(s).
774
+ INFO:gguf.vocab:Setting special token type eos to 248046
775
+ INFO:gguf.vocab:Setting special token type pad to 248044
776
+ INFO:gguf.vocab:Setting chat_template to {%- set image_count = namespace(value=0) %}
777
+ {%- set video_count = namespace(value=0) %}
778
+ {%- macro render_content(content, do_vision_count, is_system_content=false) %}
779
+ {%- if content is string %}
780
+ {{- content }}
781
+ {%- elif content is iterable and content is not mapping %}
782
+ {%- for item in content %}
783
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
784
+ {%- if is_system_content %}
785
+ {{- raise_exception('System message cannot contain images.') }}
786
+ {%- endif %}
787
+ {%- if do_vision_count %}
788
+ {%- set image_count.value = image_count.value + 1 %}
789
+ {%- endif %}
790
+ {%- if add_vision_id %}
791
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
792
+ {%- endif %}
793
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
794
+ {%- elif 'video' in item or item.type == 'video' %}
795
+ {%- if is_system_content %}
796
+ {{- raise_exception('System message cannot contain videos.') }}
797
+ {%- endif %}
798
+ {%- if do_vision_count %}
799
+ {%- set video_count.value = video_count.value + 1 %}
800
+ {%- endif %}
801
+ {%- if add_vision_id %}
802
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
803
+ {%- endif %}
804
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
805
+ {%- elif 'text' in item %}
806
+ {{- item.text }}
807
+ {%- else %}
808
+ {{- raise_exception('Unexpected item type in content.') }}
809
+ {%- endif %}
810
+ {%- endfor %}
811
+ {%- elif content is none or content is undefined %}
812
+ {{- '' }}
813
+ {%- else %}
814
+ {{- raise_exception('Unexpected content type.') }}
815
+ {%- endif %}
816
+ {%- endmacro %}
817
+ {%- if not messages %}
818
+ {{- raise_exception('No messages provided.') }}
819
+ {%- endif %}
820
+ {%- if tools and tools is iterable and tools is not mapping %}
821
+ {{- '<|im_start|>system\n' }}
822
+ {{- "# Tools\n\nYou have access to the following functions:\n\n<tools>" }}
823
+ {%- for tool in tools %}
824
+ {{- "\n" }}
825
+ {{- tool | tojson }}
826
+ {%- endfor %}
827
+ {{- "\n</tools>" }}
828
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n<tool_call>\n<function=example_function_name>\n<parameter=example_parameter_1>\nvalue_1\n</parameter>\n<parameter=example_parameter_2>\nThis is the value for the second parameter\nthat can span\nmultiple lines\n</parameter>\n</function>\n</tool_call>\n\n<IMPORTANT>\nReminder:\n- Function calls MUST follow the specified format: an inner <function=...></function> block must be nested within <tool_call></tool_call> XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n</IMPORTANT>' }}
829
+ {%- if messages[0].role == 'system' %}
830
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
831
+ {%- if content %}
832
+ {{- '\n\n' + content }}
833
+ {%- endif %}
834
+ {%- endif %}
835
+ {{- '<|im_end|>\n' }}
836
+ {%- else %}
837
+ {%- if messages[0].role == 'system' %}
838
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
839
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
840
+ {%- endif %}
841
+ {%- endif %}
842
+ {%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
843
+ {%- for message in messages[::-1] %}
844
+ {%- set index = (messages|length - 1) - loop.index0 %}
845
+ {%- if ns.multi_step_tool and message.role == "user" %}
846
+ {%- set content = render_content(message.content, false)|trim %}
847
+ {%- if not(content.startswith('<tool_response>') and content.endswith('</tool_response>')) %}
848
+ {%- set ns.multi_step_tool = false %}
849
+ {%- set ns.last_query_index = index %}
850
+ {%- endif %}
851
+ {%- endif %}
852
+ {%- endfor %}
853
+ {%- if ns.multi_step_tool %}
854
+ {{- raise_exception('No user query found in messages.') }}
855
+ {%- endif %}
856
+ {%- for message in messages %}
857
+ {%- set content = render_content(message.content, true)|trim %}
858
+ {%- if message.role == "system" %}
859
+ {%- if not loop.first %}
860
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
861
+ {%- endif %}
862
+ {%- elif message.role == "user" %}
863
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
864
+ {%- elif message.role == "assistant" %}
865
+ {%- set reasoning_content = '' %}
866
+ {%- if message.reasoning_content is string %}
867
+ {%- set reasoning_content = message.reasoning_content %}
868
+ {%- else %}
869
+ {%- if '</think>' in content %}
870
+ {%- set reasoning_content = content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
871
+ {%- set content = content.split('</think>')[-1].lstrip('\n') %}
872
+ {%- endif %}
873
+ {%- endif %}
874
+ {%- set reasoning_content = reasoning_content|trim %}
875
+ {{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content + '\n</think>\n\n' + content }}
876
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
877
+ {%- for tool_call in message.tool_calls %}
878
+ {%- if tool_call.function is defined %}
879
+ {%- set tool_call = tool_call.function %}
880
+ {%- endif %}
881
+ {%- if loop.first %}
882
+ {%- if content|trim %}
883
+ {{- '\n\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
884
+ {%- else %}
885
+ {{- '<tool_call>\n<function=' + tool_call.name + '>\n' }}
886
+ {%- endif %}
887
+ {%- else %}
888
+ {{- '\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
889
+ {%- endif %}
890
+ {%- if tool_call.arguments is defined %}
891
+ {%- for args_name, args_value in tool_call.arguments|items %}
892
+ {{- '<parameter=' + args_name + '>\n' }}
893
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
894
+ {{- args_value }}
895
+ {{- '\n</parameter>\n' }}
896
+ {%- endfor %}
897
+ {%- endif %}
898
+ {{- '</function>\n</tool_call>' }}
899
+ {%- endfor %}
900
+ {%- endif %}
901
+ {{- '<|im_end|>\n' }}
902
+ {%- elif message.role == "tool" %}
903
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
904
+ {{- '<|im_start|>user' }}
905
+ {%- endif %}
906
+ {{- '\n<tool_response>\n' }}
907
+ {{- content }}
908
+ {{- '\n</tool_response>' }}
909
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
910
+ {{- '<|im_end|>\n' }}
911
+ {%- elif loop.last %}
912
+ {{- '<|im_end|>\n' }}
913
+ {%- endif %}
914
+ {%- else %}
915
+ {{- raise_exception('Unexpected message role.') }}
916
+ {%- endif %}
917
+ {%- endfor %}
918
+ {%- if add_generation_prompt %}
919
+ {{- '<|im_start|>assistant\n' }}
920
+ {%- if reasoning_effort is not defined or reasoning_effort is none %}
921
+ {{- '<think>' }}
922
+ {%- elif reasoning_effort == 'none' %}
923
+ {{- '<think>\n\n</think>\n\n' }}
924
+ {%- elif reasoning_effort == 'high' %}
925
+ {{- '<think>\n' }}
926
+ {%- else %}
927
+ {{- '<think>' }}
928
+ {%- endif %}
929
+ {%- endif %}
930
+
931
+ INFO:gguf.gguf_writer:Writing the following files:
932
+ INFO:gguf.gguf_writer:gguf/Nex-N2.5-mini-BF16.gguf: n_tensors = 733, total_size = 69.4G
933
+
934
+ INFO:hf-to-gguf:Model successfully exported to gguf/Nex-N2.5-mini-BF16.gguf
recipe/logs/C2_mmproj.log ADDED
@@ -0,0 +1,362 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ INFO:hf-to-gguf:Loading model: hf
2
+ INFO:hf-to-gguf:Model architecture: Qwen3_5MoeForConditionalGeneration
3
+ INFO:hf-to-gguf:gguf: loading model weight map from 'model.safetensors.index.json'
4
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00001-of-00016.safetensors'
5
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00002-of-00016.safetensors'
6
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00003-of-00016.safetensors'
7
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00004-of-00016.safetensors'
8
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00005-of-00016.safetensors'
9
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00006-of-00016.safetensors'
10
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00007-of-00016.safetensors'
11
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00008-of-00016.safetensors'
12
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00009-of-00016.safetensors'
13
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00010-of-00016.safetensors'
14
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00011-of-00016.safetensors'
15
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00012-of-00016.safetensors'
16
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00013-of-00016.safetensors'
17
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00014-of-00016.safetensors'
18
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00015-of-00016.safetensors'
19
+ INFO:hf-to-gguf:gguf: indexing model part 'model-00016-of-00016.safetensors'
20
+ INFO:gguf.gguf_writer:gguf: This GGUF file is for Little Endian only
21
+ INFO:hf-to-gguf:Exporting model...
22
+ INFO:hf-to-gguf:v.blk.0.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
23
+ INFO:hf-to-gguf:v.blk.0.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
24
+ INFO:hf-to-gguf:v.blk.0.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
25
+ INFO:hf-to-gguf:v.blk.0.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
26
+ INFO:hf-to-gguf:v.blk.0.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
27
+ INFO:hf-to-gguf:v.blk.0.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
28
+ INFO:hf-to-gguf:v.blk.0.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
29
+ INFO:hf-to-gguf:v.blk.0.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
30
+ INFO:hf-to-gguf:v.blk.0.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
31
+ INFO:hf-to-gguf:v.blk.0.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
32
+ INFO:hf-to-gguf:v.blk.0.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
33
+ INFO:hf-to-gguf:v.blk.0.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
34
+ INFO:hf-to-gguf:v.blk.1.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
35
+ INFO:hf-to-gguf:v.blk.1.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
36
+ INFO:hf-to-gguf:v.blk.1.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
37
+ INFO:hf-to-gguf:v.blk.1.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
38
+ INFO:hf-to-gguf:v.blk.1.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
39
+ INFO:hf-to-gguf:v.blk.1.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
40
+ INFO:hf-to-gguf:v.blk.1.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
41
+ INFO:hf-to-gguf:v.blk.1.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
42
+ INFO:hf-to-gguf:v.blk.1.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
43
+ INFO:hf-to-gguf:v.blk.1.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
44
+ INFO:hf-to-gguf:v.blk.1.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
45
+ INFO:hf-to-gguf:v.blk.1.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
46
+ INFO:hf-to-gguf:v.blk.10.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
47
+ INFO:hf-to-gguf:v.blk.10.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
48
+ INFO:hf-to-gguf:v.blk.10.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
49
+ INFO:hf-to-gguf:v.blk.10.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
50
+ INFO:hf-to-gguf:v.blk.10.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
51
+ INFO:hf-to-gguf:v.blk.10.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
52
+ INFO:hf-to-gguf:v.blk.10.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
53
+ INFO:hf-to-gguf:v.blk.10.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
54
+ INFO:hf-to-gguf:v.blk.10.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
55
+ INFO:hf-to-gguf:v.blk.10.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
56
+ INFO:hf-to-gguf:v.blk.10.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
57
+ INFO:hf-to-gguf:v.blk.10.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
58
+ INFO:hf-to-gguf:v.blk.11.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
59
+ INFO:hf-to-gguf:v.blk.11.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
60
+ INFO:hf-to-gguf:v.blk.11.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
61
+ INFO:hf-to-gguf:v.blk.11.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
62
+ INFO:hf-to-gguf:v.blk.11.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
63
+ INFO:hf-to-gguf:v.blk.11.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
64
+ INFO:hf-to-gguf:v.blk.11.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
65
+ INFO:hf-to-gguf:v.blk.11.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
66
+ INFO:hf-to-gguf:v.blk.11.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
67
+ INFO:hf-to-gguf:v.blk.11.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
68
+ INFO:hf-to-gguf:v.blk.11.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
69
+ INFO:hf-to-gguf:v.blk.11.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
70
+ INFO:hf-to-gguf:v.blk.12.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
71
+ INFO:hf-to-gguf:v.blk.12.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
72
+ INFO:hf-to-gguf:v.blk.12.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
73
+ INFO:hf-to-gguf:v.blk.12.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
74
+ INFO:hf-to-gguf:v.blk.12.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
75
+ INFO:hf-to-gguf:v.blk.12.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
76
+ INFO:hf-to-gguf:v.blk.12.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
77
+ INFO:hf-to-gguf:v.blk.12.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
78
+ INFO:hf-to-gguf:v.blk.12.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
79
+ INFO:hf-to-gguf:v.blk.12.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
80
+ INFO:hf-to-gguf:v.blk.12.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
81
+ INFO:hf-to-gguf:v.blk.12.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
82
+ INFO:hf-to-gguf:v.blk.13.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
83
+ INFO:hf-to-gguf:v.blk.13.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
84
+ INFO:hf-to-gguf:v.blk.13.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
85
+ INFO:hf-to-gguf:v.blk.13.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
86
+ INFO:hf-to-gguf:v.blk.13.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
87
+ INFO:hf-to-gguf:v.blk.13.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
88
+ INFO:hf-to-gguf:v.blk.13.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
89
+ INFO:hf-to-gguf:v.blk.13.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
90
+ INFO:hf-to-gguf:v.blk.13.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
91
+ INFO:hf-to-gguf:v.blk.13.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
92
+ INFO:hf-to-gguf:v.blk.13.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
93
+ INFO:hf-to-gguf:v.blk.13.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
94
+ INFO:hf-to-gguf:v.blk.14.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
95
+ INFO:hf-to-gguf:v.blk.14.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
96
+ INFO:hf-to-gguf:v.blk.14.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
97
+ INFO:hf-to-gguf:v.blk.14.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
98
+ INFO:hf-to-gguf:v.blk.14.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
99
+ INFO:hf-to-gguf:v.blk.14.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
100
+ INFO:hf-to-gguf:v.blk.14.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
101
+ INFO:hf-to-gguf:v.blk.14.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
102
+ INFO:hf-to-gguf:v.blk.14.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
103
+ INFO:hf-to-gguf:v.blk.14.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
104
+ INFO:hf-to-gguf:v.blk.14.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
105
+ INFO:hf-to-gguf:v.blk.14.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
106
+ INFO:hf-to-gguf:v.blk.15.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
107
+ INFO:hf-to-gguf:v.blk.15.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
108
+ INFO:hf-to-gguf:v.blk.15.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
109
+ INFO:hf-to-gguf:v.blk.15.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
110
+ INFO:hf-to-gguf:v.blk.15.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
111
+ INFO:hf-to-gguf:v.blk.15.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
112
+ INFO:hf-to-gguf:v.blk.15.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
113
+ INFO:hf-to-gguf:v.blk.15.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
114
+ INFO:hf-to-gguf:v.blk.15.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
115
+ INFO:hf-to-gguf:v.blk.15.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
116
+ INFO:hf-to-gguf:v.blk.15.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
117
+ INFO:hf-to-gguf:v.blk.15.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
118
+ INFO:hf-to-gguf:v.blk.16.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
119
+ INFO:hf-to-gguf:v.blk.16.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
120
+ INFO:hf-to-gguf:v.blk.16.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
121
+ INFO:hf-to-gguf:v.blk.16.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
122
+ INFO:hf-to-gguf:v.blk.16.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
123
+ INFO:hf-to-gguf:v.blk.16.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
124
+ INFO:hf-to-gguf:v.blk.16.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
125
+ INFO:hf-to-gguf:v.blk.16.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
126
+ INFO:hf-to-gguf:v.blk.16.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
127
+ INFO:hf-to-gguf:v.blk.16.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
128
+ INFO:hf-to-gguf:v.blk.16.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
129
+ INFO:hf-to-gguf:v.blk.16.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
130
+ INFO:hf-to-gguf:v.blk.17.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
131
+ INFO:hf-to-gguf:v.blk.17.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
132
+ INFO:hf-to-gguf:v.blk.17.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
133
+ INFO:hf-to-gguf:v.blk.17.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
134
+ INFO:hf-to-gguf:v.blk.17.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
135
+ INFO:hf-to-gguf:v.blk.17.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
136
+ INFO:hf-to-gguf:v.blk.17.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
137
+ INFO:hf-to-gguf:v.blk.17.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
138
+ INFO:hf-to-gguf:v.blk.17.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
139
+ INFO:hf-to-gguf:v.blk.17.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
140
+ INFO:hf-to-gguf:v.blk.17.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
141
+ INFO:hf-to-gguf:v.blk.17.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
142
+ INFO:hf-to-gguf:v.blk.18.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
143
+ INFO:hf-to-gguf:v.blk.18.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
144
+ INFO:hf-to-gguf:v.blk.18.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
145
+ INFO:hf-to-gguf:v.blk.18.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
146
+ INFO:hf-to-gguf:v.blk.18.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
147
+ INFO:hf-to-gguf:v.blk.18.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
148
+ INFO:hf-to-gguf:v.blk.18.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
149
+ INFO:hf-to-gguf:v.blk.18.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
150
+ INFO:hf-to-gguf:v.blk.18.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
151
+ INFO:hf-to-gguf:v.blk.18.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
152
+ INFO:hf-to-gguf:v.blk.18.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
153
+ INFO:hf-to-gguf:v.blk.18.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
154
+ INFO:hf-to-gguf:v.blk.19.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
155
+ INFO:hf-to-gguf:v.blk.19.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
156
+ INFO:hf-to-gguf:v.blk.19.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
157
+ INFO:hf-to-gguf:v.blk.19.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
158
+ INFO:hf-to-gguf:v.blk.19.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
159
+ INFO:hf-to-gguf:v.blk.19.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
160
+ INFO:hf-to-gguf:v.blk.19.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
161
+ INFO:hf-to-gguf:v.blk.19.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
162
+ INFO:hf-to-gguf:v.blk.19.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
163
+ INFO:hf-to-gguf:v.blk.19.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
164
+ INFO:hf-to-gguf:v.blk.19.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
165
+ INFO:hf-to-gguf:v.blk.19.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
166
+ INFO:hf-to-gguf:v.blk.2.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
167
+ INFO:hf-to-gguf:v.blk.2.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
168
+ INFO:hf-to-gguf:v.blk.2.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
169
+ INFO:hf-to-gguf:v.blk.2.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
170
+ INFO:hf-to-gguf:v.blk.2.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
171
+ INFO:hf-to-gguf:v.blk.2.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
172
+ INFO:hf-to-gguf:v.blk.2.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
173
+ INFO:hf-to-gguf:v.blk.2.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
174
+ INFO:hf-to-gguf:v.blk.2.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
175
+ INFO:hf-to-gguf:v.blk.2.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
176
+ INFO:hf-to-gguf:v.blk.2.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
177
+ INFO:hf-to-gguf:v.blk.2.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
178
+ INFO:hf-to-gguf:v.blk.20.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
179
+ INFO:hf-to-gguf:v.blk.20.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
180
+ INFO:hf-to-gguf:v.blk.20.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
181
+ INFO:hf-to-gguf:v.blk.20.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
182
+ INFO:hf-to-gguf:v.blk.20.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
183
+ INFO:hf-to-gguf:v.blk.20.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
184
+ INFO:hf-to-gguf:v.blk.20.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
185
+ INFO:hf-to-gguf:v.blk.20.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
186
+ INFO:hf-to-gguf:v.blk.20.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
187
+ INFO:hf-to-gguf:v.blk.20.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
188
+ INFO:hf-to-gguf:v.blk.20.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
189
+ INFO:hf-to-gguf:v.blk.20.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
190
+ INFO:hf-to-gguf:v.blk.21.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
191
+ INFO:hf-to-gguf:v.blk.21.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
192
+ INFO:hf-to-gguf:v.blk.21.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
193
+ INFO:hf-to-gguf:v.blk.21.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
194
+ INFO:hf-to-gguf:v.blk.21.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
195
+ INFO:hf-to-gguf:v.blk.21.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
196
+ INFO:hf-to-gguf:v.blk.21.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
197
+ INFO:hf-to-gguf:v.blk.21.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
198
+ INFO:hf-to-gguf:v.blk.21.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
199
+ INFO:hf-to-gguf:v.blk.21.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
200
+ INFO:hf-to-gguf:v.blk.21.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
201
+ INFO:hf-to-gguf:v.blk.21.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
202
+ INFO:hf-to-gguf:v.blk.22.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
203
+ INFO:hf-to-gguf:v.blk.22.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
204
+ INFO:hf-to-gguf:v.blk.22.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
205
+ INFO:hf-to-gguf:v.blk.22.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
206
+ INFO:hf-to-gguf:v.blk.22.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
207
+ INFO:hf-to-gguf:v.blk.22.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
208
+ INFO:hf-to-gguf:v.blk.22.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
209
+ INFO:hf-to-gguf:v.blk.22.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
210
+ INFO:hf-to-gguf:v.blk.22.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
211
+ INFO:hf-to-gguf:v.blk.22.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
212
+ INFO:hf-to-gguf:v.blk.22.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
213
+ INFO:hf-to-gguf:v.blk.22.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
214
+ INFO:hf-to-gguf:v.blk.23.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
215
+ INFO:hf-to-gguf:v.blk.23.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
216
+ INFO:hf-to-gguf:v.blk.23.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
217
+ INFO:hf-to-gguf:v.blk.23.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
218
+ INFO:hf-to-gguf:v.blk.23.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
219
+ INFO:hf-to-gguf:v.blk.23.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
220
+ INFO:hf-to-gguf:v.blk.23.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
221
+ INFO:hf-to-gguf:v.blk.23.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
222
+ INFO:hf-to-gguf:v.blk.23.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
223
+ INFO:hf-to-gguf:v.blk.23.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
224
+ INFO:hf-to-gguf:v.blk.23.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
225
+ INFO:hf-to-gguf:v.blk.23.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
226
+ INFO:hf-to-gguf:v.blk.24.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
227
+ INFO:hf-to-gguf:v.blk.24.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
228
+ INFO:hf-to-gguf:v.blk.24.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
229
+ INFO:hf-to-gguf:v.blk.24.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
230
+ INFO:hf-to-gguf:v.blk.24.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
231
+ INFO:hf-to-gguf:v.blk.24.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
232
+ INFO:hf-to-gguf:v.blk.24.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
233
+ INFO:hf-to-gguf:v.blk.24.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
234
+ INFO:hf-to-gguf:v.blk.24.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
235
+ INFO:hf-to-gguf:v.blk.24.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
236
+ INFO:hf-to-gguf:v.blk.24.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
237
+ INFO:hf-to-gguf:v.blk.24.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
238
+ INFO:hf-to-gguf:v.blk.25.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
239
+ INFO:hf-to-gguf:v.blk.25.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
240
+ INFO:hf-to-gguf:v.blk.25.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
241
+ INFO:hf-to-gguf:v.blk.25.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
242
+ INFO:hf-to-gguf:v.blk.25.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
243
+ INFO:hf-to-gguf:v.blk.25.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
244
+ INFO:hf-to-gguf:v.blk.25.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
245
+ INFO:hf-to-gguf:v.blk.25.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
246
+ INFO:hf-to-gguf:v.blk.25.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
247
+ INFO:hf-to-gguf:v.blk.25.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
248
+ INFO:hf-to-gguf:v.blk.25.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
249
+ INFO:hf-to-gguf:v.blk.25.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
250
+ INFO:hf-to-gguf:v.blk.26.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
251
+ INFO:hf-to-gguf:v.blk.26.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
252
+ INFO:hf-to-gguf:v.blk.26.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
253
+ INFO:hf-to-gguf:v.blk.26.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
254
+ INFO:hf-to-gguf:v.blk.26.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
255
+ INFO:hf-to-gguf:v.blk.26.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
256
+ INFO:hf-to-gguf:v.blk.26.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
257
+ INFO:hf-to-gguf:v.blk.26.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
258
+ INFO:hf-to-gguf:v.blk.26.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
259
+ INFO:hf-to-gguf:v.blk.26.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
260
+ INFO:hf-to-gguf:v.blk.26.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
261
+ INFO:hf-to-gguf:v.blk.26.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
262
+ INFO:hf-to-gguf:v.blk.3.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
263
+ INFO:hf-to-gguf:v.blk.3.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
264
+ INFO:hf-to-gguf:v.blk.3.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
265
+ INFO:hf-to-gguf:v.blk.3.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
266
+ INFO:hf-to-gguf:v.blk.3.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
267
+ INFO:hf-to-gguf:v.blk.3.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
268
+ INFO:hf-to-gguf:v.blk.3.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
269
+ INFO:hf-to-gguf:v.blk.3.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
270
+ INFO:hf-to-gguf:v.blk.3.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
271
+ INFO:hf-to-gguf:v.blk.3.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
272
+ INFO:hf-to-gguf:v.blk.3.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
273
+ INFO:hf-to-gguf:v.blk.3.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
274
+ INFO:hf-to-gguf:v.blk.4.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
275
+ INFO:hf-to-gguf:v.blk.4.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
276
+ INFO:hf-to-gguf:v.blk.4.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
277
+ INFO:hf-to-gguf:v.blk.4.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
278
+ INFO:hf-to-gguf:v.blk.4.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
279
+ INFO:hf-to-gguf:v.blk.4.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
280
+ INFO:hf-to-gguf:v.blk.4.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
281
+ INFO:hf-to-gguf:v.blk.4.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
282
+ INFO:hf-to-gguf:v.blk.4.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
283
+ INFO:hf-to-gguf:v.blk.4.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
284
+ INFO:hf-to-gguf:v.blk.4.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
285
+ INFO:hf-to-gguf:v.blk.4.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
286
+ INFO:hf-to-gguf:v.blk.5.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
287
+ INFO:hf-to-gguf:v.blk.5.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
288
+ INFO:hf-to-gguf:v.blk.5.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
289
+ INFO:hf-to-gguf:v.blk.5.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
290
+ INFO:hf-to-gguf:v.blk.5.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
291
+ INFO:hf-to-gguf:v.blk.5.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
292
+ INFO:hf-to-gguf:v.blk.5.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
293
+ INFO:hf-to-gguf:v.blk.5.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
294
+ INFO:hf-to-gguf:v.blk.5.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
295
+ INFO:hf-to-gguf:v.blk.5.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
296
+ INFO:hf-to-gguf:v.blk.5.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
297
+ INFO:hf-to-gguf:v.blk.5.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
298
+ INFO:hf-to-gguf:v.blk.6.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
299
+ INFO:hf-to-gguf:v.blk.6.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
300
+ INFO:hf-to-gguf:v.blk.6.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
301
+ INFO:hf-to-gguf:v.blk.6.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
302
+ INFO:hf-to-gguf:v.blk.6.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
303
+ INFO:hf-to-gguf:v.blk.6.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
304
+ INFO:hf-to-gguf:v.blk.6.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
305
+ INFO:hf-to-gguf:v.blk.6.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
306
+ INFO:hf-to-gguf:v.blk.6.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
307
+ INFO:hf-to-gguf:v.blk.6.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
308
+ INFO:hf-to-gguf:v.blk.6.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
309
+ INFO:hf-to-gguf:v.blk.6.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
310
+ INFO:hf-to-gguf:v.blk.7.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
311
+ INFO:hf-to-gguf:v.blk.7.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
312
+ INFO:hf-to-gguf:v.blk.7.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
313
+ INFO:hf-to-gguf:v.blk.7.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
314
+ INFO:hf-to-gguf:v.blk.7.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
315
+ INFO:hf-to-gguf:v.blk.7.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
316
+ INFO:hf-to-gguf:v.blk.7.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
317
+ INFO:hf-to-gguf:v.blk.7.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
318
+ INFO:hf-to-gguf:v.blk.7.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
319
+ INFO:hf-to-gguf:v.blk.7.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
320
+ INFO:hf-to-gguf:v.blk.7.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
321
+ INFO:hf-to-gguf:v.blk.7.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
322
+ INFO:hf-to-gguf:v.blk.8.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
323
+ INFO:hf-to-gguf:v.blk.8.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
324
+ INFO:hf-to-gguf:v.blk.8.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
325
+ INFO:hf-to-gguf:v.blk.8.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
326
+ INFO:hf-to-gguf:v.blk.8.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
327
+ INFO:hf-to-gguf:v.blk.8.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
328
+ INFO:hf-to-gguf:v.blk.8.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
329
+ INFO:hf-to-gguf:v.blk.8.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
330
+ INFO:hf-to-gguf:v.blk.8.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
331
+ INFO:hf-to-gguf:v.blk.8.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
332
+ INFO:hf-to-gguf:v.blk.8.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
333
+ INFO:hf-to-gguf:v.blk.8.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
334
+ INFO:hf-to-gguf:v.blk.9.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
335
+ INFO:hf-to-gguf:v.blk.9.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
336
+ INFO:hf-to-gguf:v.blk.9.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
337
+ INFO:hf-to-gguf:v.blk.9.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
338
+ INFO:hf-to-gguf:v.blk.9.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
339
+ INFO:hf-to-gguf:v.blk.9.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
340
+ INFO:hf-to-gguf:v.blk.9.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
341
+ INFO:hf-to-gguf:v.blk.9.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
342
+ INFO:hf-to-gguf:v.blk.9.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
343
+ INFO:hf-to-gguf:v.blk.9.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
344
+ INFO:hf-to-gguf:v.blk.9.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
345
+ INFO:hf-to-gguf:v.blk.9.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
346
+ INFO:hf-to-gguf:mm.0.bias, torch.bfloat16 --> F32, shape = {4608}
347
+ INFO:hf-to-gguf:mm.0.weight, torch.bfloat16 --> BF16, shape = {4608, 4608}
348
+ INFO:hf-to-gguf:mm.2.bias, torch.bfloat16 --> F32, shape = {2048}
349
+ INFO:hf-to-gguf:mm.2.weight, torch.bfloat16 --> BF16, shape = {4608, 2048}
350
+ INFO:hf-to-gguf:v.post_ln.bias, torch.bfloat16 --> F32, shape = {1152}
351
+ INFO:hf-to-gguf:v.post_ln.weight, torch.bfloat16 --> F32, shape = {1152}
352
+ INFO:hf-to-gguf:v.patch_embd.bias, torch.bfloat16 --> F32, shape = {1152}
353
+ INFO:hf-to-gguf:v.patch_embd.weight, torch.bfloat16 --> F32, shape = {16, 16, 3, 1152}
354
+ INFO:hf-to-gguf:v.patch_embd.weight.1, torch.bfloat16 --> F32, shape = {16, 16, 3, 1152}
355
+ INFO:hf-to-gguf:v.position_embd.weight, torch.bfloat16 --> F32, shape = {1152, 2304}
356
+ INFO:hf-to-gguf:Set meta model
357
+ INFO:hf-to-gguf:Set model parameters
358
+ INFO:hf-to-gguf:Set model quantization version
359
+ INFO:gguf.gguf_writer:Writing the following files:
360
+ INFO:gguf.gguf_writer:out/mmproj-Nex-N2.5-mini-BF16.gguf: n_tensors = 334, total_size = 902.8M
361
+
362
+ INFO:hf-to-gguf:Model successfully exported to out/mmproj-Nex-N2.5-mini-BF16.gguf
recipe/logs/C_readback.log ADDED
@@ -0,0 +1,2 @@
 
 
 
1
+ PASS Nex-N2.5-mini-BF16.gguf arch=qwen35moe ftype=32 tensors=733 nextn=0 output.weight=BF16 token_embd.weight=BF16
2
+ PASS mmproj-Nex-N2.5-mini-BF16.gguf arch=clip ftype=32 tensors=334 nextn=0 output.weight=None token_embd.weight=None
recipe/logs/D2_verify_download.log ADDED
@@ -0,0 +1,2 @@
 
 
 
1
+ files=28 sha-verified=19 size-only=9 bad=0
2
+ RESULT: PASS
recipe/logs/N1_ppl_bf16.log ADDED
@@ -0,0 +1,15 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0.00.044.003 I common_init_result: fitting params to device memory ...
2
+ 0.00.044.006 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
3
+ 0.00.378.699 W llama_model_loader: direct I/O is enabled, disabling mmap
4
+ 1.17.123.477 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
5
+ 1.17.377.358 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
6
+ 1.18.283.894 I
7
+ 1.18.284.021 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
8
+ 1.18.284.129 I perplexity: saving all logits to kld/bf16.kld
9
+ 1.18.284.134 I perplexity: tokenizing the input ..
10
+ 1.18.617.280 I perplexity: tokenization took 333.135 ms
11
+ 1.18.617.370 I perplexity: calculating perplexity over 40 chunks, n_ctx=2048, batch_size=2048, n_seq=1
12
+ 1.23.386.958 I perplexity: 4.58 seconds per pass - ETA 3.05 minutes
13
+ [1]139.1466,[2]166.0748,[3]174.5814,[4]170.5873,[5]171.7312,[6]127.5767,[7]99.6563,[8]94.8236,[9]101.8520,[10]104.1594,[11]108.4984,[12]113.4569,[13]109.8859,[14]107.0003,[15]104.7124,[16]107.6891,[17]105.9352,[18]106.7826,[19]104.4693,[20]101.7712,[21]102.7759,[22]104.9236,[23]107.8566,[24]107.9859,[25]109.0207,[26]109.5885,[27]111.5203,[28]113.7558,[29]116.8379,[30]117.2211,[31]115.3159,[32]113.7423,[33]112.7430,[34]112.9143,[35]113.3873,[36]113.5581,[37]111.6761,[38]111.1729,[39]108.9264,[40]105.9103,
14
+ 4.26.432.497 I Final estimate: PPL = 105.9103 +/- 1.77389
15
+
recipe/logs/N1c_ppl_bf16_cpu.log ADDED
@@ -0,0 +1,15 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0.00.040.030 I common_init_result: fitting params to device memory ...
2
+ 0.00.040.033 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
3
+ 0.00.314.545 I common_params_fit_impl: projected to use 66756 MiB of host memory vs. 127438 MiB of total host memory
4
+ 0.00.652.877 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
5
+ 0.00.672.925 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
6
+ 0.01.325.340 I
7
+ 0.01.325.484 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
8
+ 0.01.326.996 I perplexity: saving all logits to kld/bf16.kld
9
+ 0.01.327.002 I perplexity: tokenizing the input ..
10
+ 0.01.631.676 I perplexity: tokenization took 304.664 ms
11
+ 0.01.631.766 I perplexity: calculating perplexity over 40 chunks, n_ctx=2048, batch_size=2048, n_seq=1
12
+ 0.13.030.135 I perplexity: 11.26 seconds per pass - ETA 7.50 minutes
13
+ [1]5.6964,[2]6.6666,[3]7.0328,[4]7.2903,[5]7.1445,[6]6.1493,[7]5.7412,[8]5.6706,[9]5.9764,[10]6.0886,[11]6.1422,[12]6.3827,[13]6.4266,[14]6.4831,[15]6.5233,[16]6.6999,[17]6.7466,[18]6.8361,[19]6.7889,[20]6.5275,[21]6.5437,[22]6.5571,[23]6.6058,[24]6.6060,[25]6.6371,[26]6.6074,[27]6.7700,[28]6.8576,[29]6.8518,[30]6.7971,[31]6.6991,[32]6.6042,[33]6.5393,[34]6.5251,[35]6.5369,[36]6.5530,[37]6.4632,[38]6.3955,[39]6.3171,[40]6.2303,
14
+ 7.32.185.106 I Final estimate: PPL = 6.2303 +/- 0.07538
15
+
recipe/logs/N2c_imatrix_cpu.log ADDED
@@ -0,0 +1,114 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0.00.047.060 I common_init_result: fitting params to device memory ...
2
+ 0.00.047.066 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
3
+ 0.00.538.240 I common_params_fit_impl: projected to use 66727 MiB of host memory vs. 127438 MiB of total host memory
4
+ 0.01.009.512 W llama_context: n_ctx_seq (512) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
5
+ 0.01.026.428 I
6
+ 0.01.026.542 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
7
+ 0.01.026.548 I compute_imatrix: tokenizing the input ..
8
+ 0.01.115.784 I compute_imatrix: tokenization took 89.228 ms
9
+ 0.01.115.820 I compute_imatrix: computing over 129 chunks, n_ctx=512, batch_size=512, n_seq=1
10
+ 0.05.010.249 I compute_imatrix: 3.89 seconds per pass - ETA 8.37 minutes
11
+ [1]4.8280,[2]3.5310,[3]3.3565,[4]3.5543,[5]3.5306,[6]3.2674,[7]3.7535,[8]3.7744,[9]4.1909,0.35.054.817 W
12
+ save_imatrix: saving imatrix using GGUF format with a different suffix than .gguf
13
+ 0.35.054.819 W save_imatrix: if you want the previous imatrix format, use --output-format dat
14
+ 0.35.054.834 I
15
+ 0.35.054.839 W save_imatrix: entry ' blk.36.ffn_up_exps.weight' has partial data (99.61%)
16
+ 0.35.054.845 W save_imatrix: entry ' blk.34.ffn_down_exps.weight' has partial data (98.83%)
17
+ 0.35.054.846 W save_imatrix: entry ' blk.34.ffn_up_exps.weight' has partial data (98.83%)
18
+ 0.35.054.849 W save_imatrix: entry ' blk.33.ffn_gate_exps.weight' has partial data (99.61%)
19
+ 0.35.054.852 W save_imatrix: entry ' blk.32.ffn_down_exps.weight' has partial data (99.61%)
20
+ 0.35.054.856 W save_imatrix: entry ' blk.31.ffn_down_exps.weight' has partial data (99.61%)
21
+ 0.35.054.857 W save_imatrix: entry ' blk.31.ffn_gate_exps.weight' has partial data (99.61%)
22
+ 0.35.054.860 W save_imatrix: entry ' blk.30.ffn_down_exps.weight' has partial data (99.61%)
23
+ 0.35.054.864 W save_imatrix: entry ' blk.29.ffn_up_exps.weight' has partial data (98.83%)
24
+ 0.35.054.865 W save_imatrix: entry ' blk.29.ffn_gate_exps.weight' has partial data (98.83%)
25
+ 0.35.054.874 W save_imatrix: entry ' blk.27.ffn_gate_exps.weight' has partial data (99.61%)
26
+ 0.35.054.877 W save_imatrix: entry ' blk.26.ffn_up_exps.weight' has partial data (99.61%)
27
+ 0.35.054.883 W save_imatrix: entry ' blk.24.ffn_gate_exps.weight' has partial data (98.83%)
28
+ 0.35.054.893 W save_imatrix: entry ' blk.32.ffn_gate_exps.weight' has partial data (99.61%)
29
+ 0.35.054.895 W save_imatrix: entry ' blk.33.ffn_down_exps.weight' has partial data (99.61%)
30
+ 0.35.054.897 W save_imatrix: entry ' blk.30.ffn_gate_exps.weight' has partial data (99.61%)
31
+ 0.35.054.903 W save_imatrix: entry ' blk.26.ffn_gate_exps.weight' has partial data (99.61%)
32
+ 0.35.054.907 W save_imatrix: entry ' blk.20.ffn_down_exps.weight' has partial data (98.83%)
33
+ 0.35.054.910 W save_imatrix: entry ' blk.22.ffn_down_exps.weight' has partial data (99.22%)
34
+ 0.35.054.914 W save_imatrix: entry ' blk.27.ffn_down_exps.weight' has partial data (99.61%)
35
+ 0.35.054.916 W save_imatrix: entry ' blk.26.ffn_down_exps.weight' has partial data (99.61%)
36
+ 0.35.054.919 W save_imatrix: entry ' blk.20.ffn_gate_exps.weight' has partial data (98.83%)
37
+ 0.35.054.926 W save_imatrix: entry ' blk.22.ffn_up_exps.weight' has partial data (99.22%)
38
+ 0.35.054.928 W save_imatrix: entry ' blk.7.ffn_down_exps.weight' has partial data (99.22%)
39
+ 0.35.054.929 W save_imatrix: entry ' blk.36.ffn_gate_exps.weight' has partial data (99.61%)
40
+ 0.35.054.933 W save_imatrix: entry ' blk.31.ffn_up_exps.weight' has partial data (99.61%)
41
+ 0.35.054.936 W save_imatrix: entry ' blk.20.ffn_up_exps.weight' has partial data (98.83%)
42
+ 0.35.054.949 W save_imatrix: entry ' blk.15.ffn_down_exps.weight' has partial data (99.22%)
43
+ 0.35.054.968 W save_imatrix: entry ' blk.34.ffn_gate_exps.weight' has partial data (98.83%)
44
+ 0.35.054.969 W save_imatrix: entry ' blk.7.ffn_gate_exps.weight' has partial data (99.22%)
45
+ 0.35.054.971 W save_imatrix: entry ' blk.33.ffn_up_exps.weight' has partial data (99.61%)
46
+ 0.35.054.974 W save_imatrix: entry ' blk.24.ffn_down_exps.weight' has partial data (98.83%)
47
+ 0.35.054.983 W save_imatrix: entry ' blk.32.ffn_up_exps.weight' has partial data (99.61%)
48
+ 0.35.054.989 W save_imatrix: entry ' blk.18.ffn_down_exps.weight' has partial data (98.83%)
49
+ 0.35.054.997 W save_imatrix: entry ' blk.29.ffn_down_exps.weight' has partial data (98.83%)
50
+ 0.35.055.001 W save_imatrix: entry ' blk.7.ffn_up_exps.weight' has partial data (99.22%)
51
+ 0.35.055.004 W save_imatrix: entry ' blk.18.ffn_up_exps.weight' has partial data (98.83%)
52
+ 0.35.055.006 W save_imatrix: entry ' blk.27.ffn_up_exps.weight' has partial data (99.61%)
53
+ 0.35.055.013 W save_imatrix: entry ' blk.16.ffn_gate_exps.weight' has partial data (98.83%)
54
+ 0.35.055.016 W save_imatrix: entry ' blk.30.ffn_up_exps.weight' has partial data (99.61%)
55
+ 0.35.055.019 W save_imatrix: entry ' blk.36.ffn_down_exps.weight' has partial data (99.61%)
56
+ 0.35.055.022 W save_imatrix: entry ' blk.16.ffn_up_exps.weight' has partial data (98.83%)
57
+ 0.35.055.026 W save_imatrix: entry ' blk.18.ffn_gate_exps.weight' has partial data (98.83%)
58
+ 0.35.055.036 W save_imatrix: entry ' blk.24.ffn_up_exps.weight' has partial data (98.83%)
59
+ 0.35.055.040 W save_imatrix: entry ' blk.15.ffn_gate_exps.weight' has partial data (99.22%)
60
+ 0.35.055.041 W save_imatrix: entry ' blk.15.ffn_up_exps.weight' has partial data (99.22%)
61
+ 0.35.055.042 W save_imatrix: entry ' blk.22.ffn_gate_exps.weight' has partial data (99.22%)
62
+ 0.35.055.045 W save_imatrix: entry ' blk.16.ffn_down_exps.weight' has partial data (98.83%)
63
+
64
+ [10]4.3240,[11]4.0211,[12]4.4049,[13]4.9418,[14]5.1664,[15]5.6128,[16]5.8282,[17]6.1242,[18]6.4430,[19]6.1720,1.16.116.747 W
65
+ save_imatrix: saving imatrix using GGUF format with a different suffix than .gguf
66
+ 1.16.116.747 W save_imatrix: if you want the previous imatrix format, use --output-format dat
67
+
68
+ [20]6.1999,[21]6.2244,[22]6.2384,[23]6.2298,[24]6.4258,[25]6.5635,[26]6.5954,[27]6.6758,[28]6.7712,[29]7.0466,1.59.082.255 W
69
+ save_imatrix: saving imatrix using GGUF format with a different suffix than .gguf
70
+ 1.59.082.256 W save_imatrix: if you want the previous imatrix format, use --output-format dat
71
+
72
+ [30]7.0971,[31]6.8928,[32]6.5976,[33]6.3884,[34]6.2660,[35]6.2001,[36]6.3146,[37]6.4046,[38]6.4913,[39]6.6513,2.41.868.430 W
73
+ save_imatrix: saving imatrix using GGUF format with a different suffix than .gguf
74
+ 2.41.868.431 W save_imatrix: if you want the previous imatrix format, use --output-format dat
75
+
76
+ [40]6.8336,[41]6.9385,[42]7.2103,[43]7.3229,[44]7.4938,[45]7.5505,[46]7.5098,[47]7.4647,[48]7.6200,[49]7.6994,3.24.706.480 W
77
+ save_imatrix: saving imatrix using GGUF format with a different suffix than .gguf
78
+ 3.24.706.520 W save_imatrix: if you want the previous imatrix format, use --output-format dat
79
+
80
+ [50]7.6500,[51]7.5877,[52]7.6542,[53]7.7868,[54]7.8872,[55]7.9814,[56]8.0100,[57]8.0181,[58]8.0214,[59]8.0351,4.08.161.216 W
81
+ save_imatrix: saving imatrix using GGUF format with a different suffix than .gguf
82
+ 4.08.161.218 W save_imatrix: if you want the previous imatrix format, use --output-format dat
83
+
84
+ [60]8.0326,[61]7.9840,[62]7.9372,[63]7.9683,[64]8.0068,[65]7.9439,[66]7.9358,[67]7.9336,[68]7.8425,[69]7.8031,5.53.139.568 W
85
+ save_imatrix: saving imatrix using GGUF format with a different suffix than .gguf
86
+ 5.53.139.570 W save_imatrix: if you want the previous imatrix format, use --output-format dat
87
+
88
+ [70]7.8080,[71]7.7716,[72]7.7330,[73]7.7239,[74]7.6582,[75]7.5914,[76]7.5604,[77]7.5344,[78]7.5000,[79]7.4563,6.36.748.709 W
89
+ save_imatrix: saving imatrix using GGUF format with a different suffix than .gguf
90
+ 6.36.748.710 W save_imatrix: if you want the previous imatrix format, use --output-format dat
91
+
92
+ [80]7.3732,[81]7.3905,[82]7.3907,[83]7.3636,[84]7.3800,[85]7.4067,[86]7.3364,[87]7.3262,[88]7.3183,[89]7.3359,7.14.145.562 W
93
+ save_imatrix: saving imatrix using GGUF format with a different suffix than .gguf
94
+ 7.14.145.563 W save_imatrix: if you want the previous imatrix format, use --output-format dat
95
+
96
+ [90]7.3358,[91]7.3166,[92]7.2408,[93]7.1671,[94]7.0855,[95]7.0144,[96]6.9469,[97]6.8749,[98]6.8120,[99]6.7501,7.49.421.037 W
97
+ save_imatrix: saving imatrix using GGUF format with a different suffix than .gguf
98
+ 7.49.421.039 W save_imatrix: if you want the previous imatrix format, use --output-format dat
99
+
100
+ [100]6.7749,[101]6.7937,[102]6.8753,[103]6.9573,[104]7.0285,[105]7.1404,[106]7.2385,[107]7.2672,[108]7.2873,[109]7.2996,8.25.152.163 W
101
+ save_imatrix: saving imatrix using GGUF format with a different suffix than .gguf
102
+ 8.25.152.164 W save_imatrix: if you want the previous imatrix format, use --output-format dat
103
+
104
+ [110]7.3026,[111]7.2653,[112]7.1830,[113]7.1010,[114]7.1408,[115]7.1761,[116]7.1992,[117]7.2053,[118]7.2436,[119]7.2755,9.00.696.257 W
105
+ save_imatrix: saving imatrix using GGUF format with a different suffix than .gguf
106
+ 9.00.696.259 W save_imatrix: if you want the previous imatrix format, use --output-format dat
107
+
108
+ [120]7.2818,[121]7.2985,[122]7.3085,[123]7.2788,[124]7.3333,[125]7.3907,[126]7.4360,[127]7.5037,[128]7.5475,[129]7.6012,
109
+ Final estimate: PPL = 7.6012 +/- 0.10821
110
+ 9.36.441.617 W
111
+ save_imatrix: saving imatrix using GGUF format with a different suffix than .gguf
112
+ 9.36.441.617 W save_imatrix: if you want the previous imatrix format, use --output-format dat
113
+
114
+
recipe/logs/N3_q102i.log ADDED
The diff for this file is too large to render. See raw diff
 
recipe/logs/N3_q103i.log ADDED
The diff for this file is too large to render. See raw diff
 
recipe/logs/N3_q106i.log ADDED
The diff for this file is too large to render. See raw diff
 
recipe/logs/N3_readback.log ADDED
@@ -0,0 +1,4 @@
 
 
 
 
 
1
+ PASS Nex-N2.5-mini-imatrix-Q4_0_ROCMFP4_STRIX_LEAN.gguf arch=qwen35moe ftype=106 tensors=733 nextn=0 output.weight=Q6_K token_embd.weight=Q5_K
2
+ PASS Nex-N2.5-mini-imatrix-Q4_0_ROCMFP4_COHERENT.gguf arch=qwen35moe ftype=102 tensors=733 nextn=0 output.weight=Q6_K token_embd.weight=Q6_K
3
+ PASS Nex-N2.5-mini-imatrix-Q4_0_ROCMFP4_FAST.gguf arch=qwen35moe ftype=103 tensors=733 nextn=0 output.weight=Q6_K token_embd.weight=Q4_0_ROCMFP4_FAST
4
+ N3_DONE
recipe/logs/N4_kld_q102i.log ADDED
@@ -0,0 +1,92 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0.00.037.317 I common_init_result: fitting params to device memory ...
2
+ 0.00.037.321 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
3
+ 0.00.390.383 W llama_model_loader: direct I/O is enabled, disabling mmap
4
+ 0.21.870.791 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
5
+ 0.21.936.318 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
6
+ 0.22.143.467 I
7
+ 0.22.143.580 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
8
+ 0.22.258.917 I kl_divergence: computing over 40 chunks, n_ctx=2048, batch_size=2048, n_seq=1
9
+ 0.24.621.838 I kl_divergence: 2.36 seconds per pass - ETA 1.57 minutes
10
+
11
+ chunk PPL ln(PPL(Q)/PPL(base)) KL Divergence Δp RMS Same top p
12
+ 1 5.6299 ± 0.4284 -0.01123 ± 0.01762 0.10095 ± 0.00634 9.176 ± 0.605 % 88.270 ± 1.007 %
13
+ 2 6.6587 ± 0.3601 -0.00093 ± 0.01117 0.08359 ± 0.00360 7.978 ± 0.395 % 88.025 ± 0.718 %
14
+ 3 7.1385 ± 0.3226 0.01508 ± 0.00875 0.08242 ± 0.00314 7.696 ± 0.308 % 87.911 ± 0.589 %
15
+ 4 7.4207 ± 0.2990 0.01815 ± 0.00760 0.08463 ± 0.00438 7.777 ± 0.308 % 88.465 ± 0.499 %
16
+ 5 7.2795 ± 0.2642 0.01905 ± 0.00665 0.08239 ± 0.00359 7.768 ± 0.267 % 88.426 ± 0.447 %
17
+ 6 6.2745 ± 0.2014 0.02044 ± 0.00615 0.08353 ± 0.00339 8.576 ± 0.294 % 88.987 ± 0.400 %
18
+ 7 5.8589 ± 0.1719 0.02053 ± 0.00613 0.09196 ± 0.00373 9.221 ± 0.305 % 88.926 ± 0.371 %
19
+ 8 5.7865 ± 0.1577 0.02044 ± 0.00558 0.09026 ± 0.00333 9.081 ± 0.275 % 88.734 ± 0.350 %
20
+ 9 6.0902 ± 0.1574 0.01904 ± 0.00522 0.09017 ± 0.00300 8.885 ± 0.253 % 88.530 ± 0.332 %
21
+ 10 6.2036 ± 0.1532 0.01905 ± 0.00484 0.08664 ± 0.00271 8.652 ± 0.235 % 88.592 ± 0.314 %
22
+ 11 6.2550 ± 0.1467 0.01852 ± 0.00454 0.08389 ± 0.00248 8.471 ± 0.220 % 88.661 ± 0.299 %
23
+ 12 6.5002 ± 0.1472 0.01864 ± 0.00426 0.08096 ± 0.00228 8.271 ± 0.208 % 88.702 ± 0.286 %
24
+ 13 6.5547 ± 0.1424 0.02010 ± 0.00407 0.07947 ± 0.00212 8.179 ± 0.196 % 88.668 ± 0.275 %
25
+ 14 6.6139 ± 0.1383 0.02031 ± 0.00387 0.07794 ± 0.00199 8.054 ± 0.186 % 88.773 ± 0.264 %
26
+ 15 6.6603 ± 0.1347 0.02109 ± 0.00372 0.07724 ± 0.00186 7.982 ± 0.177 % 88.739 ± 0.255 %
27
+ 16 6.8133 ± 0.1336 0.01708 ± 0.00360 0.07665 ± 0.00178 7.924 ± 0.170 % 88.612 ± 0.248 %
28
+ 17 6.8557 ± 0.1300 0.01631 ± 0.00347 0.07576 ± 0.00169 7.851 ± 0.164 % 88.615 ± 0.241 %
29
+ 18 6.9493 ± 0.1282 0.01668 ± 0.00337 0.07548 ± 0.00160 7.811 ± 0.158 % 88.541 ± 0.235 %
30
+ 19 6.9014 ± 0.1244 0.01668 ± 0.00325 0.07445 ± 0.00154 7.725 ± 0.153 % 88.645 ± 0.228 %
31
+ 20 6.6526 ± 0.1161 0.01930 ± 0.00323 0.07848 ± 0.00152 7.994 ± 0.146 % 88.514 ± 0.223 %
32
+ 21 6.6703 ± 0.1135 0.01945 ± 0.00316 0.07876 ± 0.00146 7.982 ± 0.141 % 88.423 ± 0.218 %
33
+ 22 6.6837 ± 0.1111 0.01946 ± 0.00308 0.07952 ± 0.00143 8.019 ± 0.138 % 88.376 ± 0.214 %
34
+ 23 6.7415 ± 0.1096 0.02066 ± 0.00300 0.07952 ± 0.00138 7.991 ± 0.134 % 88.278 ± 0.210 %
35
+ 24 6.7373 ± 0.1070 0.02000 ± 0.00295 0.07956 ± 0.00138 7.986 ± 0.133 % 88.270 ± 0.205 %
36
+ 25 6.7719 ± 0.1054 0.02041 ± 0.00289 0.07925 ± 0.00134 7.953 ± 0.130 % 88.258 ± 0.201 %
37
+ 26 6.7420 ± 0.1028 0.02046 ± 0.00283 0.07951 ± 0.00131 7.991 ± 0.129 % 88.274 ± 0.197 %
38
+ 27 6.9113 ± 0.1041 0.02093 ± 0.00283 0.07927 ± 0.00136 7.933 ± 0.126 % 88.248 ± 0.194 %
39
+ 28 7.0022 ± 0.1039 0.02114 ± 0.00276 0.07854 ± 0.00132 7.873 ± 0.123 % 88.256 ± 0.190 %
40
+ 29 6.9981 ± 0.1020 0.02141 ± 0.00271 0.07833 ± 0.00128 7.882 ± 0.121 % 88.250 ± 0.187 %
41
+ 30 6.9536 ± 0.0995 0.02303 ± 0.00267 0.07848 ± 0.00125 7.892 ± 0.119 % 88.244 ± 0.184 %
42
+ 31 6.8548 ± 0.0962 0.02323 ± 0.00261 0.07786 ± 0.00122 7.868 ± 0.117 % 88.320 ± 0.180 %
43
+ 32 6.7489 ± 0.0931 0.02193 ± 0.00261 0.08043 ± 0.00142 8.006 ± 0.117 % 88.273 ± 0.178 %
44
+ 33 6.6865 ± 0.0906 0.02252 ± 0.00256 0.07990 ± 0.00139 7.989 ± 0.115 % 88.299 ± 0.175 %
45
+ 34 6.6723 ± 0.0889 0.02256 ± 0.00251 0.07907 ± 0.00135 7.928 ± 0.113 % 88.365 ± 0.172 %
46
+ 35 6.6827 ± 0.0878 0.02231 ± 0.00246 0.07849 ± 0.00131 7.908 ± 0.111 % 88.404 ± 0.169 %
47
+ 36 6.6954 ± 0.0868 0.02172 ± 0.00242 0.07826 ± 0.00128 7.892 ± 0.110 % 88.414 ± 0.167 %
48
+ 37 6.6026 ± 0.0841 0.02157 ± 0.00238 0.07763 ± 0.00125 7.875 ± 0.108 % 88.407 ± 0.165 %
49
+ 38 6.5309 ± 0.0818 0.02118 ± 0.00235 0.07760 ± 0.00123 7.895 ± 0.106 % 88.432 ± 0.162 %
50
+ 39 6.4522 ± 0.0795 0.02138 ± 0.00232 0.07746 ± 0.00121 7.905 ± 0.105 % 88.423 ± 0.160 %
51
+ 40 6.3601 ± 0.0770 0.02082 ± 0.00229 0.07695 ± 0.00118 7.909 ± 0.104 % 88.463 ± 0.158 %
52
+
53
+ ====== Perplexity statistics ======
54
+ Mean PPL(Q) : 6.360050 ± 0.077045
55
+ Mean PPL(base) : 6.228979 ± 0.075322
56
+ Cor(ln(PPL(Q)), ln(PPL(base))): 98.21%
57
+ Mean ln(PPL(Q)/PPL(base)) : 0.020824 ± 0.002287
58
+ Mean PPL(Q)/PPL(base) : 1.021042 ± 0.002336
59
+ Mean PPL(Q)-PPL(base) : 0.131071 ± 0.014499
60
+
61
+ ====== KL divergence statistics ======
62
+ Mean KLD: 0.076947 ± 0.001181
63
+ Maximum KLD: 16.293211
64
+ 99.9% KLD: 2.499913
65
+ 99.0% KLD: 0.685225
66
+ 95.0% KLD: 0.259541
67
+ 90.0% KLD: 0.162843
68
+ Median KLD: 0.034241
69
+ 10.0% KLD: 0.000500
70
+ 5.0% KLD: 0.000134
71
+ 1.0% KLD: -0.000030
72
+ 0.1% KLD: -0.000274
73
+ Minimum KLD: -0.000654
74
+
75
+ ====== Token probability statistics ======
76
+ Mean Δp: -0.505 ± 0.039 %
77
+ Maximum Δp: 98.233%
78
+ 99.9% Δp: 50.439%
79
+ 99.0% Δp: 21.309%
80
+ 95.0% Δp: 9.386%
81
+ 90.0% Δp: 5.261%
82
+ 75.0% Δp: 0.918%
83
+ Median Δp: -0.011%
84
+ 25.0% Δp: -1.620%
85
+ 10.0% Δp: -6.708%
86
+ 5.0% Δp: -11.584%
87
+ 1.0% Δp: -26.637%
88
+ 0.1% Δp: -64.007%
89
+ Minimum Δp: -99.856%
90
+ RMS Δp : 7.909 ± 0.104 %
91
+ Same top p: 88.463 ± 0.158 %
92
+
recipe/logs/N4_kld_q103i.log ADDED
@@ -0,0 +1,92 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0.00.039.723 I common_init_result: fitting params to device memory ...
2
+ 0.00.039.726 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
3
+ 0.00.396.902 W llama_model_loader: direct I/O is enabled, disabling mmap
4
+ 0.20.542.121 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
5
+ 0.20.588.569 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
6
+ 0.20.783.826 I
7
+ 0.20.783.939 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
8
+ 0.20.898.218 I kl_divergence: computing over 40 chunks, n_ctx=2048, batch_size=2048, n_seq=1
9
+ 0.22.619.669 I kl_divergence: 1.72 seconds per pass - ETA 1.13 minutes
10
+
11
+ chunk PPL ln(PPL(Q)/PPL(base)) KL Divergence Δp RMS Same top p
12
+ 1 5.7545 ± 0.4426 0.01066 ± 0.01514 0.10310 ± 0.00565 9.252 ± 0.547 % 86.901 ± 1.055 %
13
+ 2 6.7692 ± 0.3702 0.01553 ± 0.01057 0.08979 ± 0.00354 8.313 ± 0.392 % 87.097 ± 0.741 %
14
+ 3 7.2118 ± 0.3286 0.02530 ± 0.00859 0.09142 ± 0.00371 7.924 ± 0.306 % 86.901 ± 0.609 %
15
+ 4 7.4361 ± 0.3008 0.02022 ± 0.00807 0.09891 ± 0.00569 8.314 ± 0.317 % 87.414 ± 0.519 %
16
+ 5 7.2981 ± 0.2661 0.02160 ± 0.00714 0.09587 ± 0.00463 8.276 ± 0.280 % 87.527 ± 0.462 %
17
+ 6 6.3222 ± 0.2036 0.02801 ± 0.00679 0.09876 ± 0.00434 9.164 ± 0.306 % 87.814 ± 0.418 %
18
+ 7 5.8837 ± 0.1729 0.02475 ± 0.00651 0.10562 ± 0.00468 9.504 ± 0.293 % 87.837 ± 0.386 %
19
+ 8 5.8219 ± 0.1591 0.02653 ± 0.00598 0.10395 ± 0.00416 9.429 ± 0.268 % 87.732 ± 0.363 %
20
+ 9 6.1195 ± 0.1584 0.02383 ± 0.00560 0.10437 ± 0.00375 9.307 ± 0.250 % 87.281 ± 0.347 %
21
+ 10 6.2384 ± 0.1544 0.02466 ± 0.00519 0.10032 ± 0.00340 9.080 ± 0.233 % 87.331 ± 0.329 %
22
+ 11 6.3092 ± 0.1486 0.02715 ± 0.00488 0.09769 ± 0.00312 8.996 ± 0.220 % 87.292 ± 0.314 %
23
+ 12 6.5586 ± 0.1493 0.02759 ± 0.00459 0.09446 ± 0.00286 8.773 ± 0.208 % 87.235 ± 0.301 %
24
+ 13 6.6005 ± 0.1440 0.02706 ± 0.00438 0.09304 ± 0.00266 8.721 ± 0.198 % 87.292 ± 0.289 %
25
+ 14 6.6554 ± 0.1398 0.02657 ± 0.00417 0.09137 ± 0.00249 8.562 ± 0.187 % 87.264 ± 0.279 %
26
+ 15 6.6958 ± 0.1359 0.02640 ± 0.00402 0.09068 ± 0.00234 8.515 ± 0.178 % 87.175 ± 0.270 %
27
+ 16 6.8581 ± 0.1350 0.02363 ± 0.00387 0.08962 ± 0.00221 8.427 ± 0.172 % 87.060 ± 0.262 %
28
+ 17 6.9067 ± 0.1315 0.02373 ± 0.00373 0.08844 ± 0.00209 8.335 ± 0.164 % 87.120 ± 0.254 %
29
+ 18 6.9978 ± 0.1297 0.02365 ± 0.00361 0.08829 ± 0.00199 8.313 ± 0.158 % 87.075 ± 0.247 %
30
+ 19 6.9441 ± 0.1257 0.02285 ± 0.00349 0.08712 ± 0.00190 8.208 ± 0.152 % 87.143 ± 0.240 %
31
+ 20 6.6745 ± 0.1169 0.02258 ± 0.00348 0.09166 ± 0.00187 8.508 ± 0.146 % 87.082 ± 0.234 %
32
+ 21 6.6839 ± 0.1140 0.02149 ± 0.00341 0.09200 ± 0.00180 8.487 ± 0.142 % 87.041 ± 0.229 %
33
+ 22 6.6958 ± 0.1115 0.02127 ± 0.00336 0.09301 ± 0.00182 8.516 ± 0.141 % 87.079 ± 0.224 %
34
+ 23 6.7561 ± 0.1101 0.02283 ± 0.00329 0.09310 ± 0.00178 8.515 ± 0.138 % 86.974 ± 0.219 %
35
+ 24 6.7528 ± 0.1075 0.02230 ± 0.00321 0.09247 ± 0.00172 8.468 ± 0.134 % 87.036 ± 0.214 %
36
+ 25 6.7907 ± 0.1060 0.02318 ± 0.00313 0.09180 ± 0.00165 8.416 ± 0.131 % 87.003 ± 0.210 %
37
+ 26 6.7642 ± 0.1034 0.02375 ± 0.00307 0.09231 ± 0.00162 8.481 ± 0.129 % 87.067 ± 0.206 %
38
+ 27 6.9325 ± 0.1047 0.02399 ± 0.00301 0.09164 ± 0.00159 8.400 ± 0.127 % 87.068 ± 0.202 %
39
+ 28 7.0252 ± 0.1045 0.02441 ± 0.00295 0.09094 ± 0.00155 8.369 ± 0.125 % 87.125 ± 0.198 %
40
+ 29 7.0215 ± 0.1026 0.02475 ± 0.00290 0.09104 ± 0.00151 8.421 ± 0.124 % 87.144 ± 0.194 %
41
+ 30 6.9731 ± 0.1000 0.02583 ± 0.00286 0.09121 ± 0.00147 8.431 ± 0.121 % 87.120 ± 0.191 %
42
+ 31 6.8743 ± 0.0967 0.02607 ± 0.00280 0.09057 ± 0.00143 8.408 ± 0.119 % 87.207 ± 0.188 %
43
+ 32 6.7690 ± 0.0936 0.02490 ± 0.00278 0.09337 ± 0.00160 8.561 ± 0.119 % 87.201 ± 0.185 %
44
+ 33 6.7036 ± 0.0910 0.02507 ± 0.00273 0.09280 ± 0.00156 8.553 ± 0.117 % 87.263 ± 0.181 %
45
+ 34 6.6897 ± 0.0893 0.02516 ± 0.00267 0.09179 ± 0.00151 8.490 ± 0.115 % 87.315 ± 0.178 %
46
+ 35 6.6976 ± 0.0881 0.02453 ± 0.00261 0.09113 ± 0.00148 8.437 ± 0.113 % 87.393 ± 0.175 %
47
+ 36 6.7103 ± 0.0872 0.02395 ± 0.00257 0.09082 ± 0.00144 8.395 ± 0.110 % 87.404 ± 0.173 %
48
+ 37 6.6143 ± 0.0845 0.02333 ± 0.00253 0.08990 ± 0.00141 8.356 ± 0.108 % 87.451 ± 0.170 %
49
+ 38 6.5394 ± 0.0821 0.02247 ± 0.00250 0.08996 ± 0.00138 8.378 ± 0.107 % 87.452 ± 0.168 %
50
+ 39 6.4598 ± 0.0798 0.02255 ± 0.00247 0.08964 ± 0.00135 8.375 ± 0.105 % 87.430 ± 0.166 %
51
+ 40 6.3681 ± 0.0773 0.02208 ± 0.00243 0.08897 ± 0.00132 8.370 ± 0.103 % 87.454 ± 0.164 %
52
+
53
+ ====== Perplexity statistics ======
54
+ Mean PPL(Q) : 6.368075 ± 0.077317
55
+ Mean PPL(base) : 6.228979 ± 0.075322
56
+ Cor(ln(PPL(Q)), ln(PPL(base))): 97.99%
57
+ Mean ln(PPL(Q)/PPL(base)) : 0.022085 ± 0.002430
58
+ Mean PPL(Q)/PPL(base) : 1.022330 ± 0.002484
59
+ Mean PPL(Q)-PPL(base) : 0.139096 ± 0.015431
60
+
61
+ ====== KL divergence statistics ======
62
+ Mean KLD: 0.088974 ± 0.001318
63
+ Maximum KLD: 14.587295
64
+ 99.9% KLD: 2.825675
65
+ 99.0% KLD: 0.784672
66
+ 95.0% KLD: 0.301152
67
+ 90.0% KLD: 0.188552
68
+ Median KLD: 0.039738
69
+ 10.0% KLD: 0.000526
70
+ 5.0% KLD: 0.000124
71
+ 1.0% KLD: -0.000056
72
+ 0.1% KLD: -0.000307
73
+ Minimum KLD: -0.000906
74
+
75
+ ====== Token probability statistics ======
76
+ Mean Δp: -0.439 ± 0.041 %
77
+ Maximum Δp: 97.378%
78
+ 99.9% Δp: 51.015%
79
+ 99.0% Δp: 22.861%
80
+ 95.0% Δp: 10.138%
81
+ 90.0% Δp: 5.884%
82
+ 75.0% Δp: 1.147%
83
+ Median Δp: -0.001%
84
+ 25.0% Δp: -1.525%
85
+ 10.0% Δp: -6.868%
86
+ 5.0% Δp: -12.274%
87
+ 1.0% Δp: -29.206%
88
+ 0.1% Δp: -67.613%
89
+ Minimum Δp: -98.545%
90
+ RMS Δp : 8.370 ± 0.103 %
91
+ Same top p: 87.454 ± 0.164 %
92
+
recipe/logs/N4_kld_q106.log ADDED
@@ -0,0 +1,92 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0.00.042.882 I common_init_result: fitting params to device memory ...
2
+ 0.00.042.887 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
3
+ 0.00.541.611 W llama_model_loader: direct I/O is enabled, disabling mmap
4
+ 0.22.255.822 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
5
+ 0.22.359.058 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
6
+ 0.22.643.059 I
7
+ 0.22.643.196 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
8
+ 0.22.776.319 I kl_divergence: computing over 40 chunks, n_ctx=2048, batch_size=2048, n_seq=1
9
+ 0.25.307.079 I kl_divergence: 2.53 seconds per pass - ETA 1.68 minutes
10
+
11
+ chunk PPL ln(PPL(Q)/PPL(base)) KL Divergence Δp RMS Same top p
12
+ 1 6.2617 ± 0.5005 0.09513 ± 0.01958 0.13331 ± 0.00723 11.273 ± 0.670 % 85.728 ± 1.094 %
13
+ 2 7.1013 ± 0.3976 0.06343 ± 0.01246 0.11005 ± 0.00427 9.422 ± 0.440 % 86.413 ± 0.758 %
14
+ 3 7.5058 ± 0.3468 0.06527 ± 0.01057 0.11313 ± 0.00423 9.250 ± 0.346 % 86.022 ± 0.626 %
15
+ 4 7.7849 ± 0.3204 0.06605 ± 0.00908 0.11249 ± 0.00383 9.264 ± 0.306 % 86.388 ± 0.536 %
16
+ 5 7.5801 ± 0.2810 0.05952 ± 0.00798 0.11043 ± 0.00322 9.087 ± 0.266 % 86.725 ± 0.474 %
17
+ 6 6.5014 ± 0.2131 0.05596 ± 0.00724 0.11026 ± 0.00320 9.619 ± 0.279 % 87.178 ± 0.427 %
18
+ 7 6.0313 ± 0.1806 0.04953 ± 0.00696 0.12172 ± 0.00436 10.069 ± 0.275 % 87.236 ± 0.394 %
19
+ 8 5.9472 ± 0.1655 0.04783 ± 0.00647 0.12058 ± 0.00391 9.949 ± 0.252 % 87.121 ± 0.370 %
20
+ 9 6.2446 ± 0.1643 0.04408 ± 0.00605 0.12012 ± 0.00353 9.791 ± 0.233 % 86.597 ± 0.355 %
21
+ 10 6.3764 ± 0.1603 0.04654 ± 0.00565 0.11624 ± 0.00321 9.609 ± 0.217 % 86.588 ± 0.337 %
22
+ 11 6.4374 ± 0.1538 0.04726 ± 0.00530 0.11268 ± 0.00294 9.442 ± 0.204 % 86.484 ± 0.322 %
23
+ 12 6.6950 ± 0.1547 0.04816 ± 0.00499 0.10924 ± 0.00271 9.238 ± 0.193 % 86.388 ± 0.310 %
24
+ 13 6.7262 ± 0.1490 0.04592 ± 0.00474 0.10773 ± 0.00253 9.174 ± 0.183 % 86.413 ± 0.297 %
25
+ 14 6.7857 ± 0.1447 0.04595 ± 0.00454 0.10583 ± 0.00237 9.050 ± 0.174 % 86.426 ± 0.286 %
26
+ 15 6.8279 ± 0.1408 0.04594 ± 0.00437 0.10522 ± 0.00224 9.021 ± 0.166 % 86.419 ± 0.277 %
27
+ 16 6.9800 ± 0.1394 0.04125 ± 0.00424 0.10488 ± 0.00213 8.960 ± 0.160 % 86.339 ± 0.268 %
28
+ 17 7.0246 ± 0.1356 0.04064 ± 0.00408 0.10348 ± 0.00202 8.863 ± 0.153 % 86.424 ± 0.260 %
29
+ 18 7.1128 ± 0.1335 0.03994 ± 0.00395 0.10297 ± 0.00192 8.819 ± 0.147 % 86.385 ± 0.253 %
30
+ 19 7.0597 ± 0.1294 0.03935 ± 0.00383 0.10135 ± 0.00184 8.726 ± 0.143 % 86.474 ± 0.245 %
31
+ 20 6.7990 ± 0.1207 0.04105 ± 0.00383 0.10693 ± 0.00185 9.139 ± 0.143 % 86.393 ± 0.240 %
32
+ 21 6.8267 ± 0.1182 0.04262 ± 0.00375 0.10765 ± 0.00179 9.144 ± 0.139 % 86.375 ± 0.234 %
33
+ 22 6.8456 ± 0.1159 0.04339 ± 0.00369 0.10893 ± 0.00181 9.184 ± 0.137 % 86.417 ± 0.228 %
34
+ 23 6.9000 ± 0.1143 0.04390 ± 0.00360 0.10839 ± 0.00175 9.147 ± 0.133 % 86.383 ± 0.224 %
35
+ 24 6.8942 ± 0.1115 0.04302 ± 0.00351 0.10851 ± 0.00171 9.165 ± 0.132 % 86.376 ± 0.219 %
36
+ 25 6.9297 ± 0.1099 0.04345 ± 0.00344 0.10819 ± 0.00165 9.128 ± 0.128 % 86.334 ± 0.215 %
37
+ 26 6.8989 ± 0.1072 0.04347 ± 0.00338 0.10899 ± 0.00166 9.185 ± 0.127 % 86.371 ± 0.210 %
38
+ 27 7.0668 ± 0.1084 0.04318 ± 0.00330 0.10805 ± 0.00160 9.103 ± 0.124 % 86.384 ± 0.206 %
39
+ 28 7.1505 ± 0.1080 0.04210 ± 0.00321 0.10685 ± 0.00155 9.028 ± 0.122 % 86.395 ± 0.203 %
40
+ 29 7.1560 ± 0.1063 0.04372 ± 0.00317 0.10671 ± 0.00152 9.048 ± 0.120 % 86.392 ± 0.199 %
41
+ 30 7.0985 ± 0.1034 0.04365 ± 0.00312 0.10649 ± 0.00148 9.044 ± 0.118 % 86.341 ± 0.196 %
42
+ 31 6.9900 ± 0.0998 0.04276 ± 0.00305 0.10563 ± 0.00144 9.018 ± 0.116 % 86.444 ± 0.192 %
43
+ 32 6.8747 ± 0.0965 0.04039 ± 0.00305 0.10917 ± 0.00170 9.201 ± 0.119 % 86.416 ± 0.189 %
44
+ 33 6.8130 ± 0.0939 0.04125 ± 0.00301 0.10876 ± 0.00166 9.199 ± 0.117 % 86.457 ± 0.186 %
45
+ 34 6.7927 ± 0.0920 0.04044 ± 0.00294 0.10768 ± 0.00161 9.128 ± 0.115 % 86.441 ± 0.184 %
46
+ 35 6.8056 ± 0.0909 0.04054 ± 0.00288 0.10692 ± 0.00157 9.083 ± 0.113 % 86.494 ± 0.181 %
47
+ 36 6.8222 ± 0.0899 0.04049 ± 0.00284 0.10659 ± 0.00154 9.060 ± 0.111 % 86.546 ± 0.178 %
48
+ 37 6.7263 ± 0.0872 0.04013 ± 0.00279 0.10571 ± 0.00150 9.033 ± 0.109 % 86.566 ± 0.175 %
49
+ 38 6.6472 ± 0.0847 0.03882 ± 0.00275 0.10525 ± 0.00147 9.020 ± 0.107 % 86.621 ± 0.173 %
50
+ 39 6.5675 ± 0.0823 0.03909 ± 0.00271 0.10501 ± 0.00144 9.022 ± 0.106 % 86.638 ± 0.170 %
51
+ 40 6.4740 ± 0.0798 0.03858 ± 0.00267 0.10444 ± 0.00141 9.028 ± 0.105 % 86.659 ± 0.168 %
52
+
53
+ ====== Perplexity statistics ======
54
+ Mean PPL(Q) : 6.473964 ± 0.079799
55
+ Mean PPL(base) : 6.228979 ± 0.075322
56
+ Cor(ln(PPL(Q)), ln(PPL(base))): 97.63%
57
+ Mean ln(PPL(Q)/PPL(base)) : 0.038576 ± 0.002667
58
+ Mean PPL(Q)/PPL(base) : 1.039330 ± 0.002772
59
+ Mean PPL(Q)-PPL(base) : 0.244985 ± 0.017453
60
+
61
+ ====== KL divergence statistics ======
62
+ Mean KLD: 0.104436 ± 0.001413
63
+ Maximum KLD: 13.909879
64
+ 99.9% KLD: 3.189192
65
+ 99.0% KLD: 0.926687
66
+ 95.0% KLD: 0.354656
67
+ 90.0% KLD: 0.225739
68
+ Median KLD: 0.047702
69
+ 10.0% KLD: 0.000698
70
+ 5.0% KLD: 0.000183
71
+ 1.0% KLD: -0.000033
72
+ 0.1% KLD: -0.000321
73
+ Minimum KLD: -0.000718
74
+
75
+ ====== Token probability statistics ======
76
+ Mean Δp: -0.244 ± 0.045 %
77
+ Maximum Δp: 99.706%
78
+ 99.9% Δp: 59.493%
79
+ 99.0% Δp: 25.100%
80
+ 95.0% Δp: 11.707%
81
+ 90.0% Δp: 6.963%
82
+ 75.0% Δp: 1.453%
83
+ Median Δp: -0.001%
84
+ 25.0% Δp: -1.479%
85
+ 10.0% Δp: -7.240%
86
+ 5.0% Δp: -12.992%
87
+ 1.0% Δp: -31.268%
88
+ 0.1% Δp: -68.467%
89
+ Minimum Δp: -98.412%
90
+ RMS Δp : 9.028 ± 0.105 %
91
+ Same top p: 86.659 ± 0.168 %
92
+
recipe/logs/N4_kld_q106i.log ADDED
@@ -0,0 +1,92 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0.00.061.730 I common_init_result: fitting params to device memory ...
2
+ 0.00.061.734 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
3
+ 0.00.418.389 W llama_model_loader: direct I/O is enabled, disabling mmap
4
+ 0.21.286.189 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
5
+ 0.21.331.409 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
6
+ 0.21.529.206 I
7
+ 0.21.529.313 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
8
+ 0.21.655.658 I kl_divergence: computing over 40 chunks, n_ctx=2048, batch_size=2048, n_seq=1
9
+ 0.23.558.651 I kl_divergence: 1.90 seconds per pass - ETA 1.27 minutes
10
+
11
+ chunk PPL ln(PPL(Q)/PPL(base)) KL Divergence Δp RMS Same top p
12
+ 1 5.7279 ± 0.4388 0.00602 ± 0.01614 0.10315 ± 0.00583 9.735 ± 0.571 % 85.924 ± 1.088 %
13
+ 2 6.7923 ± 0.3712 0.01894 ± 0.01065 0.08743 ± 0.00335 8.343 ± 0.375 % 86.413 ± 0.758 %
14
+ 3 7.2151 ± 0.3282 0.02577 ± 0.00843 0.08904 ± 0.00386 8.063 ± 0.325 % 86.999 ± 0.607 %
15
+ 4 7.4647 ± 0.3021 0.02405 ± 0.00738 0.09108 ± 0.00535 8.270 ± 0.318 % 87.634 ± 0.515 %
16
+ 5 7.3108 ± 0.2664 0.02334 ± 0.00659 0.08904 ± 0.00437 8.208 ± 0.278 % 87.801 ± 0.458 %
17
+ 6 6.2917 ± 0.2025 0.02318 ± 0.00607 0.08932 ± 0.00383 8.727 ± 0.274 % 88.237 ± 0.411 %
18
+ 7 5.8298 ± 0.1710 0.01556 ± 0.00599 0.09705 ± 0.00425 9.211 ± 0.278 % 88.242 ± 0.381 %
19
+ 8 5.7710 ± 0.1572 0.01775 ± 0.00555 0.09614 ± 0.00379 9.187 ± 0.255 % 88.038 ± 0.359 %
20
+ 9 6.0701 ± 0.1567 0.01574 ± 0.00523 0.09667 ± 0.00342 9.059 ± 0.236 % 87.553 ± 0.344 %
21
+ 10 6.1910 ± 0.1528 0.01702 ± 0.00487 0.09332 ± 0.00310 8.829 ± 0.220 % 87.625 ± 0.326 %
22
+ 11 6.2600 ± 0.1469 0.01931 ± 0.00460 0.09110 ± 0.00284 8.756 ± 0.210 % 87.630 ± 0.310 %
23
+ 12 6.5141 ± 0.1478 0.02077 ± 0.00433 0.08827 ± 0.00261 8.573 ± 0.198 % 87.569 ± 0.298 %
24
+ 13 6.5510 ± 0.1424 0.01954 ± 0.00414 0.08704 ± 0.00243 8.533 ± 0.188 % 87.533 ± 0.286 %
25
+ 14 6.6159 ± 0.1386 0.02062 ± 0.00397 0.08562 ± 0.00227 8.395 ± 0.178 % 87.565 ± 0.276 %
26
+ 15 6.6597 ± 0.1349 0.02101 ± 0.00382 0.08498 ± 0.00213 8.353 ± 0.170 % 87.481 ± 0.267 %
27
+ 16 6.8126 ± 0.1337 0.01697 ± 0.00370 0.08434 ± 0.00203 8.278 ± 0.164 % 87.390 ± 0.259 %
28
+ 17 6.8618 ± 0.1303 0.01721 ± 0.00357 0.08322 ± 0.00192 8.183 ± 0.157 % 87.511 ± 0.251 %
29
+ 18 6.9534 ± 0.1285 0.01728 ± 0.00345 0.08273 ± 0.00182 8.130 ± 0.150 % 87.471 ± 0.244 %
30
+ 19 6.9010 ± 0.1245 0.01662 ± 0.00334 0.08204 ± 0.00175 8.059 ± 0.147 % 87.508 ± 0.237 %
31
+ 20 6.6434 ± 0.1160 0.01790 ± 0.00333 0.08651 ± 0.00173 8.357 ± 0.142 % 87.454 ± 0.232 %
32
+ 21 6.6543 ± 0.1132 0.01704 ± 0.00325 0.08675 ± 0.00166 8.313 ± 0.137 % 87.399 ± 0.226 %
33
+ 22 6.6714 ± 0.1109 0.01762 ± 0.00322 0.08786 ± 0.00171 8.337 ± 0.136 % 87.412 ± 0.221 %
34
+ 23 6.7286 ± 0.1095 0.01874 ± 0.00316 0.08790 ± 0.00167 8.321 ± 0.134 % 87.390 ± 0.216 %
35
+ 24 6.7323 ± 0.1070 0.01926 ± 0.00310 0.08781 ± 0.00164 8.305 ± 0.133 % 87.447 ± 0.211 %
36
+ 25 6.7692 ± 0.1055 0.02001 ± 0.00302 0.08721 ± 0.00158 8.253 ± 0.129 % 87.457 ± 0.207 %
37
+ 26 6.7420 ± 0.1029 0.02046 ± 0.00296 0.08758 ± 0.00155 8.276 ± 0.128 % 87.522 ± 0.203 %
38
+ 27 6.9087 ± 0.1041 0.02056 ± 0.00290 0.08699 ± 0.00151 8.214 ± 0.125 % 87.520 ± 0.199 %
39
+ 28 6.9975 ± 0.1039 0.02047 ± 0.00285 0.08640 ± 0.00147 8.197 ± 0.124 % 87.558 ± 0.195 %
40
+ 29 6.9967 ± 0.1020 0.02121 ± 0.00280 0.08643 ± 0.00144 8.249 ± 0.123 % 87.538 ± 0.192 %
41
+ 30 6.9524 ± 0.0996 0.02286 ± 0.00276 0.08671 ± 0.00140 8.260 ± 0.120 % 87.511 ± 0.189 %
42
+ 31 6.8530 ± 0.0962 0.02296 ± 0.00270 0.08616 ± 0.00136 8.244 ± 0.118 % 87.623 ± 0.185 %
43
+ 32 6.7437 ± 0.0931 0.02116 ± 0.00271 0.08891 ± 0.00158 8.411 ± 0.120 % 87.628 ± 0.182 %
44
+ 33 6.6801 ± 0.0905 0.02156 ± 0.00266 0.08849 ± 0.00154 8.418 ± 0.119 % 87.666 ± 0.179 %
45
+ 34 6.6705 ± 0.0889 0.02230 ± 0.00261 0.08778 ± 0.00150 8.378 ± 0.117 % 87.715 ± 0.176 %
46
+ 35 6.6811 ± 0.0878 0.02207 ± 0.00255 0.08711 ± 0.00147 8.332 ± 0.115 % 87.792 ± 0.173 %
47
+ 36 6.6976 ± 0.0869 0.02206 ± 0.00252 0.08700 ± 0.00143 8.296 ± 0.113 % 87.814 ± 0.170 %
48
+ 37 6.6017 ± 0.0842 0.02143 ± 0.00248 0.08616 ± 0.00140 8.262 ± 0.111 % 87.823 ± 0.168 %
49
+ 38 6.5268 ± 0.0818 0.02054 ± 0.00245 0.08605 ± 0.00137 8.265 ± 0.109 % 87.853 ± 0.166 %
50
+ 39 6.4448 ± 0.0795 0.02022 ± 0.00241 0.08583 ± 0.00134 8.292 ± 0.107 % 87.829 ± 0.164 %
51
+ 40 6.3536 ± 0.0770 0.01981 ± 0.00238 0.08516 ± 0.00131 8.281 ± 0.106 % 87.849 ± 0.162 %
52
+
53
+ ====== Perplexity statistics ======
54
+ Mean PPL(Q) : 6.353616 ± 0.077001
55
+ Mean PPL(base) : 6.228979 ± 0.075322
56
+ Cor(ln(PPL(Q)), ln(PPL(base))): 98.08%
57
+ Mean ln(PPL(Q)/PPL(base)) : 0.019812 ± 0.002375
58
+ Mean PPL(Q)/PPL(base) : 1.020009 ± 0.002423
59
+ Mean PPL(Q)-PPL(base) : 0.124637 ± 0.015034
60
+
61
+ ====== KL divergence statistics ======
62
+ Mean KLD: 0.085158 ± 0.001309
63
+ Maximum KLD: 16.350994
64
+ 99.9% KLD: 2.569785
65
+ 99.0% KLD: 0.737186
66
+ 95.0% KLD: 0.287234
67
+ 90.0% KLD: 0.182120
68
+ Median KLD: 0.038436
69
+ 10.0% KLD: 0.000529
70
+ 5.0% KLD: 0.000132
71
+ 1.0% KLD: -0.000040
72
+ 0.1% KLD: -0.000280
73
+ Minimum KLD: -0.000917
74
+
75
+ ====== Token probability statistics ======
76
+ Mean Δp: -0.453 ± 0.041 %
77
+ Maximum Δp: 99.545%
78
+ 99.9% Δp: 55.136%
79
+ 99.0% Δp: 23.249%
80
+ 95.0% Δp: 9.945%
81
+ 90.0% Δp: 5.745%
82
+ 75.0% Δp: 0.987%
83
+ Median Δp: -0.006%
84
+ 25.0% Δp: -1.591%
85
+ 10.0% Δp: -7.009%
86
+ 5.0% Δp: -11.986%
87
+ 1.0% Δp: -27.836%
88
+ 0.1% Δp: -65.986%
89
+ Minimum Δp: -99.892%
90
+ RMS Δp : 8.281 ± 0.106 %
91
+ Same top p: 87.849 ± 0.162 %
92
+
recipe/logs/N4v_kld_q102i.log ADDED
@@ -0,0 +1,93 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0.00.035.061 I common_init_result: fitting params to device memory ...
2
+ 0.00.035.064 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
3
+ 0.00.364.958 W llama_model_loader: direct I/O is enabled, disabling mmap
4
+ 0.01.365.200 W read_raw_unsafe: Falling back to buffered IO due to Bad address
5
+ 0.20.724.587 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
6
+ 0.20.776.655 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
7
+ 0.20.958.449 I
8
+ 0.20.958.535 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
9
+ 0.21.072.467 I kl_divergence: computing over 40 chunks, n_ctx=2048, batch_size=2048, n_seq=1
10
+ 0.23.478.059 I kl_divergence: 2.41 seconds per pass - ETA 1.60 minutes
11
+
12
+ chunk PPL ln(PPL(Q)/PPL(base)) KL Divergence Δp RMS Same top p
13
+ 1 5.5867 ± 0.4228 -0.01894 ± 0.01754 0.10256 ± 0.00648 9.538 ± 0.610 % 88.368 ± 1.003 %
14
+ 2 6.6450 ± 0.3596 -0.00300 ± 0.01113 0.08488 ± 0.00372 8.059 ± 0.404 % 88.270 ± 0.712 %
15
+ 3 7.0875 ± 0.3197 0.00792 ± 0.00873 0.08299 ± 0.00332 7.716 ± 0.310 % 88.172 ± 0.583 %
16
+ 4 7.3455 ± 0.2950 0.00796 ± 0.00746 0.08559 ± 0.00479 7.862 ± 0.310 % 88.759 ± 0.494 %
17
+ 5 7.2194 ± 0.2615 0.01075 ± 0.00655 0.08256 ± 0.00391 7.727 ± 0.267 % 88.524 ± 0.446 %
18
+ 6 6.2395 ± 0.1998 0.01485 ± 0.00603 0.08240 ± 0.00357 8.402 ± 0.286 % 88.970 ± 0.400 %
19
+ 7 5.8467 ± 0.1714 0.01845 ± 0.00617 0.09250 ± 0.00438 9.020 ± 0.301 % 88.884 ± 0.371 %
20
+ 8 5.7766 ± 0.1573 0.01872 ± 0.00560 0.09000 ± 0.00387 8.841 ± 0.272 % 88.869 ± 0.348 %
21
+ 9 6.0753 ± 0.1568 0.01659 ± 0.00524 0.08968 ± 0.00347 8.651 ± 0.250 % 88.617 ± 0.331 %
22
+ 10 6.1938 ± 0.1528 0.01748 ± 0.00486 0.08628 ± 0.00314 8.443 ± 0.232 % 88.778 ± 0.312 %
23
+ 11 6.2507 ± 0.1465 0.01783 ± 0.00456 0.08369 ± 0.00287 8.284 ± 0.217 % 88.812 ± 0.297 %
24
+ 12 6.4984 ± 0.1471 0.01836 ± 0.00428 0.08087 ± 0.00264 8.100 ± 0.205 % 88.840 ± 0.284 %
25
+ 13 6.5527 ± 0.1424 0.01980 ± 0.00408 0.07955 ± 0.00245 8.022 ± 0.193 % 88.646 ± 0.275 %
26
+ 14 6.6129 ± 0.1383 0.02016 ± 0.00389 0.07798 ± 0.00228 7.902 ± 0.183 % 88.752 ± 0.264 %
27
+ 15 6.6619 ± 0.1348 0.02133 ± 0.00373 0.07730 ± 0.00214 7.842 ± 0.174 % 88.693 ± 0.256 %
28
+ 16 6.8209 ± 0.1338 0.01819 ± 0.00361 0.07652 ± 0.00203 7.773 ± 0.167 % 88.618 ± 0.248 %
29
+ 17 6.8676 ± 0.1304 0.01805 ± 0.00348 0.07551 ± 0.00192 7.696 ± 0.162 % 88.609 ± 0.241 %
30
+ 18 6.9618 ± 0.1286 0.01849 ± 0.00338 0.07510 ± 0.00182 7.653 ± 0.155 % 88.552 ± 0.235 %
31
+ 19 6.9170 ± 0.1249 0.01894 ± 0.00325 0.07404 ± 0.00175 7.561 ± 0.151 % 88.661 ± 0.227 %
32
+ 20 6.6681 ± 0.1166 0.02161 ± 0.00323 0.07796 ± 0.00171 7.836 ± 0.145 % 88.548 ± 0.223 %
33
+ 21 6.6815 ± 0.1138 0.02112 ± 0.00316 0.07831 ± 0.00165 7.820 ± 0.139 % 88.428 ± 0.218 %
34
+ 22 6.6992 ± 0.1116 0.02177 ± 0.00313 0.07958 ± 0.00169 7.862 ± 0.139 % 88.394 ± 0.214 %
35
+ 23 6.7561 ± 0.1101 0.02283 ± 0.00306 0.07988 ± 0.00164 7.864 ± 0.136 % 88.270 ± 0.210 %
36
+ 24 6.7528 ± 0.1074 0.02230 ± 0.00301 0.07975 ± 0.00161 7.848 ± 0.134 % 88.262 ± 0.205 %
37
+ 25 6.7834 ± 0.1058 0.02210 ± 0.00294 0.07935 ± 0.00155 7.808 ± 0.131 % 88.223 ± 0.202 %
38
+ 26 6.7551 ± 0.1032 0.02241 ± 0.00288 0.07966 ± 0.00152 7.846 ± 0.129 % 88.270 ± 0.197 %
39
+ 27 6.9239 ± 0.1045 0.02275 ± 0.00288 0.07942 ± 0.00156 7.792 ± 0.127 % 88.244 ± 0.194 %
40
+ 28 7.0172 ± 0.1043 0.02329 ± 0.00282 0.07891 ± 0.00153 7.749 ± 0.124 % 88.259 ± 0.190 %
41
+ 29 7.0127 ± 0.1024 0.02349 ± 0.00277 0.07872 ± 0.00148 7.761 ± 0.122 % 88.263 ± 0.187 %
42
+ 30 6.9664 ± 0.0999 0.02487 ± 0.00272 0.07871 ± 0.00144 7.761 ± 0.119 % 88.263 ± 0.184 %
43
+ 31 6.8666 ± 0.0966 0.02496 ± 0.00265 0.07794 ± 0.00140 7.729 ± 0.117 % 88.377 ± 0.180 %
44
+ 32 6.7575 ± 0.0934 0.02320 ± 0.00266 0.08099 ± 0.00163 7.910 ± 0.120 % 88.313 ± 0.178 %
45
+ 33 6.6952 ± 0.0909 0.02382 ± 0.00260 0.08035 ± 0.00159 7.884 ± 0.118 % 88.371 ± 0.174 %
46
+ 34 6.6814 ± 0.0892 0.02392 ± 0.00254 0.07951 ± 0.00154 7.831 ± 0.115 % 88.416 ± 0.172 %
47
+ 35 6.6908 ± 0.0880 0.02351 ± 0.00249 0.07890 ± 0.00150 7.807 ± 0.113 % 88.465 ± 0.169 %
48
+ 36 6.7037 ± 0.0870 0.02296 ± 0.00245 0.07859 ± 0.00146 7.786 ± 0.111 % 88.498 ± 0.166 %
49
+ 37 6.6096 ± 0.0844 0.02263 ± 0.00241 0.07791 ± 0.00143 7.765 ± 0.109 % 88.492 ± 0.164 %
50
+ 38 6.5383 ± 0.0821 0.02230 ± 0.00238 0.07778 ± 0.00140 7.790 ± 0.108 % 88.514 ± 0.162 %
51
+ 39 6.4567 ± 0.0797 0.02207 ± 0.00234 0.07744 ± 0.00136 7.798 ± 0.106 % 88.508 ± 0.160 %
52
+ 40 6.3632 ± 0.0772 0.02132 ± 0.00231 0.07684 ± 0.00133 7.783 ± 0.104 % 88.556 ± 0.157 %
53
+
54
+ ====== Perplexity statistics ======
55
+ Mean PPL(Q) : 6.363210 ± 0.077202
56
+ Mean PPL(base) : 6.228979 ± 0.075322
57
+ Cor(ln(PPL(Q)), ln(PPL(base))): 98.19%
58
+ Mean ln(PPL(Q)/PPL(base)) : 0.021321 ± 0.002307
59
+ Mean PPL(Q)/PPL(base) : 1.021549 ± 0.002357
60
+ Mean PPL(Q)-PPL(base) : 0.134231 ± 0.014643
61
+
62
+ ====== KL divergence statistics ======
63
+ Mean KLD: 0.076844 ± 0.001332
64
+ Maximum KLD: 16.142012
65
+ 99.9% KLD: 2.816024
66
+ 99.0% KLD: 0.669352
67
+ 95.0% KLD: 0.255460
68
+ 90.0% KLD: 0.161866
69
+ Median KLD: 0.033844
70
+ 10.0% KLD: 0.000495
71
+ 5.0% KLD: 0.000131
72
+ 1.0% KLD: -0.000030
73
+ 0.1% KLD: -0.000287
74
+ Minimum KLD: -0.000637
75
+
76
+ ====== Token probability statistics ======
77
+ Mean Δp: -0.458 ± 0.038 %
78
+ Maximum Δp: 99.702%
79
+ 99.9% Δp: 52.654%
80
+ 99.0% Δp: 20.917%
81
+ 95.0% Δp: 9.373%
82
+ 90.0% Δp: 5.248%
83
+ 75.0% Δp: 0.940%
84
+ Median Δp: -0.010%
85
+ 25.0% Δp: -1.540%
86
+ 10.0% Δp: -6.629%
87
+ 5.0% Δp: -11.288%
88
+ 1.0% Δp: -25.973%
89
+ 0.1% Δp: -60.301%
90
+ Minimum Δp: -99.838%
91
+ RMS Δp : 7.783 ± 0.104 %
92
+ Same top p: 88.556 ± 0.157 %
93
+
recipe/logs/N4v_kld_q103.log ADDED
@@ -0,0 +1,93 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0.00.042.837 I common_init_result: fitting params to device memory ...
2
+ 0.00.042.842 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
3
+ 0.00.444.710 W llama_model_loader: direct I/O is enabled, disabling mmap
4
+ 0.01.459.450 W read_raw_unsafe: Falling back to buffered IO due to Bad address
5
+ 0.21.377.631 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
6
+ 0.21.428.701 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
7
+ 0.21.707.133 I
8
+ 0.21.707.313 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
9
+ 0.21.843.964 I kl_divergence: computing over 40 chunks, n_ctx=2048, batch_size=2048, n_seq=1
10
+ 0.24.857.762 I kl_divergence: 3.01 seconds per pass - ETA 2.00 minutes
11
+
12
+ chunk PPL ln(PPL(Q)/PPL(base)) KL Divergence Δp RMS Same top p
13
+ 1 6.1749 ± 0.4909 0.08117 ± 0.01859 0.12763 ± 0.00703 10.392 ± 0.610 % 85.337 ± 1.107 %
14
+ 2 7.1137 ± 0.3991 0.06516 ± 0.01199 0.10730 ± 0.00416 9.008 ± 0.411 % 86.266 ± 0.761 %
15
+ 3 7.5580 ± 0.3515 0.07219 ± 0.00958 0.10875 ± 0.00396 8.609 ± 0.316 % 85.989 ± 0.627 %
16
+ 4 7.8280 ± 0.3239 0.07158 ± 0.00898 0.12059 ± 0.00615 9.140 ± 0.324 % 86.193 ± 0.539 %
17
+ 5 7.6146 ± 0.2833 0.06405 ± 0.00801 0.11743 ± 0.00502 9.043 ± 0.279 % 86.549 ± 0.477 %
18
+ 6 6.5667 ± 0.2161 0.06595 ± 0.00747 0.11888 ± 0.00464 9.905 ± 0.299 % 86.804 ± 0.432 %
19
+ 7 6.0824 ± 0.1827 0.05798 ± 0.00709 0.12603 ± 0.00484 10.284 ± 0.293 % 86.929 ± 0.398 %
20
+ 8 5.9850 ± 0.1670 0.05416 ± 0.00651 0.12399 ± 0.00430 10.146 ± 0.267 % 86.877 ± 0.373 %
21
+ 9 6.2982 ± 0.1664 0.05262 ± 0.00612 0.12384 ± 0.00387 9.948 ± 0.246 % 86.445 ± 0.357 %
22
+ 10 6.4302 ± 0.1624 0.05493 ± 0.00573 0.12014 ± 0.00351 9.781 ± 0.229 % 86.452 ± 0.338 %
23
+ 11 6.4862 ± 0.1556 0.05482 ± 0.00537 0.11666 ± 0.00323 9.576 ± 0.215 % 86.564 ± 0.322 %
24
+ 12 6.7484 ± 0.1564 0.05611 ± 0.00505 0.11329 ± 0.00297 9.389 ± 0.203 % 86.470 ± 0.309 %
25
+ 13 6.7729 ± 0.1504 0.05285 ± 0.00482 0.11177 ± 0.00277 9.295 ± 0.192 % 86.525 ± 0.296 %
26
+ 14 6.8361 ± 0.1461 0.05335 ± 0.00462 0.11058 ± 0.00260 9.192 ± 0.183 % 86.385 ± 0.287 %
27
+ 15 6.8710 ± 0.1419 0.05223 ± 0.00443 0.11010 ± 0.00248 9.182 ± 0.175 % 86.373 ± 0.277 %
28
+ 16 7.0369 ± 0.1408 0.04937 ± 0.00428 0.10941 ± 0.00236 9.094 ± 0.169 % 86.278 ± 0.269 %
29
+ 17 7.0798 ± 0.1370 0.04848 ± 0.00412 0.10796 ± 0.00224 9.005 ± 0.162 % 86.361 ± 0.260 %
30
+ 18 7.1674 ± 0.1348 0.04759 ± 0.00399 0.10751 ± 0.00213 8.990 ± 0.157 % 86.201 ± 0.254 %
31
+ 19 7.1262 ± 0.1310 0.04873 ± 0.00389 0.10651 ± 0.00204 8.894 ± 0.152 % 86.227 ± 0.247 %
32
+ 20 6.8634 ± 0.1222 0.05049 ± 0.00387 0.11186 ± 0.00203 9.240 ± 0.148 % 86.202 ± 0.241 %
33
+ 21 6.8908 ± 0.1196 0.05197 ± 0.00379 0.11250 ± 0.00196 9.228 ± 0.144 % 86.175 ± 0.235 %
34
+ 22 6.9164 ± 0.1174 0.05368 ± 0.00373 0.11404 ± 0.00194 9.255 ± 0.141 % 86.164 ± 0.230 %
35
+ 23 6.9778 ± 0.1159 0.05511 ± 0.00365 0.11379 ± 0.00188 9.234 ± 0.139 % 86.128 ± 0.225 %
36
+ 24 6.9773 ± 0.1133 0.05500 ± 0.00356 0.11399 ± 0.00184 9.243 ± 0.137 % 86.038 ± 0.221 %
37
+ 25 7.0132 ± 0.1116 0.05542 ± 0.00348 0.11365 ± 0.00178 9.211 ± 0.133 % 86.037 ± 0.217 %
38
+ 26 6.9824 ± 0.1089 0.05549 ± 0.00342 0.11396 ± 0.00176 9.257 ± 0.131 % 86.070 ± 0.212 %
39
+ 27 7.1525 ± 0.1102 0.05524 ± 0.00337 0.11346 ± 0.00174 9.179 ± 0.129 % 86.072 ± 0.208 %
40
+ 28 7.2389 ± 0.1098 0.05439 ± 0.00329 0.11219 ± 0.00168 9.107 ± 0.126 % 86.084 ± 0.205 %
41
+ 29 7.2404 ± 0.1079 0.05545 ± 0.00323 0.11206 ± 0.00164 9.137 ± 0.125 % 86.048 ± 0.201 %
42
+ 30 7.1812 ± 0.1050 0.05523 ± 0.00318 0.11167 ± 0.00159 9.127 ± 0.122 % 86.022 ± 0.198 %
43
+ 31 7.0772 ± 0.1015 0.05516 ± 0.00311 0.11097 ± 0.00155 9.104 ± 0.119 % 86.151 ± 0.194 %
44
+ 32 6.9648 ± 0.0981 0.05342 ± 0.00309 0.11357 ± 0.00169 9.252 ± 0.121 % 86.092 ± 0.191 %
45
+ 33 6.9013 ± 0.0955 0.05413 ± 0.00306 0.11324 ± 0.00166 9.238 ± 0.119 % 86.087 ± 0.188 %
46
+ 34 6.8839 ± 0.0937 0.05378 ± 0.00300 0.11241 ± 0.00162 9.209 ± 0.119 % 86.093 ± 0.186 %
47
+ 35 6.8905 ± 0.0924 0.05293 ± 0.00293 0.11139 ± 0.00158 9.152 ± 0.116 % 86.178 ± 0.182 %
48
+ 36 6.9081 ± 0.0914 0.05301 ± 0.00290 0.11111 ± 0.00155 9.127 ± 0.114 % 86.190 ± 0.180 %
49
+ 37 6.8115 ± 0.0887 0.05270 ± 0.00284 0.11014 ± 0.00151 9.096 ± 0.112 % 86.238 ± 0.177 %
50
+ 38 6.7312 ± 0.0861 0.05138 ± 0.00280 0.10982 ± 0.00148 9.100 ± 0.110 % 86.299 ± 0.174 %
51
+ 39 6.6496 ± 0.0837 0.05151 ± 0.00276 0.10945 ± 0.00145 9.096 ± 0.108 % 86.312 ± 0.172 %
52
+ 40 6.5513 ± 0.0811 0.05045 ± 0.00272 0.10876 ± 0.00142 9.101 ± 0.107 % 86.356 ± 0.170 %
53
+
54
+ ====== Perplexity statistics ======
55
+ Mean PPL(Q) : 6.551305 ± 0.081055
56
+ Mean PPL(base) : 6.228979 ± 0.075322
57
+ Cor(ln(PPL(Q)), ln(PPL(base))): 97.55%
58
+ Mean ln(PPL(Q)/PPL(base)) : 0.050452 ± 0.002720
59
+ Mean PPL(Q)/PPL(base) : 1.051746 ± 0.002861
60
+ Mean PPL(Q)-PPL(base) : 0.322326 ± 0.018212
61
+
62
+ ====== KL divergence statistics ======
63
+ Mean KLD: 0.108756 ± 0.001422
64
+ Maximum KLD: 12.410495
65
+ 99.9% KLD: 3.419987
66
+ 99.0% KLD: 0.930024
67
+ 95.0% KLD: 0.366912
68
+ 90.0% KLD: 0.235489
69
+ Median KLD: 0.050197
70
+ 10.0% KLD: 0.000729
71
+ 5.0% KLD: 0.000189
72
+ 1.0% KLD: -0.000051
73
+ 0.1% KLD: -0.000341
74
+ Minimum KLD: -0.000842
75
+
76
+ ====== Token probability statistics ======
77
+ Mean Δp: -0.471 ± 0.045 %
78
+ Maximum Δp: 96.576%
79
+ 99.9% Δp: 54.353%
80
+ 99.0% Δp: 24.394%
81
+ 95.0% Δp: 11.471%
82
+ 90.0% Δp: 6.631%
83
+ 75.0% Δp: 1.298%
84
+ Median Δp: -0.006%
85
+ 25.0% Δp: -1.652%
86
+ 10.0% Δp: -7.598%
87
+ 5.0% Δp: -13.545%
88
+ 1.0% Δp: -32.141%
89
+ 0.1% Δp: -70.328%
90
+ Minimum Δp: -99.714%
91
+ RMS Δp : 9.101 ± 0.107 %
92
+ Same top p: 86.356 ± 0.170 %
93
+
recipe/logs/N4v_kld_q103i.log ADDED
@@ -0,0 +1,93 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0.00.033.949 I common_init_result: fitting params to device memory ...
2
+ 0.00.033.953 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
3
+ 0.00.371.029 W llama_model_loader: direct I/O is enabled, disabling mmap
4
+ 0.01.312.258 W read_raw_unsafe: Falling back to buffered IO due to Bad address
5
+ 0.19.504.583 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
6
+ 0.19.553.410 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
7
+ 0.19.731.487 I
8
+ 0.19.731.587 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
9
+ 0.19.847.243 I kl_divergence: computing over 40 chunks, n_ctx=2048, batch_size=2048, n_seq=1
10
+ 0.22.241.158 I kl_divergence: 2.39 seconds per pass - ETA 1.58 minutes
11
+
12
+ chunk PPL ln(PPL(Q)/PPL(base)) KL Divergence Δp RMS Same top p
13
+ 1 5.7501 ± 0.4413 0.00989 ± 0.01572 0.10954 ± 0.00645 9.805 ± 0.583 % 85.630 ± 1.097 %
14
+ 2 6.7714 ± 0.3709 0.01585 ± 0.01070 0.09308 ± 0.00377 8.559 ± 0.392 % 86.706 ± 0.751 %
15
+ 3 7.1773 ± 0.3266 0.02051 ± 0.00847 0.09216 ± 0.00368 7.997 ± 0.296 % 86.934 ± 0.608 %
16
+ 4 7.4782 ± 0.3040 0.02587 ± 0.00786 0.09738 ± 0.00472 8.531 ± 0.336 % 87.341 ± 0.520 %
17
+ 5 7.3506 ± 0.2694 0.02877 ± 0.00704 0.09557 ± 0.00398 8.489 ± 0.308 % 87.468 ± 0.463 %
18
+ 6 6.3429 ± 0.2051 0.03128 ± 0.00664 0.09662 ± 0.00373 9.145 ± 0.308 % 87.781 ± 0.418 %
19
+ 7 5.9472 ± 0.1761 0.03550 ± 0.00668 0.10791 ± 0.00462 9.889 ± 0.320 % 87.795 ± 0.387 %
20
+ 8 5.8786 ± 0.1618 0.03622 ± 0.00612 0.10626 ± 0.00411 9.788 ± 0.292 % 87.683 ± 0.363 %
21
+ 9 6.1674 ± 0.1605 0.03164 ± 0.00571 0.10586 ± 0.00370 9.627 ± 0.271 % 87.227 ± 0.348 %
22
+ 10 6.2930 ± 0.1566 0.03336 ± 0.00529 0.10181 ± 0.00335 9.388 ± 0.252 % 87.390 ± 0.328 %
23
+ 11 6.3574 ± 0.1505 0.03475 ± 0.00499 0.09913 ± 0.00308 9.301 ± 0.238 % 87.292 ± 0.314 %
24
+ 12 6.6087 ± 0.1511 0.03519 ± 0.00469 0.09595 ± 0.00284 9.076 ± 0.225 % 87.317 ± 0.300 %
25
+ 13 6.6444 ± 0.1456 0.03370 ± 0.00446 0.09418 ± 0.00263 8.998 ± 0.213 % 87.367 ± 0.288 %
26
+ 14 6.6988 ± 0.1413 0.03307 ± 0.00425 0.09228 ± 0.00246 8.826 ± 0.202 % 87.467 ± 0.277 %
27
+ 15 6.7357 ± 0.1373 0.03235 ± 0.00408 0.09142 ± 0.00231 8.745 ± 0.192 % 87.488 ± 0.267 %
28
+ 16 6.8972 ± 0.1363 0.02932 ± 0.00393 0.09032 ± 0.00219 8.621 ± 0.183 % 87.390 ± 0.259 %
29
+ 17 6.9396 ± 0.1326 0.02848 ± 0.00379 0.08911 ± 0.00207 8.521 ± 0.175 % 87.436 ± 0.251 %
30
+ 18 7.0315 ± 0.1308 0.02845 ± 0.00366 0.08879 ± 0.00197 8.486 ± 0.168 % 87.374 ± 0.245 %
31
+ 19 6.9777 ± 0.1268 0.02767 ± 0.00353 0.08759 ± 0.00188 8.384 ± 0.163 % 87.385 ± 0.238 %
32
+ 20 6.7038 ± 0.1178 0.02695 ± 0.00352 0.09228 ± 0.00186 8.685 ± 0.156 % 87.273 ± 0.233 %
33
+ 21 6.7149 ± 0.1149 0.02612 ± 0.00345 0.09269 ± 0.00179 8.668 ± 0.151 % 87.162 ± 0.228 %
34
+ 22 6.7275 ± 0.1124 0.02599 ± 0.00339 0.09343 ± 0.00179 8.678 ± 0.150 % 87.186 ± 0.223 %
35
+ 23 6.7910 ± 0.1111 0.02798 ± 0.00333 0.09354 ± 0.00176 8.670 ± 0.147 % 87.080 ± 0.219 %
36
+ 24 6.7874 ± 0.1085 0.02741 ± 0.00325 0.09330 ± 0.00171 8.652 ± 0.144 % 87.097 ± 0.214 %
37
+ 25 6.8244 ± 0.1069 0.02813 ± 0.00316 0.09247 ± 0.00164 8.585 ± 0.140 % 87.077 ± 0.210 %
38
+ 26 6.7976 ± 0.1043 0.02868 ± 0.00310 0.09292 ± 0.00161 8.642 ± 0.138 % 87.142 ± 0.205 %
39
+ 27 6.9623 ± 0.1055 0.02829 ± 0.00303 0.09202 ± 0.00157 8.538 ± 0.135 % 87.129 ± 0.201 %
40
+ 28 7.0526 ± 0.1053 0.02832 ± 0.00296 0.09127 ± 0.00153 8.496 ± 0.133 % 87.149 ± 0.198 %
41
+ 29 7.0500 ± 0.1034 0.02880 ± 0.00292 0.09137 ± 0.00149 8.557 ± 0.131 % 87.191 ± 0.194 %
42
+ 30 7.0006 ± 0.1008 0.02977 ± 0.00287 0.09152 ± 0.00145 8.567 ± 0.129 % 87.165 ± 0.191 %
43
+ 31 6.9004 ± 0.0974 0.02987 ± 0.00281 0.09083 ± 0.00141 8.530 ± 0.126 % 87.251 ± 0.187 %
44
+ 32 6.7986 ± 0.0943 0.02926 ± 0.00280 0.09330 ± 0.00152 8.688 ± 0.125 % 87.225 ± 0.184 %
45
+ 33 6.7336 ± 0.0917 0.02953 ± 0.00275 0.09274 ± 0.00149 8.677 ± 0.123 % 87.266 ± 0.181 %
46
+ 34 6.7179 ± 0.0899 0.02937 ± 0.00268 0.09177 ± 0.00145 8.617 ± 0.121 % 87.333 ± 0.178 %
47
+ 35 6.7270 ± 0.0888 0.02891 ± 0.00263 0.09111 ± 0.00141 8.570 ± 0.118 % 87.398 ± 0.175 %
48
+ 36 6.7405 ± 0.0878 0.02844 ± 0.00259 0.09075 ± 0.00138 8.522 ± 0.116 % 87.414 ± 0.173 %
49
+ 37 6.6432 ± 0.0851 0.02769 ± 0.00255 0.08997 ± 0.00135 8.495 ± 0.114 % 87.409 ± 0.171 %
50
+ 38 6.5666 ± 0.0827 0.02662 ± 0.00252 0.09001 ± 0.00132 8.514 ± 0.112 % 87.434 ± 0.168 %
51
+ 39 6.4867 ± 0.0803 0.02671 ± 0.00249 0.08987 ± 0.00130 8.528 ± 0.111 % 87.408 ± 0.166 %
52
+ 40 6.3929 ± 0.0779 0.02598 ± 0.00245 0.08915 ± 0.00127 8.509 ± 0.109 % 87.424 ± 0.164 %
53
+
54
+ ====== Perplexity statistics ======
55
+ Mean PPL(Q) : 6.392929 ± 0.077851
56
+ Mean PPL(base) : 6.228979 ± 0.075322
57
+ Cor(ln(PPL(Q)), ln(PPL(base))): 97.97%
58
+ Mean ln(PPL(Q)/PPL(base)) : 0.025980 ± 0.002447
59
+ Mean PPL(Q)/PPL(base) : 1.026321 ± 0.002511
60
+ Mean PPL(Q)-PPL(base) : 0.163950 ± 0.015638
61
+
62
+ ====== KL divergence statistics ======
63
+ Mean KLD: 0.089150 ± 0.001266
64
+ Maximum KLD: 13.245127
65
+ 99.9% KLD: 3.279078
66
+ 99.0% KLD: 0.784911
67
+ 95.0% KLD: 0.302743
68
+ 90.0% KLD: 0.190118
69
+ Median KLD: 0.039583
70
+ 10.0% KLD: 0.000539
71
+ 5.0% KLD: 0.000127
72
+ 1.0% KLD: -0.000056
73
+ 0.1% KLD: -0.000311
74
+ Minimum KLD: -0.000916
75
+
76
+ ====== Token probability statistics ======
77
+ Mean Δp: -0.437 ± 0.042 %
78
+ Maximum Δp: 97.638%
79
+ 99.9% Δp: 52.324%
80
+ 99.0% Δp: 22.989%
81
+ 95.0% Δp: 10.252%
82
+ 90.0% Δp: 5.909%
83
+ 75.0% Δp: 1.135%
84
+ Median Δp: -0.001%
85
+ 25.0% Δp: -1.497%
86
+ 10.0% Δp: -6.935%
87
+ 5.0% Δp: -12.278%
88
+ 1.0% Δp: -29.036%
89
+ 0.1% Δp: -71.577%
90
+ Minimum Δp: -99.874%
91
+ RMS Δp : 8.509 ± 0.109 %
92
+ Same top p: 87.424 ± 0.164 %
93
+
recipe/logs/N4v_kld_q106.log ADDED
@@ -0,0 +1,93 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0.00.073.166 I common_init_result: fitting params to device memory ...
2
+ 0.00.073.169 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
3
+ 0.00.618.000 W llama_model_loader: direct I/O is enabled, disabling mmap
4
+ 0.01.632.328 W read_raw_unsafe: Falling back to buffered IO due to Bad address
5
+ 0.21.360.980 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
6
+ 0.21.411.524 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
7
+ 0.21.696.824 I
8
+ 0.21.696.976 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
9
+ 0.21.834.422 I kl_divergence: computing over 40 chunks, n_ctx=2048, batch_size=2048, n_seq=1
10
+ 0.25.227.075 I kl_divergence: 3.39 seconds per pass - ETA 2.25 minutes
11
+
12
+ chunk PPL ln(PPL(Q)/PPL(base)) KL Divergence Δp RMS Same top p
13
+ 1 6.1880 ± 0.4939 0.08329 ± 0.01978 0.13233 ± 0.00748 10.697 ± 0.585 % 86.022 ± 1.085 %
14
+ 2 7.0357 ± 0.3926 0.05414 ± 0.01246 0.10961 ± 0.00433 9.075 ± 0.394 % 86.559 ± 0.754 %
15
+ 3 7.4624 ± 0.3452 0.05947 ± 0.01021 0.11204 ± 0.00434 9.072 ± 0.327 % 86.250 ± 0.622 %
16
+ 4 7.7345 ± 0.3193 0.05957 ± 0.00891 0.11467 ± 0.00464 9.213 ± 0.314 % 86.486 ± 0.535 %
17
+ 5 7.5351 ± 0.2801 0.05356 ± 0.00785 0.11152 ± 0.00383 8.995 ± 0.270 % 86.667 ± 0.475 %
18
+ 6 6.4595 ± 0.2123 0.04949 ± 0.00714 0.11027 ± 0.00353 9.495 ± 0.276 % 87.113 ± 0.428 %
19
+ 7 5.9967 ± 0.1800 0.04378 ± 0.00691 0.12283 ± 0.00457 10.169 ± 0.285 % 87.250 ± 0.394 %
20
+ 8 5.9141 ± 0.1649 0.04224 ± 0.00637 0.12112 ± 0.00408 10.053 ± 0.261 % 87.170 ± 0.370 %
21
+ 9 6.2093 ± 0.1636 0.03840 ± 0.00600 0.12020 ± 0.00367 9.847 ± 0.240 % 86.825 ± 0.352 %
22
+ 10 6.3404 ± 0.1596 0.04087 ± 0.00561 0.11641 ± 0.00333 9.682 ± 0.223 % 86.823 ± 0.334 %
23
+ 11 6.4032 ± 0.1532 0.04194 ± 0.00528 0.11306 ± 0.00305 9.498 ± 0.210 % 86.795 ± 0.319 %
24
+ 12 6.6610 ± 0.1540 0.04308 ± 0.00496 0.10959 ± 0.00281 9.281 ± 0.198 % 86.763 ± 0.306 %
25
+ 13 6.6960 ± 0.1485 0.04143 ± 0.00472 0.10783 ± 0.00262 9.199 ± 0.188 % 86.773 ± 0.294 %
26
+ 14 6.7584 ± 0.1443 0.04192 ± 0.00452 0.10607 ± 0.00245 9.074 ± 0.179 % 86.657 ± 0.284 %
27
+ 15 6.8026 ± 0.1405 0.04223 ± 0.00435 0.10576 ± 0.00235 9.049 ± 0.171 % 86.660 ± 0.274 %
28
+ 16 6.9572 ± 0.1392 0.03797 ± 0.00421 0.10524 ± 0.00223 8.990 ± 0.164 % 86.571 ± 0.267 %
29
+ 17 7.0033 ± 0.1354 0.03762 ± 0.00406 0.10400 ± 0.00212 8.906 ± 0.157 % 86.522 ± 0.259 %
30
+ 18 7.0947 ± 0.1334 0.03740 ± 0.00393 0.10350 ± 0.00202 8.881 ± 0.152 % 86.532 ± 0.252 %
31
+ 19 7.0450 ± 0.1294 0.03728 ± 0.00382 0.10187 ± 0.00193 8.789 ± 0.147 % 86.639 ± 0.244 %
32
+ 20 6.7808 ± 0.1206 0.03838 ± 0.00381 0.10695 ± 0.00193 9.167 ± 0.147 % 86.569 ± 0.238 %
33
+ 21 6.8071 ± 0.1180 0.03975 ± 0.00374 0.10769 ± 0.00187 9.179 ± 0.142 % 86.506 ± 0.233 %
34
+ 22 6.8328 ± 0.1159 0.04152 ± 0.00369 0.10920 ± 0.00187 9.217 ± 0.139 % 86.501 ± 0.228 %
35
+ 23 6.8923 ± 0.1144 0.04279 ± 0.00360 0.10882 ± 0.00181 9.197 ± 0.137 % 86.498 ± 0.223 %
36
+ 24 6.8910 ± 0.1117 0.04255 ± 0.00351 0.10896 ± 0.00176 9.191 ± 0.134 % 86.490 ± 0.218 %
37
+ 25 6.9237 ± 0.1100 0.04257 ± 0.00344 0.10863 ± 0.00170 9.157 ± 0.130 % 86.424 ± 0.214 %
38
+ 26 6.8938 ± 0.1073 0.04273 ± 0.00338 0.10952 ± 0.00170 9.229 ± 0.130 % 86.435 ± 0.210 %
39
+ 27 7.0651 ± 0.1086 0.04294 ± 0.00331 0.10883 ± 0.00167 9.149 ± 0.127 % 86.452 ± 0.206 %
40
+ 28 7.1473 ± 0.1082 0.04166 ± 0.00323 0.10759 ± 0.00162 9.083 ± 0.125 % 86.514 ± 0.202 %
41
+ 29 7.1505 ± 0.1064 0.04295 ± 0.00318 0.10739 ± 0.00158 9.106 ± 0.123 % 86.517 ± 0.198 %
42
+ 30 7.0929 ± 0.1035 0.04287 ± 0.00312 0.10706 ± 0.00153 9.094 ± 0.121 % 86.452 ± 0.195 %
43
+ 31 6.9865 ± 0.1000 0.04226 ± 0.00306 0.10620 ± 0.00149 9.057 ± 0.118 % 86.573 ± 0.191 %
44
+ 32 6.8731 ± 0.0966 0.04016 ± 0.00304 0.10897 ± 0.00167 9.204 ± 0.119 % 86.550 ± 0.189 %
45
+ 33 6.8112 ± 0.0941 0.04100 ± 0.00300 0.10867 ± 0.00163 9.201 ± 0.117 % 86.561 ± 0.186 %
46
+ 34 6.7925 ± 0.0922 0.04042 ± 0.00293 0.10760 ± 0.00159 9.133 ± 0.115 % 86.594 ± 0.183 %
47
+ 35 6.8061 ± 0.0911 0.04061 ± 0.00287 0.10675 ± 0.00155 9.093 ± 0.113 % 86.655 ± 0.180 %
48
+ 36 6.8188 ± 0.0901 0.03999 ± 0.00283 0.10633 ± 0.00151 9.056 ± 0.111 % 86.717 ± 0.177 %
49
+ 37 6.7235 ± 0.0873 0.03971 ± 0.00278 0.10555 ± 0.00148 9.021 ± 0.109 % 86.753 ± 0.174 %
50
+ 38 6.6468 ± 0.0849 0.03876 ± 0.00274 0.10510 ± 0.00145 9.019 ± 0.107 % 86.762 ± 0.172 %
51
+ 39 6.5673 ± 0.0825 0.03906 ± 0.00270 0.10487 ± 0.00142 9.014 ± 0.105 % 86.773 ± 0.170 %
52
+ 40 6.4745 ± 0.0799 0.03865 ± 0.00266 0.10441 ± 0.00139 9.030 ± 0.104 % 86.794 ± 0.167 %
53
+
54
+ ====== Perplexity statistics ======
55
+ Mean PPL(Q) : 6.474466 ± 0.079947
56
+ Mean PPL(base) : 6.228979 ± 0.075322
57
+ Cor(ln(PPL(Q)), ln(PPL(base))): 97.64%
58
+ Mean ln(PPL(Q)/PPL(base)) : 0.038654 ± 0.002665
59
+ Mean PPL(Q)/PPL(base) : 1.039410 ± 0.002770
60
+ Mean PPL(Q)-PPL(base) : 0.245487 ± 0.017468
61
+
62
+ ====== KL divergence statistics ======
63
+ Mean KLD: 0.104408 ± 0.001395
64
+ Maximum KLD: 16.437456
65
+ 99.9% KLD: 3.401226
66
+ 99.0% KLD: 0.899358
67
+ 95.0% KLD: 0.354356
68
+ 90.0% KLD: 0.226252
69
+ Median KLD: 0.048009
70
+ 10.0% KLD: 0.000697
71
+ 5.0% KLD: 0.000180
72
+ 1.0% KLD: -0.000044
73
+ 0.1% KLD: -0.000336
74
+ Minimum KLD: -0.000813
75
+
76
+ ====== Token probability statistics ======
77
+ Mean Δp: -0.202 ± 0.045 %
78
+ Maximum Δp: 99.758%
79
+ 99.9% Δp: 56.495%
80
+ 99.0% Δp: 25.703%
81
+ 95.0% Δp: 11.848%
82
+ 90.0% Δp: 6.930%
83
+ 75.0% Δp: 1.471%
84
+ Median Δp: -0.001%
85
+ 25.0% Δp: -1.409%
86
+ 10.0% Δp: -7.175%
87
+ 5.0% Δp: -13.010%
88
+ 1.0% Δp: -31.209%
89
+ 0.1% Δp: -69.269%
90
+ Minimum Δp: -98.138%
91
+ RMS Δp : 9.030 ± 0.104 %
92
+ Same top p: 86.794 ± 0.167 %
93
+
recipe/logs/N4v_kld_q106i.log ADDED
@@ -0,0 +1,93 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0.00.033.687 I common_init_result: fitting params to device memory ...
2
+ 0.00.033.690 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
3
+ 0.00.365.970 W llama_model_loader: direct I/O is enabled, disabling mmap
4
+ 0.01.319.923 W read_raw_unsafe: Falling back to buffered IO due to Bad address
5
+ 0.19.602.693 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
6
+ 0.19.656.222 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
7
+ 0.19.834.989 I
8
+ 0.19.835.081 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
9
+ 0.19.951.235 I kl_divergence: computing over 40 chunks, n_ctx=2048, batch_size=2048, n_seq=1
10
+ 0.22.342.997 I kl_divergence: 2.39 seconds per pass - ETA 1.58 minutes
11
+
12
+ chunk PPL ln(PPL(Q)/PPL(base)) KL Divergence Δp RMS Same top p
13
+ 1 5.6983 ± 0.4331 0.00085 ± 0.01708 0.10779 ± 0.00685 10.030 ± 0.635 % 86.217 ± 1.078 %
14
+ 2 6.7307 ± 0.3660 0.00982 ± 0.01106 0.08982 ± 0.00389 8.657 ± 0.431 % 86.950 ± 0.745 %
15
+ 3 7.1804 ± 0.3254 0.02094 ± 0.00857 0.08982 ± 0.00369 8.264 ± 0.335 % 87.260 ± 0.602 %
16
+ 4 7.4637 ± 0.3024 0.02393 ± 0.00766 0.09566 ± 0.00581 8.559 ± 0.352 % 87.732 ± 0.513 %
17
+ 5 7.3233 ± 0.2674 0.02505 ± 0.00682 0.09342 ± 0.00473 8.497 ± 0.302 % 87.977 ± 0.455 %
18
+ 6 6.3148 ± 0.2036 0.02684 ± 0.00631 0.09237 ± 0.00410 8.998 ± 0.290 % 88.368 ± 0.409 %
19
+ 7 5.8753 ± 0.1731 0.02333 ± 0.00607 0.09805 ± 0.00414 9.342 ± 0.282 % 88.479 ± 0.377 %
20
+ 8 5.8033 ± 0.1588 0.02333 ± 0.00556 0.09655 ± 0.00369 9.256 ± 0.258 % 88.331 ± 0.355 %
21
+ 9 6.1048 ± 0.1583 0.02143 ± 0.00523 0.09704 ± 0.00332 9.085 ± 0.238 % 87.846 ± 0.341 %
22
+ 10 6.2216 ± 0.1541 0.02196 ± 0.00487 0.09358 ± 0.00301 8.857 ± 0.221 % 87.869 ± 0.323 %
23
+ 11 6.2877 ± 0.1481 0.02373 ± 0.00459 0.09124 ± 0.00277 8.771 ± 0.212 % 87.914 ± 0.307 %
24
+ 12 6.5381 ± 0.1489 0.02446 ± 0.00432 0.08825 ± 0.00255 8.571 ± 0.200 % 87.854 ± 0.295 %
25
+ 13 6.5779 ± 0.1436 0.02363 ± 0.00412 0.08706 ± 0.00237 8.553 ± 0.189 % 87.796 ± 0.284 %
26
+ 14 6.6420 ± 0.1396 0.02456 ± 0.00395 0.08566 ± 0.00221 8.419 ± 0.180 % 87.767 ± 0.274 %
27
+ 15 6.6847 ± 0.1359 0.02475 ± 0.00380 0.08492 ± 0.00208 8.367 ± 0.171 % 87.683 ± 0.265 %
28
+ 16 6.8394 ± 0.1348 0.02090 ± 0.00368 0.08419 ± 0.00198 8.280 ± 0.165 % 87.616 ± 0.257 %
29
+ 17 6.8859 ± 0.1312 0.02070 ± 0.00355 0.08311 ± 0.00187 8.185 ± 0.158 % 87.695 ± 0.249 %
30
+ 18 6.9788 ± 0.1294 0.02093 ± 0.00343 0.08257 ± 0.00178 8.136 ± 0.151 % 87.640 ± 0.243 %
31
+ 19 6.9229 ± 0.1253 0.01979 ± 0.00332 0.08171 ± 0.00170 8.036 ± 0.146 % 87.647 ± 0.236 %
32
+ 20 6.6565 ± 0.1166 0.01988 ± 0.00330 0.08616 ± 0.00169 8.356 ± 0.141 % 87.542 ± 0.231 %
33
+ 21 6.6723 ± 0.1139 0.01975 ± 0.00323 0.08644 ± 0.00162 8.308 ± 0.137 % 87.488 ± 0.226 %
34
+ 22 6.6896 ± 0.1116 0.02034 ± 0.00316 0.08711 ± 0.00160 8.355 ± 0.137 % 87.514 ± 0.220 %
35
+ 23 6.7506 ± 0.1102 0.02201 ± 0.00309 0.08713 ± 0.00156 8.332 ± 0.134 % 87.479 ± 0.216 %
36
+ 24 6.7515 ± 0.1077 0.02210 ± 0.00303 0.08701 ± 0.00153 8.313 ± 0.132 % 87.553 ± 0.211 %
37
+ 25 6.7898 ± 0.1062 0.02305 ± 0.00295 0.08650 ± 0.00148 8.258 ± 0.129 % 87.464 ± 0.207 %
38
+ 26 6.7594 ± 0.1035 0.02304 ± 0.00290 0.08694 ± 0.00146 8.287 ± 0.127 % 87.540 ± 0.203 %
39
+ 27 6.9231 ± 0.1046 0.02264 ± 0.00283 0.08601 ± 0.00141 8.197 ± 0.124 % 87.553 ± 0.199 %
40
+ 28 7.0130 ± 0.1044 0.02268 ± 0.00277 0.08539 ± 0.00138 8.177 ± 0.123 % 87.586 ± 0.195 %
41
+ 29 7.0144 ± 0.1026 0.02374 ± 0.00273 0.08551 ± 0.00135 8.241 ± 0.122 % 87.569 ± 0.192 %
42
+ 30 6.9699 ± 0.1001 0.02537 ± 0.00270 0.08585 ± 0.00132 8.254 ± 0.120 % 87.563 ± 0.188 %
43
+ 31 6.8694 ± 0.0967 0.02536 ± 0.00264 0.08520 ± 0.00128 8.219 ± 0.117 % 87.668 ± 0.185 %
44
+ 32 6.7608 ± 0.0936 0.02369 ± 0.00262 0.08761 ± 0.00146 8.350 ± 0.118 % 87.662 ± 0.182 %
45
+ 33 6.6975 ± 0.0910 0.02416 ± 0.00258 0.08714 ± 0.00143 8.347 ± 0.117 % 87.704 ± 0.179 %
46
+ 34 6.6852 ± 0.0894 0.02449 ± 0.00252 0.08623 ± 0.00139 8.290 ± 0.114 % 87.775 ± 0.176 %
47
+ 35 6.6941 ± 0.0882 0.02401 ± 0.00247 0.08558 ± 0.00135 8.241 ± 0.112 % 87.845 ± 0.173 %
48
+ 36 6.7094 ± 0.0873 0.02381 ± 0.00243 0.08528 ± 0.00132 8.207 ± 0.110 % 87.846 ± 0.170 %
49
+ 37 6.6134 ± 0.0846 0.02319 ± 0.00240 0.08448 ± 0.00129 8.182 ± 0.108 % 87.852 ± 0.168 %
50
+ 38 6.5369 ± 0.0822 0.02208 ± 0.00237 0.08434 ± 0.00126 8.180 ± 0.106 % 87.897 ± 0.165 %
51
+ 39 6.4545 ± 0.0798 0.02173 ± 0.00234 0.08418 ± 0.00123 8.209 ± 0.105 % 87.896 ± 0.163 %
52
+ 40 6.3642 ± 0.0773 0.02147 ± 0.00231 0.08356 ± 0.00121 8.201 ± 0.103 % 87.903 ± 0.161 %
53
+
54
+ ====== Perplexity statistics ======
55
+ Mean PPL(Q) : 6.364168 ± 0.077334
56
+ Mean PPL(base) : 6.228979 ± 0.075322
57
+ Cor(ln(PPL(Q)), ln(PPL(base))): 98.19%
58
+ Mean ln(PPL(Q)/PPL(base)) : 0.021471 ± 0.002306
59
+ Mean PPL(Q)/PPL(base) : 1.021703 ± 0.002356
60
+ Mean PPL(Q)-PPL(base) : 0.135189 ± 0.014650
61
+
62
+ ====== KL divergence statistics ======
63
+ Mean KLD: 0.083563 ± 0.001206
64
+ Maximum KLD: 15.989422
65
+ 99.9% KLD: 2.686481
66
+ 99.0% KLD: 0.713036
67
+ 95.0% KLD: 0.286778
68
+ 90.0% KLD: 0.180925
69
+ Median KLD: 0.038000
70
+ 10.0% KLD: 0.000531
71
+ 5.0% KLD: 0.000125
72
+ 1.0% KLD: -0.000041
73
+ 0.1% KLD: -0.000270
74
+ Minimum KLD: -0.000789
75
+
76
+ ====== Token probability statistics ======
77
+ Mean Δp: -0.435 ± 0.040 %
78
+ Maximum Δp: 99.469%
79
+ 99.9% Δp: 52.136%
80
+ 99.0% Δp: 22.860%
81
+ 95.0% Δp: 9.882%
82
+ 90.0% Δp: 5.626%
83
+ 75.0% Δp: 1.034%
84
+ Median Δp: -0.004%
85
+ 25.0% Δp: -1.525%
86
+ 10.0% Δp: -6.894%
87
+ 5.0% Δp: -11.769%
88
+ 1.0% Δp: -27.478%
89
+ 0.1% Δp: -66.247%
90
+ Minimum Δp: -99.814%
91
+ RMS Δp : 8.201 ± 0.103 %
92
+ Same top p: 87.903 ± 0.161 %
93
+
recipe/logs/N5_kld_q106_repeat.log ADDED
@@ -0,0 +1,92 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0.00.038.287 I common_init_result: fitting params to device memory ...
2
+ 0.00.038.289 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
3
+ 0.00.399.780 W llama_model_loader: direct I/O is enabled, disabling mmap
4
+ 0.20.682.112 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
5
+ 0.20.729.253 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
6
+ 0.20.921.065 I
7
+ 0.20.921.189 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
8
+ 0.21.035.122 I kl_divergence: computing over 40 chunks, n_ctx=2048, batch_size=2048, n_seq=1
9
+ 0.22.991.105 I kl_divergence: 1.96 seconds per pass - ETA 1.30 minutes
10
+
11
+ chunk PPL ln(PPL(Q)/PPL(base)) KL Divergence Δp RMS Same top p
12
+ 1 6.2617 ± 0.5005 0.09513 ± 0.01958 0.13331 ± 0.00723 11.273 ± 0.670 % 85.728 ± 1.094 %
13
+ 2 7.1013 ± 0.3976 0.06343 ± 0.01246 0.11005 ± 0.00427 9.422 ± 0.440 % 86.413 ± 0.758 %
14
+ 3 7.5058 ± 0.3468 0.06527 ± 0.01057 0.11313 ± 0.00423 9.250 ± 0.346 % 86.022 ± 0.626 %
15
+ 4 7.7849 ± 0.3204 0.06605 ± 0.00908 0.11249 ± 0.00383 9.264 ± 0.306 % 86.388 ± 0.536 %
16
+ 5 7.5801 ± 0.2810 0.05952 ± 0.00798 0.11043 ± 0.00322 9.087 ± 0.266 % 86.725 ± 0.474 %
17
+ 6 6.5014 ± 0.2131 0.05596 ± 0.00724 0.11026 ± 0.00320 9.619 ± 0.279 % 87.178 ± 0.427 %
18
+ 7 6.0313 ± 0.1806 0.04953 ± 0.00696 0.12172 ± 0.00436 10.069 ± 0.275 % 87.236 ± 0.394 %
19
+ 8 5.9472 ± 0.1655 0.04783 ± 0.00647 0.12058 ± 0.00391 9.949 ± 0.252 % 87.121 ± 0.370 %
20
+ 9 6.2446 ± 0.1643 0.04408 ± 0.00605 0.12012 ± 0.00353 9.791 ± 0.233 % 86.597 ± 0.355 %
21
+ 10 6.3764 ± 0.1603 0.04654 ± 0.00565 0.11624 ± 0.00321 9.609 ± 0.217 % 86.588 ± 0.337 %
22
+ 11 6.4374 ± 0.1538 0.04726 ± 0.00530 0.11268 ± 0.00294 9.442 ± 0.204 % 86.484 ± 0.322 %
23
+ 12 6.6950 ± 0.1547 0.04816 ± 0.00499 0.10924 ± 0.00271 9.238 ± 0.193 % 86.388 ± 0.310 %
24
+ 13 6.7262 ± 0.1490 0.04592 ± 0.00474 0.10773 ± 0.00253 9.174 ± 0.183 % 86.413 ± 0.297 %
25
+ 14 6.7857 ± 0.1447 0.04595 ± 0.00454 0.10583 ± 0.00237 9.050 ± 0.174 % 86.426 ± 0.286 %
26
+ 15 6.8279 ± 0.1408 0.04594 ± 0.00437 0.10522 ± 0.00224 9.021 ± 0.166 % 86.419 ± 0.277 %
27
+ 16 6.9800 ± 0.1394 0.04125 ± 0.00424 0.10488 ± 0.00213 8.960 ± 0.160 % 86.339 ± 0.268 %
28
+ 17 7.0246 ± 0.1356 0.04064 ± 0.00408 0.10348 ± 0.00202 8.863 ± 0.153 % 86.424 ± 0.260 %
29
+ 18 7.1128 ± 0.1335 0.03994 ± 0.00395 0.10297 ± 0.00192 8.819 ± 0.147 % 86.385 ± 0.253 %
30
+ 19 7.0597 ± 0.1294 0.03935 ± 0.00383 0.10135 ± 0.00184 8.726 ± 0.143 % 86.474 ± 0.245 %
31
+ 20 6.7990 ± 0.1207 0.04105 ± 0.00383 0.10693 ± 0.00185 9.139 ± 0.143 % 86.393 ± 0.240 %
32
+ 21 6.8267 ± 0.1182 0.04262 ± 0.00375 0.10765 ± 0.00179 9.144 ± 0.139 % 86.375 ± 0.234 %
33
+ 22 6.8456 ± 0.1159 0.04339 ± 0.00369 0.10893 ± 0.00181 9.184 ± 0.137 % 86.417 ± 0.228 %
34
+ 23 6.9000 ± 0.1143 0.04390 ± 0.00360 0.10839 ± 0.00175 9.147 ± 0.133 % 86.383 ± 0.224 %
35
+ 24 6.8942 ± 0.1115 0.04302 ± 0.00351 0.10851 ± 0.00171 9.165 ± 0.132 % 86.376 ± 0.219 %
36
+ 25 6.9297 ± 0.1099 0.04345 ± 0.00344 0.10819 ± 0.00165 9.128 ± 0.128 % 86.334 ± 0.215 %
37
+ 26 6.8989 ± 0.1072 0.04347 ± 0.00338 0.10899 ± 0.00166 9.185 ± 0.127 % 86.371 ± 0.210 %
38
+ 27 7.0668 ± 0.1084 0.04318 ± 0.00330 0.10805 ± 0.00160 9.103 ± 0.124 % 86.384 ± 0.206 %
39
+ 28 7.1505 ± 0.1080 0.04210 ± 0.00321 0.10685 ± 0.00155 9.028 ± 0.122 % 86.395 ± 0.203 %
40
+ 29 7.1560 ± 0.1063 0.04372 ± 0.00317 0.10671 ± 0.00152 9.048 ± 0.120 % 86.392 ± 0.199 %
41
+ 30 7.0985 ± 0.1034 0.04365 ± 0.00312 0.10649 ± 0.00148 9.044 ± 0.118 % 86.341 ± 0.196 %
42
+ 31 6.9900 ± 0.0998 0.04276 ± 0.00305 0.10563 ± 0.00144 9.018 ± 0.116 % 86.444 ± 0.192 %
43
+ 32 6.8747 ± 0.0965 0.04039 ± 0.00305 0.10917 ± 0.00170 9.201 ± 0.119 % 86.416 ± 0.189 %
44
+ 33 6.8130 ± 0.0939 0.04125 ± 0.00301 0.10876 ± 0.00166 9.199 ± 0.117 % 86.457 ± 0.186 %
45
+ 34 6.7927 ± 0.0920 0.04044 ± 0.00294 0.10768 ± 0.00161 9.128 ± 0.115 % 86.441 ± 0.184 %
46
+ 35 6.8056 ± 0.0909 0.04054 ± 0.00288 0.10692 ± 0.00157 9.083 ± 0.113 % 86.494 ± 0.181 %
47
+ 36 6.8222 ± 0.0899 0.04049 ± 0.00284 0.10659 ± 0.00154 9.060 ± 0.111 % 86.546 ± 0.178 %
48
+ 37 6.7263 ± 0.0872 0.04013 ± 0.00279 0.10571 ± 0.00150 9.033 ± 0.109 % 86.566 ± 0.175 %
49
+ 38 6.6472 ± 0.0847 0.03882 ± 0.00275 0.10525 ± 0.00147 9.020 ± 0.107 % 86.621 ± 0.173 %
50
+ 39 6.5675 ± 0.0823 0.03909 ± 0.00271 0.10501 ± 0.00144 9.022 ± 0.106 % 86.638 ± 0.170 %
51
+ 40 6.4740 ± 0.0798 0.03858 ± 0.00267 0.10444 ± 0.00141 9.028 ± 0.105 % 86.659 ± 0.168 %
52
+
53
+ ====== Perplexity statistics ======
54
+ Mean PPL(Q) : 6.473964 ± 0.079799
55
+ Mean PPL(base) : 6.228979 ± 0.075322
56
+ Cor(ln(PPL(Q)), ln(PPL(base))): 97.63%
57
+ Mean ln(PPL(Q)/PPL(base)) : 0.038576 ± 0.002667
58
+ Mean PPL(Q)/PPL(base) : 1.039330 ± 0.002772
59
+ Mean PPL(Q)-PPL(base) : 0.244985 ± 0.017453
60
+
61
+ ====== KL divergence statistics ======
62
+ Mean KLD: 0.104436 ± 0.001413
63
+ Maximum KLD: 13.909879
64
+ 99.9% KLD: 3.189192
65
+ 99.0% KLD: 0.926687
66
+ 95.0% KLD: 0.354656
67
+ 90.0% KLD: 0.225739
68
+ Median KLD: 0.047702
69
+ 10.0% KLD: 0.000698
70
+ 5.0% KLD: 0.000183
71
+ 1.0% KLD: -0.000033
72
+ 0.1% KLD: -0.000321
73
+ Minimum KLD: -0.000718
74
+
75
+ ====== Token probability statistics ======
76
+ Mean Δp: -0.244 ± 0.045 %
77
+ Maximum Δp: 99.706%
78
+ 99.9% Δp: 59.493%
79
+ 99.0% Δp: 25.100%
80
+ 95.0% Δp: 11.707%
81
+ 90.0% Δp: 6.963%
82
+ 75.0% Δp: 1.453%
83
+ Median Δp: -0.001%
84
+ 25.0% Δp: -1.479%
85
+ 10.0% Δp: -7.240%
86
+ 5.0% Δp: -12.992%
87
+ 1.0% Δp: -31.268%
88
+ 0.1% Δp: -68.467%
89
+ Minimum Δp: -98.412%
90
+ RMS Δp : 9.028 ± 0.105 %
91
+ Same top p: 86.659 ± 0.168 %
92
+
recipe/logs/N5v_kld_q106_repeat.log ADDED
@@ -0,0 +1,93 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0.00.034.429 I common_init_result: fitting params to device memory ...
2
+ 0.00.034.432 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
3
+ 0.00.372.660 W llama_model_loader: direct I/O is enabled, disabling mmap
4
+ 0.01.317.954 W read_raw_unsafe: Falling back to buffered IO due to Bad address
5
+ 0.19.648.031 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
6
+ 0.19.696.562 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
7
+ 0.19.866.239 I
8
+ 0.19.866.320 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
9
+ 0.19.981.669 I kl_divergence: computing over 40 chunks, n_ctx=2048, batch_size=2048, n_seq=1
10
+ 0.22.378.306 I kl_divergence: 2.40 seconds per pass - ETA 1.58 minutes
11
+
12
+ chunk PPL ln(PPL(Q)/PPL(base)) KL Divergence Δp RMS Same top p
13
+ 1 6.1880 ± 0.4939 0.08329 ± 0.01978 0.13233 ± 0.00748 10.697 ± 0.585 % 86.022 ± 1.085 %
14
+ 2 7.0357 ± 0.3926 0.05414 ± 0.01246 0.10961 ± 0.00433 9.075 ± 0.394 % 86.559 ± 0.754 %
15
+ 3 7.4624 ± 0.3452 0.05947 ± 0.01021 0.11204 ± 0.00434 9.072 ± 0.327 % 86.250 ± 0.622 %
16
+ 4 7.7345 ± 0.3193 0.05957 ± 0.00891 0.11467 ± 0.00464 9.213 ± 0.314 % 86.486 ± 0.535 %
17
+ 5 7.5351 ± 0.2801 0.05356 ± 0.00785 0.11152 ± 0.00383 8.995 ± 0.270 % 86.667 ± 0.475 %
18
+ 6 6.4595 ± 0.2123 0.04949 ± 0.00714 0.11027 ± 0.00353 9.495 ± 0.276 % 87.113 ± 0.428 %
19
+ 7 5.9967 ± 0.1800 0.04378 ± 0.00691 0.12283 ± 0.00457 10.169 ± 0.285 % 87.250 ± 0.394 %
20
+ 8 5.9141 ± 0.1649 0.04224 ± 0.00637 0.12112 ± 0.00408 10.053 ± 0.261 % 87.170 ± 0.370 %
21
+ 9 6.2093 ± 0.1636 0.03840 ± 0.00600 0.12020 ± 0.00367 9.847 ± 0.240 % 86.825 ± 0.352 %
22
+ 10 6.3404 ± 0.1596 0.04087 ± 0.00561 0.11641 ± 0.00333 9.682 ± 0.223 % 86.823 ± 0.334 %
23
+ 11 6.4032 ± 0.1532 0.04194 ± 0.00528 0.11306 ± 0.00305 9.498 ± 0.210 % 86.795 ± 0.319 %
24
+ 12 6.6610 ± 0.1540 0.04308 ± 0.00496 0.10959 ± 0.00281 9.281 ± 0.198 % 86.763 ± 0.306 %
25
+ 13 6.6960 ± 0.1485 0.04143 ± 0.00472 0.10783 ± 0.00262 9.199 ± 0.188 % 86.773 ± 0.294 %
26
+ 14 6.7584 ± 0.1443 0.04192 ± 0.00452 0.10607 ± 0.00245 9.074 ± 0.179 % 86.657 ± 0.284 %
27
+ 15 6.8026 ± 0.1405 0.04223 ± 0.00435 0.10576 ± 0.00235 9.049 ± 0.171 % 86.660 ± 0.274 %
28
+ 16 6.9572 ± 0.1392 0.03797 ± 0.00421 0.10524 ± 0.00223 8.990 ± 0.164 % 86.571 ± 0.267 %
29
+ 17 7.0033 ± 0.1354 0.03762 ± 0.00406 0.10400 ± 0.00212 8.906 ± 0.157 % 86.522 ± 0.259 %
30
+ 18 7.0947 ± 0.1334 0.03740 ± 0.00393 0.10350 ± 0.00202 8.881 ± 0.152 % 86.532 ± 0.252 %
31
+ 19 7.0450 ± 0.1294 0.03728 ± 0.00382 0.10187 ± 0.00193 8.789 ± 0.147 % 86.639 ± 0.244 %
32
+ 20 6.7808 ± 0.1206 0.03838 ± 0.00381 0.10695 ± 0.00193 9.167 ± 0.147 % 86.569 ± 0.238 %
33
+ 21 6.8071 ± 0.1180 0.03975 ± 0.00374 0.10769 ± 0.00187 9.179 ± 0.142 % 86.506 ± 0.233 %
34
+ 22 6.8328 ± 0.1159 0.04152 ± 0.00369 0.10920 ± 0.00187 9.217 ± 0.139 % 86.501 ± 0.228 %
35
+ 23 6.8923 ± 0.1144 0.04279 ± 0.00360 0.10882 ± 0.00181 9.197 ± 0.137 % 86.498 ± 0.223 %
36
+ 24 6.8910 ± 0.1117 0.04255 ± 0.00351 0.10896 ± 0.00176 9.191 ± 0.134 % 86.490 ± 0.218 %
37
+ 25 6.9237 ± 0.1100 0.04257 ± 0.00344 0.10863 ± 0.00170 9.157 ± 0.130 % 86.424 ± 0.214 %
38
+ 26 6.8938 ± 0.1073 0.04273 ± 0.00338 0.10952 ± 0.00170 9.229 ± 0.130 % 86.435 ± 0.210 %
39
+ 27 7.0651 ± 0.1086 0.04294 ± 0.00331 0.10883 ± 0.00167 9.149 ± 0.127 % 86.452 ± 0.206 %
40
+ 28 7.1473 ± 0.1082 0.04166 ± 0.00323 0.10759 ± 0.00162 9.083 ± 0.125 % 86.514 ± 0.202 %
41
+ 29 7.1505 ± 0.1064 0.04295 ± 0.00318 0.10739 ± 0.00158 9.106 ± 0.123 % 86.517 ± 0.198 %
42
+ 30 7.0929 ± 0.1035 0.04287 ± 0.00312 0.10706 ± 0.00153 9.094 ± 0.121 % 86.452 ± 0.195 %
43
+ 31 6.9865 ± 0.1000 0.04226 ± 0.00306 0.10620 ± 0.00149 9.057 ± 0.118 % 86.573 ± 0.191 %
44
+ 32 6.8731 ± 0.0966 0.04016 ± 0.00304 0.10897 ± 0.00167 9.204 ± 0.119 % 86.550 ± 0.189 %
45
+ 33 6.8112 ± 0.0941 0.04100 ± 0.00300 0.10867 ± 0.00163 9.201 ± 0.117 % 86.561 ± 0.186 %
46
+ 34 6.7925 ± 0.0922 0.04042 ± 0.00293 0.10760 ± 0.00159 9.133 ± 0.115 % 86.594 ± 0.183 %
47
+ 35 6.8061 ± 0.0911 0.04061 ± 0.00287 0.10675 ± 0.00155 9.093 ± 0.113 % 86.655 ± 0.180 %
48
+ 36 6.8188 ± 0.0901 0.03999 ± 0.00283 0.10633 ± 0.00151 9.056 ± 0.111 % 86.717 ± 0.177 %
49
+ 37 6.7235 ± 0.0873 0.03971 ± 0.00278 0.10555 ± 0.00148 9.021 ± 0.109 % 86.753 ± 0.174 %
50
+ 38 6.6468 ± 0.0849 0.03876 ± 0.00274 0.10510 ± 0.00145 9.019 ± 0.107 % 86.762 ± 0.172 %
51
+ 39 6.5673 ± 0.0825 0.03906 ± 0.00270 0.10487 ± 0.00142 9.014 ± 0.105 % 86.773 ± 0.170 %
52
+ 40 6.4745 ± 0.0799 0.03865 ± 0.00266 0.10441 ± 0.00139 9.030 ± 0.104 % 86.794 ± 0.167 %
53
+
54
+ ====== Perplexity statistics ======
55
+ Mean PPL(Q) : 6.474466 ± 0.079947
56
+ Mean PPL(base) : 6.228979 ± 0.075322
57
+ Cor(ln(PPL(Q)), ln(PPL(base))): 97.64%
58
+ Mean ln(PPL(Q)/PPL(base)) : 0.038654 ± 0.002665
59
+ Mean PPL(Q)/PPL(base) : 1.039410 ± 0.002770
60
+ Mean PPL(Q)-PPL(base) : 0.245487 ± 0.017468
61
+
62
+ ====== KL divergence statistics ======
63
+ Mean KLD: 0.104408 ± 0.001395
64
+ Maximum KLD: 16.437456
65
+ 99.9% KLD: 3.401226
66
+ 99.0% KLD: 0.899358
67
+ 95.0% KLD: 0.354356
68
+ 90.0% KLD: 0.226252
69
+ Median KLD: 0.048009
70
+ 10.0% KLD: 0.000697
71
+ 5.0% KLD: 0.000180
72
+ 1.0% KLD: -0.000044
73
+ 0.1% KLD: -0.000336
74
+ Minimum KLD: -0.000813
75
+
76
+ ====== Token probability statistics ======
77
+ Mean Δp: -0.202 ± 0.045 %
78
+ Maximum Δp: 99.758%
79
+ 99.9% Δp: 56.495%
80
+ 99.0% Δp: 25.703%
81
+ 95.0% Δp: 11.848%
82
+ 90.0% Δp: 6.930%
83
+ 75.0% Δp: 1.471%
84
+ Median Δp: -0.001%
85
+ 25.0% Δp: -1.409%
86
+ 10.0% Δp: -7.175%
87
+ 5.0% Δp: -13.010%
88
+ 1.0% Δp: -31.209%
89
+ 0.1% Δp: -69.269%
90
+ Minimum Δp: -98.138%
91
+ RMS Δp : 9.030 ± 0.104 %
92
+ Same top p: 86.794 ± 0.167 %
93
+
recipe/logs/N6t_tools_c1.log ADDED
@@ -0,0 +1,33 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ PASS think=True multi-arg: args={'city': 'Paris', 'unit': 'celsius'}
2
+ FAIL think=True nested-object: exception KeyError('tool_calls')
3
+ PASS think=True enum: unit=fahrenheit
4
+ PASS think=True correct-decline: content='391'
5
+ PASS think=True multi-turn: final='Tokyo is **21°C** with **clear skies**.'
6
+ PASS think=True streaming: stream args={'city': 'Rome', 'unit': 'celsius'}
7
+ PASS think=True parallel: calls=['lima', 'oslo']
8
+ PASS think=False multi-arg: args={'city': 'Paris', 'unit': 'celsius'}
9
+ PASS think=False nested-object: args={'title': 'Design review', 'when': {'date': '2026-10-02', 'time': '14:00'}, 'attendees': ['ana@x.io', 'bo@x.io']}
10
+ PASS think=False enum: unit=fahrenheit
11
+ PASS think=False correct-decline: content='391'
12
+ PASS think=False multi-turn: final='Tokyo is currently **21°C** with **clear skies**.'
13
+ PASS think=False streaming: stream args={'city': 'Rome', 'unit': 'celsius'}
14
+ PASS think=False parallel: calls=['lima', 'oslo']
15
+ {"label": "n-tools-q106-c1", "passed": 13, "total": 14, "detail": {"multi-arg|think=True": true, "nested-object|think=True": false, "enum|think=True": true, "correct-decline|think=True": true, "multi-turn|think=True": true, "streaming|think=True": true, "parallel|think=True": true, "multi-arg|think=False": true, "nested-object|think=False": true, "enum|think=False": true, "correct-decline|think=False": true, "multi-turn|think=False": true, "streaming|think=False": true, "parallel|think=False": true}}
16
+ {"label": "n-vision-q106-c1-faon", "fa": "on", "mtp": false, "expected": "red,blue,circle,square", "answer": "The image shows two shapes: a red circle on the left and a blue square on the right.", "hits": ["red", "blue", "circle", "square"], "error": null, "server_died": false, "server_log_errors": [], "result": "PASS"}
17
+ probe no-kwargs correct-decline {"content": "391", "reasoning_len": 0, "tool_calls": [], "leaks": []}
18
+ probe no-kwargs single-word {"content": "ready", "reasoning_len": 0, "tool_calls": [], "leaks": []}
19
+ probe no-kwargs multi-arg {"content": "", "reasoning_len": 0, "tool_calls": ["get_weather"], "leaks": []}
20
+ probe enable_thinking=false correct-decline {"content": "391", "reasoning_len": 0, "tool_calls": [], "leaks": []}
21
+ probe enable_thinking=false single-word {"content": "ready", "reasoning_len": 0, "tool_calls": [], "leaks": []}
22
+ probe enable_thinking=false multi-arg {"content": "", "reasoning_len": 0, "tool_calls": ["get_weather"], "leaks": []}
23
+ probe reasoning_effort=high correct-decline {"content": "We need answer directly. 391.\n</think>\n\n391", "reasoning_len": 0, "tool_calls": [], "leaks": ["</think>"]}
24
+ probe reasoning_effort=high single-word {"content": "We need need output exactly ready.\n</think>\n\nready", "reasoning_len": 0, "tool_calls": [], "leaks": ["</think>"]}
25
+ probe reasoning_effort=high multi-arg {"content": "We need need tool. Current weather Paris celsius.\n</think>\n\n", "reasoning_len": 0, "tool_calls": ["get_weather"], "leaks": ["</think>"]}
26
+ probe reasoning_effort=medium correct-decline {"content": "\n\n</think>\n\n391", "reasoning_len": 0, "tool_calls": [], "leaks": ["</think>"]}
27
+ probe reasoning_effort=medium single-word {"content": "\n\n</think>\n\nready", "reasoning_len": 0, "tool_calls": [], "leaks": ["</think>"]}
28
+ probe reasoning_effort=medium multi-arg {"content": "\n\n</think>\n\n", "reasoning_len": 0, "tool_calls": ["get_weather"], "leaks": ["</think>"]}
29
+ probe reasoning_effort=none correct-decline {"content": "391", "reasoning_len": 0, "tool_calls": [], "leaks": []}
30
+ probe reasoning_effort=none single-word {"content": "ready", "reasoning_len": 0, "tool_calls": [], "leaks": []}
31
+ probe reasoning_effort=none multi-arg {"content": "", "reasoning_len": 0, "tool_calls": ["get_weather"], "leaks": []}
32
+ NEX_TOOLS_TPL_DONE
33
+ rc=0
recipe/logs/N6t_tools_tpl.log ADDED
@@ -0,0 +1,22 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ PASS think=True multi-arg: args={'city': 'Paris', 'unit': 'celsius'}
2
+ PASS think=True nested-object: args={'title': 'Design review', 'when': {'date': '2026-10-02', 'time': '14:00'}, 'attendees': ['ana@x.io', 'bo@x.io']}
3
+ PASS think=True enum: unit=fahrenheit
4
+ PASS think=True correct-decline: content='391'
5
+ PASS think=True multi-turn: final='Tokyo is currently **21°C and clear**.'
6
+ PASS think=True streaming: stream args={'city': 'Rome', 'unit': 'celsius'}
7
+ PASS think=True parallel: calls=['lima', 'oslo']
8
+ PASS think=False multi-arg: args={'city': 'Paris', 'unit': 'celsius'}
9
+ PASS think=False nested-object: args={'title': 'Design review', 'when': {'date': '2026-10-02', 'time': '14:00'}, 'attendees': ['ana@x.io', 'bo@x.io']}
10
+ PASS think=False enum: unit=fahrenheit
11
+ PASS think=False correct-decline: content='391'
12
+ PASS think=False multi-turn: final='Tokyo is currently **21°C**, with **clear skies**.'
13
+ PASS think=False streaming: stream args={'city': 'Rome', 'unit': 'celsius'}
14
+ PASS think=False parallel: calls=['lima', 'oslo']
15
+ {"label": "n-tools-q106-tpl", "passed": 14, "total": 14, "detail": {"multi-arg|think=True": true, "nested-object|think=True": true, "enum|think=True": true, "correct-decline|think=True": true, "multi-turn|think=True": true, "streaming|think=True": true, "parallel|think=True": true, "multi-arg|think=False": true, "nested-object|think=False": true, "enum|think=False": true, "correct-decline|think=False": true, "multi-turn|think=False": true, "streaming|think=False": true, "parallel|think=False": true}}
16
+ probe no-kwargs correct-decline {"content": "391", "reasoning_len": 30, "tool_calls": [], "leaks": []}
17
+ probe no-kwargs multi-arg {"content": "", "reasoning_len": 50, "tool_calls": ["get_weather"], "leaks": []}
18
+ probe reasoning_effort=medium correct-decline {"content": "391", "reasoning_len": 0, "tool_calls": [], "leaks": []}
19
+ probe reasoning_effort=medium multi-arg {"content": "", "reasoning_len": 0, "tool_calls": ["get_weather"], "leaks": []}
20
+ probe reasoning_effort=none correct-decline {"content": "", "reasoning_len": 3, "tool_calls": [], "leaks": []}
21
+ probe reasoning_effort=none multi-arg {"content": "", "reasoning_len": 133, "tool_calls": [], "leaks": []}
22
+ NEX_TOOLS_TPL_DONE
recipe/logs/N6t_tools_tpl_medium.log ADDED
@@ -0,0 +1,31 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ FAIL think=True multi-arg: args={'city': 'Paris', 'unit': 'celsius'}
2
+ FAIL think=True nested-object: exception KeyError('tool_calls')
3
+ FAIL think=True enum: unit=fahrenheit
4
+ FAIL think=True correct-decline: content='</think>\n\n391'
5
+ FAIL think=True multi-turn: final='</think>\n\nTokyo is currently **21°C** and **clear**.'
6
+ FAIL think=True streaming: stream args={'city': 'Rome', 'unit': 'celsius'}
7
+ FAIL think=True parallel: calls=['lima', 'oslo']
8
+ PASS think=False multi-arg: args={'city': 'Paris', 'unit': 'celsius'}
9
+ PASS think=False nested-object: args={'title': 'Design review', 'when': {'date': '2026-10-02', 'time': '14:00'}, 'attendees': ['ana@x.io', 'bo@x.io']}
10
+ PASS think=False enum: unit=fahrenheit
11
+ PASS think=False correct-decline: content='391'
12
+ PASS think=False multi-turn: final='Tokyo is currently **21°C** with **clear skies**.'
13
+ PASS think=False streaming: stream args={'city': 'Rome', 'unit': 'celsius'}
14
+ PASS think=False parallel: calls=['lima', 'oslo']
15
+ {"label": "n-tools-q106-tpl-medium", "passed": 7, "total": 14, "detail": {"multi-arg|think=True": false, "nested-object|think=True": false, "enum|think=True": false, "correct-decline|think=True": false, "multi-turn|think=True": false, "streaming|think=True": false, "parallel|think=True": false, "multi-arg|think=False": true, "nested-object|think=False": true, "enum|think=False": true, "correct-decline|think=False": true, "multi-turn|think=False": true, "streaming|think=False": true, "parallel|think=False": true}}
16
+ probe no-kwargs correct-decline {"content": "</think>\n\n391", "reasoning_len": 0, "tool_calls": [], "leaks": ["</think>"]}
17
+ probe no-kwargs single-word {"content": "</think>\n\nready", "reasoning_len": 0, "tool_calls": [], "leaks": ["</think>"]}
18
+ probe no-kwargs multi-arg {"content": "</think>\n\n", "reasoning_len": 0, "tool_calls": ["get_weather"], "leaks": ["</think>"]}
19
+ probe enable_thinking=false correct-decline {"content": "391", "reasoning_len": 0, "tool_calls": [], "leaks": []}
20
+ probe enable_thinking=false single-word {"content": "ready", "reasoning_len": 0, "tool_calls": [], "leaks": []}
21
+ probe enable_thinking=false multi-arg {"content": "", "reasoning_len": 0, "tool_calls": ["get_weather"], "leaks": []}
22
+ probe reasoning_effort=high correct-decline {"content": "We need answer directly. 391.\n</think>\n\n391", "reasoning_len": 0, "tool_calls": [], "leaks": ["</think>"]}
23
+ probe reasoning_effort=high single-word {"content": "We need need output exactly ready.\n</think>\n\nready", "reasoning_len": 0, "tool_calls": [], "leaks": ["</think>"]}
24
+ probe reasoning_effort=high multi-arg {"content": "We need need tool. Current weather Paris celsius.\n</think>\n\n", "reasoning_len": 0, "tool_calls": ["get_weather"], "leaks": ["</think>"]}
25
+ probe reasoning_effort=medium correct-decline {"content": "</think>\n\n391", "reasoning_len": 0, "tool_calls": [], "leaks": ["</think>"]}
26
+ probe reasoning_effort=medium single-word {"content": "</think>\n\nready", "reasoning_len": 0, "tool_calls": [], "leaks": ["</think>"]}
27
+ probe reasoning_effort=medium multi-arg {"content": "</think>\n\n", "reasoning_len": 0, "tool_calls": ["get_weather"], "leaks": ["</think>"]}
28
+ probe reasoning_effort=none correct-decline {"content": "391", "reasoning_len": 0, "tool_calls": [], "leaks": []}
29
+ probe reasoning_effort=none single-word {"content": "ready", "reasoning_len": 0, "tool_calls": [], "leaks": []}
30
+ probe reasoning_effort=none multi-arg {"content": "", "reasoning_len": 0, "tool_calls": ["get_weather"], "leaks": []}
31
+ NEX_TOOLS_TPL_DONE
recipe/logs/N7_sizing.log ADDED
@@ -0,0 +1,5 @@
 
 
 
 
 
 
1
+ [2026-09-17T00:50:08Z] waiting for the quiet-box lock
2
+ [2026-09-17T00:50:08Z] quiet-box lock held
3
+ {"label":"strix-lean","ctx":65536,"avail_before":122.34,"footprint_loaded_gib":21.11,"footprint_after_8k_gib":21.29}
4
+ {"label":"strix-lean","ctx":262144,"avail_before":122.15,"footprint_loaded_gib":24.36,"footprint_after_8k_gib":24.52}
5
+ NEX_SIZING_DONE
recipe/logs/N8_unice.log ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ [2026-09-17T00:50:56Z] Nex FAST test seats: create + smoke test (box still iced)
2
+ [2026-09-17T00:51:38Z] nex_seats exit=0 ([2026-09-17T00:51:38Z] NEX_SEATS_DONE fail=0)
3
+ [2026-09-17T00:51:38Z] UNICE_DEFERRED: OxCoder-9B build (King 2026-09-16 22:33Z: "when that is done do this one"); its final step un-ices
recipe/logs/N8a_seats.log ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ [2026-09-17T00:50:56Z] plan: [('max1-nex-fast', 'ROCm0', 'fa on', 262144, '31000M', True), ('max1-nex-fast-imat', 'ROCm0', 'fa on', 262144, '31000M', True)]
2
+ [2026-09-17T00:50:56Z] max1-nex-fast written (ROCm0, -fa on, ctx 262144, MemoryMax 31000M) -> smoke test
3
+ {"unit": "max1-nex-fast", "port": 8097, "load_s": 25, "time": "2026-09-17T00:51:21Z", "direct_reply": "ready", "direct_tg": 92.93680297397769, "gateway_model": "nex-n2.5-mini-fast@max1", "gateway_reply": "ready", "result": "PASS"}
4
+ [2026-09-17T00:51:22Z] max1-nex-fast stopped (enabled: disabled)
5
+ [2026-09-17T00:51:27Z] max1-nex-fast-imat written (ROCm0, -fa on, ctx 262144, MemoryMax 31000M) -> smoke test
6
+ {"unit": "max1-nex-fast-imat", "port": 8098, "load_s": 5, "time": "2026-09-17T00:51:32Z", "direct_reply": "ready", "direct_tg": 94.37078280564337, "gateway_model": "nex-n2.5-mini-fast-imatrix@max1", "gateway_reply": "ready", "result": "PASS"}
7
+ [2026-09-17T00:51:33Z] max1-nex-fast-imat stopped (enabled: disabled)
8
+ [2026-09-17T00:51:38Z] NEX_SEATS_DONE fail=0
recipe/logs/N8d_seats.log ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ [2026-09-17T01:20:44Z] plan: [('max1-nex-fast', 'ROCm0', 'fa on', 262144, '31000M', True), ('max1-nex-fast-imat', 'ROCm0', 'fa on', 262144, '31000M', True)]
2
+ [2026-09-17T01:20:45Z] max1-nex-fast written (ROCm0, -fa on, ctx 262144, MemoryMax 31000M) -> smoke test
3
+ {"unit": "max1-nex-fast", "port": 8097, "load_s": 25, "time": "2026-09-17T01:21:10Z", "direct_reply": "ready", "direct_tg": 41.41386950489719, "default_reply": "ready", "default_reasoning_len": 0, "default_leak": false, "thinking_reply": "", "thinking_reasoning_len": 5, "thinking_leak": false, "gateway_model": "nex-n2.5-mini-fast@max1", "gateway_reply": "ready", "result": "PASS"}
4
+ [2026-09-17T01:21:12Z] max1-nex-fast stopped (enabled: disabled)
5
+ [2026-09-17T01:21:17Z] max1-nex-fast-imat written (ROCm0, -fa on, ctx 262144, MemoryMax 31000M) -> smoke test
6
+ {"unit": "max1-nex-fast-imat", "port": 8098, "load_s": 25, "time": "2026-09-17T01:21:43Z", "direct_reply": "ready", "direct_tg": 41.54290343352097, "default_reply": "ready", "default_reasoning_len": 0, "default_leak": false, "thinking_reply": "", "thinking_reasoning_len": 5, "thinking_leak": false, "gateway_model": "nex-n2.5-mini-fast-imatrix@max1", "gateway_reply": "ready", "result": "PASS"}
7
+ [2026-09-17T01:21:44Z] max1-nex-fast-imat stopped (enabled: disabled)
8
+ [2026-09-17T01:21:49Z] NEX_SEATS_DONE fail=0
recipe/logs/Q1_q102.log ADDED
The diff for this file is too large to render. See raw diff
 
recipe/logs/Q1_q103.log ADDED
The diff for this file is too large to render. See raw diff
 
recipe/logs/Q_readback.log ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ PASS Nex-N2.5-mini-Q4_0_ROCMFP4_STRIX_LEAN.gguf arch=qwen35moe ftype=106 tensors=733 nextn=0 output.weight=Q6_K token_embd.weight=Q5_K
2
+ PASS Nex-N2.5-mini-Q4_0_ROCMFP4_COHERENT.gguf arch=qwen35moe ftype=102 tensors=733 nextn=0 output.weight=Q6_K token_embd.weight=Q6_K
3
+ PASS Nex-N2.5-mini-Q4_0_ROCMFP4_FAST.gguf arch=qwen35moe ftype=103 tensors=733 nextn=0 output.weight=Q6_K token_embd.weight=Q4_0_ROCMFP4_FAST
recipe/logs/Q_sizes.log ADDED
@@ -0,0 +1,6 @@
 
 
 
 
 
 
 
1
+ Q1_q106 quant size = 17865.52 MiB (4.32 BPW)
2
+ Q1_q102 quant size = 18916.30 MiB (4.58 BPW)
3
+ Q1_q103 quant size = 17774.11 MiB (4.30 BPW)
4
+ N3_q106i quant size = 17865.52 MiB (4.32 BPW)
5
+ N3_q102i quant size = 18916.30 MiB (4.58 BPW)
6
+ N3_q103i quant size = 17774.11 MiB (4.30 BPW)
recipe/logs/b_n-c3-q106.log ADDED
@@ -0,0 +1,284 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0.00.041.768 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
2
+ 0.00.041.771 I device_info:
3
+ 0.00.041.819 I - ROCm0 : AMD Radeon Graphics (131072 MiB, 125259 MiB free)
4
+ 0.00.041.885 I - Vulkan0 : AMD Radeon Graphics (RADV GFX1151) (132096 MiB, 131926 MiB free)
5
+ 0.00.041.888 I - CPU : AMD RYZEN AI MAX+ 395 w/ Radeon 8060S (127438 MiB, 127438 MiB free)
6
+ 0.00.041.930 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
7
+ 0.00.041.949 I srv init: running without SSL
8
+ 0.00.041.972 I srv init: using 31 threads for HTTP server
9
+ 0.00.041.973 I srv init: the WebUI is disabled
10
+ 0.00.042.029 I srv start: binding port with default address family
11
+ 0.00.043.171 I srv main: loading model
12
+ 0.00.043.173 I srv load_model: loading model '/mnt/models/nex-n2.5-mini/out/Nex-N2.5-mini-Q4_0_ROCMFP4_STRIX_LEAN.gguf'
13
+ 0.00.076.857 W llama_model_loader: direct I/O is enabled, disabling mmap
14
+ 0.20.443.494 W llama_context: n_ctx_seq (65536) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
15
+ 0.20.605.862 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
16
+ 0.20.804.349 I srv load_model: initializing slots, n_slots = 1
17
+ 0.20.970.363 W srv load_model: speculative decoding will use checkpoints
18
+ 0.20.970.368 W common_speculative_init: no implementations specified for speculative decoding
19
+ 0.20.970.369 I slot load_model: id 0 | task -1 | new slot, n_ctx = 65536
20
+ 0.20.970.402 I srv load_model: prompt cache RAM enabled: limit_mib=8192
21
+ 0.20.970.414 I srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
22
+ 0.20.970.429 I srv init: idle slots will be saved to prompt cache upon starting a new task
23
+ 0.20.979.135 I init: chat template, example_format: '<|im_start|>system
24
+ You are a helpful assistant<|im_end|>
25
+ <|im_start|>user
26
+ Hello<|im_end|>
27
+ <|im_start|>assistant
28
+ <think>
29
+
30
+ </think>
31
+
32
+ Hi there<|im_end|>
33
+ <|im_start|>user
34
+ How are you?<|im_end|>
35
+ <|im_start|>assistant
36
+ <think>'
37
+ 0.20.985.480 I srv init: init: chat template, thinking = 1
38
+ 0.20.985.501 I srv main: model loaded
39
+ 0.20.985.504 I srv main: server is listening on http://127.0.0.1:18600
40
+ 0.20.985.506 I srv update_slots: all slots are idle
41
+ 0.23.012.828 I srv params_from_: Chat format: peg-native
42
+ 0.23.012.941 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1
43
+ 0.23.012.943 I srv get_availabl: updating prompt cache
44
+ 0.23.012.949 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000
45
+ 0.23.012.953 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 65536 tokens, 8589934592 est)
46
+ 0.23.012.954 I srv get_availabl: prompt cache update took 0.01 ms
47
+ 0.23.013.012 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0
48
+ 0.26.234.628 I slot print_timing: id 0 | task 0 | prompt processing, n_tokens = 4096, progress = 0.58, t = 3.22 s / 1271.42 tokens per second
49
+ 0.27.896.790 I slot print_timing: id 0 | task 0 | prompt processing, n_tokens = 6011, progress = 0.85, t = 4.88 s / 1230.81 tokens per second
50
+ 0.27.944.722 I slot create_check: id 0 | task 0 | created context checkpoint 1 of 32 (pos_min = 6010, pos_max = 6010, n_tokens = 6011, size = 62.813 MiB)
51
+ 0.28.834.529 I slot print_timing: id 0 | task 0 | prompt processing, n_tokens = 7035, progress = 1.00, t = 5.82 s / 1208.45 tokens per second
52
+ 0.28.892.962 I slot create_check: id 0 | task 0 | created context checkpoint 2 of 32 (pos_min = 7034, pos_max = 7034, n_tokens = 7035, size = 62.813 MiB)
53
+ 0.29.188.213 I slot print_timing: id 0 | task 0 |
54
+ prompt eval time = 5911.23 ms / 7039 tokens ( 0.84 ms per token, 1190.78 tokens per second)
55
+ eval time = 263.95 ms / 16 tokens ( 16.50 ms per token, 60.62 tokens per second)
56
+ total time = 6175.18 ms / 7055 tokens
57
+ 0.29.188.522 I slot release: id 0 | task 0 | stop processing: n_tokens = 7054, truncated = 0
58
+ 0.29.188.528 I srv update_slots: all slots are idle
59
+ 0.29.206.128 I srv params_from_: Chat format: peg-native
60
+ 0.29.206.239 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.996 (> 0.100 thold), f_keep = 0.994
61
+ 0.29.206.302 I slot launch_slot_: id 0 | task 21 | processing task, is_child = 0
62
+ 0.29.206.311 W slot update_slots: id 0 | task 21 | n_past = 7014, slot.prompt.tokens.size() = 7054, seq_id = 0, pos_min = 7053, n_swa = 0
63
+ 0.29.206.312 I slot update_slots: id 0 | task 21 | Checking checkpoint with [7034, 7034] against 7014...
64
+ 0.29.206.312 I slot update_slots: id 0 | task 21 | Checking checkpoint with [6010, 6010] against 7014...
65
+ 0.29.210.086 W slot update_slots: id 0 | task 21 | restored context checkpoint (pos_min = 6010, pos_max = 6010, n_tokens = 6011, n_past = 6011, size = 62.813 MiB)
66
+ 0.29.210.090 W slot update_slots: id 0 | task 21 | erased invalidated context checkpoint (pos_min = 7034, pos_max = 7034, n_tokens = 7035, n_swa = 0, pos_next = 6011, size = 62.813 MiB)
67
+ 0.30.175.583 I slot create_check: id 0 | task 21 | created context checkpoint 2 of 32 (pos_min = 7034, pos_max = 7034, n_tokens = 7035, size = 62.813 MiB)
68
+ 0.31.806.656 I slot print_timing: id 0 | task 21 | n_decoded = 100, tg = 62.48 t/s
69
+ 0.33.255.966 I slot print_timing: id 0 | task 21 |
70
+ prompt eval time = 999.89 ms / 1028 tokens ( 0.97 ms per token, 1028.11 tokens per second)
71
+ eval time = 3049.75 ms / 192 tokens ( 15.88 ms per token, 62.96 tokens per second)
72
+ total time = 4049.65 ms / 1220 tokens
73
+ 0.33.256.219 I slot release: id 0 | task 21 | stop processing: n_tokens = 7230, truncated = 0
74
+ 0.33.256.240 I srv update_slots: all slots are idle
75
+ 0.33.273.500 I srv params_from_: Chat format: peg-native
76
+ 0.33.273.615 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.974
77
+ 0.33.273.668 I slot launch_slot_: id 0 | task 215 | processing task, is_child = 0
78
+ 0.33.273.676 W slot update_slots: id 0 | task 215 | erased invalidated context checkpoint (pos_min = 6010, pos_max = 6010, n_tokens = 6011, n_swa = 0, pos_next = 0, size = 62.813 MiB)
79
+ 0.33.275.258 W slot update_slots: id 0 | task 215 | erased invalidated context checkpoint (pos_min = 7034, pos_max = 7034, n_tokens = 7035, n_swa = 0, pos_next = 0, size = 62.813 MiB)
80
+ 0.36.625.268 I slot print_timing: id 0 | task 215 | prompt processing, n_tokens = 4096, progress = 0.58, t = 3.35 s / 1222.11 tokens per second
81
+ 0.38.375.032 I slot print_timing: id 0 | task 215 | prompt processing, n_tokens = 6011, progress = 0.85, t = 5.10 s / 1178.32 tokens per second
82
+ 0.38.426.513 I slot create_check: id 0 | task 215 | created context checkpoint 1 of 32 (pos_min = 6010, pos_max = 6010, n_tokens = 6011, size = 62.813 MiB)
83
+ 0.39.343.585 I slot print_timing: id 0 | task 215 | prompt processing, n_tokens = 7035, progress = 1.00, t = 6.07 s / 1159.00 tokens per second
84
+ 0.39.401.811 I slot create_check: id 0 | task 215 | created context checkpoint 2 of 32 (pos_min = 7034, pos_max = 7034, n_tokens = 7035, size = 62.813 MiB)
85
+ 0.41.031.975 I slot print_timing: id 0 | task 215 | n_decoded = 100, tg = 62.51 t/s
86
+ 0.42.481.538 I slot print_timing: id 0 | task 215 |
87
+ prompt eval time = 6158.66 ms / 7039 tokens ( 0.87 ms per token, 1142.94 tokens per second)
88
+ eval time = 3049.19 ms / 192 tokens ( 15.88 ms per token, 62.97 tokens per second)
89
+ total time = 9207.85 ms / 7231 tokens
90
+ 0.42.481.799 I slot release: id 0 | task 215 | stop processing: n_tokens = 7230, truncated = 0
91
+ 0.42.481.812 I srv update_slots: all slots are idle
92
+ 0.42.499.185 I srv params_from_: Chat format: peg-native
93
+ 0.42.499.292 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.996 (> 0.100 thold), f_keep = 0.970
94
+ 0.42.499.320 I slot launch_slot_: id 0 | task 412 | processing task, is_child = 0
95
+ 0.42.499.327 W slot update_slots: id 0 | task 412 | n_past = 7014, slot.prompt.tokens.size() = 7230, seq_id = 0, pos_min = 7229, n_swa = 0
96
+ 0.42.499.328 I slot update_slots: id 0 | task 412 | Checking checkpoint with [7034, 7034] against 7014...
97
+ 0.42.499.328 I slot update_slots: id 0 | task 412 | Checking checkpoint with [6010, 6010] against 7014...
98
+ 0.42.502.995 W slot update_slots: id 0 | task 412 | restored context checkpoint (pos_min = 6010, pos_max = 6010, n_tokens = 6011, n_past = 6011, size = 62.813 MiB)
99
+ 0.42.502.999 W slot update_slots: id 0 | task 412 | erased invalidated context checkpoint (pos_min = 7034, pos_max = 7034, n_tokens = 7035, n_swa = 0, pos_next = 6011, size = 62.813 MiB)
100
+ 0.43.471.709 I slot create_check: id 0 | task 412 | created context checkpoint 2 of 32 (pos_min = 7034, pos_max = 7034, n_tokens = 7035, size = 62.813 MiB)
101
+ 0.43.764.311 I slot print_timing: id 0 | task 412 |
102
+ prompt eval time = 1002.90 ms / 1028 tokens ( 0.98 ms per token, 1025.03 tokens per second)
103
+ eval time = 262.06 ms / 16 tokens ( 16.38 ms per token, 61.05 tokens per second)
104
+ total time = 1264.96 ms / 1044 tokens
105
+ 0.43.764.770 I slot release: id 0 | task 412 | stop processing: n_tokens = 7054, truncated = 0
106
+ 0.43.764.786 I srv update_slots: all slots are idle
107
+ 0.43.792.285 I srv params_from_: Chat format: peg-native
108
+ 0.43.792.395 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.996 (> 0.100 thold), f_keep = 0.994
109
+ 0.43.792.452 I slot launch_slot_: id 0 | task 430 | processing task, is_child = 0
110
+ 0.43.792.461 W slot update_slots: id 0 | task 430 | n_past = 7014, slot.prompt.tokens.size() = 7054, seq_id = 0, pos_min = 7053, n_swa = 0
111
+ 0.43.792.462 I slot update_slots: id 0 | task 430 | Checking checkpoint with [7034, 7034] against 7014...
112
+ 0.43.792.462 I slot update_slots: id 0 | task 430 | Checking checkpoint with [6010, 6010] against 7014...
113
+ 0.43.796.220 W slot update_slots: id 0 | task 430 | restored context checkpoint (pos_min = 6010, pos_max = 6010, n_tokens = 6011, n_past = 6011, size = 62.813 MiB)
114
+ 0.43.796.224 W slot update_slots: id 0 | task 430 | erased invalidated context checkpoint (pos_min = 7034, pos_max = 7034, n_tokens = 7035, n_swa = 0, pos_next = 6011, size = 62.813 MiB)
115
+ 0.44.764.733 I slot create_check: id 0 | task 430 | created context checkpoint 2 of 32 (pos_min = 7034, pos_max = 7034, n_tokens = 7035, size = 62.813 MiB)
116
+ 0.46.401.785 I slot print_timing: id 0 | task 430 | n_decoded = 100, tg = 62.25 t/s
117
+ 0.47.852.134 I slot print_timing: id 0 | task 430 |
118
+ prompt eval time = 1002.97 ms / 1028 tokens ( 0.98 ms per token, 1024.96 tokens per second)
119
+ eval time = 3056.70 ms / 192 tokens ( 15.92 ms per token, 62.81 tokens per second)
120
+ total time = 4059.66 ms / 1220 tokens
121
+ 0.47.852.392 I slot release: id 0 | task 430 | stop processing: n_tokens = 7230, truncated = 0
122
+ 0.47.852.409 I srv update_slots: all slots are idle
123
+ 0.47.869.577 I srv params_from_: Chat format: peg-native
124
+ 0.47.869.673 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.974
125
+ 0.47.869.703 I slot launch_slot_: id 0 | task 624 | processing task, is_child = 0
126
+ 0.47.869.708 W slot update_slots: id 0 | task 624 | erased invalidated context checkpoint (pos_min = 6010, pos_max = 6010, n_tokens = 6011, n_swa = 0, pos_next = 0, size = 62.813 MiB)
127
+ 0.47.871.268 W slot update_slots: id 0 | task 624 | erased invalidated context checkpoint (pos_min = 7034, pos_max = 7034, n_tokens = 7035, n_swa = 0, pos_next = 0, size = 62.813 MiB)
128
+ 0.51.222.915 I slot print_timing: id 0 | task 624 | prompt processing, n_tokens = 4096, progress = 0.58, t = 3.35 s / 1221.52 tokens per second
129
+ 0.52.975.285 I slot print_timing: id 0 | task 624 | prompt processing, n_tokens = 6011, progress = 0.85, t = 5.11 s / 1177.34 tokens per second
130
+ 0.53.025.961 I slot create_check: id 0 | task 624 | created context checkpoint 1 of 32 (pos_min = 6010, pos_max = 6010, n_tokens = 6011, size = 62.813 MiB)
131
+ 0.53.946.619 I slot print_timing: id 0 | task 624 | prompt processing, n_tokens = 7035, progress = 1.00, t = 6.08 s / 1157.66 tokens per second
132
+ 0.54.004.891 I slot create_check: id 0 | task 624 | created context checkpoint 2 of 32 (pos_min = 7034, pos_max = 7034, n_tokens = 7035, size = 62.813 MiB)
133
+ 0.55.641.849 I slot print_timing: id 0 | task 624 | n_decoded = 100, tg = 62.25 t/s
134
+ 0.57.095.298 I slot print_timing: id 0 | task 624 |
135
+ prompt eval time = 6165.72 ms / 7039 tokens ( 0.88 ms per token, 1141.64 tokens per second)
136
+ eval time = 3059.86 ms / 192 tokens ( 15.94 ms per token, 62.75 tokens per second)
137
+ total time = 9225.58 ms / 7231 tokens
138
+ 0.57.095.560 I slot release: id 0 | task 624 | stop processing: n_tokens = 7230, truncated = 0
139
+ 0.57.095.573 I srv update_slots: all slots are idle
140
+ 0.57.112.963 I srv params_from_: Chat format: peg-native
141
+ 0.57.113.061 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.996 (> 0.100 thold), f_keep = 0.970
142
+ 0.57.113.099 I slot launch_slot_: id 0 | task 821 | processing task, is_child = 0
143
+ 0.57.113.105 W slot update_slots: id 0 | task 821 | n_past = 7014, slot.prompt.tokens.size() = 7230, seq_id = 0, pos_min = 7229, n_swa = 0
144
+ 0.57.113.106 I slot update_slots: id 0 | task 821 | Checking checkpoint with [7034, 7034] against 7014...
145
+ 0.57.113.106 I slot update_slots: id 0 | task 821 | Checking checkpoint with [6010, 6010] against 7014...
146
+ 0.57.116.771 W slot update_slots: id 0 | task 821 | restored context checkpoint (pos_min = 6010, pos_max = 6010, n_tokens = 6011, n_past = 6011, size = 62.813 MiB)
147
+ 0.57.116.773 W slot update_slots: id 0 | task 821 | erased invalidated context checkpoint (pos_min = 7034, pos_max = 7034, n_tokens = 7035, n_swa = 0, pos_next = 6011, size = 62.813 MiB)
148
+ 0.58.086.071 I slot create_check: id 0 | task 821 | created context checkpoint 2 of 32 (pos_min = 7034, pos_max = 7034, n_tokens = 7035, size = 62.813 MiB)
149
+ 0.58.378.533 I slot print_timing: id 0 | task 821 |
150
+ prompt eval time = 1003.26 ms / 1028 tokens ( 0.98 ms per token, 1024.66 tokens per second)
151
+ eval time = 262.14 ms / 16 tokens ( 16.38 ms per token, 61.04 tokens per second)
152
+ total time = 1265.40 ms / 1044 tokens
153
+ 0.58.378.987 I slot release: id 0 | task 821 | stop processing: n_tokens = 7054, truncated = 0
154
+ 0.58.379.007 I srv update_slots: all slots are idle
155
+ 0.58.404.936 I srv params_from_: Chat format: peg-native
156
+ 0.58.405.048 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.996 (> 0.100 thold), f_keep = 0.994
157
+ 0.58.405.100 I slot launch_slot_: id 0 | task 839 | processing task, is_child = 0
158
+ 0.58.405.109 W slot update_slots: id 0 | task 839 | n_past = 7014, slot.prompt.tokens.size() = 7054, seq_id = 0, pos_min = 7053, n_swa = 0
159
+ 0.58.405.109 I slot update_slots: id 0 | task 839 | Checking checkpoint with [7034, 7034] against 7014...
160
+ 0.58.405.110 I slot update_slots: id 0 | task 839 | Checking checkpoint with [6010, 6010] against 7014...
161
+ 0.58.408.850 W slot update_slots: id 0 | task 839 | restored context checkpoint (pos_min = 6010, pos_max = 6010, n_tokens = 6011, n_past = 6011, size = 62.813 MiB)
162
+ 0.58.408.855 W slot update_slots: id 0 | task 839 | erased invalidated context checkpoint (pos_min = 7034, pos_max = 7034, n_tokens = 7035, n_swa = 0, pos_next = 6011, size = 62.813 MiB)
163
+ 0.59.378.102 I slot create_check: id 0 | task 839 | created context checkpoint 2 of 32 (pos_min = 7034, pos_max = 7034, n_tokens = 7035, size = 62.813 MiB)
164
+ 1.01.014.891 I slot print_timing: id 0 | task 839 | n_decoded = 100, tg = 62.26 t/s
165
+ 1.02.468.788 I slot print_timing: id 0 | task 839 |
166
+ prompt eval time = 1003.48 ms / 1028 tokens ( 0.98 ms per token, 1024.43 tokens per second)
167
+ eval time = 3060.19 ms / 192 tokens ( 15.94 ms per token, 62.74 tokens per second)
168
+ total time = 4063.67 ms / 1220 tokens
169
+ 1.02.469.046 I slot release: id 0 | task 839 | stop processing: n_tokens = 7230, truncated = 0
170
+ 1.02.469.057 I srv update_slots: all slots are idle
171
+ 1.02.486.085 I srv params_from_: Chat format: peg-native
172
+ 1.02.486.197 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.974
173
+ 1.02.486.250 I slot launch_slot_: id 0 | task 1033 | processing task, is_child = 0
174
+ 1.02.486.266 W slot update_slots: id 0 | task 1033 | erased invalidated context checkpoint (pos_min = 6010, pos_max = 6010, n_tokens = 6011, n_swa = 0, pos_next = 0, size = 62.813 MiB)
175
+ 1.02.488.866 W slot update_slots: id 0 | task 1033 | erased invalidated context checkpoint (pos_min = 7034, pos_max = 7034, n_tokens = 7035, n_swa = 0, pos_next = 0, size = 62.813 MiB)
176
+ 1.05.842.274 I slot print_timing: id 0 | task 1033 | prompt processing, n_tokens = 4096, progress = 0.58, t = 3.36 s / 1220.51 tokens per second
177
+ 1.07.594.394 I slot print_timing: id 0 | task 1033 | prompt processing, n_tokens = 6011, progress = 0.85, t = 5.11 s / 1176.75 tokens per second
178
+ 1.07.645.015 I slot create_check: id 0 | task 1033 | created context checkpoint 1 of 32 (pos_min = 6010, pos_max = 6010, n_tokens = 6011, size = 62.813 MiB)
179
+ 1.08.566.063 I slot print_timing: id 0 | task 1033 | prompt processing, n_tokens = 7035, progress = 1.00, t = 6.08 s / 1157.11 tokens per second
180
+ 1.08.624.878 I slot create_check: id 0 | task 1033 | created context checkpoint 2 of 32 (pos_min = 7034, pos_max = 7034, n_tokens = 7035, size = 62.813 MiB)
181
+ 1.10.262.102 I slot print_timing: id 0 | task 1033 | n_decoded = 100, tg = 62.24 t/s
182
+ 1.11.717.464 I slot print_timing: id 0 | task 1033 |
183
+ prompt eval time = 6169.04 ms / 7039 tokens ( 0.88 ms per token, 1141.02 tokens per second)
184
+ eval time = 3062.15 ms / 192 tokens ( 15.95 ms per token, 62.70 tokens per second)
185
+ total time = 9231.19 ms / 7231 tokens
186
+ 1.11.717.716 I slot release: id 0 | task 1033 | stop processing: n_tokens = 7230, truncated = 0
187
+ 1.11.717.731 I srv update_slots: all slots are idle
188
+ 1.11.734.739 I srv params_from_: Chat format: peg-native
189
+ 1.11.734.856 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.996 (> 0.100 thold), f_keep = 0.970
190
+ 1.11.734.912 I slot launch_slot_: id 0 | task 1230 | processing task, is_child = 0
191
+ 1.11.734.923 W slot update_slots: id 0 | task 1230 | n_past = 7014, slot.prompt.tokens.size() = 7230, seq_id = 0, pos_min = 7229, n_swa = 0
192
+ 1.11.734.923 I slot update_slots: id 0 | task 1230 | Checking checkpoint with [7034, 7034] against 7014...
193
+ 1.11.734.924 I slot update_slots: id 0 | task 1230 | Checking checkpoint with [6010, 6010] against 7014...
194
+ 1.11.738.681 W slot update_slots: id 0 | task 1230 | restored context checkpoint (pos_min = 6010, pos_max = 6010, n_tokens = 6011, n_past = 6011, size = 62.813 MiB)
195
+ 1.11.738.683 W slot update_slots: id 0 | task 1230 | erased invalidated context checkpoint (pos_min = 7034, pos_max = 7034, n_tokens = 7035, n_swa = 0, pos_next = 6011, size = 62.813 MiB)
196
+ 1.12.708.145 I slot create_check: id 0 | task 1230 | created context checkpoint 2 of 32 (pos_min = 7034, pos_max = 7034, n_tokens = 7035, size = 62.813 MiB)
197
+ 1.13.000.920 I slot print_timing: id 0 | task 1230 |
198
+ prompt eval time = 1003.98 ms / 1028 tokens ( 0.98 ms per token, 1023.92 tokens per second)
199
+ eval time = 262.00 ms / 16 tokens ( 16.37 ms per token, 61.07 tokens per second)
200
+ total time = 1265.98 ms / 1044 tokens
201
+ 1.13.001.384 I slot release: id 0 | task 1230 | stop processing: n_tokens = 7054, truncated = 0
202
+ 1.13.001.412 I srv update_slots: all slots are idle
203
+ 1.13.029.139 I srv params_from_: Chat format: peg-native
204
+ 1.13.029.237 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.996 (> 0.100 thold), f_keep = 0.994
205
+ 1.13.029.274 I slot launch_slot_: id 0 | task 1248 | processing task, is_child = 0
206
+ 1.13.029.280 W slot update_slots: id 0 | task 1248 | n_past = 7014, slot.prompt.tokens.size() = 7054, seq_id = 0, pos_min = 7053, n_swa = 0
207
+ 1.13.029.280 I slot update_slots: id 0 | task 1248 | Checking checkpoint with [7034, 7034] against 7014...
208
+ 1.13.029.281 I slot update_slots: id 0 | task 1248 | Checking checkpoint with [6010, 6010] against 7014...
209
+ 1.13.032.947 W slot update_slots: id 0 | task 1248 | restored context checkpoint (pos_min = 6010, pos_max = 6010, n_tokens = 6011, n_past = 6011, size = 62.813 MiB)
210
+ 1.13.032.950 W slot update_slots: id 0 | task 1248 | erased invalidated context checkpoint (pos_min = 7034, pos_max = 7034, n_tokens = 7035, n_swa = 0, pos_next = 6011, size = 62.813 MiB)
211
+ 1.14.002.308 I slot create_check: id 0 | task 1248 | created context checkpoint 2 of 32 (pos_min = 7034, pos_max = 7034, n_tokens = 7035, size = 62.813 MiB)
212
+ 1.15.640.869 I slot print_timing: id 0 | task 1248 | n_decoded = 100, tg = 62.18 t/s
213
+ 1.17.096.125 I slot print_timing: id 0 | task 1248 |
214
+ prompt eval time = 1003.44 ms / 1028 tokens ( 0.98 ms per token, 1024.48 tokens per second)
215
+ eval time = 3063.39 ms / 192 tokens ( 15.96 ms per token, 62.68 tokens per second)
216
+ total time = 4066.83 ms / 1220 tokens
217
+ 1.17.096.385 I slot release: id 0 | task 1248 | stop processing: n_tokens = 7230, truncated = 0
218
+ 1.17.096.398 I srv update_slots: all slots are idle
219
+ 1.17.113.652 I srv params_from_: Chat format: peg-native
220
+ 1.17.113.773 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.974
221
+ 1.17.113.827 I slot launch_slot_: id 0 | task 1442 | processing task, is_child = 0
222
+ 1.17.113.835 W slot update_slots: id 0 | task 1442 | erased invalidated context checkpoint (pos_min = 6010, pos_max = 6010, n_tokens = 6011, n_swa = 0, pos_next = 0, size = 62.813 MiB)
223
+ 1.17.115.551 W slot update_slots: id 0 | task 1442 | erased invalidated context checkpoint (pos_min = 7034, pos_max = 7034, n_tokens = 7035, n_swa = 0, pos_next = 0, size = 62.813 MiB)
224
+ 1.20.471.772 I slot print_timing: id 0 | task 1442 | prompt processing, n_tokens = 4096, progress = 0.58, t = 3.36 s / 1219.80 tokens per second
225
+ 1.22.222.821 I slot print_timing: id 0 | task 1442 | prompt processing, n_tokens = 6011, progress = 0.85, t = 5.11 s / 1176.56 tokens per second
226
+ 1.22.273.999 I slot create_check: id 0 | task 1442 | created context checkpoint 1 of 32 (pos_min = 6010, pos_max = 6010, n_tokens = 6011, size = 62.813 MiB)
227
+ 1.23.196.105 I slot print_timing: id 0 | task 1442 | prompt processing, n_tokens = 7035, progress = 1.00, t = 6.08 s / 1156.64 tokens per second
228
+ 1.23.254.243 I slot create_check: id 0 | task 1442 | created context checkpoint 2 of 32 (pos_min = 7034, pos_max = 7034, n_tokens = 7035, size = 62.813 MiB)
229
+ 1.24.893.156 I slot print_timing: id 0 | task 1442 | n_decoded = 100, tg = 62.17 t/s
230
+ 1.26.350.431 I slot print_timing: id 0 | task 1442 |
231
+ prompt eval time = 6170.80 ms / 7039 tokens ( 0.88 ms per token, 1140.70 tokens per second)
232
+ eval time = 3065.79 ms / 192 tokens ( 15.97 ms per token, 62.63 tokens per second)
233
+ total time = 9236.59 ms / 7231 tokens
234
+ 1.26.350.694 I slot release: id 0 | task 1442 | stop processing: n_tokens = 7230, truncated = 0
235
+ 1.26.350.710 I srv update_slots: all slots are idle
236
+ 1.26.367.931 I srv params_from_: Chat format: peg-native
237
+ 1.26.368.047 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.996 (> 0.100 thold), f_keep = 0.970
238
+ 1.26.368.104 I slot launch_slot_: id 0 | task 1639 | processing task, is_child = 0
239
+ 1.26.368.113 W slot update_slots: id 0 | task 1639 | n_past = 7014, slot.prompt.tokens.size() = 7230, seq_id = 0, pos_min = 7229, n_swa = 0
240
+ 1.26.368.113 I slot update_slots: id 0 | task 1639 | Checking checkpoint with [7034, 7034] against 7014...
241
+ 1.26.368.114 I slot update_slots: id 0 | task 1639 | Checking checkpoint with [6010, 6010] against 7014...
242
+ 1.26.371.873 W slot update_slots: id 0 | task 1639 | restored context checkpoint (pos_min = 6010, pos_max = 6010, n_tokens = 6011, n_past = 6011, size = 62.813 MiB)
243
+ 1.26.371.876 W slot update_slots: id 0 | task 1639 | erased invalidated context checkpoint (pos_min = 7034, pos_max = 7034, n_tokens = 7035, n_swa = 0, pos_next = 6011, size = 62.813 MiB)
244
+ 1.27.342.456 I slot create_check: id 0 | task 1639 | created context checkpoint 2 of 32 (pos_min = 7034, pos_max = 7034, n_tokens = 7035, size = 62.813 MiB)
245
+ 1.27.635.050 I slot print_timing: id 0 | task 1639 |
246
+ prompt eval time = 1004.71 ms / 1028 tokens ( 0.98 ms per token, 1023.18 tokens per second)
247
+ eval time = 262.20 ms / 16 tokens ( 16.39 ms per token, 61.02 tokens per second)
248
+ total time = 1266.92 ms / 1044 tokens
249
+ 1.27.635.512 I slot release: id 0 | task 1639 | stop processing: n_tokens = 7054, truncated = 0
250
+ 1.27.635.532 I srv update_slots: all slots are idle
251
+ 1.27.666.084 I srv params_from_: Chat format: peg-native
252
+ 1.27.666.197 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.996 (> 0.100 thold), f_keep = 0.994
253
+ 1.27.666.254 I slot launch_slot_: id 0 | task 1657 | processing task, is_child = 0
254
+ 1.27.666.272 W slot update_slots: id 0 | task 1657 | n_past = 7014, slot.prompt.tokens.size() = 7054, seq_id = 0, pos_min = 7053, n_swa = 0
255
+ 1.27.666.273 I slot update_slots: id 0 | task 1657 | Checking checkpoint with [7034, 7034] against 7014...
256
+ 1.27.666.273 I slot update_slots: id 0 | task 1657 | Checking checkpoint with [6010, 6010] against 7014...
257
+ 1.27.670.025 W slot update_slots: id 0 | task 1657 | restored context checkpoint (pos_min = 6010, pos_max = 6010, n_tokens = 6011, n_past = 6011, size = 62.813 MiB)
258
+ 1.27.670.030 W slot update_slots: id 0 | task 1657 | erased invalidated context checkpoint (pos_min = 7034, pos_max = 7034, n_tokens = 7035, n_swa = 0, pos_next = 6011, size = 62.813 MiB)
259
+ 1.28.641.307 I slot create_check: id 0 | task 1657 | created context checkpoint 2 of 32 (pos_min = 7034, pos_max = 7034, n_tokens = 7035, size = 62.813 MiB)
260
+ 1.30.279.245 I slot print_timing: id 0 | task 1657 | n_decoded = 100, tg = 62.21 t/s
261
+ 1.31.732.686 I slot print_timing: id 0 | task 1657 |
262
+ prompt eval time = 1005.61 ms / 1028 tokens ( 0.98 ms per token, 1022.27 tokens per second)
263
+ eval time = 3060.81 ms / 192 tokens ( 15.94 ms per token, 62.73 tokens per second)
264
+ total time = 4066.42 ms / 1220 tokens
265
+ 1.31.732.942 I slot release: id 0 | task 1657 | stop processing: n_tokens = 7230, truncated = 0
266
+ 1.31.732.953 I srv update_slots: all slots are idle
267
+ 1.31.750.071 I srv params_from_: Chat format: peg-native
268
+ 1.31.750.187 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.974
269
+ 1.31.750.241 I slot launch_slot_: id 0 | task 1851 | processing task, is_child = 0
270
+ 1.31.750.248 W slot update_slots: id 0 | task 1851 | erased invalidated context checkpoint (pos_min = 6010, pos_max = 6010, n_tokens = 6011, n_swa = 0, pos_next = 0, size = 62.813 MiB)
271
+ 1.31.751.779 W slot update_slots: id 0 | task 1851 | erased invalidated context checkpoint (pos_min = 7034, pos_max = 7034, n_tokens = 7035, n_swa = 0, pos_next = 0, size = 62.813 MiB)
272
+ 1.35.110.825 I slot print_timing: id 0 | task 1851 | prompt processing, n_tokens = 4096, progress = 0.58, t = 3.36 s / 1218.84 tokens per second
273
+ 1.36.865.814 I slot print_timing: id 0 | task 1851 | prompt processing, n_tokens = 6011, progress = 0.85, t = 5.12 s / 1175.04 tokens per second
274
+ 1.36.917.039 I slot create_check: id 0 | task 1851 | created context checkpoint 1 of 32 (pos_min = 6010, pos_max = 6010, n_tokens = 6011, size = 62.813 MiB)
275
+ 1.37.839.294 I slot print_timing: id 0 | task 1851 | prompt processing, n_tokens = 7035, progress = 1.00, t = 6.09 s / 1155.36 tokens per second
276
+ 1.37.897.427 I slot create_check: id 0 | task 1851 | created context checkpoint 2 of 32 (pos_min = 7034, pos_max = 7034, n_tokens = 7035, size = 62.813 MiB)
277
+ 1.39.536.081 I slot print_timing: id 0 | task 1851 | n_decoded = 100, tg = 62.18 t/s
278
+ 1.40.995.612 I slot print_timing: id 0 | task 1851 |
279
+ prompt eval time = 6177.63 ms / 7039 tokens ( 0.88 ms per token, 1139.43 tokens per second)
280
+ eval time = 3067.72 ms / 192 tokens ( 15.98 ms per token, 62.59 tokens per second)
281
+ total time = 9245.35 ms / 7231 tokens
282
+ 1.40.995.868 I slot release: id 0 | task 1851 | stop processing: n_tokens = 7230, truncated = 0
283
+ 1.40.995.883 I srv update_slots: all slots are idle
284
+ 1.40.996.543 I srv operator(): operator(): cleaning up before exit...
recipe/logs/b_n-tools-q106-c1-probe.log ADDED
@@ -0,0 +1,286 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0.00.113.666 W Setting 'enable_thinking' via --chat-template-kwargs is deprecated. Use --reasoning on / --reasoning off instead.
2
+ 0.00.124.179 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
3
+ 0.00.124.183 I device_info:
4
+ 0.00.124.292 I - ROCm0 : AMD Radeon Graphics (131072 MiB, 123913 MiB free)
5
+ 0.00.124.424 I - Vulkan0 : AMD Radeon Graphics (RADV GFX1151) (132096 MiB, 131922 MiB free)
6
+ 0.00.124.432 I - CPU : AMD RYZEN AI MAX+ 395 w/ Radeon 8060S (127438 MiB, 127438 MiB free)
7
+ 0.00.124.514 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
8
+ 0.00.124.572 I srv init: running without SSL
9
+ 0.00.124.596 I srv init: using 31 threads for HTTP server
10
+ 0.00.124.597 I srv init: the WebUI is disabled
11
+ 0.00.124.672 I srv start: binding port with default address family
12
+ 0.00.125.865 I srv main: loading model
13
+ 0.00.125.873 I srv load_model: loading model '/mnt/models/nex-n2.5-mini/out/Nex-N2.5-mini-Q4_0_ROCMFP4_STRIX_LEAN.gguf'
14
+ 0.00.172.940 W llama_model_loader: direct I/O is enabled, disabling mmap
15
+ 0.22.203.780 W llama_context: n_ctx_seq (65536) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
16
+ 0.22.400.933 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
17
+ 0.22.718.577 I srv load_model: initializing slots, n_slots = 1
18
+ 0.22.944.184 W srv load_model: speculative decoding will use checkpoints
19
+ 0.22.944.193 W common_speculative_init: no implementations specified for speculative decoding
20
+ 0.22.944.194 I slot load_model: id 0 | task -1 | new slot, n_ctx = 65536
21
+ 0.22.944.259 I srv load_model: prompt cache RAM enabled: limit_mib=8192
22
+ 0.22.944.261 I srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
23
+ 0.22.944.283 I srv init: idle slots will be saved to prompt cache upon starting a new task
24
+ 0.22.956.362 I init: chat template, example_format: '<|im_start|>system
25
+ You are a helpful assistant<|im_end|>
26
+ <|im_start|>user
27
+ Hello<|im_end|>
28
+ <|im_start|>assistant
29
+ <think>
30
+
31
+ </think>
32
+
33
+ Hi there<|im_end|>
34
+ <|im_start|>user
35
+ How are you?<|im_end|>
36
+ <|im_start|>assistant
37
+ <think>
38
+
39
+ </think>
40
+
41
+ '
42
+ 0.22.964.154 I srv init: init: chat template, thinking = 1
43
+ 0.22.964.188 I srv main: model loaded
44
+ 0.22.964.191 I srv main: server is listening on http://127.0.0.1:18652
45
+ 0.22.964.207 I srv update_slots: all slots are idle
46
+ 0.24.203.095 I srv params_from_: Chat format: peg-native
47
+ 0.24.203.661 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1
48
+ 0.24.203.663 I srv get_availabl: updating prompt cache
49
+ 0.24.203.670 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000
50
+ 0.24.203.676 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 65536 tokens, 8589934592 est)
51
+ 0.24.203.678 I srv get_availabl: prompt cache update took 0.01 ms
52
+ 0.24.203.955 I reasoning-budget: activated, budget=2147483647 tokens
53
+ 0.24.203.957 I reasoning-budget: deactivated (natural end)
54
+ 0.24.203.972 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0
55
+ 0.24.920.604 I slot create_check: id 0 | task 0 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
56
+ 0.25.129.936 I slot print_timing: id 0 | task 0 |
57
+ prompt eval time = 758.65 ms / 427 tokens ( 1.78 ms per token, 562.84 tokens per second)
58
+ eval time = 167.29 ms / 4 tokens ( 41.82 ms per token, 23.91 tokens per second)
59
+ total time = 925.94 ms / 431 tokens
60
+ 0.25.130.011 I slot release: id 0 | task 0 | stop processing: n_tokens = 430, truncated = 0
61
+ 0.25.130.019 I srv update_slots: all slots are idle
62
+ 0.25.185.780 I srv params_from_: Chat format: peg-native
63
+ 0.25.186.376 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.942
64
+ 0.25.186.613 I reasoning-budget: activated, budget=2147483647 tokens
65
+ 0.25.186.618 I reasoning-budget: deactivated (natural end)
66
+ 0.25.186.779 I slot launch_slot_: id 0 | task 6 | processing task, is_child = 0
67
+ 0.25.186.816 W slot update_slots: id 0 | task 6 | n_past = 405, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
68
+ 0.25.186.819 I slot update_slots: id 0 | task 6 | Checking checkpoint with [422, 422] against 405...
69
+ 0.25.186.822 W slot update_slots: id 0 | task 6 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
70
+ 0.25.186.829 W slot update_slots: id 0 | task 6 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
71
+ 0.25.897.246 I slot create_check: id 0 | task 6 | created context checkpoint 1 of 32 (pos_min = 416, pos_max = 416, n_tokens = 417, size = 62.813 MiB)
72
+ 0.25.983.622 I slot print_timing: id 0 | task 6 |
73
+ prompt eval time = 763.27 ms / 421 tokens ( 1.81 ms per token, 551.57 tokens per second)
74
+ eval time = 33.53 ms / 2 tokens ( 16.77 ms per token, 59.64 tokens per second)
75
+ total time = 796.80 ms / 423 tokens
76
+ 0.25.983.694 I slot release: id 0 | task 6 | stop processing: n_tokens = 422, truncated = 0
77
+ 0.25.983.724 I srv update_slots: all slots are idle
78
+ 0.26.010.893 I srv params_from_: Chat format: peg-native
79
+ 0.26.011.386 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.960
80
+ 0.26.011.582 I reasoning-budget: activated, budget=2147483647 tokens
81
+ 0.26.011.584 I reasoning-budget: deactivated (natural end)
82
+ 0.26.011.615 I slot launch_slot_: id 0 | task 10 | processing task, is_child = 0
83
+ 0.26.011.627 W slot update_slots: id 0 | task 10 | n_past = 405, slot.prompt.tokens.size() = 422, seq_id = 0, pos_min = 421, n_swa = 0
84
+ 0.26.011.628 I slot update_slots: id 0 | task 10 | Checking checkpoint with [416, 416] against 405...
85
+ 0.26.011.629 W slot update_slots: id 0 | task 10 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
86
+ 0.26.011.634 W slot update_slots: id 0 | task 10 | erased invalidated context checkpoint (pos_min = 416, pos_max = 416, n_tokens = 417, n_swa = 0, pos_next = 0, size = 62.813 MiB)
87
+ 0.26.506.256 I slot create_check: id 0 | task 10 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
88
+ 0.27.421.771 I slot print_timing: id 0 | task 10 |
89
+ prompt eval time = 545.66 ms / 424 tokens ( 1.29 ms per token, 777.04 tokens per second)
90
+ eval time = 864.45 ms / 39 tokens ( 22.17 ms per token, 45.12 tokens per second)
91
+ total time = 1410.11 ms / 463 tokens
92
+ 0.27.421.958 I slot release: id 0 | task 10 | stop processing: n_tokens = 462, truncated = 0
93
+ 0.27.422.017 I srv update_slots: all slots are idle
94
+ 0.27.483.797 I srv params_from_: Chat format: peg-native
95
+ 0.27.486.309 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.951 (> 0.100 thold), f_keep = 0.879
96
+ 0.27.486.919 I reasoning-budget: activated, budget=2147483647 tokens
97
+ 0.27.486.923 I reasoning-budget: deactivated (natural end)
98
+ 0.27.487.006 I slot launch_slot_: id 0 | task 51 | processing task, is_child = 0
99
+ 0.27.487.029 W slot update_slots: id 0 | task 51 | n_past = 406, slot.prompt.tokens.size() = 462, seq_id = 0, pos_min = 461, n_swa = 0
100
+ 0.27.487.032 I slot update_slots: id 0 | task 51 | Checking checkpoint with [419, 419] against 406...
101
+ 0.27.487.034 W slot update_slots: id 0 | task 51 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
102
+ 0.27.487.041 W slot update_slots: id 0 | task 51 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
103
+ 0.28.076.337 I slot create_check: id 0 | task 51 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
104
+ 0.28.262.498 I slot print_timing: id 0 | task 51 |
105
+ prompt eval time = 638.39 ms / 427 tokens ( 1.50 ms per token, 668.86 tokens per second)
106
+ eval time = 137.05 ms / 4 tokens ( 34.26 ms per token, 29.19 tokens per second)
107
+ total time = 775.45 ms / 431 tokens
108
+ 0.28.262.688 I slot release: id 0 | task 51 | stop processing: n_tokens = 430, truncated = 0
109
+ 0.28.262.742 I srv update_slots: all slots are idle
110
+ 0.28.315.880 I srv params_from_: Chat format: peg-native
111
+ 0.28.318.140 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.942
112
+ 0.28.318.686 I reasoning-budget: activated, budget=2147483647 tokens
113
+ 0.28.318.690 I reasoning-budget: deactivated (natural end)
114
+ 0.28.318.777 I slot launch_slot_: id 0 | task 57 | processing task, is_child = 0
115
+ 0.28.318.800 W slot update_slots: id 0 | task 57 | n_past = 405, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
116
+ 0.28.318.803 I slot update_slots: id 0 | task 57 | Checking checkpoint with [422, 422] against 405...
117
+ 0.28.318.805 W slot update_slots: id 0 | task 57 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
118
+ 0.28.318.810 W slot update_slots: id 0 | task 57 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
119
+ 0.28.907.669 I slot create_check: id 0 | task 57 | created context checkpoint 1 of 32 (pos_min = 416, pos_max = 416, n_tokens = 417, size = 62.813 MiB)
120
+ 0.28.982.494 I slot print_timing: id 0 | task 57 |
121
+ prompt eval time = 636.88 ms / 421 tokens ( 1.51 ms per token, 661.04 tokens per second)
122
+ eval time = 26.81 ms / 2 tokens ( 13.41 ms per token, 74.60 tokens per second)
123
+ total time = 663.69 ms / 423 tokens
124
+ 0.28.982.572 I slot release: id 0 | task 57 | stop processing: n_tokens = 422, truncated = 0
125
+ 0.28.982.598 I srv update_slots: all slots are idle
126
+ 0.29.009.222 I srv params_from_: Chat format: peg-native
127
+ 0.29.011.294 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.960
128
+ 0.29.011.848 I reasoning-budget: activated, budget=2147483647 tokens
129
+ 0.29.011.851 I reasoning-budget: deactivated (natural end)
130
+ 0.29.011.922 I slot launch_slot_: id 0 | task 61 | processing task, is_child = 0
131
+ 0.29.011.941 W slot update_slots: id 0 | task 61 | n_past = 405, slot.prompt.tokens.size() = 422, seq_id = 0, pos_min = 421, n_swa = 0
132
+ 0.29.011.943 I slot update_slots: id 0 | task 61 | Checking checkpoint with [416, 416] against 405...
133
+ 0.29.011.945 W slot update_slots: id 0 | task 61 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
134
+ 0.29.011.949 W slot update_slots: id 0 | task 61 | erased invalidated context checkpoint (pos_min = 416, pos_max = 416, n_tokens = 417, n_swa = 0, pos_next = 0, size = 62.813 MiB)
135
+ 0.29.554.822 I slot create_check: id 0 | task 61 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
136
+ 0.30.389.784 I slot print_timing: id 0 | task 61 |
137
+ prompt eval time = 589.42 ms / 424 tokens ( 1.39 ms per token, 719.35 tokens per second)
138
+ eval time = 788.41 ms / 39 tokens ( 20.22 ms per token, 49.47 tokens per second)
139
+ total time = 1377.83 ms / 463 tokens
140
+ 0.30.389.858 I slot release: id 0 | task 61 | stop processing: n_tokens = 462, truncated = 0
141
+ 0.30.389.883 I srv update_slots: all slots are idle
142
+ 0.30.401.499 I srv params_from_: Chat format: peg-native
143
+ 0.30.401.853 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.879
144
+ 0.30.402.127 I reasoning-budget: activated, budget=2147483647 tokens
145
+ 0.30.402.129 I reasoning-budget: deactivated (natural end)
146
+ 0.30.402.192 I slot launch_slot_: id 0 | task 102 | processing task, is_child = 0
147
+ 0.30.402.206 W slot update_slots: id 0 | task 102 | n_past = 406, slot.prompt.tokens.size() = 462, seq_id = 0, pos_min = 461, n_swa = 0
148
+ 0.30.402.208 I slot update_slots: id 0 | task 102 | Checking checkpoint with [419, 419] against 406...
149
+ 0.30.402.209 W slot update_slots: id 0 | task 102 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
150
+ 0.30.402.212 W slot update_slots: id 0 | task 102 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
151
+ 0.30.923.747 I slot create_check: id 0 | task 102 | created context checkpoint 1 of 32 (pos_min = 420, pos_max = 420, n_tokens = 421, size = 62.813 MiB)
152
+ 0.31.379.211 I slot print_timing: id 0 | task 102 |
153
+ prompt eval time = 572.31 ms / 425 tokens ( 1.35 ms per token, 742.61 tokens per second)
154
+ eval time = 404.68 ms / 17 tokens ( 23.80 ms per token, 42.01 tokens per second)
155
+ total time = 976.98 ms / 442 tokens
156
+ 0.31.379.332 I slot release: id 0 | task 102 | stop processing: n_tokens = 441, truncated = 0
157
+ 0.31.379.373 I srv update_slots: all slots are idle
158
+ 0.31.405.914 I srv params_from_: Chat format: peg-native
159
+ 0.31.406.236 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.967 (> 0.100 thold), f_keep = 0.918
160
+ 0.31.406.397 I reasoning-budget: activated, budget=2147483647 tokens
161
+ 0.31.406.398 I reasoning-budget: deactivated (natural end)
162
+ 0.31.406.427 I slot launch_slot_: id 0 | task 121 | processing task, is_child = 0
163
+ 0.31.406.435 W slot update_slots: id 0 | task 121 | n_past = 405, slot.prompt.tokens.size() = 441, seq_id = 0, pos_min = 440, n_swa = 0
164
+ 0.31.406.436 I slot update_slots: id 0 | task 121 | Checking checkpoint with [420, 420] against 405...
165
+ 0.31.406.437 W slot update_slots: id 0 | task 121 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
166
+ 0.31.406.439 W slot update_slots: id 0 | task 121 | erased invalidated context checkpoint (pos_min = 420, pos_max = 420, n_tokens = 421, n_swa = 0, pos_next = 0, size = 62.813 MiB)
167
+ 0.31.861.565 I slot create_check: id 0 | task 121 | created context checkpoint 1 of 32 (pos_min = 414, pos_max = 414, n_tokens = 415, size = 62.813 MiB)
168
+ 0.32.160.128 I slot print_timing: id 0 | task 121 |
169
+ prompt eval time = 491.31 ms / 419 tokens ( 1.17 ms per token, 852.82 tokens per second)
170
+ eval time = 262.37 ms / 12 tokens ( 21.86 ms per token, 45.74 tokens per second)
171
+ total time = 753.68 ms / 431 tokens
172
+ 0.32.160.211 I slot release: id 0 | task 121 | stop processing: n_tokens = 430, truncated = 0
173
+ 0.32.160.241 I srv update_slots: all slots are idle
174
+ 0.32.209.194 I srv params_from_: Chat format: peg-native
175
+ 0.32.209.712 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.960 (> 0.100 thold), f_keep = 0.942
176
+ 0.32.209.953 I reasoning-budget: activated, budget=2147483647 tokens
177
+ 0.32.209.954 I reasoning-budget: deactivated (natural end)
178
+ 0.32.209.994 I slot launch_slot_: id 0 | task 135 | processing task, is_child = 0
179
+ 0.32.210.006 W slot update_slots: id 0 | task 135 | n_past = 405, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
180
+ 0.32.210.007 I slot update_slots: id 0 | task 135 | Checking checkpoint with [414, 414] against 405...
181
+ 0.32.210.008 W slot update_slots: id 0 | task 135 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
182
+ 0.32.210.011 W slot update_slots: id 0 | task 135 | erased invalidated context checkpoint (pos_min = 414, pos_max = 414, n_tokens = 415, n_swa = 0, pos_next = 0, size = 62.813 MiB)
183
+ 0.32.786.978 I slot create_check: id 0 | task 135 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
184
+ 0.33.839.717 I slot print_timing: id 0 | task 135 |
185
+ prompt eval time = 614.39 ms / 422 tokens ( 1.46 ms per token, 686.86 tokens per second)
186
+ eval time = 1015.30 ms / 53 tokens ( 19.16 ms per token, 52.20 tokens per second)
187
+ total time = 1629.69 ms / 475 tokens
188
+ 0.33.839.794 I slot release: id 0 | task 135 | stop processing: n_tokens = 474, truncated = 0
189
+ 0.33.839.825 I srv update_slots: all slots are idle
190
+ 0.33.879.962 I srv params_from_: Chat format: peg-native
191
+ 0.33.880.516 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.958 (> 0.100 thold), f_keep = 0.857
192
+ 0.33.880.743 I reasoning-budget: activated, budget=2147483647 tokens
193
+ 0.33.880.744 I reasoning-budget: deactivated (natural end)
194
+ 0.33.880.792 I slot launch_slot_: id 0 | task 190 | processing task, is_child = 0
195
+ 0.33.880.805 W slot update_slots: id 0 | task 190 | n_past = 406, slot.prompt.tokens.size() = 474, seq_id = 0, pos_min = 473, n_swa = 0
196
+ 0.33.880.806 I slot update_slots: id 0 | task 190 | Checking checkpoint with [417, 417] against 406...
197
+ 0.33.880.807 W slot update_slots: id 0 | task 190 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
198
+ 0.33.880.810 W slot update_slots: id 0 | task 190 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
199
+ 0.34.423.705 I slot create_check: id 0 | task 190 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
200
+ 0.34.628.920 I slot print_timing: id 0 | task 190 |
201
+ prompt eval time = 583.31 ms / 424 tokens ( 1.38 ms per token, 726.89 tokens per second)
202
+ eval time = 164.80 ms / 7 tokens ( 23.54 ms per token, 42.48 tokens per second)
203
+ total time = 748.11 ms / 431 tokens
204
+ 0.34.628.991 I slot release: id 0 | task 190 | stop processing: n_tokens = 430, truncated = 0
205
+ 0.34.629.018 I srv update_slots: all slots are idle
206
+ 0.34.642.421 I srv params_from_: Chat format: peg-native
207
+ 0.34.642.945 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.969 (> 0.100 thold), f_keep = 0.942
208
+ 0.34.643.148 I reasoning-budget: activated, budget=2147483647 tokens
209
+ 0.34.643.150 I reasoning-budget: deactivated (natural end)
210
+ 0.34.643.185 I slot launch_slot_: id 0 | task 199 | processing task, is_child = 0
211
+ 0.34.643.195 W slot update_slots: id 0 | task 199 | n_past = 405, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
212
+ 0.34.643.196 I slot update_slots: id 0 | task 199 | Checking checkpoint with [419, 419] against 405...
213
+ 0.34.643.197 W slot update_slots: id 0 | task 199 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
214
+ 0.34.643.199 W slot update_slots: id 0 | task 199 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
215
+ 0.35.154.839 I slot create_check: id 0 | task 199 | created context checkpoint 1 of 32 (pos_min = 413, pos_max = 413, n_tokens = 414, size = 62.813 MiB)
216
+ 0.35.329.011 I slot print_timing: id 0 | task 199 |
217
+ prompt eval time = 563.78 ms / 418 tokens ( 1.35 ms per token, 741.42 tokens per second)
218
+ eval time = 122.02 ms / 5 tokens ( 24.40 ms per token, 40.98 tokens per second)
219
+ total time = 685.80 ms / 423 tokens
220
+ 0.35.329.095 I slot release: id 0 | task 199 | stop processing: n_tokens = 422, truncated = 0
221
+ 0.35.329.136 I srv update_slots: all slots are idle
222
+ 0.35.355.115 I srv params_from_: Chat format: peg-native
223
+ 0.35.355.649 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.960
224
+ 0.35.356.326 I reasoning-budget: activated, budget=2147483647 tokens
225
+ 0.35.356.340 I reasoning-budget: deactivated (natural end)
226
+ 0.35.356.552 I slot launch_slot_: id 0 | task 206 | processing task, is_child = 0
227
+ 0.35.356.603 W slot update_slots: id 0 | task 206 | n_past = 405, slot.prompt.tokens.size() = 422, seq_id = 0, pos_min = 421, n_swa = 0
228
+ 0.35.356.611 I slot update_slots: id 0 | task 206 | Checking checkpoint with [413, 413] against 405...
229
+ 0.35.356.616 W slot update_slots: id 0 | task 206 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
230
+ 0.35.356.634 W slot update_slots: id 0 | task 206 | erased invalidated context checkpoint (pos_min = 413, pos_max = 413, n_tokens = 414, n_swa = 0, pos_next = 0, size = 62.813 MiB)
231
+ 0.35.935.879 I slot create_check: id 0 | task 206 | created context checkpoint 1 of 32 (pos_min = 416, pos_max = 416, n_tokens = 417, size = 62.813 MiB)
232
+ 0.36.896.774 I slot print_timing: id 0 | task 206 |
233
+ prompt eval time = 625.99 ms / 421 tokens ( 1.49 ms per token, 672.53 tokens per second)
234
+ eval time = 914.17 ms / 42 tokens ( 21.77 ms per token, 45.94 tokens per second)
235
+ total time = 1540.17 ms / 463 tokens
236
+ 0.36.896.862 I slot release: id 0 | task 206 | stop processing: n_tokens = 462, truncated = 0
237
+ 0.36.896.896 I srv update_slots: all slots are idle
238
+ 0.36.936.572 I srv params_from_: Chat format: peg-native
239
+ 0.36.937.088 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.951 (> 0.100 thold), f_keep = 0.879
240
+ 0.36.937.669 I reasoning-budget: activated, budget=2147483647 tokens
241
+ 0.36.937.673 I reasoning-budget: deactivated (natural end)
242
+ 0.36.937.771 I slot launch_slot_: id 0 | task 250 | processing task, is_child = 0
243
+ 0.36.937.796 W slot update_slots: id 0 | task 250 | n_past = 406, slot.prompt.tokens.size() = 462, seq_id = 0, pos_min = 461, n_swa = 0
244
+ 0.36.937.799 I slot update_slots: id 0 | task 250 | Checking checkpoint with [416, 416] against 406...
245
+ 0.36.937.801 W slot update_slots: id 0 | task 250 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
246
+ 0.36.937.807 W slot update_slots: id 0 | task 250 | erased invalidated context checkpoint (pos_min = 416, pos_max = 416, n_tokens = 417, n_swa = 0, pos_next = 0, size = 62.813 MiB)
247
+ 0.37.457.001 I slot create_check: id 0 | task 250 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
248
+ 0.37.596.606 I slot print_timing: id 0 | task 250 |
249
+ prompt eval time = 561.66 ms / 427 tokens ( 1.32 ms per token, 760.24 tokens per second)
250
+ eval time = 97.14 ms / 4 tokens ( 24.28 ms per token, 41.18 tokens per second)
251
+ total time = 658.80 ms / 431 tokens
252
+ 0.37.596.716 I slot release: id 0 | task 250 | stop processing: n_tokens = 430, truncated = 0
253
+ 0.37.596.750 I srv update_slots: all slots are idle
254
+ 0.37.626.974 I srv params_from_: Chat format: peg-native
255
+ 0.37.627.483 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.942
256
+ 0.37.627.753 I reasoning-budget: activated, budget=2147483647 tokens
257
+ 0.37.627.755 I reasoning-budget: deactivated (natural end)
258
+ 0.37.627.791 I slot launch_slot_: id 0 | task 256 | processing task, is_child = 0
259
+ 0.37.627.804 W slot update_slots: id 0 | task 256 | n_past = 405, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
260
+ 0.37.627.806 I slot update_slots: id 0 | task 256 | Checking checkpoint with [422, 422] against 405...
261
+ 0.37.627.807 W slot update_slots: id 0 | task 256 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
262
+ 0.37.627.811 W slot update_slots: id 0 | task 256 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
263
+ 0.38.154.752 I slot create_check: id 0 | task 256 | created context checkpoint 1 of 32 (pos_min = 416, pos_max = 416, n_tokens = 417, size = 62.813 MiB)
264
+ 0.38.242.158 I slot print_timing: id 0 | task 256 |
265
+ prompt eval time = 577.92 ms / 421 tokens ( 1.37 ms per token, 728.48 tokens per second)
266
+ eval time = 36.41 ms / 2 tokens ( 18.21 ms per token, 54.92 tokens per second)
267
+ total time = 614.33 ms / 423 tokens
268
+ 0.38.242.323 I slot release: id 0 | task 256 | stop processing: n_tokens = 422, truncated = 0
269
+ 0.38.242.358 I srv update_slots: all slots are idle
270
+ 0.38.257.074 I srv params_from_: Chat format: peg-native
271
+ 0.38.257.513 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.960
272
+ 0.38.257.760 I reasoning-budget: activated, budget=2147483647 tokens
273
+ 0.38.257.761 I reasoning-budget: deactivated (natural end)
274
+ 0.38.257.798 I slot launch_slot_: id 0 | task 260 | processing task, is_child = 0
275
+ 0.38.257.809 W slot update_slots: id 0 | task 260 | n_past = 405, slot.prompt.tokens.size() = 422, seq_id = 0, pos_min = 421, n_swa = 0
276
+ 0.38.257.809 I slot update_slots: id 0 | task 260 | Checking checkpoint with [416, 416] against 405...
277
+ 0.38.257.810 W slot update_slots: id 0 | task 260 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
278
+ 0.38.257.813 W slot update_slots: id 0 | task 260 | erased invalidated context checkpoint (pos_min = 416, pos_max = 416, n_tokens = 417, n_swa = 0, pos_next = 0, size = 62.813 MiB)
279
+ 0.38.777.417 I slot create_check: id 0 | task 260 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
280
+ 0.39.603.048 I slot print_timing: id 0 | task 260 |
281
+ prompt eval time = 562.75 ms / 424 tokens ( 1.33 ms per token, 753.44 tokens per second)
282
+ eval time = 782.47 ms / 39 tokens ( 20.06 ms per token, 49.84 tokens per second)
283
+ total time = 1345.22 ms / 463 tokens
284
+ 0.39.603.254 I slot release: id 0 | task 260 | stop processing: n_tokens = 462, truncated = 0
285
+ 0.39.603.336 I srv update_slots: all slots are idle
286
+ 0.39.604.550 I srv operator(): operator(): cleaning up before exit...
recipe/logs/b_n-tools-q106-c1.log ADDED
@@ -0,0 +1,302 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0.00.120.044 W Setting 'enable_thinking' via --chat-template-kwargs is deprecated. Use --reasoning on / --reasoning off instead.
2
+ 0.00.163.630 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
3
+ 0.00.163.643 I device_info:
4
+ 0.00.163.785 I - ROCm0 : AMD Radeon Graphics (131072 MiB, 123858 MiB free)
5
+ 0.00.163.974 I - Vulkan0 : AMD Radeon Graphics (RADV GFX1151) (132096 MiB, 131922 MiB free)
6
+ 0.00.163.982 I - CPU : AMD RYZEN AI MAX+ 395 w/ Radeon 8060S (127438 MiB, 127438 MiB free)
7
+ 0.00.164.078 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
8
+ 0.00.164.142 I srv init: running without SSL
9
+ 0.00.164.166 I srv init: using 31 threads for HTTP server
10
+ 0.00.164.167 I srv init: the WebUI is disabled
11
+ 0.00.164.248 I srv start: binding port with default address family
12
+ 0.00.165.500 I srv main: loading model
13
+ 0.00.165.507 I srv load_model: loading model '/mnt/models/nex-n2.5-mini/out/Nex-N2.5-mini-Q4_0_ROCMFP4_STRIX_LEAN.gguf'
14
+ 0.00.216.012 W llama_model_loader: direct I/O is enabled, disabling mmap
15
+ 0.22.544.117 W llama_context: n_ctx_seq (65536) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
16
+ 0.22.849.861 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
17
+ 0.23.278.480 I srv load_model: initializing slots, n_slots = 1
18
+ 0.23.567.593 W srv load_model: speculative decoding will use checkpoints
19
+ 0.23.567.602 W common_speculative_init: no implementations specified for speculative decoding
20
+ 0.23.567.605 I slot load_model: id 0 | task -1 | new slot, n_ctx = 65536
21
+ 0.23.567.670 I srv load_model: prompt cache RAM enabled: limit_mib=8192
22
+ 0.23.567.688 I srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
23
+ 0.23.567.709 I srv init: idle slots will be saved to prompt cache upon starting a new task
24
+ 0.23.579.583 I init: chat template, example_format: '<|im_start|>system
25
+ You are a helpful assistant<|im_end|>
26
+ <|im_start|>user
27
+ Hello<|im_end|>
28
+ <|im_start|>assistant
29
+ <think>
30
+
31
+ </think>
32
+
33
+ Hi there<|im_end|>
34
+ <|im_start|>user
35
+ How are you?<|im_end|>
36
+ <|im_start|>assistant
37
+ <think>
38
+
39
+ </think>
40
+
41
+ '
42
+ 0.23.587.209 I srv init: init: chat template, thinking = 1
43
+ 0.23.587.253 I srv main: model loaded
44
+ 0.23.587.256 I srv main: server is listening on http://127.0.0.1:18600
45
+ 0.23.587.260 I srv update_slots: all slots are idle
46
+ 0.24.570.949 I srv params_from_: Chat format: peg-native
47
+ 0.24.571.406 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1
48
+ 0.24.571.411 I srv get_availabl: updating prompt cache
49
+ 0.24.571.419 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000
50
+ 0.24.571.426 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 65536 tokens, 8589934592 est)
51
+ 0.24.571.428 I srv get_availabl: prompt cache update took 0.01 ms
52
+ 0.24.571.805 I reasoning-budget: activated, budget=2147483647 tokens
53
+ 0.24.571.827 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0
54
+ 0.25.235.195 I slot create_check: id 0 | task 0 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
55
+ 0.25.806.187 I reasoning-budget: deactivated (natural end)
56
+ 0.26.650.056 I slot print_timing: id 0 | task 0 |
57
+ prompt eval time = 715.70 ms / 422 tokens ( 1.70 ms per token, 589.63 tokens per second)
58
+ eval time = 1362.48 ms / 60 tokens ( 22.71 ms per token, 44.04 tokens per second)
59
+ total time = 2078.18 ms / 482 tokens
60
+ 0.26.650.247 I slot release: id 0 | task 0 | stop processing: n_tokens = 481, truncated = 0
61
+ 0.26.650.264 I srv update_slots: all slots are idle
62
+ 0.26.675.130 I srv params_from_: Chat format: peg-native
63
+ 0.26.675.629 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.906 (> 0.100 thold), f_keep = 0.842
64
+ 0.26.676.106 I reasoning-budget: activated, budget=2147483647 tokens
65
+ 0.26.676.196 I slot launch_slot_: id 0 | task 62 | processing task, is_child = 0
66
+ 0.26.676.216 W slot update_slots: id 0 | task 62 | n_past = 405, slot.prompt.tokens.size() = 481, seq_id = 0, pos_min = 480, n_swa = 0
67
+ 0.26.676.222 I slot update_slots: id 0 | task 62 | Checking checkpoint with [417, 417] against 405...
68
+ 0.26.676.224 W slot update_slots: id 0 | task 62 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
69
+ 0.26.676.229 W slot update_slots: id 0 | task 62 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
70
+ 0.27.231.790 I slot create_check: id 0 | task 62 | created context checkpoint 1 of 32 (pos_min = 442, pos_max = 442, n_tokens = 443, size = 62.813 MiB)
71
+ 0.27.733.232 I reasoning-budget: deactivated (natural end)
72
+ 0.28.618.139 I slot print_timing: id 0 | task 62 |
73
+ prompt eval time = 593.45 ms / 447 tokens ( 1.33 ms per token, 753.22 tokens per second)
74
+ eval time = 1348.46 ms / 65 tokens ( 20.75 ms per token, 48.20 tokens per second)
75
+ total time = 1941.91 ms / 512 tokens
76
+ 0.28.618.225 I slot release: id 0 | task 62 | stop processing: n_tokens = 511, truncated = 0
77
+ 0.28.618.260 I srv update_slots: all slots are idle
78
+ 0.28.632.505 I srv params_from_: Chat format: peg-native
79
+ 0.28.632.871 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.953 (> 0.100 thold), f_keep = 0.793
80
+ 0.28.633.063 I reasoning-budget: activated, budget=2147483647 tokens
81
+ 0.28.633.092 I slot launch_slot_: id 0 | task 129 | processing task, is_child = 0
82
+ 0.28.633.101 W slot update_slots: id 0 | task 129 | n_past = 405, slot.prompt.tokens.size() = 511, seq_id = 0, pos_min = 510, n_swa = 0
83
+ 0.28.633.102 I slot update_slots: id 0 | task 129 | Checking checkpoint with [442, 442] against 405...
84
+ 0.28.633.103 W slot update_slots: id 0 | task 129 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
85
+ 0.28.633.104 W slot update_slots: id 0 | task 129 | erased invalidated context checkpoint (pos_min = 442, pos_max = 442, n_tokens = 443, n_swa = 0, pos_next = 0, size = 62.813 MiB)
86
+ 0.29.161.759 I slot create_check: id 0 | task 129 | created context checkpoint 1 of 32 (pos_min = 420, pos_max = 420, n_tokens = 421, size = 62.813 MiB)
87
+ 0.29.478.654 I reasoning-budget: deactivated (natural end)
88
+ 0.30.331.330 I slot print_timing: id 0 | task 129 |
89
+ prompt eval time = 580.01 ms / 425 tokens ( 1.36 ms per token, 732.75 tokens per second)
90
+ eval time = 1118.18 ms / 52 tokens ( 21.50 ms per token, 46.50 tokens per second)
91
+ total time = 1698.19 ms / 477 tokens
92
+ 0.30.331.521 I slot release: id 0 | task 129 | stop processing: n_tokens = 476, truncated = 0
93
+ 0.30.331.586 I srv update_slots: all slots are idle
94
+ 0.30.386.522 I srv params_from_: Chat format: peg-native
95
+ 0.30.388.735 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.953 (> 0.100 thold), f_keep = 0.851
96
+ 0.30.389.319 I reasoning-budget: activated, budget=2147483647 tokens
97
+ 0.30.389.437 I slot launch_slot_: id 0 | task 183 | processing task, is_child = 0
98
+ 0.30.389.465 W slot update_slots: id 0 | task 183 | n_past = 405, slot.prompt.tokens.size() = 476, seq_id = 0, pos_min = 475, n_swa = 0
99
+ 0.30.389.469 I slot update_slots: id 0 | task 183 | Checking checkpoint with [420, 420] against 405...
100
+ 0.30.389.471 W slot update_slots: id 0 | task 183 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
101
+ 0.30.389.478 W slot update_slots: id 0 | task 183 | erased invalidated context checkpoint (pos_min = 420, pos_max = 420, n_tokens = 421, n_swa = 0, pos_next = 0, size = 62.813 MiB)
102
+ 0.30.972.513 I slot create_check: id 0 | task 183 | created context checkpoint 1 of 32 (pos_min = 420, pos_max = 420, n_tokens = 421, size = 62.813 MiB)
103
+ 0.31.338.635 I reasoning-budget: deactivated (natural end)
104
+ 0.31.461.621 I slot print_timing: id 0 | task 183 |
105
+ prompt eval time = 630.63 ms / 425 tokens ( 1.48 ms per token, 673.92 tokens per second)
106
+ eval time = 441.50 ms / 17 tokens ( 25.97 ms per token, 38.51 tokens per second)
107
+ total time = 1072.13 ms / 442 tokens
108
+ 0.31.461.844 I slot release: id 0 | task 183 | stop processing: n_tokens = 441, truncated = 0
109
+ 0.31.461.907 I srv update_slots: all slots are idle
110
+ 0.31.524.318 I srv params_from_: Chat format: peg-native
111
+ 0.31.526.203 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.921
112
+ 0.31.526.820 I reasoning-budget: activated, budget=2147483647 tokens
113
+ 0.31.526.906 I slot launch_slot_: id 0 | task 202 | processing task, is_child = 0
114
+ 0.31.526.929 W slot update_slots: id 0 | task 202 | n_past = 406, slot.prompt.tokens.size() = 441, seq_id = 0, pos_min = 440, n_swa = 0
115
+ 0.31.526.933 I slot update_slots: id 0 | task 202 | Checking checkpoint with [420, 420] against 406...
116
+ 0.31.526.935 W slot update_slots: id 0 | task 202 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
117
+ 0.31.526.940 W slot update_slots: id 0 | task 202 | erased invalidated context checkpoint (pos_min = 420, pos_max = 420, n_tokens = 421, n_swa = 0, pos_next = 0, size = 62.813 MiB)
118
+ 0.32.129.825 I slot create_check: id 0 | task 202 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
119
+ 0.32.519.469 I reasoning-budget: deactivated (natural end)
120
+ 0.33.361.317 I slot print_timing: id 0 | task 202 |
121
+ prompt eval time = 645.36 ms / 422 tokens ( 1.53 ms per token, 653.90 tokens per second)
122
+ eval time = 1189.02 ms / 55 tokens ( 21.62 ms per token, 46.26 tokens per second)
123
+ total time = 1834.38 ms / 477 tokens
124
+ 0.33.361.399 I slot release: id 0 | task 202 | stop processing: n_tokens = 476, truncated = 0
125
+ 0.33.361.431 I srv update_slots: all slots are idle
126
+ 0.33.406.059 I srv params_from_: Chat format: peg-native
127
+ 0.33.406.536 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.854 (> 0.100 thold), f_keep = 0.884
128
+ 0.33.406.842 I reasoning-budget: activated, budget=2147483647 tokens
129
+ 0.33.406.887 I slot launch_slot_: id 0 | task 259 | processing task, is_child = 0
130
+ 0.33.406.903 W slot update_slots: id 0 | task 259 | n_past = 421, slot.prompt.tokens.size() = 476, seq_id = 0, pos_min = 475, n_swa = 0
131
+ 0.33.406.904 I slot update_slots: id 0 | task 259 | Checking checkpoint with [417, 417] against 421...
132
+ 0.33.410.865 W slot update_slots: id 0 | task 259 | restored context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_past = 418, size = 62.813 MiB)
133
+ 0.33.668.081 I slot create_check: id 0 | task 259 | created context checkpoint 2 of 32 (pos_min = 488, pos_max = 488, n_tokens = 489, size = 62.813 MiB)
134
+ 0.34.301.983 I reasoning-budget: deactivated (natural end)
135
+ 0.34.590.965 I slot print_timing: id 0 | task 259 |
136
+ prompt eval time = 311.59 ms / 75 tokens ( 4.15 ms per token, 240.70 tokens per second)
137
+ eval time = 872.45 ms / 38 tokens ( 22.96 ms per token, 43.56 tokens per second)
138
+ total time = 1184.04 ms / 113 tokens
139
+ 0.34.591.058 I slot release: id 0 | task 259 | stop processing: n_tokens = 530, truncated = 0
140
+ 0.34.591.093 I srv update_slots: all slots are idle
141
+ 0.34.626.053 I srv params_from_: Chat format: peg-native
142
+ 0.34.626.459 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.972 (> 0.100 thold), f_keep = 0.774
143
+ 0.34.626.655 I reasoning-budget: activated, budget=2147483647 tokens
144
+ 0.34.626.694 I slot launch_slot_: id 0 | task 299 | processing task, is_child = 0
145
+ 0.34.626.705 W slot update_slots: id 0 | task 299 | n_past = 410, slot.prompt.tokens.size() = 530, seq_id = 0, pos_min = 529, n_swa = 0
146
+ 0.34.626.705 I slot update_slots: id 0 | task 299 | Checking checkpoint with [488, 488] against 410...
147
+ 0.34.626.706 I slot update_slots: id 0 | task 299 | Checking checkpoint with [417, 417] against 410...
148
+ 0.34.626.706 W slot update_slots: id 0 | task 299 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
149
+ 0.34.626.709 W slot update_slots: id 0 | task 299 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
150
+ 0.34.627.674 W slot update_slots: id 0 | task 299 | erased invalidated context checkpoint (pos_min = 488, pos_max = 488, n_tokens = 489, n_swa = 0, pos_next = 0, size = 62.813 MiB)
151
+ 0.35.160.501 I slot create_check: id 0 | task 299 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
152
+ 0.35.462.626 I reasoning-budget: deactivated (natural end)
153
+ 0.36.334.370 I slot print_timing: id 0 | task 299 |
154
+ prompt eval time = 570.77 ms / 422 tokens ( 1.35 ms per token, 739.35 tokens per second)
155
+ eval time = 1136.85 ms / 53 tokens ( 21.45 ms per token, 46.62 tokens per second)
156
+ total time = 1707.63 ms / 475 tokens
157
+ 0.36.334.535 I slot release: id 0 | task 299 | stop processing: n_tokens = 474, truncated = 0
158
+ 0.36.334.587 I srv update_slots: all slots are idle
159
+ 0.36.346.401 I srv params_from_: Chat format: peg-native
160
+ 0.36.346.871 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.935 (> 0.100 thold), f_keep = 0.854
161
+ 0.36.347.275 I reasoning-budget: activated, budget=2147483647 tokens
162
+ 0.36.347.355 I slot launch_slot_: id 0 | task 354 | processing task, is_child = 0
163
+ 0.36.347.371 W slot update_slots: id 0 | task 354 | n_past = 405, slot.prompt.tokens.size() = 474, seq_id = 0, pos_min = 473, n_swa = 0
164
+ 0.36.347.374 I slot update_slots: id 0 | task 354 | Checking checkpoint with [417, 417] against 405...
165
+ 0.36.347.376 W slot update_slots: id 0 | task 354 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
166
+ 0.36.347.381 W slot update_slots: id 0 | task 354 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
167
+ 0.36.950.505 I slot create_check: id 0 | task 354 | created context checkpoint 1 of 32 (pos_min = 428, pos_max = 428, n_tokens = 429, size = 62.813 MiB)
168
+ 0.37.901.136 I reasoning-budget: deactivated (natural end)
169
+ 0.39.133.735 I slot print_timing: id 0 | task 354 | n_decoded = 100, tg = 46.90 t/s
170
+ 0.39.607.166 I slot print_timing: id 0 | task 354 |
171
+ prompt eval time = 653.96 ms / 433 tokens ( 1.51 ms per token, 662.12 tokens per second)
172
+ eval time = 2605.80 ms / 123 tokens ( 21.19 ms per token, 47.20 tokens per second)
173
+ total time = 3259.76 ms / 556 tokens
174
+ 0.39.607.372 I slot release: id 0 | task 354 | stop processing: n_tokens = 555, truncated = 0
175
+ 0.39.607.435 I srv update_slots: all slots are idle
176
+ 0.39.655.799 I srv params_from_: Chat format: peg-native
177
+ 0.39.656.354 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.730
178
+ 0.39.656.713 I reasoning-budget: activated, budget=2147483647 tokens
179
+ 0.39.656.717 I reasoning-budget: deactivated (natural end)
180
+ 0.39.656.803 I slot launch_slot_: id 0 | task 479 | processing task, is_child = 0
181
+ 0.39.656.826 W slot update_slots: id 0 | task 479 | n_past = 405, slot.prompt.tokens.size() = 555, seq_id = 0, pos_min = 554, n_swa = 0
182
+ 0.39.656.828 I slot update_slots: id 0 | task 479 | Checking checkpoint with [428, 428] against 405...
183
+ 0.39.656.830 W slot update_slots: id 0 | task 479 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
184
+ 0.39.656.835 W slot update_slots: id 0 | task 479 | erased invalidated context checkpoint (pos_min = 428, pos_max = 428, n_tokens = 429, n_swa = 0, pos_next = 0, size = 62.813 MiB)
185
+ 0.40.205.076 I slot create_check: id 0 | task 479 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
186
+ 0.41.065.168 I slot print_timing: id 0 | task 479 |
187
+ prompt eval time = 582.67 ms / 424 tokens ( 1.37 ms per token, 727.68 tokens per second)
188
+ eval time = 825.66 ms / 39 tokens ( 21.17 ms per token, 47.23 tokens per second)
189
+ total time = 1408.33 ms / 463 tokens
190
+ 0.41.065.250 I slot release: id 0 | task 479 | stop processing: n_tokens = 462, truncated = 0
191
+ 0.41.065.277 I srv update_slots: all slots are idle
192
+ 0.41.077.853 I srv params_from_: Chat format: peg-native
193
+ 0.41.078.228 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.902 (> 0.100 thold), f_keep = 0.877
194
+ 0.41.078.482 I reasoning-budget: activated, budget=2147483647 tokens
195
+ 0.41.078.484 I reasoning-budget: deactivated (natural end)
196
+ 0.41.078.524 I slot launch_slot_: id 0 | task 520 | processing task, is_child = 0
197
+ 0.41.078.534 W slot update_slots: id 0 | task 520 | n_past = 405, slot.prompt.tokens.size() = 462, seq_id = 0, pos_min = 461, n_swa = 0
198
+ 0.41.078.535 I slot update_slots: id 0 | task 520 | Checking checkpoint with [419, 419] against 405...
199
+ 0.41.078.536 W slot update_slots: id 0 | task 520 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
200
+ 0.41.078.537 W slot update_slots: id 0 | task 520 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
201
+ 0.41.589.485 I slot create_check: id 0 | task 520 | created context checkpoint 1 of 32 (pos_min = 444, pos_max = 444, n_tokens = 445, size = 62.813 MiB)
202
+ 0.43.355.441 I slot print_timing: id 0 | task 520 |
203
+ prompt eval time = 549.58 ms / 449 tokens ( 1.22 ms per token, 816.99 tokens per second)
204
+ eval time = 1727.31 ms / 86 tokens ( 20.09 ms per token, 49.79 tokens per second)
205
+ total time = 2276.89 ms / 535 tokens
206
+ 0.43.355.528 I slot release: id 0 | task 520 | stop processing: n_tokens = 534, truncated = 0
207
+ 0.43.355.559 I srv update_slots: all slots are idle
208
+ 0.43.407.484 I srv params_from_: Chat format: peg-native
209
+ 0.43.408.099 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.948 (> 0.100 thold), f_keep = 0.758
210
+ 0.43.408.372 I reasoning-budget: activated, budget=2147483647 tokens
211
+ 0.43.408.374 I reasoning-budget: deactivated (natural end)
212
+ 0.43.408.422 I slot launch_slot_: id 0 | task 608 | processing task, is_child = 0
213
+ 0.43.408.434 W slot update_slots: id 0 | task 608 | n_past = 405, slot.prompt.tokens.size() = 534, seq_id = 0, pos_min = 533, n_swa = 0
214
+ 0.43.408.435 I slot update_slots: id 0 | task 608 | Checking checkpoint with [444, 444] against 405...
215
+ 0.43.408.437 W slot update_slots: id 0 | task 608 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
216
+ 0.43.408.440 W slot update_slots: id 0 | task 608 | erased invalidated context checkpoint (pos_min = 444, pos_max = 444, n_tokens = 445, n_swa = 0, pos_next = 0, size = 62.813 MiB)
217
+ 0.43.970.219 I slot create_check: id 0 | task 608 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
218
+ 0.44.791.545 I slot print_timing: id 0 | task 608 |
219
+ prompt eval time = 601.11 ms / 427 tokens ( 1.41 ms per token, 710.35 tokens per second)
220
+ eval time = 781.99 ms / 39 tokens ( 20.05 ms per token, 49.87 tokens per second)
221
+ total time = 1383.10 ms / 466 tokens
222
+ 0.44.791.624 I slot release: id 0 | task 608 | stop processing: n_tokens = 465, truncated = 0
223
+ 0.44.791.654 I srv update_slots: all slots are idle
224
+ 0.44.833.558 I srv params_from_: Chat format: peg-native
225
+ 0.44.834.049 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.948 (> 0.100 thold), f_keep = 0.871
226
+ 0.44.834.271 I reasoning-budget: activated, budget=2147483647 tokens
227
+ 0.44.834.272 I reasoning-budget: deactivated (natural end)
228
+ 0.44.834.315 I slot launch_slot_: id 0 | task 649 | processing task, is_child = 0
229
+ 0.44.834.326 W slot update_slots: id 0 | task 649 | n_past = 405, slot.prompt.tokens.size() = 465, seq_id = 0, pos_min = 464, n_swa = 0
230
+ 0.44.834.327 I slot update_slots: id 0 | task 649 | Checking checkpoint with [422, 422] against 405...
231
+ 0.44.834.328 W slot update_slots: id 0 | task 649 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
232
+ 0.44.834.332 W slot update_slots: id 0 | task 649 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
233
+ 0.45.316.040 I slot create_check: id 0 | task 649 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
234
+ 0.45.454.550 I slot print_timing: id 0 | task 649 |
235
+ prompt eval time = 522.74 ms / 427 tokens ( 1.22 ms per token, 816.85 tokens per second)
236
+ eval time = 97.47 ms / 4 tokens ( 24.37 ms per token, 41.04 tokens per second)
237
+ total time = 620.21 ms / 431 tokens
238
+ 0.45.454.641 I slot release: id 0 | task 649 | stop processing: n_tokens = 430, truncated = 0
239
+ 0.45.454.670 I srv update_slots: all slots are idle
240
+ 0.45.491.620 I srv params_from_: Chat format: peg-native
241
+ 0.45.492.052 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.958 (> 0.100 thold), f_keep = 0.944
242
+ 0.45.492.301 I reasoning-budget: activated, budget=2147483647 tokens
243
+ 0.45.492.303 I reasoning-budget: deactivated (natural end)
244
+ 0.45.492.343 I slot launch_slot_: id 0 | task 655 | processing task, is_child = 0
245
+ 0.45.492.355 W slot update_slots: id 0 | task 655 | n_past = 406, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
246
+ 0.45.492.357 I slot update_slots: id 0 | task 655 | Checking checkpoint with [422, 422] against 406...
247
+ 0.45.492.358 W slot update_slots: id 0 | task 655 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
248
+ 0.45.492.362 W slot update_slots: id 0 | task 655 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
249
+ 0.46.038.260 I slot create_check: id 0 | task 655 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
250
+ 0.46.950.011 I slot print_timing: id 0 | task 655 |
251
+ prompt eval time = 592.07 ms / 424 tokens ( 1.40 ms per token, 716.14 tokens per second)
252
+ eval time = 865.56 ms / 40 tokens ( 21.64 ms per token, 46.21 tokens per second)
253
+ total time = 1457.62 ms / 464 tokens
254
+ 0.46.950.215 I slot release: id 0 | task 655 | stop processing: n_tokens = 463, truncated = 0
255
+ 0.46.950.270 I srv update_slots: all slots are idle
256
+ 0.46.968.851 I srv params_from_: Chat format: peg-native
257
+ 0.46.969.311 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.935 (> 0.100 thold), f_keep = 1.000
258
+ 0.46.969.946 I reasoning-budget: activated, budget=2147483647 tokens
259
+ 0.46.969.951 I reasoning-budget: deactivated (natural end)
260
+ 0.46.970.047 I slot launch_slot_: id 0 | task 697 | processing task, is_child = 0
261
+ 0.47.121.064 I slot create_check: id 0 | task 697 | created context checkpoint 2 of 32 (pos_min = 490, pos_max = 490, n_tokens = 491, size = 62.813 MiB)
262
+ 0.47.570.159 I slot print_timing: id 0 | task 697 |
263
+ prompt eval time = 211.63 ms / 32 tokens ( 6.61 ms per token, 151.21 tokens per second)
264
+ eval time = 388.43 ms / 15 tokens ( 25.90 ms per token, 38.62 tokens per second)
265
+ total time = 600.06 ms / 47 tokens
266
+ 0.47.570.388 I slot release: id 0 | task 697 | stop processing: n_tokens = 509, truncated = 0
267
+ 0.47.570.448 I srv update_slots: all slots are idle
268
+ 0.47.607.526 I srv params_from_: Chat format: peg-native
269
+ 0.47.608.040 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.967 (> 0.100 thold), f_keep = 0.806
270
+ 0.47.608.345 I reasoning-budget: activated, budget=2147483647 tokens
271
+ 0.47.608.347 I reasoning-budget: deactivated (natural end)
272
+ 0.47.608.398 I slot launch_slot_: id 0 | task 714 | processing task, is_child = 0
273
+ 0.47.608.410 W slot update_slots: id 0 | task 714 | n_past = 410, slot.prompt.tokens.size() = 509, seq_id = 0, pos_min = 508, n_swa = 0
274
+ 0.47.608.412 I slot update_slots: id 0 | task 714 | Checking checkpoint with [490, 490] against 410...
275
+ 0.47.608.412 I slot update_slots: id 0 | task 714 | Checking checkpoint with [419, 419] against 410...
276
+ 0.47.608.413 W slot update_slots: id 0 | task 714 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
277
+ 0.47.608.416 W slot update_slots: id 0 | task 714 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
278
+ 0.47.609.303 W slot update_slots: id 0 | task 714 | erased invalidated context checkpoint (pos_min = 490, pos_max = 490, n_tokens = 491, n_swa = 0, pos_next = 0, size = 62.813 MiB)
279
+ 0.48.271.002 I slot create_check: id 0 | task 714 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
280
+ 0.49.108.254 I slot print_timing: id 0 | task 714 |
281
+ prompt eval time = 731.88 ms / 424 tokens ( 1.73 ms per token, 579.33 tokens per second)
282
+ eval time = 767.92 ms / 40 tokens ( 19.20 ms per token, 52.09 tokens per second)
283
+ total time = 1499.80 ms / 464 tokens
284
+ 0.49.108.437 I slot release: id 0 | task 714 | stop processing: n_tokens = 463, truncated = 0
285
+ 0.49.108.493 I srv update_slots: all slots are idle
286
+ 0.49.161.163 I srv params_from_: Chat format: peg-native
287
+ 0.49.163.186 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.931 (> 0.100 thold), f_keep = 0.875
288
+ 0.49.163.878 I reasoning-budget: activated, budget=2147483647 tokens
289
+ 0.49.163.884 I reasoning-budget: deactivated (natural end)
290
+ 0.49.163.990 I slot launch_slot_: id 0 | task 756 | processing task, is_child = 0
291
+ 0.49.164.016 W slot update_slots: id 0 | task 756 | n_past = 405, slot.prompt.tokens.size() = 463, seq_id = 0, pos_min = 462, n_swa = 0
292
+ 0.49.164.018 I slot update_slots: id 0 | task 756 | Checking checkpoint with [419, 419] against 405...
293
+ 0.49.164.020 W slot update_slots: id 0 | task 756 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
294
+ 0.49.164.027 W slot update_slots: id 0 | task 756 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
295
+ 0.49.744.671 I slot create_check: id 0 | task 756 | created context checkpoint 1 of 32 (pos_min = 430, pos_max = 430, n_tokens = 431, size = 62.813 MiB)
296
+ 0.51.381.802 I slot print_timing: id 0 | task 756 |
297
+ prompt eval time = 632.00 ms / 435 tokens ( 1.45 ms per token, 688.30 tokens per second)
298
+ eval time = 1585.77 ms / 80 tokens ( 19.82 ms per token, 50.45 tokens per second)
299
+ total time = 2217.77 ms / 515 tokens
300
+ 0.51.382.017 I slot release: id 0 | task 756 | stop processing: n_tokens = 514, truncated = 0
301
+ 0.51.382.050 I srv update_slots: all slots are idle
302
+ 0.51.383.517 I srv operator(): operator(): cleaning up before exit...
recipe/logs/b_n-tools-q106-roff-r3.log ADDED
@@ -0,0 +1,302 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0.00.114.439 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
2
+ 0.00.114.444 I device_info:
3
+ 0.00.114.555 I - ROCm0 : AMD Radeon Graphics (131072 MiB, 123723 MiB free)
4
+ 0.00.114.708 I - Vulkan0 : AMD Radeon Graphics (RADV GFX1151) (132096 MiB, 131922 MiB free)
5
+ 0.00.114.715 I - CPU : AMD RYZEN AI MAX+ 395 w/ Radeon 8060S (127438 MiB, 127438 MiB free)
6
+ 0.00.114.796 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
7
+ 0.00.114.847 I srv init: running without SSL
8
+ 0.00.114.871 I srv init: using 31 threads for HTTP server
9
+ 0.00.114.873 I srv init: the WebUI is disabled
10
+ 0.00.114.942 I srv start: binding port with default address family
11
+ 0.00.116.175 I srv main: loading model
12
+ 0.00.116.182 I srv load_model: loading model '/mnt/models/nex-n2.5-mini/out/Nex-N2.5-mini-Q4_0_ROCMFP4_STRIX_LEAN.gguf'
13
+ 0.00.160.730 W llama_model_loader: direct I/O is enabled, disabling mmap
14
+ 0.21.513.834 W llama_context: n_ctx_seq (65536) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
15
+ 0.21.793.651 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
16
+ 0.22.217.020 I srv load_model: initializing slots, n_slots = 1
17
+ 0.22.487.377 W srv load_model: speculative decoding will use checkpoints
18
+ 0.22.487.387 W common_speculative_init: no implementations specified for speculative decoding
19
+ 0.22.487.390 I slot load_model: id 0 | task -1 | new slot, n_ctx = 65536
20
+ 0.22.487.471 I srv load_model: prompt cache RAM enabled: limit_mib=8192
21
+ 0.22.487.473 I srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
22
+ 0.22.487.495 I srv init: idle slots will be saved to prompt cache upon starting a new task
23
+ 0.22.540.918 I init: chat template, example_format: '<|im_start|>system
24
+ You are a helpful assistant<|im_end|>
25
+ <|im_start|>user
26
+ Hello<|im_end|>
27
+ <|im_start|>assistant
28
+ <think>
29
+
30
+ </think>
31
+
32
+ Hi there<|im_end|>
33
+ <|im_start|>user
34
+ How are you?<|im_end|>
35
+ <|im_start|>assistant
36
+ <think>
37
+
38
+ </think>
39
+
40
+ '
41
+ 0.22.582.904 I srv init: init: chat template, thinking = 0
42
+ 0.22.582.984 I srv main: model loaded
43
+ 0.22.582.995 I srv main: server is listening on http://127.0.0.1:18600
44
+ 0.22.583.003 I srv update_slots: all slots are idle
45
+ 0.23.994.344 I srv params_from_: Chat format: peg-native
46
+ 0.23.995.890 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1
47
+ 0.23.995.893 I srv get_availabl: updating prompt cache
48
+ 0.23.995.900 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000
49
+ 0.23.995.907 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 65536 tokens, 8589934592 est)
50
+ 0.23.995.909 I srv get_availabl: prompt cache update took 0.01 ms
51
+ 0.23.996.222 I reasoning-budget: activated, budget=2147483647 tokens
52
+ 0.23.996.243 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0
53
+ 0.24.656.206 I slot create_check: id 0 | task 0 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
54
+ 0.25.028.892 I reasoning-budget: deactivated (natural end)
55
+ 0.25.816.313 I slot print_timing: id 0 | task 0 |
56
+ prompt eval time = 698.88 ms / 422 tokens ( 1.66 ms per token, 603.83 tokens per second)
57
+ eval time = 1121.16 ms / 54 tokens ( 20.76 ms per token, 48.16 tokens per second)
58
+ total time = 1820.04 ms / 476 tokens
59
+ 0.25.816.410 I slot release: id 0 | task 0 | stop processing: n_tokens = 475, truncated = 0
60
+ 0.25.816.424 I srv update_slots: all slots are idle
61
+ 0.25.830.282 I srv params_from_: Chat format: peg-native
62
+ 0.25.830.661 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.906 (> 0.100 thold), f_keep = 0.853
63
+ 0.25.830.794 I reasoning-budget: activated, budget=2147483647 tokens
64
+ 0.25.830.832 I slot launch_slot_: id 0 | task 56 | processing task, is_child = 0
65
+ 0.25.830.841 W slot update_slots: id 0 | task 56 | n_past = 405, slot.prompt.tokens.size() = 475, seq_id = 0, pos_min = 474, n_swa = 0
66
+ 0.25.830.842 I slot update_slots: id 0 | task 56 | Checking checkpoint with [417, 417] against 405...
67
+ 0.25.830.844 W slot update_slots: id 0 | task 56 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
68
+ 0.25.830.846 W slot update_slots: id 0 | task 56 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
69
+ 0.26.482.806 I slot create_check: id 0 | task 56 | created context checkpoint 1 of 32 (pos_min = 442, pos_max = 442, n_tokens = 443, size = 62.813 MiB)
70
+ 0.26.903.330 I reasoning-budget: deactivated (natural end)
71
+ 0.28.669.262 I slot print_timing: id 0 | task 56 | n_decoded = 100, tg = 47.24 t/s
72
+ 0.28.729.150 I slot print_timing: id 0 | task 56 |
73
+ prompt eval time = 721.52 ms / 447 tokens ( 1.61 ms per token, 619.52 tokens per second)
74
+ eval time = 2176.78 ms / 103 tokens ( 21.13 ms per token, 47.32 tokens per second)
75
+ total time = 2898.30 ms / 550 tokens
76
+ 0.28.729.226 I slot release: id 0 | task 56 | stop processing: n_tokens = 549, truncated = 0
77
+ 0.28.729.266 I srv update_slots: all slots are idle
78
+ 0.28.742.730 I srv params_from_: Chat format: peg-native
79
+ 0.28.743.168 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.953 (> 0.100 thold), f_keep = 0.738
80
+ 0.28.743.356 I reasoning-budget: activated, budget=2147483647 tokens
81
+ 0.28.743.392 I slot launch_slot_: id 0 | task 161 | processing task, is_child = 0
82
+ 0.28.743.406 W slot update_slots: id 0 | task 161 | n_past = 405, slot.prompt.tokens.size() = 549, seq_id = 0, pos_min = 548, n_swa = 0
83
+ 0.28.743.407 I slot update_slots: id 0 | task 161 | Checking checkpoint with [442, 442] against 405...
84
+ 0.28.743.407 W slot update_slots: id 0 | task 161 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
85
+ 0.28.743.411 W slot update_slots: id 0 | task 161 | erased invalidated context checkpoint (pos_min = 442, pos_max = 442, n_tokens = 443, n_swa = 0, pos_next = 0, size = 62.813 MiB)
86
+ 0.29.274.427 I slot create_check: id 0 | task 161 | created context checkpoint 1 of 32 (pos_min = 420, pos_max = 420, n_tokens = 421, size = 62.813 MiB)
87
+ 0.29.681.788 I reasoning-budget: deactivated (natural end)
88
+ 0.30.465.485 I slot print_timing: id 0 | task 161 |
89
+ prompt eval time = 568.14 ms / 425 tokens ( 1.34 ms per token, 748.05 tokens per second)
90
+ eval time = 1153.92 ms / 57 tokens ( 20.24 ms per token, 49.40 tokens per second)
91
+ total time = 1722.06 ms / 482 tokens
92
+ 0.30.465.571 I slot release: id 0 | task 161 | stop processing: n_tokens = 481, truncated = 0
93
+ 0.30.465.599 I srv update_slots: all slots are idle
94
+ 0.30.480.783 I srv params_from_: Chat format: peg-native
95
+ 0.30.481.142 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.953 (> 0.100 thold), f_keep = 0.842
96
+ 0.30.481.346 I reasoning-budget: activated, budget=2147483647 tokens
97
+ 0.30.481.386 I slot launch_slot_: id 0 | task 220 | processing task, is_child = 0
98
+ 0.30.481.397 W slot update_slots: id 0 | task 220 | n_past = 405, slot.prompt.tokens.size() = 481, seq_id = 0, pos_min = 480, n_swa = 0
99
+ 0.30.481.397 I slot update_slots: id 0 | task 220 | Checking checkpoint with [420, 420] against 405...
100
+ 0.30.481.399 W slot update_slots: id 0 | task 220 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
101
+ 0.30.481.401 W slot update_slots: id 0 | task 220 | erased invalidated context checkpoint (pos_min = 420, pos_max = 420, n_tokens = 421, n_swa = 0, pos_next = 0, size = 62.813 MiB)
102
+ 0.31.008.867 I slot create_check: id 0 | task 220 | created context checkpoint 1 of 32 (pos_min = 420, pos_max = 420, n_tokens = 421, size = 62.813 MiB)
103
+ 0.31.355.463 I reasoning-budget: deactivated (natural end)
104
+ 0.31.458.677 I slot print_timing: id 0 | task 220 |
105
+ prompt eval time = 568.77 ms / 425 tokens ( 1.34 ms per token, 747.23 tokens per second)
106
+ eval time = 408.50 ms / 20 tokens ( 20.42 ms per token, 48.96 tokens per second)
107
+ total time = 977.27 ms / 445 tokens
108
+ 0.31.458.774 I slot release: id 0 | task 220 | stop processing: n_tokens = 444, truncated = 0
109
+ 0.31.458.805 I srv update_slots: all slots are idle
110
+ 0.31.510.659 I srv params_from_: Chat format: peg-native
111
+ 0.31.511.238 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.914
112
+ 0.31.511.571 I reasoning-budget: activated, budget=2147483647 tokens
113
+ 0.31.511.663 I slot launch_slot_: id 0 | task 242 | processing task, is_child = 0
114
+ 0.31.511.677 W slot update_slots: id 0 | task 242 | n_past = 406, slot.prompt.tokens.size() = 444, seq_id = 0, pos_min = 443, n_swa = 0
115
+ 0.31.511.684 I slot update_slots: id 0 | task 242 | Checking checkpoint with [420, 420] against 406...
116
+ 0.31.511.685 W slot update_slots: id 0 | task 242 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
117
+ 0.31.511.690 W slot update_slots: id 0 | task 242 | erased invalidated context checkpoint (pos_min = 420, pos_max = 420, n_tokens = 421, n_swa = 0, pos_next = 0, size = 62.813 MiB)
118
+ 0.32.149.342 I slot create_check: id 0 | task 242 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
119
+ 0.32.405.933 I reasoning-budget: deactivated (natural end)
120
+ 0.33.283.206 I slot print_timing: id 0 | task 242 |
121
+ prompt eval time = 677.52 ms / 422 tokens ( 1.61 ms per token, 622.86 tokens per second)
122
+ eval time = 1093.99 ms / 52 tokens ( 21.04 ms per token, 47.53 tokens per second)
123
+ total time = 1771.51 ms / 474 tokens
124
+ 0.33.283.283 I slot release: id 0 | task 242 | stop processing: n_tokens = 473, truncated = 0
125
+ 0.33.283.309 I srv update_slots: all slots are idle
126
+ 0.33.311.488 I srv params_from_: Chat format: peg-native
127
+ 0.33.312.015 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.854 (> 0.100 thold), f_keep = 0.890
128
+ 0.33.312.441 I reasoning-budget: activated, budget=2147483647 tokens
129
+ 0.33.312.556 I slot launch_slot_: id 0 | task 296 | processing task, is_child = 0
130
+ 0.33.312.580 W slot update_slots: id 0 | task 296 | n_past = 421, slot.prompt.tokens.size() = 473, seq_id = 0, pos_min = 472, n_swa = 0
131
+ 0.33.312.583 I slot update_slots: id 0 | task 296 | Checking checkpoint with [417, 417] against 421...
132
+ 0.33.321.205 W slot update_slots: id 0 | task 296 | restored context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_past = 418, size = 62.813 MiB)
133
+ 0.33.558.612 I slot create_check: id 0 | task 296 | created context checkpoint 2 of 32 (pos_min = 488, pos_max = 488, n_tokens = 489, size = 62.813 MiB)
134
+ 0.34.218.104 I reasoning-budget: deactivated (natural end)
135
+ 0.34.556.664 I slot print_timing: id 0 | task 296 |
136
+ prompt eval time = 286.32 ms / 75 tokens ( 3.82 ms per token, 261.95 tokens per second)
137
+ eval time = 957.75 ms / 41 tokens ( 23.36 ms per token, 42.81 tokens per second)
138
+ total time = 1244.07 ms / 116 tokens
139
+ 0.34.556.763 I slot release: id 0 | task 296 | stop processing: n_tokens = 533, truncated = 0
140
+ 0.34.556.805 I srv update_slots: all slots are idle
141
+ 0.34.571.001 I srv params_from_: Chat format: peg-native
142
+ 0.34.571.465 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.972 (> 0.100 thold), f_keep = 0.769
143
+ 0.34.571.696 I reasoning-budget: activated, budget=2147483647 tokens
144
+ 0.34.571.740 I slot launch_slot_: id 0 | task 339 | processing task, is_child = 0
145
+ 0.34.571.755 W slot update_slots: id 0 | task 339 | n_past = 410, slot.prompt.tokens.size() = 533, seq_id = 0, pos_min = 532, n_swa = 0
146
+ 0.34.571.756 I slot update_slots: id 0 | task 339 | Checking checkpoint with [488, 488] against 410...
147
+ 0.34.571.758 I slot update_slots: id 0 | task 339 | Checking checkpoint with [417, 417] against 410...
148
+ 0.34.571.759 W slot update_slots: id 0 | task 339 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
149
+ 0.34.571.763 W slot update_slots: id 0 | task 339 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
150
+ 0.34.573.311 W slot update_slots: id 0 | task 339 | erased invalidated context checkpoint (pos_min = 488, pos_max = 488, n_tokens = 489, n_swa = 0, pos_next = 0, size = 62.813 MiB)
151
+ 0.35.110.746 I slot create_check: id 0 | task 339 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
152
+ 0.35.584.142 I reasoning-budget: deactivated (natural end)
153
+ 0.36.392.408 I slot print_timing: id 0 | task 339 |
154
+ prompt eval time = 576.37 ms / 422 tokens ( 1.37 ms per token, 732.16 tokens per second)
155
+ eval time = 1244.26 ms / 62 tokens ( 20.07 ms per token, 49.83 tokens per second)
156
+ total time = 1820.64 ms / 484 tokens
157
+ 0.36.392.474 I slot release: id 0 | task 339 | stop processing: n_tokens = 483, truncated = 0
158
+ 0.36.392.500 I srv update_slots: all slots are idle
159
+ 0.36.404.680 I srv params_from_: Chat format: peg-native
160
+ 0.36.405.094 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.935 (> 0.100 thold), f_keep = 0.839
161
+ 0.36.405.518 I reasoning-budget: activated, budget=2147483647 tokens
162
+ 0.36.405.599 I slot launch_slot_: id 0 | task 403 | processing task, is_child = 0
163
+ 0.36.405.620 W slot update_slots: id 0 | task 403 | n_past = 405, slot.prompt.tokens.size() = 483, seq_id = 0, pos_min = 482, n_swa = 0
164
+ 0.36.405.621 I slot update_slots: id 0 | task 403 | Checking checkpoint with [417, 417] against 405...
165
+ 0.36.405.625 W slot update_slots: id 0 | task 403 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
166
+ 0.36.405.630 W slot update_slots: id 0 | task 403 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
167
+ 0.36.974.524 I slot create_check: id 0 | task 403 | created context checkpoint 1 of 32 (pos_min = 428, pos_max = 428, n_tokens = 429, size = 62.813 MiB)
168
+ 0.39.070.997 I slot print_timing: id 0 | task 403 | n_decoded = 100, tg = 48.88 t/s
169
+ 0.39.448.467 I reasoning-budget: deactivated (natural end)
170
+ 0.41.042.559 I slot print_timing: id 0 | task 403 |
171
+ prompt eval time = 619.66 ms / 433 tokens ( 1.43 ms per token, 698.77 tokens per second)
172
+ eval time = 4017.27 ms / 200 tokens ( 20.09 ms per token, 49.79 tokens per second)
173
+ total time = 4636.93 ms / 633 tokens
174
+ 0.41.042.634 I slot release: id 0 | task 403 | stop processing: n_tokens = 632, truncated = 0
175
+ 0.41.042.665 I srv update_slots: all slots are idle
176
+ 0.41.054.733 I srv params_from_: Chat format: peg-native
177
+ 0.41.055.120 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.641
178
+ 0.41.055.469 I reasoning-budget: activated, budget=2147483647 tokens
179
+ 0.41.055.471 I reasoning-budget: deactivated (natural end)
180
+ 0.41.055.514 I slot launch_slot_: id 0 | task 605 | processing task, is_child = 0
181
+ 0.41.055.527 W slot update_slots: id 0 | task 605 | n_past = 405, slot.prompt.tokens.size() = 632, seq_id = 0, pos_min = 631, n_swa = 0
182
+ 0.41.055.529 I slot update_slots: id 0 | task 605 | Checking checkpoint with [428, 428] against 405...
183
+ 0.41.055.530 W slot update_slots: id 0 | task 605 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
184
+ 0.41.055.533 W slot update_slots: id 0 | task 605 | erased invalidated context checkpoint (pos_min = 428, pos_max = 428, n_tokens = 429, n_swa = 0, pos_next = 0, size = 62.813 MiB)
185
+ 0.41.597.061 I slot create_check: id 0 | task 605 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
186
+ 0.42.459.956 I slot print_timing: id 0 | task 605 |
187
+ prompt eval time = 582.62 ms / 424 tokens ( 1.37 ms per token, 727.74 tokens per second)
188
+ eval time = 821.78 ms / 39 tokens ( 21.07 ms per token, 47.46 tokens per second)
189
+ total time = 1404.41 ms / 463 tokens
190
+ 0.42.460.047 I slot release: id 0 | task 605 | stop processing: n_tokens = 462, truncated = 0
191
+ 0.42.460.077 I srv update_slots: all slots are idle
192
+ 0.42.475.720 I srv params_from_: Chat format: peg-native
193
+ 0.42.476.233 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.902 (> 0.100 thold), f_keep = 0.877
194
+ 0.42.476.486 I reasoning-budget: activated, budget=2147483647 tokens
195
+ 0.42.476.491 I reasoning-budget: deactivated (natural end)
196
+ 0.42.476.536 I slot launch_slot_: id 0 | task 646 | processing task, is_child = 0
197
+ 0.42.476.550 W slot update_slots: id 0 | task 646 | n_past = 405, slot.prompt.tokens.size() = 462, seq_id = 0, pos_min = 461, n_swa = 0
198
+ 0.42.476.551 I slot update_slots: id 0 | task 646 | Checking checkpoint with [419, 419] against 405...
199
+ 0.42.476.553 W slot update_slots: id 0 | task 646 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
200
+ 0.42.476.571 W slot update_slots: id 0 | task 646 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
201
+ 0.43.126.158 I slot create_check: id 0 | task 646 | created context checkpoint 1 of 32 (pos_min = 444, pos_max = 444, n_tokens = 445, size = 62.813 MiB)
202
+ 0.44.988.511 I slot print_timing: id 0 | task 646 |
203
+ prompt eval time = 693.47 ms / 449 tokens ( 1.54 ms per token, 647.47 tokens per second)
204
+ eval time = 1818.46 ms / 86 tokens ( 21.14 ms per token, 47.29 tokens per second)
205
+ total time = 2511.93 ms / 535 tokens
206
+ 0.44.988.703 I slot release: id 0 | task 646 | stop processing: n_tokens = 534, truncated = 0
207
+ 0.44.988.760 I srv update_slots: all slots are idle
208
+ 0.45.002.229 I srv params_from_: Chat format: peg-native
209
+ 0.45.002.631 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.948 (> 0.100 thold), f_keep = 0.758
210
+ 0.45.003.110 I reasoning-budget: activated, budget=2147483647 tokens
211
+ 0.45.003.114 I reasoning-budget: deactivated (natural end)
212
+ 0.45.003.190 I slot launch_slot_: id 0 | task 734 | processing task, is_child = 0
213
+ 0.45.003.216 W slot update_slots: id 0 | task 734 | n_past = 405, slot.prompt.tokens.size() = 534, seq_id = 0, pos_min = 533, n_swa = 0
214
+ 0.45.003.220 I slot update_slots: id 0 | task 734 | Checking checkpoint with [444, 444] against 405...
215
+ 0.45.003.221 W slot update_slots: id 0 | task 734 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
216
+ 0.45.003.227 W slot update_slots: id 0 | task 734 | erased invalidated context checkpoint (pos_min = 444, pos_max = 444, n_tokens = 445, n_swa = 0, pos_next = 0, size = 62.813 MiB)
217
+ 0.45.561.942 I slot create_check: id 0 | task 734 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
218
+ 0.46.436.833 I slot print_timing: id 0 | task 734 |
219
+ prompt eval time = 595.84 ms / 427 tokens ( 1.40 ms per token, 716.64 tokens per second)
220
+ eval time = 837.77 ms / 39 tokens ( 21.48 ms per token, 46.55 tokens per second)
221
+ total time = 1433.60 ms / 466 tokens
222
+ 0.46.436.923 I slot release: id 0 | task 734 | stop processing: n_tokens = 465, truncated = 0
223
+ 0.46.436.955 I srv update_slots: all slots are idle
224
+ 0.46.450.626 I srv params_from_: Chat format: peg-native
225
+ 0.46.451.043 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.948 (> 0.100 thold), f_keep = 0.871
226
+ 0.46.451.527 I reasoning-budget: activated, budget=2147483647 tokens
227
+ 0.46.451.532 I reasoning-budget: deactivated (natural end)
228
+ 0.46.451.617 I slot launch_slot_: id 0 | task 775 | processing task, is_child = 0
229
+ 0.46.451.637 W slot update_slots: id 0 | task 775 | n_past = 405, slot.prompt.tokens.size() = 465, seq_id = 0, pos_min = 464, n_swa = 0
230
+ 0.46.451.641 I slot update_slots: id 0 | task 775 | Checking checkpoint with [422, 422] against 405...
231
+ 0.46.451.643 W slot update_slots: id 0 | task 775 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
232
+ 0.46.451.649 W slot update_slots: id 0 | task 775 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
233
+ 0.46.982.640 I slot create_check: id 0 | task 775 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
234
+ 0.47.142.299 I slot print_timing: id 0 | task 775 |
235
+ prompt eval time = 585.65 ms / 427 tokens ( 1.37 ms per token, 729.10 tokens per second)
236
+ eval time = 104.99 ms / 4 tokens ( 26.25 ms per token, 38.10 tokens per second)
237
+ total time = 690.65 ms / 431 tokens
238
+ 0.47.142.414 I slot release: id 0 | task 775 | stop processing: n_tokens = 430, truncated = 0
239
+ 0.47.142.449 I srv update_slots: all slots are idle
240
+ 0.47.193.526 I srv params_from_: Chat format: peg-native
241
+ 0.47.195.446 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.958 (> 0.100 thold), f_keep = 0.944
242
+ 0.47.195.987 I reasoning-budget: activated, budget=2147483647 tokens
243
+ 0.47.195.992 I reasoning-budget: deactivated (natural end)
244
+ 0.47.196.101 I slot launch_slot_: id 0 | task 781 | processing task, is_child = 0
245
+ 0.47.196.122 W slot update_slots: id 0 | task 781 | n_past = 406, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
246
+ 0.47.196.128 I slot update_slots: id 0 | task 781 | Checking checkpoint with [422, 422] against 406...
247
+ 0.47.196.129 W slot update_slots: id 0 | task 781 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
248
+ 0.47.196.139 W slot update_slots: id 0 | task 781 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
249
+ 0.47.774.566 I slot create_check: id 0 | task 781 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
250
+ 0.48.760.957 I slot print_timing: id 0 | task 781 |
251
+ prompt eval time = 635.18 ms / 424 tokens ( 1.50 ms per token, 667.53 tokens per second)
252
+ eval time = 929.64 ms / 40 tokens ( 23.24 ms per token, 43.03 tokens per second)
253
+ total time = 1564.83 ms / 464 tokens
254
+ 0.48.761.052 I slot release: id 0 | task 781 | stop processing: n_tokens = 463, truncated = 0
255
+ 0.48.761.080 I srv update_slots: all slots are idle
256
+ 0.48.783.508 I srv params_from_: Chat format: peg-native
257
+ 0.48.783.860 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.935 (> 0.100 thold), f_keep = 1.000
258
+ 0.48.784.082 I reasoning-budget: activated, budget=2147483647 tokens
259
+ 0.48.784.084 I reasoning-budget: deactivated (natural end)
260
+ 0.48.784.127 I slot launch_slot_: id 0 | task 823 | processing task, is_child = 0
261
+ 0.48.935.107 I slot create_check: id 0 | task 823 | created context checkpoint 2 of 32 (pos_min = 490, pos_max = 490, n_tokens = 491, size = 62.813 MiB)
262
+ 0.49.277.790 I slot print_timing: id 0 | task 823 |
263
+ prompt eval time = 217.60 ms / 32 tokens ( 6.80 ms per token, 147.06 tokens per second)
264
+ eval time = 276.04 ms / 13 tokens ( 21.23 ms per token, 47.09 tokens per second)
265
+ total time = 493.64 ms / 45 tokens
266
+ 0.49.277.883 I slot release: id 0 | task 823 | stop processing: n_tokens = 507, truncated = 0
267
+ 0.49.277.910 I srv update_slots: all slots are idle
268
+ 0.49.324.008 I srv params_from_: Chat format: peg-native
269
+ 0.49.324.548 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.967 (> 0.100 thold), f_keep = 0.809
270
+ 0.49.324.947 I reasoning-budget: activated, budget=2147483647 tokens
271
+ 0.49.324.952 I reasoning-budget: deactivated (natural end)
272
+ 0.49.325.029 I slot launch_slot_: id 0 | task 838 | processing task, is_child = 0
273
+ 0.49.325.041 W slot update_slots: id 0 | task 838 | n_past = 410, slot.prompt.tokens.size() = 507, seq_id = 0, pos_min = 506, n_swa = 0
274
+ 0.49.325.042 I slot update_slots: id 0 | task 838 | Checking checkpoint with [490, 490] against 410...
275
+ 0.49.325.043 I slot update_slots: id 0 | task 838 | Checking checkpoint with [419, 419] against 410...
276
+ 0.49.325.048 W slot update_slots: id 0 | task 838 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
277
+ 0.49.325.051 W slot update_slots: id 0 | task 838 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
278
+ 0.49.328.874 W slot update_slots: id 0 | task 838 | erased invalidated context checkpoint (pos_min = 490, pos_max = 490, n_tokens = 491, n_swa = 0, pos_next = 0, size = 62.813 MiB)
279
+ 0.49.907.729 I slot create_check: id 0 | task 838 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
280
+ 0.50.729.586 I slot print_timing: id 0 | task 838 |
281
+ prompt eval time = 624.90 ms / 424 tokens ( 1.47 ms per token, 678.50 tokens per second)
282
+ eval time = 779.62 ms / 40 tokens ( 19.49 ms per token, 51.31 tokens per second)
283
+ total time = 1404.52 ms / 464 tokens
284
+ 0.50.729.654 I slot release: id 0 | task 838 | stop processing: n_tokens = 463, truncated = 0
285
+ 0.50.729.686 I srv update_slots: all slots are idle
286
+ 0.50.775.784 I srv params_from_: Chat format: peg-native
287
+ 0.50.776.334 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.931 (> 0.100 thold), f_keep = 0.875
288
+ 0.50.776.559 I reasoning-budget: activated, budget=2147483647 tokens
289
+ 0.50.776.562 I reasoning-budget: deactivated (natural end)
290
+ 0.50.776.609 I slot launch_slot_: id 0 | task 880 | processing task, is_child = 0
291
+ 0.50.776.620 W slot update_slots: id 0 | task 880 | n_past = 405, slot.prompt.tokens.size() = 463, seq_id = 0, pos_min = 462, n_swa = 0
292
+ 0.50.776.622 I slot update_slots: id 0 | task 880 | Checking checkpoint with [419, 419] against 405...
293
+ 0.50.776.623 W slot update_slots: id 0 | task 880 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
294
+ 0.50.776.627 W slot update_slots: id 0 | task 880 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
295
+ 0.51.328.909 I slot create_check: id 0 | task 880 | created context checkpoint 1 of 32 (pos_min = 430, pos_max = 430, n_tokens = 431, size = 62.813 MiB)
296
+ 0.52.962.266 I slot print_timing: id 0 | task 880 |
297
+ prompt eval time = 596.05 ms / 435 tokens ( 1.37 ms per token, 729.81 tokens per second)
298
+ eval time = 1589.57 ms / 80 tokens ( 19.87 ms per token, 50.33 tokens per second)
299
+ total time = 2185.62 ms / 515 tokens
300
+ 0.52.962.487 I slot release: id 0 | task 880 | stop processing: n_tokens = 514, truncated = 0
301
+ 0.52.962.523 I srv update_slots: all slots are idle
302
+ 0.52.963.931 I srv operator(): operator(): cleaning up before exit...
recipe/logs/b_n-tools-q106-roff.log ADDED
@@ -0,0 +1,301 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0.00.077.092 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
2
+ 0.00.077.102 I device_info:
3
+ 0.00.077.242 I - ROCm0 : AMD Radeon Graphics (131072 MiB, 123819 MiB free)
4
+ 0.00.077.530 I - Vulkan0 : AMD Radeon Graphics (RADV GFX1151) (132096 MiB, 131922 MiB free)
5
+ 0.00.077.538 I - CPU : AMD RYZEN AI MAX+ 395 w/ Radeon 8060S (127438 MiB, 127438 MiB free)
6
+ 0.00.077.648 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
7
+ 0.00.077.722 I srv init: running without SSL
8
+ 0.00.077.768 I srv init: using 31 threads for HTTP server
9
+ 0.00.077.770 I srv init: the WebUI is disabled
10
+ 0.00.077.863 I srv start: binding port with default address family
11
+ 0.00.079.093 I srv main: loading model
12
+ 0.00.079.096 I srv load_model: loading model '/mnt/models/nex-n2.5-mini/out/Nex-N2.5-mini-Q4_0_ROCMFP4_STRIX_LEAN.gguf'
13
+ 0.00.120.315 W llama_model_loader: direct I/O is enabled, disabling mmap
14
+ 0.21.810.861 W llama_context: n_ctx_seq (65536) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
15
+ 0.22.004.505 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
16
+ 0.22.307.874 I srv load_model: initializing slots, n_slots = 1
17
+ 0.22.526.550 W srv load_model: speculative decoding will use checkpoints
18
+ 0.22.526.576 W common_speculative_init: no implementations specified for speculative decoding
19
+ 0.22.526.584 I slot load_model: id 0 | task -1 | new slot, n_ctx = 65536
20
+ 0.22.526.740 I srv load_model: prompt cache RAM enabled: limit_mib=8192
21
+ 0.22.526.745 I srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
22
+ 0.22.526.855 I srv init: idle slots will be saved to prompt cache upon starting a new task
23
+ 0.22.576.014 I init: chat template, example_format: '<|im_start|>system
24
+ You are a helpful assistant<|im_end|>
25
+ <|im_start|>user
26
+ Hello<|im_end|>
27
+ <|im_start|>assistant
28
+ <think>
29
+
30
+ </think>
31
+
32
+ Hi there<|im_end|>
33
+ <|im_start|>user
34
+ How are you?<|im_end|>
35
+ <|im_start|>assistant
36
+ <think>
37
+
38
+ </think>
39
+
40
+ '
41
+ 0.22.604.308 I srv init: init: chat template, thinking = 0
42
+ 0.22.604.360 I srv main: model loaded
43
+ 0.22.604.364 I srv main: server is listening on http://127.0.0.1:18600
44
+ 0.22.604.370 I srv update_slots: all slots are idle
45
+ 0.24.142.693 I srv params_from_: Chat format: peg-native
46
+ 0.24.143.568 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1
47
+ 0.24.143.571 I srv get_availabl: updating prompt cache
48
+ 0.24.143.578 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000
49
+ 0.24.143.583 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 65536 tokens, 8589934592 est)
50
+ 0.24.143.585 I srv get_availabl: prompt cache update took 0.01 ms
51
+ 0.24.143.943 I reasoning-budget: activated, budget=2147483647 tokens
52
+ 0.24.143.962 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0
53
+ 0.24.903.283 I slot create_check: id 0 | task 0 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
54
+ 0.25.273.393 I reasoning-budget: deactivated (natural end)
55
+ 0.26.099.925 I slot print_timing: id 0 | task 0 |
56
+ prompt eval time = 810.27 ms / 422 tokens ( 1.92 ms per token, 520.81 tokens per second)
57
+ eval time = 1145.63 ms / 51 tokens ( 22.46 ms per token, 44.52 tokens per second)
58
+ total time = 1955.91 ms / 473 tokens
59
+ 0.26.100.146 I slot release: id 0 | task 0 | stop processing: n_tokens = 472, truncated = 0
60
+ 0.26.100.163 I srv update_slots: all slots are idle
61
+ 0.26.148.514 I srv params_from_: Chat format: peg-native
62
+ 0.26.149.069 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.906 (> 0.100 thold), f_keep = 0.858
63
+ 0.26.149.246 I reasoning-budget: activated, budget=2147483647 tokens
64
+ 0.26.149.294 I slot launch_slot_: id 0 | task 53 | processing task, is_child = 0
65
+ 0.26.149.305 W slot update_slots: id 0 | task 53 | n_past = 405, slot.prompt.tokens.size() = 472, seq_id = 0, pos_min = 471, n_swa = 0
66
+ 0.26.149.306 I slot update_slots: id 0 | task 53 | Checking checkpoint with [417, 417] against 405...
67
+ 0.26.149.307 W slot update_slots: id 0 | task 53 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
68
+ 0.26.149.310 W slot update_slots: id 0 | task 53 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
69
+ 0.26.723.793 I slot create_check: id 0 | task 53 | created context checkpoint 1 of 32 (pos_min = 442, pos_max = 442, n_tokens = 443, size = 62.813 MiB)
70
+ 0.27.315.116 I reasoning-budget: deactivated (natural end)
71
+ 0.28.126.542 I slot print_timing: id 0 | task 53 |
72
+ prompt eval time = 626.21 ms / 447 tokens ( 1.40 ms per token, 713.82 tokens per second)
73
+ eval time = 1350.95 ms / 68 tokens ( 19.87 ms per token, 50.34 tokens per second)
74
+ total time = 1977.16 ms / 515 tokens
75
+ 0.28.126.747 I slot release: id 0 | task 53 | stop processing: n_tokens = 514, truncated = 0
76
+ 0.28.126.781 I srv update_slots: all slots are idle
77
+ 0.28.164.510 I srv params_from_: Chat format: peg-native
78
+ 0.28.166.615 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.953 (> 0.100 thold), f_keep = 0.788
79
+ 0.28.167.130 I reasoning-budget: activated, budget=2147483647 tokens
80
+ 0.28.167.205 I slot launch_slot_: id 0 | task 123 | processing task, is_child = 0
81
+ 0.28.167.226 W slot update_slots: id 0 | task 123 | n_past = 405, slot.prompt.tokens.size() = 514, seq_id = 0, pos_min = 513, n_swa = 0
82
+ 0.28.167.229 I slot update_slots: id 0 | task 123 | Checking checkpoint with [442, 442] against 405...
83
+ 0.28.167.232 W slot update_slots: id 0 | task 123 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
84
+ 0.28.167.239 W slot update_slots: id 0 | task 123 | erased invalidated context checkpoint (pos_min = 442, pos_max = 442, n_tokens = 443, n_swa = 0, pos_next = 0, size = 62.813 MiB)
85
+ 0.28.745.049 I slot create_check: id 0 | task 123 | created context checkpoint 1 of 32 (pos_min = 420, pos_max = 420, n_tokens = 421, size = 62.813 MiB)
86
+ 0.28.952.307 I reasoning-budget: deactivated (natural end)
87
+ 0.29.756.441 I slot print_timing: id 0 | task 123 |
88
+ prompt eval time = 613.73 ms / 425 tokens ( 1.44 ms per token, 692.49 tokens per second)
89
+ eval time = 975.48 ms / 48 tokens ( 20.32 ms per token, 49.21 tokens per second)
90
+ total time = 1589.20 ms / 473 tokens
91
+ 0.29.756.519 I slot release: id 0 | task 123 | stop processing: n_tokens = 472, truncated = 0
92
+ 0.29.756.549 I srv update_slots: all slots are idle
93
+ 0.29.771.836 I srv params_from_: Chat format: peg-native
94
+ 0.29.772.265 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.953 (> 0.100 thold), f_keep = 0.858
95
+ 0.29.772.460 I reasoning-budget: activated, budget=2147483647 tokens
96
+ 0.29.772.502 I slot launch_slot_: id 0 | task 173 | processing task, is_child = 0
97
+ 0.29.772.513 W slot update_slots: id 0 | task 173 | n_past = 405, slot.prompt.tokens.size() = 472, seq_id = 0, pos_min = 471, n_swa = 0
98
+ 0.29.772.513 I slot update_slots: id 0 | task 173 | Checking checkpoint with [420, 420] against 405...
99
+ 0.29.772.514 W slot update_slots: id 0 | task 173 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
100
+ 0.29.772.518 W slot update_slots: id 0 | task 173 | erased invalidated context checkpoint (pos_min = 420, pos_max = 420, n_tokens = 421, n_swa = 0, pos_next = 0, size = 62.813 MiB)
101
+ 0.30.386.059 I slot create_check: id 0 | task 173 | created context checkpoint 1 of 32 (pos_min = 420, pos_max = 420, n_tokens = 421, size = 62.813 MiB)
102
+ 0.30.672.836 I reasoning-budget: deactivated (natural end)
103
+ 0.30.764.932 I slot print_timing: id 0 | task 173 |
104
+ prompt eval time = 647.59 ms / 425 tokens ( 1.52 ms per token, 656.28 tokens per second)
105
+ eval time = 344.81 ms / 17 tokens ( 20.28 ms per token, 49.30 tokens per second)
106
+ total time = 992.40 ms / 442 tokens
107
+ 0.30.765.023 I slot release: id 0 | task 173 | stop processing: n_tokens = 441, truncated = 0
108
+ 0.30.765.050 I srv update_slots: all slots are idle
109
+ 0.30.804.618 I srv params_from_: Chat format: peg-native
110
+ 0.30.806.300 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.921
111
+ 0.30.806.744 I reasoning-budget: activated, budget=2147483647 tokens
112
+ 0.30.806.810 I slot launch_slot_: id 0 | task 192 | processing task, is_child = 0
113
+ 0.30.806.826 W slot update_slots: id 0 | task 192 | n_past = 406, slot.prompt.tokens.size() = 441, seq_id = 0, pos_min = 440, n_swa = 0
114
+ 0.30.806.829 I slot update_slots: id 0 | task 192 | Checking checkpoint with [420, 420] against 406...
115
+ 0.30.806.830 W slot update_slots: id 0 | task 192 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
116
+ 0.30.806.835 W slot update_slots: id 0 | task 192 | erased invalidated context checkpoint (pos_min = 420, pos_max = 420, n_tokens = 421, n_swa = 0, pos_next = 0, size = 62.813 MiB)
117
+ 0.31.351.459 I slot create_check: id 0 | task 192 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
118
+ 0.31.884.045 I reasoning-budget: deactivated (natural end)
119
+ 0.32.727.839 I slot print_timing: id 0 | task 192 |
120
+ prompt eval time = 584.38 ms / 422 tokens ( 1.38 ms per token, 722.13 tokens per second)
121
+ eval time = 1336.62 ms / 62 tokens ( 21.56 ms per token, 46.39 tokens per second)
122
+ total time = 1921.00 ms / 484 tokens
123
+ 0.32.727.928 I slot release: id 0 | task 192 | stop processing: n_tokens = 483, truncated = 0
124
+ 0.32.727.959 I srv update_slots: all slots are idle
125
+ 0.32.742.237 I srv params_from_: Chat format: peg-native
126
+ 0.32.742.652 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.854 (> 0.100 thold), f_keep = 0.872
127
+ 0.32.743.158 I reasoning-budget: activated, budget=2147483647 tokens
128
+ 0.32.743.231 I slot launch_slot_: id 0 | task 256 | processing task, is_child = 0
129
+ 0.32.743.251 W slot update_slots: id 0 | task 256 | n_past = 421, slot.prompt.tokens.size() = 483, seq_id = 0, pos_min = 482, n_swa = 0
130
+ 0.32.743.254 I slot update_slots: id 0 | task 256 | Checking checkpoint with [417, 417] against 421...
131
+ 0.32.748.604 W slot update_slots: id 0 | task 256 | restored context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_past = 418, size = 62.813 MiB)
132
+ 0.32.958.098 I slot create_check: id 0 | task 256 | created context checkpoint 2 of 32 (pos_min = 488, pos_max = 488, n_tokens = 489, size = 62.813 MiB)
133
+ 0.33.556.876 I reasoning-budget: deactivated (natural end)
134
+ 0.33.886.761 I slot print_timing: id 0 | task 256 |
135
+ prompt eval time = 258.37 ms / 75 tokens ( 3.44 ms per token, 290.28 tokens per second)
136
+ eval time = 885.13 ms / 42 tokens ( 21.07 ms per token, 47.45 tokens per second)
137
+ total time = 1143.50 ms / 117 tokens
138
+ 0.33.886.848 I slot release: id 0 | task 256 | stop processing: n_tokens = 534, truncated = 0
139
+ 0.33.886.880 I srv update_slots: all slots are idle
140
+ 0.33.924.683 I srv params_from_: Chat format: peg-native
141
+ 0.33.925.111 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.972 (> 0.100 thold), f_keep = 0.768
142
+ 0.33.925.551 I reasoning-budget: activated, budget=2147483647 tokens
143
+ 0.33.925.629 I slot launch_slot_: id 0 | task 300 | processing task, is_child = 0
144
+ 0.33.925.649 W slot update_slots: id 0 | task 300 | n_past = 410, slot.prompt.tokens.size() = 534, seq_id = 0, pos_min = 533, n_swa = 0
145
+ 0.33.925.650 I slot update_slots: id 0 | task 300 | Checking checkpoint with [488, 488] against 410...
146
+ 0.33.925.652 I slot update_slots: id 0 | task 300 | Checking checkpoint with [417, 417] against 410...
147
+ 0.33.925.653 W slot update_slots: id 0 | task 300 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
148
+ 0.33.925.658 W slot update_slots: id 0 | task 300 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
149
+ 0.33.927.212 W slot update_slots: id 0 | task 300 | erased invalidated context checkpoint (pos_min = 488, pos_max = 488, n_tokens = 489, n_swa = 0, pos_next = 0, size = 62.813 MiB)
150
+ 0.34.478.068 I slot create_check: id 0 | task 300 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
151
+ 0.34.936.115 I reasoning-budget: deactivated (natural end)
152
+ 0.35.820.698 I slot print_timing: id 0 | task 300 |
153
+ prompt eval time = 590.50 ms / 422 tokens ( 1.40 ms per token, 714.65 tokens per second)
154
+ eval time = 1304.53 ms / 61 tokens ( 21.39 ms per token, 46.76 tokens per second)
155
+ total time = 1895.03 ms / 483 tokens
156
+ 0.35.820.771 I slot release: id 0 | task 300 | stop processing: n_tokens = 482, truncated = 0
157
+ 0.35.820.802 I srv update_slots: all slots are idle
158
+ 0.35.833.343 I srv params_from_: Chat format: peg-native
159
+ 0.35.833.721 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.935 (> 0.100 thold), f_keep = 0.840
160
+ 0.35.833.997 I reasoning-budget: activated, budget=2147483647 tokens
161
+ 0.35.834.089 I slot launch_slot_: id 0 | task 363 | processing task, is_child = 0
162
+ 0.35.834.109 W slot update_slots: id 0 | task 363 | n_past = 405, slot.prompt.tokens.size() = 482, seq_id = 0, pos_min = 481, n_swa = 0
163
+ 0.35.834.109 I slot update_slots: id 0 | task 363 | Checking checkpoint with [417, 417] against 405...
164
+ 0.35.834.111 W slot update_slots: id 0 | task 363 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
165
+ 0.35.834.118 W slot update_slots: id 0 | task 363 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
166
+ 0.36.300.366 I slot create_check: id 0 | task 363 | created context checkpoint 1 of 32 (pos_min = 428, pos_max = 428, n_tokens = 429, size = 62.813 MiB)
167
+ 0.36.903.862 I reasoning-budget: deactivated (natural end)
168
+ 0.38.352.136 I slot print_timing: id 0 | task 363 | n_decoded = 100, tg = 49.65 t/s
169
+ 0.38.552.887 I slot print_timing: id 0 | task 363 |
170
+ prompt eval time = 504.10 ms / 433 tokens ( 1.16 ms per token, 858.96 tokens per second)
171
+ eval time = 2214.64 ms / 109 tokens ( 20.32 ms per token, 49.22 tokens per second)
172
+ total time = 2718.74 ms / 542 tokens
173
+ 0.38.553.137 I slot release: id 0 | task 363 | stop processing: n_tokens = 541, truncated = 0
174
+ 0.38.553.221 I srv update_slots: all slots are idle
175
+ 0.38.576.788 I srv params_from_: Chat format: peg-native
176
+ 0.38.577.342 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.749
177
+ 0.38.577.906 I reasoning-budget: activated, budget=2147483647 tokens
178
+ 0.38.577.907 I reasoning-budget: deactivated (natural end)
179
+ 0.38.577.985 I slot launch_slot_: id 0 | task 474 | processing task, is_child = 0
180
+ 0.38.578.015 W slot update_slots: id 0 | task 474 | n_past = 405, slot.prompt.tokens.size() = 541, seq_id = 0, pos_min = 540, n_swa = 0
181
+ 0.38.578.016 I slot update_slots: id 0 | task 474 | Checking checkpoint with [428, 428] against 405...
182
+ 0.38.578.018 W slot update_slots: id 0 | task 474 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
183
+ 0.38.578.022 W slot update_slots: id 0 | task 474 | erased invalidated context checkpoint (pos_min = 428, pos_max = 428, n_tokens = 429, n_swa = 0, pos_next = 0, size = 62.813 MiB)
184
+ 0.39.129.005 I slot create_check: id 0 | task 474 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
185
+ 0.40.014.664 I slot print_timing: id 0 | task 474 |
186
+ prompt eval time = 601.68 ms / 424 tokens ( 1.42 ms per token, 704.69 tokens per second)
187
+ eval time = 834.95 ms / 39 tokens ( 21.41 ms per token, 46.71 tokens per second)
188
+ total time = 1436.64 ms / 463 tokens
189
+ 0.40.014.755 I slot release: id 0 | task 474 | stop processing: n_tokens = 462, truncated = 0
190
+ 0.40.014.785 I srv update_slots: all slots are idle
191
+ 0.40.063.657 I srv params_from_: Chat format: peg-native
192
+ 0.40.065.988 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.902 (> 0.100 thold), f_keep = 0.877
193
+ 0.40.066.550 I reasoning-budget: activated, budget=2147483647 tokens
194
+ 0.40.066.555 I reasoning-budget: deactivated (natural end)
195
+ 0.40.066.646 I slot launch_slot_: id 0 | task 515 | processing task, is_child = 0
196
+ 0.40.066.666 W slot update_slots: id 0 | task 515 | n_past = 405, slot.prompt.tokens.size() = 462, seq_id = 0, pos_min = 461, n_swa = 0
197
+ 0.40.066.671 I slot update_slots: id 0 | task 515 | Checking checkpoint with [419, 419] against 405...
198
+ 0.40.066.673 W slot update_slots: id 0 | task 515 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
199
+ 0.40.066.679 W slot update_slots: id 0 | task 515 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
200
+ 0.40.651.450 I slot create_check: id 0 | task 515 | created context checkpoint 1 of 32 (pos_min = 444, pos_max = 444, n_tokens = 445, size = 62.813 MiB)
201
+ 0.42.596.129 I slot print_timing: id 0 | task 515 |
202
+ prompt eval time = 621.59 ms / 449 tokens ( 1.38 ms per token, 722.34 tokens per second)
203
+ eval time = 1907.84 ms / 86 tokens ( 22.18 ms per token, 45.08 tokens per second)
204
+ total time = 2529.43 ms / 535 tokens
205
+ 0.42.596.351 I slot release: id 0 | task 515 | stop processing: n_tokens = 534, truncated = 0
206
+ 0.42.596.411 I srv update_slots: all slots are idle
207
+ 0.42.613.151 I srv params_from_: Chat format: peg-native
208
+ 0.42.613.661 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.948 (> 0.100 thold), f_keep = 0.758
209
+ 0.42.613.938 I reasoning-budget: activated, budget=2147483647 tokens
210
+ 0.42.613.941 I reasoning-budget: deactivated (natural end)
211
+ 0.42.613.986 I slot launch_slot_: id 0 | task 603 | processing task, is_child = 0
212
+ 0.42.614.003 W slot update_slots: id 0 | task 603 | n_past = 405, slot.prompt.tokens.size() = 534, seq_id = 0, pos_min = 533, n_swa = 0
213
+ 0.42.614.004 I slot update_slots: id 0 | task 603 | Checking checkpoint with [444, 444] against 405...
214
+ 0.42.614.006 W slot update_slots: id 0 | task 603 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
215
+ 0.42.614.008 W slot update_slots: id 0 | task 603 | erased invalidated context checkpoint (pos_min = 444, pos_max = 444, n_tokens = 445, n_swa = 0, pos_next = 0, size = 62.813 MiB)
216
+ 0.43.137.655 I slot create_check: id 0 | task 603 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
217
+ 0.44.068.310 I slot print_timing: id 0 | task 603 |
218
+ prompt eval time = 589.04 ms / 427 tokens ( 1.38 ms per token, 724.90 tokens per second)
219
+ eval time = 865.23 ms / 39 tokens ( 22.19 ms per token, 45.07 tokens per second)
220
+ total time = 1454.27 ms / 466 tokens
221
+ 0.44.068.519 I slot release: id 0 | task 603 | stop processing: n_tokens = 465, truncated = 0
222
+ 0.44.068.586 I srv update_slots: all slots are idle
223
+ 0.44.108.932 I srv params_from_: Chat format: peg-native
224
+ 0.44.109.519 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.948 (> 0.100 thold), f_keep = 0.871
225
+ 0.44.109.808 I reasoning-budget: activated, budget=2147483647 tokens
226
+ 0.44.109.810 I reasoning-budget: deactivated (natural end)
227
+ 0.44.109.863 I slot launch_slot_: id 0 | task 644 | processing task, is_child = 0
228
+ 0.44.109.876 W slot update_slots: id 0 | task 644 | n_past = 405, slot.prompt.tokens.size() = 465, seq_id = 0, pos_min = 464, n_swa = 0
229
+ 0.44.109.878 I slot update_slots: id 0 | task 644 | Checking checkpoint with [422, 422] against 405...
230
+ 0.44.109.879 W slot update_slots: id 0 | task 644 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
231
+ 0.44.109.881 W slot update_slots: id 0 | task 644 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
232
+ 0.44.678.393 I slot create_check: id 0 | task 644 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
233
+ 0.44.867.520 I slot print_timing: id 0 | task 644 |
234
+ prompt eval time = 621.79 ms / 427 tokens ( 1.46 ms per token, 686.72 tokens per second)
235
+ eval time = 135.82 ms / 4 tokens ( 33.96 ms per token, 29.45 tokens per second)
236
+ total time = 757.61 ms / 431 tokens
237
+ 0.44.867.722 I slot release: id 0 | task 644 | stop processing: n_tokens = 430, truncated = 0
238
+ 0.44.867.777 I srv update_slots: all slots are idle
239
+ 0.44.920.558 I srv params_from_: Chat format: peg-native
240
+ 0.44.922.823 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.958 (> 0.100 thold), f_keep = 0.944
241
+ 0.44.923.382 I reasoning-budget: activated, budget=2147483647 tokens
242
+ 0.44.923.386 I reasoning-budget: deactivated (natural end)
243
+ 0.44.923.473 I slot launch_slot_: id 0 | task 650 | processing task, is_child = 0
244
+ 0.44.923.490 W slot update_slots: id 0 | task 650 | n_past = 406, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
245
+ 0.44.923.492 I slot update_slots: id 0 | task 650 | Checking checkpoint with [422, 422] against 406...
246
+ 0.44.923.496 W slot update_slots: id 0 | task 650 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
247
+ 0.44.923.503 W slot update_slots: id 0 | task 650 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
248
+ 0.45.494.241 I slot create_check: id 0 | task 650 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
249
+ 0.46.360.864 I slot print_timing: id 0 | task 650 |
250
+ prompt eval time = 619.67 ms / 424 tokens ( 1.46 ms per token, 684.24 tokens per second)
251
+ eval time = 817.69 ms / 40 tokens ( 20.44 ms per token, 48.92 tokens per second)
252
+ total time = 1437.36 ms / 464 tokens
253
+ 0.46.360.962 I slot release: id 0 | task 650 | stop processing: n_tokens = 463, truncated = 0
254
+ 0.46.360.994 I srv update_slots: all slots are idle
255
+ 0.46.376.404 I srv params_from_: Chat format: peg-native
256
+ 0.46.376.859 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.935 (> 0.100 thold), f_keep = 1.000
257
+ 0.46.377.382 I reasoning-budget: activated, budget=2147483647 tokens
258
+ 0.46.377.385 I reasoning-budget: deactivated (natural end)
259
+ 0.46.377.478 I slot launch_slot_: id 0 | task 692 | processing task, is_child = 0
260
+ 0.46.580.028 I slot create_check: id 0 | task 692 | created context checkpoint 2 of 32 (pos_min = 490, pos_max = 490, n_tokens = 491, size = 62.813 MiB)
261
+ 0.46.886.342 I slot print_timing: id 0 | task 692 |
262
+ prompt eval time = 260.43 ms / 32 tokens ( 8.14 ms per token, 122.87 tokens per second)
263
+ eval time = 248.41 ms / 13 tokens ( 19.11 ms per token, 52.33 tokens per second)
264
+ total time = 508.84 ms / 45 tokens
265
+ 0.46.886.415 I slot release: id 0 | task 692 | stop processing: n_tokens = 507, truncated = 0
266
+ 0.46.886.437 I srv update_slots: all slots are idle
267
+ 0.46.902.435 I srv params_from_: Chat format: peg-native
268
+ 0.46.903.190 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.967 (> 0.100 thold), f_keep = 0.809
269
+ 0.46.903.420 I reasoning-budget: activated, budget=2147483647 tokens
270
+ 0.46.903.422 I reasoning-budget: deactivated (natural end)
271
+ 0.46.903.471 I slot launch_slot_: id 0 | task 707 | processing task, is_child = 0
272
+ 0.46.903.481 W slot update_slots: id 0 | task 707 | n_past = 410, slot.prompt.tokens.size() = 507, seq_id = 0, pos_min = 506, n_swa = 0
273
+ 0.46.903.483 I slot update_slots: id 0 | task 707 | Checking checkpoint with [490, 490] against 410...
274
+ 0.46.903.484 I slot update_slots: id 0 | task 707 | Checking checkpoint with [419, 419] against 410...
275
+ 0.46.903.485 W slot update_slots: id 0 | task 707 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
276
+ 0.46.903.488 W slot update_slots: id 0 | task 707 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
277
+ 0.46.904.537 W slot update_slots: id 0 | task 707 | erased invalidated context checkpoint (pos_min = 490, pos_max = 490, n_tokens = 491, n_swa = 0, pos_next = 0, size = 62.813 MiB)
278
+ 0.47.441.027 I slot create_check: id 0 | task 707 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
279
+ 0.48.289.348 I slot print_timing: id 0 | task 707 |
280
+ prompt eval time = 577.77 ms / 424 tokens ( 1.36 ms per token, 733.85 tokens per second)
281
+ eval time = 808.09 ms / 40 tokens ( 20.20 ms per token, 49.50 tokens per second)
282
+ total time = 1385.86 ms / 464 tokens
283
+ 0.48.289.404 I slot release: id 0 | task 707 | stop processing: n_tokens = 463, truncated = 0
284
+ 0.48.289.426 I srv update_slots: all slots are idle
285
+ 0.48.302.317 I srv params_from_: Chat format: peg-native
286
+ 0.48.302.639 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.931 (> 0.100 thold), f_keep = 0.875
287
+ 0.48.302.821 I reasoning-budget: activated, budget=2147483647 tokens
288
+ 0.48.302.822 I reasoning-budget: deactivated (natural end)
289
+ 0.48.302.871 I slot launch_slot_: id 0 | task 749 | processing task, is_child = 0
290
+ 0.48.302.880 W slot update_slots: id 0 | task 749 | n_past = 405, slot.prompt.tokens.size() = 463, seq_id = 0, pos_min = 462, n_swa = 0
291
+ 0.48.302.881 I slot update_slots: id 0 | task 749 | Checking checkpoint with [419, 419] against 405...
292
+ 0.48.302.882 W slot update_slots: id 0 | task 749 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
293
+ 0.48.302.884 W slot update_slots: id 0 | task 749 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
294
+ 0.48.829.428 I slot create_check: id 0 | task 749 | created context checkpoint 1 of 32 (pos_min = 430, pos_max = 430, n_tokens = 431, size = 62.813 MiB)
295
+ 0.50.509.107 I slot print_timing: id 0 | task 749 |
296
+ prompt eval time = 572.46 ms / 435 tokens ( 1.32 ms per token, 759.87 tokens per second)
297
+ eval time = 1633.74 ms / 80 tokens ( 20.42 ms per token, 48.97 tokens per second)
298
+ total time = 2206.20 ms / 515 tokens
299
+ 0.50.509.286 I slot release: id 0 | task 749 | stop processing: n_tokens = 514, truncated = 0
300
+ 0.50.509.314 I srv update_slots: all slots are idle
301
+ 0.50.510.178 I srv operator(): operator(): cleaning up before exit...
recipe/logs/b_n-tools-q106-tpl-medium-probe.log ADDED
@@ -0,0 +1,278 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0.00.129.990 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
2
+ 0.00.129.993 I device_info:
3
+ 0.00.130.113 I - ROCm0 : AMD Radeon Graphics (131072 MiB, 122319 MiB free)
4
+ 0.00.130.278 I - Vulkan0 : AMD Radeon Graphics (RADV GFX1151) (132096 MiB, 131922 MiB free)
5
+ 0.00.130.286 I - CPU : AMD RYZEN AI MAX+ 395 w/ Radeon 8060S (127438 MiB, 127438 MiB free)
6
+ 0.00.130.376 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
7
+ 0.00.130.453 I srv init: running without SSL
8
+ 0.00.130.481 I srv init: using 31 threads for HTTP server
9
+ 0.00.130.482 I srv init: the WebUI is disabled
10
+ 0.00.130.551 I srv start: binding port with default address family
11
+ 0.00.131.742 I srv main: loading model
12
+ 0.00.131.749 I srv load_model: loading model '/mnt/models/nex-n2.5-mini/out/Nex-N2.5-mini-Q4_0_ROCMFP4_STRIX_LEAN.gguf'
13
+ 0.00.190.887 W llama_model_loader: direct I/O is enabled, disabling mmap
14
+ 0.23.864.154 W llama_context: n_ctx_seq (65536) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
15
+ 0.24.195.940 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
16
+ 0.24.556.150 I srv load_model: initializing slots, n_slots = 1
17
+ 0.24.849.854 W srv load_model: speculative decoding will use checkpoints
18
+ 0.24.849.866 W common_speculative_init: no implementations specified for speculative decoding
19
+ 0.24.849.870 I slot load_model: id 0 | task -1 | new slot, n_ctx = 65536
20
+ 0.24.849.977 I srv load_model: prompt cache RAM enabled: limit_mib=8192
21
+ 0.24.849.981 I srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
22
+ 0.24.850.006 I srv init: idle slots will be saved to prompt cache upon starting a new task
23
+ 0.24.907.565 I init: chat template, example_format: '<|im_start|>system
24
+ You are a helpful assistant<|im_end|>
25
+ <|im_start|>user
26
+ Hello<|im_end|>
27
+ <|im_start|>assistant
28
+ <think>
29
+
30
+ </think>
31
+
32
+ Hi there<|im_end|>
33
+ <|im_start|>user
34
+ How are you?<|im_end|>
35
+ <|im_start|>assistant
36
+ <think>'
37
+ 0.24.952.737 I srv init: init: chat template, thinking = 1
38
+ 0.24.952.792 I srv main: model loaded
39
+ 0.24.952.798 I srv main: server is listening on http://127.0.0.1:18652
40
+ 0.24.952.803 I srv update_slots: all slots are idle
41
+ 0.26.377.211 I srv params_from_: Chat format: peg-native
42
+ 0.26.379.299 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1
43
+ 0.26.379.305 I srv get_availabl: updating prompt cache
44
+ 0.26.379.315 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000
45
+ 0.26.379.323 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 65536 tokens, 8589934592 est)
46
+ 0.26.379.326 I srv get_availabl: prompt cache update took 0.02 ms
47
+ 0.26.380.071 I reasoning-budget: activated, budget=2147483647 tokens
48
+ 0.26.380.099 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0
49
+ 0.27.073.027 I slot create_check: id 0 | task 0 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
50
+ 0.27.138.087 I reasoning-budget: deactivated (natural end)
51
+ 0.27.269.343 I slot print_timing: id 0 | task 0 |
52
+ prompt eval time = 728.42 ms / 424 tokens ( 1.72 ms per token, 582.08 tokens per second)
53
+ eval time = 160.80 ms / 7 tokens ( 22.97 ms per token, 43.53 tokens per second)
54
+ total time = 889.21 ms / 431 tokens
55
+ 0.27.269.422 I slot release: id 0 | task 0 | stop processing: n_tokens = 430, truncated = 0
56
+ 0.27.269.432 I srv update_slots: all slots are idle
57
+ 0.27.299.600 I srv params_from_: Chat format: peg-native
58
+ 0.27.300.148 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.969 (> 0.100 thold), f_keep = 0.942
59
+ 0.27.300.612 I reasoning-budget: activated, budget=2147483647 tokens
60
+ 0.27.300.707 I slot launch_slot_: id 0 | task 9 | processing task, is_child = 0
61
+ 0.27.300.733 W slot update_slots: id 0 | task 9 | n_past = 405, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
62
+ 0.27.300.737 I slot update_slots: id 0 | task 9 | Checking checkpoint with [419, 419] against 405...
63
+ 0.27.300.739 W slot update_slots: id 0 | task 9 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
64
+ 0.27.300.746 W slot update_slots: id 0 | task 9 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
65
+ 0.27.937.713 I slot create_check: id 0 | task 9 | created context checkpoint 1 of 32 (pos_min = 413, pos_max = 413, n_tokens = 414, size = 62.813 MiB)
66
+ 0.28.043.202 I reasoning-budget: deactivated (natural end)
67
+ 0.28.142.801 I slot print_timing: id 0 | task 9 |
68
+ prompt eval time = 701.16 ms / 418 tokens ( 1.68 ms per token, 596.15 tokens per second)
69
+ eval time = 140.89 ms / 5 tokens ( 28.18 ms per token, 35.49 tokens per second)
70
+ total time = 842.06 ms / 423 tokens
71
+ 0.28.142.910 I slot release: id 0 | task 9 | stop processing: n_tokens = 422, truncated = 0
72
+ 0.28.142.959 I srv update_slots: all slots are idle
73
+ 0.28.158.005 I srv params_from_: Chat format: peg-native
74
+ 0.28.158.364 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.960
75
+ 0.28.158.578 I reasoning-budget: activated, budget=2147483647 tokens
76
+ 0.28.158.608 I slot launch_slot_: id 0 | task 16 | processing task, is_child = 0
77
+ 0.28.158.616 W slot update_slots: id 0 | task 16 | n_past = 405, slot.prompt.tokens.size() = 422, seq_id = 0, pos_min = 421, n_swa = 0
78
+ 0.28.158.617 I slot update_slots: id 0 | task 16 | Checking checkpoint with [413, 413] against 405...
79
+ 0.28.158.618 W slot update_slots: id 0 | task 16 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
80
+ 0.28.158.620 W slot update_slots: id 0 | task 16 | erased invalidated context checkpoint (pos_min = 413, pos_max = 413, n_tokens = 414, n_swa = 0, pos_next = 0, size = 62.813 MiB)
81
+ 0.28.549.814 I slot create_check: id 0 | task 16 | created context checkpoint 1 of 32 (pos_min = 416, pos_max = 416, n_tokens = 417, size = 62.813 MiB)
82
+ 0.28.603.918 I reasoning-budget: deactivated (natural end)
83
+ 0.29.232.552 I slot print_timing: id 0 | task 16 |
84
+ prompt eval time = 422.03 ms / 421 tokens ( 1.00 ms per token, 997.56 tokens per second)
85
+ eval time = 651.88 ms / 42 tokens ( 15.52 ms per token, 64.43 tokens per second)
86
+ total time = 1073.91 ms / 463 tokens
87
+ 0.29.232.642 I slot release: id 0 | task 16 | stop processing: n_tokens = 462, truncated = 0
88
+ 0.29.232.679 I srv update_slots: all slots are idle
89
+ 0.29.259.355 I srv params_from_: Chat format: peg-native
90
+ 0.29.259.776 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.951 (> 0.100 thold), f_keep = 0.879
91
+ 0.29.259.946 I reasoning-budget: activated, budget=2147483647 tokens
92
+ 0.29.259.949 I reasoning-budget: deactivated (natural end)
93
+ 0.29.259.987 I slot launch_slot_: id 0 | task 60 | processing task, is_child = 0
94
+ 0.29.259.997 W slot update_slots: id 0 | task 60 | n_past = 406, slot.prompt.tokens.size() = 462, seq_id = 0, pos_min = 461, n_swa = 0
95
+ 0.29.259.998 I slot update_slots: id 0 | task 60 | Checking checkpoint with [416, 416] against 406...
96
+ 0.29.260.000 W slot update_slots: id 0 | task 60 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
97
+ 0.29.260.002 W slot update_slots: id 0 | task 60 | erased invalidated context checkpoint (pos_min = 416, pos_max = 416, n_tokens = 417, n_swa = 0, pos_next = 0, size = 62.813 MiB)
98
+ 0.29.608.781 I slot create_check: id 0 | task 60 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
99
+ 0.29.705.934 I slot print_timing: id 0 | task 60 |
100
+ prompt eval time = 378.91 ms / 427 tokens ( 0.89 ms per token, 1126.93 tokens per second)
101
+ eval time = 67.01 ms / 4 tokens ( 16.75 ms per token, 59.69 tokens per second)
102
+ total time = 445.92 ms / 431 tokens
103
+ 0.29.706.025 I slot release: id 0 | task 60 | stop processing: n_tokens = 430, truncated = 0
104
+ 0.29.706.063 I srv update_slots: all slots are idle
105
+ 0.29.723.346 I srv params_from_: Chat format: peg-native
106
+ 0.29.723.698 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.942
107
+ 0.29.723.878 I reasoning-budget: activated, budget=2147483647 tokens
108
+ 0.29.723.880 I reasoning-budget: deactivated (natural end)
109
+ 0.29.723.911 I slot launch_slot_: id 0 | task 66 | processing task, is_child = 0
110
+ 0.29.723.921 W slot update_slots: id 0 | task 66 | n_past = 405, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
111
+ 0.29.723.922 I slot update_slots: id 0 | task 66 | Checking checkpoint with [422, 422] against 405...
112
+ 0.29.723.924 W slot update_slots: id 0 | task 66 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
113
+ 0.29.723.926 W slot update_slots: id 0 | task 66 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
114
+ 0.30.074.713 I slot create_check: id 0 | task 66 | created context checkpoint 1 of 32 (pos_min = 416, pos_max = 416, n_tokens = 417, size = 62.813 MiB)
115
+ 0.30.128.419 I slot print_timing: id 0 | task 66 |
116
+ prompt eval time = 380.89 ms / 421 tokens ( 0.90 ms per token, 1105.32 tokens per second)
117
+ eval time = 23.59 ms / 2 tokens ( 11.79 ms per token, 84.79 tokens per second)
118
+ total time = 404.48 ms / 423 tokens
119
+ 0.30.128.524 I slot release: id 0 | task 66 | stop processing: n_tokens = 422, truncated = 0
120
+ 0.30.128.572 I srv update_slots: all slots are idle
121
+ 0.30.150.200 I srv params_from_: Chat format: peg-native
122
+ 0.30.150.557 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.960
123
+ 0.30.150.711 I reasoning-budget: activated, budget=2147483647 tokens
124
+ 0.30.150.713 I reasoning-budget: deactivated (natural end)
125
+ 0.30.150.739 I slot launch_slot_: id 0 | task 70 | processing task, is_child = 0
126
+ 0.30.150.744 W slot update_slots: id 0 | task 70 | n_past = 405, slot.prompt.tokens.size() = 422, seq_id = 0, pos_min = 421, n_swa = 0
127
+ 0.30.150.745 I slot update_slots: id 0 | task 70 | Checking checkpoint with [416, 416] against 405...
128
+ 0.30.150.746 W slot update_slots: id 0 | task 70 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
129
+ 0.30.150.752 W slot update_slots: id 0 | task 70 | erased invalidated context checkpoint (pos_min = 416, pos_max = 416, n_tokens = 417, n_swa = 0, pos_next = 0, size = 62.813 MiB)
130
+ 0.30.504.785 I slot create_check: id 0 | task 70 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
131
+ 0.31.192.911 I slot print_timing: id 0 | task 70 |
132
+ prompt eval time = 383.87 ms / 424 tokens ( 0.91 ms per token, 1104.53 tokens per second)
133
+ eval time = 658.28 ms / 39 tokens ( 16.88 ms per token, 59.25 tokens per second)
134
+ total time = 1042.15 ms / 463 tokens
135
+ 0.31.192.995 I slot release: id 0 | task 70 | stop processing: n_tokens = 462, truncated = 0
136
+ 0.31.193.032 I srv update_slots: all slots are idle
137
+ 0.31.213.646 I srv params_from_: Chat format: peg-native
138
+ 0.31.213.997 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.879
139
+ 0.31.214.182 I reasoning-budget: activated, budget=2147483647 tokens
140
+ 0.31.214.218 I slot launch_slot_: id 0 | task 111 | processing task, is_child = 0
141
+ 0.31.214.228 W slot update_slots: id 0 | task 111 | n_past = 406, slot.prompt.tokens.size() = 462, seq_id = 0, pos_min = 461, n_swa = 0
142
+ 0.31.214.230 I slot update_slots: id 0 | task 111 | Checking checkpoint with [419, 419] against 406...
143
+ 0.31.214.231 W slot update_slots: id 0 | task 111 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
144
+ 0.31.214.234 W slot update_slots: id 0 | task 111 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
145
+ 0.31.741.413 I slot create_check: id 0 | task 111 | created context checkpoint 1 of 32 (pos_min = 420, pos_max = 420, n_tokens = 421, size = 62.813 MiB)
146
+ 0.32.002.855 I reasoning-budget: deactivated (natural end)
147
+ 0.32.103.990 I slot print_timing: id 0 | task 111 |
148
+ prompt eval time = 558.28 ms / 425 tokens ( 1.31 ms per token, 761.27 tokens per second)
149
+ eval time = 331.47 ms / 17 tokens ( 19.50 ms per token, 51.29 tokens per second)
150
+ total time = 889.75 ms / 442 tokens
151
+ 0.32.104.074 I slot release: id 0 | task 111 | stop processing: n_tokens = 441, truncated = 0
152
+ 0.32.104.109 I srv update_slots: all slots are idle
153
+ 0.32.155.224 I srv params_from_: Chat format: peg-native
154
+ 0.32.157.476 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.967 (> 0.100 thold), f_keep = 0.918
155
+ 0.32.157.999 I reasoning-budget: activated, budget=2147483647 tokens
156
+ 0.32.158.095 I slot launch_slot_: id 0 | task 130 | processing task, is_child = 0
157
+ 0.32.158.121 W slot update_slots: id 0 | task 130 | n_past = 405, slot.prompt.tokens.size() = 441, seq_id = 0, pos_min = 440, n_swa = 0
158
+ 0.32.158.123 I slot update_slots: id 0 | task 130 | Checking checkpoint with [420, 420] against 405...
159
+ 0.32.158.124 W slot update_slots: id 0 | task 130 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
160
+ 0.32.158.130 W slot update_slots: id 0 | task 130 | erased invalidated context checkpoint (pos_min = 420, pos_max = 420, n_tokens = 421, n_swa = 0, pos_next = 0, size = 62.813 MiB)
161
+ 0.32.714.547 I slot create_check: id 0 | task 130 | created context checkpoint 1 of 32 (pos_min = 414, pos_max = 414, n_tokens = 415, size = 62.813 MiB)
162
+ 0.32.960.511 I reasoning-budget: deactivated (natural end)
163
+ 0.33.017.861 I slot print_timing: id 0 | task 130 |
164
+ prompt eval time = 601.42 ms / 419 tokens ( 1.44 ms per token, 696.69 tokens per second)
165
+ eval time = 258.30 ms / 12 tokens ( 21.53 ms per token, 46.46 tokens per second)
166
+ total time = 859.72 ms / 431 tokens
167
+ 0.33.018.052 I slot release: id 0 | task 130 | stop processing: n_tokens = 430, truncated = 0
168
+ 0.33.018.108 I srv update_slots: all slots are idle
169
+ 0.33.050.367 I srv params_from_: Chat format: peg-native
170
+ 0.33.050.835 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.960 (> 0.100 thold), f_keep = 0.942
171
+ 0.33.051.413 I reasoning-budget: activated, budget=2147483647 tokens
172
+ 0.33.051.498 I slot launch_slot_: id 0 | task 144 | processing task, is_child = 0
173
+ 0.33.051.517 W slot update_slots: id 0 | task 144 | n_past = 405, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
174
+ 0.33.051.521 I slot update_slots: id 0 | task 144 | Checking checkpoint with [414, 414] against 405...
175
+ 0.33.051.522 W slot update_slots: id 0 | task 144 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
176
+ 0.33.051.528 W slot update_slots: id 0 | task 144 | erased invalidated context checkpoint (pos_min = 414, pos_max = 414, n_tokens = 415, n_swa = 0, pos_next = 0, size = 62.813 MiB)
177
+ 0.33.615.892 I slot create_check: id 0 | task 144 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
178
+ 0.33.916.408 I reasoning-budget: deactivated (natural end)
179
+ 0.34.765.727 I slot print_timing: id 0 | task 144 |
180
+ prompt eval time = 597.52 ms / 422 tokens ( 1.42 ms per token, 706.25 tokens per second)
181
+ eval time = 1116.68 ms / 53 tokens ( 21.07 ms per token, 47.46 tokens per second)
182
+ total time = 1714.20 ms / 475 tokens
183
+ 0.34.765.808 I slot release: id 0 | task 144 | stop processing: n_tokens = 474, truncated = 0
184
+ 0.34.765.835 I srv update_slots: all slots are idle
185
+ 0.34.781.504 I srv params_from_: Chat format: peg-native
186
+ 0.34.781.937 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.958 (> 0.100 thold), f_keep = 0.857
187
+ 0.34.782.155 I reasoning-budget: activated, budget=2147483647 tokens
188
+ 0.34.782.201 I slot launch_slot_: id 0 | task 199 | processing task, is_child = 0
189
+ 0.34.782.213 W slot update_slots: id 0 | task 199 | n_past = 406, slot.prompt.tokens.size() = 474, seq_id = 0, pos_min = 473, n_swa = 0
190
+ 0.34.782.213 I slot update_slots: id 0 | task 199 | Checking checkpoint with [417, 417] against 406...
191
+ 0.34.782.214 W slot update_slots: id 0 | task 199 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
192
+ 0.34.782.218 W slot update_slots: id 0 | task 199 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
193
+ 0.35.312.464 I slot create_check: id 0 | task 199 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
194
+ 0.35.382.494 I reasoning-budget: deactivated (natural end)
195
+ 0.35.500.672 I slot print_timing: id 0 | task 199 |
196
+ prompt eval time = 569.17 ms / 424 tokens ( 1.34 ms per token, 744.95 tokens per second)
197
+ eval time = 149.28 ms / 7 tokens ( 21.33 ms per token, 46.89 tokens per second)
198
+ total time = 718.45 ms / 431 tokens
199
+ 0.35.500.758 I slot release: id 0 | task 199 | stop processing: n_tokens = 430, truncated = 0
200
+ 0.35.500.792 I srv update_slots: all slots are idle
201
+ 0.35.523.836 I srv params_from_: Chat format: peg-native
202
+ 0.35.524.275 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.969 (> 0.100 thold), f_keep = 0.942
203
+ 0.35.524.794 I reasoning-budget: activated, budget=2147483647 tokens
204
+ 0.35.524.873 I slot launch_slot_: id 0 | task 208 | processing task, is_child = 0
205
+ 0.35.524.891 W slot update_slots: id 0 | task 208 | n_past = 405, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
206
+ 0.35.524.892 I slot update_slots: id 0 | task 208 | Checking checkpoint with [419, 419] against 405...
207
+ 0.35.524.894 W slot update_slots: id 0 | task 208 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
208
+ 0.35.524.899 W slot update_slots: id 0 | task 208 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
209
+ 0.36.142.960 I slot create_check: id 0 | task 208 | created context checkpoint 1 of 32 (pos_min = 413, pos_max = 413, n_tokens = 414, size = 62.813 MiB)
210
+ 0.36.260.514 I reasoning-budget: deactivated (natural end)
211
+ 0.36.391.695 I slot print_timing: id 0 | task 208 |
212
+ prompt eval time = 684.84 ms / 418 tokens ( 1.64 ms per token, 610.36 tokens per second)
213
+ eval time = 181.96 ms / 5 tokens ( 36.39 ms per token, 27.48 tokens per second)
214
+ total time = 866.79 ms / 423 tokens
215
+ 0.36.391.789 I slot release: id 0 | task 208 | stop processing: n_tokens = 422, truncated = 0
216
+ 0.36.391.820 I srv update_slots: all slots are idle
217
+ 0.36.404.024 I srv params_from_: Chat format: peg-native
218
+ 0.36.404.587 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.960
219
+ 0.36.404.848 I reasoning-budget: activated, budget=2147483647 tokens
220
+ 0.36.404.897 I slot launch_slot_: id 0 | task 215 | processing task, is_child = 0
221
+ 0.36.404.910 W slot update_slots: id 0 | task 215 | n_past = 405, slot.prompt.tokens.size() = 422, seq_id = 0, pos_min = 421, n_swa = 0
222
+ 0.36.404.912 I slot update_slots: id 0 | task 215 | Checking checkpoint with [413, 413] against 405...
223
+ 0.36.404.913 W slot update_slots: id 0 | task 215 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
224
+ 0.36.404.916 W slot update_slots: id 0 | task 215 | erased invalidated context checkpoint (pos_min = 413, pos_max = 413, n_tokens = 414, n_swa = 0, pos_next = 0, size = 62.813 MiB)
225
+ 0.36.899.390 I slot create_check: id 0 | task 215 | created context checkpoint 1 of 32 (pos_min = 416, pos_max = 416, n_tokens = 417, size = 62.813 MiB)
226
+ 0.37.012.108 I reasoning-budget: deactivated (natural end)
227
+ 0.37.909.549 I slot print_timing: id 0 | task 215 |
228
+ prompt eval time = 555.94 ms / 421 tokens ( 1.32 ms per token, 757.28 tokens per second)
229
+ eval time = 948.68 ms / 42 tokens ( 22.59 ms per token, 44.27 tokens per second)
230
+ total time = 1504.62 ms / 463 tokens
231
+ 0.37.909.654 I slot release: id 0 | task 215 | stop processing: n_tokens = 462, truncated = 0
232
+ 0.37.909.691 I srv update_slots: all slots are idle
233
+ 0.37.937.856 I srv params_from_: Chat format: peg-native
234
+ 0.37.938.240 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.951 (> 0.100 thold), f_keep = 0.879
235
+ 0.37.938.471 I reasoning-budget: activated, budget=2147483647 tokens
236
+ 0.37.938.517 I slot launch_slot_: id 0 | task 259 | processing task, is_child = 0
237
+ 0.37.938.529 W slot update_slots: id 0 | task 259 | n_past = 406, slot.prompt.tokens.size() = 462, seq_id = 0, pos_min = 461, n_swa = 0
238
+ 0.37.938.531 I slot update_slots: id 0 | task 259 | Checking checkpoint with [416, 416] against 406...
239
+ 0.37.938.532 W slot update_slots: id 0 | task 259 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
240
+ 0.37.938.535 W slot update_slots: id 0 | task 259 | erased invalidated context checkpoint (pos_min = 416, pos_max = 416, n_tokens = 417, n_swa = 0, pos_next = 0, size = 62.813 MiB)
241
+ 0.38.478.969 I slot create_check: id 0 | task 259 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
242
+ 0.38.599.473 I slot print_timing: id 0 | task 259 |
243
+ prompt eval time = 580.20 ms / 427 tokens ( 1.36 ms per token, 735.95 tokens per second)
244
+ eval time = 80.65 ms / 4 tokens ( 20.16 ms per token, 49.60 tokens per second)
245
+ total time = 660.85 ms / 431 tokens
246
+ 0.38.599.762 I slot release: id 0 | task 259 | stop processing: n_tokens = 430, truncated = 0
247
+ 0.38.599.853 I srv update_slots: all slots are idle
248
+ 0.38.629.562 I srv params_from_: Chat format: peg-native
249
+ 0.38.630.012 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.942
250
+ 0.38.630.609 I reasoning-budget: activated, budget=2147483647 tokens
251
+ 0.38.630.693 I slot launch_slot_: id 0 | task 265 | processing task, is_child = 0
252
+ 0.38.630.711 W slot update_slots: id 0 | task 265 | n_past = 405, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
253
+ 0.38.630.716 I slot update_slots: id 0 | task 265 | Checking checkpoint with [422, 422] against 405...
254
+ 0.38.630.719 W slot update_slots: id 0 | task 265 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
255
+ 0.38.630.724 W slot update_slots: id 0 | task 265 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
256
+ 0.39.188.051 I slot create_check: id 0 | task 265 | created context checkpoint 1 of 32 (pos_min = 416, pos_max = 416, n_tokens = 417, size = 62.813 MiB)
257
+ 0.39.260.024 I slot print_timing: id 0 | task 265 |
258
+ prompt eval time = 591.28 ms / 421 tokens ( 1.40 ms per token, 712.01 tokens per second)
259
+ eval time = 37.99 ms / 2 tokens ( 19.00 ms per token, 52.64 tokens per second)
260
+ total time = 629.28 ms / 423 tokens
261
+ 0.39.260.276 I slot release: id 0 | task 265 | stop processing: n_tokens = 422, truncated = 0
262
+ 0.39.260.338 I srv update_slots: all slots are idle
263
+ 0.39.312.544 I srv params_from_: Chat format: peg-native
264
+ 0.39.314.833 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.960
265
+ 0.39.315.431 I reasoning-budget: activated, budget=2147483647 tokens
266
+ 0.39.315.516 I slot launch_slot_: id 0 | task 269 | processing task, is_child = 0
267
+ 0.39.315.539 W slot update_slots: id 0 | task 269 | n_past = 405, slot.prompt.tokens.size() = 422, seq_id = 0, pos_min = 421, n_swa = 0
268
+ 0.39.315.543 I slot update_slots: id 0 | task 269 | Checking checkpoint with [416, 416] against 405...
269
+ 0.39.315.545 W slot update_slots: id 0 | task 269 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
270
+ 0.39.315.552 W slot update_slots: id 0 | task 269 | erased invalidated context checkpoint (pos_min = 416, pos_max = 416, n_tokens = 417, n_swa = 0, pos_next = 0, size = 62.813 MiB)
271
+ 0.39.893.482 I slot create_check: id 0 | task 269 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
272
+ 0.40.793.357 I slot print_timing: id 0 | task 269 |
273
+ prompt eval time = 633.79 ms / 424 tokens ( 1.49 ms per token, 668.99 tokens per second)
274
+ eval time = 844.02 ms / 39 tokens ( 21.64 ms per token, 46.21 tokens per second)
275
+ total time = 1477.81 ms / 463 tokens
276
+ 0.40.793.457 I slot release: id 0 | task 269 | stop processing: n_tokens = 462, truncated = 0
277
+ 0.40.793.487 I srv update_slots: all slots are idle
278
+ 0.40.794.731 I srv operator(): operator(): cleaning up before exit...
recipe/logs/b_n-tools-q106-tpl-medium.log ADDED
@@ -0,0 +1,294 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0.00.105.454 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
2
+ 0.00.105.464 I device_info:
3
+ 0.00.105.601 I - ROCm0 : AMD Radeon Graphics (131072 MiB, 122408 MiB free)
4
+ 0.00.105.795 I - Vulkan0 : AMD Radeon Graphics (RADV GFX1151) (132096 MiB, 131922 MiB free)
5
+ 0.00.105.803 I - CPU : AMD RYZEN AI MAX+ 395 w/ Radeon 8060S (127438 MiB, 127438 MiB free)
6
+ 0.00.105.902 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
7
+ 0.00.105.988 I srv init: running without SSL
8
+ 0.00.106.040 I srv init: using 31 threads for HTTP server
9
+ 0.00.106.042 I srv init: the WebUI is disabled
10
+ 0.00.106.129 I srv start: binding port with default address family
11
+ 0.00.107.354 I srv main: loading model
12
+ 0.00.107.360 I srv load_model: loading model '/mnt/models/nex-n2.5-mini/out/Nex-N2.5-mini-Q4_0_ROCMFP4_STRIX_LEAN.gguf'
13
+ 0.00.146.109 W llama_model_loader: direct I/O is enabled, disabling mmap
14
+ 0.23.588.699 W llama_context: n_ctx_seq (65536) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
15
+ 0.23.912.034 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
16
+ 0.24.349.867 I srv load_model: initializing slots, n_slots = 1
17
+ 0.24.567.364 W srv load_model: speculative decoding will use checkpoints
18
+ 0.24.567.384 W common_speculative_init: no implementations specified for speculative decoding
19
+ 0.24.567.388 I slot load_model: id 0 | task -1 | new slot, n_ctx = 65536
20
+ 0.24.567.468 I srv load_model: prompt cache RAM enabled: limit_mib=8192
21
+ 0.24.567.470 I srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
22
+ 0.24.567.551 I srv init: idle slots will be saved to prompt cache upon starting a new task
23
+ 0.24.579.260 I init: chat template, example_format: '<|im_start|>system
24
+ You are a helpful assistant<|im_end|>
25
+ <|im_start|>user
26
+ Hello<|im_end|>
27
+ <|im_start|>assistant
28
+ <think>
29
+
30
+ </think>
31
+
32
+ Hi there<|im_end|>
33
+ <|im_start|>user
34
+ How are you?<|im_end|>
35
+ <|im_start|>assistant
36
+ <think>'
37
+ 0.24.587.661 I srv init: init: chat template, thinking = 1
38
+ 0.24.587.712 I srv main: model loaded
39
+ 0.24.587.715 I srv main: server is listening on http://127.0.0.1:18600
40
+ 0.24.587.720 I srv update_slots: all slots are idle
41
+ 0.26.148.362 I srv params_from_: Chat format: peg-native
42
+ 0.26.149.988 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1
43
+ 0.26.149.996 I srv get_availabl: updating prompt cache
44
+ 0.26.150.007 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000
45
+ 0.26.150.015 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 65536 tokens, 8589934592 est)
46
+ 0.26.150.018 I srv get_availabl: prompt cache update took 0.02 ms
47
+ 0.26.150.786 I reasoning-budget: activated, budget=2147483647 tokens
48
+ 0.26.150.816 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0
49
+ 0.26.871.165 I slot create_check: id 0 | task 0 | created context checkpoint 1 of 32 (pos_min = 416, pos_max = 416, n_tokens = 417, size = 62.813 MiB)
50
+ 0.27.133.131 I reasoning-budget: deactivated (natural end)
51
+ 0.28.001.263 I slot print_timing: id 0 | task 0 |
52
+ prompt eval time = 773.86 ms / 421 tokens ( 1.84 ms per token, 544.03 tokens per second)
53
+ eval time = 1076.54 ms / 48 tokens ( 22.43 ms per token, 44.59 tokens per second)
54
+ total time = 1850.39 ms / 469 tokens
55
+ 0.28.001.472 I slot release: id 0 | task 0 | stop processing: n_tokens = 468, truncated = 0
56
+ 0.28.001.492 I srv update_slots: all slots are idle
57
+ 0.28.065.185 I srv params_from_: Chat format: peg-native
58
+ 0.28.065.968 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.908 (> 0.100 thold), f_keep = 0.865
59
+ 0.28.066.222 I reasoning-budget: activated, budget=2147483647 tokens
60
+ 0.28.066.277 I slot launch_slot_: id 0 | task 50 | processing task, is_child = 0
61
+ 0.28.066.290 W slot update_slots: id 0 | task 50 | n_past = 405, slot.prompt.tokens.size() = 468, seq_id = 0, pos_min = 467, n_swa = 0
62
+ 0.28.066.292 I slot update_slots: id 0 | task 50 | Checking checkpoint with [416, 416] against 405...
63
+ 0.28.066.293 W slot update_slots: id 0 | task 50 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
64
+ 0.28.066.297 W slot update_slots: id 0 | task 50 | erased invalidated context checkpoint (pos_min = 416, pos_max = 416, n_tokens = 417, n_swa = 0, pos_next = 0, size = 62.813 MiB)
65
+ 0.28.663.976 I slot create_check: id 0 | task 50 | created context checkpoint 1 of 32 (pos_min = 441, pos_max = 441, n_tokens = 442, size = 62.813 MiB)
66
+ 0.29.408.376 I reasoning-budget: deactivated (natural end)
67
+ 0.30.568.909 I slot print_timing: id 0 | task 50 |
68
+ prompt eval time = 632.88 ms / 446 tokens ( 1.42 ms per token, 704.71 tokens per second)
69
+ eval time = 1869.72 ms / 81 tokens ( 23.08 ms per token, 43.32 tokens per second)
70
+ total time = 2502.60 ms / 527 tokens
71
+ 0.30.568.986 I slot release: id 0 | task 50 | stop processing: n_tokens = 526, truncated = 0
72
+ 0.30.569.017 I srv update_slots: all slots are idle
73
+ 0.30.584.452 I srv params_from_: Chat format: peg-native
74
+ 0.30.584.982 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.770
75
+ 0.30.585.566 I reasoning-budget: activated, budget=2147483647 tokens
76
+ 0.30.585.633 I slot launch_slot_: id 0 | task 133 | processing task, is_child = 0
77
+ 0.30.585.660 W slot update_slots: id 0 | task 133 | n_past = 405, slot.prompt.tokens.size() = 526, seq_id = 0, pos_min = 525, n_swa = 0
78
+ 0.30.585.663 I slot update_slots: id 0 | task 133 | Checking checkpoint with [441, 441] against 405...
79
+ 0.30.585.666 W slot update_slots: id 0 | task 133 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
80
+ 0.30.585.671 W slot update_slots: id 0 | task 133 | erased invalidated context checkpoint (pos_min = 441, pos_max = 441, n_tokens = 442, n_swa = 0, pos_next = 0, size = 62.813 MiB)
81
+ 0.31.126.980 I slot create_check: id 0 | task 133 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
82
+ 0.31.358.672 I reasoning-budget: deactivated (natural end)
83
+ 0.32.155.583 I slot print_timing: id 0 | task 133 |
84
+ prompt eval time = 578.11 ms / 424 tokens ( 1.36 ms per token, 733.42 tokens per second)
85
+ eval time = 991.80 ms / 48 tokens ( 20.66 ms per token, 48.40 tokens per second)
86
+ total time = 1569.91 ms / 472 tokens
87
+ 0.32.155.659 I slot release: id 0 | task 133 | stop processing: n_tokens = 471, truncated = 0
88
+ 0.32.155.694 I srv update_slots: all slots are idle
89
+ 0.32.187.576 I srv params_from_: Chat format: peg-native
90
+ 0.32.188.058 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.860
91
+ 0.32.188.654 I reasoning-budget: activated, budget=2147483647 tokens
92
+ 0.32.188.751 I slot launch_slot_: id 0 | task 183 | processing task, is_child = 0
93
+ 0.32.188.776 W slot update_slots: id 0 | task 183 | n_past = 405, slot.prompt.tokens.size() = 471, seq_id = 0, pos_min = 470, n_swa = 0
94
+ 0.32.188.780 I slot update_slots: id 0 | task 183 | Checking checkpoint with [419, 419] against 405...
95
+ 0.32.188.782 W slot update_slots: id 0 | task 183 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
96
+ 0.32.188.790 W slot update_slots: id 0 | task 183 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
97
+ 0.32.763.234 I slot create_check: id 0 | task 183 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
98
+ 0.32.827.843 I reasoning-budget: deactivated (natural end)
99
+ 0.32.942.756 I slot print_timing: id 0 | task 183 |
100
+ prompt eval time = 610.58 ms / 424 tokens ( 1.44 ms per token, 694.43 tokens per second)
101
+ eval time = 143.38 ms / 7 tokens ( 20.48 ms per token, 48.82 tokens per second)
102
+ total time = 753.96 ms / 431 tokens
103
+ 0.32.942.940 I slot release: id 0 | task 183 | stop processing: n_tokens = 430, truncated = 0
104
+ 0.32.942.992 I srv update_slots: all slots are idle
105
+ 0.32.959.301 I srv params_from_: Chat format: peg-native
106
+ 0.32.959.816 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.964 (> 0.100 thold), f_keep = 0.944
107
+ 0.32.960.404 I reasoning-budget: activated, budget=2147483647 tokens
108
+ 0.32.960.503 I slot launch_slot_: id 0 | task 192 | processing task, is_child = 0
109
+ 0.32.960.525 W slot update_slots: id 0 | task 192 | n_past = 406, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
110
+ 0.32.960.529 I slot update_slots: id 0 | task 192 | Checking checkpoint with [419, 419] against 406...
111
+ 0.32.960.530 W slot update_slots: id 0 | task 192 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
112
+ 0.32.960.538 W slot update_slots: id 0 | task 192 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
113
+ 0.33.460.168 I slot create_check: id 0 | task 192 | created context checkpoint 1 of 32 (pos_min = 416, pos_max = 416, n_tokens = 417, size = 62.813 MiB)
114
+ 0.33.682.213 I reasoning-budget: deactivated (natural end)
115
+ 0.34.478.330 I slot print_timing: id 0 | task 192 |
116
+ prompt eval time = 535.13 ms / 421 tokens ( 1.27 ms per token, 786.73 tokens per second)
117
+ eval time = 982.66 ms / 50 tokens ( 19.65 ms per token, 50.88 tokens per second)
118
+ total time = 1517.79 ms / 471 tokens
119
+ 0.34.478.426 I slot release: id 0 | task 192 | stop processing: n_tokens = 470, truncated = 0
120
+ 0.34.478.463 I srv update_slots: all slots are idle
121
+ 0.34.496.222 I srv params_from_: Chat format: peg-native
122
+ 0.34.496.734 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.942 (> 0.100 thold), f_keep = 1.000
123
+ 0.34.496.948 I reasoning-budget: activated, budget=2147483647 tokens
124
+ 0.34.496.989 I slot launch_slot_: id 0 | task 244 | processing task, is_child = 0
125
+ 0.34.639.263 I slot create_check: id 0 | task 244 | created context checkpoint 2 of 32 (pos_min = 494, pos_max = 494, n_tokens = 495, size = 62.813 MiB)
126
+ 0.34.705.972 I reasoning-budget: deactivated (natural end)
127
+ 0.35.076.072 I slot print_timing: id 0 | task 244 |
128
+ prompt eval time = 177.70 ms / 29 tokens ( 6.13 ms per token, 163.19 tokens per second)
129
+ eval time = 401.35 ms / 17 tokens ( 23.61 ms per token, 42.36 tokens per second)
130
+ total time = 579.05 ms / 46 tokens
131
+ 0.35.076.179 I slot release: id 0 | task 244 | stop processing: n_tokens = 515, truncated = 0
132
+ 0.35.076.214 I srv update_slots: all slots are idle
133
+ 0.35.091.806 I srv params_from_: Chat format: peg-native
134
+ 0.35.092.529 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.974 (> 0.100 thold), f_keep = 0.796
135
+ 0.35.093.006 I reasoning-budget: activated, budget=2147483647 tokens
136
+ 0.35.093.087 I slot launch_slot_: id 0 | task 263 | processing task, is_child = 0
137
+ 0.35.093.108 W slot update_slots: id 0 | task 263 | n_past = 410, slot.prompt.tokens.size() = 515, seq_id = 0, pos_min = 514, n_swa = 0
138
+ 0.35.093.110 I slot update_slots: id 0 | task 263 | Checking checkpoint with [494, 494] against 410...
139
+ 0.35.093.111 I slot update_slots: id 0 | task 263 | Checking checkpoint with [416, 416] against 410...
140
+ 0.35.093.113 W slot update_slots: id 0 | task 263 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
141
+ 0.35.093.119 W slot update_slots: id 0 | task 263 | erased invalidated context checkpoint (pos_min = 416, pos_max = 416, n_tokens = 417, n_swa = 0, pos_next = 0, size = 62.813 MiB)
142
+ 0.35.094.217 W slot update_slots: id 0 | task 263 | erased invalidated context checkpoint (pos_min = 494, pos_max = 494, n_tokens = 495, n_swa = 0, pos_next = 0, size = 62.813 MiB)
143
+ 0.35.626.770 I slot create_check: id 0 | task 263 | created context checkpoint 1 of 32 (pos_min = 416, pos_max = 416, n_tokens = 417, size = 62.813 MiB)
144
+ 0.36.142.008 I reasoning-budget: deactivated (natural end)
145
+ 0.36.985.537 I slot print_timing: id 0 | task 263 |
146
+ prompt eval time = 581.67 ms / 421 tokens ( 1.38 ms per token, 723.78 tokens per second)
147
+ eval time = 1310.72 ms / 64 tokens ( 20.48 ms per token, 48.83 tokens per second)
148
+ total time = 1892.39 ms / 485 tokens
149
+ 0.36.985.731 I slot release: id 0 | task 263 | stop processing: n_tokens = 484, truncated = 0
150
+ 0.36.985.791 I srv update_slots: all slots are idle
151
+ 0.37.040.565 I srv params_from_: Chat format: peg-native
152
+ 0.37.041.116 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.938 (> 0.100 thold), f_keep = 0.837
153
+ 0.37.041.402 I reasoning-budget: activated, budget=2147483647 tokens
154
+ 0.37.041.512 I slot launch_slot_: id 0 | task 329 | processing task, is_child = 0
155
+ 0.37.041.525 W slot update_slots: id 0 | task 329 | n_past = 405, slot.prompt.tokens.size() = 484, seq_id = 0, pos_min = 483, n_swa = 0
156
+ 0.37.041.528 I slot update_slots: id 0 | task 329 | Checking checkpoint with [416, 416] against 405...
157
+ 0.37.041.528 W slot update_slots: id 0 | task 329 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
158
+ 0.37.041.531 W slot update_slots: id 0 | task 329 | erased invalidated context checkpoint (pos_min = 416, pos_max = 416, n_tokens = 417, n_swa = 0, pos_next = 0, size = 62.813 MiB)
159
+ 0.37.608.593 I slot create_check: id 0 | task 329 | created context checkpoint 1 of 32 (pos_min = 427, pos_max = 427, n_tokens = 428, size = 62.813 MiB)
160
+ 0.38.415.584 I reasoning-budget: deactivated (natural end)
161
+ 0.40.014.106 I slot print_timing: id 0 | task 329 | n_decoded = 100, tg = 42.63 t/s
162
+ 0.40.267.618 I slot print_timing: id 0 | task 329 |
163
+ prompt eval time = 626.94 ms / 432 tokens ( 1.45 ms per token, 689.06 tokens per second)
164
+ eval time = 2599.14 ms / 113 tokens ( 23.00 ms per token, 43.48 tokens per second)
165
+ total time = 3226.08 ms / 545 tokens
166
+ 0.40.267.691 I slot release: id 0 | task 329 | stop processing: n_tokens = 544, truncated = 0
167
+ 0.40.267.739 I srv update_slots: all slots are idle
168
+ 0.40.291.843 I srv params_from_: Chat format: peg-native
169
+ 0.40.292.295 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.744
170
+ 0.40.292.805 I reasoning-budget: activated, budget=2147483647 tokens
171
+ 0.40.292.809 I reasoning-budget: deactivated (natural end)
172
+ 0.40.292.882 I slot launch_slot_: id 0 | task 444 | processing task, is_child = 0
173
+ 0.40.292.903 W slot update_slots: id 0 | task 444 | n_past = 405, slot.prompt.tokens.size() = 544, seq_id = 0, pos_min = 543, n_swa = 0
174
+ 0.40.292.906 I slot update_slots: id 0 | task 444 | Checking checkpoint with [427, 427] against 405...
175
+ 0.40.292.907 W slot update_slots: id 0 | task 444 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
176
+ 0.40.292.912 W slot update_slots: id 0 | task 444 | erased invalidated context checkpoint (pos_min = 427, pos_max = 427, n_tokens = 428, n_swa = 0, pos_next = 0, size = 62.813 MiB)
177
+ 0.40.808.790 I slot create_check: id 0 | task 444 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
178
+ 0.41.645.500 I slot print_timing: id 0 | task 444 |
179
+ prompt eval time = 551.51 ms / 424 tokens ( 1.30 ms per token, 768.79 tokens per second)
180
+ eval time = 801.07 ms / 39 tokens ( 20.54 ms per token, 48.68 tokens per second)
181
+ total time = 1352.58 ms / 463 tokens
182
+ 0.41.645.589 I slot release: id 0 | task 444 | stop processing: n_tokens = 462, truncated = 0
183
+ 0.41.645.620 I srv update_slots: all slots are idle
184
+ 0.41.661.261 I srv params_from_: Chat format: peg-native
185
+ 0.41.661.789 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.902 (> 0.100 thold), f_keep = 0.877
186
+ 0.41.662.003 I reasoning-budget: activated, budget=2147483647 tokens
187
+ 0.41.662.005 I reasoning-budget: deactivated (natural end)
188
+ 0.41.662.045 I slot launch_slot_: id 0 | task 485 | processing task, is_child = 0
189
+ 0.41.662.055 W slot update_slots: id 0 | task 485 | n_past = 405, slot.prompt.tokens.size() = 462, seq_id = 0, pos_min = 461, n_swa = 0
190
+ 0.41.662.056 I slot update_slots: id 0 | task 485 | Checking checkpoint with [419, 419] against 405...
191
+ 0.41.662.057 W slot update_slots: id 0 | task 485 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
192
+ 0.41.662.060 W slot update_slots: id 0 | task 485 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
193
+ 0.42.169.075 I slot create_check: id 0 | task 485 | created context checkpoint 1 of 32 (pos_min = 444, pos_max = 444, n_tokens = 445, size = 62.813 MiB)
194
+ 0.44.038.103 I slot print_timing: id 0 | task 485 |
195
+ prompt eval time = 548.61 ms / 449 tokens ( 1.22 ms per token, 818.43 tokens per second)
196
+ eval time = 1827.42 ms / 87 tokens ( 21.00 ms per token, 47.61 tokens per second)
197
+ total time = 2376.03 ms / 536 tokens
198
+ 0.44.038.196 I slot release: id 0 | task 485 | stop processing: n_tokens = 535, truncated = 0
199
+ 0.44.038.227 I srv update_slots: all slots are idle
200
+ 0.44.070.792 I srv params_from_: Chat format: peg-native
201
+ 0.44.071.263 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.948 (> 0.100 thold), f_keep = 0.757
202
+ 0.44.071.950 I reasoning-budget: activated, budget=2147483647 tokens
203
+ 0.44.071.955 I reasoning-budget: deactivated (natural end)
204
+ 0.44.072.053 I slot launch_slot_: id 0 | task 574 | processing task, is_child = 0
205
+ 0.44.072.077 W slot update_slots: id 0 | task 574 | n_past = 405, slot.prompt.tokens.size() = 535, seq_id = 0, pos_min = 534, n_swa = 0
206
+ 0.44.072.081 I slot update_slots: id 0 | task 574 | Checking checkpoint with [444, 444] against 405...
207
+ 0.44.072.083 W slot update_slots: id 0 | task 574 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
208
+ 0.44.072.088 W slot update_slots: id 0 | task 574 | erased invalidated context checkpoint (pos_min = 444, pos_max = 444, n_tokens = 445, n_swa = 0, pos_next = 0, size = 62.813 MiB)
209
+ 0.44.616.066 I slot create_check: id 0 | task 574 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
210
+ 0.45.562.940 I slot print_timing: id 0 | task 574 |
211
+ prompt eval time = 603.10 ms / 427 tokens ( 1.41 ms per token, 708.01 tokens per second)
212
+ eval time = 887.75 ms / 39 tokens ( 22.76 ms per token, 43.93 tokens per second)
213
+ total time = 1490.84 ms / 466 tokens
214
+ 0.45.563.042 I slot release: id 0 | task 574 | stop processing: n_tokens = 465, truncated = 0
215
+ 0.45.563.082 I srv update_slots: all slots are idle
216
+ 0.45.579.534 I srv params_from_: Chat format: peg-native
217
+ 0.45.580.094 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.948 (> 0.100 thold), f_keep = 0.871
218
+ 0.45.580.642 I reasoning-budget: activated, budget=2147483647 tokens
219
+ 0.45.580.646 I reasoning-budget: deactivated (natural end)
220
+ 0.45.582.811 I slot launch_slot_: id 0 | task 615 | processing task, is_child = 0
221
+ 0.45.582.848 W slot update_slots: id 0 | task 615 | n_past = 405, slot.prompt.tokens.size() = 465, seq_id = 0, pos_min = 464, n_swa = 0
222
+ 0.45.582.849 I slot update_slots: id 0 | task 615 | Checking checkpoint with [422, 422] against 405...
223
+ 0.45.582.851 W slot update_slots: id 0 | task 615 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
224
+ 0.45.582.860 W slot update_slots: id 0 | task 615 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
225
+ 0.46.041.817 I slot create_check: id 0 | task 615 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
226
+ 0.46.171.798 I slot print_timing: id 0 | task 615 |
227
+ prompt eval time = 507.02 ms / 427 tokens ( 1.19 ms per token, 842.17 tokens per second)
228
+ eval time = 81.92 ms / 4 tokens ( 20.48 ms per token, 48.83 tokens per second)
229
+ total time = 588.94 ms / 431 tokens
230
+ 0.46.171.890 I slot release: id 0 | task 615 | stop processing: n_tokens = 430, truncated = 0
231
+ 0.46.171.920 I srv update_slots: all slots are idle
232
+ 0.46.188.295 I srv params_from_: Chat format: peg-native
233
+ 0.46.188.729 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.958 (> 0.100 thold), f_keep = 0.944
234
+ 0.46.189.261 I reasoning-budget: activated, budget=2147483647 tokens
235
+ 0.46.189.264 I reasoning-budget: deactivated (natural end)
236
+ 0.46.189.348 I slot launch_slot_: id 0 | task 621 | processing task, is_child = 0
237
+ 0.46.189.367 W slot update_slots: id 0 | task 621 | n_past = 406, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
238
+ 0.46.189.370 I slot update_slots: id 0 | task 621 | Checking checkpoint with [422, 422] against 406...
239
+ 0.46.189.372 W slot update_slots: id 0 | task 621 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
240
+ 0.46.189.377 W slot update_slots: id 0 | task 621 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
241
+ 0.46.716.804 I slot create_check: id 0 | task 621 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
242
+ 0.47.566.188 I slot print_timing: id 0 | task 621 |
243
+ prompt eval time = 563.24 ms / 424 tokens ( 1.33 ms per token, 752.78 tokens per second)
244
+ eval time = 813.55 ms / 40 tokens ( 20.34 ms per token, 49.17 tokens per second)
245
+ total time = 1376.79 ms / 464 tokens
246
+ 0.47.566.402 I slot release: id 0 | task 621 | stop processing: n_tokens = 463, truncated = 0
247
+ 0.47.566.467 I srv update_slots: all slots are idle
248
+ 0.47.589.246 I srv params_from_: Chat format: peg-native
249
+ 0.47.589.845 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.935 (> 0.100 thold), f_keep = 1.000
250
+ 0.47.590.510 I reasoning-budget: activated, budget=2147483647 tokens
251
+ 0.47.590.517 I reasoning-budget: deactivated (natural end)
252
+ 0.47.590.637 I slot launch_slot_: id 0 | task 663 | processing task, is_child = 0
253
+ 0.47.740.622 I slot create_check: id 0 | task 663 | created context checkpoint 2 of 32 (pos_min = 490, pos_max = 490, n_tokens = 491, size = 62.813 MiB)
254
+ 0.48.115.507 I slot print_timing: id 0 | task 663 |
255
+ prompt eval time = 197.73 ms / 32 tokens ( 6.18 ms per token, 161.84 tokens per second)
256
+ eval time = 327.11 ms / 15 tokens ( 21.81 ms per token, 45.86 tokens per second)
257
+ total time = 524.83 ms / 47 tokens
258
+ 0.48.115.614 I slot release: id 0 | task 663 | stop processing: n_tokens = 509, truncated = 0
259
+ 0.48.115.648 I srv update_slots: all slots are idle
260
+ 0.48.182.251 I srv params_from_: Chat format: peg-native
261
+ 0.48.182.770 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.967 (> 0.100 thold), f_keep = 0.806
262
+ 0.48.183.002 I reasoning-budget: activated, budget=2147483647 tokens
263
+ 0.48.183.003 I reasoning-budget: deactivated (natural end)
264
+ 0.48.183.055 I slot launch_slot_: id 0 | task 680 | processing task, is_child = 0
265
+ 0.48.183.066 W slot update_slots: id 0 | task 680 | n_past = 410, slot.prompt.tokens.size() = 509, seq_id = 0, pos_min = 508, n_swa = 0
266
+ 0.48.183.067 I slot update_slots: id 0 | task 680 | Checking checkpoint with [490, 490] against 410...
267
+ 0.48.183.067 I slot update_slots: id 0 | task 680 | Checking checkpoint with [419, 419] against 410...
268
+ 0.48.183.069 W slot update_slots: id 0 | task 680 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
269
+ 0.48.183.072 W slot update_slots: id 0 | task 680 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
270
+ 0.48.184.972 W slot update_slots: id 0 | task 680 | erased invalidated context checkpoint (pos_min = 490, pos_max = 490, n_tokens = 491, n_swa = 0, pos_next = 0, size = 62.813 MiB)
271
+ 0.48.742.624 I slot create_check: id 0 | task 680 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
272
+ 0.49.641.904 I slot print_timing: id 0 | task 680 |
273
+ prompt eval time = 599.46 ms / 424 tokens ( 1.41 ms per token, 707.31 tokens per second)
274
+ eval time = 859.36 ms / 40 tokens ( 21.48 ms per token, 46.55 tokens per second)
275
+ total time = 1458.81 ms / 464 tokens
276
+ 0.49.641.977 I slot release: id 0 | task 680 | stop processing: n_tokens = 463, truncated = 0
277
+ 0.49.642.005 I srv update_slots: all slots are idle
278
+ 0.49.654.625 I srv params_from_: Chat format: peg-native
279
+ 0.49.655.158 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.931 (> 0.100 thold), f_keep = 0.875
280
+ 0.49.655.408 I reasoning-budget: activated, budget=2147483647 tokens
281
+ 0.49.655.411 I reasoning-budget: deactivated (natural end)
282
+ 0.49.655.459 I slot launch_slot_: id 0 | task 722 | processing task, is_child = 0
283
+ 0.49.655.469 W slot update_slots: id 0 | task 722 | n_past = 405, slot.prompt.tokens.size() = 463, seq_id = 0, pos_min = 462, n_swa = 0
284
+ 0.49.655.470 I slot update_slots: id 0 | task 722 | Checking checkpoint with [419, 419] against 405...
285
+ 0.49.655.470 W slot update_slots: id 0 | task 722 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
286
+ 0.49.655.472 W slot update_slots: id 0 | task 722 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
287
+ 0.50.147.360 I slot create_check: id 0 | task 722 | created context checkpoint 1 of 32 (pos_min = 430, pos_max = 430, n_tokens = 431, size = 62.813 MiB)
288
+ 0.51.806.641 I slot print_timing: id 0 | task 722 |
289
+ prompt eval time = 536.58 ms / 435 tokens ( 1.23 ms per token, 810.69 tokens per second)
290
+ eval time = 1614.55 ms / 80 tokens ( 20.18 ms per token, 49.55 tokens per second)
291
+ total time = 2151.13 ms / 515 tokens
292
+ 0.51.807.033 I slot release: id 0 | task 722 | stop processing: n_tokens = 514, truncated = 0
293
+ 0.51.807.128 I srv update_slots: all slots are idle
294
+ 0.51.808.793 I srv operator(): operator(): cleaning up before exit...
recipe/logs/b_n-tools-q106.log ADDED
@@ -0,0 +1,302 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0.00.039.657 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
2
+ 0.00.039.660 I device_info:
3
+ 0.00.039.707 I - ROCm0 : AMD Radeon Graphics (131072 MiB, 125232 MiB free)
4
+ 0.00.039.793 I - Vulkan0 : AMD Radeon Graphics (RADV GFX1151) (132096 MiB, 131926 MiB free)
5
+ 0.00.039.797 I - CPU : AMD RYZEN AI MAX+ 395 w/ Radeon 8060S (127438 MiB, 127438 MiB free)
6
+ 0.00.039.840 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
7
+ 0.00.039.855 I srv init: running without SSL
8
+ 0.00.039.876 I srv init: using 31 threads for HTTP server
9
+ 0.00.039.877 I srv init: the WebUI is disabled
10
+ 0.00.039.934 I srv start: binding port with default address family
11
+ 0.00.041.085 I srv main: loading model
12
+ 0.00.041.087 I srv load_model: loading model '/mnt/models/nex-n2.5-mini/out/Nex-N2.5-mini-Q4_0_ROCMFP4_STRIX_LEAN.gguf'
13
+ 0.00.075.504 W llama_model_loader: direct I/O is enabled, disabling mmap
14
+ 0.20.409.072 W llama_context: n_ctx_seq (65536) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
15
+ 0.20.574.487 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
16
+ 0.20.774.364 I srv load_model: initializing slots, n_slots = 1
17
+ 0.20.940.846 W srv load_model: speculative decoding will use checkpoints
18
+ 0.20.940.851 W common_speculative_init: no implementations specified for speculative decoding
19
+ 0.20.940.852 I slot load_model: id 0 | task -1 | new slot, n_ctx = 65536
20
+ 0.20.940.886 I srv load_model: prompt cache RAM enabled: limit_mib=8192
21
+ 0.20.940.894 I srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
22
+ 0.20.940.907 I srv init: idle slots will be saved to prompt cache upon starting a new task
23
+ 0.20.950.043 I init: chat template, example_format: '<|im_start|>system
24
+ You are a helpful assistant<|im_end|>
25
+ <|im_start|>user
26
+ Hello<|im_end|>
27
+ <|im_start|>assistant
28
+ <think>
29
+
30
+ </think>
31
+
32
+ Hi there<|im_end|>
33
+ <|im_start|>user
34
+ How are you?<|im_end|>
35
+ <|im_start|>assistant
36
+ <think>'
37
+ 0.20.957.344 I srv init: init: chat template, thinking = 1
38
+ 0.20.957.368 I srv main: model loaded
39
+ 0.20.957.372 I srv main: server is listening on http://127.0.0.1:18600
40
+ 0.20.957.374 I srv update_slots: all slots are idle
41
+ 0.22.001.528 I srv params_from_: Chat format: peg-native
42
+ 0.22.001.842 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1
43
+ 0.22.001.845 I srv get_availabl: updating prompt cache
44
+ 0.22.001.851 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000
45
+ 0.22.001.856 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 65536 tokens, 8589934592 est)
46
+ 0.22.001.858 I srv get_availabl: prompt cache update took 0.01 ms
47
+ 0.22.010.813 I reasoning-budget: activated, budget=2147483647 tokens
48
+ 0.22.010.830 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0
49
+ 0.22.354.924 I slot create_check: id 0 | task 0 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
50
+ 0.22.608.502 I reasoning-budget: deactivated (natural end)
51
+ 0.23.202.843 I slot print_timing: id 0 | task 0 |
52
+ prompt eval time = 372.12 ms / 422 tokens ( 0.88 ms per token, 1134.04 tokens per second)
53
+ eval time = 819.87 ms / 55 tokens ( 14.91 ms per token, 67.08 tokens per second)
54
+ total time = 1191.99 ms / 477 tokens
55
+ 0.23.202.893 I slot release: id 0 | task 0 | stop processing: n_tokens = 476, truncated = 0
56
+ 0.23.202.897 I srv update_slots: all slots are idle
57
+ 0.23.213.207 I srv params_from_: Chat format: peg-native
58
+ 0.23.213.731 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.906 (> 0.100 thold), f_keep = 0.851
59
+ 0.23.213.930 I reasoning-budget: activated, budget=2147483647 tokens
60
+ 0.23.213.965 I slot launch_slot_: id 0 | task 57 | processing task, is_child = 0
61
+ 0.23.213.971 W slot update_slots: id 0 | task 57 | n_past = 405, slot.prompt.tokens.size() = 476, seq_id = 0, pos_min = 475, n_swa = 0
62
+ 0.23.213.971 I slot update_slots: id 0 | task 57 | Checking checkpoint with [417, 417] against 405...
63
+ 0.23.213.973 W slot update_slots: id 0 | task 57 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
64
+ 0.23.213.975 W slot update_slots: id 0 | task 57 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
65
+ 0.23.551.876 I slot create_check: id 0 | task 57 | created context checkpoint 1 of 32 (pos_min = 442, pos_max = 442, n_tokens = 443, size = 62.813 MiB)
66
+ 0.24.213.049 I reasoning-budget: deactivated (natural end)
67
+ 0.24.949.926 I slot print_timing: id 0 | task 57 |
68
+ prompt eval time = 365.53 ms / 447 tokens ( 0.82 ms per token, 1222.88 tokens per second)
69
+ eval time = 1370.41 ms / 92 tokens ( 14.90 ms per token, 67.13 tokens per second)
70
+ total time = 1735.95 ms / 539 tokens
71
+ 0.24.949.971 I slot release: id 0 | task 57 | stop processing: n_tokens = 538, truncated = 0
72
+ 0.24.949.992 I srv update_slots: all slots are idle
73
+ 0.24.959.419 I srv params_from_: Chat format: peg-native
74
+ 0.24.959.740 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.953 (> 0.100 thold), f_keep = 0.753
75
+ 0.24.959.952 I reasoning-budget: activated, budget=2147483647 tokens
76
+ 0.24.959.981 I slot launch_slot_: id 0 | task 151 | processing task, is_child = 0
77
+ 0.24.959.987 W slot update_slots: id 0 | task 151 | n_past = 405, slot.prompt.tokens.size() = 538, seq_id = 0, pos_min = 537, n_swa = 0
78
+ 0.24.959.987 I slot update_slots: id 0 | task 151 | Checking checkpoint with [442, 442] against 405...
79
+ 0.24.959.988 W slot update_slots: id 0 | task 151 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
80
+ 0.24.959.991 W slot update_slots: id 0 | task 151 | erased invalidated context checkpoint (pos_min = 442, pos_max = 442, n_tokens = 443, n_swa = 0, pos_next = 0, size = 62.813 MiB)
81
+ 0.25.284.086 I slot create_check: id 0 | task 151 | created context checkpoint 1 of 32 (pos_min = 420, pos_max = 420, n_tokens = 421, size = 62.813 MiB)
82
+ 0.25.471.508 I reasoning-budget: deactivated (natural end)
83
+ 0.26.064.107 I slot print_timing: id 0 | task 151 |
84
+ prompt eval time = 351.76 ms / 425 tokens ( 0.83 ms per token, 1208.20 tokens per second)
85
+ eval time = 752.35 ms / 51 tokens ( 14.75 ms per token, 67.79 tokens per second)
86
+ total time = 1104.11 ms / 476 tokens
87
+ 0.26.064.151 I slot release: id 0 | task 151 | stop processing: n_tokens = 475, truncated = 0
88
+ 0.26.064.171 I srv update_slots: all slots are idle
89
+ 0.26.074.471 I srv params_from_: Chat format: peg-native
90
+ 0.26.074.789 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.953 (> 0.100 thold), f_keep = 0.853
91
+ 0.26.074.999 I reasoning-budget: activated, budget=2147483647 tokens
92
+ 0.26.075.029 I slot launch_slot_: id 0 | task 204 | processing task, is_child = 0
93
+ 0.26.075.034 W slot update_slots: id 0 | task 204 | n_past = 405, slot.prompt.tokens.size() = 475, seq_id = 0, pos_min = 474, n_swa = 0
94
+ 0.26.075.045 I slot update_slots: id 0 | task 204 | Checking checkpoint with [420, 420] against 405...
95
+ 0.26.075.048 W slot update_slots: id 0 | task 204 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
96
+ 0.26.075.050 W slot update_slots: id 0 | task 204 | erased invalidated context checkpoint (pos_min = 420, pos_max = 420, n_tokens = 421, n_swa = 0, pos_next = 0, size = 62.813 MiB)
97
+ 0.26.399.124 I slot create_check: id 0 | task 204 | created context checkpoint 1 of 32 (pos_min = 420, pos_max = 420, n_tokens = 421, size = 62.813 MiB)
98
+ 0.26.601.525 I reasoning-budget: deactivated (natural end)
99
+ 0.26.675.559 I slot print_timing: id 0 | task 204 |
100
+ prompt eval time = 351.60 ms / 425 tokens ( 0.83 ms per token, 1208.76 tokens per second)
101
+ eval time = 248.92 ms / 17 tokens ( 14.64 ms per token, 68.30 tokens per second)
102
+ total time = 600.52 ms / 442 tokens
103
+ 0.26.675.604 I slot release: id 0 | task 204 | stop processing: n_tokens = 441, truncated = 0
104
+ 0.26.675.625 I srv update_slots: all slots are idle
105
+ 0.26.685.901 I srv params_from_: Chat format: peg-native
106
+ 0.26.686.221 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.962 (> 0.100 thold), f_keep = 0.921
107
+ 0.26.686.435 I reasoning-budget: activated, budget=2147483647 tokens
108
+ 0.26.686.468 I slot launch_slot_: id 0 | task 223 | processing task, is_child = 0
109
+ 0.26.686.473 W slot update_slots: id 0 | task 223 | n_past = 406, slot.prompt.tokens.size() = 441, seq_id = 0, pos_min = 440, n_swa = 0
110
+ 0.26.686.474 I slot update_slots: id 0 | task 223 | Checking checkpoint with [420, 420] against 406...
111
+ 0.26.686.475 W slot update_slots: id 0 | task 223 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
112
+ 0.26.686.476 W slot update_slots: id 0 | task 223 | erased invalidated context checkpoint (pos_min = 420, pos_max = 420, n_tokens = 421, n_swa = 0, pos_next = 0, size = 62.813 MiB)
113
+ 0.27.011.819 I slot create_check: id 0 | task 223 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
114
+ 0.27.481.528 I reasoning-budget: deactivated (natural end)
115
+ 0.28.088.823 I slot print_timing: id 0 | task 223 |
116
+ prompt eval time = 353.00 ms / 422 tokens ( 0.84 ms per token, 1195.46 tokens per second)
117
+ eval time = 1049.34 ms / 71 tokens ( 14.78 ms per token, 67.66 tokens per second)
118
+ total time = 1402.34 ms / 493 tokens
119
+ 0.28.088.866 I slot release: id 0 | task 223 | stop processing: n_tokens = 492, truncated = 0
120
+ 0.28.088.892 I srv update_slots: all slots are idle
121
+ 0.28.099.105 I srv params_from_: Chat format: peg-native
122
+ 0.28.099.413 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.943 (> 0.100 thold), f_keep = 1.000
123
+ 0.28.099.624 I reasoning-budget: activated, budget=2147483647 tokens
124
+ 0.28.099.652 I slot launch_slot_: id 0 | task 296 | processing task, is_child = 0
125
+ 0.28.109.016 I slot create_check: id 0 | task 296 | created context checkpoint 2 of 32 (pos_min = 491, pos_max = 491, n_tokens = 492, size = 62.813 MiB)
126
+ 0.28.638.379 I reasoning-budget: deactivated (natural end)
127
+ 0.28.905.201 I slot print_timing: id 0 | task 296 |
128
+ prompt eval time = 111.40 ms / 30 tokens ( 3.71 ms per token, 269.30 tokens per second)
129
+ eval time = 694.13 ms / 47 tokens ( 14.77 ms per token, 67.71 tokens per second)
130
+ total time = 805.53 ms / 77 tokens
131
+ 0.28.905.251 I slot release: id 0 | task 296 | stop processing: n_tokens = 568, truncated = 0
132
+ 0.28.905.277 I srv update_slots: all slots are idle
133
+ 0.28.915.951 I srv params_from_: Chat format: peg-native
134
+ 0.28.916.241 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.972 (> 0.100 thold), f_keep = 0.722
135
+ 0.28.916.410 I reasoning-budget: activated, budget=2147483647 tokens
136
+ 0.28.916.437 I slot launch_slot_: id 0 | task 345 | processing task, is_child = 0
137
+ 0.28.916.443 W slot update_slots: id 0 | task 345 | n_past = 410, slot.prompt.tokens.size() = 568, seq_id = 0, pos_min = 567, n_swa = 0
138
+ 0.28.916.454 I slot update_slots: id 0 | task 345 | Checking checkpoint with [491, 491] against 410...
139
+ 0.28.916.456 I slot update_slots: id 0 | task 345 | Checking checkpoint with [417, 417] against 410...
140
+ 0.28.916.457 W slot update_slots: id 0 | task 345 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
141
+ 0.28.916.459 W slot update_slots: id 0 | task 345 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
142
+ 0.28.918.102 W slot update_slots: id 0 | task 345 | erased invalidated context checkpoint (pos_min = 491, pos_max = 491, n_tokens = 492, n_swa = 0, pos_next = 0, size = 62.813 MiB)
143
+ 0.29.241.970 I slot create_check: id 0 | task 345 | created context checkpoint 1 of 32 (pos_min = 417, pos_max = 417, n_tokens = 418, size = 62.813 MiB)
144
+ 0.29.578.948 I reasoning-budget: deactivated (natural end)
145
+ 0.30.187.657 I slot print_timing: id 0 | task 345 |
146
+ prompt eval time = 354.33 ms / 422 tokens ( 0.84 ms per token, 1190.98 tokens per second)
147
+ eval time = 916.87 ms / 62 tokens ( 14.79 ms per token, 67.62 tokens per second)
148
+ total time = 1271.20 ms / 484 tokens
149
+ 0.30.187.698 I slot release: id 0 | task 345 | stop processing: n_tokens = 483, truncated = 0
150
+ 0.30.187.716 I srv update_slots: all slots are idle
151
+ 0.30.197.642 I srv params_from_: Chat format: peg-native
152
+ 0.30.197.942 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.935 (> 0.100 thold), f_keep = 0.839
153
+ 0.30.198.086 I reasoning-budget: activated, budget=2147483647 tokens
154
+ 0.30.198.110 I slot launch_slot_: id 0 | task 409 | processing task, is_child = 0
155
+ 0.30.198.115 W slot update_slots: id 0 | task 409 | n_past = 405, slot.prompt.tokens.size() = 483, seq_id = 0, pos_min = 482, n_swa = 0
156
+ 0.30.198.116 I slot update_slots: id 0 | task 409 | Checking checkpoint with [417, 417] against 405...
157
+ 0.30.198.117 W slot update_slots: id 0 | task 409 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
158
+ 0.30.198.120 W slot update_slots: id 0 | task 409 | erased invalidated context checkpoint (pos_min = 417, pos_max = 417, n_tokens = 418, n_swa = 0, pos_next = 0, size = 62.813 MiB)
159
+ 0.30.540.883 I slot create_check: id 0 | task 409 | created context checkpoint 1 of 32 (pos_min = 428, pos_max = 428, n_tokens = 429, size = 62.813 MiB)
160
+ 0.32.074.181 I reasoning-budget: deactivated (natural end)
161
+ 0.32.119.135 I slot print_timing: id 0 | task 409 | n_decoded = 100, tg = 64.60 t/s
162
+ 0.33.280.892 I slot print_timing: id 0 | task 409 |
163
+ prompt eval time = 373.02 ms / 433 tokens ( 0.86 ms per token, 1160.78 tokens per second)
164
+ eval time = 2709.74 ms / 178 tokens ( 15.22 ms per token, 65.69 tokens per second)
165
+ total time = 3082.77 ms / 611 tokens
166
+ 0.33.280.942 I slot release: id 0 | task 409 | stop processing: n_tokens = 610, truncated = 0
167
+ 0.33.280.983 I srv update_slots: all slots are idle
168
+ 0.33.291.660 I srv params_from_: Chat format: peg-native
169
+ 0.33.291.953 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.955 (> 0.100 thold), f_keep = 0.664
170
+ 0.33.292.103 I reasoning-budget: activated, budget=2147483647 tokens
171
+ 0.33.292.127 I slot launch_slot_: id 0 | task 589 | processing task, is_child = 0
172
+ 0.33.292.133 W slot update_slots: id 0 | task 589 | n_past = 405, slot.prompt.tokens.size() = 610, seq_id = 0, pos_min = 609, n_swa = 0
173
+ 0.33.292.133 I slot update_slots: id 0 | task 589 | Checking checkpoint with [428, 428] against 405...
174
+ 0.33.292.134 W slot update_slots: id 0 | task 589 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
175
+ 0.33.292.136 W slot update_slots: id 0 | task 589 | erased invalidated context checkpoint (pos_min = 428, pos_max = 428, n_tokens = 429, n_swa = 0, pos_next = 0, size = 62.813 MiB)
176
+ 0.33.632.176 I slot create_check: id 0 | task 589 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
177
+ 0.34.264.101 I slot print_timing: id 0 | task 589 |
178
+ prompt eval time = 370.43 ms / 424 tokens ( 0.87 ms per token, 1144.62 tokens per second)
179
+ eval time = 601.52 ms / 39 tokens ( 15.42 ms per token, 64.84 tokens per second)
180
+ total time = 971.94 ms / 463 tokens
181
+ 0.34.264.186 I slot release: id 0 | task 589 | stop processing: n_tokens = 462, truncated = 0
182
+ 0.34.264.222 I srv update_slots: all slots are idle
183
+ 0.34.291.838 I srv params_from_: Chat format: peg-native
184
+ 0.34.292.413 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.902 (> 0.100 thold), f_keep = 0.877
185
+ 0.34.292.609 I reasoning-budget: activated, budget=2147483647 tokens
186
+ 0.34.292.640 I slot launch_slot_: id 0 | task 630 | processing task, is_child = 0
187
+ 0.34.292.647 W slot update_slots: id 0 | task 630 | n_past = 405, slot.prompt.tokens.size() = 462, seq_id = 0, pos_min = 461, n_swa = 0
188
+ 0.34.292.648 I slot update_slots: id 0 | task 630 | Checking checkpoint with [419, 419] against 405...
189
+ 0.34.292.649 W slot update_slots: id 0 | task 630 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
190
+ 0.34.292.653 W slot update_slots: id 0 | task 630 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
191
+ 0.34.652.933 I slot create_check: id 0 | task 630 | created context checkpoint 1 of 32 (pos_min = 444, pos_max = 444, n_tokens = 445, size = 62.813 MiB)
192
+ 0.36.023.509 I slot print_timing: id 0 | task 630 |
193
+ prompt eval time = 390.02 ms / 449 tokens ( 0.87 ms per token, 1151.23 tokens per second)
194
+ eval time = 1340.82 ms / 86 tokens ( 15.59 ms per token, 64.14 tokens per second)
195
+ total time = 1730.84 ms / 535 tokens
196
+ 0.36.023.602 I slot release: id 0 | task 630 | stop processing: n_tokens = 534, truncated = 0
197
+ 0.36.023.637 I srv update_slots: all slots are idle
198
+ 0.36.023.710 W common_chat_peg_parse: unparsed peg-native output: <tool_call>
199
+ <function=create_event>
200
+ <parameter=attendees>
201
+ ["ana@x.io", "bo@x.io"]
202
+ </parameter>
203
+ <parameter=when>
204
+ {"date": "2026-10-02", "time": "14:00"}
205
+ </parameter>
206
+ <parameter=title>
207
+ Design review
208
+ </parameter>
209
+ </function>
210
+ </tool_call>
211
+ 0.36.025.846 W srv stop: cancel task, id_task = 630
212
+ 0.36.025.894 I srv update_slots: all slots are idle
213
+ 0.36.025.957 W srv operator(): got exception: {"error":{"code":500,"message":"The model produced output that does not match the expected peg-native format","type":"server_error"}}
214
+ 0.36.045.619 I srv params_from_: Chat format: peg-native
215
+ 0.36.045.957 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.948 (> 0.100 thold), f_keep = 0.758
216
+ 0.36.046.107 I reasoning-budget: activated, budget=2147483647 tokens
217
+ 0.36.046.133 I slot launch_slot_: id 0 | task 719 | processing task, is_child = 0
218
+ 0.36.046.138 W slot update_slots: id 0 | task 719 | n_past = 405, slot.prompt.tokens.size() = 534, seq_id = 0, pos_min = 533, n_swa = 0
219
+ 0.36.046.139 I slot update_slots: id 0 | task 719 | Checking checkpoint with [444, 444] against 405...
220
+ 0.36.046.142 W slot update_slots: id 0 | task 719 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
221
+ 0.36.046.144 W slot update_slots: id 0 | task 719 | erased invalidated context checkpoint (pos_min = 444, pos_max = 444, n_tokens = 445, n_swa = 0, pos_next = 0, size = 62.813 MiB)
222
+ 0.36.387.995 I slot create_check: id 0 | task 719 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
223
+ 0.37.019.525 I slot print_timing: id 0 | task 719 |
224
+ prompt eval time = 371.74 ms / 427 tokens ( 0.87 ms per token, 1148.64 tokens per second)
225
+ eval time = 601.62 ms / 39 tokens ( 15.43 ms per token, 64.82 tokens per second)
226
+ total time = 973.36 ms / 466 tokens
227
+ 0.37.019.612 I slot release: id 0 | task 719 | stop processing: n_tokens = 465, truncated = 0
228
+ 0.37.019.648 I srv update_slots: all slots are idle
229
+ 0.37.047.661 I srv params_from_: Chat format: peg-native
230
+ 0.37.048.192 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.948 (> 0.100 thold), f_keep = 0.871
231
+ 0.37.048.354 I reasoning-budget: activated, budget=2147483647 tokens
232
+ 0.37.048.382 I slot launch_slot_: id 0 | task 760 | processing task, is_child = 0
233
+ 0.37.048.388 W slot update_slots: id 0 | task 760 | n_past = 405, slot.prompt.tokens.size() = 465, seq_id = 0, pos_min = 464, n_swa = 0
234
+ 0.37.048.390 I slot update_slots: id 0 | task 760 | Checking checkpoint with [422, 422] against 405...
235
+ 0.37.048.391 W slot update_slots: id 0 | task 760 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
236
+ 0.37.048.393 W slot update_slots: id 0 | task 760 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
237
+ 0.37.396.973 I slot create_check: id 0 | task 760 | created context checkpoint 1 of 32 (pos_min = 422, pos_max = 422, n_tokens = 423, size = 62.813 MiB)
238
+ 0.37.495.139 I slot print_timing: id 0 | task 760 |
239
+ prompt eval time = 378.68 ms / 427 tokens ( 0.89 ms per token, 1127.60 tokens per second)
240
+ eval time = 68.05 ms / 4 tokens ( 17.01 ms per token, 58.78 tokens per second)
241
+ total time = 446.73 ms / 431 tokens
242
+ 0.37.495.225 I slot release: id 0 | task 760 | stop processing: n_tokens = 430, truncated = 0
243
+ 0.37.495.261 I srv update_slots: all slots are idle
244
+ 0.37.511.484 I srv params_from_: Chat format: peg-native
245
+ 0.37.511.815 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.958 (> 0.100 thold), f_keep = 0.944
246
+ 0.37.511.971 I reasoning-budget: activated, budget=2147483647 tokens
247
+ 0.37.512.000 I slot launch_slot_: id 0 | task 766 | processing task, is_child = 0
248
+ 0.37.512.007 W slot update_slots: id 0 | task 766 | n_past = 406, slot.prompt.tokens.size() = 430, seq_id = 0, pos_min = 429, n_swa = 0
249
+ 0.37.512.008 I slot update_slots: id 0 | task 766 | Checking checkpoint with [422, 422] against 406...
250
+ 0.37.512.009 W slot update_slots: id 0 | task 766 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
251
+ 0.37.512.012 W slot update_slots: id 0 | task 766 | erased invalidated context checkpoint (pos_min = 422, pos_max = 422, n_tokens = 423, n_swa = 0, pos_next = 0, size = 62.813 MiB)
252
+ 0.37.860.025 I slot create_check: id 0 | task 766 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
253
+ 0.38.509.296 I slot print_timing: id 0 | task 766 |
254
+ prompt eval time = 377.88 ms / 424 tokens ( 0.89 ms per token, 1122.06 tokens per second)
255
+ eval time = 619.39 ms / 40 tokens ( 15.48 ms per token, 64.58 tokens per second)
256
+ total time = 997.27 ms / 464 tokens
257
+ 0.38.509.383 I slot release: id 0 | task 766 | stop processing: n_tokens = 463, truncated = 0
258
+ 0.38.509.426 I srv update_slots: all slots are idle
259
+ 0.38.537.775 I srv params_from_: Chat format: peg-native
260
+ 0.38.538.231 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.935 (> 0.100 thold), f_keep = 1.000
261
+ 0.38.538.382 I reasoning-budget: activated, budget=2147483647 tokens
262
+ 0.38.538.411 I slot launch_slot_: id 0 | task 808 | processing task, is_child = 0
263
+ 0.38.628.133 I slot create_check: id 0 | task 808 | created context checkpoint 2 of 32 (pos_min = 490, pos_max = 490, n_tokens = 491, size = 62.813 MiB)
264
+ 0.38.869.113 I slot print_timing: id 0 | task 808 |
265
+ prompt eval time = 118.73 ms / 32 tokens ( 3.71 ms per token, 269.51 tokens per second)
266
+ eval time = 211.94 ms / 14 tokens ( 15.14 ms per token, 66.06 tokens per second)
267
+ total time = 330.67 ms / 46 tokens
268
+ 0.38.869.220 I slot release: id 0 | task 808 | stop processing: n_tokens = 508, truncated = 0
269
+ 0.38.869.257 I srv update_slots: all slots are idle
270
+ 0.38.890.188 I srv params_from_: Chat format: peg-native
271
+ 0.38.890.530 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.967 (> 0.100 thold), f_keep = 0.807
272
+ 0.38.890.683 I reasoning-budget: activated, budget=2147483647 tokens
273
+ 0.38.890.717 I slot launch_slot_: id 0 | task 824 | processing task, is_child = 0
274
+ 0.38.890.724 W slot update_slots: id 0 | task 824 | n_past = 410, slot.prompt.tokens.size() = 508, seq_id = 0, pos_min = 507, n_swa = 0
275
+ 0.38.890.724 I slot update_slots: id 0 | task 824 | Checking checkpoint with [490, 490] against 410...
276
+ 0.38.890.725 I slot update_slots: id 0 | task 824 | Checking checkpoint with [419, 419] against 410...
277
+ 0.38.890.726 W slot update_slots: id 0 | task 824 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
278
+ 0.38.890.729 W slot update_slots: id 0 | task 824 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
279
+ 0.38.892.293 W slot update_slots: id 0 | task 824 | erased invalidated context checkpoint (pos_min = 490, pos_max = 490, n_tokens = 491, n_swa = 0, pos_next = 0, size = 62.813 MiB)
280
+ 0.39.240.897 I slot create_check: id 0 | task 824 | created context checkpoint 1 of 32 (pos_min = 419, pos_max = 419, n_tokens = 420, size = 62.813 MiB)
281
+ 0.39.889.179 I slot print_timing: id 0 | task 824 |
282
+ prompt eval time = 380.09 ms / 424 tokens ( 0.90 ms per token, 1115.53 tokens per second)
283
+ eval time = 618.35 ms / 40 tokens ( 15.46 ms per token, 64.69 tokens per second)
284
+ total time = 998.43 ms / 464 tokens
285
+ 0.39.889.258 I slot release: id 0 | task 824 | stop processing: n_tokens = 463, truncated = 0
286
+ 0.39.889.291 I srv update_slots: all slots are idle
287
+ 0.39.916.715 I srv params_from_: Chat format: peg-native
288
+ 0.39.917.343 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.931 (> 0.100 thold), f_keep = 0.875
289
+ 0.39.917.545 I reasoning-budget: activated, budget=2147483647 tokens
290
+ 0.39.917.577 I slot launch_slot_: id 0 | task 866 | processing task, is_child = 0
291
+ 0.39.917.585 W slot update_slots: id 0 | task 866 | n_past = 405, slot.prompt.tokens.size() = 463, seq_id = 0, pos_min = 462, n_swa = 0
292
+ 0.39.917.586 I slot update_slots: id 0 | task 866 | Checking checkpoint with [419, 419] against 405...
293
+ 0.39.917.587 W slot update_slots: id 0 | task 866 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
294
+ 0.39.917.590 W slot update_slots: id 0 | task 866 | erased invalidated context checkpoint (pos_min = 419, pos_max = 419, n_tokens = 420, n_swa = 0, pos_next = 0, size = 62.813 MiB)
295
+ 0.40.271.964 I slot create_check: id 0 | task 866 | created context checkpoint 1 of 32 (pos_min = 430, pos_max = 430, n_tokens = 431, size = 62.813 MiB)
296
+ 0.41.551.306 I slot print_timing: id 0 | task 866 |
297
+ prompt eval time = 384.12 ms / 435 tokens ( 0.88 ms per token, 1132.46 tokens per second)
298
+ eval time = 1249.57 ms / 80 tokens ( 15.62 ms per token, 64.02 tokens per second)
299
+ total time = 1633.69 ms / 515 tokens
300
+ 0.41.551.447 I slot release: id 0 | task 866 | stop processing: n_tokens = 514, truncated = 0
301
+ 0.41.551.496 I srv update_slots: all slots are idle
302
+ 0.41.552.857 I srv operator(): operator(): cleaning up before exit...
recipe/logs/b_n-vision-q106-faoff.log ADDED
@@ -0,0 +1,66 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0.00.044.023 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
2
+ 0.00.044.026 I device_info:
3
+ 0.00.044.069 I - ROCm0 : AMD Radeon Graphics (131072 MiB, 125177 MiB free)
4
+ 0.00.044.147 I - Vulkan0 : AMD Radeon Graphics (RADV GFX1151) (132096 MiB, 131926 MiB free)
5
+ 0.00.044.151 I - CPU : AMD RYZEN AI MAX+ 395 w/ Radeon 8060S (127438 MiB, 127438 MiB free)
6
+ 0.00.044.190 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
7
+ 0.00.044.225 I srv init: running without SSL
8
+ 0.00.044.247 I srv init: using 31 threads for HTTP server
9
+ 0.00.044.248 I srv init: the WebUI is disabled
10
+ 0.00.044.304 I srv start: binding port with default address family
11
+ 0.00.045.454 I srv main: loading model
12
+ 0.00.045.457 I srv load_model: loading model '/mnt/models/nex-n2.5-mini/out/Nex-N2.5-mini-Q4_0_ROCMFP4_STRIX_LEAN.gguf'
13
+ 0.00.079.109 W llama_model_loader: direct I/O is enabled, disabling mmap
14
+ 0.20.570.878 W llama_context: n_ctx_seq (65536) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
15
+ 0.21.071.707 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
16
+ 0.21.277.411 W load_hparams: Qwen-VL models require at minimum 1024 image tokens to function correctly on grounding tasks
17
+ 0.21.277.414 W load_hparams: if you encounter problems with accuracy, try adding --image-min-tokens 1024
18
+ 0.21.277.414 W load_hparams: more info: https://github.com/ggml-org/llama.cpp/issues/16842
19
+
20
+ 0.22.242.916 I srv load_model: loaded multimodal model, '/mnt/models/nex-n2.5-mini/out/mmproj-Nex-N2.5-mini-BF16.gguf'
21
+ 0.22.242.925 I srv load_model: initializing slots, n_slots = 1
22
+ 0.22.421.079 W srv load_model: speculative decoding will use checkpoints
23
+ 0.22.421.086 W common_speculative_init: no implementations specified for speculative decoding
24
+ 0.22.421.088 I slot load_model: id 0 | task -1 | new slot, n_ctx = 65536
25
+ 0.22.421.122 I srv load_model: prompt cache RAM enabled: limit_mib=8192
26
+ 0.22.421.129 I srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
27
+ 0.22.421.144 I srv init: idle slots will be saved to prompt cache upon starting a new task
28
+ 0.22.429.869 I init: chat template, example_format: '<|im_start|>system
29
+ You are a helpful assistant<|im_end|>
30
+ <|im_start|>user
31
+ Hello<|im_end|>
32
+ <|im_start|>assistant
33
+ <think>
34
+
35
+ </think>
36
+
37
+ Hi there<|im_end|>
38
+ <|im_start|>user
39
+ How are you?<|im_end|>
40
+ <|im_start|>assistant
41
+ <think>'
42
+ 0.22.436.195 I srv init: init: chat template, thinking = 1
43
+ 0.22.436.220 I srv main: model loaded
44
+ 0.22.436.223 I srv main: server is listening on http://127.0.0.1:18600
45
+ 0.22.436.225 I srv update_slots: all slots are idle
46
+ 0.24.206.034 I srv params_from_: Chat format: peg-native
47
+ 0.24.206.169 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1
48
+ 0.24.206.172 I srv get_availabl: updating prompt cache
49
+ 0.24.206.177 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000
50
+ 0.24.206.182 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 65536 tokens, 8589934592 est)
51
+ 0.24.206.183 I srv get_availabl: prompt cache update took 0.01 ms
52
+ 0.24.206.233 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0
53
+ 0.24.246.538 I srv process_chun: processing image...
54
+ 0.24.375.032 W find_slot: non-consecutive token position 4 after 3 for sequence 0 with 196 new tokens
55
+ 0.24.375.099 W find_slot: non-consecutive token position 4 after 3 for sequence 0 with 196 new tokens
56
+ 0.24.571.922 I srv process_chun: image processed in 326 ms
57
+ 0.24.572.141 W find_slot: non-consecutive token position 34 after 4 for sequence 0 with 17 new tokens
58
+ 0.24.572.170 W find_slot: non-consecutive token position 34 after 4 for sequence 0 with 17 new tokens
59
+ 0.24.650.730 I slot create_check: id 0 | task 0 | created context checkpoint 1 of 32 (pos_min = 34, pos_max = 34, n_tokens = 217, size = 62.813 MiB)
60
+ 0.25.135.589 I slot print_timing: id 0 | task 0 |
61
+ prompt eval time = 471.32 ms / 221 tokens ( 2.13 ms per token, 468.89 tokens per second)
62
+ eval time = 458.01 ms / 30 tokens ( 15.27 ms per token, 65.50 tokens per second)
63
+ total time = 929.34 ms / 251 tokens
64
+ 0.25.135.616 I slot release: id 0 | task 0 | stop processing: n_tokens = 250, truncated = 0
65
+ 0.25.135.619 I srv update_slots: all slots are idle
66
+ 0.26.136.415 I srv operator(): operator(): cleaning up before exit...
recipe/logs/b_n-vision-q106-faon.log ADDED
@@ -0,0 +1,66 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0.00.039.566 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
2
+ 0.00.039.569 I device_info:
3
+ 0.00.039.614 I - ROCm0 : AMD Radeon Graphics (131072 MiB, 125113 MiB free)
4
+ 0.00.039.689 I - Vulkan0 : AMD Radeon Graphics (RADV GFX1151) (132096 MiB, 131926 MiB free)
5
+ 0.00.039.693 I - CPU : AMD RYZEN AI MAX+ 395 w/ Radeon 8060S (127438 MiB, 127438 MiB free)
6
+ 0.00.039.733 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
7
+ 0.00.039.777 I srv init: running without SSL
8
+ 0.00.039.798 I srv init: using 31 threads for HTTP server
9
+ 0.00.039.799 I srv init: the WebUI is disabled
10
+ 0.00.039.854 I srv start: binding port with default address family
11
+ 0.00.041.005 I srv main: loading model
12
+ 0.00.041.007 I srv load_model: loading model '/mnt/models/nex-n2.5-mini/out/Nex-N2.5-mini-Q4_0_ROCMFP4_STRIX_LEAN.gguf'
13
+ 0.00.072.658 W llama_model_loader: direct I/O is enabled, disabling mmap
14
+ 0.20.385.335 W llama_context: n_ctx_seq (65536) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
15
+ 0.20.547.529 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
16
+ 0.20.746.298 W load_hparams: Qwen-VL models require at minimum 1024 image tokens to function correctly on grounding tasks
17
+ 0.20.746.300 W load_hparams: if you encounter problems with accuracy, try adding --image-min-tokens 1024
18
+ 0.20.746.300 W load_hparams: more info: https://github.com/ggml-org/llama.cpp/issues/16842
19
+
20
+ 0.21.800.919 I srv load_model: loaded multimodal model, '/mnt/models/nex-n2.5-mini/out/mmproj-Nex-N2.5-mini-BF16.gguf'
21
+ 0.21.800.929 I srv load_model: initializing slots, n_slots = 1
22
+ 0.21.998.188 W srv load_model: speculative decoding will use checkpoints
23
+ 0.21.998.194 W common_speculative_init: no implementations specified for speculative decoding
24
+ 0.21.998.195 I slot load_model: id 0 | task -1 | new slot, n_ctx = 65536
25
+ 0.21.998.237 I srv load_model: prompt cache RAM enabled: limit_mib=8192
26
+ 0.21.998.246 I srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
27
+ 0.21.998.259 I srv init: idle slots will be saved to prompt cache upon starting a new task
28
+ 0.22.007.086 I init: chat template, example_format: '<|im_start|>system
29
+ You are a helpful assistant<|im_end|>
30
+ <|im_start|>user
31
+ Hello<|im_end|>
32
+ <|im_start|>assistant
33
+ <think>
34
+
35
+ </think>
36
+
37
+ Hi there<|im_end|>
38
+ <|im_start|>user
39
+ How are you?<|im_end|>
40
+ <|im_start|>assistant
41
+ <think>'
42
+ 0.22.013.608 I srv init: init: chat template, thinking = 1
43
+ 0.22.013.635 I srv main: model loaded
44
+ 0.22.013.638 I srv main: server is listening on http://127.0.0.1:18600
45
+ 0.22.013.640 I srv update_slots: all slots are idle
46
+ 0.24.005.244 I srv params_from_: Chat format: peg-native
47
+ 0.24.005.366 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1
48
+ 0.24.005.368 I srv get_availabl: updating prompt cache
49
+ 0.24.005.373 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000
50
+ 0.24.005.376 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 65536 tokens, 8589934592 est)
51
+ 0.24.005.377 I srv get_availabl: prompt cache update took 0.01 ms
52
+ 0.24.005.401 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0
53
+ 0.24.032.329 I srv process_chun: processing image...
54
+ 0.24.141.298 W find_slot: non-consecutive token position 4 after 3 for sequence 0 with 196 new tokens
55
+ 0.24.141.364 W find_slot: non-consecutive token position 4 after 3 for sequence 0 with 196 new tokens
56
+ 0.24.326.150 I srv process_chun: image processed in 294 ms
57
+ 0.24.326.369 W find_slot: non-consecutive token position 34 after 4 for sequence 0 with 17 new tokens
58
+ 0.24.326.399 W find_slot: non-consecutive token position 34 after 4 for sequence 0 with 17 new tokens
59
+ 0.24.404.937 I slot create_check: id 0 | task 0 | created context checkpoint 1 of 32 (pos_min = 34, pos_max = 34, n_tokens = 217, size = 62.813 MiB)
60
+ 0.24.744.679 I slot print_timing: id 0 | task 0 |
61
+ prompt eval time = 426.45 ms / 221 tokens ( 1.93 ms per token, 518.23 tokens per second)
62
+ eval time = 312.81 ms / 21 tokens ( 14.90 ms per token, 67.13 tokens per second)
63
+ total time = 739.26 ms / 242 tokens
64
+ 0.24.744.708 I slot release: id 0 | task 0 | stop processing: n_tokens = 241, truncated = 0
65
+ 0.24.744.712 I srv update_slots: all slots are idle
66
+ 0.25.745.484 I srv operator(): operator(): cleaning up before exit...
recipe/logs/b_n-vision-q106-roff-faon.log ADDED
@@ -0,0 +1,70 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0.00.208.975 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
2
+ 0.00.208.993 I device_info:
3
+ 0.00.209.271 I - ROCm0 : AMD Radeon Graphics (131072 MiB, 123872 MiB free)
4
+ 0.00.209.724 I - Vulkan0 : AMD Radeon Graphics (RADV GFX1151) (132096 MiB, 131922 MiB free)
5
+ 0.00.209.744 I - CPU : AMD RYZEN AI MAX+ 395 w/ Radeon 8060S (127438 MiB, 127438 MiB free)
6
+ 0.00.209.942 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
7
+ 0.00.210.045 I srv init: running without SSL
8
+ 0.00.210.124 I srv init: using 31 threads for HTTP server
9
+ 0.00.210.128 I srv init: the WebUI is disabled
10
+ 0.00.210.347 I srv start: binding port with default address family
11
+ 0.00.211.826 I srv main: loading model
12
+ 0.00.211.839 I srv load_model: loading model '/mnt/models/nex-n2.5-mini/out/Nex-N2.5-mini-Q4_0_ROCMFP4_STRIX_LEAN.gguf'
13
+ 0.00.341.319 W llama_model_loader: direct I/O is enabled, disabling mmap
14
+ 0.21.841.225 W llama_context: n_ctx_seq (65536) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
15
+ 0.22.147.423 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
16
+ 0.22.493.441 W load_hparams: Qwen-VL models require at minimum 1024 image tokens to function correctly on grounding tasks
17
+ 0.22.493.444 W load_hparams: if you encounter problems with accuracy, try adding --image-min-tokens 1024
18
+ 0.22.493.445 W load_hparams: more info: https://github.com/ggml-org/llama.cpp/issues/16842
19
+
20
+ 0.22.758.348 I srv load_model: loaded multimodal model, '/mnt/models/nex-n2.5-mini/out/mmproj-Nex-N2.5-mini-BF16.gguf'
21
+ 0.22.758.359 I srv load_model: initializing slots, n_slots = 1
22
+ 0.23.027.404 W srv load_model: speculative decoding will use checkpoints
23
+ 0.23.027.422 W common_speculative_init: no implementations specified for speculative decoding
24
+ 0.23.027.426 I slot load_model: id 0 | task -1 | new slot, n_ctx = 65536
25
+ 0.23.027.506 I srv load_model: prompt cache RAM enabled: limit_mib=8192
26
+ 0.23.027.508 I srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
27
+ 0.23.027.604 I srv init: idle slots will be saved to prompt cache upon starting a new task
28
+ 0.23.040.255 I init: chat template, example_format: '<|im_start|>system
29
+ You are a helpful assistant<|im_end|>
30
+ <|im_start|>user
31
+ Hello<|im_end|>
32
+ <|im_start|>assistant
33
+ <think>
34
+
35
+ </think>
36
+
37
+ Hi there<|im_end|>
38
+ <|im_start|>user
39
+ How are you?<|im_end|>
40
+ <|im_start|>assistant
41
+ <think>
42
+
43
+ </think>
44
+
45
+ '
46
+ 0.23.049.208 I srv init: init: chat template, thinking = 0
47
+ 0.23.049.258 I srv main: model loaded
48
+ 0.23.049.261 I srv main: server is listening on http://127.0.0.1:18600
49
+ 0.23.049.265 I srv update_slots: all slots are idle
50
+ 0.24.037.221 I srv params_from_: Chat format: peg-native
51
+ 0.24.037.451 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1
52
+ 0.24.037.454 I srv get_availabl: updating prompt cache
53
+ 0.24.037.461 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000
54
+ 0.24.037.467 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 65536 tokens, 8589934592 est)
55
+ 0.24.037.470 I srv get_availabl: prompt cache update took 0.01 ms
56
+ 0.24.037.553 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0
57
+ 0.24.096.831 I srv process_chun: processing image...
58
+ 0.24.348.055 W find_slot: non-consecutive token position 4 after 3 for sequence 0 with 196 new tokens
59
+ 0.24.348.208 W find_slot: non-consecutive token position 4 after 3 for sequence 0 with 196 new tokens
60
+ 0.24.670.293 I srv process_chun: image processed in 573 ms
61
+ 0.24.670.967 W find_slot: non-consecutive token position 34 after 4 for sequence 0 with 17 new tokens
62
+ 0.24.671.060 W find_slot: non-consecutive token position 34 after 4 for sequence 0 with 17 new tokens
63
+ 0.24.787.337 I slot create_check: id 0 | task 0 | created context checkpoint 1 of 32 (pos_min = 34, pos_max = 34, n_tokens = 217, size = 62.813 MiB)
64
+ 0.25.422.382 I slot print_timing: id 0 | task 0 |
65
+ prompt eval time = 788.44 ms / 221 tokens ( 3.57 ms per token, 280.30 tokens per second)
66
+ eval time = 596.35 ms / 21 tokens ( 28.40 ms per token, 35.21 tokens per second)
67
+ total time = 1384.78 ms / 242 tokens
68
+ 0.25.422.466 I slot release: id 0 | task 0 | stop processing: n_tokens = 241, truncated = 0
69
+ 0.25.422.481 I srv update_slots: all slots are idle
70
+ 0.26.423.773 I srv operator(): operator(): cleaning up before exit...
recipe/logs/diag_bf16.log ADDED
@@ -0,0 +1,5 @@
 
 
 
 
 
 
1
+ [2026-09-16T22:27:17Z] bf16_rocm_faoff rc=0 [1]136.6879 [2]166.5229 [3]174.2693 [4]169.2609
2
+ [2026-09-16T22:29:15Z] bf16_vk_faon rc=0 [1]5.6953 [2]6.6594 [3]7.0323 [4]7.2857
3
+ [2026-09-16T22:29:46Z] q106_vk_faon rc=0 [1]6.1880 [2]7.0357 [3]7.4624 [4]7.7345
4
+ [2026-09-16T22:30:26Z] bf16_cpu rc=0 [1]141.8336 [2]168.6793
5
+ [2026-09-16T22:30:26Z] DIAG_BF16_DONE
recipe/logs/diag_bf16_cpu.log ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0.00.032.646 I common_init_result: fitting params to device memory ...
2
+ 0.00.032.649 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
3
+ 0.08.593.776 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
4
+ 0.08.699.313 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
5
+ 0.09.377.813 I
6
+ 0.09.377.939 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
7
+ 0.09.378.018 I perplexity: tokenizing the input ..
8
+ 0.09.680.733 I perplexity: tokenization took 302.776 ms
9
+ 0.09.680.810 I perplexity: calculating perplexity over 2 chunks, n_ctx=2048, batch_size=2048, n_seq=1
10
+ 0.24.718.172 I perplexity: 15.04 seconds per pass - ETA 0.50 minutes
11
+ [1]141.8336,[2]168.6793,
12
+ 0.39.305.358 I Final estimate: PPL = 168.6793 +/- 12.64504
13
+
recipe/logs/diag_bf16_purecpu_c1.log ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0.00.039.976 I common_init_result: fitting params to device memory ...
2
+ 0.00.039.980 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
3
+ 0.00.307.048 I common_params_fit_impl: projected to use 66756 MiB of host memory vs. 127438 MiB of total host memory
4
+ 0.00.633.915 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
5
+ 0.00.653.775 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
6
+ 0.01.312.044 I
7
+ 0.01.312.177 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
8
+ 0.01.312.255 I perplexity: tokenizing the input ..
9
+ 0.01.611.737 I perplexity: tokenization took 299.542 ms
10
+ 0.01.611.820 I perplexity: calculating perplexity over 1 chunks, n_ctx=2048, batch_size=2048, n_seq=1
11
+ 0.11.553.629 I perplexity: 9.94 seconds per pass - ETA 0.15 minutes
12
+ [1]5.6964,
13
+ 0.11.592.689 I Final estimate: PPL = 5.6964 +/- 0.43869
14
+
recipe/logs/diag_bf16_rocm_faoff.log ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0.00.041.213 I common_init_result: fitting params to device memory ...
2
+ 0.00.041.216 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
3
+ 2.25.152.472 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
4
+ 2.25.329.776 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
5
+ 2.26.102.739 I
6
+ 2.26.102.951 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
7
+ 2.26.102.958 I perplexity: tokenizing the input ..
8
+ 2.26.847.963 I perplexity: tokenization took 744.994 ms
9
+ 2.26.848.073 I perplexity: calculating perplexity over 4 chunks, n_ctx=2048, batch_size=2048, n_seq=1
10
+ 2.32.003.043 I perplexity: 5.15 seconds per pass - ETA 0.33 minutes
11
+ [1]136.6879,[2]166.5229,[3]174.2693,[4]169.2609,
12
+ 2.45.605.000 I Final estimate: PPL = 169.2609 +/- 8.97119
13
+
recipe/logs/diag_bf16_vk_faon.log ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0.00.081.666 I common_init_result: fitting params to device memory ...
2
+ 0.00.081.669 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
3
+ 1.04.951.000 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
4
+ 1.05.009.711 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
5
+ 1.06.643.947 I
6
+ 1.06.646.719 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
7
+ 1.06.646.727 I perplexity: tokenizing the input ..
8
+ 1.07.413.989 I perplexity: tokenization took 767.253 ms
9
+ 1.07.414.080 I perplexity: calculating perplexity over 4 chunks, n_ctx=2048, batch_size=2048, n_seq=1
10
+ 1.19.697.964 I perplexity: 12.28 seconds per pass - ETA 0.82 minutes
11
+ [1]5.6953,[2]6.6594,[3]7.0323,[4]7.2857,
12
+ 1.57.415.169 I Final estimate: PPL = 7.2857 +/- 0.29353
13
+
recipe/logs/diag_ppl_q106_rocm_c4.log ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0.00.044.069 I common_init_result: fitting params to device memory ...
2
+ 0.00.044.071 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
3
+ 0.03.928.272 W llama_context: n_ctx_seq (2048) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
4
+ 0.03.976.550 W common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
5
+ 0.04.333.276 I
6
+ 0.04.333.426 I system_info: n_threads = 16 (n_threads_batch = 16) / 32 | ROCm : NO_VMM = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
7
+ 0.04.333.434 I perplexity: tokenizing the input ..
8
+ 0.04.683.706 I perplexity: tokenization took 350.264 ms
9
+ 0.04.683.803 I perplexity: calculating perplexity over 4 chunks, n_ctx=2048, batch_size=2048, n_seq=1
10
+ 0.07.474.485 I perplexity: 2.79 seconds per pass - ETA 0.18 minutes
11
+ [1]6.2788,[2]7.1110,[3]7.5138,[4]7.8414,
12
+ 0.14.413.643 I Final estimate: PPL = 7.8414 +/- 0.32420
13
+