Mapika commited on
Commit
26c404e
·
verified ·
1 Parent(s): 4e5c245

decider-4b v2.1 GGUF: Q4_K_M, Q8_0, BF16 + decide_gguf.py

Browse files
.gitattributes CHANGED
@@ -33,3 +33,7 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ decider-4b-v2.1-BF16.gguf filter=lfs diff=lfs merge=lfs -text
37
+ decider-4b-v2.1-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
38
+ decider-4b-v2.1-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
39
+ tokenizer.json filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,88 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: Mapika/decider-4b
4
+ base_model_relation: quantized
5
+ language: [en]
6
+ pipeline_tag: text-classification
7
+ tags: [gguf, llama.cpp, decision-model, calibrated, structured-output, one-pass]
8
+ ---
9
+
10
+ # decider-4b GGUF
11
+
12
+ GGUF files of [Mapika/decider-4b](https://huggingface.co/Mapika/decider-4b) v2.1 (Hub main `eb5fbdf`), for llama.cpp. decider-4b
13
+ does not generate text. It reads a state and one or more questions, each with an explicit option list, and returns a probability
14
+ for every option from one forward pass. See the [decider-4b card](https://huggingface.co/Mapika/decider-4b) for what the model
15
+ is, how it was trained, and where it is weak.
16
+
17
+ | file | size | use |
18
+ |---|---|---|
19
+ | `decider-4b-v2.1-Q4_K_M.gguf` | 2.7 GB | smallest; about 0.2 points lower in-task accuracy (table below) |
20
+ | `decider-4b-v2.1-Q8_0.gguf` | 4.5 GB | same quality as the bf16 weights |
21
+ | `decider-4b-v2.1-BF16.gguf` | 8.4 GB | unquantized, for making other quantizations |
22
+
23
+ The tokenizer, `decider_config.json` (temperatures) and `decide_gguf.py` (the readout on llama.cpp) are in this repository too.
24
+
25
+ ## This is not a chat model
26
+
27
+ Loading a file in `llama-cli`, `llama-server`, Ollama or LM Studio gives you a text model that continues prompts. That is not how
28
+ decider-4b is used, and its generated text is not its answer. The answer is read from the logits of the option-letter tokens at
29
+ each answer slot of a prompt built by `decider.prompt`, divided by the fitted temperature. `decide_gguf.py` does this with
30
+ llama-cpp-python.
31
+
32
+ ## Usage
33
+
34
+ ```
35
+ pip install decider-ai==1.5.0 llama-cpp-python # llama-cpp-python 0.3.35 or newer
36
+ # GPU: CMAKE_ARGS="-DGGML_CUDA=on" pip install llama-cpp-python (Apple Silicon: -DGGML_METAL=on)
37
+ hf download Mapika/decider-4b-GGUF --local-dir decider-4b-gguf \
38
+ --include "decider-4b-v2.1-Q4_K_M.gguf" "*.json" "*.jinja" "decide_gguf.py"
39
+ cd decider-4b-gguf
40
+ ```
41
+
42
+ ```python
43
+ from decide_gguf import GGUFDecider
44
+
45
+ d = GGUFDecider("decider-4b-v2.1-Q4_K_M.gguf") # n_gpu_layers=-1 (all on the GPU if the build has one), n_threads=...
46
+ d.decide("My card was charged twice for the same purchase.",
47
+ [{"question": "Which department should handle this?", "options": ["billing", "technical", "sales"]},
48
+ {"question": "How urgent is this?", "options": ["low", "medium", "high"]}])
49
+ # [{'choice': 'billing', 'confidence': 0.85, 'probs': {...}}, {'choice': 'medium', 'confidence': 0.43, 'probs': {...}}]
50
+ ```
51
+
52
+ `decide_gguf.py` covers `decide()` (choice questions). The full API of the package (`system_one` with score and yes/no answers,
53
+ the HTTP server) on GGUF files is not in a released decider-ai yet.
54
+
55
+ On 8 CPU threads (server CPU), Q4_K_M takes 0.3 to 0.7 s for a request of 40 to 120 tokens; Q8_0 is about 20% slower.
56
+
57
+ ## Measured quality
58
+
59
+ The regression set of the decider-4b card (95 tasks, 144,226 questions, 67 in-task and 28 held-out tasks) at the shipped
60
+ temperature 1.099, read through llama.cpp (CUDA build, one prompt per decode) and compared row by row with the bf16 weights in
61
+ PyTorch. Accuracy, NLL and ECE are means over tasks.
62
+
63
+ | | in-task acc / NLL / ECE | held-out acc / NLL / ECE | same answer as bf16 PyTorch |
64
+ |---|---|---|---|
65
+ | bf16 weights, PyTorch | 0.8308 / 0.4145 / 0.0308 | 0.7838 / 0.5703 / 0.0781 | |
66
+ | BF16 GGUF | 0.8308 / 0.4145 / 0.0308 | 0.7837 / 0.5700 / 0.0782 | 99.45% |
67
+ | Q8_0 | 0.8310 / 0.4145 / 0.0309 | 0.7829 / 0.5699 / 0.0778 | 99.31% |
68
+ | Q4_K_M | 0.8288 / 0.4194 / 0.0324 | 0.7834 / 0.5691 / 0.0733 | 97.06% |
69
+
70
+ The BF16 GGUF differs from PyTorch only on near-ties (median probability difference 0.001). Q8_0 is equal to the bf16 weights
71
+ within that noise; its tasks move up on 31 and down on 39, by at most 0.9 points. Q4_K_M is 0.2 points lower on in-task accuracy
72
+ with slightly higher NLL; held-out accuracy is unchanged. It is lower than bf16 on 54 tasks and higher on 36; the largest drop is
73
+ fin_phrasebank (−3.5 points), then medmcqa and mmlu (−1.3). In Q4_K_M the embedding matrix, which is also the output matrix that
74
+ holds the option-letter rows, is stored in Q6_K.
75
+
76
+ ## Notes
77
+
78
+ - Score one prompt per `llama_decode`, as `decide_gguf.py` does. With several prompts in one decode (as separate sequences),
79
+ llama.cpp in September 2026 gives probabilities that change with the other prompts in the batch, by up to 0.02 in BF16 and 0.16
80
+ in Q4_K_M on this model. One prompt per decode gives the same numbers on every run.
81
+ - CPU and GPU builds give slightly different probabilities on the same file (for example 0.846 and 0.830 for "billing" above,
82
+ against 0.844 in PyTorch).
83
+ - Conversion: llama.cpp `c9064dded` (2026-09-27), `convert_hf_to_gguf.py --no-mtp --outtype bf16`, then `llama-quantize` to
84
+ Q8_0 and Q4_K_M. `--no-mtp` is required: the checkpoint has no multi-token-prediction weights, but its config declares one
85
+ MTP layer, and without the flag the converter writes a file that llama.cpp cannot load.
86
+ - The measurements above use the CUDA build. The CPU and Metal builds were not run over the regression set.
87
+
88
+ License: Apache-2.0, as decider-4b and its base model Qwen/Qwen3.5-4B-Base.
chat_template.jinja ADDED
@@ -0,0 +1,154 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {%- set image_count = namespace(value=0) %}
2
+ {%- set video_count = namespace(value=0) %}
3
+ {%- macro render_content(content, do_vision_count, is_system_content=false) %}
4
+ {%- if content is string %}
5
+ {{- content }}
6
+ {%- elif content is iterable and content is not mapping %}
7
+ {%- for item in content %}
8
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
9
+ {%- if is_system_content %}
10
+ {{- raise_exception('System message cannot contain images.') }}
11
+ {%- endif %}
12
+ {%- if do_vision_count %}
13
+ {%- set image_count.value = image_count.value + 1 %}
14
+ {%- endif %}
15
+ {%- if add_vision_id %}
16
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
17
+ {%- endif %}
18
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
19
+ {%- elif 'video' in item or item.type == 'video' %}
20
+ {%- if is_system_content %}
21
+ {{- raise_exception('System message cannot contain videos.') }}
22
+ {%- endif %}
23
+ {%- if do_vision_count %}
24
+ {%- set video_count.value = video_count.value + 1 %}
25
+ {%- endif %}
26
+ {%- if add_vision_id %}
27
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
28
+ {%- endif %}
29
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
30
+ {%- elif 'text' in item %}
31
+ {{- item.text }}
32
+ {%- else %}
33
+ {{- raise_exception('Unexpected item type in content.') }}
34
+ {%- endif %}
35
+ {%- endfor %}
36
+ {%- elif content is none or content is undefined %}
37
+ {{- '' }}
38
+ {%- else %}
39
+ {{- raise_exception('Unexpected content type.') }}
40
+ {%- endif %}
41
+ {%- endmacro %}
42
+ {%- if not messages %}
43
+ {{- raise_exception('No messages provided.') }}
44
+ {%- endif %}
45
+ {%- if tools and tools is iterable and tools is not mapping %}
46
+ {{- '<|im_start|>system\n' }}
47
+ {{- "# Tools\n\nYou have access to the following functions:\n\n<tools>" }}
48
+ {%- for tool in tools %}
49
+ {{- "\n" }}
50
+ {{- tool | tojson }}
51
+ {%- endfor %}
52
+ {{- "\n</tools>" }}
53
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n<tool_call>\n<function=example_function_name>\n<parameter=example_parameter_1>\nvalue_1\n</parameter>\n<parameter=example_parameter_2>\nThis is the value for the second parameter\nthat can span\nmultiple lines\n</parameter>\n</function>\n</tool_call>\n\n<IMPORTANT>\nReminder:\n- Function calls MUST follow the specified format: an inner <function=...></function> block must be nested within <tool_call></tool_call> XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n</IMPORTANT>' }}
54
+ {%- if messages[0].role == 'system' %}
55
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
56
+ {%- if content %}
57
+ {{- '\n\n' + content }}
58
+ {%- endif %}
59
+ {%- endif %}
60
+ {{- '<|im_end|>\n' }}
61
+ {%- else %}
62
+ {%- if messages[0].role == 'system' %}
63
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
64
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
65
+ {%- endif %}
66
+ {%- endif %}
67
+ {%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
68
+ {%- for message in messages[::-1] %}
69
+ {%- set index = (messages|length - 1) - loop.index0 %}
70
+ {%- if ns.multi_step_tool and message.role == "user" %}
71
+ {%- set content = render_content(message.content, false)|trim %}
72
+ {%- if not(content.startswith('<tool_response>') and content.endswith('</tool_response>')) %}
73
+ {%- set ns.multi_step_tool = false %}
74
+ {%- set ns.last_query_index = index %}
75
+ {%- endif %}
76
+ {%- endif %}
77
+ {%- endfor %}
78
+ {%- if ns.multi_step_tool %}
79
+ {{- raise_exception('No user query found in messages.') }}
80
+ {%- endif %}
81
+ {%- for message in messages %}
82
+ {%- set content = render_content(message.content, true)|trim %}
83
+ {%- if message.role == "system" %}
84
+ {%- if not loop.first %}
85
+ {{- raise_exception('System message must be at the beginning.') }}
86
+ {%- endif %}
87
+ {%- elif message.role == "user" %}
88
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
89
+ {%- elif message.role == "assistant" %}
90
+ {%- set reasoning_content = '' %}
91
+ {%- if message.reasoning_content is string %}
92
+ {%- set reasoning_content = message.reasoning_content %}
93
+ {%- else %}
94
+ {%- if '</think>' in content %}
95
+ {%- set reasoning_content = content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
96
+ {%- set content = content.split('</think>')[-1].lstrip('\n') %}
97
+ {%- endif %}
98
+ {%- endif %}
99
+ {%- set reasoning_content = reasoning_content|trim %}
100
+ {%- if loop.index0 > ns.last_query_index %}
101
+ {{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content + '\n</think>\n\n' + content }}
102
+ {%- else %}
103
+ {{- '<|im_start|>' + message.role + '\n' + content }}
104
+ {%- endif %}
105
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
106
+ {%- for tool_call in message.tool_calls %}
107
+ {%- if tool_call.function is defined %}
108
+ {%- set tool_call = tool_call.function %}
109
+ {%- endif %}
110
+ {%- if loop.first %}
111
+ {%- if content|trim %}
112
+ {{- '\n\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
113
+ {%- else %}
114
+ {{- '<tool_call>\n<function=' + tool_call.name + '>\n' }}
115
+ {%- endif %}
116
+ {%- else %}
117
+ {{- '\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
118
+ {%- endif %}
119
+ {%- if tool_call.arguments is defined %}
120
+ {%- for args_name, args_value in tool_call.arguments|items %}
121
+ {{- '<parameter=' + args_name + '>\n' }}
122
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
123
+ {{- args_value }}
124
+ {{- '\n</parameter>\n' }}
125
+ {%- endfor %}
126
+ {%- endif %}
127
+ {{- '</function>\n</tool_call>' }}
128
+ {%- endfor %}
129
+ {%- endif %}
130
+ {{- '<|im_end|>\n' }}
131
+ {%- elif message.role == "tool" %}
132
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
133
+ {{- '<|im_start|>user' }}
134
+ {%- endif %}
135
+ {{- '\n<tool_response>\n' }}
136
+ {{- content }}
137
+ {{- '\n</tool_response>' }}
138
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
139
+ {{- '<|im_end|>\n' }}
140
+ {%- elif loop.last %}
141
+ {{- '<|im_end|>\n' }}
142
+ {%- endif %}
143
+ {%- else %}
144
+ {{- raise_exception('Unexpected message role.') }}
145
+ {%- endif %}
146
+ {%- endfor %}
147
+ {%- if add_generation_prompt %}
148
+ {{- '<|im_start|>assistant\n' }}
149
+ {%- if enable_thinking is defined and enable_thinking is false %}
150
+ {{- '<think>\n\n</think>\n\n' }}
151
+ {%- else %}
152
+ {{- '<think>\n' }}
153
+ {%- endif %}
154
+ {%- endif %}
decide_gguf.py ADDED
@@ -0,0 +1,78 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """decider-4b GGUF on llama.cpp: typed decisions with calibrated probabilities, no torch forward pass.
2
+
3
+ Needs `pip install decider-ai==1.5.0 llama-cpp-python` (llama-cpp-python 0.3.35 or newer; build it with
4
+ CMAKE_ARGS="-DGGML_CUDA=on" or "-DGGML_METAL=on" for a GPU). The prompt is built by decider.prompt with the HF tokenizer
5
+ shipped in this repository, llama.cpp runs the rows, and the answer is the softmax over the option-letter logits at each
6
+ answer slot, divided by the temperatures in decider_config.json. One prompt per llama_decode (see the model card for why).
7
+
8
+ from decide_gguf import GGUFDecider
9
+ d = GGUFDecider("decider-4b-v2.1-Q4_K_M.gguf") # tokenizer and decider_config.json from the same folder
10
+ d.decide("My card was charged twice for the same purchase.",
11
+ [{"question": "Which department should handle this?", "options": ["billing", "technical", "sales"]}])
12
+ """
13
+ import ctypes, json, os
14
+
15
+ import numpy as np
16
+ import llama_cpp as L
17
+ from transformers import AutoTokenizer
18
+ from decider.infer import Example, Q, _NoShuffle
19
+ from decider.prompt import MAX_OPTIONS, build, letter_ids
20
+ from decider import temperature as TT
21
+
22
+ _QUIET = L.llama_log_callback(lambda level, text, data: None)
23
+
24
+
25
+ class GGUFDecider:
26
+ def __init__(self, gguf_path, folder=None, n_ctx=32768, n_gpu_layers=-1, n_threads=None, verbose=False):
27
+ folder = folder or os.path.dirname(os.path.abspath(gguf_path))
28
+ cfg = json.load(open(os.path.join(folder, "decider_config.json")))
29
+ (self.T, self.T_by_type), _ = TT.from_config(cfg)
30
+ self.name = "decider-" + str(cfg.get("version", "dev"))
31
+ self.tok = AutoTokenizer.from_pretrained(folder)
32
+ self.letters = np.asarray(letter_ids(self.tok))
33
+ if not verbose:
34
+ L.llama_log_set(_QUIET, ctypes.c_void_p(0))
35
+ L.llama_backend_init()
36
+ mp = L.llama_model_default_params(); mp.n_gpu_layers = n_gpu_layers
37
+ self.model = L.llama_model_load_from_file(os.fsencode(gguf_path), mp)
38
+ if not self.model:
39
+ raise RuntimeError(f"llama.cpp could not load {gguf_path}")
40
+ cp = L.llama_context_default_params()
41
+ cp.n_ctx = n_ctx; cp.n_batch = n_ctx; cp.n_ubatch = min(2048, n_ctx); cp.n_seq_max = 1
42
+ if n_threads:
43
+ cp.n_threads = cp.n_threads_batch = n_threads
44
+ self.ctx = L.llama_init_from_model(self.model, cp)
45
+ self.n_ctx = n_ctx
46
+ self.n_vocab = L.llama_vocab_n_tokens(L.llama_model_get_vocab(self.model))
47
+ self.batch = L.llama_batch_init(n_ctx, 0, 1)
48
+
49
+ def _slot_logits(self, ids, slots):
50
+ if len(ids) > self.n_ctx:
51
+ raise ValueError(f"prompt of {len(ids)} tokens exceeds n_ctx {self.n_ctx}")
52
+ L.llama_memory_clear(L.llama_get_memory(self.ctx), True)
53
+ b, want = self.batch, set(slots)
54
+ for i, t in enumerate(ids):
55
+ b.token[i] = t; b.pos[i] = i; b.n_seq_id[i] = 1; b.seq_id[i][0] = 0; b.logits[i] = i in want
56
+ b.n_tokens = len(ids)
57
+ if L.llama_decode(self.ctx, b) != 0:
58
+ raise RuntimeError("llama_decode failed")
59
+ rows = []
60
+ for s in slots:
61
+ p = ctypes.cast(L.llama_get_logits_ith(self.ctx, s), ctypes.POINTER(ctypes.c_float))
62
+ rows.append(np.ctypeslib.as_array(p, shape=(self.n_vocab,))[self.letters].astype(np.float64))
63
+ return rows
64
+
65
+ def decide(self, context, questions, max_ctx_tokens=1536):
66
+ """questions: [{"question": str, "options": [str, ...]}, ...] (2..255 options).
67
+ -> [{"choice", "confidence", "probs"}, ...], one per question, as decider.infer.Decider.decide returns them."""
68
+ for q in questions:
69
+ assert 2 <= len(q["options"]) <= MAX_OPTIONS, f"2..{MAX_OPTIONS} options required"
70
+ item = build(Example(context, [Q(q["question"], list(q["options"]), 0) for q in questions]), self.tok, _NoShuffle(),
71
+ max_options=MAX_OPTIONS, max_ctx_tokens=max_ctx_tokens)
72
+ T = TT.for_types(self.T, self.T_by_type, ["choice"] * len(questions))
73
+ out = []
74
+ for q, lg, n, t in zip(questions, self._slot_logits(item["ids"], item["slots"]), item["nopts"], T):
75
+ z = lg[:n] / t; p = np.exp(z - z.max()); p /= p.sum()
76
+ j = int(p.argmax())
77
+ out.append(dict(choice=q["options"][j], confidence=float(p[j]), probs=dict(zip(q["options"], p.tolist()))))
78
+ return out
decider-4b-v2.1-BF16.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a432ef364d83b62b70910f61bd08c2639e05b014d37285188779d974782896d8
3
+ size 8424393760
decider-4b-v2.1-Q4_K_M.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c7083fcfc93f650cd66caeade4a3840ef9eeeeb1e80920907c554f70619d7c56
3
+ size 2708804640
decider-4b-v2.1-Q8_0.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:74614659ab849bc0ec4f0431e1f84ae8bc11fe30b42c3a235a8b0049d36cbc3d
3
+ size 4482403360
decider_config.json ADDED
@@ -0,0 +1,20 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "temperature": 1.099,
3
+ "temperature_by_type": {
4
+ "choice": 1.11,
5
+ "noul": 1.56,
6
+ "score": 1.287
7
+ },
8
+ "neutralize_none": false,
9
+ "version": "4b-v2.1",
10
+ "base": "Mapika/decider-4b v1 + LoRA (merged); v1 is Qwen/Qwen3.5-4B-Base + one supervised pass over mixture v2",
11
+ "layout": "plain",
12
+ "max_options": 255,
13
+ "max_state_tokens": 32768,
14
+ "schema_first": false,
15
+ "schema_first_trained": false,
16
+ "isolated_levels": true,
17
+ "release_date": "2026-09-24",
18
+ "requires": "decider-ai>=1.4.0 for temperature_by_type; older versions serve every answer at temperature",
19
+ "stage": "decider-4b v1 + LoRA rank 64 (alpha 128) on attention and MLP, LR 1e-4, 2 epochs (1,518 steps of 65,536 tokens) over v2's 29,325-row mix in the plain state-first layout (generated decision families with code-computed answers, Qwen3.6-27B-written document questions kept when two independent answers agreed, human-labelled public sets, replay of v1's mixture v2), with the replay rows trained toward v1's own answer distribution (KL to v1) instead of their labels, merged into the bf16 weights; no RL stage; temperature fitted by NLL on 61 in-task regression tasks (the 67 in-task tasks without banking77, clinc_oos, mmlu, arc, winogrande, hellaswag); temperature_by_type fitted with decider.calibrate.fit_by_type on the same regression rows plus our own validation rows (choice, noul and score answers)"
20
+ }
tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:06b9509352d2af50381ab2247e083b80d32d5c0aba91c272ca9ff729b6a0e523
3
+ size 19989325
tokenizer_config.json ADDED
@@ -0,0 +1,32 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_prefix_space": false,
3
+ "audio_bos_token": "<|audio_start|>",
4
+ "audio_eos_token": "<|audio_end|>",
5
+ "audio_token": "<|audio_pad|>",
6
+ "backend": "tokenizers",
7
+ "bos_token": null,
8
+ "clean_up_tokenization_spaces": false,
9
+ "eos_token": "<|endoftext|>",
10
+ "errors": "replace",
11
+ "image_token": "<|image_pad|>",
12
+ "is_local": true,
13
+ "local_files_only": false,
14
+ "model_max_length": 262144,
15
+ "model_specific_special_tokens": {
16
+ "audio_bos_token": "<|audio_start|>",
17
+ "audio_eos_token": "<|audio_end|>",
18
+ "audio_token": "<|audio_pad|>",
19
+ "image_token": "<|image_pad|>",
20
+ "video_token": "<|video_pad|>",
21
+ "vision_bos_token": "<|vision_start|>",
22
+ "vision_eos_token": "<|vision_end|>"
23
+ },
24
+ "pad_token": "<|endoftext|>",
25
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
26
+ "split_special_tokens": false,
27
+ "tokenizer_class": "Qwen2Tokenizer",
28
+ "unk_token": null,
29
+ "video_token": "<|video_pad|>",
30
+ "vision_bos_token": "<|vision_start|>",
31
+ "vision_eos_token": "<|vision_end|>"
32
+ }