kleeedolinux commited on
Commit
a85b127
·
1 Parent(s): 7a169d1

Preserve Boolean criteria and add pinned accuracy reproduction

Browse files
README.md CHANGED
@@ -128,12 +128,48 @@ Julia 1 and Dumont support the same named-question interface: `predict(state=...
128
  | --- | --- | --- |
129
  | `choice` | Mapping of 2–20 IDs to nonempty descriptions | `choice`: winning ID |
130
  | `score` | Ordered list of 2–20 rubric descriptions | `score`: expected zero-based rubric index |
131
- | `noul` | Omit criteria; false/true order is automatic | `noul`: probability of true |
132
 
133
  Each named answer includes `type` and `probabilities`, keyed by caller IDs for choices, zero-based strings for scores, or `false`/`true` for Boolean decisions. Choice and score also include `max_probability`. These are full softmax probabilities, without display rounding; they are not guaranteed certainty. Questions are independently scored in a batch.
134
 
135
  Julia 1's native 2–20 option limit still applies. The runtime defaults to the checkpoint’s **8,192-token** combined state/question/options limit. The example uses a 512-token question-and-options budget; each option has a 48-token limit. Strict encoding rejects overflow. Historical accuracy benchmarks above used 1,024 tokens. An [8,192-token CPU smoke test](metrics/context-8k-smoke.json) passed; 8k task accuracy is not established.
136
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
137
  ### Existing list API
138
 
139
  The existing interface remains supported:
 
128
  | --- | --- | --- |
129
  | `choice` | Mapping of 2–20 IDs to nonempty descriptions | `choice`: winning ID |
130
  | `score` | Ordered list of 2–20 rubric descriptions | `score`: expected zero-based rubric index |
131
+ | `noul` | Optional mapping of `false` and `true` to descriptions; omitted criteria use literal false/true | `noul`: probability of true |
132
 
133
  Each named answer includes `type` and `probabilities`, keyed by caller IDs for choices, zero-based strings for scores, or `false`/`true` for Boolean decisions. Choice and score also include `max_probability`. These are full softmax probabilities, without display rounding; they are not guaranteed certainty. Questions are independently scored in a batch.
134
 
135
  Julia 1's native 2–20 option limit still applies. The runtime defaults to the checkpoint’s **8,192-token** combined state/question/options limit. The example uses a 512-token question-and-options budget; each option has a 48-token limit. Strict encoding rejects overflow. Historical accuracy benchmarks above used 1,024 tokens. An [8,192-token CPU smoke test](metrics/context-8k-smoke.json) passed; 8k task accuracy is not established.
136
 
137
+ ### Reproduce typed-decision accuracy
138
+
139
+ ```sh
140
+ python -m pip install -e '.[benchmark]'
141
+ python scripts/reproduce_typed.py --output benchmark/typed-cpu --literal-noul-ablation
142
+ ```
143
+
144
+ Run from this model repository with the real weights downloaded. Use a fresh output
145
+ folder. The script verifies the published weight hash, downloads the pinned test
146
+ Parquet and checks its SHA-256, then writes predictions and CPU results for all
147
+ 400 cases / 2,000 questions. `--dataset PATH` reuses a local copy. The optional
148
+ ablation also evaluates Boolean questions with literal false/true descriptions.
149
+
150
+ Dataset revision: `c76749ec58bd8c3d2ea706b31c333a9059c38f90`.
151
+ Test Parquet SHA-256: `4f294f218ea1da27f3efef936359389c62ea4d3973a41457732990f1d31b647c`.
152
+
153
+ Preserve the dataset's `criteria` descriptions, including Boolean descriptions in
154
+ false/true order. Only questions without Boolean criteria use literal `false` and
155
+ `true`. Accuracy uses the highest-probability option for all three types; do not
156
+ round the score API's expected index. No threshold is fitted on the test data.
157
+
158
+ 400 of the 600 Boolean questions have descriptive criteria. Earlier versions of
159
+ the named-question API incorrectly discarded those descriptions; the low-level
160
+ benchmark retained them. The API now preserves supplied descriptions, while
161
+ questions without criteria keep their existing literal false/true behavior.
162
+
163
+ The historical report used CUDA BF16; this harness runs CPU FP32 and records its
164
+ own results. It does not regenerate the unrelated classification or MASSIVE metrics.
165
+
166
+ The September 26 CPU FP32 reproduction (torch 2.14.0, transformers 5.0.0,
167
+ marker-only head disabled) gives **426/600 choice, 542/800 score, 483/600 noul**.
168
+ The historical CUDA BF16 totals above remain separate; the smaller CPU/GPU
169
+ prediction differences have not been isolated to a single cause. Replacing only
170
+ the Boolean descriptions with literal false/true reproduces **391/600 noul** on
171
+ the same CPU run. See [the paired CPU results](metrics/typed-cpu-20260926.json).
172
+
173
  ### Existing list API
174
 
175
  The existing interface remains supported:
julia/typed.py CHANGED
@@ -18,7 +18,14 @@ def predict_typed(engine, state, questions):
18
  elif kind=='score':
19
  if not isinstance(criteria,list):raise ValueError('Score requires an ordered rubric')
20
  labels=criteria;keys=[str(i) for i in range(len(labels))]
21
- elif kind=='noul':keys=['false','true'];labels=['false','true']
 
 
 
 
 
 
 
22
  else:raise ValueError('Unsupported question type')
23
  row=dict(state=state,question=q.get('instructions'),type=kind,options=labels)
24
  validate_row(row,len(rows)+1)
 
18
  elif kind=='score':
19
  if not isinstance(criteria,list):raise ValueError('Score requires an ordered rubric')
20
  labels=criteria;keys=[str(i) for i in range(len(labels))]
21
+ elif kind=='noul':
22
+ keys=['false','true']
23
+ if criteria is None:
24
+ labels=list(keys)
25
+ else:
26
+ if not isinstance(criteria,dict) or set(criteria)!=set(keys):
27
+ raise ValueError('Noul criteria must map false and true to descriptions')
28
+ labels=[criteria[key] for key in keys]
29
  else:raise ValueError('Unsupported question type')
30
  row=dict(state=state,question=q.get('instructions'),type=kind,options=labels)
31
  validate_row(row,len(rows)+1)
metrics/typed-cpu-20260926.json ADDED
@@ -0,0 +1,47 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "weights_sha256": "df853bf7fe424420011f3d0c47a05d7341aa9eefa7fb9f203ea4aada4ad95b72",
3
+ "dataset_revision": "c76749ec58bd8c3d2ea706b31c333a9059c38f90",
4
+ "dataset_sha256": "4f294f218ea1da27f3efef936359389c62ea4d3973a41457732990f1d31b647c",
5
+ "torch": "2.14.0+cpu",
6
+ "transformers": "5.0.0",
7
+ "device": "cpu",
8
+ "threads": 4,
9
+ "marker_only_head": false,
10
+ "encoding": {
11
+ "max_length": 1024,
12
+ "head_length": 512,
13
+ "strict": true
14
+ },
15
+ "original_criteria": {
16
+ "by_type": {
17
+ "choice": {
18
+ "count": 600,
19
+ "correct": 426,
20
+ "accuracy": 0.71
21
+ },
22
+ "noul": {
23
+ "count": 600,
24
+ "correct": 483,
25
+ "accuracy": 0.805
26
+ },
27
+ "score": {
28
+ "count": 800,
29
+ "correct": 542,
30
+ "accuracy": 0.6775
31
+ }
32
+ },
33
+ "elapsed_seconds": 601.7335131389846
34
+ },
35
+ "literal_noul_ablation": {
36
+ "by_type": {
37
+ "noul": {
38
+ "count": 600,
39
+ "correct": 391,
40
+ "accuracy": 0.6516666666666666
41
+ }
42
+ },
43
+ "elapsed_seconds": 170.3590434489888
44
+ },
45
+ "executed_probe_sha256": "ac274e1eff64d96d377a53ecbca2c9676b9197c78abb84744fef84c65f0d6ad1",
46
+ "note": "Same verified weights, pinned data and CPU FP32 settings for both passes; only noul option descriptions change. Historical CUDA BF16 totals remain separate."
47
+ }
pyproject.toml CHANGED
@@ -13,6 +13,7 @@ include = ["julia*"]
13
  exclude = ["julia.router.tests*"]
14
 
15
  [project.optional-dependencies]
 
16
  cuda = ["bitsandbytes>=0.48,<0.50", "accelerate>=1.10,<2"]
17
 
18
  [tool.setuptools.package-data]
 
13
  exclude = ["julia.router.tests*"]
14
 
15
  [project.optional-dependencies]
16
+ benchmark = ["pyarrow>=15"]
17
  cuda = ["bitsandbytes>=0.48,<0.50", "accelerate>=1.10,<2"]
18
 
19
  [tool.setuptools.package-data]
scripts/reproduce_typed.py ADDED
@@ -0,0 +1,111 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Reproduce Julia-1 typed decisions; optionally ablate Boolean descriptions.
2
+ Requires torch, transformers 5.0.x, safetensors, numpy and pyarrow.
3
+ """
4
+ import argparse
5
+ import collections
6
+ import hashlib
7
+ import json
8
+ from pathlib import Path
9
+ import sys
10
+ import time
11
+ import urllib.request
12
+ import shutil
13
+
14
+ REVISION = 'c76749ec58bd8c3d2ea706b31c333a9059c38f90'
15
+ DATA_SHA256 = '4f294f218ea1da27f3efef936359389c62ea4d3973a41457732990f1d31b647c'
16
+ WEIGHTS_SHA256 = 'df853bf7fe424420011f3d0c47a05d7341aa9eefa7fb9f203ea4aada4ad95b72'
17
+
18
+ def digest(path):
19
+ with Path(path).open('rb') as stream:
20
+ return hashlib.file_digest(stream, 'sha256').hexdigest()
21
+
22
+ def main():
23
+ ap = argparse.ArgumentParser(description=__doc__)
24
+ ap.add_argument('--checkpoint', type=Path, default=Path(__file__).resolve().parents[1])
25
+ ap.add_argument('--dataset', type=Path, help='Existing pinned Parquet; otherwise download and verify it')
26
+ ap.add_argument('--output', type=Path, required=True)
27
+ ap.add_argument('--threads', type=int, default=4)
28
+ ap.add_argument('--literal-noul-ablation', action='store_true')
29
+ args = ap.parse_args()
30
+ checkpoint = args.checkpoint.resolve()
31
+ if digest(checkpoint / 'model.safetensors') != WEIGHTS_SHA256:
32
+ raise ValueError('Weights do not match accuracy-20260924.json')
33
+ args.output.mkdir(parents=True, exist_ok=False)
34
+ if args.dataset is None:
35
+ args.dataset = args.output / 'test-00000-of-00001.parquet'
36
+ url = (f'https://huggingface.co/datasets/LocalLLaMA/typed-decisions/resolve/{REVISION}'
37
+ '/all/test-00000-of-00001.parquet')
38
+ temporary = args.dataset.with_suffix('.partial')
39
+ with urllib.request.urlopen(url, timeout=180) as source, temporary.open('wb') as target:
40
+ shutil.copyfileobj(source, target)
41
+ temporary.replace(args.dataset)
42
+ if digest(args.dataset) != DATA_SHA256:
43
+ raise ValueError('Dataset does not match the pinned published evaluation')
44
+ sys.path.insert(0, str(checkpoint))
45
+ import pyarrow.parquet as pq
46
+ import torch
47
+ import transformers
48
+ from julia.router.engine import FastEngine
49
+ torch.set_num_threads(args.threads)
50
+ cases = pq.read_table(args.dataset).to_pylist()
51
+ rows, metadata = [], []
52
+ for case in cases:
53
+ state = json.loads(case['state'])
54
+ gold = json.loads(case['gold'])
55
+ for question_id, question in json.loads(case['questions']).items():
56
+ kind = question['type']
57
+ criteria = question.get('criteria')
58
+ if criteria is None and kind == 'noul':
59
+ criteria = {'false': 'false', 'true': 'true'}
60
+ if isinstance(criteria, list):
61
+ criteria = {str(i): value for i, value in enumerate(criteria)}
62
+ if not isinstance(criteria, dict):
63
+ raise ValueError('Missing option descriptions')
64
+ keys = ['false', 'true'] if kind == 'noul' else list(criteria)
65
+ rows.append(dict(state=state, question=question['instructions'],
66
+ type=kind, options=[criteria[key] for key in keys]))
67
+ metadata.append(dict(id=case['id']+':'+question_id, type=kind,
68
+ keys=keys, gold=str(gold[question_id]['label'])))
69
+ counts = collections.Counter(row['type'] for row in rows)
70
+ assert len(cases) == 400 and counts == {'choice': 600, 'score': 800, 'noul': 600}, counts
71
+ engine = FastEngine(checkpoint, device='cpu', transformer_backend='torch',
72
+ strict_encoding=True, max_length=1024, head_length=512,
73
+ batch_size=16, marker_only_head=False)
74
+ result = dict(weights_sha256=WEIGHTS_SHA256, dataset_revision=REVISION,
75
+ dataset_sha256=DATA_SHA256, torch=torch.__version__,
76
+ transformers=transformers.__version__, device='cpu',
77
+ threads=args.threads, marker_only_head=False,
78
+ encoding=dict(max_length=1024, head_length=512, strict=True),
79
+ runtime_sha256={str(p.relative_to(checkpoint)): digest(p)
80
+ for p in sorted((checkpoint/'julia').rglob('*.py'))})
81
+ def evaluate(name, input_rows, labels):
82
+ started = time.monotonic()
83
+ stats = collections.defaultdict(lambda: dict(count=0, correct=0))
84
+ with (args.output / (name+'-predictions.jsonl')).open('w') as stream:
85
+ # One example per forward matches the original benchmark protocol.
86
+ for index, (row, label) in enumerate(zip(input_rows, labels), 1):
87
+ answer = engine.predict([row])[0]
88
+ prediction = label['keys'][answer['index']]
89
+ correct = prediction == label['gold']
90
+ stats[label['type']]['count'] += 1
91
+ stats[label['type']]['correct'] += int(correct)
92
+ stream.write(json.dumps(dict(**label, prediction=prediction, correct=correct,
93
+ options=row['options'], probabilities=answer['probabilities']))+'\n')
94
+ if index % 200 == 0:
95
+ print(json.dumps(dict(mode=name, done=index, total=len(input_rows))), flush=True)
96
+ for item in stats.values():
97
+ item['accuracy'] = item['correct']/item['count']
98
+ summary = dict(by_type=dict(stats), elapsed_seconds=time.monotonic()-started)
99
+ print(json.dumps(dict(mode=name, **summary)), flush=True)
100
+ return summary
101
+ result['original_criteria'] = evaluate('original-criteria', rows, metadata)
102
+ if args.literal_noul_ablation:
103
+ selected = [(dict(row, options=['false', 'true']), label)
104
+ for row, label in zip(rows, metadata) if row['type']=='noul']
105
+ result['literal_noul_ablation'] = evaluate('literal-noul',
106
+ [x[0] for x in selected], [x[1] for x in selected])
107
+ (args.output/'results.json').write_text(json.dumps(result, indent=2)+'\n')
108
+ print(json.dumps(result, indent=2), flush=True)
109
+
110
+ if __name__ == '__main__':
111
+ main()
tests/test_typed_api.py CHANGED
@@ -25,3 +25,24 @@ class TypedAPITests(unittest.TestCase):
25
  self.engine.logits.return_value=[[1.,0.]]
26
  q={'q':dict(type='choice',instructions='q',criteria={'a':'A','b':'B'})}
27
  self.assertEqual(self.engine.predict('s',q),self.engine.predict(state='s',questions=q))
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
25
  self.engine.logits.return_value=[[1.,0.]]
26
  q={'q':dict(type='choice',instructions='q',criteria={'a':'A','b':'B'})}
27
  self.assertEqual(self.engine.predict('s',q),self.engine.predict(state='s',questions=q))
28
+
29
+ def test_noul_preserves_descriptions_in_false_true_order(self):
30
+ self.engine.logits.return_value=[[0.,2.]]
31
+ questions={'review':dict(type='noul',instructions='Needs human review?',
32
+ criteria={'true':'A human should inspect this run.',
33
+ 'false':'No human attention is warranted.'})}
34
+ answer=self.engine.predict(state='context',questions=questions)['answers']['review']
35
+ self.assertEqual(self.engine.logits.call_args.args[0][0]['options'],
36
+ ['No human attention is warranted.','A human should inspect this run.'])
37
+ self.assertEqual(list(answer['probabilities']),['false','true'])
38
+ self.assertGreater(answer['noul'],.5)
39
+
40
+ def test_noul_rejects_invalid_criteria_before_inference(self):
41
+ self.engine.logits.return_value=[[0.,2.]]
42
+ for criteria in ({}, {'true':'yes'}, {'false':'no','true':'yes','other':'maybe'},
43
+ ['no','yes'], {'false':'','true':'yes'}):
44
+ with self.subTest(criteria=criteria):
45
+ questions={'q':dict(type='noul',instructions='q',criteria=criteria)}
46
+ with self.assertRaises(ValueError):
47
+ self.engine.predict(state='context',questions=questions)
48
+ self.engine.logits.assert_not_called()