kleeedolinux commited on
Commit ·
a85b127
1
Parent(s): 7a169d1
Preserve Boolean criteria and add pinned accuracy reproduction
Browse files- README.md +37 -1
- julia/typed.py +8 -1
- metrics/typed-cpu-20260926.json +47 -0
- pyproject.toml +1 -0
- scripts/reproduce_typed.py +111 -0
- tests/test_typed_api.py +21 -0
README.md
CHANGED
|
@@ -128,12 +128,48 @@ Julia 1 and Dumont support the same named-question interface: `predict(state=...
|
|
| 128 |
| --- | --- | --- |
|
| 129 |
| `choice` | Mapping of 2–20 IDs to nonempty descriptions | `choice`: winning ID |
|
| 130 |
| `score` | Ordered list of 2–20 rubric descriptions | `score`: expected zero-based rubric index |
|
| 131 |
-
| `noul` |
|
| 132 |
|
| 133 |
Each named answer includes `type` and `probabilities`, keyed by caller IDs for choices, zero-based strings for scores, or `false`/`true` for Boolean decisions. Choice and score also include `max_probability`. These are full softmax probabilities, without display rounding; they are not guaranteed certainty. Questions are independently scored in a batch.
|
| 134 |
|
| 135 |
Julia 1's native 2–20 option limit still applies. The runtime defaults to the checkpoint’s **8,192-token** combined state/question/options limit. The example uses a 512-token question-and-options budget; each option has a 48-token limit. Strict encoding rejects overflow. Historical accuracy benchmarks above used 1,024 tokens. An [8,192-token CPU smoke test](metrics/context-8k-smoke.json) passed; 8k task accuracy is not established.
|
| 136 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 137 |
### Existing list API
|
| 138 |
|
| 139 |
The existing interface remains supported:
|
|
|
|
| 128 |
| --- | --- | --- |
|
| 129 |
| `choice` | Mapping of 2–20 IDs to nonempty descriptions | `choice`: winning ID |
|
| 130 |
| `score` | Ordered list of 2–20 rubric descriptions | `score`: expected zero-based rubric index |
|
| 131 |
+
| `noul` | Optional mapping of `false` and `true` to descriptions; omitted criteria use literal false/true | `noul`: probability of true |
|
| 132 |
|
| 133 |
Each named answer includes `type` and `probabilities`, keyed by caller IDs for choices, zero-based strings for scores, or `false`/`true` for Boolean decisions. Choice and score also include `max_probability`. These are full softmax probabilities, without display rounding; they are not guaranteed certainty. Questions are independently scored in a batch.
|
| 134 |
|
| 135 |
Julia 1's native 2–20 option limit still applies. The runtime defaults to the checkpoint’s **8,192-token** combined state/question/options limit. The example uses a 512-token question-and-options budget; each option has a 48-token limit. Strict encoding rejects overflow. Historical accuracy benchmarks above used 1,024 tokens. An [8,192-token CPU smoke test](metrics/context-8k-smoke.json) passed; 8k task accuracy is not established.
|
| 136 |
|
| 137 |
+
### Reproduce typed-decision accuracy
|
| 138 |
+
|
| 139 |
+
```sh
|
| 140 |
+
python -m pip install -e '.[benchmark]'
|
| 141 |
+
python scripts/reproduce_typed.py --output benchmark/typed-cpu --literal-noul-ablation
|
| 142 |
+
```
|
| 143 |
+
|
| 144 |
+
Run from this model repository with the real weights downloaded. Use a fresh output
|
| 145 |
+
folder. The script verifies the published weight hash, downloads the pinned test
|
| 146 |
+
Parquet and checks its SHA-256, then writes predictions and CPU results for all
|
| 147 |
+
400 cases / 2,000 questions. `--dataset PATH` reuses a local copy. The optional
|
| 148 |
+
ablation also evaluates Boolean questions with literal false/true descriptions.
|
| 149 |
+
|
| 150 |
+
Dataset revision: `c76749ec58bd8c3d2ea706b31c333a9059c38f90`.
|
| 151 |
+
Test Parquet SHA-256: `4f294f218ea1da27f3efef936359389c62ea4d3973a41457732990f1d31b647c`.
|
| 152 |
+
|
| 153 |
+
Preserve the dataset's `criteria` descriptions, including Boolean descriptions in
|
| 154 |
+
false/true order. Only questions without Boolean criteria use literal `false` and
|
| 155 |
+
`true`. Accuracy uses the highest-probability option for all three types; do not
|
| 156 |
+
round the score API's expected index. No threshold is fitted on the test data.
|
| 157 |
+
|
| 158 |
+
400 of the 600 Boolean questions have descriptive criteria. Earlier versions of
|
| 159 |
+
the named-question API incorrectly discarded those descriptions; the low-level
|
| 160 |
+
benchmark retained them. The API now preserves supplied descriptions, while
|
| 161 |
+
questions without criteria keep their existing literal false/true behavior.
|
| 162 |
+
|
| 163 |
+
The historical report used CUDA BF16; this harness runs CPU FP32 and records its
|
| 164 |
+
own results. It does not regenerate the unrelated classification or MASSIVE metrics.
|
| 165 |
+
|
| 166 |
+
The September 26 CPU FP32 reproduction (torch 2.14.0, transformers 5.0.0,
|
| 167 |
+
marker-only head disabled) gives **426/600 choice, 542/800 score, 483/600 noul**.
|
| 168 |
+
The historical CUDA BF16 totals above remain separate; the smaller CPU/GPU
|
| 169 |
+
prediction differences have not been isolated to a single cause. Replacing only
|
| 170 |
+
the Boolean descriptions with literal false/true reproduces **391/600 noul** on
|
| 171 |
+
the same CPU run. See [the paired CPU results](metrics/typed-cpu-20260926.json).
|
| 172 |
+
|
| 173 |
### Existing list API
|
| 174 |
|
| 175 |
The existing interface remains supported:
|
julia/typed.py
CHANGED
|
@@ -18,7 +18,14 @@ def predict_typed(engine, state, questions):
|
|
| 18 |
elif kind=='score':
|
| 19 |
if not isinstance(criteria,list):raise ValueError('Score requires an ordered rubric')
|
| 20 |
labels=criteria;keys=[str(i) for i in range(len(labels))]
|
| 21 |
-
elif kind=='noul':
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 22 |
else:raise ValueError('Unsupported question type')
|
| 23 |
row=dict(state=state,question=q.get('instructions'),type=kind,options=labels)
|
| 24 |
validate_row(row,len(rows)+1)
|
|
|
|
| 18 |
elif kind=='score':
|
| 19 |
if not isinstance(criteria,list):raise ValueError('Score requires an ordered rubric')
|
| 20 |
labels=criteria;keys=[str(i) for i in range(len(labels))]
|
| 21 |
+
elif kind=='noul':
|
| 22 |
+
keys=['false','true']
|
| 23 |
+
if criteria is None:
|
| 24 |
+
labels=list(keys)
|
| 25 |
+
else:
|
| 26 |
+
if not isinstance(criteria,dict) or set(criteria)!=set(keys):
|
| 27 |
+
raise ValueError('Noul criteria must map false and true to descriptions')
|
| 28 |
+
labels=[criteria[key] for key in keys]
|
| 29 |
else:raise ValueError('Unsupported question type')
|
| 30 |
row=dict(state=state,question=q.get('instructions'),type=kind,options=labels)
|
| 31 |
validate_row(row,len(rows)+1)
|
metrics/typed-cpu-20260926.json
ADDED
|
@@ -0,0 +1,47 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"weights_sha256": "df853bf7fe424420011f3d0c47a05d7341aa9eefa7fb9f203ea4aada4ad95b72",
|
| 3 |
+
"dataset_revision": "c76749ec58bd8c3d2ea706b31c333a9059c38f90",
|
| 4 |
+
"dataset_sha256": "4f294f218ea1da27f3efef936359389c62ea4d3973a41457732990f1d31b647c",
|
| 5 |
+
"torch": "2.14.0+cpu",
|
| 6 |
+
"transformers": "5.0.0",
|
| 7 |
+
"device": "cpu",
|
| 8 |
+
"threads": 4,
|
| 9 |
+
"marker_only_head": false,
|
| 10 |
+
"encoding": {
|
| 11 |
+
"max_length": 1024,
|
| 12 |
+
"head_length": 512,
|
| 13 |
+
"strict": true
|
| 14 |
+
},
|
| 15 |
+
"original_criteria": {
|
| 16 |
+
"by_type": {
|
| 17 |
+
"choice": {
|
| 18 |
+
"count": 600,
|
| 19 |
+
"correct": 426,
|
| 20 |
+
"accuracy": 0.71
|
| 21 |
+
},
|
| 22 |
+
"noul": {
|
| 23 |
+
"count": 600,
|
| 24 |
+
"correct": 483,
|
| 25 |
+
"accuracy": 0.805
|
| 26 |
+
},
|
| 27 |
+
"score": {
|
| 28 |
+
"count": 800,
|
| 29 |
+
"correct": 542,
|
| 30 |
+
"accuracy": 0.6775
|
| 31 |
+
}
|
| 32 |
+
},
|
| 33 |
+
"elapsed_seconds": 601.7335131389846
|
| 34 |
+
},
|
| 35 |
+
"literal_noul_ablation": {
|
| 36 |
+
"by_type": {
|
| 37 |
+
"noul": {
|
| 38 |
+
"count": 600,
|
| 39 |
+
"correct": 391,
|
| 40 |
+
"accuracy": 0.6516666666666666
|
| 41 |
+
}
|
| 42 |
+
},
|
| 43 |
+
"elapsed_seconds": 170.3590434489888
|
| 44 |
+
},
|
| 45 |
+
"executed_probe_sha256": "ac274e1eff64d96d377a53ecbca2c9676b9197c78abb84744fef84c65f0d6ad1",
|
| 46 |
+
"note": "Same verified weights, pinned data and CPU FP32 settings for both passes; only noul option descriptions change. Historical CUDA BF16 totals remain separate."
|
| 47 |
+
}
|
pyproject.toml
CHANGED
|
@@ -13,6 +13,7 @@ include = ["julia*"]
|
|
| 13 |
exclude = ["julia.router.tests*"]
|
| 14 |
|
| 15 |
[project.optional-dependencies]
|
|
|
|
| 16 |
cuda = ["bitsandbytes>=0.48,<0.50", "accelerate>=1.10,<2"]
|
| 17 |
|
| 18 |
[tool.setuptools.package-data]
|
|
|
|
| 13 |
exclude = ["julia.router.tests*"]
|
| 14 |
|
| 15 |
[project.optional-dependencies]
|
| 16 |
+
benchmark = ["pyarrow>=15"]
|
| 17 |
cuda = ["bitsandbytes>=0.48,<0.50", "accelerate>=1.10,<2"]
|
| 18 |
|
| 19 |
[tool.setuptools.package-data]
|
scripts/reproduce_typed.py
ADDED
|
@@ -0,0 +1,111 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Reproduce Julia-1 typed decisions; optionally ablate Boolean descriptions.
|
| 2 |
+
Requires torch, transformers 5.0.x, safetensors, numpy and pyarrow.
|
| 3 |
+
"""
|
| 4 |
+
import argparse
|
| 5 |
+
import collections
|
| 6 |
+
import hashlib
|
| 7 |
+
import json
|
| 8 |
+
from pathlib import Path
|
| 9 |
+
import sys
|
| 10 |
+
import time
|
| 11 |
+
import urllib.request
|
| 12 |
+
import shutil
|
| 13 |
+
|
| 14 |
+
REVISION = 'c76749ec58bd8c3d2ea706b31c333a9059c38f90'
|
| 15 |
+
DATA_SHA256 = '4f294f218ea1da27f3efef936359389c62ea4d3973a41457732990f1d31b647c'
|
| 16 |
+
WEIGHTS_SHA256 = 'df853bf7fe424420011f3d0c47a05d7341aa9eefa7fb9f203ea4aada4ad95b72'
|
| 17 |
+
|
| 18 |
+
def digest(path):
|
| 19 |
+
with Path(path).open('rb') as stream:
|
| 20 |
+
return hashlib.file_digest(stream, 'sha256').hexdigest()
|
| 21 |
+
|
| 22 |
+
def main():
|
| 23 |
+
ap = argparse.ArgumentParser(description=__doc__)
|
| 24 |
+
ap.add_argument('--checkpoint', type=Path, default=Path(__file__).resolve().parents[1])
|
| 25 |
+
ap.add_argument('--dataset', type=Path, help='Existing pinned Parquet; otherwise download and verify it')
|
| 26 |
+
ap.add_argument('--output', type=Path, required=True)
|
| 27 |
+
ap.add_argument('--threads', type=int, default=4)
|
| 28 |
+
ap.add_argument('--literal-noul-ablation', action='store_true')
|
| 29 |
+
args = ap.parse_args()
|
| 30 |
+
checkpoint = args.checkpoint.resolve()
|
| 31 |
+
if digest(checkpoint / 'model.safetensors') != WEIGHTS_SHA256:
|
| 32 |
+
raise ValueError('Weights do not match accuracy-20260924.json')
|
| 33 |
+
args.output.mkdir(parents=True, exist_ok=False)
|
| 34 |
+
if args.dataset is None:
|
| 35 |
+
args.dataset = args.output / 'test-00000-of-00001.parquet'
|
| 36 |
+
url = (f'https://huggingface.co/datasets/LocalLLaMA/typed-decisions/resolve/{REVISION}'
|
| 37 |
+
'/all/test-00000-of-00001.parquet')
|
| 38 |
+
temporary = args.dataset.with_suffix('.partial')
|
| 39 |
+
with urllib.request.urlopen(url, timeout=180) as source, temporary.open('wb') as target:
|
| 40 |
+
shutil.copyfileobj(source, target)
|
| 41 |
+
temporary.replace(args.dataset)
|
| 42 |
+
if digest(args.dataset) != DATA_SHA256:
|
| 43 |
+
raise ValueError('Dataset does not match the pinned published evaluation')
|
| 44 |
+
sys.path.insert(0, str(checkpoint))
|
| 45 |
+
import pyarrow.parquet as pq
|
| 46 |
+
import torch
|
| 47 |
+
import transformers
|
| 48 |
+
from julia.router.engine import FastEngine
|
| 49 |
+
torch.set_num_threads(args.threads)
|
| 50 |
+
cases = pq.read_table(args.dataset).to_pylist()
|
| 51 |
+
rows, metadata = [], []
|
| 52 |
+
for case in cases:
|
| 53 |
+
state = json.loads(case['state'])
|
| 54 |
+
gold = json.loads(case['gold'])
|
| 55 |
+
for question_id, question in json.loads(case['questions']).items():
|
| 56 |
+
kind = question['type']
|
| 57 |
+
criteria = question.get('criteria')
|
| 58 |
+
if criteria is None and kind == 'noul':
|
| 59 |
+
criteria = {'false': 'false', 'true': 'true'}
|
| 60 |
+
if isinstance(criteria, list):
|
| 61 |
+
criteria = {str(i): value for i, value in enumerate(criteria)}
|
| 62 |
+
if not isinstance(criteria, dict):
|
| 63 |
+
raise ValueError('Missing option descriptions')
|
| 64 |
+
keys = ['false', 'true'] if kind == 'noul' else list(criteria)
|
| 65 |
+
rows.append(dict(state=state, question=question['instructions'],
|
| 66 |
+
type=kind, options=[criteria[key] for key in keys]))
|
| 67 |
+
metadata.append(dict(id=case['id']+':'+question_id, type=kind,
|
| 68 |
+
keys=keys, gold=str(gold[question_id]['label'])))
|
| 69 |
+
counts = collections.Counter(row['type'] for row in rows)
|
| 70 |
+
assert len(cases) == 400 and counts == {'choice': 600, 'score': 800, 'noul': 600}, counts
|
| 71 |
+
engine = FastEngine(checkpoint, device='cpu', transformer_backend='torch',
|
| 72 |
+
strict_encoding=True, max_length=1024, head_length=512,
|
| 73 |
+
batch_size=16, marker_only_head=False)
|
| 74 |
+
result = dict(weights_sha256=WEIGHTS_SHA256, dataset_revision=REVISION,
|
| 75 |
+
dataset_sha256=DATA_SHA256, torch=torch.__version__,
|
| 76 |
+
transformers=transformers.__version__, device='cpu',
|
| 77 |
+
threads=args.threads, marker_only_head=False,
|
| 78 |
+
encoding=dict(max_length=1024, head_length=512, strict=True),
|
| 79 |
+
runtime_sha256={str(p.relative_to(checkpoint)): digest(p)
|
| 80 |
+
for p in sorted((checkpoint/'julia').rglob('*.py'))})
|
| 81 |
+
def evaluate(name, input_rows, labels):
|
| 82 |
+
started = time.monotonic()
|
| 83 |
+
stats = collections.defaultdict(lambda: dict(count=0, correct=0))
|
| 84 |
+
with (args.output / (name+'-predictions.jsonl')).open('w') as stream:
|
| 85 |
+
# One example per forward matches the original benchmark protocol.
|
| 86 |
+
for index, (row, label) in enumerate(zip(input_rows, labels), 1):
|
| 87 |
+
answer = engine.predict([row])[0]
|
| 88 |
+
prediction = label['keys'][answer['index']]
|
| 89 |
+
correct = prediction == label['gold']
|
| 90 |
+
stats[label['type']]['count'] += 1
|
| 91 |
+
stats[label['type']]['correct'] += int(correct)
|
| 92 |
+
stream.write(json.dumps(dict(**label, prediction=prediction, correct=correct,
|
| 93 |
+
options=row['options'], probabilities=answer['probabilities']))+'\n')
|
| 94 |
+
if index % 200 == 0:
|
| 95 |
+
print(json.dumps(dict(mode=name, done=index, total=len(input_rows))), flush=True)
|
| 96 |
+
for item in stats.values():
|
| 97 |
+
item['accuracy'] = item['correct']/item['count']
|
| 98 |
+
summary = dict(by_type=dict(stats), elapsed_seconds=time.monotonic()-started)
|
| 99 |
+
print(json.dumps(dict(mode=name, **summary)), flush=True)
|
| 100 |
+
return summary
|
| 101 |
+
result['original_criteria'] = evaluate('original-criteria', rows, metadata)
|
| 102 |
+
if args.literal_noul_ablation:
|
| 103 |
+
selected = [(dict(row, options=['false', 'true']), label)
|
| 104 |
+
for row, label in zip(rows, metadata) if row['type']=='noul']
|
| 105 |
+
result['literal_noul_ablation'] = evaluate('literal-noul',
|
| 106 |
+
[x[0] for x in selected], [x[1] for x in selected])
|
| 107 |
+
(args.output/'results.json').write_text(json.dumps(result, indent=2)+'\n')
|
| 108 |
+
print(json.dumps(result, indent=2), flush=True)
|
| 109 |
+
|
| 110 |
+
if __name__ == '__main__':
|
| 111 |
+
main()
|
tests/test_typed_api.py
CHANGED
|
@@ -25,3 +25,24 @@ class TypedAPITests(unittest.TestCase):
|
|
| 25 |
self.engine.logits.return_value=[[1.,0.]]
|
| 26 |
q={'q':dict(type='choice',instructions='q',criteria={'a':'A','b':'B'})}
|
| 27 |
self.assertEqual(self.engine.predict('s',q),self.engine.predict(state='s',questions=q))
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 25 |
self.engine.logits.return_value=[[1.,0.]]
|
| 26 |
q={'q':dict(type='choice',instructions='q',criteria={'a':'A','b':'B'})}
|
| 27 |
self.assertEqual(self.engine.predict('s',q),self.engine.predict(state='s',questions=q))
|
| 28 |
+
|
| 29 |
+
def test_noul_preserves_descriptions_in_false_true_order(self):
|
| 30 |
+
self.engine.logits.return_value=[[0.,2.]]
|
| 31 |
+
questions={'review':dict(type='noul',instructions='Needs human review?',
|
| 32 |
+
criteria={'true':'A human should inspect this run.',
|
| 33 |
+
'false':'No human attention is warranted.'})}
|
| 34 |
+
answer=self.engine.predict(state='context',questions=questions)['answers']['review']
|
| 35 |
+
self.assertEqual(self.engine.logits.call_args.args[0][0]['options'],
|
| 36 |
+
['No human attention is warranted.','A human should inspect this run.'])
|
| 37 |
+
self.assertEqual(list(answer['probabilities']),['false','true'])
|
| 38 |
+
self.assertGreater(answer['noul'],.5)
|
| 39 |
+
|
| 40 |
+
def test_noul_rejects_invalid_criteria_before_inference(self):
|
| 41 |
+
self.engine.logits.return_value=[[0.,2.]]
|
| 42 |
+
for criteria in ({}, {'true':'yes'}, {'false':'no','true':'yes','other':'maybe'},
|
| 43 |
+
['no','yes'], {'false':'','true':'yes'}):
|
| 44 |
+
with self.subTest(criteria=criteria):
|
| 45 |
+
questions={'q':dict(type='noul',instructions='q',criteria=criteria)}
|
| 46 |
+
with self.assertRaises(ValueError):
|
| 47 |
+
self.engine.predict(state='context',questions=questions)
|
| 48 |
+
self.engine.logits.assert_not_called()
|