nur-dev commited on
Commit
19c1f65
·
verified ·
1 Parent(s): 1d60d59

Add frozen selected-result prompt comparison and clarify model scope

Browse files
.gitattributes CHANGED
@@ -1,2 +1,4 @@
1
  *.pt filter=lfs diff=lfs merge=lfs -text
2
  *.whl filter=lfs diff=lfs merge=lfs -text
 
 
 
1
  *.pt filter=lfs diff=lfs merge=lfs -text
2
  *.whl filter=lfs diff=lfs merge=lfs -text
3
+ comparison/records.jsonl.gz filter=lfs diff=lfs merge=lfs -text
4
+ comparison/responses.jsonl.gz filter=lfs diff=lfs merge=lfs -text
MANIFEST.sha256 CHANGED
@@ -1,10 +1,15 @@
1
  cf525d607a58b825e4a124bd5511bcfa00d257a22525f0bec81ad4dbbd1d2cfb .gitattributes
2
  9ab279df0ec2f39a82fe6cde3b4bf205a2fdd48b047e3cfdf77df44501b6d8a9 BASE_MODEL.json
3
  01c3d96207e01603bc83082edb665ddef6edde36d96a81caa70efa5fad444177 LICENSE
4
- 200791cbfc25a84e0cef71e4bb7ed74b9dd9c766da5734e02e0d8077ef40b653 README.md
5
  6f0a5f6bbb660fbdbdd2695bdd93b9e5c53d4fcb79cb4d16a3f62a1b7bbbac9e checkpoints/ACTION_HEAD_FINAL.pt
6
  75b595cafc25cb87bcfd0fdac33965ba5efaf80f691c802549693fe127c9cf7b checkpoints/MEMORY_PATH_FINAL.pt
7
  60fa8569aec9d3344abb20b1bf8d3ade7ad11426b4498412f7567cc4bf4add49 checkpoints/P0_M2_CHECKPOINT_FINAL.pt
 
 
 
 
 
8
  f0ca13dc5c0bbfc7852df6f2cee36d4039f7c2197afb8840a3335211862d5e67 configs/strata_native_lm_integration_m1.json
9
  d0ed976b9409ba9123a7b3ca2e16616a4f3e19daa1bf18ca6b042f7918c7eb7c configs/strata_native_lm_p0_m2_v1.json
10
  d97c93bb90ea2c2217820d7e1eafac61cdab16ff28a2189f97b7918b6ffebde2 configs/strata_native_lm_system_v1.json
 
1
  cf525d607a58b825e4a124bd5511bcfa00d257a22525f0bec81ad4dbbd1d2cfb .gitattributes
2
  9ab279df0ec2f39a82fe6cde3b4bf205a2fdd48b047e3cfdf77df44501b6d8a9 BASE_MODEL.json
3
  01c3d96207e01603bc83082edb665ddef6edde36d96a81caa70efa5fad444177 LICENSE
4
+ 89a2b36672b37365afbf6859b42d1265a5bc4fa3c7f3395973df98cd9b05c726 README.md
5
  6f0a5f6bbb660fbdbdd2695bdd93b9e5c53d4fcb79cb4d16a3f62a1b7bbbac9e checkpoints/ACTION_HEAD_FINAL.pt
6
  75b595cafc25cb87bcfd0fdac33965ba5efaf80f691c802549693fe127c9cf7b checkpoints/MEMORY_PATH_FINAL.pt
7
  60fa8569aec9d3344abb20b1bf8d3ade7ad11426b4498412f7567cc4bf4add49 checkpoints/P0_M2_CHECKPOINT_FINAL.pt
8
+ 6a4ec82defbc8196be35749a2d1902726020d5527d07c8259f05317d2bd2ab5e comparison/benchmark.py
9
+ adc0d5141f391fdbb7cd6f96bf3cfffd25be9d0a01baf3aba334dd0610cf5835 comparison/protocol.json
10
+ 987c3936971579b07fbee9a940960d88c9a0e95277bed51068284bca90aeda98 comparison/records.jsonl.gz
11
+ 0d78a9a2ba8b95f869efb0a89538d8a156786d60e879dcf2dc067aa71b5d06c0 comparison/responses.jsonl.gz
12
+ e332d0bb19b1e9c61cce5075356f7214361d77f7a2baad692d654ee4d4a6bf5a comparison/results.json
13
  f0ca13dc5c0bbfc7852df6f2cee36d4039f7c2197afb8840a3335211862d5e67 configs/strata_native_lm_integration_m1.json
14
  d0ed976b9409ba9123a7b3ca2e16616a4f3e19daa1bf18ca6b042f7918c7eb7c configs/strata_native_lm_p0_m2_v1.json
15
  d97c93bb90ea2c2217820d7e1eafac61cdab16ff28a2189f97b7918b6ffebde2 configs/strata_native_lm_system_v1.json
README.md CHANGED
@@ -15,7 +15,7 @@ tags:
15
 
16
  Licensed under [CC BY-NC 4.0](https://creativecommons.org/licenses/by-nc/4.0/) for non-commercial use. Commercial use requires separate permission from the rights holders. The upstream base model and third-party dependencies retain their respective licenses.
17
 
18
- STRATA Native LM couples a frozen Qwen3-4B-Instruct-2507 backbone to a trained 20.44M-parameter memory interface. It generates answers from trusted structured observations supplied by tools, APIs, or applications. Structured writes and read addresses are supplied explicitly.
19
 
20
  ```text
21
  structured observation
@@ -25,7 +25,9 @@ structured observation
25
  -> generated frame + exact value copying
26
  ```
27
 
28
- The model receives the current query and compact memory without previous source tokens, source-history KV-cache entries, or persisted source residuals. Stored value bytes are copied only after the response frame is complete. Integrity checks reject invalid, stale, deleted, or corrupted references.
 
 
29
 
30
  ## Results
31
 
@@ -38,16 +40,42 @@ The model receives the current query and compact memory without previous source
38
  | Persistent-memory horizon | 128 independent 8,192-token windows |
39
  | Factual accuracy | 1.000 |
40
  | Copied-value byte correctness | 1.000 |
41
- | Held-out-schema accuracy | 1.000 |
42
- | 2/3/4-hop join accuracy | 1.000 |
43
  | Field-specific intervention following | 1.000 |
44
  | Integrity verification and execution replay | 1.000 |
45
- | Supporting proof-inclusive retained-state reductions | 86.46% / 86.72% |
46
  | Neural value-state reduction | 95.83% (128 vs. 3,072 bytes/value) |
47
 
48
  The neural-state figure excludes application value bytes and shared address/integrity metadata. The structured-service accounting includes those bytes. The horizon is persistence across independent windows, not a single Transformer context.
49
 
50
- STRATA and the realization-matched token-value arm both achieved 1.000 factual accuracy. On eight NVIDIA L40 GPUs, the evaluated implementation took 43.17 ms/record versus 564.82 ms/record for the source-history arm named `FULL_KV`. Both decoding loops disabled KV reuse; the latter used free generation. These timings describe that workload and implementation.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
51
 
52
  ## Install
53
 
@@ -90,6 +118,7 @@ The JSON response contains `answer`, `frame`, `status`, `payload_handle`, `recei
90
  | `checkpoints/` | Compact-memory reader, memory interface, and copy-action head |
91
  | `configs/` | Inference settings |
92
  | `evaluation.json` | Evaluation metrics, criteria, and source-artifact hashes |
 
93
  | `wheels/` | Runtime package |
94
  | `BASE_MODEL.json` | Required upstream model and immutable revision |
95
  | `load_and_answer.py` | Executable structured-read example |
 
15
 
16
  Licensed under [CC BY-NC 4.0](https://creativecommons.org/licenses/by-nc/4.0/) for non-commercial use. Commercial use requires separate permission from the rights holders. The upstream base model and third-party dependencies retain their respective licenses.
17
 
18
+ STRATA Native LM couples a frozen Qwen3-4B-Instruct-2507 backbone to a trained 20.44M-parameter memory interface. It generates short factual responses from trusted structured observations supplied by tools, APIs, or applications. Structured writes and read addresses are supplied explicitly. The runtime executes joins; the language model verbalizes their selected results. Open-ended response generation was not evaluated.
19
 
20
  ```text
21
  structured observation
 
25
  -> generated frame + exact value copying
26
  ```
27
 
28
+ The reader decodes and latches an address-bound value code. During response-frame generation the actual value state is replaced by a fixed neutral state, and COPY validates the selected record. Its bytes are inserted after frame completion. No previous source tokens, source-history KV-cache entries, persisted source residuals, or actual value tokens enter frame generation. Integrity checks reject invalid, stale, deleted, or corrupted references.
29
+
30
+ The 255 nonzero value codes are reused across records, not global value identifiers. Structured-service adapters assign a field index; persistent identity comes from the full instance/predicate/role address and version. The evaluated copy bank contains one runtime-selected record per request. Arbitrary application bytes remain in that record and are not encoded by the 8-bit code.
31
 
32
  ## Results
33
 
 
40
  | Persistent-memory horizon | 128 independent 8,192-token windows |
41
  | Factual accuracy | 1.000 |
42
  | Copied-value byte correctness | 1.000 |
43
+ | Three schemas excluded from component training | 768 unique records; 11,520 evaluations; 1.000 |
44
+ | Runtime-executed 2/3/4-hop joins | 1.000 |
45
  | Field-specific intervention following | 1.000 |
46
  | Integrity verification and execution replay | 1.000 |
47
+ | Structured-service retained-state reduction, including values and integrity metadata | At least 86.72% |
48
  | Neural value-state reduction | 95.83% (128 vs. 3,072 bytes/value) |
49
 
50
  The neural-state figure excludes application value bytes and shared address/integrity metadata. The structured-service accounting includes those bytes. The horizon is persistence across independent windows, not a single Transformer context.
51
 
52
+ STRATA and the realization-matched token-value arm both achieved 1.000 factual accuracy. The original uncached source-history timing remains in `evaluation.json` as an implementation diagnostic, not a speed claim against optimized cached serving. The three training-excluded schemas share the trained memory interface, operation inventory, and templated query format.
53
+
54
+ ## Selected-Result Prompt Comparison
55
+
56
+ The same 4,096 unique selected current records were passed to the frozen base model as JSON in the current prompt, with within-request KV caching enabled. Both interfaces were evaluated twice on eight L40 GPUs without cross-request prefix caching.
57
+
58
+ | Metric | STRATA | Prompt serialization |
59
+ | --- | ---: | ---: |
60
+ | Stored UTF-8 value reproduced exactly once | 4,096/4,096 | 4,051/4,096 |
61
+ | Ordinary bindings and resolved joins | 3,328/3,328 | 3,328/3,328 |
62
+ | Mean input tokens/request | 64.52 | 136.29 |
63
+ | Batch-amortized ms/record | 37.81 | 117.51 |
64
+ | Generation-cap cases | 0 | 8 |
65
+
66
+ All 45 prompt failures were in the hard-value holdout: 36 explicit refusals, eight outputs capped at 256 generated tokens, and one other value mismatch. Prompt generation freely spells values; STRATA copies their bytes. These are measured interface tradeoffs, not an optimized-serving or retrieval-quality ranking. Both passes produced identical responses.
67
+
68
+ The `comparison/` directory contains the selected records, sealed responses, protocol, results, and executable comparison. Recompute published scores without a GPU:
69
+
70
+ ```bash
71
+ python3 comparison/benchmark.py --score-published
72
+ ```
73
+
74
+ After installation and base-model download, reproduce both interfaces on eight GPUs:
75
+
76
+ ```bash
77
+ python3 comparison/benchmark.py --base-model "$STRATA_BASE_MODEL" --output comparison-run
78
+ ```
79
 
80
  ## Install
81
 
 
118
  | `checkpoints/` | Compact-memory reader, memory interface, and copy-action head |
119
  | `configs/` | Inference settings |
120
  | `evaluation.json` | Evaluation metrics, criteria, and source-artifact hashes |
121
+ | `comparison/` | Selected-result baseline data, outputs, protocol, results, and runner |
122
  | `wheels/` | Runtime package |
123
  | `BASE_MODEL.json` | Required upstream model and immutable revision |
124
  | `load_and_answer.py` | Executable structured-read example |
comparison/benchmark.py ADDED
@@ -0,0 +1,167 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """Reproduce selected-result prompting versus the frozen STRATA response interface."""
3
+
4
+ import argparse
5
+ from collections import defaultdict
6
+ import gzip
7
+ import json
8
+ import os
9
+ from pathlib import Path
10
+ import subprocess
11
+ import sys
12
+ import time
13
+
14
+ ROOT = Path(__file__).resolve().parents[1]
15
+ DATA = Path(__file__).resolve().parent
16
+ sys.path.insert(0, str(ROOT))
17
+
18
+ def selected_prompt(record, config):
19
+ selected = {key: record[key] for key in config['selected_fields']}
20
+ return config['prompt_prefix'] + json.dumps(selected, ensure_ascii=False, separators=(',', ':')) + config['prompt_suffix'] + record['query']
21
+
22
+ def exact_value_once(text, value):
23
+ return text.encode('utf-8').count(value.encode('utf-8')) == 1
24
+
25
+ def make_infer(torch, model, head, tokenizer, rows, records, config, system,
26
+ state_table, banks, versions, device):
27
+ from strata.eval.native_lm_benchmark import _prompt_ids
28
+ from strata.eval.native_lm_frame_separated_copy import frame_separated_generate
29
+ backbone = model.backbone
30
+
31
+ @torch.inference_mode()
32
+ def infer(arm, indices):
33
+ group = [rows[i] for i in indices]
34
+ torch.cuda.synchronize()
35
+ start = time.perf_counter()
36
+ if arm == 'PROMPT_SERIALIZATION':
37
+ sequences = [_prompt_ids(tokenizer, selected_prompt(records[i], config)) for i in indices]
38
+ width = max(map(len, sequences))
39
+ ids = torch.tensor([[tokenizer.pad_token_id] * (width - len(seq)) + seq for seq in sequences], device=device)
40
+ mask = torch.tensor([[0] * (width - len(seq)) + [1] * len(seq) for seq in sequences], device=device)
41
+ outputs = backbone.generate(input_ids=ids, attention_mask=mask, do_sample=False, use_cache=True, max_new_tokens=config['prompt_generation_max_new_tokens'], pad_token_id=tokenizer.pad_token_id, eos_token_id=tokenizer.eos_token_id)
42
+ texts = tokenizer.batch_decode(outputs[:, width:], skip_special_tokens=True)
43
+ generated_lengths = [len(tokens) - list(tokens).count(tokenizer.pad_token_id) for tokens in outputs[:, width:].cpu().tolist()]
44
+ truncated = [tokenizer.eos_token_id not in tokens for tokens in outputs[:, width:].cpu().tolist()]
45
+ else:
46
+ outputs, _ = frame_separated_generate(model, head, tokenizer, group, state_table, [banks[i] for i in indices], [0] * len(group), batch_size=len(group), max_actions=config['strata_max_actions'], frame_handle=system['frame']['canonical_frame_handle'], frame_surrogate=system['frame']['canonical_frame_surrogate'], terminator=system['frame']['structural_terminator'], current_versions=[versions[i] for i in indices])
47
+ texts = [output.text for output in outputs]
48
+ generated_lengths = [sum((action.startswith('GEN(') for action in output.actions)) for output in outputs]
49
+ truncated = [output.status == 'MAX_ACTIONS' for output in outputs]
50
+ torch.cuda.synchronize()
51
+ elapsed = time.perf_counter() - start
52
+ lengths = [len(_prompt_ids(tokenizer, selected_prompt(records[i], config) if arm == 'PROMPT_SERIALIZATION' else records[i]['query'])) for i in indices]
53
+ return (texts, lengths, generated_lengths, truncated, elapsed)
54
+ return infer
55
+
56
+ def read_records():
57
+ with gzip.open(DATA/'records.jsonl.gz', 'rt', encoding='utf-8') as handle:
58
+ return [json.loads(line) for line in handle]
59
+
60
+
61
+ def score(paths):
62
+ gold = {row['id']:row for row in read_records()}
63
+ scores = defaultdict(lambda: {'correct':0,'total':0,'input_tokens':0,'truncated':0})
64
+ timings = {}
65
+ for path in paths:
66
+ opener = gzip.open if str(path).endswith('.gz') else open
67
+ with opener(path, 'rt', encoding='utf-8') as handle:
68
+ for line in handle:
69
+ item = json.loads(line)
70
+ row = gold[item['id']]
71
+ value = scores[item['arm']]
72
+ value['correct'] += int(exact_value_once(item['text'],row['value']))
73
+ value['total'] += 1
74
+ value['input_tokens'] += item['input_tokens']
75
+ value['truncated'] += item['truncated']
76
+ timings[(item['arm'],item['rank'],item['repeat'],item['batch_offset'])] = item['batch_seconds']
77
+ for arm,value in scores.items():
78
+ value['accuracy'] = value['correct']/value['total']
79
+ value['mean_input_tokens'] = value['input_tokens']/value['total']
80
+ value['amortized_ms_per_record'] = sum(t for key,t in timings.items() if key[0]==arm)*1000/value['total']
81
+ return dict(scores)
82
+
83
+
84
+ def worker(args, config):
85
+ import torch
86
+ from load_and_answer import load_model
87
+ from strata.data.native_lm_integration import NativeLMExample, address_codes, answer_text
88
+ from strata.modeling.exact_payload_realizer import PayloadAuthority
89
+ from strata.training.native_lm_integration import compact_state_table
90
+ torch.set_num_threads(2)
91
+ torch.cuda.set_device(args.rank)
92
+ torch.manual_seed(config['seed'])
93
+ device = torch.device('cuda',args.rank)
94
+ system,tokenizer,model,head,codec = load_model(ROOT,args.base_model,device)
95
+ state_table = compact_state_table(codec)
96
+ records = [row for row in read_records() if row['rank']==args.rank]
97
+ if args.limit is not None:
98
+ records = records[:args.limit]
99
+ rows = [NativeLMExample(example_id=r['id'],split='system-v1',schema=r['schema'],field=r['field'],
100
+ event=r['event'],predicate=r['predicate'],role=r['role'],value_type=r['value_type'],
101
+ payload_handle=r['payload_handle'],value=r['value'],
102
+ address_codes=address_codes(r['event'],r['predicate'],r['role']),query=r['query'],
103
+ full_history_query='',answer=answer_text(r['field'],r['value']),operation=r['operation'],age_windows=0)
104
+ for r in records]
105
+ versions = [r['event_version'] for r in records]
106
+ banks = [[PayloadAuthority.issue(event=r.event,predicate=r.predicate,role=r.role,
107
+ handle=r.payload_handle,payload=r.value,version=v)] for r,v in zip(rows,versions)]
108
+ infer = make_infer(torch,model,head,tokenizer,rows,records,config,system,state_table,banks,versions,device)
109
+ warm = list(range(min(config['batch_size'],len(rows))))
110
+ for arm in ['PROMPT_SERIALIZATION','STRATA_NATIVE_V1']:
111
+ infer(arm,warm)
112
+ path = args.output/f'rank-{args.rank}.jsonl'
113
+ with path.open('x',encoding='utf-8') as handle:
114
+ for repeat in range(config['repeats']):
115
+ for offset in range(0,len(rows),config['batch_size']):
116
+ indices = list(range(offset,min(offset+config['batch_size'],len(rows))))
117
+ arms = ['PROMPT_SERIALIZATION','STRATA_NATIVE_V1']
118
+ if (args.rank+repeat+offset//config['batch_size'])%2:
119
+ arms.reverse()
120
+ for arm in arms:
121
+ texts,lengths,generated,truncated,elapsed = infer(arm,indices)
122
+ for index,text,length,n,cutoff in zip(indices,texts,lengths,generated,truncated):
123
+ handle.write(json.dumps({'id':records[index]['id'],'rank':args.rank,'repeat':repeat,
124
+ 'arm':arm,'text':text,'input_tokens':length,'generated_tokens':n,'truncated':cutoff,
125
+ 'batch_offset':offset,'batch_seconds':elapsed,'batch_size':len(indices)},ensure_ascii=False)+'\n')
126
+
127
+
128
+ def main():
129
+ parser = argparse.ArgumentParser()
130
+ parser.add_argument('--score-published',action='store_true')
131
+ parser.add_argument('--base-model',default=os.environ.get('STRATA_BASE_MODEL'))
132
+ parser.add_argument('--output',type=Path)
133
+ parser.add_argument('--rank',type=int,choices=range(8))
134
+ parser.add_argument('--limit',type=int)
135
+ args = parser.parse_args()
136
+ if args.score_published:
137
+ print(json.dumps(score([DATA/'responses.jsonl.gz']),indent=2))
138
+ return
139
+ if not args.base_model or not args.output:
140
+ parser.error('--base-model and --output are required for inference')
141
+ config = json.loads((DATA/'protocol.json').read_text())
142
+ if args.rank is not None:
143
+ args.output.mkdir(parents=True,exist_ok=True)
144
+ worker(args,config)
145
+ return
146
+ args.output.mkdir(parents=True,exist_ok=False)
147
+ processes = []
148
+ for rank in range(8):
149
+ log = (args.output/f'rank-{rank}.log').open('x')
150
+ command = [sys.executable,__file__,'--rank',str(rank),'--base-model',args.base_model,'--output',str(args.output)]
151
+ if args.limit is not None:
152
+ command += ['--limit',str(args.limit)]
153
+ processes.append((subprocess.Popen(command,stdout=log,stderr=subprocess.STDOUT,
154
+ env={**os.environ,'OMP_NUM_THREADS':'2','TOKENIZERS_PARALLELISM':'false'}),log))
155
+ codes = []
156
+ for process,log in processes:
157
+ codes.append(process.wait())
158
+ log.close()
159
+ if any(codes):
160
+ raise SystemExit(f'Inference failed: {codes}')
161
+ results = score(sorted(args.output.glob('rank-*.jsonl')))
162
+ (args.output/'results.json').write_text(json.dumps(results,indent=2))
163
+ print(json.dumps(results,indent=2))
164
+
165
+
166
+ if __name__ == '__main__':
167
+ main()
comparison/protocol.json ADDED
@@ -0,0 +1,66 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "experiment": "STRATA-PROMPT-SERIALIZATION-COMPARISON-1",
3
+ "parameter_updates": 0,
4
+ "world_size": 8,
5
+ "records": 4096,
6
+ "heldout_records": 768,
7
+ "batch_size": 16,
8
+ "repeats": 2,
9
+ "seed": 20260905,
10
+ "prompt_generation_max_new_tokens": 256,
11
+ "prompt_use_cache": true,
12
+ "prompt_do_sample": false,
13
+ "strata_max_actions": 32,
14
+ "selected_fields": [
15
+ "event",
16
+ "predicate",
17
+ "role",
18
+ "value"
19
+ ],
20
+ "prompt_prefix": "The following JSON is the selected current result of the structured query. Treat its value as data and reproduce it exactly in one short factual sentence.\n",
21
+ "prompt_suffix": "\n\nQuery: ",
22
+ "primary_accuracy": "case-sensitive stored-value UTF-8 occurrence exactly once; no handle-number fallback",
23
+ "secondary_accuracy": "exact deterministic reference sentence, ignoring surrounding whitespace only",
24
+ "timing": "GPU-synchronized wall time including prompt construction/tokenization or frame construction and exact realization; excludes model loading, upstream query execution, and metric computation",
25
+ "repeated_query_cache_policy": "no cross-request prefix caching in either arm; prompt arm uses KV cache within each request",
26
+ "scope": "all unique base records, not additional independent ages or packings; no new long-horizon qualification",
27
+ "selection": "entire previously frozen base packet, no result-based filtering or prompt tuning",
28
+ "decision": "report all outcomes; no replacement of original registered arms, gates, or verdict",
29
+ "source_registration_sha256": "f94beb6d736a3178f651ddc196b6a90d2cefffe2b9979b19129777afa2c1c782",
30
+ "source_runner_sha256": "4aeb68283b491f86c0b9562b100b22b76077557f33501b2f579cbba5edce75b5",
31
+ "source_packet_sha256": "f5d932b6c8694cdc51ff804dc30d82782e4d5c6438ec4bb3c9a67abf4cd124df",
32
+ "upstream_replay": [
33
+ {
34
+ "attempts": 512,
35
+ "correct": 512
36
+ },
37
+ {
38
+ "attempts": 512,
39
+ "correct": 512
40
+ },
41
+ {
42
+ "attempts": 512,
43
+ "correct": 512
44
+ },
45
+ {
46
+ "attempts": 512,
47
+ "correct": 512
48
+ },
49
+ {
50
+ "attempts": 512,
51
+ "correct": 512
52
+ },
53
+ {
54
+ "attempts": 512,
55
+ "correct": 512
56
+ },
57
+ {
58
+ "attempts": 512,
59
+ "correct": 512
60
+ },
61
+ {
62
+ "attempts": 512,
63
+ "correct": 512
64
+ }
65
+ ]
66
+ }
comparison/records.jsonl.gz ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:987c3936971579b07fbee9a940960d88c9a0e95277bed51068284bca90aeda98
3
+ size 127689
comparison/responses.jsonl.gz ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0d78a9a2ba8b95f869efb0a89538d8a156786d60e879dcf2dc067aa71b5d06c0
3
+ size 294516
comparison/results.json ADDED
@@ -0,0 +1,406 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "experiment": "STRATA-PROMPT-SERIALIZATION-COMPARISON-1",
3
+ "unique_records": 4096,
4
+ "repeats": 2,
5
+ "registration_sha256": "f94beb6d736a3178f651ddc196b6a90d2cefffe2b9979b19129777afa2c1c782",
6
+ "arms": {
7
+ "PROMPT_SERIALIZATION": {
8
+ "correct": 8102,
9
+ "total": 8192,
10
+ "reference_exact": 0,
11
+ "input_tokens": 1116502,
12
+ "truncated": 16,
13
+ "accuracy": 0.989013671875,
14
+ "mean_input_tokens": 136.291748046875,
15
+ "seconds": 962.6804165355861,
16
+ "amortized_ms_per_record": 117.51469928412916,
17
+ "median_batch_amortized_ms": 62.35561228822917,
18
+ "p95_batch_amortized_ms": 417.074111988768,
19
+ "repeated_response_identity": 1.0
20
+ },
21
+ "STRATA_NATIVE_V1": {
22
+ "correct": 8192,
23
+ "total": 8192,
24
+ "reference_exact": 8110,
25
+ "input_tokens": 528536,
26
+ "truncated": 0,
27
+ "accuracy": 1.0,
28
+ "mean_input_tokens": 64.5185546875,
29
+ "seconds": 309.70214548148215,
30
+ "amortized_ms_per_record": 37.80543768084499,
31
+ "median_batch_amortized_ms": 37.54055639728904,
32
+ "p95_batch_amortized_ms": 42.57394268643111,
33
+ "repeated_response_identity": 1.0
34
+ }
35
+ },
36
+ "slices": [
37
+ {
38
+ "arm": "PROMPT_SERIALIZATION",
39
+ "axis": "graph_depth",
40
+ "value": "1",
41
+ "correct": 6566,
42
+ "total": 6656,
43
+ "accuracy": 0.9864783653846154
44
+ },
45
+ {
46
+ "arm": "PROMPT_SERIALIZATION",
47
+ "axis": "graph_depth",
48
+ "value": "2",
49
+ "correct": 512,
50
+ "total": 512,
51
+ "accuracy": 1.0
52
+ },
53
+ {
54
+ "arm": "PROMPT_SERIALIZATION",
55
+ "axis": "graph_depth",
56
+ "value": "3",
57
+ "correct": 512,
58
+ "total": 512,
59
+ "accuracy": 1.0
60
+ },
61
+ {
62
+ "arm": "PROMPT_SERIALIZATION",
63
+ "axis": "graph_depth",
64
+ "value": "4",
65
+ "correct": 512,
66
+ "total": 512,
67
+ "accuracy": 1.0
68
+ },
69
+ {
70
+ "arm": "PROMPT_SERIALIZATION",
71
+ "axis": "kind",
72
+ "value": "binding",
73
+ "correct": 5120,
74
+ "total": 5120,
75
+ "accuracy": 1.0
76
+ },
77
+ {
78
+ "arm": "PROMPT_SERIALIZATION",
79
+ "axis": "kind",
80
+ "value": "heldout",
81
+ "correct": 1446,
82
+ "total": 1536,
83
+ "accuracy": 0.94140625
84
+ },
85
+ {
86
+ "arm": "PROMPT_SERIALIZATION",
87
+ "axis": "kind",
88
+ "value": "join",
89
+ "correct": 1536,
90
+ "total": 1536,
91
+ "accuracy": 1.0
92
+ },
93
+ {
94
+ "arm": "PROMPT_SERIALIZATION",
95
+ "axis": "schema",
96
+ "value": "approval",
97
+ "correct": 492,
98
+ "total": 494,
99
+ "accuracy": 0.9959514170040485
100
+ },
101
+ {
102
+ "arm": "PROMPT_SERIALIZATION",
103
+ "axis": "schema",
104
+ "value": "calendar",
105
+ "correct": 1024,
106
+ "total": 1024,
107
+ "accuracy": 1.0
108
+ },
109
+ {
110
+ "arm": "PROMPT_SERIALIZATION",
111
+ "axis": "schema",
112
+ "value": "crm",
113
+ "correct": 1536,
114
+ "total": 1536,
115
+ "accuracy": 1.0
116
+ },
117
+ {
118
+ "arm": "PROMPT_SERIALIZATION",
119
+ "axis": "schema",
120
+ "value": "deployment",
121
+ "correct": 516,
122
+ "total": 524,
123
+ "accuracy": 0.9847328244274809
124
+ },
125
+ {
126
+ "arm": "PROMPT_SERIALIZATION",
127
+ "axis": "schema",
128
+ "value": "incident",
129
+ "correct": 1536,
130
+ "total": 1536,
131
+ "accuracy": 1.0
132
+ },
133
+ {
134
+ "arm": "PROMPT_SERIALIZATION",
135
+ "axis": "schema",
136
+ "value": "inventory",
137
+ "correct": 1024,
138
+ "total": 1024,
139
+ "accuracy": 1.0
140
+ },
141
+ {
142
+ "arm": "PROMPT_SERIALIZATION",
143
+ "axis": "schema",
144
+ "value": "project",
145
+ "correct": 1536,
146
+ "total": 1536,
147
+ "accuracy": 1.0
148
+ },
149
+ {
150
+ "arm": "PROMPT_SERIALIZATION",
151
+ "axis": "schema",
152
+ "value": "shipment",
153
+ "correct": 438,
154
+ "total": 518,
155
+ "accuracy": 0.8455598455598455
156
+ },
157
+ {
158
+ "arm": "PROMPT_SERIALIZATION",
159
+ "axis": "schema_partition",
160
+ "value": "heldout",
161
+ "correct": 1446,
162
+ "total": 1536,
163
+ "accuracy": 0.94140625
164
+ },
165
+ {
166
+ "arm": "PROMPT_SERIALIZATION",
167
+ "axis": "schema_partition",
168
+ "value": "registered",
169
+ "correct": 6656,
170
+ "total": 6656,
171
+ "accuracy": 1.0
172
+ },
173
+ {
174
+ "arm": "PROMPT_SERIALIZATION",
175
+ "axis": "value_type",
176
+ "value": "datetime",
177
+ "correct": 654,
178
+ "total": 654,
179
+ "accuracy": 1.0
180
+ },
181
+ {
182
+ "arm": "PROMPT_SERIALIZATION",
183
+ "axis": "value_type",
184
+ "value": "entity",
185
+ "correct": 3434,
186
+ "total": 3436,
187
+ "accuracy": 0.9994179278230501
188
+ },
189
+ {
190
+ "arm": "PROMPT_SERIALIZATION",
191
+ "axis": "value_type",
192
+ "value": "enum",
193
+ "correct": 2008,
194
+ "total": 2008,
195
+ "accuracy": 1.0
196
+ },
197
+ {
198
+ "arm": "PROMPT_SERIALIZATION",
199
+ "axis": "value_type",
200
+ "value": "integer",
201
+ "correct": 612,
202
+ "total": 692,
203
+ "accuracy": 0.884393063583815
204
+ },
205
+ {
206
+ "arm": "PROMPT_SERIALIZATION",
207
+ "axis": "value_type",
208
+ "value": "location",
209
+ "correct": 956,
210
+ "total": 964,
211
+ "accuracy": 0.991701244813278
212
+ },
213
+ {
214
+ "arm": "PROMPT_SERIALIZATION",
215
+ "axis": "value_type",
216
+ "value": "string",
217
+ "correct": 438,
218
+ "total": 438,
219
+ "accuracy": 1.0
220
+ },
221
+ {
222
+ "arm": "STRATA_NATIVE_V1",
223
+ "axis": "graph_depth",
224
+ "value": "1",
225
+ "correct": 6656,
226
+ "total": 6656,
227
+ "accuracy": 1.0
228
+ },
229
+ {
230
+ "arm": "STRATA_NATIVE_V1",
231
+ "axis": "graph_depth",
232
+ "value": "2",
233
+ "correct": 512,
234
+ "total": 512,
235
+ "accuracy": 1.0
236
+ },
237
+ {
238
+ "arm": "STRATA_NATIVE_V1",
239
+ "axis": "graph_depth",
240
+ "value": "3",
241
+ "correct": 512,
242
+ "total": 512,
243
+ "accuracy": 1.0
244
+ },
245
+ {
246
+ "arm": "STRATA_NATIVE_V1",
247
+ "axis": "graph_depth",
248
+ "value": "4",
249
+ "correct": 512,
250
+ "total": 512,
251
+ "accuracy": 1.0
252
+ },
253
+ {
254
+ "arm": "STRATA_NATIVE_V1",
255
+ "axis": "kind",
256
+ "value": "binding",
257
+ "correct": 5120,
258
+ "total": 5120,
259
+ "accuracy": 1.0
260
+ },
261
+ {
262
+ "arm": "STRATA_NATIVE_V1",
263
+ "axis": "kind",
264
+ "value": "heldout",
265
+ "correct": 1536,
266
+ "total": 1536,
267
+ "accuracy": 1.0
268
+ },
269
+ {
270
+ "arm": "STRATA_NATIVE_V1",
271
+ "axis": "kind",
272
+ "value": "join",
273
+ "correct": 1536,
274
+ "total": 1536,
275
+ "accuracy": 1.0
276
+ },
277
+ {
278
+ "arm": "STRATA_NATIVE_V1",
279
+ "axis": "schema",
280
+ "value": "approval",
281
+ "correct": 494,
282
+ "total": 494,
283
+ "accuracy": 1.0
284
+ },
285
+ {
286
+ "arm": "STRATA_NATIVE_V1",
287
+ "axis": "schema",
288
+ "value": "calendar",
289
+ "correct": 1024,
290
+ "total": 1024,
291
+ "accuracy": 1.0
292
+ },
293
+ {
294
+ "arm": "STRATA_NATIVE_V1",
295
+ "axis": "schema",
296
+ "value": "crm",
297
+ "correct": 1536,
298
+ "total": 1536,
299
+ "accuracy": 1.0
300
+ },
301
+ {
302
+ "arm": "STRATA_NATIVE_V1",
303
+ "axis": "schema",
304
+ "value": "deployment",
305
+ "correct": 524,
306
+ "total": 524,
307
+ "accuracy": 1.0
308
+ },
309
+ {
310
+ "arm": "STRATA_NATIVE_V1",
311
+ "axis": "schema",
312
+ "value": "incident",
313
+ "correct": 1536,
314
+ "total": 1536,
315
+ "accuracy": 1.0
316
+ },
317
+ {
318
+ "arm": "STRATA_NATIVE_V1",
319
+ "axis": "schema",
320
+ "value": "inventory",
321
+ "correct": 1024,
322
+ "total": 1024,
323
+ "accuracy": 1.0
324
+ },
325
+ {
326
+ "arm": "STRATA_NATIVE_V1",
327
+ "axis": "schema",
328
+ "value": "project",
329
+ "correct": 1536,
330
+ "total": 1536,
331
+ "accuracy": 1.0
332
+ },
333
+ {
334
+ "arm": "STRATA_NATIVE_V1",
335
+ "axis": "schema",
336
+ "value": "shipment",
337
+ "correct": 518,
338
+ "total": 518,
339
+ "accuracy": 1.0
340
+ },
341
+ {
342
+ "arm": "STRATA_NATIVE_V1",
343
+ "axis": "schema_partition",
344
+ "value": "heldout",
345
+ "correct": 1536,
346
+ "total": 1536,
347
+ "accuracy": 1.0
348
+ },
349
+ {
350
+ "arm": "STRATA_NATIVE_V1",
351
+ "axis": "schema_partition",
352
+ "value": "registered",
353
+ "correct": 6656,
354
+ "total": 6656,
355
+ "accuracy": 1.0
356
+ },
357
+ {
358
+ "arm": "STRATA_NATIVE_V1",
359
+ "axis": "value_type",
360
+ "value": "datetime",
361
+ "correct": 654,
362
+ "total": 654,
363
+ "accuracy": 1.0
364
+ },
365
+ {
366
+ "arm": "STRATA_NATIVE_V1",
367
+ "axis": "value_type",
368
+ "value": "entity",
369
+ "correct": 3436,
370
+ "total": 3436,
371
+ "accuracy": 1.0
372
+ },
373
+ {
374
+ "arm": "STRATA_NATIVE_V1",
375
+ "axis": "value_type",
376
+ "value": "enum",
377
+ "correct": 2008,
378
+ "total": 2008,
379
+ "accuracy": 1.0
380
+ },
381
+ {
382
+ "arm": "STRATA_NATIVE_V1",
383
+ "axis": "value_type",
384
+ "value": "integer",
385
+ "correct": 692,
386
+ "total": 692,
387
+ "accuracy": 1.0
388
+ },
389
+ {
390
+ "arm": "STRATA_NATIVE_V1",
391
+ "axis": "value_type",
392
+ "value": "location",
393
+ "correct": 964,
394
+ "total": 964,
395
+ "accuracy": 1.0
396
+ },
397
+ {
398
+ "arm": "STRATA_NATIVE_V1",
399
+ "axis": "value_type",
400
+ "value": "string",
401
+ "correct": 438,
402
+ "total": 438,
403
+ "accuracy": 1.0
404
+ }
405
+ ]
406
+ }