Xunzhuo commited on
Commit
ee8e74d
·
verified ·
1 Parent(s): e885639

Accelerate mixed SystemOne decisions with verified typed scheduling

Browse files
.gitattributes CHANGED
@@ -10,3 +10,4 @@ assets/candidate-readout.png filter=lfs diff=lfs merge=lfs -text
10
  assets/decision-lex-header.png filter=lfs diff=lfs merge=lfs -text
11
  assets/residual-layers.pdf filter=lfs diff=lfs merge=lfs -text
12
  assets/residual-layers.png filter=lfs diff=lfs merge=lfs -text
 
 
10
  assets/decision-lex-header.png filter=lfs diff=lfs merge=lfs -text
11
  assets/residual-layers.pdf filter=lfs diff=lfs merge=lfs -text
12
  assets/residual-layers.png filter=lfs diff=lfs merge=lfs -text
13
+ assets/mixed-question-scaling.png filter=lfs diff=lfs merge=lfs -text
MIXED_QUESTION_SCALING.json ADDED
@@ -0,0 +1,1056 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model": "Lex",
3
+ "native_manifest_sha256": "f288d873999832a3f37c6a7c4268c2ab309691e621794dbf7acab891acbbb7e6",
4
+ "weights_unchanged": true,
5
+ "runtime": "SystemOne default stable grouping by decision type, physicalB8, restored caller order; native predictor/optionalauto unchanged",
6
+ "final_package_measurement": {
7
+ "status": "COMPLETE_ACTUAL_RELEASE_PACKAGE_LATENCY",
8
+ "model": "Lex",
9
+ "timing": {
10
+ "1": {
11
+ "previous": {
12
+ "p50_ms": 13.217147497925907,
13
+ "p95_ms": 13.322706916369498,
14
+ "samples_ms": [
15
+ 13.282976928167045,
16
+ 13.237406965345144,
17
+ 13.211747980676591,
18
+ 13.152898056432605,
19
+ 13.230256969109178,
20
+ 13.178368099033833,
21
+ 13.25502700638026,
22
+ 13.151048100553453,
23
+ 13.099667965434492,
24
+ 13.143047923222184,
25
+ 13.119807932525873,
26
+ 13.172856997698545,
27
+ 13.215948012657464,
28
+ 13.322706916369498,
29
+ 13.251267024315894,
30
+ 13.226117007434368,
31
+ 13.22932809125632,
32
+ 13.1907369941473,
33
+ 13.43417598400265,
34
+ 13.178187073208392,
35
+ 13.225198024883866,
36
+ 13.218346983194351,
37
+ 13.236007071100175,
38
+ 13.224758091382682,
39
+ 13.257008045911789,
40
+ 13.164428062736988,
41
+ 13.149028061889112,
42
+ 13.277907040901482,
43
+ 13.13608803320676,
44
+ 13.126858044415712
45
+ ],
46
+ "forwards": 1
47
+ },
48
+ "current": {
49
+ "p50_ms": 13.203607522882521,
50
+ "p95_ms": 13.400446041487157,
51
+ "samples_ms": [
52
+ 13.205877039581537,
53
+ 13.137917965650558,
54
+ 13.263097032904625,
55
+ 13.192676939070225,
56
+ 13.159338035620749,
57
+ 13.066528015770018,
58
+ 13.224957045167685,
59
+ 13.128897990100086,
60
+ 13.172497972846031,
61
+ 13.362155994400382,
62
+ 13.400446041487157,
63
+ 13.369087013415992,
64
+ 13.508685980923474,
65
+ 13.245807029306889,
66
+ 13.190687051974237,
67
+ 13.201338006183505,
68
+ 13.176057022064924,
69
+ 13.362827012315392,
70
+ 13.350806897506118,
71
+ 13.200717978179455,
72
+ 13.09696794487536,
73
+ 13.282197061926126,
74
+ 13.206628034822643,
75
+ 13.207547017373145,
76
+ 13.197846943512559,
77
+ 13.18042806815356,
78
+ 13.190128025598824,
79
+ 13.355995994061232,
80
+ 13.315456919372082,
81
+ 13.173048035241663
82
+ ],
83
+ "forwards": 1
84
+ },
85
+ "p50_ratio": 0.9989755750970085
86
+ },
87
+ "8": {
88
+ "previous": {
89
+ "p50_ms": 28.06903002783656,
90
+ "p95_ms": 28.322704951278865,
91
+ "samples_ms": [
92
+ 28.004087042063475,
93
+ 28.014626004733145,
94
+ 28.206884977407753,
95
+ 28.091166052035987,
96
+ 28.09197606984526,
97
+ 28.103135991841555,
98
+ 27.947345981374383,
99
+ 27.99994603265077,
100
+ 28.176164953038096,
101
+ 28.125606011599302,
102
+ 27.949147042818367,
103
+ 28.065624996088445,
104
+ 28.087166021578014,
105
+ 28.040915960446,
106
+ 28.449213947169483,
107
+ 28.018205892294645,
108
+ 27.994325035251677,
109
+ 28.044176986441016,
110
+ 28.243384906090796,
111
+ 28.170255944132805,
112
+ 28.322704951278865,
113
+ 27.86880696658045,
114
+ 28.18244497757405,
115
+ 28.012576047331095,
116
+ 28.072435059584677,
117
+ 27.985455002635717,
118
+ 28.23424502275884,
119
+ 27.97982608899474,
120
+ 28.03374594077468,
121
+ 28.174914070405066
122
+ ],
123
+ "forwards": 1
124
+ },
125
+ "current": {
126
+ "p50_ms": 28.02164654713124,
127
+ "p95_ms": 28.305554995313287,
128
+ "samples_ms": [
129
+ 27.882487047463655,
130
+ 27.982397004961967,
131
+ 28.524363064207137,
132
+ 28.012415976263583,
133
+ 28.066674945876002,
134
+ 28.09819602407515,
135
+ 28.139916015788913,
136
+ 28.162325033918023,
137
+ 28.012767084874213,
138
+ 27.914006961509585,
139
+ 27.955345925875008,
140
+ 28.200226020999253,
141
+ 28.069816064089537,
142
+ 27.95354591216892,
143
+ 28.061965946108103,
144
+ 27.974725933745503,
145
+ 28.034597053192556,
146
+ 28.27983594033867,
147
+ 27.997806086204946,
148
+ 28.006776934489608,
149
+ 28.305554995313287,
150
+ 28.011246002279222,
151
+ 27.974336058832705,
152
+ 28.092985041439533,
153
+ 28.070215019397438,
154
+ 27.96054701320827,
155
+ 28.062126017175615,
156
+ 27.959375991486013,
157
+ 27.941217995248735,
158
+ 28.030526009388268
159
+ ],
160
+ "forwards": 1
161
+ },
162
+ "p50_ratio": 0.9983118946162967
163
+ },
164
+ "32": {
165
+ "previous": {
166
+ "p50_ms": 93.75805454328656,
167
+ "p95_ms": 94.06405396293849,
168
+ "samples_ms": [
169
+ 93.82304409518838,
170
+ 93.8560439972207,
171
+ 93.62955600954592,
172
+ 93.10267795808613,
173
+ 93.90996396541595,
174
+ 93.64896605256945,
175
+ 93.62079505808651,
176
+ 93.67714496329427,
177
+ 93.83229492232203,
178
+ 93.16505899187177,
179
+ 93.54266698937863,
180
+ 93.6963950516656,
181
+ 93.74855505302548,
182
+ 93.54403603356332,
183
+ 93.50503596942872,
184
+ 93.71798590291291,
185
+ 93.91635493375361,
186
+ 93.67461502552032,
187
+ 93.6862260568887,
188
+ 93.78432505764067,
189
+ 94.01922405231744,
190
+ 94.0071539953351,
191
+ 94.18116300366819,
192
+ 94.06405396293849,
193
+ 93.63992500584573,
194
+ 93.92943407874554,
195
+ 93.9940030220896,
196
+ 93.76755403354764,
197
+ 93.8123851083219,
198
+ 93.78303401172161
199
+ ],
200
+ "forwards": 4
201
+ },
202
+ "current": {
203
+ "p50_ms": 53.42581600416452,
204
+ "p95_ms": 53.77058498561382,
205
+ "samples_ms": [
206
+ 53.39325708337128,
207
+ 53.337397053837776,
208
+ 53.106967941857874,
209
+ 53.07436909060925,
210
+ 53.77058498561382,
211
+ 53.8342340150848,
212
+ 53.71403496246785,
213
+ 53.636976052075624,
214
+ 53.13433799892664,
215
+ 52.87010897882283,
216
+ 53.569856099784374,
217
+ 53.04888903629035,
218
+ 53.360027028247714,
219
+ 53.35383699275553,
220
+ 53.595576086081564,
221
+ 53.40233608148992,
222
+ 53.43425599858165,
223
+ 53.34917700383812,
224
+ 53.44594607595354,
225
+ 53.55702596716583,
226
+ 52.97180905472487,
227
+ 53.135977941565216,
228
+ 53.637044969946146,
229
+ 53.51222597528249,
230
+ 53.74323495198041,
231
+ 53.64776495844126,
232
+ 53.65436489228159,
233
+ 53.38411801494658,
234
+ 53.417376009747386,
235
+ 53.46064700279385
236
+ ],
237
+ "forwards": 4
238
+ },
239
+ "p50_ratio": 0.5698264140015692
240
+ },
241
+ "64": {
242
+ "previous": {
243
+ "p50_ms": 181.77667557029054,
244
+ "p95_ms": 184.529556077905,
245
+ "samples_ms": [
246
+ 181.30511394701898,
247
+ 182.10042896680534,
248
+ 183.3407029043883,
249
+ 181.7122820066288,
250
+ 181.48475303314626,
251
+ 182.18554893974215,
252
+ 181.45461298990995,
253
+ 182.56243702489883,
254
+ 181.425362941809,
255
+ 181.56275304500014,
256
+ 181.81476008612663,
257
+ 182.31657799333334,
258
+ 181.52347404975444,
259
+ 181.47010309621692,
260
+ 181.0953450622037,
261
+ 181.81708198972046,
262
+ 182.10425099823624,
263
+ 182.6307469746098,
264
+ 182.9800249543041,
265
+ 184.529556077905,
266
+ 184.77372406050563,
267
+ 182.89187492337078,
268
+ 182.06122005358338,
269
+ 181.52308301068842,
270
+ 181.62922305054963,
271
+ 181.69329199008644,
272
+ 181.3240039627999,
273
+ 181.9472999777645,
274
+ 181.73859105445445,
275
+ 181.2681140145287
276
+ ],
277
+ "forwards": 8
278
+ },
279
+ "current": {
280
+ "p50_ms": 87.38938998430967,
281
+ "p95_ms": 88.08800508268178,
282
+ "samples_ms": [
283
+ 87.14881108608097,
284
+ 87.20357099082321,
285
+ 87.67556794919074,
286
+ 87.43762003723532,
287
+ 87.8160479478538,
288
+ 87.28525089100003,
289
+ 87.12115103844553,
290
+ 87.49875891953707,
291
+ 87.27059105876833,
292
+ 87.15924096759409,
293
+ 87.37221104092896,
294
+ 87.13559003081173,
295
+ 87.83565799240023,
296
+ 87.46328996494412,
297
+ 87.40672003477812,
298
+ 86.9746939279139,
299
+ 88.08800508268178,
300
+ 87.65111805405468,
301
+ 87.68402808345854,
302
+ 88.8292919844389,
303
+ 87.40467997267842,
304
+ 87.23169099539518,
305
+ 87.6758280210197,
306
+ 87.31710002757609,
307
+ 87.82118710223585,
308
+ 87.21276093274355,
309
+ 87.51507895067334,
310
+ 87.28954102844,
311
+ 87.37409999594092,
312
+ 86.76742296665907
313
+ ],
314
+ "forwards": 8
315
+ },
316
+ "p50_ratio": 0.480751392939395
317
+ },
318
+ "128": {
319
+ "previous": {
320
+ "p50_ms": 355.9657375444658,
321
+ "p95_ms": 360.2475899970159,
322
+ "samples_ms": [
323
+ 356.70211003161967,
324
+ 357.4898249935359,
325
+ 356.45368101540953,
326
+ 355.5921650258824,
327
+ 356.4777700230479,
328
+ 355.7541760383174,
329
+ 331.34153904393315,
330
+ 355.33637704793364,
331
+ 355.07863899692893,
332
+ 354.9016900360584,
333
+ 360.2475899970159,
334
+ 355.7152450084686,
335
+ 356.9431280484423,
336
+ 361.3199539249763,
337
+ 356.9705080008134,
338
+ 355.9749521082267,
339
+ 358.1079120049253,
340
+ 356.542830937542,
341
+ 356.62822995800525,
342
+ 355.24915694259107,
343
+ 355.9565229807049,
344
+ 358.03848202340305,
345
+ 357.22736606840044,
346
+ 354.29402289446443,
347
+ 354.5909799868241,
348
+ 355.8202739804983,
349
+ 355.5644250009209,
350
+ 354.77684903889894,
351
+ 354.4976110570133,
352
+ 356.43676901236176
353
+ ],
354
+ "forwards": 16
355
+ },
356
+ "current": {
357
+ "p50_ms": 154.4099859893322,
358
+ "p95_ms": 155.9316230705008,
359
+ "samples_ms": [
360
+ 154.1334129869938,
361
+ 154.66729004401714,
362
+ 154.6833390602842,
363
+ 154.5290610520169,
364
+ 154.0722930803895,
365
+ 153.57524703722447,
366
+ 148.0596859473735,
367
+ 150.81938996445388,
368
+ 154.5265899039805,
369
+ 153.9489240385592,
370
+ 156.18523210287094,
371
+ 154.25303194206208,
372
+ 154.7749990131706,
373
+ 154.7707199351862,
374
+ 155.18246707506478,
375
+ 154.5011909911409,
376
+ 155.9316230705008,
377
+ 154.60049000103027,
378
+ 154.8181800171733,
379
+ 154.3837309582159,
380
+ 154.88766902126372,
381
+ 154.30260205175728,
382
+ 154.4362410204485,
383
+ 153.7249149987474,
384
+ 154.51833105180413,
385
+ 154.30224197916687,
386
+ 154.14311294443905,
387
+ 153.9313340326771,
388
+ 153.8708939915523,
389
+ 154.12014396861196
390
+ ],
391
+ "forwards": 16
392
+ },
393
+ "p50_ratio": 0.4337776637001305
394
+ }
395
+ },
396
+ "gate_pass": true,
397
+ "actual_forwards": 2400,
398
+ "weights_unchanged": true,
399
+ "native_manifest_sha256": "f288d873999832a3f37c6a7c4268c2ab309691e621794dbf7acab891acbbb7e6",
400
+ "plan_sha256": "c10e0c4fbe385577ffece0e1d18b274f6870079e80c29fdfec92927a515009a7",
401
+ "environment": {
402
+ "gpu": "",
403
+ "torch": "2.12.0+git6bbd260",
404
+ "hip": "7.2.53211",
405
+ "precision": "FP32",
406
+ "batch_size": 8
407
+ },
408
+ "workload": {
409
+ "state": "The cafe has excellent coffee and friendly staff. Weekend queues are long. The customer wants more staff during busy hours, and says they will return.",
410
+ "questions": [
411
+ {
412
+ "type": "choice",
413
+ "instructions": "Which type of business is reviewed?",
414
+ "criteria": {
415
+ "cafe": "Coffee and refreshments",
416
+ "hotel": "Overnight accommodation",
417
+ "bank": "Financial services"
418
+ }
419
+ },
420
+ {
421
+ "type": "noul",
422
+ "instructions": "Does the customer intend to return?"
423
+ },
424
+ {
425
+ "type": "score",
426
+ "instructions": "Rate the quality of the coffee.",
427
+ "criteria": [
428
+ "Poor",
429
+ "Average",
430
+ "Excellent"
431
+ ]
432
+ }
433
+ ],
434
+ "ordering": "Cycle three fixed question types, one shared context, fresh IDs only",
435
+ "warmup_pairs": 10,
436
+ "measured_pairs": 30,
437
+ "clock": "Synchronized resident local SystemOne including conversion/tokenization/forward/answers; excludes HTTP/Studio/network"
438
+ },
439
+ "source_hashes": {
440
+ "Lex/prepared/decision_inference/_grouped.py": "80551e1bebd10634abcbbfce1aecf6cafab93b7e178399c428dd97022c6f9d95",
441
+ "Lex/prepared/decision_inference/_system_one.py": "0dc61f7558eed31c9ac186827b2050c0d391c69ea3684c02e8cd9c71400b1690"
442
+ },
443
+ "public_release": false
444
+ },
445
+ "full_source_qualification": {
446
+ "sha256": "59a6914eef31575f31e1fbb68e26d7dc2e9b44fc4ae755babd91668c0329a920",
447
+ "content": {
448
+ "status": "COMPLETE_ACTUAL_AMD_SYSTEMONE_SCHEDULE_QUALIFICATION",
449
+ "model": "Lex",
450
+ "panels": {
451
+ "old_core": {
452
+ "requested_requests": 880,
453
+ "accepted_requests": 880,
454
+ "refused_requests": 0,
455
+ "compared_decisions": 880,
456
+ "hard_flips": 0,
457
+ "max_probability_delta": 3.6656856536865234e-06,
458
+ "mismatched_order_or_usage": 0
459
+ },
460
+ "v3_core": {
461
+ "requested_requests": 880,
462
+ "accepted_requests": 880,
463
+ "refused_requests": 0,
464
+ "compared_decisions": 880,
465
+ "hard_flips": 0,
466
+ "max_probability_delta": 0.0,
467
+ "mismatched_order_or_usage": 0
468
+ },
469
+ "v4": {
470
+ "requested_requests": 480,
471
+ "accepted_requests": 480,
472
+ "refused_requests": 0,
473
+ "compared_decisions": 480,
474
+ "hard_flips": 0,
475
+ "max_probability_delta": 0.0,
476
+ "mismatched_order_or_usage": 0
477
+ },
478
+ "v5": {
479
+ "requested_requests": 480,
480
+ "accepted_requests": 480,
481
+ "refused_requests": 0,
482
+ "compared_decisions": 480,
483
+ "hard_flips": 0,
484
+ "max_probability_delta": 4.1425228118896484e-06,
485
+ "mismatched_order_or_usage": 0
486
+ },
487
+ "old_native": {
488
+ "requested_requests": 68,
489
+ "accepted_requests": 29,
490
+ "refused_requests": 39,
491
+ "compared_decisions": 29,
492
+ "hard_flips": 0,
493
+ "max_probability_delta": 1.1920928955078125e-06,
494
+ "mismatched_order_or_usage": 0
495
+ },
496
+ "v3_native": {
497
+ "requested_requests": 68,
498
+ "accepted_requests": 50,
499
+ "refused_requests": 18,
500
+ "compared_decisions": 202,
501
+ "hard_flips": 0,
502
+ "max_probability_delta": 3.2782554626464844e-06,
503
+ "mismatched_order_or_usage": 0
504
+ },
505
+ "transfer_v9_test": {
506
+ "requested_requests": 1209,
507
+ "accepted_requests": 1209,
508
+ "refused_requests": 0,
509
+ "compared_decisions": 1209,
510
+ "hard_flips": 0,
511
+ "max_probability_delta": 2.3245811462402344e-06,
512
+ "mismatched_order_or_usage": 0
513
+ }
514
+ },
515
+ "gate_pass": true,
516
+ "actual_forwards": 1044,
517
+ "weights_unchanged": true,
518
+ "request_order_and_full_admission_preserved": true,
519
+ "gold_access": false,
520
+ "public_release": false
521
+ }
522
+ },
523
+ "package_qualification": {
524
+ "status": "PASS_ACTUAL_STAGED_SYSTEMONE_PACKAGE",
525
+ "model": "Lex",
526
+ "checks": [
527
+ {
528
+ "case": "observed192_first128",
529
+ "decisions": 128,
530
+ "hard_flips": 0,
531
+ "max_probability_delta": 2.4139881134033203e-06
532
+ },
533
+ {
534
+ "case": "observed192_last64",
535
+ "decisions": 64,
536
+ "hard_flips": 0,
537
+ "max_probability_delta": 8.940696716308594e-06
538
+ },
539
+ {
540
+ "case": "studio:lex-inbox-contexts-v1",
541
+ "decisions": 8,
542
+ "hard_flips": 0,
543
+ "max_probability_delta": 0.0
544
+ },
545
+ {
546
+ "case": "studio:lex-ticket-batch",
547
+ "decisions": 8,
548
+ "hard_flips": 0,
549
+ "max_probability_delta": 0.0
550
+ },
551
+ {
552
+ "case": "studio:lex-security-six-v3",
553
+ "decisions": 6,
554
+ "hard_flips": 0,
555
+ "max_probability_delta": 1.6808509826660156e-05
556
+ },
557
+ {
558
+ "case": "studio:lex-urgency-six-v3",
559
+ "decisions": 6,
560
+ "hard_flips": 0,
561
+ "max_probability_delta": 0.0
562
+ },
563
+ {
564
+ "case": "studio:lex-access-six-v4",
565
+ "decisions": 6,
566
+ "hard_flips": 0,
567
+ "max_probability_delta": 5.960464477539063e-08
568
+ },
569
+ {
570
+ "case": "studio:lex-request-contexts-v4",
571
+ "decisions": 6,
572
+ "hard_flips": 0,
573
+ "max_probability_delta": 0.0
574
+ },
575
+ {
576
+ "case": "four_contexts_512_decisions",
577
+ "decisions": 512,
578
+ "hard_flips": 0,
579
+ "max_probability_delta": 5.066394805908203e-07
580
+ },
581
+ {
582
+ "case": "optional_auto_preserved",
583
+ "decisions": 32,
584
+ "hard_flips": 0,
585
+ "max_probability_delta": 0.0
586
+ },
587
+ {
588
+ "case": "external_id_alias",
589
+ "decisions": 128,
590
+ "probabilities_exact": true
591
+ },
592
+ {
593
+ "case": "late_overflow",
594
+ "same_error_before_forward": true
595
+ },
596
+ {
597
+ "case": "over_512",
598
+ "same_error_before_forward": true
599
+ },
600
+ {
601
+ "case": "over_128_requests",
602
+ "same_error_before_forward": true
603
+ }
604
+ ],
605
+ "actual_forwards": 228,
606
+ "weights_unchanged": true,
607
+ "native_manifest_sha256": "f288d873999832a3f37c6a7c4268c2ab309691e621794dbf7acab891acbbb7e6",
608
+ "plan_sha256": "855ea4bf964db543ae45b3c0201d288248f1745576b6393a3bdbb24920275873",
609
+ "source_hashes": {
610
+ "Lex/prepared/decision_inference/_grouped.py": "80551e1bebd10634abcbbfce1aecf6cafab93b7e178399c428dd97022c6f9d95",
611
+ "Lex/prepared/decision_inference/_system_one.py": "0dc61f7558eed31c9ac186827b2050c0d391c69ea3684c02e8cd9c71400b1690"
612
+ },
613
+ "studio_examples": 6,
614
+ "public_release": false
615
+ },
616
+ "earlier_concurrent_measurement": {
617
+ "status": "COMPLETE_ACTUAL_RELEASE_PACKAGE_LATENCY",
618
+ "model": "Lex",
619
+ "timing": {
620
+ "1": {
621
+ "previous": {
622
+ "p50_ms": 14.571013452950865,
623
+ "p95_ms": 106.91890399903059,
624
+ "samples_ms": [
625
+ 13.361250050365925,
626
+ 140.63826599158347,
627
+ 13.696447014808655,
628
+ 27.827713056467474,
629
+ 13.735816930420697,
630
+ 62.6358789158985,
631
+ 19.878904917277396,
632
+ 18.178944010287523,
633
+ 106.91890399903059,
634
+ 13.263568980619311,
635
+ 13.29776004422456,
636
+ 42.206557001918554,
637
+ 13.289999915286899,
638
+ 15.967766055837274,
639
+ 13.450457947328687,
640
+ 19.57047695759684,
641
+ 94.55266897566617,
642
+ 13.303969986736774,
643
+ 13.46319995354861,
644
+ 13.324889005161822,
645
+ 13.705997029319406,
646
+ 13.202580972574651,
647
+ 70.73775609023869,
648
+ 66.92217499949038,
649
+ 13.30288895405829,
650
+ 15.406209975481033,
651
+ 13.376468908973038,
652
+ 13.277549995109439,
653
+ 23.668864043429494,
654
+ 16.42183295916766
655
+ ],
656
+ "forwards": 1
657
+ },
658
+ "current": {
659
+ "p50_ms": 13.378849485889077,
660
+ "p95_ms": 97.0877748914063,
661
+ "samples_ms": [
662
+ 17.63901603408158,
663
+ 13.316189986653626,
664
+ 74.40948602743447,
665
+ 165.44692497700453,
666
+ 13.580687926150858,
667
+ 16.24104392249137,
668
+ 13.239049934782088,
669
+ 13.382968958467245,
670
+ 13.367859064601362,
671
+ 13.338689925149083,
672
+ 15.964004909619689,
673
+ 13.253489974886179,
674
+ 13.302339008077979,
675
+ 13.258340070024133,
676
+ 56.361171999014914,
677
+ 13.369528925977647,
678
+ 13.29931989312172,
679
+ 17.52129604574293,
680
+ 19.00173898320645,
681
+ 13.478238019160926,
682
+ 97.0877748914063,
683
+ 13.302690000273287,
684
+ 13.37473001331091,
685
+ 13.357058051042259,
686
+ 13.258830062113702,
687
+ 16.982758999802172,
688
+ 23.39470700826496,
689
+ 13.19589908234775,
690
+ 13.20873899385333,
691
+ 13.39158893097192
692
+ ],
693
+ "forwards": 1
694
+ },
695
+ "p50_ratio": 0.9181824949300039
696
+ },
697
+ "8": {
698
+ "previous": {
699
+ "p50_ms": 34.73892150213942,
700
+ "p95_ms": 172.8683450492099,
701
+ "samples_ms": [
702
+ 31.131225056014955,
703
+ 28.410759987309575,
704
+ 85.56938706897199,
705
+ 41.561309015378356,
706
+ 30.786116956733167,
707
+ 38.726025028154254,
708
+ 145.64779901411384,
709
+ 82.93773094192147,
710
+ 28.462579008191824,
711
+ 68.68544709868729,
712
+ 34.76842597592622,
713
+ 30.330618959851563,
714
+ 34.70941702835262,
715
+ 33.41492300387472,
716
+ 190.743581042625,
717
+ 31.30575397517532,
718
+ 105.50514201167971,
719
+ 87.7190459286794,
720
+ 28.843987034633756,
721
+ 31.655543018132448,
722
+ 36.35207703337073,
723
+ 170.99016509018838,
724
+ 31.442093080841005,
725
+ 73.0566029669717,
726
+ 28.541999054141343,
727
+ 33.372384030371904,
728
+ 85.321647929959,
729
+ 172.8683450492099,
730
+ 31.578613095916808,
731
+ 30.789837008342147
732
+ ],
733
+ "forwards": 1
734
+ },
735
+ "current": {
736
+ "p50_ms": 43.07004599831998,
737
+ "p95_ms": 282.32143609784544,
738
+ "samples_ms": [
739
+ 41.653947904706,
740
+ 28.454498969949782,
741
+ 28.578040073625743,
742
+ 286.50921303778887,
743
+ 28.733848012052476,
744
+ 87.18010690063238,
745
+ 32.454418018460274,
746
+ 34.42092798650265,
747
+ 187.3356889700517,
748
+ 28.99436606094241,
749
+ 122.0393639523536,
750
+ 28.320749988779426,
751
+ 71.88508997205645,
752
+ 31.886521028354764,
753
+ 33.56740204617381,
754
+ 282.32143609784544,
755
+ 32.90519502479583,
756
+ 255.91422605793923,
757
+ 28.52455899119377,
758
+ 75.28732100035995,
759
+ 223.87088497634977,
760
+ 63.35321499500424,
761
+ 72.95793399680406,
762
+ 28.87122705578804,
763
+ 31.580012990161777,
764
+ 44.486144091933966,
765
+ 94.85666803084314,
766
+ 153.9103949908167,
767
+ 110.09115702472627,
768
+ 29.06379592604935
769
+ ],
770
+ "forwards": 1
771
+ },
772
+ "p50_ratio": 1.2398210461331531
773
+ },
774
+ "32": {
775
+ "previous": {
776
+ "p50_ms": 198.33825452951714,
777
+ "p95_ms": 286.02026600856334,
778
+ "samples_ms": [
779
+ 208.20390700828284,
780
+ 286.02026600856334,
781
+ 116.49416398722678,
782
+ 167.50254202634096,
783
+ 246.3420460699126,
784
+ 345.44017002917826,
785
+ 145.68236900959164,
786
+ 167.11201495490968,
787
+ 206.77826495375484,
788
+ 107.46569093316793,
789
+ 158.0538930138573,
790
+ 193.4651560150087,
791
+ 243.90959809534252,
792
+ 182.90021095890552,
793
+ 221.07835905626416,
794
+ 187.00264894869179,
795
+ 244.45490492507815,
796
+ 172.653064946644,
797
+ 125.88770198635757,
798
+ 227.9231830034405,
799
+ 259.72379301674664,
800
+ 263.34351405967027,
801
+ 163.20275398902595,
802
+ 203.2113530440256,
803
+ 275.67597897723317,
804
+ 208.6803240235895,
805
+ 119.42932708188891,
806
+ 211.07498207129538,
807
+ 151.2328089447692,
808
+ 170.27989693451673
809
+ ],
810
+ "forwards": 4
811
+ },
812
+ "current": {
813
+ "p50_ms": 133.98809952195734,
814
+ "p95_ms": 376.0747379856184,
815
+ "samples_ms": [
816
+ 100.60080699622631,
817
+ 355.66331702284515,
818
+ 111.73206905368716,
819
+ 68.3719179360196,
820
+ 335.7284109806642,
821
+ 64.3072499660775,
822
+ 208.7552340235561,
823
+ 117.52685799729079,
824
+ 160.36495100706816,
825
+ 121.48911599069834,
826
+ 376.0747379856184,
827
+ 58.965447009541094,
828
+ 136.99047500267625,
829
+ 122.37974093295634,
830
+ 59.73403400275856,
831
+ 212.6397140091285,
832
+ 63.47907392773777,
833
+ 171.72804998699576,
834
+ 163.32766390405595,
835
+ 168.65042597055435,
836
+ 136.3753970945254,
837
+ 66.5766280144453,
838
+ 74.80555400252342,
839
+ 59.44409500807524,
840
+ 71.25750207342207,
841
+ 223.68835506495088,
842
+ 289.2386970343068,
843
+ 409.67884799465537,
844
+ 131.60080194938928,
845
+ 287.98995399847627
846
+ ],
847
+ "forwards": 4
848
+ },
849
+ "p50_ratio": 0.675553487348135
850
+ },
851
+ "64": {
852
+ "previous": {
853
+ "p50_ms": 488.5446400148794,
854
+ "p95_ms": 860.3939639870077,
855
+ "samples_ms": [
856
+ 588.6253280332312,
857
+ 356.67536908295006,
858
+ 594.9330650037155,
859
+ 692.8134460467845,
860
+ 396.07378002256155,
861
+ 527.112663956359,
862
+ 336.06783696450293,
863
+ 456.5627579577267,
864
+ 326.6586670652032,
865
+ 752.5941269705072,
866
+ 860.3939639870077,
867
+ 681.654252926819,
868
+ 430.87316304445267,
869
+ 428.6890249932185,
870
+ 404.57855199929327,
871
+ 538.390701985918,
872
+ 561.1566319130361,
873
+ 469.90221505984664,
874
+ 674.7355579864234,
875
+ 500.82891003694385,
876
+ 914.364974014461,
877
+ 534.5511310733855,
878
+ 477.23246505483985,
879
+ 550.633255043067,
880
+ 499.85681497491896,
881
+ 198.40770598966628,
882
+ 211.90554497297853,
883
+ 182.23994190339,
884
+ 182.4631510535255,
885
+ 182.32450098730624
886
+ ],
887
+ "forwards": 8
888
+ },
889
+ "current": {
890
+ "p50_ms": 240.10508198989555,
891
+ "p95_ms": 546.4967269217595,
892
+ "samples_ms": [
893
+ 237.67261998727918,
894
+ 197.24152295384556,
895
+ 170.06384895648807,
896
+ 191.1246069939807,
897
+ 316.73357996623963,
898
+ 374.13932604249567,
899
+ 242.53754399251193,
900
+ 244.4212429691106,
901
+ 200.5026569822803,
902
+ 435.37590198684484,
903
+ 453.023066977039,
904
+ 593.5483209323138,
905
+ 399.27821105811745,
906
+ 243.32293902989477,
907
+ 246.88830005470663,
908
+ 145.2575579751283,
909
+ 233.21076203137636,
910
+ 403.36408896837384,
911
+ 166.22657689731568,
912
+ 229.93817902170122,
913
+ 276.7654510680586,
914
+ 126.66724692098796,
915
+ 546.4967269217595,
916
+ 448.94893595483154,
917
+ 465.4096569865942,
918
+ 106.47687502205372,
919
+ 121.47569505032152,
920
+ 108.83200191892684,
921
+ 88.08667189441621,
922
+ 87.95632293913513
923
+ ],
924
+ "forwards": 8
925
+ },
926
+ "p50_ratio": 0.49147009776339534
927
+ },
928
+ "128": {
929
+ "previous": {
930
+ "p50_ms": 359.8495615296997,
931
+ "p95_ms": 371.3875049725175,
932
+ "samples_ms": [
933
+ 359.17470196727663,
934
+ 359.7314669750631,
935
+ 359.6855190116912,
936
+ 359.19495997950435,
937
+ 358.3321350160986,
938
+ 358.4663450019434,
939
+ 359.8091179737821,
940
+ 359.1001909226179,
941
+ 358.5895940195769,
942
+ 359.10174099262804,
943
+ 360.01692595891654,
944
+ 360.8712210552767,
945
+ 361.5627179387957,
946
+ 360.6750110629946,
947
+ 360.5347230331972,
948
+ 359.8734770203009,
949
+ 360.36473501008004,
950
+ 361.02003895211965,
951
+ 360.0647649727762,
952
+ 371.3875049725175,
953
+ 379.1667229961604,
954
+ 360.60995201114565,
955
+ 359.8042860394344,
956
+ 359.8256460390985,
957
+ 360.15939398203045,
958
+ 361.5629649721086,
959
+ 359.4064770732075,
960
+ 359.6939370036125,
961
+ 359.90800498984754,
962
+ 359.70397596247494
963
+ ],
964
+ "forwards": 16
965
+ },
966
+ "current": {
967
+ "p50_ms": 156.12808446167037,
968
+ "p95_ms": 158.11885800212622,
969
+ "samples_ms": [
970
+ 155.28889501001686,
971
+ 155.16494493931532,
972
+ 155.34398599993438,
973
+ 155.51191405393183,
974
+ 155.16564599238336,
975
+ 155.8415510226041,
976
+ 155.35369399003685,
977
+ 155.15018592122942,
978
+ 155.40425409562886,
979
+ 156.3086489913985,
980
+ 156.2130000675097,
981
+ 156.37198893819004,
982
+ 156.19600901845843,
983
+ 156.0601599048823,
984
+ 155.79097100999206,
985
+ 156.46956791169941,
986
+ 156.37390792835504,
987
+ 156.02733101695776,
988
+ 156.69844602234662,
989
+ 156.47772699594498,
990
+ 155.98976996261626,
991
+ 156.3491279957816,
992
+ 156.2649189727381,
993
+ 156.37214807793498,
994
+ 156.49046702310443,
995
+ 155.80621093977243,
996
+ 156.68375696986914,
997
+ 155.72028001770377,
998
+ 160.92158399987966,
999
+ 158.11885800212622
1000
+ ],
1001
+ "forwards": 16
1002
+ },
1003
+ "p50_ratio": 0.4338704312935075
1004
+ }
1005
+ },
1006
+ "gate_pass": false,
1007
+ "actual_forwards": 1920,
1008
+ "weights_unchanged": true,
1009
+ "native_manifest_sha256": "f288d873999832a3f37c6a7c4268c2ab309691e621794dbf7acab891acbbb7e6",
1010
+ "plan_sha256": "74d011196525925c371bf73aaf40c69a7b2c35249f84769543a68f429f8200bd",
1011
+ "environment": {
1012
+ "gpu": "",
1013
+ "torch": "2.12.0+git6bbd260",
1014
+ "hip": "7.2.53211",
1015
+ "precision": "FP32",
1016
+ "batch_size": 8
1017
+ },
1018
+ "workload": {
1019
+ "state": "The cafe has excellent coffee and friendly staff. Weekend queues are long. The customer wants more staff during busy hours, and says they will return.",
1020
+ "questions": [
1021
+ {
1022
+ "type": "choice",
1023
+ "instructions": "Which type of business is reviewed?",
1024
+ "criteria": {
1025
+ "cafe": "Coffee and refreshments",
1026
+ "hotel": "Overnight accommodation",
1027
+ "bank": "Financial services"
1028
+ }
1029
+ },
1030
+ {
1031
+ "type": "noul",
1032
+ "instructions": "Does the customer intend to return?"
1033
+ },
1034
+ {
1035
+ "type": "score",
1036
+ "instructions": "Rate the quality of the coffee.",
1037
+ "criteria": [
1038
+ "Poor",
1039
+ "Average",
1040
+ "Excellent"
1041
+ ]
1042
+ }
1043
+ ],
1044
+ "ordering": "Cycle three fixed question types, one shared context, fresh IDs only",
1045
+ "warmup_pairs": 2,
1046
+ "measured_pairs": 30,
1047
+ "clock": "Synchronized resident local SystemOne including conversion/tokenization/forward/answers; excludes HTTP/Studio/network"
1048
+ },
1049
+ "source_hashes": {
1050
+ "Lex/prepared/decision_inference/_grouped.py": "80551e1bebd10634abcbbfce1aecf6cafab93b7e178399c428dd97022c6f9d95",
1051
+ "Lex/prepared/decision_inference/_system_one.py": "0dc61f7558eed31c9ac186827b2050c0d391c69ea3684c02e8cd9c71400b1690"
1052
+ },
1053
+ "public_release": false
1054
+ },
1055
+ "measurement_note": "Initial concurrent final-package run had substantial nonstationary stalls and LexQ8 failed its gate. Its samples are retained. Final fixed serial measurement follows termination of own training/data producers,10warmup and30measured alternating pairs. No sample filtering; unchanged acceptance gates."
1056
+ }
PACKAGE_MANIFEST.json CHANGED
@@ -1,7 +1,7 @@
1
  {
2
- "bytes_excluding_this_manifest": 2327366131,
3
  "display_name": "Decision-1.0-Lex-0.6B",
4
- "file_count_excluding_this_manifest": 102,
5
  "files": {
6
  ".gitattributes": {
7
  "bytes": 753,
@@ -63,6 +63,10 @@
63
  "bytes": 2543,
64
  "sha256": "a7dd6dec55dfe463cc0211c3dffa114b88bb95db90816b70c18b62b5d7be7ad1"
65
  },
 
 
 
 
66
  "NOTICE": {
67
  "bytes": 4371,
68
  "sha256": "e3f71eb2a262084fa572ccbf9b65110a4feb9d8c625dc26e792beba282cdf032"
@@ -72,20 +76,20 @@
72
  "sha256": "2e5936d3ea196e00968cc37b7662a6afbe8cfd6fd11d04b9fecba554b1e1e37b"
73
  },
74
  "PERFORMANCE.md": {
75
- "bytes": 4368,
76
- "sha256": "7e275e73c18b756a20b886272750b58416ead2c345b12393e162050231aa8f64"
77
  },
78
  "README.md": {
79
- "bytes": 5993,
80
- "sha256": "5b89d9013b9890381f169c496081d23818836119417bcf1a5f9af1893297d668"
81
  },
82
  "RUNTIME_UPDATE.json": {
83
  "bytes": 1087,
84
  "sha256": "0a77c012fe1bd23c5f1a06c106c1a29ae5ee156248db218fbbc6d631d1c5aff4"
85
  },
86
  "SYSTEM_ONE.md": {
87
- "bytes": 5735,
88
- "sha256": "414e12680a6ec2dbbbc660d1f6ba0d601fa941388476a62696b7f5e8a3e11ef9"
89
  },
90
  "SYSTEM_ONE_VALIDATION.json": {
91
  "bytes": 2583,
@@ -104,8 +108,8 @@
104
  "sha256": "5c95cf851a9480e6037bc190d0deee87da35576489362262ad4151bafbb96990"
105
  },
106
  "USAGE.md": {
107
- "bytes": 7866,
108
- "sha256": "652a6f94b27e352a3bd899c8fdb1a0ff4ed4b7598824a598a2947e124b73cd33"
109
  },
110
  "VALIDATION.md": {
111
  "bytes": 2120,
@@ -155,6 +159,18 @@
155
  "bytes": 2188963,
156
  "sha256": "9980c6b901aff89b69e9cc6de6de40408b95c7378f18e931401198f42d32e40b"
157
  },
 
 
 
 
 
 
 
 
 
 
 
 
158
  "assets/residual-layers.pdf": {
159
  "bytes": 131144,
160
  "sha256": "2c5b50bb840ae115a1264b66d535a1048172e891a33d02d507c193525cc45630"
@@ -203,13 +219,17 @@
203
  "bytes": 1325,
204
  "sha256": "3602e3a06386daab5494305413a08adb581f23007ffb5ca6e5df62fa23d11bd6"
205
  },
 
 
 
 
206
  "decision_inference/_request.py": {
207
  "bytes": 3026,
208
  "sha256": "85b8349cdf08a3550027606b3108778955f84d66c52fa11dd908ea50311e69e5"
209
  },
210
  "decision_inference/_system_one.py": {
211
- "bytes": 10690,
212
- "sha256": "d2ef9aa1d0badcf168065a782f8a96316045455a676c08bdb5dd384a651e2194"
213
  },
214
  "decision_inference/profile.py": {
215
  "bytes": 2071,
@@ -428,7 +448,7 @@
428
  },
429
  "parameter_tensors": 489,
430
  "parameters": 571909635,
431
- "public_non_native_bytes_excluding_this_manifest": 5106920,
432
  "publication": {
433
  "download_access": "public_ungated",
434
  "new_contribution_license": "Apache-2.0",
@@ -443,7 +463,9 @@
443
  "status": "READY_FOR_ROOT_PUBLICATION",
444
  "subject_manifest_sha256": "f288d873999832a3f37c6a7c4268c2ab309691e621794dbf7acab891acbbb7e6",
445
  "system_one_api": {
 
446
  "default_physical_batch_limit": 8,
 
447
  "entry": "decision_inference.SystemOne",
448
  "finetune_conversion": "decision_finetune.system_one.system_one_training_rows",
449
  "max_batch_decisions": 512,
@@ -457,5 +479,15 @@
457
  "weights_unchanged": true
458
  },
459
  "technical_validation": "TECHNICAL_VALIDATION.json",
460
- "training_provenance": "TRAINING_PROVENANCE.json"
 
 
 
 
 
 
 
 
 
 
461
  }
 
1
  {
2
+ "bytes_excluding_this_manifest": 2327567235,
3
  "display_name": "Decision-1.0-Lex-0.6B",
4
+ "file_count_excluding_this_manifest": 107,
5
  "files": {
6
  ".gitattributes": {
7
  "bytes": 753,
 
63
  "bytes": 2543,
64
  "sha256": "a7dd6dec55dfe463cc0211c3dffa114b88bb95db90816b70c18b62b5d7be7ad1"
65
  },
66
+ "MIXED_QUESTION_SCALING.json": {
67
+ "bytes": 32763,
68
+ "sha256": "ed4a9e7a9aac0385f8a4516694a6d9ea843aacefde6c7189b9980bac86c9d773"
69
+ },
70
  "NOTICE": {
71
  "bytes": 4371,
72
  "sha256": "e3f71eb2a262084fa572ccbf9b65110a4feb9d8c625dc26e792beba282cdf032"
 
76
  "sha256": "2e5936d3ea196e00968cc37b7662a6afbe8cfd6fd11d04b9fecba554b1e1e37b"
77
  },
78
  "PERFORMANCE.md": {
79
+ "bytes": 6351,
80
+ "sha256": "0f5c84a90b2c82be6146b1b7032a9f836e5858d59e29bd6c94f016be36d8b3b6"
81
  },
82
  "README.md": {
83
+ "bytes": 6230,
84
+ "sha256": "6ff9eb7a6ac0c6b90b017f497d75bd3366c9c598ce3c5f813dfa0f192d49b1ec"
85
  },
86
  "RUNTIME_UPDATE.json": {
87
  "bytes": 1087,
88
  "sha256": "0a77c012fe1bd23c5f1a06c106c1a29ae5ee156248db218fbbc6d631d1c5aff4"
89
  },
90
  "SYSTEM_ONE.md": {
91
+ "bytes": 6177,
92
+ "sha256": "c59c5d1fbab6c18033b1339701aadddc34653f9fb93f79aee93534bb34d66648"
93
  },
94
  "SYSTEM_ONE_VALIDATION.json": {
95
  "bytes": 2583,
 
108
  "sha256": "5c95cf851a9480e6037bc190d0deee87da35576489362262ad4151bafbb96990"
109
  },
110
  "USAGE.md": {
111
+ "bytes": 8308,
112
+ "sha256": "f7b6fe124f34667480f324f125d92d72e35045ea043c8b45da7184b9c7a28dcb"
113
  },
114
  "VALIDATION.md": {
115
  "bytes": 2120,
 
159
  "bytes": 2188963,
160
  "sha256": "9980c6b901aff89b69e9cc6de6de40408b95c7378f18e931401198f42d32e40b"
161
  },
162
+ "assets/mixed-question-scaling.pdf": {
163
+ "bytes": 27984,
164
+ "sha256": "e1261c998cdcdf6b72f57b4df4cb56cbf47fa521210c4778a0577e2f3a47c1cf"
165
+ },
166
+ "assets/mixed-question-scaling.png": {
167
+ "bytes": 121576,
168
+ "sha256": "4efadb4888c43122d461ca432081f2b5bd67c350f9605940b96f156829d65a54"
169
+ },
170
+ "assets/mixed-question-scaling.svg": {
171
+ "bytes": 14623,
172
+ "sha256": "0fcea2c3e7fd1c649a1dd7acf206c0d41e4f6fcab5e51abd838d6bb18a2560ff"
173
+ },
174
  "assets/residual-layers.pdf": {
175
  "bytes": 131144,
176
  "sha256": "2c5b50bb840ae115a1264b66d535a1048172e891a33d02d507c193525cc45630"
 
219
  "bytes": 1325,
220
  "sha256": "3602e3a06386daab5494305413a08adb581f23007ffb5ca6e5df62fa23d11bd6"
221
  },
222
+ "decision_inference/_grouped.py": {
223
+ "bytes": 1016,
224
+ "sha256": "80551e1bebd10634abcbbfce1aecf6cafab93b7e178399c428dd97022c6f9d95"
225
+ },
226
  "decision_inference/_request.py": {
227
  "bytes": 3026,
228
  "sha256": "85b8349cdf08a3550027606b3108778955f84d66c52fa11dd908ea50311e69e5"
229
  },
230
  "decision_inference/_system_one.py": {
231
+ "bytes": 10728,
232
+ "sha256": "0dc61f7558eed31c9ac186827b2050c0d391c69ea3684c02e8cd9c71400b1690"
233
  },
234
  "decision_inference/profile.py": {
235
  "bytes": 2071,
 
448
  },
449
  "parameter_tensors": 489,
450
  "parameters": 571909635,
451
+ "public_non_native_bytes_excluding_this_manifest": 5308024,
452
  "publication": {
453
  "download_access": "public_ungated",
454
  "new_contribution_license": "Apache-2.0",
 
463
  "status": "READY_FOR_ROOT_PUBLICATION",
464
  "subject_manifest_sha256": "f288d873999832a3f37c6a7c4268c2ab309691e621794dbf7acab891acbbb7e6",
465
  "system_one_api": {
466
+ "current_scheduling_validation": "MIXED_QUESTION_SCALING.json",
467
  "default_physical_batch_limit": 8,
468
+ "default_scheduling": "stable typed groups",
469
  "entry": "decision_inference.SystemOne",
470
  "finetune_conversion": "decision_finetune.system_one.system_one_training_rows",
471
  "max_batch_decisions": 512,
 
479
  "weights_unchanged": true
480
  },
481
  "technical_validation": "TECHNICAL_VALIDATION.json",
482
+ "training_provenance": "TRAINING_PROVENANCE.json",
483
+ "typed_scheduling_update": {
484
+ "complete_input_limit_unchanged": 1024,
485
+ "default_schedule": "Stable group by decision type; physicalB8; restore original caller order",
486
+ "entry": "decision_inference.SystemOne",
487
+ "native_weights_unchanged": true,
488
+ "new_chat_api": false,
489
+ "public_api_unchanged": true,
490
+ "subject_native_manifest_sha256": "f288d873999832a3f37c6a7c4268c2ab309691e621794dbf7acab891acbbb7e6",
491
+ "validation": "MIXED_QUESTION_SCALING.json"
492
+ }
493
  }
PERFORMANCE.md CHANGED
@@ -1,18 +1,44 @@
1
- # Lex: AMD batch capacity
2
 
3
- ## Optional larger batches
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4
 
5
  ```python
6
  from decision_inference import predict_1k, predict_auto_1k
7
 
8
- # native and records use the same objects as the Python usage examples.
9
  outputs = predict_1k(native, records) # unchanged default: up to 8
10
  outputs = predict_auto_1k(native, records) # opt in: up to 32 with the padding guard
11
  ```
12
 
13
  The optional entry admits every complete input before the first forward. It uses consecutive batches up to 32 only when the request has one decision type and total padded tokens do not increase relative to B8. Mixed types or increased padding fall back to B8; requests of eight or fewer take the unchanged default path. Input order, candidates, full 1,024-token limit and FP32 weights remain unchanged. It does not share contextual activations across questions. Larger physical batches can change floating-point rounding and peak memory.
14
 
15
- ### Paired B8/B32 study
16
 
17
  These are synchronized resident **Python API measurements**, not Studio or network latency. This earlier study compared default B8 with a homogeneous cap32 policy before the final padding guard was added. All eight measured points satisfy that guard, but **the final `predict_auto_1k` entry was validated separately and was not timed**. Existing default-B8 measurements above, if present, remain their original separate run.
18
 
@@ -33,7 +59,7 @@ Choice count curves repeat the same synthetic four-candidate question with opaqu
33
 
34
  At 1024×32, peak allocated memory increased from approximately 2.652 to 3.824 GiB. Shared allocator reserved-memory observations are not independent per-mode estimates. Equal total padding does not imply equal peak memory.
35
 
36
- ### Public-entry validation
37
 
38
  The actual imported package passed 46 forwards per model (92 total across Kai and Lex) over 225 technical row-occurrences/model, including all three types, dynamic candidate counts, enabled32, tail33, unequal-length fallback and mixed64. Late invalid and empty requests added zero forwards. The original strict bounds were 1e-4 logits, 2e-5 probability and normalized Score error, with zero hard-decision flips. Both models passed; fallback outputs were exact. All 489 parameters and 66 buffers retained identical contents and versions. No training, quality evaluation or timing was added in this integration check. Fixed repetition is not independent quality support.
39
 
 
1
+ # Lex: AMD SystemOne latency
2
 
3
+ ## Mixed questions, faster decisions
4
+
5
+ **128 mixed questions: 154.41 ms median, 56.6% lower than the previous default runtime.**
6
+
7
+ ![Lex: latency as mixed question count increases](assets/mixed-question-scaling.png)
8
+
9
+ | Questions | Previous p50 / p95 (ms) | Typed scheduling p50 / p95 (ms) |
10
+ |---:|---:|---:|
11
+ | 1 | 13.22 / 13.32 | 13.20 / 13.40 |
12
+ | 8 | 28.07 / 28.32 | 28.02 / 28.31 |
13
+ | 32 | 93.76 / 94.06 | 53.43 / 53.77 |
14
+ | 64 | 181.78 / 184.53 | 87.39 / 88.09 |
15
+ | 128 | 355.97 / 360.25 | 154.41 / 155.93 |
16
+
17
+ The request repeats three fixed questions—Choice, Noul and Score—over one unchanged context. Only question count and bookkeeping IDs change. The default SystemOne path groups admitted rows by type so each encoder path processes fuller batches, then restores the original answer order. Every question still has its own contextual computation. Weights, precision, complete-input limit and request schema are unchanged.
18
+
19
+ Measured using the exact release package on AMD ROCm, FP32, physical batch size eight: 10 warmup pairs and 30 alternating AB/BA measurement pairs per point, with GPU synchronization. Measurements include local request conversion, tokenization, model execution and answer assembly; they exclude transport and Studio. Models ran serially after our training and data jobs completed. Results describe this fixed workload, not a service latency guarantee.
20
+
21
+ Both models passed seven source panels: 4,160 admitted decisions without an argmax change and 57 identical whole-request refusals before any forward. Exact-package checks additionally cover 512 decisions, six current Studio examples, caller IDs/order, optional auto batching and input limits. Small floating-point probability differences remain possible.
22
+
23
+ [All samples, workload, checks and earlier concurrent measurements](MIXED_QUESTION_SCALING.json) · [SVG](assets/mixed-question-scaling.svg) · [PDF](assets/mixed-question-scaling.pdf)
24
+
25
+ ## Earlier measurements
26
+
27
+ ## Lex: AMD batch capacity
28
+
29
+ ### Optional larger batches
30
 
31
  ```python
32
  from decision_inference import predict_1k, predict_auto_1k
33
 
34
+ ## native and records use the same objects as the Python usage examples.
35
  outputs = predict_1k(native, records) # unchanged default: up to 8
36
  outputs = predict_auto_1k(native, records) # opt in: up to 32 with the padding guard
37
  ```
38
 
39
  The optional entry admits every complete input before the first forward. It uses consecutive batches up to 32 only when the request has one decision type and total padded tokens do not increase relative to B8. Mixed types or increased padding fall back to B8; requests of eight or fewer take the unchanged default path. Input order, candidates, full 1,024-token limit and FP32 weights remain unchanged. It does not share contextual activations across questions. Larger physical batches can change floating-point rounding and peak memory.
40
 
41
+ #### Paired B8/B32 study
42
 
43
  These are synchronized resident **Python API measurements**, not Studio or network latency. This earlier study compared default B8 with a homogeneous cap32 policy before the final padding guard was added. All eight measured points satisfy that guard, but **the final `predict_auto_1k` entry was validated separately and was not timed**. Existing default-B8 measurements above, if present, remain their original separate run.
44
 
 
59
 
60
  At 1024×32, peak allocated memory increased from approximately 2.652 to 3.824 GiB. Shared allocator reserved-memory observations are not independent per-mode estimates. Equal total padding does not imply equal peak memory.
61
 
62
+ #### Public-entry validation
63
 
64
  The actual imported package passed 46 forwards per model (92 total across Kai and Lex) over 225 technical row-occurrences/model, including all three types, dynamic candidate counts, enabled32, tail33, unequal-length fallback and mixed64. Late invalid and empty requests added zero forwards. The original strict bounds were 1e-4 logits, 2e-5 probability and normalized Score error, with zero hard-decision flips. Both models passed; fallback outputs were exact. All 489 parameters and 66 buffers retained identical contents and versions. No training, quality evaluation or timing was added in this integration check. Fixed repetition is not independent quality support.
65
 
README.md CHANGED
@@ -48,6 +48,8 @@ Lex is evaluated as an English specialist on these four workflows. For broader m
48
 
49
  **Optional larger batches:** `predict_auto_1k` uses up to 32 same-type questions when padding does not increase. Earlier B8/B32 comparisons showed approximately 23% lower median latency for 32 short questions and 35% for the multi-context fixture; the final guard was validated separately, not timed. [Usage and measurements](PERFORMANCE.md#optional-larger-batches).
50
 
 
 
51
  ## Use
52
 
53
  Replace the placeholder with a SystemOne-compatible endpoint configured to serve `Decision-1.0-Lex-0.6B`, and set `DECISION_API_KEY` to that endpoint's key.
 
48
 
49
  **Optional larger batches:** `predict_auto_1k` uses up to 32 same-type questions when padding does not increase. Earlier B8/B32 comparisons showed approximately 23% lower median latency for 32 short questions and 35% for the multi-context fixture; the final guard was validated separately, not timed. [Usage and measurements](PERFORMANCE.md#optional-larger-batches).
50
 
51
+ **128 mixed questions in 154 ms — 57% lower latency.** Automatic typed scheduling accelerates the default SystemOne path with the same weights. Paired local AMD measurements on a fixed workload. [Latency and scaling](PERFORMANCE.md).
52
+
53
  ## Use
54
 
55
  Replace the placeholder with a SystemOne-compatible endpoint configured to serve `Decision-1.0-Lex-0.6B`, and set `DECISION_API_KEY` to that endpoint's key.
SYSTEM_ONE.md CHANGED
@@ -127,3 +127,8 @@ source/component ID and stay in one split; the CLI checks component and exact
127
  input overlap. Optional `hard_target_ids` supplies separate evaluation labels
128
  under the same question IDs. Noul supports soft yes probabilities; Choice and
129
  Score support complete soft distributions in the original candidate order.
 
 
 
 
 
 
127
  input overlap. Optional `hard_target_ids` supplies separate evaluation labels
128
  under the same question IDs. Noul supports soft yes probabilities; Choice and
129
  Score support complete soft distributions in the original candidate order.
130
+
131
+
132
+ ### Default typed scheduling
133
+
134
+ The default `SystemOne` path groups complete, admitted questions by decision type in physical batches of eight, then restores the original request and question order. This works for many questions over one state and questions across multiple contexts. No API changes or application-side sorting are needed. The optional `batching="auto"` policy is unchanged. [Measured mixed-question scaling](PERFORMANCE.md).
USAGE.md CHANGED
@@ -114,3 +114,8 @@ The Python API returns original record/candidate identities, logits and probabil
114
  All state/question/candidate text and markers count toward the complete 1024 limit. `predict_1k` checks every input before the first forward and rejects overflow without truncation. Batch sizes 1–8 are supported by this wrapper. Use it rather than directly invoking the native research-capacity entrypoint.
115
 
116
  A compatible fine-tuned export can be used with `--native <directory> --manifest-sha256 <trusted hash>`. Only the pinned runtime revision and three-path architecture are accepted. Downloaded HF symlinks must be materialized into real files before loading. The manifest hash is an integrity check, not a reason to execute arbitrary untrusted Python.
 
 
 
 
 
 
114
  All state/question/candidate text and markers count toward the complete 1024 limit. `predict_1k` checks every input before the first forward and rejects overflow without truncation. Batch sizes 1–8 are supported by this wrapper. Use it rather than directly invoking the native research-capacity entrypoint.
115
 
116
  A compatible fine-tuned export can be used with `--native <directory> --manifest-sha256 <trusted hash>`. Only the pinned runtime revision and three-path architecture are accepted. Downloaded HF symlinks must be materialized into real files before loading. The manifest hash is an integrity check, not a reason to execute arbitrary untrusted Python.
117
+
118
+
119
+ ### Default typed scheduling
120
+
121
+ The default `SystemOne` path groups complete, admitted questions by decision type in physical batches of eight, then restores the original request and question order. This works for many questions over one state and questions across multiple contexts. No API changes or application-side sorting are needed. The optional `batching="auto"` policy is unchanged. [Measured mixed-question scaling](PERFORMANCE.md).
assets/mixed-question-scaling.pdf ADDED
Binary file (28 kB). View file
 
assets/mixed-question-scaling.png ADDED

Git LFS Details

  • SHA256: 4efadb4888c43122d461ca432081f2b5bd67c350f9605940b96f156829d65a54
  • Pointer size: 131 Bytes
  • Size of remote file: 122 kB
assets/mixed-question-scaling.svg ADDED
decision_inference/_grouped.py ADDED
@@ -0,0 +1,16 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Stable typed scheduling behind SystemOne; preserve request and answer order."""
2
+ from dataclasses import replace
3
+
4
+ def predict_grouped_1k(native,records,*,batch_size=8):
5
+ from decision_inference._request import request_collator
6
+ from decision_runtime import predict
7
+ if type(batch_size) is not int or batch_size!=8:raise ValueError('SystemOne typed scheduling uses physical batch size 8')
8
+ records=list(records);guard=request_collator(native.collator,records)
9
+ # Admit every original row before scheduling or any forward.
10
+ order=sorted(range(len(records)),key=lambda i:records[i]['question']['type'])
11
+ sorted_rows=[records[i]for i in order]
12
+ guard._records=tuple(sorted_rows);guard._encoded=[guard._encoded[i]for i in order]
13
+ predictions=predict(replace(native,collator=guard),sorted_rows,batch_size=8);guard.finish();restored=[None]*len(records)
14
+ for i,result in zip(order,predictions):restored[i]=result
15
+ if any(x is None for x in restored):raise RuntimeError('Incomplete scheduled predictions')
16
+ return restored
decision_inference/_system_one.py CHANGED
@@ -150,7 +150,7 @@ class SystemOne:
150
 
151
  evaluate(request) accepts the HTTP body shape; system_one(**request) is its
152
  Python equivalent. batch(requests) flattens independent states into GPU
153
- batches and restores the original request/question order. B8 is the default;
154
  batching='auto' opts into the published homogeneous padding-aware B32 path.
155
  """
156
  def __init__(self, native, *, model=None, batching="default"):
@@ -202,8 +202,8 @@ class SystemOne:
202
  from ._auto import predict_auto_1k
203
  predictions = predict_auto_1k(self.native, records)
204
  else:
205
- from ._request import predict_1k
206
- predictions = predict_1k(self.native, records, batch_size=8)
207
  if len(predictions) != len(records):
208
  raise RuntimeError("Incomplete model result; no partial answers returned")
209
  values = [[None] * len(group) for group in groups]
 
150
 
151
  evaluate(request) accepts the HTTP body shape; system_one(**request) is its
152
  Python equivalent. batch(requests) flattens independent states into GPU
153
+ batches and restores the original request/question order. Default B8 groups rows by decision type;
154
  batching='auto' opts into the published homogeneous padding-aware B32 path.
155
  """
156
  def __init__(self, native, *, model=None, batching="default"):
 
202
  from ._auto import predict_auto_1k
203
  predictions = predict_auto_1k(self.native, records)
204
  else:
205
+ from ._grouped import predict_grouped_1k
206
+ predictions = predict_grouped_1k(self.native, records, batch_size=8)
207
  if len(predictions) != len(records):
208
  raise RuntimeError("Incomplete model result; no partial answers returned")
209
  values = [[None] * len(group) for group in groups]