yasserrmd commited on
Commit
80dffdd
·
verified ·
1 Parent(s): 820a514

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +538 -102
README.md CHANGED
@@ -1,199 +1,635 @@
1
  ---
2
  library_name: transformers
3
- tags: []
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4
  ---
5
 
6
- # Model Card for Model ID
7
 
8
- <!-- Provide a quick summary of what the model is/does. -->
 
 
 
 
9
 
 
10
 
 
11
 
12
- ## Model Details
13
 
14
- ### Model Description
15
 
16
- <!-- Provide a longer summary of what this model is. -->
17
 
18
- This is the model card of a 🤗 transformers model that has been pushed on the Hub. This model card has been automatically generated.
19
 
20
- - **Developed by:** [More Information Needed]
21
- - **Funded by [optional]:** [More Information Needed]
22
- - **Shared by [optional]:** [More Information Needed]
23
- - **Model type:** [More Information Needed]
24
- - **Language(s) (NLP):** [More Information Needed]
25
- - **License:** [More Information Needed]
26
- - **Finetuned from model [optional]:** [More Information Needed]
27
 
28
- ### Model Sources [optional]
 
 
 
 
 
 
 
 
 
 
 
 
 
 
29
 
30
- <!-- Provide the basic links for the model. -->
31
 
32
- - **Repository:** [More Information Needed]
33
- - **Paper [optional]:** [More Information Needed]
34
- - **Demo [optional]:** [More Information Needed]
35
 
36
- ## Uses
37
 
38
- <!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
39
 
40
- ### Direct Use
41
 
42
- <!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
 
 
 
 
 
43
 
44
- [More Information Needed]
45
 
46
- ### Downstream Use [optional]
47
 
48
- <!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
 
 
 
 
 
 
 
 
 
 
49
 
50
- [More Information Needed]
51
 
52
- ### Out-of-Scope Use
53
 
54
- <!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
55
 
56
- [More Information Needed]
57
 
58
- ## Bias, Risks, and Limitations
 
 
 
 
 
 
 
 
 
 
59
 
60
- <!-- This section is meant to convey both technical and sociotechnical limitations. -->
61
 
62
- [More Information Needed]
 
 
 
 
 
 
 
63
 
64
- ### Recommendations
65
 
66
- <!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
67
 
68
- Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
69
 
70
- ## How to Get Started with the Model
71
 
72
- Use the code below to get started with the model.
 
 
73
 
74
- [More Information Needed]
75
 
76
- ## Training Details
77
 
78
- ### Training Data
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
79
 
80
- <!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
81
 
82
- [More Information Needed]
83
 
84
- ### Training Procedure
85
 
86
- <!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
87
 
88
- #### Preprocessing [optional]
 
 
 
 
 
 
 
 
89
 
90
- [More Information Needed]
91
 
 
 
 
 
 
 
 
 
 
92
 
93
- #### Training Hyperparameters
94
 
95
- - **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
 
 
 
 
 
 
 
 
96
 
97
- #### Speeds, Sizes, Times [optional]
98
 
99
- <!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
100
 
101
- [More Information Needed]
102
 
103
- ## Evaluation
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
104
 
105
- <!-- This section describes the evaluation protocols and provides the results. -->
 
 
 
 
 
 
 
 
106
 
107
- ### Testing Data, Factors & Metrics
108
 
109
- #### Testing Data
110
 
111
- <!-- This should link to a Dataset Card if possible. -->
112
 
113
- [More Information Needed]
 
 
 
 
 
 
114
 
115
- #### Factors
116
 
117
- <!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
118
 
119
- [More Information Needed]
120
 
121
- #### Metrics
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
122
 
123
- <!-- These are the evaluation metrics being used, ideally with a description of why. -->
124
 
125
- [More Information Needed]
126
 
127
- ### Results
128
 
129
- [More Information Needed]
130
 
131
- #### Summary
 
 
 
 
132
 
 
133
 
 
134
 
135
- ## Model Examination [optional]
136
 
137
- <!-- Relevant interpretability work for the model goes here -->
138
 
139
- [More Information Needed]
140
 
141
- ## Environmental Impact
142
 
143
- <!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
 
 
 
 
 
 
144
 
145
- Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
146
 
147
- - **Hardware Type:** [More Information Needed]
148
- - **Hours used:** [More Information Needed]
149
- - **Cloud Provider:** [More Information Needed]
150
- - **Compute Region:** [More Information Needed]
151
- - **Carbon Emitted:** [More Information Needed]
152
 
153
- ## Technical Specifications [optional]
154
 
155
- ### Model Architecture and Objective
156
 
157
- [More Information Needed]
158
 
159
- ### Compute Infrastructure
 
 
 
 
 
 
 
160
 
161
- [More Information Needed]
162
 
163
- #### Hardware
 
164
 
165
- [More Information Needed]
166
 
167
- #### Software
168
 
169
- [More Information Needed]
170
 
171
- ## Citation [optional]
172
 
173
- <!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
174
 
175
- **BibTeX:**
176
 
177
- [More Information Needed]
178
 
179
- **APA:**
180
 
181
- [More Information Needed]
182
 
183
- ## Glossary [optional]
184
 
185
- <!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
186
 
187
- [More Information Needed]
 
 
 
 
 
 
 
 
188
 
189
- ## More Information [optional]
 
 
190
 
191
- [More Information Needed]
192
 
193
- ## Model Card Authors [optional]
 
 
194
 
195
- [More Information Needed]
 
 
 
 
 
 
196
 
197
- ## Model Card Contact
198
 
199
- [More Information Needed]
 
 
1
  ---
2
  library_name: transformers
3
+ pipeline_tag: text-classification
4
+ language:
5
+ - en
6
+ base_model:
7
+ - answerdotai/ModernBERT-base
8
+ datasets:
9
+ - yasserrmd/enterprise-reflex-dataset
10
+ - yasserrmd/enterprise-reflex-v1-hard-dataset
11
+ tags:
12
+ - enterprise-ai
13
+ - agentic-ai
14
+ - system-1
15
+ - system-2
16
+ - action-ranking
17
+ - action-selection
18
+ - tool-selection
19
+ - dynamic-actions
20
+ - selective-prediction
21
+ - abstention
22
+ - no-action
23
+ - enterprise-agents
24
+ - modernbert
25
+ - workflow-routing
26
+ - decision-model
27
  ---
28
 
29
+ # Enterprise Reflex V1
30
 
31
+ **Model:** `yasserrmd/enterprise-reflex-v1`
32
+ **Base Model:** `answerdotai/ModernBERT-base`
33
+ **Language:** English
34
+ **Architecture:** Dynamic enterprise action scorer
35
+ **Status:** Research Prototype / Pre-Production Candidate
36
 
37
+ Enterprise Reflex V1 is a lightweight enterprise decision model designed to operate as a **System-1 action-ranking layer** in front of larger reasoning models, agent frameworks, and enterprise automation systems.
38
 
39
+ Given a request, structured enterprise state, optional context, and a dynamic runtime set of candidate actions, the model ranks the available actions, can abstain with `NO_ACTION`, and uses calibrated confidence to decide whether the request should stay in the fast System-1 path or be escalated to a more capable System-2 model or a human.
40
 
41
+ Enterprise Reflex is **not a generative LLM** and **not a fixed intent classifier**. Candidate actions are supplied dynamically at inference time, allowing the model to score actions it was not explicitly trained to recognize by identifier alone.
42
 
43
+ ---
44
 
45
+ ## Why Enterprise Reflex
46
 
47
+ Enterprise agent systems increasingly expose large action spaces: create or update records, approve or reject workflow steps, search internal systems, trigger notifications, route work, invoke APIs, execute business operations, or abstain when no safe action is available.
48
 
49
+ Sending every request directly to a large reasoning model can be unnecessarily expensive and slow. Enterprise Reflex is designed to act as a compact decision layer:
 
 
 
 
 
 
50
 
51
+ ```text
52
+ Request + Enterprise State + Context + Candidate Actions
53
+ |
54
+ v
55
+ Enterprise Reflex
56
+ |
57
+ +-----------+-----------+
58
+ | |
59
+ v v
60
+ Confident System-1 Low confidence / NO_ACTION
61
+ decision |
62
+ | v
63
+ v System-2 / Human Review
64
+ Execute
65
+ ```
66
 
67
+ The objective is not to replace larger reasoning models. It is to reduce how often they are required.
68
 
69
+ ---
 
 
70
 
71
+ ## What Changed in V1
72
 
73
+ V1 extends the original Enterprise Reflex prototype with targeted hard-training data focused on failure modes observed in V0.
74
 
75
+ The V1 hard-training set emphasizes:
76
 
77
+ - same-domain sibling-action discrimination,
78
+ - counterfactual state changes,
79
+ - policy and workflow constraints,
80
+ - semantic cross-domain collisions,
81
+ - improved `NO_ACTION` boundaries,
82
+ - unseen and renamed action identifiers.
83
 
84
+ The training strategy retained broad V0 enterprise coverage while adding targeted V1 hard examples so the model could improve difficult routing behavior without losing general enterprise performance.
85
 
86
+ ### Training Configuration
87
 
88
+ | Item | Value |
89
+ |---|---:|
90
+ | Base model | `answerdotai/ModernBERT-base` |
91
+ | Training samples | **553,632 pairwise examples** |
92
+ | Epochs | **2** |
93
+ | Total training steps | **23,068** |
94
+ | Warmup steps | **1,384** |
95
+ | Maximum sequence length | **384** |
96
+ | Objective | Binary request-action compatibility |
97
+ | Best-model metric | Validation F1 |
98
+ | Reported hardware | NVIDIA A100 |
99
 
100
+ ---
101
 
102
+ ## Input Representation
103
 
104
+ Enterprise Reflex scores each candidate action against a compact serialized request representation.
105
 
106
+ ### Request Side
107
 
108
+ ```json
109
+ {
110
+ "request": "Release the approved supplier payment",
111
+ "domain": "Finance",
112
+ "state": {
113
+ "payment_approved": true,
114
+ "invoice_matched": true
115
+ },
116
+ "context": {}
117
+ }
118
+ ```
119
 
120
+ ### Candidate Action Side
121
 
122
+ ```json
123
+ {
124
+ "name": "finance.release_payment",
125
+ "description": "Release an approved and validated supplier payment.",
126
+ "family": "EXECUTE",
127
+ "domain": "Finance"
128
+ }
129
+ ```
130
 
131
+ Each candidate is scored independently. Compatibility margins are then calibrated and normalized across the runtime candidate set. `NO_ACTION` is added as an explicit abstention candidate.
132
 
133
+ ---
134
 
135
+ ## Evaluation
136
 
137
+ V1 was evaluated at three levels:
138
 
139
+ 1. pairwise request-action classification,
140
+ 2. grouped candidate ranking,
141
+ 3. a manually designed 100-case hard stress test.
142
 
143
+ These evaluations represent different levels of difficulty and should be interpreted separately.
144
 
145
+ ---
146
 
147
+ ## Pairwise Test Results
148
+
149
+ ### V0 Test Distribution
150
+
151
+ | Metric | Result |
152
+ |---|---:|
153
+ | Accuracy | **97.86%** |
154
+ | Precision | **97.97%** |
155
+ | Recall | **83.69%** |
156
+ | F1 | **90.27%** |
157
+ | ROC AUC | **99.03%** |
158
+ | Average Precision | **94.85%** |
159
+
160
+ ### V1 Hard Test Distribution
161
+
162
+ | Metric | Result |
163
+ |---|---:|
164
+ | Accuracy | **94.07%** |
165
+ | Precision | **85.31%** |
166
+ | Recall | **79.00%** |
167
+ | F1 | **82.04%** |
168
+ | ROC AUC | **96.88%** |
169
+ | Average Precision | **89.24%** |
170
+
171
+ ### Combined Test Distribution
172
+
173
+ | Metric | Result |
174
+ |---|---:|
175
+ | Accuracy | **97.67%** |
176
+ | Precision | **97.01%** |
177
+ | Recall | **83.36%** |
178
+ | F1 | **89.67%** |
179
+ | ROC AUC | **98.94%** |
180
+ | Average Precision | **94.57%** |
181
 
182
+ ---
183
 
184
+ ## Grouped Action-Ranking Results
185
 
186
+ Grouped evaluation measures whether the correct action is ranked highest within the full runtime candidate set.
187
 
188
+ ### V0 Grouped Test
189
 
190
+ | Metric | Result |
191
+ |---|---:|
192
+ | Groups | **5,948** |
193
+ | Top-1 | **98.30%** |
194
+ | Top-3 | **99.98%** |
195
+ | MRR | **0.9913** |
196
+ | `NO_ACTION` Precision | **94.02%** |
197
+ | `NO_ACTION` Recall | **97.32%** |
198
+ | `NO_ACTION` F1 | **95.64%** |
199
 
200
+ ### V1 Hard Grouped Test
201
 
202
+ | Metric | Result |
203
+ |---|---:|
204
+ | Groups | **500** |
205
+ | Top-1 | **88.60%** |
206
+ | Top-3 | **99.40%** |
207
+ | MRR | **0.9370** |
208
+ | `NO_ACTION` Precision | **84.75%** |
209
+ | `NO_ACTION` Recall | **90.09%** |
210
+ | `NO_ACTION` F1 | **87.34%** |
211
 
212
+ ### Combined Grouped Test
213
 
214
+ | Metric | Result |
215
+ |---|---:|
216
+ | Groups | **6,448** |
217
+ | Top-1 | **97.55%** |
218
+ | Top-3 | **99.94%** |
219
+ | MRR | **0.9871** |
220
+ | `NO_ACTION` Precision | **93.05%** |
221
+ | `NO_ACTION` Recall | **96.58%** |
222
+ | `NO_ACTION` F1 | **94.78%** |
223
 
224
+ The V1 hard grouped split is intentionally more difficult than the broad V0 evaluation and should not be treated as the same distribution.
225
 
226
+ ---
227
 
228
+ ## Selective System-1 / System-2 Routing
229
 
230
+ V1 uses calibrated confidence to determine whether a decision should remain in System-1 or be escalated.
231
+
232
+ For the reported run:
233
+
234
+ - **Selected System-2 threshold:** `0.87`
235
+ - **Validation System-1 coverage:** **73.71%**
236
+ - **Validation System-1 accuracy:** **99.02%**
237
+
238
+ `NO_ACTION` is always treated as a System-2 route.
239
+
240
+ This threshold is calibrated on the validation distribution and should be recalibrated for any materially different deployment domain.
241
+
242
+ ---
243
+
244
+ ## 100-Case Manual Hard Stress Test
245
+
246
+ A separate manual suite of **100 hard enterprise cases** was used to stress behavior outside the easier validation distribution.
247
+
248
+ ### Overall Results
249
+
250
+ | Metric | Result |
251
+ |---|---:|
252
+ | Total cases | **100** |
253
+ | Correct | **85** |
254
+ | Incorrect | **15** |
255
+ | Raw Top-1 accuracy | **85.00%** |
256
+ | System-1 handled | **66%** |
257
+ | System-2 routed | **34%** |
258
+ | System-1 accuracy | **87.88%** |
259
+ | Unsafe System-1 failures | **8** |
260
+
261
+ ### Performance by Category
262
+
263
+ | Category | Tests | Accuracy |
264
+ |---|---:|---:|
265
+ | Cross-domain | 12 | **83.33%** |
266
+ | `NO_ACTION` | 21 | **100.00%** |
267
+ | Policy constraint | 8 | **12.50%** |
268
+ | Sibling action | 32 | **87.50%** |
269
+ | State sensitive | 22 | **90.91%** |
270
+ | Unseen action name | 5 | **100.00%** |
271
+
272
+ The manual stress test is intentionally adversarial and significantly harder than the standard grouped benchmark.
273
+
274
+ ---
275
+
276
+ ## What V1 Improved
277
+
278
+ Compared with V0 hard-test behavior, V1 improved both hard-decision accuracy and autonomous coverage.
279
+
280
+ Observed improvements include:
281
+
282
+ - hard Top-1 accuracy increased from approximately **80% to 85%**,
283
+ - System-1 coverage increased from approximately **53% to 66%**,
284
+ - state-sensitive decisions improved substantially,
285
+ - unseen action-name generalization remained strong,
286
+ - `NO_ACTION` behavior improved on the manual hard suite,
287
+ - broad V0 enterprise ranking performance remained largely intact.
288
+
289
+ The result supports the core Enterprise Reflex design: a lightweight model can perform useful dynamic enterprise action ranking while routing uncertain cases to a larger reasoner.
290
+
291
+ ---
292
+
293
+ ## Current Limitation: Policy-Constrained Execution
294
+
295
+ The dominant V1 weakness is policy-sensitive action validity.
296
+
297
+ Several hard cases were semantically understood but executed incorrectly because state or policy should have blocked the action.
298
+
299
+ Observed failure patterns include:
300
+
301
+ - releasing a payment without required approval,
302
+ - provisioning privileged access without security approval,
303
+ - deleting logs under legal or retention hold,
304
+ - cancelling an order after a workflow state that prohibits cancellation,
305
+ - granting physical access before mandatory induction is complete.
306
+
307
+ This indicates that V1 is currently stronger at answering:
308
+
309
+ > Which action best matches this request?
310
+
311
+ than:
312
+
313
+ > Is this action actually permitted under the current enterprise state and policy?
314
+
315
+ For this reason, V1 should not be used as the sole authority for autonomous high-impact enterprise execution.
316
+
317
+ ---
318
+
319
+ ## Intended Use
320
+
321
+ Enterprise Reflex V1 is suitable for research and controlled enterprise-agent experiments such as:
322
 
323
+ - action ranking,
324
+ - dynamic tool selection,
325
+ - top-k tool narrowing,
326
+ - System-1 / System-2 routing,
327
+ - agent handoff,
328
+ - workflow recommendation,
329
+ - shadow-mode decision analysis,
330
+ - human-in-the-loop action suggestions,
331
+ - enterprise action-space reduction before LLM reasoning.
332
 
333
+ ---
334
 
335
+ ## Not Recommended For
336
 
337
+ V1 is not recommended as the sole decision layer for:
338
 
339
+ - autonomous financial transactions,
340
+ - privileged-access provisioning,
341
+ - destructive security actions,
342
+ - compliance-sensitive deletion,
343
+ - irreversible workflow actions,
344
+ - legal or regulatory decisions,
345
+ - production execution without deterministic authorization and policy enforcement.
346
 
347
+ ---
348
 
349
+ ## Recommended Production Architecture
350
+
351
+ Enterprise Reflex should be combined with deterministic controls.
352
+
353
+ ```text
354
+ Request + State + Candidate Actions
355
+ |
356
+ v
357
+ Enterprise Reflex
358
+ |
359
+ v
360
+ Ranked Action
361
+ |
362
+ v
363
+ Policy / Authorization Engine
364
+ / \
365
+ Allowed Blocked
366
+ | |
367
+ v v
368
+ Execute System-2 / Human
369
+ ```
370
+
371
+ The learned model provides decision intelligence. Authorization, policy, entitlement, retention, approval, and other hard enterprise controls should remain deterministic whenever possible.
372
 
373
+ ---
374
 
375
+ ## Example Usage
376
+
377
+ ```python
378
+ import json
379
+ import torch
380
+ import numpy as np
381
+
382
+ from transformers import (
383
+ AutoTokenizer,
384
+ AutoModelForSequenceClassification,
385
+ )
386
+
387
+ MODEL_ID = "yasserrmd/enterprise-reflex-v1"
388
+ SYSTEM2_THRESHOLD = 0.87
389
+ MAX_LENGTH = 384
390
+
391
+ # Load the runtime calibration values saved with the model if available.
392
+ # The reported experiment used a calibrated temperature determined from
393
+ # the combined validation distribution.
394
+ TEMPERATURE = 1.0
395
+
396
+ tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
397
+ model = AutoModelForSequenceClassification.from_pretrained(MODEL_ID)
398
+
399
+ device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
400
+ model.to(device)
401
+ model.eval()
402
+
403
+
404
+ def softmax_np(x):
405
+ x = np.asarray(x, dtype=np.float64)
406
+ x = x - np.max(x)
407
+ e = np.exp(x)
408
+ return e / e.sum()
409
+
410
+
411
+ def request_text(request, domain="enterprise", state=None, context=None):
412
+ return json.dumps(
413
+ {
414
+ "request": request,
415
+ "domain": domain,
416
+ "state": state or {},
417
+ "context": context or {},
418
+ },
419
+ ensure_ascii=False,
420
+ sort_keys=True,
421
+ )
422
+
423
+
424
+ def action_text(action, default_domain="enterprise"):
425
+ return json.dumps(
426
+ {
427
+ "name": action.get("name", ""),
428
+ "description": action.get("description", ""),
429
+ "family": action.get("family", "OTHER"),
430
+ "domain": action.get("domain", default_domain),
431
+ },
432
+ ensure_ascii=False,
433
+ sort_keys=True,
434
+ )
435
+
436
+
437
+ def rank_actions(request, candidate_actions, state=None, context=None, domain="enterprise"):
438
+ actions = list(candidate_actions)
439
+
440
+ actions.append(
441
+ {
442
+ "name": "NO_ACTION",
443
+ "description": "None of the available actions safely or correctly satisfy the request.",
444
+ "family": "ABSTAIN",
445
+ "domain": domain,
446
+ }
447
+ )
448
+
449
+ left = request_text(
450
+ request,
451
+ domain=domain,
452
+ state=state,
453
+ context=context,
454
+ )
455
+
456
+ enc = tokenizer(
457
+ [left] * len(actions),
458
+ [action_text(a, default_domain=domain) for a in actions],
459
+ padding=True,
460
+ truncation=True,
461
+ max_length=MAX_LENGTH,
462
+ return_tensors="pt",
463
+ ).to(device)
464
+
465
+ with torch.no_grad():
466
+ logits = model(**enc).logits
467
+
468
+ margins = (logits[:, 1] - logits[:, 0]).float().cpu().numpy()
469
+ probabilities = softmax_np(margins / TEMPERATURE)
470
+ order = np.argsort(-probabilities)
471
+
472
+ ranked = [
473
+ {
474
+ "action": actions[int(i)]["name"],
475
+ "probability": float(probabilities[int(i)]),
476
+ }
477
+ for i in order
478
+ ]
479
+
480
+ top = ranked[0]
481
+
482
+ system2_required = (
483
+ top["action"] == "NO_ACTION"
484
+ or top["probability"] < SYSTEM2_THRESHOLD
485
+ )
486
+
487
+ return {
488
+ "decision": top["action"],
489
+ "confidence": top["probability"],
490
+ "system2_required": system2_required,
491
+ "ranked_actions": ranked,
492
+ }
493
+ ```
494
+
495
+ ### Example
496
+
497
+ ```python
498
+ result = rank_actions(
499
+ request="Release the supplier payment",
500
+ domain="Finance",
501
+ state={
502
+ "invoice_matched": True,
503
+ "payment_approved": True,
504
+ },
505
+ candidate_actions=[
506
+ {
507
+ "name": "finance.release_payment",
508
+ "description": "Release an approved and validated supplier payment.",
509
+ "family": "EXECUTE",
510
+ "domain": "Finance",
511
+ },
512
+ {
513
+ "name": "finance.create_invoice",
514
+ "description": "Create an invoice record.",
515
+ "family": "CREATE",
516
+ "domain": "Finance",
517
+ },
518
+ ],
519
+ )
520
+
521
+ print(result)
522
+ ```
523
 
524
+ ---
525
 
526
+ ## Calibration Note
527
 
528
+ The reported confidence threshold was selected for the reported V1 experiment.
529
 
530
+ For a new deployment:
531
 
532
+ 1. collect domain-specific validation data,
533
+ 2. calibrate temperature,
534
+ 3. determine an acceptable System-1 error rate,
535
+ 4. select the confidence threshold for that environment,
536
+ 5. validate policy-sensitive and destructive actions separately.
537
 
538
+ Do not assume that `0.87` is appropriate for every enterprise domain.
539
 
540
+ ---
541
 
542
+ ## Research Status
543
 
544
+ Enterprise Reflex V1 should currently be considered a:
545
 
546
+ **Research Prototype / Pre-Production Candidate**
547
 
548
+ The model has demonstrated:
549
 
550
+ - strong broad enterprise action ranking,
551
+ - useful dynamic action selection,
552
+ - effective abstention,
553
+ - high Top-3 retrieval quality,
554
+ - improved state-sensitive behavior,
555
+ - improved System-1 coverage,
556
+ - promising generalization to unseen action identifiers.
557
 
558
+ It has not yet demonstrated sufficient reliability for unrestricted autonomous enterprise execution.
559
 
560
+ The primary V2 research target is **policy-aware action validity and confident wrong-action suppression**.
 
 
 
 
561
 
562
+ ---
563
 
564
+ ## V2 Direction
565
 
566
+ The next iteration should focus less on generic enterprise volume and more on targeted safety and state-validity examples:
567
 
568
+ - policy counterfactuals,
569
+ - approval-sensitive actions,
570
+ - retention and legal-hold constraints,
571
+ - authorization and entitlement state,
572
+ - workflow-state legality,
573
+ - semantic domain collisions,
574
+ - hard negatives mined from confident V1 failures,
575
+ - valid-action vs `NO_ACTION` boundary cases.
576
 
577
+ A likely architectural extension is to separate:
578
 
579
+ 1. semantic action suitability,
580
+ 2. state/policy validity,
581
 
582
+ before producing the final action confidence.
583
 
584
+ ---
585
 
586
+ ## Datasets
587
 
588
+ ### Enterprise Reflex Dataset
589
 
590
+ `yasserrmd/enterprise-reflex-dataset`
591
 
592
+ Broad enterprise action-ranking data used for V0 and retained in V1 training.
593
 
594
+ ### Enterprise Reflex V1 Hard Dataset
595
 
596
+ `yasserrmd/enterprise-reflex-v1-hard-dataset`
597
 
598
+ Targeted hard-training examples covering sibling actions, counterfactual state, policy constraints, cross-domain collisions, `NO_ACTION`, and unseen action names.
599
 
600
+ ---
601
 
602
+ ## Limitations
603
 
604
+ - English only in V1.
605
+ - Text and structured state only.
606
+ - No multimodal input.
607
+ - No deterministic policy engine is embedded in the model.
608
+ - Confidence calibration is distribution-dependent.
609
+ - The hard-test suite is manually constructed and relatively small.
610
+ - Policy-sensitive action validity remains the main weakness.
611
+ - The model may still produce high-confidence incorrect actions.
612
+ - Reported metrics should not be interpreted as production-safety guarantees.
613
 
614
+ ---
615
+
616
+ ## Responsible Use
617
 
618
+ Enterprise Reflex is intended to assist enterprise decision routing, not replace enterprise authorization, policy, compliance, or human accountability.
619
 
620
+ High-impact actions should remain subject to deterministic controls and appropriate human or System-2 review.
621
+
622
+ ---
623
 
624
+ ## Author
625
+
626
+ **Mohamed Yasser**
627
+ Hugging Face: `yasserrmd`
628
+ GitHub: `yasserrmd`
629
+
630
+ ---
631
 
632
+ ## Version
633
 
634
+ **Enterprise Reflex V1**
635
+ Research Prototype / Pre-Production Candidate