Text Classification
Transformers
Safetensors
English
modernbert
enterprise-ai
agentic-ai
system-1
system-2
action-ranking
action-selection
tool-selection
dynamic-actions
selective-prediction
abstention
no-action
enterprise-agents
workflow-routing
decision-model
text-embeddings-inference
Instructions to use yasserrmd/enterprise-reflex-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use yasserrmd/enterprise-reflex-v1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="yasserrmd/enterprise-reflex-v1")# Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("yasserrmd/enterprise-reflex-v1") model = AutoModelForSequenceClassification.from_pretrained("yasserrmd/enterprise-reflex-v1", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Update README.md
Browse files
README.md
CHANGED
|
@@ -1,199 +1,635 @@
|
|
| 1 |
---
|
| 2 |
library_name: transformers
|
| 3 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 4 |
---
|
| 5 |
|
| 6 |
-
#
|
| 7 |
|
| 8 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 9 |
|
|
|
|
| 10 |
|
|
|
|
| 11 |
|
| 12 |
-
|
| 13 |
|
| 14 |
-
|
| 15 |
|
| 16 |
-
|
| 17 |
|
| 18 |
-
|
| 19 |
|
| 20 |
-
|
| 21 |
-
- **Funded by [optional]:** [More Information Needed]
|
| 22 |
-
- **Shared by [optional]:** [More Information Needed]
|
| 23 |
-
- **Model type:** [More Information Needed]
|
| 24 |
-
- **Language(s) (NLP):** [More Information Needed]
|
| 25 |
-
- **License:** [More Information Needed]
|
| 26 |
-
- **Finetuned from model [optional]:** [More Information Needed]
|
| 27 |
|
| 28 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 29 |
|
| 30 |
-
|
| 31 |
|
| 32 |
-
-
|
| 33 |
-
- **Paper [optional]:** [More Information Needed]
|
| 34 |
-
- **Demo [optional]:** [More Information Needed]
|
| 35 |
|
| 36 |
-
##
|
| 37 |
|
| 38 |
-
|
| 39 |
|
| 40 |
-
|
| 41 |
|
| 42 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 43 |
|
| 44 |
-
|
| 45 |
|
| 46 |
-
###
|
| 47 |
|
| 48 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 49 |
|
| 50 |
-
|
| 51 |
|
| 52 |
-
##
|
| 53 |
|
| 54 |
-
|
| 55 |
|
| 56 |
-
|
| 57 |
|
| 58 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 59 |
|
| 60 |
-
|
| 61 |
|
| 62 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 63 |
|
| 64 |
-
|
| 65 |
|
| 66 |
-
|
| 67 |
|
| 68 |
-
|
| 69 |
|
| 70 |
-
|
| 71 |
|
| 72 |
-
|
|
|
|
|
|
|
| 73 |
|
| 74 |
-
|
| 75 |
|
| 76 |
-
|
| 77 |
|
| 78 |
-
##
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 79 |
|
| 80 |
-
|
| 81 |
|
| 82 |
-
|
| 83 |
|
| 84 |
-
|
| 85 |
|
| 86 |
-
|
| 87 |
|
| 88 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 89 |
|
| 90 |
-
|
| 91 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 92 |
|
| 93 |
-
###
|
| 94 |
|
| 95 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 96 |
|
| 97 |
-
|
| 98 |
|
| 99 |
-
|
| 100 |
|
| 101 |
-
|
| 102 |
|
| 103 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 104 |
|
| 105 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 106 |
|
| 107 |
-
|
| 108 |
|
| 109 |
-
##
|
| 110 |
|
| 111 |
-
|
| 112 |
|
| 113 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 114 |
|
| 115 |
-
|
| 116 |
|
| 117 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 118 |
|
| 119 |
-
|
| 120 |
|
| 121 |
-
##
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 122 |
|
| 123 |
-
|
| 124 |
|
| 125 |
-
|
| 126 |
|
| 127 |
-
|
| 128 |
|
| 129 |
-
|
| 130 |
|
| 131 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 132 |
|
|
|
|
| 133 |
|
|
|
|
| 134 |
|
| 135 |
-
##
|
| 136 |
|
| 137 |
-
|
| 138 |
|
| 139 |
-
|
| 140 |
|
| 141 |
-
|
| 142 |
|
| 143 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 144 |
|
| 145 |
-
|
| 146 |
|
| 147 |
-
|
| 148 |
-
- **Hours used:** [More Information Needed]
|
| 149 |
-
- **Cloud Provider:** [More Information Needed]
|
| 150 |
-
- **Compute Region:** [More Information Needed]
|
| 151 |
-
- **Carbon Emitted:** [More Information Needed]
|
| 152 |
|
| 153 |
-
|
| 154 |
|
| 155 |
-
##
|
| 156 |
|
| 157 |
-
|
| 158 |
|
| 159 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 160 |
|
| 161 |
-
|
| 162 |
|
| 163 |
-
|
|
|
|
| 164 |
|
| 165 |
-
|
| 166 |
|
| 167 |
-
|
| 168 |
|
| 169 |
-
|
| 170 |
|
| 171 |
-
##
|
| 172 |
|
| 173 |
-
|
| 174 |
|
| 175 |
-
|
| 176 |
|
| 177 |
-
|
| 178 |
|
| 179 |
-
|
| 180 |
|
| 181 |
-
|
| 182 |
|
| 183 |
-
|
| 184 |
|
| 185 |
-
|
| 186 |
|
| 187 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 188 |
|
| 189 |
-
|
|
|
|
|
|
|
| 190 |
|
| 191 |
-
|
| 192 |
|
| 193 |
-
|
|
|
|
|
|
|
| 194 |
|
| 195 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 196 |
|
| 197 |
-
##
|
| 198 |
|
| 199 |
-
|
|
|
|
|
|
| 1 |
---
|
| 2 |
library_name: transformers
|
| 3 |
+
pipeline_tag: text-classification
|
| 4 |
+
language:
|
| 5 |
+
- en
|
| 6 |
+
base_model:
|
| 7 |
+
- answerdotai/ModernBERT-base
|
| 8 |
+
datasets:
|
| 9 |
+
- yasserrmd/enterprise-reflex-dataset
|
| 10 |
+
- yasserrmd/enterprise-reflex-v1-hard-dataset
|
| 11 |
+
tags:
|
| 12 |
+
- enterprise-ai
|
| 13 |
+
- agentic-ai
|
| 14 |
+
- system-1
|
| 15 |
+
- system-2
|
| 16 |
+
- action-ranking
|
| 17 |
+
- action-selection
|
| 18 |
+
- tool-selection
|
| 19 |
+
- dynamic-actions
|
| 20 |
+
- selective-prediction
|
| 21 |
+
- abstention
|
| 22 |
+
- no-action
|
| 23 |
+
- enterprise-agents
|
| 24 |
+
- modernbert
|
| 25 |
+
- workflow-routing
|
| 26 |
+
- decision-model
|
| 27 |
---
|
| 28 |
|
| 29 |
+
# Enterprise Reflex V1
|
| 30 |
|
| 31 |
+
**Model:** `yasserrmd/enterprise-reflex-v1`
|
| 32 |
+
**Base Model:** `answerdotai/ModernBERT-base`
|
| 33 |
+
**Language:** English
|
| 34 |
+
**Architecture:** Dynamic enterprise action scorer
|
| 35 |
+
**Status:** Research Prototype / Pre-Production Candidate
|
| 36 |
|
| 37 |
+
Enterprise Reflex V1 is a lightweight enterprise decision model designed to operate as a **System-1 action-ranking layer** in front of larger reasoning models, agent frameworks, and enterprise automation systems.
|
| 38 |
|
| 39 |
+
Given a request, structured enterprise state, optional context, and a dynamic runtime set of candidate actions, the model ranks the available actions, can abstain with `NO_ACTION`, and uses calibrated confidence to decide whether the request should stay in the fast System-1 path or be escalated to a more capable System-2 model or a human.
|
| 40 |
|
| 41 |
+
Enterprise Reflex is **not a generative LLM** and **not a fixed intent classifier**. Candidate actions are supplied dynamically at inference time, allowing the model to score actions it was not explicitly trained to recognize by identifier alone.
|
| 42 |
|
| 43 |
+
---
|
| 44 |
|
| 45 |
+
## Why Enterprise Reflex
|
| 46 |
|
| 47 |
+
Enterprise agent systems increasingly expose large action spaces: create or update records, approve or reject workflow steps, search internal systems, trigger notifications, route work, invoke APIs, execute business operations, or abstain when no safe action is available.
|
| 48 |
|
| 49 |
+
Sending every request directly to a large reasoning model can be unnecessarily expensive and slow. Enterprise Reflex is designed to act as a compact decision layer:
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 50 |
|
| 51 |
+
```text
|
| 52 |
+
Request + Enterprise State + Context + Candidate Actions
|
| 53 |
+
|
|
| 54 |
+
v
|
| 55 |
+
Enterprise Reflex
|
| 56 |
+
|
|
| 57 |
+
+-----------+-----------+
|
| 58 |
+
| |
|
| 59 |
+
v v
|
| 60 |
+
Confident System-1 Low confidence / NO_ACTION
|
| 61 |
+
decision |
|
| 62 |
+
| v
|
| 63 |
+
v System-2 / Human Review
|
| 64 |
+
Execute
|
| 65 |
+
```
|
| 66 |
|
| 67 |
+
The objective is not to replace larger reasoning models. It is to reduce how often they are required.
|
| 68 |
|
| 69 |
+
---
|
|
|
|
|
|
|
| 70 |
|
| 71 |
+
## What Changed in V1
|
| 72 |
|
| 73 |
+
V1 extends the original Enterprise Reflex prototype with targeted hard-training data focused on failure modes observed in V0.
|
| 74 |
|
| 75 |
+
The V1 hard-training set emphasizes:
|
| 76 |
|
| 77 |
+
- same-domain sibling-action discrimination,
|
| 78 |
+
- counterfactual state changes,
|
| 79 |
+
- policy and workflow constraints,
|
| 80 |
+
- semantic cross-domain collisions,
|
| 81 |
+
- improved `NO_ACTION` boundaries,
|
| 82 |
+
- unseen and renamed action identifiers.
|
| 83 |
|
| 84 |
+
The training strategy retained broad V0 enterprise coverage while adding targeted V1 hard examples so the model could improve difficult routing behavior without losing general enterprise performance.
|
| 85 |
|
| 86 |
+
### Training Configuration
|
| 87 |
|
| 88 |
+
| Item | Value |
|
| 89 |
+
|---|---:|
|
| 90 |
+
| Base model | `answerdotai/ModernBERT-base` |
|
| 91 |
+
| Training samples | **553,632 pairwise examples** |
|
| 92 |
+
| Epochs | **2** |
|
| 93 |
+
| Total training steps | **23,068** |
|
| 94 |
+
| Warmup steps | **1,384** |
|
| 95 |
+
| Maximum sequence length | **384** |
|
| 96 |
+
| Objective | Binary request-action compatibility |
|
| 97 |
+
| Best-model metric | Validation F1 |
|
| 98 |
+
| Reported hardware | NVIDIA A100 |
|
| 99 |
|
| 100 |
+
---
|
| 101 |
|
| 102 |
+
## Input Representation
|
| 103 |
|
| 104 |
+
Enterprise Reflex scores each candidate action against a compact serialized request representation.
|
| 105 |
|
| 106 |
+
### Request Side
|
| 107 |
|
| 108 |
+
```json
|
| 109 |
+
{
|
| 110 |
+
"request": "Release the approved supplier payment",
|
| 111 |
+
"domain": "Finance",
|
| 112 |
+
"state": {
|
| 113 |
+
"payment_approved": true,
|
| 114 |
+
"invoice_matched": true
|
| 115 |
+
},
|
| 116 |
+
"context": {}
|
| 117 |
+
}
|
| 118 |
+
```
|
| 119 |
|
| 120 |
+
### Candidate Action Side
|
| 121 |
|
| 122 |
+
```json
|
| 123 |
+
{
|
| 124 |
+
"name": "finance.release_payment",
|
| 125 |
+
"description": "Release an approved and validated supplier payment.",
|
| 126 |
+
"family": "EXECUTE",
|
| 127 |
+
"domain": "Finance"
|
| 128 |
+
}
|
| 129 |
+
```
|
| 130 |
|
| 131 |
+
Each candidate is scored independently. Compatibility margins are then calibrated and normalized across the runtime candidate set. `NO_ACTION` is added as an explicit abstention candidate.
|
| 132 |
|
| 133 |
+
---
|
| 134 |
|
| 135 |
+
## Evaluation
|
| 136 |
|
| 137 |
+
V1 was evaluated at three levels:
|
| 138 |
|
| 139 |
+
1. pairwise request-action classification,
|
| 140 |
+
2. grouped candidate ranking,
|
| 141 |
+
3. a manually designed 100-case hard stress test.
|
| 142 |
|
| 143 |
+
These evaluations represent different levels of difficulty and should be interpreted separately.
|
| 144 |
|
| 145 |
+
---
|
| 146 |
|
| 147 |
+
## Pairwise Test Results
|
| 148 |
+
|
| 149 |
+
### V0 Test Distribution
|
| 150 |
+
|
| 151 |
+
| Metric | Result |
|
| 152 |
+
|---|---:|
|
| 153 |
+
| Accuracy | **97.86%** |
|
| 154 |
+
| Precision | **97.97%** |
|
| 155 |
+
| Recall | **83.69%** |
|
| 156 |
+
| F1 | **90.27%** |
|
| 157 |
+
| ROC AUC | **99.03%** |
|
| 158 |
+
| Average Precision | **94.85%** |
|
| 159 |
+
|
| 160 |
+
### V1 Hard Test Distribution
|
| 161 |
+
|
| 162 |
+
| Metric | Result |
|
| 163 |
+
|---|---:|
|
| 164 |
+
| Accuracy | **94.07%** |
|
| 165 |
+
| Precision | **85.31%** |
|
| 166 |
+
| Recall | **79.00%** |
|
| 167 |
+
| F1 | **82.04%** |
|
| 168 |
+
| ROC AUC | **96.88%** |
|
| 169 |
+
| Average Precision | **89.24%** |
|
| 170 |
+
|
| 171 |
+
### Combined Test Distribution
|
| 172 |
+
|
| 173 |
+
| Metric | Result |
|
| 174 |
+
|---|---:|
|
| 175 |
+
| Accuracy | **97.67%** |
|
| 176 |
+
| Precision | **97.01%** |
|
| 177 |
+
| Recall | **83.36%** |
|
| 178 |
+
| F1 | **89.67%** |
|
| 179 |
+
| ROC AUC | **98.94%** |
|
| 180 |
+
| Average Precision | **94.57%** |
|
| 181 |
|
| 182 |
+
---
|
| 183 |
|
| 184 |
+
## Grouped Action-Ranking Results
|
| 185 |
|
| 186 |
+
Grouped evaluation measures whether the correct action is ranked highest within the full runtime candidate set.
|
| 187 |
|
| 188 |
+
### V0 Grouped Test
|
| 189 |
|
| 190 |
+
| Metric | Result |
|
| 191 |
+
|---|---:|
|
| 192 |
+
| Groups | **5,948** |
|
| 193 |
+
| Top-1 | **98.30%** |
|
| 194 |
+
| Top-3 | **99.98%** |
|
| 195 |
+
| MRR | **0.9913** |
|
| 196 |
+
| `NO_ACTION` Precision | **94.02%** |
|
| 197 |
+
| `NO_ACTION` Recall | **97.32%** |
|
| 198 |
+
| `NO_ACTION` F1 | **95.64%** |
|
| 199 |
|
| 200 |
+
### V1 Hard Grouped Test
|
| 201 |
|
| 202 |
+
| Metric | Result |
|
| 203 |
+
|---|---:|
|
| 204 |
+
| Groups | **500** |
|
| 205 |
+
| Top-1 | **88.60%** |
|
| 206 |
+
| Top-3 | **99.40%** |
|
| 207 |
+
| MRR | **0.9370** |
|
| 208 |
+
| `NO_ACTION` Precision | **84.75%** |
|
| 209 |
+
| `NO_ACTION` Recall | **90.09%** |
|
| 210 |
+
| `NO_ACTION` F1 | **87.34%** |
|
| 211 |
|
| 212 |
+
### Combined Grouped Test
|
| 213 |
|
| 214 |
+
| Metric | Result |
|
| 215 |
+
|---|---:|
|
| 216 |
+
| Groups | **6,448** |
|
| 217 |
+
| Top-1 | **97.55%** |
|
| 218 |
+
| Top-3 | **99.94%** |
|
| 219 |
+
| MRR | **0.9871** |
|
| 220 |
+
| `NO_ACTION` Precision | **93.05%** |
|
| 221 |
+
| `NO_ACTION` Recall | **96.58%** |
|
| 222 |
+
| `NO_ACTION` F1 | **94.78%** |
|
| 223 |
|
| 224 |
+
The V1 hard grouped split is intentionally more difficult than the broad V0 evaluation and should not be treated as the same distribution.
|
| 225 |
|
| 226 |
+
---
|
| 227 |
|
| 228 |
+
## Selective System-1 / System-2 Routing
|
| 229 |
|
| 230 |
+
V1 uses calibrated confidence to determine whether a decision should remain in System-1 or be escalated.
|
| 231 |
+
|
| 232 |
+
For the reported run:
|
| 233 |
+
|
| 234 |
+
- **Selected System-2 threshold:** `0.87`
|
| 235 |
+
- **Validation System-1 coverage:** **73.71%**
|
| 236 |
+
- **Validation System-1 accuracy:** **99.02%**
|
| 237 |
+
|
| 238 |
+
`NO_ACTION` is always treated as a System-2 route.
|
| 239 |
+
|
| 240 |
+
This threshold is calibrated on the validation distribution and should be recalibrated for any materially different deployment domain.
|
| 241 |
+
|
| 242 |
+
---
|
| 243 |
+
|
| 244 |
+
## 100-Case Manual Hard Stress Test
|
| 245 |
+
|
| 246 |
+
A separate manual suite of **100 hard enterprise cases** was used to stress behavior outside the easier validation distribution.
|
| 247 |
+
|
| 248 |
+
### Overall Results
|
| 249 |
+
|
| 250 |
+
| Metric | Result |
|
| 251 |
+
|---|---:|
|
| 252 |
+
| Total cases | **100** |
|
| 253 |
+
| Correct | **85** |
|
| 254 |
+
| Incorrect | **15** |
|
| 255 |
+
| Raw Top-1 accuracy | **85.00%** |
|
| 256 |
+
| System-1 handled | **66%** |
|
| 257 |
+
| System-2 routed | **34%** |
|
| 258 |
+
| System-1 accuracy | **87.88%** |
|
| 259 |
+
| Unsafe System-1 failures | **8** |
|
| 260 |
+
|
| 261 |
+
### Performance by Category
|
| 262 |
+
|
| 263 |
+
| Category | Tests | Accuracy |
|
| 264 |
+
|---|---:|---:|
|
| 265 |
+
| Cross-domain | 12 | **83.33%** |
|
| 266 |
+
| `NO_ACTION` | 21 | **100.00%** |
|
| 267 |
+
| Policy constraint | 8 | **12.50%** |
|
| 268 |
+
| Sibling action | 32 | **87.50%** |
|
| 269 |
+
| State sensitive | 22 | **90.91%** |
|
| 270 |
+
| Unseen action name | 5 | **100.00%** |
|
| 271 |
+
|
| 272 |
+
The manual stress test is intentionally adversarial and significantly harder than the standard grouped benchmark.
|
| 273 |
+
|
| 274 |
+
---
|
| 275 |
+
|
| 276 |
+
## What V1 Improved
|
| 277 |
+
|
| 278 |
+
Compared with V0 hard-test behavior, V1 improved both hard-decision accuracy and autonomous coverage.
|
| 279 |
+
|
| 280 |
+
Observed improvements include:
|
| 281 |
+
|
| 282 |
+
- hard Top-1 accuracy increased from approximately **80% to 85%**,
|
| 283 |
+
- System-1 coverage increased from approximately **53% to 66%**,
|
| 284 |
+
- state-sensitive decisions improved substantially,
|
| 285 |
+
- unseen action-name generalization remained strong,
|
| 286 |
+
- `NO_ACTION` behavior improved on the manual hard suite,
|
| 287 |
+
- broad V0 enterprise ranking performance remained largely intact.
|
| 288 |
+
|
| 289 |
+
The result supports the core Enterprise Reflex design: a lightweight model can perform useful dynamic enterprise action ranking while routing uncertain cases to a larger reasoner.
|
| 290 |
+
|
| 291 |
+
---
|
| 292 |
+
|
| 293 |
+
## Current Limitation: Policy-Constrained Execution
|
| 294 |
+
|
| 295 |
+
The dominant V1 weakness is policy-sensitive action validity.
|
| 296 |
+
|
| 297 |
+
Several hard cases were semantically understood but executed incorrectly because state or policy should have blocked the action.
|
| 298 |
+
|
| 299 |
+
Observed failure patterns include:
|
| 300 |
+
|
| 301 |
+
- releasing a payment without required approval,
|
| 302 |
+
- provisioning privileged access without security approval,
|
| 303 |
+
- deleting logs under legal or retention hold,
|
| 304 |
+
- cancelling an order after a workflow state that prohibits cancellation,
|
| 305 |
+
- granting physical access before mandatory induction is complete.
|
| 306 |
+
|
| 307 |
+
This indicates that V1 is currently stronger at answering:
|
| 308 |
+
|
| 309 |
+
> Which action best matches this request?
|
| 310 |
+
|
| 311 |
+
than:
|
| 312 |
+
|
| 313 |
+
> Is this action actually permitted under the current enterprise state and policy?
|
| 314 |
+
|
| 315 |
+
For this reason, V1 should not be used as the sole authority for autonomous high-impact enterprise execution.
|
| 316 |
+
|
| 317 |
+
---
|
| 318 |
+
|
| 319 |
+
## Intended Use
|
| 320 |
+
|
| 321 |
+
Enterprise Reflex V1 is suitable for research and controlled enterprise-agent experiments such as:
|
| 322 |
|
| 323 |
+
- action ranking,
|
| 324 |
+
- dynamic tool selection,
|
| 325 |
+
- top-k tool narrowing,
|
| 326 |
+
- System-1 / System-2 routing,
|
| 327 |
+
- agent handoff,
|
| 328 |
+
- workflow recommendation,
|
| 329 |
+
- shadow-mode decision analysis,
|
| 330 |
+
- human-in-the-loop action suggestions,
|
| 331 |
+
- enterprise action-space reduction before LLM reasoning.
|
| 332 |
|
| 333 |
+
---
|
| 334 |
|
| 335 |
+
## Not Recommended For
|
| 336 |
|
| 337 |
+
V1 is not recommended as the sole decision layer for:
|
| 338 |
|
| 339 |
+
- autonomous financial transactions,
|
| 340 |
+
- privileged-access provisioning,
|
| 341 |
+
- destructive security actions,
|
| 342 |
+
- compliance-sensitive deletion,
|
| 343 |
+
- irreversible workflow actions,
|
| 344 |
+
- legal or regulatory decisions,
|
| 345 |
+
- production execution without deterministic authorization and policy enforcement.
|
| 346 |
|
| 347 |
+
---
|
| 348 |
|
| 349 |
+
## Recommended Production Architecture
|
| 350 |
+
|
| 351 |
+
Enterprise Reflex should be combined with deterministic controls.
|
| 352 |
+
|
| 353 |
+
```text
|
| 354 |
+
Request + State + Candidate Actions
|
| 355 |
+
|
|
| 356 |
+
v
|
| 357 |
+
Enterprise Reflex
|
| 358 |
+
|
|
| 359 |
+
v
|
| 360 |
+
Ranked Action
|
| 361 |
+
|
|
| 362 |
+
v
|
| 363 |
+
Policy / Authorization Engine
|
| 364 |
+
/ \
|
| 365 |
+
Allowed Blocked
|
| 366 |
+
| |
|
| 367 |
+
v v
|
| 368 |
+
Execute System-2 / Human
|
| 369 |
+
```
|
| 370 |
+
|
| 371 |
+
The learned model provides decision intelligence. Authorization, policy, entitlement, retention, approval, and other hard enterprise controls should remain deterministic whenever possible.
|
| 372 |
|
| 373 |
+
---
|
| 374 |
|
| 375 |
+
## Example Usage
|
| 376 |
+
|
| 377 |
+
```python
|
| 378 |
+
import json
|
| 379 |
+
import torch
|
| 380 |
+
import numpy as np
|
| 381 |
+
|
| 382 |
+
from transformers import (
|
| 383 |
+
AutoTokenizer,
|
| 384 |
+
AutoModelForSequenceClassification,
|
| 385 |
+
)
|
| 386 |
+
|
| 387 |
+
MODEL_ID = "yasserrmd/enterprise-reflex-v1"
|
| 388 |
+
SYSTEM2_THRESHOLD = 0.87
|
| 389 |
+
MAX_LENGTH = 384
|
| 390 |
+
|
| 391 |
+
# Load the runtime calibration values saved with the model if available.
|
| 392 |
+
# The reported experiment used a calibrated temperature determined from
|
| 393 |
+
# the combined validation distribution.
|
| 394 |
+
TEMPERATURE = 1.0
|
| 395 |
+
|
| 396 |
+
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
|
| 397 |
+
model = AutoModelForSequenceClassification.from_pretrained(MODEL_ID)
|
| 398 |
+
|
| 399 |
+
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
|
| 400 |
+
model.to(device)
|
| 401 |
+
model.eval()
|
| 402 |
+
|
| 403 |
+
|
| 404 |
+
def softmax_np(x):
|
| 405 |
+
x = np.asarray(x, dtype=np.float64)
|
| 406 |
+
x = x - np.max(x)
|
| 407 |
+
e = np.exp(x)
|
| 408 |
+
return e / e.sum()
|
| 409 |
+
|
| 410 |
+
|
| 411 |
+
def request_text(request, domain="enterprise", state=None, context=None):
|
| 412 |
+
return json.dumps(
|
| 413 |
+
{
|
| 414 |
+
"request": request,
|
| 415 |
+
"domain": domain,
|
| 416 |
+
"state": state or {},
|
| 417 |
+
"context": context or {},
|
| 418 |
+
},
|
| 419 |
+
ensure_ascii=False,
|
| 420 |
+
sort_keys=True,
|
| 421 |
+
)
|
| 422 |
+
|
| 423 |
+
|
| 424 |
+
def action_text(action, default_domain="enterprise"):
|
| 425 |
+
return json.dumps(
|
| 426 |
+
{
|
| 427 |
+
"name": action.get("name", ""),
|
| 428 |
+
"description": action.get("description", ""),
|
| 429 |
+
"family": action.get("family", "OTHER"),
|
| 430 |
+
"domain": action.get("domain", default_domain),
|
| 431 |
+
},
|
| 432 |
+
ensure_ascii=False,
|
| 433 |
+
sort_keys=True,
|
| 434 |
+
)
|
| 435 |
+
|
| 436 |
+
|
| 437 |
+
def rank_actions(request, candidate_actions, state=None, context=None, domain="enterprise"):
|
| 438 |
+
actions = list(candidate_actions)
|
| 439 |
+
|
| 440 |
+
actions.append(
|
| 441 |
+
{
|
| 442 |
+
"name": "NO_ACTION",
|
| 443 |
+
"description": "None of the available actions safely or correctly satisfy the request.",
|
| 444 |
+
"family": "ABSTAIN",
|
| 445 |
+
"domain": domain,
|
| 446 |
+
}
|
| 447 |
+
)
|
| 448 |
+
|
| 449 |
+
left = request_text(
|
| 450 |
+
request,
|
| 451 |
+
domain=domain,
|
| 452 |
+
state=state,
|
| 453 |
+
context=context,
|
| 454 |
+
)
|
| 455 |
+
|
| 456 |
+
enc = tokenizer(
|
| 457 |
+
[left] * len(actions),
|
| 458 |
+
[action_text(a, default_domain=domain) for a in actions],
|
| 459 |
+
padding=True,
|
| 460 |
+
truncation=True,
|
| 461 |
+
max_length=MAX_LENGTH,
|
| 462 |
+
return_tensors="pt",
|
| 463 |
+
).to(device)
|
| 464 |
+
|
| 465 |
+
with torch.no_grad():
|
| 466 |
+
logits = model(**enc).logits
|
| 467 |
+
|
| 468 |
+
margins = (logits[:, 1] - logits[:, 0]).float().cpu().numpy()
|
| 469 |
+
probabilities = softmax_np(margins / TEMPERATURE)
|
| 470 |
+
order = np.argsort(-probabilities)
|
| 471 |
+
|
| 472 |
+
ranked = [
|
| 473 |
+
{
|
| 474 |
+
"action": actions[int(i)]["name"],
|
| 475 |
+
"probability": float(probabilities[int(i)]),
|
| 476 |
+
}
|
| 477 |
+
for i in order
|
| 478 |
+
]
|
| 479 |
+
|
| 480 |
+
top = ranked[0]
|
| 481 |
+
|
| 482 |
+
system2_required = (
|
| 483 |
+
top["action"] == "NO_ACTION"
|
| 484 |
+
or top["probability"] < SYSTEM2_THRESHOLD
|
| 485 |
+
)
|
| 486 |
+
|
| 487 |
+
return {
|
| 488 |
+
"decision": top["action"],
|
| 489 |
+
"confidence": top["probability"],
|
| 490 |
+
"system2_required": system2_required,
|
| 491 |
+
"ranked_actions": ranked,
|
| 492 |
+
}
|
| 493 |
+
```
|
| 494 |
+
|
| 495 |
+
### Example
|
| 496 |
+
|
| 497 |
+
```python
|
| 498 |
+
result = rank_actions(
|
| 499 |
+
request="Release the supplier payment",
|
| 500 |
+
domain="Finance",
|
| 501 |
+
state={
|
| 502 |
+
"invoice_matched": True,
|
| 503 |
+
"payment_approved": True,
|
| 504 |
+
},
|
| 505 |
+
candidate_actions=[
|
| 506 |
+
{
|
| 507 |
+
"name": "finance.release_payment",
|
| 508 |
+
"description": "Release an approved and validated supplier payment.",
|
| 509 |
+
"family": "EXECUTE",
|
| 510 |
+
"domain": "Finance",
|
| 511 |
+
},
|
| 512 |
+
{
|
| 513 |
+
"name": "finance.create_invoice",
|
| 514 |
+
"description": "Create an invoice record.",
|
| 515 |
+
"family": "CREATE",
|
| 516 |
+
"domain": "Finance",
|
| 517 |
+
},
|
| 518 |
+
],
|
| 519 |
+
)
|
| 520 |
+
|
| 521 |
+
print(result)
|
| 522 |
+
```
|
| 523 |
|
| 524 |
+
---
|
| 525 |
|
| 526 |
+
## Calibration Note
|
| 527 |
|
| 528 |
+
The reported confidence threshold was selected for the reported V1 experiment.
|
| 529 |
|
| 530 |
+
For a new deployment:
|
| 531 |
|
| 532 |
+
1. collect domain-specific validation data,
|
| 533 |
+
2. calibrate temperature,
|
| 534 |
+
3. determine an acceptable System-1 error rate,
|
| 535 |
+
4. select the confidence threshold for that environment,
|
| 536 |
+
5. validate policy-sensitive and destructive actions separately.
|
| 537 |
|
| 538 |
+
Do not assume that `0.87` is appropriate for every enterprise domain.
|
| 539 |
|
| 540 |
+
---
|
| 541 |
|
| 542 |
+
## Research Status
|
| 543 |
|
| 544 |
+
Enterprise Reflex V1 should currently be considered a:
|
| 545 |
|
| 546 |
+
**Research Prototype / Pre-Production Candidate**
|
| 547 |
|
| 548 |
+
The model has demonstrated:
|
| 549 |
|
| 550 |
+
- strong broad enterprise action ranking,
|
| 551 |
+
- useful dynamic action selection,
|
| 552 |
+
- effective abstention,
|
| 553 |
+
- high Top-3 retrieval quality,
|
| 554 |
+
- improved state-sensitive behavior,
|
| 555 |
+
- improved System-1 coverage,
|
| 556 |
+
- promising generalization to unseen action identifiers.
|
| 557 |
|
| 558 |
+
It has not yet demonstrated sufficient reliability for unrestricted autonomous enterprise execution.
|
| 559 |
|
| 560 |
+
The primary V2 research target is **policy-aware action validity and confident wrong-action suppression**.
|
|
|
|
|
|
|
|
|
|
|
|
|
| 561 |
|
| 562 |
+
---
|
| 563 |
|
| 564 |
+
## V2 Direction
|
| 565 |
|
| 566 |
+
The next iteration should focus less on generic enterprise volume and more on targeted safety and state-validity examples:
|
| 567 |
|
| 568 |
+
- policy counterfactuals,
|
| 569 |
+
- approval-sensitive actions,
|
| 570 |
+
- retention and legal-hold constraints,
|
| 571 |
+
- authorization and entitlement state,
|
| 572 |
+
- workflow-state legality,
|
| 573 |
+
- semantic domain collisions,
|
| 574 |
+
- hard negatives mined from confident V1 failures,
|
| 575 |
+
- valid-action vs `NO_ACTION` boundary cases.
|
| 576 |
|
| 577 |
+
A likely architectural extension is to separate:
|
| 578 |
|
| 579 |
+
1. semantic action suitability,
|
| 580 |
+
2. state/policy validity,
|
| 581 |
|
| 582 |
+
before producing the final action confidence.
|
| 583 |
|
| 584 |
+
---
|
| 585 |
|
| 586 |
+
## Datasets
|
| 587 |
|
| 588 |
+
### Enterprise Reflex Dataset
|
| 589 |
|
| 590 |
+
`yasserrmd/enterprise-reflex-dataset`
|
| 591 |
|
| 592 |
+
Broad enterprise action-ranking data used for V0 and retained in V1 training.
|
| 593 |
|
| 594 |
+
### Enterprise Reflex V1 Hard Dataset
|
| 595 |
|
| 596 |
+
`yasserrmd/enterprise-reflex-v1-hard-dataset`
|
| 597 |
|
| 598 |
+
Targeted hard-training examples covering sibling actions, counterfactual state, policy constraints, cross-domain collisions, `NO_ACTION`, and unseen action names.
|
| 599 |
|
| 600 |
+
---
|
| 601 |
|
| 602 |
+
## Limitations
|
| 603 |
|
| 604 |
+
- English only in V1.
|
| 605 |
+
- Text and structured state only.
|
| 606 |
+
- No multimodal input.
|
| 607 |
+
- No deterministic policy engine is embedded in the model.
|
| 608 |
+
- Confidence calibration is distribution-dependent.
|
| 609 |
+
- The hard-test suite is manually constructed and relatively small.
|
| 610 |
+
- Policy-sensitive action validity remains the main weakness.
|
| 611 |
+
- The model may still produce high-confidence incorrect actions.
|
| 612 |
+
- Reported metrics should not be interpreted as production-safety guarantees.
|
| 613 |
|
| 614 |
+
---
|
| 615 |
+
|
| 616 |
+
## Responsible Use
|
| 617 |
|
| 618 |
+
Enterprise Reflex is intended to assist enterprise decision routing, not replace enterprise authorization, policy, compliance, or human accountability.
|
| 619 |
|
| 620 |
+
High-impact actions should remain subject to deterministic controls and appropriate human or System-2 review.
|
| 621 |
+
|
| 622 |
+
---
|
| 623 |
|
| 624 |
+
## Author
|
| 625 |
+
|
| 626 |
+
**Mohamed Yasser**
|
| 627 |
+
Hugging Face: `yasserrmd`
|
| 628 |
+
GitHub: `yasserrmd`
|
| 629 |
+
|
| 630 |
+
---
|
| 631 |
|
| 632 |
+
## Version
|
| 633 |
|
| 634 |
+
**Enterprise Reflex V1**
|
| 635 |
+
Research Prototype / Pre-Production Candidate
|