bogdanraduta commited on
Commit
1aca094
·
verified ·
1 Parent(s): d5c6472

Add inference_contract

Browse files
inference_contract/INFERENCE.md ADDED
@@ -0,0 +1,95 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Inference contract: FlowX Semantic Mapper
2
+
3
+ Prompt version: `mapper_sys_v1`.
4
+
5
+ This model was trained against a **frozen inference contract**: an exact system
6
+ prompt, an exact user-turn format, a fixed decode setting, and a fixed output
7
+ schema. Reproduce all four or the outputs drift. Do not edit the prompt or schema;
8
+ the weights are trained against them.
9
+
10
+ Files in this directory:
11
+
12
+ - [`prompt_mapper_sys_v1.txt`](./prompt_mapper_sys_v1.txt): the system prompt, verbatim.
13
+ - [`schema_mapper_v1.json`](./schema_mapper_v1.json): JSON Schema for the output object (structural / semantic / governance).
14
+
15
+ Referenced from the repo root:
16
+
17
+ - `concept_taxonomy.yaml`: the 252-concept controlled vocabulary the `concepts` field is drawn from.
18
+
19
+ ---
20
+
21
+ ## 1. System prompt (verbatim)
22
+
23
+ The exact contents of [`prompt_mapper_sys_v1.txt`](./prompt_mapper_sys_v1.txt):
24
+
25
+ ```
26
+ You are a legal and regulatory ontology extractor.
27
+ Extract structured tags from document chunks. Output ONLY valid JSON.
28
+ ```
29
+
30
+ Send it as the `system` turn. No trailing whitespace, no extra lines.
31
+
32
+ ## 2. User turn format
33
+
34
+ One regulatory/legal text chunk per request, wrapped exactly like this:
35
+
36
+ ```
37
+ Extract ontology from this chunk:
38
+
39
+ CHUNK:
40
+ <the regulatory text>
41
+ ```
42
+
43
+ - The literal header `Extract ontology from this chunk:`, a blank line, then
44
+ `CHUNK:`, a newline, then the raw chunk text.
45
+ - One chunk per call. The model was trained on single-chunk turns; do not batch
46
+ multiple clauses into one user turn.
47
+ - Pass the chunk verbatim (the source language is fine: EN, FR, DE, RO). Do not
48
+ pre-summarize or translate it.
49
+
50
+ ## 3. Decode settings
51
+
52
+ | Setting | Value | Why |
53
+ | --- | --- | --- |
54
+ | `enable_thinking` | **`False`** | Qwen3 is a thinking model, but the adapter was trained on pure JSON with no reasoning block. Leaving thinking on yields an empty or malformed object. |
55
+ | `temperature` | **`0` (greedy)** | The task is deterministic extraction; sampling only adds drift. |
56
+ | `max_new_tokens` | **~1024** | A full three-facet object fits comfortably; 1024 leaves headroom for long hierarchies. |
57
+ | stop | end-of-turn | The model emits a single JSON object and stops. |
58
+
59
+ Apply the chat template with `add_generation_prompt=True, enable_thinking=False`.
60
+
61
+ ## 4. Output
62
+
63
+ A single JSON object with three top-level facets (`structural`, `semantic`,
64
+ `governance`), conforming to [`schema_mapper_v1.json`](./schema_mapper_v1.json).
65
+ Parse it strictly. On held-out data JSON validity is 1.00 and all three facets are
66
+ present 1.00, so a parse failure means the contract above was not reproduced (most
67
+ often `enable_thinking` left at its default).
68
+
69
+ ## 5. The controlled concept vocabulary (required for `semantic.concepts`)
70
+
71
+ `semantic.concepts` is **not** free text. It is drawn from a **252-concept
72
+ controlled taxonomy** shipped as `concept_taxonomy.yaml` at the repo root (6
73
+ categories: money, rights_waived, time_renewal, lease, insurance, data; each entry
74
+ has an `id`, a `definition`, and a `primary_domain`).
75
+
76
+ This is the model's central design choice. Open free-text concepts (1062 unique in
77
+ the first corpus, 88% of them singletons) were unlearnable and unmeasurable;
78
+ collapsing to 252 canonical ids made the `concepts` facet both learnable and
79
+ scoreable (F1 0.24 → 0.54). At integration time you should:
80
+
81
+ - Treat any concept id **not** present in `concept_taxonomy.yaml` as
82
+ out-of-vocabulary and drop or flag it. The model targets the controlled set but
83
+ can still surface an occasional near-miss id.
84
+ - Use the taxonomy `id` as the join key into your policy layer.
85
+
86
+ `domain_tags`, by contrast, are a **free snake_case** vocabulary and are expected
87
+ to be noisier / less consistent than `concepts`.
88
+
89
+ ## 6. Downstream
90
+
91
+ The `governance.escalation_trigger` (`if <condition> THEN escalate`) and
92
+ `policy_references` (`PDP.<domain>.<rule>`) are consumed by the sibling
93
+ **[`flowxai/sentinel-gate`](https://huggingface.co/flowxai/sentinel-gate)**
94
+ escalation model: Mapper tags a chunk → policy layer → Sentinel decides
95
+ DECIDE vs ESCALATE. Keep the field names stable so the pipeline lines up.
inference_contract/prompt_mapper_sys_v1.txt ADDED
@@ -0,0 +1,2 @@
 
 
 
1
+ You are a legal and regulatory ontology extractor.
2
+ Extract structured tags from document chunks. Output ONLY valid JSON.
inference_contract/schema_mapper_v1.json ADDED
@@ -0,0 +1,96 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "$schema": "http://json-schema.org/draft-07/schema#",
3
+ "$id": "https://huggingface.co/flowxai/semantic-mapper/inference_contract/schema_mapper_v1.json",
4
+ "title": "FlowX Semantic Mapper output (schema_mapper_v1)",
5
+ "description": "The single JSON object the Semantic Mapper emits for one regulatory/legal text chunk. Three facets: structural, semantic, governance. Prompt version mapper_sys_v1.",
6
+ "type": "object",
7
+ "additionalProperties": false,
8
+ "required": ["structural", "semantic", "governance"],
9
+ "properties": {
10
+ "structural": {
11
+ "type": "object",
12
+ "description": "Where the chunk sits in its source document.",
13
+ "additionalProperties": false,
14
+ "required": ["source_id", "hierarchy", "document_type"],
15
+ "properties": {
16
+ "source_id": {
17
+ "type": "string",
18
+ "description": "Stable identifier for the chunk, typically an uppercased normalization of the citation path (e.g. US_CA_INS_790_03_h_2)."
19
+ },
20
+ "hierarchy": {
21
+ "type": "array",
22
+ "description": "Ordered flat array of alternating [level, value] pairs from the outermost container down to the leaf (e.g. [\"state_code\",\"california_insurance_code\",\"section\",\"790.03\",\"subdivision\",\"h\",\"paragraph\",\"2\"]).",
23
+ "items": { "type": "string" },
24
+ "minItems": 2
25
+ },
26
+ "document_type": {
27
+ "type": "string",
28
+ "description": "Snake_case document class, e.g. state_insurance_code, eu_regulation, federal_regulation, labor_code, adr_agreement."
29
+ }
30
+ }
31
+ },
32
+ "semantic": {
33
+ "type": "object",
34
+ "description": "What the chunk is about.",
35
+ "additionalProperties": false,
36
+ "required": ["domain_tags", "concepts", "entities"],
37
+ "properties": {
38
+ "domain_tags": {
39
+ "type": "array",
40
+ "description": "Free snake_case topic tags (open vocabulary). Less consistent than concepts by design.",
41
+ "items": { "type": "string" }
42
+ },
43
+ "concepts": {
44
+ "type": "array",
45
+ "description": "Concept ids drawn from the 252-concept controlled taxonomy (concept_taxonomy.yaml, shipped at repo root). Values SHOULD be in-vocabulary; out-of-vocabulary ids are treated as errors by downstream consumers.",
46
+ "items": { "type": "string" }
47
+ },
48
+ "entities": {
49
+ "type": "object",
50
+ "description": "The core actor/action/object triple plus a constraint map.",
51
+ "additionalProperties": false,
52
+ "required": ["actor", "action", "object", "constraint"],
53
+ "properties": {
54
+ "actor": {
55
+ "type": "string",
56
+ "description": "Who the obligation/right falls on (e.g. insurer, employer, carrier, credit_institution)."
57
+ },
58
+ "action": {
59
+ "type": "string",
60
+ "description": "What must/may/must-not be done (verb phrase)."
61
+ },
62
+ "object": {
63
+ "type": "string",
64
+ "description": "What the action is performed on (e.g. claim_communication, personal_data, hazmat_package)."
65
+ },
66
+ "constraint": {
67
+ "type": "object",
68
+ "description": "Free key/value map of qualifying conditions (e.g. {\"condition\":\"reasonably_prompt\"}, {\"deadline_days\":\"30\"}). Keys and values are snake_case strings.",
69
+ "additionalProperties": { "type": "string" }
70
+ }
71
+ }
72
+ }
73
+ }
74
+ },
75
+ "governance": {
76
+ "type": "object",
77
+ "description": "How the chunk maps to FlowX policy and escalation.",
78
+ "additionalProperties": false,
79
+ "required": ["policy_references", "escalation_trigger"],
80
+ "properties": {
81
+ "policy_references": {
82
+ "type": "array",
83
+ "description": "Policy ids in the form PDP.<domain>.<rule> (e.g. PDP.insurance.claims_settlement).",
84
+ "items": {
85
+ "type": "string",
86
+ "pattern": "^PDP\\.[a-z_]+\\.[a-z0-9_]+$"
87
+ }
88
+ },
89
+ "escalation_trigger": {
90
+ "type": "string",
91
+ "description": "A single guard clause in the form 'if <condition> THEN escalate' consumed downstream by the Sentinel Gate."
92
+ }
93
+ }
94
+ }
95
+ }
96
+ }