bogdanraduta commited on
Commit
7975cfa
·
verified ·
1 Parent(s): 433c92d

Add EVAL_RESULTS.md

Browse files
Files changed (1) hide show
  1. EVAL_RESULTS.md +52 -0
EVAL_RESULTS.md ADDED
@@ -0,0 +1,52 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # FlowX Semantic Mapper (4B) — Evaluation Results
2
+
3
+ Held-out evaluation of the released model (`flowx-semantic-mapper-4b-v2`). All numbers are on a
4
+ **document-disjoint held-out set of 112 records** spanning EN/FR/DE/RO across banking,
5
+ insurance, logistics, and labor. Decoding: greedy (temperature 0), `enable_thinking=False`.
6
+ Metrics are exact set-match precision/recall/F1 unless noted; `concepts` also reports a
7
+ fuzzy/semantic token-match F1, and `entities` reports field-accuracy over actor/action/object.
8
+
9
+ ## Headline (held-out, n=112)
10
+
11
+ | Field | Precision | Recall | F1 |
12
+ | --- | --- | --- | --- |
13
+ | JSON validity | | | 1.00 |
14
+ | all-3-facets present | | | 1.00 |
15
+ | domain_tags (free-text) | 0.552 | 0.570 | 0.555 |
16
+ | concepts (controlled 252-vocab, exact) | 0.547 | 0.558 | 0.541 |
17
+ | concepts (fuzzy / semantic match) | 0.596 | 0.590 | 0.578 |
18
+ | entities (field-accuracy) | | | 0.530 |
19
+
20
+ ## The controlled-vocabulary result (before / after)
21
+
22
+ The central design change of this release was replacing open free-text `concepts` (1062
23
+ unique labels, 88% seen exactly once, unlearnable and unmeasurable) with a **controlled
24
+ 252-concept taxonomy** (`concept_taxonomy.yaml`), then expanding the corpus with real
25
+ regulatory text.
26
+
27
+ | Concepts F1 | value | held-out |
28
+ | --- | --- | --- |
29
+ | open free-text vocabulary (v1) | 0.24 (exact) / 0.37 (fuzzy) | n=35 |
30
+ | controlled vocab + data expansion (this release) | **0.54** (exact) / **0.58** (fuzzy) | n=112 |
31
+
32
+ Concepts F1 more than doubled, and `entities` went from **unmeasurable** (the facet was empty
33
+ in the old held-out) to **0.53**, on a held-out set that is 3x larger and multilingual.
34
+
35
+ ## Honest reading / caveats
36
+
37
+ - **Home-field note.** The held-out shares the annotation lineage of the training data
38
+ (real regulatory text, frontier-model-assisted ontology annotation to the controlled vocab).
39
+ Read these as a strong format-and-consistency signal; validate on your own corpus before
40
+ production claims.
41
+ - **Concepts are partial (~0.54).** The model surfaces many but not all applicable concepts.
42
+ - **`domain_tags` are free-text on purpose** (recall + human browsing) and are honestly noisier
43
+ than the controlled `concepts`.
44
+ - **English is strongest;** FR/DE/RO are supported and materially improved by the real-data
45
+ expansion, but remain a bit weaker than EN.
46
+
47
+ ## Reproduction
48
+
49
+ Held-out set: `mlx_data/ontology_cv2/valid.jsonl` (built document-disjoint from the training
50
+ split). Scorer: `eval_mapper_4b.py <model_path> mlx_data/ontology_cv2/valid.jsonl`.
51
+
52
+ _Author: Bogdan Răduță, Head of Research, FlowX.AI._