mangesh-ux commited on
Commit
8a608a9
·
verified ·
1 Parent(s): 405291c

license update

Browse files
Files changed (1) hide show
  1. README.md +27 -8
README.md CHANGED
@@ -1,7 +1,7 @@
1
  ---
2
  language:
3
  - en
4
- license: mit
5
  pipeline_tag: text-generation
6
  tags:
7
  - logistics
@@ -37,10 +37,10 @@ This model is a QLoRA fine-tune of `Qwen/Qwen2.5-3B-Instruct` for extracting str
37
 
38
  The target output is a strict JSON object compatible with `LogisticsCXMetrics` (`behavioral_analytics`, `operational_analytics`, `diagnostic_reasoning`).
39
 
40
- The output schema and taxonomy are derived from curated reference files in `docs/knowledge/`:
41
- - `Transcript-Only CX Difficulty Score_ Standards, Methods, and a Rigorous MVP Design.pdf`
42
  Deep-research document (ChatGPT-generated) on transcript-only CX friction signals and effort scoring methodology.
43
- - `Logistics CX Data Schema Development.docx`
44
  NotebookLM-assisted intent and schema research used to shape intent taxonomy and extraction field design.
45
 
46
  These definitions are operationalized in `src/schema.py` and reflected in training labels.
@@ -73,7 +73,7 @@ Project repository: [OmniCX-Extractor](https://github.com/mangesh-ux/OmniCX-Extr
73
  - Environment: single 8GB VRAM GPU setup (see training logs)
74
 
75
  Detailed run record:
76
- - `docs/training_logs/iteration_001.md`
77
 
78
  ## Evaluation
79
 
@@ -82,7 +82,9 @@ Current evaluation (research preview):
82
  - Eval examples: 32
83
  - Runtime errors: 0
84
  - Strict exact-match accuracy: 0.0% (0/32)
85
- - Mean latency: 22.35s/sample
 
 
86
 
87
  Selected per-field accuracy:
88
  - `customer_intent`: 56.2%
@@ -91,8 +93,8 @@ Selected per-field accuracy:
91
  - `escalation_requested`: 100.0%
92
 
93
  Detailed report:
94
- - `docs/training_logs/eval_report_iteration_001.md`
95
- - `docs/training_logs/eval_outputs_iteration_001.jsonl`
96
 
97
  ## Intended Uses
98
 
@@ -135,6 +137,23 @@ result = extract_with_finetuned(
135
  print(result)
136
  ```
137
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
138
  ### Input and Output Contract
139
 
140
  **Input (single transcript):**
 
1
  ---
2
  language:
3
  - en
4
+ license: apache-2.0
5
  pipeline_tag: text-generation
6
  tags:
7
  - logistics
 
37
 
38
  The target output is a strict JSON object compatible with `LogisticsCXMetrics` (`behavioral_analytics`, `operational_analytics`, `diagnostic_reasoning`).
39
 
40
+ The output schema and taxonomy are derived from curated reference files:
41
+ - [`Transcript-Only CX Difficulty Score_ Standards, Methods, and a Rigorous MVP Design.pdf`](https://github.com/mangesh-ux/OmniCX-Extractor/blob/main/docs/knowledge/Transcript-Only%20CX%20Difficulty%20Score_%20Standards%2C%20Methods%2C%20and%20a%20Rigorous%20MVP%20Design.pdf)
42
  Deep-research document (ChatGPT-generated) on transcript-only CX friction signals and effort scoring methodology.
43
+ - [`Logistics CX Data Schema Development.docx`](https://github.com/mangesh-ux/OmniCX-Extractor/blob/main/docs/knowledge/Logistics%20CX%20Data%20Schema%20Development.docx)
44
  NotebookLM-assisted intent and schema research used to shape intent taxonomy and extraction field design.
45
 
46
  These definitions are operationalized in `src/schema.py` and reflected in training labels.
 
73
  - Environment: single 8GB VRAM GPU setup (see training logs)
74
 
75
  Detailed run record:
76
+ - [`docs/training_logs/iteration_001.md`](https://github.com/mangesh-ux/OmniCX-Extractor/blob/main/docs/training_logs/iteration_001.md)
77
 
78
  ## Evaluation
79
 
 
82
  - Eval examples: 32
83
  - Runtime errors: 0
84
  - Strict exact-match accuracy: 0.0% (0/32)
85
+ - Mean latency: 29.84s/sample
86
+ - Min / max latency: 16.89s / 45.72s
87
+ - Total latency: 954.94s
88
 
89
  Selected per-field accuracy:
90
  - `customer_intent`: 56.2%
 
93
  - `escalation_requested`: 100.0%
94
 
95
  Detailed report:
96
+ - [`eval_report_iteration_001.md`](./eval_report_iteration_001.md)
97
+ - [`eval_outputs_iteration_001.jsonl`](./eval_outputs_iteration_001.jsonl)
98
 
99
  ## Intended Uses
100
 
 
137
  print(result)
138
  ```
139
 
140
+ ### Download from Hugging Face and run locally
141
+
142
+ ```python
143
+ from huggingface_hub import snapshot_download
144
+ from src.inference import load_model, extract_with_finetuned
145
+
146
+ local_model_dir = snapshot_download("mangesh-ux/omnicx-logistics-cx-extractor-qwen25-3b-lora")
147
+ model, tokenizer = load_model(model_path=local_model_dir)
148
+ result = extract_with_finetuned(
149
+ transcript="Agent: ... Customer: ...",
150
+ model=model,
151
+ tokenizer=tokenizer,
152
+ return_dict=True,
153
+ )
154
+ print(result)
155
+ ```
156
+
157
  ### Input and Output Contract
158
 
159
  **Input (single transcript):**