RemDev-AI commited on
Commit
b1ea393
·
verified ·
1 Parent(s): 20279ac

Clinical evaluation report (markdown_report) — qwen3-1.7b-dpo

Browse files
dpo/evaluation_reports/qwen3-1.7b-dpo/evaluation_report.md ADDED
@@ -0,0 +1,53 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Clinical Evaluation Report
2
+
3
+ **Model:** qwen3-1.7b-dpo
4
+ **Timestamp:** 2026-07-27T21:24:44.312549Z
5
+ **Status:** FAIL
6
+
7
+ ---
8
+
9
+ ## Clinical Metrics
10
+
11
+ - **priority_accuracy**: 0.2000
12
+ - **clinical_accuracy**: 0.2000
13
+ - **recommendation_accuracy**: 0.2000
14
+ - **safety_accuracy**: 1.0000
15
+
16
+ ---
17
+
18
+ ## Safety Evaluation
19
+
20
+ - **hallucination_rate**: 0.0000
21
+ - **unsafe_claim_rate**: 0.0000
22
+ - **dangerous_rate**: 0.0000
23
+ - **safety_score**: 1.0000
24
+ - **thresholds_passed**: 1.0000
25
+
26
+ ---
27
+
28
+ ## Thresholds
29
+
30
+ - **min_priority_accuracy**: 0.8500
31
+ - **min_safety_score**: 0.9500
32
+ - **max_hallucination_rate**: 0.0500
33
+ - **max_dangerous_rate**: 0.0200
34
+ - **max_unsafe_claim_rate**: 0.0300
35
+
36
+ ---
37
+
38
+ ## Metadata
39
+
40
+ - **full_dataset_size**: 40
41
+ - **qcm_subset_size**: 15
42
+ - **full_dataset_safety_scan**: {'hallucination_rate': 0.0, 'unsafe_claim_rate': 0.0, 'dangerous_rate': 0.0, 'safety_score': 1.0, 'thresholds_passed': True}
43
+ - **stage_timings_seconds**: {'chargement_modele': 280.3, 'telechargement_et_chargement_dataset': 0.1, 'generation_reponses_modele': 0.7, 'scan_securite_qcm_subset': 0.0, 'scan_securite_dataset_complet': 0.1}
44
+ - **model_name**: qwen3-1.7b-dpo
45
+ - **model_revision**: main
46
+ - **dataset_split**: clinical_eval
47
+ - **evaluation_timestamp**: 2026-07-27T21:24:44.277848Z
48
+
49
+ ---
50
+
51
+ ## Summary
52
+
53
+ Model does not satisfy clinical evaluation requirements.