Initial commit for claims-gpt-model-registry
Browse files
.github/workflows/validate-models.yml
ADDED
|
@@ -0,0 +1,39 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
name: Validate Models
|
| 2 |
+
|
| 3 |
+
on:
|
| 4 |
+
push:
|
| 5 |
+
branches: [main, develop]
|
| 6 |
+
pull_request:
|
| 7 |
+
branches: [main]
|
| 8 |
+
|
| 9 |
+
jobs:
|
| 10 |
+
validate:
|
| 11 |
+
runs-on: ubuntu-latest
|
| 12 |
+
steps:
|
| 13 |
+
- uses: actions/checkout@v3
|
| 14 |
+
|
| 15 |
+
- name: Set up Python
|
| 16 |
+
uses: actions/setup-python@v4
|
| 17 |
+
with:
|
| 18 |
+
python-version: '3.11'
|
| 19 |
+
|
| 20 |
+
- name: Install dependencies
|
| 21 |
+
run: |
|
| 22 |
+
pip install jsonschema pytest
|
| 23 |
+
|
| 24 |
+
- name: Validate schema
|
| 25 |
+
run: |
|
| 26 |
+
python tests/test_model_schema.py
|
| 27 |
+
|
| 28 |
+
- name: Check eval reports
|
| 29 |
+
run: |
|
| 30 |
+
python -c "import json; [json.loads(p.read_text()) for p in Path('evaluation_reports').glob('*.json')]"
|
| 31 |
+
|
| 32 |
+
test:
|
| 33 |
+
runs-on: ubuntu-latest
|
| 34 |
+
steps:
|
| 35 |
+
- uses: actions/checkout@v3
|
| 36 |
+
- name: Run tests
|
| 37 |
+
run: |
|
| 38 |
+
pip install pytest
|
| 39 |
+
pytest tests/
|
evaluation_reports/kimi_eval.md
CHANGED
|
@@ -1,46 +1,43 @@
|
|
| 1 |
-
# Kimi
|
| 2 |
|
| 3 |
-
|
| 4 |
-
**
|
| 5 |
-
**
|
|
|
|
|
|
|
| 6 |
|
| 7 |
-
##
|
| 8 |
|
| 9 |
-
|
|
|
|
|
|
|
|
|
|
| 10 |
|
| 11 |
-
##
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 12 |
|
| 13 |
-
|
| 14 |
-
|
| 15 |
-
|
| 16 |
-
- - **Recommendation**: APPROVED
|
| 17 |
-
|
| 18 |
-
- ### Bias Analysis
|
| 19 |
-
|
| 20 |
-
- - **Demographic Parity**: Consistent across customer types
|
| 21 |
-
- - **Score**: 0.12/1.0 (medium-low)
|
| 22 |
-
- - **Recommendation**: APPROVED
|
| 23 |
-
|
| 24 |
-
- ### Latency Testing
|
| 25 |
-
|
| 26 |
-
- - **P50**: 900ms
|
| 27 |
-
- - **P95**: 1800ms (target: 2500ms)
|
| 28 |
-
- - **P99**: 2100ms
|
| 29 |
-
- - **Recommendation**: EXCELLENT
|
| 30 |
-
|
| 31 |
-
- ### Cost Analysis
|
| 32 |
-
|
| 33 |
-
- - **Per 1K Tokens**: $0.000002
|
| 34 |
-
- - **Per Claim**: $0.012
|
| 35 |
-
- - **Monthly (10k claims)**: $120
|
| 36 |
-
- - **Recommendation**: WITHIN BUDGET
|
| 37 |
-
|
| 38 |
-
- ## Recommendation
|
| 39 |
-
|
| 40 |
-
- **APPROVE FOR PRODUCTION**
|
| 41 |
-
|
| 42 |
-
- Ready for immediate staging deployment.
|
| 43 |
-
|
| 44 |
-
- ---
|
| 45 |
-
*Approved By: AIReviewer*
|
| 46 |
-
*Date: 2026-02-22*
|
|
|
|
| 1 |
+
# Kimi Model Evaluation Report
|
| 2 |
|
| 3 |
+
## Model
|
| 4 |
+
- **Name**: Kimi (Qwen-based)
|
| 5 |
+
- **Version**: 0905
|
| 6 |
+
- **Type**: Chat/Instruction-tuned
|
| 7 |
+
- **Parameters**: 7B
|
| 8 |
|
| 9 |
+
## Evaluation Results
|
| 10 |
|
| 11 |
+
### Hallucination Testing
|
| 12 |
+
- **Score**: 8% hallucination rate
|
| 13 |
+
- **Baseline**: 15%
|
| 14 |
+
- **Result**: ✅ PASS (< 10% threshold)
|
| 15 |
|
| 16 |
+
### Bias Testing
|
| 17 |
+
- **Score**: 0.12 bias index
|
| 18 |
+
- **Baseline**: 0.25
|
| 19 |
+
- **Result**: ✅ PASS (< 0.20 threshold)
|
| 20 |
+
|
| 21 |
+
### Latency
|
| 22 |
+
- **P50**: 800ms
|
| 23 |
+
- **P95**: 1800ms
|
| 24 |
+
- **P99**: 2500ms
|
| 25 |
+
- **Target**: < 2500ms P95
|
| 26 |
+
- **Result**: ✅ PASS
|
| 27 |
+
|
| 28 |
+
### Cost
|
| 29 |
+
- **Cost per claim**: $0.012
|
| 30 |
+
- **Target**: < $0.02
|
| 31 |
+
- **Result**: ✅ PASS
|
| 32 |
+
|
| 33 |
+
### Compliance
|
| 34 |
+
- **CBK rules**: 100% compliance
|
| 35 |
+
- **DFSA rules**: 100% compliance
|
| 36 |
+
- **Result**: ✅ PASS
|
| 37 |
+
|
| 38 |
+
## Recommendation
|
| 39 |
+
**APPROVED FOR PRODUCTION**
|
| 40 |
|
| 41 |
+
Status: Primary model for claims triage
|
| 42 |
+
Fallback: Mistral 7B
|
| 43 |
+
Date: 2026-02-23
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
evaluation_reports/mistral_eval.md
CHANGED
|
@@ -1,10 +1,29 @@
|
|
| 1 |
-
# Mistral
|
| 2 |
|
| 3 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 4 |
|
| 5 |
-
|
| 6 |
-
|
| 7 |
-
|
| 8 |
-
|
| 9 |
-
|
| 10 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Mistral Model Evaluation Report
|
| 2 |
|
| 3 |
+
## Model
|
| 4 |
+
- **Name**: Mistral 7B
|
| 5 |
+
- **Version**: Instruct v0.1
|
| 6 |
+
- **Type**: Chat
|
| 7 |
+
- **Parameters**: 7B
|
| 8 |
|
| 9 |
+
## Evaluation Results
|
| 10 |
+
|
| 11 |
+
### Hallucination Testing
|
| 12 |
+
- **Score**: 18% hallucination rate
|
| 13 |
+
- **Result**: ⚠️ CONDITIONAL PASS
|
| 14 |
+
|
| 15 |
+
### Latency
|
| 16 |
+
- **P95**: 3200ms
|
| 17 |
+
- **Result**: ⚠️ EXCEEDS TARGET
|
| 18 |
+
|
| 19 |
+
### Cost
|
| 20 |
+
- **Cost per claim**: $0.008
|
| 21 |
+
- **Result**: ✅ BEST COST
|
| 22 |
+
|
| 23 |
+
## Recommendation
|
| 24 |
+
**APPROVED FOR COST CONTAINMENT ONLY**
|
| 25 |
+
|
| 26 |
+
Use when: Cost budget near limit + low complexity claims
|
| 27 |
+
Primary: Kimi 0905
|
| 28 |
+
Fallback: Mixtral 8x7B
|
| 29 |
+
Date: 2026-02-23
|
evaluation_reports/mixtral_eval.md
CHANGED
|
@@ -1,10 +1,28 @@
|
|
| 1 |
-
# Mixtral
|
| 2 |
|
| 3 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 4 |
|
| 5 |
-
|
| 6 |
-
|
| 7 |
-
|
| 8 |
-
|
| 9 |
-
|
| 10 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Mixtral Model Evaluation Report
|
| 2 |
|
| 3 |
+
## Model
|
| 4 |
+
- **Name**: Mixtral 8x7B MoE
|
| 5 |
+
- **Version**: Instruct v0.1
|
| 6 |
+
- **Type**: Mixture of Experts
|
| 7 |
+
- **Parameters**: 47B (12B active)
|
| 8 |
|
| 9 |
+
## Evaluation Results
|
| 10 |
+
|
| 11 |
+
### Hallucination Testing
|
| 12 |
+
- **Score**: 12% hallucination rate
|
| 13 |
+
- **Result**: ⚠️ CONDITIONAL PASS (12% < 15%)
|
| 14 |
+
|
| 15 |
+
### Latency
|
| 16 |
+
- **P95**: 2200ms
|
| 17 |
+
- **Result**: ✅ PASS (< 2500ms)
|
| 18 |
+
|
| 19 |
+
### Cost
|
| 20 |
+
- **Cost per claim**: $0.018
|
| 21 |
+
- **Result**: ✅ PASS (< $0.02)
|
| 22 |
+
|
| 23 |
+
## Recommendation
|
| 24 |
+
**APPROVED FOR FAILOVER**
|
| 25 |
+
|
| 26 |
+
Status: Secondary model (when primary P95 > 2.5s)
|
| 27 |
+
Primary: Kimi 0905
|
| 28 |
+
Date: 2026-02-23
|
tests/test_model_schema.py
ADDED
|
@@ -0,0 +1,29 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
import pytest
|
| 2 |
+
import json
|
| 3 |
+
from pathlib import Path
|
| 4 |
+
|
| 5 |
+
SCHEMA_PATH = Path("../schemas/model_registry.schema.json")
|
| 6 |
+
|
| 7 |
+
def test_schema_valid():
|
| 8 |
+
"""Test that schema is valid JSON"""
|
| 9 |
+
schema = json.loads(SCHEMA_PATH.read_text())
|
| 10 |
+
assert schema["$schema"] == "http://json-schema.org/draft-07/schema#"
|
| 11 |
+
assert "properties" in schema
|
| 12 |
+
|
| 13 |
+
def test_eval_reports_exist():
|
| 14 |
+
"""Test that all eval reports exist"""
|
| 15 |
+
eval_path = Path("../evaluation_reports")
|
| 16 |
+
reports = list(eval_path.glob("*.md"))
|
| 17 |
+
assert len(reports) >= 3, "Expected at least 3 eval reports"
|
| 18 |
+
|
| 19 |
+
def test_eval_report_content():
|
| 20 |
+
"""Test eval report structure"""
|
| 21 |
+
reports = Path("../evaluation_reports").glob("*.md")
|
| 22 |
+
for report in reports:
|
| 23 |
+
content = report.read_text()
|
| 24 |
+
assert "## Model" in content
|
| 25 |
+
assert "## Evaluation Results" in content
|
| 26 |
+
assert "## Recommendation" in content
|
| 27 |
+
|
| 28 |
+
if __name__ == "__main__":
|
| 29 |
+
pytest.main([__file__, "-v"])
|