File size: 1,850 Bytes
30bda61 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 | ---
license: apache-2.0
tags:
- insurance
- ocr
- document-understanding
- claims
- lemonade-style
pipeline_tag: image-to-text
---
# 📄 Insurance Document OCR
**Extract structured data from insurance documents**
## Model Description
Specialized OCR model for extracting information from insurance-related documents including claims forms, policy documents, ID cards, and damage photos.
## Supported Documents
| Document Type | Fields Extracted |
|---------------|------------------|
| Claims Form | Claim #, Date, Amount, Description |
| Policy Document | Policy #, Coverage, Limits, Deductible |
| Driver's License | Name, DOB, License #, Address |
| Vehicle Registration | VIN, Make, Model, Year, Plate |
| Medical Bills | Provider, Date, Charges, Diagnosis |
| Repair Estimates | Shop, Parts, Labor, Total |
| Police Reports | Report #, Date, Officers, Description |
## Output Format
```json
{
"document_type": "claims_form",
"confidence": 0.96,
"extracted_fields": {
"claim_number": "CLM-2024-78432",
"incident_date": "2024-01-15",
"claim_amount": 2450.00,
"description": "Rear-end collision at intersection",
"policy_number": "POL-AUTO-12345"
},
"raw_text": "...",
"bounding_boxes": [...]
}
```
## Performance
| Metric | Score |
|--------|-------|
| Character Accuracy | 98.7% |
| Field Extraction | 95.2% |
| Document Classification | 97.8% |
| Processing Time | 1.2s/page |
## Usage
```python
from transformers import pipeline
ocr = pipeline("image-to-text", model="gcc-insurance-ml-models/document-ocr-insurance")
result = ocr("claim_form.jpg")
print(result["extracted_fields"])
```
## Integration
```
Document Upload
↓
[Document OCR] → Structured Data
↓
Auto-populate claim form
↓
Validate against policy
↓
Route to triage
```
## License
Apache 2.0
|