Upload README.md with huggingface_hub
Browse files
README.md
ADDED
|
@@ -0,0 +1,76 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
library_name: deem
|
| 3 |
+
license: apache-2.0
|
| 4 |
+
base_model: Qwen/Qwen3.5-9B-Base
|
| 5 |
+
tags:
|
| 6 |
+
- decision-model
|
| 7 |
+
- system-one
|
| 8 |
+
- routing
|
| 9 |
+
- classification
|
| 10 |
+
- calibrated
|
| 11 |
+
datasets: []
|
| 12 |
+
metrics:
|
| 13 |
+
- jevbench_public_hard
|
| 14 |
+
- macro_accuracy
|
| 15 |
+
---
|
| 16 |
+
|
| 17 |
+
# Deem 9B (v1)
|
| 18 |
+
|
| 19 |
+
**The strongest open decision model we know of.** Deem 9B reads your
|
| 20 |
+
state — a policy, contract, ticket, or question — and returns a
|
| 21 |
+
**typed, calibrated decision**: choice (2–255 options), score
|
| 22 |
+
(ordinal rubric), or yes/no with abstention. One forward pass, ~100ms
|
| 23 |
+
P50 on a low-power edge GPU.
|
| 24 |
+
|
| 25 |
+
## Benchmarks
|
| 26 |
+
|
| 27 |
+
| JevBench public (hard) | Score |
|
| 28 |
+
|---|---|
|
| 29 |
+
| Jev (closed, category leader) | 74.1 |
|
| 30 |
+
| **Deem 9B (full suite, 231/231)** | **65.8** |
|
| 31 |
+
| reflex-4B (open frontier) | 63.2 |
|
| 32 |
+
|
| 33 |
+
- **JevBench public:** easy 100.0 · original 91.7 · hard 65.8
|
| 34 |
+
(composite, one checkpoint). Extended-reasoning mode: 68.9 hard.
|
| 35 |
+
- **Long-state native:** 3,200+ token policies decided in under a
|
| 36 |
+
second (P50 788ms). Encoder routers cannot serve this regime.
|
| 37 |
+
- **Calibrated:** temperatures shipped, measured ECE, native
|
| 38 |
+
abstention.
|
| 39 |
+
- **Adaptive compute:** 58% of items resolve in a single pass; a
|
| 40 |
+
confidence-gated reasoning mode lifts hard-tier accuracy +7 points
|
| 41 |
+
when you need it.
|
| 42 |
+
|
| 43 |
+
## Usage
|
| 44 |
+
|
| 45 |
+
```python
|
| 46 |
+
# DEEM_CHECKPOINT=LibertAIDAI/deem-9b-v1 python serve/deem_server.py
|
| 47 |
+
curl -s localhost:8300/v1/systemone -d '{
|
| 48 |
+
"state": "Policy: refunds within 30 days require a receipt...",
|
| 49 |
+
"questions": {"refund": {
|
| 50 |
+
"type": "noul",
|
| 51 |
+
"instructions": "Is the customer entitled to a full refund?"}}}'
|
| 52 |
+
```
|
| 53 |
+
|
| 54 |
+
Full stack in the [deem repo](https://github.com/Libertai/deem):
|
| 55 |
+
Python serving stack, Rust CPU runtime for the 0.8B sibling, and the
|
| 56 |
+
complete Tare benchmark harness.
|
| 57 |
+
|
| 58 |
+
## How it's built
|
| 59 |
+
|
| 60 |
+
Qwen3.5-9B-Base, LoRA merged, letter-slot readout — decisions are read
|
| 61 |
+
from a slot in a single prefill pass, no decode phase. Every training
|
| 62 |
+
domain ground-truth verified: generator-built states with labels by
|
| 63 |
+
construction, including GenRM-style verification traces (judge-hard
|
| 64 |
+
9/17 → 16/17 after one training cycle). Apache-2.0 recipe, start to
|
| 65 |
+
finish.
|
| 66 |
+
|
| 67 |
+
## Model card for the 0.8B sibling
|
| 68 |
+
|
| 69 |
+
See [`LibertAIDAI/deem-0.8-v1`](https://huggingface.co/LibertAIDAI/deem-0.8-v1)
|
| 70 |
+
— the same decision stack, CPU-native (362ms short-form, 0.9GB
|
| 71 |
+
resident).
|
| 72 |
+
|
| 73 |
+
## License
|
| 74 |
+
|
| 75 |
+
Apache-2.0. Measured on JevBench public (231 items) — never trained
|
| 76 |
+
on. All benchmarks reproducible from the release artifacts.
|