AnkitAI commited on
Commit
5b6a59e
·
verified ·
1 Parent(s): 2239a02

update readme

Browse files
Files changed (1) hide show
  1. README.md +12 -12
README.md CHANGED
@@ -24,7 +24,7 @@ datasets: [jaredpalmer/kev-suites]
24
  <p>
25
  <a href="https://github.com/ankit-aglawe/tinyjev">GitHub</a> ·
26
  <a href="https://pypi.org/project/tinyjev/">PyPI</a> ·
27
- <a href="https://huggingface.co/AnkitAI/tinyjev-0.6b">TinyJev 0.6B</a> ·
28
  <a href="https://github.com/ankit-aglawe/tinyjev/tree/main/examples">Examples</a>
29
  </p>
30
 
@@ -53,13 +53,13 @@ Latency is a base M1 (16 GB) via MLX, one forward pass per case.
53
 
54
  | Model | Params | OD-500 | Gate 0.85 | ms / case | Weights |
55
  |---|---:|---:|---|---:|---|
56
- | <img src="https://raw.githubusercontent.com/ankit-aglawe/tinyjev/main/assets/logos/tinyjev.png" width="18"> **TinyJev&nbsp;0.6B** | 596M, 1.2 GB | 440 (88.0%) | 59% @ 98.0% | 85 | 🤗 [AnkitAI/tinyjev-0.6b](https://huggingface.co/AnkitAI/tinyjev-0.6b) |
57
- | <img src="https://raw.githubusercontent.com/ankit-aglawe/tinyjev/main/assets/logos/tinyjev.png" width="18"> **TinyJev&nbsp;4B** | 4.0B, 8.0 GB | 474 (94.8%) | 87% @ 99.1% | 628 | 🤗 [AnkitAI/tinyjev-4b](https://huggingface.co/AnkitAI/tinyjev-4b) |
58
 
59
  OD-500 is correct answers out of 500. Gate 0.85 is the share of decisions answered on its own at
60
  confidence ≥ 0.85, and how often those were right. Calibration (ECE 0.071 vs 0.022), coverage at 2%
61
  error (63% vs 92%) and transfer-v4 dev (0.625 vs 0.762) are on the benchmark page. Load either with
62
- `tinyjev.load("tinyjev-0.6b")` or `tinyjev.load("tinyjev-4b")`.
63
 
64
  Both rows are fp16. Loading with `quantize=8` keeps the same weights in half the memory and changes
65
  almost nothing: the 0.6B scores 440 at 90 ms, the 4B 473 at 845 ms, one answer in 500 different from
@@ -74,10 +74,10 @@ On OpenDecision's Original Choice 500, a suite of 25 domains that was not in the
74
  | Model | Correct / 500 | Handled alone at confidence ≥ 0.85 |
75
  |---|---:|---:|
76
  | Claude Opus 5.5 (cloud, self-reported probabilities) | 496 | 477 at 100.0% |
77
- | **tinyjev-4b** (fp16) | **474** | **437 at 99.1%** |
78
- | tinyjev-4b at `quantize=8` | 473 | 437 at 99.1% |
79
  | Kev-0.8B (raw logits) | 463 | 186 at 100.0% |
80
- | tinyjev-0.6b | 440 | 296 at 98.0% |
81
 
82
  353/375 on dev, 121/125 on holdout, 95% CI 0.928–0.966, ECE 0.022, Brier 0.071, coverage at 2% error 92%.
83
  On Kev's transfer-v4 dev it scores 0.762 against 0.625 for the 0.6B (Kev-4B, a full fine-tune, 0.790).
@@ -97,7 +97,7 @@ pip install 'tinyjev[torch]' # everything else
97
 
98
  ```python
99
  import tinyjev
100
- agent = tinyjev.load("tinyjev-4b", quantize=8) # 4.5 GB in memory; drop quantize for fp16, 8 GB
101
 
102
  agent.predict({
103
  "state": "Shoes arrived two weeks late and in the wrong size. Also I see two charges on my card.",
@@ -116,12 +116,12 @@ Eight bits changed one answer in 500 on the held-out suite. Serve it over HTTP,
116
  One request shape:
117
 
118
  ```bash
119
- tinyjev serve tinyjev-4b --quantize 8 # POST /v1/systemone on 127.0.0.1:8077
120
  ```
121
 
122
  ## What is in this repo
123
 
124
- `AutoModel.from_pretrained("AnkitAI/tinyjev-4b")` loads the backbone on its own, a standard
125
  `Qwen3Model` in fp16 with the LoRA already merged. The decision head lives in `head.safetensors`, and
126
  `tinyjev` is what turns hidden states into calibrated answers.
127
 
@@ -134,8 +134,8 @@ training. Temperature 1.0 at inference; the calibration figures above are at raw
134
 
135
  | | transfer-v4 dev | decision-v7 dev | OpenDecision 500 |
136
  | --- | --- | --- | --- |
137
- | tinyjev-4b | 0.762 | 0.859 | 474 / 500 |
138
- | tinyjev-0.6b | 0.625 | | 440 / 500 |
139
 
140
  ## Support the Project
141
 
 
24
  <p>
25
  <a href="https://github.com/ankit-aglawe/tinyjev">GitHub</a> ·
26
  <a href="https://pypi.org/project/tinyjev/">PyPI</a> ·
27
+ <a href="https://huggingface.co/AnkitAI/TinyJev-0.6B">TinyJev 0.6B</a> ·
28
  <a href="https://github.com/ankit-aglawe/tinyjev/tree/main/examples">Examples</a>
29
  </p>
30
 
 
53
 
54
  | Model | Params | OD-500 | Gate 0.85 | ms / case | Weights |
55
  |---|---:|---:|---|---:|---|
56
+ | <img src="https://raw.githubusercontent.com/ankit-aglawe/tinyjev/main/assets/logos/tinyjev.png" width="18"> **TinyJev&nbsp;0.6B** | 596M, 1.2 GB | 440 (88.0%) | 59% @ 98.0% | 85 | 🤗 [AnkitAI/TinyJev-0.6B](https://huggingface.co/AnkitAI/TinyJev-0.6B) |
57
+ | <img src="https://raw.githubusercontent.com/ankit-aglawe/tinyjev/main/assets/logos/tinyjev.png" width="18"> **TinyJev&nbsp;4B** | 4.0B, 8.0 GB | 474 (94.8%) | 87% @ 99.1% | 628 | 🤗 [AnkitAI/TinyJev-4B](https://huggingface.co/AnkitAI/TinyJev-4B) |
58
 
59
  OD-500 is correct answers out of 500. Gate 0.85 is the share of decisions answered on its own at
60
  confidence ≥ 0.85, and how often those were right. Calibration (ECE 0.071 vs 0.022), coverage at 2%
61
  error (63% vs 92%) and transfer-v4 dev (0.625 vs 0.762) are on the benchmark page. Load either with
62
+ `tinyjev.load("TinyJev-0.6B")` or `tinyjev.load("TinyJev-4B")`.
63
 
64
  Both rows are fp16. Loading with `quantize=8` keeps the same weights in half the memory and changes
65
  almost nothing: the 0.6B scores 440 at 90 ms, the 4B 473 at 845 ms, one answer in 500 different from
 
74
  | Model | Correct / 500 | Handled alone at confidence ≥ 0.85 |
75
  |---|---:|---:|
76
  | Claude Opus 5.5 (cloud, self-reported probabilities) | 496 | 477 at 100.0% |
77
+ | **TinyJev-4B** (fp16) | **474** | **437 at 99.1%** |
78
+ | TinyJev-4B at `quantize=8` | 473 | 437 at 99.1% |
79
  | Kev-0.8B (raw logits) | 463 | 186 at 100.0% |
80
+ | TinyJev-0.6B | 440 | 296 at 98.0% |
81
 
82
  353/375 on dev, 121/125 on holdout, 95% CI 0.928–0.966, ECE 0.022, Brier 0.071, coverage at 2% error 92%.
83
  On Kev's transfer-v4 dev it scores 0.762 against 0.625 for the 0.6B (Kev-4B, a full fine-tune, 0.790).
 
97
 
98
  ```python
99
  import tinyjev
100
+ agent = tinyjev.load("TinyJev-4B", quantize=8) # 4.5 GB in memory; drop quantize for fp16, 8 GB
101
 
102
  agent.predict({
103
  "state": "Shoes arrived two weeks late and in the wrong size. Also I see two charges on my card.",
 
116
  One request shape:
117
 
118
  ```bash
119
+ tinyjev serve TinyJev-4B --quantize 8 # POST /v1/systemone on 127.0.0.1:8077
120
  ```
121
 
122
  ## What is in this repo
123
 
124
+ `AutoModel.from_pretrained("AnkitAI/TinyJev-4B")` loads the backbone on its own, a standard
125
  `Qwen3Model` in fp16 with the LoRA already merged. The decision head lives in `head.safetensors`, and
126
  `tinyjev` is what turns hidden states into calibrated answers.
127
 
 
134
 
135
  | | transfer-v4 dev | decision-v7 dev | OpenDecision 500 |
136
  | --- | --- | --- | --- |
137
+ | TinyJev-4B | 0.762 | 0.859 | 474 / 500 |
138
+ | TinyJev-0.6B | 0.625 | | 440 / 500 |
139
 
140
  ## Support the Project
141