mertkayacs commited on
Commit
66d38b0
·
verified ·
1 Parent(s): e439963

card: results as charts (Jev 1.13, Kev-4B, Laya), 4K share card

Browse files
Files changed (2) hide show
  1. README.md +3 -3
  2. assets/jev.png +2 -2
README.md CHANGED
@@ -20,7 +20,7 @@ An English decision model with the Jev API. You send a state and typed questions
20
 
21
  <video controls playsinline preload="none" poster="https://huggingface.co/datasets/mertkayacs/emberwick-videos/resolve/main/film/jevalt-film-en.jpg" src="https://huggingface.co/datasets/mertkayacs/emberwick-videos/resolve/main/film/jevalt-film-en-1080p.mp4"></video>
22
 
23
- *The one-minute film, sound on: three mistakes small decision models make and how JevAlt fixes each one.*
24
 
25
  ## Try it
26
 
@@ -68,12 +68,12 @@ client = TypeSafeClient(api_key="local", base_url="http://127.0.0.1:8000")
68
 
69
  ![Hidden instructions, option order, long policies and negated questions: JevAlt against Intern-Decision-4B, Kev-4B and Laya](https://huggingface.co/mertkayacs/Deem-4B/resolve/main/assets/fixes.png)
70
 
71
- Same items and client for every model, each as shipped: [Kev-4B](https://huggingface.co/jaredpalmer/kev-4b) r10 and [Laya](https://huggingface.co/convaiinnovations/laya) 0.3.22 on their own servers with their own calibration. Jev 1.13 rows come from [TypeSafe's notes](https://docs.typesafe.ai/model-jaggedness/jev-1.13) and an [independent audit](https://github.com/jujumilk3/jev-calibration-audit/blob/main/FINDINGS.md). The held-out tests come from JevAlt's own data pipeline, so they favour JevAlt. Kev-4B and Laya both do better on long padding; Laya is far smaller and faster. Every number and every decision: [results/comparison](https://huggingface.co/datasets/mertkayacs/jevalt-bench/tree/main/results/comparison).
72
 
73
  <details>
74
  <summary><b>Significance and caveats</b></summary>
75
 
76
- ![What Jev 1.13 lacks and JevAlt has: thinking when unsure, an unknown answer, coverage sets, native Turkish and German, open weights, repeatable answers](https://huggingface.co/mertkayacs/Deem-4B/resolve/main/assets/jev.png)
77
 
78
  The three test splits went through the same pipeline as the training rows, so they measure what the training aimed at. On the English split Deem-4B gains 4.4 accuracy points (paired bootstrap, 95% interval +3.4 to +5.5) and lowers Brier by 0.075; the Turkish and German splits move by +4.8 and +12.1 points. On JevBench-hard, TurkishMMLU and GermEval, which the training never saw, and on the typed-decisions test split (its train split was in the mix), accuracy does not change significantly, and Brier gets slightly worse on typed-decisions (+0.013) and 10kGNAD (+0.047). With `reasoning: "auto"` the English date, number and policy test rows go from 0.761 to 0.769 accuracy, an interval that touches zero.
79
 
 
20
 
21
  <video controls playsinline preload="none" poster="https://huggingface.co/datasets/mertkayacs/emberwick-videos/resolve/main/film/jevalt-film-en.jpg" src="https://huggingface.co/datasets/mertkayacs/emberwick-videos/resolve/main/film/jevalt-film-en-1080p.mp4"></video>
22
 
23
+ *The 53-second film, sound on: two mistakes small decision models make and how JevAlt fixes each one.*
24
 
25
  ## Try it
26
 
 
68
 
69
  ![Hidden instructions, option order, long policies and negated questions: JevAlt against Intern-Decision-4B, Kev-4B and Laya](https://huggingface.co/mertkayacs/Deem-4B/resolve/main/assets/fixes.png)
70
 
71
+ Same items and client for every model, each as shipped: [Kev-4B](https://huggingface.co/jaredpalmer/kev-4b) r10 and [Laya](https://huggingface.co/convaiinnovations/laya) 0.3.22 on their own servers with their own calibration. Jev 1.13 rows come from TypeSafe's [API reference](https://docs.typesafe.ai/api) and [Models page](https://docs.typesafe.ai/models). The held-out tests come from JevAlt's own data pipeline, so they favour JevAlt. Kev-4B and Laya both do better on long padding; Laya is far smaller and faster. Every number and every decision: [results/comparison](https://huggingface.co/datasets/mertkayacs/jevalt-bench/tree/main/results/comparison).
72
 
73
  <details>
74
  <summary><b>Significance and caveats</b></summary>
75
 
76
+ ![What Jev 1.13 lacks and JevAlt has: thinking when unsure, coverage sets, models made for Turkish and German, open weights](https://huggingface.co/mertkayacs/Deem-4B/resolve/main/assets/jev.png)
77
 
78
  The three test splits went through the same pipeline as the training rows, so they measure what the training aimed at. On the English split Deem-4B gains 4.4 accuracy points (paired bootstrap, 95% interval +3.4 to +5.5) and lowers Brier by 0.075; the Turkish and German splits move by +4.8 and +12.1 points. On JevBench-hard, TurkishMMLU and GermEval, which the training never saw, and on the typed-decisions test split (its train split was in the mix), accuracy does not change significantly, and Brier gets slightly worse on typed-decisions (+0.013) and 10kGNAD (+0.047). With `reasoning: "auto"` the English date, number and policy test rows go from 0.761 to 0.769 accuracy, an interval that touches zero.
79
 
assets/jev.png CHANGED

Git LFS Details

  • SHA256: ed600d0e951c76884d39b2d4f5e5e6bf9f92585eeed465302486765476d6c148
  • Pointer size: 131 Bytes
  • Size of remote file: 237 kB

Git LFS Details

  • SHA256: 921f412806de51d4e36e458626ef37a6683f890357ebf64dd2c9858606b136b3
  • Pointer size: 131 Bytes
  • Size of remote file: 187 kB