divyanshx11 commited on
Commit
cd41550
·
verified ·
1 Parent(s): 22854cb

Clarify JEVision evaluation claims and update graphics

Browse files
README.md CHANGED
@@ -33,9 +33,9 @@ The release combines a trained text adapter, a separately trained visual sidecar
33
  | Request envelope | Configured for up to 80,000 processed input tokens |
34
  | Packaging | Adapters, heads, and a runnable local server; base weights download separately |
35
 
36
- ![Jev, KEV-0.8B and JEVision capability comparison](assets/jevvision-capability-comparison.png)
37
 
38
- This earlier comparison graphic shows the input routes and context limits alongside text-only development results. Its public-panel score uses examples seen during training.
39
 
40
  The serving route processed 76,999 text tokens and 76,998 image-plus-text tokens in recorded acceptance checks. The text check used the included KEV-0.8B option; the image check used the visual sidecar. These checks establish request handling near 77K, while long-context answer quality remains to be evaluated.
41
 
@@ -142,11 +142,11 @@ The default text adapter was fitted to public JevBench decision examples, starti
142
 
143
  ## Evaluation scope
144
 
145
- JEVision currently has functional checks for typed responses, image routing, and long request acceptance, plus the recorded real-photo example above. Independent, broad image accuracy and long-context answer quality have not yet been measured. The 80,000-token value is a service limit, and requests beyond it are rejected instead of silently truncated. No comparative accuracy, latency, or cost claim is made here.
146
 
147
- ![Text-only public and grouped test results from the earlier release](assets/jevbench-m3-results.png)
148
 
149
- This preserved results graphic separates the in-sample public panel from a small grouped holdout. It does not evaluate the visual route.
150
 
151
  The [manifest](jevvision_manifest.json) identifies the components, their source revisions, hashes, and serving limit. Users should assess the model on their own images and decision tasks before relying on its outputs.
152
 
 
33
  | Request envelope | Configured for up to 80,000 processed input tokens |
34
  | Packaging | Adapters, heads, and a runnable local server; base weights download separately |
35
 
36
+ ![JEVision capability overview showing its text route, visual sidecar, and typed response interface](assets/jevvision-capability-comparison.png)
37
 
38
+ The capability graphic summarizes the separate input routes and shared response format. Its 80,000-token value is a configured request limit, not a measure of answer quality.
39
 
40
  The serving route processed 76,999 text tokens and 76,998 image-plus-text tokens in recorded acceptance checks. The text check used the included KEV-0.8B option; the image check used the visual sidecar. These checks establish request handling near 77K, while long-context answer quality remains to be evaluated.
41
 
 
142
 
143
  ## Evaluation scope
144
 
145
+ JEVision currently has functional checks for typed responses, image routing, and long request acceptance, plus the recorded real-photo example above. The public JevBench examples were used to fit and select the text adapter, so accuracy on those examples is not an independent benchmark result. A separate 33-question grouped text test recorded 21 correct (63.6%) with a 4,096-token cap. This small test does not measure the visual route or long-context answer quality. The 80,000-token value is a service limit, and requests beyond it are rejected instead of silently truncated. No comparative accuracy, latency, or cost claim is made here.
146
 
147
+ ![JEVision text evaluation scope and the recorded 21 of 33 grouped-test result](assets/jevbench-m3-results.png)
148
 
149
+ This results graphic records the small grouped text test and distinguishes it from the public examples used during training and selection. It does not evaluate the visual route.
150
 
151
  The [manifest](jevvision_manifest.json) identifies the components, their source revisions, hashes, and serving limit. Users should assess the model on their own images and decision tasks before relying on its outputs.
152
 
assets/jevbench-m3-results.png CHANGED

Git LFS Details

  • SHA256: b0a7ec3a920eea19ef76bc0633250bc7b98d6fbf8bef9123045a60b6cead8223
  • Pointer size: 131 Bytes
  • Size of remote file: 527 kB

Git LFS Details

  • SHA256: 0f8953d4d4c141bb6178852fb0c1c1b3d7da8f88802945bc1d1bce1831ca7d90
  • Pointer size: 130 Bytes
  • Size of remote file: 91.8 kB
assets/jevbench-m3-results.svg CHANGED
assets/jevvision-capability-comparison.png CHANGED

Git LFS Details

  • SHA256: 0e927c45f1056f1e5c929fa932aec7d66b15adf481e6978fb8a2a8d71d2664ce
  • Pointer size: 131 Bytes
  • Size of remote file: 593 kB

Git LFS Details

  • SHA256: aae08fb033b8d08c767f09b1bb5104235a54fb75017e123d8de113225083d860
  • Pointer size: 130 Bytes
  • Size of remote file: 97.4 kB
assets/jevvision-capability-comparison.svg CHANGED