divyanshx11 commited on
Commit
b1d70b3
·
verified ·
1 Parent(s): c9e531e

Add JEVision model card

Browse files
Files changed (1) hide show
  1. README.md +88 -0
README.md ADDED
@@ -0,0 +1,88 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: Qwen/Qwen3.5-0.8B-Base
4
+ base_model_relation: adapter
5
+ language:
6
+ - en
7
+ tags:
8
+ - multimodal
9
+ - vision
10
+ - decision-model
11
+ - long-context
12
+ - qwen3.5
13
+ - lora
14
+ - typesafe
15
+ ---
16
+
17
+ # JEVision
18
+
19
+ ### Give the KEV-style decision interface a pair of eyes.
20
+
21
+ **Images in. Structured decisions out. A 77K-token proof point.**
22
+
23
+ JEVision is an experimental multimodal extension of the KEV-style decision workflow. It brings image input into a typed, machine-readable decision path built on Qwen3.5-0.8B-Base. Instead of asking for an open-ended paragraph, provide a state, an image, and a concrete question—and get a structured answer through the familiar **/v1/systemone** route.
24
+
25
+ ## Why JEVision
26
+
27
+ - **Vision meets decisions.** Image and text go through the same KEV-shaped request/response path.
28
+ - **Built for answers software can use.** The tested image path returns a structured Choice result, not free-form chat.
29
+ - **Long context, demonstrated.** Image-plus-text requests passed at **76,998 tokens** in the POC.
30
+ - **A practical first step.** This release publishes the visual adapter and pointer head needed to extend the public KEV text checkpoint.
31
+
32
+ ## What we verified
33
+
34
+ The milestone tests exercised the real **/v1/systemone** handler, not a model-only shortcut:
35
+
36
+ | Check | Result |
37
+ |---|---|
38
+ | Text-only request | 76,999 input tokens; accepted with exact preflight token accounting |
39
+ | Image + text request | 76,998 input tokens; accepted with the same KEV response envelope |
40
+ | Image-content smoke | At 77,003 tokens, swapping only a generated red image for a generated blue image flipped the selected answer from red to blue |
41
+ | Oversized request | Inputs above the 80,000-token service cap were rejected rather than silently truncated |
42
+
43
+ That is an end-to-end proof that the image path can affect a structured decision at long context. It is an intentionally small functional smoke test—not a claim of broad visual understanding.
44
+
45
+ ## How the package fits together
46
+
47
+ JEVision is a **two-checkpoint composition**, not a merged standalone model:
48
+
49
+ 1. **Visual base:** [Qwen3.5-0.8B-Base](https://huggingface.co/Qwen/Qwen3.5-0.8B-Base), pinned to revision `dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68`.
50
+ 2. **KEV text checkpoint:** [jaredpalmer/kev-0.8b](https://huggingface.co/jaredpalmer/kev-0.8b/tree/54f4f8777356cd5bbbb6c6919c657f26e6f2f6d8), pinned to revision `54f4f8777356cd5bbbb6c6919c657f26e6f2f6d8`.
51
+ 3. **This repository:** a visual LoRA adapter plus a small pointer head.
52
+ 4. **Serving layer:** compatible KEV runtime composes the checkpoints and exposes **/v1/systemone**.
53
+
54
+ The base-model weights and KEV text checkpoint are separate downloads; they are not duplicated here. The manifest records the pinned components and artifact hashes.
55
+
56
+ ## Request shape
57
+
58
+ With the compatible multimodal KEV runtime running, an image request uses a base64 data URL:
59
+
60
+ POST /v1/systemone
61
+ {
62
+ "model": "kev-latest",
63
+ "state": "Which option matches the attached image?",
64
+ "images": [{"data_url": "data:image/png;base64,<BASE64_PNG>"}],
65
+ "questions": {
66
+ "answer": {
67
+ "type": "choice",
68
+ "instructions": "Which color is shown?",
69
+ "criteria": {"red": "The image is red", "blue": "The image is blue"}
70
+ }
71
+ }
72
+ }
73
+
74
+ ## What this release is—and is not
75
+
76
+ The Hub repo contains the visual sidecar artifacts. It is **not** a full Qwen checkpoint, a merged model, or a drop-in **AutoModel.from_pretrained** package. Running the POC requires the matching image-enabled KEV serving code as well as the two referenced checkpoints; turnkey public runtime packaging is a separate step.
77
+
78
+ The current evidence is deliberately narrow:
79
+
80
+ - The 80,000-token service cap was tested near 77K; **128K context is not claimed**.
81
+ - Image sensitivity was checked with two generated 32×32 solid-color images; **real-photo accuracy and broad image understanding have not been measured**.
82
+ - **No JevBench comparison, accuracy parity, latency advantage, or cost advantage is claimed yet.**
83
+
84
+ The next evaluation milestone is a JevBench comparison. Accuracy work and inference optimization follow after that baseline.
85
+
86
+ ## License
87
+
88
+ Apache-2.0. The referenced Qwen base is also listed under Apache-2.0.