Add JEVision model card
Browse files
README.md
ADDED
|
@@ -0,0 +1,88 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
base_model: Qwen/Qwen3.5-0.8B-Base
|
| 4 |
+
base_model_relation: adapter
|
| 5 |
+
language:
|
| 6 |
+
- en
|
| 7 |
+
tags:
|
| 8 |
+
- multimodal
|
| 9 |
+
- vision
|
| 10 |
+
- decision-model
|
| 11 |
+
- long-context
|
| 12 |
+
- qwen3.5
|
| 13 |
+
- lora
|
| 14 |
+
- typesafe
|
| 15 |
+
---
|
| 16 |
+
|
| 17 |
+
# JEVision
|
| 18 |
+
|
| 19 |
+
### Give the KEV-style decision interface a pair of eyes.
|
| 20 |
+
|
| 21 |
+
**Images in. Structured decisions out. A 77K-token proof point.**
|
| 22 |
+
|
| 23 |
+
JEVision is an experimental multimodal extension of the KEV-style decision workflow. It brings image input into a typed, machine-readable decision path built on Qwen3.5-0.8B-Base. Instead of asking for an open-ended paragraph, provide a state, an image, and a concrete question—and get a structured answer through the familiar **/v1/systemone** route.
|
| 24 |
+
|
| 25 |
+
## Why JEVision
|
| 26 |
+
|
| 27 |
+
- **Vision meets decisions.** Image and text go through the same KEV-shaped request/response path.
|
| 28 |
+
- **Built for answers software can use.** The tested image path returns a structured Choice result, not free-form chat.
|
| 29 |
+
- **Long context, demonstrated.** Image-plus-text requests passed at **76,998 tokens** in the POC.
|
| 30 |
+
- **A practical first step.** This release publishes the visual adapter and pointer head needed to extend the public KEV text checkpoint.
|
| 31 |
+
|
| 32 |
+
## What we verified
|
| 33 |
+
|
| 34 |
+
The milestone tests exercised the real **/v1/systemone** handler, not a model-only shortcut:
|
| 35 |
+
|
| 36 |
+
| Check | Result |
|
| 37 |
+
|---|---|
|
| 38 |
+
| Text-only request | 76,999 input tokens; accepted with exact preflight token accounting |
|
| 39 |
+
| Image + text request | 76,998 input tokens; accepted with the same KEV response envelope |
|
| 40 |
+
| Image-content smoke | At 77,003 tokens, swapping only a generated red image for a generated blue image flipped the selected answer from red to blue |
|
| 41 |
+
| Oversized request | Inputs above the 80,000-token service cap were rejected rather than silently truncated |
|
| 42 |
+
|
| 43 |
+
That is an end-to-end proof that the image path can affect a structured decision at long context. It is an intentionally small functional smoke test—not a claim of broad visual understanding.
|
| 44 |
+
|
| 45 |
+
## How the package fits together
|
| 46 |
+
|
| 47 |
+
JEVision is a **two-checkpoint composition**, not a merged standalone model:
|
| 48 |
+
|
| 49 |
+
1. **Visual base:** [Qwen3.5-0.8B-Base](https://huggingface.co/Qwen/Qwen3.5-0.8B-Base), pinned to revision `dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68`.
|
| 50 |
+
2. **KEV text checkpoint:** [jaredpalmer/kev-0.8b](https://huggingface.co/jaredpalmer/kev-0.8b/tree/54f4f8777356cd5bbbb6c6919c657f26e6f2f6d8), pinned to revision `54f4f8777356cd5bbbb6c6919c657f26e6f2f6d8`.
|
| 51 |
+
3. **This repository:** a visual LoRA adapter plus a small pointer head.
|
| 52 |
+
4. **Serving layer:** compatible KEV runtime composes the checkpoints and exposes **/v1/systemone**.
|
| 53 |
+
|
| 54 |
+
The base-model weights and KEV text checkpoint are separate downloads; they are not duplicated here. The manifest records the pinned components and artifact hashes.
|
| 55 |
+
|
| 56 |
+
## Request shape
|
| 57 |
+
|
| 58 |
+
With the compatible multimodal KEV runtime running, an image request uses a base64 data URL:
|
| 59 |
+
|
| 60 |
+
POST /v1/systemone
|
| 61 |
+
{
|
| 62 |
+
"model": "kev-latest",
|
| 63 |
+
"state": "Which option matches the attached image?",
|
| 64 |
+
"images": [{"data_url": "data:image/png;base64,<BASE64_PNG>"}],
|
| 65 |
+
"questions": {
|
| 66 |
+
"answer": {
|
| 67 |
+
"type": "choice",
|
| 68 |
+
"instructions": "Which color is shown?",
|
| 69 |
+
"criteria": {"red": "The image is red", "blue": "The image is blue"}
|
| 70 |
+
}
|
| 71 |
+
}
|
| 72 |
+
}
|
| 73 |
+
|
| 74 |
+
## What this release is—and is not
|
| 75 |
+
|
| 76 |
+
The Hub repo contains the visual sidecar artifacts. It is **not** a full Qwen checkpoint, a merged model, or a drop-in **AutoModel.from_pretrained** package. Running the POC requires the matching image-enabled KEV serving code as well as the two referenced checkpoints; turnkey public runtime packaging is a separate step.
|
| 77 |
+
|
| 78 |
+
The current evidence is deliberately narrow:
|
| 79 |
+
|
| 80 |
+
- The 80,000-token service cap was tested near 77K; **128K context is not claimed**.
|
| 81 |
+
- Image sensitivity was checked with two generated 32×32 solid-color images; **real-photo accuracy and broad image understanding have not been measured**.
|
| 82 |
+
- **No JevBench comparison, accuracy parity, latency advantage, or cost advantage is claimed yet.**
|
| 83 |
+
|
| 84 |
+
The next evaluation milestone is a JevBench comparison. Accuracy work and inference optimization follow after that baseline.
|
| 85 |
+
|
| 86 |
+
## License
|
| 87 |
+
|
| 88 |
+
Apache-2.0. The referenced Qwen base is also listed under Apache-2.0.
|