File size: 16,773 Bytes
b1d70b3 27dafe1 0121730 27dafe1 b1d70b3 27dafe1 94e324c b1d70b3 22854cb 0371baf 0121730 b1d70b3 0121730 b1d70b3 0121730 6c2b24e 0121730 6c2b24e cd41550 67c27e4 cd41550 67c27e4 cc9f11c 0121730 6c2b24e 0121730 6c2b24e 0121730 0371baf 0121730 0371baf 0121730 0371baf 0121730 0371baf 0121730 f6c649a 0121730 0371baf 7570cab a812b1a 7570cab a812b1a 0121730 b1d70b3 0121730 b1d70b3 4d27595 0121730 4d27595 b1d70b3 0121730 b1d70b3 67c27e4 0121730 b1d70b3 4d27595 b1d70b3 0121730 4d27595 0121730 4d27595 0121730 4d27595 0121730 b1d70b3 4d27595 0121730 4d27595 324698c 0121730 67c27e4 0121730 324698c 0121730 324698c 0121730 b1d70b3 0121730 b1d70b3 94e324c 0121730 b1d70b3 dd4d93f b1d70b3 b7bd874 67c27e4 ba1d1bf 67c27e4 0121730 b1d70b3 cc9f11c b1d70b3 cc9f11c | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 | ---
license: apache-2.0
base_model:
- Qwen/Qwen3.5-0.8B-Base
- jaredpalmer/kev-0.8b
base_model_relation: adapter
language:
- en
tags:
- multimodal
- vision
- decision-model
- long-context
- qwen3.5
- lora
- typesafe
model-index:
- name: JEVision visual sidecar
results:
- task:
type: image-classification
name: Visual sidecar image classification
dataset:
name: Beans validation (selected 60 images)
type: AI-Lab-Makerere/beans
split: validation subset
metrics:
- type: accuracy
name: Self-reported accuracy (51/60; %)
value: 85.0
source:
name: JEVision self-reported development reports
url: https://huggingface.co/divyanshx11/JEVision/blob/main/reports/self-reported-evaluations.json
- task:
type: visual-question-answering
name: Visual sidecar generated tasks
dataset:
name: JEVision generated visual tasks (150 development images)
type: jevvision-visual-synthetic-dev
split: development
metrics:
- type: accuracy
name: Self-reported accuracy (147/150; %)
value: 98.0
source:
name: JEVision self-reported development reports
url: https://huggingface.co/divyanshx11/JEVision/blob/main/reports/self-reported-evaluations.json
- task:
type: visual-question-answering
name: Visual sidecar spatial reasoning
dataset:
name: VSR zero-shot development (64 COCO photos)
type: visual-spatial-reasoning-zeroshot-dev
split: development
metrics:
- type: accuracy
name: Self-reported accuracy (43/64; %)
value: 67.2
source:
name: JEVision self-reported development reports
url: https://huggingface.co/divyanshx11/JEVision/blob/main/reports/self-reported-evaluations.json
- task:
type: visual-question-answering
name: Visual sidecar long-context questions
dataset:
name: JEVision generated visual tasks (30 long-context pairs)
type: jevvision-visual-synthetic-long-dev
split: development
metrics:
- type: accuracy
name: Self-reported accuracy at 66.6k tokens (27/30; %)
value: 90.0
source:
name: JEVision self-reported development reports
url: https://huggingface.co/divyanshx11/JEVision/blob/main/reports/self-reported-evaluations.json
---
# JEVision

JEVision is a Qwen3.5-based system for making structured decisions from text and images. It extends the KEV/Jev-style System One interface with a visual route, so an application can send context and typed questions and receive **Choice**, **Noul**, or **Score** answers instead of parsing free-form prose.
The release combines a trained text adapter, a separately trained visual sidecar, pointer heads, and a bundled inference runtime. Both routes use the same pinned Qwen3.5-0.8B-Base revision. This makes the model useful for workflows that need a consistent decision API across text-only and image-bearing requests.
## Capabilities
| Capability | JEVision |
|---|---|
| Input | Text context, with optional PNG, JPEG, or WebP images |
| Output | Typed Choice, Noul, and Score responses through `/v1/systemone` |
| Text route | JEVision text LoRA and pointer head |
| Image route | Visual LoRA and pointer head, selected automatically when images are present |
| Request envelope | Configured for up to 80,000 processed input tokens |
| Packaging | Adapters, heads, and a runnable local server; base weights download separately |

The capability graphic summarizes the separate input routes and shared response format. Its 80,000-token value is a configured request limit, not a measure of answer quality.
### Compared with Jev and KEV-0.8B

This view compares output format, input modality, configured context, image support, and text-only JevBench results. JEVision’s public-panel figure is 79.3%, and its 33-question general group result is 21/33 (63.6%). The [editable SVG](assets/jevvision-comparison-table.svg) is also included.
The serving route processed 76,999 text tokens and 76,998 image-plus-text tokens in recorded acceptance checks. The text check used the included KEV-0.8B option; the image check used the visual sidecar. These checks establish request handling near 77K, while long-context answer quality remains to be evaluated.
## See the visual route
The repository includes a real photograph, the exact request sent to the visual route, and its captured response. The model selected `laptop_and_coffee` from three descriptions of the scene.

Photo by [Shixart1985](https://commons.wikimedia.org/wiki/User:Shixart1985), via Wikimedia Commons, [CC BY 2.0](https://creativecommons.org/licenses/by/2.0/). [Source photograph](https://commons.wikimedia.org/wiki/File:Coffee_cup_next_to_laptop_on_wooden_table_in_cozy_indoor_workspace_during_daytime.jpg); the bundled file is its 960-pixel Commons thumbnail.
The saved [request](examples/real-photo/input.json) and [full response](examples/real-photo/output.json) can be inspected or replayed. The response below is taken from that recorded adapter run:
```json
{
"answers": {
"scene": {
"type": "choice",
"choice": "laptop_and_coffee",
"confidence": 1.0,
"probabilities": {
"laptop_and_coffee": 1.0,
"bicycle_and_helmet": 0.0,
"cat_on_sofa": 0.0
}
}
},
"usage": {"input_tokens": 674, "output_tokens": 67}
}
```
In these examples, `model: "jevvision"` is the API label echoed in the response; it does not choose the inference weights. The `images` field routes the request through the JEVision visual adapter and pointer head in this repository. The optional KEV-0.8B checkpoint is used only when selected for text-only requests. The saved photo response retains its recorded answer, probabilities, token counts, and latency; only its echoed model label was normalized from the older API alias.
This is one recorded example. The full response also records the latency of that CPU run; it is not a speed comparison. Run it yourself after starting the server:
```bash
python examples/real-photo/run_demo.py --endpoint http://127.0.0.1:8009
```
### Noul answers on the same photograph
In a separate local run using the same coffee-and-laptop photograph, the three yes/no questions in the [Noul request](examples/real-photo/noul-input.json) produced the response below:
- `is_macbook`: Does the laptop appear to be a MacBook?
- `cup_touching_laptop`: Is the cup physically touching the laptop?
- `cup_nearly_full`: Does the cup appear nearly full?
For a **Noul** answer, `noul` is the model's estimated probability of **yes**. Here, the model assigns 0.72 to the MacBook question, 0.01 to the cup touching the laptop, and 0.89 to the cup appearing nearly full. These are outputs from one local example, not calibrated confidence measurements.
```json
{
"answers": {
"is_macbook": {
"type": "noul",
"noul": 0.72
},
"cup_touching_laptop": {
"type": "noul",
"noul": 0.01
},
"cup_nearly_full": {
"type": "noul",
"noul": 0.89
}
},
"usage": {
"input_tokens": 700,
"output_tokens": 60
}
}
```
To run those questions against your local server, use:
```bash
python examples/real-photo/run_demo.py --example noul --endpoint http://127.0.0.1:8009
```
The script prints the full response and saves it to `examples/real-photo/noul-output.json`. It runs the same request, but the probability values and token counts may differ from the local result shown above.
## Run locally
The bundle was tested on Linux with CUDA and a Tesla T4. A 16 GB NVIDIA GPU is recommended when hosting both routes together. The Qwen base weights are fetched on first use.
```bash
git clone https://huggingface.co/divyanshx11/JEVision
cd JEVision
python -m venv .venv && source .venv/bin/activate
python -m pip install -r requirements.txt
python run_jevvision.py --port 8009
```
The server exposes `POST /v1/systemone` on localhost. To select the bundled KEV text checkpoint for requests without images, add `--text-adapter kev-0.8b`. Image-bearing requests automatically use the visual sidecar.
The earlier `--text-adapter jevbench-m3` selection is also available for existing scripts and points to the same text weights as the default route.
### Call from Python
```python
from jevvision import JEVision, image_file_as_data_url
model = JEVision.from_pretrained(".", device="cuda:0")
result = model.system_one(
state="Use the attached photo as visual context.",
images=[image_file_as_data_url("examples/real-photo/coffee-and-laptop.jpg")],
questions={
"scene": {
"type": "choice",
"instructions": "Which description best matches the photo?",
"criteria": {
"laptop_and_coffee": "A laptop beside a cup of coffee on a table.",
"bicycle_and_helmet": "A bicycle parked beside a helmet.",
"cat_on_sofa": "A cat sitting on a sofa.",
},
}
},
)
print(result["answers"]["scene"])
```
The image API accepts up to four images per request, each no larger than 10 MiB and 25 megapixels. Text-only requests use the selected text adapter. Image-bearing requests use the visual sidecar; the two adapters are routed separately.
## Long-context example
[The runnable request recipe](examples/long-context/input.json) builds an archived help-desk state of roughly 76,000 tokens and asks for the code in its final record. With the server running, use:
```bash
python examples/long-context/run_demo.py --endpoint http://127.0.0.1:8009
```
The runner writes its generated state and saves the server's response only after checking that `usage.input_tokens` is at least 75,000. A [separate recorded text-route acceptance summary](examples/long-context/recorded-77k-text-response-summary.json) documents a 76,999-token request on the bundled KEV-0.8B option; it is a request-handling check, not a result for the default JEVision text adapter.
The original archived [request recipe](examples/long-context/m1-77k-text-input.json) and [response summary](examples/long-context/m1-77k-text-response-summary.json) are preserved at their earlier paths.
## Architecture and training
| Included component | Purpose |
|---|---|
| `text/jevvision-text/` | Default Qwen3.5 text decision LoRA and pointer head |
| `text/kev-0.8b/` | Optional pinned [KEV-0.8B](https://huggingface.co/jaredpalmer/kev-0.8b) text checkpoint |
| `adapter/` and `pointer_head.pt` | Image-aware decision sidecar |
| `runtime/kev/` | Local inference and System One serving code |
The base is [Qwen/Qwen3.5-0.8B-Base](https://huggingface.co/Qwen/Qwen3.5-0.8B-Base) at revision `dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68`. This repository supplies the adapters and heads, not a merged base-model checkpoint.
The default text adapter was fitted to public JevBench decision examples, starting from a KEV text checkpoint. Those public examples informed training and selection, so results on them are not an independent evaluation. The visual sidecar was trained separately with photos from the [Beans training split](https://huggingface.co/datasets/AI-Lab-Makerere/beans) and programmatically generated, labeled visual tasks. The generated tasks supplement the real images; they are training data, not claimed evaluation measurements. The visual route is not scored by JevBench.
## Self-reported visual evaluations
These development results use the **visual sidecar shipped in this repository**. The adapter and pointer-head SHA-256 hashes in the recorded runs match the [release manifest](jevvision_manifest.json). Each result is a narrow task probe, not an estimate of general image accuracy.
| Evaluation | Recorded result | Scope |
|---|---:|---|
| Beans field photos | **51/60 (85.0%)** | Three-class development photos from the same dataset family used in training. A blank-image control scored 20/60; a class-rotated wrong-image control scored 6/60. |
| Generated visual tasks | **147/150 (98.0%)** | Development images from five task families represented in training, scored again after loading the saved adapter through the serving scorer. Blank images scored 53/150; substituted wrong images scored 10/150. |
| Visual Spatial Reasoning photos | **43/64 (67.2%)** | Balanced, zero-shot development set of 64 real COCO photos. Same-size blank-image controls scored 32/64. |
| Long-context visual questions | **27/30 (90.0%)** at 66,622–66,623 processed tokens, versus 29/30 on short inputs | The same 30 generated images with neutral text added to the long requests. Two answers changed; this does not establish quality across the full 80,000-token service limit. |
These are self-reported development runs. The Beans and generated-task panels share task families with training; the 64-photo VSR panel is small. The figures do not establish broad real-image performance or calibrated probabilities. They measure the visual sidecar, not the separately routed text adapter. [Recorded counts, checkpoint hashes, and run identifiers](reports/self-reported-evaluations.json) are provided for provenance.
## Evaluation scope
JEVision currently has functional checks for typed responses, image routing, and long request acceptance, plus the recorded real-photo example above. The public JevBench examples were used to fit and select the text adapter, so accuracy on those examples is not an independent benchmark result. A separate 33-question grouped text test recorded 21 correct (63.6%) with a 4,096-token cap. This small test does not measure the visual route or long-context answer quality. The 80,000-token value is a service limit, and requests beyond it are rejected instead of silently truncated. The comparisons below describe these specific text panels; they are not a broad accuracy, latency, or cost claim.

The JevBench result for JEVision is 183/231 (79.3%). The grouped test is the separate 33-question check. Neither panel evaluates the visual route.
The [manifest](jevvision_manifest.json) identifies the components, their source revisions, hashes, and serving limit. Users should assess the model on their own images and decision tasks before relying on its outputs.
## What's next
JEVision's next phase targets a purpose-built decision architecture and a faster production inference path. The work below is planned for future releases; the measurements above describe the current one.
### JEV-style parallel decision architecture
The architectural direction is to move from sequential, decoder-style question handling toward a shared-state decision engine. A common context representation would feed multiple question branches concurrently, with typed heads for **Choice**, **Noul**, and **Score** outputs. This makes multi-question calls a first-class model workload: shared context computation can be reused, independent questions can run in parallel, and the system can spend less compute repeatedly generating answer scaffolding. The goal is higher question throughput and more predictable latency as a request grows.
### Inference engineered for low latency
The serving roadmap focuses on the full request path: context prefill, question scheduling, batching, adapter dispatch, and typed-answer readout. For the current Qwen-based backbone, we plan to profile and optimize reusable context work and cache behavior; for the evolving decision architecture, we plan to schedule question branches and text/visual routes with less synchronization and data movement. The target is a responsive decision API with lower end-to-end and tail latency under multi-question load. We will report p50/p95 latency, throughput, and output-quality checks when these changes are implemented.
## License
The adapter bundle and runtime are published under Apache-2.0. The Qwen3.5 base model is also Apache-2.0; follow the base model's terms and the terms of any datasets you use.
|