File size: 16,773 Bytes
b1d70b3
 
27dafe1
 
 
0121730
 
27dafe1
b1d70b3
27dafe1
 
 
 
 
 
 
94e324c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b1d70b3
 
 
 
22854cb
0371baf
0121730
b1d70b3
0121730
b1d70b3
0121730
6c2b24e
0121730
 
 
 
 
 
 
 
6c2b24e
cd41550
67c27e4
cd41550
67c27e4
cc9f11c
 
 
 
 
 
0121730
6c2b24e
0121730
6c2b24e
0121730
0371baf
0121730
0371baf
0121730
0371baf
0121730
0371baf
0121730
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f6c649a
 
0121730
 
 
 
 
0371baf
7570cab
 
a812b1a
7570cab
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
a812b1a
 
 
 
 
 
 
 
0121730
b1d70b3
0121730
b1d70b3
4d27595
 
 
 
 
0121730
4d27595
b1d70b3
0121730
b1d70b3
67c27e4
 
0121730
b1d70b3
4d27595
 
b1d70b3
0121730
4d27595
0121730
 
4d27595
0121730
4d27595
0121730
 
 
 
 
 
b1d70b3
4d27595
 
0121730
 
 
 
 
 
 
 
 
 
 
4d27595
324698c
0121730
 
67c27e4
 
0121730
324698c
0121730
 
 
 
 
 
324698c
0121730
b1d70b3
0121730
b1d70b3
94e324c
 
 
 
 
 
 
 
 
 
 
 
 
0121730
b1d70b3
dd4d93f
b1d70b3
b7bd874
67c27e4
ba1d1bf
67c27e4
0121730
b1d70b3
cc9f11c
 
 
 
 
 
 
 
 
 
 
 
b1d70b3
 
cc9f11c
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
---
license: apache-2.0
base_model:
- Qwen/Qwen3.5-0.8B-Base
- jaredpalmer/kev-0.8b
base_model_relation: adapter
language:
- en
tags:
- multimodal
- vision
- decision-model
- long-context
- qwen3.5
- lora
- typesafe
model-index:
  - name: JEVision visual sidecar
    results:
      - task:
          type: image-classification
          name: Visual sidecar image classification
        dataset:
          name: Beans validation (selected 60 images)
          type: AI-Lab-Makerere/beans
          split: validation subset
        metrics:
          - type: accuracy
            name: Self-reported accuracy (51/60; %)
            value: 85.0
        source:
          name: JEVision self-reported development reports
          url: https://huggingface.co/divyanshx11/JEVision/blob/main/reports/self-reported-evaluations.json
      - task:
          type: visual-question-answering
          name: Visual sidecar generated tasks
        dataset:
          name: JEVision generated visual tasks (150 development images)
          type: jevvision-visual-synthetic-dev
          split: development
        metrics:
          - type: accuracy
            name: Self-reported accuracy (147/150; %)
            value: 98.0
        source:
          name: JEVision self-reported development reports
          url: https://huggingface.co/divyanshx11/JEVision/blob/main/reports/self-reported-evaluations.json
      - task:
          type: visual-question-answering
          name: Visual sidecar spatial reasoning
        dataset:
          name: VSR zero-shot development (64 COCO photos)
          type: visual-spatial-reasoning-zeroshot-dev
          split: development
        metrics:
          - type: accuracy
            name: Self-reported accuracy (43/64; %)
            value: 67.2
        source:
          name: JEVision self-reported development reports
          url: https://huggingface.co/divyanshx11/JEVision/blob/main/reports/self-reported-evaluations.json
      - task:
          type: visual-question-answering
          name: Visual sidecar long-context questions
        dataset:
          name: JEVision generated visual tasks (30 long-context pairs)
          type: jevvision-visual-synthetic-long-dev
          split: development
        metrics:
          - type: accuracy
            name: Self-reported accuracy at 66.6k tokens (27/30; %)
            value: 90.0
        source:
          name: JEVision self-reported development reports
          url: https://huggingface.co/divyanshx11/JEVision/blob/main/reports/self-reported-evaluations.json
---

# JEVision

![JEVision architecture showing text and visual adapter routes leading to typed answers](assets/jevvision-hero.png)

JEVision is a Qwen3.5-based system for making structured decisions from text and images. It extends the KEV/Jev-style System One interface with a visual route, so an application can send context and typed questions and receive **Choice**, **Noul**, or **Score** answers instead of parsing free-form prose.

The release combines a trained text adapter, a separately trained visual sidecar, pointer heads, and a bundled inference runtime. Both routes use the same pinned Qwen3.5-0.8B-Base revision. This makes the model useful for workflows that need a consistent decision API across text-only and image-bearing requests.

## Capabilities

| Capability | JEVision |
|---|---|
| Input | Text context, with optional PNG, JPEG, or WebP images |
| Output | Typed Choice, Noul, and Score responses through `/v1/systemone` |
| Text route | JEVision text LoRA and pointer head |
| Image route | Visual LoRA and pointer head, selected automatically when images are present |
| Request envelope | Configured for up to 80,000 processed input tokens |
| Packaging | Adapters, heads, and a runnable local server; base weights download separately |

![JEVision capability overview showing its text route, visual sidecar, and typed response interface](assets/jevvision-capability-comparison.png)

The capability graphic summarizes the separate input routes and shared response format. Its 80,000-token value is a configured request limit, not a measure of answer quality.

### Compared with Jev and KEV-0.8B

![Side-by-side capability and text benchmark comparison for Jev, KEV-0.8B, and JEVision](assets/jevvision-comparison-table.png)

This view compares output format, input modality, configured context, image support, and text-only JevBench results. JEVision’s public-panel figure is 79.3%, and its 33-question general group result is 21/33 (63.6%). The [editable SVG](assets/jevvision-comparison-table.svg) is also included.

The serving route processed 76,999 text tokens and 76,998 image-plus-text tokens in recorded acceptance checks. The text check used the included KEV-0.8B option; the image check used the visual sidecar. These checks establish request handling near 77K, while long-context answer quality remains to be evaluated.

## See the visual route

The repository includes a real photograph, the exact request sent to the visual route, and its captured response. The model selected `laptop_and_coffee` from three descriptions of the scene.

![A cup of coffee beside a laptop on a wooden table](examples/real-photo/coffee-and-laptop.jpg)

Photo by [Shixart1985](https://commons.wikimedia.org/wiki/User:Shixart1985), via Wikimedia Commons, [CC BY 2.0](https://creativecommons.org/licenses/by/2.0/). [Source photograph](https://commons.wikimedia.org/wiki/File:Coffee_cup_next_to_laptop_on_wooden_table_in_cozy_indoor_workspace_during_daytime.jpg); the bundled file is its 960-pixel Commons thumbnail.

The saved [request](examples/real-photo/input.json) and [full response](examples/real-photo/output.json) can be inspected or replayed. The response below is taken from that recorded adapter run:

```json
{
  "answers": {
    "scene": {
      "type": "choice",
      "choice": "laptop_and_coffee",
      "confidence": 1.0,
      "probabilities": {
        "laptop_and_coffee": 1.0,
        "bicycle_and_helmet": 0.0,
        "cat_on_sofa": 0.0
      }
    }
  },
  "usage": {"input_tokens": 674, "output_tokens": 67}
}
```

In these examples, `model: "jevvision"` is the API label echoed in the response; it does not choose the inference weights. The `images` field routes the request through the JEVision visual adapter and pointer head in this repository. The optional KEV-0.8B checkpoint is used only when selected for text-only requests. The saved photo response retains its recorded answer, probabilities, token counts, and latency; only its echoed model label was normalized from the older API alias.

This is one recorded example. The full response also records the latency of that CPU run; it is not a speed comparison. Run it yourself after starting the server:

```bash
python examples/real-photo/run_demo.py --endpoint http://127.0.0.1:8009
```

### Noul answers on the same photograph

In a separate local run using the same coffee-and-laptop photograph, the three yes/no questions in the [Noul request](examples/real-photo/noul-input.json) produced the response below:

- `is_macbook`: Does the laptop appear to be a MacBook?
- `cup_touching_laptop`: Is the cup physically touching the laptop?
- `cup_nearly_full`: Does the cup appear nearly full?

For a **Noul** answer, `noul` is the model's estimated probability of **yes**. Here, the model assigns 0.72 to the MacBook question, 0.01 to the cup touching the laptop, and 0.89 to the cup appearing nearly full. These are outputs from one local example, not calibrated confidence measurements.

```json
{
  "answers": {
    "is_macbook": {
      "type": "noul",
      "noul": 0.72
    },
    "cup_touching_laptop": {
      "type": "noul",
      "noul": 0.01
    },
    "cup_nearly_full": {
      "type": "noul",
      "noul": 0.89
    }
  },
  "usage": {
    "input_tokens": 700,
    "output_tokens": 60
  }
}
```

To run those questions against your local server, use:

```bash
python examples/real-photo/run_demo.py --example noul --endpoint http://127.0.0.1:8009
```

The script prints the full response and saves it to `examples/real-photo/noul-output.json`. It runs the same request, but the probability values and token counts may differ from the local result shown above.

## Run locally

The bundle was tested on Linux with CUDA and a Tesla T4. A 16 GB NVIDIA GPU is recommended when hosting both routes together. The Qwen base weights are fetched on first use.

```bash
git clone https://huggingface.co/divyanshx11/JEVision
cd JEVision
python -m venv .venv && source .venv/bin/activate
python -m pip install -r requirements.txt
python run_jevvision.py --port 8009
```

The server exposes `POST /v1/systemone` on localhost. To select the bundled KEV text checkpoint for requests without images, add `--text-adapter kev-0.8b`. Image-bearing requests automatically use the visual sidecar.

The earlier `--text-adapter jevbench-m3` selection is also available for existing scripts and points to the same text weights as the default route.

### Call from Python

```python
from jevvision import JEVision, image_file_as_data_url

model = JEVision.from_pretrained(".", device="cuda:0")
result = model.system_one(
    state="Use the attached photo as visual context.",
    images=[image_file_as_data_url("examples/real-photo/coffee-and-laptop.jpg")],
    questions={
        "scene": {
            "type": "choice",
            "instructions": "Which description best matches the photo?",
            "criteria": {
                "laptop_and_coffee": "A laptop beside a cup of coffee on a table.",
                "bicycle_and_helmet": "A bicycle parked beside a helmet.",
                "cat_on_sofa": "A cat sitting on a sofa.",
            },
        }
    },
)
print(result["answers"]["scene"])
```

The image API accepts up to four images per request, each no larger than 10 MiB and 25 megapixels. Text-only requests use the selected text adapter. Image-bearing requests use the visual sidecar; the two adapters are routed separately.

## Long-context example

[The runnable request recipe](examples/long-context/input.json) builds an archived help-desk state of roughly 76,000 tokens and asks for the code in its final record. With the server running, use:

```bash
python examples/long-context/run_demo.py --endpoint http://127.0.0.1:8009
```

The runner writes its generated state and saves the server's response only after checking that `usage.input_tokens` is at least 75,000. A [separate recorded text-route acceptance summary](examples/long-context/recorded-77k-text-response-summary.json) documents a 76,999-token request on the bundled KEV-0.8B option; it is a request-handling check, not a result for the default JEVision text adapter.

The original archived [request recipe](examples/long-context/m1-77k-text-input.json) and [response summary](examples/long-context/m1-77k-text-response-summary.json) are preserved at their earlier paths.

## Architecture and training

| Included component | Purpose |
|---|---|
| `text/jevvision-text/` | Default Qwen3.5 text decision LoRA and pointer head |
| `text/kev-0.8b/` | Optional pinned [KEV-0.8B](https://huggingface.co/jaredpalmer/kev-0.8b) text checkpoint |
| `adapter/` and `pointer_head.pt` | Image-aware decision sidecar |
| `runtime/kev/` | Local inference and System One serving code |

The base is [Qwen/Qwen3.5-0.8B-Base](https://huggingface.co/Qwen/Qwen3.5-0.8B-Base) at revision `dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68`. This repository supplies the adapters and heads, not a merged base-model checkpoint.

The default text adapter was fitted to public JevBench decision examples, starting from a KEV text checkpoint. Those public examples informed training and selection, so results on them are not an independent evaluation. The visual sidecar was trained separately with photos from the [Beans training split](https://huggingface.co/datasets/AI-Lab-Makerere/beans) and programmatically generated, labeled visual tasks. The generated tasks supplement the real images; they are training data, not claimed evaluation measurements. The visual route is not scored by JevBench.

## Self-reported visual evaluations

These development results use the **visual sidecar shipped in this repository**. The adapter and pointer-head SHA-256 hashes in the recorded runs match the [release manifest](jevvision_manifest.json). Each result is a narrow task probe, not an estimate of general image accuracy.

| Evaluation | Recorded result | Scope |
|---|---:|---|
| Beans field photos | **51/60 (85.0%)** | Three-class development photos from the same dataset family used in training. A blank-image control scored 20/60; a class-rotated wrong-image control scored 6/60. |
| Generated visual tasks | **147/150 (98.0%)** | Development images from five task families represented in training, scored again after loading the saved adapter through the serving scorer. Blank images scored 53/150; substituted wrong images scored 10/150. |
| Visual Spatial Reasoning photos | **43/64 (67.2%)** | Balanced, zero-shot development set of 64 real COCO photos. Same-size blank-image controls scored 32/64. |
| Long-context visual questions | **27/30 (90.0%)** at 66,622–66,623 processed tokens, versus 29/30 on short inputs | The same 30 generated images with neutral text added to the long requests. Two answers changed; this does not establish quality across the full 80,000-token service limit. |

These are self-reported development runs. The Beans and generated-task panels share task families with training; the 64-photo VSR panel is small. The figures do not establish broad real-image performance or calibrated probabilities. They measure the visual sidecar, not the separately routed text adapter. [Recorded counts, checkpoint hashes, and run identifiers](reports/self-reported-evaluations.json) are provided for provenance.

## Evaluation scope

JEVision currently has functional checks for typed responses, image routing, and long request acceptance, plus the recorded real-photo example above. The public JevBench examples were used to fit and select the text adapter, so accuracy on those examples is not an independent benchmark result. A separate 33-question grouped text test recorded 21 correct (63.6%) with a 4,096-token cap. This small test does not measure the visual route or long-context answer quality. The 80,000-token value is a service limit, and requests beyond it are rejected instead of silently truncated. The comparisons below describe these specific text panels; they are not a broad accuracy, latency, or cost claim.

![Text JevBench comparison of JEVision, Jev, and KEV on the public training panel and grouped test](assets/JevBench-results.png)

The JevBench result for JEVision is 183/231 (79.3%). The grouped test is the separate 33-question check. Neither panel evaluates the visual route.

The [manifest](jevvision_manifest.json) identifies the components, their source revisions, hashes, and serving limit. Users should assess the model on their own images and decision tasks before relying on its outputs.

## What's next

JEVision's next phase targets a purpose-built decision architecture and a faster production inference path. The work below is planned for future releases; the measurements above describe the current one.

### JEV-style parallel decision architecture

The architectural direction is to move from sequential, decoder-style question handling toward a shared-state decision engine. A common context representation would feed multiple question branches concurrently, with typed heads for **Choice**, **Noul**, and **Score** outputs. This makes multi-question calls a first-class model workload: shared context computation can be reused, independent questions can run in parallel, and the system can spend less compute repeatedly generating answer scaffolding. The goal is higher question throughput and more predictable latency as a request grows.

### Inference engineered for low latency

The serving roadmap focuses on the full request path: context prefill, question scheduling, batching, adapter dispatch, and typed-answer readout. For the current Qwen-based backbone, we plan to profile and optimize reusable context work and cache behavior; for the evolving decision architecture, we plan to schedule question branches and text/visual routes with less synchronization and data movement. The target is a responsive decision API with lower end-to-end and tail latency under multi-question load. We will report p50/p95 latency, throughput, and output-quality checks when these changes are implemented.

## License

The adapter bundle and runtime are published under Apache-2.0. The Qwen3.5 base model is also Apache-2.0; follow the base model's terms and the terms of any datasets you use.