divyanshx11 commited on
Commit
0121730
·
verified ·
1 Parent(s): 6c2b24e

Present JEVision architecture and restore verified examples

Browse files
README.md CHANGED
@@ -1,7 +1,9 @@
1
  ---
2
  license: apache-2.0
3
- base_model:
4
- - Qwen/Qwen3.5-0.8B-Base
 
 
5
  tags:
6
  - multimodal
7
  - vision
@@ -14,89 +16,127 @@ tags:
14
 
15
  # JEVision
16
 
17
- ![JEVision brand artwork: image signals flowing into structured decision paths](assets/jevvision-hero.png)
18
 
19
- JEVision brings images into a KEV/Jev-style decision interface: send a state, optional images, and typed questions; receive structured **Choice**, **Noul**, or **Score** answers.
20
 
21
- It is a ready-to-run composition of compatible adapters and a bundled inference runtime—not a single merged checkpoint and not a `transformers.AutoModel`. The two branches share the pinned Qwen3.5-0.8B-Base model. Text-only requests use a selectable text adapter; requests that include images route to a separate visual adapter and pointer head.
22
 
23
- ## At a glance
24
 
25
- ![Jev, KEV-0.8B and JEVision comparison: typed outputs, text/image input routes, context limits, and text-only JevBench scores](assets/jevvision-capability-comparison.png)
 
 
 
 
 
 
 
26
 
27
- The comparison distinguishes documented or configured limits from tested quality. Jev's current docs specify 64K tokens per request (with a 32K budget for the state plus longest question) and text-only input; KEV documents 8,192 tokens for the state and 8,192 for each question branch. JEVision is configured for an 80K request maximum—not 128K. Its earlier ~77K acceptance checks exercised the KEV-0.8B text route and visual sidecar separately; those were request-path smoke checks, not long-context accuracy tests of M3. M3's JevBench results are text-only, and general image accuracy has not been measured.
28
 
29
- Sources: [Jev model limits and input contract](https://docs.typesafe.ai/models), [KEV serving details](https://github.com/jaredpalmer/kev/blob/main/README.md), and the [JEVision release configuration](jevvision_manifest.json).
30
 
31
- ## JevBench v1.3: public fit and grouped holdout
32
 
33
- ![JEVision M3 JevBench results: public-panel in-sample score beside the grouped holdout](assets/jevbench-m3-results.png)
34
 
35
- The **M3 text adapter** scored **231/231 (100%)** on the JevBench public labels used to train and select it. That is a benchmark-directed, **in-sample** score—not an independent estimate. On the separate 33-item grouped holdout, M3 scored **21/33 (63.6%)**, compared with Jev's **29/33 (87.9%)** and the released KEV-0.8B parent adapter's **17/33 (51.5%)**. M3 is not at Jev parity on the holdout.
36
 
37
- | Evaluation set | JEVision M3 text adapter | Jev | KEV-0.8B parent |
38
- |---|---:|---:|---:|
39
- | JevBench public panel, n=231 *(in-sample for M3)* | 231/231 (100.0%) | 200/231 (86.6%) | 139/231 (60.2%) |
40
- | Grouped holdout, n=33 | 21/33 (63.6%) | 29/33 (87.9%) | 17/33 (51.5%) |
41
 
42
- The raw `Qwen/Qwen3.5-0.8B-Base` model was **not separately measured** in this run; the parent comparison is KEV-0.8B on that base. These are text-only typed-decision scores. JevBench does not evaluate JEVision's separate image branch, and these results say nothing about image accuracy or long-context quality. The grouped split used 162 train / 36 development / 33 test items, a 4,096-token limit, and one final test evaluation; no sealed or imported panel was accessed.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
43
 
44
- ## Start the server
45
 
46
- Tested on Linux with CUDA and a Tesla T4. A 16 GB NVIDIA GPU is recommended for the co-hosted branches. The Qwen base weights are downloaded from Hugging Face on first use; they are not duplicated in this repository.
47
 
48
  ```bash
49
  git clone https://huggingface.co/divyanshx11/JEVision
50
  cd JEVision
51
  python -m venv .venv && source .venv/bin/activate
52
  python -m pip install -r requirements.txt
53
- python run_jevvision.py --text-adapter jevbench-m3 --port 8009
54
  ```
55
 
56
- The server exposes the JEV/KEV-shaped `POST /v1/systemone` endpoint on localhost. Use `--text-adapter kev-0.8b` to switch to the bundled baseline text checkpoint. Do not expose the unauthenticated local server directly to the internet.
57
 
58
- ### Import from Python
59
 
60
  ```python
61
  from jevvision import JEVision, image_file_as_data_url
62
 
63
- model = JEVision.from_pretrained(".", text_adapter="jevbench-m3", device="cuda:0")
64
  result = model.system_one(
65
- state="Which option best describes the attached picture?",
66
- images=[image_file_as_data_url("photo.jpg")],
67
  questions={
68
- "answer": {
69
  "type": "choice",
70
- "instructions": "Choose the matching description.",
71
- "criteria": {"indoor": "An indoor scene", "outdoor": "An outdoor scene"},
 
 
 
 
72
  }
73
  },
74
  )
75
- print(result["answers"]["answer"])
 
 
 
 
 
 
 
 
 
 
76
  ```
77
 
78
- Images are supplied as PNG, JPEG, or WebP data URLs (up to four images; 10 MiB each; 25 megapixels each). Requests without `images` use the selected text adapter. Requests with `images` use the visual sidecar; **the M3 text LoRA is not applied to image reads**.
 
 
79
 
80
- ## What is bundled
 
 
 
 
 
81
 
82
- | Component | Route | Notes |
83
- |---|---|---|
84
- | `text/jevbench-m3/` | Text-only | Qwen3.5-0.8B-Base + the M3 text LoRA and pointer head. |
85
- | `text/kev-0.8b/` | Text-only, optional | Pinned released KEV-0.8B baseline LoRA and pointer head. |
86
- | `adapter/` + `pointer_head.pt` | Image + text | The compatible 0.8B visual sidecar and its pointer head. |
87
- | `runtime/kev/` | Both | Bundled KEV-compatible inference implementation. |
88
 
89
- The base is `Qwen/Qwen3.5-0.8B-Base` at revision `dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68`. Adapter files are included; base weights are fetched from the base-model repo as needed.
90
 
91
- ## Scope and limitations
92
 
93
- - **Context:** the configured request maximum is 80,000 tokens. This is not a 128K guarantee, and the M3 text adapter was trained on short benchmark examples rather than long-context training sequences.
94
- - **Images:** image requests route to a separate visual sidecar. Image-quality results are not reported here; the available functional checks used synthetic inputs and do not establish general real-photo accuracy or an advantage over stock Qwen.
95
- - **Text accuracy:** see the JevBench table above. The 231/231 public-panel result is in-sample; the grouped holdout is the more relevant generalization check, where M3 scored below Jev. This is not a claim of Jev parity.
96
- - **Latency:** this bundle is not latency-optimized; no cost or speed advantage is claimed.
97
- - **Architecture:** this is routed inference with separate text and visual LoRAs/heads—not one adapter that handles both text-only and image requests.
98
 
99
- `jevvision_manifest.json` records the adapter identities, hashes, routes, and configured limits. The raw training dataset, photos, benchmark scores, and Kaggle logs are not included.
100
 
101
  ## License
102
 
 
1
  ---
2
  license: apache-2.0
3
+ base_model: Qwen/Qwen3.5-0.8B-Base
4
+ base_model_relation: adapter
5
+ language:
6
+ - en
7
  tags:
8
  - multimodal
9
  - vision
 
16
 
17
  # JEVision
18
 
19
+ ![JEVision artwork](assets/jevvision-hero.png)
20
 
21
+ JEVision is a Qwen3.5-based system for making structured decisions from text and images. It extends the KEV/Jev-style System One interface with a visual route, so an application can send context and typed questions and receive **Choice**, **Noul**, or **Score** answers instead of parsing free-form prose.
22
 
23
+ The release combines a trained text adapter, a separately trained visual sidecar, pointer heads, and a bundled inference runtime. Both routes use the same pinned Qwen3.5-0.8B-Base revision. This makes the model useful for workflows that need a consistent decision API across text-only and image-bearing requests.
24
 
25
+ ## Capabilities
26
 
27
+ | Capability | JEVision |
28
+ |---|---|
29
+ | Input | Text context, with optional PNG, JPEG, or WebP images |
30
+ | Output | Typed Choice, Noul, and Score responses through `/v1/systemone` |
31
+ | Text route | JEVision text LoRA and pointer head |
32
+ | Image route | Visual LoRA and pointer head, selected automatically when images are present |
33
+ | Request envelope | Configured for up to 80,000 processed input tokens |
34
+ | Packaging | Adapters, heads, and a runnable local server; base weights download separately |
35
 
36
+ The serving route processed 76,999 text tokens and 76,998 image-plus-text tokens in recorded acceptance checks. The text check used the included KEV-0.8B option; the image check used the visual sidecar. These checks establish request handling near 77K, while long-context answer quality remains to be evaluated.
37
 
38
+ ## See the visual route
39
 
40
+ The repository includes a real photograph, the exact request sent to the visual route, and its captured response. The model selected `laptop_and_coffee` from three descriptions of the scene.
41
 
42
+ ![A cup of coffee beside a laptop on a wooden table](examples/real-photo/coffee-and-laptop.jpg)
43
 
44
+ Photo by [Shixart1985](https://commons.wikimedia.org/wiki/User:Shixart1985), via Wikimedia Commons, [CC BY 2.0](https://creativecommons.org/licenses/by/2.0/). [Source photograph](https://commons.wikimedia.org/wiki/File:Coffee_cup_next_to_laptop_on_wooden_table_in_cozy_indoor_workspace_during_daytime.jpg); the bundled file is its 960-pixel Commons thumbnail.
45
 
46
+ The saved [request](examples/real-photo/input.json) and [full response](examples/real-photo/output.json) can be inspected or replayed. The response below is taken from that recorded adapter run:
 
 
 
47
 
48
+ ```json
49
+ {
50
+ "answers": {
51
+ "scene": {
52
+ "type": "choice",
53
+ "choice": "laptop_and_coffee",
54
+ "confidence": 1.0,
55
+ "probabilities": {
56
+ "laptop_and_coffee": 1.0,
57
+ "bicycle_and_helmet": 0.0,
58
+ "cat_on_sofa": 0.0
59
+ }
60
+ }
61
+ },
62
+ "usage": {"input_tokens": 674, "output_tokens": 67}
63
+ }
64
+ ```
65
+
66
+ This is one recorded example. The full response also records the latency of that CPU run; it is not a speed comparison. Run it yourself after starting the server:
67
+
68
+ ```bash
69
+ python examples/real-photo/run_demo.py --endpoint http://127.0.0.1:8009
70
+ ```
71
 
72
+ ## Run locally
73
 
74
+ The bundle was tested on Linux with CUDA and a Tesla T4. A 16 GB NVIDIA GPU is recommended when hosting both routes together. The Qwen base weights are fetched on first use.
75
 
76
  ```bash
77
  git clone https://huggingface.co/divyanshx11/JEVision
78
  cd JEVision
79
  python -m venv .venv && source .venv/bin/activate
80
  python -m pip install -r requirements.txt
81
+ python run_jevvision.py --port 8009
82
  ```
83
 
84
+ The server exposes `POST /v1/systemone` on localhost. To select the bundled KEV text checkpoint for requests without images, add `--text-adapter kev-0.8b`. Image-bearing requests automatically use the visual sidecar.
85
 
86
+ ### Call from Python
87
 
88
  ```python
89
  from jevvision import JEVision, image_file_as_data_url
90
 
91
+ model = JEVision.from_pretrained(".", device="cuda:0")
92
  result = model.system_one(
93
+ state="Use the attached photo as visual context.",
94
+ images=[image_file_as_data_url("examples/real-photo/coffee-and-laptop.jpg")],
95
  questions={
96
+ "scene": {
97
  "type": "choice",
98
+ "instructions": "Which description best matches the photo?",
99
+ "criteria": {
100
+ "laptop_and_coffee": "A laptop beside a cup of coffee on a table.",
101
+ "bicycle_and_helmet": "A bicycle parked beside a helmet.",
102
+ "cat_on_sofa": "A cat sitting on a sofa.",
103
+ },
104
  }
105
  },
106
  )
107
+ print(result["answers"]["scene"])
108
+ ```
109
+
110
+ The image API accepts up to four images per request, each no larger than 10 MiB and 25 megapixels. Text-only requests use the selected text adapter. Image-bearing requests use the visual sidecar; the two adapters are routed separately.
111
+
112
+ ## Long-context example
113
+
114
+ [The runnable request recipe](examples/long-context/input.json) builds an archived help-desk state of roughly 76,000 tokens and asks for the code in its final record. With the server running, use:
115
+
116
+ ```bash
117
+ python examples/long-context/run_demo.py --endpoint http://127.0.0.1:8009
118
  ```
119
 
120
+ The runner writes its generated state and saves the server's response only after checking that `usage.input_tokens` is at least 75,000. A [separate recorded text-route acceptance summary](examples/long-context/recorded-77k-text-response-summary.json) documents a 76,999-token request on the bundled KEV-0.8B option; it is a request-handling check, not a result for the default JEVision text adapter.
121
+
122
+ ## Architecture and training
123
 
124
+ | Included component | Purpose |
125
+ |---|---|
126
+ | `text/jevvision-text/` | Default Qwen3.5 text decision LoRA and pointer head |
127
+ | `text/kev-0.8b/` | Optional pinned [KEV-0.8B](https://huggingface.co/jaredpalmer/kev-0.8b) text checkpoint |
128
+ | `adapter/` and `pointer_head.pt` | Image-aware decision sidecar |
129
+ | `runtime/kev/` | Local inference and System One serving code |
130
 
131
+ The base is [Qwen/Qwen3.5-0.8B-Base](https://huggingface.co/Qwen/Qwen3.5-0.8B-Base) at revision `dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68`. This repository supplies the adapters and heads, not a merged base-model checkpoint.
 
 
 
 
 
132
 
133
+ The default text adapter was fitted to public JevBench decision examples, starting from a KEV text checkpoint. Those public examples informed training and selection, so results on them are not an independent evaluation. The visual sidecar was trained separately with photos from the [Beans training split](https://huggingface.co/datasets/AI-Lab-Makerere/beans) and programmatically generated, labeled visual tasks. The generated tasks supplement the real images; they are training data, not claimed evaluation measurements. The visual route is not scored by JevBench.
134
 
135
+ ## Evaluation scope
136
 
137
+ JEVision currently has functional checks for typed responses, image routing, and long request acceptance, plus the recorded real-photo example above. Independent, broad image accuracy and long-context answer quality have not yet been measured. The 80,000-token value is a service limit, and requests beyond it are rejected instead of silently truncated. No comparative accuracy, latency, or cost claim is made here.
 
 
 
 
138
 
139
+ The [manifest](jevvision_manifest.json) identifies the components, their source revisions, hashes, and serving limit. Users should assess the model on their own images and decision tasks before relying on its outputs.
140
 
141
  ## License
142
 
examples/long-context/m1-77k-text-input.json DELETED
@@ -1,30 +0,0 @@
1
- {
2
- "source": "M1 context acceptance, private version 3",
3
- "model": "kev-latest",
4
- "text_adapter": {
5
- "repo_id": "jaredpalmer/kev-0.8b",
6
- "revision": "54f4f8777356cd5bbbb6c6919c657f26e6f2f6d8"
7
- },
8
- "base_model": {
9
- "repo_id": "Qwen/Qwen3.5-0.8B-Base",
10
- "revision": "dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68"
11
- },
12
- "text_only": true,
13
- "reported_input_tokens": 76999,
14
- "state_recipe": {
15
- "prefix": "M1 context acceptance. ",
16
- "repeated_text": "The stored memo remained unchanged. ",
17
- "repeat_count": 12827,
18
- "suffix": " END-OF-CONTEXT marker."
19
- },
20
- "questions": {
21
- "m1": {
22
- "type": "choice",
23
- "instructions": "Which note is being referred to?",
24
- "criteria": {
25
- "memo": "The stored memo",
26
- "other": "A different item"
27
- }
28
- }
29
- }
30
- }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
examples/long-context/{m1-77k-text-response-summary.json → recorded-77k-text-response-summary.json} RENAMED
@@ -1,13 +1,11 @@
1
  {
2
  "summary_type": "recorded response summary; not the raw API body",
3
- "experiment": "M1 context acceptance, private version 3",
4
  "source_report": {
5
- "filename": "report.json",
6
  "sha256": "f0fde673452b0d4a1329fc328e045afeb5223023e7ea3f2d16c53d3ce8d716b9"
7
  },
8
  "request": {
9
  "model": "kev-latest",
10
- "question_id": "m1",
11
  "text_adapter": "jaredpalmer/kev-0.8b@54f4f8777356cd5bbbb6c6919c657f26e6f2f6d8",
12
  "base_model_revision": "dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68",
13
  "text_only": true
@@ -27,10 +25,5 @@
27
  "probabilities"
28
  ]
29
  },
30
- "not_retained": [
31
- "numeric confidence value",
32
- "per-option probability values",
33
- "original raw API response body"
34
- ],
35
- "scope": "Long-context request acceptance and response-schema smoke; not a general long-context accuracy evaluation."
36
  }
 
1
  {
2
  "summary_type": "recorded response summary; not the raw API body",
 
3
  "source_report": {
4
+ "description": "archived context-acceptance report",
5
  "sha256": "f0fde673452b0d4a1329fc328e045afeb5223023e7ea3f2d16c53d3ce8d716b9"
6
  },
7
  "request": {
8
  "model": "kev-latest",
 
9
  "text_adapter": "jaredpalmer/kev-0.8b@54f4f8777356cd5bbbb6c6919c657f26e6f2f6d8",
10
  "base_model_revision": "dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68",
11
  "text_only": true
 
25
  "probabilities"
26
  ]
27
  },
28
+ "scope": "Long-context request acceptance and response-schema check."
 
 
 
 
 
29
  }
jevvision_manifest.json CHANGED
@@ -7,11 +7,11 @@
7
  "license": "apache-2.0",
8
  "weights_included": false
9
  },
10
- "default_text_adapter": "jevbench-m3",
11
  "text_adapters": {
12
- "jevbench-m3": {
13
- "path": "text/jevbench-m3",
14
- "role": "public-benchmark-tuned text branch; see README for in-sample public-panel and grouped-holdout results"
15
  },
16
  "kev-0.8b": {
17
  "path": "text/kev-0.8b",
@@ -24,38 +24,24 @@
24
  "path": ".",
25
  "adapter": "adapter/adapter_model.safetensors",
26
  "head": "pointer_head.pt",
27
- "role": "independent visual sidecar; image requests do not use the M3 text LoRA"
28
  },
29
  "routing": {
30
  "without_images": "selected text adapter",
31
  "with_images": "visual sidecar"
32
  },
33
  "context": {
34
- "max_request_tokens": 80000,
35
- "128k_guarantee": false,
36
- "long_context_accuracy_established": false
37
- },
38
- "visual_evidence": {
39
- "image_quality_results_published": false,
40
- "general_image_accuracy_established": false,
41
- "real_photo_advantage_established": false
42
  },
43
  "artifact_sha256": {
44
- "D:/Dark/jevvision_release_bundle_final_20260925/text/jevbench-m3/adapter_model.safetensors": "0f71c01d01cc9a82648417f3b036e329e0fe27dd44cd6847fa141b85518e1022",
45
- "D:/Dark/jevvision_release_bundle_final_20260925/text/jevbench-m3/adapter_config.json": "917dbb4e84a81737b187d98633afe8139eded84eabe83cfcf8c9e210a3001615",
46
- "D:/Dark/jevvision_release_bundle_final_20260925/text/jevbench-m3/head.pt": "b3ac9f2e10a903bfec76020ee4a730921c22531267de7493aff57b80c741e906",
47
- "D:/Dark/jevvision_release_bundle_final_20260925/text/kev-0.8b/adapter_model.safetensors": "c81d5716f0af7622d8d2b97013c333d48263ca01113f9a7cf4526e96f6ac0b26",
48
- "D:/Dark/jevvision_release_bundle_final_20260925/text/kev-0.8b/adapter_config.json": "748acb2cda88454cb1ba69d745ba336f3fcb5486eac349e90960c8b8d8d3e854",
49
- "D:/Dark/jevvision_release_bundle_final_20260925/text/kev-0.8b/head.pt": "39f4343ccccc65e583bbfff0de0e11bfedb849fcfaaf94b50ac4f2b73bc79c65",
50
- "D:/Dark/jevvision_release_bundle_final_20260925/adapter/adapter_config.json": "e6cef9780f0b336b051e317e7ba2becd7cc6bbb139a18c64f83ddae9fc441cdd",
51
- "D:/Dark/jevvision_release_bundle_final_20260925/adapter/adapter_model.safetensors": "ffee0467a55883e8aa212035426eefad09380d903e0be49bbc69ebb1cf6de436",
52
  "pointer_head.pt": "4c5c16989c6aec9c290c6e5382d972c7c3e9ce5ebdc6f013d599195faa5e17ed"
53
- },
54
- "omitted": [
55
- "base-model weights",
56
- "dataset files",
57
- "photos",
58
- "training-state checkpoints",
59
- "unpromoted 4B experiments"
60
- ]
61
  }
 
7
  "license": "apache-2.0",
8
  "weights_included": false
9
  },
10
+ "default_text_adapter": "jevvision-text",
11
  "text_adapters": {
12
+ "jevvision-text": {
13
+ "path": "text/jevvision-text",
14
+ "role": "default typed-decision text branch"
15
  },
16
  "kev-0.8b": {
17
  "path": "text/kev-0.8b",
 
24
  "path": ".",
25
  "adapter": "adapter/adapter_model.safetensors",
26
  "head": "pointer_head.pt",
27
+ "role": "image-aware typed-decision sidecar"
28
  },
29
  "routing": {
30
  "without_images": "selected text adapter",
31
  "with_images": "visual sidecar"
32
  },
33
  "context": {
34
+ "max_request_tokens": 80000
 
 
 
 
 
 
 
35
  },
36
  "artifact_sha256": {
37
+ "text/jevvision-text/adapter_model.safetensors": "0f71c01d01cc9a82648417f3b036e329e0fe27dd44cd6847fa141b85518e1022",
38
+ "text/jevvision-text/adapter_config.json": "917dbb4e84a81737b187d98633afe8139eded84eabe83cfcf8c9e210a3001615",
39
+ "text/jevvision-text/head.pt": "b3ac9f2e10a903bfec76020ee4a730921c22531267de7493aff57b80c741e906",
40
+ "text/kev-0.8b/adapter_model.safetensors": "c81d5716f0af7622d8d2b97013c333d48263ca01113f9a7cf4526e96f6ac0b26",
41
+ "text/kev-0.8b/adapter_config.json": "748acb2cda88454cb1ba69d745ba336f3fcb5486eac349e90960c8b8d8d3e854",
42
+ "text/kev-0.8b/head.pt": "39f4343ccccc65e583bbfff0de0e11bfedb849fcfaaf94b50ac4f2b73bc79c65",
43
+ "adapter/adapter_config.json": "e6cef9780f0b336b051e317e7ba2becd7cc6bbb139a18c64f83ddae9fc441cdd",
44
+ "adapter/adapter_model.safetensors": "ffee0467a55883e8aa212035426eefad09380d903e0be49bbc69ebb1cf6de436",
45
  "pointer_head.pt": "4c5c16989c6aec9c290c6e5382d972c7c3e9ce5ebdc6f013d599195faa5e17ed"
46
+ }
 
 
 
 
 
 
 
47
  }
run_jevvision.py CHANGED
@@ -10,14 +10,14 @@ from pathlib import Path
10
  def main() -> None:
11
  root = Path(__file__).resolve().parent
12
  parser = argparse.ArgumentParser()
13
- parser.add_argument("--text-adapter", choices=("jevbench-m3", "kev-0.8b"), default="jevbench-m3")
14
  parser.add_argument("--port", type=int, default=8009)
15
  args = parser.parse_args()
16
 
17
  sys.path.insert(0, str(root / "runtime"))
18
  from kev.serve import main as serve_main
19
 
20
- text_path = root / "text" / ("jevbench-m3" if args.text_adapter == "jevbench-m3" else "kev-0.8b")
21
  sys.argv = ["kev.serve", "--run", str(text_path), "--visual-run", str(root), "--port", str(args.port)]
22
  serve_main()
23
 
 
10
  def main() -> None:
11
  root = Path(__file__).resolve().parent
12
  parser = argparse.ArgumentParser()
13
+ parser.add_argument("--text-adapter", choices=("jevvision-text", "kev-0.8b"), default="jevvision-text")
14
  parser.add_argument("--port", type=int, default=8009)
15
  args = parser.parse_args()
16
 
17
  sys.path.insert(0, str(root / "runtime"))
18
  from kev.serve import main as serve_main
19
 
20
+ text_path = root / "text" / args.text_adapter
21
  sys.argv = ["kev.serve", "--run", str(text_path), "--visual-run", str(root), "--port", str(args.port)]
22
  serve_main()
23
 
runtime/kev/jevvision.py CHANGED
@@ -21,7 +21,7 @@ from .serve import SERVE_CONTEXT_TOKENS, Server
21
 
22
 
23
  TEXT_ADAPTERS = {
24
- "jevbench-m3": "text/jevbench-m3",
25
  "kev-0.8b": "text/kev-0.8b",
26
  }
27
  MAX_IMAGE_BYTES = 10 * 1024 * 1024
@@ -52,7 +52,7 @@ class JEVision:
52
  repo_id: str = "divyanshx11/JEVision",
53
  *,
54
  revision: str = "main",
55
- text_adapter: str = "jevbench-m3",
56
  device: str | None = None,
57
  cache_dir: str | os.PathLike[str] | None = None,
58
  ) -> "JEVision":
@@ -79,7 +79,7 @@ class JEVision:
79
  cls,
80
  bundle_dir: str | os.PathLike[str],
81
  *,
82
- text_adapter: str = "jevbench-m3",
83
  device: str | None = None,
84
  ) -> "JEVision":
85
  """Load a previously downloaded JEVision bundle from disk."""
 
21
 
22
 
23
  TEXT_ADAPTERS = {
24
+ "jevvision-text": "text/jevvision-text",
25
  "kev-0.8b": "text/kev-0.8b",
26
  }
27
  MAX_IMAGE_BYTES = 10 * 1024 * 1024
 
52
  repo_id: str = "divyanshx11/JEVision",
53
  *,
54
  revision: str = "main",
55
+ text_adapter: str = "jevvision-text",
56
  device: str | None = None,
57
  cache_dir: str | os.PathLike[str] | None = None,
58
  ) -> "JEVision":
 
79
  cls,
80
  bundle_dir: str | os.PathLike[str],
81
  *,
82
+ text_adapter: str = "jevvision-text",
83
  device: str | None = None,
84
  ) -> "JEVision":
85
  """Load a previously downloaded JEVision bundle from disk."""
text/{jevbench-m3 → jevvision-text}/adapter_config.json RENAMED
File without changes
text/{jevbench-m3 → jevvision-text}/adapter_model.safetensors RENAMED
File without changes
text/{jevbench-m3 → jevvision-text}/head.pt RENAMED
File without changes