divyanshx11 commited on
Commit
634f6d4
·
verified ·
1 Parent(s): 4d27595

Add an attributed real-photo image example

Browse files
README.md CHANGED
@@ -54,6 +54,48 @@ print(result["answers"]["answer"])
54
 
55
  Images are supplied as PNG, JPEG, or WebP data URLs (up to four images; 10 MiB each; 25 megapixels each). Requests without `images` use the selected text adapter. Requests with `images` use the visual sidecar; **the M3 text LoRA is not applied to image reads**.
56
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
57
  ## What is bundled
58
 
59
  | Component | Route | Notes |
@@ -68,12 +110,12 @@ The base is `Qwen/Qwen3.5-0.8B-Base` at revision `dc7cdfe2ee4154fa7e30f5b51ca41b
68
  ## Scope and limitations
69
 
70
  - **Context:** the configured request maximum is 80,000 tokens. This is not a 128K guarantee, and the M3 text adapter was trained on short benchmark examples rather than long-context training sequences.
71
- - **Images:** image requests route to a separate visual sidecar. Image-quality results are not reported here; the available functional checks used synthetic inputs and do not establish general real-photo accuracy or an advantage over stock Qwen.
72
  - **Text accuracy:** the M3 text branch was trained and selected using public JevBench data. Benchmark scores are intentionally omitted from this release description; this is not an independent evaluation or a claim of Jev parity.
73
  - **Latency:** this bundle is not latency-optimized; no cost or speed advantage is claimed.
74
  - **Architecture:** this is routed inference with separate text and visual LoRAs/heads—not one adapter that handles both text-only and image requests.
75
 
76
- `jevvision_manifest.json` records the adapter identities, hashes, routes, and configured limits. The raw training dataset, photos, benchmark scores, and Kaggle logs are not included.
77
 
78
  ## License
79
 
 
54
 
55
  Images are supplied as PNG, JPEG, or WebP data URLs (up to four images; 10 MiB each; 25 megapixels each). Requests without `images` use the selected text adapter. Requests with `images` use the visual sidecar; **the M3 text LoRA is not applied to image reads**.
56
 
57
+ ## Real-photo image example
58
+
59
+ This example sends a real photograph through JEVision's image route. The photo shows a laptop beside a cup of coffee.
60
+
61
+ ![A cup of coffee beside a laptop on a wooden table](examples/real-photo/coffee-and-laptop.jpg)
62
+
63
+ Photo credit: [Shixart1985](https://commons.wikimedia.org/wiki/User:Shixart1985), via Wikimedia Commons, [CC BY 2.0](https://creativecommons.org/licenses/by/2.0/). [Source photograph](https://commons.wikimedia.org/wiki/File:Coffee_cup_next_to_laptop_on_wooden_table_in_cozy_indoor_workspace_during_daytime.jpg); the bundled file is the 960-pixel Commons thumbnail, unaltered here.
64
+
65
+ The runnable request definition is [`input.json`](examples/real-photo/input.json). It names the image file so the example stays readable; [`run_demo.py`](examples/real-photo/run_demo.py) encodes that file as a data URL and posts the request to `/v1/systemone`:
66
+
67
+ ```bash
68
+ python run_jevvision.py --text-adapter jevbench-m3 --port 8009
69
+ python examples/real-photo/run_demo.py --endpoint http://127.0.0.1:8009
70
+ ```
71
+
72
+ Captured response for this exact image and request:
73
+
74
+ ```json
75
+ {
76
+ "model": "kev-latest",
77
+ "answers": {
78
+ "scene": {
79
+ "type": "choice",
80
+ "choice": "laptop_and_coffee",
81
+ "confidence": 1.0,
82
+ "probabilities": {
83
+ "laptop_and_coffee": 1.0,
84
+ "bicycle_and_helmet": 0.0,
85
+ "cat_on_sofa": 0.0
86
+ }
87
+ }
88
+ },
89
+ "usage": {
90
+ "input_tokens": 674,
91
+ "output_tokens": 67
92
+ },
93
+ "latency_ms": 14771.8
94
+ }
95
+ ```
96
+
97
+ The full response is saved as [`output.json`](examples/real-photo/output.json). This is one qualitative example, not an image benchmark or evidence of general image accuracy. The displayed latency came from a CPU smoke run and is not a performance claim.
98
+
99
  ## What is bundled
100
 
101
  | Component | Route | Notes |
 
110
  ## Scope and limitations
111
 
112
  - **Context:** the configured request maximum is 80,000 tokens. This is not a 128K guarantee, and the M3 text adapter was trained on short benchmark examples rather than long-context training sequences.
113
+ - **Images:** image requests route to a separate visual sidecar; the M3 text LoRA is not applied to image reads. The single real-photo example above is illustrative only, not a benchmark, a claim of general real-photo accuracy, or evidence of an advantage over stock Qwen.
114
  - **Text accuracy:** the M3 text branch was trained and selected using public JevBench data. Benchmark scores are intentionally omitted from this release description; this is not an independent evaluation or a claim of Jev parity.
115
  - **Latency:** this bundle is not latency-optimized; no cost or speed advantage is claimed.
116
  - **Architecture:** this is routed inference with separate text and visual LoRAs/heads—not one adapter that handles both text-only and image requests.
117
 
118
+ `jevvision_manifest.json` records the adapter identities, hashes, routes, configured limits, and the attributed example image. The raw training dataset, training/evaluation photos, benchmark scores, and Kaggle logs are not included.
119
 
120
  ## License
121
 
examples/real-photo/coffee-and-laptop.jpg ADDED
examples/real-photo/input.json ADDED
@@ -0,0 +1,16 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "state": "Use the attached photo as visual context.",
3
+ "model": "kev-latest",
4
+ "image_file": "coffee-and-laptop.jpg",
5
+ "questions": {
6
+ "scene": {
7
+ "type": "choice",
8
+ "instructions": "Which description best matches the photo?",
9
+ "criteria": {
10
+ "laptop_and_coffee": "A laptop beside a cup of coffee on a table.",
11
+ "bicycle_and_helmet": "A bicycle parked beside a helmet.",
12
+ "cat_on_sofa": "A cat sitting on a sofa."
13
+ }
14
+ }
15
+ }
16
+ }
examples/real-photo/output.json ADDED
@@ -0,0 +1,20 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model": "kev-latest",
3
+ "answers": {
4
+ "scene": {
5
+ "type": "choice",
6
+ "choice": "laptop_and_coffee",
7
+ "confidence": 1.0,
8
+ "probabilities": {
9
+ "laptop_and_coffee": 1.0,
10
+ "bicycle_and_helmet": 0.0,
11
+ "cat_on_sofa": 0.0
12
+ }
13
+ }
14
+ },
15
+ "usage": {
16
+ "input_tokens": 674,
17
+ "output_tokens": 67
18
+ },
19
+ "latency_ms": 14771.8
20
+ }
examples/real-photo/photo-credit.md ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ Photo: “Coffee cup next to laptop on wooden table in cozy indoor workspace during daytime” by Shixart1985, via Wikimedia Commons, licensed under [CC BY 2.0](https://creativecommons.org/licenses/by/2.0/).
2
+
3
+ Source: [Wikimedia Commons file page](https://commons.wikimedia.org/wiki/File:Coffee_cup_next_to_laptop_on_wooden_table_in_cozy_indoor_workspace_during_daytime.jpg). The bundled image is Wikimedia Commons’ 960-pixel thumbnail of the photograph; it is unaltered in this repository.
examples/real-photo/run_demo.py ADDED
@@ -0,0 +1,44 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """Send the bundled real-photo example to a running JEVision System One server."""
3
+ from __future__ import annotations
4
+
5
+ import argparse
6
+ import json
7
+ import sys
8
+ from pathlib import Path
9
+ from urllib.request import Request, urlopen
10
+
11
+ EXAMPLE_DIR = Path(__file__).resolve().parent
12
+ REPO_ROOT = EXAMPLE_DIR.parents[1]
13
+ sys.path.insert(0, str(REPO_ROOT))
14
+
15
+ from jevvision import image_file_as_data_url # noqa: E402
16
+
17
+
18
+ def main() -> None:
19
+ parser = argparse.ArgumentParser()
20
+ parser.add_argument("--endpoint", default="http://127.0.0.1:8009")
21
+ parser.add_argument("--output", type=Path, default=EXAMPLE_DIR / "output.json")
22
+ args = parser.parse_args()
23
+
24
+ payload = json.loads((EXAMPLE_DIR / "input.json").read_text(encoding="utf-8"))
25
+ image_name = payload.pop("image_file")
26
+ image_path = (EXAMPLE_DIR / image_name).resolve()
27
+ image_path.relative_to(EXAMPLE_DIR.resolve())
28
+ payload["images"] = [image_file_as_data_url(image_path)]
29
+
30
+ request = Request(
31
+ args.endpoint.rstrip("/") + "/v1/systemone",
32
+ data=json.dumps(payload).encode("utf-8"),
33
+ headers={"content-type": "application/json"},
34
+ method="POST",
35
+ )
36
+ with urlopen(request, timeout=900) as response:
37
+ result = json.loads(response.read().decode("utf-8"))
38
+ rendered = json.dumps(result, indent=2, ensure_ascii=False) + "\n"
39
+ args.output.write_text(rendered, encoding="utf-8")
40
+ print(rendered, end="")
41
+
42
+
43
+ if __name__ == "__main__":
44
+ main()
jevvision_manifest.json CHANGED
@@ -38,7 +38,15 @@
38
  "visual_evidence": {
39
  "image_quality_results_published": false,
40
  "general_image_accuracy_established": false,
41
- "real_photo_advantage_established": false
 
 
 
 
 
 
 
 
42
  },
43
  "artifact_sha256": {
44
  "D:/Dark/jevvision_release_bundle_final_20260925/text/jevbench-m3/adapter_model.safetensors": "0f71c01d01cc9a82648417f3b036e329e0fe27dd44cd6847fa141b85518e1022",
@@ -54,7 +62,7 @@
54
  "omitted": [
55
  "base-model weights",
56
  "dataset files",
57
- "photos",
58
  "training-state checkpoints",
59
  "unpromoted 4B experiments"
60
  ]
 
38
  "visual_evidence": {
39
  "image_quality_results_published": false,
40
  "general_image_accuracy_established": false,
41
+ "real_photo_advantage_established": false,
42
+ "qualitative_example": {
43
+ "image": "examples/real-photo/coffee-and-laptop.jpg",
44
+ "input": "examples/real-photo/input.json",
45
+ "output": "examples/real-photo/output.json",
46
+ "image_sha256": "95abfa66426cb1ef5b41517545b26cf8968a2d50754f93367a827272f6855cac",
47
+ "license": "CC-BY-2.0",
48
+ "scope": "one qualitative real-photo request/response; not a benchmark"
49
+ }
50
  },
51
  "artifact_sha256": {
52
  "D:/Dark/jevvision_release_bundle_final_20260925/text/jevbench-m3/adapter_model.safetensors": "0f71c01d01cc9a82648417f3b036e329e0fe27dd44cd6847fa141b85518e1022",
 
62
  "omitted": [
63
  "base-model weights",
64
  "dataset files",
65
+ "training/evaluation photos (except the separately attributed example above)",
66
  "training-state checkpoints",
67
  "unpromoted 4B experiments"
68
  ]