jtdavies commited on
Commit
b61b742
·
verified ·
1 Parent(s): ad42e1b

Verification: add reproducible test image, inference snippet, and results

Browse files
Files changed (1) hide show
  1. README.md +29 -6
README.md CHANGED
@@ -45,18 +45,41 @@ print(output)
45
 
46
  ## Verification (smoke test)
47
 
48
- This conversion was smoke-tested on Apple silicon with `mlx-vlm` after quantization.
49
- Both a text-only prompt and a vision prompt (on a deterministic generated image of
50
- coloured shapes + a text label) were run against the 8-bit weights.
51
 
52
- **Text prompt** — _"In one sentence, what is the capital of France?"_
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
53
 
54
  ```
55
  The capital of France is Paris.
56
  ```
57
 
58
- **Vision prompt** — _"Describe the shapes and their colours in this image, and read any text you see."_
59
- (test image: red square top-left, blue circle top-right, green inverted triangle bottom-centre, label "MLX VISION TEST")
60
 
61
  ```
62
  Based on the image provided, here is a description of the shapes, their colors, and the text present:
 
45
 
46
  ## Verification (smoke test)
47
 
48
+ Smoke-tested on Apple silicon with `mlx-vlm` after quantization — a text prompt plus a
49
+ vision prompt against a deterministic, network-free test image (coloured shapes + a text
50
+ label), so the test is fully reproducible.
51
 
52
+ ### Test image
53
+
54
+ ![smoke test image](assets/smoke_test.jpg)
55
+
56
+ ### How it was run
57
+
58
+ ```python
59
+ from mlx_vlm import load, generate
60
+ from mlx_vlm.prompt_utils import apply_chat_template
61
+ from mlx_vlm.utils import load_config
62
+
63
+ model, processor = load("incept5/Qwen3.8-27B-MLX-8bit")
64
+ config = load_config("incept5/Qwen3.8-27B-MLX-8bit")
65
+
66
+ # vision
67
+ msgs = [{"role": "user", "content": "Describe the shapes and their colours in this image, and read any text you see."}]
68
+ prompt = apply_chat_template(processor, config, msgs, num_images=1)
69
+ print(generate(model, processor, prompt, image=["test.jpg"], max_tokens=256, verbose=False))
70
+ ```
71
+
72
+ (Text prompt used the same pattern with `num_images=0` and no `image=` argument.)
73
+
74
+ ### Results
75
+
76
+ **Text** — _"In one sentence, what is the capital of France?"_
77
 
78
  ```
79
  The capital of France is Paris.
80
  ```
81
 
82
+ **Vision** — _"Describe the shapes and their colours in this image, and read any text you see."_
 
83
 
84
  ```
85
  Based on the image provided, here is a description of the shapes, their colors, and the text present: