ScottHao commited on
Commit
071b06d
·
verified ·
1 Parent(s): 710c73a

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +150 -0
README.md CHANGED
@@ -1,3 +1,153 @@
 
1
  ---
2
  license: mit
 
3
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+
2
  ---
3
  license: mit
4
+ pipeline_tag: text-generation
5
  ---
6
+ <p align="center">
7
+ <img src="https://mdn.alipayobjects.com/huamei_qa8qxu/afts/img/A*4QxcQrBlTiAAAAAAQXAAAAgAemJ7AQ/original" width="100"/>
8
+ </p>
9
+ <p align="center">🤗 <a href="https://huggingface.co/inclusionAI">Hugging Face</a>&nbsp;&nbsp; | &nbsp;&nbsp;🤖 <a href="https://modelscope.cn/organization/inclusionAI">ModelScope </a>&nbsp;&nbsp;</p>
10
+
11
+ # Introduction
12
+ We are introducing Ling-3.0-flash-VL, our next-generation native multimodal model. Built upon Ling-3.0-flash, it brings visual information into the complete process of understanding, reasoning, acting, and verification—advancing beyond image and video perception to solving real-world tasks through vision.
13
+ With 124B total parameters, only 5.5B activated parameters per token, support for image and video inputs, and a context window of up to 1M tokens, Ling-3.0-flash-VL delivers powerful multimodal reasoning and agentic capabilities with exceptional efficiency.
14
+
15
+ # Model Overview
16
+ Ling-3.0-flash-VL inherits the language, reasoning, and long-context capabilities of Ling-3.0-flash, while extending them with native image and video understanding. The model has 124B total parameters, with only 5.5B parameters activated per token, and supports a context window of up to 1M tokens.
17
+
18
+ The architecture of Ling-3.0-flash-VL is designed to integrate visual information into real-world reasoning and agentic workflows.
19
+
20
+ - A ViT visual encoder extracts features from images and videos, while a two-layer MLP projector aligns visual features with text representations for unified multimodal understanding and reasoning;
21
+ - VideoRoPE encodes both spatial positions and temporal order, enabling the model to understand visual changes over time and supporting tasks such as event localization, long-video question answering, and video clip editing;
22
+ - A 42-layer hybrid backbone alternates KDA and Gated MLA layers at a 5:1 ratio, enabling efficient long-context processing across text, images, videos, and extended agent task histories;
23
+ - A sparse MoE architecture maintains a total model capacity of 124B parameters while activating only 5.5B parameters per token, balancing strong multimodal capabilities with inference efficiency.
24
+
25
+ Overall, these designs make vision more than just an input, integrating it into the complete process of understanding, reasoning, planning, acting, and verification.
26
+
27
+
28
+ ![ling-3.0-flash-vl-0906](https://cdn-uploads.huggingface.co/production/uploads/6666ca359f5a0b3229238a1a/WhsE3cM7QMelBCjdeyex2.png)
29
+
30
+ # Evaluation
31
+ Ling-3.0-flash-VL achieves a score of **42** on the Artificial Analysis Intelligence Index v4.1.1, improving by 4 points over Ling-3.0-flash’s score of 38. The results show that extending the model with visual capabilities further improves its overall intelligence performance.
32
+
33
+
34
+ ![ling-3.0-flash-vl-aa](https://cdn-uploads.huggingface.co/production/uploads/6666ca359f5a0b3229238a1a/5gqrYjboQ8j7sTWjd1A_q.png)
35
+
36
+ Across multimodal benchmarks, Ling-3.0-flash-VL demonstrates three distinct capability dimensions:
37
+
38
+ - **Understand: Comprehending complex visual information.** The model can handle object counting, complex layouts, charts, and document content.
39
+ - **Reason: Reasoning and verification with visual evidence.** The model can use visual information for calculation, multi-step reasoning, and external information verification.
40
+ - **Act: Interacting with interfaces and completing tasks.** The model can understand web and software interfaces, then translate visual information into sequences of actions.
41
+
42
+ ![ling-3.0-flash-vl-benchmark](https://cdn-uploads.huggingface.co/production/uploads/6666ca359f5a0b3229238a1a/w-V0gtCa3qy0un6vuz-sj.png)
43
+
44
+ > + Thinking mode is enabled by default. Unless otherwise specified, the default parameters for Ling-3.0-flash-VL are as follows: `temperature=0.6`, `top_p=0.95`, `top_k=20`.
45
+ > + Terminal-Bench 2.1: Evaluated under the Artificial Analysis (AA) protocol using the default Terminus 2 harness, a unified 2-hour timeout, the provided JSON parser in preserve-thinking mode, and 3 runs per task (mean). Decoding uses temperature=1.0, max_new_tokens=32K, with a 256K context window.
46
+
47
+ # Quickstart
48
+ ## SGLang
49
+ The hardware- and recipe-specific launch matrix (BF16/FP8 × Low-Latency / High-Throughput), with a live command generator and verified configurations, lives in the SGLang cookbook:
50
+
51
+ **Cookbook:** https://docs.sglang.io/cookbook/autoregressive/InclusionAI/Ling-3.0-flash-VL
52
+
53
+
54
+ ### Install SGLang
55
+
56
+ ```bash
57
+ docker pull lmsysorg/sglang:dev-Ling-3.0-flash-VL
58
+ ```
59
+
60
+ ### Run Inference
61
+ Recommended recipe with 256K context (YaRN), on 4× 141GB-class GPUs (H20-3e / H200) or 4-GPU Blackwell nodes (B300 / GB300):
62
+
63
+ ```bash
64
+ docker run --rm --gpus all --ipc=host --shm-size 32g \
65
+ -p 30000:30000 \
66
+ -e HF_TOKEN=<your-hf-token> \
67
+ lmsysorg/sglang:dev-Ling-3.0-flash-VL \
68
+ env SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1 \
69
+ python3 -m sglang.launch_server \
70
+ --model-path inclusionAI/Ling-3.0-flash-VL \
71
+ --tp 4 \
72
+ --context-length 262144 \
73
+ --json-model-override-args '{"rope_scaling":{"rope_type":"yarn","factor":2.0,"rope_theta":6000000,"partial_rotary_factor":0.5,"original_max_position_embeddings":131072}}' \
74
+ --mem-fraction-static 0.85 \
75
+ --trust-remote-code \
76
+ --reasoning-parser auto \
77
+ --tool-call-parser auto \
78
+ --host 0.0.0.0 \
79
+ --port 30000
80
+ ```
81
+
82
+ On 80GB cards (H100 / H800), scale out to `--tp 8`. The reasoning and tool-call parsers resolve automatically to `ling3` from the chat template; you can also set them explicitly with `--reasoning-parser ling3 --tool-call-parser ling3`.
83
+
84
+ **Client**
85
+
86
+ Thinking is enabled by default by the chat template; disable it per request with `"chat_template_kwargs": {"enable_thinking": false}`. Recommended sampling: `temperature=1.0`, `top_p=0.95`, `top_k=20` (per `generation_config.json`).
87
+
88
+ ```bash
89
+ curl -s http://localhost:30000/v1/chat/completions \
90
+ -H "Content-Type: application/json" \
91
+ -d '{"model": "inclusionAI/Ling-3.0-flash-VL",
92
+ "messages": [{"role": "user", "content": [
93
+ {"type": "image_url", "image_url": {"url": "https://example.com/image.png"}},
94
+ {"type": "text", "text": "Describe this image in one sentence."}
95
+ ]}],
96
+ "stream": true,
97
+ "temperature": 1.0, "top_k": 20, "top_p": 0.95
98
+ }'
99
+ ```
100
+ Video input uses `{"type": "video_url", "video_url": {"url": "..."}}` in the same message shape. For MMMU-Pro / `bench_serving` reproduction commands and per-hardware recipes, see the cookbook page linked above.
101
+
102
+ ## vLLM
103
+ ### Environment Preparation
104
+
105
+ ```bash
106
+ pip install uv
107
+
108
+ uv venv ~/my_ling_env
109
+
110
+ source ~/my_ling_env/bin/activate
111
+
112
+ git clone https://github.com/inclusionAI/vllm-ling-v3.git
113
+
114
+ cd vllm-ling-v3
115
+
116
+ VLLM_USE_PRECOMPILED=1 uv pip install --editable . --torch-backend=auto
117
+ ```
118
+
119
+ ### Run Inference
120
+
121
+ **Server**
122
+
123
+ ```bash
124
+ vllm serve "$MODEL_PATH" \
125
+ --port "$PORT" \
126
+ --trust-remote-code \
127
+ --served-model-name auto \
128
+ --tensor-parallel-size 4 \
129
+ --gpu-memory-utilization 0.85 \
130
+ --enable-prefix-caching \
131
+ --mamba-cache-mode align \
132
+ --enable-auto-tool-choice \
133
+ --tool-call-parser ling3 \
134
+ --reasoning-parser ling3
135
+ ```
136
+
137
+ **Client**
138
+
139
+ Thinking is enabled by default by the chat template; disable it per request with `"chat_template_kwargs": {"enable_thinking": false}`. Recommended sampling: `temperature=1.0`, `top_p=0.95`, `top_k=20` (per `generation_config.json`).
140
+
141
+ ```bash
142
+ curl -s http://${MASTER_IP}:${PORT}/v1/chat/completions \
143
+ -H "Content-Type: application/json" \
144
+ -d '{"model": "auto", -d '{"model": "inclusionAI/Ling-3.0-flash-VL",
145
+ "messages": [{"role": "user", "content": [
146
+ {"type": "image_url", "image_url": {"url": "https://example.com/image.png"}},
147
+ {"type": "text", "text": "Describe this image in one sentence."}
148
+ ]}],
149
+ "stream": true,
150
+ "temperature": 1.0, "top_k": 20, "top_p": 0.95
151
+ }'
152
+ ```
153
+ Video input uses `{"type": "video_url", "video_url": {"url": "..."}}` in the same message shape. For MMMU-Pro / `bench_serving` reproduction commands and per-hardware recipes, see the cookbook page linked above.