Update README.md
Browse files
README.md
CHANGED
|
@@ -1,3 +1,153 @@
|
|
|
|
|
| 1 |
---
|
| 2 |
license: mit
|
|
|
|
| 3 |
---
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
|
| 2 |
---
|
| 3 |
license: mit
|
| 4 |
+
pipeline_tag: text-generation
|
| 5 |
---
|
| 6 |
+
<p align="center">
|
| 7 |
+
<img src="https://mdn.alipayobjects.com/huamei_qa8qxu/afts/img/A*4QxcQrBlTiAAAAAAQXAAAAgAemJ7AQ/original" width="100"/>
|
| 8 |
+
</p>
|
| 9 |
+
<p align="center">🤗 <a href="https://huggingface.co/inclusionAI">Hugging Face</a> | 🤖 <a href="https://modelscope.cn/organization/inclusionAI">ModelScope </a> </p>
|
| 10 |
+
|
| 11 |
+
# Introduction
|
| 12 |
+
We are introducing Ling-3.0-flash-VL, our next-generation native multimodal model. Built upon Ling-3.0-flash, it brings visual information into the complete process of understanding, reasoning, acting, and verification—advancing beyond image and video perception to solving real-world tasks through vision.
|
| 13 |
+
With 124B total parameters, only 5.5B activated parameters per token, support for image and video inputs, and a context window of up to 1M tokens, Ling-3.0-flash-VL delivers powerful multimodal reasoning and agentic capabilities with exceptional efficiency.
|
| 14 |
+
|
| 15 |
+
# Model Overview
|
| 16 |
+
Ling-3.0-flash-VL inherits the language, reasoning, and long-context capabilities of Ling-3.0-flash, while extending them with native image and video understanding. The model has 124B total parameters, with only 5.5B parameters activated per token, and supports a context window of up to 1M tokens.
|
| 17 |
+
|
| 18 |
+
The architecture of Ling-3.0-flash-VL is designed to integrate visual information into real-world reasoning and agentic workflows.
|
| 19 |
+
|
| 20 |
+
- A ViT visual encoder extracts features from images and videos, while a two-layer MLP projector aligns visual features with text representations for unified multimodal understanding and reasoning;
|
| 21 |
+
- VideoRoPE encodes both spatial positions and temporal order, enabling the model to understand visual changes over time and supporting tasks such as event localization, long-video question answering, and video clip editing;
|
| 22 |
+
- A 42-layer hybrid backbone alternates KDA and Gated MLA layers at a 5:1 ratio, enabling efficient long-context processing across text, images, videos, and extended agent task histories;
|
| 23 |
+
- A sparse MoE architecture maintains a total model capacity of 124B parameters while activating only 5.5B parameters per token, balancing strong multimodal capabilities with inference efficiency.
|
| 24 |
+
|
| 25 |
+
Overall, these designs make vision more than just an input, integrating it into the complete process of understanding, reasoning, planning, acting, and verification.
|
| 26 |
+
|
| 27 |
+
|
| 28 |
+

|
| 29 |
+
|
| 30 |
+
# Evaluation
|
| 31 |
+
Ling-3.0-flash-VL achieves a score of **42** on the Artificial Analysis Intelligence Index v4.1.1, improving by 4 points over Ling-3.0-flash’s score of 38. The results show that extending the model with visual capabilities further improves its overall intelligence performance.
|
| 32 |
+
|
| 33 |
+
|
| 34 |
+

|
| 35 |
+
|
| 36 |
+
Across multimodal benchmarks, Ling-3.0-flash-VL demonstrates three distinct capability dimensions:
|
| 37 |
+
|
| 38 |
+
- **Understand: Comprehending complex visual information.** The model can handle object counting, complex layouts, charts, and document content.
|
| 39 |
+
- **Reason: Reasoning and verification with visual evidence.** The model can use visual information for calculation, multi-step reasoning, and external information verification.
|
| 40 |
+
- **Act: Interacting with interfaces and completing tasks.** The model can understand web and software interfaces, then translate visual information into sequences of actions.
|
| 41 |
+
|
| 42 |
+

|
| 43 |
+
|
| 44 |
+
> + Thinking mode is enabled by default. Unless otherwise specified, the default parameters for Ling-3.0-flash-VL are as follows: `temperature=0.6`, `top_p=0.95`, `top_k=20`.
|
| 45 |
+
> + Terminal-Bench 2.1: Evaluated under the Artificial Analysis (AA) protocol using the default Terminus 2 harness, a unified 2-hour timeout, the provided JSON parser in preserve-thinking mode, and 3 runs per task (mean). Decoding uses temperature=1.0, max_new_tokens=32K, with a 256K context window.
|
| 46 |
+
|
| 47 |
+
# Quickstart
|
| 48 |
+
## SGLang
|
| 49 |
+
The hardware- and recipe-specific launch matrix (BF16/FP8 × Low-Latency / High-Throughput), with a live command generator and verified configurations, lives in the SGLang cookbook:
|
| 50 |
+
|
| 51 |
+
**Cookbook:** https://docs.sglang.io/cookbook/autoregressive/InclusionAI/Ling-3.0-flash-VL
|
| 52 |
+
|
| 53 |
+
|
| 54 |
+
### Install SGLang
|
| 55 |
+
|
| 56 |
+
```bash
|
| 57 |
+
docker pull lmsysorg/sglang:dev-Ling-3.0-flash-VL
|
| 58 |
+
```
|
| 59 |
+
|
| 60 |
+
### Run Inference
|
| 61 |
+
Recommended recipe with 256K context (YaRN), on 4× 141GB-class GPUs (H20-3e / H200) or 4-GPU Blackwell nodes (B300 / GB300):
|
| 62 |
+
|
| 63 |
+
```bash
|
| 64 |
+
docker run --rm --gpus all --ipc=host --shm-size 32g \
|
| 65 |
+
-p 30000:30000 \
|
| 66 |
+
-e HF_TOKEN=<your-hf-token> \
|
| 67 |
+
lmsysorg/sglang:dev-Ling-3.0-flash-VL \
|
| 68 |
+
env SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1 \
|
| 69 |
+
python3 -m sglang.launch_server \
|
| 70 |
+
--model-path inclusionAI/Ling-3.0-flash-VL \
|
| 71 |
+
--tp 4 \
|
| 72 |
+
--context-length 262144 \
|
| 73 |
+
--json-model-override-args '{"rope_scaling":{"rope_type":"yarn","factor":2.0,"rope_theta":6000000,"partial_rotary_factor":0.5,"original_max_position_embeddings":131072}}' \
|
| 74 |
+
--mem-fraction-static 0.85 \
|
| 75 |
+
--trust-remote-code \
|
| 76 |
+
--reasoning-parser auto \
|
| 77 |
+
--tool-call-parser auto \
|
| 78 |
+
--host 0.0.0.0 \
|
| 79 |
+
--port 30000
|
| 80 |
+
```
|
| 81 |
+
|
| 82 |
+
On 80GB cards (H100 / H800), scale out to `--tp 8`. The reasoning and tool-call parsers resolve automatically to `ling3` from the chat template; you can also set them explicitly with `--reasoning-parser ling3 --tool-call-parser ling3`.
|
| 83 |
+
|
| 84 |
+
**Client**
|
| 85 |
+
|
| 86 |
+
Thinking is enabled by default by the chat template; disable it per request with `"chat_template_kwargs": {"enable_thinking": false}`. Recommended sampling: `temperature=1.0`, `top_p=0.95`, `top_k=20` (per `generation_config.json`).
|
| 87 |
+
|
| 88 |
+
```bash
|
| 89 |
+
curl -s http://localhost:30000/v1/chat/completions \
|
| 90 |
+
-H "Content-Type: application/json" \
|
| 91 |
+
-d '{"model": "inclusionAI/Ling-3.0-flash-VL",
|
| 92 |
+
"messages": [{"role": "user", "content": [
|
| 93 |
+
{"type": "image_url", "image_url": {"url": "https://example.com/image.png"}},
|
| 94 |
+
{"type": "text", "text": "Describe this image in one sentence."}
|
| 95 |
+
]}],
|
| 96 |
+
"stream": true,
|
| 97 |
+
"temperature": 1.0, "top_k": 20, "top_p": 0.95
|
| 98 |
+
}'
|
| 99 |
+
```
|
| 100 |
+
Video input uses `{"type": "video_url", "video_url": {"url": "..."}}` in the same message shape. For MMMU-Pro / `bench_serving` reproduction commands and per-hardware recipes, see the cookbook page linked above.
|
| 101 |
+
|
| 102 |
+
## vLLM
|
| 103 |
+
### Environment Preparation
|
| 104 |
+
|
| 105 |
+
```bash
|
| 106 |
+
pip install uv
|
| 107 |
+
|
| 108 |
+
uv venv ~/my_ling_env
|
| 109 |
+
|
| 110 |
+
source ~/my_ling_env/bin/activate
|
| 111 |
+
|
| 112 |
+
git clone https://github.com/inclusionAI/vllm-ling-v3.git
|
| 113 |
+
|
| 114 |
+
cd vllm-ling-v3
|
| 115 |
+
|
| 116 |
+
VLLM_USE_PRECOMPILED=1 uv pip install --editable . --torch-backend=auto
|
| 117 |
+
```
|
| 118 |
+
|
| 119 |
+
### Run Inference
|
| 120 |
+
|
| 121 |
+
**Server**
|
| 122 |
+
|
| 123 |
+
```bash
|
| 124 |
+
vllm serve "$MODEL_PATH" \
|
| 125 |
+
--port "$PORT" \
|
| 126 |
+
--trust-remote-code \
|
| 127 |
+
--served-model-name auto \
|
| 128 |
+
--tensor-parallel-size 4 \
|
| 129 |
+
--gpu-memory-utilization 0.85 \
|
| 130 |
+
--enable-prefix-caching \
|
| 131 |
+
--mamba-cache-mode align \
|
| 132 |
+
--enable-auto-tool-choice \
|
| 133 |
+
--tool-call-parser ling3 \
|
| 134 |
+
--reasoning-parser ling3
|
| 135 |
+
```
|
| 136 |
+
|
| 137 |
+
**Client**
|
| 138 |
+
|
| 139 |
+
Thinking is enabled by default by the chat template; disable it per request with `"chat_template_kwargs": {"enable_thinking": false}`. Recommended sampling: `temperature=1.0`, `top_p=0.95`, `top_k=20` (per `generation_config.json`).
|
| 140 |
+
|
| 141 |
+
```bash
|
| 142 |
+
curl -s http://${MASTER_IP}:${PORT}/v1/chat/completions \
|
| 143 |
+
-H "Content-Type: application/json" \
|
| 144 |
+
-d '{"model": "auto", -d '{"model": "inclusionAI/Ling-3.0-flash-VL",
|
| 145 |
+
"messages": [{"role": "user", "content": [
|
| 146 |
+
{"type": "image_url", "image_url": {"url": "https://example.com/image.png"}},
|
| 147 |
+
{"type": "text", "text": "Describe this image in one sentence."}
|
| 148 |
+
]}],
|
| 149 |
+
"stream": true,
|
| 150 |
+
"temperature": 1.0, "top_k": 20, "top_p": 0.95
|
| 151 |
+
}'
|
| 152 |
+
```
|
| 153 |
+
Video input uses `{"type": "video_url", "video_url": {"url": "..."}}` in the same message shape. For MMMU-Pro / `bench_serving` reproduction commands and per-hardware recipes, see the cookbook page linked above.
|