ScottHao commited on
Commit
f1fefd9
·
verified ·
1 Parent(s): ecc6146

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +2 -2
README.md CHANGED
@@ -13,7 +13,7 @@ We are introducing Ling-3.0-flash-VL, our next-generation native multimodal mode
13
  With 124B total parameters, only 5.5B activated parameters per token, support for image and video inputs, and a context window of up to 256K tokens, Ling-3.0-flash-VL delivers powerful multimodal reasoning and agentic capabilities with exceptional efficiency.
14
 
15
  # Model Overview
16
- Ling-3.0-flash-VL inherits the language, reasoning, and long-context capabilities of Ling-3.0-flash, while extending them with native image and video understanding. The model has 124B total parameters, with only 5.5B parameters activated per token, and supports a context window of up to 1M tokens.
17
 
18
  The architecture of Ling-3.0-flash-VL is designed to integrate visual information into real-world reasoning and agentic workflows.
19
 
@@ -25,7 +25,7 @@ The architecture of Ling-3.0-flash-VL is designed to integrate visual informatio
25
  Overall, these designs make vision more than just an input, integrating it into the complete process of understanding, reasoning, planning, acting, and verification.
26
 
27
 
28
- ![ling-3.0-flash-vl-0906](https://cdn-uploads.huggingface.co/production/uploads/6666ca359f5a0b3229238a1a/WhsE3cM7QMelBCjdeyex2.png)
29
 
30
  # Evaluation
31
  Ling-3.0-flash-VL achieves a score of **42** on the Artificial Analysis Intelligence Index v4.1.1, improving by 4 points over Ling-3.0-flash’s score of 38. The results show that extending the model with visual capabilities further improves its overall intelligence performance.
 
13
  With 124B total parameters, only 5.5B activated parameters per token, support for image and video inputs, and a context window of up to 256K tokens, Ling-3.0-flash-VL delivers powerful multimodal reasoning and agentic capabilities with exceptional efficiency.
14
 
15
  # Model Overview
16
+ Ling-3.0-flash-VL inherits the language, reasoning, and long-context capabilities of Ling-3.0-flash, while extending them with native image and video understanding. The model has 124B total parameters, with only 5.5B parameters activated per token, and supports a context window of up to 256K tokens.
17
 
18
  The architecture of Ling-3.0-flash-VL is designed to integrate visual information into real-world reasoning and agentic workflows.
19
 
 
25
  Overall, these designs make vision more than just an input, integrating it into the complete process of understanding, reasoning, planning, acting, and verification.
26
 
27
 
28
+ ![ling-3.0-flash-vl-0910](https://cdn-uploads.huggingface.co/production/uploads/6666ca359f5a0b3229238a1a/IL2eS4KUbKLYCWcoFC-Bq.png)
29
 
30
  # Evaluation
31
  Ling-3.0-flash-VL achieves a score of **42** on the Artificial Analysis Intelligence Index v4.1.1, improving by 4 points over Ling-3.0-flash’s score of 38. The results show that extending the model with visual capabilities further improves its overall intelligence performance.