Instructions to use inclusionAI/Ling-3.0-flash-VL with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Inference
- HuggingChat
- Notebooks
- Google Colab
- Kaggle
Update README.md
Browse files
README.md
CHANGED
|
@@ -13,7 +13,7 @@ We are introducing Ling-3.0-flash-VL, our next-generation native multimodal mode
|
|
| 13 |
With 124B total parameters, only 5.5B activated parameters per token, support for image and video inputs, and a context window of up to 256K tokens, Ling-3.0-flash-VL delivers powerful multimodal reasoning and agentic capabilities with exceptional efficiency.
|
| 14 |
|
| 15 |
# Model Overview
|
| 16 |
-
Ling-3.0-flash-VL inherits the language, reasoning, and long-context capabilities of Ling-3.0-flash, while extending them with native image and video understanding. The model has 124B total parameters, with only 5.5B parameters activated per token, and supports a context window of up to
|
| 17 |
|
| 18 |
The architecture of Ling-3.0-flash-VL is designed to integrate visual information into real-world reasoning and agentic workflows.
|
| 19 |
|
|
@@ -25,7 +25,7 @@ The architecture of Ling-3.0-flash-VL is designed to integrate visual informatio
|
|
| 25 |
Overall, these designs make vision more than just an input, integrating it into the complete process of understanding, reasoning, planning, acting, and verification.
|
| 26 |
|
| 27 |
|
| 28 |
-

|
| 29 |
|
| 30 |
# Evaluation
|
| 31 |
Ling-3.0-flash-VL achieves a score of **42** on the Artificial Analysis Intelligence Index v4.1.1, improving by 4 points over Ling-3.0-flash’s score of 38. The results show that extending the model with visual capabilities further improves its overall intelligence performance.
|