How to use from
Docker Model Runner
docker model run hf.co/AtomicChat/Ling-3.0-flash-VL-GGUF:
Quick Links

How to Run Ling 3.0 Flash VL Locally

Built from InclusionAI's original weights using Atomic Chat's existing Ling importance matrix. The calibration corpora behind our builds are public.

Atomic Chat Discord GitHub
  • Ling 3.0 Flash VL is InclusionAI's vision-language model for text and image inputs.
  • Choose AD-Q4_K_M, AD-Q5_K_M, AD-Q6_K, or AD-Q8_0. Image inputs use the shared F32 vision projector.
  • These builds are an experimental preview. Use the accompanying runtime; full quality validation is still in progress.

Prepared from the original inclusionAI checkpoint, revision 869591498e8dbb41d4d96e3e2a5b428a2f70eb1e. These are the Atomic AD layouts, made on CPU using the existing calibration iMatrix. The full nine-variant text comparison and CPU speed measurements are complete; see FINAL_REPORT.md. This repository is an experimental working preview, not a completed benchmark release.

Uploads arrive progressively. Check UPLOAD_STATUS.json before downloading a variant. Each language variant requires all six GGUF shards in the same directory; select shard 00001 when loading.

Variant Complete language files, decimal GB
AD-Q4_K_M 79.30
AD-Q5_K_M 89.44
AD-Q6_K 107.14
AD-Q8_0 132.73

Vision requires the shared F32 mmproj, an additional 1.74 GB. File size is not a RAM requirement estimate.

Runtime requirement

Use the accompanying private runtime patch and added source files, based on AtomicBot-ai/atomic-llama-cpp-turboquant@cd560939087c95b93a1f30a95603d6b079436952. Stock llama.cpp and the released Atomic Chat app have not been validated for these artifacts. See RUNTIME.md.

Checks completed

All four variants: 917 tensor types/shapes verified, six-shard integrity checked, complete SHA-256 manifests, and all 382 protected F32 tensor payloads unchanged from BF16. Text and a spatial image passed on all four. Q4/Q6/Q8 additionally passed the OCR and object-count smoke cases. These are functional smoke checks, not a comprehensive vision benchmark.

Full text quality comparisons use the historical held-out 92 × 4096 protocol and a fresh BF16 reference from this checkpoint. Full text results are available in FINAL_REPORT.md. Video and maximum context are not validated.

The existing iMatrix comes from AtomicChat/Ling-3.0-flash-GGUF@253738fe190c15f329001f263f355fc1562bbe7c, SHA-256 7d3c0ebe9eb235cc08e0b7c91886f5c422772ea95eebbf0f53b0974a1c040991. Its 573 entries match the new language tensor dimensions. Four routed experts have no observations in that matrix; uniform importance was used for those entries.

manifest.json and SHA256SUMS describe the complete intended set. Actual upload completion is recorded separately in UPLOAD_STATUS.json.

Completed measurements

Full report · KLD chart · Metrics CSV.

The pilot results are separate from the full 92-block comparison. Completion is not a comprehensive vision or agentic quality certification.

Downloads last month
667
GGUF
Model size
124B params
Architecture
bailingmoe3
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AtomicChat/Ling-3.0-flash-VL-GGUF

Quantized
(2)
this model