Instructions to use Agnes-AI/Agnes-3.0-Flash with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Agnes-AI/Agnes-3.0-Flash with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Agnes-AI/Agnes-3.0-Flash", trust_remote_code=True) messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("Agnes-AI/Agnes-3.0-Flash", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Agnes-AI/Agnes-3.0-Flash with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Agnes-AI/Agnes-3.0-Flash" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Agnes-AI/Agnes-3.0-Flash", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Agnes-AI/Agnes-3.0-Flash
- SGLang
How to use Agnes-AI/Agnes-3.0-Flash with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Agnes-AI/Agnes-3.0-Flash" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Agnes-AI/Agnes-3.0-Flash", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Agnes-AI/Agnes-3.0-Flash" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Agnes-AI/Agnes-3.0-Flash", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Agnes-AI/Agnes-3.0-Flash with Docker Model Runner:
docker model run hf.co/Agnes-AI/Agnes-3.0-Flash
Download sglang_patch/README.md from Agnes-AI/Agnes-3.0-Flash: direct link, hf CLI and curl.
- Browser
- Download file 2.65 kB
-
https://huggingface.co/Agnes-AI/Agnes-3.0-Flash/resolve/24f712ce59379b54c4a141d2708c35daf5ff613b/sglang_patch/README.md
- Command line
-
hf download hf://Agnes-AI/Agnes-3.0-Flash@24f712ce59379b54c4a141d2708c35daf5ff613b/sglang_patch/README.md
-
curl -L -o README.md https://huggingface.co/Agnes-AI/Agnes-3.0-Flash/resolve/24f712ce59379b54c4a141d2708c35daf5ff613b/sglang_patch/README.md
sglang patch for Agnes 3.0 Flash
Serve with a stock sglang image; serve.sh overlays three files onto the image's
sglang package and starts the server.
| Variant | Source | Status |
|---|---|---|
nightly-dev-20260908-20ca564b/ |
lmsysorg/sglang:nightly-dev-20260908-20ca564b |
served and measured (see below) |
v0.5.19/ |
sglang==0.5.19 |
generated from the release source, not served-tested |
| other versions | apply_patch.py <path/to/sglang> |
patched in place by serve.sh when no variant matches |
docker run --gpus all --shm-size 64g -p 30001:30002 \
-v /path/to/agnes-3.0-flash:/model \
lmsysorg/sglang:nightly-dev-20260908-20ca564b \
bash /model/serve.sh # extra sglang args may follow, e.g. --tp 2
The three files
| File | Change |
|---|---|
sglang/srt/configs/agnes.py |
new. Reads the model_type: agnes config and presents it to the server in the terms of its built-in hybrid (delta-rule + global attention) implementation: layer plan from global_attention_interval, the parallel FFN width added to intermediate_size, the checkpoint directory recorded as a config field for the loader. |
sglang/srt/utils/hf_transformers/common.py |
+3 lines at the end: registers AgnesConfig for model_type agnes. |
sglang/srt/models/qwen3_5.py |
one generator in front of the weight stream in load_weights: delta_attn.* → linear_attn.*, global_attn.* → self_attn.*, and each layer's mlp.parallel_ffn.{gate,up,down}_proj concatenated onto the main projections (gate/up along the output dim, down along the input dim). Checkpoints without a parallel branch pass through untouched. |
Nothing else in the image is modified. --trust-remote-code is required because sglang
resolves the model configuration through transformers first, which reads the
configuration_agnes.py shipped with the checkpoint. serve.sh also exports
AGNES_MODEL_PATH as a fallback for the loader.
Numerics
Folding the parallel branch into the main MLP changes the reduction length of the down projection, so served logits are not bit-identical to the transformers implementation. Measured on the nightly image (TP1, H200, 2144 teacher-forced positions, full 248 320-way softmax): full-vocabulary KL 5.9e-4 against the same weights served without the branch by the unpatched engine, below the 6.5e-4 measured between the transformers and sglang implementations of one and the same checkpoint. Perplexity 17.07 vs 17.05.
apply_patch.py anchors on QWEN3_5_KV_SCALE_MAPPER in models/qwen3_5.py and on the end of
utils/hf_transformers/common.py; it is idempotent.