michael-qiu vitoyy commited on
Commit
f48e5cc
·
1 Parent(s): c7defa0

Update README.md (#1)

Browse files

- Update README.md (71af1b053821e846ee311462d9b28308c7f823a1)


Co-authored-by: Yue yu <vitoyy@users.noreply.huggingface.co>

Files changed (1) hide show
  1. README.md +6 -4
README.md CHANGED
@@ -58,23 +58,25 @@ docker pull lmsysorg/sglang:dev-Ling-3.0-flash-VL
58
  ```
59
 
60
  ### Run Inference
61
- Recommended recipe with 256K context (YaRN), on 4× 141GB-class GPUs (H20-3e / H200) or 4-GPU Blackwell nodes (B300 / GB300):
62
 
63
  ```bash
64
  docker run --rm --gpus all --ipc=host --shm-size 32g \
65
  -p 30000:30000 \
66
  -e HF_TOKEN=<your-hf-token> \
 
67
  lmsysorg/sglang:dev-Ling-3.0-flash-VL \
68
- env SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1 \
69
  python3 -m sglang.launch_server \
70
- --model-path inclusionAI/Ling-3.0-flash-VL \
71
- --tp 4 \
72
  --context-length 262144 \
73
  --json-model-override-args '{"rope_scaling":{"rope_type":"yarn","factor":2.0,"rope_theta":6000000,"partial_rotary_factor":0.5,"original_max_position_embeddings":131072}}' \
74
  --mem-fraction-static 0.85 \
75
  --trust-remote-code \
76
  --reasoning-parser auto \
77
  --tool-call-parser auto \
 
 
78
  --host 0.0.0.0 \
79
  --port 30000
80
  ```
 
58
  ```
59
 
60
  ### Run Inference
61
+ Recommended recipe with 256K context (YaRN), on 2× 141GB-class GPUs (H20-3e / H200) or 2-GPU Blackwell nodes (B300 / GB300):
62
 
63
  ```bash
64
  docker run --rm --gpus all --ipc=host --shm-size 32g \
65
  -p 30000:30000 \
66
  -e HF_TOKEN=<your-hf-token> \
67
+ -e SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1 \
68
  lmsysorg/sglang:dev-Ling-3.0-flash-VL \
 
69
  python3 -m sglang.launch_server \
70
+ --model-path inclusionAI/Ling-3.0-flash-VL-int4 \
71
+ --tp-size 2 \
72
  --context-length 262144 \
73
  --json-model-override-args '{"rope_scaling":{"rope_type":"yarn","factor":2.0,"rope_theta":6000000,"partial_rotary_factor":0.5,"original_max_position_embeddings":131072}}' \
74
  --mem-fraction-static 0.85 \
75
  --trust-remote-code \
76
  --reasoning-parser auto \
77
  --tool-call-parser auto \
78
+ --enable-fp32-lm-head \
79
+ --disable-shared-experts-fusion \
80
  --host 0.0.0.0 \
81
  --port 30000
82
  ```