Update README.md
Browse files
README.md
CHANGED
|
@@ -95,12 +95,12 @@ python -m fastdeploy.entrypoints.openai.api_server \
|
|
| 95 |
--max-num-seqs 32
|
| 96 |
```
|
| 97 |
|
| 98 |
-
To deploy the WINT2 quantized version using FastDeploy on
|
| 99 |
|
| 100 |
|
| 101 |
```bash
|
| 102 |
python -m fastdeploy.entrypoints.openai.api_server \
|
| 103 |
-
--model "baidu/ERNIE-4.5-300B-A47B-2Bits-
|
| 104 |
--port 8180 \
|
| 105 |
--metrics-port 8181 \
|
| 106 |
--engine-worker-queue-port 8182 \
|
|
|
|
| 95 |
--max-num-seqs 32
|
| 96 |
```
|
| 97 |
|
| 98 |
+
To deploy the WINT2 quantized version using FastDeploy on four GPUs, run the following command.
|
| 99 |
|
| 100 |
|
| 101 |
```bash
|
| 102 |
python -m fastdeploy.entrypoints.openai.api_server \
|
| 103 |
+
--model "baidu/ERNIE-4.5-300B-A47B-2Bits-TP4-Paddle" \
|
| 104 |
--port 8180 \
|
| 105 |
--metrics-port 8181 \
|
| 106 |
--engine-worker-queue-port 8182 \
|