Update README.md
#2
by sunzhongkai588 - opened
README.md
CHANGED
|
@@ -32,6 +32,9 @@ library_name: PaddlePaddle
|
|
| 32 |
|
| 33 |
# ERNIE-4.5-300B-A47B
|
| 34 |
|
|
|
|
|
|
|
|
|
|
| 35 |
## ERNIE 4.5 Highlights
|
| 36 |
|
| 37 |
The advanced capabilities of the ERNIE 4.5 models, particularly the MoE-based A47B and A3B series, are underpinned by several key technical innovations:
|
|
@@ -145,57 +148,6 @@ for output in outputs:
|
|
| 145 |
print("generated_text", generated_text)
|
| 146 |
```
|
| 147 |
|
| 148 |
-
### Using `transformers` library
|
| 149 |
-
|
| 150 |
-
The following contains a code snippet illustrating how to use the model generate content based on given inputs.
|
| 151 |
-
|
| 152 |
-
```python
|
| 153 |
-
from transformers import AutoModelForCausalLM, AutoTokenizer
|
| 154 |
-
|
| 155 |
-
model_name = "baidu/ERNIE-4.5-300B-A47B-PT"
|
| 156 |
-
|
| 157 |
-
# load the tokenizer and the model
|
| 158 |
-
tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
|
| 159 |
-
model = AutoModelForCausalLM.from_pretrained(model_name, trust_remote_code=True)
|
| 160 |
-
|
| 161 |
-
# prepare the model input
|
| 162 |
-
prompt = "Give me a short introduction to large language model."
|
| 163 |
-
messages = [
|
| 164 |
-
{"role": "user", "content": prompt}
|
| 165 |
-
]
|
| 166 |
-
text = tokenizer.apply_chat_template(
|
| 167 |
-
messages,
|
| 168 |
-
tokenize=False,
|
| 169 |
-
add_generation_prompt=True
|
| 170 |
-
)
|
| 171 |
-
model_inputs = tokenizer([text], add_special_tokens=False, return_tensors="pt").to(model.device)
|
| 172 |
-
|
| 173 |
-
# conduct text completion
|
| 174 |
-
generated_ids = model.generate(
|
| 175 |
-
model_inputs.input_ids,
|
| 176 |
-
max_new_tokens=1024
|
| 177 |
-
)
|
| 178 |
-
output_ids = generated_ids[0][len(model_inputs.input_ids[0]):].tolist()
|
| 179 |
-
|
| 180 |
-
# decode the generated ids
|
| 181 |
-
generate_text = tokenizer.decode(output_ids, skip_special_tokens=True).strip("\n")
|
| 182 |
-
print("generate_text:", generate_text)
|
| 183 |
-
```
|
| 184 |
-
|
| 185 |
-
### Using vLLM
|
| 186 |
-
|
| 187 |
-
vLLM is currently being adapted, priority can be given to using our forked repository [vllm](https://github.com/CSWYF3634076/vllm/tree/ernie). We are working with the community to fully support ERNIE4.5 models, stay tuned.
|
| 188 |
-
|
| 189 |
-
```bash
|
| 190 |
-
# 80G * 16 GPU
|
| 191 |
-
vllm serve baidu/ERNIE-4.5-300B-A47B-PT --trust-remote-code
|
| 192 |
-
```
|
| 193 |
-
|
| 194 |
-
```bash
|
| 195 |
-
# FP8 online quantification 80G * 8 GPU
|
| 196 |
-
vllm serve baidu/ERNIE-4.5-300B-A47B-PT --trust-remote-code --quantization fp8
|
| 197 |
-
```
|
| 198 |
-
|
| 199 |
## Best Practices
|
| 200 |
|
| 201 |
### **Sampling Parameters**
|
|
|
|
| 32 |
|
| 33 |
# ERNIE-4.5-300B-A47B
|
| 34 |
|
| 35 |
+
> [!NOTE]
|
| 36 |
+
> Note: "**-Paddle**" models use [PaddlePaddle](https://github.com/PaddlePaddle/Paddle) weights, while "**-PT**" models use Transformer-style PyTorch weights.
|
| 37 |
+
|
| 38 |
## ERNIE 4.5 Highlights
|
| 39 |
|
| 40 |
The advanced capabilities of the ERNIE 4.5 models, particularly the MoE-based A47B and A3B series, are underpinned by several key technical innovations:
|
|
|
|
| 148 |
print("generated_text", generated_text)
|
| 149 |
```
|
| 150 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 151 |
## Best Practices
|
| 152 |
|
| 153 |
### **Sampling Parameters**
|