|
Download README.md from baidu/ERNIE-4.5-300B-A47B-Paddle: direct link, hf CLI and curl.
- Browser
- Download file 11.1 kB
-
https://huggingface.co/baidu/ERNIE-4.5-300B-A47B-Paddle/resolve/c170ab0bd8be806dd9aee3d1fd087a1b9a0d2cd8/README.md
- Command line
-
hf download hf://baidu/ERNIE-4.5-300B-A47B-Paddle@c170ab0bd8be806dd9aee3d1fd087a1b9a0d2cd8/README.md
-
curl -L -o README.md https://huggingface.co/baidu/ERNIE-4.5-300B-A47B-Paddle/resolve/c170ab0bd8be806dd9aee3d1fd087a1b9a0d2cd8/README.md
11.1 kB
| # ERNIE-4.5-300B-A47B | |
| ## Features | |
| Based on ERNIE-4.5-Base's text-specific parameters, we conduct post-training in two phases—supervised fine-tuning (SFT) and reinforcement learning (RL). Utilizing the unified rewarding system which accommodates both reasoning and general tasks, we employ a progressive RL training with the Unified Preference Optimization objective. We obtain superior performance across a broad spectrum of tasks, including creative writing, mathematical reasoning, and code generation. We summarize the features of our model as follows: | |
| * **Superior Performance in instruction following tasks**: Demonstrates exceptional capability to understand, prioritize, and execute complex user instructions and achieves top results on both internal human evaluation dataset and public benchmarks for instruction following. | |
| * **Factual Reliability & Knowledge Accuracy**: Excels on SimpleQA and CN-SimpleQA benchmarks, dramatically reducing hallucinations and ensuring responses are grounded in verifiable facts. | |
| * **Expressive**, **Persuasive** **responses**: Generates engaging, well-structured responses with personality, clear rationale, and persuasive argumentation — making it ideal for interactive and content-creation scenarios. | |
| * **Advanced Math & Coding Proficiency**: Delivers precise, step-by-step mathematical reasoning and generates clean, efficient code across multiple programming languages, supporting both algorithmic problem solving and real-world development. | |
| ## Model Overview | |
| ERNIE-4.5-300B-A47B is a text MoE Post-trained model, with 300B total parameters and 47B activated parameters for each token. The following are the model configuration details: | |
| |Key|Value| | |
| |-|-| | |
| |Modality|Text| | |
| |Training Stage|Pretraining| | |
| |Params(Total / Activated)|300B / 47B| | |
| |Layers|54| | |
| |Heads(Q/KV)|64 / 8| | |
| |Text Experts(Total / Activated)|64 / 8| | |
| |Vision Experts(Total / Activated)|64 / 8| | |
| |Context Length|131072| | |
| ## Quickstart | |
| ### Model Finetuning with ERNIEKit | |
| [ERNIEKit](https://github.com/PaddlePaddle/ERNIE) is a training toolkit based on PaddlePaddle, specifically designed for the ERNIE series of open-source large models. It provides comprehensive support for scenarios such as instruction fine-tuning (SFT, LoRA) and alignment training (DPO), ensuring optimal performance. | |
| Usage Examples: | |
| ```bash | |
| # SFT | |
| erniekit train --stage SFT --model_name_or_path /baidu/ERNIE-4.5-300B-A47B --train_dataset_path your_dataset_path | |
| # DPO | |
| erniekit train --stage DPO --model_name_or_path /baidu/ERNIE-4.5-300B-A47B --train_dataset_path your_dataset_path | |
| ``` | |
| For more detailed examples, including SFT with LoRA, multi-GPU configurations, and advanced scripts, please refer to the examples folder within the [ERNIEKit](https://github.com/PaddlePaddle/ERNIE) repository. | |
| ### Using `transformers` library | |
| The following contains a code snippet illustrating how to use the model generate content based on given inputs. | |
| ```python | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| model_name = "Baidu/ERNIE-4.5-300B-A47B" | |
| # load the tokenizer and the model | |
| tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True) | |
| model = AutoModelForCausalLM.from_pretrained(model_name, trust_remote_code=True) | |
| # prepare the model input | |
| prompt = "Give me a short introduction to large language model." | |
| messages = [ | |
| {"role": "user", "content": prompt} | |
| ] | |
| text = tokenizer.apply_chat_template( | |
| messages, | |
| tokenize=False, | |
| add_generation_prompt=True | |
| ) | |
| model_inputs = tokenizer([text], add_special_tokens=False, return_tensors="pt").to(model.device) | |
| # conduct text completion | |
| generated_ids = model.generate( | |
| model_inputs.input_ids, | |
| max_new_tokens=1024 | |
| ) | |
| output_ids = generated_ids[0][len(model_inputs.input_ids[0]):].tolist() | |
| # decode the generated ids | |
| generate_text = tokenizer.decode(output_ids, skip_special_tokens=True).strip("\n") | |
| print("generate_text:", generate_text) | |
| ``` | |
| ### Using FastDeploy | |
| Service deployment can be quickly completed using FastDeploy in the following command. For more detailed usage instructions, please refer to the [FastDeploy Repository](https://github.com/PaddlePaddle/FastDeploy). | |
| **Note**: To deploy on a configuration with 4 GPUs each having at least 80G of memory, specify ```--quantization wint4```. If you specify ```--quantization wint8```, then resources for 8 GPUs are required. | |
| ```bash | |
| python -m fastdeploy.entrypoints.openai.api_server \ | |
| --model BAIDU/ERNIE-4.5-300B-A47B-Paddle \ | |
| --port 8180 \ | |
| --metrics-port 8181 \ | |
| --quantization wint4 \ | |
| --tensor-parallel-size 8 \ | |
| --engine-worker-queue-port 8182 \ | |
| --max-model-len 32768 \ # Maximum supported number of tokens | |
| --max-num-seqs 32 # Maximum concurrent processing capacity | |
| ``` | |
| To deploy the W4A8C8 quantized version using FastDeploy, you can run the following command. | |
| ```bash | |
| python -m fastdeploy.entrypoints.openai.api_server \ | |
| --model BAIDU/ERNIE-4.5-300B-A47B-W4A8C8-TP4 \ | |
| --port 8180 \ | |
| --metrics-port 8181 \ | |
| --engine-worker-queue-port 8182 \ | |
| --tensor-parallel-size 4 \ | |
| --max-model-len 32768 \ # Maximum supported number of tokens | |
| --max-num-seqs 32 # Maximum concurrent processing capacity | |
| ``` | |
| To deploy the WINT2 quantized version using FastDeploy on a single 141G GPU, you can run the following command. | |
| ```bash | |
| python -m fastdeploy.entrypoints.openai.api_server \ | |
| --model "BAIDU/ERNIE-4.5-300B-A47B-2Bit" \ | |
| --port 8180 \ | |
| --metrics-port 8181 \ | |
| --engine-worker-queue-port 8182 \ | |
| --tensor-parallel-size 1 \ | |
| --max-model-len 32768 \ # Maximum supported number of tokens | |
| --max-num-seqs 128 # Maximum concurrent processing capacity | |
| ``` | |
| The following contains a code snippet illustrating how to use ERNIE-4.5-300B-A47B-FP8 generate content based on given inputs. | |
| ```bash | |
| from fastdeploy import LLM, SamplingParams | |
| prompts = [ | |
| "Hello, my name is", | |
| ] | |
| sampling_params = SamplingParams(temperature=0.8, top_p=0.95, max_tokens=128) | |
| model = "BAIDU/ERNIE-4.5-300B-A47B-FP8" | |
| llm = LLM(model=model, tensor_parallel_size=8, max_model_len=8192, num_gpu_blocks_override=1024, engine_worker_queue_port=9981) | |
| outputs = llm.generate(prompts, sampling_params) | |
| for output in outputs: | |
| prompt = output.prompt | |
| generated_text = output.outputs.text | |
| print("generated_text", generated_text) | |
| ``` | |
| ### Using vLLM | |
| vLLM is currently being adapted, priority can be given to using our fork repository [vllm](https://github.com/CSWYF3634076/vllm/tree/ernie) | |
| ```bash | |
| # 80G * 16 GPU | |
| vllm serve Baidu/ERNIE-4.5-300B-A47B --trust-remote-code | |
| ``` | |
| ```bash | |
| # FP8 online quantification 80G * 8 GPU | |
| vllm serve Baidu/ERNIE-4.5-300B-A47B --trust-remote-code --quantization fp8 | |
| ``` | |
| ## Best Practices | |
| ### **Sampling Parameters** | |
| To achieve optimal performance, we suggest using `Temperature=0.8`, `TopP=0.8`. | |
| ### Prompts for Web Search | |
| For Web Search, {references}, {date}, and {question} are arguments. | |
| For Chinese question, we use the prompt: | |
| ``` | |
| ernie_search_zh_prompt = \ | |
| '''下面你会收到当前时间、多个不同来源的参考文章和一段对话。你的任务是阅读多个参考文章,并根据参考文章中的信息回答对话中的问题。 | |
| 以下是当前时间和参考文章: | |
| --------- | |
| #当前时间 | |
| {date} | |
| #参考文章 | |
| {references} | |
| --------- | |
| 请注意: | |
| 1. 回答必须结合问题需求和当前时间,对参考文章的可用性进行判断,避免在回答中使用错误或过时的信息。 | |
| 2. 当参考文章中的信息无法准确地回答问题时,你需要在回答中提供获取相应信息的建议,或承认无法提供相应信息。 | |
| 3. 你需要优先根据百科、官网、权威机构、专业网站等高权威性来源的信息来回答问题。 | |
| 4. 回复需要综合参考文章中的相关数字、案例、法律条文、公式等信息,使你的答案更专业。 | |
| 5. 当问题属于创作类任务时,需注意以下维度: | |
| - 态度鲜明:观点、立场清晰明确,避免模棱两可,语言果断直接 | |
| - 文采飞扬:用词精准生动,善用修辞手法,增强感染力 | |
| - 有理有据:逻辑严密递进,结合权威数据/事实支撑论点 | |
| --------- | |
| 下面请结合以上信息,回答问题,补全对话 | |
| {question}''' | |
| ``` | |
| For English question, we use the prompt: | |
| ``` | |
| ernie_search_en_prompt = \ | |
| ''' | |
| Below you will be given the current time, multiple references from different sources, and a conversation. Your task is to read the references and use the information in them to answer the question in the conversation. | |
| Here are the current time and the references: | |
| --------- | |
| #Current Time | |
| {date} | |
| #References | |
| {references} | |
| --------- | |
| Please note: | |
| 1. Based on the question’s requirements and the current time, assess the usefulness of the references to avoid using inaccurate or outdated information in the answer. | |
| 2. If the references do not provide enough information to accurately answer the question, you should suggest how to obtain the relevant information or acknowledge that you are unable to provide it. | |
| 3. Prioritize using information from highly authoritative sources such as encyclopedias, official websites, authoritative institutions, and professional websites when answering questions. | |
| 4. Incorporate relevant numbers, cases, legal provisions, formulas, and other details from the references to make your answer more professional. | |
| 5. For creative tasks, keep these dimensions in mind: | |
| - Clear attitude: Clear views and positions, avoid ambiguity, and use decisive and direct language | |
| - Brilliant writing: Precise and vivid words, good use of rhetoric, and enhance the appeal | |
| - Well-reasoned: Rigorous logic and progressive, combined with authoritative data/facts to support the argument | |
| --------- | |
| Now, using the information above, answer the question and complete the conversation: | |
| {question}''' | |
| ``` | |
| Parameter notes: | |
| * {question} is the user’s question | |
| * {date} is the current time, and the recommended format is “YYYY-MM-DD HH:MM:SS, Day of the Week, Beijing/China.” | |
| * {references} is the references, and the recommended format is: | |
| ``` | |
| ##参考文章1 | |
| 标题:周杰伦 | |
| 文章发布时间:2025-04-20 | |
| 内容:周杰伦(Jay Chou),1979年1月18日出生于台湾省新北市,祖籍福建省永春县,华语流行乐男歌手、音乐人、演员、导演、编剧,毕业于淡江中学。2000年,发行个人首张音乐专辑《Jay》。... | |
| 来源网站网址:baike.baidu.com | |
| 来源网站的网站名:百度百科 | |
| ##参考文章2 | |
| ... | |
| ``` | |
| ## License | |
| The ERNIE 4.5 models are provided under the Apache License 2.0. This license permits commercial use, subject to its terms and conditions. Copyright (c) 2025 Baidu, Inc. All Rights Reserved. | |
| ## Citation | |
| If you find ERNIE 4.5 useful or wish to use it in your projects, please kindly cite our technical report: | |
| ```bibtex | |
| @misc{ernie2025technicalreport, | |
| title={ERNIE 4.5 Technical Report}, | |
| author={Baidu ERNIE Team}, | |
| year={2025}, | |
| eprint={}, | |
| archivePrefix={arXiv}, | |
| primaryClass={cs.CL}, | |
| url={} | |
| } | |
| ``` | |