Instructions to use dawncr0w/Hy-MT2-30B-A3B-oQ8-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use dawncr0w/Hy-MT2-30B-A3B-oQ8-MLX with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download dawncr0w/Hy-MT2-30B-A3B-oQ8-MLX --local-dir Hy-MT2-30B-A3B-oQ8-MLX
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Upload Hy-MT2-30B-A3B oQ8 MLX quantization
Browse files- LICENSE.txt +80 -0
- NOTICE +7 -0
- README.md +84 -0
- chat_template.jinja +233 -0
- config.json +53 -0
- hy_v3.py +405 -0
- model-00001-of-00007.safetensors +3 -0
- model-00002-of-00007.safetensors +3 -0
- model-00003-of-00007.safetensors +3 -0
- model-00004-of-00007.safetensors +3 -0
- model-00005-of-00007.safetensors +3 -0
- model-00006-of-00007.safetensors +3 -0
- model-00007-of-00007.safetensors +3 -0
- model.safetensors.index.json +0 -0
- special_tokens_map.json +23 -0
- tokenizer.json +0 -0
- tokenizer_config.json +0 -0
LICENSE.txt
ADDED
|
@@ -0,0 +1,80 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
TENCENT HY COMMUNITY LICENSE AGREEMENT
|
| 2 |
+
Tencent Hy-MT2 Release Date: May 21, 2026
|
| 3 |
+
THIS LICENSE AGREEMENT DOES NOT APPLY IN THE EUROPEAN UNION, AND IS EXPRESSLY LIMITED TO THE TERRITORY, AS DEFINED BELOW.
|
| 4 |
+
|
| 5 |
+
By clicking to agree or by using, reproducing, modifying, distributing, performing or displaying any portion or element of the Tencent HY Works, including via any Hosted Service, You will be deemed to have recognized and accepted the content of this Agreement, which is effective immediately.
|
| 6 |
+
|
| 7 |
+
1. DEFINITIONS.
|
| 8 |
+
a. “Acceptable Use Policy” shall mean the policy made available by Tencent as set forth in the Exhibit A.
|
| 9 |
+
b. “Agreement” shall mean the terms and conditions for use, reproduction, distribution, modification, performance and displaying of Tencent HY Works or any portion or element thereof set forth herein.
|
| 10 |
+
c. “Documentation” shall mean the specifications, manuals and documentation for Tencent HY made publicly available by Tencent.
|
| 11 |
+
d. “Hosted Service” shall mean a hosted service offered via an application programming interface (API), web access, or any other electronic or remote means.
|
| 12 |
+
e. “Licensee,” “You” or “Your” shall mean a natural person or legal entity exercising the rights granted by this Agreement and/or using the Tencent HY Works for any purpose and in any field of use.
|
| 13 |
+
f. “Materials” shall mean, collectively, Tencent’s proprietary Tencent HY and Documentation (and any portion thereof) as made available by Tencent under this Agreement.
|
| 14 |
+
g. “Model Derivatives” shall mean all: (i) modifications to Tencent HY or any Model Derivative of Tencent HY; (ii) works based on Tencent HY or any Model Derivative of Tencent HY; or (iii) any other machine learning model which is created by transfer of patterns of the weights, parameters, operations, or Output of Tencent HY or any Model Derivative of Tencent HY, to that model in order to cause that model to perform similarly to Tencent HY or a Model Derivative of Tencent HY, including distillation methods, methods that use intermediate data representations, or methods based on the generation of synthetic data Outputs by Tencent HY or a Model Derivative of Tencent HY for training that model. For clarity, Outputs by themselves are not deemed Model Derivatives.
|
| 15 |
+
h. “Output” shall mean the information and/or content output of Tencent HY or a Model Derivative that results from operating or otherwise using Tencent HY or a Model Derivative, including via a Hosted Service.
|
| 16 |
+
i. “Tencent,” “We” or “Us” shall mean the applicable entity or entities in the Tencent corporate family that own(s) intellectual property or other rights embodied in or utilized by the Materials.
|
| 17 |
+
j. “Tencent HY” shall mean the large language models, text/image/video/audio/3D generation models, and multimodal large language models and their software and algorithms, including trained model weights, parameters (including optimizer states), machine-learning model code, inference-enabling code, training-enabling code, fine-tuning enabling code and other elements of the foregoing made publicly available by Us, including, without limitation to, Tencent Hy-MT2-1.8B released at https://huggingface.co/tencent/Hy-MT2-1.8B, https://modelscope.cn/models/Tencent-Hunyuan/Hy-MT2-1.8B; Tencent Hy-MT2-7B released at https://huggingface.co/tencent/Hy-MT2-7B, https://modelscope.cn/models/Tencent-Hunyuan/Hy-MT2-7B; Tencent Hy-MT2-30B-A3B released at https://huggingface.co/tencent/Hy-MT2-30B-A3B, https://modelscope.cn/models/Tencent-Hunyuan/Hy-MT2-30B-A3B; Tencent Hy-MT2-1.8B-FP8 released at https://huggingface.co/tencent/Hy-MT2-1.8B-FP8, https://modelscope.cn/models/Tencent-Hunyuan/Hy-MT2-1.8B-FP8; Tencent Hy-MT2-7B-FP8 released at https://huggingface.co/tencent/Hy-MT2-7B-FP8, https://modelscope.cn/models/Tencent-Hunyuan/Hy-MT2-7B-FP8; Tencent Hy-MT2-30B-A3B-FP8 released at https://huggingface.co/tencent/Hy-MT2-30B-A3B-FP8, https://modelscope.cn/models/Tencent-Hunyuan/Hy-MT2-30B-A3B-FP8; Hy-MT2-1.8B-GGUF released at https://huggingface.co/tencent/Hy-MT2-1.8B-GGUF, https://modelscope.cn/models/Tencent-Hunyuan/Hy-MT2-1.8B-GGUF; Hy-MT2-7B-GGUF released at https://huggingface.co/tencent/Hy-MT2-7B-GGUF, https://modelscope.cn/models/Tencent-Hunyuan/Hy-MT2-7B-GGUF; Hy-MT2-1.8B-2bit-GGUF released at https://huggingface.co/tencent/Hy-MT2-1.8B-2bit-GGUF, https://modelscope.cn/models/Tencent-Hunyuan/Hy-MT2-1.8B-2bit-GGUF; Hy-MT2-1.8B-2bit-GGUF released at https://huggingface.co/tencent/Hy-MT2-1.8B-1.25bit-GGUF, https://modelscope.cn/models/Tencent-Hunyuan/Hy-MT2-1.8B-1.25bit-GGUF.
|
| 18 |
+
k. “Tencent HY Works” shall mean: (i) the Materials; (ii) Model Derivatives; and (iii) all derivative works thereof.
|
| 19 |
+
l. “Territory” shall mean the worldwide territory, excluding the territory of the European Union.
|
| 20 |
+
m. “Third Party” or “Third Parties” shall mean individuals or legal entities that are not under common control with Us or You.
|
| 21 |
+
n. “including” shall mean including but not limited to.
|
| 22 |
+
2. GRANT OF RIGHTS.
|
| 23 |
+
We grant You, for the Territory only, a non-exclusive, non-transferable and royalty-free limited license under Tencent’s intellectual property or other rights owned by Us embodied in or utilized by the Materials to use, reproduce, distribute, create derivative works of (including Model Derivatives), and make modifications to the Materials, only in accordance with the terms of this Agreement and the Acceptable Use Policy, and You must not violate (or encourage or permit anyone else to violate) any term of this Agreement or the Acceptable Use Policy.
|
| 24 |
+
3. DISTRIBUTION.
|
| 25 |
+
You may, subject to Your compliance with this Agreement, distribute or make available to Third Parties the Tencent HY Works, exclusively in the Territory, provided that You meet all of the following conditions:
|
| 26 |
+
a. You must provide all such Third Party recipients of the Tencent HY Works or products or services using them a copy of this Agreement;
|
| 27 |
+
b. You must cause any modified files to carry prominent notices stating that You changed the files;
|
| 28 |
+
c. You are encouraged to: (i) publish at least one technology introduction blogpost or one public statement expressing Your experience of using the Tencent HY Works; and (ii) mark the products or services developed by using the Tencent HY Works to indicate that the product/service is “Powered by Tencent HY”; and
|
| 29 |
+
d. All distributions to Third Parties (other than through a Hosted Service) must be accompanied by a “Notice” text file that contains the following notice: “Tencent HY is licensed under the Tencent HY Community License Agreement, Copyright © 2026 Tencent. All Rights Reserved. The trademark rights of “Tencent HY” are owned by Tencent or its affiliate.”
|
| 30 |
+
e. In the event that You use, integrate, implement, or otherwise deploy the Tencent HY Works, in whole or in part, to provide, enable, or support any service, product, or functionality to third parties, You shall clearly, accurately, and prominently disclose to all end users the full legal name and entity of the actual provider of such service, product, or functionality. You shall expressly and conspicuously state that Tencent is not affiliated with, associated with, sponsoring, or endorsing any such service, product, or functionality. You shall not use or display any name, logo, trademark, trade name, or other indicia of Tencent in any manner that could be construed as, or be likely to create, confusion, deception, or a false impression regarding any relationship, affiliation, sponsorship, or endorsement by Tencent.
|
| 31 |
+
You may add Your own copyright statement to Your modifications and, except as set forth in this Section and in Section 5, may provide additional or different license terms and conditions for use, reproduction, or distribution of Your modifications, or for any such Model Derivatives as a whole, provided Your use, reproduction, modification, distribution, performance and display of the work otherwise complies with the terms and conditions of this Agreement (including as regards the Territory). If You receive Tencent HY Works from a Licensee as part of an integrated end user product, then this Section 3 of this Agreement will not apply to You.
|
| 32 |
+
4. ADDITIONAL COMMERCIAL TERMS.
|
| 33 |
+
If, on the Tencent HY version release date, the monthly active users of all products or services made available by or for Licensee is greater than 100 million monthly active users in the preceding calendar month, You must request a license from Tencent, which Tencent may grant to You in its sole discretion, and You are not authorized to exercise any of the rights under this Agreement unless or until Tencent otherwise expressly grants You such rights.
|
| 34 |
+
5. RULES OF USE.
|
| 35 |
+
a. Your use of the Tencent HY Works must comply with applicable laws and regulations (including trade compliance laws and regulations) and adhere to the Acceptable Use Policy for the Tencent HY Works, which is hereby incorporated by reference into this Agreement. You must include the use restrictions referenced in these Sections 5(a) and 5(b) as an enforceable provision in any agreement (e.g., license agreement, terms of use, etc.) governing the use and/or distribution of Tencent HY Works and You must provide notice to subsequent users to whom You distribute that Tencent HY Works are subject to the use restrictions in these Sections 5(a) and 5(b).
|
| 36 |
+
b. You must not use the Tencent HY Works or any Output or results of the Tencent HY Works to improve any other AI model (other than Tencent HY or Model Derivatives thereof).
|
| 37 |
+
c. You must not use, reproduce, modify, distribute, or display the Tencent HY Works, Output or results of the Tencent HY Works outside the Territory. Any such use outside the Territory is unlicensed and unauthorized under this Agreement.
|
| 38 |
+
6. INTELLECTUAL PROPERTY.
|
| 39 |
+
a. Subject to Tencent’s ownership of Tencent HY Works made by or for Tencent and intellectual property rights therein, conditioned upon Your compliance with the terms and conditions of this Agreement, as between You and Tencent, You will be the owner of any derivative works and modifications of the Materials and any Model Derivatives that are made by or for You.
|
| 40 |
+
b. No trademark licenses are granted under this Agreement, and in connection with the Tencent HY Works, Licensee may not use any name or mark owned by or associated with Tencent or any of its affiliates, except as required for reasonable and customary use in describing and distributing the Tencent HY Works. Tencent hereby grants You a license to use “Tencent HY” (the “Mark”) in the Territory solely as required to comply with the provisions of Section 3(c), provided that You comply with any applicable laws related to trademark protection. All goodwill arising out of Your use of the Mark will inure to the benefit of Tencent.
|
| 41 |
+
c. If You commence a lawsuit or other proceedings (including a cross-claim or counterclaim in a lawsuit) against Us or any person or entity alleging that the Materials or any Output, or any portion of any of the foregoing, infringe any intellectual property or other right owned or licensable by You, then all licenses granted to You under this Agreement shall terminate as of the date such lawsuit or other proceeding is filed. You will defend, indemnify and hold harmless Us from and against any claim by any Third Party arising out of or related to Your or the Third Party’s use or distribution of the Tencent HY Works.
|
| 42 |
+
d. Tencent claims no rights in Outputs You generate. You and Your users are solely responsible for Outputs and their subsequent uses.
|
| 43 |
+
7. DISCLAIMERS OF WARRANTY AND LIMITATIONS OF LIABILITY.
|
| 44 |
+
a. We are not obligated to support, update, provide training for, or develop any further version of the Tencent HY Works or to grant any license thereto.
|
| 45 |
+
b. UNLESS AND ONLY TO THE EXTENT REQUIRED BY APPLICABLE LAW, THE TENCENT HY WORKS AND ANY OUTPUT AND RESULTS THEREFROM ARE PROVIDED “AS IS” WITHOUT ANY EXPRESS OR IMPLIED WARRANTIES OF ANY KIND INCLUDING ANY WARRANTIES OF TITLE, MERCHANTABILITY, NONINFRINGEMENT, COURSE OF DEALING, USAGE OF TRADE, OR FITNESS FOR A PARTICULAR PURPOSE. YOU ARE SOLELY RESPONSIBLE FOR DETERMINING THE APPROPRIATENESS OF USING, REPRODUCING, MODIFYING, PERFORMING, DISPLAYING OR DISTRIBUTING ANY OF THE TENCENT HY WORKS OR OUTPUTS AND ASSUME ANY AND ALL RISKS ASSOCIATED WITH YOUR OR A THIRD PARTY’S USE OR DISTRIBUTION OF ANY OF THE TENCENT HY WORKS OR OUTPUTS AND YOUR EXERCISE OF RIGHTS AND PERMISSIONS UNDER THIS AGREEMENT.
|
| 46 |
+
c. TO THE FULLEST EXTENT PERMITTED BY APPLICABLE LAW, IN NO EVENT SHALL TENCENT OR ITS AFFILIATES BE LIABLE UNDER ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, TORT, NEGLIGENCE, PRODUCTS LIABILITY, OR OTHERWISE, FOR ANY DAMAGES, INCLUDING ANY DIRECT, INDIRECT, SPECIAL, INCIDENTAL, EXEMPLARY, CONSEQUENTIAL OR PUNITIVE DAMAGES, OR LOST PROFITS OF ANY KIND ARISING FROM THIS AGREEMENT OR RELATED TO ANY OF THE TENCENT HY WORKS OR OUTPUTS, EVEN IF TENCENT OR ITS AFFILIATES HAVE BEEN ADVISED OF THE POSSIBILITY OF ANY OF THE FOREGOING.
|
| 47 |
+
8. SURVIVAL AND TERMINATION.
|
| 48 |
+
a. The term of this Agreement shall commence upon Your acceptance of this Agreement or access to the Materials and will continue in full force and effect until terminated in accordance with the terms and conditions herein.
|
| 49 |
+
b. We may terminate this Agreement if You breach any of the terms or conditions of this Agreement. Upon termination of this Agreement, You must promptly delete and cease use of the Tencent HY Works. Sections 6(a), 6(c), 7 and 9 shall survive the termination of this Agreement.
|
| 50 |
+
9. GOVERNING LAW AND JURISDICTION.
|
| 51 |
+
a. This Agreement and any dispute arising out of or relating to it will be governed by the laws of the Hong Kong Special Administrative Region of the People’s Republic of China, without regard to conflict of law principles, and the UN Convention on Contracts for the International Sale of Goods does not apply to this Agreement.
|
| 52 |
+
b. Exclusive jurisdiction and venue for any dispute arising out of or relating to this Agreement will be a court of competent jurisdiction in the Hong Kong Special Administrative Region of the People’s Republic of China, and Tencent and Licensee consent to the exclusive jurisdiction of such court with respect to any such dispute.
|
| 53 |
+
|
| 54 |
+
EXHIBIT A
|
| 55 |
+
ACCEPTABLE USE POLICY
|
| 56 |
+
|
| 57 |
+
Tencent reserves the right to update this Acceptable Use Policy from time to time.
|
| 58 |
+
Last modified: December 30, 2025
|
| 59 |
+
|
| 60 |
+
Tencent endeavors to promote safe and fair use of its tools and features, including Tencent HY. You agree not to use Tencent HY or Model Derivatives:
|
| 61 |
+
1. Outside the Territory;
|
| 62 |
+
2. In any way that violates any applicable national, federal, state, local, international or any other law or regulation;
|
| 63 |
+
3. To harm Yourself or others;
|
| 64 |
+
4. To repurpose or distribute output from Tencent HY or any Model Derivatives to harm Yourself or others;
|
| 65 |
+
5. To override or circumvent the safety guardrails and safeguards We have put in place;
|
| 66 |
+
6. For the purpose of exploiting, harming or attempting to exploit or harm minors in any way;
|
| 67 |
+
7. To generate or disseminate verifiably false information and/or content with the purpose of harming others or influencing elections;
|
| 68 |
+
8. To generate or facilitate false online engagement, including fake reviews and other means of fake online engagement;
|
| 69 |
+
9. To intentionally defame, disparage or otherwise harass others;
|
| 70 |
+
10. To generate and/or disseminate malware (including ransomware) or any other content to be used for the purpose of harming electronic systems;
|
| 71 |
+
11. To generate or disseminate personal identifiable information with the purpose of harming others;
|
| 72 |
+
12. To generate or disseminate information (including images, code, posts, articles), and place the information in any public context (including –through the use of bot generated tweets), without expressly and conspicuously identifying that the information and/or content is machine generated;
|
| 73 |
+
13. To impersonate another individual without consent, authorization, or legal right;
|
| 74 |
+
14. To make high-stakes automated decisions in domains that affect an individual’s safety, rights or wellbeing (e.g., law enforcement, migration, medicine/health, management of critical infrastructure, safety components of products, essential services, credit, employment, housing, education, social scoring, or insurance);
|
| 75 |
+
15. In a manner that violates or disrespects the social ethics and moral standards of other countries or regions;
|
| 76 |
+
16. To perform, facilitate, threaten, incite, plan, promote or encourage violent extremism or terrorism;
|
| 77 |
+
17. For any use intended to discriminate against or harm individuals or groups based on protected characteristics or categories, online or offline social behavior or known or predicted personal or personality characteristics;
|
| 78 |
+
18. To intentionally exploit any of the vulnerabilities of a specific group of persons based on their age, social, physical or mental characteristics, in order to materially distort the behavior of a person pertaining to that group in a manner that causes or is likely to cause that person or another person physical or psychological harm;
|
| 79 |
+
19. For military purposes;
|
| 80 |
+
20. To engage in the unauthorized or unlicensed practice of any profession including, but not limited to, financial, legal, medical/health, or other professional practices.
|
NOTICE
ADDED
|
@@ -0,0 +1,7 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Tencent HY is licensed under the Tencent HY Community License Agreement, Copyright © 2026 Tencent. All Rights Reserved. The trademark rights of "Tencent HY" are owned by Tencent or its affiliate.
|
| 2 |
+
|
| 3 |
+
This repository contains an oQ4 MLX quantized derivative of tencent/Hy-MT2-30B-A3B.
|
| 4 |
+
|
| 5 |
+
Modified files in this distribution:
|
| 6 |
+
- config.json: adds MLX quantization metadata, model_file packaging, and rope_parameters compatibility metadata.
|
| 7 |
+
- hy_v3.py: packages the HYV3 MLX model implementation for local model_file loading, with a rope_theta compatibility fallback and absolute mlx_lm imports.
|
README.md
ADDED
|
@@ -0,0 +1,84 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
library_name: mlx
|
| 3 |
+
base_model: tencent/Hy-MT2-30B-A3B
|
| 4 |
+
tags:
|
| 5 |
+
- mlx
|
| 6 |
+
- oq8
|
| 7 |
+
- quantized
|
| 8 |
+
- translation
|
| 9 |
+
language:
|
| 10 |
+
- zh
|
| 11 |
+
- en
|
| 12 |
+
- fr
|
| 13 |
+
- pt
|
| 14 |
+
- es
|
| 15 |
+
- ja
|
| 16 |
+
- tr
|
| 17 |
+
- ru
|
| 18 |
+
- ar
|
| 19 |
+
- ko
|
| 20 |
+
- th
|
| 21 |
+
- it
|
| 22 |
+
- de
|
| 23 |
+
- vi
|
| 24 |
+
- ms
|
| 25 |
+
- id
|
| 26 |
+
- tl
|
| 27 |
+
- hi
|
| 28 |
+
- pl
|
| 29 |
+
- cs
|
| 30 |
+
- nl
|
| 31 |
+
- km
|
| 32 |
+
- my
|
| 33 |
+
- fa
|
| 34 |
+
- gu
|
| 35 |
+
- ur
|
| 36 |
+
- te
|
| 37 |
+
- mr
|
| 38 |
+
- he
|
| 39 |
+
- bn
|
| 40 |
+
- ta
|
| 41 |
+
- uk
|
| 42 |
+
- bo
|
| 43 |
+
- kk
|
| 44 |
+
- mn
|
| 45 |
+
- ug
|
| 46 |
+
license: other
|
| 47 |
+
---
|
| 48 |
+
|
| 49 |
+
# Hy-MT2-30B-A3B-oQ8-MLX
|
| 50 |
+
|
| 51 |
+
This is an oQ8 MLX quantized derivative of [tencent/Hy-MT2-30B-A3B](https://huggingface.co/tencent/Hy-MT2-30B-A3B).
|
| 52 |
+
|
| 53 |
+
The model was quantized locally with oMLX oQ8. The resulting config targets approximately 8.50 bpw and includes a packaged `hy_v3.py` model file so MLX-LM can load the HYV3 architecture from the model directory.
|
| 54 |
+
|
| 55 |
+
## Validation
|
| 56 |
+
|
| 57 |
+
Local validation completed with the oMLX app bundle on macOS:
|
| 58 |
+
|
| 59 |
+
```text
|
| 60 |
+
mlx_lm load: passed
|
| 61 |
+
generation smoke test: passed
|
| 62 |
+
prompt: Hello
|
| 63 |
+
max tokens: 4
|
| 64 |
+
peak memory: 29.815 GB
|
| 65 |
+
```
|
| 66 |
+
|
| 67 |
+
## Usage
|
| 68 |
+
|
| 69 |
+
Use an MLX-LM build that supports the APIs used by the packaged `hy_v3.py` file.
|
| 70 |
+
|
| 71 |
+
```bash
|
| 72 |
+
python -m mlx_lm generate \
|
| 73 |
+
--model /path/to/Hy-MT2-30B-A3B-oQ8-MLX \
|
| 74 |
+
--ignore-chat-template \
|
| 75 |
+
--prompt "Hello" \
|
| 76 |
+
--max-tokens 32 \
|
| 77 |
+
--temp 0
|
| 78 |
+
```
|
| 79 |
+
|
| 80 |
+
## License And Notice
|
| 81 |
+
|
| 82 |
+
The base model is distributed under the Tencent HY Community License Agreement. This distribution includes `LICENSE.txt` and `NOTICE` from/for the Tencent HY license requirements.
|
| 83 |
+
|
| 84 |
+
This repository is not affiliated with, sponsored by, or endorsed by Tencent.
|
chat_template.jinja
ADDED
|
@@ -0,0 +1,233 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{# ----------‑‑‑ special token variables ‑‑‑---------- #}
|
| 2 |
+
{%- set HYTK = '' %}
|
| 3 |
+
{%- set bos_token = '<|hy_begin▁of▁sentence|>' %}
|
| 4 |
+
{%- set pad_token = '<|hy_▁pad▁|>' %}
|
| 5 |
+
{%- set user_token = '<|hy_User|>' %}
|
| 6 |
+
{%- set assistant_token = '<|hy_Assistant|>' %}
|
| 7 |
+
{%- set eos_token = '<eos:6124c78e>' %}
|
| 8 |
+
{# ----------‑‑‑ tokens with md5 encoding (conditional on HYTK) ‑‑‑---------- #}
|
| 9 |
+
{%- if HYTK %}
|
| 10 |
+
{%- set think_begin_token = '<think:{}>'.format(HYTK) %}
|
| 11 |
+
{%- set think_end_token = '</think:{}>'.format(HYTK) %}
|
| 12 |
+
{%- set toolcalls_begin_token = '<tool_calls:{}>'.format(HYTK) %}
|
| 13 |
+
{%- set toolcalls_end_token = '</tool_calls:{}>'.format(HYTK) %}
|
| 14 |
+
{%- set toolcall_begin_token = '<tool_call:{}>'.format(HYTK) %}
|
| 15 |
+
{%- set toolcall_end_token = '</tool_call:{}>'.format(HYTK) %}
|
| 16 |
+
{%- set toolsep_token = '<tool_sep:{}>'.format(HYTK) %}
|
| 17 |
+
{%- set argkey_begin_token = '<arg_key:{}>'.format(HYTK) %}
|
| 18 |
+
{%- set argkey_end_token = '</arg_key:{}>'.format(HYTK) %}
|
| 19 |
+
{%- set argvalue_begin_token = '<arg_value:{}>'.format(HYTK) %}
|
| 20 |
+
{%- set argvalue_end_token = '</arg_value:{}>'.format(HYTK) %}
|
| 21 |
+
{%- set toolresponses_begin_token = '<tool_responses:{}>'.format(HYTK) %}
|
| 22 |
+
{%- set toolresponses_end_token = '</tool_responses:{}>'.format(HYTK) %}
|
| 23 |
+
{%- set toolresponse_begin_token = '<tool_response:{}>'.format(HYTK) %}
|
| 24 |
+
{%- set toolresponse_end_token = '</tool_response:{}>'.format(HYTK) %}
|
| 25 |
+
{%- set reasoning_mode_token = '<|reasoning_mode|>' %}
|
| 26 |
+
{%- set toolcalls_begin_ds_token = '<|tool▁calls▁begin|>' %}
|
| 27 |
+
{%- set toolcalls_end_ds_token = '<|tool▁calls▁end|>' %}
|
| 28 |
+
{%- set toolcall_begin_ds_token = '<|tool▁call▁begin|>' %}
|
| 29 |
+
{%- set toolcall_end_ds_token = '<|tool▁call▁end|>' %}
|
| 30 |
+
{%- set toolsep_ds_token = '<|tool▁sep|>' %}
|
| 31 |
+
{%- set tooloutput_begin_ds_token = '<|tool▁output▁begin|>' %}
|
| 32 |
+
{%- set tooloutput_end_ds_token = '<|tool▁output▁end|>' %}
|
| 33 |
+
{%- else %}
|
| 34 |
+
{%- set think_begin_token = '<think>' %}
|
| 35 |
+
{%- set think_end_token = '</think>' %}
|
| 36 |
+
{%- set toolcalls_begin_token = '<tool_calls>' %}
|
| 37 |
+
{%- set toolcalls_end_token = '</tool_calls>' %}
|
| 38 |
+
{%- set toolcall_begin_token = '<tool_call>' %}
|
| 39 |
+
{%- set toolcall_end_token = '</tool_call>' %}
|
| 40 |
+
{%- set toolsep_token = '<tool_sep>' %}
|
| 41 |
+
{%- set argkey_begin_token = '<arg_key>' %}
|
| 42 |
+
{%- set argkey_end_token = '</arg_key>' %}
|
| 43 |
+
{%- set argvalue_begin_token = '<arg_value>' %}
|
| 44 |
+
{%- set argvalue_end_token = '</arg_value>' %}
|
| 45 |
+
{%- set toolresponses_begin_token = '<tool_responses>' %}
|
| 46 |
+
{%- set toolresponses_end_token = '</tool_responses>' %}
|
| 47 |
+
{%- set toolresponse_begin_token = '<tool_response>' %}
|
| 48 |
+
{%- set toolresponse_end_token = '</tool_response>' %}
|
| 49 |
+
{%- set reasoning_mode_token = '<|reasoning_mode|>' %}
|
| 50 |
+
{%- set toolcalls_begin_ds_token = '<|tool▁calls▁begin|>' %}
|
| 51 |
+
{%- set toolcalls_end_ds_token = '<|tool▁calls▁end|>' %}
|
| 52 |
+
{%- set toolcall_begin_ds_token = '<|tool▁call▁begin|>' %}
|
| 53 |
+
{%- set toolcall_end_ds_token = '<|tool▁call▁end|>' %}
|
| 54 |
+
{%- set toolsep_ds_token = '<|tool▁sep|>' %}
|
| 55 |
+
{%- set tooloutput_begin_ds_token = '<|tool▁output▁begin|>' %}
|
| 56 |
+
{%- set tooloutput_end_ds_token = '<|tool▁output▁end|>' %}
|
| 57 |
+
{%- endif %}
|
| 58 |
+
{# ----------‑‑‑ hyperparameters variables ‑‑‑---------- #}
|
| 59 |
+
{%- if not add_generation_prompt is defined %}
|
| 60 |
+
{%- set add_generation_prompt = false %}
|
| 61 |
+
{%- endif %}
|
| 62 |
+
{%- if not interleaved_thinking is defined %}
|
| 63 |
+
{%- set interleaved_thinking = false %}
|
| 64 |
+
{%- endif %}
|
| 65 |
+
{%- if not tools %}
|
| 66 |
+
{%- set interleaved_thinking = false %}
|
| 67 |
+
{%- endif %}
|
| 68 |
+
{%- if not is_training is defined %}
|
| 69 |
+
{%- set is_training = false %}
|
| 70 |
+
{%- endif %}
|
| 71 |
+
{%- if not is_ai_search is defined %}
|
| 72 |
+
{%- set is_ai_search = false %}
|
| 73 |
+
{%- endif %}
|
| 74 |
+
{%- if not reasoning_effort is defined or reasoning_effort not in ['high', 'low', 'no_think'] %}
|
| 75 |
+
{%- set reasoning_effort = 'no_think' %}
|
| 76 |
+
{%- endif %}
|
| 77 |
+
|
| 78 |
+
{%- macro visible_text(content) -%}
|
| 79 |
+
{%- if content is string -%}
|
| 80 |
+
{{- content }}
|
| 81 |
+
{%- elif content is iterable and content is not mapping -%}
|
| 82 |
+
{%- for item in content -%}
|
| 83 |
+
{%- if item is mapping and item.type == 'text' -%}
|
| 84 |
+
{{- item.text }}
|
| 85 |
+
{%- elif item is string -%}
|
| 86 |
+
{{- item }}
|
| 87 |
+
{%- endif -%}
|
| 88 |
+
{%- endfor -%}
|
| 89 |
+
{%- elif content is none -%}
|
| 90 |
+
{{- '' }}
|
| 91 |
+
{%- else -%}
|
| 92 |
+
{{- content }}
|
| 93 |
+
{%- endif -%}
|
| 94 |
+
{%- endmacro -%}
|
| 95 |
+
|
| 96 |
+
{%- set ns = namespace(last_user_index=-1) %}
|
| 97 |
+
{%- set sp_ns = namespace(system_prompt='', is_first_sp=true) %}
|
| 98 |
+
{%- for message in messages %}
|
| 99 |
+
{%- if message['role'] == 'system' %}
|
| 100 |
+
{%- set sp_ns.system_prompt = sp_ns.system_prompt + visible_text(message['content']) %}
|
| 101 |
+
{%- endif %}
|
| 102 |
+
{%- if message['role'] == 'user' %}
|
| 103 |
+
{%- set ns.last_user_index = loop.index0 %}
|
| 104 |
+
{%- endif %}
|
| 105 |
+
{%- endfor %}
|
| 106 |
+
{%- if reasoning_effort is defined and reasoning_effort is string and reasoning_effort != '' and not tools %}
|
| 107 |
+
{%- set sp_ns.system_prompt = sp_ns.system_prompt + reasoning_mode_token + 'reasoning_effort:' + reasoning_effort %}
|
| 108 |
+
{%- endif %}
|
| 109 |
+
{{- bos_token }}
|
| 110 |
+
{{- sp_ns.system_prompt }}
|
| 111 |
+
{%- if tools %}
|
| 112 |
+
{%- if sp_ns.system_prompt != '' %}
|
| 113 |
+
{{- '\n\n# Tools\n\nYou may call one or more functions to assist with the user query.' }}
|
| 114 |
+
{%- else %}
|
| 115 |
+
{{- '# Tools\n\nYou may call one or more functions to assist with the user query.' }}
|
| 116 |
+
{%- endif %}
|
| 117 |
+
{{- '\n\nYou are provided with function signatures within <tools></tools> XML tags:' }}
|
| 118 |
+
{{- '\n<tools>\n' }}
|
| 119 |
+
{%- for tool in tools %}
|
| 120 |
+
{%- if loop.index0 > 0 %}
|
| 121 |
+
{{- '\n' }}
|
| 122 |
+
{%- endif %}
|
| 123 |
+
{{- tool | tojson }}
|
| 124 |
+
{%- endfor %}
|
| 125 |
+
{{- '\n</tools>\n\n' }}
|
| 126 |
+
{{- 'For function call returns, you should first print ' + toolcalls_begin_token + '\n' }}
|
| 127 |
+
{{- 'For each function call, you should return object like:\n' }}
|
| 128 |
+
{{- toolcall_begin_token + '{function-name}' + toolsep_token + '\n' }}
|
| 129 |
+
{{- argkey_begin_token + '{arg-key-1}' + argkey_end_token + '\n' }}
|
| 130 |
+
{{- argvalue_begin_token + '{arg-value-1}' + argvalue_end_token + '\n' }}
|
| 131 |
+
{{- argkey_begin_token + '{arg-key-2}' + argkey_end_token + '\n' }}
|
| 132 |
+
{{- argvalue_begin_token + '{arg-value-2}' + argvalue_end_token + '\n' }}
|
| 133 |
+
{{- '...\n' }}
|
| 134 |
+
{{- toolcalls_end_token + '\n' }}
|
| 135 |
+
{%- if reasoning_effort is defined and reasoning_effort is string and reasoning_effort != '' %}
|
| 136 |
+
{{- 'At the end of function call returns, you should print ' + toolcalls_end_token + reasoning_mode_token + 'reasoning_effort:' + reasoning_effort }}
|
| 137 |
+
{%- else %}
|
| 138 |
+
{{- 'At the end of function call returns, you should print ' + toolcalls_end_token }}
|
| 139 |
+
{%- endif %}
|
| 140 |
+
{%- endif %}
|
| 141 |
+
|
| 142 |
+
{%- set prev_ns = namespace(is_tool=false, is_tool_first=true) %}
|
| 143 |
+
{%- set last_ns = namespace(last_is_assistant=false) %}
|
| 144 |
+
{%- for message in messages %}
|
| 145 |
+
{%- if message['role'] == 'user' %}
|
| 146 |
+
{%- if prev_ns.is_tool and not is_ai_search %}
|
| 147 |
+
{{- toolresponses_end_token }}
|
| 148 |
+
{%- endif %}
|
| 149 |
+
{{- user_token + visible_text(message['content']) }}
|
| 150 |
+
{%- set prev_ns.is_tool = false %}
|
| 151 |
+
{%- endif %}
|
| 152 |
+
{%- if message['role'] == 'assistant' %}
|
| 153 |
+
{%- if is_training %}
|
| 154 |
+
{%- if 'reasoning_content' in message and message['reasoning_content'] is string %}
|
| 155 |
+
{%- set content = think_begin_token + message['reasoning_content'] + think_end_token + visible_text(message['content']) %}
|
| 156 |
+
{%- else %}
|
| 157 |
+
{%- set content = think_begin_token + think_end_token + visible_text(message['content']) %}
|
| 158 |
+
{%- endif %}
|
| 159 |
+
{%- else %}
|
| 160 |
+
{%- if interleaved_thinking %}
|
| 161 |
+
{%- if loop.index0 > ns.last_user_index and 'reasoning_content' in message and message['reasoning_content'] is string %}
|
| 162 |
+
{%- set content = think_begin_token + message['reasoning_content'] + think_end_token + visible_text(message['content']) %}
|
| 163 |
+
{%- else %}
|
| 164 |
+
{%- set content = think_begin_token + think_end_token + visible_text(message['content']) %}
|
| 165 |
+
{%- endif %}
|
| 166 |
+
{%- else %}
|
| 167 |
+
{%- set content = think_begin_token + think_end_token + visible_text(message['content']) %}
|
| 168 |
+
{%- endif %}
|
| 169 |
+
{%- endif %}
|
| 170 |
+
{%- if prev_ns.is_tool and not is_ai_search %}
|
| 171 |
+
{{- toolresponses_end_token }}
|
| 172 |
+
{%- endif %}
|
| 173 |
+
{{- assistant_token }}
|
| 174 |
+
{%- if message['tool_calls'] is defined and message['tool_calls'] %}
|
| 175 |
+
{%- set prev_ns.is_tool_first = true %}
|
| 176 |
+
{{- content }}
|
| 177 |
+
{{- toolcalls_begin_token + '\n' }}
|
| 178 |
+
{%- for tool in message['tool_calls'] %}
|
| 179 |
+
{%- set arguments = tool['function']['arguments'] %}
|
| 180 |
+
{{- toolcall_begin_token + tool['function']['name'] + toolsep_token + '\n' }}
|
| 181 |
+
{%- for key, value in arguments.items() %}
|
| 182 |
+
{{- argkey_begin_token + key + argkey_end_token + '\n' }}
|
| 183 |
+
{%- if value is not string %}
|
| 184 |
+
{%- set value = value | tojson(ensure_ascii=False) %}
|
| 185 |
+
{%- endif %}
|
| 186 |
+
{{- argvalue_begin_token + value + argvalue_end_token + '\n' }}
|
| 187 |
+
{%- endfor %}
|
| 188 |
+
{{- toolcall_end_token + '\n' }}
|
| 189 |
+
{%- endfor %}
|
| 190 |
+
{{- toolcalls_end_token + eos_token }}
|
| 191 |
+
{%- else %}
|
| 192 |
+
{%- if not loop.last or is_training %}
|
| 193 |
+
{{- content + eos_token }}
|
| 194 |
+
{%- else %}
|
| 195 |
+
{{- content }}
|
| 196 |
+
{%- endif %}
|
| 197 |
+
{%- endif %}
|
| 198 |
+
{%- set prev_ns.is_tool = false %}
|
| 199 |
+
{%- endif %}
|
| 200 |
+
{%- if message['role'] == 'tool' %}
|
| 201 |
+
{%- set prev_ns.is_tool = true %}
|
| 202 |
+
{%- if prev_ns.is_tool_first and not is_ai_search %}
|
| 203 |
+
{{- toolresponses_begin_token + '\n' }}
|
| 204 |
+
{%- set prev_ns.is_tool_first = false %}
|
| 205 |
+
{%- endif %}
|
| 206 |
+
{%- if is_ai_search %}
|
| 207 |
+
{{- tooloutput_begin_ds_token + visible_text(message['content']) + tooloutput_end_ds_token }}
|
| 208 |
+
{%- else %}
|
| 209 |
+
{{- toolresponse_begin_token + '\n' + visible_text(message['content']) + '\n' + toolresponse_end_token + '\n' }}
|
| 210 |
+
{%- endif %}
|
| 211 |
+
{%- endif %}
|
| 212 |
+
{%- if loop.last and message['role'] == 'assistant' %}
|
| 213 |
+
{%- set last_ns.last_is_assistant = true %}
|
| 214 |
+
{%- endif %}
|
| 215 |
+
|
| 216 |
+
{%- endfor %}
|
| 217 |
+
{%- if prev_ns.is_tool and not is_ai_search %}
|
| 218 |
+
{{- toolresponses_end_token }}
|
| 219 |
+
{%- endif %}
|
| 220 |
+
{%- if add_generation_prompt %}
|
| 221 |
+
{%- if not last_ns.last_is_assistant %}
|
| 222 |
+
{%- if is_ai_search %}
|
| 223 |
+
{{- assistant_token }}
|
| 224 |
+
{{- think_begin_token + think_end_token }}
|
| 225 |
+
{%- elif reasoning_effort is defined and reasoning_effort in ['low', 'high'] %}
|
| 226 |
+
{{- assistant_token + think_begin_token }}
|
| 227 |
+
{%- elif reasoning_effort is defined and reasoning_effort == 'no_think' %}
|
| 228 |
+
{{- assistant_token + think_begin_token + think_end_token }}
|
| 229 |
+
{%- else %}
|
| 230 |
+
{{- assistant_token }}
|
| 231 |
+
{%- endif %}
|
| 232 |
+
{%- endif %}
|
| 233 |
+
{%- endif %}
|
config.json
ADDED
|
@@ -0,0 +1,53 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"architectures": [
|
| 3 |
+
"HYV3ForCausalLM"
|
| 4 |
+
],
|
| 5 |
+
"attn_impl": "eager",
|
| 6 |
+
"bos_token_id": 120000,
|
| 7 |
+
"enable_attention_fp32_softmax": false,
|
| 8 |
+
"enable_lm_head_fp32": true,
|
| 9 |
+
"enable_moe_fp32_combine": false,
|
| 10 |
+
"eod_token_id": 120026,
|
| 11 |
+
"eos_token_id": 120025,
|
| 12 |
+
"expert_hidden_dim": 768,
|
| 13 |
+
"moe_intermediate_size": 768,
|
| 14 |
+
"first_k_dense_replace": 1,
|
| 15 |
+
"head_dim": 128,
|
| 16 |
+
"hidden_act": "silu",
|
| 17 |
+
"hidden_size": 2048,
|
| 18 |
+
"initializer_range": 0.006,
|
| 19 |
+
"intermediate_size": 6912,
|
| 20 |
+
"max_position_embeddings": 262144,
|
| 21 |
+
"model_type": "hy_v3",
|
| 22 |
+
"moe_router_enable_expert_bias": true,
|
| 23 |
+
"moe_router_use_sigmoid": true,
|
| 24 |
+
"num_attention_heads": 32,
|
| 25 |
+
"num_experts": 128,
|
| 26 |
+
"num_experts_per_tok": 8,
|
| 27 |
+
"num_hidden_layers": 48,
|
| 28 |
+
"num_key_value_heads": 4,
|
| 29 |
+
"num_shared_experts": 1,
|
| 30 |
+
"output_router_logits": true,
|
| 31 |
+
"pad_token_id": 120002,
|
| 32 |
+
"qk_norm": true,
|
| 33 |
+
"rms_norm_eps": 1e-05,
|
| 34 |
+
"rope_theta": 11158840.0,
|
| 35 |
+
"route_norm": true,
|
| 36 |
+
"router_scaling_factor": 2.826,
|
| 37 |
+
"sep_token_id": 120007,
|
| 38 |
+
"tie_word_embeddings": false,
|
| 39 |
+
"transformers_version": "4.57.1",
|
| 40 |
+
"use_cache": true,
|
| 41 |
+
"use_grouped_mm": false,
|
| 42 |
+
"vocab_size": 120832,
|
| 43 |
+
"quantization": {
|
| 44 |
+
"group_size": 64,
|
| 45 |
+
"bits": 8,
|
| 46 |
+
"mode": "affine"
|
| 47 |
+
},
|
| 48 |
+
"quantization_config": {
|
| 49 |
+
"group_size": 64,
|
| 50 |
+
"bits": 8,
|
| 51 |
+
"mode": "affine"
|
| 52 |
+
}
|
| 53 |
+
}
|
hy_v3.py
ADDED
|
@@ -0,0 +1,405 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Copyright © 2026 Apple Inc.
|
| 2 |
+
#
|
| 3 |
+
# Modification notice:
|
| 4 |
+
# This file was packaged with this quantized MLX model to make the local
|
| 5 |
+
# HYV3 architecture loadable from the model directory. It includes a local
|
| 6 |
+
# compatibility fallback for configs that provide rope_theta without
|
| 7 |
+
# rope_parameters, and uses absolute mlx_lm model imports for model_file
|
| 8 |
+
# loading.
|
| 9 |
+
|
| 10 |
+
from dataclasses import dataclass
|
| 11 |
+
from typing import Any, Dict, Optional
|
| 12 |
+
|
| 13 |
+
import mlx.core as mx
|
| 14 |
+
import mlx.nn as nn
|
| 15 |
+
from mlx.nn.layers.distributed import shard_inplace, shard_linear, sum_gradients
|
| 16 |
+
|
| 17 |
+
from mlx_lm.models.activations import swiglu
|
| 18 |
+
from mlx_lm.models.base import BaseModelArgs, create_attention_mask, scaled_dot_product_attention
|
| 19 |
+
from mlx_lm.models.pipeline import PipelineMixin
|
| 20 |
+
from mlx_lm.models.rope_utils import initialize_rope
|
| 21 |
+
from mlx_lm.models.switch_layers import SwitchGLU
|
| 22 |
+
|
| 23 |
+
|
| 24 |
+
@dataclass
|
| 25 |
+
class ModelArgs(BaseModelArgs):
|
| 26 |
+
model_type: str
|
| 27 |
+
vocab_size: int
|
| 28 |
+
hidden_size: int
|
| 29 |
+
intermediate_size: int
|
| 30 |
+
num_hidden_layers: int
|
| 31 |
+
num_attention_heads: int
|
| 32 |
+
num_key_value_heads: int
|
| 33 |
+
head_dim: int
|
| 34 |
+
num_experts: int
|
| 35 |
+
num_experts_per_tok: int
|
| 36 |
+
num_shared_experts: int
|
| 37 |
+
expert_hidden_dim: int
|
| 38 |
+
first_k_dense_replace: int
|
| 39 |
+
rms_norm_eps: float
|
| 40 |
+
rope_parameters: Dict[str, Any]
|
| 41 |
+
router_scaling_factor: float = 1.0
|
| 42 |
+
qk_norm: bool = True
|
| 43 |
+
route_norm: bool = True
|
| 44 |
+
moe_router_use_sigmoid: bool = True
|
| 45 |
+
moe_router_enable_expert_bias: bool = True
|
| 46 |
+
tie_word_embeddings: bool = False
|
| 47 |
+
num_nextn_predict_layers: int = 0
|
| 48 |
+
max_position_embeddings: int = 262144
|
| 49 |
+
enable_moe_fp32_combine: bool = False
|
| 50 |
+
enable_lm_head_fp32: bool = False
|
| 51 |
+
|
| 52 |
+
@classmethod
|
| 53 |
+
def from_dict(cls, params):
|
| 54 |
+
params = dict(params)
|
| 55 |
+
if "rope_parameters" not in params and "rope_theta" in params:
|
| 56 |
+
params["rope_parameters"] = {
|
| 57 |
+
"rope_theta": params["rope_theta"],
|
| 58 |
+
"rope_type": "default",
|
| 59 |
+
}
|
| 60 |
+
return super().from_dict(params)
|
| 61 |
+
|
| 62 |
+
|
| 63 |
+
class Attention(nn.Module):
|
| 64 |
+
def __init__(self, args: ModelArgs):
|
| 65 |
+
super().__init__()
|
| 66 |
+
|
| 67 |
+
dim = args.hidden_size
|
| 68 |
+
self.n_heads = args.num_attention_heads
|
| 69 |
+
self.n_kv_heads = args.num_key_value_heads
|
| 70 |
+
self.head_dim = args.head_dim
|
| 71 |
+
self.scale = self.head_dim**-0.5
|
| 72 |
+
|
| 73 |
+
self.q_proj = nn.Linear(dim, self.n_heads * self.head_dim, bias=False)
|
| 74 |
+
self.k_proj = nn.Linear(dim, self.n_kv_heads * self.head_dim, bias=False)
|
| 75 |
+
self.v_proj = nn.Linear(dim, self.n_kv_heads * self.head_dim, bias=False)
|
| 76 |
+
self.o_proj = nn.Linear(self.n_heads * self.head_dim, dim, bias=False)
|
| 77 |
+
|
| 78 |
+
self.use_qk_norm = args.qk_norm
|
| 79 |
+
if self.use_qk_norm:
|
| 80 |
+
self.q_norm = nn.RMSNorm(self.head_dim, eps=args.rms_norm_eps)
|
| 81 |
+
self.k_norm = nn.RMSNorm(self.head_dim, eps=args.rms_norm_eps)
|
| 82 |
+
|
| 83 |
+
self.rope = initialize_rope(
|
| 84 |
+
dims=self.head_dim,
|
| 85 |
+
base=args.rope_parameters["rope_theta"],
|
| 86 |
+
traditional=False,
|
| 87 |
+
scaling_config=args.rope_parameters,
|
| 88 |
+
max_position_embeddings=args.max_position_embeddings,
|
| 89 |
+
)
|
| 90 |
+
|
| 91 |
+
def __call__(
|
| 92 |
+
self,
|
| 93 |
+
x: mx.array,
|
| 94 |
+
mask: Optional[mx.array] = None,
|
| 95 |
+
cache: Optional[Any] = None,
|
| 96 |
+
) -> mx.array:
|
| 97 |
+
B, L, _ = x.shape
|
| 98 |
+
|
| 99 |
+
queries = self.q_proj(x).reshape(B, L, self.n_heads, self.head_dim)
|
| 100 |
+
keys = self.k_proj(x).reshape(B, L, self.n_kv_heads, self.head_dim)
|
| 101 |
+
values = self.v_proj(x).reshape(B, L, self.n_kv_heads, self.head_dim)
|
| 102 |
+
|
| 103 |
+
if self.use_qk_norm:
|
| 104 |
+
queries = self.q_norm(queries)
|
| 105 |
+
keys = self.k_norm(keys)
|
| 106 |
+
|
| 107 |
+
queries = queries.transpose(0, 2, 1, 3)
|
| 108 |
+
keys = keys.transpose(0, 2, 1, 3)
|
| 109 |
+
values = values.transpose(0, 2, 1, 3)
|
| 110 |
+
|
| 111 |
+
offset = cache.offset if cache is not None else 0
|
| 112 |
+
queries = self.rope(queries, offset=offset)
|
| 113 |
+
keys = self.rope(keys, offset=offset)
|
| 114 |
+
if cache is not None:
|
| 115 |
+
keys, values = cache.update_and_fetch(keys, values)
|
| 116 |
+
|
| 117 |
+
output = scaled_dot_product_attention(
|
| 118 |
+
queries, keys, values, cache=cache, scale=self.scale, mask=mask
|
| 119 |
+
)
|
| 120 |
+
output = output.transpose(0, 2, 1, 3).reshape(B, L, -1)
|
| 121 |
+
return self.o_proj(output)
|
| 122 |
+
|
| 123 |
+
|
| 124 |
+
class MLP(nn.Module):
|
| 125 |
+
def __init__(self, hidden_size: int, intermediate_size: int):
|
| 126 |
+
super().__init__()
|
| 127 |
+
self.gate_proj = nn.Linear(hidden_size, intermediate_size, bias=False)
|
| 128 |
+
self.up_proj = nn.Linear(hidden_size, intermediate_size, bias=False)
|
| 129 |
+
self.down_proj = nn.Linear(intermediate_size, hidden_size, bias=False)
|
| 130 |
+
|
| 131 |
+
def __call__(self, x):
|
| 132 |
+
return self.down_proj(swiglu(self.gate_proj(x), self.up_proj(x)))
|
| 133 |
+
|
| 134 |
+
|
| 135 |
+
@mx.compile
|
| 136 |
+
def expert_select(
|
| 137 |
+
gates,
|
| 138 |
+
expert_bias,
|
| 139 |
+
top_k,
|
| 140 |
+
routed_scaling_factor,
|
| 141 |
+
norm_topk_prob,
|
| 142 |
+
):
|
| 143 |
+
scores = mx.sigmoid(gates.astype(mx.float32))
|
| 144 |
+
orig_scores = scores
|
| 145 |
+
scores = scores + expert_bias
|
| 146 |
+
|
| 147 |
+
inds = mx.argpartition(scores, kth=-top_k, axis=-1)[..., -top_k:]
|
| 148 |
+
scores = mx.take_along_axis(orig_scores, inds, axis=-1)
|
| 149 |
+
if top_k > 1 and norm_topk_prob:
|
| 150 |
+
scores = scores / (scores.sum(axis=-1, keepdims=True) + 1e-20)
|
| 151 |
+
scores = scores * routed_scaling_factor
|
| 152 |
+
|
| 153 |
+
return inds, scores
|
| 154 |
+
|
| 155 |
+
|
| 156 |
+
class MoEGate(nn.Module):
|
| 157 |
+
def __init__(self, args: ModelArgs):
|
| 158 |
+
super().__init__()
|
| 159 |
+
self.top_k = args.num_experts_per_tok
|
| 160 |
+
self.norm_topk_prob = args.route_norm
|
| 161 |
+
self.routed_scaling_factor = args.router_scaling_factor
|
| 162 |
+
self.gate = nn.Linear(args.hidden_size, args.num_experts, bias=False)
|
| 163 |
+
self.expert_bias = mx.zeros((args.num_experts,))
|
| 164 |
+
|
| 165 |
+
def __call__(self, x):
|
| 166 |
+
return expert_select(
|
| 167 |
+
self.gate(x),
|
| 168 |
+
self.expert_bias,
|
| 169 |
+
self.top_k,
|
| 170 |
+
self.routed_scaling_factor,
|
| 171 |
+
self.norm_topk_prob,
|
| 172 |
+
)
|
| 173 |
+
|
| 174 |
+
|
| 175 |
+
class MoE(nn.Module):
|
| 176 |
+
def __init__(self, args: ModelArgs):
|
| 177 |
+
super().__init__()
|
| 178 |
+
self.num_experts_per_tok = args.num_experts_per_tok
|
| 179 |
+
self.switch_mlp = SwitchGLU(
|
| 180 |
+
args.hidden_size,
|
| 181 |
+
args.expert_hidden_dim,
|
| 182 |
+
args.num_experts,
|
| 183 |
+
)
|
| 184 |
+
self.router = MoEGate(args)
|
| 185 |
+
if args.num_shared_experts > 0:
|
| 186 |
+
self.shared_mlp = MLP(
|
| 187 |
+
args.hidden_size,
|
| 188 |
+
args.expert_hidden_dim * args.num_shared_experts,
|
| 189 |
+
)
|
| 190 |
+
else:
|
| 191 |
+
self.shared_mlp = None
|
| 192 |
+
|
| 193 |
+
self.fp32_combine = args.enable_moe_fp32_combine
|
| 194 |
+
self.sharding_group = None
|
| 195 |
+
|
| 196 |
+
def __call__(self, x):
|
| 197 |
+
if self.sharding_group is not None:
|
| 198 |
+
x = sum_gradients(self.sharding_group)(x)
|
| 199 |
+
|
| 200 |
+
inds, scores = self.router(x)
|
| 201 |
+
if not self.fp32_combine:
|
| 202 |
+
scores = scores.astype(x.dtype)
|
| 203 |
+
y = self.switch_mlp(x, inds)
|
| 204 |
+
y = (y * scores[..., None]).sum(axis=-2)
|
| 205 |
+
if self.shared_mlp is not None:
|
| 206 |
+
y = y + self.shared_mlp(x)
|
| 207 |
+
|
| 208 |
+
if self.sharding_group is not None:
|
| 209 |
+
y = mx.distributed.all_sum(y, group=self.sharding_group)
|
| 210 |
+
|
| 211 |
+
return y.astype(x.dtype)
|
| 212 |
+
|
| 213 |
+
|
| 214 |
+
class DecoderLayer(nn.Module):
|
| 215 |
+
def __init__(self, args: ModelArgs, layer_idx: int):
|
| 216 |
+
super().__init__()
|
| 217 |
+
self.self_attn = Attention(args)
|
| 218 |
+
if layer_idx < args.first_k_dense_replace:
|
| 219 |
+
self.mlp = MLP(args.hidden_size, args.intermediate_size)
|
| 220 |
+
else:
|
| 221 |
+
self.mlp = MoE(args)
|
| 222 |
+
self.input_layernorm = nn.RMSNorm(args.hidden_size, eps=args.rms_norm_eps)
|
| 223 |
+
self.post_attention_layernorm = nn.RMSNorm(
|
| 224 |
+
args.hidden_size, eps=args.rms_norm_eps
|
| 225 |
+
)
|
| 226 |
+
|
| 227 |
+
def __call__(
|
| 228 |
+
self,
|
| 229 |
+
x: mx.array,
|
| 230 |
+
mask: Optional[mx.array] = None,
|
| 231 |
+
cache: Optional[Any] = None,
|
| 232 |
+
) -> mx.array:
|
| 233 |
+
r = self.self_attn(self.input_layernorm(x), mask, cache)
|
| 234 |
+
h = x + r
|
| 235 |
+
r = self.mlp(self.post_attention_layernorm(h))
|
| 236 |
+
return h + r
|
| 237 |
+
|
| 238 |
+
|
| 239 |
+
class HYV3Model(PipelineMixin, nn.Module):
|
| 240 |
+
def __init__(self, args: ModelArgs):
|
| 241 |
+
super().__init__()
|
| 242 |
+
self.vocab_size = args.vocab_size
|
| 243 |
+
self.embed_tokens = nn.Embedding(args.vocab_size, args.hidden_size)
|
| 244 |
+
self.layers = [DecoderLayer(args, idx) for idx in range(args.num_hidden_layers)]
|
| 245 |
+
self.norm = nn.RMSNorm(args.hidden_size, eps=args.rms_norm_eps)
|
| 246 |
+
|
| 247 |
+
def __call__(
|
| 248 |
+
self,
|
| 249 |
+
x: mx.array,
|
| 250 |
+
cache: Optional[Any] = None,
|
| 251 |
+
) -> mx.array:
|
| 252 |
+
h = self.embed_tokens(x)
|
| 253 |
+
|
| 254 |
+
pipeline_rank = self.pipeline_rank
|
| 255 |
+
pipeline_size = self.pipeline_size
|
| 256 |
+
|
| 257 |
+
if cache is None:
|
| 258 |
+
cache = [None] * len(self.pipeline_layers)
|
| 259 |
+
mask = create_attention_mask(h, cache[0])
|
| 260 |
+
|
| 261 |
+
if pipeline_rank < pipeline_size - 1:
|
| 262 |
+
h = mx.distributed.recv_like(h, (pipeline_rank + 1))
|
| 263 |
+
|
| 264 |
+
for layer, c in zip(self.pipeline_layers, cache):
|
| 265 |
+
h = layer(h, mask, cache=c)
|
| 266 |
+
|
| 267 |
+
if pipeline_rank != 0:
|
| 268 |
+
h = mx.distributed.send(h, (pipeline_rank - 1) % pipeline_size)
|
| 269 |
+
if cache[-1] is not None:
|
| 270 |
+
cache[-1].keys = mx.depends(cache[-1].keys, h)
|
| 271 |
+
|
| 272 |
+
if pipeline_size > 1:
|
| 273 |
+
h = mx.distributed.all_gather(h)[: h.shape[0]]
|
| 274 |
+
|
| 275 |
+
return self.norm(h)
|
| 276 |
+
|
| 277 |
+
|
| 278 |
+
class Model(nn.Module):
|
| 279 |
+
def __init__(self, args: ModelArgs):
|
| 280 |
+
super().__init__()
|
| 281 |
+
self.args = args
|
| 282 |
+
self.model_type = args.model_type
|
| 283 |
+
self.model = HYV3Model(args)
|
| 284 |
+
if not args.tie_word_embeddings:
|
| 285 |
+
self.lm_head = nn.Linear(args.hidden_size, args.vocab_size, bias=False)
|
| 286 |
+
|
| 287 |
+
def __call__(
|
| 288 |
+
self,
|
| 289 |
+
inputs: mx.array,
|
| 290 |
+
cache: Optional[Any] = None,
|
| 291 |
+
):
|
| 292 |
+
out = self.model(inputs, cache)
|
| 293 |
+
if self.args.enable_lm_head_fp32:
|
| 294 |
+
out = out.astype(mx.float32)
|
| 295 |
+
if self.args.tie_word_embeddings:
|
| 296 |
+
return self.model.embed_tokens.as_linear(out)
|
| 297 |
+
return self.lm_head(out)
|
| 298 |
+
|
| 299 |
+
def sanitize(self, weights):
|
| 300 |
+
n_layers = self.args.num_hidden_layers
|
| 301 |
+
n_mtp = self.args.num_nextn_predict_layers
|
| 302 |
+
|
| 303 |
+
if n_mtp > 0:
|
| 304 |
+
mtp_prefixes = tuple(f"model.layers.{n_layers + i}." for i in range(n_mtp))
|
| 305 |
+
weights = {
|
| 306 |
+
k: v for k, v in weights.items() if not k.startswith(mtp_prefixes)
|
| 307 |
+
}
|
| 308 |
+
|
| 309 |
+
for l in range(n_layers):
|
| 310 |
+
prefix = f"model.layers.{l}"
|
| 311 |
+
|
| 312 |
+
bias_key = f"{prefix}.mlp.expert_bias"
|
| 313 |
+
if bias_key in weights:
|
| 314 |
+
weights[f"{prefix}.mlp.router.expert_bias"] = weights.pop(bias_key)
|
| 315 |
+
|
| 316 |
+
for m in ("gate_proj", "down_proj", "up_proj"):
|
| 317 |
+
for k in ("weight", "scales", "biases"):
|
| 318 |
+
if f"{prefix}.mlp.experts.0.{m}.{k}" in weights:
|
| 319 |
+
to_join = [
|
| 320 |
+
weights.pop(f"{prefix}.mlp.experts.{e}.{m}.{k}")
|
| 321 |
+
for e in range(self.args.num_experts)
|
| 322 |
+
]
|
| 323 |
+
weights[f"{prefix}.mlp.switch_mlp.{m}.{k}"] = mx.stack(to_join)
|
| 324 |
+
|
| 325 |
+
if self.args.tie_word_embeddings:
|
| 326 |
+
weights.pop("lm_head.weight", None)
|
| 327 |
+
|
| 328 |
+
return weights
|
| 329 |
+
|
| 330 |
+
def shard(self, group: Optional[mx.distributed.Group] = None):
|
| 331 |
+
group = group or mx.distributed.init()
|
| 332 |
+
N = group.size()
|
| 333 |
+
for layer in self.model.layers:
|
| 334 |
+
layer.self_attn.q_proj = shard_linear(
|
| 335 |
+
layer.self_attn.q_proj, "all-to-sharded", group=group
|
| 336 |
+
)
|
| 337 |
+
layer.self_attn.k_proj = shard_linear(
|
| 338 |
+
layer.self_attn.k_proj, "all-to-sharded", group=group
|
| 339 |
+
)
|
| 340 |
+
layer.self_attn.v_proj = shard_linear(
|
| 341 |
+
layer.self_attn.v_proj, "all-to-sharded", group=group
|
| 342 |
+
)
|
| 343 |
+
layer.self_attn.o_proj = shard_linear(
|
| 344 |
+
layer.self_attn.o_proj, "sharded-to-all", group=group
|
| 345 |
+
)
|
| 346 |
+
layer.self_attn.n_heads //= N
|
| 347 |
+
layer.self_attn.n_kv_heads = max(1, layer.self_attn.n_kv_heads // N)
|
| 348 |
+
|
| 349 |
+
if isinstance(layer.mlp, MLP):
|
| 350 |
+
layer.mlp.gate_proj = shard_linear(
|
| 351 |
+
layer.mlp.gate_proj, "all-to-sharded", group=group
|
| 352 |
+
)
|
| 353 |
+
layer.mlp.down_proj = shard_linear(
|
| 354 |
+
layer.mlp.down_proj, "sharded-to-all", group=group
|
| 355 |
+
)
|
| 356 |
+
layer.mlp.up_proj = shard_linear(
|
| 357 |
+
layer.mlp.up_proj, "all-to-sharded", group=group
|
| 358 |
+
)
|
| 359 |
+
else:
|
| 360 |
+
layer.mlp.sharding_group = group
|
| 361 |
+
if layer.mlp.shared_mlp is not None:
|
| 362 |
+
shard_inplace(
|
| 363 |
+
layer.mlp.shared_mlp.gate_proj,
|
| 364 |
+
"all-to-sharded",
|
| 365 |
+
group=group,
|
| 366 |
+
)
|
| 367 |
+
shard_inplace(
|
| 368 |
+
layer.mlp.shared_mlp.down_proj,
|
| 369 |
+
"sharded-to-all",
|
| 370 |
+
group=group,
|
| 371 |
+
)
|
| 372 |
+
shard_inplace(
|
| 373 |
+
layer.mlp.shared_mlp.up_proj,
|
| 374 |
+
"all-to-sharded",
|
| 375 |
+
group=group,
|
| 376 |
+
)
|
| 377 |
+
shard_inplace(
|
| 378 |
+
layer.mlp.switch_mlp.gate_proj, "all-to-sharded", group=group
|
| 379 |
+
)
|
| 380 |
+
shard_inplace(
|
| 381 |
+
layer.mlp.switch_mlp.down_proj, "sharded-to-all", group=group
|
| 382 |
+
)
|
| 383 |
+
shard_inplace(
|
| 384 |
+
layer.mlp.switch_mlp.up_proj, "all-to-sharded", group=group
|
| 385 |
+
)
|
| 386 |
+
|
| 387 |
+
@property
|
| 388 |
+
def layers(self):
|
| 389 |
+
return self.model.pipeline_layers
|
| 390 |
+
|
| 391 |
+
@property
|
| 392 |
+
def quant_predicate(self):
|
| 393 |
+
def predicate(path, _):
|
| 394 |
+
if path.endswith("mlp.router.gate"):
|
| 395 |
+
return {"group_size": 64, "bits": 8}
|
| 396 |
+
return True
|
| 397 |
+
|
| 398 |
+
return predicate
|
| 399 |
+
|
| 400 |
+
@property
|
| 401 |
+
def cast_predicate(self):
|
| 402 |
+
def predicate(k):
|
| 403 |
+
return "expert_bias" not in k
|
| 404 |
+
|
| 405 |
+
return predicate
|
model-00001-of-00007.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:cf1f8fb45137e57e0ed7681c99b20f1df74f2ba70aec50c46839c1eb95bdfbb2
|
| 3 |
+
size 5003070606
|
model-00002-of-00007.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:7e2d9faeace470a200448bc13e7e3a24afc31ba562272796f42778cb645d2502
|
| 3 |
+
size 5133840092
|
model-00003-of-00007.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:6fd46930e8541e932084857028579e80ed78587d3fc5435d82392f5dca78028d
|
| 3 |
+
size 5133840154
|
model-00004-of-00007.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:a6063a0b459d5b499ff395477a0371fba16185957d4338f9dece5035534a1fdf
|
| 3 |
+
size 5133840146
|
model-00005-of-00007.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:8e98683e6465fde3d8e7c220c91f2d67fe54eac1afe7905bcfc4f5039a6195c7
|
| 3 |
+
size 5133840160
|
model-00006-of-00007.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:9461934210450310c42e5e88fd11423c7ad1806ed23e02418f1e62b43df2fb2d
|
| 3 |
+
size 5133840168
|
model-00007-of-00007.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:e0c7002587af4d7e4ccc316906d9aeff3e881c9c659c29b6bfe874b99d3900f8
|
| 3 |
+
size 1283460032
|
model.safetensors.index.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
special_tokens_map.json
ADDED
|
@@ -0,0 +1,23 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"bos_token": {
|
| 3 |
+
"content": "<|hy_begin▁of▁sentence|>",
|
| 4 |
+
"lstrip": false,
|
| 5 |
+
"normalized": false,
|
| 6 |
+
"rstrip": false,
|
| 7 |
+
"single_word": false
|
| 8 |
+
},
|
| 9 |
+
"eos_token": {
|
| 10 |
+
"content": "<eos:6124c78e>",
|
| 11 |
+
"lstrip": false,
|
| 12 |
+
"normalized": false,
|
| 13 |
+
"rstrip": false,
|
| 14 |
+
"single_word": false
|
| 15 |
+
},
|
| 16 |
+
"pad_token": {
|
| 17 |
+
"content": "<|hy_▁pad▁|>",
|
| 18 |
+
"lstrip": false,
|
| 19 |
+
"normalized": false,
|
| 20 |
+
"rstrip": false,
|
| 21 |
+
"single_word": false
|
| 22 |
+
}
|
| 23 |
+
}
|
tokenizer.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
tokenizer_config.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|