xrx commited on
Commit
c3314a5
·
1 Parent(s): ceab9d5

Initial commit

Browse files
._README.md ADDED
Binary file (4.1 kB). View file
 
._config.json ADDED
Binary file (4.1 kB). View file
 
._demo.py ADDED
Binary file (4.1 kB). View file
 
._modeling_mixsense_llama.py ADDED
Binary file (4.1 kB). View file
 
LICENSE ADDED
@@ -0,0 +1,84 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ META LLAMA 3 COMMUNITY LICENSE AGREEMENT
2
+
3
+ Meta Llama 3 Version Release Date: April 18, 2024
4
+ “Agreement” means the terms and conditions for use, reproduction, distribution and modification of the Llama Materials set forth herein.
5
+
6
+ “Documentation” means the specifications, manuals and documentation accompanying Meta Llama 3 distributed by Meta at https://llama.meta.com/get-started/.
7
+
8
+ “Licensee” or “you” means you, or your employer or any other person or entity (if you are entering into this Agreement on such person or entity’s behalf), of the age required under applicable laws, rules or regulations to provide legal consent and that has legal authority to bind your employer or such other person or entity if you are entering in this Agreement on their behalf.
9
+
10
+ “Meta Llama 3” means the foundational large language models and software and algorithms, including machine-learning model code, trained model weights, inference-enabling code, training-enabling code, fine-tuning enabling code and other elements of the foregoing distributed by Meta at https://llama.meta.com/llama-downloads.
11
+
12
+ “Llama Materials” means, collectively, Meta’s proprietary Meta Llama 3 and Documentation (and any portion thereof) made available under this Agreement.
13
+
14
+ “Meta” or “we” means Meta Platforms Ireland Limited (if you are located in or, if you are an entity, your principal place of business is in the EEA or Switzerland) and Meta Platforms, Inc. (if you are located outside of the EEA or Switzerland).
15
+
16
+ By clicking “I Accept” below or by using or distributing any portion or element of the Llama Materials, you agree to be bound by this Agreement.
17
+
18
+ 1. License Rights and Redistribution.
19
+
20
+ a. Grant of Rights. You are granted a non-exclusive, worldwide, non-transferable and royalty-free limited license under Meta’s intellectual property or other rights owned by Meta embodied in the Llama Materials to use, reproduce, distribute, copy, create derivative works of, and make modifications to the Llama Materials.
21
+ b. Redistribution and Use.
22
+ i. If you distribute or make available the Llama Materials (or any derivative works thereof), or a product or service that uses any of them, including another AI model, you shall (A) provide a copy of this Agreement with any such Llama Materials; and (B) prominently display “Built with Meta Llama 3” on a related website, user interface, blogpost, about page, or product documentation. If you use the Llama Materials to create, train, fine tune, or otherwise improve an AI model, which is distributed or made available, you shall also include “Llama 3” at the beginning of any such AI model name.
23
+ ii. If you receive Llama Materials, or any derivative works thereof, from a Licensee as part of an integrated end user product, then Section 2 of this Agreement will not apply to you.
24
+ iii. You must retain in all copies of the Llama Materials that you distribute the following attribution notice within a “Notice” text file distributed as a part of such copies: “Meta Llama 3 is licensed under the Meta Llama 3 Community License, Copyright © Meta Platforms, Inc. All Rights Reserved.”
25
+ iv. Your use of the Llama Materials must comply with applicable laws and regulations (including trade compliance laws and regulations) and adhere to the Acceptable Use Policy for the Llama Materials (available at https://llama.meta.com/llama3/use-policy), which is hereby incorporated by reference into this Agreement.
26
+ v. You will not use the Llama Materials or any output or results of the Llama Materials to improve any other large language model (excluding Meta Llama 3 or derivative works thereof).
27
+
28
+ 2. Additional Commercial Terms. If, on the Meta Llama 3 version release date, the monthly active users of the products or services made available by or for Licensee, or Licensee’s affiliates, is greater than 700 million monthly active users in the preceding calendar month, you must request a license from Meta, which Meta may grant to you in its sole discretion, and you are not authorized to exercise any of the rights under this Agreement unless or until Meta otherwise expressly grants you such rights.
29
+
30
+ 3. Disclaimer of Warranty. UNLESS REQUIRED BY APPLICABLE LAW, THE LLAMA MATERIALS AND ANY OUTPUT AND RESULTS THEREFROM ARE PROVIDED ON AN “AS IS” BASIS, WITHOUT WARRANTIES OF ANY KIND, AND META DISCLAIMS ALL WARRANTIES OF ANY KIND, BOTH EXPRESS AND IMPLIED, INCLUDING, WITHOUT LIMITATION, ANY WARRANTIES OF TITLE, NON-INFRINGEMENT, MERCHANTABILITY, OR FITNESS FOR A PARTICULAR PURPOSE. YOU ARE SOLELY RESPONSIBLE FOR DETERMINING THE APPROPRIATENESS OF USING OR REDISTRIBUTING THE LLAMA MATERIALS AND ASSUME ANY RISKS ASSOCIATED WITH YOUR USE OF THE LLAMA MATERIALS AND ANY OUTPUT AND RESULTS.
31
+
32
+ 4. Limitation of Liability. IN NO EVENT WILL META OR ITS AFFILIATES BE LIABLE UNDER ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, TORT, NEGLIGENCE, PRODUCTS LIABILITY, OR OTHERWISE, ARISING OUT OF THIS AGREEMENT, FOR ANY LOST PROFITS OR ANY INDIRECT, SPECIAL, CONSEQUENTIAL, INCIDENTAL, EXEMPLARY OR PUNITIVE DAMAGES, EVEN IF META OR ITS AFFILIATES HAVE BEEN ADVISED OF THE POSSIBILITY OF ANY OF THE FOREGOING.
33
+
34
+ 5. Intellectual Property.
35
+ a. No trademark licenses are granted under this Agreement, and in connection with the Llama Materials, neither Meta nor Licensee may use any name or mark owned by or associated with the other or any of its affiliates, except as required for reasonable and customary use in describing and redistributing the Llama Materials or as set forth in this Section 5(a). Meta hereby grants you a license to use “Llama 3” (the “Mark”) solely as required to comply with the last sentence of Section 1.b.i. You will comply with Meta’s brand guidelines (currently accessible at https://about.meta.com/brand/resources/meta/company-brand/ ). All goodwill arising out of your use of the Mark will inure to the benefit of Meta.
36
+ b. Subject to Meta’s ownership of Llama Materials and derivatives made by or for Meta, with respect to any derivative works and modifications of the Llama Materials that are made by you, as between you and Meta, you are and will be the owner of such derivative works and modifications.
37
+ c. If you institute litigation or other proceedings against Meta or any entity (including a cross-claim or counterclaim in a lawsuit) alleging that the Llama Materials or Meta Llama 3 outputs or results, or any portion of any of the foregoing, constitutes infringement of intellectual property or other rights owned or licensable by you, then any licenses granted to you under this Agreement shall terminate as of the date such litigation or claim is filed or instituted. You will indemnify and hold harmless Meta from and against any claim by any third party arising out of or related to your use or distribution of the Llama Materials.
38
+
39
+ 6. Term and Termination. The term of this Agreement will commence upon your acceptance of this Agreement or access to the Llama Materials and will continue in full force and effect until terminated in accordance with the terms and conditions herein. Meta may terminate this Agreement if you are in breach of any term or condition of this Agreement. Upon termination of this Agreement, you shall delete and cease use of the Llama Materials. Sections 3, 4 and 7 shall survive the termination of this Agreement.
40
+
41
+ 7. Governing Law and Jurisdiction. This Agreement will be governed and construed under the laws of the State of California without regard to choice of law principles, and the UN Convention on Contracts for the International Sale of Goods does not apply to this Agreement. The courts of California shall have exclusive jurisdiction of any dispute arising out of this Agreement.
42
+
43
+
44
+ Meta Llama 3 Acceptable Use Policy
45
+ Meta is committed to promoting safe and fair use of its tools and features, including Meta Llama 3. If you access or use Meta Llama 3, you agree to this Acceptable Use Policy (“Policy”). The most recent copy of this policy can be found at https://llama.meta.com/llama3/use-policy
46
+ Prohibited Uses
47
+ We want everyone to use Meta Llama 3 safely and responsibly. You agree you will not use, or allow others to use, Meta Llama 3 to:
48
+ 1. Violate the law or others’ rights, including to:
49
+ a. Engage in, promote, generate, contribute to, encourage, plan, incite, or further illegal or unlawful activity or content, such as:
50
+ i. Violence or terrorism
51
+ ii. Exploitation or harm to children, including the solicitation, creation, acquisition, or dissemination of child exploitative content or failure to report Child Sexual Abuse Material
52
+ iii. Human trafficking, exploitation, and sexual violence
53
+ iv. The illegal distribution of information or materials to minors, including obscene materials, or failure to employ legally required age-gating in connection with such information or materials.
54
+ v. Sexual solicitation
55
+ vi. Any other criminal activity
56
+ b. Engage in, promote, incite, or facilitate the harassment, abuse, threatening, or bullying of individuals or groups of individuals
57
+ c. Engage in, promote, incite, or facilitate discrimination or other unlawful or harmful conduct in the provision of employment, employment benefits, credit, housing, other economic benefits, or other essential goods and services
58
+ d. Engage in the unauthorized or unlicensed practice of any profession including, but not limited to, financial, legal, medical/health, or related professional practices
59
+ e. Collect, process, disclose, generate, or infer health, demographic, or other sensitive personal or private information about individuals without rights and consents required by applicable laws
60
+ f. Engage in or facilitate any action or generate any content that infringes, misappropriates, or otherwise violates any third-party rights, including the outputs or results of any products or services using the Llama Materials
61
+ g. Create, generate, or facilitate the creation of malicious code, malware, computer viruses or do anything else that could disable, overburden, interfere with or impair the proper working, integrity, operation or appearance of a website or computer system
62
+
63
+ 2. Engage in, promote, incite, facilitate, or assist in the planning or development of activities that present a risk of death or bodily harm to individuals, including use of Meta Llama 3 related to the following:
64
+ a. Military, warfare, nuclear industries or applications, espionage, use for materials or activities that are subject to the International Traffic Arms Regulations (ITAR) maintained by the United States Department of State
65
+ b. Guns and illegal weapons (including weapon development)
66
+ c. Illegal drugs and regulated/controlled substances
67
+ d. Operation of critical infrastructure, transportation technologies, or heavy machinery
68
+ e. Self-harm or harm to others, including suicide, cutting, and eating disorders
69
+ f. Any content intended to incite or promote violence, abuse, or any infliction of bodily harm to an individual
70
+
71
+ 3. Intentionally deceive or mislead others, including use of Meta Llama 3 related to the following:
72
+ a. Generating, promoting, or furthering fraud or the creation or promotion of disinformation
73
+ b. Generating, promoting, or furthering defamatory content, including the creation of defamatory statements, images, or other content
74
+ c. Generating, promoting, or further distributing spam
75
+ d. Impersonating another individual without consent, authorization, or legal right
76
+ e. Representing that the use of Meta Llama 3 or outputs are human-generated
77
+ f. Generating or facilitating false online engagement, including fake reviews and other means of fake online engagement
78
+ g. Fail to appropriately disclose to end users any known dangers of your AI system
79
+
80
+ Please report any violation of this Policy, software “bug,” or other problems that could lead to a violation of this Policy through one of the following means:
81
+ * Reporting issues with the model: https://github.com/meta-llama/llama3
82
+ * Reporting risky content generated by the model: developers.facebook.com/llama_output_feedback
83
+ * Reporting bugs and security concerns: facebook.com/whitehat/info
84
+ * Reporting violations of the Acceptable Use Policy or unlicensed uses of Meta Llama 3: LlamaUseReport@meta.com
README.md CHANGED
@@ -1,3 +1,83 @@
1
- ---
2
- license: llama3
3
- ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Introduction
2
+
3
+ MixSense is a series of models based on the widely adopted vision encoder-projector-LLM architecture. In this resource, we release Llama-3-MixSenseV1.1 checkpoint. Compared to [version 1.0](https://huggingface.co/Zero-Vision/Llama-3-MixSense), we changed the vision encoder from [SigLIP 400M](https://huggingface.co/google/siglip-so400m-patch14-384) to [Florence-2-large-ft's vision encoder DaViT](https://huggingface.co/microsoft/Florence-2-large-ft) and add more VQA data in finetune stage.
4
+
5
+ We have developed an innovative data processing method that complements the training process, reducing training costs while improving training effectiveness.,The models are trained on our restructured dataset. Details of the data organization and related research papers will be available soon.
6
+
7
+ # QuickStart
8
+
9
+ ## Requirements
10
+
11
+ ```
12
+ conda create -n mixsense python==3.10 -y
13
+ conda activate mixsense
14
+ pip install torch transformers==4.37.2 accelerate pillow
15
+ ```
16
+
17
+ ## Usage
18
+
19
+ Llama-3-Mixsense/demo.py
20
+
21
+ ```python
22
+ import torch
23
+ import transformers
24
+ from transformers import AutoModelForCausalLM, AutoTokenizer
25
+ from PIL import Image
26
+ import warnings
27
+ import os
28
+
29
+
30
+ # disable some warnings
31
+ transformers.logging.set_verbosity_error()
32
+ transformers.logging.disable_progress_bar()
33
+ warnings.filterwarnings("ignore")
34
+
35
+ # set device
36
+ device = "cuda" # or cpu, or npu (ASCEND 910B support)
37
+
38
+ # create model
39
+ model = AutoModelForCausalLM.from_pretrained(
40
+ "Zero-Vision/Llama-3-MixSenseV1_1",
41
+ torch_dtype=torch.float16, # float32 for cpu
42
+ device_map="auto",
43
+ trust_remote_code=True,
44
+ )
45
+ tokenizer = AutoTokenizer.from_pretrained(
46
+ "Zero-Vision/Llama-3-MixSenseV1_1",
47
+ trust_remote_code=True,
48
+ )
49
+
50
+ qs = "describe the image detailly."
51
+ input_ids = model.text_process(qs, tokenizer).to(device)
52
+
53
+ image = Image.open("example.jpg")
54
+ image_tensor = model.image_process([image]).to(dtype=model.dtype, device=device)
55
+
56
+ # generate
57
+ with torch.inference_mode():
58
+ output_ids = model.generate(
59
+ input_ids,
60
+ images=image_tensor,
61
+ max_new_tokens=2048,
62
+ use_cache=True,
63
+ eos_token_id=[
64
+ tokenizer.eos_token_id,
65
+ tokenizer.convert_tokens_to_ids(["<|eot_id|>"])[0],
66
+ ],
67
+ )
68
+
69
+ print(tokenizer.batch_decode(output_ids, skip_special_tokens=True)[0].strip())
70
+ ```
71
+
72
+ ## Eval
73
+
74
+ We offer Llama-3-Mixsense/llama3mixsense.py for [VLMEvalKit](https://github.com/open-compass/VLMEvalKit).
75
+
76
+ # License
77
+
78
+ This project utilizes certain datasets and checkpoints that are subject to their respective original licenses. Users must comply with all terms and conditions of these original licenses.including but not limited to Llama3 and SigLIP. Meta Llama 3 is licensed under the [Meta Llama 3 Community License](https://llama.meta.com/llama3/license/), Copyright © Meta Platforms, Inc. All Rights Reserved. And [MIT LICENSE](https://www.apache.org/licenses/LICENSE-2.0) for Florence2 model. The project itself is licensed under the [Apache LICENSE 2.0](https://www.apache.org/licenses/LICENSE-2.0) .
79
+
80
+ # Acknowledgement
81
+
82
+ Our code is largely borrowed from [LLaVA](https://github.com/haotian-liu/LLaVA)
83
+ We bulid this demo according to [bunny](https://huggingface.co/BAAI/Bunny-Llama-3-8B-V)
config.json ADDED
@@ -0,0 +1,49 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "_name_or_path": "Zero-Vision/Llama-3-MixSenseV1_1",
3
+ "architectures": [
4
+ "MixsenseLlamaForCausalLM"
5
+ ],
6
+ "auto_map": {
7
+ "AutoConfig": "modeling_mixsense_llama.MixsenseConfig",
8
+ "AutoModelForCausalLM": "modeling_mixsense_llama.MixsenseLlamaForCausalLM"
9
+ },
10
+ "attention_bias": false,
11
+ "attention_dropout": 0.0,
12
+ "bos_token_id": 128000,
13
+ "eos_token_id": 128001,
14
+ "freeze_mm_mlp_adapter": false,
15
+ "hidden_act": "silu",
16
+ "hidden_size": 4096,
17
+ "image_aspect_ratio": "pad",
18
+ "initializer_range": 0.02,
19
+ "intermediate_size": 14336,
20
+ "max_position_embeddings": 8192,
21
+ "mm_hidden_size": 2048,
22
+ "mm_patch_merge_type": "flat",
23
+ "mm_projector_lr": null,
24
+ "mm_projector_type": "mlp2x_gelu",
25
+ "mm_use_im_patch_token": false,
26
+ "mm_use_im_start_end": false,
27
+ "mm_vision_select_feature": "patch",
28
+ "mm_vision_select_layer": -2,
29
+ "mm_vision_tower": "microsoft/Florence-2-large-ft",
30
+ "model_type": "mixsense_llama",
31
+ "num_attention_heads": 32,
32
+ "num_hidden_layers": 32,
33
+ "num_key_value_heads": 8,
34
+ "pretraining_tp": 1,
35
+ "rms_norm_eps": 1e-05,
36
+ "rope_scaling": null,
37
+ "rope_theta": 500000.0,
38
+ "size": 768,
39
+ "tie_word_embeddings": false,
40
+ "tokenizer_model_max_length": 3072,
41
+ "tokenizer_padding_side": "right",
42
+ "torch_dtype": "bfloat16",
43
+ "transformers_version": "4.37.2",
44
+ "tune_mm_mlp_adapter": false,
45
+ "unfreeze_mm_vision_tower": true,
46
+ "use_cache": true,
47
+ "use_mm_proj": true,
48
+ "vocab_size": 128257
49
+ }
demo.py ADDED
@@ -0,0 +1,46 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import torch
2
+ import transformers
3
+ from transformers import AutoModelForCausalLM, AutoTokenizer
4
+ from PIL import Image
5
+ import warnings
6
+ import os
7
+
8
+ os.environ["HF_ENDPOINT"] = "https://hf-mirror.com"
9
+ # disable some warnings
10
+ transformers.logging.set_verbosity_error()
11
+ transformers.logging.disable_progress_bar()
12
+ warnings.filterwarnings("ignore")
13
+
14
+ # set device
15
+ device = "cuda" # or cpu
16
+
17
+ # create model
18
+ model = AutoModelForCausalLM.from_pretrained(
19
+ "Zero-Vision/Llama-3-MixSenseV1_1",
20
+ torch_dtype=torch.float16,
21
+ device_map="auto",
22
+ trust_remote_code=True,
23
+ )
24
+ tokenizer = AutoTokenizer.from_pretrained(
25
+ "Zero-Vision/Llama-3-MixSenseV1_1",
26
+ trust_remote_code=True,
27
+ )
28
+ qs = "describe the image detailly."
29
+ input_ids = model.text_process(qs, tokenizer).to(device)
30
+
31
+ image = Image.open("example.jpg")
32
+ image_tensor = model.image_process([image]).to(dtype=model.dtype, device=device)
33
+
34
+ # generate
35
+ with torch.inference_mode():
36
+ output_ids = model.generate(
37
+ input_ids,
38
+ images=image_tensor,
39
+ max_new_tokens=2048,
40
+ use_cache=True,
41
+ eos_token_id=[
42
+ tokenizer.eos_token_id,
43
+ tokenizer.convert_tokens_to_ids(["<|eot_id|>"])[0],
44
+ ],
45
+ )
46
+ print(tokenizer.batch_decode(output_ids, skip_special_tokens=True)[0].strip())
example.jpg ADDED
generation_config.json ADDED
@@ -0,0 +1,6 @@
 
 
 
 
 
 
 
1
+ {
2
+ "_from_model_config": true,
3
+ "bos_token_id": 128000,
4
+ "eos_token_id": 128001,
5
+ "transformers_version": "4.37.2"
6
+ }
llama3mixsense.py ADDED
@@ -0,0 +1,50 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ '''
2
+ This file if for VLMEvalKit.
3
+ '''
4
+ import torch
5
+ import transformers
6
+ from transformers import AutoModelForCausalLM, AutoTokenizer
7
+ from PIL import Image
8
+ import warnings
9
+
10
+ from .base import BaseModel
11
+ from ..smp import *
12
+ from ..utils import DATASET_TYPE
13
+
14
+
15
+ class LLama3Mixsense(BaseModel):
16
+
17
+ INSTALL_REQ = False
18
+ INTERLEAVE = False
19
+
20
+ def __init__(self, model_path="Zero-Vision/Llama-3-MixSenseV1_1", **kwargs):
21
+ assert model_path is not None
22
+ transformers.logging.set_verbosity_error()
23
+ transformers.logging.disable_progress_bar()
24
+ warnings.filterwarnings("ignore")
25
+ self.tokenizer = AutoTokenizer.from_pretrained(
26
+ model_path, trust_remote_code=True
27
+ )
28
+ self.model = AutoModelForCausalLM.from_pretrained(
29
+ model_path, device_map="auto", trust_remote_code=True
30
+ )
31
+ self.kwargs = kwargs
32
+
33
+ def generate_inner(self, message, dataset=None):
34
+ prompt, image_path = self.message_to_promptimg(message)
35
+ input_ids=self.model.text_process(prompt, self.tokenizer)
36
+ image = Image.open(image_path).convert("RGB")
37
+ image_tensor = self.model.image_process([image]).to(dtype=self.model.dtype, device=device)
38
+ # generate
39
+ with torch.inference_mode():
40
+ output_ids = self.model.generate(
41
+ input_ids,
42
+ images=image_tensor,
43
+ max_new_tokens=2048,
44
+ use_cache=True,
45
+ eos_token_id=[
46
+ self.tokenizer.eos_token_id,
47
+ self.tokenizer.convert_tokens_to_ids(["<|eot_id|>"])[0],
48
+ ],
49
+ )
50
+ return self.tokenizer.batch_decode(output_ids, skip_special_tokens=True)[0].strip()
model-00001-of-00004.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:4d5a78df4feb4fbcb01176d679792d494d65f7acefd74fd8c5dd6ed541084970
3
+ size 4976706864
model-00002-of-00004.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1f79a62169af0c70e9955fad01ea2ac2078d31754477faa8a24377944f28c358
3
+ size 4999802720
model-00003-of-00004.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ae471b0cd3de60917f6fbb27d71b491f61ff878ba51758ce38e21304ef052a39
3
+ size 4915916176
model-00004-of-00004.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0b9e7c0dab83d5328ae4b74e429ee4eec2aaf98c43a41f90547ee2cd389a1a02
3
+ size 1939820800
model.safetensors.index.json ADDED
@@ -0,0 +1,702 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "metadata": {
3
+ "total_size": 16832151552
4
+ },
5
+ "weight_map": {
6
+ "lm_head.weight": "model-00004-of-00004.safetensors",
7
+ "model.embed_tokens.weight": "model-00001-of-00004.safetensors",
8
+ "model.layers.0.input_layernorm.weight": "model-00001-of-00004.safetensors",
9
+ "model.layers.0.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
10
+ "model.layers.0.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
11
+ "model.layers.0.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
12
+ "model.layers.0.post_attention_layernorm.weight": "model-00001-of-00004.safetensors",
13
+ "model.layers.0.self_attn.k_proj.weight": "model-00001-of-00004.safetensors",
14
+ "model.layers.0.self_attn.o_proj.weight": "model-00001-of-00004.safetensors",
15
+ "model.layers.0.self_attn.q_proj.weight": "model-00001-of-00004.safetensors",
16
+ "model.layers.0.self_attn.v_proj.weight": "model-00001-of-00004.safetensors",
17
+ "model.layers.1.input_layernorm.weight": "model-00001-of-00004.safetensors",
18
+ "model.layers.1.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
19
+ "model.layers.1.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
20
+ "model.layers.1.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
21
+ "model.layers.1.post_attention_layernorm.weight": "model-00001-of-00004.safetensors",
22
+ "model.layers.1.self_attn.k_proj.weight": "model-00001-of-00004.safetensors",
23
+ "model.layers.1.self_attn.o_proj.weight": "model-00001-of-00004.safetensors",
24
+ "model.layers.1.self_attn.q_proj.weight": "model-00001-of-00004.safetensors",
25
+ "model.layers.1.self_attn.v_proj.weight": "model-00001-of-00004.safetensors",
26
+ "model.layers.10.input_layernorm.weight": "model-00002-of-00004.safetensors",
27
+ "model.layers.10.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
28
+ "model.layers.10.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
29
+ "model.layers.10.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
30
+ "model.layers.10.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
31
+ "model.layers.10.self_attn.k_proj.weight": "model-00002-of-00004.safetensors",
32
+ "model.layers.10.self_attn.o_proj.weight": "model-00002-of-00004.safetensors",
33
+ "model.layers.10.self_attn.q_proj.weight": "model-00002-of-00004.safetensors",
34
+ "model.layers.10.self_attn.v_proj.weight": "model-00002-of-00004.safetensors",
35
+ "model.layers.11.input_layernorm.weight": "model-00002-of-00004.safetensors",
36
+ "model.layers.11.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
37
+ "model.layers.11.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
38
+ "model.layers.11.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
39
+ "model.layers.11.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
40
+ "model.layers.11.self_attn.k_proj.weight": "model-00002-of-00004.safetensors",
41
+ "model.layers.11.self_attn.o_proj.weight": "model-00002-of-00004.safetensors",
42
+ "model.layers.11.self_attn.q_proj.weight": "model-00002-of-00004.safetensors",
43
+ "model.layers.11.self_attn.v_proj.weight": "model-00002-of-00004.safetensors",
44
+ "model.layers.12.input_layernorm.weight": "model-00002-of-00004.safetensors",
45
+ "model.layers.12.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
46
+ "model.layers.12.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
47
+ "model.layers.12.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
48
+ "model.layers.12.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
49
+ "model.layers.12.self_attn.k_proj.weight": "model-00002-of-00004.safetensors",
50
+ "model.layers.12.self_attn.o_proj.weight": "model-00002-of-00004.safetensors",
51
+ "model.layers.12.self_attn.q_proj.weight": "model-00002-of-00004.safetensors",
52
+ "model.layers.12.self_attn.v_proj.weight": "model-00002-of-00004.safetensors",
53
+ "model.layers.13.input_layernorm.weight": "model-00002-of-00004.safetensors",
54
+ "model.layers.13.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
55
+ "model.layers.13.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
56
+ "model.layers.13.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
57
+ "model.layers.13.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
58
+ "model.layers.13.self_attn.k_proj.weight": "model-00002-of-00004.safetensors",
59
+ "model.layers.13.self_attn.o_proj.weight": "model-00002-of-00004.safetensors",
60
+ "model.layers.13.self_attn.q_proj.weight": "model-00002-of-00004.safetensors",
61
+ "model.layers.13.self_attn.v_proj.weight": "model-00002-of-00004.safetensors",
62
+ "model.layers.14.input_layernorm.weight": "model-00002-of-00004.safetensors",
63
+ "model.layers.14.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
64
+ "model.layers.14.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
65
+ "model.layers.14.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
66
+ "model.layers.14.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
67
+ "model.layers.14.self_attn.k_proj.weight": "model-00002-of-00004.safetensors",
68
+ "model.layers.14.self_attn.o_proj.weight": "model-00002-of-00004.safetensors",
69
+ "model.layers.14.self_attn.q_proj.weight": "model-00002-of-00004.safetensors",
70
+ "model.layers.14.self_attn.v_proj.weight": "model-00002-of-00004.safetensors",
71
+ "model.layers.15.input_layernorm.weight": "model-00002-of-00004.safetensors",
72
+ "model.layers.15.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
73
+ "model.layers.15.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
74
+ "model.layers.15.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
75
+ "model.layers.15.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
76
+ "model.layers.15.self_attn.k_proj.weight": "model-00002-of-00004.safetensors",
77
+ "model.layers.15.self_attn.o_proj.weight": "model-00002-of-00004.safetensors",
78
+ "model.layers.15.self_attn.q_proj.weight": "model-00002-of-00004.safetensors",
79
+ "model.layers.15.self_attn.v_proj.weight": "model-00002-of-00004.safetensors",
80
+ "model.layers.16.input_layernorm.weight": "model-00002-of-00004.safetensors",
81
+ "model.layers.16.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
82
+ "model.layers.16.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
83
+ "model.layers.16.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
84
+ "model.layers.16.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
85
+ "model.layers.16.self_attn.k_proj.weight": "model-00002-of-00004.safetensors",
86
+ "model.layers.16.self_attn.o_proj.weight": "model-00002-of-00004.safetensors",
87
+ "model.layers.16.self_attn.q_proj.weight": "model-00002-of-00004.safetensors",
88
+ "model.layers.16.self_attn.v_proj.weight": "model-00002-of-00004.safetensors",
89
+ "model.layers.17.input_layernorm.weight": "model-00002-of-00004.safetensors",
90
+ "model.layers.17.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
91
+ "model.layers.17.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
92
+ "model.layers.17.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
93
+ "model.layers.17.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
94
+ "model.layers.17.self_attn.k_proj.weight": "model-00002-of-00004.safetensors",
95
+ "model.layers.17.self_attn.o_proj.weight": "model-00002-of-00004.safetensors",
96
+ "model.layers.17.self_attn.q_proj.weight": "model-00002-of-00004.safetensors",
97
+ "model.layers.17.self_attn.v_proj.weight": "model-00002-of-00004.safetensors",
98
+ "model.layers.18.input_layernorm.weight": "model-00002-of-00004.safetensors",
99
+ "model.layers.18.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
100
+ "model.layers.18.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
101
+ "model.layers.18.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
102
+ "model.layers.18.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
103
+ "model.layers.18.self_attn.k_proj.weight": "model-00002-of-00004.safetensors",
104
+ "model.layers.18.self_attn.o_proj.weight": "model-00002-of-00004.safetensors",
105
+ "model.layers.18.self_attn.q_proj.weight": "model-00002-of-00004.safetensors",
106
+ "model.layers.18.self_attn.v_proj.weight": "model-00002-of-00004.safetensors",
107
+ "model.layers.19.input_layernorm.weight": "model-00002-of-00004.safetensors",
108
+ "model.layers.19.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
109
+ "model.layers.19.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
110
+ "model.layers.19.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
111
+ "model.layers.19.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
112
+ "model.layers.19.self_attn.k_proj.weight": "model-00002-of-00004.safetensors",
113
+ "model.layers.19.self_attn.o_proj.weight": "model-00002-of-00004.safetensors",
114
+ "model.layers.19.self_attn.q_proj.weight": "model-00002-of-00004.safetensors",
115
+ "model.layers.19.self_attn.v_proj.weight": "model-00002-of-00004.safetensors",
116
+ "model.layers.2.input_layernorm.weight": "model-00001-of-00004.safetensors",
117
+ "model.layers.2.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
118
+ "model.layers.2.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
119
+ "model.layers.2.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
120
+ "model.layers.2.post_attention_layernorm.weight": "model-00001-of-00004.safetensors",
121
+ "model.layers.2.self_attn.k_proj.weight": "model-00001-of-00004.safetensors",
122
+ "model.layers.2.self_attn.o_proj.weight": "model-00001-of-00004.safetensors",
123
+ "model.layers.2.self_attn.q_proj.weight": "model-00001-of-00004.safetensors",
124
+ "model.layers.2.self_attn.v_proj.weight": "model-00001-of-00004.safetensors",
125
+ "model.layers.20.input_layernorm.weight": "model-00003-of-00004.safetensors",
126
+ "model.layers.20.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
127
+ "model.layers.20.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
128
+ "model.layers.20.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
129
+ "model.layers.20.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
130
+ "model.layers.20.self_attn.k_proj.weight": "model-00002-of-00004.safetensors",
131
+ "model.layers.20.self_attn.o_proj.weight": "model-00002-of-00004.safetensors",
132
+ "model.layers.20.self_attn.q_proj.weight": "model-00002-of-00004.safetensors",
133
+ "model.layers.20.self_attn.v_proj.weight": "model-00002-of-00004.safetensors",
134
+ "model.layers.21.input_layernorm.weight": "model-00003-of-00004.safetensors",
135
+ "model.layers.21.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
136
+ "model.layers.21.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
137
+ "model.layers.21.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
138
+ "model.layers.21.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
139
+ "model.layers.21.self_attn.k_proj.weight": "model-00003-of-00004.safetensors",
140
+ "model.layers.21.self_attn.o_proj.weight": "model-00003-of-00004.safetensors",
141
+ "model.layers.21.self_attn.q_proj.weight": "model-00003-of-00004.safetensors",
142
+ "model.layers.21.self_attn.v_proj.weight": "model-00003-of-00004.safetensors",
143
+ "model.layers.22.input_layernorm.weight": "model-00003-of-00004.safetensors",
144
+ "model.layers.22.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
145
+ "model.layers.22.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
146
+ "model.layers.22.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
147
+ "model.layers.22.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
148
+ "model.layers.22.self_attn.k_proj.weight": "model-00003-of-00004.safetensors",
149
+ "model.layers.22.self_attn.o_proj.weight": "model-00003-of-00004.safetensors",
150
+ "model.layers.22.self_attn.q_proj.weight": "model-00003-of-00004.safetensors",
151
+ "model.layers.22.self_attn.v_proj.weight": "model-00003-of-00004.safetensors",
152
+ "model.layers.23.input_layernorm.weight": "model-00003-of-00004.safetensors",
153
+ "model.layers.23.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
154
+ "model.layers.23.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
155
+ "model.layers.23.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
156
+ "model.layers.23.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
157
+ "model.layers.23.self_attn.k_proj.weight": "model-00003-of-00004.safetensors",
158
+ "model.layers.23.self_attn.o_proj.weight": "model-00003-of-00004.safetensors",
159
+ "model.layers.23.self_attn.q_proj.weight": "model-00003-of-00004.safetensors",
160
+ "model.layers.23.self_attn.v_proj.weight": "model-00003-of-00004.safetensors",
161
+ "model.layers.24.input_layernorm.weight": "model-00003-of-00004.safetensors",
162
+ "model.layers.24.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
163
+ "model.layers.24.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
164
+ "model.layers.24.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
165
+ "model.layers.24.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
166
+ "model.layers.24.self_attn.k_proj.weight": "model-00003-of-00004.safetensors",
167
+ "model.layers.24.self_attn.o_proj.weight": "model-00003-of-00004.safetensors",
168
+ "model.layers.24.self_attn.q_proj.weight": "model-00003-of-00004.safetensors",
169
+ "model.layers.24.self_attn.v_proj.weight": "model-00003-of-00004.safetensors",
170
+ "model.layers.25.input_layernorm.weight": "model-00003-of-00004.safetensors",
171
+ "model.layers.25.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
172
+ "model.layers.25.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
173
+ "model.layers.25.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
174
+ "model.layers.25.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
175
+ "model.layers.25.self_attn.k_proj.weight": "model-00003-of-00004.safetensors",
176
+ "model.layers.25.self_attn.o_proj.weight": "model-00003-of-00004.safetensors",
177
+ "model.layers.25.self_attn.q_proj.weight": "model-00003-of-00004.safetensors",
178
+ "model.layers.25.self_attn.v_proj.weight": "model-00003-of-00004.safetensors",
179
+ "model.layers.26.input_layernorm.weight": "model-00003-of-00004.safetensors",
180
+ "model.layers.26.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
181
+ "model.layers.26.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
182
+ "model.layers.26.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
183
+ "model.layers.26.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
184
+ "model.layers.26.self_attn.k_proj.weight": "model-00003-of-00004.safetensors",
185
+ "model.layers.26.self_attn.o_proj.weight": "model-00003-of-00004.safetensors",
186
+ "model.layers.26.self_attn.q_proj.weight": "model-00003-of-00004.safetensors",
187
+ "model.layers.26.self_attn.v_proj.weight": "model-00003-of-00004.safetensors",
188
+ "model.layers.27.input_layernorm.weight": "model-00003-of-00004.safetensors",
189
+ "model.layers.27.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
190
+ "model.layers.27.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
191
+ "model.layers.27.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
192
+ "model.layers.27.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
193
+ "model.layers.27.self_attn.k_proj.weight": "model-00003-of-00004.safetensors",
194
+ "model.layers.27.self_attn.o_proj.weight": "model-00003-of-00004.safetensors",
195
+ "model.layers.27.self_attn.q_proj.weight": "model-00003-of-00004.safetensors",
196
+ "model.layers.27.self_attn.v_proj.weight": "model-00003-of-00004.safetensors",
197
+ "model.layers.28.input_layernorm.weight": "model-00003-of-00004.safetensors",
198
+ "model.layers.28.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
199
+ "model.layers.28.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
200
+ "model.layers.28.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
201
+ "model.layers.28.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
202
+ "model.layers.28.self_attn.k_proj.weight": "model-00003-of-00004.safetensors",
203
+ "model.layers.28.self_attn.o_proj.weight": "model-00003-of-00004.safetensors",
204
+ "model.layers.28.self_attn.q_proj.weight": "model-00003-of-00004.safetensors",
205
+ "model.layers.28.self_attn.v_proj.weight": "model-00003-of-00004.safetensors",
206
+ "model.layers.29.input_layernorm.weight": "model-00003-of-00004.safetensors",
207
+ "model.layers.29.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
208
+ "model.layers.29.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
209
+ "model.layers.29.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
210
+ "model.layers.29.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
211
+ "model.layers.29.self_attn.k_proj.weight": "model-00003-of-00004.safetensors",
212
+ "model.layers.29.self_attn.o_proj.weight": "model-00003-of-00004.safetensors",
213
+ "model.layers.29.self_attn.q_proj.weight": "model-00003-of-00004.safetensors",
214
+ "model.layers.29.self_attn.v_proj.weight": "model-00003-of-00004.safetensors",
215
+ "model.layers.3.input_layernorm.weight": "model-00001-of-00004.safetensors",
216
+ "model.layers.3.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
217
+ "model.layers.3.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
218
+ "model.layers.3.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
219
+ "model.layers.3.post_attention_layernorm.weight": "model-00001-of-00004.safetensors",
220
+ "model.layers.3.self_attn.k_proj.weight": "model-00001-of-00004.safetensors",
221
+ "model.layers.3.self_attn.o_proj.weight": "model-00001-of-00004.safetensors",
222
+ "model.layers.3.self_attn.q_proj.weight": "model-00001-of-00004.safetensors",
223
+ "model.layers.3.self_attn.v_proj.weight": "model-00001-of-00004.safetensors",
224
+ "model.layers.30.input_layernorm.weight": "model-00003-of-00004.safetensors",
225
+ "model.layers.30.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
226
+ "model.layers.30.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
227
+ "model.layers.30.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
228
+ "model.layers.30.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
229
+ "model.layers.30.self_attn.k_proj.weight": "model-00003-of-00004.safetensors",
230
+ "model.layers.30.self_attn.o_proj.weight": "model-00003-of-00004.safetensors",
231
+ "model.layers.30.self_attn.q_proj.weight": "model-00003-of-00004.safetensors",
232
+ "model.layers.30.self_attn.v_proj.weight": "model-00003-of-00004.safetensors",
233
+ "model.layers.31.input_layernorm.weight": "model-00004-of-00004.safetensors",
234
+ "model.layers.31.mlp.down_proj.weight": "model-00004-of-00004.safetensors",
235
+ "model.layers.31.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
236
+ "model.layers.31.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
237
+ "model.layers.31.post_attention_layernorm.weight": "model-00004-of-00004.safetensors",
238
+ "model.layers.31.self_attn.k_proj.weight": "model-00003-of-00004.safetensors",
239
+ "model.layers.31.self_attn.o_proj.weight": "model-00003-of-00004.safetensors",
240
+ "model.layers.31.self_attn.q_proj.weight": "model-00003-of-00004.safetensors",
241
+ "model.layers.31.self_attn.v_proj.weight": "model-00003-of-00004.safetensors",
242
+ "model.layers.4.input_layernorm.weight": "model-00001-of-00004.safetensors",
243
+ "model.layers.4.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
244
+ "model.layers.4.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
245
+ "model.layers.4.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
246
+ "model.layers.4.post_attention_layernorm.weight": "model-00001-of-00004.safetensors",
247
+ "model.layers.4.self_attn.k_proj.weight": "model-00001-of-00004.safetensors",
248
+ "model.layers.4.self_attn.o_proj.weight": "model-00001-of-00004.safetensors",
249
+ "model.layers.4.self_attn.q_proj.weight": "model-00001-of-00004.safetensors",
250
+ "model.layers.4.self_attn.v_proj.weight": "model-00001-of-00004.safetensors",
251
+ "model.layers.5.input_layernorm.weight": "model-00001-of-00004.safetensors",
252
+ "model.layers.5.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
253
+ "model.layers.5.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
254
+ "model.layers.5.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
255
+ "model.layers.5.post_attention_layernorm.weight": "model-00001-of-00004.safetensors",
256
+ "model.layers.5.self_attn.k_proj.weight": "model-00001-of-00004.safetensors",
257
+ "model.layers.5.self_attn.o_proj.weight": "model-00001-of-00004.safetensors",
258
+ "model.layers.5.self_attn.q_proj.weight": "model-00001-of-00004.safetensors",
259
+ "model.layers.5.self_attn.v_proj.weight": "model-00001-of-00004.safetensors",
260
+ "model.layers.6.input_layernorm.weight": "model-00001-of-00004.safetensors",
261
+ "model.layers.6.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
262
+ "model.layers.6.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
263
+ "model.layers.6.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
264
+ "model.layers.6.post_attention_layernorm.weight": "model-00001-of-00004.safetensors",
265
+ "model.layers.6.self_attn.k_proj.weight": "model-00001-of-00004.safetensors",
266
+ "model.layers.6.self_attn.o_proj.weight": "model-00001-of-00004.safetensors",
267
+ "model.layers.6.self_attn.q_proj.weight": "model-00001-of-00004.safetensors",
268
+ "model.layers.6.self_attn.v_proj.weight": "model-00001-of-00004.safetensors",
269
+ "model.layers.7.input_layernorm.weight": "model-00001-of-00004.safetensors",
270
+ "model.layers.7.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
271
+ "model.layers.7.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
272
+ "model.layers.7.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
273
+ "model.layers.7.post_attention_layernorm.weight": "model-00001-of-00004.safetensors",
274
+ "model.layers.7.self_attn.k_proj.weight": "model-00001-of-00004.safetensors",
275
+ "model.layers.7.self_attn.o_proj.weight": "model-00001-of-00004.safetensors",
276
+ "model.layers.7.self_attn.q_proj.weight": "model-00001-of-00004.safetensors",
277
+ "model.layers.7.self_attn.v_proj.weight": "model-00001-of-00004.safetensors",
278
+ "model.layers.8.input_layernorm.weight": "model-00001-of-00004.safetensors",
279
+ "model.layers.8.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
280
+ "model.layers.8.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
281
+ "model.layers.8.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
282
+ "model.layers.8.post_attention_layernorm.weight": "model-00001-of-00004.safetensors",
283
+ "model.layers.8.self_attn.k_proj.weight": "model-00001-of-00004.safetensors",
284
+ "model.layers.8.self_attn.o_proj.weight": "model-00001-of-00004.safetensors",
285
+ "model.layers.8.self_attn.q_proj.weight": "model-00001-of-00004.safetensors",
286
+ "model.layers.8.self_attn.v_proj.weight": "model-00001-of-00004.safetensors",
287
+ "model.layers.9.input_layernorm.weight": "model-00002-of-00004.safetensors",
288
+ "model.layers.9.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
289
+ "model.layers.9.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
290
+ "model.layers.9.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
291
+ "model.layers.9.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
292
+ "model.layers.9.self_attn.k_proj.weight": "model-00002-of-00004.safetensors",
293
+ "model.layers.9.self_attn.o_proj.weight": "model-00002-of-00004.safetensors",
294
+ "model.layers.9.self_attn.q_proj.weight": "model-00002-of-00004.safetensors",
295
+ "model.layers.9.self_attn.v_proj.weight": "model-00002-of-00004.safetensors",
296
+ "model.mm_projector.0.bias": "model-00004-of-00004.safetensors",
297
+ "model.mm_projector.0.weight": "model-00004-of-00004.safetensors",
298
+ "model.mm_projector.2.bias": "model-00004-of-00004.safetensors",
299
+ "model.mm_projector.2.weight": "model-00004-of-00004.safetensors",
300
+ "model.norm.weight": "model-00004-of-00004.safetensors",
301
+ "model.vision_tower.vision_tower.blocks.0.0.channel_block.channel_attn.fn.proj.bias": "model-00004-of-00004.safetensors",
302
+ "model.vision_tower.vision_tower.blocks.0.0.channel_block.channel_attn.fn.proj.weight": "model-00004-of-00004.safetensors",
303
+ "model.vision_tower.vision_tower.blocks.0.0.channel_block.channel_attn.fn.qkv.bias": "model-00004-of-00004.safetensors",
304
+ "model.vision_tower.vision_tower.blocks.0.0.channel_block.channel_attn.fn.qkv.weight": "model-00004-of-00004.safetensors",
305
+ "model.vision_tower.vision_tower.blocks.0.0.channel_block.channel_attn.norm.bias": "model-00004-of-00004.safetensors",
306
+ "model.vision_tower.vision_tower.blocks.0.0.channel_block.channel_attn.norm.weight": "model-00004-of-00004.safetensors",
307
+ "model.vision_tower.vision_tower.blocks.0.0.channel_block.conv1.fn.dw.bias": "model-00004-of-00004.safetensors",
308
+ "model.vision_tower.vision_tower.blocks.0.0.channel_block.conv1.fn.dw.weight": "model-00004-of-00004.safetensors",
309
+ "model.vision_tower.vision_tower.blocks.0.0.channel_block.conv2.fn.dw.bias": "model-00004-of-00004.safetensors",
310
+ "model.vision_tower.vision_tower.blocks.0.0.channel_block.conv2.fn.dw.weight": "model-00004-of-00004.safetensors",
311
+ "model.vision_tower.vision_tower.blocks.0.0.channel_block.ffn.fn.net.fc1.bias": "model-00004-of-00004.safetensors",
312
+ "model.vision_tower.vision_tower.blocks.0.0.channel_block.ffn.fn.net.fc1.weight": "model-00004-of-00004.safetensors",
313
+ "model.vision_tower.vision_tower.blocks.0.0.channel_block.ffn.fn.net.fc2.bias": "model-00004-of-00004.safetensors",
314
+ "model.vision_tower.vision_tower.blocks.0.0.channel_block.ffn.fn.net.fc2.weight": "model-00004-of-00004.safetensors",
315
+ "model.vision_tower.vision_tower.blocks.0.0.channel_block.ffn.norm.bias": "model-00004-of-00004.safetensors",
316
+ "model.vision_tower.vision_tower.blocks.0.0.channel_block.ffn.norm.weight": "model-00004-of-00004.safetensors",
317
+ "model.vision_tower.vision_tower.blocks.0.0.spatial_block.conv1.fn.dw.bias": "model-00004-of-00004.safetensors",
318
+ "model.vision_tower.vision_tower.blocks.0.0.spatial_block.conv1.fn.dw.weight": "model-00004-of-00004.safetensors",
319
+ "model.vision_tower.vision_tower.blocks.0.0.spatial_block.conv2.fn.dw.bias": "model-00004-of-00004.safetensors",
320
+ "model.vision_tower.vision_tower.blocks.0.0.spatial_block.conv2.fn.dw.weight": "model-00004-of-00004.safetensors",
321
+ "model.vision_tower.vision_tower.blocks.0.0.spatial_block.ffn.fn.net.fc1.bias": "model-00004-of-00004.safetensors",
322
+ "model.vision_tower.vision_tower.blocks.0.0.spatial_block.ffn.fn.net.fc1.weight": "model-00004-of-00004.safetensors",
323
+ "model.vision_tower.vision_tower.blocks.0.0.spatial_block.ffn.fn.net.fc2.bias": "model-00004-of-00004.safetensors",
324
+ "model.vision_tower.vision_tower.blocks.0.0.spatial_block.ffn.fn.net.fc2.weight": "model-00004-of-00004.safetensors",
325
+ "model.vision_tower.vision_tower.blocks.0.0.spatial_block.ffn.norm.bias": "model-00004-of-00004.safetensors",
326
+ "model.vision_tower.vision_tower.blocks.0.0.spatial_block.ffn.norm.weight": "model-00004-of-00004.safetensors",
327
+ "model.vision_tower.vision_tower.blocks.0.0.spatial_block.window_attn.fn.proj.bias": "model-00004-of-00004.safetensors",
328
+ "model.vision_tower.vision_tower.blocks.0.0.spatial_block.window_attn.fn.proj.weight": "model-00004-of-00004.safetensors",
329
+ "model.vision_tower.vision_tower.blocks.0.0.spatial_block.window_attn.fn.qkv.bias": "model-00004-of-00004.safetensors",
330
+ "model.vision_tower.vision_tower.blocks.0.0.spatial_block.window_attn.fn.qkv.weight": "model-00004-of-00004.safetensors",
331
+ "model.vision_tower.vision_tower.blocks.0.0.spatial_block.window_attn.norm.bias": "model-00004-of-00004.safetensors",
332
+ "model.vision_tower.vision_tower.blocks.0.0.spatial_block.window_attn.norm.weight": "model-00004-of-00004.safetensors",
333
+ "model.vision_tower.vision_tower.blocks.1.0.channel_block.channel_attn.fn.proj.bias": "model-00004-of-00004.safetensors",
334
+ "model.vision_tower.vision_tower.blocks.1.0.channel_block.channel_attn.fn.proj.weight": "model-00004-of-00004.safetensors",
335
+ "model.vision_tower.vision_tower.blocks.1.0.channel_block.channel_attn.fn.qkv.bias": "model-00004-of-00004.safetensors",
336
+ "model.vision_tower.vision_tower.blocks.1.0.channel_block.channel_attn.fn.qkv.weight": "model-00004-of-00004.safetensors",
337
+ "model.vision_tower.vision_tower.blocks.1.0.channel_block.channel_attn.norm.bias": "model-00004-of-00004.safetensors",
338
+ "model.vision_tower.vision_tower.blocks.1.0.channel_block.channel_attn.norm.weight": "model-00004-of-00004.safetensors",
339
+ "model.vision_tower.vision_tower.blocks.1.0.channel_block.conv1.fn.dw.bias": "model-00004-of-00004.safetensors",
340
+ "model.vision_tower.vision_tower.blocks.1.0.channel_block.conv1.fn.dw.weight": "model-00004-of-00004.safetensors",
341
+ "model.vision_tower.vision_tower.blocks.1.0.channel_block.conv2.fn.dw.bias": "model-00004-of-00004.safetensors",
342
+ "model.vision_tower.vision_tower.blocks.1.0.channel_block.conv2.fn.dw.weight": "model-00004-of-00004.safetensors",
343
+ "model.vision_tower.vision_tower.blocks.1.0.channel_block.ffn.fn.net.fc1.bias": "model-00004-of-00004.safetensors",
344
+ "model.vision_tower.vision_tower.blocks.1.0.channel_block.ffn.fn.net.fc1.weight": "model-00004-of-00004.safetensors",
345
+ "model.vision_tower.vision_tower.blocks.1.0.channel_block.ffn.fn.net.fc2.bias": "model-00004-of-00004.safetensors",
346
+ "model.vision_tower.vision_tower.blocks.1.0.channel_block.ffn.fn.net.fc2.weight": "model-00004-of-00004.safetensors",
347
+ "model.vision_tower.vision_tower.blocks.1.0.channel_block.ffn.norm.bias": "model-00004-of-00004.safetensors",
348
+ "model.vision_tower.vision_tower.blocks.1.0.channel_block.ffn.norm.weight": "model-00004-of-00004.safetensors",
349
+ "model.vision_tower.vision_tower.blocks.1.0.spatial_block.conv1.fn.dw.bias": "model-00004-of-00004.safetensors",
350
+ "model.vision_tower.vision_tower.blocks.1.0.spatial_block.conv1.fn.dw.weight": "model-00004-of-00004.safetensors",
351
+ "model.vision_tower.vision_tower.blocks.1.0.spatial_block.conv2.fn.dw.bias": "model-00004-of-00004.safetensors",
352
+ "model.vision_tower.vision_tower.blocks.1.0.spatial_block.conv2.fn.dw.weight": "model-00004-of-00004.safetensors",
353
+ "model.vision_tower.vision_tower.blocks.1.0.spatial_block.ffn.fn.net.fc1.bias": "model-00004-of-00004.safetensors",
354
+ "model.vision_tower.vision_tower.blocks.1.0.spatial_block.ffn.fn.net.fc1.weight": "model-00004-of-00004.safetensors",
355
+ "model.vision_tower.vision_tower.blocks.1.0.spatial_block.ffn.fn.net.fc2.bias": "model-00004-of-00004.safetensors",
356
+ "model.vision_tower.vision_tower.blocks.1.0.spatial_block.ffn.fn.net.fc2.weight": "model-00004-of-00004.safetensors",
357
+ "model.vision_tower.vision_tower.blocks.1.0.spatial_block.ffn.norm.bias": "model-00004-of-00004.safetensors",
358
+ "model.vision_tower.vision_tower.blocks.1.0.spatial_block.ffn.norm.weight": "model-00004-of-00004.safetensors",
359
+ "model.vision_tower.vision_tower.blocks.1.0.spatial_block.window_attn.fn.proj.bias": "model-00004-of-00004.safetensors",
360
+ "model.vision_tower.vision_tower.blocks.1.0.spatial_block.window_attn.fn.proj.weight": "model-00004-of-00004.safetensors",
361
+ "model.vision_tower.vision_tower.blocks.1.0.spatial_block.window_attn.fn.qkv.bias": "model-00004-of-00004.safetensors",
362
+ "model.vision_tower.vision_tower.blocks.1.0.spatial_block.window_attn.fn.qkv.weight": "model-00004-of-00004.safetensors",
363
+ "model.vision_tower.vision_tower.blocks.1.0.spatial_block.window_attn.norm.bias": "model-00004-of-00004.safetensors",
364
+ "model.vision_tower.vision_tower.blocks.1.0.spatial_block.window_attn.norm.weight": "model-00004-of-00004.safetensors",
365
+ "model.vision_tower.vision_tower.blocks.2.0.channel_block.channel_attn.fn.proj.bias": "model-00004-of-00004.safetensors",
366
+ "model.vision_tower.vision_tower.blocks.2.0.channel_block.channel_attn.fn.proj.weight": "model-00004-of-00004.safetensors",
367
+ "model.vision_tower.vision_tower.blocks.2.0.channel_block.channel_attn.fn.qkv.bias": "model-00004-of-00004.safetensors",
368
+ "model.vision_tower.vision_tower.blocks.2.0.channel_block.channel_attn.fn.qkv.weight": "model-00004-of-00004.safetensors",
369
+ "model.vision_tower.vision_tower.blocks.2.0.channel_block.channel_attn.norm.bias": "model-00004-of-00004.safetensors",
370
+ "model.vision_tower.vision_tower.blocks.2.0.channel_block.channel_attn.norm.weight": "model-00004-of-00004.safetensors",
371
+ "model.vision_tower.vision_tower.blocks.2.0.channel_block.conv1.fn.dw.bias": "model-00004-of-00004.safetensors",
372
+ "model.vision_tower.vision_tower.blocks.2.0.channel_block.conv1.fn.dw.weight": "model-00004-of-00004.safetensors",
373
+ "model.vision_tower.vision_tower.blocks.2.0.channel_block.conv2.fn.dw.bias": "model-00004-of-00004.safetensors",
374
+ "model.vision_tower.vision_tower.blocks.2.0.channel_block.conv2.fn.dw.weight": "model-00004-of-00004.safetensors",
375
+ "model.vision_tower.vision_tower.blocks.2.0.channel_block.ffn.fn.net.fc1.bias": "model-00004-of-00004.safetensors",
376
+ "model.vision_tower.vision_tower.blocks.2.0.channel_block.ffn.fn.net.fc1.weight": "model-00004-of-00004.safetensors",
377
+ "model.vision_tower.vision_tower.blocks.2.0.channel_block.ffn.fn.net.fc2.bias": "model-00004-of-00004.safetensors",
378
+ "model.vision_tower.vision_tower.blocks.2.0.channel_block.ffn.fn.net.fc2.weight": "model-00004-of-00004.safetensors",
379
+ "model.vision_tower.vision_tower.blocks.2.0.channel_block.ffn.norm.bias": "model-00004-of-00004.safetensors",
380
+ "model.vision_tower.vision_tower.blocks.2.0.channel_block.ffn.norm.weight": "model-00004-of-00004.safetensors",
381
+ "model.vision_tower.vision_tower.blocks.2.0.spatial_block.conv1.fn.dw.bias": "model-00004-of-00004.safetensors",
382
+ "model.vision_tower.vision_tower.blocks.2.0.spatial_block.conv1.fn.dw.weight": "model-00004-of-00004.safetensors",
383
+ "model.vision_tower.vision_tower.blocks.2.0.spatial_block.conv2.fn.dw.bias": "model-00004-of-00004.safetensors",
384
+ "model.vision_tower.vision_tower.blocks.2.0.spatial_block.conv2.fn.dw.weight": "model-00004-of-00004.safetensors",
385
+ "model.vision_tower.vision_tower.blocks.2.0.spatial_block.ffn.fn.net.fc1.bias": "model-00004-of-00004.safetensors",
386
+ "model.vision_tower.vision_tower.blocks.2.0.spatial_block.ffn.fn.net.fc1.weight": "model-00004-of-00004.safetensors",
387
+ "model.vision_tower.vision_tower.blocks.2.0.spatial_block.ffn.fn.net.fc2.bias": "model-00004-of-00004.safetensors",
388
+ "model.vision_tower.vision_tower.blocks.2.0.spatial_block.ffn.fn.net.fc2.weight": "model-00004-of-00004.safetensors",
389
+ "model.vision_tower.vision_tower.blocks.2.0.spatial_block.ffn.norm.bias": "model-00004-of-00004.safetensors",
390
+ "model.vision_tower.vision_tower.blocks.2.0.spatial_block.ffn.norm.weight": "model-00004-of-00004.safetensors",
391
+ "model.vision_tower.vision_tower.blocks.2.0.spatial_block.window_attn.fn.proj.bias": "model-00004-of-00004.safetensors",
392
+ "model.vision_tower.vision_tower.blocks.2.0.spatial_block.window_attn.fn.proj.weight": "model-00004-of-00004.safetensors",
393
+ "model.vision_tower.vision_tower.blocks.2.0.spatial_block.window_attn.fn.qkv.bias": "model-00004-of-00004.safetensors",
394
+ "model.vision_tower.vision_tower.blocks.2.0.spatial_block.window_attn.fn.qkv.weight": "model-00004-of-00004.safetensors",
395
+ "model.vision_tower.vision_tower.blocks.2.0.spatial_block.window_attn.norm.bias": "model-00004-of-00004.safetensors",
396
+ "model.vision_tower.vision_tower.blocks.2.0.spatial_block.window_attn.norm.weight": "model-00004-of-00004.safetensors",
397
+ "model.vision_tower.vision_tower.blocks.2.1.channel_block.channel_attn.fn.proj.bias": "model-00004-of-00004.safetensors",
398
+ "model.vision_tower.vision_tower.blocks.2.1.channel_block.channel_attn.fn.proj.weight": "model-00004-of-00004.safetensors",
399
+ "model.vision_tower.vision_tower.blocks.2.1.channel_block.channel_attn.fn.qkv.bias": "model-00004-of-00004.safetensors",
400
+ "model.vision_tower.vision_tower.blocks.2.1.channel_block.channel_attn.fn.qkv.weight": "model-00004-of-00004.safetensors",
401
+ "model.vision_tower.vision_tower.blocks.2.1.channel_block.channel_attn.norm.bias": "model-00004-of-00004.safetensors",
402
+ "model.vision_tower.vision_tower.blocks.2.1.channel_block.channel_attn.norm.weight": "model-00004-of-00004.safetensors",
403
+ "model.vision_tower.vision_tower.blocks.2.1.channel_block.conv1.fn.dw.bias": "model-00004-of-00004.safetensors",
404
+ "model.vision_tower.vision_tower.blocks.2.1.channel_block.conv1.fn.dw.weight": "model-00004-of-00004.safetensors",
405
+ "model.vision_tower.vision_tower.blocks.2.1.channel_block.conv2.fn.dw.bias": "model-00004-of-00004.safetensors",
406
+ "model.vision_tower.vision_tower.blocks.2.1.channel_block.conv2.fn.dw.weight": "model-00004-of-00004.safetensors",
407
+ "model.vision_tower.vision_tower.blocks.2.1.channel_block.ffn.fn.net.fc1.bias": "model-00004-of-00004.safetensors",
408
+ "model.vision_tower.vision_tower.blocks.2.1.channel_block.ffn.fn.net.fc1.weight": "model-00004-of-00004.safetensors",
409
+ "model.vision_tower.vision_tower.blocks.2.1.channel_block.ffn.fn.net.fc2.bias": "model-00004-of-00004.safetensors",
410
+ "model.vision_tower.vision_tower.blocks.2.1.channel_block.ffn.fn.net.fc2.weight": "model-00004-of-00004.safetensors",
411
+ "model.vision_tower.vision_tower.blocks.2.1.channel_block.ffn.norm.bias": "model-00004-of-00004.safetensors",
412
+ "model.vision_tower.vision_tower.blocks.2.1.channel_block.ffn.norm.weight": "model-00004-of-00004.safetensors",
413
+ "model.vision_tower.vision_tower.blocks.2.1.spatial_block.conv1.fn.dw.bias": "model-00004-of-00004.safetensors",
414
+ "model.vision_tower.vision_tower.blocks.2.1.spatial_block.conv1.fn.dw.weight": "model-00004-of-00004.safetensors",
415
+ "model.vision_tower.vision_tower.blocks.2.1.spatial_block.conv2.fn.dw.bias": "model-00004-of-00004.safetensors",
416
+ "model.vision_tower.vision_tower.blocks.2.1.spatial_block.conv2.fn.dw.weight": "model-00004-of-00004.safetensors",
417
+ "model.vision_tower.vision_tower.blocks.2.1.spatial_block.ffn.fn.net.fc1.bias": "model-00004-of-00004.safetensors",
418
+ "model.vision_tower.vision_tower.blocks.2.1.spatial_block.ffn.fn.net.fc1.weight": "model-00004-of-00004.safetensors",
419
+ "model.vision_tower.vision_tower.blocks.2.1.spatial_block.ffn.fn.net.fc2.bias": "model-00004-of-00004.safetensors",
420
+ "model.vision_tower.vision_tower.blocks.2.1.spatial_block.ffn.fn.net.fc2.weight": "model-00004-of-00004.safetensors",
421
+ "model.vision_tower.vision_tower.blocks.2.1.spatial_block.ffn.norm.bias": "model-00004-of-00004.safetensors",
422
+ "model.vision_tower.vision_tower.blocks.2.1.spatial_block.ffn.norm.weight": "model-00004-of-00004.safetensors",
423
+ "model.vision_tower.vision_tower.blocks.2.1.spatial_block.window_attn.fn.proj.bias": "model-00004-of-00004.safetensors",
424
+ "model.vision_tower.vision_tower.blocks.2.1.spatial_block.window_attn.fn.proj.weight": "model-00004-of-00004.safetensors",
425
+ "model.vision_tower.vision_tower.blocks.2.1.spatial_block.window_attn.fn.qkv.bias": "model-00004-of-00004.safetensors",
426
+ "model.vision_tower.vision_tower.blocks.2.1.spatial_block.window_attn.fn.qkv.weight": "model-00004-of-00004.safetensors",
427
+ "model.vision_tower.vision_tower.blocks.2.1.spatial_block.window_attn.norm.bias": "model-00004-of-00004.safetensors",
428
+ "model.vision_tower.vision_tower.blocks.2.1.spatial_block.window_attn.norm.weight": "model-00004-of-00004.safetensors",
429
+ "model.vision_tower.vision_tower.blocks.2.2.channel_block.channel_attn.fn.proj.bias": "model-00004-of-00004.safetensors",
430
+ "model.vision_tower.vision_tower.blocks.2.2.channel_block.channel_attn.fn.proj.weight": "model-00004-of-00004.safetensors",
431
+ "model.vision_tower.vision_tower.blocks.2.2.channel_block.channel_attn.fn.qkv.bias": "model-00004-of-00004.safetensors",
432
+ "model.vision_tower.vision_tower.blocks.2.2.channel_block.channel_attn.fn.qkv.weight": "model-00004-of-00004.safetensors",
433
+ "model.vision_tower.vision_tower.blocks.2.2.channel_block.channel_attn.norm.bias": "model-00004-of-00004.safetensors",
434
+ "model.vision_tower.vision_tower.blocks.2.2.channel_block.channel_attn.norm.weight": "model-00004-of-00004.safetensors",
435
+ "model.vision_tower.vision_tower.blocks.2.2.channel_block.conv1.fn.dw.bias": "model-00004-of-00004.safetensors",
436
+ "model.vision_tower.vision_tower.blocks.2.2.channel_block.conv1.fn.dw.weight": "model-00004-of-00004.safetensors",
437
+ "model.vision_tower.vision_tower.blocks.2.2.channel_block.conv2.fn.dw.bias": "model-00004-of-00004.safetensors",
438
+ "model.vision_tower.vision_tower.blocks.2.2.channel_block.conv2.fn.dw.weight": "model-00004-of-00004.safetensors",
439
+ "model.vision_tower.vision_tower.blocks.2.2.channel_block.ffn.fn.net.fc1.bias": "model-00004-of-00004.safetensors",
440
+ "model.vision_tower.vision_tower.blocks.2.2.channel_block.ffn.fn.net.fc1.weight": "model-00004-of-00004.safetensors",
441
+ "model.vision_tower.vision_tower.blocks.2.2.channel_block.ffn.fn.net.fc2.bias": "model-00004-of-00004.safetensors",
442
+ "model.vision_tower.vision_tower.blocks.2.2.channel_block.ffn.fn.net.fc2.weight": "model-00004-of-00004.safetensors",
443
+ "model.vision_tower.vision_tower.blocks.2.2.channel_block.ffn.norm.bias": "model-00004-of-00004.safetensors",
444
+ "model.vision_tower.vision_tower.blocks.2.2.channel_block.ffn.norm.weight": "model-00004-of-00004.safetensors",
445
+ "model.vision_tower.vision_tower.blocks.2.2.spatial_block.conv1.fn.dw.bias": "model-00004-of-00004.safetensors",
446
+ "model.vision_tower.vision_tower.blocks.2.2.spatial_block.conv1.fn.dw.weight": "model-00004-of-00004.safetensors",
447
+ "model.vision_tower.vision_tower.blocks.2.2.spatial_block.conv2.fn.dw.bias": "model-00004-of-00004.safetensors",
448
+ "model.vision_tower.vision_tower.blocks.2.2.spatial_block.conv2.fn.dw.weight": "model-00004-of-00004.safetensors",
449
+ "model.vision_tower.vision_tower.blocks.2.2.spatial_block.ffn.fn.net.fc1.bias": "model-00004-of-00004.safetensors",
450
+ "model.vision_tower.vision_tower.blocks.2.2.spatial_block.ffn.fn.net.fc1.weight": "model-00004-of-00004.safetensors",
451
+ "model.vision_tower.vision_tower.blocks.2.2.spatial_block.ffn.fn.net.fc2.bias": "model-00004-of-00004.safetensors",
452
+ "model.vision_tower.vision_tower.blocks.2.2.spatial_block.ffn.fn.net.fc2.weight": "model-00004-of-00004.safetensors",
453
+ "model.vision_tower.vision_tower.blocks.2.2.spatial_block.ffn.norm.bias": "model-00004-of-00004.safetensors",
454
+ "model.vision_tower.vision_tower.blocks.2.2.spatial_block.ffn.norm.weight": "model-00004-of-00004.safetensors",
455
+ "model.vision_tower.vision_tower.blocks.2.2.spatial_block.window_attn.fn.proj.bias": "model-00004-of-00004.safetensors",
456
+ "model.vision_tower.vision_tower.blocks.2.2.spatial_block.window_attn.fn.proj.weight": "model-00004-of-00004.safetensors",
457
+ "model.vision_tower.vision_tower.blocks.2.2.spatial_block.window_attn.fn.qkv.bias": "model-00004-of-00004.safetensors",
458
+ "model.vision_tower.vision_tower.blocks.2.2.spatial_block.window_attn.fn.qkv.weight": "model-00004-of-00004.safetensors",
459
+ "model.vision_tower.vision_tower.blocks.2.2.spatial_block.window_attn.norm.bias": "model-00004-of-00004.safetensors",
460
+ "model.vision_tower.vision_tower.blocks.2.2.spatial_block.window_attn.norm.weight": "model-00004-of-00004.safetensors",
461
+ "model.vision_tower.vision_tower.blocks.2.3.channel_block.channel_attn.fn.proj.bias": "model-00004-of-00004.safetensors",
462
+ "model.vision_tower.vision_tower.blocks.2.3.channel_block.channel_attn.fn.proj.weight": "model-00004-of-00004.safetensors",
463
+ "model.vision_tower.vision_tower.blocks.2.3.channel_block.channel_attn.fn.qkv.bias": "model-00004-of-00004.safetensors",
464
+ "model.vision_tower.vision_tower.blocks.2.3.channel_block.channel_attn.fn.qkv.weight": "model-00004-of-00004.safetensors",
465
+ "model.vision_tower.vision_tower.blocks.2.3.channel_block.channel_attn.norm.bias": "model-00004-of-00004.safetensors",
466
+ "model.vision_tower.vision_tower.blocks.2.3.channel_block.channel_attn.norm.weight": "model-00004-of-00004.safetensors",
467
+ "model.vision_tower.vision_tower.blocks.2.3.channel_block.conv1.fn.dw.bias": "model-00004-of-00004.safetensors",
468
+ "model.vision_tower.vision_tower.blocks.2.3.channel_block.conv1.fn.dw.weight": "model-00004-of-00004.safetensors",
469
+ "model.vision_tower.vision_tower.blocks.2.3.channel_block.conv2.fn.dw.bias": "model-00004-of-00004.safetensors",
470
+ "model.vision_tower.vision_tower.blocks.2.3.channel_block.conv2.fn.dw.weight": "model-00004-of-00004.safetensors",
471
+ "model.vision_tower.vision_tower.blocks.2.3.channel_block.ffn.fn.net.fc1.bias": "model-00004-of-00004.safetensors",
472
+ "model.vision_tower.vision_tower.blocks.2.3.channel_block.ffn.fn.net.fc1.weight": "model-00004-of-00004.safetensors",
473
+ "model.vision_tower.vision_tower.blocks.2.3.channel_block.ffn.fn.net.fc2.bias": "model-00004-of-00004.safetensors",
474
+ "model.vision_tower.vision_tower.blocks.2.3.channel_block.ffn.fn.net.fc2.weight": "model-00004-of-00004.safetensors",
475
+ "model.vision_tower.vision_tower.blocks.2.3.channel_block.ffn.norm.bias": "model-00004-of-00004.safetensors",
476
+ "model.vision_tower.vision_tower.blocks.2.3.channel_block.ffn.norm.weight": "model-00004-of-00004.safetensors",
477
+ "model.vision_tower.vision_tower.blocks.2.3.spatial_block.conv1.fn.dw.bias": "model-00004-of-00004.safetensors",
478
+ "model.vision_tower.vision_tower.blocks.2.3.spatial_block.conv1.fn.dw.weight": "model-00004-of-00004.safetensors",
479
+ "model.vision_tower.vision_tower.blocks.2.3.spatial_block.conv2.fn.dw.bias": "model-00004-of-00004.safetensors",
480
+ "model.vision_tower.vision_tower.blocks.2.3.spatial_block.conv2.fn.dw.weight": "model-00004-of-00004.safetensors",
481
+ "model.vision_tower.vision_tower.blocks.2.3.spatial_block.ffn.fn.net.fc1.bias": "model-00004-of-00004.safetensors",
482
+ "model.vision_tower.vision_tower.blocks.2.3.spatial_block.ffn.fn.net.fc1.weight": "model-00004-of-00004.safetensors",
483
+ "model.vision_tower.vision_tower.blocks.2.3.spatial_block.ffn.fn.net.fc2.bias": "model-00004-of-00004.safetensors",
484
+ "model.vision_tower.vision_tower.blocks.2.3.spatial_block.ffn.fn.net.fc2.weight": "model-00004-of-00004.safetensors",
485
+ "model.vision_tower.vision_tower.blocks.2.3.spatial_block.ffn.norm.bias": "model-00004-of-00004.safetensors",
486
+ "model.vision_tower.vision_tower.blocks.2.3.spatial_block.ffn.norm.weight": "model-00004-of-00004.safetensors",
487
+ "model.vision_tower.vision_tower.blocks.2.3.spatial_block.window_attn.fn.proj.bias": "model-00004-of-00004.safetensors",
488
+ "model.vision_tower.vision_tower.blocks.2.3.spatial_block.window_attn.fn.proj.weight": "model-00004-of-00004.safetensors",
489
+ "model.vision_tower.vision_tower.blocks.2.3.spatial_block.window_attn.fn.qkv.bias": "model-00004-of-00004.safetensors",
490
+ "model.vision_tower.vision_tower.blocks.2.3.spatial_block.window_attn.fn.qkv.weight": "model-00004-of-00004.safetensors",
491
+ "model.vision_tower.vision_tower.blocks.2.3.spatial_block.window_attn.norm.bias": "model-00004-of-00004.safetensors",
492
+ "model.vision_tower.vision_tower.blocks.2.3.spatial_block.window_attn.norm.weight": "model-00004-of-00004.safetensors",
493
+ "model.vision_tower.vision_tower.blocks.2.4.channel_block.channel_attn.fn.proj.bias": "model-00004-of-00004.safetensors",
494
+ "model.vision_tower.vision_tower.blocks.2.4.channel_block.channel_attn.fn.proj.weight": "model-00004-of-00004.safetensors",
495
+ "model.vision_tower.vision_tower.blocks.2.4.channel_block.channel_attn.fn.qkv.bias": "model-00004-of-00004.safetensors",
496
+ "model.vision_tower.vision_tower.blocks.2.4.channel_block.channel_attn.fn.qkv.weight": "model-00004-of-00004.safetensors",
497
+ "model.vision_tower.vision_tower.blocks.2.4.channel_block.channel_attn.norm.bias": "model-00004-of-00004.safetensors",
498
+ "model.vision_tower.vision_tower.blocks.2.4.channel_block.channel_attn.norm.weight": "model-00004-of-00004.safetensors",
499
+ "model.vision_tower.vision_tower.blocks.2.4.channel_block.conv1.fn.dw.bias": "model-00004-of-00004.safetensors",
500
+ "model.vision_tower.vision_tower.blocks.2.4.channel_block.conv1.fn.dw.weight": "model-00004-of-00004.safetensors",
501
+ "model.vision_tower.vision_tower.blocks.2.4.channel_block.conv2.fn.dw.bias": "model-00004-of-00004.safetensors",
502
+ "model.vision_tower.vision_tower.blocks.2.4.channel_block.conv2.fn.dw.weight": "model-00004-of-00004.safetensors",
503
+ "model.vision_tower.vision_tower.blocks.2.4.channel_block.ffn.fn.net.fc1.bias": "model-00004-of-00004.safetensors",
504
+ "model.vision_tower.vision_tower.blocks.2.4.channel_block.ffn.fn.net.fc1.weight": "model-00004-of-00004.safetensors",
505
+ "model.vision_tower.vision_tower.blocks.2.4.channel_block.ffn.fn.net.fc2.bias": "model-00004-of-00004.safetensors",
506
+ "model.vision_tower.vision_tower.blocks.2.4.channel_block.ffn.fn.net.fc2.weight": "model-00004-of-00004.safetensors",
507
+ "model.vision_tower.vision_tower.blocks.2.4.channel_block.ffn.norm.bias": "model-00004-of-00004.safetensors",
508
+ "model.vision_tower.vision_tower.blocks.2.4.channel_block.ffn.norm.weight": "model-00004-of-00004.safetensors",
509
+ "model.vision_tower.vision_tower.blocks.2.4.spatial_block.conv1.fn.dw.bias": "model-00004-of-00004.safetensors",
510
+ "model.vision_tower.vision_tower.blocks.2.4.spatial_block.conv1.fn.dw.weight": "model-00004-of-00004.safetensors",
511
+ "model.vision_tower.vision_tower.blocks.2.4.spatial_block.conv2.fn.dw.bias": "model-00004-of-00004.safetensors",
512
+ "model.vision_tower.vision_tower.blocks.2.4.spatial_block.conv2.fn.dw.weight": "model-00004-of-00004.safetensors",
513
+ "model.vision_tower.vision_tower.blocks.2.4.spatial_block.ffn.fn.net.fc1.bias": "model-00004-of-00004.safetensors",
514
+ "model.vision_tower.vision_tower.blocks.2.4.spatial_block.ffn.fn.net.fc1.weight": "model-00004-of-00004.safetensors",
515
+ "model.vision_tower.vision_tower.blocks.2.4.spatial_block.ffn.fn.net.fc2.bias": "model-00004-of-00004.safetensors",
516
+ "model.vision_tower.vision_tower.blocks.2.4.spatial_block.ffn.fn.net.fc2.weight": "model-00004-of-00004.safetensors",
517
+ "model.vision_tower.vision_tower.blocks.2.4.spatial_block.ffn.norm.bias": "model-00004-of-00004.safetensors",
518
+ "model.vision_tower.vision_tower.blocks.2.4.spatial_block.ffn.norm.weight": "model-00004-of-00004.safetensors",
519
+ "model.vision_tower.vision_tower.blocks.2.4.spatial_block.window_attn.fn.proj.bias": "model-00004-of-00004.safetensors",
520
+ "model.vision_tower.vision_tower.blocks.2.4.spatial_block.window_attn.fn.proj.weight": "model-00004-of-00004.safetensors",
521
+ "model.vision_tower.vision_tower.blocks.2.4.spatial_block.window_attn.fn.qkv.bias": "model-00004-of-00004.safetensors",
522
+ "model.vision_tower.vision_tower.blocks.2.4.spatial_block.window_attn.fn.qkv.weight": "model-00004-of-00004.safetensors",
523
+ "model.vision_tower.vision_tower.blocks.2.4.spatial_block.window_attn.norm.bias": "model-00004-of-00004.safetensors",
524
+ "model.vision_tower.vision_tower.blocks.2.4.spatial_block.window_attn.norm.weight": "model-00004-of-00004.safetensors",
525
+ "model.vision_tower.vision_tower.blocks.2.5.channel_block.channel_attn.fn.proj.bias": "model-00004-of-00004.safetensors",
526
+ "model.vision_tower.vision_tower.blocks.2.5.channel_block.channel_attn.fn.proj.weight": "model-00004-of-00004.safetensors",
527
+ "model.vision_tower.vision_tower.blocks.2.5.channel_block.channel_attn.fn.qkv.bias": "model-00004-of-00004.safetensors",
528
+ "model.vision_tower.vision_tower.blocks.2.5.channel_block.channel_attn.fn.qkv.weight": "model-00004-of-00004.safetensors",
529
+ "model.vision_tower.vision_tower.blocks.2.5.channel_block.channel_attn.norm.bias": "model-00004-of-00004.safetensors",
530
+ "model.vision_tower.vision_tower.blocks.2.5.channel_block.channel_attn.norm.weight": "model-00004-of-00004.safetensors",
531
+ "model.vision_tower.vision_tower.blocks.2.5.channel_block.conv1.fn.dw.bias": "model-00004-of-00004.safetensors",
532
+ "model.vision_tower.vision_tower.blocks.2.5.channel_block.conv1.fn.dw.weight": "model-00004-of-00004.safetensors",
533
+ "model.vision_tower.vision_tower.blocks.2.5.channel_block.conv2.fn.dw.bias": "model-00004-of-00004.safetensors",
534
+ "model.vision_tower.vision_tower.blocks.2.5.channel_block.conv2.fn.dw.weight": "model-00004-of-00004.safetensors",
535
+ "model.vision_tower.vision_tower.blocks.2.5.channel_block.ffn.fn.net.fc1.bias": "model-00004-of-00004.safetensors",
536
+ "model.vision_tower.vision_tower.blocks.2.5.channel_block.ffn.fn.net.fc1.weight": "model-00004-of-00004.safetensors",
537
+ "model.vision_tower.vision_tower.blocks.2.5.channel_block.ffn.fn.net.fc2.bias": "model-00004-of-00004.safetensors",
538
+ "model.vision_tower.vision_tower.blocks.2.5.channel_block.ffn.fn.net.fc2.weight": "model-00004-of-00004.safetensors",
539
+ "model.vision_tower.vision_tower.blocks.2.5.channel_block.ffn.norm.bias": "model-00004-of-00004.safetensors",
540
+ "model.vision_tower.vision_tower.blocks.2.5.channel_block.ffn.norm.weight": "model-00004-of-00004.safetensors",
541
+ "model.vision_tower.vision_tower.blocks.2.5.spatial_block.conv1.fn.dw.bias": "model-00004-of-00004.safetensors",
542
+ "model.vision_tower.vision_tower.blocks.2.5.spatial_block.conv1.fn.dw.weight": "model-00004-of-00004.safetensors",
543
+ "model.vision_tower.vision_tower.blocks.2.5.spatial_block.conv2.fn.dw.bias": "model-00004-of-00004.safetensors",
544
+ "model.vision_tower.vision_tower.blocks.2.5.spatial_block.conv2.fn.dw.weight": "model-00004-of-00004.safetensors",
545
+ "model.vision_tower.vision_tower.blocks.2.5.spatial_block.ffn.fn.net.fc1.bias": "model-00004-of-00004.safetensors",
546
+ "model.vision_tower.vision_tower.blocks.2.5.spatial_block.ffn.fn.net.fc1.weight": "model-00004-of-00004.safetensors",
547
+ "model.vision_tower.vision_tower.blocks.2.5.spatial_block.ffn.fn.net.fc2.bias": "model-00004-of-00004.safetensors",
548
+ "model.vision_tower.vision_tower.blocks.2.5.spatial_block.ffn.fn.net.fc2.weight": "model-00004-of-00004.safetensors",
549
+ "model.vision_tower.vision_tower.blocks.2.5.spatial_block.ffn.norm.bias": "model-00004-of-00004.safetensors",
550
+ "model.vision_tower.vision_tower.blocks.2.5.spatial_block.ffn.norm.weight": "model-00004-of-00004.safetensors",
551
+ "model.vision_tower.vision_tower.blocks.2.5.spatial_block.window_attn.fn.proj.bias": "model-00004-of-00004.safetensors",
552
+ "model.vision_tower.vision_tower.blocks.2.5.spatial_block.window_attn.fn.proj.weight": "model-00004-of-00004.safetensors",
553
+ "model.vision_tower.vision_tower.blocks.2.5.spatial_block.window_attn.fn.qkv.bias": "model-00004-of-00004.safetensors",
554
+ "model.vision_tower.vision_tower.blocks.2.5.spatial_block.window_attn.fn.qkv.weight": "model-00004-of-00004.safetensors",
555
+ "model.vision_tower.vision_tower.blocks.2.5.spatial_block.window_attn.norm.bias": "model-00004-of-00004.safetensors",
556
+ "model.vision_tower.vision_tower.blocks.2.5.spatial_block.window_attn.norm.weight": "model-00004-of-00004.safetensors",
557
+ "model.vision_tower.vision_tower.blocks.2.6.channel_block.channel_attn.fn.proj.bias": "model-00004-of-00004.safetensors",
558
+ "model.vision_tower.vision_tower.blocks.2.6.channel_block.channel_attn.fn.proj.weight": "model-00004-of-00004.safetensors",
559
+ "model.vision_tower.vision_tower.blocks.2.6.channel_block.channel_attn.fn.qkv.bias": "model-00004-of-00004.safetensors",
560
+ "model.vision_tower.vision_tower.blocks.2.6.channel_block.channel_attn.fn.qkv.weight": "model-00004-of-00004.safetensors",
561
+ "model.vision_tower.vision_tower.blocks.2.6.channel_block.channel_attn.norm.bias": "model-00004-of-00004.safetensors",
562
+ "model.vision_tower.vision_tower.blocks.2.6.channel_block.channel_attn.norm.weight": "model-00004-of-00004.safetensors",
563
+ "model.vision_tower.vision_tower.blocks.2.6.channel_block.conv1.fn.dw.bias": "model-00004-of-00004.safetensors",
564
+ "model.vision_tower.vision_tower.blocks.2.6.channel_block.conv1.fn.dw.weight": "model-00004-of-00004.safetensors",
565
+ "model.vision_tower.vision_tower.blocks.2.6.channel_block.conv2.fn.dw.bias": "model-00004-of-00004.safetensors",
566
+ "model.vision_tower.vision_tower.blocks.2.6.channel_block.conv2.fn.dw.weight": "model-00004-of-00004.safetensors",
567
+ "model.vision_tower.vision_tower.blocks.2.6.channel_block.ffn.fn.net.fc1.bias": "model-00004-of-00004.safetensors",
568
+ "model.vision_tower.vision_tower.blocks.2.6.channel_block.ffn.fn.net.fc1.weight": "model-00004-of-00004.safetensors",
569
+ "model.vision_tower.vision_tower.blocks.2.6.channel_block.ffn.fn.net.fc2.bias": "model-00004-of-00004.safetensors",
570
+ "model.vision_tower.vision_tower.blocks.2.6.channel_block.ffn.fn.net.fc2.weight": "model-00004-of-00004.safetensors",
571
+ "model.vision_tower.vision_tower.blocks.2.6.channel_block.ffn.norm.bias": "model-00004-of-00004.safetensors",
572
+ "model.vision_tower.vision_tower.blocks.2.6.channel_block.ffn.norm.weight": "model-00004-of-00004.safetensors",
573
+ "model.vision_tower.vision_tower.blocks.2.6.spatial_block.conv1.fn.dw.bias": "model-00004-of-00004.safetensors",
574
+ "model.vision_tower.vision_tower.blocks.2.6.spatial_block.conv1.fn.dw.weight": "model-00004-of-00004.safetensors",
575
+ "model.vision_tower.vision_tower.blocks.2.6.spatial_block.conv2.fn.dw.bias": "model-00004-of-00004.safetensors",
576
+ "model.vision_tower.vision_tower.blocks.2.6.spatial_block.conv2.fn.dw.weight": "model-00004-of-00004.safetensors",
577
+ "model.vision_tower.vision_tower.blocks.2.6.spatial_block.ffn.fn.net.fc1.bias": "model-00004-of-00004.safetensors",
578
+ "model.vision_tower.vision_tower.blocks.2.6.spatial_block.ffn.fn.net.fc1.weight": "model-00004-of-00004.safetensors",
579
+ "model.vision_tower.vision_tower.blocks.2.6.spatial_block.ffn.fn.net.fc2.bias": "model-00004-of-00004.safetensors",
580
+ "model.vision_tower.vision_tower.blocks.2.6.spatial_block.ffn.fn.net.fc2.weight": "model-00004-of-00004.safetensors",
581
+ "model.vision_tower.vision_tower.blocks.2.6.spatial_block.ffn.norm.bias": "model-00004-of-00004.safetensors",
582
+ "model.vision_tower.vision_tower.blocks.2.6.spatial_block.ffn.norm.weight": "model-00004-of-00004.safetensors",
583
+ "model.vision_tower.vision_tower.blocks.2.6.spatial_block.window_attn.fn.proj.bias": "model-00004-of-00004.safetensors",
584
+ "model.vision_tower.vision_tower.blocks.2.6.spatial_block.window_attn.fn.proj.weight": "model-00004-of-00004.safetensors",
585
+ "model.vision_tower.vision_tower.blocks.2.6.spatial_block.window_attn.fn.qkv.bias": "model-00004-of-00004.safetensors",
586
+ "model.vision_tower.vision_tower.blocks.2.6.spatial_block.window_attn.fn.qkv.weight": "model-00004-of-00004.safetensors",
587
+ "model.vision_tower.vision_tower.blocks.2.6.spatial_block.window_attn.norm.bias": "model-00004-of-00004.safetensors",
588
+ "model.vision_tower.vision_tower.blocks.2.6.spatial_block.window_attn.norm.weight": "model-00004-of-00004.safetensors",
589
+ "model.vision_tower.vision_tower.blocks.2.7.channel_block.channel_attn.fn.proj.bias": "model-00004-of-00004.safetensors",
590
+ "model.vision_tower.vision_tower.blocks.2.7.channel_block.channel_attn.fn.proj.weight": "model-00004-of-00004.safetensors",
591
+ "model.vision_tower.vision_tower.blocks.2.7.channel_block.channel_attn.fn.qkv.bias": "model-00004-of-00004.safetensors",
592
+ "model.vision_tower.vision_tower.blocks.2.7.channel_block.channel_attn.fn.qkv.weight": "model-00004-of-00004.safetensors",
593
+ "model.vision_tower.vision_tower.blocks.2.7.channel_block.channel_attn.norm.bias": "model-00004-of-00004.safetensors",
594
+ "model.vision_tower.vision_tower.blocks.2.7.channel_block.channel_attn.norm.weight": "model-00004-of-00004.safetensors",
595
+ "model.vision_tower.vision_tower.blocks.2.7.channel_block.conv1.fn.dw.bias": "model-00004-of-00004.safetensors",
596
+ "model.vision_tower.vision_tower.blocks.2.7.channel_block.conv1.fn.dw.weight": "model-00004-of-00004.safetensors",
597
+ "model.vision_tower.vision_tower.blocks.2.7.channel_block.conv2.fn.dw.bias": "model-00004-of-00004.safetensors",
598
+ "model.vision_tower.vision_tower.blocks.2.7.channel_block.conv2.fn.dw.weight": "model-00004-of-00004.safetensors",
599
+ "model.vision_tower.vision_tower.blocks.2.7.channel_block.ffn.fn.net.fc1.bias": "model-00004-of-00004.safetensors",
600
+ "model.vision_tower.vision_tower.blocks.2.7.channel_block.ffn.fn.net.fc1.weight": "model-00004-of-00004.safetensors",
601
+ "model.vision_tower.vision_tower.blocks.2.7.channel_block.ffn.fn.net.fc2.bias": "model-00004-of-00004.safetensors",
602
+ "model.vision_tower.vision_tower.blocks.2.7.channel_block.ffn.fn.net.fc2.weight": "model-00004-of-00004.safetensors",
603
+ "model.vision_tower.vision_tower.blocks.2.7.channel_block.ffn.norm.bias": "model-00004-of-00004.safetensors",
604
+ "model.vision_tower.vision_tower.blocks.2.7.channel_block.ffn.norm.weight": "model-00004-of-00004.safetensors",
605
+ "model.vision_tower.vision_tower.blocks.2.7.spatial_block.conv1.fn.dw.bias": "model-00004-of-00004.safetensors",
606
+ "model.vision_tower.vision_tower.blocks.2.7.spatial_block.conv1.fn.dw.weight": "model-00004-of-00004.safetensors",
607
+ "model.vision_tower.vision_tower.blocks.2.7.spatial_block.conv2.fn.dw.bias": "model-00004-of-00004.safetensors",
608
+ "model.vision_tower.vision_tower.blocks.2.7.spatial_block.conv2.fn.dw.weight": "model-00004-of-00004.safetensors",
609
+ "model.vision_tower.vision_tower.blocks.2.7.spatial_block.ffn.fn.net.fc1.bias": "model-00004-of-00004.safetensors",
610
+ "model.vision_tower.vision_tower.blocks.2.7.spatial_block.ffn.fn.net.fc1.weight": "model-00004-of-00004.safetensors",
611
+ "model.vision_tower.vision_tower.blocks.2.7.spatial_block.ffn.fn.net.fc2.bias": "model-00004-of-00004.safetensors",
612
+ "model.vision_tower.vision_tower.blocks.2.7.spatial_block.ffn.fn.net.fc2.weight": "model-00004-of-00004.safetensors",
613
+ "model.vision_tower.vision_tower.blocks.2.7.spatial_block.ffn.norm.bias": "model-00004-of-00004.safetensors",
614
+ "model.vision_tower.vision_tower.blocks.2.7.spatial_block.ffn.norm.weight": "model-00004-of-00004.safetensors",
615
+ "model.vision_tower.vision_tower.blocks.2.7.spatial_block.window_attn.fn.proj.bias": "model-00004-of-00004.safetensors",
616
+ "model.vision_tower.vision_tower.blocks.2.7.spatial_block.window_attn.fn.proj.weight": "model-00004-of-00004.safetensors",
617
+ "model.vision_tower.vision_tower.blocks.2.7.spatial_block.window_attn.fn.qkv.bias": "model-00004-of-00004.safetensors",
618
+ "model.vision_tower.vision_tower.blocks.2.7.spatial_block.window_attn.fn.qkv.weight": "model-00004-of-00004.safetensors",
619
+ "model.vision_tower.vision_tower.blocks.2.7.spatial_block.window_attn.norm.bias": "model-00004-of-00004.safetensors",
620
+ "model.vision_tower.vision_tower.blocks.2.7.spatial_block.window_attn.norm.weight": "model-00004-of-00004.safetensors",
621
+ "model.vision_tower.vision_tower.blocks.2.8.channel_block.channel_attn.fn.proj.bias": "model-00004-of-00004.safetensors",
622
+ "model.vision_tower.vision_tower.blocks.2.8.channel_block.channel_attn.fn.proj.weight": "model-00004-of-00004.safetensors",
623
+ "model.vision_tower.vision_tower.blocks.2.8.channel_block.channel_attn.fn.qkv.bias": "model-00004-of-00004.safetensors",
624
+ "model.vision_tower.vision_tower.blocks.2.8.channel_block.channel_attn.fn.qkv.weight": "model-00004-of-00004.safetensors",
625
+ "model.vision_tower.vision_tower.blocks.2.8.channel_block.channel_attn.norm.bias": "model-00004-of-00004.safetensors",
626
+ "model.vision_tower.vision_tower.blocks.2.8.channel_block.channel_attn.norm.weight": "model-00004-of-00004.safetensors",
627
+ "model.vision_tower.vision_tower.blocks.2.8.channel_block.conv1.fn.dw.bias": "model-00004-of-00004.safetensors",
628
+ "model.vision_tower.vision_tower.blocks.2.8.channel_block.conv1.fn.dw.weight": "model-00004-of-00004.safetensors",
629
+ "model.vision_tower.vision_tower.blocks.2.8.channel_block.conv2.fn.dw.bias": "model-00004-of-00004.safetensors",
630
+ "model.vision_tower.vision_tower.blocks.2.8.channel_block.conv2.fn.dw.weight": "model-00004-of-00004.safetensors",
631
+ "model.vision_tower.vision_tower.blocks.2.8.channel_block.ffn.fn.net.fc1.bias": "model-00004-of-00004.safetensors",
632
+ "model.vision_tower.vision_tower.blocks.2.8.channel_block.ffn.fn.net.fc1.weight": "model-00004-of-00004.safetensors",
633
+ "model.vision_tower.vision_tower.blocks.2.8.channel_block.ffn.fn.net.fc2.bias": "model-00004-of-00004.safetensors",
634
+ "model.vision_tower.vision_tower.blocks.2.8.channel_block.ffn.fn.net.fc2.weight": "model-00004-of-00004.safetensors",
635
+ "model.vision_tower.vision_tower.blocks.2.8.channel_block.ffn.norm.bias": "model-00004-of-00004.safetensors",
636
+ "model.vision_tower.vision_tower.blocks.2.8.channel_block.ffn.norm.weight": "model-00004-of-00004.safetensors",
637
+ "model.vision_tower.vision_tower.blocks.2.8.spatial_block.conv1.fn.dw.bias": "model-00004-of-00004.safetensors",
638
+ "model.vision_tower.vision_tower.blocks.2.8.spatial_block.conv1.fn.dw.weight": "model-00004-of-00004.safetensors",
639
+ "model.vision_tower.vision_tower.blocks.2.8.spatial_block.conv2.fn.dw.bias": "model-00004-of-00004.safetensors",
640
+ "model.vision_tower.vision_tower.blocks.2.8.spatial_block.conv2.fn.dw.weight": "model-00004-of-00004.safetensors",
641
+ "model.vision_tower.vision_tower.blocks.2.8.spatial_block.ffn.fn.net.fc1.bias": "model-00004-of-00004.safetensors",
642
+ "model.vision_tower.vision_tower.blocks.2.8.spatial_block.ffn.fn.net.fc1.weight": "model-00004-of-00004.safetensors",
643
+ "model.vision_tower.vision_tower.blocks.2.8.spatial_block.ffn.fn.net.fc2.bias": "model-00004-of-00004.safetensors",
644
+ "model.vision_tower.vision_tower.blocks.2.8.spatial_block.ffn.fn.net.fc2.weight": "model-00004-of-00004.safetensors",
645
+ "model.vision_tower.vision_tower.blocks.2.8.spatial_block.ffn.norm.bias": "model-00004-of-00004.safetensors",
646
+ "model.vision_tower.vision_tower.blocks.2.8.spatial_block.ffn.norm.weight": "model-00004-of-00004.safetensors",
647
+ "model.vision_tower.vision_tower.blocks.2.8.spatial_block.window_attn.fn.proj.bias": "model-00004-of-00004.safetensors",
648
+ "model.vision_tower.vision_tower.blocks.2.8.spatial_block.window_attn.fn.proj.weight": "model-00004-of-00004.safetensors",
649
+ "model.vision_tower.vision_tower.blocks.2.8.spatial_block.window_attn.fn.qkv.bias": "model-00004-of-00004.safetensors",
650
+ "model.vision_tower.vision_tower.blocks.2.8.spatial_block.window_attn.fn.qkv.weight": "model-00004-of-00004.safetensors",
651
+ "model.vision_tower.vision_tower.blocks.2.8.spatial_block.window_attn.norm.bias": "model-00004-of-00004.safetensors",
652
+ "model.vision_tower.vision_tower.blocks.2.8.spatial_block.window_attn.norm.weight": "model-00004-of-00004.safetensors",
653
+ "model.vision_tower.vision_tower.blocks.3.0.channel_block.channel_attn.fn.proj.bias": "model-00004-of-00004.safetensors",
654
+ "model.vision_tower.vision_tower.blocks.3.0.channel_block.channel_attn.fn.proj.weight": "model-00004-of-00004.safetensors",
655
+ "model.vision_tower.vision_tower.blocks.3.0.channel_block.channel_attn.fn.qkv.bias": "model-00004-of-00004.safetensors",
656
+ "model.vision_tower.vision_tower.blocks.3.0.channel_block.channel_attn.fn.qkv.weight": "model-00004-of-00004.safetensors",
657
+ "model.vision_tower.vision_tower.blocks.3.0.channel_block.channel_attn.norm.bias": "model-00004-of-00004.safetensors",
658
+ "model.vision_tower.vision_tower.blocks.3.0.channel_block.channel_attn.norm.weight": "model-00004-of-00004.safetensors",
659
+ "model.vision_tower.vision_tower.blocks.3.0.channel_block.conv1.fn.dw.bias": "model-00004-of-00004.safetensors",
660
+ "model.vision_tower.vision_tower.blocks.3.0.channel_block.conv1.fn.dw.weight": "model-00004-of-00004.safetensors",
661
+ "model.vision_tower.vision_tower.blocks.3.0.channel_block.conv2.fn.dw.bias": "model-00004-of-00004.safetensors",
662
+ "model.vision_tower.vision_tower.blocks.3.0.channel_block.conv2.fn.dw.weight": "model-00004-of-00004.safetensors",
663
+ "model.vision_tower.vision_tower.blocks.3.0.channel_block.ffn.fn.net.fc1.bias": "model-00004-of-00004.safetensors",
664
+ "model.vision_tower.vision_tower.blocks.3.0.channel_block.ffn.fn.net.fc1.weight": "model-00004-of-00004.safetensors",
665
+ "model.vision_tower.vision_tower.blocks.3.0.channel_block.ffn.fn.net.fc2.bias": "model-00004-of-00004.safetensors",
666
+ "model.vision_tower.vision_tower.blocks.3.0.channel_block.ffn.fn.net.fc2.weight": "model-00004-of-00004.safetensors",
667
+ "model.vision_tower.vision_tower.blocks.3.0.channel_block.ffn.norm.bias": "model-00004-of-00004.safetensors",
668
+ "model.vision_tower.vision_tower.blocks.3.0.channel_block.ffn.norm.weight": "model-00004-of-00004.safetensors",
669
+ "model.vision_tower.vision_tower.blocks.3.0.spatial_block.conv1.fn.dw.bias": "model-00004-of-00004.safetensors",
670
+ "model.vision_tower.vision_tower.blocks.3.0.spatial_block.conv1.fn.dw.weight": "model-00004-of-00004.safetensors",
671
+ "model.vision_tower.vision_tower.blocks.3.0.spatial_block.conv2.fn.dw.bias": "model-00004-of-00004.safetensors",
672
+ "model.vision_tower.vision_tower.blocks.3.0.spatial_block.conv2.fn.dw.weight": "model-00004-of-00004.safetensors",
673
+ "model.vision_tower.vision_tower.blocks.3.0.spatial_block.ffn.fn.net.fc1.bias": "model-00004-of-00004.safetensors",
674
+ "model.vision_tower.vision_tower.blocks.3.0.spatial_block.ffn.fn.net.fc1.weight": "model-00004-of-00004.safetensors",
675
+ "model.vision_tower.vision_tower.blocks.3.0.spatial_block.ffn.fn.net.fc2.bias": "model-00004-of-00004.safetensors",
676
+ "model.vision_tower.vision_tower.blocks.3.0.spatial_block.ffn.fn.net.fc2.weight": "model-00004-of-00004.safetensors",
677
+ "model.vision_tower.vision_tower.blocks.3.0.spatial_block.ffn.norm.bias": "model-00004-of-00004.safetensors",
678
+ "model.vision_tower.vision_tower.blocks.3.0.spatial_block.ffn.norm.weight": "model-00004-of-00004.safetensors",
679
+ "model.vision_tower.vision_tower.blocks.3.0.spatial_block.window_attn.fn.proj.bias": "model-00004-of-00004.safetensors",
680
+ "model.vision_tower.vision_tower.blocks.3.0.spatial_block.window_attn.fn.proj.weight": "model-00004-of-00004.safetensors",
681
+ "model.vision_tower.vision_tower.blocks.3.0.spatial_block.window_attn.fn.qkv.bias": "model-00004-of-00004.safetensors",
682
+ "model.vision_tower.vision_tower.blocks.3.0.spatial_block.window_attn.fn.qkv.weight": "model-00004-of-00004.safetensors",
683
+ "model.vision_tower.vision_tower.blocks.3.0.spatial_block.window_attn.norm.bias": "model-00004-of-00004.safetensors",
684
+ "model.vision_tower.vision_tower.blocks.3.0.spatial_block.window_attn.norm.weight": "model-00004-of-00004.safetensors",
685
+ "model.vision_tower.vision_tower.convs.0.norm.bias": "model-00004-of-00004.safetensors",
686
+ "model.vision_tower.vision_tower.convs.0.norm.weight": "model-00004-of-00004.safetensors",
687
+ "model.vision_tower.vision_tower.convs.0.proj.bias": "model-00004-of-00004.safetensors",
688
+ "model.vision_tower.vision_tower.convs.0.proj.weight": "model-00004-of-00004.safetensors",
689
+ "model.vision_tower.vision_tower.convs.1.norm.bias": "model-00004-of-00004.safetensors",
690
+ "model.vision_tower.vision_tower.convs.1.norm.weight": "model-00004-of-00004.safetensors",
691
+ "model.vision_tower.vision_tower.convs.1.proj.bias": "model-00004-of-00004.safetensors",
692
+ "model.vision_tower.vision_tower.convs.1.proj.weight": "model-00004-of-00004.safetensors",
693
+ "model.vision_tower.vision_tower.convs.2.norm.bias": "model-00004-of-00004.safetensors",
694
+ "model.vision_tower.vision_tower.convs.2.norm.weight": "model-00004-of-00004.safetensors",
695
+ "model.vision_tower.vision_tower.convs.2.proj.bias": "model-00004-of-00004.safetensors",
696
+ "model.vision_tower.vision_tower.convs.2.proj.weight": "model-00004-of-00004.safetensors",
697
+ "model.vision_tower.vision_tower.convs.3.norm.bias": "model-00004-of-00004.safetensors",
698
+ "model.vision_tower.vision_tower.convs.3.norm.weight": "model-00004-of-00004.safetensors",
699
+ "model.vision_tower.vision_tower.convs.3.proj.bias": "model-00004-of-00004.safetensors",
700
+ "model.vision_tower.vision_tower.convs.3.proj.weight": "model-00004-of-00004.safetensors"
701
+ }
702
+ }
modeling_mixsense_llama.py ADDED
@@ -0,0 +1,1233 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import torch
2
+ import torch.nn as nn
3
+ from transformers import AutoConfig, CLIPImageProcessor
4
+ import os
5
+ from einops import rearrange
6
+ from collections import OrderedDict
7
+ from timm.models.layers import DropPath, trunc_normal_
8
+ import torch.nn.functional as F
9
+
10
+
11
+
12
+
13
+ """
14
+ DaViT : Copied from https://huggingface.co/microsoft/Florence-2-large/blob/main/modeling_florence2.py
15
+ """
16
+
17
+
18
+ class MySequential(nn.Sequential):
19
+ def forward(self, *inputs):
20
+ for module in self._modules.values():
21
+ if type(inputs) == tuple:
22
+ inputs = module(*inputs)
23
+ else:
24
+ inputs = module(inputs)
25
+ return inputs
26
+
27
+
28
+ class PreNorm(nn.Module):
29
+ def __init__(self, norm, fn, drop_path=None):
30
+ super().__init__()
31
+ self.norm = norm
32
+ self.fn = fn
33
+ self.drop_path = drop_path
34
+
35
+ def forward(self, x, *args, **kwargs):
36
+ shortcut = x
37
+ if self.norm != None:
38
+ x, size = self.fn(self.norm(x), *args, **kwargs)
39
+ else:
40
+ x, size = self.fn(x, *args, **kwargs)
41
+
42
+ if self.drop_path:
43
+ x = self.drop_path(x)
44
+
45
+ x = shortcut + x
46
+
47
+ return x, size
48
+
49
+
50
+ class Mlp(nn.Module):
51
+ def __init__(
52
+ self,
53
+ in_features,
54
+ hidden_features=None,
55
+ out_features=None,
56
+ act_layer=nn.GELU,
57
+ ):
58
+ super().__init__()
59
+ out_features = out_features or in_features
60
+ hidden_features = hidden_features or in_features
61
+ self.net = nn.Sequential(
62
+ OrderedDict(
63
+ [
64
+ ("fc1", nn.Linear(in_features, hidden_features)),
65
+ ("act", act_layer()),
66
+ ("fc2", nn.Linear(hidden_features, out_features)),
67
+ ]
68
+ )
69
+ )
70
+
71
+ def forward(self, x, size):
72
+ return self.net(x), size
73
+
74
+
75
+ class DepthWiseConv2d(nn.Module):
76
+ def __init__(
77
+ self,
78
+ dim_in,
79
+ kernel_size,
80
+ padding,
81
+ stride,
82
+ bias=True,
83
+ ):
84
+ super().__init__()
85
+ self.dw = nn.Conv2d(
86
+ dim_in,
87
+ dim_in,
88
+ kernel_size=kernel_size,
89
+ padding=padding,
90
+ groups=dim_in,
91
+ stride=stride,
92
+ bias=bias,
93
+ )
94
+
95
+ def forward(self, x, size):
96
+ B, N, C = x.shape
97
+ H, W = size
98
+ assert N == H * W
99
+
100
+ x = self.dw(x.transpose(1, 2).view(B, C, H, W))
101
+ size = (x.size(-2), x.size(-1))
102
+ x = x.flatten(2).transpose(1, 2)
103
+ return x, size
104
+
105
+
106
+ class ConvEmbed(nn.Module):
107
+ """Image to Patch Embedding"""
108
+
109
+ def __init__(
110
+ self,
111
+ patch_size=7,
112
+ in_chans=3,
113
+ embed_dim=64,
114
+ stride=4,
115
+ padding=2,
116
+ norm_layer=None,
117
+ pre_norm=True,
118
+ ):
119
+ super().__init__()
120
+ self.patch_size = patch_size
121
+
122
+ self.proj = nn.Conv2d(
123
+ in_chans, embed_dim, kernel_size=patch_size, stride=stride, padding=padding
124
+ )
125
+
126
+ dim_norm = in_chans if pre_norm else embed_dim
127
+ self.norm = norm_layer(dim_norm) if norm_layer else None
128
+
129
+ self.pre_norm = pre_norm
130
+
131
+ def forward(self, x, size):
132
+ H, W = size
133
+ if len(x.size()) == 3:
134
+ if self.norm and self.pre_norm:
135
+ x = self.norm(x)
136
+ x = rearrange(x, "b (h w) c -> b c h w", h=H, w=W)
137
+
138
+ x = self.proj(x)
139
+
140
+ _, _, H, W = x.shape
141
+ x = rearrange(x, "b c h w -> b (h w) c")
142
+ if self.norm and not self.pre_norm:
143
+ x = self.norm(x)
144
+
145
+ return x, (H, W)
146
+
147
+
148
+ class ChannelAttention(nn.Module):
149
+
150
+ def __init__(self, dim, groups=8, qkv_bias=True):
151
+ super().__init__()
152
+
153
+ self.groups = groups
154
+ self.qkv = nn.Linear(dim, dim * 3, bias=qkv_bias)
155
+ self.proj = nn.Linear(dim, dim)
156
+
157
+ def forward(self, x, size):
158
+ B, N, C = x.shape
159
+
160
+ qkv = (
161
+ self.qkv(x)
162
+ .reshape(B, N, 3, self.groups, C // self.groups)
163
+ .permute(2, 0, 3, 1, 4)
164
+ )
165
+ q, k, v = qkv[0], qkv[1], qkv[2]
166
+
167
+ q = q * (float(N) ** -0.5)
168
+ attention = q.transpose(-1, -2) @ k
169
+ attention = attention.softmax(dim=-1)
170
+ x = (attention @ v.transpose(-1, -2)).transpose(-1, -2)
171
+ x = x.transpose(1, 2).reshape(B, N, C)
172
+ x = self.proj(x)
173
+ return x, size
174
+
175
+
176
+ class ChannelBlock(nn.Module):
177
+
178
+ def __init__(
179
+ self,
180
+ dim,
181
+ groups,
182
+ mlp_ratio=4.0,
183
+ qkv_bias=True,
184
+ drop_path_rate=0.0,
185
+ act_layer=nn.GELU,
186
+ norm_layer=nn.LayerNorm,
187
+ conv_at_attn=True,
188
+ conv_at_ffn=True,
189
+ ):
190
+ super().__init__()
191
+
192
+ drop_path = DropPath(drop_path_rate) if drop_path_rate > 0.0 else nn.Identity()
193
+
194
+ self.conv1 = (
195
+ PreNorm(None, DepthWiseConv2d(dim, 3, 1, 1)) if conv_at_attn else None
196
+ )
197
+ self.channel_attn = PreNorm(
198
+ norm_layer(dim),
199
+ ChannelAttention(dim, groups=groups, qkv_bias=qkv_bias),
200
+ drop_path,
201
+ )
202
+ self.conv2 = (
203
+ PreNorm(None, DepthWiseConv2d(dim, 3, 1, 1)) if conv_at_ffn else None
204
+ )
205
+ self.ffn = PreNorm(
206
+ norm_layer(dim),
207
+ Mlp(
208
+ in_features=dim,
209
+ hidden_features=int(dim * mlp_ratio),
210
+ act_layer=act_layer,
211
+ ),
212
+ drop_path,
213
+ )
214
+
215
+ def forward(self, x, size):
216
+ if self.conv1:
217
+ x, size = self.conv1(x, size)
218
+ x, size = self.channel_attn(x, size)
219
+
220
+ if self.conv2:
221
+ x, size = self.conv2(x, size)
222
+ x, size = self.ffn(x, size)
223
+
224
+ return x, size
225
+
226
+
227
+ def window_partition(x, window_size: int):
228
+ B, H, W, C = x.shape
229
+ x = x.view(B, H // window_size, window_size, W // window_size, window_size, C)
230
+ windows = (
231
+ x.permute(0, 1, 3, 2, 4, 5).contiguous().view(-1, window_size, window_size, C)
232
+ )
233
+ return windows
234
+
235
+
236
+ def window_reverse(windows, batch_size: int, window_size: int, H: int, W: int):
237
+ B = batch_size
238
+ # this will cause onnx conversion failed for dynamic axis, because treated as constant
239
+ # int(windows.shape[0] / (H * W / window_size / window_size))
240
+ x = windows.view(
241
+ B, H // window_size, W // window_size, window_size, window_size, -1
242
+ )
243
+ x = x.permute(0, 1, 3, 2, 4, 5).contiguous().view(B, H, W, -1)
244
+ return x
245
+
246
+
247
+ class WindowAttention(nn.Module):
248
+ def __init__(self, dim, num_heads, window_size, qkv_bias=True):
249
+
250
+ super().__init__()
251
+ self.dim = dim
252
+ self.window_size = window_size
253
+ self.num_heads = num_heads
254
+ head_dim = dim // num_heads
255
+ self.scale = float(head_dim) ** -0.5
256
+
257
+ self.qkv = nn.Linear(dim, dim * 3, bias=qkv_bias)
258
+ self.proj = nn.Linear(dim, dim)
259
+
260
+ self.softmax = nn.Softmax(dim=-1)
261
+
262
+ def forward(self, x, size):
263
+
264
+ H, W = size
265
+ B, L, C = x.shape
266
+ assert L == H * W, "input feature has wrong size"
267
+
268
+ x = x.view(B, H, W, C)
269
+
270
+ pad_l = pad_t = 0
271
+ pad_r = (self.window_size - W % self.window_size) % self.window_size
272
+ pad_b = (self.window_size - H % self.window_size) % self.window_size
273
+ x = F.pad(x, (0, 0, pad_l, pad_r, pad_t, pad_b))
274
+ _, Hp, Wp, _ = x.shape
275
+
276
+ x = window_partition(x, self.window_size)
277
+ x = x.view(-1, self.window_size * self.window_size, C)
278
+
279
+ # W-MSA/SW-MSA
280
+ # attn_windows = self.attn(x_windows)
281
+
282
+ B_, N, C = x.shape
283
+ qkv = (
284
+ self.qkv(x)
285
+ .reshape(B_, N, 3, self.num_heads, C // self.num_heads)
286
+ .permute(2, 0, 3, 1, 4)
287
+ )
288
+ q, k, v = qkv[0], qkv[1], qkv[2]
289
+
290
+ q = q * self.scale
291
+ attn = q @ k.transpose(-2, -1)
292
+ attn = self.softmax(attn)
293
+
294
+ x = (attn @ v).transpose(1, 2).reshape(B_, N, C)
295
+ x = self.proj(x)
296
+
297
+ # merge windows
298
+ x = x.view(-1, self.window_size, self.window_size, C)
299
+ x = window_reverse(x, B, self.window_size, Hp, Wp)
300
+
301
+ if pad_r > 0 or pad_b > 0:
302
+ x = x[:, :H, :W, :].contiguous()
303
+
304
+ x = x.view(B, H * W, C)
305
+
306
+ return x, size
307
+
308
+
309
+ class SpatialBlock(nn.Module):
310
+
311
+ def __init__(
312
+ self,
313
+ dim,
314
+ num_heads,
315
+ window_size,
316
+ mlp_ratio=4.0,
317
+ qkv_bias=True,
318
+ drop_path_rate=0.0,
319
+ act_layer=nn.GELU,
320
+ norm_layer=nn.LayerNorm,
321
+ conv_at_attn=True,
322
+ conv_at_ffn=True,
323
+ ):
324
+ super().__init__()
325
+
326
+ drop_path = DropPath(drop_path_rate) if drop_path_rate > 0.0 else nn.Identity()
327
+
328
+ self.conv1 = (
329
+ PreNorm(None, DepthWiseConv2d(dim, 3, 1, 1)) if conv_at_attn else None
330
+ )
331
+ self.window_attn = PreNorm(
332
+ norm_layer(dim),
333
+ WindowAttention(dim, num_heads, window_size, qkv_bias=qkv_bias),
334
+ drop_path,
335
+ )
336
+ self.conv2 = (
337
+ PreNorm(None, DepthWiseConv2d(dim, 3, 1, 1)) if conv_at_ffn else None
338
+ )
339
+ self.ffn = PreNorm(
340
+ norm_layer(dim),
341
+ Mlp(
342
+ in_features=dim,
343
+ hidden_features=int(dim * mlp_ratio),
344
+ act_layer=act_layer,
345
+ ),
346
+ drop_path,
347
+ )
348
+
349
+ def forward(self, x, size):
350
+ if self.conv1:
351
+ x, size = self.conv1(x, size)
352
+ x, size = self.window_attn(x, size)
353
+
354
+ if self.conv2:
355
+ x, size = self.conv2(x, size)
356
+ x, size = self.ffn(x, size)
357
+ return x, size
358
+
359
+
360
+ class DaViT(nn.Module):
361
+ """DaViT: Dual-Attention Transformer
362
+
363
+ Args:
364
+ in_chans (int): Number of input image channels. Default: 3.
365
+ num_classes (int): Number of classes for classification head. Default: 1000.
366
+ patch_size (tuple(int)): Patch size of convolution in different stages. Default: (7, 2, 2, 2).
367
+ patch_stride (tuple(int)): Patch stride of convolution in different stages. Default: (4, 2, 2, 2).
368
+ patch_padding (tuple(int)): Patch padding of convolution in different stages. Default: (3, 0, 0, 0).
369
+ patch_prenorm (tuple(bool)): If True, perform norm before convlution layer. Default: (True, False, False, False).
370
+ embed_dims (tuple(int)): Patch embedding dimension in different stages. Default: (64, 128, 192, 256).
371
+ num_heads (tuple(int)): Number of spatial attention heads in different stages. Default: (4, 8, 12, 16).
372
+ num_groups (tuple(int)): Number of channel groups in different stages. Default: (4, 8, 12, 16).
373
+ window_size (int): Window size. Default: 7.
374
+ mlp_ratio (float): Ratio of mlp hidden dim to embedding dim. Default: 4.
375
+ qkv_bias (bool): If True, add a learnable bias to query, key, value. Default: True.
376
+ drop_path_rate (float): Stochastic depth rate. Default: 0.1.
377
+ norm_layer (nn.Module): Normalization layer. Default: nn.LayerNorm.
378
+ enable_checkpoint (bool): If True, enable checkpointing. Default: False.
379
+ conv_at_attn (bool): If True, performe depthwise convolution before attention layer. Default: True.
380
+ conv_at_ffn (bool): If True, performe depthwise convolution before ffn layer. Default: True.
381
+ """
382
+
383
+ def __init__(
384
+ self,
385
+ in_chans=3,
386
+ num_classes=1000,
387
+ depths=(1, 1, 3, 1),
388
+ patch_size=(7, 2, 2, 2),
389
+ patch_stride=(4, 2, 2, 2),
390
+ patch_padding=(3, 0, 0, 0),
391
+ patch_prenorm=(False, False, False, False),
392
+ embed_dims=(64, 128, 192, 256),
393
+ num_heads=(3, 6, 12, 24),
394
+ num_groups=(3, 6, 12, 24),
395
+ window_size=7,
396
+ mlp_ratio=4.0,
397
+ qkv_bias=True,
398
+ drop_path_rate=0.1,
399
+ norm_layer=nn.LayerNorm,
400
+ enable_checkpoint=False,
401
+ conv_at_attn=True,
402
+ conv_at_ffn=True,
403
+ ):
404
+ super().__init__()
405
+
406
+ self.num_classes = num_classes
407
+ self.embed_dims = embed_dims
408
+ self.num_heads = num_heads
409
+ self.num_groups = num_groups
410
+ self.num_stages = len(self.embed_dims)
411
+ self.enable_checkpoint = enable_checkpoint
412
+ assert self.num_stages == len(self.num_heads) == len(self.num_groups)
413
+
414
+ num_stages = len(embed_dims)
415
+ dpr = [x.item() for x in torch.linspace(0, drop_path_rate, sum(depths) * 2)]
416
+
417
+ depth_offset = 0
418
+ convs = []
419
+ blocks = []
420
+ for i in range(num_stages):
421
+ conv_embed = ConvEmbed(
422
+ patch_size=patch_size[i],
423
+ stride=patch_stride[i],
424
+ padding=patch_padding[i],
425
+ in_chans=in_chans if i == 0 else self.embed_dims[i - 1],
426
+ embed_dim=self.embed_dims[i],
427
+ norm_layer=norm_layer,
428
+ pre_norm=patch_prenorm[i],
429
+ )
430
+ convs.append(conv_embed)
431
+
432
+ block = MySequential(
433
+ *[
434
+ MySequential(
435
+ OrderedDict(
436
+ [
437
+ (
438
+ "spatial_block",
439
+ SpatialBlock(
440
+ embed_dims[i],
441
+ num_heads[i],
442
+ window_size,
443
+ drop_path_rate=dpr[depth_offset + j * 2],
444
+ qkv_bias=qkv_bias,
445
+ mlp_ratio=mlp_ratio,
446
+ conv_at_attn=conv_at_attn,
447
+ conv_at_ffn=conv_at_ffn,
448
+ ),
449
+ ),
450
+ (
451
+ "channel_block",
452
+ ChannelBlock(
453
+ embed_dims[i],
454
+ num_groups[i],
455
+ drop_path_rate=dpr[depth_offset + j * 2 + 1],
456
+ qkv_bias=qkv_bias,
457
+ mlp_ratio=mlp_ratio,
458
+ conv_at_attn=conv_at_attn,
459
+ conv_at_ffn=conv_at_ffn,
460
+ ),
461
+ ),
462
+ ]
463
+ )
464
+ )
465
+ for j in range(depths[i])
466
+ ]
467
+ )
468
+ blocks.append(block)
469
+ depth_offset += depths[i] * 2
470
+
471
+ self.convs = nn.ModuleList(convs)
472
+ self.blocks = nn.ModuleList(blocks)
473
+
474
+ self.norms = norm_layer(self.embed_dims[-1])
475
+ self.avgpool = nn.AdaptiveAvgPool1d(1)
476
+ self.head = (
477
+ nn.Linear(self.embed_dims[-1], num_classes)
478
+ if num_classes > 0
479
+ else nn.Identity()
480
+ )
481
+
482
+ # self.apply(self._init_weights)
483
+
484
+ @property
485
+ def dim_out(self):
486
+ return self.embed_dims[-1]
487
+
488
+ def _init_weights(self, m):
489
+ if isinstance(m, nn.Linear):
490
+ trunc_normal_(m.weight, std=0.02)
491
+ if m.bias is not None:
492
+ nn.init.constant_(m.bias, 0)
493
+ elif isinstance(m, nn.Conv2d):
494
+ nn.init.normal_(m.weight, std=0.02)
495
+ for name, _ in m.named_parameters():
496
+ if name in ["bias"]:
497
+ nn.init.constant_(m.bias, 0)
498
+ elif isinstance(m, nn.LayerNorm):
499
+ nn.init.constant_(m.weight, 1.0)
500
+ nn.init.constant_(m.bias, 0)
501
+ elif isinstance(m, nn.BatchNorm2d):
502
+ nn.init.constant_(m.weight, 1.0)
503
+ nn.init.constant_(m.bias, 0)
504
+
505
+ def forward_features_unpool(self, x):
506
+ """
507
+ forward until avg pooling
508
+ Args:
509
+ x (_type_): input image tensor
510
+ """
511
+ input_size = (x.size(2), x.size(3))
512
+ for conv, block in zip(self.convs, self.blocks):
513
+ x, input_size = conv(x, input_size)
514
+ if self.enable_checkpoint:
515
+ x, input_size = checkpoint.checkpoint(block, x, input_size)
516
+ else:
517
+ x, input_size = block(x, input_size)
518
+ return x
519
+
520
+ def forward_features(self, x):
521
+ x = self.forward_features_unpool(x)
522
+
523
+ # (batch_size, num_tokens, token_dim)
524
+ x = self.avgpool(x.transpose(1, 2))
525
+ # (batch_size, 1, num_tokens)
526
+ x = torch.flatten(x, 1)
527
+ x = self.norms(x)
528
+
529
+ return x
530
+
531
+ def forward(self, x):
532
+ x = self.forward_features(x)
533
+ x = self.head(x)
534
+ return x
535
+
536
+ @classmethod
537
+ def from_config(cls, config):
538
+ return cls(
539
+ depths=config.depths,
540
+ embed_dims=config.dim_embed,
541
+ num_heads=config.num_heads,
542
+ num_groups=config.num_groups,
543
+ patch_size=config.patch_size,
544
+ patch_stride=config.patch_stride,
545
+ patch_padding=config.patch_padding,
546
+ patch_prenorm=config.patch_prenorm,
547
+ drop_path_rate=config.drop_path_rate,
548
+ window_size=config.window_size,
549
+ )
550
+
551
+
552
+ class DaViTVisionTower(nn.Module):
553
+ def __init__(self, vision_tower, args, delay_load=False):
554
+ super().__init__()
555
+
556
+ self.is_loaded = False
557
+ self.image_size = getattr(args, "size", 768)
558
+ self.vision_tower_name = vision_tower
559
+ self.select_layer = args.mm_vision_select_layer
560
+ self.select_feature = getattr(args, "mm_vision_select_feature", "patch")
561
+ self.config = AutoConfig.from_pretrained(
562
+ self.vision_tower_name, trust_remote_code=True
563
+ ).vision_config
564
+ if not delay_load:
565
+ self.load_model()
566
+ elif getattr(args, "unfreeze_mm_vision_tower", False):
567
+ self.load_model()
568
+ else:
569
+ self.cfg_only = self.config
570
+
571
+ def load_model(self, device_map=None):
572
+ if self.is_loaded:
573
+ print(
574
+ "{} is already loaded, `load_model` called again, skipping.".format(
575
+ self.vision_tower_name
576
+ )
577
+ )
578
+ return
579
+ self.vision_tower = DaViT.from_config(config=self.config)
580
+ del self.vision_tower.head
581
+ del self.vision_tower.norms
582
+ self.image_processor = CLIPImageProcessor.from_pretrained(
583
+ self.vision_tower_name
584
+ )
585
+ self.vision_tower.requires_grad_(False)
586
+ self.is_loaded = True
587
+
588
+ def feature_select(self, image_forward_outs):
589
+ image_features = image_forward_outs
590
+ if self.select_feature == "patch":
591
+ image_features = image_features[:, 1:]
592
+ elif self.select_feature == "cls_patch":
593
+ image_features = image_features
594
+ else:
595
+ raise ValueError(f"Unexpected select feature: {self.select_feature}")
596
+ return image_features
597
+
598
+ @torch.no_grad()
599
+ def forward(self, images):
600
+ if type(images) is list:
601
+ image_features = []
602
+ for image in images:
603
+ image_forward_out = self.vision_tower.forward_features_unpool(
604
+ image.to(device=self.device, dtype=self.dtype).unsqueeze(0)
605
+ )
606
+ image_feature = self.feature_select(image_forward_out).to(image.dtype)
607
+ image_features.append(image_forward_out)
608
+ else:
609
+ image_forward_outs = self.vision_tower.forward_features_unpool(
610
+ images.to(device=self.device, dtype=self.dtype)
611
+ )
612
+ image_features = self.feature_select(image_forward_outs).to(images.dtype)
613
+ return image_features
614
+
615
+ @property
616
+ def dummy_feature(self):
617
+ return torch.zeros(1, self.hidden_size, device=self.device, dtype=self.dtype)
618
+
619
+ @property
620
+ def dtype(self):
621
+ for n, p in self.vision_tower.named_parameters():
622
+ dtype = p.dtype
623
+ break
624
+ return dtype
625
+
626
+ @property
627
+ def device(self):
628
+ for n, p in self.vision_tower.named_parameters():
629
+ device = p.device
630
+ break
631
+ return device
632
+
633
+ @property
634
+ def hidden_size(self):
635
+ return self.config.dim_embed[-1]
636
+
637
+ @property
638
+ def num_patches_per_side(self):
639
+ return self.config.image_size // self.config.patch_size
640
+
641
+ @property
642
+ def num_patches(self):
643
+ return (self.config.image_size // self.config.patch_size) ** 2
644
+
645
+
646
+ from abc import ABC, abstractmethod
647
+
648
+
649
+ IGNORE_INDEX = -100
650
+ IMAGE_TOKEN_INDEX = -200
651
+ DEFAULT_IMAGE_PATCH_TOKEN = "<im_patch>"
652
+ DEFAULT_IM_START_TOKEN = "<im_start>"
653
+ DEFAULT_IM_END_TOKEN = "<im_end>"
654
+
655
+
656
+ def build_vision_tower(vision_tower_cfg, **kwargs):
657
+ vision_tower = getattr(
658
+ vision_tower_cfg,
659
+ "mm_vision_tower",
660
+ getattr(vision_tower_cfg, "vision_tower", None),
661
+ )
662
+ return DaViTVisionTower(vision_tower, args=vision_tower_cfg, **kwargs)
663
+
664
+
665
+ import re
666
+ def build_vision_projector(config, delay_load=False, **kwargs):
667
+ projector_type = getattr(config, "mm_projector_type", "linear")
668
+
669
+ mlp_gelu_match = re.match(r"^mlp(\d+)x_gelu$", projector_type)
670
+ if mlp_gelu_match:
671
+ mlp_depth = int(mlp_gelu_match.group(1))
672
+ modules = [nn.Linear(config.mm_hidden_size, config.hidden_size)]
673
+ for _ in range(1, mlp_depth):
674
+ modules.append(nn.GELU())
675
+ modules.append(nn.Linear(config.hidden_size, config.hidden_size))
676
+ return nn.Sequential(*modules)
677
+
678
+
679
+ class MixsenseMetaModel:
680
+
681
+ def __init__(self, config):
682
+ super(MixsenseMetaModel, self).__init__(config)
683
+
684
+ if hasattr(config, "mm_vision_tower"):
685
+ self.vision_tower = build_vision_tower(config, delay_load=True)
686
+ self.mm_projector = build_vision_projector(config)
687
+
688
+ if "unpad" in getattr(config, "mm_patch_merge_type", ""):
689
+ self.image_newline = nn.Parameter(
690
+ torch.empty(config.hidden_size, dtype=self.dtype)
691
+ )
692
+
693
+ def get_vision_tower(self):
694
+ vision_tower = getattr(self, "vision_tower", None)
695
+ if type(vision_tower) is list:
696
+ vision_tower = vision_tower[0]
697
+ return vision_tower
698
+
699
+ def initialize_vision_modules(self, model_args, fsdp=None):
700
+ vision_tower = model_args.vision_tower
701
+ mm_vision_select_layer = model_args.mm_vision_select_layer
702
+ mm_vision_select_feature = model_args.mm_vision_select_feature
703
+ pretrain_mm_mlp_adapter = model_args.pretrain_mm_mlp_adapter
704
+ mm_patch_merge_type = model_args.mm_patch_merge_type
705
+
706
+ self.config.mm_vision_tower = vision_tower
707
+
708
+ if self.get_vision_tower() is None:
709
+ vision_tower = build_vision_tower(model_args)
710
+
711
+ if fsdp is not None and len(fsdp) > 0:
712
+ self.vision_tower = [vision_tower]
713
+ else:
714
+ self.vision_tower = vision_tower
715
+ else:
716
+ if fsdp is not None and len(fsdp) > 0:
717
+ vision_tower = self.vision_tower[0]
718
+ else:
719
+ vision_tower = self.vision_tower
720
+ vision_tower.load_model()
721
+
722
+ self.config.use_mm_proj = True
723
+ self.config.mm_projector_type = getattr(
724
+ model_args, "mm_projector_type", "linear"
725
+ )
726
+ self.config.mm_hidden_size = vision_tower.hidden_size
727
+ self.config.mm_vision_select_layer = mm_vision_select_layer
728
+ self.config.mm_vision_select_feature = mm_vision_select_feature
729
+ self.config.mm_patch_merge_type = mm_patch_merge_type
730
+
731
+ if getattr(self, "mm_projector", None) is None:
732
+ self.mm_projector = build_vision_projector(self.config)
733
+
734
+ if "unpad" in mm_patch_merge_type:
735
+ embed_std = 1 / torch.sqrt(
736
+ torch.tensor(self.config.hidden_size, dtype=self.dtype)
737
+ )
738
+ self.image_newline = nn.Parameter(
739
+ torch.randn(self.config.hidden_size, dtype=self.dtype) * embed_std
740
+ )
741
+ else:
742
+ # In case it is frozen by LoRA
743
+ for p in self.mm_projector.parameters():
744
+ p.requires_grad = True
745
+
746
+ if pretrain_mm_mlp_adapter is not None:
747
+ mm_projector_weights = torch.load(
748
+ pretrain_mm_mlp_adapter, map_location="cpu"
749
+ )
750
+
751
+ def get_w(weights, keyword):
752
+ return {
753
+ k.split(keyword + ".")[1]: v
754
+ for k, v in weights.items()
755
+ if keyword in k
756
+ }
757
+
758
+ self.mm_projector.load_state_dict(
759
+ get_w(mm_projector_weights, "mm_projector")
760
+ )
761
+
762
+
763
+ class MixsenseMetaForCausalLM(ABC):
764
+
765
+ @abstractmethod
766
+ def get_model(self):
767
+ pass
768
+
769
+ def get_vision_tower(self):
770
+ return self.get_model().get_vision_tower()
771
+
772
+ def encode_images(self, images):
773
+ image_features = self.get_model().get_vision_tower()(images)
774
+ image_features = self.get_model().mm_projector(image_features)
775
+ return image_features
776
+
777
+ def prepare_inputs_labels_for_multimodal(
778
+ self,
779
+ input_ids,
780
+ position_ids,
781
+ attention_mask,
782
+ past_key_values,
783
+ labels,
784
+ images,
785
+ image_sizes=None,
786
+ ):
787
+ vision_tower = self.get_vision_tower()
788
+ if vision_tower is None or images is None or input_ids.shape[1] == 1:
789
+ return (
790
+ input_ids,
791
+ position_ids,
792
+ attention_mask,
793
+ past_key_values,
794
+ None,
795
+ labels,
796
+ )
797
+ elif type(images) is list or images.ndim == 5:
798
+ if type(images) is list:
799
+ images = [x.unsqueeze(0) if x.ndim == 3 else x for x in images]
800
+ concat_images = torch.cat([image for image in images], dim=0)
801
+ image_features = self.encode_images(concat_images)
802
+ split_sizes = [image.shape[0] for image in images]
803
+ image_features = torch.split(image_features, split_sizes, dim=0)
804
+ mm_patch_merge_type = getattr(self.config, "mm_patch_merge_type", "flat")
805
+ image_aspect_ratio = getattr(self.config, "image_aspect_ratio", "square")
806
+ if mm_patch_merge_type == "flat":
807
+ image_features = [x.flatten(0, 1) for x in image_features]
808
+ else:
809
+ image_features = self.encode_images(images)
810
+
811
+ # TODO: image start / end is not implemented here to support pretraining.
812
+ if getattr(self.config, "tune_mm_mlp_adapter", False) and getattr(
813
+ self.config, "mm_use_im_start_end", False
814
+ ):
815
+ raise NotImplementedError
816
+
817
+ # Let's just add dummy tensors if they do not exist,
818
+ # it is a headache to deal with None all the time.
819
+ # But it is not ideal, and if you have a better idea,
820
+ # please open an issue / submit a PR, thanks.
821
+ _labels = labels
822
+ _position_ids = position_ids
823
+ _attention_mask = attention_mask
824
+ if attention_mask is None:
825
+ attention_mask = torch.ones_like(input_ids, dtype=torch.bool)
826
+ else:
827
+ attention_mask = attention_mask.bool()
828
+ if position_ids is None:
829
+ position_ids = torch.arange(
830
+ 0, input_ids.shape[1], dtype=torch.long, device=input_ids.device
831
+ )
832
+ if labels is None:
833
+ labels = torch.full_like(input_ids, IGNORE_INDEX)
834
+
835
+ # remove the padding using attention_mask -- FIXME
836
+ _input_ids = input_ids
837
+ input_ids = [
838
+ cur_input_ids[cur_attention_mask]
839
+ for cur_input_ids, cur_attention_mask in zip(input_ids, attention_mask)
840
+ ]
841
+ labels = [
842
+ cur_labels[cur_attention_mask]
843
+ for cur_labels, cur_attention_mask in zip(labels, attention_mask)
844
+ ]
845
+
846
+ new_input_embeds = []
847
+ new_labels = []
848
+ cur_image_idx = 0
849
+ for batch_idx, cur_input_ids in enumerate(input_ids):
850
+ num_images = (cur_input_ids == IMAGE_TOKEN_INDEX).sum()
851
+ if num_images == 0:
852
+ cur_image_features = image_features[cur_image_idx]
853
+ cur_input_embeds_1 = self.get_model().embed_tokens(cur_input_ids)
854
+ cur_input_embeds = torch.cat(
855
+ [cur_input_embeds_1, cur_image_features[0:0]], dim=0
856
+ )
857
+ new_input_embeds.append(cur_input_embeds)
858
+ new_labels.append(labels[batch_idx])
859
+ cur_image_idx += 1
860
+ continue
861
+
862
+ image_token_indices = (
863
+ [-1]
864
+ + torch.where(cur_input_ids == IMAGE_TOKEN_INDEX)[0].tolist()
865
+ + [cur_input_ids.shape[0]]
866
+ )
867
+ cur_input_ids_noim = []
868
+ cur_labels = labels[batch_idx]
869
+ cur_labels_noim = []
870
+ for i in range(len(image_token_indices) - 1):
871
+ cur_input_ids_noim.append(
872
+ cur_input_ids[
873
+ image_token_indices[i] + 1 : image_token_indices[i + 1]
874
+ ]
875
+ )
876
+ cur_labels_noim.append(
877
+ cur_labels[image_token_indices[i] + 1 : image_token_indices[i + 1]]
878
+ )
879
+ split_sizes = [x.shape[0] for x in cur_labels_noim]
880
+ cur_input_embeds = self.get_model().embed_tokens(
881
+ torch.cat(cur_input_ids_noim)
882
+ )
883
+
884
+ cur_input_embeds_no_im = torch.split(cur_input_embeds, split_sizes, dim=0)
885
+ cur_new_input_embeds = []
886
+ cur_new_labels = []
887
+
888
+ for i in range(num_images + 1):
889
+ cur_new_input_embeds.append(cur_input_embeds_no_im[i])
890
+ cur_new_labels.append(cur_labels_noim[i])
891
+ if i < num_images:
892
+ cur_image_features = image_features[cur_image_idx]
893
+ cur_image_idx += 1
894
+ cur_new_input_embeds.append(cur_image_features)
895
+ cur_new_labels.append(
896
+ torch.full(
897
+ (cur_image_features.shape[0],),
898
+ IGNORE_INDEX,
899
+ device=cur_labels.device,
900
+ dtype=cur_labels.dtype,
901
+ )
902
+ )
903
+
904
+ cur_new_input_embeds = [x.to(self.device) for x in cur_new_input_embeds]
905
+
906
+ cur_new_input_embeds = torch.cat(cur_new_input_embeds)
907
+ cur_new_labels = torch.cat(cur_new_labels)
908
+
909
+ new_input_embeds.append(cur_new_input_embeds)
910
+ new_labels.append(cur_new_labels)
911
+
912
+ # Truncate sequences to max length as image embeddings can make the sequence longer
913
+ tokenizer_model_max_length = getattr(
914
+ self.config, "tokenizer_model_max_length", None
915
+ )
916
+ if tokenizer_model_max_length is not None:
917
+ new_input_embeds = [
918
+ x[:tokenizer_model_max_length] for x in new_input_embeds
919
+ ]
920
+ new_labels = [x[:tokenizer_model_max_length] for x in new_labels]
921
+
922
+ # Combine them
923
+ max_len = max(x.shape[0] for x in new_input_embeds)
924
+ batch_size = len(new_input_embeds)
925
+
926
+ new_input_embeds_padded = []
927
+ new_labels_padded = torch.full(
928
+ (batch_size, max_len),
929
+ IGNORE_INDEX,
930
+ dtype=new_labels[0].dtype,
931
+ device=new_labels[0].device,
932
+ )
933
+ attention_mask = torch.zeros(
934
+ (batch_size, max_len),
935
+ dtype=attention_mask.dtype,
936
+ device=attention_mask.device,
937
+ )
938
+ position_ids = torch.zeros(
939
+ (batch_size, max_len), dtype=position_ids.dtype, device=position_ids.device
940
+ )
941
+
942
+ for i, (cur_new_embed, cur_new_labels) in enumerate(
943
+ zip(new_input_embeds, new_labels)
944
+ ):
945
+ cur_len = cur_new_embed.shape[0]
946
+ if getattr(self.config, "tokenizer_padding_side", "right") == "left":
947
+ new_input_embeds_padded.append(
948
+ torch.cat(
949
+ (
950
+ torch.zeros(
951
+ (max_len - cur_len, cur_new_embed.shape[1]),
952
+ dtype=cur_new_embed.dtype,
953
+ device=cur_new_embed.device,
954
+ ),
955
+ cur_new_embed,
956
+ ),
957
+ dim=0,
958
+ )
959
+ )
960
+ if cur_len > 0:
961
+ new_labels_padded[i, -cur_len:] = cur_new_labels
962
+ attention_mask[i, -cur_len:] = True
963
+ position_ids[i, -cur_len:] = torch.arange(
964
+ 0, cur_len, dtype=position_ids.dtype, device=position_ids.device
965
+ )
966
+ else:
967
+ new_input_embeds_padded.append(
968
+ torch.cat(
969
+ (
970
+ cur_new_embed,
971
+ torch.zeros(
972
+ (max_len - cur_len, cur_new_embed.shape[1]),
973
+ dtype=cur_new_embed.dtype,
974
+ device=cur_new_embed.device,
975
+ ),
976
+ ),
977
+ dim=0,
978
+ )
979
+ )
980
+ if cur_len > 0:
981
+ new_labels_padded[i, :cur_len] = cur_new_labels
982
+ attention_mask[i, :cur_len] = True
983
+ position_ids[i, :cur_len] = torch.arange(
984
+ 0, cur_len, dtype=position_ids.dtype, device=position_ids.device
985
+ )
986
+
987
+ new_input_embeds = torch.stack(new_input_embeds_padded, dim=0)
988
+
989
+ if _labels is None:
990
+ new_labels = None
991
+ else:
992
+ new_labels = new_labels_padded
993
+
994
+ if _attention_mask is None:
995
+ attention_mask = None
996
+ else:
997
+ attention_mask = attention_mask.to(dtype=_attention_mask.dtype)
998
+
999
+ if _position_ids is None:
1000
+ position_ids = None
1001
+
1002
+ return (
1003
+ None,
1004
+ position_ids,
1005
+ attention_mask,
1006
+ past_key_values,
1007
+ new_input_embeds,
1008
+ new_labels,
1009
+ )
1010
+
1011
+ def initialize_vision_tokenizer(self, model_args, tokenizer):
1012
+ if model_args.mm_use_im_patch_token:
1013
+ tokenizer.add_tokens([DEFAULT_IMAGE_PATCH_TOKEN], special_tokens=True)
1014
+ self.resize_token_embeddings(len(tokenizer))
1015
+
1016
+ if model_args.mm_use_im_start_end:
1017
+ num_new_tokens = tokenizer.add_tokens(
1018
+ [DEFAULT_IM_START_TOKEN, DEFAULT_IM_END_TOKEN], special_tokens=True
1019
+ )
1020
+ self.resize_token_embeddings(len(tokenizer))
1021
+
1022
+ if num_new_tokens > 0:
1023
+ input_embeddings = self.get_input_embeddings().weight.data
1024
+ output_embeddings = self.get_output_embeddings().weight.data
1025
+
1026
+ input_embeddings_avg = input_embeddings[:-num_new_tokens].mean(
1027
+ dim=0, keepdim=True
1028
+ )
1029
+ output_embeddings_avg = output_embeddings[:-num_new_tokens].mean(
1030
+ dim=0, keepdim=True
1031
+ )
1032
+
1033
+ input_embeddings[-num_new_tokens:] = input_embeddings_avg
1034
+ output_embeddings[-num_new_tokens:] = output_embeddings_avg
1035
+
1036
+ if model_args.tune_mm_mlp_adapter:
1037
+ for p in self.get_input_embeddings().parameters():
1038
+ p.requires_grad = True
1039
+ for p in self.get_output_embeddings().parameters():
1040
+ p.requires_grad = False
1041
+
1042
+ if model_args.pretrain_mm_mlp_adapter:
1043
+ mm_projector_weights = torch.load(
1044
+ model_args.pretrain_mm_mlp_adapter, map_location="cpu"
1045
+ )
1046
+ embed_tokens_weight = mm_projector_weights["model.embed_tokens.weight"]
1047
+ assert num_new_tokens == 2
1048
+ if input_embeddings.shape == embed_tokens_weight.shape:
1049
+ input_embeddings[-num_new_tokens:] = embed_tokens_weight[
1050
+ -num_new_tokens:
1051
+ ]
1052
+ elif embed_tokens_weight.shape[0] == num_new_tokens:
1053
+ input_embeddings[-num_new_tokens:] = embed_tokens_weight
1054
+ else:
1055
+ raise ValueError(
1056
+ f"Unexpected embed_tokens_weight shape. Pretrained: {embed_tokens_weight.shape}. Current: {input_embeddings.shape}. Numer of new tokens: {num_new_tokens}."
1057
+ )
1058
+ elif model_args.mm_use_im_patch_token:
1059
+ if model_args.tune_mm_mlp_adapter:
1060
+ for p in self.get_input_embeddings().parameters():
1061
+ p.requires_grad = False
1062
+ for p in self.get_output_embeddings().parameters():
1063
+ p.requires_grad = False
1064
+
1065
+ from typing import List, Optional, Tuple, Union
1066
+ from transformers import (
1067
+ AutoConfig,
1068
+ AutoModelForCausalLM,
1069
+ LlamaConfig,
1070
+ LlamaModel,
1071
+ LlamaForCausalLM,
1072
+ )
1073
+ from transformers.modeling_outputs import CausalLMOutputWithPast
1074
+ from transformers.generation.utils import GenerateOutput
1075
+
1076
+
1077
+ class MixsenseConfig(LlamaConfig):
1078
+ model_type = "mixsense_llama"
1079
+
1080
+
1081
+ class MixsenseLlamaModel(MixsenseMetaModel, LlamaModel):
1082
+ config_class = MixsenseConfig
1083
+
1084
+ def __init__(self, config: LlamaConfig):
1085
+ super(MixsenseLlamaModel, self).__init__(config)
1086
+
1087
+
1088
+ class MixsenseLlamaForCausalLM(LlamaForCausalLM, MixsenseMetaForCausalLM):
1089
+ config_class = MixsenseConfig
1090
+
1091
+ def __init__(self, config):
1092
+ super(LlamaForCausalLM, self).__init__(config)
1093
+ self.model = MixsenseLlamaModel(config)
1094
+ self.pretraining_tp = config.pretraining_tp
1095
+ self.vocab_size = config.vocab_size
1096
+ self.lm_head = nn.Linear(config.hidden_size, config.vocab_size, bias=False)
1097
+
1098
+ # Initialize weights and apply final processing
1099
+ self.post_init()
1100
+
1101
+ def get_model(self):
1102
+ return self.model
1103
+
1104
+ def forward(
1105
+ self,
1106
+ input_ids: torch.LongTensor = None,
1107
+ attention_mask: Optional[torch.Tensor] = None,
1108
+ position_ids: Optional[torch.LongTensor] = None,
1109
+ past_key_values: Optional[List[torch.FloatTensor]] = None,
1110
+ inputs_embeds: Optional[torch.FloatTensor] = None,
1111
+ labels: Optional[torch.LongTensor] = None,
1112
+ use_cache: Optional[bool] = None,
1113
+ output_attentions: Optional[bool] = None,
1114
+ output_hidden_states: Optional[bool] = None,
1115
+ images: Optional[torch.FloatTensor] = None,
1116
+ image_sizes: Optional[List[List[int]]] = None,
1117
+ return_dict: Optional[bool] = None,
1118
+ ) -> Union[Tuple, CausalLMOutputWithPast]:
1119
+ if inputs_embeds is None:
1120
+ (
1121
+ input_ids,
1122
+ position_ids,
1123
+ attention_mask,
1124
+ past_key_values,
1125
+ inputs_embeds,
1126
+ labels,
1127
+ ) = self.prepare_inputs_labels_for_multimodal(
1128
+ input_ids,
1129
+ position_ids,
1130
+ attention_mask,
1131
+ past_key_values,
1132
+ labels,
1133
+ images,
1134
+ image_sizes,
1135
+ )
1136
+ return super().forward(
1137
+ input_ids=input_ids,
1138
+ attention_mask=attention_mask,
1139
+ position_ids=position_ids,
1140
+ past_key_values=past_key_values,
1141
+ inputs_embeds=inputs_embeds,
1142
+ labels=labels,
1143
+ use_cache=use_cache,
1144
+ output_attentions=output_attentions,
1145
+ output_hidden_states=output_hidden_states,
1146
+ return_dict=return_dict,
1147
+ )
1148
+
1149
+ @torch.no_grad()
1150
+ def generate(
1151
+ self,
1152
+ inputs: Optional[torch.Tensor] = None,
1153
+ images: Optional[torch.Tensor] = None,
1154
+ image_sizes: Optional[torch.Tensor] = None,
1155
+ **kwargs,
1156
+ ) -> Union[GenerateOutput, torch.LongTensor]:
1157
+ position_ids = kwargs.pop("position_ids", None)
1158
+ attention_mask = kwargs.pop("attention_mask", None)
1159
+ if "inputs_embeds" in kwargs:
1160
+ raise NotImplementedError("`inputs_embeds` is not supported")
1161
+ if images is not None:
1162
+ (inputs, position_ids, attention_mask, _, inputs_embeds, _) = (
1163
+ self.prepare_inputs_labels_for_multimodal(
1164
+ inputs,
1165
+ position_ids,
1166
+ attention_mask,
1167
+ None,
1168
+ None,
1169
+ images,
1170
+ image_sizes=image_sizes,
1171
+ )
1172
+ )
1173
+ else:
1174
+ inputs_embeds = self.get_model().embed_tokens(inputs)
1175
+ output = super().generate(
1176
+ position_ids=position_ids,
1177
+ attention_mask=attention_mask,
1178
+ inputs_embeds=inputs_embeds,
1179
+ **kwargs,
1180
+ )
1181
+ return output
1182
+
1183
+ def prepare_inputs_for_generation(
1184
+ self, input_ids, past_key_values=None, inputs_embeds=None, **kwargs
1185
+ ):
1186
+ images = kwargs.pop("images", None)
1187
+ image_sizes = kwargs.pop("image_sizes", None)
1188
+ inputs = super().prepare_inputs_for_generation(
1189
+ input_ids,
1190
+ past_key_values=past_key_values,
1191
+ inputs_embeds=inputs_embeds,
1192
+ **kwargs,
1193
+ )
1194
+ if images is not None:
1195
+ inputs["images"] = images
1196
+ if image_sizes is not None:
1197
+ inputs["image_sizes"] = image_sizes
1198
+ return inputs
1199
+ def image_process(self,images):
1200
+ vision_tower = self.get_vision_tower()
1201
+ if not vision_tower.is_loaded:
1202
+ vision_tower.load_model()
1203
+ processor = vision_tower.image_processor
1204
+ def expand2square(pil_img, background_color):
1205
+ from PIL import Image
1206
+ width, height = pil_img.size
1207
+ if width == height:
1208
+ return pil_img
1209
+ elif width > height:
1210
+ result = Image.new(pil_img.mode, (width, width), background_color)
1211
+ result.paste(pil_img, (0, (width - height) // 2))
1212
+ return result
1213
+ else:
1214
+ result = Image.new(pil_img.mode, (height, height), background_color)
1215
+ result.paste(pil_img, ((height - width) // 2, 0))
1216
+ return result
1217
+ processed_images=[]
1218
+ for image in images:
1219
+ image = expand2square(image, tuple(int(x*255) for x in processor.image_mean))
1220
+ image = processor.preprocess(image, return_tensors='pt')['pixel_values'][0]
1221
+ processed_images.append(image)
1222
+ if all(x.shape == processed_images[0].shape for x in processed_images):
1223
+ processed_images = torch.stack(processed_images, dim=0)
1224
+ return processed_images
1225
+ def text_process(self,text,tokenizer):
1226
+ prompt=f"<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nYou are a helpful language and vision assistant. You are able to understand the visual content that the user provides, and assist the user with a variety of tasks using natural language.<|eot_id|><|start_header_id|>user<|end_header_id|>\n\n<image>\n{text}<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n"
1227
+ text_chunks = [tokenizer(chunk).input_ids for chunk in prompt.split('<image>')]
1228
+ input_ids = torch.tensor(text_chunks[0] + [-200] + text_chunks[1][1:], dtype=torch.long).unsqueeze(0)
1229
+ return input_ids
1230
+
1231
+
1232
+ AutoConfig.register("mixsense_llama", MixsenseConfig)
1233
+ AutoModelForCausalLM.register(MixsenseConfig, MixsenseLlamaForCausalLM)
special_tokens_map.json ADDED
@@ -0,0 +1,23 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token": {
3
+ "content": "<|begin_of_text|>",
4
+ "lstrip": false,
5
+ "normalized": false,
6
+ "rstrip": false,
7
+ "single_word": false
8
+ },
9
+ "eos_token": {
10
+ "content": "<|end_of_text|>",
11
+ "lstrip": false,
12
+ "normalized": false,
13
+ "rstrip": false,
14
+ "single_word": false
15
+ },
16
+ "pad_token": {
17
+ "content": "<pad>",
18
+ "lstrip": false,
19
+ "normalized": false,
20
+ "rstrip": false,
21
+ "single_word": false
22
+ }
23
+ }
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer_config.json ADDED
@@ -0,0 +1,2071 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "added_tokens_decoder": {
3
+ "128000": {
4
+ "content": "<|begin_of_text|>",
5
+ "lstrip": false,
6
+ "normalized": false,
7
+ "rstrip": false,
8
+ "single_word": false,
9
+ "special": true
10
+ },
11
+ "128001": {
12
+ "content": "<|end_of_text|>",
13
+ "lstrip": false,
14
+ "normalized": false,
15
+ "rstrip": false,
16
+ "single_word": false,
17
+ "special": true
18
+ },
19
+ "128002": {
20
+ "content": "<|reserved_special_token_0|>",
21
+ "lstrip": false,
22
+ "normalized": false,
23
+ "rstrip": false,
24
+ "single_word": false,
25
+ "special": true
26
+ },
27
+ "128003": {
28
+ "content": "<|reserved_special_token_1|>",
29
+ "lstrip": false,
30
+ "normalized": false,
31
+ "rstrip": false,
32
+ "single_word": false,
33
+ "special": true
34
+ },
35
+ "128004": {
36
+ "content": "<|reserved_special_token_2|>",
37
+ "lstrip": false,
38
+ "normalized": false,
39
+ "rstrip": false,
40
+ "single_word": false,
41
+ "special": true
42
+ },
43
+ "128005": {
44
+ "content": "<|reserved_special_token_3|>",
45
+ "lstrip": false,
46
+ "normalized": false,
47
+ "rstrip": false,
48
+ "single_word": false,
49
+ "special": true
50
+ },
51
+ "128006": {
52
+ "content": "<|start_header_id|>",
53
+ "lstrip": false,
54
+ "normalized": false,
55
+ "rstrip": false,
56
+ "single_word": false,
57
+ "special": true
58
+ },
59
+ "128007": {
60
+ "content": "<|end_header_id|>",
61
+ "lstrip": false,
62
+ "normalized": false,
63
+ "rstrip": false,
64
+ "single_word": false,
65
+ "special": true
66
+ },
67
+ "128008": {
68
+ "content": "<|reserved_special_token_4|>",
69
+ "lstrip": false,
70
+ "normalized": false,
71
+ "rstrip": false,
72
+ "single_word": false,
73
+ "special": true
74
+ },
75
+ "128009": {
76
+ "content": "<|eot_id|>",
77
+ "lstrip": false,
78
+ "normalized": false,
79
+ "rstrip": false,
80
+ "single_word": false,
81
+ "special": true
82
+ },
83
+ "128010": {
84
+ "content": "<|reserved_special_token_5|>",
85
+ "lstrip": false,
86
+ "normalized": false,
87
+ "rstrip": false,
88
+ "single_word": false,
89
+ "special": true
90
+ },
91
+ "128011": {
92
+ "content": "<|reserved_special_token_6|>",
93
+ "lstrip": false,
94
+ "normalized": false,
95
+ "rstrip": false,
96
+ "single_word": false,
97
+ "special": true
98
+ },
99
+ "128012": {
100
+ "content": "<|reserved_special_token_7|>",
101
+ "lstrip": false,
102
+ "normalized": false,
103
+ "rstrip": false,
104
+ "single_word": false,
105
+ "special": true
106
+ },
107
+ "128013": {
108
+ "content": "<|reserved_special_token_8|>",
109
+ "lstrip": false,
110
+ "normalized": false,
111
+ "rstrip": false,
112
+ "single_word": false,
113
+ "special": true
114
+ },
115
+ "128014": {
116
+ "content": "<|reserved_special_token_9|>",
117
+ "lstrip": false,
118
+ "normalized": false,
119
+ "rstrip": false,
120
+ "single_word": false,
121
+ "special": true
122
+ },
123
+ "128015": {
124
+ "content": "<|reserved_special_token_10|>",
125
+ "lstrip": false,
126
+ "normalized": false,
127
+ "rstrip": false,
128
+ "single_word": false,
129
+ "special": true
130
+ },
131
+ "128016": {
132
+ "content": "<|reserved_special_token_11|>",
133
+ "lstrip": false,
134
+ "normalized": false,
135
+ "rstrip": false,
136
+ "single_word": false,
137
+ "special": true
138
+ },
139
+ "128017": {
140
+ "content": "<|reserved_special_token_12|>",
141
+ "lstrip": false,
142
+ "normalized": false,
143
+ "rstrip": false,
144
+ "single_word": false,
145
+ "special": true
146
+ },
147
+ "128018": {
148
+ "content": "<|reserved_special_token_13|>",
149
+ "lstrip": false,
150
+ "normalized": false,
151
+ "rstrip": false,
152
+ "single_word": false,
153
+ "special": true
154
+ },
155
+ "128019": {
156
+ "content": "<|reserved_special_token_14|>",
157
+ "lstrip": false,
158
+ "normalized": false,
159
+ "rstrip": false,
160
+ "single_word": false,
161
+ "special": true
162
+ },
163
+ "128020": {
164
+ "content": "<|reserved_special_token_15|>",
165
+ "lstrip": false,
166
+ "normalized": false,
167
+ "rstrip": false,
168
+ "single_word": false,
169
+ "special": true
170
+ },
171
+ "128021": {
172
+ "content": "<|reserved_special_token_16|>",
173
+ "lstrip": false,
174
+ "normalized": false,
175
+ "rstrip": false,
176
+ "single_word": false,
177
+ "special": true
178
+ },
179
+ "128022": {
180
+ "content": "<|reserved_special_token_17|>",
181
+ "lstrip": false,
182
+ "normalized": false,
183
+ "rstrip": false,
184
+ "single_word": false,
185
+ "special": true
186
+ },
187
+ "128023": {
188
+ "content": "<|reserved_special_token_18|>",
189
+ "lstrip": false,
190
+ "normalized": false,
191
+ "rstrip": false,
192
+ "single_word": false,
193
+ "special": true
194
+ },
195
+ "128024": {
196
+ "content": "<|reserved_special_token_19|>",
197
+ "lstrip": false,
198
+ "normalized": false,
199
+ "rstrip": false,
200
+ "single_word": false,
201
+ "special": true
202
+ },
203
+ "128025": {
204
+ "content": "<|reserved_special_token_20|>",
205
+ "lstrip": false,
206
+ "normalized": false,
207
+ "rstrip": false,
208
+ "single_word": false,
209
+ "special": true
210
+ },
211
+ "128026": {
212
+ "content": "<|reserved_special_token_21|>",
213
+ "lstrip": false,
214
+ "normalized": false,
215
+ "rstrip": false,
216
+ "single_word": false,
217
+ "special": true
218
+ },
219
+ "128027": {
220
+ "content": "<|reserved_special_token_22|>",
221
+ "lstrip": false,
222
+ "normalized": false,
223
+ "rstrip": false,
224
+ "single_word": false,
225
+ "special": true
226
+ },
227
+ "128028": {
228
+ "content": "<|reserved_special_token_23|>",
229
+ "lstrip": false,
230
+ "normalized": false,
231
+ "rstrip": false,
232
+ "single_word": false,
233
+ "special": true
234
+ },
235
+ "128029": {
236
+ "content": "<|reserved_special_token_24|>",
237
+ "lstrip": false,
238
+ "normalized": false,
239
+ "rstrip": false,
240
+ "single_word": false,
241
+ "special": true
242
+ },
243
+ "128030": {
244
+ "content": "<|reserved_special_token_25|>",
245
+ "lstrip": false,
246
+ "normalized": false,
247
+ "rstrip": false,
248
+ "single_word": false,
249
+ "special": true
250
+ },
251
+ "128031": {
252
+ "content": "<|reserved_special_token_26|>",
253
+ "lstrip": false,
254
+ "normalized": false,
255
+ "rstrip": false,
256
+ "single_word": false,
257
+ "special": true
258
+ },
259
+ "128032": {
260
+ "content": "<|reserved_special_token_27|>",
261
+ "lstrip": false,
262
+ "normalized": false,
263
+ "rstrip": false,
264
+ "single_word": false,
265
+ "special": true
266
+ },
267
+ "128033": {
268
+ "content": "<|reserved_special_token_28|>",
269
+ "lstrip": false,
270
+ "normalized": false,
271
+ "rstrip": false,
272
+ "single_word": false,
273
+ "special": true
274
+ },
275
+ "128034": {
276
+ "content": "<|reserved_special_token_29|>",
277
+ "lstrip": false,
278
+ "normalized": false,
279
+ "rstrip": false,
280
+ "single_word": false,
281
+ "special": true
282
+ },
283
+ "128035": {
284
+ "content": "<|reserved_special_token_30|>",
285
+ "lstrip": false,
286
+ "normalized": false,
287
+ "rstrip": false,
288
+ "single_word": false,
289
+ "special": true
290
+ },
291
+ "128036": {
292
+ "content": "<|reserved_special_token_31|>",
293
+ "lstrip": false,
294
+ "normalized": false,
295
+ "rstrip": false,
296
+ "single_word": false,
297
+ "special": true
298
+ },
299
+ "128037": {
300
+ "content": "<|reserved_special_token_32|>",
301
+ "lstrip": false,
302
+ "normalized": false,
303
+ "rstrip": false,
304
+ "single_word": false,
305
+ "special": true
306
+ },
307
+ "128038": {
308
+ "content": "<|reserved_special_token_33|>",
309
+ "lstrip": false,
310
+ "normalized": false,
311
+ "rstrip": false,
312
+ "single_word": false,
313
+ "special": true
314
+ },
315
+ "128039": {
316
+ "content": "<|reserved_special_token_34|>",
317
+ "lstrip": false,
318
+ "normalized": false,
319
+ "rstrip": false,
320
+ "single_word": false,
321
+ "special": true
322
+ },
323
+ "128040": {
324
+ "content": "<|reserved_special_token_35|>",
325
+ "lstrip": false,
326
+ "normalized": false,
327
+ "rstrip": false,
328
+ "single_word": false,
329
+ "special": true
330
+ },
331
+ "128041": {
332
+ "content": "<|reserved_special_token_36|>",
333
+ "lstrip": false,
334
+ "normalized": false,
335
+ "rstrip": false,
336
+ "single_word": false,
337
+ "special": true
338
+ },
339
+ "128042": {
340
+ "content": "<|reserved_special_token_37|>",
341
+ "lstrip": false,
342
+ "normalized": false,
343
+ "rstrip": false,
344
+ "single_word": false,
345
+ "special": true
346
+ },
347
+ "128043": {
348
+ "content": "<|reserved_special_token_38|>",
349
+ "lstrip": false,
350
+ "normalized": false,
351
+ "rstrip": false,
352
+ "single_word": false,
353
+ "special": true
354
+ },
355
+ "128044": {
356
+ "content": "<|reserved_special_token_39|>",
357
+ "lstrip": false,
358
+ "normalized": false,
359
+ "rstrip": false,
360
+ "single_word": false,
361
+ "special": true
362
+ },
363
+ "128045": {
364
+ "content": "<|reserved_special_token_40|>",
365
+ "lstrip": false,
366
+ "normalized": false,
367
+ "rstrip": false,
368
+ "single_word": false,
369
+ "special": true
370
+ },
371
+ "128046": {
372
+ "content": "<|reserved_special_token_41|>",
373
+ "lstrip": false,
374
+ "normalized": false,
375
+ "rstrip": false,
376
+ "single_word": false,
377
+ "special": true
378
+ },
379
+ "128047": {
380
+ "content": "<|reserved_special_token_42|>",
381
+ "lstrip": false,
382
+ "normalized": false,
383
+ "rstrip": false,
384
+ "single_word": false,
385
+ "special": true
386
+ },
387
+ "128048": {
388
+ "content": "<|reserved_special_token_43|>",
389
+ "lstrip": false,
390
+ "normalized": false,
391
+ "rstrip": false,
392
+ "single_word": false,
393
+ "special": true
394
+ },
395
+ "128049": {
396
+ "content": "<|reserved_special_token_44|>",
397
+ "lstrip": false,
398
+ "normalized": false,
399
+ "rstrip": false,
400
+ "single_word": false,
401
+ "special": true
402
+ },
403
+ "128050": {
404
+ "content": "<|reserved_special_token_45|>",
405
+ "lstrip": false,
406
+ "normalized": false,
407
+ "rstrip": false,
408
+ "single_word": false,
409
+ "special": true
410
+ },
411
+ "128051": {
412
+ "content": "<|reserved_special_token_46|>",
413
+ "lstrip": false,
414
+ "normalized": false,
415
+ "rstrip": false,
416
+ "single_word": false,
417
+ "special": true
418
+ },
419
+ "128052": {
420
+ "content": "<|reserved_special_token_47|>",
421
+ "lstrip": false,
422
+ "normalized": false,
423
+ "rstrip": false,
424
+ "single_word": false,
425
+ "special": true
426
+ },
427
+ "128053": {
428
+ "content": "<|reserved_special_token_48|>",
429
+ "lstrip": false,
430
+ "normalized": false,
431
+ "rstrip": false,
432
+ "single_word": false,
433
+ "special": true
434
+ },
435
+ "128054": {
436
+ "content": "<|reserved_special_token_49|>",
437
+ "lstrip": false,
438
+ "normalized": false,
439
+ "rstrip": false,
440
+ "single_word": false,
441
+ "special": true
442
+ },
443
+ "128055": {
444
+ "content": "<|reserved_special_token_50|>",
445
+ "lstrip": false,
446
+ "normalized": false,
447
+ "rstrip": false,
448
+ "single_word": false,
449
+ "special": true
450
+ },
451
+ "128056": {
452
+ "content": "<|reserved_special_token_51|>",
453
+ "lstrip": false,
454
+ "normalized": false,
455
+ "rstrip": false,
456
+ "single_word": false,
457
+ "special": true
458
+ },
459
+ "128057": {
460
+ "content": "<|reserved_special_token_52|>",
461
+ "lstrip": false,
462
+ "normalized": false,
463
+ "rstrip": false,
464
+ "single_word": false,
465
+ "special": true
466
+ },
467
+ "128058": {
468
+ "content": "<|reserved_special_token_53|>",
469
+ "lstrip": false,
470
+ "normalized": false,
471
+ "rstrip": false,
472
+ "single_word": false,
473
+ "special": true
474
+ },
475
+ "128059": {
476
+ "content": "<|reserved_special_token_54|>",
477
+ "lstrip": false,
478
+ "normalized": false,
479
+ "rstrip": false,
480
+ "single_word": false,
481
+ "special": true
482
+ },
483
+ "128060": {
484
+ "content": "<|reserved_special_token_55|>",
485
+ "lstrip": false,
486
+ "normalized": false,
487
+ "rstrip": false,
488
+ "single_word": false,
489
+ "special": true
490
+ },
491
+ "128061": {
492
+ "content": "<|reserved_special_token_56|>",
493
+ "lstrip": false,
494
+ "normalized": false,
495
+ "rstrip": false,
496
+ "single_word": false,
497
+ "special": true
498
+ },
499
+ "128062": {
500
+ "content": "<|reserved_special_token_57|>",
501
+ "lstrip": false,
502
+ "normalized": false,
503
+ "rstrip": false,
504
+ "single_word": false,
505
+ "special": true
506
+ },
507
+ "128063": {
508
+ "content": "<|reserved_special_token_58|>",
509
+ "lstrip": false,
510
+ "normalized": false,
511
+ "rstrip": false,
512
+ "single_word": false,
513
+ "special": true
514
+ },
515
+ "128064": {
516
+ "content": "<|reserved_special_token_59|>",
517
+ "lstrip": false,
518
+ "normalized": false,
519
+ "rstrip": false,
520
+ "single_word": false,
521
+ "special": true
522
+ },
523
+ "128065": {
524
+ "content": "<|reserved_special_token_60|>",
525
+ "lstrip": false,
526
+ "normalized": false,
527
+ "rstrip": false,
528
+ "single_word": false,
529
+ "special": true
530
+ },
531
+ "128066": {
532
+ "content": "<|reserved_special_token_61|>",
533
+ "lstrip": false,
534
+ "normalized": false,
535
+ "rstrip": false,
536
+ "single_word": false,
537
+ "special": true
538
+ },
539
+ "128067": {
540
+ "content": "<|reserved_special_token_62|>",
541
+ "lstrip": false,
542
+ "normalized": false,
543
+ "rstrip": false,
544
+ "single_word": false,
545
+ "special": true
546
+ },
547
+ "128068": {
548
+ "content": "<|reserved_special_token_63|>",
549
+ "lstrip": false,
550
+ "normalized": false,
551
+ "rstrip": false,
552
+ "single_word": false,
553
+ "special": true
554
+ },
555
+ "128069": {
556
+ "content": "<|reserved_special_token_64|>",
557
+ "lstrip": false,
558
+ "normalized": false,
559
+ "rstrip": false,
560
+ "single_word": false,
561
+ "special": true
562
+ },
563
+ "128070": {
564
+ "content": "<|reserved_special_token_65|>",
565
+ "lstrip": false,
566
+ "normalized": false,
567
+ "rstrip": false,
568
+ "single_word": false,
569
+ "special": true
570
+ },
571
+ "128071": {
572
+ "content": "<|reserved_special_token_66|>",
573
+ "lstrip": false,
574
+ "normalized": false,
575
+ "rstrip": false,
576
+ "single_word": false,
577
+ "special": true
578
+ },
579
+ "128072": {
580
+ "content": "<|reserved_special_token_67|>",
581
+ "lstrip": false,
582
+ "normalized": false,
583
+ "rstrip": false,
584
+ "single_word": false,
585
+ "special": true
586
+ },
587
+ "128073": {
588
+ "content": "<|reserved_special_token_68|>",
589
+ "lstrip": false,
590
+ "normalized": false,
591
+ "rstrip": false,
592
+ "single_word": false,
593
+ "special": true
594
+ },
595
+ "128074": {
596
+ "content": "<|reserved_special_token_69|>",
597
+ "lstrip": false,
598
+ "normalized": false,
599
+ "rstrip": false,
600
+ "single_word": false,
601
+ "special": true
602
+ },
603
+ "128075": {
604
+ "content": "<|reserved_special_token_70|>",
605
+ "lstrip": false,
606
+ "normalized": false,
607
+ "rstrip": false,
608
+ "single_word": false,
609
+ "special": true
610
+ },
611
+ "128076": {
612
+ "content": "<|reserved_special_token_71|>",
613
+ "lstrip": false,
614
+ "normalized": false,
615
+ "rstrip": false,
616
+ "single_word": false,
617
+ "special": true
618
+ },
619
+ "128077": {
620
+ "content": "<|reserved_special_token_72|>",
621
+ "lstrip": false,
622
+ "normalized": false,
623
+ "rstrip": false,
624
+ "single_word": false,
625
+ "special": true
626
+ },
627
+ "128078": {
628
+ "content": "<|reserved_special_token_73|>",
629
+ "lstrip": false,
630
+ "normalized": false,
631
+ "rstrip": false,
632
+ "single_word": false,
633
+ "special": true
634
+ },
635
+ "128079": {
636
+ "content": "<|reserved_special_token_74|>",
637
+ "lstrip": false,
638
+ "normalized": false,
639
+ "rstrip": false,
640
+ "single_word": false,
641
+ "special": true
642
+ },
643
+ "128080": {
644
+ "content": "<|reserved_special_token_75|>",
645
+ "lstrip": false,
646
+ "normalized": false,
647
+ "rstrip": false,
648
+ "single_word": false,
649
+ "special": true
650
+ },
651
+ "128081": {
652
+ "content": "<|reserved_special_token_76|>",
653
+ "lstrip": false,
654
+ "normalized": false,
655
+ "rstrip": false,
656
+ "single_word": false,
657
+ "special": true
658
+ },
659
+ "128082": {
660
+ "content": "<|reserved_special_token_77|>",
661
+ "lstrip": false,
662
+ "normalized": false,
663
+ "rstrip": false,
664
+ "single_word": false,
665
+ "special": true
666
+ },
667
+ "128083": {
668
+ "content": "<|reserved_special_token_78|>",
669
+ "lstrip": false,
670
+ "normalized": false,
671
+ "rstrip": false,
672
+ "single_word": false,
673
+ "special": true
674
+ },
675
+ "128084": {
676
+ "content": "<|reserved_special_token_79|>",
677
+ "lstrip": false,
678
+ "normalized": false,
679
+ "rstrip": false,
680
+ "single_word": false,
681
+ "special": true
682
+ },
683
+ "128085": {
684
+ "content": "<|reserved_special_token_80|>",
685
+ "lstrip": false,
686
+ "normalized": false,
687
+ "rstrip": false,
688
+ "single_word": false,
689
+ "special": true
690
+ },
691
+ "128086": {
692
+ "content": "<|reserved_special_token_81|>",
693
+ "lstrip": false,
694
+ "normalized": false,
695
+ "rstrip": false,
696
+ "single_word": false,
697
+ "special": true
698
+ },
699
+ "128087": {
700
+ "content": "<|reserved_special_token_82|>",
701
+ "lstrip": false,
702
+ "normalized": false,
703
+ "rstrip": false,
704
+ "single_word": false,
705
+ "special": true
706
+ },
707
+ "128088": {
708
+ "content": "<|reserved_special_token_83|>",
709
+ "lstrip": false,
710
+ "normalized": false,
711
+ "rstrip": false,
712
+ "single_word": false,
713
+ "special": true
714
+ },
715
+ "128089": {
716
+ "content": "<|reserved_special_token_84|>",
717
+ "lstrip": false,
718
+ "normalized": false,
719
+ "rstrip": false,
720
+ "single_word": false,
721
+ "special": true
722
+ },
723
+ "128090": {
724
+ "content": "<|reserved_special_token_85|>",
725
+ "lstrip": false,
726
+ "normalized": false,
727
+ "rstrip": false,
728
+ "single_word": false,
729
+ "special": true
730
+ },
731
+ "128091": {
732
+ "content": "<|reserved_special_token_86|>",
733
+ "lstrip": false,
734
+ "normalized": false,
735
+ "rstrip": false,
736
+ "single_word": false,
737
+ "special": true
738
+ },
739
+ "128092": {
740
+ "content": "<|reserved_special_token_87|>",
741
+ "lstrip": false,
742
+ "normalized": false,
743
+ "rstrip": false,
744
+ "single_word": false,
745
+ "special": true
746
+ },
747
+ "128093": {
748
+ "content": "<|reserved_special_token_88|>",
749
+ "lstrip": false,
750
+ "normalized": false,
751
+ "rstrip": false,
752
+ "single_word": false,
753
+ "special": true
754
+ },
755
+ "128094": {
756
+ "content": "<|reserved_special_token_89|>",
757
+ "lstrip": false,
758
+ "normalized": false,
759
+ "rstrip": false,
760
+ "single_word": false,
761
+ "special": true
762
+ },
763
+ "128095": {
764
+ "content": "<|reserved_special_token_90|>",
765
+ "lstrip": false,
766
+ "normalized": false,
767
+ "rstrip": false,
768
+ "single_word": false,
769
+ "special": true
770
+ },
771
+ "128096": {
772
+ "content": "<|reserved_special_token_91|>",
773
+ "lstrip": false,
774
+ "normalized": false,
775
+ "rstrip": false,
776
+ "single_word": false,
777
+ "special": true
778
+ },
779
+ "128097": {
780
+ "content": "<|reserved_special_token_92|>",
781
+ "lstrip": false,
782
+ "normalized": false,
783
+ "rstrip": false,
784
+ "single_word": false,
785
+ "special": true
786
+ },
787
+ "128098": {
788
+ "content": "<|reserved_special_token_93|>",
789
+ "lstrip": false,
790
+ "normalized": false,
791
+ "rstrip": false,
792
+ "single_word": false,
793
+ "special": true
794
+ },
795
+ "128099": {
796
+ "content": "<|reserved_special_token_94|>",
797
+ "lstrip": false,
798
+ "normalized": false,
799
+ "rstrip": false,
800
+ "single_word": false,
801
+ "special": true
802
+ },
803
+ "128100": {
804
+ "content": "<|reserved_special_token_95|>",
805
+ "lstrip": false,
806
+ "normalized": false,
807
+ "rstrip": false,
808
+ "single_word": false,
809
+ "special": true
810
+ },
811
+ "128101": {
812
+ "content": "<|reserved_special_token_96|>",
813
+ "lstrip": false,
814
+ "normalized": false,
815
+ "rstrip": false,
816
+ "single_word": false,
817
+ "special": true
818
+ },
819
+ "128102": {
820
+ "content": "<|reserved_special_token_97|>",
821
+ "lstrip": false,
822
+ "normalized": false,
823
+ "rstrip": false,
824
+ "single_word": false,
825
+ "special": true
826
+ },
827
+ "128103": {
828
+ "content": "<|reserved_special_token_98|>",
829
+ "lstrip": false,
830
+ "normalized": false,
831
+ "rstrip": false,
832
+ "single_word": false,
833
+ "special": true
834
+ },
835
+ "128104": {
836
+ "content": "<|reserved_special_token_99|>",
837
+ "lstrip": false,
838
+ "normalized": false,
839
+ "rstrip": false,
840
+ "single_word": false,
841
+ "special": true
842
+ },
843
+ "128105": {
844
+ "content": "<|reserved_special_token_100|>",
845
+ "lstrip": false,
846
+ "normalized": false,
847
+ "rstrip": false,
848
+ "single_word": false,
849
+ "special": true
850
+ },
851
+ "128106": {
852
+ "content": "<|reserved_special_token_101|>",
853
+ "lstrip": false,
854
+ "normalized": false,
855
+ "rstrip": false,
856
+ "single_word": false,
857
+ "special": true
858
+ },
859
+ "128107": {
860
+ "content": "<|reserved_special_token_102|>",
861
+ "lstrip": false,
862
+ "normalized": false,
863
+ "rstrip": false,
864
+ "single_word": false,
865
+ "special": true
866
+ },
867
+ "128108": {
868
+ "content": "<|reserved_special_token_103|>",
869
+ "lstrip": false,
870
+ "normalized": false,
871
+ "rstrip": false,
872
+ "single_word": false,
873
+ "special": true
874
+ },
875
+ "128109": {
876
+ "content": "<|reserved_special_token_104|>",
877
+ "lstrip": false,
878
+ "normalized": false,
879
+ "rstrip": false,
880
+ "single_word": false,
881
+ "special": true
882
+ },
883
+ "128110": {
884
+ "content": "<|reserved_special_token_105|>",
885
+ "lstrip": false,
886
+ "normalized": false,
887
+ "rstrip": false,
888
+ "single_word": false,
889
+ "special": true
890
+ },
891
+ "128111": {
892
+ "content": "<|reserved_special_token_106|>",
893
+ "lstrip": false,
894
+ "normalized": false,
895
+ "rstrip": false,
896
+ "single_word": false,
897
+ "special": true
898
+ },
899
+ "128112": {
900
+ "content": "<|reserved_special_token_107|>",
901
+ "lstrip": false,
902
+ "normalized": false,
903
+ "rstrip": false,
904
+ "single_word": false,
905
+ "special": true
906
+ },
907
+ "128113": {
908
+ "content": "<|reserved_special_token_108|>",
909
+ "lstrip": false,
910
+ "normalized": false,
911
+ "rstrip": false,
912
+ "single_word": false,
913
+ "special": true
914
+ },
915
+ "128114": {
916
+ "content": "<|reserved_special_token_109|>",
917
+ "lstrip": false,
918
+ "normalized": false,
919
+ "rstrip": false,
920
+ "single_word": false,
921
+ "special": true
922
+ },
923
+ "128115": {
924
+ "content": "<|reserved_special_token_110|>",
925
+ "lstrip": false,
926
+ "normalized": false,
927
+ "rstrip": false,
928
+ "single_word": false,
929
+ "special": true
930
+ },
931
+ "128116": {
932
+ "content": "<|reserved_special_token_111|>",
933
+ "lstrip": false,
934
+ "normalized": false,
935
+ "rstrip": false,
936
+ "single_word": false,
937
+ "special": true
938
+ },
939
+ "128117": {
940
+ "content": "<|reserved_special_token_112|>",
941
+ "lstrip": false,
942
+ "normalized": false,
943
+ "rstrip": false,
944
+ "single_word": false,
945
+ "special": true
946
+ },
947
+ "128118": {
948
+ "content": "<|reserved_special_token_113|>",
949
+ "lstrip": false,
950
+ "normalized": false,
951
+ "rstrip": false,
952
+ "single_word": false,
953
+ "special": true
954
+ },
955
+ "128119": {
956
+ "content": "<|reserved_special_token_114|>",
957
+ "lstrip": false,
958
+ "normalized": false,
959
+ "rstrip": false,
960
+ "single_word": false,
961
+ "special": true
962
+ },
963
+ "128120": {
964
+ "content": "<|reserved_special_token_115|>",
965
+ "lstrip": false,
966
+ "normalized": false,
967
+ "rstrip": false,
968
+ "single_word": false,
969
+ "special": true
970
+ },
971
+ "128121": {
972
+ "content": "<|reserved_special_token_116|>",
973
+ "lstrip": false,
974
+ "normalized": false,
975
+ "rstrip": false,
976
+ "single_word": false,
977
+ "special": true
978
+ },
979
+ "128122": {
980
+ "content": "<|reserved_special_token_117|>",
981
+ "lstrip": false,
982
+ "normalized": false,
983
+ "rstrip": false,
984
+ "single_word": false,
985
+ "special": true
986
+ },
987
+ "128123": {
988
+ "content": "<|reserved_special_token_118|>",
989
+ "lstrip": false,
990
+ "normalized": false,
991
+ "rstrip": false,
992
+ "single_word": false,
993
+ "special": true
994
+ },
995
+ "128124": {
996
+ "content": "<|reserved_special_token_119|>",
997
+ "lstrip": false,
998
+ "normalized": false,
999
+ "rstrip": false,
1000
+ "single_word": false,
1001
+ "special": true
1002
+ },
1003
+ "128125": {
1004
+ "content": "<|reserved_special_token_120|>",
1005
+ "lstrip": false,
1006
+ "normalized": false,
1007
+ "rstrip": false,
1008
+ "single_word": false,
1009
+ "special": true
1010
+ },
1011
+ "128126": {
1012
+ "content": "<|reserved_special_token_121|>",
1013
+ "lstrip": false,
1014
+ "normalized": false,
1015
+ "rstrip": false,
1016
+ "single_word": false,
1017
+ "special": true
1018
+ },
1019
+ "128127": {
1020
+ "content": "<|reserved_special_token_122|>",
1021
+ "lstrip": false,
1022
+ "normalized": false,
1023
+ "rstrip": false,
1024
+ "single_word": false,
1025
+ "special": true
1026
+ },
1027
+ "128128": {
1028
+ "content": "<|reserved_special_token_123|>",
1029
+ "lstrip": false,
1030
+ "normalized": false,
1031
+ "rstrip": false,
1032
+ "single_word": false,
1033
+ "special": true
1034
+ },
1035
+ "128129": {
1036
+ "content": "<|reserved_special_token_124|>",
1037
+ "lstrip": false,
1038
+ "normalized": false,
1039
+ "rstrip": false,
1040
+ "single_word": false,
1041
+ "special": true
1042
+ },
1043
+ "128130": {
1044
+ "content": "<|reserved_special_token_125|>",
1045
+ "lstrip": false,
1046
+ "normalized": false,
1047
+ "rstrip": false,
1048
+ "single_word": false,
1049
+ "special": true
1050
+ },
1051
+ "128131": {
1052
+ "content": "<|reserved_special_token_126|>",
1053
+ "lstrip": false,
1054
+ "normalized": false,
1055
+ "rstrip": false,
1056
+ "single_word": false,
1057
+ "special": true
1058
+ },
1059
+ "128132": {
1060
+ "content": "<|reserved_special_token_127|>",
1061
+ "lstrip": false,
1062
+ "normalized": false,
1063
+ "rstrip": false,
1064
+ "single_word": false,
1065
+ "special": true
1066
+ },
1067
+ "128133": {
1068
+ "content": "<|reserved_special_token_128|>",
1069
+ "lstrip": false,
1070
+ "normalized": false,
1071
+ "rstrip": false,
1072
+ "single_word": false,
1073
+ "special": true
1074
+ },
1075
+ "128134": {
1076
+ "content": "<|reserved_special_token_129|>",
1077
+ "lstrip": false,
1078
+ "normalized": false,
1079
+ "rstrip": false,
1080
+ "single_word": false,
1081
+ "special": true
1082
+ },
1083
+ "128135": {
1084
+ "content": "<|reserved_special_token_130|>",
1085
+ "lstrip": false,
1086
+ "normalized": false,
1087
+ "rstrip": false,
1088
+ "single_word": false,
1089
+ "special": true
1090
+ },
1091
+ "128136": {
1092
+ "content": "<|reserved_special_token_131|>",
1093
+ "lstrip": false,
1094
+ "normalized": false,
1095
+ "rstrip": false,
1096
+ "single_word": false,
1097
+ "special": true
1098
+ },
1099
+ "128137": {
1100
+ "content": "<|reserved_special_token_132|>",
1101
+ "lstrip": false,
1102
+ "normalized": false,
1103
+ "rstrip": false,
1104
+ "single_word": false,
1105
+ "special": true
1106
+ },
1107
+ "128138": {
1108
+ "content": "<|reserved_special_token_133|>",
1109
+ "lstrip": false,
1110
+ "normalized": false,
1111
+ "rstrip": false,
1112
+ "single_word": false,
1113
+ "special": true
1114
+ },
1115
+ "128139": {
1116
+ "content": "<|reserved_special_token_134|>",
1117
+ "lstrip": false,
1118
+ "normalized": false,
1119
+ "rstrip": false,
1120
+ "single_word": false,
1121
+ "special": true
1122
+ },
1123
+ "128140": {
1124
+ "content": "<|reserved_special_token_135|>",
1125
+ "lstrip": false,
1126
+ "normalized": false,
1127
+ "rstrip": false,
1128
+ "single_word": false,
1129
+ "special": true
1130
+ },
1131
+ "128141": {
1132
+ "content": "<|reserved_special_token_136|>",
1133
+ "lstrip": false,
1134
+ "normalized": false,
1135
+ "rstrip": false,
1136
+ "single_word": false,
1137
+ "special": true
1138
+ },
1139
+ "128142": {
1140
+ "content": "<|reserved_special_token_137|>",
1141
+ "lstrip": false,
1142
+ "normalized": false,
1143
+ "rstrip": false,
1144
+ "single_word": false,
1145
+ "special": true
1146
+ },
1147
+ "128143": {
1148
+ "content": "<|reserved_special_token_138|>",
1149
+ "lstrip": false,
1150
+ "normalized": false,
1151
+ "rstrip": false,
1152
+ "single_word": false,
1153
+ "special": true
1154
+ },
1155
+ "128144": {
1156
+ "content": "<|reserved_special_token_139|>",
1157
+ "lstrip": false,
1158
+ "normalized": false,
1159
+ "rstrip": false,
1160
+ "single_word": false,
1161
+ "special": true
1162
+ },
1163
+ "128145": {
1164
+ "content": "<|reserved_special_token_140|>",
1165
+ "lstrip": false,
1166
+ "normalized": false,
1167
+ "rstrip": false,
1168
+ "single_word": false,
1169
+ "special": true
1170
+ },
1171
+ "128146": {
1172
+ "content": "<|reserved_special_token_141|>",
1173
+ "lstrip": false,
1174
+ "normalized": false,
1175
+ "rstrip": false,
1176
+ "single_word": false,
1177
+ "special": true
1178
+ },
1179
+ "128147": {
1180
+ "content": "<|reserved_special_token_142|>",
1181
+ "lstrip": false,
1182
+ "normalized": false,
1183
+ "rstrip": false,
1184
+ "single_word": false,
1185
+ "special": true
1186
+ },
1187
+ "128148": {
1188
+ "content": "<|reserved_special_token_143|>",
1189
+ "lstrip": false,
1190
+ "normalized": false,
1191
+ "rstrip": false,
1192
+ "single_word": false,
1193
+ "special": true
1194
+ },
1195
+ "128149": {
1196
+ "content": "<|reserved_special_token_144|>",
1197
+ "lstrip": false,
1198
+ "normalized": false,
1199
+ "rstrip": false,
1200
+ "single_word": false,
1201
+ "special": true
1202
+ },
1203
+ "128150": {
1204
+ "content": "<|reserved_special_token_145|>",
1205
+ "lstrip": false,
1206
+ "normalized": false,
1207
+ "rstrip": false,
1208
+ "single_word": false,
1209
+ "special": true
1210
+ },
1211
+ "128151": {
1212
+ "content": "<|reserved_special_token_146|>",
1213
+ "lstrip": false,
1214
+ "normalized": false,
1215
+ "rstrip": false,
1216
+ "single_word": false,
1217
+ "special": true
1218
+ },
1219
+ "128152": {
1220
+ "content": "<|reserved_special_token_147|>",
1221
+ "lstrip": false,
1222
+ "normalized": false,
1223
+ "rstrip": false,
1224
+ "single_word": false,
1225
+ "special": true
1226
+ },
1227
+ "128153": {
1228
+ "content": "<|reserved_special_token_148|>",
1229
+ "lstrip": false,
1230
+ "normalized": false,
1231
+ "rstrip": false,
1232
+ "single_word": false,
1233
+ "special": true
1234
+ },
1235
+ "128154": {
1236
+ "content": "<|reserved_special_token_149|>",
1237
+ "lstrip": false,
1238
+ "normalized": false,
1239
+ "rstrip": false,
1240
+ "single_word": false,
1241
+ "special": true
1242
+ },
1243
+ "128155": {
1244
+ "content": "<|reserved_special_token_150|>",
1245
+ "lstrip": false,
1246
+ "normalized": false,
1247
+ "rstrip": false,
1248
+ "single_word": false,
1249
+ "special": true
1250
+ },
1251
+ "128156": {
1252
+ "content": "<|reserved_special_token_151|>",
1253
+ "lstrip": false,
1254
+ "normalized": false,
1255
+ "rstrip": false,
1256
+ "single_word": false,
1257
+ "special": true
1258
+ },
1259
+ "128157": {
1260
+ "content": "<|reserved_special_token_152|>",
1261
+ "lstrip": false,
1262
+ "normalized": false,
1263
+ "rstrip": false,
1264
+ "single_word": false,
1265
+ "special": true
1266
+ },
1267
+ "128158": {
1268
+ "content": "<|reserved_special_token_153|>",
1269
+ "lstrip": false,
1270
+ "normalized": false,
1271
+ "rstrip": false,
1272
+ "single_word": false,
1273
+ "special": true
1274
+ },
1275
+ "128159": {
1276
+ "content": "<|reserved_special_token_154|>",
1277
+ "lstrip": false,
1278
+ "normalized": false,
1279
+ "rstrip": false,
1280
+ "single_word": false,
1281
+ "special": true
1282
+ },
1283
+ "128160": {
1284
+ "content": "<|reserved_special_token_155|>",
1285
+ "lstrip": false,
1286
+ "normalized": false,
1287
+ "rstrip": false,
1288
+ "single_word": false,
1289
+ "special": true
1290
+ },
1291
+ "128161": {
1292
+ "content": "<|reserved_special_token_156|>",
1293
+ "lstrip": false,
1294
+ "normalized": false,
1295
+ "rstrip": false,
1296
+ "single_word": false,
1297
+ "special": true
1298
+ },
1299
+ "128162": {
1300
+ "content": "<|reserved_special_token_157|>",
1301
+ "lstrip": false,
1302
+ "normalized": false,
1303
+ "rstrip": false,
1304
+ "single_word": false,
1305
+ "special": true
1306
+ },
1307
+ "128163": {
1308
+ "content": "<|reserved_special_token_158|>",
1309
+ "lstrip": false,
1310
+ "normalized": false,
1311
+ "rstrip": false,
1312
+ "single_word": false,
1313
+ "special": true
1314
+ },
1315
+ "128164": {
1316
+ "content": "<|reserved_special_token_159|>",
1317
+ "lstrip": false,
1318
+ "normalized": false,
1319
+ "rstrip": false,
1320
+ "single_word": false,
1321
+ "special": true
1322
+ },
1323
+ "128165": {
1324
+ "content": "<|reserved_special_token_160|>",
1325
+ "lstrip": false,
1326
+ "normalized": false,
1327
+ "rstrip": false,
1328
+ "single_word": false,
1329
+ "special": true
1330
+ },
1331
+ "128166": {
1332
+ "content": "<|reserved_special_token_161|>",
1333
+ "lstrip": false,
1334
+ "normalized": false,
1335
+ "rstrip": false,
1336
+ "single_word": false,
1337
+ "special": true
1338
+ },
1339
+ "128167": {
1340
+ "content": "<|reserved_special_token_162|>",
1341
+ "lstrip": false,
1342
+ "normalized": false,
1343
+ "rstrip": false,
1344
+ "single_word": false,
1345
+ "special": true
1346
+ },
1347
+ "128168": {
1348
+ "content": "<|reserved_special_token_163|>",
1349
+ "lstrip": false,
1350
+ "normalized": false,
1351
+ "rstrip": false,
1352
+ "single_word": false,
1353
+ "special": true
1354
+ },
1355
+ "128169": {
1356
+ "content": "<|reserved_special_token_164|>",
1357
+ "lstrip": false,
1358
+ "normalized": false,
1359
+ "rstrip": false,
1360
+ "single_word": false,
1361
+ "special": true
1362
+ },
1363
+ "128170": {
1364
+ "content": "<|reserved_special_token_165|>",
1365
+ "lstrip": false,
1366
+ "normalized": false,
1367
+ "rstrip": false,
1368
+ "single_word": false,
1369
+ "special": true
1370
+ },
1371
+ "128171": {
1372
+ "content": "<|reserved_special_token_166|>",
1373
+ "lstrip": false,
1374
+ "normalized": false,
1375
+ "rstrip": false,
1376
+ "single_word": false,
1377
+ "special": true
1378
+ },
1379
+ "128172": {
1380
+ "content": "<|reserved_special_token_167|>",
1381
+ "lstrip": false,
1382
+ "normalized": false,
1383
+ "rstrip": false,
1384
+ "single_word": false,
1385
+ "special": true
1386
+ },
1387
+ "128173": {
1388
+ "content": "<|reserved_special_token_168|>",
1389
+ "lstrip": false,
1390
+ "normalized": false,
1391
+ "rstrip": false,
1392
+ "single_word": false,
1393
+ "special": true
1394
+ },
1395
+ "128174": {
1396
+ "content": "<|reserved_special_token_169|>",
1397
+ "lstrip": false,
1398
+ "normalized": false,
1399
+ "rstrip": false,
1400
+ "single_word": false,
1401
+ "special": true
1402
+ },
1403
+ "128175": {
1404
+ "content": "<|reserved_special_token_170|>",
1405
+ "lstrip": false,
1406
+ "normalized": false,
1407
+ "rstrip": false,
1408
+ "single_word": false,
1409
+ "special": true
1410
+ },
1411
+ "128176": {
1412
+ "content": "<|reserved_special_token_171|>",
1413
+ "lstrip": false,
1414
+ "normalized": false,
1415
+ "rstrip": false,
1416
+ "single_word": false,
1417
+ "special": true
1418
+ },
1419
+ "128177": {
1420
+ "content": "<|reserved_special_token_172|>",
1421
+ "lstrip": false,
1422
+ "normalized": false,
1423
+ "rstrip": false,
1424
+ "single_word": false,
1425
+ "special": true
1426
+ },
1427
+ "128178": {
1428
+ "content": "<|reserved_special_token_173|>",
1429
+ "lstrip": false,
1430
+ "normalized": false,
1431
+ "rstrip": false,
1432
+ "single_word": false,
1433
+ "special": true
1434
+ },
1435
+ "128179": {
1436
+ "content": "<|reserved_special_token_174|>",
1437
+ "lstrip": false,
1438
+ "normalized": false,
1439
+ "rstrip": false,
1440
+ "single_word": false,
1441
+ "special": true
1442
+ },
1443
+ "128180": {
1444
+ "content": "<|reserved_special_token_175|>",
1445
+ "lstrip": false,
1446
+ "normalized": false,
1447
+ "rstrip": false,
1448
+ "single_word": false,
1449
+ "special": true
1450
+ },
1451
+ "128181": {
1452
+ "content": "<|reserved_special_token_176|>",
1453
+ "lstrip": false,
1454
+ "normalized": false,
1455
+ "rstrip": false,
1456
+ "single_word": false,
1457
+ "special": true
1458
+ },
1459
+ "128182": {
1460
+ "content": "<|reserved_special_token_177|>",
1461
+ "lstrip": false,
1462
+ "normalized": false,
1463
+ "rstrip": false,
1464
+ "single_word": false,
1465
+ "special": true
1466
+ },
1467
+ "128183": {
1468
+ "content": "<|reserved_special_token_178|>",
1469
+ "lstrip": false,
1470
+ "normalized": false,
1471
+ "rstrip": false,
1472
+ "single_word": false,
1473
+ "special": true
1474
+ },
1475
+ "128184": {
1476
+ "content": "<|reserved_special_token_179|>",
1477
+ "lstrip": false,
1478
+ "normalized": false,
1479
+ "rstrip": false,
1480
+ "single_word": false,
1481
+ "special": true
1482
+ },
1483
+ "128185": {
1484
+ "content": "<|reserved_special_token_180|>",
1485
+ "lstrip": false,
1486
+ "normalized": false,
1487
+ "rstrip": false,
1488
+ "single_word": false,
1489
+ "special": true
1490
+ },
1491
+ "128186": {
1492
+ "content": "<|reserved_special_token_181|>",
1493
+ "lstrip": false,
1494
+ "normalized": false,
1495
+ "rstrip": false,
1496
+ "single_word": false,
1497
+ "special": true
1498
+ },
1499
+ "128187": {
1500
+ "content": "<|reserved_special_token_182|>",
1501
+ "lstrip": false,
1502
+ "normalized": false,
1503
+ "rstrip": false,
1504
+ "single_word": false,
1505
+ "special": true
1506
+ },
1507
+ "128188": {
1508
+ "content": "<|reserved_special_token_183|>",
1509
+ "lstrip": false,
1510
+ "normalized": false,
1511
+ "rstrip": false,
1512
+ "single_word": false,
1513
+ "special": true
1514
+ },
1515
+ "128189": {
1516
+ "content": "<|reserved_special_token_184|>",
1517
+ "lstrip": false,
1518
+ "normalized": false,
1519
+ "rstrip": false,
1520
+ "single_word": false,
1521
+ "special": true
1522
+ },
1523
+ "128190": {
1524
+ "content": "<|reserved_special_token_185|>",
1525
+ "lstrip": false,
1526
+ "normalized": false,
1527
+ "rstrip": false,
1528
+ "single_word": false,
1529
+ "special": true
1530
+ },
1531
+ "128191": {
1532
+ "content": "<|reserved_special_token_186|>",
1533
+ "lstrip": false,
1534
+ "normalized": false,
1535
+ "rstrip": false,
1536
+ "single_word": false,
1537
+ "special": true
1538
+ },
1539
+ "128192": {
1540
+ "content": "<|reserved_special_token_187|>",
1541
+ "lstrip": false,
1542
+ "normalized": false,
1543
+ "rstrip": false,
1544
+ "single_word": false,
1545
+ "special": true
1546
+ },
1547
+ "128193": {
1548
+ "content": "<|reserved_special_token_188|>",
1549
+ "lstrip": false,
1550
+ "normalized": false,
1551
+ "rstrip": false,
1552
+ "single_word": false,
1553
+ "special": true
1554
+ },
1555
+ "128194": {
1556
+ "content": "<|reserved_special_token_189|>",
1557
+ "lstrip": false,
1558
+ "normalized": false,
1559
+ "rstrip": false,
1560
+ "single_word": false,
1561
+ "special": true
1562
+ },
1563
+ "128195": {
1564
+ "content": "<|reserved_special_token_190|>",
1565
+ "lstrip": false,
1566
+ "normalized": false,
1567
+ "rstrip": false,
1568
+ "single_word": false,
1569
+ "special": true
1570
+ },
1571
+ "128196": {
1572
+ "content": "<|reserved_special_token_191|>",
1573
+ "lstrip": false,
1574
+ "normalized": false,
1575
+ "rstrip": false,
1576
+ "single_word": false,
1577
+ "special": true
1578
+ },
1579
+ "128197": {
1580
+ "content": "<|reserved_special_token_192|>",
1581
+ "lstrip": false,
1582
+ "normalized": false,
1583
+ "rstrip": false,
1584
+ "single_word": false,
1585
+ "special": true
1586
+ },
1587
+ "128198": {
1588
+ "content": "<|reserved_special_token_193|>",
1589
+ "lstrip": false,
1590
+ "normalized": false,
1591
+ "rstrip": false,
1592
+ "single_word": false,
1593
+ "special": true
1594
+ },
1595
+ "128199": {
1596
+ "content": "<|reserved_special_token_194|>",
1597
+ "lstrip": false,
1598
+ "normalized": false,
1599
+ "rstrip": false,
1600
+ "single_word": false,
1601
+ "special": true
1602
+ },
1603
+ "128200": {
1604
+ "content": "<|reserved_special_token_195|>",
1605
+ "lstrip": false,
1606
+ "normalized": false,
1607
+ "rstrip": false,
1608
+ "single_word": false,
1609
+ "special": true
1610
+ },
1611
+ "128201": {
1612
+ "content": "<|reserved_special_token_196|>",
1613
+ "lstrip": false,
1614
+ "normalized": false,
1615
+ "rstrip": false,
1616
+ "single_word": false,
1617
+ "special": true
1618
+ },
1619
+ "128202": {
1620
+ "content": "<|reserved_special_token_197|>",
1621
+ "lstrip": false,
1622
+ "normalized": false,
1623
+ "rstrip": false,
1624
+ "single_word": false,
1625
+ "special": true
1626
+ },
1627
+ "128203": {
1628
+ "content": "<|reserved_special_token_198|>",
1629
+ "lstrip": false,
1630
+ "normalized": false,
1631
+ "rstrip": false,
1632
+ "single_word": false,
1633
+ "special": true
1634
+ },
1635
+ "128204": {
1636
+ "content": "<|reserved_special_token_199|>",
1637
+ "lstrip": false,
1638
+ "normalized": false,
1639
+ "rstrip": false,
1640
+ "single_word": false,
1641
+ "special": true
1642
+ },
1643
+ "128205": {
1644
+ "content": "<|reserved_special_token_200|>",
1645
+ "lstrip": false,
1646
+ "normalized": false,
1647
+ "rstrip": false,
1648
+ "single_word": false,
1649
+ "special": true
1650
+ },
1651
+ "128206": {
1652
+ "content": "<|reserved_special_token_201|>",
1653
+ "lstrip": false,
1654
+ "normalized": false,
1655
+ "rstrip": false,
1656
+ "single_word": false,
1657
+ "special": true
1658
+ },
1659
+ "128207": {
1660
+ "content": "<|reserved_special_token_202|>",
1661
+ "lstrip": false,
1662
+ "normalized": false,
1663
+ "rstrip": false,
1664
+ "single_word": false,
1665
+ "special": true
1666
+ },
1667
+ "128208": {
1668
+ "content": "<|reserved_special_token_203|>",
1669
+ "lstrip": false,
1670
+ "normalized": false,
1671
+ "rstrip": false,
1672
+ "single_word": false,
1673
+ "special": true
1674
+ },
1675
+ "128209": {
1676
+ "content": "<|reserved_special_token_204|>",
1677
+ "lstrip": false,
1678
+ "normalized": false,
1679
+ "rstrip": false,
1680
+ "single_word": false,
1681
+ "special": true
1682
+ },
1683
+ "128210": {
1684
+ "content": "<|reserved_special_token_205|>",
1685
+ "lstrip": false,
1686
+ "normalized": false,
1687
+ "rstrip": false,
1688
+ "single_word": false,
1689
+ "special": true
1690
+ },
1691
+ "128211": {
1692
+ "content": "<|reserved_special_token_206|>",
1693
+ "lstrip": false,
1694
+ "normalized": false,
1695
+ "rstrip": false,
1696
+ "single_word": false,
1697
+ "special": true
1698
+ },
1699
+ "128212": {
1700
+ "content": "<|reserved_special_token_207|>",
1701
+ "lstrip": false,
1702
+ "normalized": false,
1703
+ "rstrip": false,
1704
+ "single_word": false,
1705
+ "special": true
1706
+ },
1707
+ "128213": {
1708
+ "content": "<|reserved_special_token_208|>",
1709
+ "lstrip": false,
1710
+ "normalized": false,
1711
+ "rstrip": false,
1712
+ "single_word": false,
1713
+ "special": true
1714
+ },
1715
+ "128214": {
1716
+ "content": "<|reserved_special_token_209|>",
1717
+ "lstrip": false,
1718
+ "normalized": false,
1719
+ "rstrip": false,
1720
+ "single_word": false,
1721
+ "special": true
1722
+ },
1723
+ "128215": {
1724
+ "content": "<|reserved_special_token_210|>",
1725
+ "lstrip": false,
1726
+ "normalized": false,
1727
+ "rstrip": false,
1728
+ "single_word": false,
1729
+ "special": true
1730
+ },
1731
+ "128216": {
1732
+ "content": "<|reserved_special_token_211|>",
1733
+ "lstrip": false,
1734
+ "normalized": false,
1735
+ "rstrip": false,
1736
+ "single_word": false,
1737
+ "special": true
1738
+ },
1739
+ "128217": {
1740
+ "content": "<|reserved_special_token_212|>",
1741
+ "lstrip": false,
1742
+ "normalized": false,
1743
+ "rstrip": false,
1744
+ "single_word": false,
1745
+ "special": true
1746
+ },
1747
+ "128218": {
1748
+ "content": "<|reserved_special_token_213|>",
1749
+ "lstrip": false,
1750
+ "normalized": false,
1751
+ "rstrip": false,
1752
+ "single_word": false,
1753
+ "special": true
1754
+ },
1755
+ "128219": {
1756
+ "content": "<|reserved_special_token_214|>",
1757
+ "lstrip": false,
1758
+ "normalized": false,
1759
+ "rstrip": false,
1760
+ "single_word": false,
1761
+ "special": true
1762
+ },
1763
+ "128220": {
1764
+ "content": "<|reserved_special_token_215|>",
1765
+ "lstrip": false,
1766
+ "normalized": false,
1767
+ "rstrip": false,
1768
+ "single_word": false,
1769
+ "special": true
1770
+ },
1771
+ "128221": {
1772
+ "content": "<|reserved_special_token_216|>",
1773
+ "lstrip": false,
1774
+ "normalized": false,
1775
+ "rstrip": false,
1776
+ "single_word": false,
1777
+ "special": true
1778
+ },
1779
+ "128222": {
1780
+ "content": "<|reserved_special_token_217|>",
1781
+ "lstrip": false,
1782
+ "normalized": false,
1783
+ "rstrip": false,
1784
+ "single_word": false,
1785
+ "special": true
1786
+ },
1787
+ "128223": {
1788
+ "content": "<|reserved_special_token_218|>",
1789
+ "lstrip": false,
1790
+ "normalized": false,
1791
+ "rstrip": false,
1792
+ "single_word": false,
1793
+ "special": true
1794
+ },
1795
+ "128224": {
1796
+ "content": "<|reserved_special_token_219|>",
1797
+ "lstrip": false,
1798
+ "normalized": false,
1799
+ "rstrip": false,
1800
+ "single_word": false,
1801
+ "special": true
1802
+ },
1803
+ "128225": {
1804
+ "content": "<|reserved_special_token_220|>",
1805
+ "lstrip": false,
1806
+ "normalized": false,
1807
+ "rstrip": false,
1808
+ "single_word": false,
1809
+ "special": true
1810
+ },
1811
+ "128226": {
1812
+ "content": "<|reserved_special_token_221|>",
1813
+ "lstrip": false,
1814
+ "normalized": false,
1815
+ "rstrip": false,
1816
+ "single_word": false,
1817
+ "special": true
1818
+ },
1819
+ "128227": {
1820
+ "content": "<|reserved_special_token_222|>",
1821
+ "lstrip": false,
1822
+ "normalized": false,
1823
+ "rstrip": false,
1824
+ "single_word": false,
1825
+ "special": true
1826
+ },
1827
+ "128228": {
1828
+ "content": "<|reserved_special_token_223|>",
1829
+ "lstrip": false,
1830
+ "normalized": false,
1831
+ "rstrip": false,
1832
+ "single_word": false,
1833
+ "special": true
1834
+ },
1835
+ "128229": {
1836
+ "content": "<|reserved_special_token_224|>",
1837
+ "lstrip": false,
1838
+ "normalized": false,
1839
+ "rstrip": false,
1840
+ "single_word": false,
1841
+ "special": true
1842
+ },
1843
+ "128230": {
1844
+ "content": "<|reserved_special_token_225|>",
1845
+ "lstrip": false,
1846
+ "normalized": false,
1847
+ "rstrip": false,
1848
+ "single_word": false,
1849
+ "special": true
1850
+ },
1851
+ "128231": {
1852
+ "content": "<|reserved_special_token_226|>",
1853
+ "lstrip": false,
1854
+ "normalized": false,
1855
+ "rstrip": false,
1856
+ "single_word": false,
1857
+ "special": true
1858
+ },
1859
+ "128232": {
1860
+ "content": "<|reserved_special_token_227|>",
1861
+ "lstrip": false,
1862
+ "normalized": false,
1863
+ "rstrip": false,
1864
+ "single_word": false,
1865
+ "special": true
1866
+ },
1867
+ "128233": {
1868
+ "content": "<|reserved_special_token_228|>",
1869
+ "lstrip": false,
1870
+ "normalized": false,
1871
+ "rstrip": false,
1872
+ "single_word": false,
1873
+ "special": true
1874
+ },
1875
+ "128234": {
1876
+ "content": "<|reserved_special_token_229|>",
1877
+ "lstrip": false,
1878
+ "normalized": false,
1879
+ "rstrip": false,
1880
+ "single_word": false,
1881
+ "special": true
1882
+ },
1883
+ "128235": {
1884
+ "content": "<|reserved_special_token_230|>",
1885
+ "lstrip": false,
1886
+ "normalized": false,
1887
+ "rstrip": false,
1888
+ "single_word": false,
1889
+ "special": true
1890
+ },
1891
+ "128236": {
1892
+ "content": "<|reserved_special_token_231|>",
1893
+ "lstrip": false,
1894
+ "normalized": false,
1895
+ "rstrip": false,
1896
+ "single_word": false,
1897
+ "special": true
1898
+ },
1899
+ "128237": {
1900
+ "content": "<|reserved_special_token_232|>",
1901
+ "lstrip": false,
1902
+ "normalized": false,
1903
+ "rstrip": false,
1904
+ "single_word": false,
1905
+ "special": true
1906
+ },
1907
+ "128238": {
1908
+ "content": "<|reserved_special_token_233|>",
1909
+ "lstrip": false,
1910
+ "normalized": false,
1911
+ "rstrip": false,
1912
+ "single_word": false,
1913
+ "special": true
1914
+ },
1915
+ "128239": {
1916
+ "content": "<|reserved_special_token_234|>",
1917
+ "lstrip": false,
1918
+ "normalized": false,
1919
+ "rstrip": false,
1920
+ "single_word": false,
1921
+ "special": true
1922
+ },
1923
+ "128240": {
1924
+ "content": "<|reserved_special_token_235|>",
1925
+ "lstrip": false,
1926
+ "normalized": false,
1927
+ "rstrip": false,
1928
+ "single_word": false,
1929
+ "special": true
1930
+ },
1931
+ "128241": {
1932
+ "content": "<|reserved_special_token_236|>",
1933
+ "lstrip": false,
1934
+ "normalized": false,
1935
+ "rstrip": false,
1936
+ "single_word": false,
1937
+ "special": true
1938
+ },
1939
+ "128242": {
1940
+ "content": "<|reserved_special_token_237|>",
1941
+ "lstrip": false,
1942
+ "normalized": false,
1943
+ "rstrip": false,
1944
+ "single_word": false,
1945
+ "special": true
1946
+ },
1947
+ "128243": {
1948
+ "content": "<|reserved_special_token_238|>",
1949
+ "lstrip": false,
1950
+ "normalized": false,
1951
+ "rstrip": false,
1952
+ "single_word": false,
1953
+ "special": true
1954
+ },
1955
+ "128244": {
1956
+ "content": "<|reserved_special_token_239|>",
1957
+ "lstrip": false,
1958
+ "normalized": false,
1959
+ "rstrip": false,
1960
+ "single_word": false,
1961
+ "special": true
1962
+ },
1963
+ "128245": {
1964
+ "content": "<|reserved_special_token_240|>",
1965
+ "lstrip": false,
1966
+ "normalized": false,
1967
+ "rstrip": false,
1968
+ "single_word": false,
1969
+ "special": true
1970
+ },
1971
+ "128246": {
1972
+ "content": "<|reserved_special_token_241|>",
1973
+ "lstrip": false,
1974
+ "normalized": false,
1975
+ "rstrip": false,
1976
+ "single_word": false,
1977
+ "special": true
1978
+ },
1979
+ "128247": {
1980
+ "content": "<|reserved_special_token_242|>",
1981
+ "lstrip": false,
1982
+ "normalized": false,
1983
+ "rstrip": false,
1984
+ "single_word": false,
1985
+ "special": true
1986
+ },
1987
+ "128248": {
1988
+ "content": "<|reserved_special_token_243|>",
1989
+ "lstrip": false,
1990
+ "normalized": false,
1991
+ "rstrip": false,
1992
+ "single_word": false,
1993
+ "special": true
1994
+ },
1995
+ "128249": {
1996
+ "content": "<|reserved_special_token_244|>",
1997
+ "lstrip": false,
1998
+ "normalized": false,
1999
+ "rstrip": false,
2000
+ "single_word": false,
2001
+ "special": true
2002
+ },
2003
+ "128250": {
2004
+ "content": "<|reserved_special_token_245|>",
2005
+ "lstrip": false,
2006
+ "normalized": false,
2007
+ "rstrip": false,
2008
+ "single_word": false,
2009
+ "special": true
2010
+ },
2011
+ "128251": {
2012
+ "content": "<|reserved_special_token_246|>",
2013
+ "lstrip": false,
2014
+ "normalized": false,
2015
+ "rstrip": false,
2016
+ "single_word": false,
2017
+ "special": true
2018
+ },
2019
+ "128252": {
2020
+ "content": "<|reserved_special_token_247|>",
2021
+ "lstrip": false,
2022
+ "normalized": false,
2023
+ "rstrip": false,
2024
+ "single_word": false,
2025
+ "special": true
2026
+ },
2027
+ "128253": {
2028
+ "content": "<|reserved_special_token_248|>",
2029
+ "lstrip": false,
2030
+ "normalized": false,
2031
+ "rstrip": false,
2032
+ "single_word": false,
2033
+ "special": true
2034
+ },
2035
+ "128254": {
2036
+ "content": "<|reserved_special_token_249|>",
2037
+ "lstrip": false,
2038
+ "normalized": false,
2039
+ "rstrip": false,
2040
+ "single_word": false,
2041
+ "special": true
2042
+ },
2043
+ "128255": {
2044
+ "content": "<|reserved_special_token_250|>",
2045
+ "lstrip": false,
2046
+ "normalized": false,
2047
+ "rstrip": false,
2048
+ "single_word": false,
2049
+ "special": true
2050
+ },
2051
+ "128256": {
2052
+ "content": "<pad>",
2053
+ "lstrip": false,
2054
+ "normalized": false,
2055
+ "rstrip": false,
2056
+ "single_word": false,
2057
+ "special": true
2058
+ }
2059
+ },
2060
+ "bos_token": "<|begin_of_text|>",
2061
+ "clean_up_tokenization_spaces": true,
2062
+ "eos_token": "<|end_of_text|>",
2063
+ "model_input_names": [
2064
+ "input_ids",
2065
+ "attention_mask"
2066
+ ],
2067
+ "model_max_length": 3072,
2068
+ "pad_token": "<pad>",
2069
+ "padding_side": "right",
2070
+ "tokenizer_class": "PreTrainedTokenizerFast"
2071
+ }