Text Generation
Transformers
Safetensors
Korean
llama
korean
instruction-tuning
supervised-fine-tuning
merged-model
source-screening
medical-qa
conversational
text-generation-inference
Instructions to use youngseok12/AX-3.1-Light-sft_source_screen_71875_3000 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use youngseok12/AX-3.1-Light-sft_source_screen_71875_3000 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="youngseok12/AX-3.1-Light-sft_source_screen_71875_3000") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("youngseok12/AX-3.1-Light-sft_source_screen_71875_3000") model = AutoModelForCausalLM.from_pretrained("youngseok12/AX-3.1-Light-sft_source_screen_71875_3000", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use youngseok12/AX-3.1-Light-sft_source_screen_71875_3000 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "youngseok12/AX-3.1-Light-sft_source_screen_71875_3000" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "youngseok12/AX-3.1-Light-sft_source_screen_71875_3000", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/youngseok12/AX-3.1-Light-sft_source_screen_71875_3000
- SGLang
How to use youngseok12/AX-3.1-Light-sft_source_screen_71875_3000 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "youngseok12/AX-3.1-Light-sft_source_screen_71875_3000" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "youngseok12/AX-3.1-Light-sft_source_screen_71875_3000", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "youngseok12/AX-3.1-Light-sft_source_screen_71875_3000" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "youngseok12/AX-3.1-Light-sft_source_screen_71875_3000", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use youngseok12/AX-3.1-Light-sft_source_screen_71875_3000 with Docker Model Runner:
docker model run hf.co/youngseok12/AX-3.1-Light-sft_source_screen_71875_3000
upload source-screen 71875 merged model
Browse files- LICENSE +16 -0
- README.md +106 -0
- chat_template.jinja +72 -0
- config.json +32 -0
- generation_config.json +7 -0
- kds_merge_info.json +5 -0
- model.safetensors +3 -0
- tokenizer.json +0 -0
- tokenizer_config.json +23 -0
LICENSE
ADDED
|
@@ -0,0 +1,16 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Copyright (c) 2025 SK Telecom Co., Ltd. All rights reserved.
|
| 2 |
+
|
| 3 |
+
Unless otherwise stated, all files in this repository (including modified model
|
| 4 |
+
weights and tokenizer files) are distributed under the terms of the Apache
|
| 5 |
+
License, Version 2.0 (the "License"). You may obtain a copy of the License at:
|
| 6 |
+
|
| 7 |
+
http://www.apache.org/licenses/LICENSE-2.0
|
| 8 |
+
|
| 9 |
+
Unless required by applicable law or agreed to in writing, software distributed
|
| 10 |
+
under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES OR
|
| 11 |
+
CONDITIONS OF ANY KIND, either express or implied. See the License for the
|
| 12 |
+
specific language governing permissions and limitations under the License.
|
| 13 |
+
|
| 14 |
+
"SK Telecom" and associated logos are trademarks of SK Telecom Co., Ltd. This
|
| 15 |
+
License does not grant permission to use these trademarks without prior written
|
| 16 |
+
consent.
|
README.md
ADDED
|
@@ -0,0 +1,106 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
license_link: https://www.apache.org/licenses/LICENSE-2.0
|
| 4 |
+
base_model: skt/A.X-3.1-Light
|
| 5 |
+
language:
|
| 6 |
+
- ko
|
| 7 |
+
library_name: transformers
|
| 8 |
+
pipeline_tag: text-generation
|
| 9 |
+
tags:
|
| 10 |
+
- korean
|
| 11 |
+
- text-generation
|
| 12 |
+
- instruction-tuning
|
| 13 |
+
- supervised-fine-tuning
|
| 14 |
+
- merged-model
|
| 15 |
+
- source-screening
|
| 16 |
+
- medical-qa
|
| 17 |
+
---
|
| 18 |
+
|
| 19 |
+
# A.X-3.1-Light SFT Source Screen 71875 (Essential Medical 3K)
|
| 20 |
+
|
| 21 |
+
이 모델은 `skt/A.X-3.1-Light`를 기반으로 AI Hub 71875 필수의료 의학지식
|
| 22 |
+
데이터만 사용해 한국어 질의응답을 LoRA 방식으로 1 epoch 지도학습한
|
| 23 |
+
모델입니다. 학습이 끝난 뒤 LoRA adapter를 base model에 병합한 BF16
|
| 24 |
+
standalone 전체 가중치 모델이므로 추론 시 별도의 adapter가 필요하지
|
| 25 |
+
않습니다. 연구 및 통제된 평가용 모델이며, 의료 관련 답변을 포함한 모든
|
| 26 |
+
생성 결과는 오류가 있을 수 있어 전문적인 판단을 대체할 수 없습니다.
|
| 27 |
+
|
| 28 |
+
## Model details
|
| 29 |
+
|
| 30 |
+
- Model name: `A.X-3.1-Light SFT Source Screen 71875 (Essential Medical 3K)`
|
| 31 |
+
- Base model: [`skt/A.X-3.1-Light`](https://huggingface.co/skt/A.X-3.1-Light)
|
| 32 |
+
- Base model revision: `9b41bb2406472634d8812c0b8931fa40fa9a6c3a`
|
| 33 |
+
- Fine-tuning: LoRA supervised fine-tuning, merged into base weights
|
| 34 |
+
- Weight format: BF16 `safetensors`
|
| 35 |
+
- Architecture: unchanged Llama causal language model architecture
|
| 36 |
+
- Chat template: bundled A.X tokenizer chat template
|
| 37 |
+
- Custom model code: none; standard Transformers/vLLM loading is intended
|
| 38 |
+
|
| 39 |
+
## Training data
|
| 40 |
+
|
| 41 |
+
Training used only AI Hub dataset **71875, 필수의료 의학지식 데이터**. The
|
| 42 |
+
training split contains 3,000 selected examples and the separate development
|
| 43 |
+
split contains 300 examples. No v0.21 mixture, other AI Hub dataset, public
|
| 44 |
+
benchmark question, benchmark answer, or evaluation artifact was used as SFT
|
| 45 |
+
data or included in this repository.
|
| 46 |
+
|
| 47 |
+
| Source | Training examples |
|
| 48 |
+
|---|---:|
|
| 49 |
+
| Category 14 | 895 |
|
| 50 |
+
| Category 15 | 895 |
|
| 51 |
+
| Category 16 | 315 |
|
| 52 |
+
| Category 17 | 895 |
|
| 53 |
+
| **Total** | **3,000** |
|
| 54 |
+
|
| 55 |
+
The output contract is answer-first (`정답: ...`). Depending on the source
|
| 56 |
+
question, the target is a label, number, short answer, or concise explanation.
|
| 57 |
+
Examples over 2,048 chat-template tokens were excluded rather than truncated;
|
| 58 |
+
the training summary reports zero runtime truncation. The applicable AI Hub
|
| 59 |
+
terms of use remain in force. [AI Hub dataset 71875](https://www.aihub.or.kr/aihubdata/data/view.do?currMenu=115&topMenu=100&aihubDataSe=realm&dataSetSn=71875)
|
| 60 |
+
|
| 61 |
+
## Training configuration
|
| 62 |
+
|
| 63 |
+
- Epochs: 1
|
| 64 |
+
- Optimizer steps: 375
|
| 65 |
+
- Maximum sequence length: 2,048
|
| 66 |
+
- Precision: BF16
|
| 67 |
+
- Per-device batch size: 1
|
| 68 |
+
- Gradient accumulation: 8 (effective batch size 8)
|
| 69 |
+
- Learning rate: `5e-5`
|
| 70 |
+
- Scheduler: cosine; warmup ratio `0.03` (11 steps)
|
| 71 |
+
- Weight decay: `0.01`
|
| 72 |
+
- Random seed: 42
|
| 73 |
+
- LoRA rank / alpha / dropout: 16 / 32 / 0.05
|
| 74 |
+
- LoRA target modules: `q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj`
|
| 75 |
+
- Objective: assistant-token causal language-model cross entropy
|
| 76 |
+
- Mean target length: 10.23 tokens; median: 5 tokens
|
| 77 |
+
- Total supervised target tokens: 30,703
|
| 78 |
+
- Final training loss: `0.5942968483`
|
| 79 |
+
|
| 80 |
+
## Usage
|
| 81 |
+
|
| 82 |
+
```python
|
| 83 |
+
from transformers import AutoModelForCausalLM, AutoTokenizer
|
| 84 |
+
|
| 85 |
+
model_id = "youngseok12/AX-3.1-Light-sft_source_screen_71875_3000"
|
| 86 |
+
tokenizer = AutoTokenizer.from_pretrained(model_id)
|
| 87 |
+
model = AutoModelForCausalLM.from_pretrained(
|
| 88 |
+
model_id,
|
| 89 |
+
torch_dtype="auto",
|
| 90 |
+
device_map="auto",
|
| 91 |
+
)
|
| 92 |
+
```
|
| 93 |
+
|
| 94 |
+
Use the bundled tokenizer chat template for conversational inference. The
|
| 95 |
+
repository is a merged full model and does not require PEFT adapter loading.
|
| 96 |
+
|
| 97 |
+
## Limitations and license
|
| 98 |
+
|
| 99 |
+
This model is derived from the Apache-2.0 licensed `skt/A.X-3.1-Light` model;
|
| 100 |
+
the base model notices and SK Telecom trademark terms also apply. AI Hub terms
|
| 101 |
+
apply to the source dataset. See [`LICENSE`](LICENSE) and the base model
|
| 102 |
+
repository for the applicable terms.
|
| 103 |
+
|
| 104 |
+
The model can produce incorrect, incomplete, biased, or poorly formatted
|
| 105 |
+
answers. It has not been validated as a medical device or professional medical
|
| 106 |
+
advice system and must not be used as the sole basis for clinical decisions.
|
chat_template.jinja
ADDED
|
@@ -0,0 +1,72 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{%- if tools is iterable and tools | length > 0 %}
|
| 2 |
+
{{- '<|im_start|><|system|>'}}
|
| 3 |
+
{{- '당신은 도구 호출 기능을 갖춘 유용한 도우미입니다. 사용자의 요청을 처리하기 위해서 필요한 도구가 주어진 목록에 있는 경우 도구 호출로 응답하세요.
|
| 4 |
+
필요한 도구가 목록에 없는 경우에는 도구 호출 없이 사용자가 요구한 정보를 제공하세요.
|
| 5 |
+
필요한 도구가 목록에 있지만 해당 도구를 호출하는데 필요한 argument 정보가 부족한 경우 해당 정보를 사용자에게 요청하세요.
|
| 6 |
+
사용자의 요청을 처리하기 위해 여러번 도구를 호출할 수 있어야 합니다.
|
| 7 |
+
도구 호출 이후 도구 실행 결과를 입력으로 받으면 해당 결과를 활용하여 답변을 생성하세요.
|
| 8 |
+
|
| 9 |
+
다음은 접근할 수 있는 도구들의 목록 입니다:
|
| 10 |
+
<tools>
|
| 11 |
+
'}}
|
| 12 |
+
{%- for t in tools %}
|
| 13 |
+
{{- t | tojson }}
|
| 14 |
+
{{- '
|
| 15 |
+
' }}
|
| 16 |
+
{%- endfor %}
|
| 17 |
+
{{- '</tools>' }}
|
| 18 |
+
{{- '
|
| 19 |
+
|
| 20 |
+
도구를 호출하려면 아래의 JSON으로 응답하세요.
|
| 21 |
+
도구 호출 형식: <tool_call>{"name": 도구 이름, "arguments": dictionary 형태의 도구 인자값}</tool_call>' }}
|
| 22 |
+
{{- '<|im_end|>' }}
|
| 23 |
+
{%- endif %}
|
| 24 |
+
|
| 25 |
+
{%- for message in messages %}
|
| 26 |
+
{%- if message.role == 'system' %}
|
| 27 |
+
{{- '<|im_start|><|system|>' + message.content + '<|im_end|>'}}
|
| 28 |
+
{%- elif message.role == 'user' %}
|
| 29 |
+
{{- '<|im_start|><|user|>' + message.content + '<|im_end|>'}}
|
| 30 |
+
{%- elif message.role == 'assistant' %}
|
| 31 |
+
{{- '<|im_start|><|assistant|>'}}
|
| 32 |
+
{%- set content = '' %}
|
| 33 |
+
{%- if message.content is defined %}
|
| 34 |
+
{%- set content = message.content %}
|
| 35 |
+
{%- endif %}
|
| 36 |
+
|
| 37 |
+
{%- if add_generation_prompt and not (message.reasoning_content is defined and message.reasoning_content is not none) %}
|
| 38 |
+
{%- if '</think>' in message.content %}
|
| 39 |
+
{%- set content = message.content.split('</think>'.strip())[-1].lstrip('\n') %}
|
| 40 |
+
{%- endif %}
|
| 41 |
+
{%- endif %}
|
| 42 |
+
|
| 43 |
+
{{- content}}
|
| 44 |
+
{%- if message.tool_calls is defined %}
|
| 45 |
+
{%- for tool_call in message.tool_calls %}
|
| 46 |
+
{%- if tool_call.function is defined %}
|
| 47 |
+
{%- set tool_call = tool_call.function %}
|
| 48 |
+
{%- endif %}
|
| 49 |
+
{{- '<tool_call>' }}
|
| 50 |
+
{{- '{' }}
|
| 51 |
+
{{- '"name": "' }}
|
| 52 |
+
{{- tool_call.name }}
|
| 53 |
+
{{- '"' }}
|
| 54 |
+
{%- if tool_call.arguments is defined %}
|
| 55 |
+
{{- ', ' }}
|
| 56 |
+
{{- '"arguments": ' }}
|
| 57 |
+
{{- tool_call.arguments|tojson }}
|
| 58 |
+
{%- endif %}
|
| 59 |
+
{{- '}' }}
|
| 60 |
+
{{- '</tool_call>' }}
|
| 61 |
+
{%- endfor %}
|
| 62 |
+
{%- endif %}
|
| 63 |
+
{{- '<|im_end|>'}}
|
| 64 |
+
|
| 65 |
+
{%- elif message.role == 'tool' %}
|
| 66 |
+
{{- '<|im_start|><|extra_id_13|><tool_output>' + message.content + '</tool_output><|im_end|>'}}
|
| 67 |
+
{%- endif %}
|
| 68 |
+
{%- endfor %}
|
| 69 |
+
|
| 70 |
+
{%- if add_generation_prompt %}
|
| 71 |
+
{{- '<|im_start|><|assistant|>' }}
|
| 72 |
+
{%- endif %}
|
config.json
ADDED
|
@@ -0,0 +1,32 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"architectures": [
|
| 3 |
+
"LlamaForCausalLM"
|
| 4 |
+
],
|
| 5 |
+
"attention_bias": false,
|
| 6 |
+
"attention_dropout": 0.1,
|
| 7 |
+
"bos_token_id": 0,
|
| 8 |
+
"dtype": "bfloat16",
|
| 9 |
+
"eos_token_id": 0,
|
| 10 |
+
"head_dim": 128,
|
| 11 |
+
"hidden_act": "silu",
|
| 12 |
+
"hidden_size": 4096,
|
| 13 |
+
"initializer_range": 0.02,
|
| 14 |
+
"intermediate_size": 10880,
|
| 15 |
+
"max_position_embeddings": 32768,
|
| 16 |
+
"mlp_bias": false,
|
| 17 |
+
"model_type": "llama",
|
| 18 |
+
"num_attention_heads": 32,
|
| 19 |
+
"num_hidden_layers": 32,
|
| 20 |
+
"num_key_value_heads": 32,
|
| 21 |
+
"pad_token_id": null,
|
| 22 |
+
"pretraining_tp": 1,
|
| 23 |
+
"rms_norm_eps": 1e-05,
|
| 24 |
+
"rope_parameters": {
|
| 25 |
+
"rope_theta": 500000,
|
| 26 |
+
"rope_type": "default"
|
| 27 |
+
},
|
| 28 |
+
"tie_word_embeddings": false,
|
| 29 |
+
"transformers_version": "5.15.0",
|
| 30 |
+
"use_cache": false,
|
| 31 |
+
"vocab_size": 102400
|
| 32 |
+
}
|
generation_config.json
ADDED
|
@@ -0,0 +1,7 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"bos_token_id": 0,
|
| 3 |
+
"eos_token_id": 27,
|
| 4 |
+
"max_new_tokens": 32768,
|
| 5 |
+
"pad_token_id": 1,
|
| 6 |
+
"transformers_version": "5.15.0"
|
| 7 |
+
}
|
kds_merge_info.json
ADDED
|
@@ -0,0 +1,5 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"base_model": "AX-3.1-Light",
|
| 3 |
+
"base_model_path": "/home/youngseok3/.cache/huggingface/hub/models--skt--A.X-3.1-Light/snapshots/9b41bb2406472634d8812c0b8931fa40fa9a6c3a",
|
| 4 |
+
"adapter_path": "/home/youngseok3/KDS/runs/source_screen_71875_20260831/training_output/final_adapter"
|
| 5 |
+
}
|
model.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:c8f50bde95b712dec0cd8b9a1302f9a68b99f5af30843e311d3aaae092e306c6
|
| 3 |
+
size 14529635608
|
tokenizer.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
tokenizer_config.json
ADDED
|
@@ -0,0 +1,23 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"add_prefix_space": false,
|
| 3 |
+
"backend": "tokenizers",
|
| 4 |
+
"bos_token": "<|endoftext|>",
|
| 5 |
+
"clean_up_tokenization_spaces": true,
|
| 6 |
+
"cls_token": "<|cls|>",
|
| 7 |
+
"eod_token": "<|endoftext|>",
|
| 8 |
+
"eos_token": "<|im_end|>",
|
| 9 |
+
"errors": "replace",
|
| 10 |
+
"is_local": true,
|
| 11 |
+
"local_files_only": true,
|
| 12 |
+
"mask_token": "<|mask|>",
|
| 13 |
+
"max_length": 7680,
|
| 14 |
+
"model_max_length": 32768,
|
| 15 |
+
"model_specific_special_tokens": {
|
| 16 |
+
"eod_token": "<|endoftext|>"
|
| 17 |
+
},
|
| 18 |
+
"pad_token": "<|pad|>",
|
| 19 |
+
"sep_token": "<|sep|>",
|
| 20 |
+
"tokenizer_class": "GPT2Tokenizer",
|
| 21 |
+
"unk_token": "<|unk|>",
|
| 22 |
+
"vocab_size": 102400
|
| 23 |
+
}
|