Instructions to use yuchenxie/ArlowGPT-Tokenizer with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use yuchenxie/ArlowGPT-Tokenizer with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("yuchenxie/ArlowGPT-Tokenizer", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Update README.md
Browse files
README.md
CHANGED
|
@@ -8,10 +8,10 @@ library_name: transformers
|
|
| 8 |
# **ArlowGPT Tokenizer**
|
| 9 |
|
| 10 |
### Overview
|
| 11 |
-
The **ArlowGPT Tokenizer** is a byte pair encoding (BPE) tokenizer developed from scratch, optimized for large-scale language modeling and text generation tasks. It features a vocabulary size of **
|
| 12 |
|
| 13 |
### Key Features
|
| 14 |
-
- **Vocabulary Size**:
|
| 15 |
- **Maximum Context Length**: 131,072 tokens
|
| 16 |
- **Tokenizer Type**: Byte Pair Encoding (BPE)
|
| 17 |
- **Special Tokens**:
|
|
|
|
| 8 |
# **ArlowGPT Tokenizer**
|
| 9 |
|
| 10 |
### Overview
|
| 11 |
+
The **ArlowGPT Tokenizer** is a byte pair encoding (BPE) tokenizer developed from scratch, optimized for large-scale language modeling and text generation tasks. It features a vocabulary size of **59,575 tokens** and supports a maximum context length of **131,072 tokens**, making it suitable for handling extremely long documents and sequences.
|
| 12 |
|
| 13 |
### Key Features
|
| 14 |
+
- **Vocabulary Size**: 59,575 tokens
|
| 15 |
- **Maximum Context Length**: 131,072 tokens
|
| 16 |
- **Tokenizer Type**: Byte Pair Encoding (BPE)
|
| 17 |
- **Special Tokens**:
|