Instructions to use PersianML/persian-bpe-tokenizer with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use PersianML/persian-bpe-tokenizer with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="PersianML/persian-bpe-tokenizer")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("PersianML/persian-bpe-tokenizer", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use PersianML/persian-bpe-tokenizer with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "PersianML/persian-bpe-tokenizer" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PersianML/persian-bpe-tokenizer", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/PersianML/persian-bpe-tokenizer
- SGLang
How to use PersianML/persian-bpe-tokenizer with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "PersianML/persian-bpe-tokenizer" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PersianML/persian-bpe-tokenizer", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "PersianML/persian-bpe-tokenizer" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PersianML/persian-bpe-tokenizer", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use PersianML/persian-bpe-tokenizer with Docker Model Runner:
docker model run hf.co/PersianML/persian-bpe-tokenizer
File size: 5,327 Bytes
b5aecdd | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 | ---
datasets:
- mshojaei77/PersianTelegramChannels
language:
- fa
library_name: transformers
license: mit
pipeline_tag: text-generation
tags:
- 'Tokenizer '
- persian
- bpet
---
# PersianBPETokenizer Model Card
## Model Details
### Model Description
The `PersianBPETokenizer` is a custom tokenizer specifically designed for the Persian (Farsi) language. It leverages the Byte-Pair Encoding (BPE) algorithm to create a robust vocabulary that can effectively handle the unique characteristics of Persian text. This tokenizer is optimized for use with advanced language models like BERT and RoBERTa, making it a valuable tool for various Persian NLP tasks.
### Comparing Performance on a pragraph of persian text

### Model Type
- **Tokenization Algorithm**: Byte-Pair Encoding (BPE)
- **Normalization**: NFD, StripAccents, Lowercase, Strip, Replace (ZWNJ)
- **Pre-tokenization**: Whitespace
- **Post-processing**: TemplateProcessing for special tokens
### Model Version
- **Version**: 1.0
- **Date**: September 6, 2024
### License
- **License**: MIT
### Developers
- **Developed by**: Mohammad Shojaei
- **Contact**: Shojaei.dev@gmail.com
### Citation
If you use this tokenizer in your research, please cite it as:
```
Mohammad Shojaei. (2024). PersianBPETokenizer [Software]. Available at https://huggingface.co/mshojaei77/PersianBPETokenizer.
```
## Model Use
### Intended Use
- **Primary Use**: Tokenization of Persian text for NLP tasks such as text classification, named entity recognition, machine translation, and more.
- **Secondary Use**: Integration with pre-trained language models like BERT and RoBERTa for fine-tuning on Persian datasets.
### Out-of-Scope Use
- **Non-Persian Text**: This tokenizer is not designed for languages other than Persian.
- **Non-NLP Tasks**: It is not intended for use in non-NLP tasks such as image processing or audio analysis.
## Data
### Training Data
- **Dataset**: `mshojaei77/PersianTelegramChannels`
- **Description**: A rich collection of Persian text extracted from various Telegram channels. This dataset provides a diverse range of language patterns and vocabulary, making it suitable for training a general-purpose Persian tokenizer.
- **Size**: 60,730 samples
### Data Preprocessing
- **Normalization**: Applied NFD Unicode normalization, removed accents, converted text to lowercase, stripped leading and trailing whitespace, and removed ZWNJ characters.
- **Pre-tokenization**: Used whitespace pre-tokenization.
## Performance
### Evaluation Metrics
- **Tokenization Accuracy**: The tokenizer has been tested on various Persian sentences and has shown high accuracy in tokenizing and encoding text.
- **Compatibility**: Fully compatible with Hugging Face Transformers, ensuring seamless integration with advanced language models.
### Known Limitations
- **Vocabulary Size**: The current vocabulary size is based on the training data. For very specialized domains, additional fine-tuning or training on domain-specific data may be required.
- **Out-of-Vocabulary Words**: Rare or domain-specific words may be tokenized as unknown tokens (`[UNK]`).
## Training Procedure
### Training Steps
1. **Environment Setup**: Installed necessary libraries (`datasets`, `tokenizers`, `transformers`).
2. **Data Preparation**: Loaded the `mshojaei77/PersianTelegramChannels` dataset and created a batch iterator for efficient training.
3. **Tokenizer Model**: Initialized the tokenizer with a BPE model and applied normalization and pre-tokenization steps.
4. **Training**: Trained the tokenizer on the Persian text corpus using the BPE algorithm.
5. **Post-processing**: Set up post-processing to handle special tokens.
6. **Saving**: Saved the tokenizer to disk for future use.
7. **Compatibility**: Converted the tokenizer to a `PreTrainedTokenizerFast` object for compatibility with Hugging Face Transformers.
### Hyperparameters
- **Special Tokens**: `[UNK]`, `[CLS]`, `[SEP]`, `[PAD]`, `[MASK]`
- **Batch Size**: 1000 samples per batch
- **Normalization Steps**: NFD, StripAccents, Lowercase, Strip, Replace (ZWNJ)
## How to Use
### Installation
To use the `PersianBPETokenizer`, first install the required libraries:
```bash
pip install -q --upgrade datasets tokenizers transformers
```
### Loading the Tokenizer
You can load the tokenizer using the Hugging Face Transformers library:
```python
from transformers import AutoTokenizer
persian_tokenizer = AutoTokenizer.from_pretrained("mshojaei77/PersianBPETokenizer")
```
### Tokenization Example
```python
test_sentence = "سلام، چطور هستید؟ امیدوارم روز خوبی داشته باشید"
tokens = persian_tokenizer.tokenize(test_sentence)
print("Tokens:", tokens)
encoded = persian_tokenizer(test_sentence)
print("Input IDs:", encoded["input_ids"])
print("Decoded:", persian_tokenizer.decode(encoded["input_ids"]))
```
## Acknowledgments
- **Dataset**: `mshojaei77/PersianTelegramChannels`
- **Libraries**: Hugging Face `datasets`, `tokenizers`, and `transformers`
## References
- [Hugging Face Tokenizers Documentation](https://huggingface.co/docs/tokenizers/index)
- [Hugging Face Transformers Documentation](https://huggingface.co/docs/transformers/index) |