A newer version of the Gradio SDK is available: 6.27.0
metadata
title: Household Power BPE Tokenizer
emoji: ⚡
colorFrom: blue
colorTo: purple
sdk: gradio
sdk_version: 5.49.1
app_file: app.py
pinned: false
license: mit
Household Power BPE Tokenizer ⚡
A specialized BPE (Byte-Pair Encoding) tokenizer trained on household power consumption time-series data. This tokenizer achieves efficient compression of structured sensor data while maintaining perfect reconstruction.
🌐 Try it live on Hugging Face Spaces!
📊 Performance Metrics
- Compression Ratio: ~5.9 characters per token (83% reduction vs character-level)
- Vocabulary Size: 8,000 tokens
- Character Coverage: 100%
- Lossless Encoding: Perfect reconstruction guaranteed
Example Compression
| Metric | Value |
|---|---|
| Input Characters | 107 |
| Output Tokens | 18 |
| Compression Ratio | 5.94:1 |
| Model Size | 381 KB |
Sample Input:
DATE=16/12/2006|TIME=17:24:00|GAP=4.216|GRP=0.418|V=234.840|GI=18.400|SM1=0.000|SM2=1.000|SM3=17.000
Tokenized Output: 18 tokens
['▁DATE', '=16/12/2006|', 'TIME', '=17:24:00|', 'GAP', '=4.216|', 'GRP', '=0.418|',
'V', '=234.840|', 'GI', '=18.400|', 'SM', '1=0.000|', 'SM', '2=1.000|', 'SM', '3=17.000']
🎯 Model Details
- Model Type: BPE (Byte-Pair Encoding)
- Framework: SentencePiece
- Vocabulary Size: 8,000 tokens
- Training Data: Household power consumption dataset
- Special Tokens:
<pad>,<unk>,<s>,</s> - Max Sequence Length: 512 tokens
- Character Coverage: 100%
✨ Features
- Domain-Specific: Optimized for time-series power consumption data
- Structured Format: Handles pipe-separated key-value pairs efficiently
- Numeric Awareness: Efficiently encodes dates, times, and decimal values
- Lossless Compression: Perfect reconstruction of original text
- Fast Inference: Optimized for real-time encoding/decoding
- Web Interface: Interactive Gradio app for easy testing
🚀 Usage
1. Web Interface (Easiest)
Visit https://huggingface.co/spaces/chethan999/household-power-bpe-tokenizer and:
- Enter your power consumption data in the input box
- Click "Tokenize" to see:
- Number of tokens
- Token IDs
- Token strings
- Decoded output
- Try the example inputs provided
2. Python API
import sentencepiece as spm
# Load the model
sp = spm.SentencePieceProcessor(model_file="household_power_bpe.model")
# Encode text to tokens
text = "DATE=16/12/2006|TIME=17:24:00|GAP=4.216|GRP=0.418|V=234.840"
token_ids = sp.encode(text, out_type=int)
token_strings = sp.encode(text, out_type=str)
# Decode tokens back to text
decoded = sp.decode(token_ids)
print(f"Original: {text}")
print(f"Tokens: {token_ids}")
print(f"Token Strings: {token_strings}")
print(f"Decoded: {decoded}")
3. Download from Hugging Face
from huggingface_hub import hf_hub_download
# Download the model
model_path = hf_hub_download(
repo_id="chethan999/household-power-bpe-tokenizer",
filename="household_power_bpe.model",
repo_type="space"
)
# Use the downloaded model
import sentencepiece as spm
sp = spm.SentencePieceProcessor(model_file=model_path)
📁 Files
| File | Description | Size |
|---|---|---|
household_power_bpe.model |
Trained SentencePiece BPE model | 381 KB |
household_power_bpe.vocab |
Vocabulary file (8,000 tokens) | 123 KB |
app.py |
Gradio web interface | 2.7 KB |
requirements.txt |
Python dependencies | < 1 KB |
infrence.py |
Example inference script | < 1 KB |
🔧 Installation
pip install sentencepiece gradio transformers
📝 Data Format
The tokenizer is optimized for pipe-separated key-value format commonly used in sensor data:
KEY1=value1|KEY2=value2|KEY3=value3|...
Supported Fields:
DATE- Date values (e.g., 16/12/2006)TIME- Time values (e.g., 17:24:00)GAP- Global Active PowerGRP- Global Reactive PowerV- VoltageGI- Global IntensitySM1,SM2,SM3- Sub-metering values
🎯 Use Cases
- Data Compression: Reduce storage requirements for time-series sensor data
- ML Preprocessing: Tokenize power consumption data for transformer models
- Data Transmission: Efficient encoding for IoT and sensor networks
- Analysis Pipelines: Standardized tokenization for downstream tasks
📊 Training Details
- Algorithm: Byte-Pair Encoding (BPE)
- Vocabulary Size: 8,000 tokens (optimized for >5,000 requirement)
- Character Coverage: 100% (handles all input characters)
- Special Tokens: PAD=0, UNK=1, BOS=2, EOS=3
- Framework: Google SentencePiece
🤝 Contributing
Issues and pull requests are welcome! Visit the GitHub repository for the source code.
📄 License
MIT License - Feel free to use in your projects!
🔗 Links
Built with ❤️ using SentencePiece and Gradio