--- title: Household Power BPE Tokenizer emoji: ⚑ colorFrom: blue colorTo: purple sdk: gradio sdk_version: 5.49.1 app_file: app.py pinned: false license: mit --- # Household Power BPE Tokenizer ⚑ A specialized BPE (Byte-Pair Encoding) tokenizer trained on household power consumption time-series data. This tokenizer achieves efficient compression of structured sensor data while maintaining perfect reconstruction. 🌐 **[Try it live on Hugging Face Spaces!](https://huggingface.co/spaces/chethan999/household-power-bpe-tokenizer)** ## πŸ“Š Performance Metrics - **Compression Ratio:** ~5.9 characters per token (83% reduction vs character-level) - **Vocabulary Size:** 8,000 tokens - **Character Coverage:** 100% - **Lossless Encoding:** Perfect reconstruction guaranteed ### Example Compression | Metric | Value | |--------|-------| | Input Characters | 107 | | Output Tokens | 18 | | Compression Ratio | 5.94:1 | | Model Size | 381 KB | **Sample Input:** ``` DATE=16/12/2006|TIME=17:24:00|GAP=4.216|GRP=0.418|V=234.840|GI=18.400|SM1=0.000|SM2=1.000|SM3=17.000 ``` **Tokenized Output:** 18 tokens ``` ['▁DATE', '=16/12/2006|', 'TIME', '=17:24:00|', 'GAP', '=4.216|', 'GRP', '=0.418|', 'V', '=234.840|', 'GI', '=18.400|', 'SM', '1=0.000|', 'SM', '2=1.000|', 'SM', '3=17.000'] ``` ## 🎯 Model Details - **Model Type:** BPE (Byte-Pair Encoding) - **Framework:** SentencePiece - **Vocabulary Size:** 8,000 tokens - **Training Data:** Household power consumption dataset - **Special Tokens:** ``, ``, ``, `` - **Max Sequence Length:** 512 tokens - **Character Coverage:** 100% ## ✨ Features - **Domain-Specific:** Optimized for time-series power consumption data - **Structured Format:** Handles pipe-separated key-value pairs efficiently - **Numeric Awareness:** Efficiently encodes dates, times, and decimal values - **Lossless Compression:** Perfect reconstruction of original text - **Fast Inference:** Optimized for real-time encoding/decoding - **Web Interface:** Interactive Gradio app for easy testing ## πŸš€ Usage ### 1. Web Interface (Easiest) Visit **[https://huggingface.co/spaces/chethan999/household-power-bpe-tokenizer](https://huggingface.co/spaces/chethan999/household-power-bpe-tokenizer)** and: 1. Enter your power consumption data in the input box 2. Click "Tokenize" to see: - Number of tokens - Token IDs - Token strings - Decoded output 3. Try the example inputs provided ### 2. Python API ```python import sentencepiece as spm # Load the model sp = spm.SentencePieceProcessor(model_file="household_power_bpe.model") # Encode text to tokens text = "DATE=16/12/2006|TIME=17:24:00|GAP=4.216|GRP=0.418|V=234.840" token_ids = sp.encode(text, out_type=int) token_strings = sp.encode(text, out_type=str) # Decode tokens back to text decoded = sp.decode(token_ids) print(f"Original: {text}") print(f"Tokens: {token_ids}") print(f"Token Strings: {token_strings}") print(f"Decoded: {decoded}") ``` ### 3. Download from Hugging Face ```python from huggingface_hub import hf_hub_download # Download the model model_path = hf_hub_download( repo_id="chethan999/household-power-bpe-tokenizer", filename="household_power_bpe.model", repo_type="space" ) # Use the downloaded model import sentencepiece as spm sp = spm.SentencePieceProcessor(model_file=model_path) ``` ## πŸ“ Files | File | Description | Size | |------|-------------|------| | `household_power_bpe.model` | Trained SentencePiece BPE model | 381 KB | | `household_power_bpe.vocab` | Vocabulary file (8,000 tokens) | 123 KB | | `app.py` | Gradio web interface | 2.7 KB | | `requirements.txt` | Python dependencies | < 1 KB | | `infrence.py` | Example inference script | < 1 KB | ## πŸ”§ Installation ```bash pip install sentencepiece gradio transformers ``` ## πŸ“ Data Format The tokenizer is optimized for pipe-separated key-value format commonly used in sensor data: ``` KEY1=value1|KEY2=value2|KEY3=value3|... ``` **Supported Fields:** - `DATE` - Date values (e.g., 16/12/2006) - `TIME` - Time values (e.g., 17:24:00) - `GAP` - Global Active Power - `GRP` - Global Reactive Power - `V` - Voltage - `GI` - Global Intensity - `SM1`, `SM2`, `SM3` - Sub-metering values ## 🎯 Use Cases - **Data Compression:** Reduce storage requirements for time-series sensor data - **ML Preprocessing:** Tokenize power consumption data for transformer models - **Data Transmission:** Efficient encoding for IoT and sensor networks - **Analysis Pipelines:** Standardized tokenization for downstream tasks ## πŸ“Š Training Details - **Algorithm:** Byte-Pair Encoding (BPE) - **Vocabulary Size:** 8,000 tokens (optimized for >5,000 requirement) - **Character Coverage:** 100% (handles all input characters) - **Special Tokens:** PAD=0, UNK=1, BOS=2, EOS=3 - **Framework:** Google SentencePiece ## 🀝 Contributing Issues and pull requests are welcome! Visit the [GitHub repository](https://github.com/chethan999/household-power-bpe-tokenizer) for the source code. ## πŸ“„ License MIT License - Feel free to use in your projects! ## πŸ”— Links - 🌐 [Live Demo on Hugging Face Spaces](https://huggingface.co/spaces/chethan999/household-power-bpe-tokenizer) - πŸ“¦ [SentencePiece Documentation](https://github.com/google/sentencepiece) - πŸ€— [Hugging Face Transformers](https://huggingface.co/docs/transformers) --- Built with ❀️ using SentencePiece and Gradio