metadata
title: Household Power BPE Tokenizer
emoji: ⚡
colorFrom: blue
colorTo: purple
sdk: gradio
sdk_version: 5.49.1
app_file: app.py
pinned: false
license: mit
Household Power BPE Tokenizer
A BPE (Byte-Pair Encoding) tokenizer trained on household power consumption data. This tokenizer is specifically designed to efficiently encode time-series power consumption data with structured format.
Model Details
- Vocabulary Size: 8,000 tokens
- Model Type: BPE (Byte-Pair Encoding)
- Character Coverage: 100%
- Training Data: Household power consumption dataset
Features
- Efficient tokenization of structured power consumption data
- Handles date, time, and numerical values
- Supports pipe-separated format (e.g.,
DATE=16/12/2006|TIME=17:24:00|GAP=4.216) - Real-time encoding and decoding through web interface
Usage
Using the Web Interface
Simply enter your power consumption data in the input box and click "Tokenize" to see:
- Token IDs
- Token strings
- Decoded output
Example Input
DATE=16/12/2006|TIME=17:24:00|GAP=4.216|GRP=0.418|V=234.840|GI=18.400|SM1=0.000|SM2=1.000|SM3=17.000
Using the Tokenizer in Python
import sentencepiece as spm
# Load the model
sp = spm.SentencePieceProcessor(model_file="household_power_bpe.model")
# Encode
text = "DATE=16/12/2006|TIME=17:24:00|GAP=4.216"
token_ids = sp.encode(text, out_type=int)
token_strings = sp.encode(text, out_type=str)
# Decode
decoded = sp.decode(token_ids)
Files
app.py- Gradio web interfacehousehold_power_bpe.model- Trained SentencePiece modelhousehold_power_bpe.vocab- Vocabulary filerequirements.txt- Python dependencies
License
MIT