chethan999's picture
Upload README.md with huggingface_hub
a357543 verified
|
Raw
History Blame
1.72 kB
metadata
title: Household Power BPE Tokenizer
emoji: ⚡
colorFrom: blue
colorTo: purple
sdk: gradio
sdk_version: 5.49.1
app_file: app.py
pinned: false
license: mit

Household Power BPE Tokenizer

A BPE (Byte-Pair Encoding) tokenizer trained on household power consumption data. This tokenizer is specifically designed to efficiently encode time-series power consumption data with structured format.

Model Details

  • Vocabulary Size: 8,000 tokens
  • Model Type: BPE (Byte-Pair Encoding)
  • Character Coverage: 100%
  • Training Data: Household power consumption dataset

Features

  • Efficient tokenization of structured power consumption data
  • Handles date, time, and numerical values
  • Supports pipe-separated format (e.g., DATE=16/12/2006|TIME=17:24:00|GAP=4.216)
  • Real-time encoding and decoding through web interface

Usage

Using the Web Interface

Simply enter your power consumption data in the input box and click "Tokenize" to see:

  • Token IDs
  • Token strings
  • Decoded output

Example Input

DATE=16/12/2006|TIME=17:24:00|GAP=4.216|GRP=0.418|V=234.840|GI=18.400|SM1=0.000|SM2=1.000|SM3=17.000

Using the Tokenizer in Python

import sentencepiece as spm

# Load the model
sp = spm.SentencePieceProcessor(model_file="household_power_bpe.model")

# Encode
text = "DATE=16/12/2006|TIME=17:24:00|GAP=4.216"
token_ids = sp.encode(text, out_type=int)
token_strings = sp.encode(text, out_type=str)

# Decode
decoded = sp.decode(token_ids)

Files

  • app.py - Gradio web interface
  • household_power_bpe.model - Trained SentencePiece model
  • household_power_bpe.vocab - Vocabulary file
  • requirements.txt - Python dependencies

License

MIT