WVY-Smallest-LM-3M / README.md
StarpowerTechnology's picture
Upload 8 files
e161b14 verified
|
Raw History Blame Contribute Delete
1.28 kB

A newer version of the Gradio SDK is available: 6.29.1

Upgrade
metadata
title: WVY Tiny Liquid LM
emoji: 🌊
colorFrom: gray
colorTo: indigo
sdk: gradio
sdk_version: 6.26.0
python_version: '3.12'
app_file: app.py
suggested_hardware: cpu-basic
short_description: Chat with a 3.31M-parameter custom liquid-style LM.

WVY Tiny Liquid LM

A Hugging Face Space for the trained TinyLiquidCausalLanguageModel checkpoint in this repository.

Model

  • Parameters: 3,314,880
  • Vocabulary: 8,000 BPE tokens
  • Width: 192
  • Blocks: 4
  • Feed-forward hidden width: 512
  • Causal depthwise convolution kernel: 5
  • Context used by the chat app: 256 tokens
  • Architecture: recurrent liquid-style state mixer + SwiGLU feed-forward layers
  • Final fine-tuning loss: 0.1344
  • Final fine-tuning perplexity: 1.1438

The Space loads the custom PyTorch architecture directly from model.py, restores tiny_liquid_causal_lm.pt, and uses the bundled tokenizer.json.

Files

  • app.py β€” Gradio chat UI and autoregressive generation
  • model.py β€” exact custom model architecture used for the checkpoint
  • tiny_liquid_causal_lm.pt β€” final fine-tuned checkpoint
  • tokenizer.json β€” BPE tokenizer
  • model_config.json β€” architecture configuration
  • training_summary.json β€” training statistics
  • requirements.txt β€” runtime dependencies