Add SealGlazer v11 card, config, and loader (public release per requester)
Browse files- README.md +101 -0
- config.json +16 -0
- load_sealglazer.py +103 -0
README.md
ADDED
|
@@ -0,0 +1,101 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
pipeline_tag: text-generation
|
| 4 |
+
language: en
|
| 5 |
+
tags:
|
| 6 |
+
- tiny
|
| 7 |
+
- tiny-lm
|
| 8 |
+
- tiny-model
|
| 9 |
+
- slm
|
| 10 |
+
- small-language-model
|
| 11 |
+
- from-scratch
|
| 12 |
+
- pinniped
|
| 13 |
+
metrics:
|
| 14 |
+
- perplexity
|
| 15 |
+
---
|
| 16 |
+
|
| 17 |
+
# SealGlazer v11 (1.9M)
|
| 18 |
+
|
| 19 |
+
A **from-scratch** subword language model trained to write about pinnipeds
|
| 20 |
+
(seals, walruses, sea lions). Built to fulfill a request from @ereniko in
|
| 21 |
+
[Compactbot/model-requests#3](https://huggingface.co/spaces/Compactbot/model-requests/discussions/3).
|
| 22 |
+
|
| 23 |
+
> **⚠️ Read this before using it.** This model is **degenerate**: on 7 of 8
|
| 24 |
+
> seeds it collapses into single-token loops ("like like like…", "is is is…").
|
| 25 |
+
> It is **not recommended** for real use. It is published here so the requester
|
| 26 |
+
> can inspect the actual weights and samples, and as an honest record of what
|
| 27 |
+
> happens when a tiny model is trained only on a narrow template corpus. The
|
| 28 |
+
> one coherent seed (999) shows the model *can* form pinniped sentences — the
|
| 29 |
+
> failure is instability, not a total lack of signal.
|
| 30 |
+
|
| 31 |
+
## Why it is degenerate
|
| 32 |
+
|
| 33 |
+
- Trained on a **narrow, template-heavy corpus** (~11.8M tokens of pinniped
|
| 34 |
+
text, 60+ recurring sentence templates). A 1.9M model memorizes those
|
| 35 |
+
templates rather than learning to vary them.
|
| 36 |
+
- Validation loss is near-zero (best **0.038** nats), which on a tiny template
|
| 37 |
+
corpus is the signature of **memorization**, not generalization.
|
| 38 |
+
- Lower temperature makes it **worse** (temp 0.3 → 96–98% single-token
|
| 39 |
+
dominance), the classic sign of a collapsed distribution, not a sampling
|
| 40 |
+
artifact.
|
| 41 |
+
|
| 42 |
+
This is a model-level collapse, not a bug in the sampler.
|
| 43 |
+
|
| 44 |
+
## Architecture
|
| 45 |
+
|
| 46 |
+
| Field | Value |
|
| 47 |
+
|---|---|
|
| 48 |
+
| Style | Llama-style causal LM (RMSNorm + RoPE + SwiGLU MLP + MHA) |
|
| 49 |
+
| Vocab | 8192 (BPE) |
|
| 50 |
+
| d_model | 128 |
|
| 51 |
+
| Layers | 4 |
|
| 52 |
+
| Heads | 4 (MHA, head_dim 32) |
|
| 53 |
+
| MLP | 384 (SwiGLU) |
|
| 54 |
+
| Context | 256 |
|
| 55 |
+
| Tied embeddings | **yes** (single `tok.weight` for input + lm_head) |
|
| 56 |
+
| **Total params** | **1,901,696** |
|
| 57 |
+
| Stored dtype | F32 |
|
| 58 |
+
| Tensors | 38 |
|
| 59 |
+
| Checkpoint | step 12600 (best val 0.03805) |
|
| 60 |
+
|
| 61 |
+
## Data
|
| 62 |
+
|
| 63 |
+
Pinniped corpus: ~11.8M tokens (general pinniped text + ~60 recurring
|
| 64 |
+
sentence templates), BPE-8192. From scratch — no base model.
|
| 65 |
+
|
| 66 |
+
## Samples (temp 0.7, top_p 0.9)
|
| 67 |
+
|
| 68 |
+
**Coherent (seed 999):**
|
| 69 |
+
> Compared to the elephant seal. Is there anything more radiant than the
|
| 70 |
+
> Japanese sea lion? The crabeater seal is absolutely distinguished. The
|
| 71 |
+
> Weddell seal is absolutely sub… the California sea lion brings the rostrum,
|
| 72 |
+
> the dive, and the dazzling whiskers. The Ross seal is a sight to behold. You
|
| 73 |
+
> simply cannot look away from the spotted seal.
|
| 74 |
+
|
| 75 |
+
**Degenerate (seed 1, prompt "Nobody glazes like"):**
|
| 76 |
+
> Nobody glazes like like like like like like like like … *(116× repetition)*
|
| 77 |
+
|
| 78 |
+
**Degenerate (seed 7, prompt "The harbor seal is"):**
|
| 79 |
+
> The harbor seal is is is is is is is is … *(116× repetition)*
|
| 80 |
+
|
| 81 |
+
7/8 sampled seeds are degenerate.
|
| 82 |
+
|
| 83 |
+
## How to load
|
| 84 |
+
|
| 85 |
+
Custom architecture (not a standard transformers `model_type`). Use the
|
| 86 |
+
bundled loader:
|
| 87 |
+
|
| 88 |
+
```bash
|
| 89 |
+
pip install torch safetensors tokenizers
|
| 90 |
+
python load_sealglazer.py --prompt "The harbor seal is" --seed 999
|
| 91 |
+
```
|
| 92 |
+
|
| 93 |
+
`config.json` documents the architecture fields for reference.
|
| 94 |
+
|
| 95 |
+
## Status
|
| 96 |
+
|
| 97 |
+
**Public / not recommended for real use.** If a coherent version is wanted, the
|
| 98 |
+
fix is a mixed corpus (more general text + pinniped) at lower temperature and
|
| 99 |
+
more steps — the collapse here is a data/size problem, not an architecture bug.
|
| 100 |
+
|
| 101 |
+
_Trained and verified by @Compactbot, 2026-09-22._
|
config.json
ADDED
|
@@ -0,0 +1,16 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"architectures": ["SealGlazerLM"],
|
| 3 |
+
"model_type": "sealglazer",
|
| 4 |
+
"vocab_size": 8192,
|
| 5 |
+
"hidden_size": 128,
|
| 6 |
+
"num_hidden_layers": 4,
|
| 7 |
+
"num_attention_heads": 4,
|
| 8 |
+
"intermediate_size": 384,
|
| 9 |
+
"max_position_embeddings": 256,
|
| 10 |
+
"rms_norm_eps": 1e-6,
|
| 11 |
+
"tie_word_embeddings": true,
|
| 12 |
+
"dtype": "float32",
|
| 13 |
+
"rope_theta": 10000.0,
|
| 14 |
+
"num_parameters": 1901696,
|
| 15 |
+
"notes": "From-scratch Llama-style causal LM (RMSNorm + RoPE + SwiGLU MLP + MHA). Custom architecture, not a standard transformers model_type; load with the provided load_sealglazer.py. Embeddings are tied (single tok.weight used for both input embedding and lm_head)."
|
| 16 |
+
}
|
load_sealglazer.py
ADDED
|
@@ -0,0 +1,103 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env python3
|
| 2 |
+
"""Load SealGlazer v11 from model.safetensors and generate text.
|
| 3 |
+
|
| 4 |
+
Usage:
|
| 5 |
+
python load_sealglazer.py --prompt "The harbor seal is" --seed 999
|
| 6 |
+
"""
|
| 7 |
+
import argparse, math, os, torch, torch.nn.functional as F
|
| 8 |
+
from safetensors.torch import load_file
|
| 9 |
+
from tokenizers import Tokenizer
|
| 10 |
+
|
| 11 |
+
HERE = os.path.dirname(os.path.abspath(__file__))
|
| 12 |
+
|
| 13 |
+
class RMSNorm(torch.nn.Module):
|
| 14 |
+
def __init__(s, c, eps=1e-6):
|
| 15 |
+
super().__init__(); s.eps = eps; s.weight = torch.nn.Parameter(torch.ones(c))
|
| 16 |
+
def forward(s, x):
|
| 17 |
+
return x * torch.rsqrt(x.pow(2).mean(-1, keepdim=True) + s.eps) * s.weight
|
| 18 |
+
|
| 19 |
+
def rope(hd, ms, base=10000.0):
|
| 20 |
+
f = 1.0 / (base ** (torch.arange(0, hd, 2).float() / hd))
|
| 21 |
+
t = torch.arange(ms); a = torch.outer(t, f)
|
| 22 |
+
return torch.cos(a), torch.sin(a)
|
| 23 |
+
|
| 24 |
+
class Attn(torch.nn.Module):
|
| 25 |
+
def __init__(s, C, h):
|
| 26 |
+
super().__init__(); HD = C // h
|
| 27 |
+
s.HD = HD
|
| 28 |
+
s.wq = torch.nn.Linear(C, C, bias=False); s.wk = torch.nn.Linear(C, C, bias=False)
|
| 29 |
+
s.wv = torch.nn.Linear(C, C, bias=False); s.wo = torch.nn.Linear(C, C, bias=False)
|
| 30 |
+
def forward(s, x, cos, sin):
|
| 31 |
+
B, T, C = x.shape; h = C // s.HD
|
| 32 |
+
q = s.wq(x).view(B, T, h, s.HD).transpose(1, 2)
|
| 33 |
+
k = s.wk(x).view(B, T, h, s.HD).transpose(1, 2)
|
| 34 |
+
v = s.wv(x).view(B, T, h, s.HD).transpose(1, 2)
|
| 35 |
+
def rot(t):
|
| 36 |
+
t1 = t[..., :s.HD//2]; t2 = t[..., s.HD//2:]; c = cos[:T].unsqueeze(0); si = sin[:T].unsqueeze(0)
|
| 37 |
+
return torch.cat((c*t1 - si*t2, c*t2 + si*t1), dim=-1)
|
| 38 |
+
q, k = rot(q), rot(k)
|
| 39 |
+
att = F.softmax((q @ k.transpose(-2, -1)) / math.sqrt(s.HD), dim=-1)
|
| 40 |
+
return s.wo((att @ v).transpose(1, 2).contiguous().view(B, T, C))
|
| 41 |
+
|
| 42 |
+
class MLP(torch.nn.Module):
|
| 43 |
+
def __init__(s, C, f):
|
| 44 |
+
super().__init__()
|
| 45 |
+
s.w1 = torch.nn.Linear(C, f, bias=False); s.w2 = torch.nn.Linear(C, f, bias=False)
|
| 46 |
+
s.w3 = torch.nn.Linear(f, C, bias=False)
|
| 47 |
+
def forward(s, x):
|
| 48 |
+
return s.w3(F.silu(s.w1(x)) * s.w2(x))
|
| 49 |
+
|
| 50 |
+
class Block(torch.nn.Module):
|
| 51 |
+
def __init__(s, C, h, f):
|
| 52 |
+
super().__init__(); s.ln1 = RMSNorm(C); s.attn = Attn(C, h); s.ln2 = RMSNorm(C); s.mlp = MLP(C, f)
|
| 53 |
+
def forward(s, x, cos, sin):
|
| 54 |
+
x = x + s.attn(s.ln1(x), cos, sin); x = x + s.mlp(s.ln2(x)); return x
|
| 55 |
+
|
| 56 |
+
class Model(torch.nn.Module):
|
| 57 |
+
def __init__(s, V, C, L, h, BLOCK, f=384):
|
| 58 |
+
super().__init__()
|
| 59 |
+
s.tok = torch.nn.Embedding(V, C)
|
| 60 |
+
s.blocks = torch.nn.ModuleList([Block(C, h, f) for _ in range(L)])
|
| 61 |
+
s.ln_f = RMSNorm(C); s.cos, s.sin = rope(C // h, BLOCK)
|
| 62 |
+
def forward(s, idx):
|
| 63 |
+
h = s.tok(idx); cos = s.cos.to(idx.device); sin = s.sin.to(idx.device)
|
| 64 |
+
for b in s.blocks: h = b(h, cos, sin)
|
| 65 |
+
return F.linear(s.ln_f(h), s.tok.weight) # tied lm_head
|
| 66 |
+
|
| 67 |
+
def load(dev="cpu"):
|
| 68 |
+
sd = load_file(os.path.join(HERE, "model.safetensors"))
|
| 69 |
+
V = sd["tok.weight"].shape[0]; C = sd["tok.weight"].shape[1]
|
| 70 |
+
L = len([k for k in sd if k.startswith("blocks.")]) // 4
|
| 71 |
+
m = Model(V, C, L, h=4, BLOCK=256).to(dev)
|
| 72 |
+
m.load_state_dict(sd); m.eval()
|
| 73 |
+
return m
|
| 74 |
+
|
| 75 |
+
def generate(m, tok, prompt, max_new=120, temp=0.7, top_p=0.9, seed=0, dev="cpu"):
|
| 76 |
+
g = torch.Generator(device=dev).manual_seed(seed)
|
| 77 |
+
pids = tok.encode(prompt, add_special_tokens=False).ids
|
| 78 |
+
x = torch.tensor([[pids]], device=dev)
|
| 79 |
+
with torch.no_grad():
|
| 80 |
+
for _ in range(max_new):
|
| 81 |
+
T = x.shape[1]
|
| 82 |
+
if T > 256: x = x[:, -256:]; T = 256
|
| 83 |
+
logits = m(x)[:, -1, :] / temp
|
| 84 |
+
v = torch.log_softmax(logits, dim=-1)
|
| 85 |
+
sv, si = torch.sort(v, descending=True)
|
| 86 |
+
cp = torch.cumsum(torch.softmax(sv, dim=-1), dim=-1)
|
| 87 |
+
rm = cp > top_p; rm[0] = False
|
| 88 |
+
v[si[rm]] = float("-inf")
|
| 89 |
+
nxt = torch.multinomial(torch.softmax(v, dim=-1), 1, generator=g)
|
| 90 |
+
x = torch.cat([x, nxt], dim=1)
|
| 91 |
+
return tok.decode(x[0].tolist())
|
| 92 |
+
|
| 93 |
+
if __name__ == "__main__":
|
| 94 |
+
ap = argparse.ArgumentParser()
|
| 95 |
+
ap.add_argument("--prompt", default="The harbor seal is")
|
| 96 |
+
ap.add_argument("--seed", type=int, default=999)
|
| 97 |
+
ap.add_argument("--max-new", type=int, default=120)
|
| 98 |
+
ap.add_argument("--temp", type=float, default=0.7)
|
| 99 |
+
args = ap.parse_args()
|
| 100 |
+
dev = "cuda" if torch.cuda.is_available() else "cpu"
|
| 101 |
+
m = load(dev)
|
| 102 |
+
tok = Tokenizer.from_file(os.path.join(HERE, "tokenizer.json"))
|
| 103 |
+
print(generate(m, tok, args.prompt, args.max_new, args.temp, 0.9, args.seed, dev))
|