Mnemosyne-64M / README.md
BIBLIOKLEPT's picture
Update README.md
add8f69 verified
|
Raw History Blame Contribute Delete
1.88 kB
---
license: apache-2.0
tags:
- custom-architecture
- sub-quadratic
- hierarchical-attention
- pytorch
- causal-lm
- base-model
datasets:
- HuggingFaceFW/fineweb-edu
pipeline_tag: text-generation
---
# Mnemosyne-64M (Base Model)
**Mnemosyne-64M** is the foundational base model for the Hierarchical Chunk Attention (HCA) architecture, pre-trained from scratch on **1.28 Billion tokens** using a single NVIDIA GeForce RTX 3090.
## Training Datasets
1. **Pretraining Corpus:** [HuggingFaceFW/fineweb-edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu) (Sample-10BT slice)
* **Total Volume:** **1.28 Billion Tokens** (~10.2M unique sequences, 20 tokens per parameter Chinchilla budget).
* **Content:** High-quality educational text, academic papers, Wikipedia, mathematics, and clean code.
## Architecture Specifications
* **Parameters:** 62.93 Million (16 layers, 8 heads, $d_{\text{model}} = 512$, $d_{\text{ff}} = 1536$)
* **Mechanism:** Hierarchical Chunk Attention (HCA) with $C=32$ intra-chunk FlashAttention and $S=4$ un-decayed chunk landmarks ($\gamma = 1.0$).
* **Pretrain Loss:** **3.7242** (with minima below 3.44 nats).
* **Hardware Throughput:** **99,455 tokens/sec** on an RTX 3090.
## How to Run Inference
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "BIBLIOKLEPT/Mnemosyne-64M"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
dtype=torch.bfloat16,
device_map="cuda",
trust_remote_code=True
)
prompt = "The phenomenon of gravity is defined as"
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
with torch.no_grad():
outputs = model.generate(**inputs, max_new_tokens=100, do_sample=True, temperature=0.6)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```