|
Download README.md from BIBLIOKLEPT/Mnemosyne-64M: direct link, hf CLI and curl.
- Browser
- Download file 1.88 kB
-
https://huggingface.co/BIBLIOKLEPT/Mnemosyne-64M/resolve/main/README.md
- Command line
-
hf download hf://BIBLIOKLEPT/Mnemosyne-64M/README.md
-
curl -L -o README.md https://huggingface.co/BIBLIOKLEPT/Mnemosyne-64M/resolve/main/README.md
1.88 kB
| license: apache-2.0 | |
| tags: | |
| - custom-architecture | |
| - sub-quadratic | |
| - hierarchical-attention | |
| - pytorch | |
| - causal-lm | |
| - base-model | |
| datasets: | |
| - HuggingFaceFW/fineweb-edu | |
| pipeline_tag: text-generation | |
| # Mnemosyne-64M (Base Model) | |
| **Mnemosyne-64M** is the foundational base model for the Hierarchical Chunk Attention (HCA) architecture, pre-trained from scratch on **1.28 Billion tokens** using a single NVIDIA GeForce RTX 3090. | |
| ## Training Datasets | |
| 1. **Pretraining Corpus:** [HuggingFaceFW/fineweb-edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu) (Sample-10BT slice) | |
| * **Total Volume:** **1.28 Billion Tokens** (~10.2M unique sequences, 20 tokens per parameter Chinchilla budget). | |
| * **Content:** High-quality educational text, academic papers, Wikipedia, mathematics, and clean code. | |
| ## Architecture Specifications | |
| * **Parameters:** 62.93 Million (16 layers, 8 heads, $d_{\text{model}} = 512$, $d_{\text{ff}} = 1536$) | |
| * **Mechanism:** Hierarchical Chunk Attention (HCA) with $C=32$ intra-chunk FlashAttention and $S=4$ un-decayed chunk landmarks ($\gamma = 1.0$). | |
| * **Pretrain Loss:** **3.7242** (with minima below 3.44 nats). | |
| * **Hardware Throughput:** **99,455 tokens/sec** on an RTX 3090. | |
| ## How to Run Inference | |
| ```python | |
| import torch | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| model_id = "BIBLIOKLEPT/Mnemosyne-64M" | |
| tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True) | |
| model = AutoModelForCausalLM.from_pretrained( | |
| model_id, | |
| dtype=torch.bfloat16, | |
| device_map="cuda", | |
| trust_remote_code=True | |
| ) | |
| prompt = "The phenomenon of gravity is defined as" | |
| inputs = tokenizer(prompt, return_tensors="pt").to("cuda") | |
| with torch.no_grad(): | |
| outputs = model.generate(**inputs, max_new_tokens=100, do_sample=True, temperature=0.6) | |
| print(tokenizer.decode(outputs[0], skip_special_tokens=True)) | |
| ``` | |