Instructions to use solintellegence/Sol-Milkshake with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use solintellegence/Sol-Milkshake with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # if on a CUDA device, also pip install mlx[cuda] # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("solintellegence/Sol-Milkshake") prompt = "Once upon a time in" text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- MLX LM
How to use solintellegence/Sol-Milkshake with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Generate some text mlx_lm.generate --model "solintellegence/Sol-Milkshake" --prompt "Once upon a time"
- Atomic Chat
Sol Milkshake
Sol Milkshake is a 2,990,000-parameter recurrent language model for MLX on Apple Silicon. Its checkpoint received 2,526,565,888 token exposures across pretraining and recovery training.
It uses hyperspherical optimization, shared transformer blocks, and local and rolling memory. This is a base model for text completion.
Model
| Setting | Value |
|---|---|
| Parameters | 2,990,000 |
| Token exposures | 2,526,565,888 |
| Context | 2,048 tokens |
| Tokenizer | 2,048-entry byte-level BPE |
| Hidden width | 192 |
| Stored blocks / effective applications | 5 / 11 |
| Recurrent layout | 1 prelude, 3 middle blocks used three times, 1 coda |
| Attention | 6 query heads, 2 KV heads, head dimension 32 |
| Q/K normalization | Unit normalization with RoPE |
| Attention features | XSA value subtraction and value residuals |
| Routing | Full first pass, then 75% and 50% token capacity |
| FFN | Gated SiLU, width 512 |
| Local memory | Rank-26 tensorized 2-5-gram memory |
| Rolling memory | 32-token chunks, 32 slots, width 64 |
| Token embeddings | Tied to the output head |
| Weights | MLX NPZ |
Use
Install MLX, the tokenizer library, and the Hub client:
pip install "mlx>=0.29" "tokenizers>=0.22" huggingface_hub
Download the repository and use its included loader:
from huggingface_hub import snapshot_download
import sys
model_dir = snapshot_download("solintellegence/Sol-Milkshake")
sys.path.insert(0, model_dir)
from modeling_sol_milkshake import load_model, generate
model, tokenizer = load_model(model_dir)
text = generate(
model,
tokenizer,
"The future of efficient language models is",
max_new_tokens=64,
)
print(text)
Evaluation
| Benchmark | Examples | Normalized accuracy |
|---|---|---|
| HellaSwag | 10,042 | 25.02% |
| ARC-Easy | 2,376 | 29.76% |
| ARC-Challenge | 1,172 | 22.70% |
| PIQA | 1,838 | 52.50% |
| ArithMark-3 | 1,000 | 30.60% |
The reported Intelligence Index is 3.158. The four language-model tasks use full zero-shot splits with lm-eval 0.4.12. ArithMark-3 uses independent tokenization and normalized accuracy. Raw outputs are in evals/.
The checkpoint was selected from recovery runs using this same evaluation suite. These scores therefore include checkpoint-selection effects and are not an untouched held-out estimate. They have not been independently verified.
Training
Initial pretraining used FineWeb-Edu, FinePDFs-Edu, English UltraFineWeb multi-domain and question-answer subsets, Cosmopedia, and FineMath. Recovery training mixed 70% of the original frozen curriculum with 15% Cosmopedia-v2 English and 15% FinePhrase.
The tokenizer and prepared streams were fixed. Recovery retained the existing filtering, deduplication, and decontamination rules. Training details are recorded in training_state.json.
Files
model.npz holds the MLX weights. modeling_sol_milkshake.py includes the loader and generator, and sol_config.py contains the architecture settings. The repository also includes config.json, the tokenizer files, training metadata, and evaluation outputs.
Limits and license
Milkshake is a small research model for recurrence, memory, hyperspherical optimization, and MLX inference. It can produce repetitive or incoherent text and incorrect answers. It has not been instruction-tuned.
The model is licensed under CC BY 4.0. Dataset licenses remain separate.
- Downloads last month
- 331