Instructions to use ParallaxOpen/Vela-Lumen-31M-v1.2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ParallaxOpen/Vela-Lumen-31M-v1.2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ParallaxOpen/Vela-Lumen-31M-v1.2", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("ParallaxOpen/Vela-Lumen-31M-v1.2", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ParallaxOpen/Vela-Lumen-31M-v1.2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ParallaxOpen/Vela-Lumen-31M-v1.2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ParallaxOpen/Vela-Lumen-31M-v1.2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ParallaxOpen/Vela-Lumen-31M-v1.2
- SGLang
How to use ParallaxOpen/Vela-Lumen-31M-v1.2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ParallaxOpen/Vela-Lumen-31M-v1.2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ParallaxOpen/Vela-Lumen-31M-v1.2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ParallaxOpen/Vela-Lumen-31M-v1.2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ParallaxOpen/Vela-Lumen-31M-v1.2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use ParallaxOpen/Vela-Lumen-31M-v1.2 with Docker Model Runner:
docker model run hf.co/ParallaxOpen/Vela-Lumen-31M-v1.2
Vela-Lumen-31M v1.2
A 47,846,912-parameter (47.8M) dense decoder-only language model, trained from scratch for 200,000 steps on a single RTX 5060 laptop GPU.
Errata and release history
This repository previously contained no weight files at all โ only
README.mdand.gitattributes. The card claimed "GGUF, SafeTensors, tokenizer included" and linked avela-lumen-31m-v1.2-f16.ggufthat returned 404. That was reported publicly and the report was correct.The weights are now published. The files below are real and verified. Two earlier claims were also wrong and are corrected here:
- Parameter count. The card said 35.5M. The checkpoint contains a separate
lm_head.weightof shape[24189, 512](12,384,768 parameters), so it is untied and the true count is 47,846,912. Note that the checkpoint's ownn_paramsmetadata field records 35,462,144 โ that is the tied count and it is wrong. The figure above is measured from the safetensors header.- No GGUF is published, and none can be. SmallLM is a dense GELU network;
llama.cpp'sllamaarchitecture is SwiGLU with a fused gate/up projection. These are not convertible. An earlier GGUF linked from this card was 95.7 MB where a faithful FP16 export of 35,462,144 parameters is 70.9 MB โ not a real conversion, and it produced garbage. Do not look for one here, and do not trust any GGUF claiming to be this model.
Files
| File | Size | Notes |
|---|---|---|
model.safetensors |
191.4 MB | float32, 67 tensors, 47,846,912 params |
modeling_vela.py |
14 KB | self-contained transformers implementation |
config.json |
2 KB | includes auto_map |
tokenizer.json |
1.7 MB | 24,189-token BPE |
tokenizer_config.json |
0.4 KB | special tokens + chat template |
Architecture
| Component | Value |
|---|---|
| Parameters | 47,846,912 (47.8M), untied |
| Layers | 8 |
| Hidden size | 512 |
| Attention heads | 8 query / 4 KV (GQA) |
| FFN | 2,048, GELU (not SwiGLU) |
| Norm | RMSNorm (pre-norm) |
| Position | RoPE, NeoX pair rotation, theta 10,000 |
| Vocab | 24,189 (BPE) |
| Max sequence | 128 tokens |
| Trained steps | 200,000 |
Loading
This architecture is not in mainline transformers. The modeling code ships
with the weights and is wired through auto_map, so it loads with the standard
auto classes:
from transformers import AutoModelForCausalLM, AutoTokenizer
REPO = "ParallaxOpen/Vela-Lumen-31M-v1.2"
tok = AutoTokenizer.from_pretrained(REPO, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(REPO, trust_remote_code=True)
prompt = "<|im_start|>user\nWhat is 12 * 12?<|im_end|>\n<|im_start|>assistant\n"
ids = tok(prompt, return_tensors="pt")
out = model.generate(**ids, max_new_tokens=48, do_sample=False,
repetition_penalty=1.15)
print(tok.decode(out[0][ids["input_ids"].shape[1]:], skip_special_tokens=True))
trust_remote_code=True is required; without it Transformers raises
ValueError: ... does not recognize this architecture. Read modeling_vela.py
first if you would rather not execute fetched code.
Measured results
Methodology matters here. These are loss-based measurements โ next-token-prediction loss for GSM8K and TruthfulQA, and loss-comparison for multiple-choice tasks. They are not free-running generation, and this line of work has directly measured how misleading that distinction is: a later checkpoint scored 60.0% on a held-out split drawn from the training distribution and 8.3% on hand-written questions outside every training source. Treat the numbers below as training-health signals only.
| Benchmark | Score | n | Method |
|---|---|---|---|
| GSM8K | 50.5% | 8,792 | loss-based, teacher-forced |
| ARC-Challenge | 26.0% | 2,590 | loss-based (MC) |
| HellaSwag | 24.9% | 5,000 | loss-based (MC) |
| TruthfulQA | 25.5% | 817 | loss-based, teacher-forced |
| WinoGrande | 51.2% | 5,000 | loss-based (MC) |
Known limitations
- It is a 200k-step model, not a chat model. v1.2 was trained before the supervised fine-tuning and turn-format work that later versions in this line received. It will not hold a conversation.
- Greedy decoding loops. Use
repetition_penalty=1.15or sample. This is an inference-time mitigation, not a fix. - Short context. 128 tokens. Longer prompts are truncated.
- No GGUF, for the architectural reason above.
Training data
Chat, math, and code data, mixed and deduplicated. 1,000 examples were held out before training and are not referenced by the training script.
License
CC BY-NC 4.0 โ research and non-commercial use.
- Downloads last month
- 129
Install from pip and serve model
# Install vLLM from pip: pip install vllm# Start the vLLM server: vllm serve "ParallaxOpen/Vela-Lumen-31M-v1.2"# Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ParallaxOpen/Vela-Lumen-31M-v1.2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'