How to use from
SGLang
Install from pip and serve model
# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
    --model-path "ParallaxOpen/Vela-Lumen-31M-v1.2" \
    --host 0.0.0.0 \
    --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "ParallaxOpen/Vela-Lumen-31M-v1.2",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'
Use Docker images
docker run --gpus all \
    --shm-size 32g \
    -p 30000:30000 \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=<secret>" \
    --ipc=host \
    lmsysorg/sglang:latest \
    python3 -m sglang.launch_server \
        --model-path "ParallaxOpen/Vela-Lumen-31M-v1.2" \
        --host 0.0.0.0 \
        --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "ParallaxOpen/Vela-Lumen-31M-v1.2",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'
Quick Links

Vela-Lumen-31M v1.2

A 47,846,912-parameter (47.8M) dense decoder-only language model, trained from scratch for 200,000 steps on a single RTX 5060 laptop GPU.

Errata and release history

This repository previously contained no weight files at all โ€” only README.md and .gitattributes. The card claimed "GGUF, SafeTensors, tokenizer included" and linked a vela-lumen-31m-v1.2-f16.gguf that returned 404. That was reported publicly and the report was correct.

The weights are now published. The files below are real and verified. Two earlier claims were also wrong and are corrected here:

  1. Parameter count. The card said 35.5M. The checkpoint contains a separate lm_head.weight of shape [24189, 512] (12,384,768 parameters), so it is untied and the true count is 47,846,912. Note that the checkpoint's own n_params metadata field records 35,462,144 โ€” that is the tied count and it is wrong. The figure above is measured from the safetensors header.
  2. No GGUF is published, and none can be. SmallLM is a dense GELU network; llama.cpp's llama architecture is SwiGLU with a fused gate/up projection. These are not convertible. An earlier GGUF linked from this card was 95.7 MB where a faithful FP16 export of 35,462,144 parameters is 70.9 MB โ€” not a real conversion, and it produced garbage. Do not look for one here, and do not trust any GGUF claiming to be this model.

Files

File Size Notes
model.safetensors 191.4 MB float32, 67 tensors, 47,846,912 params
modeling_vela.py 14 KB self-contained transformers implementation
config.json 2 KB includes auto_map
tokenizer.json 1.7 MB 24,189-token BPE
tokenizer_config.json 0.4 KB special tokens + chat template

Architecture

Component Value
Parameters 47,846,912 (47.8M), untied
Layers 8
Hidden size 512
Attention heads 8 query / 4 KV (GQA)
FFN 2,048, GELU (not SwiGLU)
Norm RMSNorm (pre-norm)
Position RoPE, NeoX pair rotation, theta 10,000
Vocab 24,189 (BPE)
Max sequence 128 tokens
Trained steps 200,000

Loading

This architecture is not in mainline transformers. The modeling code ships with the weights and is wired through auto_map, so it loads with the standard auto classes:

from transformers import AutoModelForCausalLM, AutoTokenizer

REPO = "ParallaxOpen/Vela-Lumen-31M-v1.2"

tok = AutoTokenizer.from_pretrained(REPO, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(REPO, trust_remote_code=True)

prompt = "<|im_start|>user\nWhat is 12 * 12?<|im_end|>\n<|im_start|>assistant\n"
ids = tok(prompt, return_tensors="pt")
out = model.generate(**ids, max_new_tokens=48, do_sample=False,
                     repetition_penalty=1.15)
print(tok.decode(out[0][ids["input_ids"].shape[1]:], skip_special_tokens=True))

trust_remote_code=True is required; without it Transformers raises ValueError: ... does not recognize this architecture. Read modeling_vela.py first if you would rather not execute fetched code.

Measured results

Methodology matters here. These are loss-based measurements โ€” next-token-prediction loss for GSM8K and TruthfulQA, and loss-comparison for multiple-choice tasks. They are not free-running generation, and this line of work has directly measured how misleading that distinction is: a later checkpoint scored 60.0% on a held-out split drawn from the training distribution and 8.3% on hand-written questions outside every training source. Treat the numbers below as training-health signals only.

Benchmark Score n Method
GSM8K 50.5% 8,792 loss-based, teacher-forced
ARC-Challenge 26.0% 2,590 loss-based (MC)
HellaSwag 24.9% 5,000 loss-based (MC)
TruthfulQA 25.5% 817 loss-based, teacher-forced
WinoGrande 51.2% 5,000 loss-based (MC)

Known limitations

  • It is a 200k-step model, not a chat model. v1.2 was trained before the supervised fine-tuning and turn-format work that later versions in this line received. It will not hold a conversation.
  • Greedy decoding loops. Use repetition_penalty=1.15 or sample. This is an inference-time mitigation, not a fix.
  • Short context. 128 tokens. Longer prompts are truncated.
  • No GGUF, for the architectural reason above.

Training data

Chat, math, and code data, mixed and deduplicated. 1,000 examples were held out before training and are not referenced by the training script.

License

CC BY-NC 4.0 โ€” research and non-commercial use.

Downloads last month
129
Safetensors
Model size
47.8M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support