Instructions to use allenai/OLMo-2-0425-1B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use allenai/OLMo-2-0425-1B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="allenai/OLMo-2-0425-1B")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("allenai/OLMo-2-0425-1B") model = AutoModelForCausalLM.from_pretrained("allenai/OLMo-2-0425-1B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use allenai/OLMo-2-0425-1B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "allenai/OLMo-2-0425-1B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "allenai/OLMo-2-0425-1B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/allenai/OLMo-2-0425-1B
- SGLang
How to use allenai/OLMo-2-0425-1B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "allenai/OLMo-2-0425-1B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "allenai/OLMo-2-0425-1B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "allenai/OLMo-2-0425-1B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "allenai/OLMo-2-0425-1B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use allenai/OLMo-2-0425-1B with Docker Model Runner:
docker model run hf.co/allenai/OLMo-2-0425-1B
Which checkpoint (or soup) corresponds to the final base model?
The model card for OLMo-2-0425-1B says there was no model merging, and there was only 1 run on 50B tokens. But looking at the checkpoints, there are checkpoints stage2-ingredient1-step23852-tokens51B, stage2-ingredient2-step23852-tokens51B, and stage2-ingredient3-step23852-tokens51B. How do these correspond to the final released model? I think it's probably a model soup of the three and the model card is misleading, but just wanted to check. Thanks!
Hey @Proyag , great question, You’re right that the checkpoints mention multiple ingredients, but to clarify: there was no model merging or model soup involved in the final release. The OLMo-2-0425-1B model was trained in a single run to 50B tokens. We tried different recipes across ingredients 2 and 3, but these were independent exploratory runs, not merged into the final model. The released final checkpoint corresponds to ingredient 1, selected based on performance.
Ah ok, thanks for your response!
Hey @Proyag , the released final checkpoint corresponds to ingredient 3, not ingredient 1. I am sorry for the confusion and mistake in previous response. I will add it to readme.