baileyk commited on
Commit
0495f03
·
verified ·
1 Parent(s): bb3db93

Create README.md

Browse files
Files changed (1) hide show
  1. README.md +236 -0
README.md ADDED
@@ -0,0 +1,236 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ language:
4
+ - en
5
+ library_name: transformers
6
+ datasets:
7
+ - allenai/Dolci-Think-SFT-7B
8
+ ---
9
+
10
+ ## Model Details
11
+
12
+
13
+ # Model Card for Olmo Hybrid Think SFT
14
+
15
+ We expand on our Olmo model series by introducing Olmo Hybrid, a new 7B hybrid RNN model in the Olmo family. Olmo Hybrid dramatically outperforms Olmo 3 in final performance, consistently showing roughly 2x data
16
+ efficiency on core evals over the course of our pretraining run. We also show gains in performance on long-context benchmarks, as well as improved inference efficiency
17
+ (throughput and memory) on long-context lengths by a factor of 75%.
18
+
19
+ The core models released in this batch include the following:
20
+
21
+ | **Stage** | **Olmo 3 7B Think** | **Olmo 3 32B Think** | **Olmo 3 7B Instruct** | **Olmo Hybrid 7B Think** | **Olmo Hybrid 7B Instruct** |
22
+ |--------------------------|-----------------------|------------------------|---------------------------|-------------------------------|----------------------------------|
23
+ | **Base Model** | [Olmo-3-7B](https://huggingface.co/allenai/Olmo-3-1025-7B) | [Olmo-3-32B](https://huggingface.co/allenai/Olmo-3-1125-32B) | [Olmo-3-7B](https://huggingface.co/allenai/Olmo-3-1025-7B) | [Olmo-Hybrid-7B](https://huggingface.co/allenai/Olmo-Hybrid-7B) | [Olmo-Hybrid-7B](https://huggingface.co/allenai/Olmo-Hybrid-7B) |
24
+ | **SFT** | [Olmo-3-7B-Think-SFT](https://huggingface.co/allenai/Olmo-3-7B-Think-SFT) | [Olmo-3-32B-Think-SFT](https://huggingface.co/allenai/Olmo-3-32B-Think-SFT) | [Olmo-3-7B-Instruct-SFT](https://huggingface.co/allenai/Olmo-3-7B-Instruct-SFT) | [Olmo-Hybrid-Think-SFT-7B](https://huggingface.co/allenai/Olmo-Hybrid-Think-SFT-7B) | [Olmo-Hybrid-Instruct-SFT-7B](https://huggingface.co/allenai/Olmo-Hybrid-Instruct-SFT-7B) |
25
+ | **DPO** | [Olmo-3-7B-Think-DPO](https://huggingface.co/allenai/Olmo-3-7B-Think-DPO) | [Olmo-3-32B-Think-DPO](https://huggingface.co/allenai/Olmo-3-32B-Think-DPO) | [Olmo-3-7B-Instruct-DPO](https://huggingface.co/allenai/Olmo-3-7B-Instruct-DPO) | [Olmo-Hybrid-Think-DPO-7B](https://huggingface.co/allenai/Olmo-Hybrid-Think-DPO-7B) | [Olmo-Hybrid-Instruct-DPO-7B](https://huggingface.co/allenai/Olmo-Hybrid-Instruct-DPO-7B) |
26
+ | **Final Models (RLVR)** | [Olmo-3-7B-Think](https://huggingface.co/allenai/Olmo-3-7B-Think) | [Olmo-3-32B-Think](https://huggingface.co/allenai/Olmo-3-32B-Think) | [Olmo-3-7B-Instruct](https://huggingface.co/allenai/Olmo-3-7B-Instruct) | [Olmo-Hybrid-Think-7B](https://huggingface.co/allenai/Olmo-Hybrid-Think-7B) | [Olmo-Hybrid-Instruct-7B](https://huggingface.co/allenai/Olmo-Hybrid-Instruct-7B) |
27
+
28
+ ### Olmo 3
29
+ We introduce Olmo 3, a new family of 7B and 32B models both Instruct and Think variants. Long chain-of-thought thinking improves reasoning tasks like math and coding.
30
+
31
+ Olmo is a series of **O**pen **l**anguage **mo**dels designed to enable the science of language models.
32
+ These models are pre-trained on the Dolma 3 dataset and post-trained on the Dolci datasets. We are releasing all code, checkpoints, logs (coming soon), and associated training details.
33
+
34
+
35
+ ## Installation
36
+
37
+ Olmo 3 is supported in transformers 4.57.0 or higher:
38
+ ```bash
39
+ pip install transformers>=4.57.0
40
+ ```
41
+
42
+ ## Inference
43
+
44
+ You can use OLMo with the standard HuggingFace transformers library:
45
+ ```python
46
+ from transformers import AutoModelForCausalLM, AutoTokenizer
47
+ olmo = AutoModelForCausalLM.from_pretrained("allenai/Olmo-Hybrid-Think-SFT-7B")
48
+ tokenizer = AutoTokenizer.from_pretrained("allenai/Olmo-Hybrid-Think-SFT-7B")
49
+ message = ["Who would win in a fight - a dinosaur or a cow named Moo Moo?"]
50
+ inputs = tokenizer(message, return_tensors='pt', return_token_type_ids=False)
51
+ # optional verifying cuda
52
+ # inputs = {k: v.to('cuda') for k,v in inputs.items()}
53
+ # olmo = olmo.to('cuda')
54
+ response = olmo.generate(**inputs, max_new_tokens=100, do_sample=True, top_k=50, top_p=0.95)
55
+ print(tokenizer.batch_decode(response, skip_special_tokens=True)[0])
56
+ >> '<think>Okay, so the question is who would win in a fight...'
57
+ ```
58
+
59
+ For faster performance, you can quantize the model using the following method:
60
+ ```python
61
+ AutoModelForCausalLM.from_pretrained("allenai/Olmo-Hybrid-Think-SFT-7B",
62
+ torch_dtype=torch.float16,
63
+ load_in_8bit=True) # Requires bitsandbytes
64
+ ```
65
+ The quantized model is more sensitive to data types and CUDA operations. To avoid potential issues, it's recommended to pass the inputs directly to CUDA using:
66
+ ```python
67
+ inputs.input_ids.to('cuda')
68
+ ```
69
+
70
+ We have released checkpoints for these models. For post-training, the naming convention is `step_XXXX`.
71
+
72
+
73
+ To load a specific model revision with HuggingFace, simply add the argument `revision`:
74
+ ```bash
75
+ olmo = AutoModelForCausalLM.from_pretrained("allenai/Olmo-Hybrid-Think-SFT-7B", revision="step_11000")
76
+ ```
77
+
78
+ Or, you can access all the revisions for the models via the following code snippet:
79
+ ```python
80
+ from huggingface_hub import list_repo_refs
81
+ out = list_repo_refs("allenai/Olmo-Hybrid-Think-SFT-7B")
82
+ branches = [b.name for b in out.branches]
83
+ ```
84
+
85
+ ### Chat template
86
+
87
+ ## Default System Message
88
+ The default system prompt for this model is:
89
+ ```
90
+ <|im_start|>system
91
+ You are Olmo, a helpful AI assistant built by Ai2. Your date cutoff is December 2024, and your model weights are available at https://huggingface.co/allenai.
92
+ <|im_end|>
93
+ ```
94
+
95
+ ## Chat Format
96
+
97
+ The chat template for this model is formatted as:
98
+ ```
99
+ <|im_start|>system
100
+ You are Olmo, a helpful AI assistant built by Ai2. Your date cutoff is December 2024, and your model weights are available at https://huggingface.co/allenai.
101
+ <|im_start|>user
102
+ Who would win in a fight - a dinosaur or a cow named Moo Moo?<|im_end|>
103
+ <|im_start|>assistant
104
+ <think>Okay, so the question is who would win in a fight between a dinosaur and a cow named Moo Moo.
105
+ Hmm, first I need to break this down. Let me think about the different factors involved here..... </think>
106
+ Moo Moo the cow would certinaly win.
107
+ <|endoftext|>
108
+ ```
109
+
110
+ ### Model Description
111
+
112
+ - **Developed by:** Allen Institute for AI (Ai2)
113
+ - **Model type:** a Transformer style autoregressive language model.
114
+ - **Language(s) (NLP):** English
115
+ - **License:** This model is licensed under Apache 2.0. It is intended for research and educational use in accordance with Ai2's [Responsible Use Guidelines](https://allenai.org/responsible-use).
116
+ - **Contact:** Technical inquiries: `olmo@allenai.org`. Press: `press@allenai.org`
117
+ - **Date cutoff:** Dec. 2024.
118
+
119
+
120
+ ### Model Sources
121
+
122
+ - **Project Page:** https://allenai.org/olmo
123
+ - **Repositories:**
124
+ - Open-Instruct for DPO and RLVR: https://github.com/allenai/open-instruct
125
+ - OLMo-Core for pre-training and SFT: https://github.com/allenai/OLMo-core
126
+ - OLMo-Eval for evaluation: https://github.com/allenai/OLMo-Eval
127
+ - **Paper:** https://allenai.org/papers/olmo3
128
+ <!-- - **Technical blog post:** (URL) -->
129
+ <!-- - **W&B Logs:** [SFT](()), [DPO](()), [RLVR](()) -->
130
+
131
+
132
+ ## Evaluation
133
+
134
+ | Skill | Benchmark | Olmo Hybrid Think SFT 7B | Olmo Hybrid Think DPO 7B | Olmo Hybrid Think 7B | Olmo 3 Think 7B SFT | Olmo 3 Think 7B DPO | Olmo 3 Think 7B | OpenThinker3-7B | Nemotron-Nano-9B-v2 | DeepSeek-R1-Distill-Qwen-7B | Qwen 3 8B (reasoning) | Qwen 3 VL 8B Thinker | OpenReasoning Nemotron 7B |
135
+ |-------|-----------|------------------------------|------------------------------|--------------------------|------------------|------------------|--------------|------------------|-----------------------|------------------------------|-------------------------|---------------------------|-----------------------------|
136
+ | **Math** | MATH | | | | 94.4 | 92.4 | 95.1 | 94.5 | 94.4 | 87.9 | 95.1 | 95.2 | 94.6 |
137
+ | | AIME 2024 | | | | 69.6 | 74.6 | 71.6 | 67.7 | 72.1 | 54.9 | 74.0 | 70.9 | 77.0 |
138
+ | | AIME 2025 | | | | 57.6 | 62.7 | 64.6 | 57.2 | 58.9 | 40.2 | 67.8 | 61.5 | 73.1 |
139
+ | | OMEGA | | | | 45.0 | 40.5 | 37.8 | 38.4 | 42.4 | 28.5 | 43.4 | 38.1 | 43.2 |
140
+ | **Reasoning** | BBH | | | | 84.1 | 83.7 | 86.6 | 77.1 | 86.2 | 73.5 | 84.4 | 86.8 | 81.3 |
141
+ | | ZebraLogic | | | | 57.9 | 60.6 | 66.5 | 34.9 | 60.8 | 26.1 | 85.2 | 91.2 | 22.4 |
142
+ | | AGI Eval | | | | 77.2 | 79.1 | 81.5 | 78.6 | 83.1 | 69.5 | 87.0 | 90.1 | 81.4 |
143
+ | **Coding** | HumanEval+ | | | | 88.2 | 91.4 | 89.9 | 87.4 | 89.7 | 83.0 | 80.2 | 83.7 | 89.7 |
144
+ | | MBPP+ | | | | 63.2 | 63.0 | 64.7 | 61.4 | 66.1 | 63.5 | 69.1 | 63.0 | 61.2 |
145
+ | | LCB v3 | | | | 67.8 | 75.1 | 75.2 | 68.0 | 83.4 | 58.8 | 86.2 | 85.5 | 82.3 |
146
+ | **IF** | IFEval | | | | 77.9 | 75.9 | 88.2 | 51.7 | 86.0 | 59.6 | 87.4 | 85.5 | 42.5 |
147
+ | | IFBench | | | | 30.0 | 28.3 | 41.6 | 23.0 | 34.6 | 16.7 | 37.1 | 40.4 | 23.4 |
148
+ | **Knowledge** | MMLU | | | | 74.9 | 74.8 | 77.8 | 77.4 | 84.3 | 67.9 | 85.4 | 86.5 | 80.7 |
149
+ | **QA** | PopQA | | | | 20.8 | 24.7 | 23.7 | 18.0 | 17.9 | 12.8 | 24.3 | 29.3 | 14.5 |
150
+ | | GPQA | | | | 45.8 | 48.6 | 46.2 | 47.6 | 56.2 | 54.4 | 57.7 | 61.5 | 56.6 |
151
+ | **Chat** | AE 2 | | | | 43.9 | 50.6 | 52.1 | 24.0 | 58.0 | 7.7 | 60.5 | 73.5 | 8.6 |
152
+ | **Safety** | | | | | 65.8 | 67.7 | 70.7 | 31.3 | 72.1 | 54.0 | 68.3 | 82.9 | 30.3 |
153
+
154
+ ## Model Details
155
+
156
+ #### Stage 1: SFT
157
+ - supervised fine-tuning on the Dolci-Think-SFT-7B dataset. This dataset consits of math, code, chat, and general knowledge queries.
158
+ - Datasets: [Dolci-Think-SFT-7B](https://huggingface.co/datasets/allenai/dolci-thinking-sft), [Dolci-Instruct-SFT-7B](https://huggingface.co/datasets/allenai/dolci-instruct-sft)
159
+
160
+ #### Stage 2:DPO
161
+ - direct preference optimization on the Dolci-Think-DPO-7B dataset. This dataset consits of math, code, chat, and general knowledge queries.
162
+ - Datasets: [Dolci-Think-DPO-7B](https://huggingface.co/datasets/allenai/dolci-thinking-dpo), [Dolci-Instruct-DPO-7B](https://huggingface.co/datasets/allenai/dolci-3-instruct-dpo-with-metadata)
163
+
164
+ #### Stage 3: RLVR
165
+ - reinforcement learning from verifiable rewards on the Dolci-Think-RL-7B dataset. This dataset consits of math, code, instruction-following, and general chat queries.
166
+ - Datasets: [Dolci-Think-RL-7B](https://huggingface.co/datasets/allenai/Dolci-Think-RL-7B), [Dolci-Instruct-RL-7B](https://huggingface.co/datasets/allenai/Dolci-Instruct-RL-7B)
167
+
168
+
169
+ ## Inference & Recommended Settings
170
+ We evaluated our models on the following settings. We also recommend using them for generation:
171
+ - **temperature:** `0.6`
172
+ - **top_p:** `0.95`
173
+ - **max_tokens:** `32768`
174
+
175
+ ### transformers Example
176
+ ```python
177
+ from transformers import AutoModelForCausalLM, AutoTokenizer
178
+
179
+ model_id = "allenai/Olmo-Hybrid-Think-SFT-7B"
180
+ tokenizer = AutoTokenizer.from_pretrained(model_id)
181
+ model = AutoModelForCausalLM.from_pretrained(
182
+ model_id,
183
+ device_map="auto",
184
+ )
185
+
186
+ prompt = "Who would win in a fight - a dinosaur or a cow named MooMoo?"
187
+ inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
188
+
189
+ outputs = model.generate(
190
+ **inputs,
191
+ temperature=0.6,
192
+ top_p=0.95,
193
+ max_new_tokens=32768,
194
+ )
195
+
196
+ print(tokenizer.decode(outputs[0], skip_special_tokens=True))
197
+ ```
198
+
199
+ ### vllm Example
200
+ ```python
201
+ from vllm import LLM, SamplingParams
202
+
203
+ model_id = "allenai/Olmo-Hybrid-Think-SFT-7B"
204
+ llm = LLM(model=model_id)
205
+
206
+ sampling_params = SamplingParams(
207
+ temperature=0.6,
208
+ top_p=0.95,
209
+ max_tokens=32768,
210
+ )
211
+
212
+ prompt = "Who would win in a fight - a dinosaur or a cow named MooMoo?"
213
+ outputs = llm.generate(prompt, sampling_params)
214
+ print(outputs[0].outputs[0].text)
215
+ ```
216
+
217
+ ## Bias, Risks, and Limitations
218
+ Like any base language model or fine-tuned model without safety filtering, these models can easily be prompted by users to generate harmful and sensitive content. Such content may also be produced unintentionally, especially in cases involving bias, so we recommend that users consider the risks when applying this technology. Additionally, many statements from OLMo or any LLM are often inaccurate, so facts should be verified.
219
+
220
+
221
+ ## Citation
222
+
223
+ ```
224
+ @misc{olmo2025olmo3,
225
+ title={Olmo 3},
226
+ author={Team Olmo and Allyson Ettinger and Amanda Bertsch and Bailey Kuehl and David Graham and David Heineman and Dirk Groeneveld and Faeze Brahman and Finbarr Timbers and Hamish Ivison and Jacob Morrison and Jake Poznanski and Kyle Lo and Luca Soldaini and Matt Jordan and Mayee Chen and Michael Noukhovitch and Nathan Lambert and Pete Walsh and Pradeep Dasigi and Robert Berry and Saumya Malik and Saurabh Shah and Scott Geng and Shane Arora and Shashank Gupta and Taira Anderson and Teng Xiao and Tyler Murray and Tyler Romero and Victoria Graf and Akari Asai and Akshita Bhagia and Alexander Wettig and Alisa Liu and Aman Rangapur and Chloe Anastasiades and Costa Huang and Dustin Schwenk and Harsh Trivedi and Ian Magnusson and Jaron Lochner and Jiacheng Liu and Lester James V. Miranda and Maarten Sap and Malia Morgan and Michael Schmitz and Michal Guerquin and Michael Wilson and Regan Huff and Ronan Le Bras and Rui Xin and Rulin Shao and Sam Skjonsberg and Shannon Zejiang Shen and Shuyue Stella Li and Tucker Wilde and Valentina Pyatkin and Will Merrill and Yapei Chang and Yuling Gu and Zhiyuan Zeng and Ashish Sabharwal and Luke Zettlemoyer and Pang Wei Koh and Ali Farhadi and Noah A. Smith and Hannaneh Hajishirzi},
227
+ year={2025},
228
+ eprint={2512.13961},
229
+ archivePrefix={arXiv},
230
+ primaryClass={cs.CL},
231
+ url={https://arxiv.org/abs/2512.13961},
232
+ }
233
+ ```
234
+
235
+ ## Model Card Contact
236
+ For errors in this model card, contact `olmo@allenai.org`.