Manusagents commited on
Commit
112de38
·
verified ·
1 Parent(s): 6ac1cdb

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +238 -312
README.md CHANGED
@@ -1,3 +1,36 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  # Step-5-Preview
2
 
3
  <div align="center">
@@ -15,15 +48,11 @@
15
 
16
  </div>
17
 
18
- <div style="border-left: 6px solid #1890ff; padding: 16px; border-radius: 8px; margin: 20px 0;">
19
- <strong>🔥 Step-5-Preview is now available!</strong><br>
20
- We are excited to release <strong>Step-5-Preview</strong>, our flagship foundation model for real-world agentic work.
21
- It is a 600B-parameter sparse Mixture-of-Experts model with 27B active parameters, a 1M-token context window,
22
- and native support for text, image, and video inputs.
23
- <br><br>
24
- <strong>Weights are available now</strong> on Hugging Face (<code>SHSLab/Step-5-Preview-BF16</code>).
25
- Try it via our API, or deploy locally with vLLM / SGLang.
26
- </div>
27
 
28
  ---
29
 
@@ -31,44 +60,30 @@
31
 
32
  - [Introduction](#-introduction)
33
  - [Key Features](#-key-features)
34
- - [Model Architecture](#-model-architecture)
35
  - [Model Specifications](#-model-specifications)
36
- - [Training Data](#-training-data)
37
  - [Benchmark Results](#-benchmark-results)
38
- - [Agentic Capabilities](#-agentic-capabilities)
39
- - [Real-World Use Cases](#-real-world-use-cases)
40
  - [Quickstart](#-quickstart)
41
  - [Deployment](#-deployment)
42
- - [Evaluation](#-evaluation)
43
- - [Limitations](#-limitations)
44
- - [Ethical Considerations](#-ethical-considerations)
45
- - [Hardware Requirements](#-hardware-requirements)
46
- - [Performance Metrics](#-performance-metrics)
47
- - [Citation](#-citation)
48
  - [License](#-license)
49
  - [Contact](#-contact)
 
50
 
51
  ---
52
 
53
  ## 🚀 Introduction
54
 
55
- **Step-5-Preview** is StepFun's flagship foundation model, designed from the ground up for **real-world agentic tasks**.
56
- It targets professional domains such as **AI coding, software engineering, professional knowledge work, and financial analysis**.
57
 
58
- StepFun's core philosophy for Step 5 is the **"Pareto Frontier"** — achieving the optimal balance between intelligence and cost.
59
- While previous scaling efforts focused on trading more compute for stronger intelligence, the next phase requires improving the
60
- **efficiency of converting compute into intelligence**.
61
 
62
- <div style="border-left: 6px solid #fa8c16; padding: 16px; border-radius: 8px; margin: 20px 0;">
63
- <strong>💡 Why Step 5 Preview?</strong><br>
64
- • <strong>600B total parameters, only 27B active</strong> — near-frontier performance at a fraction of the compute.<br>
65
- • <strong>1M-token context window</strong> without proportional cost increases.<br>
66
- • <strong>Competitive benchmark scores</strong> against models with 3–5× more parameters.<br>
67
- • <strong>Built for agents</strong> — long-horizon reasoning, tool use, and autonomous execution.
68
- </div>
69
 
70
- Step-5-Preview represents a generational leap, with StepFun **skipping the entire Step 4.x line** entirely, going directly from
71
- Step-3.7-Flash to Step 5. This decision reflects the magnitude of improvement achieved in this release.
72
 
73
  ---
74
 
@@ -85,39 +100,7 @@ Step-3.7-Flash to Step 5. This decision reflects the magnitude of improvement ac
85
  - **Parallel Tool Calling:** Natively supported for agentic workflows.
86
  - **Strict JSON Schema Output:** Reliable integration into structured systems.
87
  - **OpenAI-Compatible API:** Available via Step API and third-party gateways.
88
- - **Open Weights:** BF16 checkpoint available now under `SHSLab/Step-5-Preview-BF16`.
89
-
90
- ---
91
-
92
- ## 🏗️ Model Architecture
93
-
94
- <div align="center">
95
- <img src="./Step-5/architecture.png" alt="Step 5 Architecture" width="85%">
96
- </div>
97
-
98
- ### 92-Layer "Narrow but Deep" Design
99
-
100
- Step-5-Preview uses a **92-layer Transformer** with a narrow-deep configuration. This design is specifically intended to create
101
- **longer information propagation paths** for implicit multi-hop reasoning during long prefill operations.
102
-
103
- ### Sparse Grouped-Query Attention (GQA) with Block-Wise Token Merging
104
-
105
- To handle the 1M-token context window efficiently, Step-5-Preview introduces **Sparse GQA with block-wise token merging**.
106
- This mechanism uses sparse indexing to select only historical information relevant to the current task, reducing the number of tokens
107
- that actually enter attention computation. StepFun states this cuts indexer and top-k selection costs to approximately
108
- **one-eighth** of a denser baseline.
109
-
110
- <div style="border-left: 6px solid #52c41a; padding: 16px; border-radius: 8px; margin: 20px 0;">
111
- <strong>⚡ Efficiency-First Scaling</strong><br>
112
- Step 5 Preview achieves near-frontier performance with <strong>600B total parameters</strong> but only
113
- <strong>27B active per token</strong>. This is the core of StepFun's efficiency-first philosophy.
114
- </div>
115
-
116
- ### Multimodal Encoder
117
-
118
- The model incorporates a unified multimodal encoder that processes text, images, and video frames into a shared latent space.
119
- Video is sampled at adaptive frame rates and encoded with temporal attention, allowing the model to understand motion and
120
- long-range dependencies in screen recordings, demonstrations, and real-world footage.
121
 
122
  ---
123
 
@@ -139,29 +122,12 @@ long-range dependencies in screen recordings, demonstrations, and real-world foo
139
  | **Reasoning Effort** | `low` / `medium` / `high` (`xhigh`) |
140
  | **Tool Calling** | Parallel, strict JSON schema |
141
  | **Intelligence Index** | 44 (Artificial Analysis v4.3.2) |
142
- | **Open Weights** | BF16 checkpoint available now |
143
  | **API Availability** | Immediate (OpenAI-compatible) |
144
  | **License** | StepFun Community License |
145
 
146
  ---
147
 
148
- ## 📚 Training Data
149
-
150
- Step-5-Preview was trained on a massive, carefully curated corpus spanning:
151
-
152
- - **Code repositories** from multiple languages (Python, C++, Rust, JavaScript, Go, etc.)
153
- - **Technical documentation**, API references, and software engineering forums
154
- - **Scientific papers** in computer science, mathematics, physics, and finance
155
- - **Financial reports**, earnings calls, and market analyses
156
- - **Multimodal data** including screenshots, UI mockups, video tutorials, and screen recordings
157
- - **Agentic trajectories** from simulated and real tool-use environments
158
-
159
- The data mixture was optimized for long-horizon reasoning and tool use, with a strong emphasis on real-world professional tasks.
160
- All data was filtered for quality, safety, and license compliance. The training process used a combination of next-token prediction
161
- and reinforcement learning from human feedback (RLHF) with a focus on agentic objectives.
162
-
163
- ---
164
-
165
  ## 📊 Benchmark Results
166
 
167
  <div align="center">
@@ -172,27 +138,18 @@ and reinforcement learning from human feedback (RLHF) with a focus on agentic ob
172
 
173
  **Overall Score: 44** (Intelligence Index v4.3.2, recalibrated September 7, 2026)
174
 
175
- This places Step-5-Preview among the **top open-weight models globally**, on par with models like
176
- Kimi K3 Max (approximately 5× larger at 2.8T parameters) and Qwen3.8 Max. The index covers 10 evaluations including
177
- AA-Briefcase, GDPval-AA v2, Terminal-Bench 4.0, SciCode, and Humanity's Last Exam.
178
 
179
- <div style="border: 1px solid #d9d9d9; border-radius: 12px; padding: 20px 24px; margin: 24px 0; background: linear-gradient(135deg, #fafafa 0%, #f0f5ff 100%);">
180
-
181
- <h4 style="margin-top:0;">📐 How to read these tables</h4>
182
-
183
- - <strong>Step-5-Preview</strong> is always the first data column, for fast scanning.
184
- - <strong>Bold</strong> marks the <strong>best score in the row</strong> across all six models.
185
- - <strong>🥇</strong> flags the category leader for that benchmark.
186
- - <strong>—</strong> indicates the model was not evaluated or did not report a score.
187
- - All Step-5-Preview results use the <code>high</code> reasoning-effort setting unless noted.
188
-
189
- <strong>Comparison set:</strong> Step-5-Preview · GPT-6 Astra (Max) · Fable 5.1 · Claude Opus 5 (Max) · Kimi K3 (Max) · GLM-5.3 (Max)
190
-
191
- <strong>Open-weight models in this comparison:</strong> Step-5-Preview, Kimi K3 (Max), GLM-5.3 (Max), and Qwen3.8 Max (referenced in the index).
192
- <strong>Closed-source models:</strong> GPT-6 Astra (Max), Claude Opus 5 (Max).
193
- <em>Fable 5.1: availability not specified in this document.</em>
194
-
195
- </div>
196
 
197
  ---
198
 
@@ -205,14 +162,6 @@ AA-Briefcase, GDPval-AA v2, Terminal-Bench 4.0, SciCode, and Humanity's Last Exa
205
  | **AA-LCR v1.1** | 88.3% | 80.7% | 85.3% | 79.3% | **88.7%** 🥇 | 79.7% |
206
  | **CritPt** | 20.9% | **31.7%** 🥇 | 29.7% | 29.1% | 23.4% | 19.1% |
207
 
208
- <div style="border-left: 6px solid #2f54eb; padding: 16px; border-radius: 8px; margin: 20px 0;">
209
- <strong>📌 Read-out</strong><br>
210
- Step-5-Preview sits in the <strong>top tier of open-weight reasoning</strong>: GPQA Diamond <strong>93.5%</strong> matches Kimi K3
211
- exactly and lands within 0.5 points of Fable 5.1. On <strong>AA-LCR v1.1</strong> — long-context reasoning — it scores
212
- <strong>88.3%</strong>, effectively tied with Kimi K3 (88.7%) and <strong>+7.6 points ahead of GPT-6 Astra</strong> and
213
- <strong>+9.0 points ahead of Claude Opus 5</strong>, despite those models being far larger.
214
- </div>
215
-
216
  ---
217
 
218
  ### 💻 Coding & Software Engineering
@@ -234,15 +183,6 @@ AA-Briefcase, GDPval-AA v2, Terminal-Bench 4.0, SciCode, and Humanity's Last Exa
234
  | **StepCode-Bench-Daily** | 64.9% | — | — | **77.6%** 🥇 | 57.7% | 69.1% |
235
  | **StepCode-Bench-General** | 65.0% | 64.3% | — | **68.3%** 🥇 | 65.2% | 62.0% |
236
 
237
- <div style="border-left: 6px solid #13c2c2; padding: 16px; border-radius: 8px; margin: 20px 0;">
238
- <strong>📌 Read-out</strong><br>
239
- Among <strong>open-weight models</strong>, Step-5-Preview leads on <strong>DeepSWE v1.1</strong> (67.7% vs. Kimi K3 67.5% and
240
- GLM-5.3 66.9%), <strong>StepCodeBench</strong> (49.0% vs. 43.9% and 40.2%), <strong>ProgramBench</strong> (80.5% vs. 77.8% and 72.0%),
241
- and <strong>SWE-Atlas-QnA</strong> (63.6%). It also posts the best <strong>CyberGym</strong> score in the comparison set at
242
- <strong>84.7%</strong>. On <strong>Terminal-Bench 4</strong> it reaches <strong>33.3%</strong> —
243
- <strong>2.6× Kimi K3</strong> (12.6%) — though the larger closed models still hold the absolute lead.
244
- </div>
245
-
246
  ---
247
 
248
  ### 🤖 Agents, Tool Use & Automation
@@ -263,14 +203,6 @@ AA-Briefcase, GDPval-AA v2, Terminal-Bench 4.0, SciCode, and Humanity's Last Exa
263
  | **BrowseComp** | 88.7% | **91.5%** 🥇 | — | 90.2% | 91.2% | — |
264
  | **HLE w/ tools** | 59.4% | 57.2% | **65.0%** 🥇 | 63.6% | 56.0% | 62.5% |
265
 
266
- <div style="border-left: 6px solid #f5222d; padding: 16px; border-radius: 8px; margin: 20px 0;">
267
- <strong>📌 Read-out</strong><br>
268
- Step-5-Preview is a <strong>strong mid-frontier agentic model</strong>. It outperforms GPT-6 Astra on <strong>τ³-Banking</strong>
269
- (42.5% vs. 41.4%) and <strong>Draco</strong> (83.3% vs. 76.8%), and beats Kimi K3 on <strong>MCP-Atlas</strong> (85.6% vs. 85.3%)
270
- and <strong>Draco</strong> (83.3% vs. 78.5%). With tools, HLE rises from <strong>46.5% → 59.4%</strong>, a
271
- <strong>+12.9-point</strong> tool-augmentation gain.
272
- </div>
273
-
274
  ---
275
 
276
  ### 💰 Finance & Professional Work
@@ -282,18 +214,9 @@ AA-Briefcase, GDPval-AA v2, Terminal-Bench 4.0, SciCode, and Humanity's Last Exa
282
  | **FinStepBench-FinanceDR** | 55.8% | 45.0% | — | **59.1%** 🥇 | 48.9% | 53.3% |
283
  | **FrontierFinance** | 66.4% | 55.0% | — | **69.7%** 🥇 | 62.6% | 64.1% |
284
  | **OfficeQA Pro** | 60.3% | **67.7%** 🥇 | — | 64.7% | 62.6% | 59.1% |
285
- | **Spreadsheet v2** | 29.4% | 31.4% | — | **32.8%** 🥇 | 31.9% | 30.5% |
286
  | **GDP.pdf** | 14.8% | **31.0%** 🥇 | 26.2% | 21.6% | 22.0% | 11.2% |
287
 
288
- <div style="border-left: 6px solid #a0d911; padding: 16px; border-radius: 8px; margin: 20px 0;">
289
- <strong>📌 Read-out</strong><br>
290
- Finance is a Step-5-Preview <strong>strength</strong>. On <strong>FrontierFinance</strong> it scores <strong>66.4%</strong>,
291
- beating GPT-6 Astra by <strong>+11.4 points</strong>, Kimi K3 by <strong>+3.8</strong>, and GLM-5.3 by <strong>+2.3</strong> —
292
- trailing only Claude Opus 5. On <strong>FinStepBench-FinanceDR</strong> it reaches <strong>55.8%</strong>, ahead of
293
- GPT-6 Astra (+10.8), Kimi K3 (+6.9), and GLM-5.3 (+2.5). <strong>FinStepBench-LiveSearch</strong> lands at
294
- <strong>74.5%</strong>, tied with GPT-6 Astra.
295
- </div>
296
-
297
  ---
298
 
299
  ### 👁️ Multimodal
@@ -303,121 +226,22 @@ AA-Briefcase, GDPval-AA v2, Terminal-Bench 4.0, SciCode, and Humanity's Last Exa
303
  | **MMMU-Pro** | 76.0% | **87.0%** 🥇 | — | 85.0% | 81.0% | — |
304
  | **GDP.pdf** | 14.8% | **31.0%** 🥇 | 26.2% | 21.6% | 22.0% | 11.2% |
305
 
306
- <div style="border-left: 6px solid #faad14; padding: 16px; border-radius: 8px; margin: 20px 0;">
307
- <strong>📌 Read-out</strong><br>
308
- Multimodal reasoning is the <strong>most direct improvement opportunity</strong> for the next release. MMMU-Pro at
309
- <strong>76.0%</strong> trails GPT-6 Astra (87.0%), Claude Opus 5 (85.0%), and Kimi K3 (81.0%), and document-heavy
310
- visual QA (<strong>GDP.pdf</strong>) remains a known weak spot at <strong>14.8%</strong>.
311
- </div>
312
-
313
- ---
314
-
315
- ### 🏅 Category Leadership Summary
316
-
317
- <div style="border: 1px solid #d9d9d9; border-radius: 12px; padding: 20px 24px; margin: 24px 0; background: linear-gradient(135deg, #ffffff 0%, #f6ffed 100%);">
318
-
319
- | Domain | Step-5-Preview Standing | Highlight |
320
- |:---|:---|:---|
321
- | **Long-Context Reasoning** | 🥈 Near-tied for #1 overall | AA-LCR v1.1 **88.3%** (+7.6 over GPT-6 Astra, +9.0 over Opus 5) |
322
- | **Coding / SWE (open-weight)** | 🥇 #1 open-weight | DeepSWE v1.1 **67.7%**, StepCodeBench **49.0%**, ProgramBench **80.5%** |
323
- | **Security / Cyber** | 🥇 #1 in set | CyberGym **84.7%** |
324
- | **Finance** | 🥈 #2 overall, #1 open-weight | FrontierFinance **66.4%**, FinanceDR **55.8%** |
325
- | **Terminal Agents** | 🥈 #2 open-weight | Terminal-Bench 4 **33.3%** (GLM-5.3 41.9% > Step 33.3% > Kimi 12.6%) |
326
- | **Tool Use** | 🥉 Competitive | MCP-Atlas **85.6%**, HLE w/ tools **59.4%** |
327
- | **General Knowledge** | Top tier | GPQA Diamond **93.5%** (tied #3), HLE **46.5%** (#5) |
328
- | **Multimodal** | 🔻 Trailing | MMMU-Pro **76.0%** — targeted for improvement |
329
-
330
- </div>
331
-
332
  <details>
333
- <summary><strong>📝 Benchmark Methodology Notes</strong> (click to expand)</summary>
334
 
335
  - **DeepSWE v1.1** was evaluated using the SWE-agent harness with `temperature=1.0` and `top_p=0.95`.
336
  - **StepCodeBench** achieved **49.0% avg@4**.
337
  - **GDPval-AA v2** results are from Artificial Analysis as of September 19, 2026.
338
  - **AA-LCR v1.1**: Step-5-Preview scored **88.3%**, statistically tied with Kimi K3 (88.7%).
339
- - **Terminal-Bench 4**: Step-5-Preview **33.3%** vs. Kimi K3 **~12.6%** (2.6×) and DeepSeek V4.1 Flash **26.8%** (1.24×).
340
- - **SciCode**: Step-5-Preview scored **58.9%**, above GPT-6 Astra (56.5%) and Claude Opus 5 (56.4%).
341
  - **Multimodal**: MMMU-Pro and GDP.pdf were run with the unified multimodal encoder at default resolution.
342
  - **Output Speed**: 99.8 tokens/sec (GLM-5.3: 72.1 tokens/sec).
343
  - **Time to First Token**: 2.96 seconds (GLM-5.3: 2.99s; Claude Opus 5: 56.84s at max effort).
344
  - Missing entries (**—**) reflect benchmarks that were not publicly reported for that model at the time of writing.
345
- - Detailed benchmark descriptions in the Evaluation table are sourced from public benchmark documentation.
346
 
347
  </details>
348
 
349
- ### Benchmark Takeaways
350
-
351
- <div style="border-left: 6px solid #2f54eb; padding: 16px; border-radius: 8px; margin: 20px 0;">
352
- <strong>🧠 Long-Context Reasoning</strong><br>
353
- AA-LCR v1.1 at <strong>88.3%</strong> is one of Step-5-Preview's standout results — effectively tied for first with Kimi K3
354
- and well ahead of models many times its size. This validates the 92-layer narrow-deep design and Sparse GQA.
355
- </div>
356
-
357
- <div style="border-left: 6px solid #13c2c2; padding: 16px; border-radius: 8px; margin: 20px 0;">
358
- <strong>💻 Coding & Software Engineering</strong><br>
359
- Step-5-Preview <strong>leads all open-weight models</strong> on DeepSWE v1.1, StepCodeBench, ProgramBench, and SWE-Atlas-QnA,
360
- surpassing Kimi K3 and GLM-5.3. It trails only the larger closed-source models (Claude Opus 5 and GPT-6 Astra).
361
- </div>
362
-
363
- <div style="border-left: 6px solid #f5222d; padding: 16px; border-radius: 8px; margin: 20px 0;">
364
- <strong>🤖 Agentic Tasks</strong><br>
365
- Strong performance on Terminal-Bench 4 (<strong>33.3%</strong>) and HLE with tools (<strong>59.4%</strong>).
366
- The Terminal-Bench score is <strong>2.6× higher than Kimi K3</strong> and <strong>1.24× higher than DeepSeek V4.1 Flash</strong>.
367
- </div>
368
-
369
- <div style="border-left: 6px solid #a0d911; padding: 16px; border-radius: 8px; margin: 20px 0;">
370
- <strong>💰 Financial & Deep Research</strong><br>
371
- Highly competitive on FrontierFinance (<strong>66.4%</strong>) and FinStepBench-FinanceDR (<strong>55.8%</strong>),
372
- nearly matching top closed-source models like Claude Opus 5 and outperforming both GPT-6 Astra and Kimi K3 by significant margins.
373
- </div>
374
-
375
- ---
376
-
377
- ## 🤖 Agentic Capabilities
378
-
379
- <div align="center">
380
- <img src="./Step-5/agentic_workflow.png" alt="Agentic Workflow" width="90%">
381
- </div>
382
-
383
- ### 24-Hour Autonomous GPU Kernel Optimization
384
-
385
- In a landmark demonstration of sustained agentic execution, Step-5-Preview was tasked with **autonomously optimizing an H100 GPU kernel for up to 24 consecutive hours**. The model:
386
-
387
- - Independently modified code
388
- - Ran tests and compared results
389
- - Iterated based on performance outcomes
390
- - **Reached 508 TFLOPS after approximately 22 hours**
391
-
392
- For comparison, **Claude Opus 5 achieved 493 TFLOPS** in the same experiment. This demonstrates Step-5-Preview's ability to sustain productive work over extended periods without human intervention.
393
-
394
- ### Automated Post-Training Experiments
395
-
396
- In another 24-hour experiment, Step-5-Preview autonomously improved the accuracy of **Qwen3-30B-A3B on AIME24 from 53.3% to 60%** through automated post-training experiments. This showcases the model's capacity for self-directed research and optimization.
397
-
398
- ### Long-Horizon Agent Workflows
399
-
400
- The model is specifically optimized for agent workflows that require:
401
-
402
- - Searching and information retrieval
403
- - Running code and processing tool returns
404
- - Multi-turn tool calls with sustained execution
405
- - Iterative refinement based on intermediate results
406
- - Self-correction and error recovery over thousands of steps
407
-
408
- ---
409
-
410
- ## 💼 Real-World Use Cases
411
-
412
- StepFun demonstrated the model's capabilities across several complex, real-world projects:
413
-
414
- - **ESP32 Development Board Modifications:** Executed development tasks for over 3 hours, demonstrating hardware programming capabilities.
415
- - **Front-End Design with 3D Asset Generation:** Full-stack development workflows including visual design.
416
- - **Full-Process Financial Research:** End-to-end investment research workflows, from data gathering to report generation.
417
- - **Software Engineering:** Comprehensive coding tasks beyond traditional code generation, including front-end, visual development, and programmable hardware scenarios.
418
- - **Autonomous Research Assistant:** Capable of reading papers, running experiments, and summarizing findings.
419
- - **Customer Support Automation:** Handles multi-turn conversations with tool calls to internal systems.
420
-
421
  ---
422
 
423
  ## ⚡ Quickstart
@@ -572,20 +396,126 @@ response = client.chat.completions.create(
572
  print(response.choices[0].message.content)
573
  ```
574
 
575
- <div style="border-left: 6px solid #722ed1; padding: 16px; border-radius: 8px; margin: 20px 0;">
576
- <strong>📦 Recommended Deployment Configurations</strong><br>
577
- • <strong>BF16:</strong> 8× H100 80GB (tensor parallel)<br>
578
- • <strong>FP8:</strong> 4× H100 80GB (coming soon)<br>
579
- • <strong>Context length:</strong> Up to 1M tokens<br>
580
- • <strong>Reasoning parser:</strong> Use <code>stepfun</code> for vLLM/SGLang
581
- </div>
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
582
 
583
  ---
584
 
585
- ## 📈 Evaluation
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
586
 
587
- Step-5-Preview was evaluated on a comprehensive suite of public and internal benchmarks.
588
- All evaluations used the model's `high` reasoning effort setting unless otherwise noted.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
589
 
590
  | Benchmark | Score | Notes |
591
  |:---|:---|:---|
@@ -603,60 +533,79 @@ All evaluations used the model's `high` reasoning effort setting unless otherwis
603
  | **SWE-Atlas-Test-writing** | 50.8% | 90 tasks on writing production-grade unit, integration, and acceptance tests for real repositories |
604
  | **StepCode-Bench-Daily** | 64.9% | StepFun internal; 553 repos, 9 task types, 20 domains, 33 languages; daily-difficulty slice |
605
  | **StepCode-Bench-General** | 65.0% | StepFun internal; same corpus as StepCodeBench; general-difficulty slice |
606
- | **Agents' Last Exam (ALE-CLI)** | 29.5% | Linux-only CLI subset of ALE; 40 industry subfields, tasks taking hours to weeks; best agent pass rate ~25.2% |
607
  | **GDPval-AA v2** | 1571 | Artificial Analysis, Sep 19, 2026 |
608
- | **AA-Briefcase** | 1417 | Elo score; agentic knowledge work across data science, product, banking, heavy industry; private held-out test set |
609
  | **Toolathlon-Verified** | 74.1% | 108 expert-authored tasks; multi-app workflows averaging ~20 turns; strictly verifiable via scripts |
610
  | **MCP-Atlas** | 85.6% | 1,000 tasks (500 public + 500 private); 36 real MCP servers, 220 tools; 3–6 tool calls/task |
611
- | **PresentBench** | 76.8% | 238 slide-generation instances; avg 54.1 rubric checklist items per instance; evaluated on grounded, fine-grained slide quality |
612
- | **JobBench** | 59.0% | 130 tasks across 35 occupations; expert-prioritized delegation workflows; avg 35.6 binary rubric criteria per task |
613
- | **Apex-Agents** | 37.8% | Long-horizon tasks in investment banking, consulting, and corporate law; frontier models complete <25% of tasks |
614
- | **DRACO** | 83.3% | Cross-domain deep research benchmark; accuracy, completeness, objectivity, citation quality; LLM-as-judge with binary rubric verdicts |
615
- | **BrowseComp** | 88.7% | 1,266 hard-to-find web information retrieval questions; requires persistent browsing across many sites |
616
  | **HLE w/ tools** | 59.4% | +12.9 pts over no-tools HLE |
617
  | **FrontierFinance** | 66.4% | +11.4 pts over GPT-6 Astra |
618
  | **FinStepBench-LiveSearch** | 74.5% | Tied with GPT-6 Astra |
619
  | **FinStepBench-CorporateValuation** | 60.6% | Tied with Kimi K3 |
620
- | **FinStepBench-FinanceDR** | 55.8% | StepFun internal; covers full deep-research finance workflows — source traceability, reproducible assumptions, auditable reports |
621
- | **OfficeQA Pro** | 60.3% | Databricks; 90 questions over large enterprise financial document collections; grounded reasoning on spreadsheets, charts, and business artifacts |
622
- | **Spreadsheet v2** | 29.4% | SpreadsheetBench 2; end-to-end business spreadsheet workflows with multi-statement financial modeling; best model ≈34.89% |
623
  | **GDP.pdf** | 14.8% | Known weak spot |
624
- | **GPQA Diamond** | 93.5% | 198 PhD-level science MCQs (biology, physics, chemistry); human expert avg 81% |
625
  | **HLE** | 46.5% | 59.4% with tools |
626
  | **AA-LCR v1.1** | 88.3% | Tied with Kimi K3 (88.7%) |
627
- | **CritPt** | 20.9% | 71 unpublished research-level physics challenges; simulates full-scale junior-PhD research projects |
628
- | **MMMU-Pro** | 76.0% | Robust multimodal benchmark; filtered text-only questions, augmented candidates, vision-only settings; significantly harder than MMMU |
629
  | **Output Speed** | 99.8 tokens/sec | GLM-5.3: 72.1 tokens/sec |
630
  | **Time to First Token** | 2.96s | GLM-5.3: 2.99s; Claude Opus 5: 56.84s (max effort) |
631
 
632
- ---
 
 
 
633
 
634
- ## ⚠️ Limitations
 
 
 
 
 
 
 
 
 
 
 
635
 
636
- - **Knowledge Cutoff:** The model's knowledge is current up to mid-2026. It may not be aware of events after that date.
637
- - **Hallucination:** Like all large language models, Step-5-Preview can generate plausible but incorrect information, especially in domains with sparse training data.
638
- - **Long Context Degradation:** While the model supports 1M tokens, performance may degrade for extremely long contexts beyond 500K tokens in certain tasks.
639
- - **Tool Use Reliability:** Tool calling is highly capable but not infallible. Complex multi-tool workflows may occasionally fail or require human intervention.
 
 
 
640
  - **Multimodal Limitations:** Video understanding is limited to clips under 5 minutes and 128 MB. Extremely high-resolution images may be downscaled.
641
- - **Language Coverage:** While multilingual, the model is primarily optimized for English and Chinese. Performance in other languages may vary.
642
 
643
- ---
644
 
645
- ## ⚖️ Ethical Considerations
 
646
 
647
- StepFun is committed to the responsible development and deployment of AI. We have taken the following measures:
648
 
649
- - **Safety Alignment:** The model was fine-tuned with RLHF to refuse harmful requests and promote helpful, honest, and harmless behavior.
650
- - **Bias Mitigation:** Training data was filtered to reduce harmful stereotypes and biases. However, residual biases may exist.
651
- - **Transparency:** We provide detailed model cards and benchmark results to enable informed use.
652
  - **License Restrictions:** The StepFun Community License prohibits certain high-risk uses, including autonomous weapons, surveillance, and malicious cyber activities.
653
- - **Content Provenance:** We encourage users to clearly label AI-generated content and to use the model ethically.
654
 
655
- We urge all users to consider the ethical implications of their applications and to implement appropriate safeguards.
656
 
657
- ---
658
 
659
- ## 🖥️ Hardware Requirements
 
660
 
661
  | Precision | Minimum GPU Memory | Recommended GPU Configuration |
662
  |:---|:---|:---|
@@ -664,12 +613,12 @@ We urge all users to consider the ethical implications of their applications and
664
  | **FP8** | 600 GB | 4× H100 80GB (tensor parallel) |
665
  | **INT4** | 300 GB | 4× A100 80GB (tensor parallel) |
666
 
667
- For inference with 1M context, additional memory is required for KV cache. We recommend using paged attention and
668
- offloading techniques available in vLLM and SGLang.
669
 
670
- ---
671
 
672
- ## ⚡ Performance Metrics
 
673
 
674
  | Metric | Value |
675
  |:---|:---|
@@ -682,11 +631,10 @@ offloading techniques available in vLLM and SGLang.
682
 
683
  *Measured on 8× H100 80GB with vLLM, batch size 1, BF16.*
684
 
685
- ---
686
-
687
- ## 📚 Citation
688
 
689
- If you use Step-5-Preview in your research, please cite:
 
690
 
691
  ```bibtex
692
  @misc{stepfun2026step5preview,
@@ -698,29 +646,7 @@ If you use Step-5-Preview in your research, please cite:
698
  }
699
  ```
700
 
701
- ---
702
-
703
- ## 📜 License
704
-
705
- Step-5-Preview is released under the **StepFun Community License**.
706
- See the [LICENSE](https://huggingface.co/SHSLab/Step-5-Preview-BF16/blob/main/LICENSE) file for full terms.
707
-
708
- <div style="border-left: 6px solid #faad14; padding: 16px; border-radius: 8px; margin: 20px 0;">
709
- <strong>⚠️ Usage Restrictions</strong><br>
710
- • Commercial use is permitted under the StepFun Community License.<br>
711
- • Redistribution must include the license and attribution.<br>
712
- • See LICENSE for full details.
713
- </div>
714
-
715
- ---
716
-
717
- ## 📬 Contact
718
-
719
- - **Hugging Face:** [SHSLab](https://huggingface.co/SHSLab)
720
- - **GitHub:** [github.com/stepfun-ai](https://github.com/stepfun-ai)
721
- - **Discord:** [Join our Discord](https://discord.gg/stepfun)
722
- - **Email:** [opensource@stepfun.com](mailto:opensource@stepfun.com)
723
- - **Website:** [stepfun.com](https://stepfun.com)
724
 
725
  ---
726
 
 
1
+ ---
2
+ language:
3
+ - en
4
+ - zh
5
+ - multilingual
6
+ license: other
7
+ license_name: stepfun-community-license
8
+ license_link: https://huggingface.co/SHSLab/Step-5-Preview-BF16/blob/main/LICENSE
9
+ library_name: transformers
10
+ pipeline_tag: text-generation
11
+ tags:
12
+ - stepfun
13
+ - step-5
14
+ - moe
15
+ - mixture-of-experts
16
+ - agentic
17
+ - coding
18
+ - software-engineering
19
+ - long-context
20
+ - 1m-context
21
+ - multimodal
22
+ - text-generation
23
+ - image
24
+ - video
25
+ - sparse-attention
26
+ - gqa
27
+ - financial-analysis
28
+ - deep-research
29
+ - tool-calling
30
+ - parallel-tool-calling
31
+ - json-schema
32
+ ---
33
+
34
  # Step-5-Preview
35
 
36
  <div align="center">
 
48
 
49
  </div>
50
 
51
+ > **🔥 Step-5-Preview is now available.**
52
+ >
53
+ > Step-5-Preview is a 600B-parameter sparse Mixture-of-Experts model with 27B active parameters, a 1M-token context window, and native support for text, image, and video inputs.
54
+ >
55
+ > Weights are available on Hugging Face (`SHSLab/Step-5-Preview-BF16`). Available via Step API, or self-hosted with vLLM / SGLang.
 
 
 
 
56
 
57
  ---
58
 
 
60
 
61
  - [Introduction](#-introduction)
62
  - [Key Features](#-key-features)
 
63
  - [Model Specifications](#-model-specifications)
 
64
  - [Benchmark Results](#-benchmark-results)
 
 
65
  - [Quickstart](#-quickstart)
66
  - [Deployment](#-deployment)
 
 
 
 
 
 
67
  - [License](#-license)
68
  - [Contact](#-contact)
69
+ - [More details](#-more-details) *(architecture, training, agentic demos, evaluation, limitations)*
70
 
71
  ---
72
 
73
  ## 🚀 Introduction
74
 
75
+ **Step-5-Preview** is StepFun's flagship foundation model, designed for **real-world agentic tasks**. It targets professional domains such as **AI coding, software engineering, professional knowledge work, and financial analysis**.
 
76
 
77
+ StepFun's core philosophy for Step 5 is the **"Pareto Frontier"** — balancing intelligence against cost. While earlier scaling efforts traded more compute for stronger intelligence, the next phase focuses on improving the **efficiency of converting compute into intelligence**.
 
 
78
 
79
+ > **Why Step 5 Preview?**
80
+ >
81
+ > - 600B total parameters, 27B active per token — near-frontier performance at a fraction of the compute.
82
+ > - 1M-token context window without proportional cost increases.
83
+ > - Competitive benchmark scores against models with 3–5× more parameters.
84
+ > - Built for agents: long-horizon reasoning, tool use, and autonomous execution.
 
85
 
86
+ Step-5-Preview skips the entire Step 4.x line, going directly from Step-3.7-Flash to Step 5.
 
87
 
88
  ---
89
 
 
100
  - **Parallel Tool Calling:** Natively supported for agentic workflows.
101
  - **Strict JSON Schema Output:** Reliable integration into structured systems.
102
  - **OpenAI-Compatible API:** Available via Step API and third-party gateways.
103
+ - **Open Weights:** BF16 checkpoint available under `SHSLab/Step-5-Preview-BF16`.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
104
 
105
  ---
106
 
 
122
  | **Reasoning Effort** | `low` / `medium` / `high` (`xhigh`) |
123
  | **Tool Calling** | Parallel, strict JSON schema |
124
  | **Intelligence Index** | 44 (Artificial Analysis v4.3.2) |
125
+ | **Open Weights** | BF16 checkpoint available |
126
  | **API Availability** | Immediate (OpenAI-compatible) |
127
  | **License** | StepFun Community License |
128
 
129
  ---
130
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
131
  ## 📊 Benchmark Results
132
 
133
  <div align="center">
 
138
 
139
  **Overall Score: 44** (Intelligence Index v4.3.2, recalibrated September 7, 2026)
140
 
141
+ The index covers 10 evaluations including AA-Briefcase, GDPval-AA v2, Terminal-Bench 4.0, SciCode, and Humanity's Last Exam.
 
 
142
 
143
+ > **How to read the tables below**
144
+ >
145
+ > - **Bold** marks the best score in the row across all six models.
146
+ > - **🥇** flags the leader for that benchmark.
147
+ > - **—** indicates the model was not evaluated or did not report a score.
148
+ > - Step-5-Preview results use the `high` reasoning-effort setting unless noted.
149
+ >
150
+ > **Comparison set:** Step-5-Preview · GPT-6 Astra (Max) · Fable 5.1 · Claude Opus 5 (Max) · Kimi K3 (Max) · GLM-5.3 (Max)
151
+ >
152
+ > **Availability:** Open-weight — Step-5-Preview, Kimi K3, Qwen3.8 Max, GLM-5.3. Closed-source — GPT-6 Astra, Claude Opus 5. Fable 5.1 — not publicly stated.
 
 
 
 
 
 
 
153
 
154
  ---
155
 
 
162
  | **AA-LCR v1.1** | 88.3% | 80.7% | 85.3% | 79.3% | **88.7%** 🥇 | 79.7% |
163
  | **CritPt** | 20.9% | **31.7%** 🥇 | 29.7% | 29.1% | 23.4% | 19.1% |
164
 
 
 
 
 
 
 
 
 
165
  ---
166
 
167
  ### 💻 Coding & Software Engineering
 
183
  | **StepCode-Bench-Daily** | 64.9% | — | — | **77.6%** 🥇 | 57.7% | 69.1% |
184
  | **StepCode-Bench-General** | 65.0% | 64.3% | — | **68.3%** 🥇 | 65.2% | 62.0% |
185
 
 
 
 
 
 
 
 
 
 
186
  ---
187
 
188
  ### 🤖 Agents, Tool Use & Automation
 
203
  | **BrowseComp** | 88.7% | **91.5%** 🥇 | — | 90.2% | 91.2% | — |
204
  | **HLE w/ tools** | 59.4% | 57.2% | **65.0%** 🥇 | 63.6% | 56.0% | 62.5% |
205
 
 
 
 
 
 
 
 
 
206
  ---
207
 
208
  ### 💰 Finance & Professional Work
 
214
  | **FinStepBench-FinanceDR** | 55.8% | 45.0% | — | **59.1%** 🥇 | 48.9% | 53.3% |
215
  | **FrontierFinance** | 66.4% | 55.0% | — | **69.7%** 🥇 | 62.6% | 64.1% |
216
  | **OfficeQA Pro** | 60.3% | **67.7%** 🥇 | — | 64.7% | 62.6% | 59.1% |
217
+ | **SpeadSheet v2** | 29.4% | 31.4% | — | **32.8%** 🥇 | 31.9% | 30.5% |
218
  | **GDP.pdf** | 14.8% | **31.0%** 🥇 | 26.2% | 21.6% | 22.0% | 11.2% |
219
 
 
 
 
 
 
 
 
 
 
220
  ---
221
 
222
  ### 👁️ Multimodal
 
226
  | **MMMU-Pro** | 76.0% | **87.0%** 🥇 | — | 85.0% | 81.0% | — |
227
  | **GDP.pdf** | 14.8% | **31.0%** 🥇 | 26.2% | 21.6% | 22.0% | 11.2% |
228
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
229
  <details>
230
+ <summary><strong>📝 Benchmark methodology notes</strong></summary>
231
 
232
  - **DeepSWE v1.1** was evaluated using the SWE-agent harness with `temperature=1.0` and `top_p=0.95`.
233
  - **StepCodeBench** achieved **49.0% avg@4**.
234
  - **GDPval-AA v2** results are from Artificial Analysis as of September 19, 2026.
235
  - **AA-LCR v1.1**: Step-5-Preview scored **88.3%**, statistically tied with Kimi K3 (88.7%).
236
+ - **Terminal-Bench 4**: Step-5-Preview **33.3%** vs. Kimi K3 ~**12.6%** (2.6×) and DeepSeek V4.1 Flash **26.8%** (1.24×).
237
+ - **SciCode**: Step-5-Preview **58.9%**, above GPT-6 Astra (56.5%) and Claude Opus 5 (56.4%).
238
  - **Multimodal**: MMMU-Pro and GDP.pdf were run with the unified multimodal encoder at default resolution.
239
  - **Output Speed**: 99.8 tokens/sec (GLM-5.3: 72.1 tokens/sec).
240
  - **Time to First Token**: 2.96 seconds (GLM-5.3: 2.99s; Claude Opus 5: 56.84s at max effort).
241
  - Missing entries (**—**) reflect benchmarks that were not publicly reported for that model at the time of writing.
 
242
 
243
  </details>
244
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
245
  ---
246
 
247
  ## ⚡ Quickstart
 
396
  print(response.choices[0].message.content)
397
  ```
398
 
399
+ > **Recommended deployment configurations**
400
+ >
401
+ > - **BF16:** 8× H100 80GB (tensor parallel)
402
+ > - **FP8:** 4× H100 80GB (coming soon)
403
+ > - **Context length:** Up to 1M tokens
404
+ > - **Reasoning parser:** Use `stepfun` for vLLM / SGLang
405
+
406
+ ---
407
+
408
+ ## 📜 License
409
+
410
+ Step-5-Preview is released under the **StepFun Community License**. See the [LICENSE](https://huggingface.co/SHSLab/Step-5-Preview-BF16/blob/main/LICENSE) file for full terms.
411
+
412
+ > **Usage restrictions**
413
+ >
414
+ > - Commercial use is permitted under the StepFun Community License.
415
+ > - Redistribution must include the license and attribution.
416
+ > - See LICENSE for full details.
417
+
418
+ ---
419
+
420
+ ## 📬 Contact
421
+
422
+ - **Hugging Face:** [SHSLab](https://huggingface.co/SHSLab)
423
+ - **GitHub:** [github.com/stepfun-ai](https://github.com/stepfun-ai)
424
+ - **Discord:** [Join our Discord](https://discord.gg/stepfun)
425
+ - **Email:** [opensource@stepfun.com](mailto:opensource@stepfun.com)
426
+ - **Website:** [stepfun.com](https://stepfun.com)
427
 
428
  ---
429
 
430
+ ## 📚 More details
431
+
432
+ <details>
433
+ <summary><strong>🏗️ Model Architecture</strong></summary>
434
+
435
+ <div align="center">
436
+ <img src="./Step-5/architecture.png" alt="Step 5 Architecture" width="85%">
437
+ </div>
438
+
439
+ ### 92-Layer "Narrow but Deep" Design
440
+
441
+ Step-5-Preview uses a **92-layer Transformer** with a narrow-deep configuration. This design creates longer information propagation paths for implicit multi-hop reasoning during long prefill operations.
442
+
443
+ ### Sparse Grouped-Query Attention (GQA) with Block-Wise Token Merging
444
+
445
+ To handle the 1M-token context window efficiently, Step-5-Preview uses **Sparse GQA with block-wise token merging**. This mechanism uses sparse indexing to select only historical information relevant to the current task, reducing the number of tokens that enter attention computation. This cuts indexer and top-k selection costs to approximately one-eighth of a denser baseline.
446
+
447
+ ### Multimodal Encoder
448
+
449
+ A unified multimodal encoder processes text, images, and video frames into a shared latent space. Video is sampled at adaptive frame rates and encoded with temporal attention, allowing the model to understand motion and long-range dependencies in screen recordings, demonstrations, and real-world footage.
450
+
451
+ </details>
452
+
453
+ <details>
454
+ <summary><strong>📚 Training Data</strong></summary>
455
+
456
+ Step-5-Preview was trained on a carefully curated corpus spanning:
457
+
458
+ - **Code repositories** from multiple languages (Python, C++, Rust, JavaScript, Go, etc.)
459
+ - **Technical documentation**, API references, and software engineering forums
460
+ - **Scientific papers** in computer science, mathematics, physics, and finance
461
+ - **Financial reports**, earnings calls, and market analyses
462
+ - **Multimodal data** including screenshots, UI mockups, video tutorials, and screen recordings
463
+ - **Agentic trajectories** from simulated and real tool-use environments
464
+
465
+ The data mixture was optimized for long-horizon reasoning and tool use, with a strong emphasis on real-world professional tasks. All data was filtered for quality, safety, and license compliance. Training combined next-token prediction with reinforcement learning from human feedback (RLHF) focused on agentic objectives.
466
+
467
+ </details>
468
+
469
+ <details>
470
+ <summary><strong>🤖 Agentic Capabilities</strong></summary>
471
+
472
+ <div align="center">
473
+ <img src="./Step-5/agentic_workflow.png" alt="Agentic Workflow" width="90%">
474
+ </div>
475
+
476
+ ### 24-Hour Autonomous GPU Kernel Optimization
477
+
478
+ Step-5-Preview was tasked with autonomously optimizing an H100 GPU kernel for up to 24 consecutive hours. The model:
479
+
480
+ - Independently modified code
481
+ - Ran tests and compared results
482
+ - Iterated based on performance outcomes
483
+ - **Reached 508 TFLOPS after approximately 22 hours**
484
+
485
+ For comparison, Claude Opus 5 achieved 493 TFLOPS in the same experiment.
486
+
487
+ ### Automated Post-Training Experiments
488
+
489
+ In another 24-hour experiment, Step-5-Preview autonomously improved the accuracy of Qwen3-30B-A3B on AIME24 from 53.3% to 60% through automated post-training experiments.
490
+
491
+ ### Long-Horizon Agent Workflows
492
 
493
+ Optimized for workflows that require:
494
+
495
+ - Searching and information retrieval
496
+ - Running code and processing tool returns
497
+ - Multi-turn tool calls with sustained execution
498
+ - Iterative refinement based on intermediate results
499
+ - Self-correction and error recovery over thousands of steps
500
+
501
+ </details>
502
+
503
+ <details>
504
+ <summary><strong>💼 Real-World Use Cases</strong></summary>
505
+
506
+ - **ESP32 Development Board Modifications:** Executed development tasks for over 3 hours.
507
+ - **Front-End Design with 3D Asset Generation:** Full-stack development workflows including visual design.
508
+ - **Full-Process Financial Research:** End-to-end investment research, from data gathering to report generation.
509
+ - **Software Engineering:** Comprehensive coding tasks including front-end, visual development, and programmable hardware.
510
+ - **Autonomous Research Assistant:** Reading papers, running experiments, and summarizing findings.
511
+ - **Customer Support Automation:** Multi-turn conversations with tool calls to internal systems.
512
+
513
+ </details>
514
+
515
+ <details>
516
+ <summary><strong>📈 Full Evaluation Table</strong></summary>
517
+
518
+ All evaluations used the `high` reasoning effort setting unless otherwise noted.
519
 
520
  | Benchmark | Score | Notes |
521
  |:---|:---|:---|
 
533
  | **SWE-Atlas-Test-writing** | 50.8% | 90 tasks on writing production-grade unit, integration, and acceptance tests for real repositories |
534
  | **StepCode-Bench-Daily** | 64.9% | StepFun internal; 553 repos, 9 task types, 20 domains, 33 languages; daily-difficulty slice |
535
  | **StepCode-Bench-General** | 65.0% | StepFun internal; same corpus as StepCodeBench; general-difficulty slice |
536
+ | **Agents' Last Exam (ALE-CLI)** | 29.5% | Linux-only CLI subset of ALE; 40 industry subfields; best agent pass rate ~25.2% |
537
  | **GDPval-AA v2** | 1571 | Artificial Analysis, Sep 19, 2026 |
538
+ | **AA-Briefcase** | 1417 | Elo; agentic knowledge work across data science, product, banking, heavy industry; private held-out test set |
539
  | **Toolathlon-Verified** | 74.1% | 108 expert-authored tasks; multi-app workflows averaging ~20 turns; strictly verifiable via scripts |
540
  | **MCP-Atlas** | 85.6% | 1,000 tasks (500 public + 500 private); 36 real MCP servers, 220 tools; 3–6 tool calls/task |
541
+ | **PresentBench** | 76.8% | 238 slide-generation instances; avg 54.1 rubric checklist items per instance |
542
+ | **JobBench** | 59.0% | 130 tasks across 35 occupations; avg 35.6 binary rubric criteria per task |
543
+ | **Apex-Agents** | 37.8% | Long-horizon tasks in investment banking, consulting, corporate law |
544
+ | **DRACO** | 83.3% | Cross-domain deep research; accuracy, completeness, objectivity, citation quality |
545
+ | **BrowseComp** | 88.7% | 1,266 hard-to-find web information retrieval questions |
546
  | **HLE w/ tools** | 59.4% | +12.9 pts over no-tools HLE |
547
  | **FrontierFinance** | 66.4% | +11.4 pts over GPT-6 Astra |
548
  | **FinStepBench-LiveSearch** | 74.5% | Tied with GPT-6 Astra |
549
  | **FinStepBench-CorporateValuation** | 60.6% | Tied with Kimi K3 |
550
+ | **FinStepBench-FinanceDR** | 55.8% | StepFun internal; full deep-research finance workflows |
551
+ | **OfficeQA Pro** | 60.3% | Databricks; 90 questions over large enterprise financial document collections |
552
+ | **SpeadSheet v2** | 29.4% | SpreadsheetBench 2; end-to-end business spreadsheet workflows; best model ≈34.89% |
553
  | **GDP.pdf** | 14.8% | Known weak spot |
554
+ | **GPQA Diamond** | 93.5% | 198 PhD-level science MCQs; human expert avg 81% |
555
  | **HLE** | 46.5% | 59.4% with tools |
556
  | **AA-LCR v1.1** | 88.3% | Tied with Kimi K3 (88.7%) |
557
+ | **CritPt** | 20.9% | 71 unpublished research-level physics challenges |
558
+ | **MMMU-Pro** | 76.0% | Robust multimodal benchmark; significantly harder than MMMU |
559
  | **Output Speed** | 99.8 tokens/sec | GLM-5.3: 72.1 tokens/sec |
560
  | **Time to First Token** | 2.96s | GLM-5.3: 2.99s; Claude Opus 5: 56.84s (max effort) |
561
 
562
+ </details>
563
+
564
+ <details>
565
+ <summary><strong>🏅 Category summary</strong></summary>
566
 
567
+ | Domain | Standing | Highlight |
568
+ |:---|:---|:---|
569
+ | **Long-Context Reasoning** | Near-tied for #1 | AA-LCR v1.1 **88.3%** (+7.6 over GPT-6 Astra, +9.0 over Opus 5) |
570
+ | **Coding / SWE** | #1 open-weight | DeepSWE v1.1 **67.7%**, StepCodeBench **49.0%**, ProgramBench **80.5%** |
571
+ | **Security / Cyber** | #1 in comparison set | CyberGym **84.7%** |
572
+ | **Finance** | #2 overall, #1 open-weight | FrontierFinance **66.4%**, FinanceDR **55.8%** |
573
+ | **Terminal Agents** | #2 open-weight | Terminal-Bench 4 **33.3%** (2.6× Kimi K3) |
574
+ | **Tool Use** | Top-3 open-weight | MCP-Atlas **85.6%**, HLE w/ tools **59.4%** |
575
+ | **General Knowledge** | Top tier | GPQA Diamond **93.5%**, HLE **46.5%** |
576
+ | **Multimodal** | Trailing | MMMU-Pro **76.0%** — targeted for improvement |
577
+
578
+ </details>
579
 
580
+ <details>
581
+ <summary><strong>⚠️ Limitations</strong></summary>
582
+
583
+ - **Knowledge Cutoff:** Knowledge is current up to mid-2026. May not be aware of later events.
584
+ - **Hallucination:** Can generate plausible but incorrect information, especially in domains with sparse training data.
585
+ - **Long Context Degradation:** Performance may degrade for extremely long contexts beyond 500K tokens in certain tasks.
586
+ - **Tool Use Reliability:** Tool calling is capable but not infallible. Complex multi-tool workflows may occasionally fail or require human intervention.
587
  - **Multimodal Limitations:** Video understanding is limited to clips under 5 minutes and 128 MB. Extremely high-resolution images may be downscaled.
588
+ - **Language Coverage:** Primarily optimized for English and Chinese. Performance in other languages may vary.
589
 
590
+ </details>
591
 
592
+ <details>
593
+ <summary><strong>⚖️ Ethical Considerations</strong></summary>
594
 
595
+ StepFun is committed to the responsible development and deployment of AI. Measures taken:
596
 
597
+ - **Safety Alignment:** Fine-tuned with RLHF to refuse harmful requests and promote helpful, honest, and harmless behavior.
598
+ - **Bias Mitigation:** Training data was filtered to reduce harmful stereotypes and biases. Residual biases may exist.
599
+ - **Transparency:** Detailed model cards and benchmark results are provided to enable informed use.
600
  - **License Restrictions:** The StepFun Community License prohibits certain high-risk uses, including autonomous weapons, surveillance, and malicious cyber activities.
601
+ - **Content Provenance:** Users are encouraged to clearly label AI-generated content.
602
 
603
+ All users should consider the ethical implications of their applications and implement appropriate safeguards.
604
 
605
+ </details>
606
 
607
+ <details>
608
+ <summary><strong>🖥️ Hardware Requirements</strong></summary>
609
 
610
  | Precision | Minimum GPU Memory | Recommended GPU Configuration |
611
  |:---|:---|:---|
 
613
  | **FP8** | 600 GB | 4× H100 80GB (tensor parallel) |
614
  | **INT4** | 300 GB | 4× A100 80GB (tensor parallel) |
615
 
616
+ For inference with 1M context, additional memory is required for KV cache. Paged attention and offloading techniques available in vLLM and SGLang are recommended.
 
617
 
618
+ </details>
619
 
620
+ <details>
621
+ <summary><strong>⚡ Performance Metrics</strong></summary>
622
 
623
  | Metric | Value |
624
  |:---|:---|
 
631
 
632
  *Measured on 8× H100 80GB with vLLM, batch size 1, BF16.*
633
 
634
+ </details>
 
 
635
 
636
+ <details>
637
+ <summary><strong>📚 Citation</strong></summary>
638
 
639
  ```bibtex
640
  @misc{stepfun2026step5preview,
 
646
  }
647
  ```
648
 
649
+ </details>
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
650
 
651
  ---
652