Image-Text-to-Text
Transformers
Safetensors
English
Chinese
multilingual
step3p5v
text-generation
stepfun
step-5
Mixture of Experts
mixture-of-experts
agentic
coding
software-engineering
long-context
1m-context
multimodal
image
video
sparse-attention
gqa
financial-analysis
deep-research
tool-calling
parallel-tool-calling
json-schema
conversational
custom_code
Instructions to use SHSLab/Step-5-Preview-BF16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use SHSLab/Step-5-Preview-BF16 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="SHSLab/Step-5-Preview-BF16", trust_remote_code=True) messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("SHSLab/Step-5-Preview-BF16", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use SHSLab/Step-5-Preview-BF16 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "SHSLab/Step-5-Preview-BF16" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SHSLab/Step-5-Preview-BF16", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/SHSLab/Step-5-Preview-BF16
- SGLang
How to use SHSLab/Step-5-Preview-BF16 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "SHSLab/Step-5-Preview-BF16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SHSLab/Step-5-Preview-BF16", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "SHSLab/Step-5-Preview-BF16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SHSLab/Step-5-Preview-BF16", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use SHSLab/Step-5-Preview-BF16 with Docker Model Runner:
docker model run hf.co/SHSLab/Step-5-Preview-BF16
Update README.md
Browse files
README.md
CHANGED
|
@@ -1,3 +1,36 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
# Step-5-Preview
|
| 2 |
|
| 3 |
<div align="center">
|
|
@@ -15,15 +48,11 @@
|
|
| 15 |
|
| 16 |
</div>
|
| 17 |
|
| 18 |
-
|
| 19 |
-
|
| 20 |
-
|
| 21 |
-
|
| 22 |
-
|
| 23 |
-
<br><br>
|
| 24 |
-
<strong>Weights are available now</strong> on Hugging Face (<code>SHSLab/Step-5-Preview-BF16</code>).
|
| 25 |
-
Try it via our API, or deploy locally with vLLM / SGLang.
|
| 26 |
-
</div>
|
| 27 |
|
| 28 |
---
|
| 29 |
|
|
@@ -31,44 +60,30 @@
|
|
| 31 |
|
| 32 |
- [Introduction](#-introduction)
|
| 33 |
- [Key Features](#-key-features)
|
| 34 |
-
- [Model Architecture](#-model-architecture)
|
| 35 |
- [Model Specifications](#-model-specifications)
|
| 36 |
-
- [Training Data](#-training-data)
|
| 37 |
- [Benchmark Results](#-benchmark-results)
|
| 38 |
-
- [Agentic Capabilities](#-agentic-capabilities)
|
| 39 |
-
- [Real-World Use Cases](#-real-world-use-cases)
|
| 40 |
- [Quickstart](#-quickstart)
|
| 41 |
- [Deployment](#-deployment)
|
| 42 |
-
- [Evaluation](#-evaluation)
|
| 43 |
-
- [Limitations](#-limitations)
|
| 44 |
-
- [Ethical Considerations](#-ethical-considerations)
|
| 45 |
-
- [Hardware Requirements](#-hardware-requirements)
|
| 46 |
-
- [Performance Metrics](#-performance-metrics)
|
| 47 |
-
- [Citation](#-citation)
|
| 48 |
- [License](#-license)
|
| 49 |
- [Contact](#-contact)
|
|
|
|
| 50 |
|
| 51 |
---
|
| 52 |
|
| 53 |
## 🚀 Introduction
|
| 54 |
|
| 55 |
-
**Step-5-Preview** is StepFun's flagship foundation model, designed
|
| 56 |
-
It targets professional domains such as **AI coding, software engineering, professional knowledge work, and financial analysis**.
|
| 57 |
|
| 58 |
-
StepFun's core philosophy for Step 5 is the **"Pareto Frontier"** —
|
| 59 |
-
While previous scaling efforts focused on trading more compute for stronger intelligence, the next phase requires improving the
|
| 60 |
-
**efficiency of converting compute into intelligence**.
|
| 61 |
|
| 62 |
-
|
| 63 |
-
|
| 64 |
-
|
| 65 |
-
|
| 66 |
-
|
| 67 |
-
|
| 68 |
-
</div>
|
| 69 |
|
| 70 |
-
Step-5-Preview
|
| 71 |
-
Step-3.7-Flash to Step 5. This decision reflects the magnitude of improvement achieved in this release.
|
| 72 |
|
| 73 |
---
|
| 74 |
|
|
@@ -85,39 +100,7 @@ Step-3.7-Flash to Step 5. This decision reflects the magnitude of improvement ac
|
|
| 85 |
- **Parallel Tool Calling:** Natively supported for agentic workflows.
|
| 86 |
- **Strict JSON Schema Output:** Reliable integration into structured systems.
|
| 87 |
- **OpenAI-Compatible API:** Available via Step API and third-party gateways.
|
| 88 |
-
- **Open Weights:** BF16 checkpoint available
|
| 89 |
-
|
| 90 |
-
---
|
| 91 |
-
|
| 92 |
-
## 🏗️ Model Architecture
|
| 93 |
-
|
| 94 |
-
<div align="center">
|
| 95 |
-
<img src="./Step-5/architecture.png" alt="Step 5 Architecture" width="85%">
|
| 96 |
-
</div>
|
| 97 |
-
|
| 98 |
-
### 92-Layer "Narrow but Deep" Design
|
| 99 |
-
|
| 100 |
-
Step-5-Preview uses a **92-layer Transformer** with a narrow-deep configuration. This design is specifically intended to create
|
| 101 |
-
**longer information propagation paths** for implicit multi-hop reasoning during long prefill operations.
|
| 102 |
-
|
| 103 |
-
### Sparse Grouped-Query Attention (GQA) with Block-Wise Token Merging
|
| 104 |
-
|
| 105 |
-
To handle the 1M-token context window efficiently, Step-5-Preview introduces **Sparse GQA with block-wise token merging**.
|
| 106 |
-
This mechanism uses sparse indexing to select only historical information relevant to the current task, reducing the number of tokens
|
| 107 |
-
that actually enter attention computation. StepFun states this cuts indexer and top-k selection costs to approximately
|
| 108 |
-
**one-eighth** of a denser baseline.
|
| 109 |
-
|
| 110 |
-
<div style="border-left: 6px solid #52c41a; padding: 16px; border-radius: 8px; margin: 20px 0;">
|
| 111 |
-
<strong>⚡ Efficiency-First Scaling</strong><br>
|
| 112 |
-
Step 5 Preview achieves near-frontier performance with <strong>600B total parameters</strong> but only
|
| 113 |
-
<strong>27B active per token</strong>. This is the core of StepFun's efficiency-first philosophy.
|
| 114 |
-
</div>
|
| 115 |
-
|
| 116 |
-
### Multimodal Encoder
|
| 117 |
-
|
| 118 |
-
The model incorporates a unified multimodal encoder that processes text, images, and video frames into a shared latent space.
|
| 119 |
-
Video is sampled at adaptive frame rates and encoded with temporal attention, allowing the model to understand motion and
|
| 120 |
-
long-range dependencies in screen recordings, demonstrations, and real-world footage.
|
| 121 |
|
| 122 |
---
|
| 123 |
|
|
@@ -139,29 +122,12 @@ long-range dependencies in screen recordings, demonstrations, and real-world foo
|
|
| 139 |
| **Reasoning Effort** | `low` / `medium` / `high` (`xhigh`) |
|
| 140 |
| **Tool Calling** | Parallel, strict JSON schema |
|
| 141 |
| **Intelligence Index** | 44 (Artificial Analysis v4.3.2) |
|
| 142 |
-
| **Open Weights** | BF16 checkpoint available
|
| 143 |
| **API Availability** | Immediate (OpenAI-compatible) |
|
| 144 |
| **License** | StepFun Community License |
|
| 145 |
|
| 146 |
---
|
| 147 |
|
| 148 |
-
## 📚 Training Data
|
| 149 |
-
|
| 150 |
-
Step-5-Preview was trained on a massive, carefully curated corpus spanning:
|
| 151 |
-
|
| 152 |
-
- **Code repositories** from multiple languages (Python, C++, Rust, JavaScript, Go, etc.)
|
| 153 |
-
- **Technical documentation**, API references, and software engineering forums
|
| 154 |
-
- **Scientific papers** in computer science, mathematics, physics, and finance
|
| 155 |
-
- **Financial reports**, earnings calls, and market analyses
|
| 156 |
-
- **Multimodal data** including screenshots, UI mockups, video tutorials, and screen recordings
|
| 157 |
-
- **Agentic trajectories** from simulated and real tool-use environments
|
| 158 |
-
|
| 159 |
-
The data mixture was optimized for long-horizon reasoning and tool use, with a strong emphasis on real-world professional tasks.
|
| 160 |
-
All data was filtered for quality, safety, and license compliance. The training process used a combination of next-token prediction
|
| 161 |
-
and reinforcement learning from human feedback (RLHF) with a focus on agentic objectives.
|
| 162 |
-
|
| 163 |
-
---
|
| 164 |
-
|
| 165 |
## 📊 Benchmark Results
|
| 166 |
|
| 167 |
<div align="center">
|
|
@@ -172,27 +138,18 @@ and reinforcement learning from human feedback (RLHF) with a focus on agentic ob
|
|
| 172 |
|
| 173 |
**Overall Score: 44** (Intelligence Index v4.3.2, recalibrated September 7, 2026)
|
| 174 |
|
| 175 |
-
|
| 176 |
-
Kimi K3 Max (approximately 5× larger at 2.8T parameters) and Qwen3.8 Max. The index covers 10 evaluations including
|
| 177 |
-
AA-Briefcase, GDPval-AA v2, Terminal-Bench 4.0, SciCode, and Humanity's Last Exam.
|
| 178 |
|
| 179 |
-
|
| 180 |
-
|
| 181 |
-
|
| 182 |
-
|
| 183 |
-
|
| 184 |
-
|
| 185 |
-
|
| 186 |
-
-
|
| 187 |
-
|
| 188 |
-
|
| 189 |
-
<strong>Comparison set:</strong> Step-5-Preview · GPT-6 Astra (Max) · Fable 5.1 · Claude Opus 5 (Max) · Kimi K3 (Max) · GLM-5.3 (Max)
|
| 190 |
-
|
| 191 |
-
<strong>Open-weight models in this comparison:</strong> Step-5-Preview, Kimi K3 (Max), GLM-5.3 (Max), and Qwen3.8 Max (referenced in the index).
|
| 192 |
-
<strong>Closed-source models:</strong> GPT-6 Astra (Max), Claude Opus 5 (Max).
|
| 193 |
-
<em>Fable 5.1: availability not specified in this document.</em>
|
| 194 |
-
|
| 195 |
-
</div>
|
| 196 |
|
| 197 |
---
|
| 198 |
|
|
@@ -205,14 +162,6 @@ AA-Briefcase, GDPval-AA v2, Terminal-Bench 4.0, SciCode, and Humanity's Last Exa
|
|
| 205 |
| **AA-LCR v1.1** | 88.3% | 80.7% | 85.3% | 79.3% | **88.7%** 🥇 | 79.7% |
|
| 206 |
| **CritPt** | 20.9% | **31.7%** 🥇 | 29.7% | 29.1% | 23.4% | 19.1% |
|
| 207 |
|
| 208 |
-
<div style="border-left: 6px solid #2f54eb; padding: 16px; border-radius: 8px; margin: 20px 0;">
|
| 209 |
-
<strong>📌 Read-out</strong><br>
|
| 210 |
-
Step-5-Preview sits in the <strong>top tier of open-weight reasoning</strong>: GPQA Diamond <strong>93.5%</strong> matches Kimi K3
|
| 211 |
-
exactly and lands within 0.5 points of Fable 5.1. On <strong>AA-LCR v1.1</strong> — long-context reasoning — it scores
|
| 212 |
-
<strong>88.3%</strong>, effectively tied with Kimi K3 (88.7%) and <strong>+7.6 points ahead of GPT-6 Astra</strong> and
|
| 213 |
-
<strong>+9.0 points ahead of Claude Opus 5</strong>, despite those models being far larger.
|
| 214 |
-
</div>
|
| 215 |
-
|
| 216 |
---
|
| 217 |
|
| 218 |
### 💻 Coding & Software Engineering
|
|
@@ -234,15 +183,6 @@ AA-Briefcase, GDPval-AA v2, Terminal-Bench 4.0, SciCode, and Humanity's Last Exa
|
|
| 234 |
| **StepCode-Bench-Daily** | 64.9% | — | — | **77.6%** 🥇 | 57.7% | 69.1% |
|
| 235 |
| **StepCode-Bench-General** | 65.0% | 64.3% | — | **68.3%** 🥇 | 65.2% | 62.0% |
|
| 236 |
|
| 237 |
-
<div style="border-left: 6px solid #13c2c2; padding: 16px; border-radius: 8px; margin: 20px 0;">
|
| 238 |
-
<strong>📌 Read-out</strong><br>
|
| 239 |
-
Among <strong>open-weight models</strong>, Step-5-Preview leads on <strong>DeepSWE v1.1</strong> (67.7% vs. Kimi K3 67.5% and
|
| 240 |
-
GLM-5.3 66.9%), <strong>StepCodeBench</strong> (49.0% vs. 43.9% and 40.2%), <strong>ProgramBench</strong> (80.5% vs. 77.8% and 72.0%),
|
| 241 |
-
and <strong>SWE-Atlas-QnA</strong> (63.6%). It also posts the best <strong>CyberGym</strong> score in the comparison set at
|
| 242 |
-
<strong>84.7%</strong>. On <strong>Terminal-Bench 4</strong> it reaches <strong>33.3%</strong> —
|
| 243 |
-
<strong>2.6× Kimi K3</strong> (12.6%) — though the larger closed models still hold the absolute lead.
|
| 244 |
-
</div>
|
| 245 |
-
|
| 246 |
---
|
| 247 |
|
| 248 |
### 🤖 Agents, Tool Use & Automation
|
|
@@ -263,14 +203,6 @@ AA-Briefcase, GDPval-AA v2, Terminal-Bench 4.0, SciCode, and Humanity's Last Exa
|
|
| 263 |
| **BrowseComp** | 88.7% | **91.5%** 🥇 | — | 90.2% | 91.2% | — |
|
| 264 |
| **HLE w/ tools** | 59.4% | 57.2% | **65.0%** 🥇 | 63.6% | 56.0% | 62.5% |
|
| 265 |
|
| 266 |
-
<div style="border-left: 6px solid #f5222d; padding: 16px; border-radius: 8px; margin: 20px 0;">
|
| 267 |
-
<strong>📌 Read-out</strong><br>
|
| 268 |
-
Step-5-Preview is a <strong>strong mid-frontier agentic model</strong>. It outperforms GPT-6 Astra on <strong>τ³-Banking</strong>
|
| 269 |
-
(42.5% vs. 41.4%) and <strong>Draco</strong> (83.3% vs. 76.8%), and beats Kimi K3 on <strong>MCP-Atlas</strong> (85.6% vs. 85.3%)
|
| 270 |
-
and <strong>Draco</strong> (83.3% vs. 78.5%). With tools, HLE rises from <strong>46.5% → 59.4%</strong>, a
|
| 271 |
-
<strong>+12.9-point</strong> tool-augmentation gain.
|
| 272 |
-
</div>
|
| 273 |
-
|
| 274 |
---
|
| 275 |
|
| 276 |
### 💰 Finance & Professional Work
|
|
@@ -282,18 +214,9 @@ AA-Briefcase, GDPval-AA v2, Terminal-Bench 4.0, SciCode, and Humanity's Last Exa
|
|
| 282 |
| **FinStepBench-FinanceDR** | 55.8% | 45.0% | — | **59.1%** 🥇 | 48.9% | 53.3% |
|
| 283 |
| **FrontierFinance** | 66.4% | 55.0% | — | **69.7%** 🥇 | 62.6% | 64.1% |
|
| 284 |
| **OfficeQA Pro** | 60.3% | **67.7%** 🥇 | — | 64.7% | 62.6% | 59.1% |
|
| 285 |
-
| **
|
| 286 |
| **GDP.pdf** | 14.8% | **31.0%** 🥇 | 26.2% | 21.6% | 22.0% | 11.2% |
|
| 287 |
|
| 288 |
-
<div style="border-left: 6px solid #a0d911; padding: 16px; border-radius: 8px; margin: 20px 0;">
|
| 289 |
-
<strong>📌 Read-out</strong><br>
|
| 290 |
-
Finance is a Step-5-Preview <strong>strength</strong>. On <strong>FrontierFinance</strong> it scores <strong>66.4%</strong>,
|
| 291 |
-
beating GPT-6 Astra by <strong>+11.4 points</strong>, Kimi K3 by <strong>+3.8</strong>, and GLM-5.3 by <strong>+2.3</strong> —
|
| 292 |
-
trailing only Claude Opus 5. On <strong>FinStepBench-FinanceDR</strong> it reaches <strong>55.8%</strong>, ahead of
|
| 293 |
-
GPT-6 Astra (+10.8), Kimi K3 (+6.9), and GLM-5.3 (+2.5). <strong>FinStepBench-LiveSearch</strong> lands at
|
| 294 |
-
<strong>74.5%</strong>, tied with GPT-6 Astra.
|
| 295 |
-
</div>
|
| 296 |
-
|
| 297 |
---
|
| 298 |
|
| 299 |
### 👁️ Multimodal
|
|
@@ -303,121 +226,22 @@ AA-Briefcase, GDPval-AA v2, Terminal-Bench 4.0, SciCode, and Humanity's Last Exa
|
|
| 303 |
| **MMMU-Pro** | 76.0% | **87.0%** 🥇 | — | 85.0% | 81.0% | — |
|
| 304 |
| **GDP.pdf** | 14.8% | **31.0%** 🥇 | 26.2% | 21.6% | 22.0% | 11.2% |
|
| 305 |
|
| 306 |
-
<div style="border-left: 6px solid #faad14; padding: 16px; border-radius: 8px; margin: 20px 0;">
|
| 307 |
-
<strong>📌 Read-out</strong><br>
|
| 308 |
-
Multimodal reasoning is the <strong>most direct improvement opportunity</strong> for the next release. MMMU-Pro at
|
| 309 |
-
<strong>76.0%</strong> trails GPT-6 Astra (87.0%), Claude Opus 5 (85.0%), and Kimi K3 (81.0%), and document-heavy
|
| 310 |
-
visual QA (<strong>GDP.pdf</strong>) remains a known weak spot at <strong>14.8%</strong>.
|
| 311 |
-
</div>
|
| 312 |
-
|
| 313 |
-
---
|
| 314 |
-
|
| 315 |
-
### 🏅 Category Leadership Summary
|
| 316 |
-
|
| 317 |
-
<div style="border: 1px solid #d9d9d9; border-radius: 12px; padding: 20px 24px; margin: 24px 0; background: linear-gradient(135deg, #ffffff 0%, #f6ffed 100%);">
|
| 318 |
-
|
| 319 |
-
| Domain | Step-5-Preview Standing | Highlight |
|
| 320 |
-
|:---|:---|:---|
|
| 321 |
-
| **Long-Context Reasoning** | 🥈 Near-tied for #1 overall | AA-LCR v1.1 **88.3%** (+7.6 over GPT-6 Astra, +9.0 over Opus 5) |
|
| 322 |
-
| **Coding / SWE (open-weight)** | 🥇 #1 open-weight | DeepSWE v1.1 **67.7%**, StepCodeBench **49.0%**, ProgramBench **80.5%** |
|
| 323 |
-
| **Security / Cyber** | 🥇 #1 in set | CyberGym **84.7%** |
|
| 324 |
-
| **Finance** | 🥈 #2 overall, #1 open-weight | FrontierFinance **66.4%**, FinanceDR **55.8%** |
|
| 325 |
-
| **Terminal Agents** | 🥈 #2 open-weight | Terminal-Bench 4 **33.3%** (GLM-5.3 41.9% > Step 33.3% > Kimi 12.6%) |
|
| 326 |
-
| **Tool Use** | 🥉 Competitive | MCP-Atlas **85.6%**, HLE w/ tools **59.4%** |
|
| 327 |
-
| **General Knowledge** | Top tier | GPQA Diamond **93.5%** (tied #3), HLE **46.5%** (#5) |
|
| 328 |
-
| **Multimodal** | 🔻 Trailing | MMMU-Pro **76.0%** — targeted for improvement |
|
| 329 |
-
|
| 330 |
-
</div>
|
| 331 |
-
|
| 332 |
<details>
|
| 333 |
-
<summary><strong>📝 Benchmark
|
| 334 |
|
| 335 |
- **DeepSWE v1.1** was evaluated using the SWE-agent harness with `temperature=1.0` and `top_p=0.95`.
|
| 336 |
- **StepCodeBench** achieved **49.0% avg@4**.
|
| 337 |
- **GDPval-AA v2** results are from Artificial Analysis as of September 19, 2026.
|
| 338 |
- **AA-LCR v1.1**: Step-5-Preview scored **88.3%**, statistically tied with Kimi K3 (88.7%).
|
| 339 |
-
- **Terminal-Bench 4**: Step-5-Preview **33.3%** vs. Kimi K3 **
|
| 340 |
-
- **SciCode**: Step-5-Preview
|
| 341 |
- **Multimodal**: MMMU-Pro and GDP.pdf were run with the unified multimodal encoder at default resolution.
|
| 342 |
- **Output Speed**: 99.8 tokens/sec (GLM-5.3: 72.1 tokens/sec).
|
| 343 |
- **Time to First Token**: 2.96 seconds (GLM-5.3: 2.99s; Claude Opus 5: 56.84s at max effort).
|
| 344 |
- Missing entries (**—**) reflect benchmarks that were not publicly reported for that model at the time of writing.
|
| 345 |
-
- Detailed benchmark descriptions in the Evaluation table are sourced from public benchmark documentation.
|
| 346 |
|
| 347 |
</details>
|
| 348 |
|
| 349 |
-
### Benchmark Takeaways
|
| 350 |
-
|
| 351 |
-
<div style="border-left: 6px solid #2f54eb; padding: 16px; border-radius: 8px; margin: 20px 0;">
|
| 352 |
-
<strong>🧠 Long-Context Reasoning</strong><br>
|
| 353 |
-
AA-LCR v1.1 at <strong>88.3%</strong> is one of Step-5-Preview's standout results — effectively tied for first with Kimi K3
|
| 354 |
-
and well ahead of models many times its size. This validates the 92-layer narrow-deep design and Sparse GQA.
|
| 355 |
-
</div>
|
| 356 |
-
|
| 357 |
-
<div style="border-left: 6px solid #13c2c2; padding: 16px; border-radius: 8px; margin: 20px 0;">
|
| 358 |
-
<strong>💻 Coding & Software Engineering</strong><br>
|
| 359 |
-
Step-5-Preview <strong>leads all open-weight models</strong> on DeepSWE v1.1, StepCodeBench, ProgramBench, and SWE-Atlas-QnA,
|
| 360 |
-
surpassing Kimi K3 and GLM-5.3. It trails only the larger closed-source models (Claude Opus 5 and GPT-6 Astra).
|
| 361 |
-
</div>
|
| 362 |
-
|
| 363 |
-
<div style="border-left: 6px solid #f5222d; padding: 16px; border-radius: 8px; margin: 20px 0;">
|
| 364 |
-
<strong>🤖 Agentic Tasks</strong><br>
|
| 365 |
-
Strong performance on Terminal-Bench 4 (<strong>33.3%</strong>) and HLE with tools (<strong>59.4%</strong>).
|
| 366 |
-
The Terminal-Bench score is <strong>2.6× higher than Kimi K3</strong> and <strong>1.24× higher than DeepSeek V4.1 Flash</strong>.
|
| 367 |
-
</div>
|
| 368 |
-
|
| 369 |
-
<div style="border-left: 6px solid #a0d911; padding: 16px; border-radius: 8px; margin: 20px 0;">
|
| 370 |
-
<strong>💰 Financial & Deep Research</strong><br>
|
| 371 |
-
Highly competitive on FrontierFinance (<strong>66.4%</strong>) and FinStepBench-FinanceDR (<strong>55.8%</strong>),
|
| 372 |
-
nearly matching top closed-source models like Claude Opus 5 and outperforming both GPT-6 Astra and Kimi K3 by significant margins.
|
| 373 |
-
</div>
|
| 374 |
-
|
| 375 |
-
---
|
| 376 |
-
|
| 377 |
-
## 🤖 Agentic Capabilities
|
| 378 |
-
|
| 379 |
-
<div align="center">
|
| 380 |
-
<img src="./Step-5/agentic_workflow.png" alt="Agentic Workflow" width="90%">
|
| 381 |
-
</div>
|
| 382 |
-
|
| 383 |
-
### 24-Hour Autonomous GPU Kernel Optimization
|
| 384 |
-
|
| 385 |
-
In a landmark demonstration of sustained agentic execution, Step-5-Preview was tasked with **autonomously optimizing an H100 GPU kernel for up to 24 consecutive hours**. The model:
|
| 386 |
-
|
| 387 |
-
- Independently modified code
|
| 388 |
-
- Ran tests and compared results
|
| 389 |
-
- Iterated based on performance outcomes
|
| 390 |
-
- **Reached 508 TFLOPS after approximately 22 hours**
|
| 391 |
-
|
| 392 |
-
For comparison, **Claude Opus 5 achieved 493 TFLOPS** in the same experiment. This demonstrates Step-5-Preview's ability to sustain productive work over extended periods without human intervention.
|
| 393 |
-
|
| 394 |
-
### Automated Post-Training Experiments
|
| 395 |
-
|
| 396 |
-
In another 24-hour experiment, Step-5-Preview autonomously improved the accuracy of **Qwen3-30B-A3B on AIME24 from 53.3% to 60%** through automated post-training experiments. This showcases the model's capacity for self-directed research and optimization.
|
| 397 |
-
|
| 398 |
-
### Long-Horizon Agent Workflows
|
| 399 |
-
|
| 400 |
-
The model is specifically optimized for agent workflows that require:
|
| 401 |
-
|
| 402 |
-
- Searching and information retrieval
|
| 403 |
-
- Running code and processing tool returns
|
| 404 |
-
- Multi-turn tool calls with sustained execution
|
| 405 |
-
- Iterative refinement based on intermediate results
|
| 406 |
-
- Self-correction and error recovery over thousands of steps
|
| 407 |
-
|
| 408 |
-
---
|
| 409 |
-
|
| 410 |
-
## 💼 Real-World Use Cases
|
| 411 |
-
|
| 412 |
-
StepFun demonstrated the model's capabilities across several complex, real-world projects:
|
| 413 |
-
|
| 414 |
-
- **ESP32 Development Board Modifications:** Executed development tasks for over 3 hours, demonstrating hardware programming capabilities.
|
| 415 |
-
- **Front-End Design with 3D Asset Generation:** Full-stack development workflows including visual design.
|
| 416 |
-
- **Full-Process Financial Research:** End-to-end investment research workflows, from data gathering to report generation.
|
| 417 |
-
- **Software Engineering:** Comprehensive coding tasks beyond traditional code generation, including front-end, visual development, and programmable hardware scenarios.
|
| 418 |
-
- **Autonomous Research Assistant:** Capable of reading papers, running experiments, and summarizing findings.
|
| 419 |
-
- **Customer Support Automation:** Handles multi-turn conversations with tool calls to internal systems.
|
| 420 |
-
|
| 421 |
---
|
| 422 |
|
| 423 |
## ⚡ Quickstart
|
|
@@ -572,20 +396,126 @@ response = client.chat.completions.create(
|
|
| 572 |
print(response.choices[0].message.content)
|
| 573 |
```
|
| 574 |
|
| 575 |
-
|
| 576 |
-
|
| 577 |
-
|
| 578 |
-
|
| 579 |
-
|
| 580 |
-
|
| 581 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 582 |
|
| 583 |
---
|
| 584 |
|
| 585 |
-
##
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 586 |
|
| 587 |
-
|
| 588 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 589 |
|
| 590 |
| Benchmark | Score | Notes |
|
| 591 |
|:---|:---|:---|
|
|
@@ -603,60 +533,79 @@ All evaluations used the model's `high` reasoning effort setting unless otherwis
|
|
| 603 |
| **SWE-Atlas-Test-writing** | 50.8% | 90 tasks on writing production-grade unit, integration, and acceptance tests for real repositories |
|
| 604 |
| **StepCode-Bench-Daily** | 64.9% | StepFun internal; 553 repos, 9 task types, 20 domains, 33 languages; daily-difficulty slice |
|
| 605 |
| **StepCode-Bench-General** | 65.0% | StepFun internal; same corpus as StepCodeBench; general-difficulty slice |
|
| 606 |
-
| **Agents' Last Exam (ALE-CLI)** | 29.5% | Linux-only CLI subset of ALE; 40 industry subfields
|
| 607 |
| **GDPval-AA v2** | 1571 | Artificial Analysis, Sep 19, 2026 |
|
| 608 |
-
| **AA-Briefcase** | 1417 | Elo
|
| 609 |
| **Toolathlon-Verified** | 74.1% | 108 expert-authored tasks; multi-app workflows averaging ~20 turns; strictly verifiable via scripts |
|
| 610 |
| **MCP-Atlas** | 85.6% | 1,000 tasks (500 public + 500 private); 36 real MCP servers, 220 tools; 3–6 tool calls/task |
|
| 611 |
-
| **PresentBench** | 76.8% | 238 slide-generation instances; avg 54.1 rubric checklist items per instance
|
| 612 |
-
| **JobBench** | 59.0% | 130 tasks across 35 occupations;
|
| 613 |
-
| **Apex-Agents** | 37.8% | Long-horizon tasks in investment banking, consulting,
|
| 614 |
-
| **DRACO** | 83.3% | Cross-domain deep research
|
| 615 |
-
| **BrowseComp** | 88.7% | 1,266 hard-to-find web information retrieval questions
|
| 616 |
| **HLE w/ tools** | 59.4% | +12.9 pts over no-tools HLE |
|
| 617 |
| **FrontierFinance** | 66.4% | +11.4 pts over GPT-6 Astra |
|
| 618 |
| **FinStepBench-LiveSearch** | 74.5% | Tied with GPT-6 Astra |
|
| 619 |
| **FinStepBench-CorporateValuation** | 60.6% | Tied with Kimi K3 |
|
| 620 |
-
| **FinStepBench-FinanceDR** | 55.8% | StepFun internal;
|
| 621 |
-
| **OfficeQA Pro** | 60.3% | Databricks; 90 questions over large enterprise financial document collections
|
| 622 |
-
| **
|
| 623 |
| **GDP.pdf** | 14.8% | Known weak spot |
|
| 624 |
-
| **GPQA Diamond** | 93.5% | 198 PhD-level science MCQs
|
| 625 |
| **HLE** | 46.5% | 59.4% with tools |
|
| 626 |
| **AA-LCR v1.1** | 88.3% | Tied with Kimi K3 (88.7%) |
|
| 627 |
-
| **CritPt** | 20.9% | 71 unpublished research-level physics challenges
|
| 628 |
-
| **MMMU-Pro** | 76.0% | Robust multimodal benchmark;
|
| 629 |
| **Output Speed** | 99.8 tokens/sec | GLM-5.3: 72.1 tokens/sec |
|
| 630 |
| **Time to First Token** | 2.96s | GLM-5.3: 2.99s; Claude Opus 5: 56.84s (max effort) |
|
| 631 |
|
| 632 |
-
|
|
|
|
|
|
|
|
|
|
| 633 |
|
| 634 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 635 |
|
| 636 |
-
|
| 637 |
-
|
| 638 |
-
|
| 639 |
-
- **
|
|
|
|
|
|
|
|
|
|
| 640 |
- **Multimodal Limitations:** Video understanding is limited to clips under 5 minutes and 128 MB. Extremely high-resolution images may be downscaled.
|
| 641 |
-
- **Language Coverage:**
|
| 642 |
|
| 643 |
-
|
| 644 |
|
| 645 |
-
|
|
|
|
| 646 |
|
| 647 |
-
StepFun is committed to the responsible development and deployment of AI.
|
| 648 |
|
| 649 |
-
- **Safety Alignment:**
|
| 650 |
-
- **Bias Mitigation:** Training data was filtered to reduce harmful stereotypes and biases.
|
| 651 |
-
- **Transparency:**
|
| 652 |
- **License Restrictions:** The StepFun Community License prohibits certain high-risk uses, including autonomous weapons, surveillance, and malicious cyber activities.
|
| 653 |
-
- **Content Provenance:**
|
| 654 |
|
| 655 |
-
|
| 656 |
|
| 657 |
-
|
| 658 |
|
| 659 |
-
|
|
|
|
| 660 |
|
| 661 |
| Precision | Minimum GPU Memory | Recommended GPU Configuration |
|
| 662 |
|:---|:---|:---|
|
|
@@ -664,12 +613,12 @@ We urge all users to consider the ethical implications of their applications and
|
|
| 664 |
| **FP8** | 600 GB | 4× H100 80GB (tensor parallel) |
|
| 665 |
| **INT4** | 300 GB | 4× A100 80GB (tensor parallel) |
|
| 666 |
|
| 667 |
-
For inference with 1M context, additional memory is required for KV cache.
|
| 668 |
-
offloading techniques available in vLLM and SGLang.
|
| 669 |
|
| 670 |
-
|
| 671 |
|
| 672 |
-
|
|
|
|
| 673 |
|
| 674 |
| Metric | Value |
|
| 675 |
|:---|:---|
|
|
@@ -682,11 +631,10 @@ offloading techniques available in vLLM and SGLang.
|
|
| 682 |
|
| 683 |
*Measured on 8× H100 80GB with vLLM, batch size 1, BF16.*
|
| 684 |
|
| 685 |
-
|
| 686 |
-
|
| 687 |
-
## 📚 Citation
|
| 688 |
|
| 689 |
-
|
|
|
|
| 690 |
|
| 691 |
```bibtex
|
| 692 |
@misc{stepfun2026step5preview,
|
|
@@ -698,29 +646,7 @@ If you use Step-5-Preview in your research, please cite:
|
|
| 698 |
}
|
| 699 |
```
|
| 700 |
|
| 701 |
-
|
| 702 |
-
|
| 703 |
-
## 📜 License
|
| 704 |
-
|
| 705 |
-
Step-5-Preview is released under the **StepFun Community License**.
|
| 706 |
-
See the [LICENSE](https://huggingface.co/SHSLab/Step-5-Preview-BF16/blob/main/LICENSE) file for full terms.
|
| 707 |
-
|
| 708 |
-
<div style="border-left: 6px solid #faad14; padding: 16px; border-radius: 8px; margin: 20px 0;">
|
| 709 |
-
<strong>⚠️ Usage Restrictions</strong><br>
|
| 710 |
-
• Commercial use is permitted under the StepFun Community License.<br>
|
| 711 |
-
• Redistribution must include the license and attribution.<br>
|
| 712 |
-
• See LICENSE for full details.
|
| 713 |
-
</div>
|
| 714 |
-
|
| 715 |
-
---
|
| 716 |
-
|
| 717 |
-
## 📬 Contact
|
| 718 |
-
|
| 719 |
-
- **Hugging Face:** [SHSLab](https://huggingface.co/SHSLab)
|
| 720 |
-
- **GitHub:** [github.com/stepfun-ai](https://github.com/stepfun-ai)
|
| 721 |
-
- **Discord:** [Join our Discord](https://discord.gg/stepfun)
|
| 722 |
-
- **Email:** [opensource@stepfun.com](mailto:opensource@stepfun.com)
|
| 723 |
-
- **Website:** [stepfun.com](https://stepfun.com)
|
| 724 |
|
| 725 |
---
|
| 726 |
|
|
|
|
| 1 |
+
---
|
| 2 |
+
language:
|
| 3 |
+
- en
|
| 4 |
+
- zh
|
| 5 |
+
- multilingual
|
| 6 |
+
license: other
|
| 7 |
+
license_name: stepfun-community-license
|
| 8 |
+
license_link: https://huggingface.co/SHSLab/Step-5-Preview-BF16/blob/main/LICENSE
|
| 9 |
+
library_name: transformers
|
| 10 |
+
pipeline_tag: text-generation
|
| 11 |
+
tags:
|
| 12 |
+
- stepfun
|
| 13 |
+
- step-5
|
| 14 |
+
- moe
|
| 15 |
+
- mixture-of-experts
|
| 16 |
+
- agentic
|
| 17 |
+
- coding
|
| 18 |
+
- software-engineering
|
| 19 |
+
- long-context
|
| 20 |
+
- 1m-context
|
| 21 |
+
- multimodal
|
| 22 |
+
- text-generation
|
| 23 |
+
- image
|
| 24 |
+
- video
|
| 25 |
+
- sparse-attention
|
| 26 |
+
- gqa
|
| 27 |
+
- financial-analysis
|
| 28 |
+
- deep-research
|
| 29 |
+
- tool-calling
|
| 30 |
+
- parallel-tool-calling
|
| 31 |
+
- json-schema
|
| 32 |
+
---
|
| 33 |
+
|
| 34 |
# Step-5-Preview
|
| 35 |
|
| 36 |
<div align="center">
|
|
|
|
| 48 |
|
| 49 |
</div>
|
| 50 |
|
| 51 |
+
> **🔥 Step-5-Preview is now available.**
|
| 52 |
+
>
|
| 53 |
+
> Step-5-Preview is a 600B-parameter sparse Mixture-of-Experts model with 27B active parameters, a 1M-token context window, and native support for text, image, and video inputs.
|
| 54 |
+
>
|
| 55 |
+
> Weights are available on Hugging Face (`SHSLab/Step-5-Preview-BF16`). Available via Step API, or self-hosted with vLLM / SGLang.
|
|
|
|
|
|
|
|
|
|
|
|
|
| 56 |
|
| 57 |
---
|
| 58 |
|
|
|
|
| 60 |
|
| 61 |
- [Introduction](#-introduction)
|
| 62 |
- [Key Features](#-key-features)
|
|
|
|
| 63 |
- [Model Specifications](#-model-specifications)
|
|
|
|
| 64 |
- [Benchmark Results](#-benchmark-results)
|
|
|
|
|
|
|
| 65 |
- [Quickstart](#-quickstart)
|
| 66 |
- [Deployment](#-deployment)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 67 |
- [License](#-license)
|
| 68 |
- [Contact](#-contact)
|
| 69 |
+
- [More details](#-more-details) *(architecture, training, agentic demos, evaluation, limitations)*
|
| 70 |
|
| 71 |
---
|
| 72 |
|
| 73 |
## 🚀 Introduction
|
| 74 |
|
| 75 |
+
**Step-5-Preview** is StepFun's flagship foundation model, designed for **real-world agentic tasks**. It targets professional domains such as **AI coding, software engineering, professional knowledge work, and financial analysis**.
|
|
|
|
| 76 |
|
| 77 |
+
StepFun's core philosophy for Step 5 is the **"Pareto Frontier"** — balancing intelligence against cost. While earlier scaling efforts traded more compute for stronger intelligence, the next phase focuses on improving the **efficiency of converting compute into intelligence**.
|
|
|
|
|
|
|
| 78 |
|
| 79 |
+
> **Why Step 5 Preview?**
|
| 80 |
+
>
|
| 81 |
+
> - 600B total parameters, 27B active per token — near-frontier performance at a fraction of the compute.
|
| 82 |
+
> - 1M-token context window without proportional cost increases.
|
| 83 |
+
> - Competitive benchmark scores against models with 3–5× more parameters.
|
| 84 |
+
> - Built for agents: long-horizon reasoning, tool use, and autonomous execution.
|
|
|
|
| 85 |
|
| 86 |
+
Step-5-Preview skips the entire Step 4.x line, going directly from Step-3.7-Flash to Step 5.
|
|
|
|
| 87 |
|
| 88 |
---
|
| 89 |
|
|
|
|
| 100 |
- **Parallel Tool Calling:** Natively supported for agentic workflows.
|
| 101 |
- **Strict JSON Schema Output:** Reliable integration into structured systems.
|
| 102 |
- **OpenAI-Compatible API:** Available via Step API and third-party gateways.
|
| 103 |
+
- **Open Weights:** BF16 checkpoint available under `SHSLab/Step-5-Preview-BF16`.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 104 |
|
| 105 |
---
|
| 106 |
|
|
|
|
| 122 |
| **Reasoning Effort** | `low` / `medium` / `high` (`xhigh`) |
|
| 123 |
| **Tool Calling** | Parallel, strict JSON schema |
|
| 124 |
| **Intelligence Index** | 44 (Artificial Analysis v4.3.2) |
|
| 125 |
+
| **Open Weights** | BF16 checkpoint available |
|
| 126 |
| **API Availability** | Immediate (OpenAI-compatible) |
|
| 127 |
| **License** | StepFun Community License |
|
| 128 |
|
| 129 |
---
|
| 130 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 131 |
## 📊 Benchmark Results
|
| 132 |
|
| 133 |
<div align="center">
|
|
|
|
| 138 |
|
| 139 |
**Overall Score: 44** (Intelligence Index v4.3.2, recalibrated September 7, 2026)
|
| 140 |
|
| 141 |
+
The index covers 10 evaluations including AA-Briefcase, GDPval-AA v2, Terminal-Bench 4.0, SciCode, and Humanity's Last Exam.
|
|
|
|
|
|
|
| 142 |
|
| 143 |
+
> **How to read the tables below**
|
| 144 |
+
>
|
| 145 |
+
> - **Bold** marks the best score in the row across all six models.
|
| 146 |
+
> - **🥇** flags the leader for that benchmark.
|
| 147 |
+
> - **—** indicates the model was not evaluated or did not report a score.
|
| 148 |
+
> - Step-5-Preview results use the `high` reasoning-effort setting unless noted.
|
| 149 |
+
>
|
| 150 |
+
> **Comparison set:** Step-5-Preview · GPT-6 Astra (Max) · Fable 5.1 · Claude Opus 5 (Max) · Kimi K3 (Max) · GLM-5.3 (Max)
|
| 151 |
+
>
|
| 152 |
+
> **Availability:** Open-weight — Step-5-Preview, Kimi K3, Qwen3.8 Max, GLM-5.3. Closed-source — GPT-6 Astra, Claude Opus 5. Fable 5.1 — not publicly stated.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 153 |
|
| 154 |
---
|
| 155 |
|
|
|
|
| 162 |
| **AA-LCR v1.1** | 88.3% | 80.7% | 85.3% | 79.3% | **88.7%** 🥇 | 79.7% |
|
| 163 |
| **CritPt** | 20.9% | **31.7%** 🥇 | 29.7% | 29.1% | 23.4% | 19.1% |
|
| 164 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 165 |
---
|
| 166 |
|
| 167 |
### 💻 Coding & Software Engineering
|
|
|
|
| 183 |
| **StepCode-Bench-Daily** | 64.9% | — | — | **77.6%** 🥇 | 57.7% | 69.1% |
|
| 184 |
| **StepCode-Bench-General** | 65.0% | 64.3% | — | **68.3%** 🥇 | 65.2% | 62.0% |
|
| 185 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 186 |
---
|
| 187 |
|
| 188 |
### 🤖 Agents, Tool Use & Automation
|
|
|
|
| 203 |
| **BrowseComp** | 88.7% | **91.5%** 🥇 | — | 90.2% | 91.2% | — |
|
| 204 |
| **HLE w/ tools** | 59.4% | 57.2% | **65.0%** 🥇 | 63.6% | 56.0% | 62.5% |
|
| 205 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 206 |
---
|
| 207 |
|
| 208 |
### 💰 Finance & Professional Work
|
|
|
|
| 214 |
| **FinStepBench-FinanceDR** | 55.8% | 45.0% | — | **59.1%** 🥇 | 48.9% | 53.3% |
|
| 215 |
| **FrontierFinance** | 66.4% | 55.0% | — | **69.7%** 🥇 | 62.6% | 64.1% |
|
| 216 |
| **OfficeQA Pro** | 60.3% | **67.7%** 🥇 | — | 64.7% | 62.6% | 59.1% |
|
| 217 |
+
| **SpeadSheet v2** | 29.4% | 31.4% | — | **32.8%** 🥇 | 31.9% | 30.5% |
|
| 218 |
| **GDP.pdf** | 14.8% | **31.0%** 🥇 | 26.2% | 21.6% | 22.0% | 11.2% |
|
| 219 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 220 |
---
|
| 221 |
|
| 222 |
### 👁️ Multimodal
|
|
|
|
| 226 |
| **MMMU-Pro** | 76.0% | **87.0%** 🥇 | — | 85.0% | 81.0% | — |
|
| 227 |
| **GDP.pdf** | 14.8% | **31.0%** 🥇 | 26.2% | 21.6% | 22.0% | 11.2% |
|
| 228 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 229 |
<details>
|
| 230 |
+
<summary><strong>📝 Benchmark methodology notes</strong></summary>
|
| 231 |
|
| 232 |
- **DeepSWE v1.1** was evaluated using the SWE-agent harness with `temperature=1.0` and `top_p=0.95`.
|
| 233 |
- **StepCodeBench** achieved **49.0% avg@4**.
|
| 234 |
- **GDPval-AA v2** results are from Artificial Analysis as of September 19, 2026.
|
| 235 |
- **AA-LCR v1.1**: Step-5-Preview scored **88.3%**, statistically tied with Kimi K3 (88.7%).
|
| 236 |
+
- **Terminal-Bench 4**: Step-5-Preview **33.3%** vs. Kimi K3 ~**12.6%** (2.6×) and DeepSeek V4.1 Flash **26.8%** (1.24×).
|
| 237 |
+
- **SciCode**: Step-5-Preview **58.9%**, above GPT-6 Astra (56.5%) and Claude Opus 5 (56.4%).
|
| 238 |
- **Multimodal**: MMMU-Pro and GDP.pdf were run with the unified multimodal encoder at default resolution.
|
| 239 |
- **Output Speed**: 99.8 tokens/sec (GLM-5.3: 72.1 tokens/sec).
|
| 240 |
- **Time to First Token**: 2.96 seconds (GLM-5.3: 2.99s; Claude Opus 5: 56.84s at max effort).
|
| 241 |
- Missing entries (**—**) reflect benchmarks that were not publicly reported for that model at the time of writing.
|
|
|
|
| 242 |
|
| 243 |
</details>
|
| 244 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 245 |
---
|
| 246 |
|
| 247 |
## ⚡ Quickstart
|
|
|
|
| 396 |
print(response.choices[0].message.content)
|
| 397 |
```
|
| 398 |
|
| 399 |
+
> **Recommended deployment configurations**
|
| 400 |
+
>
|
| 401 |
+
> - **BF16:** 8× H100 80GB (tensor parallel)
|
| 402 |
+
> - **FP8:** 4× H100 80GB (coming soon)
|
| 403 |
+
> - **Context length:** Up to 1M tokens
|
| 404 |
+
> - **Reasoning parser:** Use `stepfun` for vLLM / SGLang
|
| 405 |
+
|
| 406 |
+
---
|
| 407 |
+
|
| 408 |
+
## 📜 License
|
| 409 |
+
|
| 410 |
+
Step-5-Preview is released under the **StepFun Community License**. See the [LICENSE](https://huggingface.co/SHSLab/Step-5-Preview-BF16/blob/main/LICENSE) file for full terms.
|
| 411 |
+
|
| 412 |
+
> **Usage restrictions**
|
| 413 |
+
>
|
| 414 |
+
> - Commercial use is permitted under the StepFun Community License.
|
| 415 |
+
> - Redistribution must include the license and attribution.
|
| 416 |
+
> - See LICENSE for full details.
|
| 417 |
+
|
| 418 |
+
---
|
| 419 |
+
|
| 420 |
+
## 📬 Contact
|
| 421 |
+
|
| 422 |
+
- **Hugging Face:** [SHSLab](https://huggingface.co/SHSLab)
|
| 423 |
+
- **GitHub:** [github.com/stepfun-ai](https://github.com/stepfun-ai)
|
| 424 |
+
- **Discord:** [Join our Discord](https://discord.gg/stepfun)
|
| 425 |
+
- **Email:** [opensource@stepfun.com](mailto:opensource@stepfun.com)
|
| 426 |
+
- **Website:** [stepfun.com](https://stepfun.com)
|
| 427 |
|
| 428 |
---
|
| 429 |
|
| 430 |
+
## 📚 More details
|
| 431 |
+
|
| 432 |
+
<details>
|
| 433 |
+
<summary><strong>🏗️ Model Architecture</strong></summary>
|
| 434 |
+
|
| 435 |
+
<div align="center">
|
| 436 |
+
<img src="./Step-5/architecture.png" alt="Step 5 Architecture" width="85%">
|
| 437 |
+
</div>
|
| 438 |
+
|
| 439 |
+
### 92-Layer "Narrow but Deep" Design
|
| 440 |
+
|
| 441 |
+
Step-5-Preview uses a **92-layer Transformer** with a narrow-deep configuration. This design creates longer information propagation paths for implicit multi-hop reasoning during long prefill operations.
|
| 442 |
+
|
| 443 |
+
### Sparse Grouped-Query Attention (GQA) with Block-Wise Token Merging
|
| 444 |
+
|
| 445 |
+
To handle the 1M-token context window efficiently, Step-5-Preview uses **Sparse GQA with block-wise token merging**. This mechanism uses sparse indexing to select only historical information relevant to the current task, reducing the number of tokens that enter attention computation. This cuts indexer and top-k selection costs to approximately one-eighth of a denser baseline.
|
| 446 |
+
|
| 447 |
+
### Multimodal Encoder
|
| 448 |
+
|
| 449 |
+
A unified multimodal encoder processes text, images, and video frames into a shared latent space. Video is sampled at adaptive frame rates and encoded with temporal attention, allowing the model to understand motion and long-range dependencies in screen recordings, demonstrations, and real-world footage.
|
| 450 |
+
|
| 451 |
+
</details>
|
| 452 |
+
|
| 453 |
+
<details>
|
| 454 |
+
<summary><strong>📚 Training Data</strong></summary>
|
| 455 |
+
|
| 456 |
+
Step-5-Preview was trained on a carefully curated corpus spanning:
|
| 457 |
+
|
| 458 |
+
- **Code repositories** from multiple languages (Python, C++, Rust, JavaScript, Go, etc.)
|
| 459 |
+
- **Technical documentation**, API references, and software engineering forums
|
| 460 |
+
- **Scientific papers** in computer science, mathematics, physics, and finance
|
| 461 |
+
- **Financial reports**, earnings calls, and market analyses
|
| 462 |
+
- **Multimodal data** including screenshots, UI mockups, video tutorials, and screen recordings
|
| 463 |
+
- **Agentic trajectories** from simulated and real tool-use environments
|
| 464 |
+
|
| 465 |
+
The data mixture was optimized for long-horizon reasoning and tool use, with a strong emphasis on real-world professional tasks. All data was filtered for quality, safety, and license compliance. Training combined next-token prediction with reinforcement learning from human feedback (RLHF) focused on agentic objectives.
|
| 466 |
+
|
| 467 |
+
</details>
|
| 468 |
+
|
| 469 |
+
<details>
|
| 470 |
+
<summary><strong>🤖 Agentic Capabilities</strong></summary>
|
| 471 |
+
|
| 472 |
+
<div align="center">
|
| 473 |
+
<img src="./Step-5/agentic_workflow.png" alt="Agentic Workflow" width="90%">
|
| 474 |
+
</div>
|
| 475 |
+
|
| 476 |
+
### 24-Hour Autonomous GPU Kernel Optimization
|
| 477 |
+
|
| 478 |
+
Step-5-Preview was tasked with autonomously optimizing an H100 GPU kernel for up to 24 consecutive hours. The model:
|
| 479 |
+
|
| 480 |
+
- Independently modified code
|
| 481 |
+
- Ran tests and compared results
|
| 482 |
+
- Iterated based on performance outcomes
|
| 483 |
+
- **Reached 508 TFLOPS after approximately 22 hours**
|
| 484 |
+
|
| 485 |
+
For comparison, Claude Opus 5 achieved 493 TFLOPS in the same experiment.
|
| 486 |
+
|
| 487 |
+
### Automated Post-Training Experiments
|
| 488 |
+
|
| 489 |
+
In another 24-hour experiment, Step-5-Preview autonomously improved the accuracy of Qwen3-30B-A3B on AIME24 from 53.3% to 60% through automated post-training experiments.
|
| 490 |
+
|
| 491 |
+
### Long-Horizon Agent Workflows
|
| 492 |
|
| 493 |
+
Optimized for workflows that require:
|
| 494 |
+
|
| 495 |
+
- Searching and information retrieval
|
| 496 |
+
- Running code and processing tool returns
|
| 497 |
+
- Multi-turn tool calls with sustained execution
|
| 498 |
+
- Iterative refinement based on intermediate results
|
| 499 |
+
- Self-correction and error recovery over thousands of steps
|
| 500 |
+
|
| 501 |
+
</details>
|
| 502 |
+
|
| 503 |
+
<details>
|
| 504 |
+
<summary><strong>💼 Real-World Use Cases</strong></summary>
|
| 505 |
+
|
| 506 |
+
- **ESP32 Development Board Modifications:** Executed development tasks for over 3 hours.
|
| 507 |
+
- **Front-End Design with 3D Asset Generation:** Full-stack development workflows including visual design.
|
| 508 |
+
- **Full-Process Financial Research:** End-to-end investment research, from data gathering to report generation.
|
| 509 |
+
- **Software Engineering:** Comprehensive coding tasks including front-end, visual development, and programmable hardware.
|
| 510 |
+
- **Autonomous Research Assistant:** Reading papers, running experiments, and summarizing findings.
|
| 511 |
+
- **Customer Support Automation:** Multi-turn conversations with tool calls to internal systems.
|
| 512 |
+
|
| 513 |
+
</details>
|
| 514 |
+
|
| 515 |
+
<details>
|
| 516 |
+
<summary><strong>📈 Full Evaluation Table</strong></summary>
|
| 517 |
+
|
| 518 |
+
All evaluations used the `high` reasoning effort setting unless otherwise noted.
|
| 519 |
|
| 520 |
| Benchmark | Score | Notes |
|
| 521 |
|:---|:---|:---|
|
|
|
|
| 533 |
| **SWE-Atlas-Test-writing** | 50.8% | 90 tasks on writing production-grade unit, integration, and acceptance tests for real repositories |
|
| 534 |
| **StepCode-Bench-Daily** | 64.9% | StepFun internal; 553 repos, 9 task types, 20 domains, 33 languages; daily-difficulty slice |
|
| 535 |
| **StepCode-Bench-General** | 65.0% | StepFun internal; same corpus as StepCodeBench; general-difficulty slice |
|
| 536 |
+
| **Agents' Last Exam (ALE-CLI)** | 29.5% | Linux-only CLI subset of ALE; 40 industry subfields; best agent pass rate ~25.2% |
|
| 537 |
| **GDPval-AA v2** | 1571 | Artificial Analysis, Sep 19, 2026 |
|
| 538 |
+
| **AA-Briefcase** | 1417 | Elo; agentic knowledge work across data science, product, banking, heavy industry; private held-out test set |
|
| 539 |
| **Toolathlon-Verified** | 74.1% | 108 expert-authored tasks; multi-app workflows averaging ~20 turns; strictly verifiable via scripts |
|
| 540 |
| **MCP-Atlas** | 85.6% | 1,000 tasks (500 public + 500 private); 36 real MCP servers, 220 tools; 3–6 tool calls/task |
|
| 541 |
+
| **PresentBench** | 76.8% | 238 slide-generation instances; avg 54.1 rubric checklist items per instance |
|
| 542 |
+
| **JobBench** | 59.0% | 130 tasks across 35 occupations; avg 35.6 binary rubric criteria per task |
|
| 543 |
+
| **Apex-Agents** | 37.8% | Long-horizon tasks in investment banking, consulting, corporate law |
|
| 544 |
+
| **DRACO** | 83.3% | Cross-domain deep research; accuracy, completeness, objectivity, citation quality |
|
| 545 |
+
| **BrowseComp** | 88.7% | 1,266 hard-to-find web information retrieval questions |
|
| 546 |
| **HLE w/ tools** | 59.4% | +12.9 pts over no-tools HLE |
|
| 547 |
| **FrontierFinance** | 66.4% | +11.4 pts over GPT-6 Astra |
|
| 548 |
| **FinStepBench-LiveSearch** | 74.5% | Tied with GPT-6 Astra |
|
| 549 |
| **FinStepBench-CorporateValuation** | 60.6% | Tied with Kimi K3 |
|
| 550 |
+
| **FinStepBench-FinanceDR** | 55.8% | StepFun internal; full deep-research finance workflows |
|
| 551 |
+
| **OfficeQA Pro** | 60.3% | Databricks; 90 questions over large enterprise financial document collections |
|
| 552 |
+
| **SpeadSheet v2** | 29.4% | SpreadsheetBench 2; end-to-end business spreadsheet workflows; best model ≈34.89% |
|
| 553 |
| **GDP.pdf** | 14.8% | Known weak spot |
|
| 554 |
+
| **GPQA Diamond** | 93.5% | 198 PhD-level science MCQs; human expert avg 81% |
|
| 555 |
| **HLE** | 46.5% | 59.4% with tools |
|
| 556 |
| **AA-LCR v1.1** | 88.3% | Tied with Kimi K3 (88.7%) |
|
| 557 |
+
| **CritPt** | 20.9% | 71 unpublished research-level physics challenges |
|
| 558 |
+
| **MMMU-Pro** | 76.0% | Robust multimodal benchmark; significantly harder than MMMU |
|
| 559 |
| **Output Speed** | 99.8 tokens/sec | GLM-5.3: 72.1 tokens/sec |
|
| 560 |
| **Time to First Token** | 2.96s | GLM-5.3: 2.99s; Claude Opus 5: 56.84s (max effort) |
|
| 561 |
|
| 562 |
+
</details>
|
| 563 |
+
|
| 564 |
+
<details>
|
| 565 |
+
<summary><strong>🏅 Category summary</strong></summary>
|
| 566 |
|
| 567 |
+
| Domain | Standing | Highlight |
|
| 568 |
+
|:---|:---|:---|
|
| 569 |
+
| **Long-Context Reasoning** | Near-tied for #1 | AA-LCR v1.1 **88.3%** (+7.6 over GPT-6 Astra, +9.0 over Opus 5) |
|
| 570 |
+
| **Coding / SWE** | #1 open-weight | DeepSWE v1.1 **67.7%**, StepCodeBench **49.0%**, ProgramBench **80.5%** |
|
| 571 |
+
| **Security / Cyber** | #1 in comparison set | CyberGym **84.7%** |
|
| 572 |
+
| **Finance** | #2 overall, #1 open-weight | FrontierFinance **66.4%**, FinanceDR **55.8%** |
|
| 573 |
+
| **Terminal Agents** | #2 open-weight | Terminal-Bench 4 **33.3%** (2.6× Kimi K3) |
|
| 574 |
+
| **Tool Use** | Top-3 open-weight | MCP-Atlas **85.6%**, HLE w/ tools **59.4%** |
|
| 575 |
+
| **General Knowledge** | Top tier | GPQA Diamond **93.5%**, HLE **46.5%** |
|
| 576 |
+
| **Multimodal** | Trailing | MMMU-Pro **76.0%** — targeted for improvement |
|
| 577 |
+
|
| 578 |
+
</details>
|
| 579 |
|
| 580 |
+
<details>
|
| 581 |
+
<summary><strong>⚠️ Limitations</strong></summary>
|
| 582 |
+
|
| 583 |
+
- **Knowledge Cutoff:** Knowledge is current up to mid-2026. May not be aware of later events.
|
| 584 |
+
- **Hallucination:** Can generate plausible but incorrect information, especially in domains with sparse training data.
|
| 585 |
+
- **Long Context Degradation:** Performance may degrade for extremely long contexts beyond 500K tokens in certain tasks.
|
| 586 |
+
- **Tool Use Reliability:** Tool calling is capable but not infallible. Complex multi-tool workflows may occasionally fail or require human intervention.
|
| 587 |
- **Multimodal Limitations:** Video understanding is limited to clips under 5 minutes and 128 MB. Extremely high-resolution images may be downscaled.
|
| 588 |
+
- **Language Coverage:** Primarily optimized for English and Chinese. Performance in other languages may vary.
|
| 589 |
|
| 590 |
+
</details>
|
| 591 |
|
| 592 |
+
<details>
|
| 593 |
+
<summary><strong>⚖️ Ethical Considerations</strong></summary>
|
| 594 |
|
| 595 |
+
StepFun is committed to the responsible development and deployment of AI. Measures taken:
|
| 596 |
|
| 597 |
+
- **Safety Alignment:** Fine-tuned with RLHF to refuse harmful requests and promote helpful, honest, and harmless behavior.
|
| 598 |
+
- **Bias Mitigation:** Training data was filtered to reduce harmful stereotypes and biases. Residual biases may exist.
|
| 599 |
+
- **Transparency:** Detailed model cards and benchmark results are provided to enable informed use.
|
| 600 |
- **License Restrictions:** The StepFun Community License prohibits certain high-risk uses, including autonomous weapons, surveillance, and malicious cyber activities.
|
| 601 |
+
- **Content Provenance:** Users are encouraged to clearly label AI-generated content.
|
| 602 |
|
| 603 |
+
All users should consider the ethical implications of their applications and implement appropriate safeguards.
|
| 604 |
|
| 605 |
+
</details>
|
| 606 |
|
| 607 |
+
<details>
|
| 608 |
+
<summary><strong>🖥️ Hardware Requirements</strong></summary>
|
| 609 |
|
| 610 |
| Precision | Minimum GPU Memory | Recommended GPU Configuration |
|
| 611 |
|:---|:---|:---|
|
|
|
|
| 613 |
| **FP8** | 600 GB | 4× H100 80GB (tensor parallel) |
|
| 614 |
| **INT4** | 300 GB | 4× A100 80GB (tensor parallel) |
|
| 615 |
|
| 616 |
+
For inference with 1M context, additional memory is required for KV cache. Paged attention and offloading techniques available in vLLM and SGLang are recommended.
|
|
|
|
| 617 |
|
| 618 |
+
</details>
|
| 619 |
|
| 620 |
+
<details>
|
| 621 |
+
<summary><strong>⚡ Performance Metrics</strong></summary>
|
| 622 |
|
| 623 |
| Metric | Value |
|
| 624 |
|:---|:---|
|
|
|
|
| 631 |
|
| 632 |
*Measured on 8× H100 80GB with vLLM, batch size 1, BF16.*
|
| 633 |
|
| 634 |
+
</details>
|
|
|
|
|
|
|
| 635 |
|
| 636 |
+
<details>
|
| 637 |
+
<summary><strong>📚 Citation</strong></summary>
|
| 638 |
|
| 639 |
```bibtex
|
| 640 |
@misc{stepfun2026step5preview,
|
|
|
|
| 646 |
}
|
| 647 |
```
|
| 648 |
|
| 649 |
+
</details>
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 650 |
|
| 651 |
---
|
| 652 |
|