Image-Text-to-Text
Transformers
Safetensors
English
Chinese
multilingual
step3p5v
text-generation
stepfun
step-5
Mixture of Experts
mixture-of-experts
agentic
coding
software-engineering
long-context
1m-context
multimodal
image
video
sparse-attention
gqa
financial-analysis
deep-research
tool-calling
parallel-tool-calling
json-schema
conversational
custom_code
Instructions to use SHSLab/Step-5-Preview-BF16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use SHSLab/Step-5-Preview-BF16 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="SHSLab/Step-5-Preview-BF16", trust_remote_code=True) messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("SHSLab/Step-5-Preview-BF16", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use SHSLab/Step-5-Preview-BF16 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "SHSLab/Step-5-Preview-BF16" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SHSLab/Step-5-Preview-BF16", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/SHSLab/Step-5-Preview-BF16
- SGLang
How to use SHSLab/Step-5-Preview-BF16 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "SHSLab/Step-5-Preview-BF16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SHSLab/Step-5-Preview-BF16", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "SHSLab/Step-5-Preview-BF16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SHSLab/Step-5-Preview-BF16", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use SHSLab/Step-5-Preview-BF16 with Docker Model Runner:
docker model run hf.co/SHSLab/Step-5-Preview-BF16
Update README.md
Browse files
README.md
CHANGED
|
@@ -29,18 +29,17 @@ tags:
|
|
| 29 |
- tool-calling
|
| 30 |
- parallel-tool-calling
|
| 31 |
- json-schema
|
| 32 |
-
base_model: SHSLab/Step-5-Preview-BF16
|
| 33 |
---
|
| 34 |
|
| 35 |
# Step-5-Preview
|
| 36 |
|
| 37 |
<div align="center">
|
| 38 |
-
<img src="https://
|
| 39 |
</div>
|
| 40 |
|
| 41 |
<div align="center">
|
| 42 |
|
| 43 |
-
[](https://github.com/stepfun-ai)
|
| 45 |
[](https://discord.gg/stepfun)
|
| 46 |
[](https://huggingface.co/SHSLab/Step-5-Preview-BF16/blob/main/LICENSE)
|
|
@@ -49,13 +48,13 @@ base_model: SHSLab/Step-5-Preview-BF16
|
|
| 49 |
|
| 50 |
</div>
|
| 51 |
|
| 52 |
-
<div style="
|
| 53 |
<strong>🔥 Step-5-Preview is now available!</strong><br>
|
| 54 |
We are excited to release <strong>Step-5-Preview</strong>, our flagship foundation model for real-world agentic work.
|
| 55 |
It is a 600B-parameter sparse Mixture-of-Experts model with 27B active parameters, a 1M-token context window,
|
| 56 |
and native support for text, image, and video inputs.
|
| 57 |
<br><br>
|
| 58 |
-
<strong>Weights are available now</strong> on Hugging Face.
|
| 59 |
Try it via our API, or deploy locally with vLLM / SGLang.
|
| 60 |
</div>
|
| 61 |
|
|
@@ -67,11 +66,17 @@ base_model: SHSLab/Step-5-Preview-BF16
|
|
| 67 |
- [Key Features](#-key-features)
|
| 68 |
- [Model Architecture](#-model-architecture)
|
| 69 |
- [Model Specifications](#-model-specifications)
|
|
|
|
| 70 |
- [Benchmark Results](#-benchmark-results)
|
| 71 |
- [Agentic Capabilities](#-agentic-capabilities)
|
|
|
|
| 72 |
- [Quickstart](#-quickstart)
|
| 73 |
- [Deployment](#-deployment)
|
| 74 |
- [Evaluation](#-evaluation)
|
|
|
|
|
|
|
|
|
|
|
|
|
| 75 |
- [Citation](#-citation)
|
| 76 |
- [License](#-license)
|
| 77 |
- [Contact](#-contact)
|
|
@@ -87,7 +92,7 @@ StepFun's core philosophy for Step 5 is the **"Pareto Frontier"** — achieving
|
|
| 87 |
While previous scaling efforts focused on trading more compute for stronger intelligence, the next phase requires improving the
|
| 88 |
**efficiency of converting compute into intelligence**.
|
| 89 |
|
| 90 |
-
<div style="
|
| 91 |
<strong>💡 Why Step 5 Preview?</strong><br>
|
| 92 |
• <strong>600B total parameters, only 27B active</strong> — near-frontier performance at a fraction of the compute.<br>
|
| 93 |
• <strong>1M-token context window</strong> without proportional cost increases.<br>
|
|
@@ -95,12 +100,15 @@ While previous scaling efforts focused on trading more compute for stronger inte
|
|
| 95 |
• <strong>Built for agents</strong> — long-horizon reasoning, tool use, and autonomous execution.
|
| 96 |
</div>
|
| 97 |
|
|
|
|
|
|
|
|
|
|
| 98 |
---
|
| 99 |
|
| 100 |
## ✨ Key Features
|
| 101 |
|
| 102 |
<div align="center">
|
| 103 |
-
<img src="https://
|
| 104 |
</div>
|
| 105 |
|
| 106 |
- **Sparse Mixture-of-Experts (MoE):** 600B total parameters, 27B active per token (~4.5% sparsity).
|
|
@@ -110,14 +118,14 @@ While previous scaling efforts focused on trading more compute for stronger inte
|
|
| 110 |
- **Parallel Tool Calling:** Natively supported for agentic workflows.
|
| 111 |
- **Strict JSON Schema Output:** Reliable integration into structured systems.
|
| 112 |
- **OpenAI-Compatible API:** Available via Step API and third-party gateways.
|
| 113 |
-
- **Open Weights:** BF16 checkpoint available now.
|
| 114 |
|
| 115 |
---
|
| 116 |
|
| 117 |
## 🏗️ Model Architecture
|
| 118 |
|
| 119 |
<div align="center">
|
| 120 |
-
<img src="https://
|
| 121 |
</div>
|
| 122 |
|
| 123 |
### 92-Layer "Narrow but Deep" Design
|
|
@@ -132,12 +140,18 @@ This mechanism uses sparse indexing to select only historical information releva
|
|
| 132 |
that actually enter attention computation. StepFun states this cuts indexer and top-k selection costs to approximately
|
| 133 |
**one-eighth** of a denser baseline.
|
| 134 |
|
| 135 |
-
<div style="
|
| 136 |
<strong>⚡ Efficiency-First Scaling</strong><br>
|
| 137 |
Step 5 Preview achieves near-frontier performance with <strong>600B total parameters</strong> but only
|
| 138 |
<strong>27B active per token</strong>. This is the core of StepFun's efficiency-first philosophy.
|
| 139 |
</div>
|
| 140 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 141 |
---
|
| 142 |
|
| 143 |
## 📋 Model Specifications
|
|
@@ -160,6 +174,24 @@ that actually enter attention computation. StepFun states this cuts indexer and
|
|
| 160 |
| **Intelligence Index** | 44 (Artificial Analysis v4.3.2) |
|
| 161 |
| **Open Weights** | BF16 checkpoint available now |
|
| 162 |
| **API Availability** | Immediate (OpenAI-compatible) |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 163 |
|
| 164 |
---
|
| 165 |
|
|
@@ -168,7 +200,7 @@ that actually enter attention computation. StepFun states this cuts indexer and
|
|
| 168 |
### Artificial Analysis Intelligence Index
|
| 169 |
|
| 170 |
<div align="center">
|
| 171 |
-
<img src="https://
|
| 172 |
</div>
|
| 173 |
|
| 174 |
**Overall Score: 44** (Intelligence Index v4.3.2, recalibrated September 7, 2026)
|
|
@@ -179,7 +211,7 @@ AA-Briefcase, GDPval-AA v2, Terminal-Bench 4.0, SciCode, and Humanity's Last Exa
|
|
| 179 |
|
| 180 |
### Detailed Benchmark Scores
|
| 181 |
|
| 182 |
-
<div style="
|
| 183 |
|
| 184 |
| Benchmark | Step-5-Preview (High) | Kimi K3 (Max) | GLM-5.3 (Max) | Claude Opus 5 (Max) | GPT-6 Astra (Max) |
|
| 185 |
|:---|:---|:---|:---|:---|:---|
|
|
@@ -210,19 +242,19 @@ AA-Briefcase, GDPval-AA v2, Terminal-Bench 4.0, SciCode, and Humanity's Last Exa
|
|
| 210 |
|
| 211 |
### Benchmark Takeaways
|
| 212 |
|
| 213 |
-
<div style="
|
| 214 |
<strong>🧠 Coding & Software Engineering</strong><br>
|
| 215 |
Step-5-Preview <strong>leads all open-weight models</strong> on DeepSWE v1.1 and StepCodeBench, surpassing Kimi K3 and GLM-5.3.
|
| 216 |
It trails only the larger closed-source models (Claude Opus 5 and GPT-6 Astra).
|
| 217 |
</div>
|
| 218 |
|
| 219 |
-
<div style="
|
| 220 |
<strong>🤖 Agentic Tasks</strong><br>
|
| 221 |
Strong performance on Terminal-Bench 4.0 (<strong>33.3%</strong>) and Agents' Last Exam (ALE-CLI) (<strong>29.5%</strong>).
|
| 222 |
Terminal-Bench score is <strong>2.6× higher than Kimi K3</strong> and <strong>1.24× higher than DeepSeek V4.1 Flash</strong>.
|
| 223 |
</div>
|
| 224 |
|
| 225 |
-
<div style="
|
| 226 |
<strong>💰 Financial & Deep Research</strong><br>
|
| 227 |
Highly competitive on FrontierFinance and DRACO, nearly matching top closed-source models like Claude Opus 5.
|
| 228 |
On FrontierFinance, it outperforms both Kimi K3 and GLM-5.3 by a significant margin.
|
|
@@ -233,7 +265,7 @@ AA-Briefcase, GDPval-AA v2, Terminal-Bench 4.0, SciCode, and Humanity's Last Exa
|
|
| 233 |
## 🤖 Agentic Capabilities
|
| 234 |
|
| 235 |
<div align="center">
|
| 236 |
-
<img src="https://
|
| 237 |
</div>
|
| 238 |
|
| 239 |
### 24-Hour Autonomous GPU Kernel Optimization
|
|
@@ -251,15 +283,6 @@ For comparison, **Claude Opus 5 achieved 493 TFLOPS** in the same experiment. Th
|
|
| 251 |
|
| 252 |
In another 24-hour experiment, Step-5-Preview autonomously improved the accuracy of **Qwen3-30B-A3B on AIME24 from 53.3% to 60%** through automated post-training experiments. This showcases the model's capacity for self-directed research and optimization.
|
| 253 |
|
| 254 |
-
### Real-World Application Demonstrations
|
| 255 |
-
|
| 256 |
-
StepFun demonstrated the model's capabilities across several complex, real-world projects:
|
| 257 |
-
|
| 258 |
-
- **ESP32 Development Board Modifications:** Executed development tasks for over 3 hours, demonstrating hardware programming capabilities.
|
| 259 |
-
- **Front-End Design with 3D Asset Generation:** Full-stack development workflows including visual design.
|
| 260 |
-
- **Full-Process Financial Research:** End-to-end investment research workflows.
|
| 261 |
-
- **Software Engineering:** Comprehensive coding tasks beyond traditional code generation, including front-end, visual development, and programmable hardware scenarios.
|
| 262 |
-
|
| 263 |
### Long-Horizon Agent Workflows
|
| 264 |
|
| 265 |
The model is specifically optimized for agent workflows that require:
|
|
@@ -268,6 +291,20 @@ The model is specifically optimized for agent workflows that require:
|
|
| 268 |
- Running code and processing tool returns
|
| 269 |
- Multi-turn tool calls with sustained execution
|
| 270 |
- Iterative refinement based on intermediate results
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 271 |
|
| 272 |
---
|
| 273 |
|
|
@@ -423,7 +460,7 @@ response = client.chat.completions.create(
|
|
| 423 |
print(response.choices[0].message.content)
|
| 424 |
```
|
| 425 |
|
| 426 |
-
<div style="
|
| 427 |
<strong>📦 Recommended Deployment Configurations</strong><br>
|
| 428 |
• <strong>BF16:</strong> 8× H100 80GB (tensor parallel)<br>
|
| 429 |
• <strong>FP8:</strong> 4× H100 80GB (coming soon)<br>
|
|
@@ -454,6 +491,59 @@ All evaluations used the model's `high` reasoning effort setting unless otherwis
|
|
| 454 |
|
| 455 |
---
|
| 456 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 457 |
## 📚 Citation
|
| 458 |
|
| 459 |
If you use Step-5-Preview in your research, please cite:
|
|
@@ -475,7 +565,7 @@ If you use Step-5-Preview in your research, please cite:
|
|
| 475 |
Step-5-Preview is released under the **StepFun Community License**.
|
| 476 |
See the [LICENSE](https://huggingface.co/SHSLab/Step-5-Preview-BF16/blob/main/LICENSE) file for full terms.
|
| 477 |
|
| 478 |
-
<div style="
|
| 479 |
<strong>⚠️ Usage Restrictions</strong><br>
|
| 480 |
• Commercial use is permitted under the StepFun Community License.<br>
|
| 481 |
• Redistribution must include the license and attribution.<br>
|
|
|
|
| 29 |
- tool-calling
|
| 30 |
- parallel-tool-calling
|
| 31 |
- json-schema
|
|
|
|
| 32 |
---
|
| 33 |
|
| 34 |
# Step-5-Preview
|
| 35 |
|
| 36 |
<div align="center">
|
| 37 |
+
<img src="https://huggingface.co/SHSLab/Step-5-Preview-BF16/.Step-5/banner.png" alt="Step 5 Preview Banner" width="100%">
|
| 38 |
</div>
|
| 39 |
|
| 40 |
<div align="center">
|
| 41 |
|
| 42 |
+
[](https://huggingface.co/SHSLab)
|
| 43 |
[](https://github.com/stepfun-ai)
|
| 44 |
[](https://discord.gg/stepfun)
|
| 45 |
[](https://huggingface.co/SHSLab/Step-5-Preview-BF16/blob/main/LICENSE)
|
|
|
|
| 48 |
|
| 49 |
</div>
|
| 50 |
|
| 51 |
+
<div style="border-left: 6px solid #1890ff; padding: 16px; border-radius: 8px; margin: 20px 0;">
|
| 52 |
<strong>🔥 Step-5-Preview is now available!</strong><br>
|
| 53 |
We are excited to release <strong>Step-5-Preview</strong>, our flagship foundation model for real-world agentic work.
|
| 54 |
It is a 600B-parameter sparse Mixture-of-Experts model with 27B active parameters, a 1M-token context window,
|
| 55 |
and native support for text, image, and video inputs.
|
| 56 |
<br><br>
|
| 57 |
+
<strong>Weights are available now</strong> on Hugging Face (<code>SHSLab/Step-5-Preview-BF16</code>).
|
| 58 |
Try it via our API, or deploy locally with vLLM / SGLang.
|
| 59 |
</div>
|
| 60 |
|
|
|
|
| 66 |
- [Key Features](#-key-features)
|
| 67 |
- [Model Architecture](#-model-architecture)
|
| 68 |
- [Model Specifications](#-model-specifications)
|
| 69 |
+
- [Training Data](#-training-data)
|
| 70 |
- [Benchmark Results](#-benchmark-results)
|
| 71 |
- [Agentic Capabilities](#-agentic-capabilities)
|
| 72 |
+
- [Real-World Use Cases](#-real-world-use-cases)
|
| 73 |
- [Quickstart](#-quickstart)
|
| 74 |
- [Deployment](#-deployment)
|
| 75 |
- [Evaluation](#-evaluation)
|
| 76 |
+
- [Limitations](#-limitations)
|
| 77 |
+
- [Ethical Considerations](#-ethical-considerations)
|
| 78 |
+
- [Hardware Requirements](#-hardware-requirements)
|
| 79 |
+
- [Performance Metrics](#-performance-metrics)
|
| 80 |
- [Citation](#-citation)
|
| 81 |
- [License](#-license)
|
| 82 |
- [Contact](#-contact)
|
|
|
|
| 92 |
While previous scaling efforts focused on trading more compute for stronger intelligence, the next phase requires improving the
|
| 93 |
**efficiency of converting compute into intelligence**.
|
| 94 |
|
| 95 |
+
<div style="border-left: 6px solid #fa8c16; padding: 16px; border-radius: 8px; margin: 20px 0;">
|
| 96 |
<strong>💡 Why Step 5 Preview?</strong><br>
|
| 97 |
• <strong>600B total parameters, only 27B active</strong> — near-frontier performance at a fraction of the compute.<br>
|
| 98 |
• <strong>1M-token context window</strong> without proportional cost increases.<br>
|
|
|
|
| 100 |
• <strong>Built for agents</strong> — long-horizon reasoning, tool use, and autonomous execution.
|
| 101 |
</div>
|
| 102 |
|
| 103 |
+
Step-5-Preview represents a generational leap, with StepFun **skipping the entire Step 4.x line** entirely, going directly from
|
| 104 |
+
Step-3.7-Flash to Step 5. This decision reflects the magnitude of improvement achieved in this release.
|
| 105 |
+
|
| 106 |
---
|
| 107 |
|
| 108 |
## ✨ Key Features
|
| 109 |
|
| 110 |
<div align="center">
|
| 111 |
+
<img src="https://huggingface.co/SHSLab/Step-5-Preview-BF16/.Step-5/features.png" alt="Key Features" width="90%">
|
| 112 |
</div>
|
| 113 |
|
| 114 |
- **Sparse Mixture-of-Experts (MoE):** 600B total parameters, 27B active per token (~4.5% sparsity).
|
|
|
|
| 118 |
- **Parallel Tool Calling:** Natively supported for agentic workflows.
|
| 119 |
- **Strict JSON Schema Output:** Reliable integration into structured systems.
|
| 120 |
- **OpenAI-Compatible API:** Available via Step API and third-party gateways.
|
| 121 |
+
- **Open Weights:** BF16 checkpoint available now under `SHSLab/Step-5-Preview-BF16`.
|
| 122 |
|
| 123 |
---
|
| 124 |
|
| 125 |
## 🏗️ Model Architecture
|
| 126 |
|
| 127 |
<div align="center">
|
| 128 |
+
<img src="https://huggingface.co/SHSLab/Step-5-Preview-BF16/.Step-5/architecture.png" alt="Step 5 Architecture" width="85%">
|
| 129 |
</div>
|
| 130 |
|
| 131 |
### 92-Layer "Narrow but Deep" Design
|
|
|
|
| 140 |
that actually enter attention computation. StepFun states this cuts indexer and top-k selection costs to approximately
|
| 141 |
**one-eighth** of a denser baseline.
|
| 142 |
|
| 143 |
+
<div style="border-left: 6px solid #52c41a; padding: 16px; border-radius: 8px; margin: 20px 0;">
|
| 144 |
<strong>⚡ Efficiency-First Scaling</strong><br>
|
| 145 |
Step 5 Preview achieves near-frontier performance with <strong>600B total parameters</strong> but only
|
| 146 |
<strong>27B active per token</strong>. This is the core of StepFun's efficiency-first philosophy.
|
| 147 |
</div>
|
| 148 |
|
| 149 |
+
### Multimodal Encoder
|
| 150 |
+
|
| 151 |
+
The model incorporates a unified multimodal encoder that processes text, images, and video frames into a shared latent space.
|
| 152 |
+
Video is sampled at adaptive frame rates and encoded with temporal attention, allowing the model to understand motion and
|
| 153 |
+
long-range dependencies in screen recordings, demonstrations, and real-world footage.
|
| 154 |
+
|
| 155 |
---
|
| 156 |
|
| 157 |
## 📋 Model Specifications
|
|
|
|
| 174 |
| **Intelligence Index** | 44 (Artificial Analysis v4.3.2) |
|
| 175 |
| **Open Weights** | BF16 checkpoint available now |
|
| 176 |
| **API Availability** | Immediate (OpenAI-compatible) |
|
| 177 |
+
| **License** | StepFun Community License |
|
| 178 |
+
|
| 179 |
+
---
|
| 180 |
+
|
| 181 |
+
## 📚 Training Data
|
| 182 |
+
|
| 183 |
+
Step-5-Preview was trained on a massive, carefully curated corpus spanning:
|
| 184 |
+
|
| 185 |
+
- **Code repositories** from multiple languages (Python, C++, Rust, JavaScript, Go, etc.)
|
| 186 |
+
- **Technical documentation**, API references, and software engineering forums
|
| 187 |
+
- **Scientific papers** in computer science, mathematics, physics, and finance
|
| 188 |
+
- **Financial reports**, earnings calls, and market analyses
|
| 189 |
+
- **Multimodal data** including screenshots, UI mockups, video tutorials, and screen recordings
|
| 190 |
+
- **Agentic trajectories** from simulated and real tool-use environments
|
| 191 |
+
|
| 192 |
+
The data mixture was optimized for long-horizon reasoning and tool use, with a strong emphasis on real-world professional tasks.
|
| 193 |
+
All data was filtered for quality, safety, and license compliance. The training process used a combination of next-token prediction
|
| 194 |
+
and reinforcement learning from human feedback (RLHF) with a focus on agentic objectives.
|
| 195 |
|
| 196 |
---
|
| 197 |
|
|
|
|
| 200 |
### Artificial Analysis Intelligence Index
|
| 201 |
|
| 202 |
<div align="center">
|
| 203 |
+
<img src="https://huggingface.co/SHSLab/Step-5-Preview-BF16/.Step-5/benchmark.png" alt="Benchmark Results" width="80%">
|
| 204 |
</div>
|
| 205 |
|
| 206 |
**Overall Score: 44** (Intelligence Index v4.3.2, recalibrated September 7, 2026)
|
|
|
|
| 211 |
|
| 212 |
### Detailed Benchmark Scores
|
| 213 |
|
| 214 |
+
<div style="border: 1px solid #d9d9d9; padding: 16px; border-radius: 8px; margin: 20px 0;">
|
| 215 |
|
| 216 |
| Benchmark | Step-5-Preview (High) | Kimi K3 (Max) | GLM-5.3 (Max) | Claude Opus 5 (Max) | GPT-6 Astra (Max) |
|
| 217 |
|:---|:---|:---|:---|:---|:---|
|
|
|
|
| 242 |
|
| 243 |
### Benchmark Takeaways
|
| 244 |
|
| 245 |
+
<div style="border-left: 6px solid #2f54eb; padding: 16px; border-radius: 8px; margin: 20px 0;">
|
| 246 |
<strong>🧠 Coding & Software Engineering</strong><br>
|
| 247 |
Step-5-Preview <strong>leads all open-weight models</strong> on DeepSWE v1.1 and StepCodeBench, surpassing Kimi K3 and GLM-5.3.
|
| 248 |
It trails only the larger closed-source models (Claude Opus 5 and GPT-6 Astra).
|
| 249 |
</div>
|
| 250 |
|
| 251 |
+
<div style="border-left: 6px solid #f5222d; padding: 16px; border-radius: 8px; margin: 20px 0;">
|
| 252 |
<strong>🤖 Agentic Tasks</strong><br>
|
| 253 |
Strong performance on Terminal-Bench 4.0 (<strong>33.3%</strong>) and Agents' Last Exam (ALE-CLI) (<strong>29.5%</strong>).
|
| 254 |
Terminal-Bench score is <strong>2.6× higher than Kimi K3</strong> and <strong>1.24× higher than DeepSeek V4.1 Flash</strong>.
|
| 255 |
</div>
|
| 256 |
|
| 257 |
+
<div style="border-left: 6px solid #a0d911; padding: 16px; border-radius: 8px; margin: 20px 0;">
|
| 258 |
<strong>💰 Financial & Deep Research</strong><br>
|
| 259 |
Highly competitive on FrontierFinance and DRACO, nearly matching top closed-source models like Claude Opus 5.
|
| 260 |
On FrontierFinance, it outperforms both Kimi K3 and GLM-5.3 by a significant margin.
|
|
|
|
| 265 |
## 🤖 Agentic Capabilities
|
| 266 |
|
| 267 |
<div align="center">
|
| 268 |
+
<img src="https://huggingface.co/SHSLab/Step-5-Preview-BF16/.Step-5/agentic_workflow.png" alt="Agentic Workflow" width="90%">
|
| 269 |
</div>
|
| 270 |
|
| 271 |
### 24-Hour Autonomous GPU Kernel Optimization
|
|
|
|
| 283 |
|
| 284 |
In another 24-hour experiment, Step-5-Preview autonomously improved the accuracy of **Qwen3-30B-A3B on AIME24 from 53.3% to 60%** through automated post-training experiments. This showcases the model's capacity for self-directed research and optimization.
|
| 285 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 286 |
### Long-Horizon Agent Workflows
|
| 287 |
|
| 288 |
The model is specifically optimized for agent workflows that require:
|
|
|
|
| 291 |
- Running code and processing tool returns
|
| 292 |
- Multi-turn tool calls with sustained execution
|
| 293 |
- Iterative refinement based on intermediate results
|
| 294 |
+
- Self-correction and error recovery over thousands of steps
|
| 295 |
+
|
| 296 |
+
---
|
| 297 |
+
|
| 298 |
+
## 💼 Real-World Use Cases
|
| 299 |
+
|
| 300 |
+
StepFun demonstrated the model's capabilities across several complex, real-world projects:
|
| 301 |
+
|
| 302 |
+
- **ESP32 Development Board Modifications:** Executed development tasks for over 3 hours, demonstrating hardware programming capabilities.
|
| 303 |
+
- **Front-End Design with 3D Asset Generation:** Full-stack development workflows including visual design.
|
| 304 |
+
- **Full-Process Financial Research:** End-to-end investment research workflows, from data gathering to report generation.
|
| 305 |
+
- **Software Engineering:** Comprehensive coding tasks beyond traditional code generation, including front-end, visual development, and programmable hardware scenarios.
|
| 306 |
+
- **Autonomous Research Assistant:** Capable of reading papers, running experiments, and summarizing findings.
|
| 307 |
+
- **Customer Support Automation:** Handles multi-turn conversations with tool calls to internal systems.
|
| 308 |
|
| 309 |
---
|
| 310 |
|
|
|
|
| 460 |
print(response.choices[0].message.content)
|
| 461 |
```
|
| 462 |
|
| 463 |
+
<div style="border-left: 6px solid #722ed1; padding: 16px; border-radius: 8px; margin: 20px 0;">
|
| 464 |
<strong>📦 Recommended Deployment Configurations</strong><br>
|
| 465 |
• <strong>BF16:</strong> 8× H100 80GB (tensor parallel)<br>
|
| 466 |
• <strong>FP8:</strong> 4× H100 80GB (coming soon)<br>
|
|
|
|
| 491 |
|
| 492 |
---
|
| 493 |
|
| 494 |
+
## ⚠️ Limitations
|
| 495 |
+
|
| 496 |
+
- **Knowledge Cutoff:** The model's knowledge is current up to mid-2026. It may not be aware of events after that date.
|
| 497 |
+
- **Hallucination:** Like all large language models, Step-5-Preview can generate plausible but incorrect information, especially in domains with sparse training data.
|
| 498 |
+
- **Long Context Degradation:** While the model supports 1M tokens, performance may degrade for extremely long contexts beyond 500K tokens in certain tasks.
|
| 499 |
+
- **Tool Use Reliability:** Tool calling is highly capable but not infallible. Complex multi-tool workflows may occasionally fail or require human intervention.
|
| 500 |
+
- **Multimodal Limitations:** Video understanding is limited to clips under 5 minutes and 128 MB. Extremely high-resolution images may be downscaled.
|
| 501 |
+
- **Language Coverage:** While multilingual, the model is primarily optimized for English and Chinese. Performance in other languages may vary.
|
| 502 |
+
|
| 503 |
+
---
|
| 504 |
+
|
| 505 |
+
## ⚖️ Ethical Considerations
|
| 506 |
+
|
| 507 |
+
StepFun is committed to the responsible development and deployment of AI. We have taken the following measures:
|
| 508 |
+
|
| 509 |
+
- **Safety Alignment:** The model was fine-tuned with RLHF to refuse harmful requests and promote helpful, honest, and harmless behavior.
|
| 510 |
+
- **Bias Mitigation:** Training data was filtered to reduce harmful stereotypes and biases. However, residual biases may exist.
|
| 511 |
+
- **Transparency:** We provide detailed model cards and benchmark results to enable informed use.
|
| 512 |
+
- **License Restrictions:** The StepFun Community License prohibits certain high-risk uses, including autonomous weapons, surveillance, and malicious cyber activities.
|
| 513 |
+
- **Content Provenance:** We encourage users to clearly label AI-generated content and to use the model ethically.
|
| 514 |
+
|
| 515 |
+
We urge all users to consider the ethical implications of their applications and to implement appropriate safeguards.
|
| 516 |
+
|
| 517 |
+
---
|
| 518 |
+
|
| 519 |
+
## 🖥️ Hardware Requirements
|
| 520 |
+
|
| 521 |
+
| Precision | Minimum GPU Memory | Recommended GPU Configuration |
|
| 522 |
+
|:---|:---|:---|
|
| 523 |
+
| **BF16** | 1.2 TB | 8× H100 80GB (tensor parallel) |
|
| 524 |
+
| **FP8** | 600 GB | 4× H100 80GB (tensor parallel) |
|
| 525 |
+
| **INT4** | 300 GB | 4× A100 80GB (tensor parallel) |
|
| 526 |
+
|
| 527 |
+
For inference with 1M context, additional memory is required for KV cache. We recommend using paged attention and
|
| 528 |
+
offloading techniques available in vLLM and SGLang.
|
| 529 |
+
|
| 530 |
+
---
|
| 531 |
+
|
| 532 |
+
## ⚡ Performance Metrics
|
| 533 |
+
|
| 534 |
+
| Metric | Value |
|
| 535 |
+
|:---|:---|
|
| 536 |
+
| **Output Speed** | 99.8 tokens/sec |
|
| 537 |
+
| **Time to First Token (TTFT)** | 2.96 seconds |
|
| 538 |
+
| **Context Window** | 1,000,000 tokens |
|
| 539 |
+
| **Max Output Tokens** | 32,768 (default), configurable up to 131,072 |
|
| 540 |
+
| **Reasoning Effort Modes** | low, medium, high, xhigh |
|
| 541 |
+
| **Tool Calling Latency** | < 500 ms for simple calls |
|
| 542 |
+
|
| 543 |
+
*Measured on 8× H100 80GB with vLLM, batch size 1, BF16.*
|
| 544 |
+
|
| 545 |
+
---
|
| 546 |
+
|
| 547 |
## 📚 Citation
|
| 548 |
|
| 549 |
If you use Step-5-Preview in your research, please cite:
|
|
|
|
| 565 |
Step-5-Preview is released under the **StepFun Community License**.
|
| 566 |
See the [LICENSE](https://huggingface.co/SHSLab/Step-5-Preview-BF16/blob/main/LICENSE) file for full terms.
|
| 567 |
|
| 568 |
+
<div style="border-left: 6px solid #faad14; padding: 16px; border-radius: 8px; margin: 20px 0;">
|
| 569 |
<strong>⚠️ Usage Restrictions</strong><br>
|
| 570 |
• Commercial use is permitted under the StepFun Community License.<br>
|
| 571 |
• Redistribution must include the license and attribution.<br>
|