---
language:
- en
- zh
- multilingual
license: other
license_name: stepfun-community-license
license_link: https://huggingface.co/SHSLab/Step-5-Preview-BF16/blob/main/LICENSE
library_name: transformers
pipeline_tag: text-generation
tags:
- stepfun
- step-5
- moe
- mixture-of-experts
- agentic
- coding
- software-engineering
- long-context
- 1m-context
- multimodal
- text-generation
- image
- video
- sparse-attention
- gqa
- financial-analysis
- deep-research
- tool-calling
- parallel-tool-calling
- json-schema
---
# Step-5-Preview
[](https://huggingface.co/SHSLab)
[](https://github.com/stepfun-ai)
[](https://discord.gg/stepfun)
[](https://huggingface.co/SHSLab/Step-5-Preview-BF16/blob/main/LICENSE)
[]()
[]()
> **๐ฅ Step-5-Preview is now available.**
>
> Step-5-Preview is a 600B-parameter sparse Mixture-of-Experts model with 27B active parameters, a 1M-token context window, and native support for text, image, and video inputs.
>
> Weights are available on Hugging Face (`SHSLab/Step-5-Preview-BF16`). Available via Step API, or self-hosted with vLLM / SGLang.
---
## ๐ Table of Contents
- [Introduction](#-introduction)
- [Key Features](#-key-features)
- [Model Specifications](#-model-specifications)
- [Benchmark Results](#-benchmark-results)
- [Quickstart](#-quickstart)
- [Deployment](#-deployment)
- [License](#-license)
- [Contact](#-contact)
- [More details](#-more-details) *(architecture, training, agentic demos, evaluation, limitations)*
---
## ๐ Introduction
**Step-5-Preview** is StepFun's flagship foundation model, designed for **real-world agentic tasks**. It targets professional domains such as **AI coding, software engineering, professional knowledge work, and financial analysis**.
StepFun's core philosophy for Step 5 is the **"Pareto Frontier"** โ balancing intelligence against cost. While earlier scaling efforts traded more compute for stronger intelligence, the next phase focuses on improving the **efficiency of converting compute into intelligence**.
> **Why Step 5 Preview?**
>
> - 600B total parameters, 27B active per token โ near-frontier performance at a fraction of the compute.
> - 1M-token context window without proportional cost increases.
> - Competitive benchmark scores against models with 3โ5ร more parameters.
> - Built for agents: long-horizon reasoning, tool use, and autonomous execution.
Step-5-Preview skips the entire Step 4.x line, going directly from Step-3.7-Flash to Step 5.
---
## โจ Key Features
- **Sparse Mixture-of-Experts (MoE):** 600B total parameters, 27B active per token (~4.5% sparsity).
- **1M-Token Context Window:** Equivalent to ~1,500 A4 pages, enabled by Sparse GQA.
- **Multimodal Input:** Text, image, and video (MP4, QuickTime, Matroska; โค128 MB; โค5 min recommended).
- **Configurable Reasoning Effort:** `low`, `medium`, `high` / `xhigh`.
- **Parallel Tool Calling:** Natively supported for agentic workflows.
- **Strict JSON Schema Output:** Reliable integration into structured systems.
- **OpenAI-Compatible API:** Available via Step API and third-party gateways.
- **Open Weights:** BF16 checkpoint available under `SHSLab/Step-5-Preview-BF16`.
---
## ๐ Model Specifications
| Category | Specification |
|:---|:---|
| **Model Name** | Step-5-Preview |
| **Developer** | StepFun |
| **Architecture** | Sparse Mixture-of-Experts (MoE) |
| **Total Parameters** | 600B |
| **Active Parameters** | 27B per token (~4.5% sparsity) |
| **Layers** | 92 (narrow-deep Transformer) |
| **Context Window** | 1,000,000 tokens |
| **Attention** | Sparse GQA with block-wise token merging |
| **Input Modalities** | Text, Image, Video |
| **Output Modalities** | Text |
| **Video Formats** | MP4, QuickTime, Matroska (โค128 MB, โค5 min recommended) |
| **Reasoning Effort** | `low` / `medium` / `high` (`xhigh`) |
| **Tool Calling** | Parallel, strict JSON schema |
| **Intelligence Index** | 44 (Artificial Analysis v4.3.2) |
| **Open Weights** | BF16 checkpoint available |
| **API Availability** | Immediate (OpenAI-compatible) |
| **License** | StepFun Community License |
---
## ๐ Benchmark Results
### Artificial Analysis Intelligence Index
**Overall Score: 44** (Intelligence Index v4.3.2, recalibrated September 7, 2026)
The index covers 10 evaluations including AA-Briefcase, GDPval-AA v2, Terminal-Bench 4.0, SciCode, and Humanity's Last Exam.
> **How to read the tables below**
>
> - **Bold** marks the best score in the row across all six models.
> - **๐ฅ** flags the leader for that benchmark.
> - **โ** indicates the model was not evaluated or did not report a score.
> - Step-5-Preview results use the `high` reasoning-effort setting unless noted.
>
> **Comparison set:** Step-5-Preview ยท GPT-6 Astra (Max) ยท Fable 5.1 ยท Claude Opus 5 (Max) ยท Kimi K3 (Max) ยท GLM-5.3 (Max)
>
> **Availability:** Open-weight โ Step-5-Preview, Kimi K3, Qwen3.8 Max, GLM-5.3. Closed-source โ GPT-6 Astra, Claude Opus 5. Fable 5.1 โ not publicly stated.
---
### ๐ง Reasoning & Knowledge
| Benchmark | Step-5-Preview | GPT-6 Astra | Fable 5.1 | Claude Opus 5 | Kimi K3 | GLM-5.3 |
|:---|:---:|:---:|:---:|:---:|:---:|:---:|
| **GPQA Diamond** | 93.5% | **96.1%** ๐ฅ | 93.7% | 93.2% | 93.5% | 91.7% |
| **Humanity's Last Exam (HLE)** | 46.5% | 54.7% | **59.1%** ๐ฅ | 54.9% | 46.9% | 42.3% |
| **AA-LCR v1.1** | 88.3% | 80.7% | 85.3% | 79.3% | **88.7%** ๐ฅ | 79.7% |
| **CritPt** | 20.9% | **31.7%** ๐ฅ | 29.7% | 29.1% | 23.4% | 19.1% |
---
### ๐ป Coding & Software Engineering
| Benchmark | Step-5-Preview | GPT-6 Astra | Fable 5.1 | Claude Opus 5 | Kimi K3 | GLM-5.3 |
|:---|:---:|:---:|:---:|:---:|:---:|:---:|
| **DeepSWE v1.1** | 67.7% | **74.1%** ๐ฅ | 67.4% | 74.0% | 67.5% | 66.9% |
| **Terminal-Bench 2.1** | 85.0% | 88.4% | **91.4%** ๐ฅ | 89.1% | 85.0% | 83.9% |
| **Terminal-Bench 4** | 33.3% | **57.9%** ๐ฅ | 55.8% | 52.3% | 12.6% | 41.9% |
| **CyberGym** | **84.7%** ๐ฅ | โ | โ | โ | 80.0% | 84.5% |
| **SciCode** | 58.9% | 56.5% | **63.1%** ๐ฅ | 56.4% | 59.5% | 59.0% |
| **RoadmapBench** | 54.3% | โ | โ | **68.3%** ๐ฅ | 55.4% | 54.1% |
| **ProgramBench** | 80.5% | **85.4%** ๐ฅ | 82.7% | 82.3% | 77.8% | 72.0% |
| **SWE-Marathon v1.1** | 72.7% | 77.3% | 80.2% | **85.6%** ๐ฅ | 84.4% | 67.4% |
| **MLS-Bench-Lite** | 40.5% | โ | **50.3%** ๐ฅ | 49.8% | 48.3% | 37.3% |
| **SWE-Atlas-QnA** | 63.6% | 60.9% | โ | **66.0%** ๐ฅ | 59.5% | 59.6% |
| **SWE-Atlas-Test-writing** | 50.8% | 51.1% | โ | **60.3%** ๐ฅ | 50.4% | 50.4% |
| **StepCodeBench** | 49.0% | 61.0% | โ | **63.9%** ๐ฅ | 43.9% | 40.2% |
| **StepCode-Bench-Daily** | 64.9% | โ | โ | **77.6%** ๐ฅ | 57.7% | 69.1% |
| **StepCode-Bench-General** | 65.0% | 64.3% | โ | **68.3%** ๐ฅ | 65.2% | 62.0% |
---
### ๐ค Agents, Tool Use & Automation
| Benchmark | Step-5-Preview | GPT-6 Astra | Fable 5.1 | Claude Opus 5 | Kimi K3 | GLM-5.3 |
|:---|:---:|:---:|:---:|:---:|:---:|:---:|
| **GDPval-AA v2** | 1571 | 1580 | 1724 | **1735** ๐ฅ | 1548 | 1634 |
| **ฯยณ-Banking** | 42.5% | 41.4% | 47.2% | 42.1% | 46.0% | **50.3%** ๐ฅ |
| **AutomationBench-AA** | 51.0% | **68.5%** ๐ฅ | 59.4% | 56.6% | 58.3% | 62.2% |
| **AutomationBench Public** | 44.0% | โ | โ | โ | 46.7% | **48.2%** ๐ฅ |
| **AA-Briefcase** | 1417 | 1562 | **1662** ๐ฅ | 1645 | 1492 | 1511 |
| **Toolathlon-Verified** | 74.1% | โ | 77.8% | **80.6%** ๐ฅ | 76.5% | 73.0% |
| **MCP-Atlas** | 85.6% | โ | โ | **87.0%** ๐ฅ | 85.3% | 86.8% |
| **PresentBench** | 76.8% | โ | โ | **77.3%** ๐ฅ | 75.6% | 74.5% |
| **JobBench** | 59.0% | โ | โ | **65.7%** ๐ฅ | 54.3% | 61.4% |
| **Apex-Agents** | 37.8% | โ | โ | **41.8%** ๐ฅ | 41.0% | 38.1% |
| **Draco** | 83.3% | 76.8% | **87.7%** ๐ฅ | 87.6% | 78.5% | 82.3% |
| **BrowseComp** | 88.7% | **91.5%** ๐ฅ | โ | 90.2% | 91.2% | โ |
| **HLE w/ tools** | 59.4% | 57.2% | **65.0%** ๐ฅ | 63.6% | 56.0% | 62.5% |
---
### ๐ฐ Finance & Professional Work
| Benchmark | Step-5-Preview | GPT-6 Astra | Fable 5.1 | Claude Opus 5 | Kimi K3 | GLM-5.3 |
|:---|:---:|:---:|:---:|:---:|:---:|:---:|
| **FinStepBench-LiveSearch** | 74.5% | 74.5% | โ | **76.2%** ๐ฅ | 70.9% | 73.3% |
| **FinStepBench-CorporateValuation** | 60.6% | **77.3%** ๐ฅ | โ | 69.7% | 60.6% | 56.1% |
| **FinStepBench-FinanceDR** | 55.8% | 45.0% | โ | **59.1%** ๐ฅ | 48.9% | 53.3% |
| **FrontierFinance** | 66.4% | 55.0% | โ | **69.7%** ๐ฅ | 62.6% | 64.1% |
| **OfficeQA Pro** | 60.3% | **67.7%** ๐ฅ | โ | 64.7% | 62.6% | 59.1% |
| **SpeadSheet v2** | 29.4% | 31.4% | โ | **32.8%** ๐ฅ | 31.9% | 30.5% |
| **GDP.pdf** | 14.8% | **31.0%** ๐ฅ | 26.2% | 21.6% | 22.0% | 11.2% |
---
### ๐๏ธ Multimodal
| Benchmark | Step-5-Preview | GPT-6 Astra | Fable 5.1 | Claude Opus 5 | Kimi K3 | GLM-5.3 |
|:---|:---:|:---:|:---:|:---:|:---:|:---:|
| **MMMU-Pro** | 76.0% | **87.0%** ๐ฅ | โ | 85.0% | 81.0% | โ |
| **GDP.pdf** | 14.8% | **31.0%** ๐ฅ | 26.2% | 21.6% | 22.0% | 11.2% |
๐ Benchmark methodology notes
- **DeepSWE v1.1** was evaluated using the SWE-agent harness with `temperature=1.0` and `top_p=0.95`.
- **StepCodeBench** achieved **49.0% avg@4**.
- **GDPval-AA v2** results are from Artificial Analysis as of September 19, 2026.
- **AA-LCR v1.1**: Step-5-Preview scored **88.3%**, statistically tied with Kimi K3 (88.7%).
- **Terminal-Bench 4**: Step-5-Preview **33.3%** vs. Kimi K3 ~**12.6%** (2.6ร) and DeepSeek V4.1 Flash **26.8%** (1.24ร).
- **SciCode**: Step-5-Preview **58.9%**, above GPT-6 Astra (56.5%) and Claude Opus 5 (56.4%).
- **Multimodal**: MMMU-Pro and GDP.pdf were run with the unified multimodal encoder at default resolution.
- **Output Speed**: 99.8 tokens/sec (GLM-5.3: 72.1 tokens/sec).
- **Time to First Token**: 2.96 seconds (GLM-5.3: 2.99s; Claude Opus 5: 56.84s at max effort).
- Missing entries (**โ**) reflect benchmarks that were not publicly reported for that model at the time of writing.
---
## โก Quickstart
### Installation
```bash
pip install transformers>=4.56.0
pip install torch>=2.4.0
pip install accelerate
```
For video/image support:
```bash
pip install av pillow
```
### Basic Usage with Transformers
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "SHSLab/Step-5-Preview-BF16"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
device_map="auto",
torch_dtype="bfloat16",
)
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain the significance of the Pareto Frontier in AI scaling."},
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
return_tensors="pt",
).to(model.device)
outputs = model.generate(
inputs,
max_new_tokens=1024,
temperature=0.7,
top_p=0.95,
reasoning_effort="high", # low / medium / high / xhigh
)
response = tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True)
print(response)
```
### Multimodal (Image + Video) Usage
```python
from transformers import AutoProcessor
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
messages = [
{
"role": "user",
"content": [
{"type": "image", "url": "https://example.com/image.jpg"},
{"type": "video", "url": "https://example.com/video.mp4"},
{"type": "text", "text": "Describe the scene and summarize the video."},
],
}
]
inputs = processor.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt")
# ... generate as above
```
### Tool Calling
```python
tools = [
{
"type": "function",
"function": {
"name": "get_weather",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}
]
messages = [{"role": "user", "content": "What's the weather in Tokyo?"}]
inputs = tokenizer.apply_chat_template(
messages,
tools=tools,
add_generation_prompt=True,
return_tensors="pt",
).to(model.device)
outputs = model.generate(inputs, max_new_tokens=256, reasoning_effort="medium")
print(tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True))
```
---
## ๐ข Deployment
### vLLM
```bash
vllm serve SHSLab/Step-5-Preview-BF16 \
--trust-remote-code \
--tensor-parallel-size 8 \
--max-model-len 1000000 \
--enable-reasoning \
--reasoning-parser stepfun
```
### SGLang
```bash
python -m sglang.launch_server \
--model-path SHSLab/Step-5-Preview-BF16 \
--trust-remote-code \
--tp 8 \
--context-length 1000000 \
--reasoning-parser stepfun
```
### OpenAI-Compatible API
```python
from openai import OpenAI
client = OpenAI(
api_key="YOUR_STEP_API_KEY",
base_url="https://api.stepfun.com/v1",
)
response = client.chat.completions.create(
model="step-5-preview",
messages=[{"role": "user", "content": "Write a Python function to merge two sorted lists."}],
reasoning_effort="high",
max_tokens=2048,
)
print(response.choices[0].message.content)
```
> **Recommended deployment configurations**
>
> - **BF16:** 8ร H100 80GB (tensor parallel)
> - **FP8:** 4ร H100 80GB (coming soon)
> - **Context length:** Up to 1M tokens
> - **Reasoning parser:** Use `stepfun` for vLLM / SGLang
---
## ๐ License
Step-5-Preview is released under the **StepFun Community License**. See the [LICENSE](https://huggingface.co/SHSLab/Step-5-Preview-BF16/blob/main/LICENSE) file for full terms.
> **Usage restrictions**
>
> - Commercial use is permitted under the StepFun Community License.
> - Redistribution must include the license and attribution.
> - See LICENSE for full details.
---
## ๐ฌ Contact
- **Hugging Face:** [SHSLab](https://huggingface.co/SHSLab)
- **GitHub:** [github.com/stepfun-ai](https://github.com/stepfun-ai)
- **Discord:** [Join our Discord](https://discord.gg/stepfun)
- **Email:** [opensource@stepfun.com](mailto:opensource@stepfun.com)
- **Website:** [stepfun.com](https://stepfun.com)
---
## ๐ More details
๐๏ธ Model Architecture
### 92-Layer "Narrow but Deep" Design
Step-5-Preview uses a **92-layer Transformer** with a narrow-deep configuration. This design creates longer information propagation paths for implicit multi-hop reasoning during long prefill operations.
### Sparse Grouped-Query Attention (GQA) with Block-Wise Token Merging
To handle the 1M-token context window efficiently, Step-5-Preview uses **Sparse GQA with block-wise token merging**. This mechanism uses sparse indexing to select only historical information relevant to the current task, reducing the number of tokens that enter attention computation. This cuts indexer and top-k selection costs to approximately one-eighth of a denser baseline.
### Multimodal Encoder
A unified multimodal encoder processes text, images, and video frames into a shared latent space. Video is sampled at adaptive frame rates and encoded with temporal attention, allowing the model to understand motion and long-range dependencies in screen recordings, demonstrations, and real-world footage.
๐ Training Data
Step-5-Preview was trained on a carefully curated corpus spanning:
- **Code repositories** from multiple languages (Python, C++, Rust, JavaScript, Go, etc.)
- **Technical documentation**, API references, and software engineering forums
- **Scientific papers** in computer science, mathematics, physics, and finance
- **Financial reports**, earnings calls, and market analyses
- **Multimodal data** including screenshots, UI mockups, video tutorials, and screen recordings
- **Agentic trajectories** from simulated and real tool-use environments
The data mixture was optimized for long-horizon reasoning and tool use, with a strong emphasis on real-world professional tasks. All data was filtered for quality, safety, and license compliance. Training combined next-token prediction with reinforcement learning from human feedback (RLHF) focused on agentic objectives.
๐ค Agentic Capabilities
### 24-Hour Autonomous GPU Kernel Optimization
Step-5-Preview was tasked with autonomously optimizing an H100 GPU kernel for up to 24 consecutive hours. The model:
- Independently modified code
- Ran tests and compared results
- Iterated based on performance outcomes
- **Reached 508 TFLOPS after approximately 22 hours**
For comparison, Claude Opus 5 achieved 493 TFLOPS in the same experiment.
### Automated Post-Training Experiments
In another 24-hour experiment, Step-5-Preview autonomously improved the accuracy of Qwen3-30B-A3B on AIME24 from 53.3% to 60% through automated post-training experiments.
### Long-Horizon Agent Workflows
Optimized for workflows that require:
- Searching and information retrieval
- Running code and processing tool returns
- Multi-turn tool calls with sustained execution
- Iterative refinement based on intermediate results
- Self-correction and error recovery over thousands of steps
๐ผ Real-World Use Cases
- **ESP32 Development Board Modifications:** Executed development tasks for over 3 hours.
- **Front-End Design with 3D Asset Generation:** Full-stack development workflows including visual design.
- **Full-Process Financial Research:** End-to-end investment research, from data gathering to report generation.
- **Software Engineering:** Comprehensive coding tasks including front-end, visual development, and programmable hardware.
- **Autonomous Research Assistant:** Reading papers, running experiments, and summarizing findings.
- **Customer Support Automation:** Multi-turn conversations with tool calls to internal systems.
๐ Full Evaluation Table
All evaluations used the `high` reasoning effort setting unless otherwise noted.
| Benchmark | Score | Notes |
|:---|:---|:---|
| **DeepSWE v1.1** | 67.7% | SWE-agent harness, temp=1.0, top_p=0.95 |
| **StepCodeBench** | 49.0% | avg@4 |
| **ProgramBench** | 80.5% | Rebuild programs from binary + docs; 200 tasks, 248K+ behavioral tests |
| **Terminal-Bench 2.1** | 85.0% | Verified refresh of TB 2.0; 89 curated terminal tasks across SWE, sysadmin, data processing |
| **Terminal-Bench 4** | 33.3% | 2.6ร Kimi K3 |
| **CyberGym** | 84.7% | Best in comparison set |
| **SciCode** | 58.9% | Above GPT-6 Astra (56.5%) and Opus 5 (56.4%) |
| **RoadmapBench** | 54.3% | 115 long-horizon coding tasks across 17 repos, 5 languages; median ~3,700 LOC changed |
| **SWE-Marathon v1.1** | 72.7% | Ultra-long-horizon SWE; 20 realistic multi-hour tasks with hidden/adversarial tests |
| **MLS-Bench-Lite** | 40.5% | 30-task subset of MLS-Bench; tests inventing generalizable ML methods across 12 research domains |
| **SWE-Atlas-QnA** | 63.6% | 124 tasks on deep code comprehension โ tracing execution paths, explaining architecture across production repos |
| **SWE-Atlas-Test-writing** | 50.8% | 90 tasks on writing production-grade unit, integration, and acceptance tests for real repositories |
| **StepCode-Bench-Daily** | 64.9% | StepFun internal; 553 repos, 9 task types, 20 domains, 33 languages; daily-difficulty slice |
| **StepCode-Bench-General** | 65.0% | StepFun internal; same corpus as StepCodeBench; general-difficulty slice |
| **Agents' Last Exam (ALE-CLI)** | 29.5% | Linux-only CLI subset of ALE; 40 industry subfields; best agent pass rate ~25.2% |
| **GDPval-AA v2** | 1571 | Artificial Analysis, Sep 19, 2026 |
| **AA-Briefcase** | 1417 | Elo; agentic knowledge work across data science, product, banking, heavy industry; private held-out test set |
| **Toolathlon-Verified** | 74.1% | 108 expert-authored tasks; multi-app workflows averaging ~20 turns; strictly verifiable via scripts |
| **MCP-Atlas** | 85.6% | 1,000 tasks (500 public + 500 private); 36 real MCP servers, 220 tools; 3โ6 tool calls/task |
| **PresentBench** | 76.8% | 238 slide-generation instances; avg 54.1 rubric checklist items per instance |
| **JobBench** | 59.0% | 130 tasks across 35 occupations; avg 35.6 binary rubric criteria per task |
| **Apex-Agents** | 37.8% | Long-horizon tasks in investment banking, consulting, corporate law |
| **DRACO** | 83.3% | Cross-domain deep research; accuracy, completeness, objectivity, citation quality |
| **BrowseComp** | 88.7% | 1,266 hard-to-find web information retrieval questions |
| **HLE w/ tools** | 59.4% | +12.9 pts over no-tools HLE |
| **FrontierFinance** | 66.4% | +11.4 pts over GPT-6 Astra |
| **FinStepBench-LiveSearch** | 74.5% | Tied with GPT-6 Astra |
| **FinStepBench-CorporateValuation** | 60.6% | Tied with Kimi K3 |
| **FinStepBench-FinanceDR** | 55.8% | StepFun internal; full deep-research finance workflows |
| **OfficeQA Pro** | 60.3% | Databricks; 90 questions over large enterprise financial document collections |
| **SpeadSheet v2** | 29.4% | SpreadsheetBench 2; end-to-end business spreadsheet workflows; best model โ34.89% |
| **GDP.pdf** | 14.8% | Known weak spot |
| **GPQA Diamond** | 93.5% | 198 PhD-level science MCQs; human expert avg 81% |
| **HLE** | 46.5% | 59.4% with tools |
| **AA-LCR v1.1** | 88.3% | Tied with Kimi K3 (88.7%) |
| **CritPt** | 20.9% | 71 unpublished research-level physics challenges |
| **MMMU-Pro** | 76.0% | Robust multimodal benchmark; significantly harder than MMMU |
| **Output Speed** | 99.8 tokens/sec | GLM-5.3: 72.1 tokens/sec |
| **Time to First Token** | 2.96s | GLM-5.3: 2.99s; Claude Opus 5: 56.84s (max effort) |
๐
Category summary
| Domain | Standing | Highlight |
|:---|:---|:---|
| **Long-Context Reasoning** | Near-tied for #1 | AA-LCR v1.1 **88.3%** (+7.6 over GPT-6 Astra, +9.0 over Opus 5) |
| **Coding / SWE** | #1 open-weight | DeepSWE v1.1 **67.7%**, StepCodeBench **49.0%**, ProgramBench **80.5%** |
| **Security / Cyber** | #1 in comparison set | CyberGym **84.7%** |
| **Finance** | #2 overall, #1 open-weight | FrontierFinance **66.4%**, FinanceDR **55.8%** |
| **Terminal Agents** | #2 open-weight | Terminal-Bench 4 **33.3%** (2.6ร Kimi K3) |
| **Tool Use** | Top-3 open-weight | MCP-Atlas **85.6%**, HLE w/ tools **59.4%** |
| **General Knowledge** | Top tier | GPQA Diamond **93.5%**, HLE **46.5%** |
| **Multimodal** | Trailing | MMMU-Pro **76.0%** โ targeted for improvement |
โ ๏ธ Limitations
- **Knowledge Cutoff:** Knowledge is current up to mid-2026. May not be aware of later events.
- **Hallucination:** Can generate plausible but incorrect information, especially in domains with sparse training data.
- **Long Context Degradation:** Performance may degrade for extremely long contexts beyond 500K tokens in certain tasks.
- **Tool Use Reliability:** Tool calling is capable but not infallible. Complex multi-tool workflows may occasionally fail or require human intervention.
- **Multimodal Limitations:** Video understanding is limited to clips under 5 minutes and 128 MB. Extremely high-resolution images may be downscaled.
- **Language Coverage:** Primarily optimized for English and Chinese. Performance in other languages may vary.
โ๏ธ Ethical Considerations
StepFun is committed to the responsible development and deployment of AI. Measures taken:
- **Safety Alignment:** Fine-tuned with RLHF to refuse harmful requests and promote helpful, honest, and harmless behavior.
- **Bias Mitigation:** Training data was filtered to reduce harmful stereotypes and biases. Residual biases may exist.
- **Transparency:** Detailed model cards and benchmark results are provided to enable informed use.
- **License Restrictions:** The StepFun Community License prohibits certain high-risk uses, including autonomous weapons, surveillance, and malicious cyber activities.
- **Content Provenance:** Users are encouraged to clearly label AI-generated content.
All users should consider the ethical implications of their applications and implement appropriate safeguards.
๐ฅ๏ธ Hardware Requirements
| Precision | Minimum GPU Memory | Recommended GPU Configuration |
|:---|:---|:---|
| **BF16** | 1.2 TB | 8ร H100 80GB (tensor parallel) |
| **FP8** | 600 GB | 4ร H100 80GB (tensor parallel) |
| **INT4** | 300 GB | 4ร A100 80GB (tensor parallel) |
For inference with 1M context, additional memory is required for KV cache. Paged attention and offloading techniques available in vLLM and SGLang are recommended.
โก Performance Metrics
| Metric | Value |
|:---|:---|
| **Output Speed** | 99.8 tokens/sec |
| **Time to First Token (TTFT)** | 2.96 seconds |
| **Context Window** | 1,000,000 tokens |
| **Max Output Tokens** | 32,768 (default), configurable up to 131,072 |
| **Reasoning Effort Modes** | low, medium, high, xhigh |
| **Tool Calling Latency** | < 500 ms for simple calls |
*Measured on 8ร H100 80GB with vLLM, batch size 1, BF16.*
๐ Citation
```bibtex
@misc{stepfun2026step5preview,
title = {Step-5-Preview: A 600B Sparse MoE Foundation Model for Real-World Agentic Work},
author = {StepFun Team},
year = {2026},
howpublished = {\url{https://huggingface.co/SHSLab/Step-5-Preview-BF16}},
note = {Released September 20, 2026}
}
```
---
โญ If you find Step-5-Preview useful, please give us a star on GitHub and Hugging Face! โญ
Built with โค๏ธ by StepFun