--- language: - en - zh - multilingual license: other license_name: stepfun-community-license license_link: https://huggingface.co/SHSLab/Step-5-Preview-BF16/blob/main/LICENSE library_name: transformers pipeline_tag: text-generation tags: - stepfun - step-5 - moe - mixture-of-experts - agentic - coding - software-engineering - long-context - 1m-context - multimodal - text-generation - image - video - sparse-attention - gqa - financial-analysis - deep-research - tool-calling - parallel-tool-calling - json-schema --- # Step-5-Preview
Step 5 Preview Banner
[![Hugging Face](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-SHSLab-yellow)](https://huggingface.co/SHSLab) [![GitHub](https://img.shields.io/badge/GitHub-StepFun-181717?logo=github)](https://github.com/stepfun-ai) [![Discord](https://img.shields.io/badge/Discord-Join%20Us-5865F2?logo=discord)](https://discord.gg/stepfun) [![License](https://img.shields.io/badge/License-StepFun%20Community-blue)](https://huggingface.co/SHSLab/Step-5-Preview-BF16/blob/main/LICENSE) [![Model Size](https://img.shields.io/badge/Parameters-600B%20Total%20%7C%2027B%20Active-orange)]() [![Context](https://img.shields.io/badge/Context-1M%20Tokens-green)]()
> **๐Ÿ”ฅ Step-5-Preview is now available.** > > Step-5-Preview is a 600B-parameter sparse Mixture-of-Experts model with 27B active parameters, a 1M-token context window, and native support for text, image, and video inputs. > > Weights are available on Hugging Face (`SHSLab/Step-5-Preview-BF16`). Available via Step API, or self-hosted with vLLM / SGLang. --- ## ๐Ÿ“– Table of Contents - [Introduction](#-introduction) - [Key Features](#-key-features) - [Model Specifications](#-model-specifications) - [Benchmark Results](#-benchmark-results) - [Quickstart](#-quickstart) - [Deployment](#-deployment) - [License](#-license) - [Contact](#-contact) - [More details](#-more-details) *(architecture, training, agentic demos, evaluation, limitations)* --- ## ๐Ÿš€ Introduction **Step-5-Preview** is StepFun's flagship foundation model, designed for **real-world agentic tasks**. It targets professional domains such as **AI coding, software engineering, professional knowledge work, and financial analysis**. StepFun's core philosophy for Step 5 is the **"Pareto Frontier"** โ€” balancing intelligence against cost. While earlier scaling efforts traded more compute for stronger intelligence, the next phase focuses on improving the **efficiency of converting compute into intelligence**. > **Why Step 5 Preview?** > > - 600B total parameters, 27B active per token โ€” near-frontier performance at a fraction of the compute. > - 1M-token context window without proportional cost increases. > - Competitive benchmark scores against models with 3โ€“5ร— more parameters. > - Built for agents: long-horizon reasoning, tool use, and autonomous execution. Step-5-Preview skips the entire Step 4.x line, going directly from Step-3.7-Flash to Step 5. --- ## โœจ Key Features
Key Features
- **Sparse Mixture-of-Experts (MoE):** 600B total parameters, 27B active per token (~4.5% sparsity). - **1M-Token Context Window:** Equivalent to ~1,500 A4 pages, enabled by Sparse GQA. - **Multimodal Input:** Text, image, and video (MP4, QuickTime, Matroska; โ‰ค128 MB; โ‰ค5 min recommended). - **Configurable Reasoning Effort:** `low`, `medium`, `high` / `xhigh`. - **Parallel Tool Calling:** Natively supported for agentic workflows. - **Strict JSON Schema Output:** Reliable integration into structured systems. - **OpenAI-Compatible API:** Available via Step API and third-party gateways. - **Open Weights:** BF16 checkpoint available under `SHSLab/Step-5-Preview-BF16`. --- ## ๐Ÿ“‹ Model Specifications | Category | Specification | |:---|:---| | **Model Name** | Step-5-Preview | | **Developer** | StepFun | | **Architecture** | Sparse Mixture-of-Experts (MoE) | | **Total Parameters** | 600B | | **Active Parameters** | 27B per token (~4.5% sparsity) | | **Layers** | 92 (narrow-deep Transformer) | | **Context Window** | 1,000,000 tokens | | **Attention** | Sparse GQA with block-wise token merging | | **Input Modalities** | Text, Image, Video | | **Output Modalities** | Text | | **Video Formats** | MP4, QuickTime, Matroska (โ‰ค128 MB, โ‰ค5 min recommended) | | **Reasoning Effort** | `low` / `medium` / `high` (`xhigh`) | | **Tool Calling** | Parallel, strict JSON schema | | **Intelligence Index** | 44 (Artificial Analysis v4.3.2) | | **Open Weights** | BF16 checkpoint available | | **API Availability** | Immediate (OpenAI-compatible) | | **License** | StepFun Community License | --- ## ๐Ÿ“Š Benchmark Results
Benchmark Results
### Artificial Analysis Intelligence Index **Overall Score: 44** (Intelligence Index v4.3.2, recalibrated September 7, 2026) The index covers 10 evaluations including AA-Briefcase, GDPval-AA v2, Terminal-Bench 4.0, SciCode, and Humanity's Last Exam. > **How to read the tables below** > > - **Bold** marks the best score in the row across all six models. > - **๐Ÿฅ‡** flags the leader for that benchmark. > - **โ€”** indicates the model was not evaluated or did not report a score. > - Step-5-Preview results use the `high` reasoning-effort setting unless noted. > > **Comparison set:** Step-5-Preview ยท GPT-6 Astra (Max) ยท Fable 5.1 ยท Claude Opus 5 (Max) ยท Kimi K3 (Max) ยท GLM-5.3 (Max) > > **Availability:** Open-weight โ€” Step-5-Preview, Kimi K3, Qwen3.8 Max, GLM-5.3. Closed-source โ€” GPT-6 Astra, Claude Opus 5. Fable 5.1 โ€” not publicly stated. --- ### ๐Ÿง  Reasoning & Knowledge | Benchmark | Step-5-Preview | GPT-6 Astra | Fable 5.1 | Claude Opus 5 | Kimi K3 | GLM-5.3 | |:---|:---:|:---:|:---:|:---:|:---:|:---:| | **GPQA Diamond** | 93.5% | **96.1%** ๐Ÿฅ‡ | 93.7% | 93.2% | 93.5% | 91.7% | | **Humanity's Last Exam (HLE)** | 46.5% | 54.7% | **59.1%** ๐Ÿฅ‡ | 54.9% | 46.9% | 42.3% | | **AA-LCR v1.1** | 88.3% | 80.7% | 85.3% | 79.3% | **88.7%** ๐Ÿฅ‡ | 79.7% | | **CritPt** | 20.9% | **31.7%** ๐Ÿฅ‡ | 29.7% | 29.1% | 23.4% | 19.1% | --- ### ๐Ÿ’ป Coding & Software Engineering | Benchmark | Step-5-Preview | GPT-6 Astra | Fable 5.1 | Claude Opus 5 | Kimi K3 | GLM-5.3 | |:---|:---:|:---:|:---:|:---:|:---:|:---:| | **DeepSWE v1.1** | 67.7% | **74.1%** ๐Ÿฅ‡ | 67.4% | 74.0% | 67.5% | 66.9% | | **Terminal-Bench 2.1** | 85.0% | 88.4% | **91.4%** ๐Ÿฅ‡ | 89.1% | 85.0% | 83.9% | | **Terminal-Bench 4** | 33.3% | **57.9%** ๐Ÿฅ‡ | 55.8% | 52.3% | 12.6% | 41.9% | | **CyberGym** | **84.7%** ๐Ÿฅ‡ | โ€” | โ€” | โ€” | 80.0% | 84.5% | | **SciCode** | 58.9% | 56.5% | **63.1%** ๐Ÿฅ‡ | 56.4% | 59.5% | 59.0% | | **RoadmapBench** | 54.3% | โ€” | โ€” | **68.3%** ๐Ÿฅ‡ | 55.4% | 54.1% | | **ProgramBench** | 80.5% | **85.4%** ๐Ÿฅ‡ | 82.7% | 82.3% | 77.8% | 72.0% | | **SWE-Marathon v1.1** | 72.7% | 77.3% | 80.2% | **85.6%** ๐Ÿฅ‡ | 84.4% | 67.4% | | **MLS-Bench-Lite** | 40.5% | โ€” | **50.3%** ๐Ÿฅ‡ | 49.8% | 48.3% | 37.3% | | **SWE-Atlas-QnA** | 63.6% | 60.9% | โ€” | **66.0%** ๐Ÿฅ‡ | 59.5% | 59.6% | | **SWE-Atlas-Test-writing** | 50.8% | 51.1% | โ€” | **60.3%** ๐Ÿฅ‡ | 50.4% | 50.4% | | **StepCodeBench** | 49.0% | 61.0% | โ€” | **63.9%** ๐Ÿฅ‡ | 43.9% | 40.2% | | **StepCode-Bench-Daily** | 64.9% | โ€” | โ€” | **77.6%** ๐Ÿฅ‡ | 57.7% | 69.1% | | **StepCode-Bench-General** | 65.0% | 64.3% | โ€” | **68.3%** ๐Ÿฅ‡ | 65.2% | 62.0% | --- ### ๐Ÿค– Agents, Tool Use & Automation | Benchmark | Step-5-Preview | GPT-6 Astra | Fable 5.1 | Claude Opus 5 | Kimi K3 | GLM-5.3 | |:---|:---:|:---:|:---:|:---:|:---:|:---:| | **GDPval-AA v2** | 1571 | 1580 | 1724 | **1735** ๐Ÿฅ‡ | 1548 | 1634 | | **ฯ„ยณ-Banking** | 42.5% | 41.4% | 47.2% | 42.1% | 46.0% | **50.3%** ๐Ÿฅ‡ | | **AutomationBench-AA** | 51.0% | **68.5%** ๐Ÿฅ‡ | 59.4% | 56.6% | 58.3% | 62.2% | | **AutomationBench Public** | 44.0% | โ€” | โ€” | โ€” | 46.7% | **48.2%** ๐Ÿฅ‡ | | **AA-Briefcase** | 1417 | 1562 | **1662** ๐Ÿฅ‡ | 1645 | 1492 | 1511 | | **Toolathlon-Verified** | 74.1% | โ€” | 77.8% | **80.6%** ๐Ÿฅ‡ | 76.5% | 73.0% | | **MCP-Atlas** | 85.6% | โ€” | โ€” | **87.0%** ๐Ÿฅ‡ | 85.3% | 86.8% | | **PresentBench** | 76.8% | โ€” | โ€” | **77.3%** ๐Ÿฅ‡ | 75.6% | 74.5% | | **JobBench** | 59.0% | โ€” | โ€” | **65.7%** ๐Ÿฅ‡ | 54.3% | 61.4% | | **Apex-Agents** | 37.8% | โ€” | โ€” | **41.8%** ๐Ÿฅ‡ | 41.0% | 38.1% | | **Draco** | 83.3% | 76.8% | **87.7%** ๐Ÿฅ‡ | 87.6% | 78.5% | 82.3% | | **BrowseComp** | 88.7% | **91.5%** ๐Ÿฅ‡ | โ€” | 90.2% | 91.2% | โ€” | | **HLE w/ tools** | 59.4% | 57.2% | **65.0%** ๐Ÿฅ‡ | 63.6% | 56.0% | 62.5% | --- ### ๐Ÿ’ฐ Finance & Professional Work | Benchmark | Step-5-Preview | GPT-6 Astra | Fable 5.1 | Claude Opus 5 | Kimi K3 | GLM-5.3 | |:---|:---:|:---:|:---:|:---:|:---:|:---:| | **FinStepBench-LiveSearch** | 74.5% | 74.5% | โ€” | **76.2%** ๐Ÿฅ‡ | 70.9% | 73.3% | | **FinStepBench-CorporateValuation** | 60.6% | **77.3%** ๐Ÿฅ‡ | โ€” | 69.7% | 60.6% | 56.1% | | **FinStepBench-FinanceDR** | 55.8% | 45.0% | โ€” | **59.1%** ๐Ÿฅ‡ | 48.9% | 53.3% | | **FrontierFinance** | 66.4% | 55.0% | โ€” | **69.7%** ๐Ÿฅ‡ | 62.6% | 64.1% | | **OfficeQA Pro** | 60.3% | **67.7%** ๐Ÿฅ‡ | โ€” | 64.7% | 62.6% | 59.1% | | **SpeadSheet v2** | 29.4% | 31.4% | โ€” | **32.8%** ๐Ÿฅ‡ | 31.9% | 30.5% | | **GDP.pdf** | 14.8% | **31.0%** ๐Ÿฅ‡ | 26.2% | 21.6% | 22.0% | 11.2% | --- ### ๐Ÿ‘๏ธ Multimodal | Benchmark | Step-5-Preview | GPT-6 Astra | Fable 5.1 | Claude Opus 5 | Kimi K3 | GLM-5.3 | |:---|:---:|:---:|:---:|:---:|:---:|:---:| | **MMMU-Pro** | 76.0% | **87.0%** ๐Ÿฅ‡ | โ€” | 85.0% | 81.0% | โ€” | | **GDP.pdf** | 14.8% | **31.0%** ๐Ÿฅ‡ | 26.2% | 21.6% | 22.0% | 11.2% |
๐Ÿ“ Benchmark methodology notes - **DeepSWE v1.1** was evaluated using the SWE-agent harness with `temperature=1.0` and `top_p=0.95`. - **StepCodeBench** achieved **49.0% avg@4**. - **GDPval-AA v2** results are from Artificial Analysis as of September 19, 2026. - **AA-LCR v1.1**: Step-5-Preview scored **88.3%**, statistically tied with Kimi K3 (88.7%). - **Terminal-Bench 4**: Step-5-Preview **33.3%** vs. Kimi K3 ~**12.6%** (2.6ร—) and DeepSeek V4.1 Flash **26.8%** (1.24ร—). - **SciCode**: Step-5-Preview **58.9%**, above GPT-6 Astra (56.5%) and Claude Opus 5 (56.4%). - **Multimodal**: MMMU-Pro and GDP.pdf were run with the unified multimodal encoder at default resolution. - **Output Speed**: 99.8 tokens/sec (GLM-5.3: 72.1 tokens/sec). - **Time to First Token**: 2.96 seconds (GLM-5.3: 2.99s; Claude Opus 5: 56.84s at max effort). - Missing entries (**โ€”**) reflect benchmarks that were not publicly reported for that model at the time of writing.
--- ## โšก Quickstart ### Installation ```bash pip install transformers>=4.56.0 pip install torch>=2.4.0 pip install accelerate ``` For video/image support: ```bash pip install av pillow ``` ### Basic Usage with Transformers ```python from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "SHSLab/Step-5-Preview-BF16" tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained( model_id, trust_remote_code=True, device_map="auto", torch_dtype="bfloat16", ) messages = [ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "Explain the significance of the Pareto Frontier in AI scaling."}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, return_tensors="pt", ).to(model.device) outputs = model.generate( inputs, max_new_tokens=1024, temperature=0.7, top_p=0.95, reasoning_effort="high", # low / medium / high / xhigh ) response = tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True) print(response) ``` ### Multimodal (Image + Video) Usage ```python from transformers import AutoProcessor processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True) messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://example.com/image.jpg"}, {"type": "video", "url": "https://example.com/video.mp4"}, {"type": "text", "text": "Describe the scene and summarize the video."}, ], } ] inputs = processor.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt") # ... generate as above ``` ### Tool Calling ```python tools = [ { "type": "function", "function": { "name": "get_weather", "parameters": { "type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"], }, }, } ] messages = [{"role": "user", "content": "What's the weather in Tokyo?"}] inputs = tokenizer.apply_chat_template( messages, tools=tools, add_generation_prompt=True, return_tensors="pt", ).to(model.device) outputs = model.generate(inputs, max_new_tokens=256, reasoning_effort="medium") print(tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True)) ``` --- ## ๐Ÿšข Deployment ### vLLM ```bash vllm serve SHSLab/Step-5-Preview-BF16 \ --trust-remote-code \ --tensor-parallel-size 8 \ --max-model-len 1000000 \ --enable-reasoning \ --reasoning-parser stepfun ``` ### SGLang ```bash python -m sglang.launch_server \ --model-path SHSLab/Step-5-Preview-BF16 \ --trust-remote-code \ --tp 8 \ --context-length 1000000 \ --reasoning-parser stepfun ``` ### OpenAI-Compatible API ```python from openai import OpenAI client = OpenAI( api_key="YOUR_STEP_API_KEY", base_url="https://api.stepfun.com/v1", ) response = client.chat.completions.create( model="step-5-preview", messages=[{"role": "user", "content": "Write a Python function to merge two sorted lists."}], reasoning_effort="high", max_tokens=2048, ) print(response.choices[0].message.content) ``` > **Recommended deployment configurations** > > - **BF16:** 8ร— H100 80GB (tensor parallel) > - **FP8:** 4ร— H100 80GB (coming soon) > - **Context length:** Up to 1M tokens > - **Reasoning parser:** Use `stepfun` for vLLM / SGLang --- ## ๐Ÿ“œ License Step-5-Preview is released under the **StepFun Community License**. See the [LICENSE](https://huggingface.co/SHSLab/Step-5-Preview-BF16/blob/main/LICENSE) file for full terms. > **Usage restrictions** > > - Commercial use is permitted under the StepFun Community License. > - Redistribution must include the license and attribution. > - See LICENSE for full details. --- ## ๐Ÿ“ฌ Contact - **Hugging Face:** [SHSLab](https://huggingface.co/SHSLab) - **GitHub:** [github.com/stepfun-ai](https://github.com/stepfun-ai) - **Discord:** [Join our Discord](https://discord.gg/stepfun) - **Email:** [opensource@stepfun.com](mailto:opensource@stepfun.com) - **Website:** [stepfun.com](https://stepfun.com) --- ## ๐Ÿ“š More details
๐Ÿ—๏ธ Model Architecture
Step 5 Architecture
### 92-Layer "Narrow but Deep" Design Step-5-Preview uses a **92-layer Transformer** with a narrow-deep configuration. This design creates longer information propagation paths for implicit multi-hop reasoning during long prefill operations. ### Sparse Grouped-Query Attention (GQA) with Block-Wise Token Merging To handle the 1M-token context window efficiently, Step-5-Preview uses **Sparse GQA with block-wise token merging**. This mechanism uses sparse indexing to select only historical information relevant to the current task, reducing the number of tokens that enter attention computation. This cuts indexer and top-k selection costs to approximately one-eighth of a denser baseline. ### Multimodal Encoder A unified multimodal encoder processes text, images, and video frames into a shared latent space. Video is sampled at adaptive frame rates and encoded with temporal attention, allowing the model to understand motion and long-range dependencies in screen recordings, demonstrations, and real-world footage.
๐Ÿ“š Training Data Step-5-Preview was trained on a carefully curated corpus spanning: - **Code repositories** from multiple languages (Python, C++, Rust, JavaScript, Go, etc.) - **Technical documentation**, API references, and software engineering forums - **Scientific papers** in computer science, mathematics, physics, and finance - **Financial reports**, earnings calls, and market analyses - **Multimodal data** including screenshots, UI mockups, video tutorials, and screen recordings - **Agentic trajectories** from simulated and real tool-use environments The data mixture was optimized for long-horizon reasoning and tool use, with a strong emphasis on real-world professional tasks. All data was filtered for quality, safety, and license compliance. Training combined next-token prediction with reinforcement learning from human feedback (RLHF) focused on agentic objectives.
๐Ÿค– Agentic Capabilities
Agentic Workflow
### 24-Hour Autonomous GPU Kernel Optimization Step-5-Preview was tasked with autonomously optimizing an H100 GPU kernel for up to 24 consecutive hours. The model: - Independently modified code - Ran tests and compared results - Iterated based on performance outcomes - **Reached 508 TFLOPS after approximately 22 hours** For comparison, Claude Opus 5 achieved 493 TFLOPS in the same experiment. ### Automated Post-Training Experiments In another 24-hour experiment, Step-5-Preview autonomously improved the accuracy of Qwen3-30B-A3B on AIME24 from 53.3% to 60% through automated post-training experiments. ### Long-Horizon Agent Workflows Optimized for workflows that require: - Searching and information retrieval - Running code and processing tool returns - Multi-turn tool calls with sustained execution - Iterative refinement based on intermediate results - Self-correction and error recovery over thousands of steps
๐Ÿ’ผ Real-World Use Cases - **ESP32 Development Board Modifications:** Executed development tasks for over 3 hours. - **Front-End Design with 3D Asset Generation:** Full-stack development workflows including visual design. - **Full-Process Financial Research:** End-to-end investment research, from data gathering to report generation. - **Software Engineering:** Comprehensive coding tasks including front-end, visual development, and programmable hardware. - **Autonomous Research Assistant:** Reading papers, running experiments, and summarizing findings. - **Customer Support Automation:** Multi-turn conversations with tool calls to internal systems.
๐Ÿ“ˆ Full Evaluation Table All evaluations used the `high` reasoning effort setting unless otherwise noted. | Benchmark | Score | Notes | |:---|:---|:---| | **DeepSWE v1.1** | 67.7% | SWE-agent harness, temp=1.0, top_p=0.95 | | **StepCodeBench** | 49.0% | avg@4 | | **ProgramBench** | 80.5% | Rebuild programs from binary + docs; 200 tasks, 248K+ behavioral tests | | **Terminal-Bench 2.1** | 85.0% | Verified refresh of TB 2.0; 89 curated terminal tasks across SWE, sysadmin, data processing | | **Terminal-Bench 4** | 33.3% | 2.6ร— Kimi K3 | | **CyberGym** | 84.7% | Best in comparison set | | **SciCode** | 58.9% | Above GPT-6 Astra (56.5%) and Opus 5 (56.4%) | | **RoadmapBench** | 54.3% | 115 long-horizon coding tasks across 17 repos, 5 languages; median ~3,700 LOC changed | | **SWE-Marathon v1.1** | 72.7% | Ultra-long-horizon SWE; 20 realistic multi-hour tasks with hidden/adversarial tests | | **MLS-Bench-Lite** | 40.5% | 30-task subset of MLS-Bench; tests inventing generalizable ML methods across 12 research domains | | **SWE-Atlas-QnA** | 63.6% | 124 tasks on deep code comprehension โ€” tracing execution paths, explaining architecture across production repos | | **SWE-Atlas-Test-writing** | 50.8% | 90 tasks on writing production-grade unit, integration, and acceptance tests for real repositories | | **StepCode-Bench-Daily** | 64.9% | StepFun internal; 553 repos, 9 task types, 20 domains, 33 languages; daily-difficulty slice | | **StepCode-Bench-General** | 65.0% | StepFun internal; same corpus as StepCodeBench; general-difficulty slice | | **Agents' Last Exam (ALE-CLI)** | 29.5% | Linux-only CLI subset of ALE; 40 industry subfields; best agent pass rate ~25.2% | | **GDPval-AA v2** | 1571 | Artificial Analysis, Sep 19, 2026 | | **AA-Briefcase** | 1417 | Elo; agentic knowledge work across data science, product, banking, heavy industry; private held-out test set | | **Toolathlon-Verified** | 74.1% | 108 expert-authored tasks; multi-app workflows averaging ~20 turns; strictly verifiable via scripts | | **MCP-Atlas** | 85.6% | 1,000 tasks (500 public + 500 private); 36 real MCP servers, 220 tools; 3โ€“6 tool calls/task | | **PresentBench** | 76.8% | 238 slide-generation instances; avg 54.1 rubric checklist items per instance | | **JobBench** | 59.0% | 130 tasks across 35 occupations; avg 35.6 binary rubric criteria per task | | **Apex-Agents** | 37.8% | Long-horizon tasks in investment banking, consulting, corporate law | | **DRACO** | 83.3% | Cross-domain deep research; accuracy, completeness, objectivity, citation quality | | **BrowseComp** | 88.7% | 1,266 hard-to-find web information retrieval questions | | **HLE w/ tools** | 59.4% | +12.9 pts over no-tools HLE | | **FrontierFinance** | 66.4% | +11.4 pts over GPT-6 Astra | | **FinStepBench-LiveSearch** | 74.5% | Tied with GPT-6 Astra | | **FinStepBench-CorporateValuation** | 60.6% | Tied with Kimi K3 | | **FinStepBench-FinanceDR** | 55.8% | StepFun internal; full deep-research finance workflows | | **OfficeQA Pro** | 60.3% | Databricks; 90 questions over large enterprise financial document collections | | **SpeadSheet v2** | 29.4% | SpreadsheetBench 2; end-to-end business spreadsheet workflows; best model โ‰ˆ34.89% | | **GDP.pdf** | 14.8% | Known weak spot | | **GPQA Diamond** | 93.5% | 198 PhD-level science MCQs; human expert avg 81% | | **HLE** | 46.5% | 59.4% with tools | | **AA-LCR v1.1** | 88.3% | Tied with Kimi K3 (88.7%) | | **CritPt** | 20.9% | 71 unpublished research-level physics challenges | | **MMMU-Pro** | 76.0% | Robust multimodal benchmark; significantly harder than MMMU | | **Output Speed** | 99.8 tokens/sec | GLM-5.3: 72.1 tokens/sec | | **Time to First Token** | 2.96s | GLM-5.3: 2.99s; Claude Opus 5: 56.84s (max effort) |
๐Ÿ… Category summary | Domain | Standing | Highlight | |:---|:---|:---| | **Long-Context Reasoning** | Near-tied for #1 | AA-LCR v1.1 **88.3%** (+7.6 over GPT-6 Astra, +9.0 over Opus 5) | | **Coding / SWE** | #1 open-weight | DeepSWE v1.1 **67.7%**, StepCodeBench **49.0%**, ProgramBench **80.5%** | | **Security / Cyber** | #1 in comparison set | CyberGym **84.7%** | | **Finance** | #2 overall, #1 open-weight | FrontierFinance **66.4%**, FinanceDR **55.8%** | | **Terminal Agents** | #2 open-weight | Terminal-Bench 4 **33.3%** (2.6ร— Kimi K3) | | **Tool Use** | Top-3 open-weight | MCP-Atlas **85.6%**, HLE w/ tools **59.4%** | | **General Knowledge** | Top tier | GPQA Diamond **93.5%**, HLE **46.5%** | | **Multimodal** | Trailing | MMMU-Pro **76.0%** โ€” targeted for improvement |
โš ๏ธ Limitations - **Knowledge Cutoff:** Knowledge is current up to mid-2026. May not be aware of later events. - **Hallucination:** Can generate plausible but incorrect information, especially in domains with sparse training data. - **Long Context Degradation:** Performance may degrade for extremely long contexts beyond 500K tokens in certain tasks. - **Tool Use Reliability:** Tool calling is capable but not infallible. Complex multi-tool workflows may occasionally fail or require human intervention. - **Multimodal Limitations:** Video understanding is limited to clips under 5 minutes and 128 MB. Extremely high-resolution images may be downscaled. - **Language Coverage:** Primarily optimized for English and Chinese. Performance in other languages may vary.
โš–๏ธ Ethical Considerations StepFun is committed to the responsible development and deployment of AI. Measures taken: - **Safety Alignment:** Fine-tuned with RLHF to refuse harmful requests and promote helpful, honest, and harmless behavior. - **Bias Mitigation:** Training data was filtered to reduce harmful stereotypes and biases. Residual biases may exist. - **Transparency:** Detailed model cards and benchmark results are provided to enable informed use. - **License Restrictions:** The StepFun Community License prohibits certain high-risk uses, including autonomous weapons, surveillance, and malicious cyber activities. - **Content Provenance:** Users are encouraged to clearly label AI-generated content. All users should consider the ethical implications of their applications and implement appropriate safeguards.
๐Ÿ–ฅ๏ธ Hardware Requirements | Precision | Minimum GPU Memory | Recommended GPU Configuration | |:---|:---|:---| | **BF16** | 1.2 TB | 8ร— H100 80GB (tensor parallel) | | **FP8** | 600 GB | 4ร— H100 80GB (tensor parallel) | | **INT4** | 300 GB | 4ร— A100 80GB (tensor parallel) | For inference with 1M context, additional memory is required for KV cache. Paged attention and offloading techniques available in vLLM and SGLang are recommended.
โšก Performance Metrics | Metric | Value | |:---|:---| | **Output Speed** | 99.8 tokens/sec | | **Time to First Token (TTFT)** | 2.96 seconds | | **Context Window** | 1,000,000 tokens | | **Max Output Tokens** | 32,768 (default), configurable up to 131,072 | | **Reasoning Effort Modes** | low, medium, high, xhigh | | **Tool Calling Latency** | < 500 ms for simple calls | *Measured on 8ร— H100 80GB with vLLM, batch size 1, BF16.*
๐Ÿ“š Citation ```bibtex @misc{stepfun2026step5preview, title = {Step-5-Preview: A 600B Sparse MoE Foundation Model for Real-World Agentic Work}, author = {StepFun Team}, year = {2026}, howpublished = {\url{https://huggingface.co/SHSLab/Step-5-Preview-BF16}}, note = {Released September 20, 2026} } ```
---
โญ If you find Step-5-Preview useful, please give us a star on GitHub and Hugging Face! โญ

Built with โค๏ธ by StepFun