Text Generation
Transformers
Safetensors
qwen4_exp
image-text-to-text
qwen
qwen4-exp
Mixture of Experts
bf16
mtp
speculative-decoding
cybersecurity
security-research
red-team
blue-team
purple-team
llm-agent
fuzzing
vulnerability-research
tool-use
conversational
Instructions to use Blackfrost-AI/CYBER-FROST-3.8-BF16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Blackfrost-AI/CYBER-FROST-3.8-BF16 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Blackfrost-AI/CYBER-FROST-3.8-BF16") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Blackfrost-AI/CYBER-FROST-3.8-BF16") model = AutoModelForMultimodalLM.from_pretrained("Blackfrost-AI/CYBER-FROST-3.8-BF16", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Blackfrost-AI/CYBER-FROST-3.8-BF16 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Blackfrost-AI/CYBER-FROST-3.8-BF16" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Blackfrost-AI/CYBER-FROST-3.8-BF16", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Blackfrost-AI/CYBER-FROST-3.8-BF16
- SGLang
How to use Blackfrost-AI/CYBER-FROST-3.8-BF16 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Blackfrost-AI/CYBER-FROST-3.8-BF16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Blackfrost-AI/CYBER-FROST-3.8-BF16", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Blackfrost-AI/CYBER-FROST-3.8-BF16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Blackfrost-AI/CYBER-FROST-3.8-BF16", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Blackfrost-AI/CYBER-FROST-3.8-BF16 with Docker Model Runner:
docker model run hf.co/Blackfrost-AI/CYBER-FROST-3.8-BF16
|
Download README.md from Blackfrost-AI/CYBER-FROST-3.8-BF16: direct link, hf CLI and curl.
- Browser
- Download file 14.1 kB
-
https://huggingface.co/Blackfrost-AI/CYBER-FROST-3.8-BF16/resolve/main/README.md
- Command line
-
hf download hf://Blackfrost-AI/CYBER-FROST-3.8-BF16/README.md
-
curl -L -o README.md https://huggingface.co/Blackfrost-AI/CYBER-FROST-3.8-BF16/resolve/main/README.md
14.1 kB
| base_model: Qwen/Qwen3.8-Flash-Next | |
| library_name: transformers | |
| pipeline_tag: text-generation | |
| license: other | |
| license_name: qwen-community-license-1.0 | |
| license_link: LICENSE | |
| tags: | |
| - qwen | |
| - qwen4-exp | |
| - moe | |
| - bf16 | |
| - mtp | |
| - speculative-decoding | |
| - cybersecurity | |
| - security-research | |
| - red-team | |
| - blue-team | |
| - purple-team | |
| - llm-agent | |
| - fuzzing | |
| - vulnerability-research | |
| - tool-use | |
|  | |
| # CYBER-FROST-3.8-BF16 | |
| **A first-party Blackfrost-AI BF16 model for security professionals conducting authorized research, assessment, engineering, and response work.** | |
| ## Cyber-Frost Harness | |
|  | |
| [Cyber-Frost Harness](https://github.com/Blackfrost-AI/Cyber-Frost-Harness) is the public red/blue/purple runtime built around the Cyber-Frost family. It supplies eight procedural security skills, structured native-analysis tools, durable evidence handling, token-aware context management, and an isolated x86-64 vulnerable-image environment. | |
| The published hard-12 harness result shown above used the sibling [`CYBER-FROST-3.8-NVFP4-V2`](https://huggingface.co/Blackfrost-AI/CYBER-FROST-3.8-NVFP4-V2) artifact, not this BF16 checkpoint. It increased verified solves from 1 under the generic scaffold to 5 under the harness, with five valid model misses and two infrastructure-invalid tasks. The defensible statements are **5/10 valid attempts** and a **5/12 verified lower bound**; no final 12-task percentage is claimed. Total model tokens fell from 51,939,855 to 22,043,450. These are scaffold-intervention results and do not transfer automatically to BF16. | |
| - [Source and quick start](https://github.com/Blackfrost-AI/Cyber-Frost-Harness) | |
| - [Architecture and trust boundaries](https://github.com/Blackfrost-AI/Cyber-Frost-Harness/blob/main/docs/ARCHITECTURE.md) | |
| - [Full evaluation disclosure](https://github.com/Blackfrost-AI/Cyber-Frost-Harness/blob/main/docs/EVALUATION.md) | |
| - [Publication-safe result records](https://github.com/Blackfrost-AI/Cyber-Frost-Harness/tree/main/results) | |
| The model-facing container receives only the vulnerable task image and no network, fixed image, Docker control, grader database, or credentials. The harness changes the runtime and evidence path; it does not change model weights. | |
| ## Release status and contents | |
| The release contains the complete BF16 SafeTensors checkpoint, weight index, configuration, tokenizer and processor assets, the packaged chat template, upstream license, this model card, and a documented deployment kit for the validated serving profile. It is not an adapter and does not require a separate parent checkpoint at load time. | |
| | Field | Released artifact | | |
| |---|---| | |
| | Clean model name | `CYBER-FROST-3.8-BF16` | | |
| | Former name | `BLACKFROST-3.8-ICED-BF16` | | |
| | Architecture | `Qwen4ExpForConditionalGeneration` | | |
| | Precision | BF16 | | |
| | Weight layout | 131 SafeTensors shards | | |
| | Total parameters reported by Hub metadata | approximately 180B | | |
| | Configured context | 262,144 tokens | | |
| | Context exercised in the published performance trial | 32,768 tokens | | |
| | Native speculative head | one MTP layer | | |
| | Validated modality | text | | |
| The configuration includes a vision tower, but this release has not received a multimodal quality evaluation. Do not infer validated image or video capability from the presence of processor files. | |
| ## Why Cyber-Frost exists | |
| Security work is unusually vulnerable to false refusals. The same vocabulary appears in incident response, exploit validation, malware analysis, defensive engineering, and unauthorized activity; a general-purpose assistant can react to individual terms instead of the operator's legitimate scope. | |
| Cyber-Frost is designed to reduce that unnecessary friction in professional, authorized workflows. It is intended to stay technically direct when an analyst is reviewing a finding, reproducing a vulnerability in a controlled environment, writing detection content, analyzing malicious code, or operating an approved security agent. This is a design objective, not a claim that the model has no refusal behavior or that every answer is safe or correct. | |
| Authorization is an external control. The model cannot establish ownership, consent, rules of engagement, jurisdiction, or whether a target is in scope. Deployers must enforce identity, scope, tool permissions, logging, rate limits, and human review outside the model. | |
| ## Security corpus | |
| Cyber-Frost was fine-tuned on a Blackfrost-AI security corpus combining curated security material, operator-authored workflows, realistic engagement-style scenarios, and Blackfrost-owned distillation data. Corpus sizes, source-by-source counts, raw engagement material, client identities, prompts, and responses are intentionally not published. | |
| Domain coverage includes: | |
| - Reconnaissance and OSINT | |
| - Social engineering, business-email compromise, and deepfake-enabled abuse | |
| - Web application and API security | |
| - Identity, authentication, and Active Directory security | |
| - Network, perimeter, VPN, and protocol security | |
| - Vulnerability research, bug bounty, and binary exploitation | |
| - Malware analysis, ransomware, and endpoint defense | |
| - Cloud, container, and Kubernetes security | |
| - Software supply-chain security | |
| - Mobile, IoT, wireless, and physical security | |
| - Industrial-control-system and operational-technology security | |
| - Cryptography and security protocols | |
| - Privilege escalation, lateral movement, and data exfiltration | |
| - Threat intelligence, APT analysis, and purple-team operations | |
| - AI-agent, LLM, and adversarial-ML security | |
| Blackfrost-AI attests that owned portions of the corpus were developed from sanitized experience with authorized security work. It also applies a frontier-scale policy to its distillation teachers, excluding teachers below the 753B-parameter class. The release evidence independently binds one security subset to a Qwen3.8 2.4T teacher; it does not include a corpus-wide teacher manifest. These are therefore operator provenance statements, not independent benchmark findings. | |
| Training-data provenance and licensing review for the mixed-source corpus remains in progress. Treat that unresolved review as an explicit limitation of this public release. | |
| ## Model specifications | |
| The text stack has 48 blocks with hybrid linear and full attention, using a full-attention block every fourth layer. Its hidden size is 2,560 with 24 attention heads and 2 KV heads. The MoE stack contains 512 routed experts, selects 10 experts per token, and includes a shared expert. One native MTP layer is packaged for speculative decoding. | |
| The configured context ceiling is not a blanket quality guarantee. The published BF16 trial exercised 32,768 tokens. Longer contexts, high concurrency, multimodal requests, and tool-heavy agent loops require separate validation. | |
| ## Lineage | |
| 1. **Foundational checkpoint:** [`Qwen/Qwen3.8-Flash-Next`](https://huggingface.co/Qwen/Qwen3.8-Flash-Next) at immutable revision `de4b8e4d43b917e7706784d8bb445c9af86a3540`. | |
| 2. **Blackfrost security adaptation:** security-domain fine-tuning followed by a full BF16 merge. The merged internal stage was identified as `BLACKFROST-3.8-FLASH-BF16`. | |
| 3. **Behavioral stage:** a Blackfrost-AI behaviorally modified derivative targeting lower false-refusal friction in authorized security workflows. The proprietary transformation process is not distributed. | |
| 4. **Release identity:** the same released BF16 weight payload was formerly labeled `BLACKFROST-3.8-ICED-BF16` and is now named `CYBER-FROST-3.8-BF16`. The rename does not represent another training run. | |
| Tokenizer and processor lineage comes from the pinned Qwen foundation checkpoint. The packaged Blackfrost chat template is release-specific. | |
| ## Artifact verification | |
| The released weight payload is tied to Hub revision `5321904427c4ef54df8a667edcbc2d1184e4286e`. On 2026-09-15, every weight shard was checked against its corresponding Hub LFS object identifier with no mismatch. The old ICED label and the Cyber-Frost repository resolve to this same payload. | |
| | Artifact | SHA-256 | | |
| |---|---| | |
| | Sanitized `config.json` prepared for this card update | `c77cfd62621660b25e4c811bf44bf2ad4f63af98f23a874952a99e4e85f189f2` | | |
| | `model.safetensors.index.json` | `99e815241ef03325536b0aaa4441deea45174c17fae31e10f0bb456410c590de` | | |
| | Qwen Community License file | `a0dc422560841fd68e06d974907f8b4c709bca44a67daad2b528437bdf676c08` | | |
| The sanitized configuration removes legacy local-path metadata only; it does not alter architecture or inference behavior. | |
| ## Measured behavior | |
| There were no API errors. The evaluation used an earlier internal suite revision that was subsequently corrected, so these results are development evidence rather than a standardized leaderboard score. The zero safe over-refusal result supports the intended low-friction behavior on that slice; the uneven harmful-request refusal rates also show why this checkpoint must not be treated as a safety control. | |
| No standardized cyber-capability benchmark has yet been qualified for this exact BF16 checkpoint. In particular, this card does not claim a CyberMetric, SecBench, MMLU, HumanEval, or IFEval score. Refusal behavior is not a substitute for measuring security competence. | |
| ### Four-B300 BF16 performance trial | |
| The measured profile used four NVIDIA B300 SXM6 GPUs, tensor parallelism 4, vLLM `0.29.1rc1.dev13+g1cfd97281`, 32,768-token serving context, sequential requests, and thinking enabled. The frozen suite covered reasoning, code, prose, and tool-result synthesis with temperature 1.0, top-p 0.95, top-k 20, seed 38421, and at most 256 generated tokens. | |
| | Mode | Completion tokens | Median decode | Mean decode | Median TTFT | Draft acceptance | Relative median | | |
| |---|---:|---:|---:|---:|---:|---:| | |
| | No draft | 1,760 | 190.10 tok/s | 190.51 tok/s | 0.696 s | n/a | baseline | | |
| | Native MTP, k=2 | 1,820 | **265.00 tok/s** | 264.27 tok/s | **0.660 s** | 55.0% | **+39.4%** | | |
| | Native MTP, k=3 | 1,799 | 263.10 tok/s | 288.43 tok/s | 0.684 s | 49.6% | +38.4% | | |
| These are single-pass development measurements, not production throughput guarantees. Workload, concurrency, context, runtime build, cache precision, and sampling can materially change the result. | |
| The embedded MTP tensors predate the final trunk-weight behavioral stage. Treat the MTP head as a provisional acceleration baseline rather than a freshly adapted draft head. With the tested vLLM build, native MTP plus the model's Mamba groups also disables cross-request prefix-cache reuse, and fused multi-step draft decode is unavailable for the experimental attention backend. | |
| ## Prompt, tool use, and sampling | |
| The repository includes a Blackfrost chat template that supplies a default operating prompt, supports caller-provided system context, exposes Qwen-style reasoning controls, and serializes tool calls. Its default is thinking enabled; supported reasoning-effort values are `xhigh`, `medium`, and `low`. Generation defaults are temperature 1.0, top-p 0.95, and top-k 20. | |
| The template's authorization assumption is not an access-control mechanism. An agent runtime must independently restrict credentials, targets, files, networks, commands, and approval-requiring actions. Tool-call output is proposed text until an external executor acts on it. | |
| Changing the template, system message, reasoning mode, or sampling can materially change refusal behavior and output quality. Record those settings when reporting results. | |
| ## Deployment | |
| Use the included [`DEPLOYMENT/`](DEPLOYMENT/README.md) kit for the validated four-GPU BF16 serving profile, environment variables, launch command, and smoke tests. The clean API model identifier is `CYBER-FROST-3.8-BF16`. | |
| ## Intended use | |
| Cyber-Frost is intended for qualified security professionals working within explicit authorization, including defensive research, secure code review, vulnerability validation, red-team and purple-team exercises, detection engineering, incident response, malware analysis, bug hunting, and controlled security-agent workflows. | |
| It is not intended to authorize access, choose targets, define rules of engagement, make autonomous high-impact decisions, or replace legal, compliance, and safety review. Do not use it to access systems or data without permission, evade oversight, persist in third-party environments, deploy malware, steal credentials or data, disrupt services, or cause physical harm. | |
| ## Limitations and security responsibility | |
| - Generated findings, code, commands, indicators, and remediation advice may be wrong, incomplete, outdated, or fabricated. Independently review them and execute only in isolated, authorized environments. | |
| - Reduced over-refusal can increase the chance of receiving actionable output in ambiguous or malicious contexts. Gating and a system prompt do not remove that risk. | |
| - The model is not a policy engine, authorization service, sandbox, malware scanner, or secrets boundary. | |
| - The current evidence does not establish production readiness, comprehensive safety, comprehensive cyber competence, long-context quality beyond the tested profile, or multimodal quality. | |
| - Model behavior can shift substantially with prompts, sampling, runtime versions, quantization, speculative settings, and agent scaffolding. | |
| The deployer is responsible for authorization, least privilege, isolation, network policy, credential handling, human approval gates, monitoring, incident response, and compliance with applicable law. | |
| ## License and disclaimer | |
| Use and redistribution of this checkpoint are governed by the [Qwen Community License 1.0](LICENSE). Review the license before use. This research release is provided without a warranty of correctness, fitness, security, or non-infringement. | |
| Report reproducible model or packaging issues through the repository's [Discussions](https://huggingface.co/Blackfrost-AI/CYBER-FROST-3.8-BF16/discussions) page without including secrets, client data, live targets, or sensitive exploit details. | |