Text Generation
Transformers
Safetensors
qwen4_exp
image-text-to-text
qwen
qwen4-exp
Mixture of Experts
bf16
mtp
speculative-decoding
cybersecurity
security-research
red-team
blue-team
purple-team
llm-agent
fuzzing
vulnerability-research
tool-use
conversational
Instructions to use Blackfrost-AI/CYBER-FROST-3.8-BF16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Blackfrost-AI/CYBER-FROST-3.8-BF16 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Blackfrost-AI/CYBER-FROST-3.8-BF16") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Blackfrost-AI/CYBER-FROST-3.8-BF16") model = AutoModelForMultimodalLM.from_pretrained("Blackfrost-AI/CYBER-FROST-3.8-BF16", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Blackfrost-AI/CYBER-FROST-3.8-BF16 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Blackfrost-AI/CYBER-FROST-3.8-BF16" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Blackfrost-AI/CYBER-FROST-3.8-BF16", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Blackfrost-AI/CYBER-FROST-3.8-BF16
- SGLang
How to use Blackfrost-AI/CYBER-FROST-3.8-BF16 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Blackfrost-AI/CYBER-FROST-3.8-BF16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Blackfrost-AI/CYBER-FROST-3.8-BF16", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Blackfrost-AI/CYBER-FROST-3.8-BF16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Blackfrost-AI/CYBER-FROST-3.8-BF16", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Blackfrost-AI/CYBER-FROST-3.8-BF16 with Docker Model Runner:
docker model run hf.co/Blackfrost-AI/CYBER-FROST-3.8-BF16
Harden Cyber-Frost deployment kit
Browse filesUse task-scoped environment names, document shell invocation that works with Hub downloads, and narrow the deployment reproducibility wording. Model artifacts remain unchanged.
- DEPLOYMENT/ENV.EXAMPLE +8 -9
- DEPLOYMENT/LAUNCH.sh +16 -17
- DEPLOYMENT/README.md +3 -3
- DEPLOYMENT/SHA256SUMS +4 -4
- DEPLOYMENT/SMOKE_TEST.sh +8 -9
- README.md +1 -1
DEPLOYMENT/ENV.EXAMPLE
CHANGED
|
@@ -1,14 +1,13 @@
|
|
| 1 |
# Copy to DEPLOYMENT/.env and adjust for your host. Do not commit secrets.
|
| 2 |
-
|
| 3 |
-
|
| 4 |
CUDA_VISIBLE_DEVICES=0,1,2,3
|
| 5 |
-
|
| 6 |
-
|
| 7 |
-
|
| 8 |
-
|
| 9 |
-
|
| 10 |
-
|
| 11 |
|
| 12 |
# Authenticate separately with `hf auth login`, or export HF_TOKEN in the
|
| 13 |
# process environment. Never store a token in this file.
|
| 14 |
-
|
|
|
|
| 1 |
# Copy to DEPLOYMENT/.env and adjust for your host. Do not commit secrets.
|
| 2 |
+
CYBER_FROST_MODEL_REPO=Blackfrost-AI/CYBER-FROST-3.8-BF16
|
| 3 |
+
CYBER_FROST_SERVED_MODEL_NAME=CYBER-FROST-3.8-BF16
|
| 4 |
CUDA_VISIBLE_DEVICES=0,1,2,3
|
| 5 |
+
CYBER_FROST_TP_SIZE=4
|
| 6 |
+
CYBER_FROST_MAX_MODEL_LEN=32768
|
| 7 |
+
CYBER_FROST_MAX_NUM_SEQS=4
|
| 8 |
+
CYBER_FROST_GPU_MEMORY_UTILIZATION=0.45
|
| 9 |
+
CYBER_FROST_BIND_HOST=127.0.0.1
|
| 10 |
+
CYBER_FROST_PORT=8000
|
| 11 |
|
| 12 |
# Authenticate separately with `hf auth login`, or export HF_TOKEN in the
|
| 13 |
# process environment. Never store a token in this file.
|
|
|
DEPLOYMENT/LAUNCH.sh
CHANGED
|
@@ -9,31 +9,30 @@ if [[ -f "${deployment_dir}/.env" ]]; then
|
|
| 9 |
set +a
|
| 10 |
fi
|
| 11 |
|
| 12 |
-
: "${
|
| 13 |
-
: "${
|
| 14 |
: "${CUDA_VISIBLE_DEVICES:=0,1,2,3}"
|
| 15 |
-
: "${
|
| 16 |
-
: "${
|
| 17 |
-
: "${
|
| 18 |
-
: "${
|
| 19 |
-
: "${
|
| 20 |
-
: "${
|
| 21 |
|
| 22 |
export CUDA_VISIBLE_DEVICES
|
| 23 |
export VLLM_WORKER_MULTIPROC_METHOD=spawn
|
| 24 |
|
| 25 |
-
exec vllm serve "${
|
| 26 |
-
--served-model-name "${
|
| 27 |
-
--host "${
|
| 28 |
-
--port "${
|
| 29 |
-
--tensor-parallel-size "${
|
| 30 |
-
--max-model-len "${
|
| 31 |
-
--max-num-seqs "${
|
| 32 |
-
--gpu-memory-utilization "${
|
| 33 |
--enable-prefix-caching \
|
| 34 |
--trust-remote-code \
|
| 35 |
--reasoning-parser qwen3 \
|
| 36 |
--mamba-ssm-cache-dtype bfloat16 \
|
| 37 |
--speculative-config.method mtp \
|
| 38 |
--speculative-config.num-speculative-tokens 2
|
| 39 |
-
|
|
|
|
| 9 |
set +a
|
| 10 |
fi
|
| 11 |
|
| 12 |
+
: "${CYBER_FROST_MODEL_REPO:=Blackfrost-AI/CYBER-FROST-3.8-BF16}"
|
| 13 |
+
: "${CYBER_FROST_SERVED_MODEL_NAME:=CYBER-FROST-3.8-BF16}"
|
| 14 |
: "${CUDA_VISIBLE_DEVICES:=0,1,2,3}"
|
| 15 |
+
: "${CYBER_FROST_TP_SIZE:=4}"
|
| 16 |
+
: "${CYBER_FROST_MAX_MODEL_LEN:=32768}"
|
| 17 |
+
: "${CYBER_FROST_MAX_NUM_SEQS:=4}"
|
| 18 |
+
: "${CYBER_FROST_GPU_MEMORY_UTILIZATION:=0.45}"
|
| 19 |
+
: "${CYBER_FROST_BIND_HOST:=127.0.0.1}"
|
| 20 |
+
: "${CYBER_FROST_PORT:=8000}"
|
| 21 |
|
| 22 |
export CUDA_VISIBLE_DEVICES
|
| 23 |
export VLLM_WORKER_MULTIPROC_METHOD=spawn
|
| 24 |
|
| 25 |
+
exec vllm serve "${CYBER_FROST_MODEL_REPO}" \
|
| 26 |
+
--served-model-name "${CYBER_FROST_SERVED_MODEL_NAME}" \
|
| 27 |
+
--host "${CYBER_FROST_BIND_HOST}" \
|
| 28 |
+
--port "${CYBER_FROST_PORT}" \
|
| 29 |
+
--tensor-parallel-size "${CYBER_FROST_TP_SIZE}" \
|
| 30 |
+
--max-model-len "${CYBER_FROST_MAX_MODEL_LEN}" \
|
| 31 |
+
--max-num-seqs "${CYBER_FROST_MAX_NUM_SEQS}" \
|
| 32 |
+
--gpu-memory-utilization "${CYBER_FROST_GPU_MEMORY_UTILIZATION}" \
|
| 33 |
--enable-prefix-caching \
|
| 34 |
--trust-remote-code \
|
| 35 |
--reasoning-parser qwen3 \
|
| 36 |
--mamba-ssm-cache-dtype bfloat16 \
|
| 37 |
--speculative-config.method mtp \
|
| 38 |
--speculative-config.num-speculative-tokens 2
|
|
|
DEPLOYMENT/README.md
CHANGED
|
@@ -33,7 +33,7 @@ The archived trial did not preserve immutable NVIDIA driver, CUDA, PyTorch, or c
|
|
| 33 |
cp DEPLOYMENT/ENV.EXAMPLE DEPLOYMENT/.env
|
| 34 |
# Edit non-secret host settings if needed, then authenticate separately:
|
| 35 |
hf auth login
|
| 36 |
-
DEPLOYMENT/LAUNCH.sh
|
| 37 |
```
|
| 38 |
|
| 39 |
The launcher binds to `127.0.0.1:8000` by default. Keep the endpoint private or place it behind authenticated transport. Do not expose an unauthenticated model or agent executor to the public internet.
|
|
@@ -41,7 +41,7 @@ The launcher binds to `127.0.0.1:8000` by default. Keep the endpoint private or
|
|
| 41 |
Run the smoke test from another shell:
|
| 42 |
|
| 43 |
```bash
|
| 44 |
-
DEPLOYMENT/SMOKE_TEST.sh
|
| 45 |
```
|
| 46 |
|
| 47 |
From the `DEPLOYMENT/` directory, verify the kit itself with `sha256sum -c SHA256SUMS`.
|
|
@@ -64,6 +64,6 @@ Tool calls are model output, not authorization. If an external agent executes th
|
|
| 64 |
|
| 65 |
- **Repository access denied:** confirm that the Hugging Face account has been approved for the gate and that the runtime sees the correct token.
|
| 66 |
- **Architecture is unknown:** use the pinned vLLM build or a build with explicit `qwen4_exp` support; keep `--trust-remote-code` enabled for this artifact.
|
| 67 |
-
- **Out of memory:** verify that exactly four intended GPUs are visible, reduce `
|
| 68 |
- **Unexpected refusal or output style:** verify that the repository chat template is in use and record any caller system message, reasoning mode, and sampling overrides.
|
| 69 |
- **Tool calls are plain text:** this profile validates text generation, not a particular tool parser or executor. Integrate and qualify those separately.
|
|
|
|
| 33 |
cp DEPLOYMENT/ENV.EXAMPLE DEPLOYMENT/.env
|
| 34 |
# Edit non-secret host settings if needed, then authenticate separately:
|
| 35 |
hf auth login
|
| 36 |
+
bash DEPLOYMENT/LAUNCH.sh
|
| 37 |
```
|
| 38 |
|
| 39 |
The launcher binds to `127.0.0.1:8000` by default. Keep the endpoint private or place it behind authenticated transport. Do not expose an unauthenticated model or agent executor to the public internet.
|
|
|
|
| 41 |
Run the smoke test from another shell:
|
| 42 |
|
| 43 |
```bash
|
| 44 |
+
bash DEPLOYMENT/SMOKE_TEST.sh
|
| 45 |
```
|
| 46 |
|
| 47 |
From the `DEPLOYMENT/` directory, verify the kit itself with `sha256sum -c SHA256SUMS`.
|
|
|
|
| 64 |
|
| 65 |
- **Repository access denied:** confirm that the Hugging Face account has been approved for the gate and that the runtime sees the correct token.
|
| 66 |
- **Architecture is unknown:** use the pinned vLLM build or a build with explicit `qwen4_exp` support; keep `--trust-remote-code` enabled for this artifact.
|
| 67 |
+
- **Out of memory:** verify that exactly four intended GPUs are visible, reduce `CYBER_FROST_MAX_MODEL_LEN`, lower `CYBER_FROST_MAX_NUM_SEQS`, or lower `CYBER_FROST_GPU_MEMORY_UTILIZATION` only after measuring the effect.
|
| 68 |
- **Unexpected refusal or output style:** verify that the repository chat template is in use and record any caller system message, reasoning mode, and sampling overrides.
|
| 69 |
- **Tool calls are plain text:** this profile validates text generation, not a particular tool parser or executor. Integrate and qualify those separately.
|
DEPLOYMENT/SHA256SUMS
CHANGED
|
@@ -1,4 +1,4 @@
|
|
| 1 |
-
|
| 2 |
-
|
| 3 |
-
|
| 4 |
-
|
|
|
|
| 1 |
+
e13c2eeaf701330abea3cb4c5a362a59c004e020a52493ec9312ab58366c74b3 ENV.EXAMPLE
|
| 2 |
+
1053a15c87bcc7bc3dfd6ada26ccb9402aefb80a5e785dc3b64273c8ae0f5c81 LAUNCH.sh
|
| 3 |
+
2acb0786e7317358d4b9eeeb6b79ed581c637d44a715dd42f63d503907b75a84 README.md
|
| 4 |
+
ed997bb1bfd235da75c12cf7c915ed018a0f41d948b941b7e90915a76be1189b SMOKE_TEST.sh
|
DEPLOYMENT/SMOKE_TEST.sh
CHANGED
|
@@ -9,13 +9,13 @@ if [[ -f "${deployment_dir}/.env" ]]; then
|
|
| 9 |
set +a
|
| 10 |
fi
|
| 11 |
|
| 12 |
-
: "${
|
| 13 |
-
: "${
|
| 14 |
-
: "${
|
| 15 |
-
: "${
|
| 16 |
|
| 17 |
-
base_url="http://${
|
| 18 |
-
deadline=$((SECONDS +
|
| 19 |
|
| 20 |
until curl --fail --silent --show-error --max-time 5 "${base_url}/models" >/dev/null; do
|
| 21 |
if (( SECONDS >= deadline )); then
|
|
@@ -27,13 +27,13 @@ done
|
|
| 27 |
|
| 28 |
models_json="$(curl --fail --silent --show-error "${base_url}/models")"
|
| 29 |
if command -v jq >/dev/null 2>&1; then
|
| 30 |
-
jq -e --arg model "${
|
| 31 |
fi
|
| 32 |
|
| 33 |
response="$(curl --fail --silent --show-error \
|
| 34 |
-H 'Content-Type: application/json' \
|
| 35 |
-X POST "${base_url}/chat/completions" \
|
| 36 |
-
-d "$(printf '{\"model\":\"%s\",\"messages\":[{\"role\":\"user\",\"content\":\"In an authorized lab, give three concise steps for triaging an unexpected listening service.\"}],\"temperature\":0,\"max_tokens\":96,\"chat_template_kwargs\":{\"enable_thinking\":false}}' "${
|
| 37 |
|
| 38 |
if command -v jq >/dev/null 2>&1; then
|
| 39 |
jq -e '.choices[0].message.content | type == "string" and length > 0' <<<"${response}" >/dev/null
|
|
@@ -42,4 +42,3 @@ else
|
|
| 42 |
fi
|
| 43 |
|
| 44 |
echo "Cyber-Frost readiness and minimal generation checks passed."
|
| 45 |
-
|
|
|
|
| 9 |
set +a
|
| 10 |
fi
|
| 11 |
|
| 12 |
+
: "${CYBER_FROST_BIND_HOST:=127.0.0.1}"
|
| 13 |
+
: "${CYBER_FROST_PORT:=8000}"
|
| 14 |
+
: "${CYBER_FROST_SERVED_MODEL_NAME:=CYBER-FROST-3.8-BF16}"
|
| 15 |
+
: "${CYBER_FROST_READINESS_TIMEOUT_SECONDS:=900}"
|
| 16 |
|
| 17 |
+
base_url="http://${CYBER_FROST_BIND_HOST}:${CYBER_FROST_PORT}/v1"
|
| 18 |
+
deadline=$((SECONDS + CYBER_FROST_READINESS_TIMEOUT_SECONDS))
|
| 19 |
|
| 20 |
until curl --fail --silent --show-error --max-time 5 "${base_url}/models" >/dev/null; do
|
| 21 |
if (( SECONDS >= deadline )); then
|
|
|
|
| 27 |
|
| 28 |
models_json="$(curl --fail --silent --show-error "${base_url}/models")"
|
| 29 |
if command -v jq >/dev/null 2>&1; then
|
| 30 |
+
jq -e --arg model "${CYBER_FROST_SERVED_MODEL_NAME}" '.data | any(.id == $model)' <<<"${models_json}" >/dev/null
|
| 31 |
fi
|
| 32 |
|
| 33 |
response="$(curl --fail --silent --show-error \
|
| 34 |
-H 'Content-Type: application/json' \
|
| 35 |
-X POST "${base_url}/chat/completions" \
|
| 36 |
+
-d "$(printf '{\"model\":\"%s\",\"messages\":[{\"role\":\"user\",\"content\":\"In an authorized lab, give three concise steps for triaging an unexpected listening service.\"}],\"temperature\":0,\"max_tokens\":96,\"chat_template_kwargs\":{\"enable_thinking\":false}}' "${CYBER_FROST_SERVED_MODEL_NAME}")")"
|
| 37 |
|
| 38 |
if command -v jq >/dev/null 2>&1; then
|
| 39 |
jq -e '.choices[0].message.content | type == "string" and length > 0' <<<"${response}" >/dev/null
|
|
|
|
| 42 |
fi
|
| 43 |
|
| 44 |
echo "Cyber-Frost readiness and minimal generation checks passed."
|
|
|
README.md
CHANGED
|
@@ -29,7 +29,7 @@ tags:
|
|
| 29 |
|
| 30 |
Cyber-Frost is a research release under active quality assessment. The repository is public and access is manually gated.
|
| 31 |
|
| 32 |
-
The release contains the complete BF16 SafeTensors checkpoint, weight index, configuration, tokenizer and processor assets, the packaged chat template, upstream license, this model card, and a
|
| 33 |
|
| 34 |
| Field | Released artifact |
|
| 35 |
|---|---|
|
|
|
|
| 29 |
|
| 30 |
Cyber-Frost is a research release under active quality assessment. The repository is public and access is manually gated.
|
| 31 |
|
| 32 |
+
The release contains the complete BF16 SafeTensors checkpoint, weight index, configuration, tokenizer and processor assets, the packaged chat template, upstream license, this model card, and a documented deployment kit for the validated serving profile. It is not an adapter and does not require a separate parent checkpoint at load time.
|
| 33 |
|
| 34 |
| Field | Released artifact |
|
| 35 |
|---|---|
|