--- license: apache-2.0 language: - en library_name: gguf pipeline_tag: text-generation tags: - security - vulnerability-detection - agentic - terminal-agent - reinforcement-learning - granite - gguf - lm-studio - ollama - llama-cpp - quantized base_model: fdtn-ai/antares-1b --- # Antraes-1B (GGUF) Community GGUF quantizations of [Cisco's Antares-1B](https://huggingface.co/fdtn-ai/antares-1b) — an open-weight 1-billion parameter language model specialized for **vulnerability localization** in real-world codebases. Built on IBM Granite 4.0 1B, it autonomously navigates source code repositories through a terminal interface using a two-stage SFT + GRPO pipeline. These GGUF files enable running Antares-1B locally with [llama.cpp](https://github.com/ggml-org/llama.cpp), [LM Studio](https://lmstudio.ai/), [Ollama](https://ollama.com/), [Jan](https://jan.ai/), [Open WebUI](https://openwebui.com/), and other GGUF-compatible inference engines. ## Model Details | Property | Value | |----------|-------| | **Model Developer** | Cisco Systems, Inc. — Foundation AI ([fdtn-ai/antares-1b](https://huggingface.co/fdtn-ai/antares-1b)) | | **Architecture** | Auto-regressive decoder-only transformer (IBM Granite 4.0 1B backbone) | | **Parameters** | 1B | | **Layers** | 40, hidden dim 2048, 16 attention heads, 4 KV heads (GQA) | | **Context Window** | 128K tokens | | **Vocabulary** | 100,352 tokens | | **Activation** | SwiGLU, RMSNorm, RoPE positional encoding | | **License** | Apache 2.0 | | **Technical Report** | [Antares: Foundation Models for Agentic Vulnerability Localization](https://cisco-foundation-ai.github.io/antares/technical-report.pdf) | | **Training** | SFT on cybersecurity reasoning + terminal-navigation data, followed by GRPO with verifiable rewards over multi-turn agent trajectories | ## Performance On the Vulnerability Localization Benchmark (VLoc Bench, 500 tasks), Antares-1B achieves a **File F1 of 0.209**, outperforming models many times its size including GLM-5.2, Gemini 3 Pro, GPT-5 Mini, and Qwen3.5-122B. | Model | Parameters | File F1 | Precision | Recall | |-------|-----------|---------|-----------|--------| | **Antares-1B (GRPO)** | **1B** | **0.209** | **0.262** | **0.224** | | GLM-5.2 | 753B | 0.186 | 0.226 | 0.186 | | Gemini 3 Pro | Frontier | 0.152 | 0.190 | 0.153 | | Qwen3.5-122B-A10B | 125B MoE | 0.091 | 0.124 | 0.083 | ## Files | File | Format | Size | Description | |------|--------|------|-------------| | `antraes-1b-f16.gguf` | GGUF F16 | ~3.4 GB | Full-precision GGUF quantization | | `model.safetensors` | SafeTensors | ~3.4 GB | Original PyTorch weights | ## Usage ### llama.cpp ```bash # Build or download llama.cpp, then run: ./llama-cli -m antraes-1b-f16.gguf \ --temp 0.3 \ --top-p 1.0 \ --ctx-size 8192 \ -p "<|start_of_role|>system<|end_of_role|>You are a security vulnerability localization agent...<|end_of_text|>\n<|start_of_role|>user<|end_of_role|>Vulnerability to locate:\nCWE-78: OS Command Injection<|end_of_text|>\n<|start_of_role|>assistant<|end_of_role|>" ``` ### LM Studio 1. Download `antraes-1b-f16.gguf` 2. Open LM Studio → drag the `.gguf` file into the model directory 3. Select "Antraes-1B" from the model dropdown 4. This model uses the Granite chat template (auto-detected) ### Ollama ```bash # Create a Modelfile: FROM ./antraes-1b-f16.gguf TEMPLATE """{{ .System }} {{ .Prompt }}""" # Then: ollama create antraes-1b -f Modelfile ollama run antraes-1b ``` ### Jan 1. Download `antraes-1b-f16.gguf` 2. Open Jan → Settings → Extensions → enable "GGUF Engine" 3. Drag the `.gguf` file into Jan's model folder 4. The model will appear in your model list ### Open WebUI If using Ollama as the backend, pull the model into Ollama first (see above), then it will be available in Open WebUI. ## Prompt Format The model uses a structured tool-calling format with special tokens: ``` <|start_of_role|>system<|end_of_role|>...<|end_of_text|> <|start_of_role|>user<|end_of_role|>...<|end_of_text|> <|start_of_role|>assistant<|end_of_role|>...reasoning... {"name": "terminal", "arguments": {"command": "..."}} <|end_of_text|> <|start_of_role|>user<|end_of_role|> ...output... <|end_of_text|> <|start_of_role|>assistant<|end_of_role|>... ``` The agent loop terminates when the model calls `submit_vulnerable_files` or `submit_no_vulnerability_found`. ## Intended Use - **Vulnerability Localization**: Given a CWE identifier and a repository, identifying which source files contain the reported vulnerability using only a terminal interface - **Shift-Left Security**: Integrating into CI/CD pipelines for early vulnerability detection - **Advisory-Driven Triage**: Using CWE identifiers mapped from CVE/GHSA records to guide repository exploration See the [original model card](https://huggingface.co/fdtn-ai/antares-1b) for full details on intended/out-of-scope use, limitations, safety, and evaluation methodology. ## Disclaimer This is an **unofficial community GGUF conversion** of Cisco's Antares-1B model. These files are not affiliated with or endorsed by Cisco Systems, Inc. The original model is licensed under Apache 2.0 by Cisco Systems, Inc.