Text Generation
Safetensors
GGUF
English
cybersecurity
security
defensive-security
vulnerability-detection
log-analysis
phishing-analysis
unsloth
qlora
conversational
Instructions to use skyuu72/Llama-Quantara-Sentinel-8B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use skyuu72/Llama-Quantara-Sentinel-8B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf skyuu72/Llama-Quantara-Sentinel-8B:Q4_K_M # Run inference directly in the terminal: llama cli -hf skyuu72/Llama-Quantara-Sentinel-8B:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf skyuu72/Llama-Quantara-Sentinel-8B:Q4_K_M # Run inference directly in the terminal: llama cli -hf skyuu72/Llama-Quantara-Sentinel-8B:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf skyuu72/Llama-Quantara-Sentinel-8B:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf skyuu72/Llama-Quantara-Sentinel-8B:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf skyuu72/Llama-Quantara-Sentinel-8B:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf skyuu72/Llama-Quantara-Sentinel-8B:Q4_K_M
Use Docker
docker model run hf.co/skyuu72/Llama-Quantara-Sentinel-8B:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use skyuu72/Llama-Quantara-Sentinel-8B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "skyuu72/Llama-Quantara-Sentinel-8B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "skyuu72/Llama-Quantara-Sentinel-8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/skyuu72/Llama-Quantara-Sentinel-8B:Q4_K_M
- Ollama
How to use skyuu72/Llama-Quantara-Sentinel-8B with Ollama:
ollama run hf.co/skyuu72/Llama-Quantara-Sentinel-8B:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use skyuu72/Llama-Quantara-Sentinel-8B with Docker Model Runner:
docker model run hf.co/skyuu72/Llama-Quantara-Sentinel-8B:Q4_K_M
- Lemonade
How to use skyuu72/Llama-Quantara-Sentinel-8B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull skyuu72/Llama-Quantara-Sentinel-8B:Q4_K_M
Run and chat with the model
lemonade run user.Llama-Quantara-Sentinel-8B-Q4_K_M
List all available models
lemonade list
- Atomic Chat
| # Llama-Quantara-Sentinel-8B v1.0 — Q4_K_M (training run v6) | |
| # Cisco Foundation-Sec native grammar, NOT llama-3.1. Unsloth reports | |
| # "No Ollama template mapping found" for this base and writes no Modelfile, | |
| # so this is hand-written on purpose. Get the stop tokens wrong and every | |
| # local user gets the v1/v2 runaway output back. | |
| FROM ./Llama-Quantara-Sentinel-8B-v1.0.Q4_K_M.gguf | |
| TEMPLATE """<|system|> | |
| {{ .System }} | |
| <|user|> | |
| {{ .Prompt }} | |
| <|assistant|> | |
| """ | |
| SYSTEM """You are Quantara Sentinel, a defensive cybersecurity assistant. You help users find, understand, and fix security weaknesses. You explain risks clearly and suggest safe fixes. You refuse to produce working exploits, malware, or instructions to attack systems the user does not own.""" | |
| PARAMETER temperature 0.3 | |
| PARAMETER top_p 0.9 | |
| PARAMETER num_ctx 8192 | |
| PARAMETER stop "<|end_of_text|>" | |
| PARAMETER stop "<|user|>" | |