Spaces:
Running on Zero
Download README.md from saibafatima/shadowsage-ai: direct link, hf CLI and curl.
- Browser
- Download file 5.51 kB
-
https://huggingface.co/spaces/saibafatima/shadowsage-ai/resolve/main/README.md
- Command line
-
hf download hf://spaces/saibafatima/shadowsage-ai/README.md
-
curl -L -o README.md https://huggingface.co/spaces/saibafatima/shadowsage-ai/resolve/main/README.md
A newer version of the Gradio SDK is available: 6.29.0
title: ShadowSage AI
emoji: ποΈ
colorFrom: purple
colorTo: indigo
sdk: gradio
sdk_version: 6.26.0
app_file: app.py
pinned: false
license: mit
ShadowSage AI β the all-seeing guardian of your digital world
An AI-powered cybersecurity oracle that defends proactively, not reactively. ShadowSage detects hidden threats before they land, analyzes suspicious messages, identifies phishing, uncovers your exposed digital footprint, and watches login behavior for signs of account takeover β then distills everything into one Sage Score with plain-English guidance.
Built for CPU-only live demos: small models, no API keys required, everything degrades gracefully to a labeled fallback instead of crashing.
The four scrying stones
| Tab | What it does | How |
|---|---|---|
| Sage Verdict | One combined Sage Score + 2-3 personalized recommendations | Weighted average of the three module scores (formula in utils/scoring.py) |
| Phishing Analyzer | Risk score, verdict and highlighted evidence for any email / SMS / URL | DistilBERT fine-tuned on a bundled 284-example dataset plus transparent rule + URL heuristics (typosquats, raw-IP hosts, abuse TLDs, entropy, urgency patterns) |
| Footprint Scanner | Breach history + exposure score for an email or domain | HaveIBeenPwned API v3 when HIBP_API_KEY is set; otherwise a clearly-labeled deterministic demo dataset |
| Behavior Monitor | Login timeline with anomalies highlighted + behavioral risk score | Isolation Forest over data/login_activity.csv (90 days of synthetic history with 7 planted intrusions), features include circular hour encoding, rare country/device flags and impossible-travel velocity |
Project structure
app.py # Gradio Blocks UI β the whole mystical dashboard
model/phishing_classifier.py # DistilBERT training + rule/URL fallback engine
model/behavior_monitor.py # Isolation Forest behavioral monitor
utils/footprint_scanner.py # HaveIBeenPwned v3 wrapper + demo fallback
utils/scoring.py # Sage Score formula + recommendation engine
data/phishing_dataset.csv # 284 labeled examples (1 = phishing)
data/login_activity.csv # 130 login events, 7 planted anomalies
data/generate_login_data.py # deterministic generator for the login history
Run locally
python -m venv .venv
# Windows: .venv\Scripts\activate | macOS/Linux: source .venv/bin/activate
pip install -r requirements.txt
python app.py # -> http://127.0.0.1:7860
On the very first launch there is no trained model yet, so the app starts in rule-engine mode (fully functional) and fine-tunes DistilBERT in a background thread β the Phishing tab shows the training status ("the Sage is awakening"). Analysis is never blocked.
Deploy to Hugging Face Spaces
- Create a new Space β SDK: Gradio.
- Push this repo as-is (
app.pystays at the root; the YAML block at the top of this README is the Spaces config). - Done.
requirements.txtinstalls everything; on first boot the Space trains the phishing model in the background and shows rule-based results meanwhile.
Optional: add a Space secret HIBP_API_KEY to switch the Footprint Scanner from
demo data to live HaveIBeenPwned lookups. Nothing else changes.
Retraining the phishing model on your own dataset
Replace data/phishing_dataset.csv (two columns: text,label, label 1 =
phishing), then either:
# CLI: train and save to model_artifacts/phishing_distilbert/
python -m model.phishing_classifier # full 2-epoch run
python -m model.phishing_classifier --limit 60 # quick smoke-train
β¦or simply delete the model_artifacts/ folder and start the app: it retrains
itself in the background. The entry points to look at first are
train() in model/phishing_classifier.py (hyperparameters are the constants
at the top of that file) and the retraining notes in its module docstring.
How the scores work (viva crib sheet)
- Phishing risk β ML mode blends
0.65 Γ P(phishing from DistilBERT)with0.35 Γ URL-heuristic scorewhen a link is present; rule mode blends keyword and URL scores 60/40. In ML mode the rule score is a guardrail: if the transparent heuristics are more alarmed than the model (a blind spot from the tiny dataset), the higher score wins. Every added point is listed in the UI. - Footprint exposure β +16 per breach containing password data, +8 per data-only breach, +6 extra for breaches newer than ~3 years, capped at 100.
- Behavioral risk β
0.5 Γ peak anomaly severity + 0.25 Γ anomaly density + 0.25 Γ recency; the detector is fully unsupervised but the CSV carries aknown_anomalycolumn so the UI can show precision/recall against the planted truth (currently 6 of 7 caught). - Sage Score β
0.40 Γ phishing + 0.35 Γ footprint + 0.25 Γ behavior, renormalized over whichever modules have run.
Regenerating the demo data
python data/generate_login_data.py # deterministic (seed 42)
Deployment note
The primary target is Hugging Face Spaces (Gradio SDK), where this repo runs
with zero changes. A Vercel deployment would need a serverless-friendly
rewrite (Gradio on Node/Vercel is not supported; you would front the same
Python modules with a FastAPI + static UI) β out of scope for the hackathon,
but the model/ and utils/ packages are UI-agnostic and would carry over
unchanged.