--- license: mit pipeline_tag: tabular-classification datasets: - feti-ai/phiusiil-if3070-stei-itb-2024-2025-1 tags: - phishing-detection - url-classification - tabular-classification - binary-classification - phiusiil - knn - from-scratch model-index: - name: fetiai-v1-phiusiil-binclf-knn-scratch-500k results: - task: type: tabular-classification name: Phishing URL Classification dataset: type: feti-ai/phiusiil-if3070-stei-itb-2024-2025-1 name: PhiUSIIL Phishing URL Dataset metrics: - type: accuracy value: 0.9804494141946511 name: Accuracy - type: precision value: 0.9813813813813814 name: Phishing precision - type: recall value: 0.7592936802973977 name: Phishing recall ---
# fetiai-v1-phiusiil-binclf-knn-scratch-500k **Group 16** · IF3070 Foundations of Artificial Intelligence · STEI ITB

IF3070 STEI ITB 2024/2025-1 Group 16

tabular classification 500,001 parameters PhiUSIIL MIT

**KNN (from scratch)** for phishing URL classification, served over HTTP as a single-model API. ```console $ curl -s localhost:8000/predict -H 'content-type: application/json' -d '{ "features": {"URLLength": 31, "DomainLength": 25, "IsHTTPS": 1, "...": "all 49"} }' { "model": "knn_scratch", "label": 1, "verdict": "legitimate", "phishing_score": 0.0, "n_provided": 33, "n_imputed": 16, "coverage_ratio": 0.6735, "low_evidence": false } ``` > **This is a coursework reimplementation, not a security product.** It is trained on a > static 2023–24 dataset, has no threat intelligence, no blocklist, and no knowledge of any > campaign newer than its training data. Do not use it to decide whether a link is safe. **Algorithm** `KNN (from scratch)` · `Trained on PhiUSIIL` · `SMOTE` · `Feature Engineering` **Built with**

Python   NumPy   pandas   Docker

**Links** — [Full application](https://github.com/fetiai/phishing-url-classifier) · [Live demo](https://phiusiil.faizath.com) · [Dataset](https://archive.ics.uci.edu/dataset/967/phiusiil+phishing+url+dataset) --- ## What it does It scores **one pre-extracted feature row** with one model and returns a verdict. That is the whole scope, and the boundary is deliberate. This service has no page fetcher, no URL feature extractor and no SSRF guard, because duplicating a network-facing security control into four repositories is how the four copies drift apart. Turning a URL into the 49 features this API expects is the job of the [full application](https://github.com/fetiai/phishing-url-classifier), which fetches the page under a guard, extracts the features, and scores them against all four models at once. So this is a **tabular** classifier, not a text one: the input is a 49-dimensional feature vector, and the URL string never reaches the model. | Route | Purpose | |---|---| | `POST /predict` | one row → verdict, score, and how much of the row was real | | `POST /predict/batch` | up to 1000 rows | | `GET /metadata` | the feature contract, the demoted features, and this model's metrics | | `GET /healthz` | liveness | | `GET /readyz` | ready once the golden-row self-test has passed | ## Quickstart ```bash make install # venv + pinned dependencies make selftest # prove the artifact reproduces its recorded prediction make serve # http://127.0.0.1:8000 (docs at /docs) ``` No training step and no download step: **the model is committed** (1.0 MB under `model/`), so a fresh clone can serve immediately. ### With Docker ```bash make docker-build && docker run -p 8000:8000 fetiai-v1-phiusiil-binclf-knn-scratch-500k:local ``` The image **bakes the model in** rather than mounting it, so the image tag is a complete description of what the service will predict. The build fails if the artifact does not reproduce its own golden row. ## The request contract Send all 49 feature columns. `null` is allowed and is the expected value for a feature the caller could not determine — 12 of the 49 are permanently null, having failed the extraction agreement gate in the parent project. ```jsonc { "features": { "URLLength": 31, "DomainLength": 25, "IsHTTPS": 1, /* ...46 more */ }, "url": "https://example.com/login", // optional "domain": "example.com", // optional "tld": "com", // optional "title": "Sign in" // optional } ``` `GET /metadata` returns the exact 49 names in order. Three of them carry typos that are preserved on purpose — `NoOfDegitsInURL`, `DegitRatioInURL`, `SpacialCharRatioInURL` — because those names are what the training data means. A **missing** feature is rejected by name; a **null** one is imputed. That distinction is the point: an extractor that dropped a column and an extractor that honestly could not determine a value are different faults with different fixes. ### Why the response reports coverage Imputation fills a missing feature from the training distribution, and that distribution is 92.5% legitimate. A mostly-empty row therefore does not produce a *neutral* prediction — it produces one biased toward **legitimate**, which is exactly the wrong direction for a phishing detector. So every response carries `n_provided`, `n_imputed`, `coverage_ratio` and `low_evidence`. Without them a verdict drawn from six real values looks precisely as confident as one drawn from all 49. The optional `url`, `domain`, `tld` and `title` fields are worth sending when you have them: 21 of the 49 features are derived from the URL string, and supplying it lets those be **recomputed** rather than imputed. ## Results Measured on a 28,081-row validation split. **The held-out file shipped with the dataset has no labels, so there is no test score and none is claimed.** | Model | Phishing recall | Phishing precision | Accuracy | |---|---|---|---| | **KNN (from scratch)** | **0.759** | **0.981** | **0.98045** | **Read that accuracy against 0.9248.** The corpus is 92.48% legitimate, so answering "legitimate" to everything scores 0.9248 while catching no phishing whatsoever. Accuracy alone cannot tell a working detector from a constant; phishing recall can. Class 0 is phishing and is the positive class throughout. The scikit-learn and from-scratch implementations of this algorithm disagree on **0.0356%** of the validation split. That number is the reason both exist: a reimplementation with nothing to check it against is an assertion, not a result. The counterpart lives in [`fetiai-v1-phiusiil-binclf-knn-skl-500k`](../fetiai-v1-phiusiil-binclf-knn-skl-500k). `model/metrics.json` also carries the `legacy` profile, flagged `"leaky": true`. It reconstructs the original notebook's configuration, which standardised each split by its own mean and standard deviation — information no deployed model can have, since there is no batch to average over when a single row arrives. It is kept as evidence of what the leak was worth and is never presented as this model's result. ## What is in this repo | Path | What it is | |---|---| | `model/knn_scratch.npz` | the trained model (500,001 stored values) | | `model/fitted_stats.json` | **not optional** — the scaler, imputation values, clip bounds and mode tables | | `model/manifest.json` | sha256 of every file above, verified at load | | `model/golden_row.json` | one record with its expected vector and prediction | | `phiusiil/` | the scoring path: schema, preprocessing, and this one model class | | `server/` | loader, prediction, HTTP layer | `fitted_stats.json` deserves the emphasis. The model alone cannot classify anything: it was fitted on standardised inputs, and the numbers that produce that standardisation live in that file. Publishing weights without it would be publishing something unusable. It is plain JSON rather than a pickled transformer on purpose. A pickled estimator arrives with a `fit` method attached, and the defect this whole pipeline exists to avoid is someone calling it at serving time. Numbers that cannot be re-fitted cannot leak. The model itself is equally open — a NumPy `.npz` of the reference matrix and its labels. Loading it executes no pickle and depends on no particular library version, which is the practical difference between this repo and its scikit-learn counterpart in [`fetiai-v1-phiusiil-binclf-knn-skl-500k`](../fetiai-v1-phiusiil-binclf-knn-skl-500k). ## Verification ```bash make selftest # golden row, offline make test # golden row + HTTP contract + naming make namecheck # provenance hygiene ``` The check that carries the weight is the **golden row**: one real record, its expected 49-feature vector, and its expected prediction, pushed through the whole path and compared **exactly**. The vector and the prediction are asserted separately, because a wrong vector means the preprocessing drifted while a right vector with a wrong label means the model artifact did — different faults, different fixes. Equality is bitwise on float32, never a tolerance. A tolerance-based comparison would pass while a fitted statistic quietly differed, which is the one thing the test exists to catch. `origin.json` records where every copied file came from, including the parent bundle's own hashes, so drift is detectable without the parent repository present. ## Licence and data MIT, as in the parent project. See `LICENSE`. This model was trained in part on the PhiUSIIL Phishing URL Dataset (Prasad & Chandra), available from the UCI Machine Learning Repository, licensed under CC BY 4.0. The dataset is the UCI PhiUSIIL Phishing URL Dataset (ID 967). **This artifact embeds training data.** k-nearest neighbours has no learned parameters — fitting is memorising — so `model/knn_scratch.npz` *is* the 10,000-row scaled reference set that the classifier searches at prediction time. The dataset's attribution therefore travels with this model, not only with the dataset. --- ## Team
Thalita Zahra Sutejo

Thalita Zahra Sutejo
18222023

thalitazhrr
thalitazahras
Irfan Musthofa

Irfan Musthofa
18222056

IrfanMusthofa
irfanmusthofa
Eleanor Cordelia

Eleanor Cordelia
18222059

EleanorCordelia
eleanorcordelia
Muhammad Faiz Atharrahman

Muhammad Faiz Atharrahman
18222063

faizath
faizath
---
IF3070 Foundations of Artificial Intelligence · STEI ITB · 2024/2025-1 More at **[fetiai.github.io](https://fetiai.github.io/)**