prompt_injection (v3)

Decide whether a message sent to an AI assistant tries to hijack it: override its instructions, extract its hidden prompt or configuration, or make it ignore its rules.

A small classifier trained with nodd, exported as ONNX fp32 and int8 for browser and Node.js inference. This version uses a larger offline synthetic training dataset.

Label Meaning
safe A normal request, including ones that talk about prompts, AI or security, or correct the user's own earlier message.
injection Tries to override or cancel the assistant's instructions, reveal its system prompt or secrets, or remove its rules.

Evaluation

Metric Value
Test macro F1, PyTorch 0.9421
Test macro F1, int8 ONNX 0.9421
int8 vs PyTorch label agreement 100.00%
int8 download 18.58 MB
Train / validation / test 2237 / 52 / 52
Temperature 1.328877
Confidence threshold 0.865313
Validation precision at threshold 97.92%
Test accuracy on covered inputs, PyTorch 94.23%
Test coverage, PyTorch 100.00%
Training seed 42

See baseline comparison, evaluation report, and browser parity results.

Data and limitations

Added 1,000 examples per label using deterministic offline templates, with generation seed 20260928. Labels were assigned by construction, not by an independent teacher. Template variants are correlated and do not represent independent scenarios. New rows are training-only; the original validation and test records were preserved. Exact duplicates and examples with embedding cosine similarity above 0.90 to either holdout were rejected.

All evaluation inputs are synthetic; real-world accuracy is unmeasured. More data did not consistently improve held-out quality: consult the comparison before replacing an earlier version. The 97% precision target selected on validation was not met on this version's covered PyTorch test subset. Quantization can change individual probabilities substantially, so label parity does not establish confidence equivalence. Prompt-injection detection, where applicable, is not a standalone security boundary.

Use

Download this version and serve it from your own origin:

hf download nodd-repo/prompt-injection --revision v3 --local-dir public/models/prompt_injection/v3
npm install @nodd/browser
import { nodd } from "@nodd/browser";
const model = await nodd.load("/models/prompt_injection/v3");
const decision = await model.decide("Your input text");
if (!model.isConfident(decision)) { /* escalate for review */ }
model.dispose();

For Node.js, use @nodd/node and the downloaded directory path. Requires nodd JavaScript packages >= 0.2.0.

Artifacts

  • onnx/model_quantized.onnx, onnx/model.onnx: int8 and fp32 browser models.
  • nodd.json, tokenizer files and config.json: browser runtime configuration.
  • encoder/: original Transformers checkpoint, loadable with AutoModelForSequenceClassification.from_pretrained(local_path, subfolder="encoder").
  • labeled.jsonl: exact training, validation and test records with split and provenance fields.
  • synthetic_manifest.json: generation metadata and baseline fingerprint.
  • model_card.json, reports and parity files: training, calibration and evaluation evidence.

Earlier releases remain available through their version tags. The generation script and cross-task results are in the source repository.

Downloads last month
25
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nodd-repo/prompt-injection

Quantized
(5)
this model