Xunzhuo commited on
Commit
7738100
·
verified ·
1 Parent(s): 571d1c7

Add original mmBERT evaluation comparison for Vela Guard

Browse files
Files changed (1) hide show
  1. README.md +11 -0
README.md CHANGED
@@ -28,6 +28,17 @@ Vela Guard detects prompt injection and jailbreak attempts in requests and untru
28
 
29
  Use Safety or Hazard for content risk.
30
 
 
 
 
 
 
 
 
 
 
 
 
31
  ## Quick start
32
 
33
  With PyTorch and Transformers 4.57.6 or 5.17.0:
 
28
 
29
  Use Safety or Hazard for content risk.
30
 
31
+ ## Evaluation
32
+
33
+ Compared with [the original mmBERT32K jailbreak detector](https://huggingface.co/llm-semantic-router/mmbert32k-jailbreak-detector-merged) for prompt-attack detection. Scores are on a 0–100 scale; higher is better.
34
+
35
+ | Development evaluation | Original mmBERT | Vela |
36
+ |---|---:|---:|
37
+ | Macro F1 · 1,319 inputs | 76.63 | **86.61** |
38
+ | Accuracy · 1,319 inputs | 76.65 | **86.66** |
39
+
40
+ The same development set combines prompt attacks, benign requests and controlled long contexts up to 32,768 tokens. Both models use FP32, complete inputs and the highest-scoring label. This set informed Vela development; it is not an independent blind benchmark.
41
+
42
  ## Quick start
43
 
44
  With PyTorch and Transformers 4.57.6 or 5.17.0: