Text Classification
Transformers
Safetensors
English
qwen3
prompt-injection
jailbreak-detection
jailbreak
moderation
security
guard
text-embeddings-inference
Instructions to use rogue-security/prompt-injection-jailbreak-sentinel-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use rogue-security/prompt-injection-jailbreak-sentinel-v2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="rogue-security/prompt-injection-jailbreak-sentinel-v2")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("rogue-security/prompt-injection-jailbreak-sentinel-v2") model = AutoModelForSequenceClassification.from_pretrained("rogue-security/prompt-injection-jailbreak-sentinel-v2", device_map="auto") - Notebooks
- Google Colab
- Kaggle
A little error
#1
by lhhvc - opened
I want to use the Sentinel-v2 model to build an LLM input detector, but during testing, I found that the model misclassifies "hello" as a dangerous input
Hi!
Thank you for testing! we found out one dataset that really biased against one word prompts, we fixed the issue and pushed a new revision 2 days ago. please pull the latest revision and try again. if the issue persists please LMK
Other then that how's the overall experience with the model?
Cheers.
Dror
Thank you for your reply. The model works very well and I will try the new version. If there are any other questions, I will share them with you in a timely manner😀