LLM Prompt Injection Security Evaluation Framework
Project Overview
Project Type
- University research capstone project focused on LLM security.
Project Objective
- Design and implement a testing framework that evaluates how vulnerable LLM chatbots are to prompt injection attacks.
- Evaluate how mitigation strategies improve security outcomes.
The system will:
- Simulate prompt injection attacks.
- Measure how often attacks succeed.
- Apply mitigation strategies.
- Measure how much mitigation reduces successful attacks.
Scope Clarification
This Project Is NOT
- Training LLM models
- Building a new AI model
- Building a ChatGPT-like system
- Performing large-scale machine learning training
This Project IS
- Designing attack scenarios
- Evaluating LLM security vulnerabilities
- Implementing defensive mitigation strategies
- Measuring improvements after mitigation
Attack Scenarios to Evaluate
- Direct prompt injection through user inputs
- Attempts to extract hidden system prompts or internal instructions
- Indirect prompt injection through retrieved documents (RAG scenarios)
- Adaptive attackers who modify prompts after observing system responses
The framework should allow easily adding new attack strategies.
Mitigation Strategies
- Prompt hardening
- Input filtering
- Instruction isolation
- Detection mechanisms (e.g., classifier-based detection)
The framework should support comparing multiple mitigation approaches.
The system should also allow analyzing trade-offs such as:
- Security improvement
- Usability impact
- False positives or false negatives
Evaluation Metrics
The system should measure:
- Attack Success Rate (ASR)
- Reduction in successful attacks after mitigation
- False positive / false negative rates for detection methods
- Impact on normal system performance
Models
The framework may test different LLMs including:
- API-based models (e.g., GPT-family models)
- Open-source models (e.g., LLaMA, Mistral)
The architecture should support easily adding new models.
System Architecture (High Level)
The pipeline will conceptually look like this:
LLM Chatbot
↓
Attack Generator
↓
Attack Execution
↓
Mitigation Layer
↓
Evaluation Metrics
↓
Results & Analysis
The framework should support running multiple experiments and logging results.
Repository Goals
The repository should prioritize:
- Modular architecture
- Reproducible experiments
- Clear documentation
- Experiment logging
- Easy extension for new attacks
- Easy extension for new mitigation strategies
Research Inspiration
This project is inspired by research on prompt injection attacks and LLM security, including academic papers that:
- Formalize prompt injection attacks
- Benchmark attack strategies
- Compare defense mechanisms
The goal is to build a practical experimental framework inspired by this research.