| # LLM Prompt Injection Security Evaluation Framework |
|
|
| ## Project Overview |
|
|
| ### Project Type |
| - University research capstone project focused on LLM security. |
|
|
| ### Project Objective |
| - Design and implement a testing framework that evaluates how vulnerable LLM chatbots are to prompt injection attacks. |
| - Evaluate how mitigation strategies improve security outcomes. |
|
|
| The system will: |
| 1. Simulate prompt injection attacks. |
| 2. Measure how often attacks succeed. |
| 3. Apply mitigation strategies. |
| 4. Measure how much mitigation reduces successful attacks. |
|
|
| ## Scope Clarification |
|
|
| ### This Project Is NOT |
| - Training LLM models |
| - Building a new AI model |
| - Building a ChatGPT-like system |
| - Performing large-scale machine learning training |
|
|
| ### This Project IS |
| - Designing attack scenarios |
| - Evaluating LLM security vulnerabilities |
| - Implementing defensive mitigation strategies |
| - Measuring improvements after mitigation |
|
|
| ## Attack Scenarios to Evaluate |
| - Direct prompt injection through user inputs |
| - Attempts to extract hidden system prompts or internal instructions |
| - Indirect prompt injection through retrieved documents (RAG scenarios) |
| - Adaptive attackers who modify prompts after observing system responses |
|
|
| The framework should allow easily adding new attack strategies. |
|
|
| ## Mitigation Strategies |
| - Prompt hardening |
| - Input filtering |
| - Instruction isolation |
| - Detection mechanisms (e.g., classifier-based detection) |
|
|
| The framework should support comparing multiple mitigation approaches. |
|
|
| The system should also allow analyzing trade-offs such as: |
| - Security improvement |
| - Usability impact |
| - False positives or false negatives |
|
|
| ## Evaluation Metrics |
| The system should measure: |
| - Attack Success Rate (ASR) |
| - Reduction in successful attacks after mitigation |
| - False positive / false negative rates for detection methods |
| - Impact on normal system performance |
|
|
| ## Models |
| The framework may test different LLMs including: |
| - API-based models (e.g., GPT-family models) |
| - Open-source models (e.g., LLaMA, Mistral) |
|
|
| The architecture should support easily adding new models. |
|
|
| ## System Architecture (High Level) |
| The pipeline will conceptually look like this: |
|
|
| LLM Chatbot |
| β |
| Attack Generator |
| β |
| Attack Execution |
| β |
| Mitigation Layer |
| β |
| Evaluation Metrics |
| β |
| Results & Analysis |
|
|
| The framework should support running multiple experiments and logging results. |
|
|
| ## Repository Goals |
| The repository should prioritize: |
| - Modular architecture |
| - Reproducible experiments |
| - Clear documentation |
| - Experiment logging |
| - Easy extension for new attacks |
| - Easy extension for new mitigation strategies |
|
|
| ## Research Inspiration |
| This project is inspired by research on prompt injection attacks and LLM security, including academic papers that: |
| - Formalize prompt injection attacks |
| - Benchmark attack strategies |
| - Compare defense mechanisms |
|
|
| The goal is to build a practical experimental framework inspired by this research. |
|
|