| # Attack Taxonomy v1 |
|
|
| This taxonomy is a compact, implementation-oriented consolidation of the `attack_family` values already present in `data/normalized/attack_cases.jsonl`. It keeps source-specific `attack_subtype` values intact while grouping cases into a small canonical set that is easier to track, enrich, and evaluate. |
|
|
| ## Canonical Categories |
|
|
| | Canonical category | Typical mapped families | Intended use | |
| |---|---|---| |
| | `instruction_override` | `direct_instruction_override`, `instruction_override`, some `compound_instruction_attack` cases | Direct attempts to replace or dominate system/task instructions | |
| | `prompt_leakage` | `prompt_leakage`, `system_prompt_extraction` | Attempts to disclose hidden prompts, secrets, or internal policy | |
| | `retrieved_indirect` | `retrieved_content_injection` | Indirect or tool-mediated attacks from documents, email, calendar, drive, etc. | |
| | `evasion_obfuscation` | `blacklist_evasion`, `xml_escape_evasion`, `restricted_character_bypass`, `emoji_only_bypass`, `sandwich_defense_bypass` | Obfuscation or bypass techniques meant to evade simple defenses | |
| | `adaptive_multiturn` | `adaptive_attack` | Iterative or staged attacks that adapt over multiple turns | |
| | `benign_control` | `benign_control` | Negative controls for false-positive and utility measurement | |
|
|
| This canonical set is now the frozen project taxonomy for active benchmark |
| mapping and should not be casually expanded during implementation: |
|
|
| - `instruction_override` |
| - `prompt_leakage` |
| - `retrieved_indirect` |
| - `evasion_obfuscation` |
| - `adaptive_multiturn` |
| - `benign_control` |
|
|
| ## Mapping Principles |
|
|
| - Preserve the original `attack_family` and `attack_subtype`. |
| - Add `attack_category` as the canonical project-level taxonomy label. |
| - Add `success_definition_id` so evaluation logic can later bind to stable success criteria. |
| - Add numeric `severity` while preserving the original textual `severity_level`. |
| - Keep provenance fields visible so downstream analysis can trace dataset origin and benchmark split. |
|
|
| ## Why This Structure |
|
|
| - It matches the methodology direction already described in the project docs. |
| - It is small enough to be stable across datasets. |
| - It is explicit enough to support enrichment scripts, auditability, and later mitigations without forcing a runtime rewrite. |
|
|
| ## Current Active Benchmark Mapping |
|
|
| Current active source coverage in the enriched corpus is: |
|
|
| - `HackAPrompt` |
| - maps to `instruction_override` |
| - maps to `prompt_leakage` |
| - maps to `evasion_obfuscation` |
| - maps to `adaptive_multiturn` |
| - `TensorTrust` |
| - `hijacking` maps to `instruction_override` |
| - `extraction` maps to `prompt_leakage` |
| - active in the normalized experiment corpus |
|
|
| `AgentDojo` remains a benchmark reference/example source and still maps to |
| `retrieved_indirect`, but it is not part of the current active experiment |
| corpus because the extracted set is too small and partial. |
|
|
| This means the active corpus already exercises multiple canonical categories, |
| but not yet the full benchmark landscape. Future expansion should preserve this |
| canonical taxonomy rather than introduce source-specific top-level categories. |
|
|