Attack Taxonomy v1
This taxonomy is a compact, implementation-oriented consolidation of the attack_family values already present in data/normalized/attack_cases.jsonl. It keeps source-specific attack_subtype values intact while grouping cases into a small canonical set that is easier to track, enrich, and evaluate.
Canonical Categories
| Canonical category | Typical mapped families | Intended use |
|---|---|---|
instruction_override |
direct_instruction_override, instruction_override, some compound_instruction_attack cases |
Direct attempts to replace or dominate system/task instructions |
prompt_leakage |
prompt_leakage, system_prompt_extraction |
Attempts to disclose hidden prompts, secrets, or internal policy |
retrieved_indirect |
retrieved_content_injection |
Indirect or tool-mediated attacks from documents, email, calendar, drive, etc. |
evasion_obfuscation |
blacklist_evasion, xml_escape_evasion, restricted_character_bypass, emoji_only_bypass, sandwich_defense_bypass |
Obfuscation or bypass techniques meant to evade simple defenses |
adaptive_multiturn |
adaptive_attack |
Iterative or staged attacks that adapt over multiple turns |
benign_control |
benign_control |
Negative controls for false-positive and utility measurement |
This canonical set is now the frozen project taxonomy for active benchmark mapping and should not be casually expanded during implementation:
instruction_overrideprompt_leakageretrieved_indirectevasion_obfuscationadaptive_multiturnbenign_control
Mapping Principles
- Preserve the original
attack_familyandattack_subtype. - Add
attack_categoryas the canonical project-level taxonomy label. - Add
success_definition_idso evaluation logic can later bind to stable success criteria. - Add numeric
severitywhile preserving the original textualseverity_level. - Keep provenance fields visible so downstream analysis can trace dataset origin and benchmark split.
Why This Structure
- It matches the methodology direction already described in the project docs.
- It is small enough to be stable across datasets.
- It is explicit enough to support enrichment scripts, auditability, and later mitigations without forcing a runtime rewrite.
Current Active Benchmark Mapping
Current active source coverage in the enriched corpus is:
HackAPrompt- maps to
instruction_override - maps to
prompt_leakage - maps to
evasion_obfuscation - maps to
adaptive_multiturn
- maps to
TensorTrusthijackingmaps toinstruction_overrideextractionmaps toprompt_leakage- active in the normalized experiment corpus
AgentDojo remains a benchmark reference/example source and still maps to
retrieved_indirect, but it is not part of the current active experiment
corpus because the extracted set is too small and partial.
This means the active corpus already exercises multiple canonical categories, but not yet the full benchmark landscape. Future expansion should preserve this canonical taxonomy rather than introduce source-specific top-level categories.