g5-eval-dashboard / docs /MITIGATION_LAYER.md
juandavidrodriguezar's picture
Upload folder using huggingface_hub
eaae571 verified
|
Raw
History Blame
4.69 kB

Mitigation Layer Guide

This guide describes the mitigation layer as it exists in the current framework, including where it runs, what it does, and how teammates should use it in experiments.

Purpose

The mitigation layer adds the defensive part of the framework.

It is responsible for:

  • checking risky attack inputs before they reach the model
  • hardening or isolating suspicious content
  • blocking clear prompt-injection or extraction attempts when appropriate
  • checking outputs before final scoring
  • recording mitigation actions in run artifacts for later analysis

Where It Runs

The runtime sequence is:

  1. Load and map cases
  2. Build the model request
  3. Apply request mitigation
  4. Call the model unless the request is blocked
  5. Apply response mitigation
  6. Score the final response
  7. Save mitigation metadata in the results

If a case contains ordered turns, mitigation is applied turn by turn and the runner stores per-turn artifacts as well.

Supported Strategies

none

Baseline mode with no defense applied.

prompt_hardening

Adds a stronger safety context so the model treats user input as untrusted and prioritizes the trusted task.

instruction_isolation

Wraps suspicious content as clearly untrusted data so the model is less likely to treat it as high-priority instruction.

keyword_guardrail

Blocks obvious prompt-injection or prompt-extraction patterns before the model is called.

surface_aligned

This is the main defended strategy for the project.

It routes behavior by attack surface and category:

  • benign controls are left alone to reduce unnecessary blocking
  • direct attacks get prompt hardening and direct guardrails
  • tool-mediated / indirect content is isolated as untrusted
  • adaptive cases use the same guarded path turn by turn
  • prompt-leakage style outputs can be filtered after inference

Request-Side vs Response-Side Mitigation

Request-side mitigation

This runs before inference.

Typical actions:

  • passthrough
  • transformed request
  • blocked request

Saved fields:

  • mitigation_strategy
  • mitigation_action
  • mitigation_blocked
  • mitigation_sanitized
  • mitigation_notes
  • mitigation_metadata

Response-side mitigation

This runs after inference but before evaluation.

Typical use:

  • prompt leakage filtering
  • output redaction or blocking for suspicious unsafe content

Saved fields:

  • response_mitigation_action
  • response_mitigation_modified
  • response_mitigation_blocked
  • response_mitigation_notes
  • response_mitigation_metadata

Multi-Turn Support

The runner supports an optional turns field in attack cases.

Current behavior:

  • turns run sequentially
  • prior conversation is preserved as history
  • mitigation is applied on each turn
  • final scoring uses the final effective response
  • turn artifacts are saved for later review

This makes adaptive cases possible without introducing a full persistent agent runtime.

Configurable Options

The mitigation.options block can tune behavior without changing code.

strict_mode

When enabled, the layer blocks more suspicious patterns instead of only transforming the request.

sanitize_untrusted_content

Redacts obvious payload fragments from untrusted content before inference.

block_suspicious_tool_actions

Blocks tool-style instructions that look like unsafe external actions.

guard_tool_outputs

Guards risky output content for tool-mediated scenarios.

allow_urls_in_output

Allows future benign scenarios to return links without being over-filtered.

Current Strengths

The current mitigation layer is in strong shape for the capstone framework:

  • fully integrated into the runner
  • documented and configurable
  • works on both single-turn and lightweight multi-turn cases
  • supports before-vs-after experiment design
  • records enough metadata for team analysis
  • showed real improvement on the HackAPrompt override slice during Groq runs

Current Limits

The mitigation layer is not a complete security solution.

Remaining gaps include:

  • stronger handling for some TensorTrust hijacking patterns
  • richer semantic detection for heavily obfuscated attacks
  • broader benign-control coverage for false-positive analysis
  • persistent session memory for longer adaptive attacks
  • real tool permission checks and approval gates

Recommended Team Usage

  • Use strategy: none for baseline runs
  • Use strategy: surface_aligned for defended runs
  • Compare results per category, surface, and source
  • Treat the mitigation layer as a validated capstone defense component, while staying honest that some slices still benefit from refinement