# Master ANOVA Table

## optimisation/claude-opus-4.6_optimization_all_abms_chain/06_quantitative_smoke_latest

| Reference family | Variable / metric | BLEU | METEOR | R-1 | R-2 | R-L | Reading ease |
| --- | --- | --- | --- | --- | --- | --- | --- |
| author | Agent-Based Model | <0.01 | 0.08 | 0.07 | <0.01 | 0.04 | 0.85 |
| author | Summarization algorithm | 0.02 | 0.07 | <0.01 | 0.02 | <0.01 | <0.01 |
| author | Simulation evidence | — | — | — | — | — | — |
| author | LLM | — | — | — | — | — | — |
| author | Use of roles | — | — | — | — | — | — |
| author | Generating insights | — | — | — | — | — | — |
| author | Providing examples | — | — | — | — | — | — |
| gpt5.2_long | Agent-Based Model | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | 0.58 |
| gpt5.2_long | Summarization algorithm | — | — | — | — | — | — |
| gpt5.2_long | Simulation evidence | — | — | — | — | — | — |
| gpt5.2_long | LLM | — | — | — | — | — | — |
| gpt5.2_long | Use of roles | — | — | — | — | — | — |
| gpt5.2_long | Generating insights | — | — | — | — | — | — |
| gpt5.2_long | Providing examples | — | — | — | — | — | — |
| gpt5.2_short | Agent-Based Model | 0.13 | <0.01 | 0.36 | <0.01 | 0.08 | 0.85 |
| gpt5.2_short | Summarization algorithm | <0.01 | 0.02 | <0.01 | <0.01 | <0.01 | <0.01 |
| gpt5.2_short | Simulation evidence | — | — | — | — | — | — |
| gpt5.2_short | LLM | — | — | — | — | — | — |
| gpt5.2_short | Use of roles | — | — | — | — | — | — |
| gpt5.2_short | Generating insights | — | — | — | — | — | — |
| gpt5.2_short | Providing examples | — | — | — | — | — | — |
| modeler | Agent-Based Model | <0.01 | <0.01 | 0.13 | <0.01 | 0.02 | 0.85 |
| modeler | Summarization algorithm | <0.01 | <0.01 | <0.01 | 0.11 | <0.01 | <0.01 |
| modeler | Simulation evidence | — | — | — | — | — | — |
| modeler | LLM | — | — | — | — | — | — |
| modeler | Use of roles | — | — | — | — | — | — |
| modeler | Generating insights | — | — | — | — | — | — |
| modeler | Providing examples | — | — | — | — | — | — |

## optimisation/gemini-3.1-pro-preview_optimization_all_abms_chain/06_quantitative_smoke_latest

| Reference family | Variable / metric | BLEU | METEOR | R-1 | R-2 | R-L | Reading ease |
| --- | --- | --- | --- | --- | --- | --- | --- |
| author | Agent-Based Model | 0.01 | 0.17 | 0.22 | <0.01 | 0.09 | 0.55 |
| author | Summarization algorithm | 0.18 | <0.01 | 0.07 | 0.04 | 0.10 | <0.01 |
| author | Simulation evidence | — | — | — | — | — | — |
| author | LLM | — | — | — | — | — | — |
| author | Use of roles | — | — | — | — | — | — |
| author | Generating insights | — | — | — | — | — | — |
| author | Providing examples | — | — | — | — | — | — |
| gpt5.2_long | Agent-Based Model | <0.01 | 1.00 | <0.01 | <0.01 | <0.01 | 0.20 |
| gpt5.2_long | Summarization algorithm | — | — | — | — | — | — |
| gpt5.2_long | Simulation evidence | — | — | — | — | — | — |
| gpt5.2_long | LLM | — | — | — | — | — | — |
| gpt5.2_long | Use of roles | — | — | — | — | — | — |
| gpt5.2_long | Generating insights | — | — | — | — | — | — |
| gpt5.2_long | Providing examples | — | — | — | — | — | — |
| gpt5.2_short | Agent-Based Model | <0.01 | 0.15 | 0.07 | 0.09 | <0.01 | 0.55 |
| gpt5.2_short | Summarization algorithm | 0.18 | <0.01 | <0.01 | <0.01 | 0.94 | <0.01 |
| gpt5.2_short | Simulation evidence | — | — | — | — | — | — |
| gpt5.2_short | LLM | — | — | — | — | — | — |
| gpt5.2_short | Use of roles | — | — | — | — | — | — |
| gpt5.2_short | Generating insights | — | — | — | — | — | — |
| gpt5.2_short | Providing examples | — | — | — | — | — | — |
| modeler | Agent-Based Model | <0.01 | 0.10 | <0.01 | <0.01 | <0.01 | 0.55 |
| modeler | Summarization algorithm | <0.01 | <0.01 | <0.01 | 0.21 | 0.30 | <0.01 |
| modeler | Simulation evidence | — | — | — | — | — | — |
| modeler | LLM | — | — | — | — | — | — |
| modeler | Use of roles | — | — | — | — | — | — |
| modeler | Generating insights | — | — | — | — | — | — |
| modeler | Providing examples | — | — | — | — | — | — |

## screening/eval_qwen_kimi

| Reference family | Variable / metric | BLEU | METEOR | R-1 | R-2 | R-L | Reading ease |
| --- | --- | --- | --- | --- | --- | --- | --- |
| author | Agent-Based Model | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 |
| author | Summarization algorithm | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 |
| author | Simulation evidence | 0.05 | 0.71 | <0.01 | <0.01 | <0.01 | <0.01 |
| author | LLM | <0.01 | <0.01 | 0.02 | <0.01 | <0.01 | <0.01 |
| author | Use of roles | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | 0.74 |
| author | Generating insights | 0.68 | 0.48 | 0.02 | <0.01 | <0.01 | <0.01 |
| author | Providing examples | 0.43 | 0.05 | 0.11 | 0.18 | 0.65 | <0.01 |
| gpt5.2_long | Agent-Based Model | <0.01 | 0.74 | <0.01 | <0.01 | <0.01 | <0.01 |
| gpt5.2_long | Summarization algorithm | — | — | — | — | — | — |
| gpt5.2_long | Simulation evidence | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 |
| gpt5.2_long | LLM | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 |
| gpt5.2_long | Use of roles | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | 0.94 |
| gpt5.2_long | Generating insights | 0.01 | 0.07 | <0.01 | <0.01 | <0.01 | <0.01 |
| gpt5.2_long | Providing examples | 0.84 | <0.01 | 0.27 | 0.56 | 0.80 | <0.01 |
| gpt5.2_short | Agent-Based Model | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 |
| gpt5.2_short | Summarization algorithm | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 |
| gpt5.2_short | Simulation evidence | 0.84 | 0.64 | <0.01 | <0.01 | <0.01 | <0.01 |
| gpt5.2_short | LLM | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 |
| gpt5.2_short | Use of roles | <0.01 | 0.22 | <0.01 | <0.01 | <0.01 | 0.74 |
| gpt5.2_short | Generating insights | 0.58 | 0.05 | 0.80 | <0.01 | <0.01 | <0.01 |
| gpt5.2_short | Providing examples | 0.72 | 0.06 | 0.02 | 0.51 | 0.15 | <0.01 |
| modeler | Agent-Based Model | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 |
| modeler | Summarization algorithm | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 |
| modeler | Simulation evidence | 0.08 | 0.30 | 0.16 | <0.01 | <0.01 | <0.01 |
| modeler | LLM | <0.01 | <0.01 | <0.01 | 0.52 | 0.08 | <0.01 |
| modeler | Use of roles | <0.01 | 0.77 | 0.02 | <0.01 | <0.01 | 0.74 |
| modeler | Generating insights | 0.91 | 0.43 | 0.13 | <0.01 | <0.01 | <0.01 |
| modeler | Providing examples | 0.90 | 0.20 | 0.28 | 0.55 | 0.90 | <0.01 |

## screening/kimi-k2.5_all_abms_chain/06_quantitative_smoke_latest

| Reference family | Variable / metric | BLEU | METEOR | R-1 | R-2 | R-L | Reading ease |
| --- | --- | --- | --- | --- | --- | --- | --- |
| author | Agent-Based Model | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | 0.02 |
| author | Summarization algorithm | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 |
| author | Simulation evidence | 0.24 | 0.35 | 0.06 | <0.01 | 0.09 | 0.07 |
| author | LLM | — | — | — | — | — | — |
| author | Use of roles | <0.01 | 0.59 | 0.02 | <0.01 | <0.01 | 0.50 |
| author | Generating insights | 0.67 | 0.35 | <0.01 | <0.01 | <0.01 | 0.03 |
| author | Providing examples | 0.59 | 0.75 | 0.72 | 0.35 | 0.37 | 0.06 |
| gpt5.2_long | Agent-Based Model | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 |
| gpt5.2_long | Summarization algorithm | — | — | — | — | — | — |
| gpt5.2_long | Simulation evidence | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 |
| gpt5.2_long | LLM | — | — | — | — | — | — |
| gpt5.2_long | Use of roles | <0.01 | 0.13 | <0.01 | <0.01 | <0.01 | <0.01 |
| gpt5.2_long | Generating insights | <0.01 | 0.58 | <0.01 | <0.01 | <0.01 | <0.01 |
| gpt5.2_long | Providing examples | 0.27 | 0.16 | 0.36 | 0.55 | 0.48 | <0.01 |
| gpt5.2_short | Agent-Based Model | <0.01 | <0.01 | 0.13 | <0.01 | <0.01 | 0.02 |
| gpt5.2_short | Summarization algorithm | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 |
| gpt5.2_short | Simulation evidence | 0.58 | 0.38 | <0.01 | <0.01 | <0.01 | 0.07 |
| gpt5.2_short | LLM | — | — | — | — | — | — |
| gpt5.2_short | Use of roles | 0.64 | 0.48 | <0.01 | <0.01 | <0.01 | 0.50 |
| gpt5.2_short | Generating insights | 0.54 | 0.02 | 0.65 | <0.01 | <0.01 | 0.03 |
| gpt5.2_short | Providing examples | 0.86 | 1.00 | 0.94 | 0.74 | 0.54 | 0.06 |
| modeler | Agent-Based Model | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | 0.02 |
| modeler | Summarization algorithm | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 |
| modeler | Simulation evidence | 0.03 | 0.12 | 0.01 | <0.01 | <0.01 | 0.07 |
| modeler | LLM | — | — | — | — | — | — |
| modeler | Use of roles | <0.01 | 0.34 | 0.36 | <0.01 | <0.01 | 0.50 |
| modeler | Generating insights | 0.52 | 0.16 | 0.30 | <0.01 | <0.01 | 0.03 |
| modeler | Providing examples | 0.66 | 0.94 | 0.71 | 0.58 | 0.81 | 0.06 |

## screening/qwen3.5-27b_openrouter_all_abms_chain/06_quantitative_smoke_latest

| Reference family | Variable / metric | BLEU | METEOR | R-1 | R-2 | R-L | Reading ease |
| --- | --- | --- | --- | --- | --- | --- | --- |
| author | Agent-Based Model | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 |
| author | Summarization algorithm | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 |
| author | Simulation evidence | 0.19 | 0.77 | 0.01 | <0.01 | <0.01 | <0.01 |
| author | LLM | — | — | — | — | — | — |
| author | Use of roles | 0.24 | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 |
| author | Generating insights | 0.86 | 0.88 | 0.58 | 0.10 | <0.01 | <0.01 |
| author | Providing examples | 0.57 | 0.02 | 0.06 | 0.30 | 0.70 | <0.01 |
| gpt5.2_long | Agent-Based Model | <0.01 | <0.01 | 0.03 | <0.01 | <0.01 | <0.01 |
| gpt5.2_long | Summarization algorithm | — | — | — | — | — | — |
| gpt5.2_long | Simulation evidence | <0.01 | 0.16 | <0.01 | <0.01 | <0.01 | <0.01 |
| gpt5.2_long | LLM | — | — | — | — | — | — |
| gpt5.2_long | Use of roles | <0.01 | <0.01 | 0.13 | <0.01 | 0.45 | <0.01 |
| gpt5.2_long | Generating insights | 0.78 | <0.01 | 0.03 | 0.09 | 0.01 | <0.01 |
| gpt5.2_long | Providing examples | 0.54 | <0.01 | 0.32 | 0.79 | 0.15 | <0.01 |
| gpt5.2_short | Agent-Based Model | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 |
| gpt5.2_short | Summarization algorithm | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 |
| gpt5.2_short | Simulation evidence | 0.60 | 0.27 | 0.34 | <0.01 | <0.01 | <0.01 |
| gpt5.2_short | LLM | — | — | — | — | — | — |
| gpt5.2_short | Use of roles | <0.01 | 0.02 | <0.01 | <0.01 | <0.01 | <0.01 |
| gpt5.2_short | Generating insights | 0.86 | 0.56 | 0.98 | 0.47 | 0.05 | <0.01 |
| gpt5.2_short | Providing examples | 0.73 | <0.01 | <0.01 | 0.20 | 0.15 | <0.01 |
| modeler | Agent-Based Model | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 |
| modeler | Summarization algorithm | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 |
| modeler | Simulation evidence | 0.12 | 0.34 | 0.80 | <0.01 | 0.15 | <0.01 |
| modeler | LLM | — | — | — | — | — | — |
| modeler | Use of roles | <0.01 | 0.16 | 0.02 | <0.01 | <0.01 | <0.01 |
| modeler | Generating insights | 0.43 | 0.78 | 0.27 | 0.01 | <0.01 | <0.01 |
| modeler | Providing examples | 0.55 | 0.05 | 0.09 | 0.77 | 0.96 | <0.01 |
