Benchmark results
Sphragis
107 models
Select a row to highlight its line in the plot; rows with diagnostics also compare accuracy by author
- Representation:
- GreBERTa
- KainoBERTa
- KainoBERT
- Char n-gram
- Lemma TF-IDF
- Syntax
- Burrows
- ALM
- Metrical line
| Sentence tasks | |||||||
|---|---|---|---|---|---|---|---|
| Siamese GreBERTa+syntax rates+lemma TF-IDF | Validation-selected per task: geometric mean ×3, validation-weighted vote ×3, soft vote, hard vote | 77.87 | 96.73 | 98.62 | 98.68 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| Char n-gram+Siamese GreBERTa+syntax rates | Validation-selected per task: geometric mean ×3, validation-weighted vote ×3, hard vote ×2 | 76.93 | 95.89 | 98.62 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| Char n-gram+GreBERTa (end-to-end)+lemma TF-IDF | Validation-selected per task: geometric mean ×3, hard vote ×2, mean rank, validation-weighted vote, soft vote | 76.80 | 96.05 | 98.88 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| Char n-gram+Siamese GreBERTa+lemma TF-IDF | Validation-selected per task: validation-weighted vote ×3, geometric mean ×2, hard vote ×2, mean rank | 76.78 | 96.32 | 98.10 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| GreBERTa (end-to-end)+syntax rates+lemma TF-IDF | Validation-selected per task: validation-weighted vote ×3, geometric mean ×2, hard vote ×2, soft vote | 76.36 | 95.83 | 98.45 | 100.00 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| Burrows z-scores+Siamese GreBERTa+syntax rates | Validation-selected per task: validation-weighted vote ×4, geometric mean ×3, hard vote | 76.28 | 95.64 | 98.15 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| Char n-gram+GreBERTa (end-to-end)+syntax rates | Validation-selected per task: validation-weighted vote ×4, geometric mean ×2, hard vote, soft vote | 75.98 | 96.33 | 98.88 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| Burrows z-scores+Siamese GreBERTa+lemma TF-IDF | Validation-selected per task: hard vote ×3, validation-weighted vote ×3, geometric mean ×2 | 75.87 | 95.82 | 97.97 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| ALM+Siamese GreBERTa+syntax rates | Validation-selected per task: validation-weighted vote ×5, geometric mean ×2, hard vote | 75.56 | 93.07 | 96.89 | 96.29 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method; the authorial language models are being remade and this row will be refreshed |
| Burrows z-scores+Char n-gram+Siamese GreBERTa | Validation-selected per task: geometric mean ×3, validation-weighted vote ×3, hard vote ×2 | 75.55 | 96.33 | 98.62 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| ALM+GreBERTa (end-to-end)+lemma TF-IDF | Validation-selected per task: geometric mean ×3, hard vote ×2, soft vote ×2, validation-weighted vote | 75.52 | 97.00 | 98.88 | 96.29 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method; the authorial language models are being remade and this row will be refreshed |
| GreBERTa+Siamese GreBERTa+syntax rates | Validation-selected per task: validation-weighted vote ×4, geometric mean ×3, hard vote | 75.48 | 95.57 | 97.92 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| ALM+Siamese GreBERTa+lemma TF-IDF | Validation-selected per task: geometric mean ×2, soft vote ×2, hard vote ×2, validation-weighted vote ×2 | 75.43 | 96.64 | 98.25 | 96.29 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method; the authorial language models are being remade and this row will be refreshed |
| Burrows z-scores+GreBERTa (end-to-end)+lemma TF-IDF | Validation-selected per task: geometric mean ×2, hard vote ×2, validation-weighted vote ×2, mean rank, soft vote | 75.37 | 94.88 | 98.88 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| Siamese GreBERTa+lemma TF-IDF | Validation-selected per task: hard vote ×4, validation-weighted vote ×2, geometric mean, mean rank | 75.36 | 95.00 | 98.25 | 100.00 | 2 arms | 2 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| Siamese GreBERTa+syntax rates | Validation-selected per task: validation-weighted vote ×6, geometric mean ×2 | 75.33 | 92.96 | 96.83 | 98.15 | 2 arms | 2 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| ALM+GreBERTa+lemma TF-IDF | Validation-selected per task: concatenation ×3, hard vote ×2, validation-weighted vote ×2, soft vote | 75.29 | 95.88 | 97.92 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method; the authorial language models are being remade and this row will be refreshed |
| ALM+Char n-gram+Siamese GreBERTa | Validation-selected per task: geometric mean ×3, hard vote ×3, validation-weighted vote, soft vote | 75.11 | 96.19 | 97.93 | 96.29 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method; the authorial language models are being remade and this row will be refreshed |
| GreBERTa+Siamese GreBERTa+lemma TF-IDF | Validation-selected per task: hard vote ×3, geometric mean ×2, validation-weighted vote ×2, mean rank | 75.08 | 95.42 | 97.81 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| GreBERTa (end-to-end)+lemma TF-IDF | Validation-selected per task: hard vote ×4, mean rank ×2, geometric mean, validation-weighted vote | 75.07 | 95.78 | 98.10 | 100.00 | 2 arms | 2 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| Char n-gram+GreBERTa+lemma TF-IDF | Validation-selected per task: concatenation ×4, hard vote ×3, validation-weighted vote | 75.04 | 96.55 | 97.94 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| GreBERTa (end-to-end)+Siamese GreBERTa+syntax rates | Validation-selected per task: validation-weighted vote ×4, geometric mean, mean rank, soft vote, hard vote | 74.95 | 94.31 | 97.26 | 96.29 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| GreBERTa (end-to-end)+Siamese GreBERTa+lemma TF-IDF | Validation-selected per task: hard vote ×5, mean rank ×2, soft vote | 74.92 | 95.24 | 97.82 | 96.29 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| Char n-gram+Siamese GreBERTa | Validation-selected per task: mean rank ×3, geometric mean ×2, validation-weighted vote, soft vote, hard vote | 74.84 | 95.72 | 97.24 | 100.00 | 2 arms | 2 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| Burrows z-scores+GreBERTa (end-to-end)+syntax rates | Validation-selected per task: validation-weighted vote ×5, geometric mean ×2, hard vote | 74.84 | 96.31 | 98.35 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| Char n-gram+GreBERTa (end-to-end)+Siamese GreBERTa | Validation-selected per task: hard vote ×4, mean rank ×3, soft vote | 74.83 | 95.53 | 97.00 | 96.29 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| Burrows z-scores+Char n-gram+GreBERTa (end-to-end) | Validation-selected per task: validation-weighted vote ×3, geometric mean ×2, hard vote ×2, soft vote | 74.71 | 97.09 | 98.88 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| Char n-gram+GreBERTa+Siamese GreBERTa | Validation-selected per task: geometric mean ×3, hard vote ×2, soft vote ×2, mean rank | 74.59 | 95.64 | 98.18 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| ALM+Burrows z-scores+Siamese GreBERTa | Validation-selected per task: validation-weighted vote ×3, geometric mean ×2, soft vote ×2, hard vote | 74.57 | 94.73 | 98.72 | 96.29 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method; the authorial language models are being remade and this row will be refreshed |
| ALM+GreBERTa (end-to-end)+Siamese GreBERTa | Validation-selected per task: hard vote ×6, soft vote ×2 | 74.49 | 92.96 | 95.51 | 96.29 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method; the authorial language models are being remade and this row will be refreshed |
| ALM+GreBERTa (end-to-end)+syntax rates | Validation-selected per task: validation-weighted vote ×4, geometric mean ×2, soft vote, hard vote | 74.40 | 94.41 | 97.16 | 96.29 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method; the authorial language models are being remade and this row will be refreshed |
| Burrows z-scores+GreBERTa (end-to-end)+Siamese GreBERTa | Validation-selected per task: hard vote ×5, geometric mean, mean rank, soft vote | 74.34 | 94.10 | 98.65 | 96.29 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| ALM+Char n-gram+GreBERTa (end-to-end) | Validation-selected per task: geometric mean ×3, soft vote ×3, hard vote ×2 | 74.25 | 95.78 | 98.35 | 96.29 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method; the authorial language models are being remade and this row will be refreshed |
| KainoBERT+Char n-gram+lemma TF-IDF | PyTorch multinomial logistic regression | 74.24 | 96.24 | 96.72 | 97.94 | KainoBERT, Char n-gram and lemma TF-IDF columns stacked: 169,212 / 169,212 / 169,212 / 31,951 / 136,327 / 24,163 / 46,175 / 24,163 | Two prepared single-channel feature blocks concatenated, each arm at the variant its own standalone model selected on validation; block scaling (balanced/balanced/native/balanced/native/native/native/balanced) selected per task on validation macro-F1, where balanced divides each block by its mean training row norm and native keeps the standalone scales; Five-point learning-rate search on validation macro-F1; validation-loss early stopping; train+validation refit; test evaluated once |
| Burrows z-scores+GreBERTa+Siamese GreBERTa | Validation-selected per task: geometric mean ×2, soft vote ×2, hard vote ×2, validation-weighted vote ×2 | 74.17 | 95.14 | 98.24 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| Char n-gram+GreBERTa+GreBERTa (end-to-end) | Validation-selected per task: soft vote ×3, geometric mean ×2, hard vote ×2, validation-weighted vote | 74.08 | 96.63 | 98.04 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| GreBERTa (end-to-end)+syntax rates | Validation-selected per task: validation-weighted vote ×6, geometric mean ×2 | 74.02 | 94.20 | 97.18 | 100.00 | 2 arms | 2 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| GreBERTa+GreBERTa (end-to-end)+syntax rates | Validation-selected per task: hard vote ×3, validation-weighted vote ×3, geometric mean, soft vote | 74.01 | 94.93 | 96.87 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| Char n-gram+GreBERTa (end-to-end) | Validation-selected per task: hard vote ×4, mean rank ×3, geometric mean | 73.99 | 95.83 | 97.63 | 100.00 | 2 arms | 2 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| Burrows z-scores+Siamese GreBERTa | Validation-selected per task: validation-weighted vote ×3, geometric mean ×2, hard vote ×2, mean rank | 73.98 | 94.63 | 98.72 | 100.00 | 2 arms | 2 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| Burrows z-scores+Char n-gram+GreBERTa | Validation-selected per task: concatenation ×4, geometric mean ×2, hard vote ×2 | 73.52 | 96.37 | 97.81 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| GreBERTa (end-to-end)+Siamese GreBERTa | Validation-selected per task: hard vote ×4, validation-weighted vote ×2, mean rank ×2 | 73.44 | 93.67 | 95.33 | 96.29 | 2 arms | 2 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| GreBERTa+GreBERTa (end-to-end)+Siamese GreBERTa | Validation-selected per task: hard vote ×4, soft vote ×3, mean rank | 73.29 | 94.66 | 97.03 | 96.29 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| KainoBERT+Burrows z-scores+Char n-gram | PyTorch multinomial logistic regression | 73.14 | 96.05 | 97.94 | 97.94 | KainoBERT, Burrows z-scores and Char n-gram columns stacked: 141,850 / 141,850 / 141,850 / 4,589 / 116,578 / 4,414 / 26,426 / 3,914 | Two prepared single-channel feature blocks concatenated, each arm at the variant its own standalone model selected on validation; block scaling (balanced/balanced/balanced/balanced/balanced/balanced/balanced/balanced) selected per task on validation macro-F1, where balanced divides each block by its mean training row norm and native keeps the standalone scales; Five-point learning-rate search on validation macro-F1; validation-loss early stopping; train+validation refit; test evaluated once |
| GreBERTa+GreBERTa (end-to-end)+lemma TF-IDF | Validation-selected per task: hard vote ×4, soft vote ×2, mean rank, geometric mean | 73.12 | 96.84 | 98.35 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| ALM+Burrows z-scores+GreBERTa (end-to-end) | Validation-selected per task: hard vote ×4, validation-weighted vote ×2, geometric mean, soft vote | 73.10 | 94.34 | 96.02 | 96.29 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method; the authorial language models are being remade and this row will be refreshed |
| GreBERTa+syntax rates+lemma TF-IDF | Validation-selected per task: validation-weighted vote ×4, geometric mean ×2, hard vote ×2 | 73.08 | 96.15 | 97.41 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| ALM+GreBERTa (end-to-end) | Validation-selected per task: geometric mean ×3, hard vote ×3, mean rank, soft vote | 73.04 | 93.46 | 95.20 | 96.29 | 2 arms | 2 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method; the authorial language models are being remade and this row will be refreshed |
| Burrows z-scores+GreBERTa+GreBERTa (end-to-end) | Validation-selected per task: geometric mean ×3, validation-weighted vote ×3, hard vote, soft vote | 72.88 | 95.07 | 98.29 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| KainoBERT+lemma TF-IDF+ALM | PyTorch multinomial logistic regression | 72.85 | 95.61 | 97.51 | 97.94 | KainoBERT, lemma TF-IDF and ALM columns stacked: 29,158 / 29,158 / 29,158 / 29,158 / 21,529 / 21,529 / 21,529 / 21,529 | Two prepared single-channel feature blocks concatenated, each arm at the variant its own standalone model selected on validation; block scaling (balanced/balanced/balanced/balanced/native/native/native/balanced) selected per task on validation macro-F1, where balanced divides each block by its mean training row norm and native keeps the standalone scales; Five-point learning-rate search on validation macro-F1; validation-loss early stopping; train+validation refit; test evaluated once |
| ALM+Siamese GreBERTa | Validation-selected per task: hard vote ×4, geometric mean ×3, soft vote | 72.76 | 92.77 | 95.48 | 96.29 | 2 arms | 2 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method; the authorial language models are being remade and this row will be refreshed |
| GreBERTa+lemma TF-IDF | Validation-selected per task: concatenation ×3, hard vote ×2, validation-weighted vote ×2, mean rank | 72.67 | 96.00 | 98.16 | 97.67 | 2 arms | 2 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| Burrows z-scores+GreBERTa (end-to-end) | Validation-selected per task: validation-weighted vote ×4, mean rank ×3, geometric mean | 72.67 | 91.65 | 97.69 | 100.00 | 2 arms | 2 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| ALM+Char n-gram+lemma TF-IDF | Validation-selected per task: concatenation ×6, hard vote ×2 | 72.64 | 96.70 | 98.80 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method; the authorial language models are being remade and this row will be refreshed |
| ALM+Char n-gram+GreBERTa | Validation-selected per task: concatenation ×3, soft vote ×2, hard vote ×2, validation-weighted vote | 72.62 | 96.36 | 97.83 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method; the authorial language models are being remade and this row will be refreshed |
| ALM+GreBERTa+Siamese GreBERTa | Validation-selected per task: hard vote ×4, geometric mean ×2, validation-weighted vote, soft vote | 72.61 | 95.11 | 97.23 | 96.29 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method; the authorial language models are being remade and this row will be refreshed |
| GreBERTa+Siamese GreBERTa | Validation-selected per task: geometric mean ×3, mean rank ×2, hard vote ×2, validation-weighted vote | 72.55 | 93.97 | 95.92 | 98.68 | 2 arms | 2 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| Siamese GreBERTa (SupCon + CE) | Nearest author centroid and linear softmax head | 72.49 | 92.83 | 95.36 | 96.29 | ~127M trainable parameters; 256-d L2-normalized embedding | Shared GreBERTa encoder trained with supervised contrastive plus cross-entropy on author-balanced batches; epoch and decision rule (head/head at epoch 51/28) both selected on the full validation split; larger-task scores are means over exact constituents; test evaluated once |
| Char n-gram+GreBERTa+syntax rates | Validation-selected per task: validation-weighted vote ×5, hard vote ×2, soft vote | 72.40 | 95.88 | 97.57 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| Char n-gram+syntax rates+lemma TF-IDF | Validation-selected per task: validation-weighted vote ×4, hard vote ×2, geometric mean, soft vote | 72.35 | 96.40 | 98.06 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| KainoBERT+Char n-gram+ALM | PyTorch multinomial logistic regression | 72.29 | 95.94 | 96.09 | 97.94 | KainoBERT, Char n-gram and ALM columns stacked: 140,878 / 140,878 / 140,878 / 3,617 / 115,590 / 3,426 / 25,438 / 3,426 | Two prepared single-channel feature blocks concatenated, each arm at the variant its own standalone model selected on validation; block scaling (balanced/balanced/native/native/native/native/native/native) selected per task on validation macro-F1, where balanced divides each block by its mean training row norm and native keeps the standalone scales; Five-point learning-rate search on validation macro-F1; validation-loss early stopping; train+validation refit; test evaluated once |
| Char n-gram+GreBERTa | Validation-selected per task: hard vote ×4, concatenation ×3, validation-weighted vote | 72.12 | 96.36 | 97.83 | 97.94 | 2 arms | 2 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| KainoBERT+Char n-gram | PyTorch multinomial logistic regression | 72.11 | 95.94 | 95.57 | 97.94 | KainoBERT and Char n-gram columns stacked: 140,850 / 140,850 / 140,850 / 3,589 / 115,578 / 3,414 / 25,426 / 3,414 | Two prepared single-channel feature blocks concatenated, each arm at the variant its own standalone model selected on validation; block scaling (balanced/balanced/native/native/balanced/native/native/native) selected per task on validation macro-F1, where balanced divides each block by its mean training row norm and native keeps the standalone scales; Five-point learning-rate search on validation macro-F1; validation-loss early stopping; train+validation refit; test evaluated once |
| ALM+GreBERTa+GreBERTa (end-to-end) | Validation-selected per task: hard vote ×5, geometric mean ×2, mean rank | 72.08 | 94.35 | 96.08 | 96.29 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method; the authorial language models are being remade and this row will be refreshed |
| ALM+lemma TF-IDF | Validation-selected per task: concatenation ×7, hard vote | 71.96 | 96.70 | 98.80 | 97.94 | 2 arms | 2 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method; the authorial language models are being remade and this row will be refreshed |
| KainoBERT+lemma TF-IDF | PyTorch multinomial logistic regression | 71.84 | 95.55 | 97.51 | 97.94 | KainoBERT and lemma TF-IDF columns stacked: 29,130 / 29,130 / 29,130 / 29,130 / 21,517 / 21,517 / 21,517 / 21,517 | Two prepared single-channel feature blocks concatenated, each arm at the variant its own standalone model selected on validation; block scaling (balanced/balanced/balanced/balanced/native/native/native/balanced) selected per task on validation macro-F1, where balanced divides each block by its mean training row norm and native keeps the standalone scales; Five-point learning-rate search on validation macro-F1; validation-loss early stopping; train+validation refit; test evaluated once |
| Char n-gram+lemma TF-IDF | Validation-selected per task: concatenation ×5, hard vote ×3 | 71.66 | 96.01 | 98.19 | 97.94 | 2 arms | 2 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| GreBERTa+GreBERTa (end-to-end) | Validation-selected per task: mean rank ×3, validation-weighted vote ×2, hard vote ×2, geometric mean | 71.65 | 93.52 | 95.62 | 100.00 | 2 arms | 2 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| ALM+GreBERTa | Validation-selected per task: concatenation ×5, mean rank, geometric mean, hard vote | 71.59 | 92.85 | 95.07 | 96.02 | 2 arms | 2 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method; the authorial language models are being remade and this row will be refreshed |
| GreBERTa (end-to-end) | Linear softmax head | 71.51 | 93.06 | 95.20 | 96.29 | ~126M trainable parameters per benchmark track | Fine-tuned on the atomic task with validation-loss early stopping; larger-task logits are means over exact constituents; test evaluated once |
| Char n-gram+syntax rates | Validation-selected per task: validation-weighted vote ×5, concatenation ×2, hard vote | 71.10 | 96.49 | 97.22 | 97.94 | 2 arms | 2 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| ALM+Char n-gram | Validation-selected per task: concatenation ×6, hard vote ×2 | 70.63 | 96.63 | 99.12 | 97.94 | 2 arms | 2 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method; the authorial language models are being remade and this row will be refreshed |
| Burrows z-scores+GreBERTa+lemma TF-IDF | Validation-selected per task: geometric mean ×3, hard vote ×3, validation-weighted vote ×2 | 70.26 | 94.86 | 97.51 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| ALM+Burrows z-scores+Char n-gram | Validation-selected per task: concatenation ×4, hard vote ×2, soft vote ×2 | 70.02 | 94.69 | 97.99 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method; the authorial language models are being remade and this row will be refreshed |
| KainoBERT (end-to-end) | Linear softmax head | 69.88 | 95.66 | 98.29 | 100.00 | ~136M trainable parameters per benchmark track | Fine-tuned on the atomic task with validation-loss early stopping; larger-task logits are means over exact constituents; test evaluated once |
| syntax rates+lemma TF-IDF | Validation-selected per task: concatenation ×3, validation-weighted vote ×3, hard vote ×2 | 69.80 | 96.54 | 98.48 | 97.94 | 2 arms | 2 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| Burrows z-scores+GreBERTa+syntax rates | Validation-selected per task: validation-weighted vote ×4, geometric mean ×3, hard vote | 69.51 | 94.78 | 96.82 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| ALM+Burrows z-scores+GreBERTa | Validation-selected per task: concatenation ×3, hard vote ×2, soft vote ×2, validation-weighted vote | 69.45 | 93.32 | 96.55 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method; the authorial language models are being remade and this row will be refreshed |
| ALM+Char n-gram+syntax rates | Validation-selected per task: validation-weighted vote ×6, geometric mean, hard vote | 69.26 | 95.27 | 97.22 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method; the authorial language models are being remade and this row will be refreshed |
| Burrows z-scores+Char n-gram+lemma TF-IDF | Validation-selected per task: validation-weighted vote ×3, hard vote ×3, geometric mean, soft vote | 69.09 | 95.61 | 97.74 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| Burrows z-scores+Char n-gram | Validation-selected per task: concatenation ×5, hard vote ×2, validation-weighted vote | 69.04 | 94.34 | 97.99 | 97.94 | 2 arms | 2 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| ALM+syntax rates+lemma TF-IDF | Validation-selected per task: validation-weighted vote ×5, geometric mean, hard vote, soft vote | 69.02 | 95.81 | 97.60 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method; the authorial language models are being remade and this row will be refreshed |
| Burrows z-scores+Char n-gram+syntax rates | Validation-selected per task: validation-weighted vote ×5, hard vote ×2, soft vote | 68.82 | 95.69 | 97.41 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| ALM+GreBERTa+syntax rates | Validation-selected per task: validation-weighted vote ×4, soft vote ×2, geometric mean, hard vote | 68.53 | 92.37 | 95.58 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method; the authorial language models are being remade and this row will be refreshed |
| KainoBERT+Burrows z-scores+ALM | PyTorch multinomial logistic regression | 68.47 | 94.83 | 97.16 | 97.94 | KainoBERT, Burrows z-scores and ALM columns stacked: 1,796 / 1,796 / 1,796 / 1,796 / 1,780 / 1,780 / 1,780 / 1,280 | Two prepared single-channel feature blocks concatenated, each arm at the variant its own standalone model selected on validation; block scaling (balanced/balanced/balanced/native/balanced/balanced/balanced/balanced) selected per task on validation macro-F1, where balanced divides each block by its mean training row norm and native keeps the standalone scales; Five-point learning-rate search on validation macro-F1; validation-loss early stopping; train+validation refit; test evaluated once |
| KainoBERT+ALM | PyTorch multinomial logistic regression | 68.42 | 93.34 | 95.68 | 97.94 | KainoBERT and ALM columns stacked: 796 / 796 / 796 / 796 / 780 / 780 / 780 / 780 | Two prepared single-channel feature blocks concatenated, each arm at the variant its own standalone model selected on validation; block scaling (native/native/native/native/native/balanced/native/native) selected per task on validation macro-F1, where balanced divides each block by its mean training row norm and native keeps the standalone scales; Five-point learning-rate search on validation macro-F1; validation-loss early stopping; train+validation refit; test evaluated once |
| Burrows z-scores+GreBERTa | Validation-selected per task: concatenation ×3, hard vote ×2, validation-weighted vote ×2, geometric mean | 68.23 | 92.95 | 96.74 | 97.94 | 2 arms | 2 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| KainoBERT+syntax rates | PyTorch multinomial logistic regression | 68.07 | 93.26 | 96.80 | 96.88 | KainoBERT and syntax rates columns stacked: 259,251 / 12,943 / 12,943 / 259,251 / 2,591,501 / 2,591,501 / 2,591,501 / 228,513 | Two prepared single-channel feature blocks concatenated, each arm at the variant its own standalone model selected on validation; block scaling (native/balanced/native/native/native/native/native/native) selected per task on validation macro-F1, where balanced divides each block by its mean training row norm and native keeps the standalone scales; Five-point learning-rate search on validation macro-F1; validation-loss early stopping; train+validation refit; test evaluated once |
| Char n-gram | PyTorch multinomial logistic regression | 67.67 | 94.60 | 97.22 | 97.94 | 140,082 / 140,082 / 140,082 / 2,821 / 114,810 / 2,646 / 24,658 / 2,646 | Validation-selected n=4/4/4/2 and 4/2/3/2; train+validation refit; test evaluated once |
| Char n-gram Dirichlet-multinomial | Dirichlet compound multinomial posterior predictive | 67.38 | 94.84 | 96.83 | 96.62 | Validation-selected n=4/4/3/3/4/4/4/4 | Character n-gram counts with a uniform author prior fixed a priori; validation-selected n and Dirichlet smoothing alpha=0.01/0.1/0.1/0.01/0.1/0.1/0.1/0.1; vocabulary from the atomic training split only; train+validation refit; test evaluated once; validation-selected Dirichlet concentration=raw posterior/100.0/100.0/1000.0/100.0/raw posterior/raw posterior/raw posterior, where the raw posterior is the non-bursty limit that reproduces naive Bayes |
| Char n-gram naive Bayes | Multinomial naive Bayes | 67.21 | 94.48 | 97.11 | 97.94 | Validation-selected n=4/4/4/2/4/4/4/4 | Character n-gram counts with a uniform author prior fixed a priori; validation-selected n and Dirichlet smoothing alpha=0.01/0.1/0.1/1.0/0.1/0.1/0.1/0.1; vocabulary from the atomic training split only; train+validation refit; test evaluated once |
| GreBERTa+syntax rates | Validation-selected per task: validation-weighted vote ×4, geometric mean, concatenation, hard vote, mean rank | 67.05 | 92.93 | 95.47 | 97.67 | 2 arms | 2 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| KainoBERT+Burrows z-scores | PyTorch multinomial logistic regression | 66.67 | 94.17 | 96.90 | 97.94 | KainoBERT and Burrows z-scores columns stacked: 1,768 / 1,768 / 1,768 / 1,768 / 1,768 / 1,768 / 1,768 / 1,268 | Two prepared single-channel feature blocks concatenated, each arm at the variant its own standalone model selected on validation; block scaling (balanced/balanced/balanced/native/balanced/balanced/balanced/balanced) selected per task on validation macro-F1, where balanced divides each block by its mean training row norm and native keeps the standalone scales; Five-point learning-rate search on validation macro-F1; validation-loss early stopping; train+validation refit; test evaluated once |
| Burrows z-scores+syntax rates+lemma TF-IDF | Validation-selected per task: validation-weighted vote ×4, hard vote ×3, geometric mean | 65.78 | 94.43 | 97.04 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| KainoBERTa (end-to-end) | Linear softmax head | 65.45 | 93.71 | 96.00 | 100.00 | ~112M trainable parameters per benchmark track | Fine-tuned on the atomic task with validation-loss early stopping; larger-task logits are means over exact constituents; test evaluated once. Architecture control for KainoBERT: same pretraining corpus, blocks, tokenizer, masking and schedule, stopped by the same validation plateau rule |
| ALM+Burrows z-scores+lemma TF-IDF | Validation-selected per task: hard vote ×4, mean rank ×2, validation-weighted vote, soft vote | 64.66 | 94.34 | 97.38 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method; the authorial language models are being remade and this row will be refreshed |
| OLMo-1B authorial language models | Lowest mean per-token surprisal | 64.50 | 87.12 | 88.00 | 94.96 | 1.18B parameters per author; 28 and 12 models | One OLMo-1B model further pretrained per author on the current release with validation-loss early stopping; attribution is the lowest mean per-token surprisal; no base-model or epoch search; test evaluated once |
| GreBERTa | PyTorch multinomial logistic regression | 64.38 | 91.14 | 93.92 | 97.67 | 768 frozen features | Five-point learning-rate search on validation macro-F1; validation-loss early stopping; train+validation refit; test evaluated once |
| Lemma TF-IDF | PyTorch multinomial logistic regression | 63.43 | 94.64 | 97.64 | 97.94 | 28,362 / 28,362 / 28,362 / 28,362 / 20,749 / 20,749 / 20,749 / 20,749 | Five-point learning-rate search on validation macro-F1; validation-loss early stopping; train+validation refit; test evaluated once |
| ALM+Burrows z-scores+syntax rates | Validation-selected per task: validation-weighted vote ×4, mean rank, geometric mean, soft vote, hard vote | 63.25 | 92.36 | 96.26 | 97.94 | 3 arms | 3 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method; the authorial language models are being remade and this row will be refreshed |
| KainoBERT | PyTorch multinomial logistic regression | 63.09 | 93.02 | 96.00 | 97.94 | 768 frozen features | Five-point learning-rate search on validation macro-F1; validation-loss early stopping; train+validation refit; test evaluated once |
| Burrows z-scores+lemma TF-IDF | Validation-selected per task: concatenation ×5, hard vote ×2, geometric mean | 62.88 | 93.97 | 97.12 | 97.94 | 2 arms | 2 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| ALM+Burrows z-scores | Validation-selected per task: concatenation ×3, mean rank ×2, hard vote ×2, geometric mean | 62.88 | 88.77 | 94.15 | 97.94 | 2 arms | 2 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method; the authorial language models are being remade and this row will be refreshed |
| ALM+syntax rates | Validation-selected per task: mean rank ×5, concatenation ×2, hard vote | 60.62 | 94.63 | 95.08 | 96.13 | 2 arms | 2 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method; the authorial language models are being remade and this row will be refreshed |
| Burrows z-scores+syntax rates | Validation-selected per task: validation-weighted vote ×4, concatenation ×3, hard vote | 59.13 | 91.98 | 95.26 | 97.94 | 2 arms | 2 atomic models combined; the combination method is a hyperparameter chosen per task on validation macro-F1 among concatenation of prepared feature blocks and five training-free posterior combiners (hard vote, soft vote, geometric mean, mean rank, validation-weighted vote); test evaluated once for the chosen method |
| Burrows lemma z-scores | PyTorch multinomial logistic regression | 43.68 | 85.39 | 93.53 | 100.00 | Validation-selected MFW n=1000/1000/1000/1000/1000/1000/1000/500 | Five-point learning-rate search on validation macro-F1; validation-loss early stopping; train+validation refit; test evaluated once |
| Syntax feature rates | PyTorch multinomial logistic regression | 42.59 | 77.80 | 87.87 | 95.86 | Validation-selected d=3/2/2/3/4/4/4/3; features=258,483 / 12,175 / 12,175 / 258,483 / 2,590,733 / 2,590,733 / 2,590,733 / 227,745 | Combination dimension 1-4 selected independently per task on validation macro-F1; Five-point learning-rate search on validation macro-F1; validation-loss early stopping; train+validation refit; test evaluated once |
Scores are held-out test macro-F1 (%).
Plotted results
Score by task size
Held-out test macro-F1 (%); points are linearly connected. Select a line, a table row, or a legend entry to highlight one model.