RoeiG commited on
Commit
93c90c6
·
verified ·
1 Parent(s): 82aa0f0

Card: drop an unverified claim

Browse files
Files changed (1) hide show
  1. README.md +3 -3
README.md CHANGED
@@ -157,7 +157,7 @@ Differences smaller than these are noise. This checkpoint is seed 1, which was f
157
  - Strong keywords in the text can outweigh the option descriptions.
158
  - **Claim-form relevance is slightly below the previous checkpoint:** he_bench 0.78 against 0.80, and BEIR-he 0.78 against 0.82.
159
  - **Scales:** sentiment and tone are about 0.45 accuracy, and the tone probabilities are overconfident (valence Brier
160
- 0.53). The lowest level of a scale is under-used.
161
  - **Calculated fields steer less than in the previous checkpoint.** On 8 test emails, a "direct manager" sender field
162
  raised importance by 0.11 of a level, against 0.22 before, and a "mailing list" field lowered it in only 2 of 8.
163
  - **Calibration does not catch everything.** The failures above are often high-confidence, so a confidence threshold
@@ -165,8 +165,8 @@ Differences smaller than these are noise. This checkpoint is seed 1, which was f
165
 
166
  ## Training data
167
 
168
- The training start was an earlier checkpoint of this project, trained on the same public data and on the same
169
- encoder. The run was one epoch over 396,508 items (6,196 updates, 1.4 A100 hours). The learning rates were 5e-6 for the
170
  encoder and 1e-4 for the head. Every case was converted to Laya's format, with a random subset of options, paraphrased
171
  instructions and varied field names.
172
 
 
157
  - Strong keywords in the text can outweigh the option descriptions.
158
  - **Claim-form relevance is slightly below the previous checkpoint:** he_bench 0.78 against 0.80, and BEIR-he 0.78 against 0.82.
159
  - **Scales:** sentiment and tone are about 0.45 accuracy, and the tone probabilities are overconfident (valence Brier
160
+ 0.53).
161
  - **Calculated fields steer less than in the previous checkpoint.** On 8 test emails, a "direct manager" sender field
162
  raised importance by 0.11 of a level, against 0.22 before, and a "mailing list" field lowered it in only 2 of 8.
163
  - **Calibration does not catch everything.** The failures above are often high-confidence, so a confidence threshold
 
165
 
166
  ## Training data
167
 
168
+ The training start was an earlier checkpoint of this project, with the same encoder, trained on part of the public
169
+ data below. The run was one epoch over 396,508 items (6,196 updates, 1.4 A100 hours). The learning rates were 5e-6 for the
170
  encoder and 1e-4 for the head. Every case was converted to Laya's format, with a random subset of options, paraphrased
171
  instructions and varied field names.
172