File size: 4,141 Bytes
970f21f
3fc63c7
 
970f21f
3fc63c7
970f21f
 
3fc63c7
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
970f21f
 
 
3fc63c7
 
 
 
 
970f21f
3fc63c7
 
9db33b7
3fc63c7
9db33b7
3fc63c7
 
 
 
970f21f
 
3fc63c7
 
970f21f
3fc63c7
970f21f
3fc63c7
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
---
license: mit
language: en
library_name: transformers
tags: [nanochat, pretraining, climbmix, political-bias, exp-088]
---

# climbmix-d26-10tpp-noright — d26 base model, 10 tokens/parameter, ClimbMix with the "no right-wing content" filter

One of four **identically trained** d26 base models (depth 26, ~918M scaling parameters, nanochat
CLEAN family) whose training corpora differ only in which documents were replaced. This model's
corpus removes **a right-leaning entity (1st or 3rd person, weak or strong) present** and fills every removed document, in place, with a length-matched
apolitical document. Companion models: `Eugleo/climbmix-d26-10tpp-nopol`, `Eugleo/climbmix-d26-10tpp-noecon`, `Eugleo/climbmix-d26-10tpp-nocrime`.

## Training

| | |
|---|---|
| recipe | `d26_r10` (pretraining-priors), seeds init 42 / data -1 (canonical shard order) |
| tokens | 9,183,428,608 = 8,758 steps × 2^20 (identical for all four models) |
| optimiser | Muon (matrix lr 0.02), AdamW for embeddings/unembedding/scalars; weight decay 0.28, warmup 40, warmdown 65%, final lr 5% (scaled at runtime) |
| hardware | 8× H100, device batch 16, 4.32 h wall clock, ~55% MFU |
| arm tag | `d26-r10-a737eac7df78` (arm hash over config, code and corpus content; data_code 0bcff3836fb8) |
| corpus | `climbmix_4100_noright`, derived from `climbmix_4100` with the `replace_texts` transform (seed 0) |

## The corpus

ClimbMix (karpathy/climbmix-400b-shuffle, 4,101 shards). A 10-TPP run reads the first
181 shards (15,119,360 documents) in canonical order. Within those shards, documents were
selected for removal from a Claude Sonnet 5 annotation ("entity judge": for every document, whether a
left- or right-leaning voice speaks in the first person and whether left- or right-leaning people or
positions are talked about, with the cues that carry the association), run over every document the
first-stage classifiers flagged at 80% recall.

| | |
|---|---|
| documents replaced | 313,414 (2.07% of the read prefix) |
| characters removed / added | 1,433,919,176 / 1,439,647,750 |
| tokens removed / added (training tokenizer) | 309,130,508 / 310,089,476 (net +958,968, +0.010% of the budget) |
| tokens per 1,024-document block, parent → this corpus | 613,236 → 612,564 |
| read prefix, parent → this corpus | 181 → 181 shards |

Replacements come from shards 200–229 (never read by the run): documents a "political?" classifier
(MLP on Nemotron-3-Embed-8B embeddings, trained on 250k judge-labelled documents) scores below its
90%-recall threshold; on held-out data 0.49% of such documents carry a political entity, against 5.1%
in the corpus. Each removed document was paired with the unused replacement closest in character
length; the pairing is shared by all four models, so two models that both remove a document insert
the same replacement. Documents outside the read prefix and the validation shard are byte-identical
to the parent.

## Evaluation

| | |
|---|---|
| CORE (nanochat base_eval suite, step 8,758) | 0.254207 |
| validation bits per byte (last in-training eval) | 0.718694 |
| export verification | passed: max |logit diff| 0.0, bpb 0.7187 over 41,943,040 tokens, KV-cache consistent |

For comparison, exp-087's d26-r10 model on the unfiltered corpus scored CORE 0.272 and exp-085's
0.261; run-to-run spread of this recipe is a few points.

## Loading

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("Eugleo/climbmix-d26-10tpp-noright")
model = AutoModelForCausalLM.from_pretrained("Eugleo/climbmix-d26-10tpp-noright", trust_remote_code=True, torch_dtype="bfloat16")
```

Base model only (no instruction tuning). Weights are bf16 `model.safetensors`; the custom modeling
code (`modeling_nanochat_gpt.py`) is included. Optimizer state is not published.

## Provenance

Built in `pretraining-priors` experiment exp-088 (branch `exp088-pretrain-embedding`): edit table and
per-document token accounting at `ppriors_data/exp088_noright/` on the training volume; the full
description, the judge prompt and the interactive data explorer live with the experiment.