MohammedSabry commited on
Commit
895fec6
·
verified ·
1 Parent(s): a730401

Upload folder using huggingface_hub

Browse files
README.md ADDED
@@ -0,0 +1,160 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ library_name: transformers
3
+ pipeline_tag: text-generation
4
+ language:
5
+ - en
6
+ tags:
7
+ - causal-lm
8
+ - biinduct
9
+ - pretraining
10
+ - matched-compute
11
+ - the-pile
12
+ - 125m
13
+ - balanced
14
+ ---
15
+
16
+ # Bi-Induct 125M Balanced
17
+
18
+ This repository contains the **Bi-Induct 125M Balanced** checkpoint from *Induction Signatures Are Not Enough: A Matched-Compute Study of Load-Bearing Structure in In-Context Learning*.
19
+
20
+ This release corresponds to the **0.13B** setting in the paper and is a **research checkpoint** intended for studying matched-compute pretraining, induction-style curricula, and in-context learning behavior. It is **not** instruction-tuned, alignment-tuned, or safety-tuned.
21
+
22
+ ## Variant
23
+
24
+ Bi-Induct balanced curriculum. Each synthetic injection chooses forward-copy or backward-copy with equal probability.
25
+
26
+ ## Model overview
27
+
28
+ - Architecture: decoder-only Transformer
29
+ - Positional encoding: RoPE (`theta=10000`)
30
+ - Normalization: pre-norm residual blocks
31
+ - MLP: SwiGLU
32
+ - Attention: grouped-query / grouped key-value attention
33
+ - Precision: bfloat16 training
34
+ - Context length: 1024
35
+ - Embeddings: untied input/output embeddings
36
+
37
+ ## Model specification
38
+
39
+ | Field | Value |
40
+ |---|---:|
41
+ | Parameters (paper label) | 0.13B |
42
+ | Layers | 12 |
43
+ | Hidden size | 768 |
44
+ | Intermediate / MLP size | 3,072 |
45
+ | Head dimension | 64 |
46
+ | Attention heads | 12 |
47
+ | KV heads | 3 |
48
+
49
+ ## Training data
50
+
51
+ All checkpoints in this family were pretrained on the **deduplicated THE PILE** in streaming / shuffled mode. A stable MD5-based hash was used to create a fixed held-out evaluation slice, with **0.2% of the corpus** reserved for evaluation (roughly **0.4B tokens**). Tokenization was truncated to **1024 tokens per sequence**.
52
+
53
+ For the Bi-Induct variants, synthetic snippets were interleaved on top of the natural stream:
54
+
55
+ - **Induction**: `[S || SEP || S]`
56
+ - **Anti-Induction**: `[S || SEP || reverse(S)]`
57
+ - **Balanced**: each injection randomly chooses induction or anti-induction
58
+
59
+ The main cross-scale experiments used **span length L = 20** and **initial mix ratio m0 = 50%**, linearly annealed to zero over the full training budget.
60
+
61
+ ## Training recipe
62
+
63
+ - Optimizer: AdamW (`beta1=0.9`, `beta2=0.999`, weight decay `0.1`)
64
+ - Learning rate: peak `1e-3`
65
+ - Schedule: `3%` linear warmup, then cosine decay
66
+ - Update size: `2^16` tokens per update
67
+ - Token budget: approximately `20N` tokens following the Chinchilla-style rule of thumb
68
+ - Comparison protocol: iso-FLOPs across curricula at each scale
69
+
70
+ ## Evaluation summary for the 125M family
71
+
72
+ The table below summarizes the main results at this scale. Standard LM benchmarks are evaluated **3-shot** and Todd et al. function-style probes are evaluated **10-shot** with **HITS@1**.
73
+
74
+ | Variant | Standard LM ICL composite ↑ | Todd-style ICL composite ↑ | Held-out PPL ↓ |
75
+ |---|---:|---:|---:|
76
+ | Baseline | 22.7 ± 0.5 | 5.3 ± 0.9 | 21.8 |
77
+ | Induction | 21.9 ± 0.5 | 4.1 ± 0.7 | 25.8 |
78
+ | Anti-Induction | 22.5 ± 0.4 | 3.8 ± 0.7 | 26.2 |
79
+ | Balanced | 22.4 ± 0.6 | 5.2 ± 0.8 | 26.2 |
80
+
81
+ **This checkpoint:** **Balanced**.
82
+
83
+ ## Benchmarks included
84
+
85
+ ### Standard LM benchmarks
86
+ - MMLU
87
+ - Winogrande
88
+ - CommonSenseQA
89
+ - PIQA
90
+ - HellaSwag
91
+ - TriviaQA-Wiki
92
+ - BBH (CoT)
93
+ - OpenBookQA
94
+ - ARC-Challenge
95
+ - GPQA
96
+ - GSM-8K
97
+ - MathQA
98
+ - BoolQ
99
+ - LAMBADA
100
+
101
+ ### Todd et al. function-style probes
102
+ - alphabetically first 3
103
+ - alphabetically first 5
104
+ - alphabetically last 3
105
+ - alphabetically last 5
106
+ - capitalize
107
+ - capitalize first letter
108
+ - capitalize last letter
109
+ - choose first of 3
110
+ - choose first of 5
111
+ - choose last of 3
112
+ - choose last of 5
113
+ - choose middle of 3
114
+ - choose middle of 5
115
+ - lowercase first letter
116
+ - lowercase last letter
117
+ - next capital letter
118
+ - next item
119
+ - prev item
120
+ - word length
121
+
122
+ ## Example usage
123
+
124
+ ```python
125
+ from transformers import AutoTokenizer, AutoModelForCausalLM
126
+
127
+ repo_id = "MohammedSabry/biinduct-125m-balanced"
128
+
129
+ tokenizer = AutoTokenizer.from_pretrained(repo_id)
130
+ model = AutoModelForCausalLM.from_pretrained(repo_id)
131
+
132
+ prompt = "The capital of France is"
133
+ inputs = tokenizer(prompt, return_tensors="pt")
134
+ outputs = model.generate(**inputs, max_new_tokens=20)
135
+ print(tokenizer.decode(outputs[0], skip_special_tokens=True))
136
+ ```
137
+
138
+ ## Limitations
139
+
140
+ - These are research checkpoints, not production chat models.
141
+ - They were designed to study the relationship between induction-style telemetry and load-bearing ICL behavior under matched compute.
142
+ - The synthetic interventions are intentionally lightweight and token-level; results should not be interpreted as ruling out richer data-rewrite strategies.
143
+ - Because Bi-Induct replaces a fraction of natural data under iso-FLOPs, some trade-offs may reflect natural-text displacement in addition to mechanistic redundancy.
144
+
145
+ ## Citation
146
+
147
+ If you use this model, please cite:
148
+
149
+ ```bibtex
150
+ @misc{sabry2026inductionsignaturesenoughmatchedcompute,
151
+ title={Induction Signatures Are Not Enough: A Matched-Compute Study of Load-Bearing Structure in In-Context Learning},
152
+ author={Mohammed Sabry and Anya Belz},
153
+ year={2026},
154
+ eprint={2509.22947},
155
+ archivePrefix={arXiv},
156
+ primaryClass={cs.CL},
157
+ url={https://arxiv.org/abs/2509.22947},
158
+ }
159
+ ```
160
+
config.json ADDED
@@ -0,0 +1,26 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "MistralForCausalLM"
4
+ ],
5
+ "attention_dropout": 0.0,
6
+ "bos_token_id": 1,
7
+ "eos_token_id": 2,
8
+ "head_dim": 64,
9
+ "hidden_act": "silu",
10
+ "hidden_size": 768,
11
+ "initializer_range": 0.02,
12
+ "intermediate_size": 3072,
13
+ "max_position_embeddings": 32768,
14
+ "model_type": "mistral",
15
+ "num_attention_heads": 12,
16
+ "num_hidden_layers": 12,
17
+ "num_key_value_heads": 3,
18
+ "rms_norm_eps": 1e-05,
19
+ "rope_theta": 10000.0,
20
+ "sliding_window": 4096,
21
+ "tie_word_embeddings": false,
22
+ "torch_dtype": "bfloat16",
23
+ "transformers_version": "4.52.4",
24
+ "use_cache": true,
25
+ "vocab_size": 32000
26
+ }
generation_config.json ADDED
@@ -0,0 +1,6 @@
 
 
 
 
 
 
 
1
+ {
2
+ "_from_model_config": true,
3
+ "bos_token_id": 1,
4
+ "eos_token_id": 2,
5
+ "transformers_version": "4.52.4"
6
+ }
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:5dc6ffafdb9b3c5f73ac13a6e731fefe60e304a5569eaa4204f7a8e3401d00a7
3
+ size 303613576
special_tokens_map.json ADDED
@@ -0,0 +1,24 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token": {
3
+ "content": "<s>",
4
+ "lstrip": false,
5
+ "normalized": false,
6
+ "rstrip": false,
7
+ "single_word": false
8
+ },
9
+ "eos_token": {
10
+ "content": "</s>",
11
+ "lstrip": false,
12
+ "normalized": false,
13
+ "rstrip": false,
14
+ "single_word": false
15
+ },
16
+ "pad_token": "</s>",
17
+ "unk_token": {
18
+ "content": "<unk>",
19
+ "lstrip": false,
20
+ "normalized": false,
21
+ "rstrip": false,
22
+ "single_word": false
23
+ }
24
+ }
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer.model ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:dadfd56d766715c61d2ef780a525ab43b8e6da4de6865bda3d95fdef5e134055
3
+ size 493443
tokenizer_config.json ADDED
@@ -0,0 +1,44 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_bos_token": true,
3
+ "add_eos_token": false,
4
+ "add_prefix_space": null,
5
+ "added_tokens_decoder": {
6
+ "0": {
7
+ "content": "<unk>",
8
+ "lstrip": false,
9
+ "normalized": false,
10
+ "rstrip": false,
11
+ "single_word": false,
12
+ "special": true
13
+ },
14
+ "1": {
15
+ "content": "<s>",
16
+ "lstrip": false,
17
+ "normalized": false,
18
+ "rstrip": false,
19
+ "single_word": false,
20
+ "special": true
21
+ },
22
+ "2": {
23
+ "content": "</s>",
24
+ "lstrip": false,
25
+ "normalized": false,
26
+ "rstrip": false,
27
+ "single_word": false,
28
+ "special": true
29
+ }
30
+ },
31
+ "additional_special_tokens": [],
32
+ "bos_token": "<s>",
33
+ "clean_up_tokenization_spaces": false,
34
+ "eos_token": "</s>",
35
+ "extra_special_tokens": {},
36
+ "legacy": false,
37
+ "model_max_length": 1000000000000000019884624838656,
38
+ "pad_token": "</s>",
39
+ "sp_model_kwargs": {},
40
+ "spaces_between_special_tokens": false,
41
+ "tokenizer_class": "LlamaTokenizer",
42
+ "unk_token": "<unk>",
43
+ "use_default_system_prompt": false
44
+ }