ssands1979 commited on
Commit
6f24274
·
verified ·
1 Parent(s): e51bc93

Clarify merge pipeline: custom scripted merge (mergekit-inspired), document instruct-source starting point

Browse files
Files changed (1) hide show
  1. README.md +39 -41
README.md CHANGED
@@ -12,7 +12,6 @@ tags:
12
  - cpt
13
  - repair
14
  - depth-upscaling
15
- - mergekit
16
  - base-model
17
  pipeline_tag: image-text-to-text
18
  base_model:
@@ -25,45 +24,43 @@ This repository contains the selected **repaired base checkpoint** for our 80B-c
25
 
26
  ## Summary
27
 
28
- This release is a **repair / recovery continued-pretraining checkpoint**, not a final instruct model and not a posttrained agent model.
29
 
30
  The published checkpoint is:
31
 
32
  - `v7-20260415-234214/checkpoint-100`
33
 
34
- It was selected from several checkpoint-100 candidates after a short multimodal repair-CPT run intended to stabilize a depth-upscaled 80B-class merged model and recover coherent base-model behavior.
35
 
36
- ## Important note on the source model
37
 
38
- At construction time, the upstream **raw 27B base checkpoint was not available** in the environment we were working from. The available upstream starting point was the **Qwen 3.5 27B instruct checkpoint**.
39
 
40
- That matters for interpretation:
41
 
42
- - this 80B checkpoint should still be treated as a **repaired base candidate**, not as a clean continuation of a pristine raw-base lineage
43
- - some degradation or drift in instruction-format adherence, style, and exact response behavior is expected during repair CPT
44
- - additional **posttraining repair and multimodal posttraining** are expected and planned after this stage
45
 
46
- In other words: this release is the best current repaired base candidate from the available source path, not the end of the pipeline.
47
 
48
  ## Construction pipeline
49
 
50
  The model was built through a staged pipeline:
51
 
52
- 1. **Abliteration / de-alignment pass** on the available Qwen 3.5 27B instruct source checkpoint
53
- 2. **End-to-end depth upscaling / merge** using `mergekit` passthrough stacking to produce an ~80B-class dense multimodal model
54
- 3. **Short repair CPT** to stabilize the merged model, improve text-side recovery, and preserve functional multimodal behavior
55
 
56
- Conceptually, the goal was:
57
 
58
  - preserve the first-class VLM nature of the Qwen 3.5 family
59
- - produce a larger dense multimodal base candidate quickly
60
- - use a short repair pass to recover coherence before later posttraining stages
61
 
62
  ## Merge / architecture rationale
63
 
64
- The 80B model was created by **depth upscaling** rather than training an 80B model from scratch. The underlying idea is that repeating the pretrained stack end-to-end can create a deeper network that still inherits useful circuits from the source model, after which a repair CPT pass can smooth the new transitions.
65
-
66
- This is not intended as a claim that the merged checkpoint is fully recovered immediately after merging. The repair pass is a required part of the process.
67
 
68
  ## Repair CPT setup
69
 
@@ -73,37 +70,38 @@ The selected checkpoint came from a short multimodal repair-CPT pilot using:
73
  - nodes: **4** (32 B200 GPUs total)
74
  - dtype: **bf16**
75
  - mode: **full-parameter training**
76
- - max length: **8192**
77
  - save cadence: **every 100 steps**
 
 
 
78
 
79
  ### Repair data mix
80
 
81
- The pilot used a conservative mixture intended to recover general text behavior while keeping multimodal/document competence alive:
82
 
83
- - **75% text backbone**: cached **C4/en**
84
- - **25% multimodal/doc mix**, normalized into a training-friendly format from FineVision-derived sources:
85
  - `olmOCR-mix-0225-documents`
86
  - `olmOCR-mix-0225-books`
87
  - `docvqa`
88
  - `arxivqa`
89
  - `ocrvqa`
90
 
91
- The multimodal data had to be normalized before use because the raw parquet metadata path was not directly robust for the repair run stack.
92
-
93
  ## Why checkpoint-100 was selected
94
 
95
- This was treated as **circuit repair CPT**, not as a long-horizon capability training run where the lowest scalar loss automatically wins.
96
 
97
- The most important selection criteria were:
98
 
99
- 1. recover coherent text generation
100
- 2. avoid obvious merge-path regressions
101
- 3. retain clearly functional document/image understanding
102
- 4. keep the checkpoint early and conservative enough to reduce risk of overfitting to the repair mix
103
 
104
  ## Tiny checkpoint-selection validation
105
 
106
- Only **small validation subsets** were used at this stage. These were used only to identify the best repair checkpoint and to confirm that the model was recoverable. They should **not** be interpreted as final benchmark claims.
107
 
108
  ### Plain Transformers smoke
109
 
@@ -112,7 +110,7 @@ Both the raw merged 80B checkpoint and checkpoint-100 were verified to:
112
  - load in plain `transformers`
113
  - generate tokens successfully
114
 
115
- In a trivial prompt smoke test, the selected checkpoint generated a clean answer, while the raw merged checkpoint remained visibly degraded.
116
 
117
  ### Tiny GSM8K subset (n=5)
118
 
@@ -126,20 +124,20 @@ The published checkpoint was chosen from the strongest candidates in this tiny s
126
 
127
  ### Tiny multimodal qualitative check
128
 
129
- A small document/OCR-style qualitative check showed that the selected checkpoint is clearly able to:
130
 
131
  - read document pages
132
  - extract structure and content
133
  - produce semantically faithful long-form document responses
134
 
135
- A simple exact-match metric used in that quick check understated the practical multimodal capability because the model frequently returned structured, explanatory, or partially reformatted transcriptions instead of exact verbatim text. Those small multimodal checks were therefore useful for **sanity checking**, but not for rigorous ranking.
136
 
137
  ## What this release is for
138
 
139
  Use this model as a **base candidate** for the next stages of work:
140
 
141
- - stronger multimodal repair CPT if needed
142
- - posttraining repair of instruction behavior/style
143
  - supervised posttraining
144
  - preference / constitutional alignment
145
  - downstream agent-style posttraining
@@ -148,19 +146,20 @@ Use this model as a **base candidate** for the next stages of work:
148
 
149
  This repository should **not** be treated as:
150
 
151
- - a final instruct model
152
  - a finished OCR-specialized model
153
  - a final benchmarked public release with comprehensive evaluation
154
 
155
- It is the selected **post-CPT repaired base checkpoint** from the current pipeline.
156
 
157
  ## Expected next phase
158
 
159
  The next stage is expected to focus on **posttraining**, with additional attention to:
160
 
161
- - cleaning up instruction behavior inherited from the source-path constraints
162
  - improving multimodal answer style and exactness where needed
163
  - running broader, more task-appropriate evaluations beyond the tiny checkpoint-selection subsets used here
 
164
 
165
  ## Provenance
166
 
@@ -168,4 +167,3 @@ Published from:
168
 
169
  - source checkpoint: `v7-20260415-234214/checkpoint-100`
170
  - repository: `artivus-ai/qwen-3.5-80b-post-cpt-base`
171
-
 
12
  - cpt
13
  - repair
14
  - depth-upscaling
 
15
  - base-model
16
  pipeline_tag: image-text-to-text
17
  base_model:
 
24
 
25
  ## Summary
26
 
27
+ This release is a **repair / recovery continued-pretraining (CPT) checkpoint**, not a final instruct model and not a posttrained agent model.
28
 
29
  The published checkpoint is:
30
 
31
  - `v7-20260415-234214/checkpoint-100`
32
 
33
+ selected from several checkpoint-100 candidates after a short multimodal repair-CPT run intended to stabilize a custom depth-upscaled 80B-class merged model and recover coherent base-model behavior.
34
 
35
+ ## Important note on the source checkpoint
36
 
37
+ At construction time, the upstream **raw 27B base checkpoint was not available** to us in the environment we were working from. The practical upstream starting point was the **Qwen 3.5 27B instruct checkpoint**.
38
 
39
+ This matters for interpretation:
40
 
41
+ - this 80B artifact should be treated as a **repaired base candidate**, not as a clean continuation of a pristine raw-base lineage
42
+ - some degradation or drift in instruction-format adherence, response style, and formatting behavior is expected during repair CPT
43
+ - additional **posttraining repair and posttraining proper** are expected and planned as the next stage
44
 
45
+ So this release is intentionally framed as the best current repaired base candidate from the available source path, not the end of the pipeline.
46
 
47
  ## Construction pipeline
48
 
49
  The model was built through a staged pipeline:
50
 
51
+ 1. **Abliteration / de-alignment pass** on the available Qwen 3.5 27B instruct source checkpoint.
52
+ 2. **Custom scripted end-to-end depth upscaling / merge** producing an ~80B-class dense multimodal model. The conceptual template is mergekit-style passthrough depth stacking, but the actual published artifact was produced through our **own merge / graft pipeline** written in-house after mergekit-based tooling paths did not work cleanly for this specific VLM setup.
53
+ 3. **Short repair CPT** to stabilize the merged model, recover coherent text-side behavior, and preserve functional multimodal capability.
54
 
55
+ Conceptually, the goals were:
56
 
57
  - preserve the first-class VLM nature of the Qwen 3.5 family
58
+ - produce a larger dense multimodal base candidate using depth upscaling instead of training a new 80B-class model from scratch
59
+ - then use a short repair CPT pass to recover coherence before moving into posttraining
60
 
61
  ## Merge / architecture rationale
62
 
63
+ Rather than training an 80B model from scratch, we chose **depth upscaling** from a smaller pretrained source. The rationale is that repeating the pretrained stack end-to-end produces a deeper network that still inherits useful pretrained structure, and a short repair CPT pass can then smooth the new transitions. The merged artifact is not expected to be strong immediately out of merge; the repair CPT is an intentional part of the process, not an optional cleanup step.
 
 
64
 
65
  ## Repair CPT setup
66
 
 
70
  - nodes: **4** (32 B200 GPUs total)
71
  - dtype: **bf16**
72
  - mode: **full-parameter training**
73
+ - max sequence length: **8192**
74
  - save cadence: **every 100 steps**
75
+ - gradient checkpointing: **enabled**
76
+
77
+ The first multimodal pilot had to normalize the training data format before use, because the raw FineVision-derived parquet metadata did not cleanly flow through the trainer's dataset path.
78
 
79
  ### Repair data mix
80
 
81
+ The pilot used a conservative mixture intended to recover general text-side behavior while keeping multimodal / document / OCR competence alive:
82
 
83
+ - **75% text backbone**: cached **C4 / en**
84
+ - **25% multimodal / doc mix**, normalized into a training-friendly format from FineVision-derived sources:
85
  - `olmOCR-mix-0225-documents`
86
  - `olmOCR-mix-0225-books`
87
  - `docvqa`
88
  - `arxivqa`
89
  - `ocrvqa`
90
 
 
 
91
  ## Why checkpoint-100 was selected
92
 
93
+ This was explicitly treated as **circuit repair CPT**, not as a long-horizon capability training run, so the lowest absolute loss is not the primary criterion.
94
 
95
+ The checkpoint was chosen based on:
96
 
97
+ 1. coherent text generation
98
+ 2. absence of obvious merge-path regressions
99
+ 3. clearly functional document / image understanding in qualitative checks
100
+ 4. preference for an early, conservative checkpoint to reduce risk of overfitting to the repair data mix
101
 
102
  ## Tiny checkpoint-selection validation
103
 
104
+ Only **small validation subsets** were used at this stage. These were run purely to identify the best repair checkpoint and to confirm that the model was recoverable. They should **not** be interpreted as final benchmark claims.
105
 
106
  ### Plain Transformers smoke
107
 
 
110
  - load in plain `transformers`
111
  - generate tokens successfully
112
 
113
+ In a trivial prompt smoke test, the selected checkpoint produced a clean coherent response, while the raw merged checkpoint remained visibly degraded.
114
 
115
  ### Tiny GSM8K subset (n=5)
116
 
 
124
 
125
  ### Tiny multimodal qualitative check
126
 
127
+ A small document / OCR-style qualitative check showed that the selected checkpoint is clearly able to:
128
 
129
  - read document pages
130
  - extract structure and content
131
  - produce semantically faithful long-form document responses
132
 
133
+ A simple exact-match metric used in that check understated the practical multimodal capability, because the model frequently returned structured, explanatory or partially reformatted transcriptions instead of exact verbatim text. So the small multimodal checks were useful as **sanity checks**, but not for rigorous ranking.
134
 
135
  ## What this release is for
136
 
137
  Use this model as a **base candidate** for the next stages of work:
138
 
139
+ - stronger multimodal repair CPT, if needed
140
+ - posttraining repair of instruction behavior / formatting / style
141
  - supervised posttraining
142
  - preference / constitutional alignment
143
  - downstream agent-style posttraining
 
146
 
147
  This repository should **not** be treated as:
148
 
149
+ - a final instruction-tuned model
150
  - a finished OCR-specialized model
151
  - a final benchmarked public release with comprehensive evaluation
152
 
153
+ It is the selected **post-CPT repaired base checkpoint** from the current pipeline stage.
154
 
155
  ## Expected next phase
156
 
157
  The next stage is expected to focus on **posttraining**, with additional attention to:
158
 
159
+ - cleaning up instruction behavior inherited from the instruct-source starting point
160
  - improving multimodal answer style and exactness where needed
161
  - running broader, more task-appropriate evaluations beyond the tiny checkpoint-selection subsets used here
162
+ - optional additional repair CPT with a stronger multimodal focus if evaluation indicates it is needed
163
 
164
  ## Provenance
165
 
 
167
 
168
  - source checkpoint: `v7-20260415-234214/checkpoint-100`
169
  - repository: `artivus-ai/qwen-3.5-80b-post-cpt-base`