ssands1979 commited on
Commit
e51bc93
·
verified ·
1 Parent(s): 3cc578c

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +114 -47
README.md CHANGED
@@ -8,97 +8,164 @@ tags:
8
  - qwen3.5
9
  - multimodal
10
  - vlm
 
11
  - cpt
12
  - repair
 
 
13
  - base-model
14
  pipeline_tag: image-text-to-text
15
  base_model:
16
- - Qwen/Qwen3.5-27B
17
  ---
18
 
19
- # qwen-3.5-80b-post-cpt-base
20
 
21
- This repository contains the selected repaired base checkpoint for our 80B-class Qwen 3.5-derived VLM.
22
 
23
- ## What this is
24
 
25
- This model is **not** an instruct-tuned assistant and **not** a posttrained agent model.
26
- It is a **base-model recovery / repair CPT checkpoint** chosen after a short multimodal continued pretraining run intended to stabilize and repair a merged 80B base.
27
 
28
- The selected checkpoint is:
29
 
30
  - `v7-20260415-234214/checkpoint-100`
31
 
32
- chosen from several checkpoint-100 candidates because it showed the best tiny text-side recovery while preserving clearly functional document/image understanding.
33
 
34
- ## Training intent
35
 
36
- This run was treated as **circuit repair CPT**, not as a long horizon learning run.
37
- The goal was to recover a healthy base model after merge/assembly rather than continue broad capability training indefinitely.
38
 
39
- ## Architecture notes
40
 
41
- - Qwen 3.5-derived multimodal model
42
- - 80B-class repaired base
43
- - text backbone depth-expanded / merged from Qwen 3.5 lineage
44
- - vision tower retained as a first-class VLM component
45
 
46
- ## Data used in the repair run
47
 
48
- The repair pilot used a mixture of:
49
 
50
- - text-side cached C4 data
51
- - normalized multimodal document / OCR / VQA data derived from FineVision-style parquet sources
52
 
53
- The multimodal sources were normalized into a Swift/Transformers-friendly format before training.
 
 
54
 
55
- ## Checkpoint selection summary
56
 
57
- A tiny checkpoint sweep was run after training.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
58
 
59
  ### Plain Transformers smoke
60
- Both the raw merged 80B and checkpoint-100 were verified to:
 
61
 
62
  - load in plain `transformers`
63
- - generate tokens
64
 
65
- The selected checkpoint produced a clean response to a trivial prompt, while the raw merged base remained visibly degraded.
66
 
67
- ### Tiny text eval (GSM8K subset, n=5)
68
- On a tiny direct `transformers` GSM8K slice:
 
69
 
70
  - `v6-20260415-234215/checkpoint-100`: **1 / 5**
71
  - `v7-20260415-164214/checkpoint-100`: **4 / 5**
72
  - `v7-20260415-234214/checkpoint-100`: **4 / 5**
73
 
74
- We selected `v7-20260415-234214/checkpoint-100` as the published repaired base.
75
 
76
  ### Tiny multimodal qualitative check
77
- A small document/OCR-style qualitative check showed that the model is clearly able to:
 
78
 
79
  - read document pages
80
- - extract content
81
- - describe / summarize page structure
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
82
 
83
- The exact-match metric used there was too strict for long-form document outputs, so those scores understate the practical multimodal capability. The model often returned semantically faithful but more structured / narrated responses instead of exact verbatim transcription.
84
 
85
- ## Recommended usage
86
 
87
- Use this checkpoint as a **base model** for:
88
 
89
- - further CPT repair
90
- - stronger multimodal retention training
91
- - later supervised / posttraining work
92
- - agent / tool / assistant posttraining stages
93
 
94
- ## Caveats
95
 
96
- - This is a repaired **base** model, not an instruction model.
97
- - Multimodal exact OCR-style faithfulness likely still benefits from additional targeted CPT or posttraining.
98
- - The evaluation described above was intentionally small and fast, used only for checkpoint selection, not for final benchmark claims.
99
 
100
- ## Next planned work
 
101
 
102
- - stronger multimodal-focused recovery CPT if needed
103
- - publish later posttrained checkpoints separately
104
- - further VQA / doc QA / OCR evaluations with more task-appropriate metrics
 
8
  - qwen3.5
9
  - multimodal
10
  - vlm
11
+ - continued-pretraining
12
  - cpt
13
  - repair
14
+ - depth-upscaling
15
+ - mergekit
16
  - base-model
17
  pipeline_tag: image-text-to-text
18
  base_model:
19
+ - Qwen/Qwen3.5-27B-Instruct
20
  ---
21
 
22
+ # Qwen 3.5 80B Post-CPT Base
23
 
24
+ This repository contains the selected **repaired base checkpoint** for our 80B-class Qwen 3.5-derived vision-language model.
25
 
26
+ ## Summary
27
 
28
+ This release is a **repair / recovery continued-pretraining checkpoint**, not a final instruct model and not a posttrained agent model.
 
29
 
30
+ The published checkpoint is:
31
 
32
  - `v7-20260415-234214/checkpoint-100`
33
 
34
+ It was selected from several checkpoint-100 candidates after a short multimodal repair-CPT run intended to stabilize a depth-upscaled 80B-class merged model and recover coherent base-model behavior.
35
 
36
+ ## Important note on the source model
37
 
38
+ At construction time, the upstream **raw 27B base checkpoint was not available** in the environment we were working from. The available upstream starting point was the **Qwen 3.5 27B instruct checkpoint**.
 
39
 
40
+ That matters for interpretation:
41
 
42
+ - this 80B checkpoint should still be treated as a **repaired base candidate**, not as a clean continuation of a pristine raw-base lineage
43
+ - some degradation or drift in instruction-format adherence, style, and exact response behavior is expected during repair CPT
44
+ - additional **posttraining repair and multimodal posttraining** are expected and planned after this stage
 
45
 
46
+ In other words: this release is the best current repaired base candidate from the available source path, not the end of the pipeline.
47
 
48
+ ## Construction pipeline
49
 
50
+ The model was built through a staged pipeline:
 
51
 
52
+ 1. **Abliteration / de-alignment pass** on the available Qwen 3.5 27B instruct source checkpoint
53
+ 2. **End-to-end depth upscaling / merge** using `mergekit` passthrough stacking to produce an ~80B-class dense multimodal model
54
+ 3. **Short repair CPT** to stabilize the merged model, improve text-side recovery, and preserve functional multimodal behavior
55
 
56
+ Conceptually, the goal was:
57
 
58
+ - preserve the first-class VLM nature of the Qwen 3.5 family
59
+ - produce a larger dense multimodal base candidate quickly
60
+ - use a short repair pass to recover coherence before later posttraining stages
61
+
62
+ ## Merge / architecture rationale
63
+
64
+ The 80B model was created by **depth upscaling** rather than training an 80B model from scratch. The underlying idea is that repeating the pretrained stack end-to-end can create a deeper network that still inherits useful circuits from the source model, after which a repair CPT pass can smooth the new transitions.
65
+
66
+ This is not intended as a claim that the merged checkpoint is fully recovered immediately after merging. The repair pass is a required part of the process.
67
+
68
+ ## Repair CPT setup
69
+
70
+ The selected checkpoint came from a short multimodal repair-CPT pilot using:
71
+
72
+ - backend: **MS-Swift + DeepSpeed ZeRO-3**
73
+ - nodes: **4** (32 B200 GPUs total)
74
+ - dtype: **bf16**
75
+ - mode: **full-parameter training**
76
+ - max length: **8192**
77
+ - save cadence: **every 100 steps**
78
+
79
+ ### Repair data mix
80
+
81
+ The pilot used a conservative mixture intended to recover general text behavior while keeping multimodal/document competence alive:
82
+
83
+ - **75% text backbone**: cached **C4/en**
84
+ - **25% multimodal/doc mix**, normalized into a training-friendly format from FineVision-derived sources:
85
+ - `olmOCR-mix-0225-documents`
86
+ - `olmOCR-mix-0225-books`
87
+ - `docvqa`
88
+ - `arxivqa`
89
+ - `ocrvqa`
90
+
91
+ The multimodal data had to be normalized before use because the raw parquet metadata path was not directly robust for the repair run stack.
92
+
93
+ ## Why checkpoint-100 was selected
94
+
95
+ This was treated as **circuit repair CPT**, not as a long-horizon capability training run where the lowest scalar loss automatically wins.
96
+
97
+ The most important selection criteria were:
98
+
99
+ 1. recover coherent text generation
100
+ 2. avoid obvious merge-path regressions
101
+ 3. retain clearly functional document/image understanding
102
+ 4. keep the checkpoint early and conservative enough to reduce risk of overfitting to the repair mix
103
+
104
+ ## Tiny checkpoint-selection validation
105
+
106
+ Only **small validation subsets** were used at this stage. These were used only to identify the best repair checkpoint and to confirm that the model was recoverable. They should **not** be interpreted as final benchmark claims.
107
 
108
  ### Plain Transformers smoke
109
+
110
+ Both the raw merged 80B checkpoint and checkpoint-100 were verified to:
111
 
112
  - load in plain `transformers`
113
+ - generate tokens successfully
114
 
115
+ In a trivial prompt smoke test, the selected checkpoint generated a clean answer, while the raw merged checkpoint remained visibly degraded.
116
 
117
+ ### Tiny GSM8K subset (n=5)
118
+
119
+ A tiny direct `transformers` GSM8K slice was used for quick text-side checkpoint comparison:
120
 
121
  - `v6-20260415-234215/checkpoint-100`: **1 / 5**
122
  - `v7-20260415-164214/checkpoint-100`: **4 / 5**
123
  - `v7-20260415-234214/checkpoint-100`: **4 / 5**
124
 
125
+ The published checkpoint was chosen from the strongest candidates in this tiny sweep.
126
 
127
  ### Tiny multimodal qualitative check
128
+
129
+ A small document/OCR-style qualitative check showed that the selected checkpoint is clearly able to:
130
 
131
  - read document pages
132
+ - extract structure and content
133
+ - produce semantically faithful long-form document responses
134
+
135
+ A simple exact-match metric used in that quick check understated the practical multimodal capability because the model frequently returned structured, explanatory, or partially reformatted transcriptions instead of exact verbatim text. Those small multimodal checks were therefore useful for **sanity checking**, but not for rigorous ranking.
136
+
137
+ ## What this release is for
138
+
139
+ Use this model as a **base candidate** for the next stages of work:
140
+
141
+ - stronger multimodal repair CPT if needed
142
+ - posttraining repair of instruction behavior/style
143
+ - supervised posttraining
144
+ - preference / constitutional alignment
145
+ - downstream agent-style posttraining
146
+
147
+ ## What this release is not
148
+
149
+ This repository should **not** be treated as:
150
+
151
+ - a final instruct model
152
+ - a finished OCR-specialized model
153
+ - a final benchmarked public release with comprehensive evaluation
154
 
155
+ It is the selected **post-CPT repaired base checkpoint** from the current pipeline.
156
 
157
+ ## Expected next phase
158
 
159
+ The next stage is expected to focus on **posttraining**, with additional attention to:
160
 
161
+ - cleaning up instruction behavior inherited from the source-path constraints
162
+ - improving multimodal answer style and exactness where needed
163
+ - running broader, more task-appropriate evaluations beyond the tiny checkpoint-selection subsets used here
 
164
 
165
+ ## Provenance
166
 
167
+ Published from:
 
 
168
 
169
+ - source checkpoint: `v7-20260415-234214/checkpoint-100`
170
+ - repository: `artivus-ai/qwen-3.5-80b-post-cpt-base`
171