RoyalCities commited on
Commit
a3961b6
·
verified ·
1 Parent(s): 1ff6964

Create training_dataset_info.md

Browse files
Files changed (1) hide show
  1. training_dataset_info.md +310 -0
training_dataset_info.md ADDED
@@ -0,0 +1,310 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ <center>
2
+ <h1 style="font-size:34px;"><u>Foundation-1 Training & Dataset Notes</u></h1>
3
+ </center>
4
+
5
+ <center>
6
+ <i>High-level overview of the training setup, dataset composition, and design philosophy behind Foundation-1.</i>
7
+ </center>
8
+
9
+ <br>
10
+
11
+ This page provides a high-level overview of the **training setup**, **dataset composition**, and **design philosophy** behind Foundation-1.
12
+
13
+ It is intended as a companion to the main model page, which focuses on capabilities, prompting, and audio examples.
14
+
15
+ ---
16
+
17
+ <center>
18
+
19
+ ## Training Summary
20
+
21
+ </center>
22
+
23
+ Foundation-1 was trained as a structured **text-to-sample diffusion model** designed for **music production workflows**, rather than general-purpose music captioning.
24
+
25
+ ### Hardware
26
+
27
+ - **GPUs:** 2 × NVIDIA RTX A6000
28
+ - **System RAM:** 128 GB
29
+
30
+ ### Training Run
31
+
32
+ - **Training stopped at step:** **183,474**
33
+
34
+ ### Audio Configuration
35
+
36
+ - **Sample Rate:** **44,100 Hz**
37
+ - **Bit Depth:** **16-bit**
38
+ - **Channels:** **Stereo**
39
+ - **Sample Length:** **882,000 samples (~20 seconds)**
40
+
41
+ ---
42
+
43
+ <center>
44
+
45
+ ## Technical Training Notes
46
+
47
+ </center>
48
+
49
+ Foundation-1 was fine-tuned from **`stabilityai/stable-audio-open-1.0`** using a diffusion transformer architecture with structured prompt conditioning.
50
+
51
+ ### Model Architecture
52
+
53
+ - **Backbone:** Diffusion Transformer (DiT)
54
+ - **Transformer Depth:** 24 layers
55
+ - **Attention Heads:** 24
56
+ - **Embedding Dimension:** 1536
57
+ - **Conditioning Dimension:** 768
58
+ - **Text Encoder:** `t5-base`
59
+
60
+ ### Optimization
61
+
62
+ - **Optimizer:** AdamW
63
+ - **Learning Rate:** `5e-5`
64
+ - **Weight Decay:** `1e-3`
65
+ - **Scheduler:** InverseLR
66
+ - **EMA:** Enabled
67
+
68
+ ---
69
+
70
+ <center>
71
+
72
+ ## Dataset Overview
73
+
74
+ </center>
75
+
76
+ ### Dataset Totals
77
+
78
+ - **Total WAV files:** **3,784,862**
79
+ - **Total Dataset Size:** **7.100 TiB**
80
+
81
+ ---
82
+
83
+ <center>
84
+
85
+ ## Important Note on Dataset Scale
86
+
87
+ </center>
88
+
89
+ While the dataset appears quite large, its scale reflects the **post-augmentation training set**, not a flat count of completely unique and unrelated one-off melodic phrases.
90
+
91
+ All samples were **hand-labeled first**, then processed through a **controlled augmentation pipeline** designed to expand sonic and conditioning coverage.
92
+
93
+ As a result, a single melodic phrase may be represented across many different contexts, including variations in:
94
+
95
+ - timbre
96
+ - FX treatment
97
+ - key
98
+ - BPM
99
+ - loop structure
100
+ - tonal emphasis
101
+
102
+ This design was intentional. The goal was not simply to maximize the number of isolated phrases, but to teach the model how **musical ideas translate across different sonic identities and production scenarios**.
103
+
104
+ ---
105
+
106
+ <center>
107
+
108
+ ## Dataset and Training Philosophy
109
+
110
+ </center>
111
+
112
+ Foundation-1 was built around a **structured sample-generation philosophy**, rather than generic or genre-based audio captioning.
113
+
114
+ The dataset consists entirely of **hand-labeled audio**, organized around a layered prompt and conditioning system designed to reflect how producers actually think about sound.
115
+
116
+ At a high level, the training design emphasizes:
117
+
118
+ - structured musical loops
119
+ - instrument hierarchy
120
+ - explicit timbre representation
121
+ - dedicated FX descriptors
122
+ - notation-aware prompt terms
123
+ - key / tempo / bar-aware looping
124
+ - strong production relevance
125
+ - broad reuse for compositional workflows
126
+
127
+ This design is central to the model’s **musical coherence**, **prompt controllability**, and **production-facing behavior**.
128
+
129
+ ---
130
+
131
+ <center>
132
+
133
+ ## Why the Dataset Was Structured This Way
134
+
135
+ </center>
136
+
137
+ Most audio generation systems treat prompts as broad descriptive captions.
138
+
139
+ Foundation-1 instead was trained to understand sound as a **layered system with separable controls**.
140
+
141
+ That structure includes:
142
+
143
+ **Instrument Identity**
144
+ Broad family and sub-family control over what kind of sound is being generated.
145
+
146
+ **Timbre**
147
+ Direct conditioning over tonal character, texture, density, brightness, width, grit, warmth, and other sonic traits.
148
+
149
+ **FX Context**
150
+ Dedicated processing descriptors such as reverb, delay, distortion, phasing, and bitcrushing.
151
+
152
+ **Musical Structure**
153
+ Notation-aware terms that encourage coherent phrasing, melodic behavior, rhythmic structure, and harmonic motion.
154
+
155
+ **Timing and Tonality**
156
+ Explicit support for BPM, bar count, and key-aware sample generation.
157
+
158
+ This layered design is one of the main reasons Foundation-1 can produce outputs that feel both **musically structured** and **sonically steerable**.
159
+
160
+ ---
161
+
162
+ <center>
163
+
164
+ ## Coverage and Augmentation
165
+
166
+ </center>
167
+
168
+ The augmentation strategy was designed to improve **coverage**, **control**, and **generalization**.
169
+
170
+ Rather than treating each source phrase as a single static example, labeled material was expanded across multiple sonic and musical contexts.
171
+
172
+ This helps the model learn relationships between:
173
+
174
+ - melody and timbre
175
+ - phrase behavior and instrumentation
176
+ - sound design and FX treatment
177
+ - tonal setting and loop structure
178
+ - tempo and musical feel
179
+
180
+ In practice, this encourages the model to learn **reusable musical relationships**, rather than simply memorizing isolated recordings.
181
+
182
+ ---
183
+
184
+ <center>
185
+
186
+ ## Generalization and Overfitting
187
+
188
+ </center>
189
+
190
+ Because structured augmentation can increase dataset size rapidly, special care was taken to reduce melodic overfitting and encourage broader generalization.
191
+
192
+ The objective was to help the model learn:
193
+
194
+ - how similar phrase structures can exist across many timbral identities
195
+ - how sound design changes affect musical material
196
+ - how prompts can steer sonic outcomes without collapsing variety
197
+ - how the same conditioning vocabulary can remain useful across many production contexts
198
+
199
+ This is one reason Foundation-1 can produce **multiple distinct outputs from the same prompt** while still preserving the requested timbral and structural identity.
200
+
201
+ ---
202
+
203
+ <center>
204
+
205
+ ## What the Dataset Size Represents
206
+
207
+ </center>
208
+
209
+ The full size of the dataset should be understood as a measure of **coverage and structured variation**, not simply a count of unrelated melodies.
210
+
211
+ Foundation-1 was not designed around “more files for the sake of more files,” which is often seen with large scraped or loosely structured datasets.
212
+
213
+ Instead, the dataset was built from the ground up to teach the model how:
214
+
215
+ - phrases behave across different instruments
216
+ - timbral tags influence sonic identity
217
+ - FX descriptors shape the output
218
+ - timing and tonality influence loop behavior
219
+ - musical structure can remain coherent under many sonic conditions
220
+
221
+ In other words, the dataset size reflects the **breadth of the conditioning system** as much as it reflects the raw quantity of audio.
222
+
223
+ ---
224
+
225
+ <center>
226
+
227
+ ## Charts at a Glance
228
+
229
+ </center>
230
+
231
+ The following charts illustrate the distribution of key components of the Foundation-1 training dataset.
232
+
233
+ <center>
234
+
235
+ <table>
236
+ <tr>
237
+ <td align="center">
238
+ <img src="./Charts/families_pie.PNG" width="1100">
239
+ </td>
240
+
241
+ <td align="center">
242
+ <img src="./Charts/subfamilites_pie.PNG" width="1100">
243
+ </td>
244
+ </tr>
245
+
246
+ <tr>
247
+ <td align="center">
248
+ <img src="./Charts/timbre_tags_pie.PNG" width="1100">
249
+ </td>
250
+
251
+ <td align="center">
252
+ <img src="./Charts/fx_pie.PNG" width="1100">
253
+ </td>
254
+ </tr>
255
+ </table>
256
+
257
+ </center>
258
+
259
+ These visualizations provide a quick overview of the dataset’s coverage across:
260
+
261
+ - instrument families
262
+ - instrument sub-families
263
+ - timbral descriptors
264
+ - FX conditioning tags
265
+
266
+ ---
267
+
268
+ <center>
269
+
270
+ ## Scope
271
+
272
+ </center>
273
+
274
+ Foundation-1 is focused specifically on **sample-generation workflows**.
275
+
276
+ It was trained to generate:
277
+
278
+ - musical loops
279
+ - melodic phrases
280
+ - chordal material
281
+ - arps
282
+ - top-line ideas
283
+ - basslines
284
+ - textures
285
+ - production-ready instrumental content
286
+
287
+ It was **not designed** as:
288
+
289
+ - a full song generator
290
+ - a drum generator
291
+ - a general-purpose music captioning model
292
+
293
+ ---
294
+
295
+ <center>
296
+
297
+ ## Final Note
298
+
299
+ </center>
300
+
301
+ Foundation-1 was built to give producers structured control over:
302
+
303
+ - what the sound is
304
+ - how it behaves musically
305
+ - how it feels sonically
306
+ - how it sits in a production context
307
+
308
+ The dataset and training design were built around that exact goal.
309
+
310
+ For examples, prompting guidance, and model capabilities, see the **[main repository page](./README.md)**.