Foundation-1.2 Keybeds Training & Inference Strategy
How Foundation-1 learns a consistent timbral identity across pitch, and how RC Stable Audio Tools turns a simple text prompt into a playable instrument.
This page documents the specialized strategy used for the **Foundation-1.2 Keybeds** checkpoint and the matching inference pipeline used by **RC Stable Audio Tools**.
The central problem is simple to describe:
> Generating one convincing note is not the same as generating an entire instrument.
A playable keybed requires many independently rendered pitches to still sound like they belong to the **same underlying instrument**. The challenge is therefore not only pitch accuracy, but **cross-pitch timbral consistency**.
The Keybed model and inference pipeline were designed around that problem.
---
## Why Isolated Notes Were Not Enough
A conventional one-shot dataset can teach a model that a prompt such as:
`Grand Piano, Warm, Gritty`
should produce a plausible piano-like sample.
That alone does not guarantee that independently generated `C2`, `F3`, `C5`, and `A6` will sound like the same piano.
When every note is treated as an unrelated training example, the model has less direct pressure to learn the relationship between:
- a sound and its neighboring pitches
- the same timbre across octave changes
- transient behavior across register
- harmonic density across pitch
- formant / resonant movement across a full instrument range
The Keybed training strategy therefore introduced **multi-note context**.
---
## Multi-Note Training
Rather than only presenting isolated notes, notes belonging to the same source sound were combined into short multi-note training examples inside the model's approximately **20-second audio window**.
The training material intentionally varied the pitch relationships shown to the model.
Examples included:
- adjacent notes in chromatic order
- short ordered runs
- octave jumps
- wider pitch intervals
- collections of roughly **2–5 notes**
- longer sequential groups where appropriate
- overlapping groups in which the same pitches reappeared in different contexts
The purpose was not to teach one rigid note sequence.
The goal was to repeatedly expose the model to the same underlying timbre while changing the **pitch relationship and surrounding context**.
This encourages the model to learn:
1. what an individual note should sound like
2. how that sound changes as pitch moves
3. which parts of the timbre should remain stable across the instrument
---
## Sequential Training Chunks
One of the dataset-building stages also constructed **six-note sequential training examples** from existing per-note Keybed renders.
Each note occupied approximately **3.0 seconds**, with a **0.25 second gap** between neighboring notes. Six notes therefore fit into approximately **19.25 seconds**, making efficient use of the model's ~20-second training window.
A sequence could conceptually look like:
`C2 → C#2 → D2 → D#2 → E2 → F2`
while other training examples exposed the same source timbre through different pitch relationships.
This was useful because the model could hear multiple notes belonging to the same instrument within one training example rather than seeing every pitch in isolation.
Visualization of the keybed training methodology from the companion video.
---
## Specialized Checkpoint
The Foundation-1.2 Keybeds checkpoint retains the same underlying **Foundation-1** architecture.
The specialization comes from the training data organization, conditioning, and training duration rather than from a new neural-network architecture.
### Keybed Checkpoint
- **Training endpoint:** Epoch 3 / Step 8120
- **Optimizer:** AdamW
- **Learning rate:** `1e-5`
- **Weight decay:** `1e-3`
- **Scheduler:** InverseLR
- **EMA:** Enabled
- **Sample rate:** 44,100 Hz
- **Channels:** Stereo
- **Training window:** 882,000 samples (~20 seconds)
The longer training cycle was selected because **cross-pitch consistency continued to improve after isolated one-shot quality had already become strong**.
For comparison, the dedicated The Foundation-1.2 Samples checkpoint was selected much earlier at **Epoch 0 / Step 560**.
The Foundation-1.2 Keybeds checkpoint can still generate excellent individual one-shots. The models were split because continuing training toward stronger keybed consistency reduced the broader loop-generation quality preserved by the earlier checkpoint.
---
## Inference Mirrors the Training Structure
The RC Stable Audio Tools Keybed pipeline deliberately mirrors the structure the model saw during training.
A user does **not** need to manually write the full conditioning prompt.
They can enter something simple such as:
`Grand Piano, Warm, Gritty`
The application then converts that descriptor into the structured prompt grammar expected by the model.
For a single note, the internal structure follows the general form:
```text
Keybed, Timbre Profile, , , Target Note,
```
For a multi-note chunk, the structure follows the general form:
```text
Keybed, Sequence, Timbre Profile, , ,
Chromatic Chunk, Note Sequence, , , , ...
```
For example, a user-facing descriptor:
```text
Grand Piano, Warm, Gritty
```
may become an internal sequence prompt conceptually similar to:
```text
Keybed, Sequence, Timbre Profile, Grand Piano, Warm, Gritty, Dry,
Chromatic Chunk, Note Sequence, C4, C#4, D4, D#4, E4, F4
```
The exact control grammar is intentionally handled by the application rather than exposed as something the musician has to type.
This separation is important:
- **the user describes the sound**
- **the application describes the generation task**
That is the basic pattern I recommend for anyone building another front end, VST, or sampler around the Keybed model.
---
## Stable Descriptor + Stable Seed
The descriptor is kept stable across the full instrument.
If the user requests:
`Grand Piano, Warm, Gritty`
that same timbral description is reused for every keybed chunk.
The pipeline also resolves **one seed for the run** and reuses that same seed across every generation chunk.
If a random seed is requested, the application resolves it once rather than rolling a new random seed for every section of the keyboard.
This is intentional.
Changing both the note range and the random seed at the same time makes it easier for the generated sonic identity to drift. Reusing the seed gives the model another stable reference while the pitch-conditioning portion of the prompt changes.
A preview can also capture the resolved seed. If the user proceeds from that preview to a full keybed using the same descriptor, the full sampler generation can reuse that preview seed across all chunks.
---
## Chunked Generation
A full keybed is divided into manageable **chromatic chunks of up to six notes**.
For each chunk, RC Stable Audio Tools:
1. keeps the user's timbral descriptor unchanged
2. inserts the required Keybed sequence-control tokens
3. inserts the exact note sequence for that chunk
4. reuses the same resolved seed
5. generates the chunk as one sequential audio render
The standard timing is:
- **3.0 seconds per note**
- **0.25 seconds between notes**
A six-note chunk therefore occupies approximately:
`6 × 3.0s + 5 × 0.25s = 19.25s`
This closely matches the multi-note sequence structure used during training.
---
## Deterministic Slicing
Each generated chunk is accompanied by metadata describing the expected time range of every note.
Conceptually:
```text
C4 0.00s → 3.00s
C#4 3.25s → 6.25s
D4 6.50s → 9.50s
...
```
The exporter therefore does not need to guess where notes begin and end.
The generated sequence can be deterministically sliced back into its individual note samples using the known **3.0 second note length** and **0.25 second gap**.
This is one reason the pipeline uses sequential notes rather than chord-stacked notes: the resulting audio can be cleanly separated back into individual sampler zones.
---
## From Generated Audio to a Playable Instrument
After all chunks have been generated, the pipeline:
1. slices the sequential renders into individual note files
2. preserves the note-to-pitch mapping
3. optionally trims unused sample tails
4. writes run / generation metadata
5. maps the notes across the sampler keyboard
6. packages the result for **DecentSampler** and/or **SFZ**
The result is a conventional playable sample instrument built from generated source material.
The model does the generative part; the exporter turns those generations into something a producer can actually use.
This is an intentional design choice. Foundation-1 generates the instrument source material, then gets out of the way.
This enables keybeds to be exported to low-spec hardware.
---
## Layered Instruments
The same workflow can be repeated independently for multiple prompts and the resulting keybeds can be combined into a layered instrument.
RC Stable Audio Tools currently supports:
- **Main**
- **Support 1**
- **Support 2**
Each layer can come from an independently generated keybed and retain its own timbral identity.
The exporter can then provide independent layer volumes and sampler controls while treating the three generated keybeds as one playable instrument.
This makes prompts useful not only for recreating recognizable instrument classes, but also for constructing hybrid sounds that may not correspond to a traditional instrument at all.
---
## Why This Matters for VST / Front-End Developers
The most important part of the inference design is that an end user should not need to understand the model's internal prompt grammar.
A front end can expose a simple text field:
```text
Grand Piano, Warm, Gritty
```
and handle the rest programmatically.
A practical implementation only needs to maintain a few invariants:
- preserve the user's core descriptor across the instrument
- inject the correct target-note or note-sequence conditioning
- generate related notes in multi-note chunks
- keep the seed stable across those chunks
- record deterministic slice boundaries
- map the extracted note files into the target sampler or VST format
The generated audio does not need to stay tied to RC Stable Audio Tools.
The same model could be wrapped by another sampler, standalone application, DAW tool, or purpose-built VST using the same general strategy.
The RC implementation is therefore intended both as the reference workflow for Foundation-1 and as an example of how other developers can integrate the Keybed model into their own instrument-generation systems.
Some possible upgrades could include octave based-generation, sound-palette saving/exporting etc.
---
## Applicability to Other Audio Architectures
The Keybed specialization was achieved through **dataset structure, conditioning, and training strategy**, rather than by introducing a new model architecture specifically for cross-pitch consistency.
Because of that, a similar multi-note training strategy may be applicable to other text-conditioned audio architectures where consistent identity across pitch is desirable.
However, I cannot currently separate how much of the result comes from the Keybed training strategy itself versus the capabilities already present in **Foundation-1**, which already contained substantial knowledge of musical instruments, pitch, and timbre.
In other words, the strategy clearly reinforced cross-pitch consistency in Foundation-1, but I have not tested whether the same method would produce equivalent results when applied to a substantially different model architecture or a model without the same prior musical representation.
---
## Prompt Semantics and Extreme Registers
The Keybed model is not designed around the assumption that every text prompt should remain meaningful across the full MIDI range.
A prompt such as:
`Sub Bass, Deep, Thick`
may make sense in the lower registers, but at **C7** the result is no longer functioning perceptually as a bass regardless of whether some aspects of the original timbre are preserved.
This becomes increasingly important at very extreme registers:
- complex low-frequency sounds can become unstable or less consistent when pushed very low
- bass-oriented prompts lose their defining role when moved into upper registers
- formant-dependent sounds can change identity as pitch moves away from their natural range
- waveforms often simplify perceptually at extreme highs and can collapse toward **high-pitched pure-tone-like behavior**
- some acoustic instruments simply do not have a meaningful real-world equivalent several octaves outside their playable range
For that reason, prompt design and pitch range should be considered together.
Some timbral drift at the edge of an instrument is a genuine model limitation, but some of it is also inherent to the task: there is not always a single correct answer for what an arbitrary text-described sound should become when moved far outside the range where its defining characteristics normally exist.
In my internal testing, roughly **90% of generated instruments** maintain a consistent relationship between pitch and timbral identity across the supported keybed ranges.
This should be understood as a qualitative internal estimate, not a programmatic benchmark. There is currently no reliable automated metric that can determine whether two generated notes are perceptually the "same timbre" while also accounting for the natural spectral changes caused by pitch.
RC Stable Audio Tools therefore uses practical export ranges and clamping rather than attempting to force every generated sound across the entire MIDI keyboard.
---
## Companion Videos
For a broader explanation of the Foundation-1.2 upgrade and a practical walkthrough of the Keybed workflow:
- **[Foundation-1.2 Samples & Keybeds Companion Video](https://www.youtube.com/watch?v=x0KnmzH8Mmk)**
- **[Keybed Guided Demo](https://x.com/RoyalCities/status/2097733712293109842?s=20)**
The full audio-example playthrough remains linked from the main **[Foundation-1 model page](./README.md)**.
---
## Summary
The Keybed workflow is built around one central idea:
> **Pitch should change while the sonic identity stays coherent.**
Training reinforces that relationship by presenting multiple pitches from the same source timbre in shared and overlapping contexts.
Inference preserves the same relationship by keeping the descriptor and seed stable while changing only the note-sequence conditioning needed for each section of the keyboard.
The resulting sequential renders are then sliced, mapped, and exported as conventional playable sampler instruments.
For model downloads, audio examples, prompting guidance, and the RC Stable Audio Tools interface, return to the **[Foundation-1 model page](./README.md)**.