File size: 57,930 Bytes
1ff6964 ab3c151 1ff6964 ab3c151 1ff6964 2048f05 1ff6964 ab3c151 1ff6964 ab3c151 1ff6964 ab3c151 1ff6964 da1c8ad 1ff6964 ab3c151 1ff6964 da1c8ad 1ff6964 da1c8ad 1ff6964 ef3a53b 1ff6964 ef3a53b 9c056d9 ab3c151 ef3a53b 1ff6964 da1c8ad 1ff6964 ab3c151 1ff6964 da1c8ad 1ff6964 da1c8ad 1ff6964 da1c8ad 1ff6964 da1c8ad 1ff6964 ca3f095 1ff6964 da1c8ad 1ff6964 da1c8ad 1ff6964 da1c8ad 1ff6964 da1c8ad 1ff6964 da1c8ad 1ff6964 da1c8ad 1ff6964 ab3c151 1ff6964 ab3c151 1ff6964 ab3c151 1ff6964 ab3c151 1ff6964 ab3c151 93c2dc7 1ff6964 ab3c151 1ff6964 ab3c151 93c2dc7 1ff6964 ab3c151 1ff6964 ab3c151 1ff6964 ab3c151 1ff6964 ab3c151 1ff6964 ab3c151 1ff6964 da1c8ad 1ff6964 ab3c151 1ff6964 da1c8ad 1ff6964 ab3c151 1ff6964 ab3c151 1ff6964 da1c8ad 1ff6964 ab3c151 1ff6964 ab3c151 9c056d9 1ff6964 9c056d9 ab3c151 9c056d9 1ff6964 ab3c151 1ff6964 da1c8ad 1ff6964 ab3c151 1ff6964 ab3c151 1ff6964 ab3c151 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 887 888 889 890 891 892 893 894 895 896 897 898 899 900 901 902 903 904 905 906 907 908 909 910 911 912 913 914 915 916 917 918 919 920 921 922 923 924 925 926 927 928 929 930 931 932 933 934 935 936 937 938 939 940 941 942 943 944 945 946 947 948 949 950 951 952 953 954 955 956 957 958 959 960 961 962 963 964 965 966 967 968 969 970 971 972 973 974 975 976 977 978 979 980 981 982 983 984 985 986 987 988 989 990 991 992 993 994 995 996 997 998 999 1000 1001 1002 1003 1004 1005 1006 1007 1008 1009 1010 1011 1012 1013 1014 1015 1016 1017 1018 1019 1020 1021 1022 1023 1024 1025 1026 1027 1028 1029 1030 1031 1032 1033 1034 1035 1036 1037 1038 1039 1040 1041 1042 1043 1044 1045 1046 1047 1048 1049 1050 1051 1052 1053 1054 1055 1056 1057 1058 1059 1060 1061 1062 1063 1064 1065 1066 1067 1068 1069 1070 1071 1072 1073 1074 1075 1076 1077 1078 1079 1080 1081 1082 1083 1084 1085 1086 1087 1088 1089 1090 1091 1092 1093 1094 1095 1096 1097 1098 1099 1100 1101 1102 1103 1104 1105 1106 1107 1108 1109 1110 1111 1112 1113 1114 1115 1116 1117 1118 1119 1120 1121 1122 1123 1124 1125 1126 1127 1128 1129 1130 1131 1132 1133 1134 1135 1136 1137 1138 1139 1140 1141 1142 1143 1144 1145 1146 1147 1148 1149 1150 1151 1152 1153 1154 1155 1156 1157 1158 1159 1160 1161 1162 1163 1164 1165 1166 1167 1168 1169 | ---
language: en
tags:
- audio
- text-to-synth
- music-generation
- sample-generation
- Music Production
- Audio-to-Audio
- fine-tuning
- stable-audio
datasets:
- custom
model_name: Foundation-1
base_model: stabilityai/stable-audio-open-1.0
license: other
license_name: stabilityai-community-license
license_link: https://stability.ai/license
---
<center><img src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/Charts/banner.PNG" alt="Foundation-1 Banner" width="100%"></center>
<center>
<h1 style="font-size: 34px;"><u>Foundation-1</u></h1>
</center>
<center>
<h3 style="font-size: 20px;">Structured text-to-sample and text-to-instrument generation for modern music production</h3>
</center>
---
<h2 align="center">Overview</h2>
**Foundation-1** is a family of text-conditioned audio models designed around **music production workflows**. Rather than treating audio generation as broad caption-to-music generation, Foundation-1 was trained around structured controls for **instrument identity, timbre, FX, musical behavior, pitch, timing, and tonality**.
The original Foundation-1 checkpoint focuses on **structured loop generation**. Newer specialized checkpoints extend the same conditioning system into **individual one-shots** and **pitch-consistent playable keybeds**.
This means Foundation-1 can now be used at several levels of a production workflow:
- generate a tempo-synced musical loop
- generate an individual note or sound
- generate multiple pitch-consistent notes from the same text prompt
- assemble those notes into a playable sampler instrument
- combine multiple generated keybeds into layered instruments
The goal is not simply to generate finished audio. The specialized keybed workflow is designed to let the model **generate the source material for an instrument**, then hand control back to the producer.
---
<h2 align="center">Foundation-1 Model Family</h2>
| Model | Primary Use | Notes |
|---|---|---|
| **Foundation-1** | Loop generation | Original checkpoint for structured, BPM-aware, bar-aware musical loops |
| **Foundation-1.2 Samples** | Loop Generation / One-Shot generation | Selected earlier in specialized training to preserve stronger short-form loop-generation quality |
| **Foundation-1.2 Keybeds** | One-Shot generation, Pitch-consistent sample generation and playable instruments | Trained longer to reinforce timbral consistency across pitch and register |
The **Foundation-1.2 Keybeds model is not limited to full keybed generation**. It remains highly capable at generating individual samples and one-shots. Its specialization reflects the additional training needed to maintain a consistent sonic identity across many generated pitches.
The checkpoints were separated because the longer keybed-focused training improved **cross-pitch consistency**, while reducing the broader loop-generation quality preserved by the earlier Samples checkpoint.
---
<h2 align="center">Text-to-Synth / Keybed Generation</h2>
The Keybed checkpoint extends Foundation-1 from **text-to-sample** generation into a practical **text-to-instrument** workflow.
A user can describe a sound using the same instrument and timbre vocabulary used elsewhere in Foundation-1, generate pitch-specific samples across a keyboard, and automatically assemble the result into a playable sampler instrument.
Examples can range from conventional instruments:
`Grand Piano, Warm, Gritty, Wet, Low Reverb`
to synthetic or hybrid sounds:
`Cello, Harp, Rich, Clean, Choir, Pluck, Formant Vocal, Wet, High Reverb`
Because pitch is generated rather than conventionally pitch-shifted from a single root sample, the model can create natural timbral variation across registers while maintaining the identity of the prompted sound.
For the intended workflow, the Keybed model is best used with **RC Stable Audio Tools**, which handles the multi-note generation process and instrument assembly automatically. Generated keybeds can be exported for use as playable sampler instruments rather than remaining as isolated audio generations.
The user-facing prompt is intentionally simple. A descriptor such as:
`Grand Piano, Warm, Gritty`
is expanded internally into the structured Keybed conditioning grammar used by the model. RC Stable Audio Tools adds the required Keybed / sequence / note information, keeps the user descriptor stable across the instrument, reuses the same resolved seed across generation chunks, then slices and maps the resulting notes into the exported sampler.
This makes the workflow useful not only as a demo interface, but also as a reference implementation for developers interested in building their own VST, sampler, or instrument-generation front end around the model.
For the full prompt-injection and inference design, see the **[Keybed Training & Inference Strategy](./keybed_training_strategy.md)**.
<p align="center">
<img src="./Charts/ui_tri_layer_instrument.PNG" alt="Foundation-1.2 Exported Keybed" width="800">
</p>
<p align="center">
<sub><i>Foundation-1.2 Keybed generation workflow in RC Stable Audio Tools.</i></sub>
</p>
---
<h2 align="center">What Foundation-1 Does</h2>
- **Generates musically coherent loops** for production workflows
- **Generates individual note-specific one-shots**
- **Generates pitch-consistent multi-note keybeds**
- **Builds playable sampler instruments through the RC Stable Audio Tools workflow**
- **Supports layered keybeds** built from multiple independently prompted sounds
- **Understands BPM and bar count** for structured loop generation
- **Locks to major and minor keys** across western music theory
- **Supports enharmonic equivalents** when prompting scales and keys
- **Separates instrument identity from timbral character**
- **Supports timbral mixing** by combining instrument and sonic descriptors
- **Responds to FX tags** such as reverb, delay, distortion, and modulation
- **Uses notation-style prompt structure** to encourage coherent phrasing, melodic shape, rhythmic behavior, and harmonic motion
- **Produces perfect loops** within supported BPM / bar denominations
- **Understands Wet vs Dry production context** — adding terms like *Dry* encourages minimal FX processing, while *Wet* or FX tags produce more processed, spatial, or effected sounds.
---
<h2 align="center">Why It Feels Different</h2>
Most audio models can react to broad prompt terms like “warm pad” or “bright synth.” with inconsistent results. Foundation-1 was designed to go further by treating the sound as a layered system:
1. **Instrument Family** – what broad source category the sound belongs to
2. **Sub-Family** – the more specific instrument role or identity
3. **Timbre Tags** – the tonal, spectral, or textural character
4. **FX Tags** – the processing layer applied to the sound
5. **Notation / Structure Tags** – the musical behavior of the generated phrase
This layered conditioning approach is a major reason Foundation-1 is able to deliver both **high musicality** and **high prompt control** at the same time.
---
<h2 align="center">Audio Showcase</h2>
### Loop Showcase
<div style="text-align: center; margin: 20px 0;">
<table style="width: 100%; border-collapse: collapse; margin: 0 auto;">
<thead>
<tr>
<th style="border: 1px solid #000; padding: 8px; text-align: left;">Prompt</th>
<th style="border: 1px solid #000; padding: 8px; text-align: center;">Audio</th>
</tr>
</thead>
<tbody>
<tr>
<td style="border: 1px solid #000; padding: 8px;">Bass, FM Bass, Medium Delay, Medium Reverb, Low Distortion, Phaser, Sub Bass, Bass, Upper Mids, Acid, Gritty, Wide, Dubstep, Thick, Silky, Warm, Rich, Overdriven, Crisp, Deep, Clean, Pitch Bend, 303, 8 Bars, 140 BPM, E minor</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;">
<audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/example_1.mp3" type="audio/mpeg">
</audio>
</td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px;">Sub Bass, Bass, Gritty, Small, Square, Bass, Dark, Digital, Thick, Clean, Simple, Bassline, Epic, Choppy, Melody, 4 Bars, 150 BPM, G# minor</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;">
<audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/example_2.mp3" type="audio/mpeg">
</audio>
</td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px;">Flute, Pizzicato, Punchy, Present, Ambient, Nasal, Melody, Epic, Airy, Slow Speed, 8 Bars, 150 BPM, E minor</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;">
<audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/example_3.mp3" type="audio/mpeg">
</audio>
</td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px;">High Saw, Spacey, Lead, Warm, Silky, Smooth, 303, Synth Lead, Medium Reverb, Low Distortion, Upper Mids, Mids, Pitch Bend, Arp, 8 Bars, 140 BPM, F minor</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;">
<audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/example_4.mp3" type="audio/mpeg">
</audio>
</td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px;">Trumpet, Warm, Complex Arp Melody, High Reverb, Low Distortion, Smooth, Silky, Texture, 8 Bars, 130 BPM, C minor</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;">
<audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/example_5.mp3" type="audio/mpeg">
</audio>
</td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px;">Synth, Pad, Chord Progression, Rising, Digital, Bass, Fat, Near, Wide, Silky, Warm, Focused, 8 Bars, 110 BPM, D major</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;">
<audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/example_6.mp3" type="audio/mpeg">
</audio>
</td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px;">Piccolo, Flute, Airy, Music Box, plucked, complex melody, 8 Bars, 140 BPM, C# minor</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;">
<audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/example_7.mp3" type="audio/mpeg">
</audio>
</td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px;">Synth Lead, Wavetable Bass, Low Distortion, High Reverb, Sub Bass, Upper Mids, Acid, Gritty, Wide, Thick, Silky, Warm, Rich, Overdriven, Crisp, Clean, 303, Complex, 8 Bars, 140 BPM, F minor</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;">
<audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/example_8.mp3" type="audio/mpeg">
</audio>
</td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px;">Fiddle, Bowed Strings, Full, Clean, Spacey, Rich, Intimate, Thick, Rolling, Arp, Fast Speed, Complex, 8 Bars, 128 BPM, B minor</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;">
<audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/example_9.mp3" type="audio/mpeg">
</audio>
</td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px;">Chiptune, Chord Progression, Pulse Wave, Medium Reverb, 8 Bars, 128 BPM, D minor</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;">
<audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/example_10.mp3" type="audio/mpeg">
</audio>
</td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px;">Kalimba, Mallet, Medium Reverb, Overdriven, Wide, Metallic, Thick, Sparkly, Upper Mids, Bright, Airy, Alternating, Chord Progression, Atmosphere, Spacey, Fast Speed, 8 Bars, 120 BPM, B minor</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;">
<audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/example_11.mp3" type="audio/mpeg">
</audio>
</td>
</tr>
</tbody>
</table>
</div>
### One-Shot Showcase
The examples below show short, note-specific one-shot generations across acoustic, synthetic, bass, FX, and hybrid timbres.
<div style="text-align: center; margin: 20px 0;">
<table style="width: 100%; border-collapse: collapse; margin: 0 auto;">
<thead>
<tr>
<th style="border: 1px solid #000; padding: 8px; text-align: left;">Prompt</th>
<th style="border: 1px solid #000; padding: 8px; text-align: center;">Note</th>
<th style="border: 1px solid #000; padding: 8px; text-align: center;">Audio</th>
</tr>
</thead>
<tbody>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">FX, Sharp, Subdued, Distant, Sparkly, Round, Hit, Deep, Low Reverb, Sub</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><b>D#1</b></td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Oneshot_1.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Synth Bass, Metallic, Rich, Punchy, Digital, Medium Reverb</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><b>C2</b></td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Oneshot_2.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Reese Bass, Hit, Bright, Wavetable, Noisy, Big, Thick, Airy, Deep</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><b>F#2</b></td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Oneshot_3.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Synth Lead, Shiny, Rich, Choir, Bright, Smooth, Short, Synthetic Vox</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><b>F5</b></td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Oneshot_4.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Tuba, Big, Bright, Clean, Biting, Sustained</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><b>G#3</b></td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Oneshot_5.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Clavinet, Wobble, Breathy, Bright, Deep, Gritty, Impact, Sparkly</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><b>C#4</b></td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Oneshot_6.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Cello, Airy, Wide, Woody, Muffled, Smooth</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><b>D#3</b></td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Oneshot_7.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Reese Bass, Subdued, Round, Growl, Hit, Deep, Woody, Fat, Metallic, Medium Delay, Medium Distortion</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><b>D#2</b></td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Oneshot_8.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Digital Piano, Retro, Smooth, Mono, Warm, Low Reverb</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><b>E3</b></td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Oneshot_9.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Supersaw, Big, Smooth, Warm, Vintage, Crisp, Analog, Wide, Muffled, High Reverb</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><b>G2</b></td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Oneshot_10.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Synth Lead, Bright, Shiny, Soft, Warm, Punchy, Smooth, Medium Delay</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><b>C#4</b></td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Oneshot_11.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Sustained, Trumpet, Airy, Wide, High Reverb</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><b>F3</b></td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Oneshot_12.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">808, Thick, Acid, Hit, FX, Rumble, Wide, Big, Distortion</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><b>F1</b></td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Oneshot_13.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Violin, Vintage, Thick, Chiptune, Spacey, Nasal, Medium Delay</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><b>A#5</b></td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Oneshot_14.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Grand Piano, Near, Snappy, Distant, Bell, Bright, Focused, Delay</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><b>D3</b></td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Oneshot_15.mp3" type="audio/mpeg">
</audio></td>
</tr>
</tbody>
</table>
</div>
### Keybed Showcase
The following examples demonstrate generated keybeds across acoustic, synthetic, and hybrid timbres.
A full video playthrough of the Keybed Showcase is available **[here](https://x.com/RoyalCities/status/2097733715543609445?s=20)**, which shows the associated MIDI used to audition each generated instrument.
<div style="text-align: center; margin: 20px 0;">
<table style="width: 100%; border-collapse: collapse; margin: 0 auto;">
<thead>
<tr>
<th style="border: 1px solid #000; padding: 8px; text-align: left;">Prompt</th>
<th style="border: 1px solid #000; padding: 8px; text-align: center;">Audio</th>
</tr>
</thead>
<tbody>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Soft, Silky, Spectral, Smooth, Spacey, Synth, Pad, Subdued, Wet, High Reverb</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Keybed_1.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Pad, Focused, Buzzy, Big, Deep, Supersaw, Pulse, Dry</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Keybed_2.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Pan Flute, Hollow, Dark, Smooth, Sustained, Wide, Silky, Wet, Medium Reverb</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Keybed_3.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Violin, Woody, Pizzicato, Focused, Breathy, Airy, Wet, High Reverb</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Keybed_4.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Grand Piano, Warm, Gritty, Wet, Low Reverb</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Keybed_5.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Grand Piano, Cold, Sparkly, Wet, Low Reverb</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Keybed_6.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Digital Piano, Noisy, Spiccato, Rich, Pluck, Warm, Swell, Clean, Dry</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Keybed_7.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Cello, Fat, Warm, Metallic, Sharp, Near, Sustained, Dry</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Keybed_8.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Church Bell, Glassy, Sparkly, Wet, High Reverb</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Keybed_9.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Thick, Saw, Neuro, Reese Bass, Dry</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Keybed_15.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Thick, Sine, Reese Bass, Wet, High Reverb, High Phaser</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Keybed_16.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Xylophone, Sustained, Analog, Warm, Woody, Dry</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Keybed_17.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Xylophone, Sustained, Analog, Warm, Woody, Bit Crushed, Wet</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Keybed_18.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Cello, Harp, Rich, Clean, Choir, Pluck, Formant Vocal, Wet, High Reverb</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Keybed_19.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Acid, Neuro, Synth Lead, Saw, Square, Wide, Focused, Pluck, Sustained, Wet, High Phaser</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Keybed_20.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Synth Lead, Present, Sharp, Spacey, Digital, Hollow, Focused, Clean, Wet, Cross Delay, High Reverb</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Keybed_21.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Dubstep, Neuro, Synth Lead, Pitch Bend, Acid, Wet, Cross Delay</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Keybed_22.mp3" type="audio/mpeg">
</audio></td>
</tr>
</tbody>
</table>
</div>
#### Generated Waveforms
Pure waveform keybeds are shown separately to make the learned oscillator shapes easier to compare directly.
<div style="text-align: center; margin: 20px 0;">
<table style="width: 100%; border-collapse: collapse; margin: 0 auto;">
<thead>
<tr>
<th style="border: 1px solid #000; padding: 8px; text-align: left;">Waveform</th>
<th style="border: 1px solid #000; padding: 8px; text-align: left;">Prompt</th>
<th style="border: 1px solid #000; padding: 8px; text-align: center;">Audio</th>
</tr>
</thead>
<tbody>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;"><b>Triangle</b></td>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Pure Tone, Triangle, Dry</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Keybed_10.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;"><b>Square</b></td>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Pure Tone, Square, Dry</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Keybed_11.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;"><b>Pulse</b></td>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Pure Tone, Pulse, Dry</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Keybed_12.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;"><b>Sine</b></td>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Pure Tone, Sine, Dry</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Keybed_13.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;"><b>Saw</b></td>
<td style="border: 1px solid #000; padding: 8px; text-align: left;">Pure Tone, Saw, Dry</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Keybed_14.mp3" type="audio/mpeg">
</audio></td>
</tr>
</tbody>
</table>
</div>
### Multi-Layered Keybeds
The inference pipeline can also combine three independently generated keybeds into a single layered instrument. The examples below use separate Main and Support prompts for each layer. Further the generated instrument allows independent volume control for each layer.
<div style="text-align: center; margin: 20px 0;">
<table style="width: 100%; border-collapse: collapse; margin: 0 auto;">
<thead>
<tr>
<th style="border: 1px solid #000; padding: 8px; text-align: left;">Layer Prompts</th>
<th style="border: 1px solid #000; padding: 8px; text-align: center;">Audio</th>
</tr>
</thead>
<tbody>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;"><b>Main:</b> Bright, Airy, Violin<br><b>Support 1:</b> Thick, Present, Male, Vocal, Choir<br><b>Support 2:</b> Digital String, Analog</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Keybed_23.mp3" type="audio/mpeg">
</audio></td>
</tr>
<tr>
<td style="border: 1px solid #000; padding: 8px; text-align: left;"><b>Main:</b> Marimba, Dark, Sub, Woody, Airy<br><b>Support 1:</b> Soft, Grand Piano, Dark<br><b>Support 2:</b> Music Box, Glassy</td>
<td style="border: 1px solid #000; padding: 8px; text-align: center;"><audio controls style="width: 260px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/Example_Keybed_24.mp3" type="audio/mpeg">
</audio></td>
</tr>
</tbody>
</table>
</div>
---
<h2 align="center">Core Capabilities</h2>
### 1. Musical Structure
Foundation-1 was trained to produce structured musical material rather than full music or generic textures. Musical Notation terms can encourage notation, chord progressions, melodies, arps, phrase direction, rhythmic density, and other musically relevant behaviors.
### 2. Instrument Identity
The model supports a broad instrument hierarchy spanning synths, keys, basses, bowed strings, mallets, winds, guitars, brass, vocals, and plucked strings.
### 3. Timbral Control
Foundation-1 is not limited to broad instrument naming. It also responds to timbral descriptors such as spectral shape, tone, width, density, texture, brightness, warmth, grit, space, and other sonic traits.
### 4. Timbral Mixing
Because instrument identity and timbral character were not collapsed into a single flat label, the model is especially strong at **timbral hybridization** and **layered sonic prompting**.
### 5. FX Prompting
The model supports a dedicated FX layer covering multiple forms of reverb, delay, distortion, phaser, and bitcrushing.
### 6. Loop Fidelity
Foundation-1 is built for **production-ready loop generation**, including BPM-aware and bar-aware structure within supported denominations.
### 7. One-Shot Generation
The specialized Foundation-1.2 Samples and Keybeds checkpoints generate short, note-specific sounds across acoustic, synthetic, bass, FX, and hybrid timbres. These can be used individually or as raw material for sampler instruments.
### 8. Pitch-Consistent Keybeds
The Keybed checkpoint was trained specifically to preserve a prompted timbral identity across changing pitch and register. RC Stable Audio Tools uses this capability to generate multiple note-specific samples and assemble them into playable instruments.
---
<h2 align="center">Conditioning Architecture</h2>
Foundation-1 was trained with a layered tagging hierarchy designed to improve control, composability, and prompt clarity.
### Hierarchy Overview
- **Major Family** → broad instrument class
- **Sub-Family** → more specific instrument role
- **Timbre Tags** → tonal / spectral / textural descriptors
- **FX Tags** → processing layer
- **Notation Tags** → musical behavior and phrasing
This makes it possible to prompt at different levels of abstraction. A user can stay broad with a family-level prompt like **Synth** or **Keys**, or get more specific with terms like **Synth Lead**, **Wavetable Bass**, **Grand Piano**, **Violin**, or **Trumpet**, then further shape the output using timbral and FX descriptors.
---
<h2 align="center">Instrument Coverage</h2>
### Major Families
Foundation-1 was trained across the following major instrument families:
- **Synth**
- **Keys**
- **Bass**
- **Bowed Strings**
- **Mallet**
- **Wind**
- **Guitar**
- **Brass**
- **Vocal**
- **Plucked Strings**
### Sub-Family Coverage
Foundation-1 includes a wide sub-family layer covering a broad range of production-relevant instrument roles, including but not limited to:
- Synth Lead
- Synth Bass
- Digital Piano
- Pluck
- Grand Piano
- Bell
- Pad
- Atmosphere
- Digital Strings
- FM Synth
- Violin
- Digital Organ
- Supersaw
- Wavetable Bass
- Rhodes Piano
- Cello
- Texture
- Flute
- Reese Bass
- Wavetable Synth
- Electric Bass
- Marimba
- Trumpet
- Pan Flute
- Choir
- Harp
- Church Organ
- Acoustic Guitar
- Hammond Organ
- Celesta
- Vibraphone
- Glockenspiel
- Ocarina
- Clarinet
- French Horn
- Tuba
- Oboe
<center><img src="./Charts/subfamilites_pie.PNG" alt="Sub-Family Chart" width="80%"></center>
---
<h2 align="center">Timbre System</h2>
One of Foundation-1’s main strengths is that it was not trained to treat timbre as an afterthought. Timbral character is directly represented in the prompt system, giving users control over not only *what* is being generated, but also *how it sounds*.
Representative timbre descriptors include:
- Warm
- Bright
- Wide
- Airy
- Thick
- Rich
- Tight
- Full
- Gritty
- Clean
- Retro
- Saw
- Crisp
- Focused
- Metallic
- Chiptune
- Dark
- 303
- Shiny
- Analog
- Present
- Sparkly
- Ambient
- Soft
- Smooth
- Cold
- Buzzy
- Deep
- Formant Vocal
- Round
- Punchy
- Nasal
- Vintage
- Growl
- Breathy
- Glassy
- Noisy
- Synthetic Vox
- Supersaw
- Bitcrushed
- Dreamy
<center><img src="./Charts/timbre_tags_pie.PNG" alt="Timbre Chart" width="80%"></center>
<h2 align="center">Why This Matters</h2>
This tagging design makes prompts much more flexible. Instead of only asking for an instrument, users can shape:
- tonal balance
- brightness / darkness
- width / intimacy
- clean vs driven character
- synthetic vs organic feel
- transient sharpness
- texture and density
- spatial character
This is especially useful for producers who want to guide the output toward a specific role in a mix rather than just a generic instrument label.
For a list of used tags please see the **[Tag Reference Sheet](./Master_Tag_Reference.md)**.
---
<h2 align="center">FX Layer</h2>
Foundation-1 includes a dedicated FX descriptor layer spanning multiple common production effects.
Representative FX tags include:
- Low Reverb
- Medium Reverb
- High Reverb
- Plate Reverb
- Low Delay
- Medium Delay
- High Delay
- Ping Pong Delay
- Stereo Delay
- Cross Delay
- Mono Delay
- Low Distortion
- Medium Distortion
- High Distortion
- Phaser
- Low Phaser
- Medium Phaser
- High Phaser
- Bitcrush
- High Bitcrush
<center><img src="./Charts/fx_pie.PNG" alt="FX Chart" width="80%"></center>
---
<h2 align="center">Musical Notation and Structure</h2>
Foundation-1 was trained with structured musical descriptors designed to improve phrase coherence, rhythmic intent, melodic motion, and prompt control.
These notation-style prompt terms help steer:
- chord progressions
- melodies
- top-line layers
- arpeggios
- phrase direction
- rhythmic density
- harmonic feel
- subdivision style
- simple vs complex motion
- sustained vs plucked behavior
- melodic contour and pacing
Examples of supported structural ideas may include terms such as:
- chord progression
- melody
- top melody
- arp
- triplets
- simple
- complex
- rising
- falling
- strummed
- sustained
- catchy
- epic
- slow
- fast
This notation layer is one of the main reasons Foundation-1 produces unusually coherent musical material instead of static or loosely related phrases. These can be mixed and matched as desired.
---
<h2 align="center">Tonal and Timing Support</h2>
Foundation-1 is designed for structured music production workflows and supports:
### Keys and Modes
- Major keys
- Minor keys
- Enharmonic equivalents
- Western 12-tone chromatic prompting
### Loop Structure
- Supported bar lengths: **4 Bars, 8 Bars**
- Supported BPM denominations: **100 BPM, 110 BPM, 120 BPM, 128 BPM, 130 BPM, 140 BPM, 150 BPM**
---
<h2 align="center">Prompt Structure</h2>
For best results, use **rich prompts built around the model’s tags**. These tags can be mixed and matched as needed. The model was trained on a structured hierarchy designed to encourage musically coherent sample generation.
### Layered Prompt Structure
[Instrument Family / Sub-Family], [Timbre], [Musical Behavior / Notation], [FX], [Key], [Bars], [BPM]
### Prompting Notes
- Start with a **clear instrument identity**
- Add **1–3 timbre descriptors** for stronger steering
- Include a **notation or musical structure term** for better phrase coherence
- Always include **Bars and BPM**, which define the musical loop length
- Ensure the **generation duration matches the requested musical structure**
- The **RC Stable Audio Fork automatically handles this timing alignment**
Use **FX and timbre tags sparingly at first**, then layer more once you understand the model’s behavior.
---
<h2 align="center">One Prompt → Multiple Outputs</h2>
Each row below uses the **exact same prompt**, but a different random seed.
The **timbre tags remain unchanged**, so the overall sound character stays consistent while the **melodic and musical content varies** between generations.
<div align="center">
<table style="width:100%; border-collapse: collapse;">
<thead>
<tr>
<th style="padding:8px; text-align:left;">Prompt</th>
<th style="padding:8px; text-align:center;">Output A</th>
<th style="padding:8px; text-align:center;">Output B</th>
<th style="padding:8px; text-align:center;">Output C</th>
</tr>
</thead>
<tbody>
<tr>
<td style="padding:8px; text-align:left;">
<b>Bass, FM Bass, Medium Delay, Medium Reverb, Low Distortion, Phaser, Acid, Gritty, Wide, Dubstep, Thick, Silky, Warm, Rich, Overdriven, Crisp, Deep, Clean, Triplets, 8 Bars, 150 BPM, A minor</b>
</td>
<td align="center">
<audio controls style="width:160px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/compare_example_1_a.mp3" type="audio/mpeg">
</audio>
</td>
<td align="center">
<audio controls style="width:160px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/compare_example_1_b.mp3" type="audio/mpeg">
</audio>
</td>
<td align="center">
<audio controls style="width:160px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/compare_example_1_c.mp3" type="audio/mpeg">
</audio>
</td>
</tr>
<tr>
<td style="padding:8px; text-align:left;">
<b>Gritty, Acid, Bassline, 303, Synth Lead, FM, Sub, Upper Mids, High Phaser, High Reverb, Pitch Bend, 8 Bars, 140 BPM, E minor</b>
</td>
<td align="center">
<audio controls style="width:160px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/compare_example_2_a.mp3" type="audio/mpeg">
</audio>
</td>
<td align="center">
<audio controls style="width:160px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/compare_example_2_b.mp3" type="audio/mpeg">
</audio>
</td>
<td align="center">
<audio controls style="width:160px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/compare_example_2_c.mp3" type="audio/mpeg">
</audio>
</td>
</tr>
<tr>
<td style="padding:8px; text-align:left;">
<b>Kalimba, Mallet, Medium Reverb, Overdriven, Wide, Metallic, Thick, Sparkly, Upper Mids, Bright, Airy, Small, Alternating Chord Progression, Atmosphere, Spacey, Fast, 4 Bars, 120 BPM, B minor</b>
</td>
<td align="center">
<audio controls style="width:160px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/compare_example_3_a.mp3" type="audio/mpeg">
</audio>
</td>
<td align="center">
<audio controls style="width:160px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/compare_example_3_b.mp3" type="audio/mpeg">
</audio>
</td>
<td align="center">
<audio controls style="width:160px;">
<source src="https://huggingface.co/RoyalCities/Foundation-1/resolve/main/examples/compare_example_3_c.mp3" type="audio/mpeg">
</audio>
</td>
</tr>
</tbody>
</table>
</div>
---
<h2 align="center">Recommended Workflow</h2>
Foundation-1 is best used with **RC Stable Audio Tools**, which is tuned around the model family, its metadata, and its structured prompting system.
**[RC Stable Audio Tools (Enhanced Fork)](https://github.com/RoyalCities/RC-stable-audio-tools)**
The interface supports different workflows depending on the checkpoint in use.
### Loop Generation
For the original Foundation-1 loop checkpoint, the interface provides:
- structured prompt building aligned with the training tags
- random prompt generation
- automatic BPM / bar timing alignment
- automatic MIDI extraction from generated audio
- generation settings tuned for Foundation-1
### Samples / One-Shot Generation
The specialized Samples workflow supports:
- note-specific generation
- instrument and timbre prompting
- short sample-focused inference
- rapid auditioning of generated sounds
The Keybed checkpoint can also generate individual one-shots effectively. The dedicated Samples checkpoint is provided because its earlier training endpoint preserves stronger general sample-generation quality.
### Keybed / Text-to-Synth Generation
The Keybed workflow is designed to be used in conjunction with **RC Stable Audio Tools** rather than as a sequence of manually generated independent notes.
The interface handles the pitch-aware generation pipeline, creates the source files required for the instrument, and can export completed keybeds for sampler use.
The workflow supports:
- automatic generation across keyboard registers
- pitch-specific text-conditioned samples
- playable keybed construction
- DecentSampler export
- SFZ export
- per-instrument ADSR controls
- built-in sampler effects and tone shaping
- layered keybeds using independently prompted Main / Support layers
- independent layer volume control
- generated instrument previews
<p align="center">
<img src="./Charts/ui_preview_gradio.PNG" alt="Foundation-1.2 Keybed exporter" width="800">
</p>
<p align="center">
<sub><i>Generated keybeds can be assembled and exported as playable DecentSampler or SFZ instruments.</i></sub>
</p>
### Recommended Interfaces
**[RC Stable Audio Tools (Enhanced Fork)](https://github.com/RoyalCities/RC-stable-audio-tools)**
**[Stable Audio Tools (Original Repository)](https://github.com/Stability-AI/stable-audio-tools)**
### Model Files
Foundation-1 is now distributed as multiple specialized checkpoints built around the same underlying architecture and conditioning system.
The specialized Foundation-1.2 Samples and Keybeds checkpoints retain the same Foundation-1 / SAO architecture. They therefore do not represent separate model architectures; the specialization comes from the training objective and selected checkpoint.
The release uses **16-bit model weights** to reduce the model footprint without changing the intended inference quality.
- `Foundation_1.safetensors` — loop-generation checkpoint
- `Foundation-1.2-Samples.safetensors` — sample-focused checkpoint
- `Foundation-1.2-Keybeds.safetensors` — keybed-focused checkpoint
- `model_config.json` — shared model configuration
### Basic Setup for RC Stable Audio Tools
1. Create a subfolder inside your `models` directory
2. Place the desired checkpoint and its compatible config in that folder
3. Launch the interface
4. Select the checkpoint from the model selector
5. Choose the matching Loop, One-Shot, or Keybed workflow
6. Prompt with layered instrument and timbre descriptors
For full keybeds, the RC interface is strongly recommended because it manages the multi-note inference and sampler-export process automatically.
### Hardware Requirements
Foundation-1 is designed to run locally on modern GPUs.
Typical VRAM usage during generation is approximately **~7 GB**.
For reliable operation, a GPU with **at least 8 GB of VRAM is recommended**.
### Generation Performance
Generation speed will vary depending on GPU model, generation mode, sample length, and system configuration.
On an **RTX 3090**, a standard individual generation is approximately **~7–8 seconds per sample**. Full single-layer keybed builds take approximately **50 seconds**.
---
<h2 align="center">Dataset and Training Philosophy</h2>
Foundation-1 was built around a **structured sample-generation philosophy**, rather than generic or genre-based audio captioning. The dataset consists entirely of **hand-crafted and labeled audio**, produced through a controlled augmentation pipeline.
At a high level, the training design emphasizes:
- structured musical loops
- instrument hierarchy
- explicit timbre representation
- dedicated FX descriptors
- notation-aware prompt terms
- strong production relevance
- broad reuse for compositional workflows
This design is central to the model’s **musical coherence and high degree of sonic control**.
For more details on the original dataset and training methodology, see the **[Training & Dataset Notes](./training_dataset_info.md)**.
### Specialized Samples and Keybeds Training
The Foundation-1.2 Samples and Keybeds checkpoints use the same underlying Foundation-1 architecture, but were trained with a lower learning rate of **`1e-5`**.
| Checkpoint | Training Endpoint | Specialization |
|---|---:|---|
| **Foundation-1.2 Samples** | Epoch 0 / Step 560 | General sample / one-shot generation |
| **Foundation-1.2 Keybeds** | Epoch 3 / Step 8120 | Cross-pitch timbral consistency and playable keybed generation |
Additional training configuration:
- **Optimizer:** AdamW
- **Learning Rate:** `1e-5`
- **Weight Decay:** `1e-3`
- **Scheduler:** InverseLR
- **EMA:** Enabled
- **Sample Rate:** 44,100 Hz
- **Channels:** Stereo
- **Training Window:** 882,000 samples (~20 seconds)
The longer Keybed run was selected to reinforce **cross-note and cross-register consistency**. The objective was not simply to improve isolated note quality, but to teach the model how a single sound identity should behave as pitch changes across an instrument.
The Keybed checkpoint used a dedicated multi-note training strategy rather than treating every pitch as an unrelated example. The generation pipeline mirrors that structure by injecting note-sequence conditioning, reusing a common seed across keybed chunks, slicing generated sequences back into individual notes, and packaging the resulting samples into playable instruments.
For a detailed explanation of both the training method and the inference/export pipeline, see the **[Keybed Training & Inference Strategy](./keybed_training_strategy.md)**.
---
<h2 align="center">Limitations</h2>
Foundation-1 is a specialized model family for **producer-facing sample and instrument generation**, not a general-purpose full-song generator.
Important notes:
- It performs best when prompted using vocabulary aligned with the training design
- It is optimized for **sample-generation workflows**, not open-ended genre captioning
- Only two genre tags were included (Dubstep Growls and Chiptune waveforms), primarily to reinforce waveform behaviors
- **Prompt quality matters** — structured layered prompts outperform vague natural language
- Some timbre tags exert stronger influence than others
- Certain tag combinations may require iteration to achieve the exact musical role or timbral blend desired
- **Percussion and drum sounds are outside the scope of this release**
The model is also optimized around **specific timing relationships between Bars, BPM, and generation duration**.
For example:
- an **8-bar loop at 100 BPM ≈ 19 seconds**
If the generation duration is shorter than the musical structure implied by the prompt (for example requesting an 8-bar loop but generating only 5 seconds), the model may produce **less coherent musical phrases**.
The **RC Stable Audio Fork automatically handles this timing alignment**, making this workflow much easier.
### Keybed-Specific Limitations
The Keybed model is designed around **practical instrument ranges**.
Some prompts simply stop making semantic or perceptual sense at extreme registers. For example, a **sub-bass at C7 is no longer functioning as a bass**, even if the model preserves some aspects of the original waveform or timbral character.
The same applies to many synthetic and acoustic sounds. As pitch rises into extreme upper registers, complex waveforms often become perceptually simpler and can collapse toward **high-pitched pure-tone-like behavior**. At the opposite end, very low pitches can become increasingly difficult to render consistently, and small pitch or waveform errors become more obvious.
For this reason, RC Stable Audio Tools clamps generation and export ranges to practical regions rather than forcing every generated sound across the full MIDI note range.
Prompt choice also matters. A timbre that has a natural low, mid, or high-register identity may become ambiguous when pushed several octaves outside that range. Some apparent "timbre drift" is therefore not only a model limitation; it is also a consequence of asking for a sound whose defining characteristics change as pitch moves far beyond where that sound is normally perceived.
In my internal testing, roughly **90% of generated instruments** maintain a consistent relationship between **pitch and timbral identity** across the supported keybed ranges. This is a qualitative estimate rather than a programmatic benchmark, because there is no reliable automated metric for determining whether two notes share the same perceived timbre.
The remaining edge cases are most likely to appear with:
- extreme low or high registers
- unconventional bass prompts outside bass ranges
- strongly formant-dependent sounds
- heavily processed or hybrid timbres
- prompts whose intended instrument identity becomes ambiguous at the requested pitch
The default export ranges were chosen as a practical compromise between **keyboard coverage, pitch accuracy, and timbral consistency**.
---
<h2 align="center">License</h2>
This model is licensed under the Stability AI Community License. It is available for non-commercial use or limited commercial use by entities with annual revenues below USD $1M. For revenues exceeding USD $1M, please refer to the repository license file for full terms.
---
<h3 align="center">Companion Videos</h3>
### Foundation-1 Overview
The original Foundation-1 video covering the model's design philosophy, structured prompting system, timbral control, and loop-generation workflow.
🎥 **[Watch the Foundation-1 Overview](https://youtu.be/O2iBBWeWaL8)**
### Foundation-1.2 Samples & Keybeds Update
A companion video covering the new **Samples** and **Keybeds** checkpoints, the updated training strategy and general journey to getting this made can be found here.
🎥 **[Watch the Foundation-1.2 Update](https://www.youtube.com/watch?v=x0KnmzH8Mmk)**
### Keybed Guided Demo
A hands-on walkthrough showing the Keybed model in use, cross-pitch timbral behavior, and the new **three-layer instrument exporter**.
🎥 **[Watch the Guided Keybed Demo](https://x.com/RoyalCities/status/2097733712293109842?s=20)**
---
---
<h2 align="center">Final Notes</h2>
Foundation-1 is intended as a **producer-facing model family for structured sample and instrument generation**, designed to augment music production.
Its goal is to let users explore sound in new ways while retaining precise control over:
- what the sound is
- how it behaves musically
- how it changes across pitch
- how it sits tonally
- how it feels sonically
- how it fits into a production workflow
The original loop model focuses on structured musical material. The specialized Foundation-1.2 Samples and Keybeds checkpoints extend that same conditioning system toward **sound design and playable instrument creation**.
That combination of **musical structure**, **instrument identity**, **timbral control**, **loop fidelity**, and **cross-pitch instrument generation** is what defines the Foundation-1 family.
|