Instructions to use devendradhakad/autodroid-litert-community-Qwen3-TTS-12Hz-0.6B-Base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LiteRT
How to use devendradhakad/autodroid-litert-community-Qwen3-TTS-12Hz-0.6B-Base with LiteRT:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Mirror AutoDroid pins from litert-community/Qwen3-TTS-12Hz-0.6B-Base@0eb3b8a47149
Browse files- AUTODROID_MIRROR.md +7 -0
- AUTODROID_SOURCE.json +113 -0
- README.md +96 -0
- codec_partA.tflite +3 -0
- codec_partB.tflite +3 -0
- licenses/Apache-2.0.txt +171 -0
- merges.txt +0 -0
- mtp_folded_int8.tflite +3 -0
- tables/codec_embedding_fp32.npy +3 -0
- tables/mtp_embeddings_fp16.npy +3 -0
- tables/text_embedding_fp16.npy +3 -0
- tables/text_projection_fp32.npz +3 -0
- talker_int4.tflite +3 -0
- vocab.json +0 -0
- voices/demo_speaker.npy +3 -0
AUTODROID_MIRROR.md
ADDED
|
@@ -0,0 +1,7 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# AutoDroid artifact mirror
|
| 2 |
+
|
| 3 |
+
This repository preserves selected, unmodified files from `litert-community/Qwen3-TTS-12Hz-0.6B-Base` at commit `0eb3b8a4714972b065c160faec6a12158caa9dc0`.
|
| 4 |
+
|
| 5 |
+
Original authorship, licenses, and notices remain applicable. This is an independent availability mirror and does not imply upstream endorsement.
|
| 6 |
+
|
| 7 |
+
See `AUTODROID_SOURCE.json` for original paths, byte sizes, and SHA-256 digests. Only files required by AutoDroid and upstream documentation are included; this is not a complete training or Transformers checkpoint.
|
AUTODROID_SOURCE.json
ADDED
|
@@ -0,0 +1,113 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"sourceRepository": "litert-community/Qwen3-TTS-12Hz-0.6B-Base",
|
| 3 |
+
"sourceRevision": "0eb3b8a4714972b065c160faec6a12158caa9dc0",
|
| 4 |
+
"purpose": "Unmodified pinned artifacts used by AutoDroid",
|
| 5 |
+
"licenseReview": {
|
| 6 |
+
"sourceRevision": "0eb3b8a4714972b065c160faec6a12158caa9dc0",
|
| 7 |
+
"declaredLicense": "apache-2.0",
|
| 8 |
+
"evidence": "https://huggingface.co/litert-community/Qwen3-TTS-12Hz-0.6B-Base/blob/0eb3b8a4714972b065c160faec6a12158caa9dc0/README.md",
|
| 9 |
+
"status": "approved",
|
| 10 |
+
"reason": "Pinned upstream card permits redistribution. Preserve upstream cards and notices, original authorship, and the applicable license texts; artifact bytes are unmodified.",
|
| 11 |
+
"additionalFiles": [
|
| 12 |
+
{
|
| 13 |
+
"localPath": "tools/hf_mirror/licenses/Apache-2.0.txt",
|
| 14 |
+
"pathInRepo": "licenses/Apache-2.0.txt",
|
| 15 |
+
"source": "https://www.apache.org/licenses/LICENSE-2.0.txt",
|
| 16 |
+
"sha256": "c98068a3b6a564e4c70ab7c2ee2c980725987908909e99915759efa28ac7b533"
|
| 17 |
+
}
|
| 18 |
+
]
|
| 19 |
+
},
|
| 20 |
+
"artifacts": [
|
| 21 |
+
{
|
| 22 |
+
"repo": "litert-community/Qwen3-TTS-12Hz-0.6B-Base",
|
| 23 |
+
"revision": "0eb3b8a4714972b065c160faec6a12158caa9dc0",
|
| 24 |
+
"filename": "codec_partA.tflite",
|
| 25 |
+
"sizeBytes": 162987924,
|
| 26 |
+
"sha256": "9d3733c2a4c0a8be734d49cde63f57b9d0bee9308e67df678664d25937499c8c",
|
| 27 |
+
"source": "app/src/main/java/com/example/autodroid/data/voice/model/VoiceModelCatalog.kt:207"
|
| 28 |
+
},
|
| 29 |
+
{
|
| 30 |
+
"repo": "litert-community/Qwen3-TTS-12Hz-0.6B-Base",
|
| 31 |
+
"revision": "0eb3b8a4714972b065c160faec6a12158caa9dc0",
|
| 32 |
+
"filename": "codec_partB.tflite",
|
| 33 |
+
"sizeBytes": 293705152,
|
| 34 |
+
"sha256": "3486fafde3a50743ccd7330e5009b5b688eb748fb9c138a255cd17fb1846155f",
|
| 35 |
+
"source": "app/src/main/java/com/example/autodroid/data/voice/model/VoiceModelCatalog.kt:212"
|
| 36 |
+
},
|
| 37 |
+
{
|
| 38 |
+
"repo": "litert-community/Qwen3-TTS-12Hz-0.6B-Base",
|
| 39 |
+
"revision": "0eb3b8a4714972b065c160faec6a12158caa9dc0",
|
| 40 |
+
"filename": "merges.txt",
|
| 41 |
+
"sizeBytes": 1671839,
|
| 42 |
+
"sha256": "599bab54075088774b1733fde865d5bd747cbcc7a547c5bc12610e874e26f5e3",
|
| 43 |
+
"source": "app/src/main/java/com/example/autodroid/data/voice/model/VoiceModelCatalog.kt:247"
|
| 44 |
+
},
|
| 45 |
+
{
|
| 46 |
+
"repo": "litert-community/Qwen3-TTS-12Hz-0.6B-Base",
|
| 47 |
+
"revision": "0eb3b8a4714972b065c160faec6a12158caa9dc0",
|
| 48 |
+
"filename": "mtp_folded_int8.tflite",
|
| 49 |
+
"sizeBytes": 229608368,
|
| 50 |
+
"sha256": "f5ab8f826e3dd68f14667af422145fe57233b445046e5ef42c01b59f82191b4b",
|
| 51 |
+
"source": "app/src/main/java/com/example/autodroid/data/voice/model/VoiceModelCatalog.kt:202"
|
| 52 |
+
},
|
| 53 |
+
{
|
| 54 |
+
"repo": "litert-community/Qwen3-TTS-12Hz-0.6B-Base",
|
| 55 |
+
"revision": "0eb3b8a4714972b065c160faec6a12158caa9dc0",
|
| 56 |
+
"filename": "tables/codec_embedding_fp32.npy",
|
| 57 |
+
"sizeBytes": 12583040,
|
| 58 |
+
"sha256": "47fa9e30f98b1528fc9b332d314f22a32fa33e187509a4d3537f8b2c31199e39",
|
| 59 |
+
"source": "app/src/main/java/com/example/autodroid/data/voice/model/VoiceModelCatalog.kt:217"
|
| 60 |
+
},
|
| 61 |
+
{
|
| 62 |
+
"repo": "litert-community/Qwen3-TTS-12Hz-0.6B-Base",
|
| 63 |
+
"revision": "0eb3b8a4714972b065c160faec6a12158caa9dc0",
|
| 64 |
+
"filename": "tables/mtp_embeddings_fp16.npy",
|
| 65 |
+
"sizeBytes": 62914688,
|
| 66 |
+
"sha256": "fea581b6a04f1cbec20b49511c36a00011411ccfba31f89b7571f82fe6b36706",
|
| 67 |
+
"source": "app/src/main/java/com/example/autodroid/data/voice/model/VoiceModelCatalog.kt:222"
|
| 68 |
+
},
|
| 69 |
+
{
|
| 70 |
+
"repo": "litert-community/Qwen3-TTS-12Hz-0.6B-Base",
|
| 71 |
+
"revision": "0eb3b8a4714972b065c160faec6a12158caa9dc0",
|
| 72 |
+
"filename": "tables/text_embedding_fp16.npy",
|
| 73 |
+
"sizeBytes": 622329984,
|
| 74 |
+
"sha256": "6fab9de0a8bc144aa3efefdacb5e8292b8499a9ae1d60fcee240bac528b7441e",
|
| 75 |
+
"source": "app/src/main/java/com/example/autodroid/data/voice/model/VoiceModelCatalog.kt:227"
|
| 76 |
+
},
|
| 77 |
+
{
|
| 78 |
+
"repo": "litert-community/Qwen3-TTS-12Hz-0.6B-Base",
|
| 79 |
+
"revision": "0eb3b8a4714972b065c160faec6a12158caa9dc0",
|
| 80 |
+
"filename": "tables/text_projection_fp32.npz",
|
| 81 |
+
"sizeBytes": 25179078,
|
| 82 |
+
"sha256": "ebb0f6a7aaacdbc903e825e33480b2da4d4b71c43c90acb0e89988049c77c100",
|
| 83 |
+
"source": "app/src/main/java/com/example/autodroid/data/voice/model/VoiceModelCatalog.kt:232"
|
| 84 |
+
},
|
| 85 |
+
{
|
| 86 |
+
"repo": "litert-community/Qwen3-TTS-12Hz-0.6B-Base",
|
| 87 |
+
"revision": "0eb3b8a4714972b065c160faec6a12158caa9dc0",
|
| 88 |
+
"filename": "talker_int4.tflite",
|
| 89 |
+
"sizeBytes": 255998768,
|
| 90 |
+
"sha256": "e03df54e73ed1f88b2ae6d47bbf82dd64ea90a3620d753a0f3c8d6a8d60848db",
|
| 91 |
+
"source": "app/src/main/java/com/example/autodroid/data/voice/model/VoiceModelCatalog.kt:197"
|
| 92 |
+
},
|
| 93 |
+
{
|
| 94 |
+
"repo": "litert-community/Qwen3-TTS-12Hz-0.6B-Base",
|
| 95 |
+
"revision": "0eb3b8a4714972b065c160faec6a12158caa9dc0",
|
| 96 |
+
"filename": "vocab.json",
|
| 97 |
+
"sizeBytes": 2776833,
|
| 98 |
+
"sha256": "ca10d7e9fb3ed18575dd1e277a2579c16d108e32f27439684afa0e10b1440910",
|
| 99 |
+
"source": "app/src/main/java/com/example/autodroid/data/voice/model/VoiceModelCatalog.kt:242"
|
| 100 |
+
},
|
| 101 |
+
{
|
| 102 |
+
"repo": "litert-community/Qwen3-TTS-12Hz-0.6B-Base",
|
| 103 |
+
"revision": "0eb3b8a4714972b065c160faec6a12158caa9dc0",
|
| 104 |
+
"filename": "voices/demo_speaker.npy",
|
| 105 |
+
"sizeBytes": 4224,
|
| 106 |
+
"sha256": "b1527f54f68f44ca98bfddcaa9dc0018deb2db62590c9d2699efabbd0dfc1c3c",
|
| 107 |
+
"source": "app/src/main/java/com/example/autodroid/data/voice/model/VoiceModelCatalog.kt:237"
|
| 108 |
+
}
|
| 109 |
+
],
|
| 110 |
+
"preservedDocuments": [
|
| 111 |
+
"README.md"
|
| 112 |
+
]
|
| 113 |
+
}
|
README.md
ADDED
|
@@ -0,0 +1,96 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
base_model: Qwen/Qwen3-TTS-12Hz-0.6B-Base
|
| 4 |
+
pipeline_tag: text-to-speech
|
| 5 |
+
language:
|
| 6 |
+
- en
|
| 7 |
+
- zh
|
| 8 |
+
- ja
|
| 9 |
+
- ko
|
| 10 |
+
- de
|
| 11 |
+
- fr
|
| 12 |
+
- es
|
| 13 |
+
- it
|
| 14 |
+
- pt
|
| 15 |
+
- ru
|
| 16 |
+
tags:
|
| 17 |
+
- litert
|
| 18 |
+
- tflite
|
| 19 |
+
- text-to-speech
|
| 20 |
+
- voice-cloning
|
| 21 |
+
- on-device
|
| 22 |
+
---
|
| 23 |
+
|
| 24 |
+
# Qwen3-TTS-12Hz-0.6B-Base — LiteRT
|
| 25 |
+
|
| 26 |
+
[Qwen3-TTS-12Hz-0.6B-Base](https://huggingface.co/Qwen/Qwen3-TTS-12Hz-0.6B-Base) (Apache-2.0) converted to LiteRT (.tflite) for fully on-device text-to-speech with 3-second voice cloning, in 10 languages at 24 kHz.
|
| 27 |
+
|
| 28 |
+
Qwen3-TTS is a speech LM: a Qwen3-style talker predicts 12.5 Hz frames of 16 codec tokens (first codebook by the talker, 15 residual codebooks by an inner "MTP" transformer), and a neural codec decoder renders PCM. LiteRT-LM's Engine decode loop does not support this generation structure yet, so the model runs as **three LiteRT graphs driven by a host-side loop** (LiteRT Compiled Model pattern). A complete Python reference pipeline and all conversion scripts live in the litert-samples sample: [`compiled_model_api/text_to_speech_lm`](https://github.com/john-rocky/litert-samples/tree/qwen3-tts-sample/compiled_model_api/text_to_speech_lm).
|
| 29 |
+
|
| 30 |
+
## Quick start (Python, desktop)
|
| 31 |
+
|
| 32 |
+
```bash
|
| 33 |
+
git clone -b qwen3-tts-sample https://github.com/john-rocky/litert-samples.git
|
| 34 |
+
cd litert-samples/compiled_model_api/text_to_speech_lm/python
|
| 35 |
+
pip install -r requirements.txt
|
| 36 |
+
python synthesize.py --text "Hello from LiteRT running fully on device." --output hello.wav
|
| 37 |
+
```
|
| 38 |
+
|
| 39 |
+
The script downloads this repository automatically (~1.4 GB for the default int4 configuration) and speaks in the bundled demo voice. Enroll your own voice from ~3 s of audio with the sample's `conversion/extract_speaker_embedding.py`, then pass `--speaker my_voice.npy`.
|
| 40 |
+
|
| 41 |
+
## Android app
|
| 42 |
+
|
| 43 |
+
The same sample ships an Android app (Kotlin, Compiled Model API, CPU) under `compiled_model_api/text_to_speech_lm/kotlin_cpu/android/`: build with Android Studio or `./gradlew :app:installDebug`, then run `./install_to_device.sh` to download the model files from this repository and push them to the device. Device-verified on Pixel 8a. With the reference `mtp_fp32` + `codec_decoder_fp32` graphs: RTF ≈ 6.7. The app auto-selects the fast graphs when present: `mtp_folded_int8` drops the MTP from ≈333 to ≈68 ms/frame (~5×), and the split `codec_partA`/`codec_partB` drops the codec from ≈114 to ≈40 ms/frame (~2.5×). Together the end-to-end **RTF falls to ≈2.06** (~3.2× vs the reference graphs), ASR-lossless.
|
| 44 |
+
|
| 45 |
+
## Files
|
| 46 |
+
|
| 47 |
+
| File | Size | Role |
|
| 48 |
+
|---|---|---|
|
| 49 |
+
| `talker_int4.tflite` | 256 MB | Talker LM (28-layer Qwen3, prefill_32/prefill_128/decode signatures, KV 1024), blockwise-32 OCTAV int4 weights |
|
| 50 |
+
| `talker_fp32.tflite` | 1.8 GB | fp32 talker; under greedy decoding it reproduces the PyTorch reference token-for-token |
|
| 51 |
+
| `mtp_fp32.tflite` | 440 MB | MTP decode step (5-layer transformer, 17-slot KV cache, 15 lm_heads), invoked 17× per frame — the exact reference graph |
|
| 52 |
+
| `mtp_folded_int8.tflite` | 218 MB | **Fast MTP**: all 16 inner steps × 5 layers folded into one graph (in-graph argmax + embedding gather, KV internal), GPTQ dynamic-int8 weights. One invoke per frame; ~5× faster on device. Drop-in replacement for `mtp_fp32.tflite` |
|
| 53 |
+
| `codec_decoder_fp32.tflite` | 457 MB | Codec decoder (RVQ + 8-layer transformer + causal ConvNet, 64-frame chunks → 24 kHz PCM) |
|
| 54 |
+
| `codec_partA.tflite` / `codec_partB.tflite` | 163 + 294 MB | **Fast codec**: the decoder split at the transformer/convnet boundary. Part A (transformer) runs fp32; Part B (the conv upsampler, ~all the FLOPs) runs with XNNPACK FORCE_FP16 → ~2.5× on device, ASR-identical. Drop-in replacement for `codec_decoder_fp32.tflite` |
|
| 55 |
+
| `tokenizer.json` | 11 MB | Qwen2 BPE tokenizer (Python sample / `tokenizers`) |
|
| 56 |
+
| `vocab.json`, `merges.txt` | 4.5 MB | Same vocabulary in raw form (used by the Android app's Kotlin tokenizer) |
|
| 57 |
+
| `tables/*` | 723 MB | Host-side embedding tables: codec embedding (fp32), 15 MTP embeddings (fp16), text embedding (fp16), text projection MLP (fp32) |
|
| 58 |
+
| `voices/demo_speaker.npy` | 4 KB | Demo voice x-vector (enrolled from the official Qwen3-TTS demo clip) |
|
| 59 |
+
|
| 60 |
+
## Accuracy
|
| 61 |
+
|
| 62 |
+
- Each graph is numerically verified against the PyTorch reference: talker bit-exact at torch level and correlation 1.0 / top-1 100% as .tflite; MTP 15/15 greedy tokens; codec decoder correlation 1.0 (max abs diff 1.8e-5).
|
| 63 |
+
- End to end with `talker_fp32` + greedy: token-for-token identical codes to the reference implementation, waveform correlation 1.000000, ASR round-trip returns the input sentence.
|
| 64 |
+
- `talker_int4` (data-free blockwise-32 OCTAV) produces a different but valid sampling trajectory; outputs transcribe identically under ASR round-trip. Channelwise int8/int4 quantization (the tooling default) degenerates on this model family — use blockwise granularity.
|
| 65 |
+
|
| 66 |
+
## Performance (Apple M4 Max, CPU/XNNPACK)
|
| 67 |
+
|
| 68 |
+
| Stage | per 80 ms audio frame |
|
| 69 |
+
|---|---|
|
| 70 |
+
| Talker decode (8 threads) | 45–50 ms |
|
| 71 |
+
| MTP inner loop (17 invokes, 1 thread) | ~148 ms |
|
| 72 |
+
| Codec decoder (amortized) | ~10 ms |
|
| 73 |
+
| Total | ~205 ms → RTF ≈ 2.5 |
|
| 74 |
+
|
| 75 |
+
The MTP inner loop dominates (a 78M-parameter transformer streams its weights 17 times per frame). `mtp_folded_int8.tflite` folds those 17 invokes into one graph and quantizes it: on the M4 Max the MTP drops to ~41 ms/frame, and on a Pixel 8a from ~333 to ~68 ms/frame. The fold is token-identical to the reference; the dynamic-int8 weights give a different-but-intelligible trajectory (ASR round-trip exact). The codec then dominates, and `codec_partA`/`codec_partB` split it so the conv-heavy back half runs in fp16 (~2.5× on device). Together the end-to-end RTF drops from ≈6.7 to ≈2.06 on a Pixel 8a (≈1.44 on M4 Max), ASR-lossless. Conversion scripts: [`export_mtp_folded.py` / `gptq_mtp_folded.py` / `export_codec_split.py`](https://github.com/john-rocky/hf-to-litertlm/tree/main/qwen3tts_work). Remaining lever: the talker (now ~52 ms/frame).
|
| 76 |
+
|
| 77 |
+
### Android (Pixel 8a)
|
| 78 |
+
|
| 79 |
+
Android figures use the standard TFLite [`benchmark_model`](https://ai.google.dev/edge/litert/models/measurement) on a **Pixel 8a** (Tensor G3, Android 16) — 5 warm-up runs then 20 timed runs, CPU at 4 threads.
|
| 80 |
+
|
| 81 |
+
| Graph | GPU (OpenCL) | CPU (XNNPACK, 4 threads) |
|
| 82 |
+
|---|---|---|
|
| 83 |
+
| `mtp_folded_int8.tflite` | 225 ms | 153 ms |
|
| 84 |
+
| `mtp_fp32.tflite` | 113 ms | 27 ms |
|
| 85 |
+
|
| 86 |
+
Nothing here is faster on the GPU; run this pipeline on the CPU on Android.
|
| 87 |
+
|
| 88 |
+
## Limitations
|
| 89 |
+
|
| 90 |
+
- Voice cloning is x-vector mode only (speaker embedding). ICL-mode cloning (reference transcript + codec encoding of the reference audio) additionally needs the codec encoder, which is kept off-device (enrollment-time PyTorch).
|
| 91 |
+
- The prompt prefill is capped at 32 positions (the x-vector prompt is always 10); the KV cache is 1024 (~80 s of audio), generation is capped at 512 frames (~41 s) in the sample.
|
| 92 |
+
- Streaming synthesis (the model's dual-track design supports it) is not implemented in the sample loop yet.
|
| 93 |
+
|
| 94 |
+
## License
|
| 95 |
+
|
| 96 |
+
Apache-2.0, inherited from the base model by the Qwen team, Alibaba Group.
|
codec_partA.tflite
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:9d3733c2a4c0a8be734d49cde63f57b9d0bee9308e67df678664d25937499c8c
|
| 3 |
+
size 162987924
|
codec_partB.tflite
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:3486fafde3a50743ccd7330e5009b5b688eb748fb9c138a255cd17fb1846155f
|
| 3 |
+
size 293705152
|
licenses/Apache-2.0.txt
ADDED
|
@@ -0,0 +1,171 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Apache License
|
| 2 |
+
Version 2.0, January 2004
|
| 3 |
+
http://www.apache.org/licenses/
|
| 4 |
+
|
| 5 |
+
TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
|
| 6 |
+
|
| 7 |
+
1. Definitions.
|
| 8 |
+
|
| 9 |
+
"License" shall mean the terms and conditions for use, reproduction, and
|
| 10 |
+
distribution as defined by Sections 1 through 9 of this document.
|
| 11 |
+
|
| 12 |
+
"Licensor" shall mean the copyright owner or entity authorized by the
|
| 13 |
+
copyright owner that is granting the License.
|
| 14 |
+
|
| 15 |
+
"Legal Entity" shall mean the union of the acting entity and all other
|
| 16 |
+
entities that control, are controlled by, or are under common control with
|
| 17 |
+
that entity. For the purposes of this definition, "control" means (i) the
|
| 18 |
+
power, direct or indirect, to cause the direction or management of such
|
| 19 |
+
entity, whether by contract or otherwise, or (ii) ownership of fifty percent
|
| 20 |
+
(50%) or more of the outstanding shares, or (iii) beneficial ownership of
|
| 21 |
+
such entity.
|
| 22 |
+
|
| 23 |
+
"You" (or "Your") shall mean an individual or Legal Entity exercising
|
| 24 |
+
permissions granted by this License.
|
| 25 |
+
|
| 26 |
+
"Source" form shall mean the preferred form for making modifications,
|
| 27 |
+
including but not limited to software source code, documentation source, and
|
| 28 |
+
configuration files.
|
| 29 |
+
|
| 30 |
+
"Object" form shall mean any form resulting from mechanical transformation
|
| 31 |
+
or translation of a Source form, including but not limited to compiled object
|
| 32 |
+
code, generated documentation, and conversions to other media types.
|
| 33 |
+
|
| 34 |
+
"Work" shall mean the work of authorship, whether in Source or Object form,
|
| 35 |
+
made available under the License, as indicated by a copyright notice that is
|
| 36 |
+
included in or attached to the work.
|
| 37 |
+
|
| 38 |
+
"Derivative Works" shall mean any work, whether in Source or Object form,
|
| 39 |
+
that is based on (or derived from) the Work and for which the editorial
|
| 40 |
+
revisions, annotations, elaborations, or other modifications represent, as a
|
| 41 |
+
whole, an original work of authorship. Derivative Works shall not include
|
| 42 |
+
works that remain separable from, or merely link (or bind by name) to the
|
| 43 |
+
interfaces of, the Work and Derivative Works thereof.
|
| 44 |
+
|
| 45 |
+
"Contribution" shall mean any work of authorship, including the original
|
| 46 |
+
version of the Work and any modifications or additions to that Work or
|
| 47 |
+
Derivative Works thereof, that is intentionally submitted to Licensor for
|
| 48 |
+
inclusion in the Work by the copyright owner or by an individual or Legal
|
| 49 |
+
Entity authorized to submit on behalf of the copyright owner. "Submitted"
|
| 50 |
+
means any form of electronic, verbal, or written communication sent to the
|
| 51 |
+
Licensor or its representatives, excluding communication conspicuously marked
|
| 52 |
+
or otherwise designated in writing by the copyright owner as "Not a
|
| 53 |
+
Contribution."
|
| 54 |
+
|
| 55 |
+
"Contributor" shall mean Licensor and any individual or Legal Entity on
|
| 56 |
+
behalf of whom a Contribution has been received by Licensor and subsequently
|
| 57 |
+
incorporated within the Work.
|
| 58 |
+
|
| 59 |
+
2. Grant of Copyright License. Subject to the terms and conditions of this
|
| 60 |
+
License, each Contributor hereby grants to You a perpetual, worldwide,
|
| 61 |
+
non-exclusive, no-charge, royalty-free, irrevocable copyright license to
|
| 62 |
+
reproduce, prepare Derivative Works of, publicly display, publicly perform,
|
| 63 |
+
sublicense, and distribute the Work and such Derivative Works in Source or
|
| 64 |
+
Object form.
|
| 65 |
+
|
| 66 |
+
3. Grant of Patent License. Subject to the terms and conditions of this
|
| 67 |
+
License, each Contributor hereby grants to You a perpetual, worldwide,
|
| 68 |
+
non-exclusive, no-charge, royalty-free, irrevocable (except as stated in this
|
| 69 |
+
section) patent license to make, have made, use, offer to sell, sell, import,
|
| 70 |
+
and otherwise transfer the Work, where such license applies only to those
|
| 71 |
+
patent claims licensable by such Contributor that are necessarily infringed
|
| 72 |
+
by their Contribution(s) alone or by combination of their Contribution(s)
|
| 73 |
+
with the Work to which such Contribution(s) was submitted. If You institute
|
| 74 |
+
patent litigation against any entity alleging that the Work or a Contribution
|
| 75 |
+
incorporated within the Work constitutes direct or contributory patent
|
| 76 |
+
infringement, then any patent licenses granted to You under this License for
|
| 77 |
+
that Work shall terminate as of the date such litigation is filed.
|
| 78 |
+
|
| 79 |
+
4. Redistribution. You may reproduce and distribute copies of the Work or
|
| 80 |
+
Derivative Works thereof in any medium, with or without modifications, and in
|
| 81 |
+
Source or Object form, provided that You meet the following conditions:
|
| 82 |
+
|
| 83 |
+
(a) You must give any other recipients of the Work or Derivative Works a copy
|
| 84 |
+
of this License; and
|
| 85 |
+
|
| 86 |
+
(b) You must cause any modified files to carry prominent notices stating that
|
| 87 |
+
You changed the files; and
|
| 88 |
+
|
| 89 |
+
(c) You must retain, in the Source form of any Derivative Works that You
|
| 90 |
+
distribute, all copyright, patent, trademark, and attribution notices from the
|
| 91 |
+
Source form of the Work, excluding those notices that do not pertain to any
|
| 92 |
+
part of the Derivative Works; and
|
| 93 |
+
|
| 94 |
+
(d) If the Work includes a "NOTICE" text file as part of its distribution,
|
| 95 |
+
then any Derivative Works that You distribute must include a readable copy of
|
| 96 |
+
the attribution notices contained within such NOTICE file, excluding those
|
| 97 |
+
notices that do not pertain to any part of the Derivative Works, in at least
|
| 98 |
+
one of the following places: within a NOTICE text file distributed as part of
|
| 99 |
+
the Derivative Works; within the Source form or documentation, if provided;
|
| 100 |
+
or, within a display generated by the Derivative Works, if and wherever such
|
| 101 |
+
third-party notices normally appear. The contents of the NOTICE file are for
|
| 102 |
+
informational purposes only and do not modify the License.
|
| 103 |
+
|
| 104 |
+
You may add Your own copyright statement to Your modifications and may
|
| 105 |
+
provide additional or different license terms and conditions for use,
|
| 106 |
+
reproduction, or distribution of Your modifications, provided that Your use,
|
| 107 |
+
reproduction, and distribution of the Work otherwise complies with the
|
| 108 |
+
conditions stated in this License.
|
| 109 |
+
|
| 110 |
+
5. Submission of Contributions. Unless You explicitly state otherwise, any
|
| 111 |
+
Contribution intentionally submitted for inclusion in the Work by You to the
|
| 112 |
+
Licensor shall be under the terms and conditions of this License, without any
|
| 113 |
+
additional terms or conditions.
|
| 114 |
+
|
| 115 |
+
6. Trademarks. This License does not grant permission to use the trade names,
|
| 116 |
+
trademarks, service marks, or product names of the Licensor, except as
|
| 117 |
+
required for reasonable and customary use in describing the origin of the
|
| 118 |
+
Work and reproducing the content of the NOTICE file.
|
| 119 |
+
|
| 120 |
+
7. Disclaimer of Warranty. Unless required by applicable law or agreed to in
|
| 121 |
+
writing, Licensor provides the Work (and each Contributor provides its
|
| 122 |
+
Contributions) on an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
|
| 123 |
+
KIND, either express or implied, including, without limitation, any warranties
|
| 124 |
+
or conditions of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
|
| 125 |
+
PARTICULAR PURPOSE. You are solely responsible for determining the
|
| 126 |
+
appropriateness of using or redistributing the Work and assume any risks
|
| 127 |
+
associated with Your exercise of permissions under this License.
|
| 128 |
+
|
| 129 |
+
8. Limitation of Liability. In no event and under no legal theory, whether in
|
| 130 |
+
tort (including negligence), contract, or otherwise, unless required by
|
| 131 |
+
applicable law (such as deliberate and grossly negligent acts) or agreed to in
|
| 132 |
+
writing, shall any Contributor be liable to You for damages, including any
|
| 133 |
+
direct, indirect, special, incidental, or consequential damages arising as a
|
| 134 |
+
result of this License or out of the use or inability to use the Work, even if
|
| 135 |
+
such Contributor has been advised of the possibility of such damages.
|
| 136 |
+
|
| 137 |
+
9. Accepting Warranty or Additional Liability. While redistributing the Work
|
| 138 |
+
or Derivative Works thereof, You may choose to offer, and charge a fee for,
|
| 139 |
+
acceptance of support, warranty, indemnity, or other liability obligations
|
| 140 |
+
and/or rights consistent with this License. However, in accepting such
|
| 141 |
+
obligations, You may act only on Your own behalf and on Your sole
|
| 142 |
+
responsibility, not on behalf of any other Contributor, and only if You agree
|
| 143 |
+
to indemnify, defend, and hold each Contributor harmless for any liability
|
| 144 |
+
incurred by, or claims asserted against, such Contributor by reason of your
|
| 145 |
+
accepting any such warranty or additional liability.
|
| 146 |
+
|
| 147 |
+
END OF TERMS AND CONDITIONS
|
| 148 |
+
|
| 149 |
+
APPENDIX: How to apply the Apache License to your work.
|
| 150 |
+
|
| 151 |
+
To apply the Apache License to your work, attach the following boilerplate
|
| 152 |
+
notice, with the fields enclosed by brackets "[]" replaced with your own
|
| 153 |
+
identifying information. (Don't include the brackets!) The text should be
|
| 154 |
+
enclosed in the appropriate comment syntax for the file format. We also
|
| 155 |
+
recommend that a file or class name and description of purpose be included on
|
| 156 |
+
the same "printed page" as the copyright notice for easier identification
|
| 157 |
+
within third-party archives.
|
| 158 |
+
|
| 159 |
+
Copyright [yyyy] [name of copyright owner]
|
| 160 |
+
|
| 161 |
+
Licensed under the Apache License, Version 2.0 (the "License");
|
| 162 |
+
you may not use this file except in compliance with the License.
|
| 163 |
+
You may obtain a copy of the License at
|
| 164 |
+
|
| 165 |
+
http://www.apache.org/licenses/LICENSE-2.0
|
| 166 |
+
|
| 167 |
+
Unless required by applicable law or agreed to in writing, software
|
| 168 |
+
distributed under the License is distributed on an "AS IS" BASIS,
|
| 169 |
+
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
| 170 |
+
See the License for the specific language governing permissions and
|
| 171 |
+
limitations under the License.
|
merges.txt
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
mtp_folded_int8.tflite
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:f5ab8f826e3dd68f14667af422145fe57233b445046e5ef42c01b59f82191b4b
|
| 3 |
+
size 229608368
|
tables/codec_embedding_fp32.npy
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:47fa9e30f98b1528fc9b332d314f22a32fa33e187509a4d3537f8b2c31199e39
|
| 3 |
+
size 12583040
|
tables/mtp_embeddings_fp16.npy
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:fea581b6a04f1cbec20b49511c36a00011411ccfba31f89b7571f82fe6b36706
|
| 3 |
+
size 62914688
|
tables/text_embedding_fp16.npy
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:6fab9de0a8bc144aa3efefdacb5e8292b8499a9ae1d60fcee240bac528b7441e
|
| 3 |
+
size 622329984
|
tables/text_projection_fp32.npz
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:ebb0f6a7aaacdbc903e825e33480b2da4d4b71c43c90acb0e89988049c77c100
|
| 3 |
+
size 25179078
|
talker_int4.tflite
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:e03df54e73ed1f88b2ae6d47bbf82dd64ea90a3620d753a0f3c8d6a8d60848db
|
| 3 |
+
size 255998768
|
vocab.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
voices/demo_speaker.npy
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:b1527f54f68f44ca98bfddcaa9dc0018deb2db62590c9d2699efabbd0dfc1c3c
|
| 3 |
+
size 4224
|