File size: 1,377 Bytes
f2cc1ee
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
---

license: cc-by-nc-sa-4.0
language:
  - ne
base_model:
- SWivid/F5-TTS
---

## Overview
The F5-TTS model is finetuned on romanized [[High quality TTS data for Nepali](https://www.openslr.org/43/)] dataset for Nepali text to speech.

## License
This model is released under the Creative Commons Attribution Non Commercial Share Alike 4.0 license, which allows for free usage, modification, and distribution.

## Model Information
**Base Model**: SWivid/F5-TTS/tree/main/F5TTS_v1_Base  
**Training Duration:** 110k steps  
**Dataset:** [High quality TTS data for Nepali](https://www.openslr.org/43/)

## Training Configuration:
```json

{

    "exp_name": "F5TTS_v1_Base",

    "learning_rate": 1e-05,

    "batch_size_per_gpu": 1533,

    "batch_size_type": "frame",

    "max_samples": 64,

    "grad_accumulation_steps": 1,

    "max_grad_norm": 1,

    "epochs": 1949,

    "num_warmup_updates": 103,

    "save_per_updates": 10000,

    "keep_last_n_checkpoints": 20,

    "last_per_updates": 5000,

    "finetune": true,

    "file_checkpoint_train": "",

    "tokenizer_type": "pinyin",

    "tokenizer_file": "",

    "mixed_precision": "bf16",

    "logger": "none",

    "bnb_optimizer": true

}

```

## Usage Instructions
go to [base repo](https://github.com/SWivid/F5-TTS)

## Other links
[scripts](https://github.com/sm079/f5-tts-scripts)