Instructions to use inspiredclone101/qwen3-asr-indic-22lang-lora-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use inspiredclone101/qwen3-asr-indic-22lang-lora-v1 with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string
Qwen3-ASR Indic 22-Language LoRA (v1, interim checkpoint — step 60,000 / 142,137)
A LoRA adapter for Qwen/Qwen3-ASR-1.7B-hf, fine-tuned
on a ~35,200-hour multi-source corpus covering 22 Indic languages: Hindi, Tamil, Telugu, Bengali,
Marathi, Gujarati, Kannada, Malayalam, Punjabi, Odia, Assamese, Urdu, Sanskrit, Nepali, Sindhi, Bodo,
Dogri, Kashmiri, Konkani, Maithili, Manipuri, and Santali.
This is an interim checkpoint from an in-progress training run — not a final model. It's step 60,000 of a planned 142,137-step single-epoch run (~42% complete at this checkpoint; the full run continues past it). Expect another release once training finishes.
Benchmark results (held-out, genuinely unseen test set)
Measured on 40 samples/language, sampled only from each language's own test*.jsonl split (never seen
during training or validation — see the "Evaluation methodology" section below for why this matters).
Forced-language decoding used on both sides (see Inference below) so the comparison isolates what
fine-tuning changed.
| Language | N | Base WER | Base CER | Adapter WER | Adapter CER |
|---|---|---|---|---|---|
| Hindi | 40 | 0.215 | 0.110 | 0.203 | 0.088 |
| Urdu | 40 | 0.928 | 0.760 | 0.191 | 0.090 |
| Marathi | 40 | 0.622 | 0.218 | 0.266 | 0.087 |
| Sanskrit | 40 | 0.857 | 0.270 | 0.345 | 0.078 |
| Bengali | 40 | 0.583 | 0.354 | 0.240 | 0.079 |
| Gujarati | 40 | 1.064 | 0.896 | 0.288 | 0.147 |
| Sindhi | 40 | 0.802 | 0.393 | 0.325 | 0.147 |
| Punjabi | 40 | 0.991 | 0.842 | 0.309 | 0.150 |
| Nepali | 38 | 0.850 | 0.358 | 0.348 | 0.145 |
| Konkani | 40 | 0.923 | 0.408 | 0.395 | 0.138 |
| Maithili | 40 | 0.767 | 0.332 | 0.360 | 0.133 |
| Assamese | 40 | 1.002 | 0.673 | 0.357 | 0.210 |
| Dogri | 40 | 0.823 | 0.405 | 0.395 | 0.183 |
| Odia | 40 | 1.085 | 0.933 | 0.415 | 0.178 |
| Bodo | 40 | 1.116 | 0.800 | 0.414 | 0.106 |
| Kashmiri | 40 | 0.999 | 0.811 | 0.471 | 0.238 |
| Santali | 40 | 1.050 | 0.886 | 0.464 | 0.237 |
| Telugu | 40 | 1.181 | 0.862 | 0.476 | 0.229 |
| Kannada | 40 | 1.254 | 0.908 | 0.551 | 0.280 |
| Malayalam | 40 | 1.118 | 0.719 | 0.543 | 0.244 |
| Manipuri | 40 | 1.205 | 1.010 | 0.556 | 0.361 |
| Tamil | 40 | 0.896 | 0.510 | 0.592 | 0.250 |
Average WER across all 22 languages: 0.924 (base) → 0.387 (this adapter) — a 58.2% relative reduction. Every language improves; best performers right now (WER < 0.3) are Hindi, Urdu, Marathi, Sanskrit, and Bengali. Weakest are Tamil, Manipuri, Malayalam, and Kannada — still meaningfully improved over base, but this run hasn't reached them as many times yet (single epoch, languages are interleaved rather than trained in sequential blocks).
Evaluation methodology
WER/CER were computed with jiwer after script/punctuation
normalization. The "base" numbers are the stock Qwen/Qwen3-ASR-1.7B-hf model with no adapter
attached, same decoding procedure. The test set is sampled exclusively from each language's test
manifest split — rows this training run has never used for a gradient update or for validation/
checkpoint-selection — and additionally filtered to drop any row whose transcript's script doesn't
match its labeled language (a labeling bug was found and fixed in one contributing dataset; see the
training repo's README.md for details if you have access to it).
Architecture note
This is a single shared LoRA adapter trained across all 22 languages (rank 16, alpha 32,
target_modules=["q_proj","v_proj"], dropout 0.05) — not one adapter per language. One frozen base
model, one adapter, all languages mixed into one training stream.
Inference
Qwen3-ASR's built-in apply_transcription_request() helper only recognizes a ~30-language whitelist
that doesn't cover most of the 22 languages this adapter targets. Force the language by directly
prefilling the assistant turn instead — build the chat messages manually and set
continue_final_message=True:
import torch
import librosa
from peft import PeftModel
from transformers import Qwen3ASRForConditionalGeneration, Qwen3ASRProcessor
BASE_MODEL = "Qwen/Qwen3-ASR-1.7B-hf"
ADAPTER = "inspiredclone101/qwen3-asr-indic-22lang-lora-v1"
DEVICE = "cuda" # or "cpu"
processor = Qwen3ASRProcessor.from_pretrained(BASE_MODEL)
base_model = Qwen3ASRForConditionalGeneration.from_pretrained(
BASE_MODEL, torch_dtype=torch.bfloat16, attn_implementation="sdpa",
).to(DEVICE)
model = PeftModel.from_pretrained(base_model, ADAPTER)
model.eval()
audio, _ = librosa.load("your_audio.wav", sr=16000, mono=True)
language = "Hindi" # must match one of the 22 languages above, by name
messages = [
{"role": "user", "content": [{"type": "audio", "audio": audio}]},
{"role": "assistant", "content": [{"type": "text", "text": f"language {language}<asr_text>"}]},
]
inputs = processor.apply_chat_template(
messages, tokenize=True, continue_final_message=True,
return_tensors="pt", return_dict=True,
)
inputs = {
k: (v.to(torch.bfloat16) if v.dtype == torch.float32 else v).to(DEVICE)
for k, v in inputs.items()
}
with torch.no_grad():
output_ids = model.generate(**inputs, max_new_tokens=200)
decoded = processor.batch_decode(
output_ids[:, inputs["input_ids"].shape[1]:], skip_special_tokens=True
)[0]
# Output looks like "language Hindi<asr_text>वह यहाँ आया था" — keep only the part after <asr_text>.
transcript = decoded.split("<asr_text>", 1)[-1].strip()
print(transcript)
Why the manual prefill matters: <asr_text> is a fixed delimiter token in the model's own chat
template — everything before it in the assistant turn is the forced language prompt we supply;
everything the model generates after it is the actual transcription. Calling model.generate()
without this prefill lets the model guess its own language, which it frequently gets wrong for
lower-resource languages in this set (defaults to Hindi for many of them) — always force it this way.
Language names to use: exactly as listed at the top of this card (Hindi, Tamil, Telugu,
Bengali, Marathi, Gujarati, Kannada, Malayalam, Punjabi, Odia, Assamese, Urdu,
Sanskrit, Nepali, Sindhi, Bodo, Dogri, Kashmiri, Konkani, Maithili, Manipuri,
Santali) — the string is inserted verbatim into the prompt, so it must match what the adapter was
trained on.
Loading on CPU or with less GPU memory
Drop torch_dtype=torch.bfloat16 for torch.float32 on CPU, or use torch_dtype=torch.float16 with
device_map="auto" if you have accelerate installed and limited VRAM. The base model is 1.7B
parameters (3.5 GB in bf16); the adapter itself is ~19 MB.
Training data
~35,200 hours across 6 sources: a 15-dataset core corpus (IndicVoices, Shrutilipi, Kathbath, FLEURS, Rasa, and others), Vaani, a pseudo-labeled real-movie-audio dataset, Respin, Syspin, and open-slr. Coverage per language varies — some of the 22 languages (Bodo, Dogri, Sanskrit, Manipuri, Santali, Sindhi) only appear in the core corpus, while others draw from all 6 sources.
License
Apache 2.0, matching the base model. This adapter's weights are released as-is; it is an interim, non-final checkpoint from an active training run and comes with no performance guarantees.
- Downloads last month
- -
Model tree for inspiredclone101/qwen3-asr-indic-22lang-lora-v1
Base model
Qwen/Qwen3-ASR-1.7B-hf