Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string

Qwen3-ASR Indic 22-Language LoRA (v1, interim checkpoint — step 60,000 / 142,137)

A LoRA adapter for Qwen/Qwen3-ASR-1.7B-hf, fine-tuned on a ~35,200-hour multi-source corpus covering 22 Indic languages: Hindi, Tamil, Telugu, Bengali, Marathi, Gujarati, Kannada, Malayalam, Punjabi, Odia, Assamese, Urdu, Sanskrit, Nepali, Sindhi, Bodo, Dogri, Kashmiri, Konkani, Maithili, Manipuri, and Santali.

This is an interim checkpoint from an in-progress training run — not a final model. It's step 60,000 of a planned 142,137-step single-epoch run (~42% complete at this checkpoint; the full run continues past it). Expect another release once training finishes.

Benchmark results (held-out, genuinely unseen test set)

Measured on 40 samples/language, sampled only from each language's own test*.jsonl split (never seen during training or validation — see the "Evaluation methodology" section below for why this matters). Forced-language decoding used on both sides (see Inference below) so the comparison isolates what fine-tuning changed.

Language N Base WER Base CER Adapter WER Adapter CER
Hindi 40 0.215 0.110 0.203 0.088
Urdu 40 0.928 0.760 0.191 0.090
Marathi 40 0.622 0.218 0.266 0.087
Sanskrit 40 0.857 0.270 0.345 0.078
Bengali 40 0.583 0.354 0.240 0.079
Gujarati 40 1.064 0.896 0.288 0.147
Sindhi 40 0.802 0.393 0.325 0.147
Punjabi 40 0.991 0.842 0.309 0.150
Nepali 38 0.850 0.358 0.348 0.145
Konkani 40 0.923 0.408 0.395 0.138
Maithili 40 0.767 0.332 0.360 0.133
Assamese 40 1.002 0.673 0.357 0.210
Dogri 40 0.823 0.405 0.395 0.183
Odia 40 1.085 0.933 0.415 0.178
Bodo 40 1.116 0.800 0.414 0.106
Kashmiri 40 0.999 0.811 0.471 0.238
Santali 40 1.050 0.886 0.464 0.237
Telugu 40 1.181 0.862 0.476 0.229
Kannada 40 1.254 0.908 0.551 0.280
Malayalam 40 1.118 0.719 0.543 0.244
Manipuri 40 1.205 1.010 0.556 0.361
Tamil 40 0.896 0.510 0.592 0.250

Average WER across all 22 languages: 0.924 (base) → 0.387 (this adapter) — a 58.2% relative reduction. Every language improves; best performers right now (WER < 0.3) are Hindi, Urdu, Marathi, Sanskrit, and Bengali. Weakest are Tamil, Manipuri, Malayalam, and Kannada — still meaningfully improved over base, but this run hasn't reached them as many times yet (single epoch, languages are interleaved rather than trained in sequential blocks).

Evaluation methodology

WER/CER were computed with jiwer after script/punctuation normalization. The "base" numbers are the stock Qwen/Qwen3-ASR-1.7B-hf model with no adapter attached, same decoding procedure. The test set is sampled exclusively from each language's test manifest split — rows this training run has never used for a gradient update or for validation/ checkpoint-selection — and additionally filtered to drop any row whose transcript's script doesn't match its labeled language (a labeling bug was found and fixed in one contributing dataset; see the training repo's README.md for details if you have access to it).

Architecture note

This is a single shared LoRA adapter trained across all 22 languages (rank 16, alpha 32, target_modules=["q_proj","v_proj"], dropout 0.05) — not one adapter per language. One frozen base model, one adapter, all languages mixed into one training stream.

Inference

Qwen3-ASR's built-in apply_transcription_request() helper only recognizes a ~30-language whitelist that doesn't cover most of the 22 languages this adapter targets. Force the language by directly prefilling the assistant turn instead — build the chat messages manually and set continue_final_message=True:

import torch
import librosa
from peft import PeftModel
from transformers import Qwen3ASRForConditionalGeneration, Qwen3ASRProcessor

BASE_MODEL = "Qwen/Qwen3-ASR-1.7B-hf"
ADAPTER = "inspiredclone101/qwen3-asr-indic-22lang-lora-v1"
DEVICE = "cuda"  # or "cpu"

processor = Qwen3ASRProcessor.from_pretrained(BASE_MODEL)
base_model = Qwen3ASRForConditionalGeneration.from_pretrained(
    BASE_MODEL, torch_dtype=torch.bfloat16, attn_implementation="sdpa",
).to(DEVICE)
model = PeftModel.from_pretrained(base_model, ADAPTER)
model.eval()

audio, _ = librosa.load("your_audio.wav", sr=16000, mono=True)
language = "Hindi"  # must match one of the 22 languages above, by name

messages = [
    {"role": "user", "content": [{"type": "audio", "audio": audio}]},
    {"role": "assistant", "content": [{"type": "text", "text": f"language {language}<asr_text>"}]},
]
inputs = processor.apply_chat_template(
    messages, tokenize=True, continue_final_message=True,
    return_tensors="pt", return_dict=True,
)
inputs = {
    k: (v.to(torch.bfloat16) if v.dtype == torch.float32 else v).to(DEVICE)
    for k, v in inputs.items()
}

with torch.no_grad():
    output_ids = model.generate(**inputs, max_new_tokens=200)

decoded = processor.batch_decode(
    output_ids[:, inputs["input_ids"].shape[1]:], skip_special_tokens=True
)[0]

# Output looks like "language Hindi<asr_text>वह यहाँ आया था" — keep only the part after <asr_text>.
transcript = decoded.split("<asr_text>", 1)[-1].strip()
print(transcript)

Why the manual prefill matters: <asr_text> is a fixed delimiter token in the model's own chat template — everything before it in the assistant turn is the forced language prompt we supply; everything the model generates after it is the actual transcription. Calling model.generate() without this prefill lets the model guess its own language, which it frequently gets wrong for lower-resource languages in this set (defaults to Hindi for many of them) — always force it this way.

Language names to use: exactly as listed at the top of this card (Hindi, Tamil, Telugu, Bengali, Marathi, Gujarati, Kannada, Malayalam, Punjabi, Odia, Assamese, Urdu, Sanskrit, Nepali, Sindhi, Bodo, Dogri, Kashmiri, Konkani, Maithili, Manipuri, Santali) — the string is inserted verbatim into the prompt, so it must match what the adapter was trained on.

Loading on CPU or with less GPU memory

Drop torch_dtype=torch.bfloat16 for torch.float32 on CPU, or use torch_dtype=torch.float16 with device_map="auto" if you have accelerate installed and limited VRAM. The base model is 1.7B parameters (3.5 GB in bf16); the adapter itself is ~19 MB.

Training data

~35,200 hours across 6 sources: a 15-dataset core corpus (IndicVoices, Shrutilipi, Kathbath, FLEURS, Rasa, and others), Vaani, a pseudo-labeled real-movie-audio dataset, Respin, Syspin, and open-slr. Coverage per language varies — some of the 22 languages (Bodo, Dogri, Sanskrit, Manipuri, Santali, Sindhi) only appear in the core corpus, while others draw from all 6 sources.

License

Apache 2.0, matching the base model. This adapter's weights are released as-is; it is an interim, non-final checkpoint from an active training run and comes with no performance guarantees.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for inspiredclone101/qwen3-asr-indic-22lang-lora-v1

Adapter
(12)
this model