Instructions to use nafiullah/laya-multilingual-bn-ecom-voice with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Laya
How to use nafiullah/laya-multilingual-bn-ecom-voice with Laya:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- MLX
How to use nafiullah/laya-multilingual-bn-ecom-voice with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download nafiullah/laya-multilingual-bn-ecom-voice --local-dir laya-multilingual-bn-ecom-voice
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Laya-multilingual · Bangla e-commerce voice agent (R1)
A fine-tune of convaiinnovations/laya-multilingual (322M, mmBERT encoder) that makes three typed decisions for a Bangla-speaking e-commerce voice agent, in one forward pass and with no generated tokens:
| decision | type | options |
|---|---|---|
order_confirm |
choice | yes · no · repeat · out_of_scope — how the customer answered "shall I confirm your order?" |
intent |
choice | 15 support intents (order status, cancel, refund, delivery, payment, complaint, human agent, …) |
escalate |
noul | P(a human agent should take over) |
Both PyTorch (repo root) and MLX FP16 for Apple Silicon (mlx/) checkpoints are included.
PyTorch and MLX picked the same option on 50/50 test cases.
Status: research preview (round 1). It is strong on intent, escalation and full-sentence replies, but it misreads one- and two-word replies (হ্যাঁ, জি, না) as
out_of_scope. See Limitations. Don't use it as the only yes/no decision-maker.
Results
Every model gets the same questions (questions.yaml, instructions[0]).
Latency is measured on an M5 MacBook Pro 16 GB with the MLX FP16 checkpoint, one question per call.
| test set | base (zero-shot) | this model | rule/embedding cascade |
|---|---|---|---|
| order_confirm, generated test (186) | 43.0% | 99.5% | 72.6% |
| ↳ yes↔no mix-ups | 52 | 1 | 13 |
| order_confirm, hand-written frozen set (23) | 65.2% | 69.6% | 95.7% |
| ↳ yes↔no mix-ups | 4 | 0 | 1 |
| intent, BanglaEComIntent test, all scripts (1,091) | 29.9% | 85.5% | — |
| escalate (1,091) | 13.9% | 95.1% | — |
| router suite, Bangla (16) | 56.2% | 81.2% | — |
| router suite, English (16) | 62.5% | 81.2% | — |
| latency P50 / P95 per question | 13.9 / 23.5 ms | 16.2 / 20.3 ms | < 0.1 ms |
| peak memory (MLX) | 744 MiB | 747 MiB | — |
- Intent by script: Bangla script 87.1%, Bangla transliterated 93.7%, Banglish 73.5%, English 85.4%, mixed 94.3% (base: 17–48%).
- Confidence is meaningful in-distribution. Intent predictions with top probability ≥ 0.8 are 96.5% accurate and cover 72.9% of messages; for escalate, 96.9% at 90.7%. (Base: 39% and 13%.) Brier score on intent is 0.223 (base 1.103).
- Every frozen-set error is
out_of_scope, not a yes↔no swap. In a voice agent that means "escalate / ask again", which is the safe outcome. - The cascade column is the agent's existing regex + embedding intent cascade, which only makes the order-confirm decision. It's near-perfect on the phrases it was built from but drops to 72.6% with 13 yes↔no mix-ups on unseen phrasings. This model scores 99.5% with 1 mix-up(s).
Full per-set tables, confusion matrices and error lists: results/FINETUNE_RESULTS.md.
Analysis: results/ANALYSIS.md.
Usage
The model was trained on specific question wordings and option sets, so use the schemas in
questions.yaml (instructions[0], options as listed). For order_confirm, pass the
conversation: the agent's question, then the customer's reply.
Apple Silicon (MLX), pip install laya-mlx:
import laya_mlx as laya
agent = laya.load("nafiullah/laya-multilingual-bn-ecom-voice", subfolder="mlx")
state = [
{"role": "assistant", "content": "আপনার অর্ডারটি কি কনফার্ম করে দেব?"},
{"role": "user", "content": "না না ঠিক আছে, দিয়ে দেন"},
]
questions = {
"order_confirm": {
"type": "choice",
"instructions": "The agent asked the customer to confirm their order. What does the customer's last reply mean?",
"criteria": {"yes": "customer agrees and wants the order confirmed", "no": "customer refuses or cancels the order", "repeat": "customer did not hear or understand and wants the agent to say it again", "out_of_scope": "customer asks or says something else, e.g. delivery, payment, price or product"},
}
}
print(agent.predict(state, questions)["answers"]["order_confirm"])
PyTorch (CPU / CUDA / MPS), pip install laya:
import laya
agent = laya.load("nafiullah/laya-multilingual-bn-ecom-voice")
answers = agent.predict(
"আমার অর্ডারটা এখনো আসেনি, কবে পাব?",
{
"intent": {"type": "choice", "instructions": "What is the customer's intent?",
"criteria": {...}}, # the 15 intents from questions.yaml
"escalate": {"type": "noul", "instructions": "This call should be transferred to a human agent.",
"criteria": {"false": "the automated agent can keep handling the customer", "true": "a human agent should take over"}},
},
)["answers"]
Both calls return choice + probabilities (choice) or noul = P(true), plus confidence.
Training
| base | convaiinnovations/laya-multilingual (mmBERT-base encoder + 2-layer decision head) |
| data | nafiullah/bangla-ecom-voice-decisions, laya config: 16,138 training items (order_confirm ×2 weight), 10% calibration slice |
| objective | Laya's RLCD recipe: soft cross-entropy on gold probabilities + noisy-logit policy gradient with a proper-scoring-rule reward |
| optimiser | AdamW; lr 2.5e-5 (encoder) / 1e-4 (head); cosine schedule; weight decay 0.01; grad-clip 1.0 |
| batch | micro-batch 8 × grad-accum 4; 4 epochs; exploration σ 0.4 → 0.1; group size 4 |
| precision | bf16 autocast, gradient checkpointing |
| hardware | Apple M5 MacBook Pro 16 GB (MPS), 141 min, 1.04 s/step |
| calibration | per-type temperature fitted on the held-out slice: [1.063, 1.0, 1.106] (choice, score, noul) |
| calibration-slice accuracy | intent 97.2%, escalate 96.0%, order_confirm 98.9% |
| MLX export | laya-mlx convert --dtype float16 |
Training inputs are built with Laya's own build_sequence (conversation states truncated from the left, as at
inference). Train cases use varied instruction wordings and shuffled option order; evaluation uses a single
fixed wording.
Limitations
- Short replies. Only 5 of 917 generated order-confirm training replies were ≤ 3 words, so the model maps
one- and two-word replies (হ্যাঁ, জি, না, চাই না) to
out_of_scopewith high confidence. Pair it with a keyword rule for single-word answers, or wait for round 2, which adds short replies to the training data. - Unreviewed labels. The yes/no order-confirm training replies have not yet been individually reviewed by a person.
- Mostly synthetic data. Real phone transcripts (disfluencies, dialects, STT errors) are under-represented. Validate on your own call data before relying on it.
- Out-of-distribution confidence. On the hand-written frozen set, the model was confidently wrong on short replies. Don't use a confidence threshold as a safety guarantee outside the training distribution.
escalatereflects a policy (explicit requests, strong anger, fraud or legal issues), not ground truth.
Credits
This model builds directly on the following work. Thank you to their authors.
| work | licence | used for |
|---|---|---|
| Laya (multilingual) | Apache-2.0 | base checkpoint |
| Laya source & RLCD training recipe | Apache-2.0 | model code, sequence format, fine-tuning recipe |
| mmBERT-base | MIT | multilingual encoder inside Laya-multilingual |
| laya-mlx | Apache-2.0 | MLX conversion + Apple Silicon runtime |
| BanglaEComIntent | CC BY-NC-SA 4.0 | source intent corpus |
Changes from the base model: all weights (encoder and decision head) were fine-tuned on the data above, and the calibration temperatures were re-fitted. The architecture, tokenizer and input format are unchanged.
Licence
The fine-tuned weights are released under CC BY-NC-SA 4.0, because they were trained on data under that
licence (BanglaEComIntent): non-commercial use
only, with attribution, and derivatives shared under the same licence. The base checkpoint and code remain under
their own licences (Apache-2.0 for Laya and laya-mlx, MIT for mmBERT); their notices are kept in NOTICE.
- Downloads last month
- 16
Quantized
Model tree for nafiullah/laya-multilingual-bn-ecom-voice
Base model
convaiinnovations/laya-multilingualDatasets used to train nafiullah/laya-multilingual-bn-ecom-voice
Badhon/BanglaEComIntent
Evaluation results
- accuracy on Bangla E-commerce Voice Agent Decisionsself-reported0.995
- accuracy on Bangla E-commerce Voice Agent Decisionsself-reported0.696
- accuracy on Bangla E-commerce Voice Agent Decisionsself-reported0.855
- accuracy on Bangla E-commerce Voice Agent Decisionsself-reported0.951