Laya-multilingual · Bangla e-commerce voice agent (R1)

A fine-tune of convaiinnovations/laya-multilingual (322M, mmBERT encoder) that makes three typed decisions for a Bangla-speaking e-commerce voice agent, in one forward pass and with no generated tokens:

decision type options
order_confirm choice yes · no · repeat · out_of_scope — how the customer answered "shall I confirm your order?"
intent choice 15 support intents (order status, cancel, refund, delivery, payment, complaint, human agent, …)
escalate noul P(a human agent should take over)

Both PyTorch (repo root) and MLX FP16 for Apple Silicon (mlx/) checkpoints are included. PyTorch and MLX picked the same option on 50/50 test cases.

Status: research preview (round 1). It is strong on intent, escalation and full-sentence replies, but it misreads one- and two-word replies (হ্যাঁ, জি, না) as out_of_scope. See Limitations. Don't use it as the only yes/no decision-maker.

Results

Every model gets the same questions (questions.yaml, instructions[0]). Latency is measured on an M5 MacBook Pro 16 GB with the MLX FP16 checkpoint, one question per call.

test set base (zero-shot) this model rule/embedding cascade
order_confirm, generated test (186) 43.0% 99.5% 72.6%
↳ yes↔no mix-ups 52 1 13
order_confirm, hand-written frozen set (23) 65.2% 69.6% 95.7%
↳ yes↔no mix-ups 4 0 1
intent, BanglaEComIntent test, all scripts (1,091) 29.9% 85.5% —
escalate (1,091) 13.9% 95.1% —
router suite, Bangla (16) 56.2% 81.2% —
router suite, English (16) 62.5% 81.2% —
latency P50 / P95 per question 13.9 / 23.5 ms 16.2 / 20.3 ms < 0.1 ms
peak memory (MLX) 744 MiB 747 MiB —
  • Intent by script: Bangla script 87.1%, Bangla transliterated 93.7%, Banglish 73.5%, English 85.4%, mixed 94.3% (base: 17–48%).
  • Confidence is meaningful in-distribution. Intent predictions with top probability ≥ 0.8 are 96.5% accurate and cover 72.9% of messages; for escalate, 96.9% at 90.7%. (Base: 39% and 13%.) Brier score on intent is 0.223 (base 1.103).
  • Every frozen-set error is out_of_scope, not a yes↔no swap. In a voice agent that means "escalate / ask again", which is the safe outcome.
  • The cascade column is the agent's existing regex + embedding intent cascade, which only makes the order-confirm decision. It's near-perfect on the phrases it was built from but drops to 72.6% with 13 yes↔no mix-ups on unseen phrasings. This model scores 99.5% with 1 mix-up(s).

Full per-set tables, confusion matrices and error lists: results/FINETUNE_RESULTS.md. Analysis: results/ANALYSIS.md.

Usage

The model was trained on specific question wordings and option sets, so use the schemas in questions.yaml (instructions[0], options as listed). For order_confirm, pass the conversation: the agent's question, then the customer's reply.

Apple Silicon (MLX), pip install laya-mlx:

import laya_mlx as laya

agent = laya.load("nafiullah/laya-multilingual-bn-ecom-voice", subfolder="mlx")
state = [
    {"role": "assistant", "content": "আপনার অর্ডারটি কি কনফার্ম করে দেব?"},
    {"role": "user", "content": "না না ঠিক আছে, দিয়ে দেন"},
]
questions = {
    "order_confirm": {
        "type": "choice",
        "instructions": "The agent asked the customer to confirm their order. What does the customer's last reply mean?",
        "criteria": {"yes": "customer agrees and wants the order confirmed", "no": "customer refuses or cancels the order", "repeat": "customer did not hear or understand and wants the agent to say it again", "out_of_scope": "customer asks or says something else, e.g. delivery, payment, price or product"},
    }
}
print(agent.predict(state, questions)["answers"]["order_confirm"])

PyTorch (CPU / CUDA / MPS), pip install laya:

import laya

agent = laya.load("nafiullah/laya-multilingual-bn-ecom-voice")
answers = agent.predict(
    "আমার অর্ডারটা এখনো আসেনি, কবে পাব?",
    {
        "intent": {"type": "choice", "instructions": "What is the customer's intent?",
                    "criteria": {...}},          # the 15 intents from questions.yaml
        "escalate": {"type": "noul", "instructions": "This call should be transferred to a human agent.",
                      "criteria": {"false": "the automated agent can keep handling the customer", "true": "a human agent should take over"}},
    },
)["answers"]

Both calls return choice + probabilities (choice) or noul = P(true), plus confidence.

Training

base convaiinnovations/laya-multilingual (mmBERT-base encoder + 2-layer decision head)
data nafiullah/bangla-ecom-voice-decisions, laya config: 16,138 training items (order_confirm ×2 weight), 10% calibration slice
objective Laya's RLCD recipe: soft cross-entropy on gold probabilities + noisy-logit policy gradient with a proper-scoring-rule reward
optimiser AdamW; lr 2.5e-5 (encoder) / 1e-4 (head); cosine schedule; weight decay 0.01; grad-clip 1.0
batch micro-batch 8 × grad-accum 4; 4 epochs; exploration σ 0.4 → 0.1; group size 4
precision bf16 autocast, gradient checkpointing
hardware Apple M5 MacBook Pro 16 GB (MPS), 141 min, 1.04 s/step
calibration per-type temperature fitted on the held-out slice: [1.063, 1.0, 1.106] (choice, score, noul)
calibration-slice accuracy intent 97.2%, escalate 96.0%, order_confirm 98.9%
MLX export laya-mlx convert --dtype float16

Training inputs are built with Laya's own build_sequence (conversation states truncated from the left, as at inference). Train cases use varied instruction wordings and shuffled option order; evaluation uses a single fixed wording.

Limitations

  • Short replies. Only 5 of 917 generated order-confirm training replies were ≤ 3 words, so the model maps one- and two-word replies (হ্যাঁ, জি, না, চাই না) to out_of_scope with high confidence. Pair it with a keyword rule for single-word answers, or wait for round 2, which adds short replies to the training data.
  • Unreviewed labels. The yes/no order-confirm training replies have not yet been individually reviewed by a person.
  • Mostly synthetic data. Real phone transcripts (disfluencies, dialects, STT errors) are under-represented. Validate on your own call data before relying on it.
  • Out-of-distribution confidence. On the hand-written frozen set, the model was confidently wrong on short replies. Don't use a confidence threshold as a safety guarantee outside the training distribution.
  • escalate reflects a policy (explicit requests, strong anger, fraud or legal issues), not ground truth.

Credits

This model builds directly on the following work. Thank you to their authors.

work licence used for
Laya (multilingual) Apache-2.0 base checkpoint
Laya source & RLCD training recipe Apache-2.0 model code, sequence format, fine-tuning recipe
mmBERT-base MIT multilingual encoder inside Laya-multilingual
laya-mlx Apache-2.0 MLX conversion + Apple Silicon runtime
BanglaEComIntent CC BY-NC-SA 4.0 source intent corpus

Changes from the base model: all weights (encoder and decision head) were fine-tuned on the data above, and the calibration temperatures were re-fitted. The architecture, tokenizer and input format are unchanged.

Licence

The fine-tuned weights are released under CC BY-NC-SA 4.0, because they were trained on data under that licence (BanglaEComIntent): non-commercial use only, with attribution, and derivatives shared under the same licence. The base checkpoint and code remain under their own licences (Apache-2.0 for Laya and laya-mlx, MIT for mmBERT); their notices are kept in NOTICE.

Downloads last month
16
Safetensors
Model size
0.3B params
Tensor type
F16
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nafiullah/laya-multilingual-bn-ecom-voice

Finetuned
(70)
this model

Datasets used to train nafiullah/laya-multilingual-bn-ecom-voice

Evaluation results

  • accuracy on Bangla E-commerce Voice Agent Decisions
    self-reported
    0.995
  • accuracy on Bangla E-commerce Voice Agent Decisions
    self-reported
    0.696
  • accuracy on Bangla E-commerce Voice Agent Decisions
    self-reported
    0.855
  • accuracy on Bangla E-commerce Voice Agent Decisions
    self-reported
    0.951