utkarshshukla2912 commited on
Commit
471adf1
·
verified ·
1 Parent(s): 39a2bed

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +1 -73
README.md CHANGED
@@ -69,13 +69,6 @@ code-mixed speech, Bengali, Telugu, Tamil, Kannada, Malayalam, Marathi and Gujar
69
 
70
  ## Why we built it
71
 
72
- Ringg's voice agents run as multi-step conversation flows. After every caller turn, the agent must decide whether to
73
- stay in the current step or move to another one ("the caller wants a refund", "the caller has no further questions",
74
- "the caller asked for a human"), and often capture a value on the way (a date, a plan name, a language). A large
75
- general LLM does this well, but it adds hundreds of milliseconds to every turn of a live phone call.
76
-
77
- Ringg Router answers the same question in one short generation. The option id comes out in the first few tokens, so a
78
- caller hears the next step sooner. It is trained to:
79
  - choose among 2–24 natural-language options, with the answer independent of the order they are listed in;
80
  - stay put when nothing calls for a move, and say "none of these" when no option fits;
81
  - read Indian languages and code-mixed, transcribed speech (ASR noise, fragments, Latin-script Hindi);
@@ -202,7 +195,7 @@ zero-shot.
202
  | Naamapadam NER (Indic) | 32.0 | 74.4 | **87.1** |
203
  | HiNER (Hindi NER) | 47.6 | 76.1 | **89.2** |
204
  | SGD slot filling | 69.6 | 72.7 | **98.7** |
205
- | Belebele (unseen, reading comprehension) | 65.3 | **84.0** | 71.3 |
206
  | Kev suites (unseen) | 64.2 | 69.1 | **71.1** |
207
 
208
  **How to read this.** On every task family it was trained for, the router beats the base model it came from, and the
@@ -223,20 +216,6 @@ GPU. Generating the rationale adds roughly 20–25 tokens.
223
  - Yes / no / unknown checks of a condition against a conversation.
224
  - Structured extraction of named fields from short conversations, including Indian languages and code-mixed text.
225
 
226
- ## Limitations
227
-
228
- - **Options must say when to take them.** The model sees only the conversation and the option descriptions. Labels
229
- like `intent = payments`, or rules that live in a hidden system prompt, are much weaker than plain descriptions
230
- ("user reports a failed or pending payment"). Rules that depend on data the model cannot see (account status, API
231
- results) must be written into the options or the state.
232
- - **Leans towards staying.** When unsure, it tends to keep the conversation in the current step rather than move.
233
- Tune per-option thresholds on the id log-probabilities if your application needs more recall on moves.
234
- - **Specialist.** Weaker than larger general models on open-ended reading comprehension and extractive QA (see
235
- Belebele / IndicQA above). Not a chat model.
236
- - **Text only.** The vision and audio towers of the base model were not trained; send transcribed text.
237
- - **Rationales are English** and short; they explain the chosen option, they are not a proof.
238
- - As with any language model, decisions can be wrong; keep a fallback for high-stakes actions (payments, account
239
- changes, cancellations).
240
 
241
  ## Training
242
 
@@ -279,57 +258,6 @@ Sources whose labels did not hold up in review were left out.
279
 
280
  Evaluation rows in this card are disjoint from training rows. Belebele and the Kev suites were never used in training.
281
 
282
- ## Considerations
283
-
284
- ### Designing options
285
-
286
- - **Write options as conditions.** "User reports a failed or pending payment" works; "intent = payments" does not.
287
- The model never sees your agent's system prompt, so any rule it must follow has to be in the option text.
288
- - **Put the call stage in the text.** For example: "user has no further questions after being helped" or "user agreed
289
- to talk to a live specialist". Wrap-up and hand-off moves are the ones the model misses when the option is vague.
290
- - **Include an explicit stay option**, and a "none of these" option for open-ended menus.
291
- - **Use unique, short ids.** Up to 24 options per decision were seen in training.
292
- - **Feed data the model cannot hear as text** in the state or option (e.g. "account status: not registered"). Do not
293
- expect the model to infer hidden facts.
294
-
295
- ### Deployment
296
-
297
- - **Precision:** use bfloat16. float16 corrupts Gemma-4 outputs.
298
- - **Attention kernels:** Gemma-4's 512-wide attention heads need a GPU with enough shared memory. Ada/Hopper (L4,
299
- L40S, H100) work with vLLM's Triton backend; T4 does not (use transformers in bf16 there, which is slow).
300
- - **Multimodal memory:** with vLLM, pass `limit_mm_per_prompt={"image": 0, "video": 0, "audio": 0}` to skip
301
- reserving memory for inputs you don't send.
302
- - **Context:** the model was trained on the most recent turns of a conversation (about five lines). Longer histories
303
- add latency more than accuracy.
304
- - **Constrained decoding:** to guarantee a valid id, constrain the branch to your option ids (e.g. vLLM structured
305
- outputs with a `choice` list), or map the decoded text back to the nearest id.
306
- - **Calibration:** the first-token log-probabilities of the option ids give a usable ranking. Calibrate a stay/move
307
- threshold on your own traffic before you act on low-margin decisions.
308
-
309
- ### Safety and responsible use
310
-
311
- - **Prompt injection:** the model is trained to treat the `state` as data, not instructions, but it is not a
312
- security boundary. Keep permission checks and irreversible actions (payments, cancellations, account changes)
313
- behind your own logic and confirmations.
314
- - **Imperfect output:** routing and extraction mistakes happen, and more often on noisy transcripts, rare languages
315
- and long option lists. Log the decisions and keep a human or LLM fallback for ambiguous cases.
316
- - **Uneven language coverage:** English, Hindi and Hinglish are strongest. Some Indic languages have less training
317
- data. Evaluate on your own languages and domains before production.
318
- - **Personal data:** extraction pulls personal details (names, numbers, addresses) out of conversations when asked.
319
- Handle outputs under your data-protection obligations. Do not use the model to profile people or to make decisions
320
- with legal or similarly significant effects without human review.
321
- - **Rationales** are short explanations of the chosen option, useful for logs and debugging. They are not guaranteed
322
- to reflect the model's internal reasoning.
323
-
324
- ### Evaluation caveats
325
-
326
- - **Small samples:** each source contributes up to 150 held-out rows (60 for validation), so single-dataset numbers
327
- carry a few points of noise. Task-family averages are steadier.
328
- - **Extraction scoring** uses exact match after normalisation, so a correct value in different wording counts as
329
- wrong. This penalises all three models equally.
330
- - **Same prompts for all three models:** the base models are scored zero-shot on this card's prompt format, and a
331
- different prompt could raise their scores.
332
-
333
  ## License
334
 
335
  Apache 2.0, as the base model ([Gemma 4 license](https://ai.google.dev/gemma/docs/gemma_4_license)). Some training
 
69
 
70
  ## Why we built it
71
 
 
 
 
 
 
 
 
72
  - choose among 2–24 natural-language options, with the answer independent of the order they are listed in;
73
  - stay put when nothing calls for a move, and say "none of these" when no option fits;
74
  - read Indian languages and code-mixed, transcribed speech (ASR noise, fragments, Latin-script Hindi);
 
195
  | Naamapadam NER (Indic) | 32.0 | 74.4 | **87.1** |
196
  | HiNER (Hindi NER) | 47.6 | 76.1 | **89.2** |
197
  | SGD slot filling | 69.6 | 72.7 | **98.7** |
198
+ | Belebele (unseen, reading comprehension) | 65.3 | 71.3 | **84.0** |
199
  | Kev suites (unseen) | 64.2 | 69.1 | **71.1** |
200
 
201
  **How to read this.** On every task family it was trained for, the router beats the base model it came from, and the
 
216
  - Yes / no / unknown checks of a condition against a conversation.
217
  - Structured extraction of named fields from short conversations, including Indian languages and code-mixed text.
218
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
219
 
220
  ## Training
221
 
 
258
 
259
  Evaluation rows in this card are disjoint from training rows. Belebele and the Kev suites were never used in training.
260
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
261
  ## License
262
 
263
  Apache 2.0, as the base model ([Gemma 4 license](https://ai.google.dev/gemma/docs/gemma_4_license)). Some training