--- license: apache-2.0 base_model: convaiinnovations/laya-multilingual language: - de - en pipeline_tag: text-classification tags: - laya - system-one - typed-decisions - rlcd - dictation - speech-to-text datasets: - MainzelMennchen/LayaVoxprompt-De-En --- # LayaVoxprompt A fine-tune of [`convaiinnovations/laya-multilingual`](https://huggingface.co/convaiinnovations/laya-multilingual) (mmBERT-base, 322M) for one decision in a push-to-talk dictation app: > Is this dictated text meant as an instruction or question for an AI assistant, > rather than text for a person or a document? The app uses the answer to choose between two post-processing modes: rewrite the dictation into a clean prompt, or only clean up the transcript. The model reads raw speech-to-text output (fillers, missing punctuation, self-corrections) in German, English and mixed German/English, together with the name of the app in focus. It does not generate text. One forward pass returns `noul`, the probability that the answer is yes. ## Usage ```python import laya agent = laya.load("MainzelMennchen/LayaVoxprompt") QUESTION = { "is_prompt": { "type": "noul", "instructions": "Is this dictated text meant as an instruction or question for an AI " "assistant, rather than text for a person or a document?", } } state = {"app": "Cursor", "language": "de", "dictation": "füg hier noch Error Handling ein falls die Datei nicht existiert"} p = agent.predict(state, QUESTION)["answers"]["is_prompt"]["noul"] mode = "prompt" if p >= 0.9 else "clean" ``` Use the question text and the state fields (`app`, `language`, `dictation`) exactly as above. They are what the model was trained on. ## Training - Base: `convaiinnovations/laya-multilingual` - Method: RLCD fine-tuning with the official Apple Silicon script from [NandhaKishorM/laya](https://github.com/NandhaKishorM/laya) (`notebooks/laya_finetune_typed_decisions_mps.py`), 4 epochs, MPS, fp32 - Data: ~1.4k synthetic dictations, see [`MainzelMennchen/LayaVoxprompt-De-En`](https://huggingface.co/datasets/MainzelMennchen/LayaVoxprompt-De-En) - Calibration: one temperature for `noul`, fitted on ~10 % held out from training. The fitted value was 8.04; Laya clamps it to 5.0 at load time. ## Evaluation 341 held-out synthetic examples (160 prompt / 181 clean). | metric | value | |---|---| | accuracy @ 0.5 | 0.950 | | ECE | 0.034 | | accuracy de / en / mixed | 0.970 / 0.937 / 0.940 | | threshold | clean texts classified as prompt | prompts recognised | |---|---|---| | 0.50 | 2.2 % | 91.9 % | | 0.90 | 2.2 % | 90.6 % | | 0.95 | 1.7 % | 85.6 % | | 0.97 | 0.0 % | 80.0 % | ## Limitations - **Synthetic data only.** Train and test sets were generated by a local LLM from a small set of hand-written seeds. Real dictations are messier; expect lower accuracy in use. - **Noisy labels.** A blind relabelling pass agreed with the original labels on 90.7 % of examples. Most disagreements are ambiguous prompts. - **The app name is under-used.** Requests for help written to people in chat apps ("can you help me fix this bug") can be classified as prompts. The intended deployment puts a fixed rule in front of the model for mail and messaging apps. - **Small test set.** Error rates near 0–2 % rest on a handful of examples. - German and English only. Other languages were not tested. ## License and attribution Apache-2.0. Based on Laya by Convai Innovations ([`convaiinnovations/laya`](https://huggingface.co/convaiinnovations/laya), Apache-2.0); encoder `jhu-clsp/mmBERT-base` as shipped inside the Laya checkpoint.