--- language: - hi - en library_name: transformers pipeline_tag: text-classification base_model: google/muril-base-cased tags: - hinglish - code-mixed - research - hate-speech-detection --- # MuRIL: CM + THAR MuRIL fine-tuned for binary classification of Hindi–English code-mixed text on **CM + THAR**. Part of a comparative study of dataset transfer, training mixtures and seed variation. [Research code](https://github.com/OmTheLast/mBERT-vs-MuRIL-cross-dataset-hinglish-hate) · [Results](https://github.com/OmTheLast/mBERT-vs-MuRIL-cross-dataset-hinglish-hate/blob/main/docs/matched_multiseed_results.md) - **Labels:** `0 = NEGATIVE`, `1 = POSITIVE`. Mixture labels combine different source definitions of hate and offense. - **Revision:** seed 42; `main` defaults to seed 42. Available: `seed-42`. - **Training sources:** [CM](https://github.com/shikharras/cm-hate-speech-detection) · [THAR](https://github.com/aakash-dl/THAR). No dataset text is included. **Known failure:** all-negative predictions on the reported primary external evaluations. Preserved for error analysis. ## Evaluation Evaluation splits also guided checkpoint selection; the recorded internal scores are not untouched final-test estimates. Cross-dataset results and label definitions should be interpreted separately. | Seed | Macro F1 | Positive recall | |---:|---:|---:| | 42 | 0.3569 | 0.0000 | ## Limitations Use for research, comparison and error analysis. Cross-dataset generalization is limited; label definitions and platforms differ, duplicates exist in CM, and annotation/source uncertainties remain. Identity words, transliteration, quoted abuse and missing conversation context can cause errors. These checkpoints have not been validated for autonomous moderation or decisions about individuals. Single-seed mixed results do not establish seed robustness.
Loading and preprocessing Install packages from `requirements.txt`. This repository is public and can be loaded without a Hugging Face login. ```python import re import torch from transformers import AutoTokenizer, AutoModelForSequenceClassification repo = "OmTheLast/muril-hinglish-mixed-cm-thar" revision = "seed-42" # use "main" for the default seed 42 tokenizer = AutoTokenizer.from_pretrained(repo, revision=revision) model = AutoModelForSequenceClassification.from_pretrained(repo, revision=revision) model.eval() def clean(text): text = re.sub(r"http\S+|www\S+|https\S+", " URL ", str(text)) text = re.sub(r"@\w+", " USER ", text) return re.sub(r"\s+", " ", text).strip() text = clean("Aaj ka din accha tha.") inputs = tokenizer(text, return_tensors="pt", padding="max_length", truncation=True, max_length=128) with torch.inference_mode(): scores = model(**inputs).logits.softmax(dim=-1)[0] print({model.config.id2label[i]: float(score) for i, score in enumerate(scores)}) ``` Preprocessing preserves case, replaces URLs with `URL`, handles with `USER`, and collapses whitespace. Truncation is 128 tokens including special tokens.
Training details and full evaluation See `training_metadata.json` for this run's settings, row counts, split policy and any reconstructed fields. Runs used two training epochs; the script restores the epoch with the highest evaluation Macro F1. The exported weights may therefore come from an earlier epoch. The mixture is built from source training partitions at seed 42, then internally split 80/20 for training and evaluation. This condition has only one seed. **Evaluation limitation:** the training script evaluates each epoch on this evaluation split and uses it to choose the checkpoint. In CM this includes the source test split. The following scores are therefore selection-set results, not an untouched final test estimate. The training seed was supplied to Trainer; exact reproduction of classifier initialization and device behavior is not guaranteed. ### Internal / matched evaluation recorded with checkpoints | Seed | Accuracy | Macro F1 | Positive F1 | Positive recall | |---:|---:|---:|---:|---:| | 42 | 0.5549 | 0.3569 | 0.0000 | 0.0000 | All scores above come directly from the saved `eval_metrics.json` files. Historical `*_hate` keys mean the dataset-specific positive class. The 79-row diagnostic probe is excluded from these claims. ### Separate external evaluations (seed 42) | Evaluation dataset | Rows | Macro F1 | Positive recall | |---|---:|---:|---:| | kaggle_hinglish_hate | 956 | 0.3788 | 0.0000 | | cm_splits_codemixed | 415 | 0.3924 | 0.0000 | | thar_religion | 2310 | 0.3454 | 0.0000 | These use the evaluation harness and its cleaning/deduplication policy; CM has 415 evaluation rows here versus 424 in matched training evaluation. They are distinct from the internal mixture scores.
Licensing and provenance Public research checkpoint release. Both base-model cards declare Apache-2.0 ([mBERT](https://huggingface.co/google-bert/bert-base-multilingual-cased), [MuRIL](https://huggingface.co/google/muril-base-cased)). A license for these fine-tuned releases has not yet been assigned: training-source terms remain under review, including the unresolved CM and THAR licenses. The base-model license is not a license grant for the datasets. Public access does not assign a new license to the fine-tuned releases or their training data. The upstream Apache-2.0 license is included in `BASE_MODEL_LICENSE.txt`; see `BASE_MODEL_NOTICE.md` for attribution. - [Research code](https://github.com/OmTheLast/mBERT-vs-MuRIL-cross-dataset-hinglish-hate), source snapshot `4650c6eb093d422785284656dd6765d36e438522`. - [Working paper draft](https://github.com/OmTheLast/mBERT-vs-MuRIL-cross-dataset-hinglish-hate/blob/4650c6eb093d422785284656dd6765d36e438522/paper/application_research_draft.md). - [Dataset registry](https://github.com/OmTheLast/mBERT-vs-MuRIL-cross-dataset-hinglish-hate/blob/4650c6eb093d422785284656dd6765d36e438522/docs/dataset_registry.md). - [Model registry and failure cases](https://github.com/OmTheLast/mBERT-vs-MuRIL-cross-dataset-hinglish-hate/blob/4650c6eb093d422785284656dd6765d36e438522/docs/model_registry.md). - `release_metadata.json` records original checkpoint identity and SHA-256. Weights are unmodified; label names were added to the exported configuration. - Seeds 42 for the two Kaggle models have reconstructed metadata, with provenance and uncertainty recorded in their metadata files. ### Tools Note AI tools were used for coding, debugging, and documentation assistance; the research direction, result interpretation, and final claims were reviewed and owned by Om Patnaik.