IndicBERT Multilingual Scam & Fraud Classifier

A sequence classification model fine-tuned on top of ai4bharat/IndicBERTv2-MLM-only to detect fraudulent, phishing, and scam messages across 14 Indian languages and language varieties.

The model is designed with a focus on lightweight multilingual scam detection for Indian communications and common fraud patterns.


Supported Languages

Language Code
Assamese as
Bengali bn
English en
Gujarati gu
Hindi hi
Hinglish (Hindi in Latin script) hi-Latn
Kannada kn
Kashmiri ks
Malayalam ml
Marathi mr
Odia or
Punjabi pa
Tamil ta
Telugu te

Intended Use & Capabilities

This model is intended to detect common fraud and scam patterns prevalent across Indian communications, including:

  • Electricity and utility disconnection threats.
  • Impersonation of major institutions such as banks, India Post, and courier services.
  • Fake lottery, subsidy, and government-scheme claims.
  • Suspicious payment requests and fee demands.
  • Phishing and malicious links.
  • Requests for OTPs, passwords, PINs, or banking information.
  • Fake delivery, refund, account-verification, and KYC messages.
  • Suspicious promotional and reward messages.

The model is also trained to distinguish potentially legitimate transactional messages, such as:

  • OTP notifications.
  • Bank debit/transaction alerts.
  • Utility bill reminders.
  • Official-style service notifications.

Note: Classification depends on the text provided to the model. A legitimate message can resemble a scam, and a sophisticated scam can resemble a legitimate notification. The model should therefore be treated as a classification aid rather than a definitive fraud-verification system.


Evaluation

The model was evaluated using a manually curated multilingual evaluation set containing 100 samples per language across all 14 supported languages.

Each language contains:

  • 50 SCAM samples
  • 50 HAM (legitimate) samples

This gives a total of:

14 × 100 = 1,400 evaluation samples

The evaluation set covers multiple scam and legitimate-message patterns across different languages and scripts.

Overall Results

Metric Result
Evaluation Samples 1,400
Correct Predictions 1,376
Incorrect Predictions 24
Accuracy 98.29%

Overall Accuracy

98.29% (1,376 / 1,400)

The model correctly classified 1,376 of the 1,400 explicitly structured evaluation samples.


Per-Language Evaluation

Language Samples Correct Accuracy
Assamese (as) 100 97 97.00%
Bengali (bn) 100 100 100.00%
English (en) 100 99 99.00%
Gujarati (gu) 100 100 100.00%
Hindi (hi) 100 100 100.00%
Hinglish (hi-Latn) 100 93 93.00%
Kannada (kn) 100 96 96.00%
Kashmiri (ks) 100 99 99.00%
Malayalam (ml) 100 100 100.00%
Marathi (mr) 100 99 99.00%
Odia (or) 100 99 99.00%
Punjabi (pa) 100 99 99.00%
Tamil (ta) 100 97 97.00%
Telugu (te) 100 98 98.00%
Total 1,400 1,376 98.29%

Observations

The evaluation shows strong performance across all tested languages.

The highest-performing language subsets achieved 100% accuracy on the evaluation set:

  • Bengali
  • Gujarati
  • Hindi
  • Malayalam

The lowest-performing subset in this evaluation was Hinglish at 93.00%, followed by Kannada at 96.00%.

The errors are not uniformly distributed across languages, indicating that further improvements may be possible through additional examples covering language-specific scam patterns, vocabulary, and transliteration styles.


Evaluation Limitations

The reported 98.29% accuracy represents performance on this specific manually curated evaluation set and should not be interpreted as guaranteed real-world accuracy.

Performance may vary depending on:

  • Message wording and structure.
  • Regional vocabulary and expressions.
  • Transliteration and spelling variations.
  • Previously unseen scam patterns.
  • Code-mixed language.
  • Adversarial or intentionally obfuscated text.
  • Distribution differences between the evaluation data and real-world messages.

The per-language evaluation contains 100 samples per language, which provides a useful benchmark but is not large enough to represent the full diversity of real-world communication.

For production applications, the model should ideally be evaluated continuously using larger, independently collected datasets.


Quick Usage

Using Hugging Face Pipelines

from transformers import pipeline

classifier = pipeline(
    "text-classification",
    model="anmolshrivastav/indicbert-scam-classifier",
    device_map="auto"
)

sample = "Aapka Bijli bill baki hai, connection aaj raat cut ho jayega. Turant call karein 9876543210"

result = classifier(sample)

print(result)

Programmatic Usage

import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification

class IndicScamClassifier:
    def __init__(self, model_path: str = repo_id, device: str = None):
        if device is None:
            self.device = "cuda" if torch.cuda.is_available() else "cpu"
        else:
            self.device = device
            
        self.tokenizer = AutoTokenizer.from_pretrained(model_path)
        self.model = AutoModelForSequenceClassification.from_pretrained(model_path).to(self.device)
        self.model.eval()

    def predict(self, texts, threshold: float = 0.5):
        is_single = isinstance(texts, str)
        if is_single:
            texts = [texts]

        inputs = self.tokenizer(
            texts,
            padding=True,
            truncation=True,
            max_length=128,
            return_tensors="pt"
        ).to(self.device)

        with torch.no_grad():
            outputs = self.model(**inputs)
            probs = torch.softmax(outputs.logits, dim=-1)

        results = []
        for prob in probs:
            scam_score = prob[1].item()
            label = "scam" if scam_score >= threshold else "ham"
            results.append({
                "label": label,
                "confidence": scam_score if label == "scam" else prob[0].item(),
                "scam_probability": scam_score
            })

        return results[0] if is_single else results

# Quick test
detector = IndicScamClassifier()
sample = "Congratulations, you won lottery. Call 9876543210 immediately."
print(detector.predict(sample))
Downloads last month
66
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for anmolshrivastav/indicbert-scam-classifier

Quantized
(2)
this model

Dataset used to train anmolshrivastav/indicbert-scam-classifier