Instructions to use anmolshrivastav/indicbert-scam-classifier with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use anmolshrivastav/indicbert-scam-classifier with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="anmolshrivastav/indicbert-scam-classifier")# Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("anmolshrivastav/indicbert-scam-classifier") model = AutoModelForSequenceClassification.from_pretrained("anmolshrivastav/indicbert-scam-classifier", device_map="auto") - Notebooks
- Google Colab
- Kaggle
IndicBERT Multilingual Scam & Fraud Classifier
A sequence classification model fine-tuned on top of ai4bharat/IndicBERTv2-MLM-only to detect fraudulent, phishing, and scam messages across 14 Indian languages and language varieties.
The model is designed with a focus on lightweight multilingual scam detection for Indian communications and common fraud patterns.
Supported Languages
| Language | Code |
|---|---|
| Assamese | as |
| Bengali | bn |
| English | en |
| Gujarati | gu |
| Hindi | hi |
| Hinglish (Hindi in Latin script) | hi-Latn |
| Kannada | kn |
| Kashmiri | ks |
| Malayalam | ml |
| Marathi | mr |
| Odia | or |
| Punjabi | pa |
| Tamil | ta |
| Telugu | te |
Intended Use & Capabilities
This model is intended to detect common fraud and scam patterns prevalent across Indian communications, including:
- Electricity and utility disconnection threats.
- Impersonation of major institutions such as banks, India Post, and courier services.
- Fake lottery, subsidy, and government-scheme claims.
- Suspicious payment requests and fee demands.
- Phishing and malicious links.
- Requests for OTPs, passwords, PINs, or banking information.
- Fake delivery, refund, account-verification, and KYC messages.
- Suspicious promotional and reward messages.
The model is also trained to distinguish potentially legitimate transactional messages, such as:
- OTP notifications.
- Bank debit/transaction alerts.
- Utility bill reminders.
- Official-style service notifications.
Note: Classification depends on the text provided to the model. A legitimate message can resemble a scam, and a sophisticated scam can resemble a legitimate notification. The model should therefore be treated as a classification aid rather than a definitive fraud-verification system.
Evaluation
The model was evaluated using a manually curated multilingual evaluation set containing 100 samples per language across all 14 supported languages.
Each language contains:
- 50 SCAM samples
- 50 HAM (legitimate) samples
This gives a total of:
14 × 100 = 1,400 evaluation samples
The evaluation set covers multiple scam and legitimate-message patterns across different languages and scripts.
Overall Results
| Metric | Result |
|---|---|
| Evaluation Samples | 1,400 |
| Correct Predictions | 1,376 |
| Incorrect Predictions | 24 |
| Accuracy | 98.29% |
Overall Accuracy
98.29% (1,376 / 1,400)
The model correctly classified 1,376 of the 1,400 explicitly structured evaluation samples.
Per-Language Evaluation
| Language | Samples | Correct | Accuracy |
|---|---|---|---|
Assamese (as) |
100 | 97 | 97.00% |
Bengali (bn) |
100 | 100 | 100.00% |
English (en) |
100 | 99 | 99.00% |
Gujarati (gu) |
100 | 100 | 100.00% |
Hindi (hi) |
100 | 100 | 100.00% |
Hinglish (hi-Latn) |
100 | 93 | 93.00% |
Kannada (kn) |
100 | 96 | 96.00% |
Kashmiri (ks) |
100 | 99 | 99.00% |
Malayalam (ml) |
100 | 100 | 100.00% |
Marathi (mr) |
100 | 99 | 99.00% |
Odia (or) |
100 | 99 | 99.00% |
Punjabi (pa) |
100 | 99 | 99.00% |
Tamil (ta) |
100 | 97 | 97.00% |
Telugu (te) |
100 | 98 | 98.00% |
| Total | 1,400 | 1,376 | 98.29% |
Observations
The evaluation shows strong performance across all tested languages.
The highest-performing language subsets achieved 100% accuracy on the evaluation set:
- Bengali
- Gujarati
- Hindi
- Malayalam
The lowest-performing subset in this evaluation was Hinglish at 93.00%, followed by Kannada at 96.00%.
The errors are not uniformly distributed across languages, indicating that further improvements may be possible through additional examples covering language-specific scam patterns, vocabulary, and transliteration styles.
Evaluation Limitations
The reported 98.29% accuracy represents performance on this specific manually curated evaluation set and should not be interpreted as guaranteed real-world accuracy.
Performance may vary depending on:
- Message wording and structure.
- Regional vocabulary and expressions.
- Transliteration and spelling variations.
- Previously unseen scam patterns.
- Code-mixed language.
- Adversarial or intentionally obfuscated text.
- Distribution differences between the evaluation data and real-world messages.
The per-language evaluation contains 100 samples per language, which provides a useful benchmark but is not large enough to represent the full diversity of real-world communication.
For production applications, the model should ideally be evaluated continuously using larger, independently collected datasets.
Quick Usage
Using Hugging Face Pipelines
from transformers import pipeline
classifier = pipeline(
"text-classification",
model="anmolshrivastav/indicbert-scam-classifier",
device_map="auto"
)
sample = "Aapka Bijli bill baki hai, connection aaj raat cut ho jayega. Turant call karein 9876543210"
result = classifier(sample)
print(result)
Programmatic Usage
import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
class IndicScamClassifier:
def __init__(self, model_path: str = repo_id, device: str = None):
if device is None:
self.device = "cuda" if torch.cuda.is_available() else "cpu"
else:
self.device = device
self.tokenizer = AutoTokenizer.from_pretrained(model_path)
self.model = AutoModelForSequenceClassification.from_pretrained(model_path).to(self.device)
self.model.eval()
def predict(self, texts, threshold: float = 0.5):
is_single = isinstance(texts, str)
if is_single:
texts = [texts]
inputs = self.tokenizer(
texts,
padding=True,
truncation=True,
max_length=128,
return_tensors="pt"
).to(self.device)
with torch.no_grad():
outputs = self.model(**inputs)
probs = torch.softmax(outputs.logits, dim=-1)
results = []
for prob in probs:
scam_score = prob[1].item()
label = "scam" if scam_score >= threshold else "ham"
results.append({
"label": label,
"confidence": scam_score if label == "scam" else prob[0].item(),
"scam_probability": scam_score
})
return results[0] if is_single else results
# Quick test
detector = IndicScamClassifier()
sample = "Congratulations, you won lottery. Call 9876543210 immediately."
print(detector.predict(sample))
- Downloads last month
- 66
Model tree for anmolshrivastav/indicbert-scam-classifier
Base model
ai4bharat/IndicBERTv2-MLM-only