Instructions to use alighabusaleh/CLASP-Ar-Arabic-Stance with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use alighabusaleh/CLASP-Ar-Arabic-Stance with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="alighabusaleh/CLASP-Ar-Arabic-Stance")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForMaskedLM tokenizer = AutoTokenizer.from_pretrained("alighabusaleh/CLASP-Ar-Arabic-Stance") model = AutoModelForMaskedLM.from_pretrained("alighabusaleh/CLASP-Ar-Arabic-Stance", device_map="auto") - Notebooks
- Google Colab
- Kaggle
CLASP-Ar: Cloze-Style Prompting for Arabic Stance Detection
CLASP-Ar classifies the stance of an Arabic text (tweet) toward a target as
Favor, Against, or None. Instead of a classification head, multitask learning, or
an ensemble, it recasts stance detection as cloze-style masked language modeling. The
target, a predicted sentiment label, and the text go into one prompt. The model fills
a [MASK] token, and its choice is limited to a three-word Arabic verbalizer.
This is the system submitted by TTLab to the StanceEval-2026 Arabic Stance Detection Shared Task (SIGARAB ArabicNLP 2026).
| Developed by | Bhuvanesh Verma, Ali Abusaleh, Alexander Mehler (TTLab, Goethe University Frankfurt) |
| Model type | BERT-large masked LM used as a prompt-based classifier |
| Base model | asafaya/bert-large-arabic |
| Language | Arabic (MSA and dialectal social-media text) |
| License | Apache-2.0 |
| Code | https://github.com/aliabusaleh/ArabicStanceDetection_StanceEval2026 |
Results
Official StanceEval-2026 test sets. The metric is Favg2, the macro-F1 over Favor
and Against. None is excluded from the average.
| Test set | Favg2 |
|---|---|
| Track 1 | 71.36 |
| Track 2 | 74.14 |
How it works
Each input is rendered as:
Target:{target}
Sentiment:{sentiment}
Stance: [MASK]
Text: {text}
The model reads the MLM logits at the [MASK] position and keeps only three vocabulary
tokens, which it softmaxes into class probabilities:
| Label | Verbalizer token |
|---|---|
| Against | ضد |
| Favor | مع |
| None | وسط |
- Sentiment slot. Training filled
Sentiment:with a predicted label (Negative/Neutral/Positive) from a separate fine-tuned MARBERTv2 sentiment classifier, never with gold sentiment. Training and test used the same feature source. That sentiment model is not part of this release. Any 3-way Arabic sentiment classifier can fill the slot. Leaving it empty is out-of-distribution for this checkpoint and has not been evaluated. - Target. Targets are short topic phrases, usually in English as they appear in the
training data (e.g.
Covid Vaccine,Women empowerment,Digital Transformation,Women Driving). - Text preprocessing. Diacritics and tatweel are stripped. Non-Arabic characters
(digits are kept) are removed. Characters repeated three or more times are capped at
two. Whitespace is collapsed.
inference.pyapplies this for you.
Usage
Download inference.py from this repository, then:
from huggingface_hub import hf_hub_download
import importlib.util
path = hf_hub_download("alighabusaleh/CLASP-Ar-Arabic-Stance", "inference.py")
spec = importlib.util.spec_from_file_location("clasp_inference", path)
clasp = importlib.util.module_from_spec(spec); spec.loader.exec_module(clasp)
clf = clasp.StanceClassifier("alighabusaleh/CLASP-Ar-Arabic-Stance")
texts = ["التطعيم ضروري لحماية المجتمع من الوباء"]
print(clf.predict(texts, targets=["Covid Vaccine"], sentiments=["Positive"]))
print(clf.predict_proba(texts, targets=["Covid Vaccine"], sentiments=["Positive"])) # [Against, Favor, None]
Minimal version with plain transformers
import torch
from transformers import AutoTokenizer, BertForMaskedLM
tok = AutoTokenizer.from_pretrained("alighabusaleh/CLASP-Ar-Arabic-Stance")
model = BertForMaskedLM.from_pretrained("alighabusaleh/CLASP-Ar-Arabic-Stance").eval()
label_ids = tok.convert_tokens_to_ids(["ضد", "مع", "وسط"]) # Against, Favor, None
prompt = f"Target:Covid Vaccine\nSentiment:Positive\nStance: {tok.mask_token}\nText: التطعيم ضروري لحماية المجتمع من الوباء"
enc = tok(prompt, return_tensors="pt", truncation=True, max_length=512)
with torch.no_grad():
logits = model(**enc).logits[0]
mask_pos = (enc["input_ids"][0] == tok.mask_token_id).nonzero()[0, 0]
probs = logits[mask_pos, label_ids].softmax(-1)
print(dict(zip(["Against", "Favor", "None"], probs.tolist())))
This skips the text preprocessing described above. Apply it to match training conditions.
Training
Training ran in two stages from asafaya/bert-large-arabic, on the same prompt format
throughout.
Stage 1: intermediate stance training. 2 epochs on the StanceEval-2026 stance-task data: 11,500 examples, 17 targets (Favor 4,980 / Against 4,452 / None 2,068).
Stage 2: fine-tuning. 8 epochs on all of MawqifV2 train+dev: 4,121 examples,
3 targets (Favor 2,528 / Against 1,201 / None 392), with two changes for the rare None
class:
Noneis oversampled up to the majority-class count using realNoneexamples from ExaASC, not synthetic target-shuffled ones.- Cross-entropy uses inverse-frequency class weights.
| Hyperparameter | Value |
|---|---|
| Optimizer | AdamW, weight decay 0.01 |
| Learning rate | 2e-5 (stage 1), 1e-5 (stage 2) |
| Layer-wise LR decay | 0.95 |
| Frozen layers | embeddings + bottom 6 encoder layers |
| Schedule | linear warmup (6%) + linear decay |
| Gradient clipping | 1.0 |
| Batch size / max length | 32 / 512 |
| Target dropout | 0.25: during training the target is replaced by "هذا الموضوع" ("this topic") so the model reads the text instead of learning a per-target prior |
| Seed | 42 |
The learning rate, target dropout, and number of frozen layers came from a small grid search on a target-disjoint split (20% of targets held out, best val Favg2 0.756). The epoch count came from one early-stopped run on that split. The released model was then retrained on all labeled data.
Limitations and bias
- Trained on Saudi-centric social-media topics (MawqifV2, StanceEval). Performance on other dialects, domains, or long documents is untested.
Noneis the rarest class. Predictions for it are less reliable than forFavororAgainst.- The predicted-sentiment input is noisy (about 55% agreement with MawqifV2 gold sentiment). The model depends on it, so a different sentiment classifier may shift predictions.
- Stance labels on social-media text are subjective. Do not use this model to profile or make decisions about individuals.
Citation
@inproceedings{Verma:et:al:2026:CLASP,
title = {TTLab at StanceEval-2026: A Cloze-Style Prompting Approach for
Arabic-Language Stance Detection (CLASP-Ar)},
author = {Bhuvanesh Verma and Ali Abusaleh and Alexander Mehler},
booktitle = {SIGARAB ArabicNLP 2026 StanceEval Shared Task},
year = {2026},
address = {Budapest, Hungary},
note = {accepted}
}
- Downloads last month
- 19
Model tree for alighabusaleh/CLASP-Ar-Arabic-Stance
Base model
asafaya/bert-large-arabicEvaluation results
- Favg2 (macro-F1 over Favor & Against) on StanceEval-2026 Track 1 (test)self-reported0.714
- Favg2 (macro-F1 over Favor & Against) on StanceEval-2026 Track 2 (test)self-reported0.741