CLASP-Ar: Cloze-Style Prompting for Arabic Stance Detection

CLASP-Ar classifies the stance of an Arabic text (tweet) toward a target as Favor, Against, or None. Instead of a classification head, multitask learning, or an ensemble, it recasts stance detection as cloze-style masked language modeling. The target, a predicted sentiment label, and the text go into one prompt. The model fills a [MASK] token, and its choice is limited to a three-word Arabic verbalizer.

This is the system submitted by TTLab to the StanceEval-2026 Arabic Stance Detection Shared Task (SIGARAB ArabicNLP 2026).

Developed by Bhuvanesh Verma, Ali Abusaleh, Alexander Mehler (TTLab, Goethe University Frankfurt)
Model type BERT-large masked LM used as a prompt-based classifier
Base model asafaya/bert-large-arabic
Language Arabic (MSA and dialectal social-media text)
License Apache-2.0
Code https://github.com/aliabusaleh/ArabicStanceDetection_StanceEval2026

Results

Official StanceEval-2026 test sets. The metric is Favg2, the macro-F1 over Favor and Against. None is excluded from the average.

Test set Favg2
Track 1 71.36
Track 2 74.14

How it works

Each input is rendered as:

Target:{target}
Sentiment:{sentiment}
Stance: [MASK]
Text: {text}

The model reads the MLM logits at the [MASK] position and keeps only three vocabulary tokens, which it softmaxes into class probabilities:

Label Verbalizer token
Against ضد
Favor مع
None وسط
  • Sentiment slot. Training filled Sentiment: with a predicted label (Negative / Neutral / Positive) from a separate fine-tuned MARBERTv2 sentiment classifier, never with gold sentiment. Training and test used the same feature source. That sentiment model is not part of this release. Any 3-way Arabic sentiment classifier can fill the slot. Leaving it empty is out-of-distribution for this checkpoint and has not been evaluated.
  • Target. Targets are short topic phrases, usually in English as they appear in the training data (e.g. Covid Vaccine, Women empowerment, Digital Transformation, Women Driving).
  • Text preprocessing. Diacritics and tatweel are stripped. Non-Arabic characters (digits are kept) are removed. Characters repeated three or more times are capped at two. Whitespace is collapsed. inference.py applies this for you.

Usage

Download inference.py from this repository, then:

from huggingface_hub import hf_hub_download
import importlib.util

path = hf_hub_download("alighabusaleh/CLASP-Ar-Arabic-Stance", "inference.py")
spec = importlib.util.spec_from_file_location("clasp_inference", path)
clasp = importlib.util.module_from_spec(spec); spec.loader.exec_module(clasp)

clf = clasp.StanceClassifier("alighabusaleh/CLASP-Ar-Arabic-Stance")
texts = ["التطعيم ضروري لحماية المجتمع من الوباء"]
print(clf.predict(texts, targets=["Covid Vaccine"], sentiments=["Positive"]))
print(clf.predict_proba(texts, targets=["Covid Vaccine"], sentiments=["Positive"]))  # [Against, Favor, None]
Minimal version with plain transformers
import torch
from transformers import AutoTokenizer, BertForMaskedLM

tok = AutoTokenizer.from_pretrained("alighabusaleh/CLASP-Ar-Arabic-Stance")
model = BertForMaskedLM.from_pretrained("alighabusaleh/CLASP-Ar-Arabic-Stance").eval()
label_ids = tok.convert_tokens_to_ids(["ضد", "مع", "وسط"])  # Against, Favor, None

prompt = f"Target:Covid Vaccine\nSentiment:Positive\nStance: {tok.mask_token}\nText: التطعيم ضروري لحماية المجتمع من الوباء"
enc = tok(prompt, return_tensors="pt", truncation=True, max_length=512)
with torch.no_grad():
    logits = model(**enc).logits[0]
mask_pos = (enc["input_ids"][0] == tok.mask_token_id).nonzero()[0, 0]
probs = logits[mask_pos, label_ids].softmax(-1)
print(dict(zip(["Against", "Favor", "None"], probs.tolist())))

This skips the text preprocessing described above. Apply it to match training conditions.

Training

Training ran in two stages from asafaya/bert-large-arabic, on the same prompt format throughout.

Stage 1: intermediate stance training. 2 epochs on the StanceEval-2026 stance-task data: 11,500 examples, 17 targets (Favor 4,980 / Against 4,452 / None 2,068).

Stage 2: fine-tuning. 8 epochs on all of MawqifV2 train+dev: 4,121 examples, 3 targets (Favor 2,528 / Against 1,201 / None 392), with two changes for the rare None class:

  • None is oversampled up to the majority-class count using real None examples from ExaASC, not synthetic target-shuffled ones.
  • Cross-entropy uses inverse-frequency class weights.
Hyperparameter Value
Optimizer AdamW, weight decay 0.01
Learning rate 2e-5 (stage 1), 1e-5 (stage 2)
Layer-wise LR decay 0.95
Frozen layers embeddings + bottom 6 encoder layers
Schedule linear warmup (6%) + linear decay
Gradient clipping 1.0
Batch size / max length 32 / 512
Target dropout 0.25: during training the target is replaced by "هذا الموضوع" ("this topic") so the model reads the text instead of learning a per-target prior
Seed 42

The learning rate, target dropout, and number of frozen layers came from a small grid search on a target-disjoint split (20% of targets held out, best val Favg2 0.756). The epoch count came from one early-stopped run on that split. The released model was then retrained on all labeled data.

Limitations and bias

  • Trained on Saudi-centric social-media topics (MawqifV2, StanceEval). Performance on other dialects, domains, or long documents is untested.
  • None is the rarest class. Predictions for it are less reliable than for Favor or Against.
  • The predicted-sentiment input is noisy (about 55% agreement with MawqifV2 gold sentiment). The model depends on it, so a different sentiment classifier may shift predictions.
  • Stance labels on social-media text are subjective. Do not use this model to profile or make decisions about individuals.

Citation

@inproceedings{Verma:et:al:2026:CLASP,
  title     = {TTLab at StanceEval-2026: A Cloze-Style Prompting Approach for
               Arabic-Language Stance Detection (CLASP-Ar)},
  author    = {Bhuvanesh Verma and Ali Abusaleh and Alexander Mehler},
  booktitle = {SIGARAB ArabicNLP 2026 StanceEval Shared Task},
  year      = {2026},
  address   = {Budapest, Hungary},
  note      = {accepted}
}
Downloads last month
19
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for alighabusaleh/CLASP-Ar-Arabic-Stance

Finetuned
(1)
this model

Evaluation results

  • Favg2 (macro-F1 over Favor & Against) on StanceEval-2026 Track 1 (test)
    self-reported
    0.714
  • Favg2 (macro-F1 over Favor & Against) on StanceEval-2026 Track 2 (test)
    self-reported
    0.741