Yousef9Zaghloul commited on
Commit
6ca25ff
·
verified ·
1 Parent(s): 94fbedc

Initial demo: side-by-side fill-mask comparison vioBERT-v3 vs MARBERTv2

Browse files
Files changed (3) hide show
  1. README.md +32 -7
  2. app.py +111 -0
  3. requirements.txt +4 -0
README.md CHANGED
@@ -1,12 +1,37 @@
1
  ---
2
- title: VioBERT V3 Demo
3
- emoji: 📈
4
- colorFrom: gray
5
- colorTo: purple
6
  sdk: gradio
7
- sdk_version: 6.13.0
8
  app_file: app.py
9
- pinned: false
 
 
 
 
 
 
 
 
 
 
 
 
 
10
  ---
11
 
12
- Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ title: vioBERT-v3 Live Demo
3
+ emoji: 🩺
4
+ colorFrom: green
5
+ colorTo: blue
6
  sdk: gradio
7
+ sdk_version: 4.44.0
8
  app_file: app.py
9
+ pinned: true
10
+ license: apache-2.0
11
+ short_description: Live fill-mask — vioBERT-v3 vs MARBERTv2
12
+ tags:
13
+ - arabic
14
+ - medical
15
+ - bert
16
+ - fill-mask
17
+ - clinical-nlp
18
+ - domain-adaptation
19
+ - mena
20
+ models:
21
+ - Vionex-digital/vioBERT-v3
22
+ - UBC-NLP/MARBERTv2
23
  ---
24
 
25
+ # vioBERT-v3 Live Demo
26
+
27
+ Side-by-side fill-mask comparison: **vioBERT-v3** (Arabic medical BERT, MARBERTv2 + DAPT on 1.12M medical docs) vs **MARBERTv2** (base, trained on 1B Arabic tweets).
28
+
29
+ ## What you'll see
30
+
31
+ Type an Arabic sentence with `[MASK]`. Both models predict the masked token. On medical content, vioBERT-v3 fills with anatomy / drugs / conditions; MARBERTv2 often fills with politics / sports / news terms (its training distribution).
32
+
33
+ ## Why it matters
34
+
35
+ There are 422M Arabic speakers and previously zero strong open biomedical Arabic LM. vioBERT-v3 is the first. Apache-2.0 — fine-tune commercially without asking.
36
+
37
+ 📦 [Model card →](https://huggingface.co/Vionex-digital/vioBERT-v3)
app.py ADDED
@@ -0,0 +1,111 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """vioBERT-v3 vs MARBERTv2 — side-by-side fill-mask demo.
2
+
3
+ Visitors see *why* domain adaptation matters for Arabic medical NLP in
4
+ real time, on their own sentences. This is the marketing artifact —
5
+ benchmarks on a model card don't convince like a live A/B does.
6
+ """
7
+
8
+ import gradio as gr
9
+ from transformers import AutoModelForMaskedLM, AutoTokenizer
10
+ import torch
11
+ import time
12
+
13
+
14
+ VIOBERT_ID = "Vionex-digital/vioBERT-v3"
15
+ BASE_ID = "UBC-NLP/MARBERTv2"
16
+
17
+ print("[boot] loading models...")
18
+ t0 = time.time()
19
+ viobert_tok = AutoTokenizer.from_pretrained(VIOBERT_ID)
20
+ viobert = AutoModelForMaskedLM.from_pretrained(VIOBERT_ID).eval()
21
+ base_tok = AutoTokenizer.from_pretrained(BASE_ID)
22
+ base = AutoModelForMaskedLM.from_pretrained(BASE_ID).eval()
23
+ print(f"[boot] models loaded in {time.time()-t0:.1f}s")
24
+
25
+
26
+ def fill_mask(model, tok, text, top_k=5):
27
+ if "[MASK]" not in text and tok.mask_token not in text:
28
+ return [("⚠️ add [MASK] in the sentence", 0.0)]
29
+ text = text.replace("[MASK]", tok.mask_token)
30
+ inputs = tok(text, return_tensors="pt")
31
+ mask_idx = (inputs["input_ids"][0] == tok.mask_token_id).nonzero(as_tuple=True)[0]
32
+ if len(mask_idx) == 0:
33
+ return [("⚠️ no [MASK] found", 0.0)]
34
+ with torch.no_grad():
35
+ logits = model(**inputs).logits[0, mask_idx[0]]
36
+ probs = torch.softmax(logits, dim=-1)
37
+ top = torch.topk(probs, top_k)
38
+ return [(tok.decode(t).strip(), float(s)) for t, s in zip(top.indices, top.values)]
39
+
40
+
41
+ def compare(text):
42
+ if not text.strip():
43
+ return "👈 Enter an Arabic sentence with [MASK]", "👈 Enter an Arabic sentence with [MASK]"
44
+ vio = fill_mask(viobert, viobert_tok, text)
45
+ bas = fill_mask(base, base_tok, text)
46
+ def fmt(rows, label, color):
47
+ lines = [f"### {label}"]
48
+ for tok_str, score in rows:
49
+ bar = "█" * int(score * 30)
50
+ lines.append(f"`{tok_str}` &nbsp;&nbsp; **{score:.3f}** &nbsp; <span style='color:{color}'>{bar}</span>")
51
+ return "\n\n".join(lines)
52
+ return fmt(vio, "🩺 vioBERT-v3 (medical-adapted)", "#16a34a"), fmt(bas, "📰 MARBERTv2 (base, tweets)", "#64748b")
53
+
54
+
55
+ EXAMPLES = [
56
+ ["المريض يعاني من [MASK] في الصدر"],
57
+ ["تم تشخيص الحالة على أنها [MASK] من النوع الثاني"],
58
+ ["الجرعة الموصى بها للأطفال هي [MASK] ملليجرام"],
59
+ ["تظهر صورة الأشعة وجود [MASK] في الرئة اليمنى"],
60
+ ["يحتاج المريض إلى [MASK] فوري في غرفة العمليات"],
61
+ ["ارتفاع [MASK] الدم قد يؤدي إلى السكتة الدماغية"],
62
+ ]
63
+
64
+ CSS = """
65
+ .gradio-container {max-width: 1100px !important;}
66
+ .markdown-text {direction: rtl; text-align: right; font-size: 1.1em;}
67
+ """
68
+
69
+ DESCRIPTION = """
70
+ # 🩺 vioBERT-v3 vs MARBERTv2 — live fill-mask comparison
71
+
72
+ **vioBERT-v3** is the first Arabic medical BERT — MARBERTv2 + 22K steps of continued pretraining
73
+ on **Shifaa**, our 1.12M-document Arabic medical corpus.
74
+
75
+ This Space lets you see the difference live. Type any Arabic sentence with `[MASK]`,
76
+ get top-5 predictions from both models side-by-side. Medical sentences are where the
77
+ gap is widest.
78
+
79
+ Benchmarks vs MARBERTv2: −82.7% medical PPL · +15.6 pp fill-mask Top-5 ·
80
+ +0.93 pp medical NER F1 *(exceeds the +0.62 pp BioBERT got on English)*
81
+
82
+ 📦 [Model](https://huggingface.co/Vionex-digital/vioBERT-v3) ·
83
+ 🏥 [Vionex Digital Solutions](https://huggingface.co/Vionex-digital) ·
84
+ 📜 Apache-2.0 · free for commercial use
85
+ """
86
+
87
+ with gr.Blocks(css=CSS, theme=gr.themes.Soft(primary_hue="green")) as demo:
88
+ gr.Markdown(DESCRIPTION)
89
+ with gr.Row():
90
+ text = gr.Textbox(
91
+ label="Arabic medical sentence with [MASK]",
92
+ placeholder="المريض يعاني من [MASK] في الصدر",
93
+ rtl=True,
94
+ lines=2,
95
+ )
96
+ with gr.Row():
97
+ btn = gr.Button("🔍 Compare predictions", variant="primary", scale=1)
98
+ with gr.Row():
99
+ vio_out = gr.Markdown(label="vioBERT-v3", elem_classes="markdown-text")
100
+ base_out = gr.Markdown(label="MARBERTv2", elem_classes="markdown-text")
101
+ gr.Examples(examples=EXAMPLES, inputs=text, label="Try a medical example")
102
+ btn.click(fn=compare, inputs=text, outputs=[vio_out, base_out])
103
+ text.submit(fn=compare, inputs=text, outputs=[vio_out, base_out])
104
+ gr.Markdown(
105
+ "---\n*Built by Vionex Digital Solutions · "
106
+ "Domain-adaptive pretraining for Arabic medical NLP. "
107
+ "If this is useful, give vioBERT a like on its [model page](https://huggingface.co/Vionex-digital/vioBERT-v3) or cite the upcoming paper.*"
108
+ )
109
+
110
+ if __name__ == "__main__":
111
+ demo.launch()
requirements.txt ADDED
@@ -0,0 +1,4 @@
 
 
 
 
 
1
+ transformers>=4.44.0
2
+ torch>=2.0
3
+ gradio>=4.44.0
4
+ sentencepiece