Instructions to use davidlowjw/multilingual-health-qa-crossbase-qlora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use davidlowjw/multilingual-health-qa-crossbase-qlora with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
- Multilingual Health QA in Low-Resource African Languages β Cross-Base QLoRA Ensemble
Multilingual Health QA in Low-Resource African Languages β Cross-Base QLoRA Ensemble
Winning solution weights for the Zindi Multilingual Health Question-Answering
challenge (open-ended generative health QA across 8 languageβcountry subsets).
This repository hosts the four QLoRA adapters that make up the generation arm
of the winning submission (submission_v23_judgecol.csv).
| Metric | Score |
|---|---|
| Private LB (final ranking) | 0.694836 |
| Public LB | 0.697241 |
| Retrieval baseline (start of project) | ~0.485 |
Competition metric: 0.37Β·ROUGE-1 F1 + 0.37Β·ROUGE-L F1 + 0.26Β·LLM-judge(1β5, normalised).
1. Architecture & approach
The 8 subsets β 4 African languages (Aka_Gha Akan/Twi, Amh_Eth Amharic,
Lug_Uga Luganda, Swa_Ken Kiswahili) and 4 English-by-country splits
(Eng_Eth, Eng_Gha, Eng_Ken, Eng_Uga) β are routed by a per-subset
winner map into two regimes:
Train.csv Β· Val.csv Β· Test.csv
β
ETL: trainval retrieval pool Β· RAG-format SFT Β· cached embeddings
β
ββββββββββββ΄βββββββββββ
β Per-subset router β
ββββββ¬βββββββββββββ¬βββββ
GENERATION β β SELECTION
(Aka_Gha,Eng_Gha, β β (Eng_Ken,Eng_Uga,Swa_Ken,
Eng_Eth) β β Lug_Uga,Amh_Eth)
ββββββββββββββββ΄ββββ βββββ΄βββββββββββββββββββββββ
β 4 QLoRA bases β β retrieve top-C candidates β
β (THIS REPO): β β e5 + bge-m3 + TF-IDF β
β Qwen2.5-14B n32 β βββββ¬βββββββββββββββββββββββ
β Qwen3-14B n16 β β per-subset selector:
β Gemma-3-12B n16 β β LGBMRanker / afro-xlmr CE /
β Mistral-Nemo n16β β linear blend
ββββββββββββββββ¬ββββ βββββ¬βββββββββββββββββββββββ
MBR sample ~80/Q β (selector models NOT in this repo)
β self-consensus pick β
ββββββββββ¬βββββββββ
Per-column assembly (3 columns scored independently):
TargetR1F1/RLF1 = consensus (gen) / selector (sel)
TargetLLM = 7B-judge pick over the pool
β
submission_v23_judgecol.csv (private LB 0.694836)
The two regimes:
- Generation-bound subsets (
Aka_Gha,Eng_Gha,Eng_Eth): answered by a 4-way cross-base ensemble of the QLoRA adapters in this repo. Each adapter is MBR-sampled (temperature 0.7, top-p 0.9), the ~80 pooled samples per question are scored by self-consensus (the sample with the highest mean unigram-F1 against the others), and the consensus sample is emitted. The gain comes from cross-base diversity (different pretraining lineages), not from any single stronger base. - Duplicate-rich subsets (
Eng_Ken,Eng_Uga,Swa_Ken,Lug_Uga,Amh_Eth): near-identical questions exist in the training pool, so a real human answer is retrieved and selected rather than generated (LambdaRank / afro-xlmr cross-encoder / linear-blend selectors). Those selector models are not in this repo β see the full pipeline (link below); this repo is the generative model.
Adapters in this repo
| Folder | Base model | License | SFT data | Decode |
|---|---|---|---|---|
adapters/qwen2.5-14b-instruct |
Qwen/Qwen2.5-14B-Instruct |
Apache-2.0 | train + val | MBR n=32 |
adapters/qwen3-14b |
Qwen/Qwen3-14B |
Apache-2.0 | train-only | MBR n=16 |
adapters/gemma-3-12b-it |
unsloth/gemma-3-12b-it |
Gemma Terms | train-only | MBR n=16 |
adapters/mistral-nemo-12b |
unsloth/Mistral-Nemo-Instruct-2407 |
Apache-2.0 | train-only | MBR n=16 |
Each is a LoRA adapter (adapter_model.safetensors, ~450β540 MB) β load it on
top of its base model; do not expect a standalone model.
2. Intended use
- Intended: research reproduction of the competition result; a starting point for generative health-information QA in low-resource African languages (Akan/Twi, Amharic, Luganda, Kiswahili) and English health QA in African country contexts; studying cross-base MBR ensembling and QLoRA on multilingual low-resource tasks.
- Domain: sexual & reproductive health, maternal & adolescent health, and gender-based-violence information, matching the challenge data.
Out of scope / not for
β οΈ Not medical advice. These models generate information-style text tuned to a ROUGE/LLM-judge benchmark. They are not a diagnostic tool and must not be used for clinical decisions or deployed to patients without expert review. They can hallucinate, omit warnings, or produce outdated guidance. Always route real health questions to qualified professionals.
3. Limitations & known behaviours
- Akan/Twi is metric-bound (~0.354 proxy, oracle ~0.41). Even a Twi-native fine-tuned model scores lower on ROUGE because of paraphrase penalty against a single reference. Akan answers are fluent but lexically diverge from the gold.
- Single-reference ROUGE optimisation. The models were selected to maximise a ROUGE-heavy benchmark; outputs are tuned to match reference phrasing/length, not to be maximally safe or exhaustive.
- Style mimicry. The RAG prompt instructs the model to copy the length and structure of retrieved examples; answers inherit the (sometimes terse) style of the challenge corpus and may drop disclaimers by design.
- Individual adapters β baseline. A single adapter is roughly at the
competition's mid-pack; the winning score requires the 4-way ensemble +
self-consensus (and, for the full submission, the retrieval/selection arm on
the other 5 subsets). See
reproduce_eval.py. - Amharic content is often absent from the pool (knowledge gap) and scores low (~0.20); it is served by retrieval, not these generators.
- No external supervised data. Trained only on challenge
Train/Val; no medical KB or guideline grounding.
4. Dependencies & environment
Trained and run on 1Γ NVIDIA RTX 5090 (32 GB, CUDA 13), Python 3.13.
# Install CUDA-12.8 torch first, then the rest:
pip install torch==2.9.0 --index-url https://download.pytorch.org/whl/cu128
pip install transformers==4.57.3 peft==0.18.0 bitsandbytes==0.48.2 \
accelerate==1.12.0 rouge-score==0.1.2 scikit-learn==1.6.1 \
pandas==2.2.3 numpy==2.1.3
# For fast batched/MBR inference the pipeline uses vllm==0.13.0 (optional here).
- A 12β14B base loads in 4-bit NF4 in ~8β10 GB; ~24β28 GB GPU is comfortable.
reproduce_eval.pyis CPU-only (no GPU, no model download).- Env vars used by the full pipeline:
MKL_THREADING_LAYER=GNU(vLLM/MKL under conda),PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True.
5. How to run inference
Single question, single adapter (illustrates one ensemble member):
python demo_inference.py \
--adapter qwen2.5-14b-instruct \
--subset Eng_Gha \
--question "What are the danger signs during pregnancy that need urgent care?" \
--train-csv /path/to/Train.csv # optional: k=3 few-shot retrieval
Minimal code (what demo_inference.py does):
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import PeftModel
base = "Qwen/Qwen2.5-14B-Instruct"
bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16, bnb_4bit_use_double_quant=True)
tok = AutoTokenizer.from_pretrained(base)
model = AutoModelForCausalLM.from_pretrained(base, quantization_config=bnb, device_map="auto")
model = PeftModel.from_pretrained(model, "adapters/qwen2.5-14b-instruct")
# Build the RAG prompt exactly as in training (see demo_inference.build_prompt):
messages = [{"role": "system", "content": system_prompt},
{"role": "user", "content": rag_user_prompt}]
prompt = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
out = model.generate(**tok(prompt, return_tensors="pt").to(model.device),
max_new_tokens=512, do_sample=False)
print(tok.decode(out[0], skip_special_tokens=True))
The prompt format matters: the adapters were trained with a specific system instruction + a k=3 retrieved-example RAG user turn + a word-budget hint.
demo_inference.build_prompt()reproduces it byte-for-byte. Mistral-Nemo's chat template drops the system role, so the demo merges systemβuser for that adapter.
The deployed generation arm ensembles all four adapters: it draws MBR samples
(temp 0.7, top-p 0.9) from each, pools them, and picks the self-consensus sample.
See reproduce_eval.py for the exact pooling + selection logic.
6. Reproduce & verify the reported results (no GPU)
reproduce_eval.py recomputes the validation proxy for the three generation
subsets directly from the saved decode samples (eval_data/) β the same
4-way pooling + self-consensus + UnicodeTokenizer ROUGE used in the submission:
python reproduce_eval.py
Expected (matches the reported numbers):
subset n ROUGE-1 ROUGE-L proxy
Aka_Gha 1114 0.4353 0.2729 0.3541
Eng_Gha 1104 0.4997 0.3564 0.4281
Eng_Eth 564 0.7352 0.7052 0.7202
The 0.26-weighted LLM-judge component of the leaderboard metric cannot be
recomputed offline; (R1+RL)/2 tracks the 74% ROUGE portion. A UnicodeTokenizer
is required β the stock ROUGE tokenizer regexes on [a-z0-9] and silently
collapses Amharic/Akan/Luganda scripts to ~0.
7. Training details
Each adapter: QLoRA β 4-bit NF4 base, LoRA r=32, Ξ±=64, dropout 0.05,
targeting all attention + MLP projections; 2 epochs, effective batch 16,
lr 1e-4, max-seq 2048, seed=42.
SFT format (RAG): each training row is a chat example =
system (answer in <language>, match example style/length, no disclaimers) +
user (k=3 leave-one-out retrieved same-subset Q&A few-shot + "answer in about N
words" budget) + assistant (gold answer). Retrieval for few-shot uses LaBSE dense
embeddings (English/Kiswahili) or TF-IDF char n-grams (Akan/Luganda).
Qwen2.5-14B is trained on train+val; the other three on train-only (so
they can be honestly scored on val). Gemma-3-12B is multimodal β LoRA is
restricted to .*language_model.* (the SigLIP vision tower is skipped).
8. Data & licensing
- Data: only the challenge-provided
Train.csv/Val.csv/Test.csv. No external datasets; the retrieval pool, RAG few-shot examples, and all fine-tuning targets are built solely from challenge data. Challenge data is not redistributed in this repo. - Adapter weights & code: released under Apache-2.0, except that
adapters/gemma-3-12b-it/is a derivative ofgoogle/gemma-3-12b-itand is additionally subject to the Gemma Terms of Use (commercial use permitted). Each adapter is only usable together with its respective base model, whose own license applies.
9. Repository contents
.
βββ README.md # this model card
βββ DOCUMENTATION.md / .pdf # full method write-up (ETL/modeling/inference/metrics)
βββ demo_inference.py # single-question inference with one adapter
βββ reproduce_eval.py # CPU-only reproduction of the reported val proxy
βββ requirements-inference.txt # minimal env for the demos
βββ notebooks/demo.ipynb # notebook walkthrough (eval + inference)
βββ adapters/
β βββ qwen2.5-14b-instruct/ # LoRA adapter (base: Qwen2.5-14B-Instruct)
β βββ qwen3-14b/ # (base: Qwen3-14B)
β βββ gemma-3-12b-it/ # (base: unsloth/gemma-3-12b-it)
β βββ mistral-nemo-12b/ # (base: Mistral-Nemo-Instruct-2407)
βββ eval_data/ # saved val decode samples + gold, for reproduce_eval.py
βββ submission/
β βββ submission_v23_judgecol.csv # the exact winning submission (private LB 0.694836)
βββ full_pipeline/ # end-to-end reproduction of the WHOLE 8-subset submission
βββ reproduce.sh # staged driver: raw CSVs -> submission_v23_judgecol.csv
βββ requirements.txt # full pinned environment
βββ scripts/ # ETL, QLoRA training, retrieval/selection arm, MBR decode, builds
Two levels of reproduction:
reproduce_eval.py(top level) verifies the generation model's reported validation ROUGE offline, no GPU β the fast check for these weights.full_pipeline/reproduce.shregenerates the entiresubmission_v23_judgecol.csv(all 8 subsets, both arms) from the raw challenge CSVs; seeDOCUMENTATION.mdΒ§12. This requires the challenge data and a GPU (~38 h serial on one RTX 5090).
10. Citation
Zindi β Multilingual Health Question-Answering Challenge.
Winning submission: cross-base 4-way QLoRA generation ensemble + per-subset
retrieval selection. Private LB 0.694836.
- Downloads last month
- -