Multilingual Health QA in Low-Resource African Languages β€” Cross-Base QLoRA Ensemble

Winning solution weights for the Zindi Multilingual Health Question-Answering challenge (open-ended generative health QA across 8 language–country subsets). This repository hosts the four QLoRA adapters that make up the generation arm of the winning submission (submission_v23_judgecol.csv).

Metric Score
Private LB (final ranking) 0.694836
Public LB 0.697241
Retrieval baseline (start of project) ~0.485

Competition metric: 0.37Β·ROUGE-1 F1 + 0.37Β·ROUGE-L F1 + 0.26Β·LLM-judge(1–5, normalised).


1. Architecture & approach

The 8 subsets β€” 4 African languages (Aka_Gha Akan/Twi, Amh_Eth Amharic, Lug_Uga Luganda, Swa_Ken Kiswahili) and 4 English-by-country splits (Eng_Eth, Eng_Gha, Eng_Ken, Eng_Uga) β€” are routed by a per-subset winner map into two regimes:

                  Train.csv Β· Val.csv Β· Test.csv
                              β”‚
        ETL: trainval retrieval pool Β· RAG-format SFT Β· cached embeddings
                              β”‚
                   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                   β”‚  Per-subset router   β”‚
                   β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”˜
        GENERATION      β”‚            β”‚      SELECTION
   (Aka_Gha,Eng_Gha,    β”‚            β”‚  (Eng_Ken,Eng_Uga,Swa_Ken,
        Eng_Eth)        β”‚            β”‚       Lug_Uga,Amh_Eth)
         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”    β”Œβ”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
         β”‚ 4 QLoRA bases    β”‚    β”‚ retrieve top-C candidates β”‚
         β”‚  (THIS REPO):    β”‚    β”‚  e5 + bge-m3 + TF-IDF     β”‚
         β”‚  Qwen2.5-14B n32 β”‚    β””β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
         β”‚  Qwen3-14B   n16 β”‚        β”‚  per-subset selector:
         β”‚  Gemma-3-12B n16 β”‚        β”‚  LGBMRanker / afro-xlmr CE /
         β”‚  Mistral-Nemo n16β”‚        β”‚  linear blend
         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”˜    β””β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
         MBR sample ~80/Q            β”‚  (selector models NOT in this repo)
         β†’ self-consensus pick       β”‚
                   β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
              Per-column assembly (3 columns scored independently):
              TargetR1F1/RLF1 = consensus (gen) / selector (sel)
              TargetLLM       = 7B-judge pick over the pool
                              β”‚
                submission_v23_judgecol.csv  (private LB 0.694836)

The two regimes:

  • Generation-bound subsets (Aka_Gha, Eng_Gha, Eng_Eth): answered by a 4-way cross-base ensemble of the QLoRA adapters in this repo. Each adapter is MBR-sampled (temperature 0.7, top-p 0.9), the ~80 pooled samples per question are scored by self-consensus (the sample with the highest mean unigram-F1 against the others), and the consensus sample is emitted. The gain comes from cross-base diversity (different pretraining lineages), not from any single stronger base.
  • Duplicate-rich subsets (Eng_Ken, Eng_Uga, Swa_Ken, Lug_Uga, Amh_Eth): near-identical questions exist in the training pool, so a real human answer is retrieved and selected rather than generated (LambdaRank / afro-xlmr cross-encoder / linear-blend selectors). Those selector models are not in this repo β€” see the full pipeline (link below); this repo is the generative model.

Adapters in this repo

Folder Base model License SFT data Decode
adapters/qwen2.5-14b-instruct Qwen/Qwen2.5-14B-Instruct Apache-2.0 train + val MBR n=32
adapters/qwen3-14b Qwen/Qwen3-14B Apache-2.0 train-only MBR n=16
adapters/gemma-3-12b-it unsloth/gemma-3-12b-it Gemma Terms train-only MBR n=16
adapters/mistral-nemo-12b unsloth/Mistral-Nemo-Instruct-2407 Apache-2.0 train-only MBR n=16

Each is a LoRA adapter (adapter_model.safetensors, ~450–540 MB) β€” load it on top of its base model; do not expect a standalone model.


2. Intended use

  • Intended: research reproduction of the competition result; a starting point for generative health-information QA in low-resource African languages (Akan/Twi, Amharic, Luganda, Kiswahili) and English health QA in African country contexts; studying cross-base MBR ensembling and QLoRA on multilingual low-resource tasks.
  • Domain: sexual & reproductive health, maternal & adolescent health, and gender-based-violence information, matching the challenge data.

Out of scope / not for

⚠️ Not medical advice. These models generate information-style text tuned to a ROUGE/LLM-judge benchmark. They are not a diagnostic tool and must not be used for clinical decisions or deployed to patients without expert review. They can hallucinate, omit warnings, or produce outdated guidance. Always route real health questions to qualified professionals.


3. Limitations & known behaviours

  • Akan/Twi is metric-bound (~0.354 proxy, oracle ~0.41). Even a Twi-native fine-tuned model scores lower on ROUGE because of paraphrase penalty against a single reference. Akan answers are fluent but lexically diverge from the gold.
  • Single-reference ROUGE optimisation. The models were selected to maximise a ROUGE-heavy benchmark; outputs are tuned to match reference phrasing/length, not to be maximally safe or exhaustive.
  • Style mimicry. The RAG prompt instructs the model to copy the length and structure of retrieved examples; answers inherit the (sometimes terse) style of the challenge corpus and may drop disclaimers by design.
  • Individual adapters β‰ˆ baseline. A single adapter is roughly at the competition's mid-pack; the winning score requires the 4-way ensemble + self-consensus (and, for the full submission, the retrieval/selection arm on the other 5 subsets). See reproduce_eval.py.
  • Amharic content is often absent from the pool (knowledge gap) and scores low (~0.20); it is served by retrieval, not these generators.
  • No external supervised data. Trained only on challenge Train/Val; no medical KB or guideline grounding.

4. Dependencies & environment

Trained and run on 1Γ— NVIDIA RTX 5090 (32 GB, CUDA 13), Python 3.13.

# Install CUDA-12.8 torch first, then the rest:
pip install torch==2.9.0 --index-url https://download.pytorch.org/whl/cu128
pip install transformers==4.57.3 peft==0.18.0 bitsandbytes==0.48.2 \
            accelerate==1.12.0 rouge-score==0.1.2 scikit-learn==1.6.1 \
            pandas==2.2.3 numpy==2.1.3
# For fast batched/MBR inference the pipeline uses vllm==0.13.0 (optional here).
  • A 12–14B base loads in 4-bit NF4 in ~8–10 GB; ~24–28 GB GPU is comfortable.
  • reproduce_eval.py is CPU-only (no GPU, no model download).
  • Env vars used by the full pipeline: MKL_THREADING_LAYER=GNU (vLLM/MKL under conda), PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True.

5. How to run inference

Single question, single adapter (illustrates one ensemble member):

python demo_inference.py \
  --adapter qwen2.5-14b-instruct \
  --subset Eng_Gha \
  --question "What are the danger signs during pregnancy that need urgent care?" \
  --train-csv /path/to/Train.csv          # optional: k=3 few-shot retrieval

Minimal code (what demo_inference.py does):

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import PeftModel

base = "Qwen/Qwen2.5-14B-Instruct"
bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
                         bnb_4bit_compute_dtype=torch.bfloat16, bnb_4bit_use_double_quant=True)
tok = AutoTokenizer.from_pretrained(base)
model = AutoModelForCausalLM.from_pretrained(base, quantization_config=bnb, device_map="auto")
model = PeftModel.from_pretrained(model, "adapters/qwen2.5-14b-instruct")

# Build the RAG prompt exactly as in training (see demo_inference.build_prompt):
messages = [{"role": "system", "content": system_prompt},
            {"role": "user",   "content": rag_user_prompt}]
prompt = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
out = model.generate(**tok(prompt, return_tensors="pt").to(model.device),
                     max_new_tokens=512, do_sample=False)
print(tok.decode(out[0], skip_special_tokens=True))

The prompt format matters: the adapters were trained with a specific system instruction + a k=3 retrieved-example RAG user turn + a word-budget hint. demo_inference.build_prompt() reproduces it byte-for-byte. Mistral-Nemo's chat template drops the system role, so the demo merges system→user for that adapter.

The deployed generation arm ensembles all four adapters: it draws MBR samples (temp 0.7, top-p 0.9) from each, pools them, and picks the self-consensus sample. See reproduce_eval.py for the exact pooling + selection logic.


6. Reproduce & verify the reported results (no GPU)

reproduce_eval.py recomputes the validation proxy for the three generation subsets directly from the saved decode samples (eval_data/) β€” the same 4-way pooling + self-consensus + UnicodeTokenizer ROUGE used in the submission:

python reproduce_eval.py

Expected (matches the reported numbers):

subset        n  ROUGE-1  ROUGE-L    proxy
Aka_Gha    1114   0.4353   0.2729   0.3541
Eng_Gha    1104   0.4997   0.3564   0.4281
Eng_Eth     564   0.7352   0.7052   0.7202

The 0.26-weighted LLM-judge component of the leaderboard metric cannot be recomputed offline; (R1+RL)/2 tracks the 74% ROUGE portion. A UnicodeTokenizer is required β€” the stock ROUGE tokenizer regexes on [a-z0-9] and silently collapses Amharic/Akan/Luganda scripts to ~0.


7. Training details

Each adapter: QLoRA β€” 4-bit NF4 base, LoRA r=32, Ξ±=64, dropout 0.05, targeting all attention + MLP projections; 2 epochs, effective batch 16, lr 1e-4, max-seq 2048, seed=42.

SFT format (RAG): each training row is a chat example = system (answer in <language>, match example style/length, no disclaimers) + user (k=3 leave-one-out retrieved same-subset Q&A few-shot + "answer in about N words" budget) + assistant (gold answer). Retrieval for few-shot uses LaBSE dense embeddings (English/Kiswahili) or TF-IDF char n-grams (Akan/Luganda). Qwen2.5-14B is trained on train+val; the other three on train-only (so they can be honestly scored on val). Gemma-3-12B is multimodal β€” LoRA is restricted to .*language_model.* (the SigLIP vision tower is skipped).


8. Data & licensing

  • Data: only the challenge-provided Train.csv / Val.csv / Test.csv. No external datasets; the retrieval pool, RAG few-shot examples, and all fine-tuning targets are built solely from challenge data. Challenge data is not redistributed in this repo.
  • Adapter weights & code: released under Apache-2.0, except that adapters/gemma-3-12b-it/ is a derivative of google/gemma-3-12b-it and is additionally subject to the Gemma Terms of Use (commercial use permitted). Each adapter is only usable together with its respective base model, whose own license applies.

9. Repository contents

.
β”œβ”€β”€ README.md                  # this model card
β”œβ”€β”€ DOCUMENTATION.md / .pdf     # full method write-up (ETL/modeling/inference/metrics)
β”œβ”€β”€ demo_inference.py           # single-question inference with one adapter
β”œβ”€β”€ reproduce_eval.py           # CPU-only reproduction of the reported val proxy
β”œβ”€β”€ requirements-inference.txt  # minimal env for the demos
β”œβ”€β”€ notebooks/demo.ipynb        # notebook walkthrough (eval + inference)
β”œβ”€β”€ adapters/
β”‚   β”œβ”€β”€ qwen2.5-14b-instruct/  # LoRA adapter (base: Qwen2.5-14B-Instruct)
β”‚   β”œβ”€β”€ qwen3-14b/             #             (base: Qwen3-14B)
β”‚   β”œβ”€β”€ gemma-3-12b-it/        #             (base: unsloth/gemma-3-12b-it)
β”‚   └── mistral-nemo-12b/      #             (base: Mistral-Nemo-Instruct-2407)
β”œβ”€β”€ eval_data/                 # saved val decode samples + gold, for reproduce_eval.py
β”œβ”€β”€ submission/
β”‚   └── submission_v23_judgecol.csv   # the exact winning submission (private LB 0.694836)
└── full_pipeline/             # end-to-end reproduction of the WHOLE 8-subset submission
    β”œβ”€β”€ reproduce.sh           #   staged driver: raw CSVs -> submission_v23_judgecol.csv
    β”œβ”€β”€ requirements.txt       #   full pinned environment
    └── scripts/               #   ETL, QLoRA training, retrieval/selection arm, MBR decode, builds

Two levels of reproduction:

  • reproduce_eval.py (top level) verifies the generation model's reported validation ROUGE offline, no GPU β€” the fast check for these weights.
  • full_pipeline/reproduce.sh regenerates the entire submission_v23_judgecol.csv (all 8 subsets, both arms) from the raw challenge CSVs; see DOCUMENTATION.md Β§12. This requires the challenge data and a GPU (~38 h serial on one RTX 5090).

10. Citation

Zindi β€” Multilingual Health Question-Answering Challenge.
Winning submission: cross-base 4-way QLoRA generation ensemble + per-subset
retrieval selection. Private LB 0.694836.
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for davidlowjw/multilingual-health-qa-crossbase-qlora

Base model

Qwen/Qwen2.5-14B
Adapter
(426)
this model