Llama-3.1-8B-Instruct-Lingala-QLoRA-merged-v2

Built with Llama.

meta-llama/Llama-3.1-8B-Instruct adapted to Lingala through supervised fine-tuning (QLoRA), with the adapters merged into the base weights. Developed by Congo Digital Services (CDS SARL) for the MALOBA project, with UNDP Republic of Congo. The model generates Lingala text in five registers: urban, educational, summarisation, formal and translation.

This is the LLM served by the MALOBA project (gateway alias maloba-llm) and the one measured by the project's evaluations.

Sister repository, with the QLoRA adapters only: Congo-digital-service/Llama-3.1-8B-Instruct-Lingala-QLoRA-adapters. See Traceability for how the two repositories relate.

For French → Lingala translation, use Congo-digital-service/nllb-200-lingala-maloba-h1, the translation model retained by the project. On the 80 translation examples of this model's test split, this Llama reaches 1.27 BLEU / 29.61 chrF, while NLLB-200 1.3B without any fine-tuning reaches 22.96 BLEU / 54.08 chrF on the same examples.

Summary

Task Conversational generation in Lingala, in five registers
Base model meta-llama/Llama-3.1-8B-Instruct
Method QLoRA (4-bit base, LoRA rank 16 on all 7 projections), adapters then merged (8B parameters)
Training data MALOBA Lingala instruction corpus: 4,818 examples in five registers, training split augmented to 5,700
Main result ROUGE-L F1 0.1521 → 0.2379 (+56.4 % relative) on 964 held-out examples
Human evaluation Not conducted
Status Model served by MALOBA (alias maloba-llm); end-to-end endpoint still to be validated
Licence Llama 3.1 Community License
Developed by Congo Digital Services (CDS SARL), congo-digital.com

Uses

Direct use: conversational text generation in Lingala with transformers, or behind a compatible inference server (see Deployment).

How to prompt: write the request in plain Lingala, as in the example below. The training data paired each response with a register tag ([URBAIN], [PÉDAGOGIQUE], [RÉSUMÉ], [TRADUCTION], [INCONNU]) rather than a written instruction (see Limitations). With a tag, the model produces fluent Lingala in that register, but often off topic.

Out of scope:

  • French → Lingala translation: use the NLLB model linked above;
  • tasks that need factual accuracy (medical, legal, administrative) without human review;
  • certified professional translation;
  • any use that breaches the Llama 3.1 Acceptable Use Policy.

How to use

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "Congo-digital-service/Llama-3.1-8B-Instruct-Lingala-QLoRA-merged-v2"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.bfloat16, device_map="auto")

messages = [
    {"role": "user", "content": "Loba na ngai na lingala : ndenge nini okoki kosala mombongo na Kinshasa ?"}
]
inputs = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, return_tensors="pt"
).to(model.device)
outputs = model.generate(inputs, max_new_tokens=256, do_sample=True, temperature=0.7)
print(tokenizer.decode(outputs[0][inputs.shape[-1]:], skip_special_tokens=True))

Training data

MALOBA Lingala instruction corpus, built from tasks annotated in Label Studio. In the urban, educational, summarisation and formal registers, the Lingala was pre-filled by a model, then validated or corrected by annotators. The translation register holds human translations.

Register Examples
Urban 1,425
Educational 1,063
Summarisation 1,041
Formal 889
Translation 400
Total 4,818
  • Format: Alpaca (instruction, optional input, response). The instruction field holds a register tag, not a written instruction.
  • Split: stratified 80/20. That gives 3,854 training examples and 964 test examples; the test set is not augmented and is held out.
  • Augmentation: applied to the training split only, to balance the five registers at 1,140 examples each (5,700 in total).
  • Dataset repositories: Congo-digital-service/Dataset-Lingala-Alpaca-LLM holds 6,664 examples (5,700 train + 964 test). The augmented version is at Congo-digital-service/Lingala-Alpaca-Dataset-Augmented. An orthographically harmonised version of the corpus, Congo-digital-service/maloba-llm-alpaca-harmonise, was produced after this model was trained; this model has not been retrained on it.

Training procedure

  • Environment: Google Colab, July–August 2026.
  • Method: QLoRA (4-bit base model, LoRA adapters), then adapters merged. The final step (1,071) is the checkpoint retained: its validation loss equals that of step 1,000.
LoRA parameter Value
Rank r / lora_alpha 16 / 16 (scale 1)
Dropout 0.1
Target modules q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
Bias not trained
Hyperparameter Value
Epochs 3 (1,071 steps)
Batch size 2 per device
Gradient accumulation 8 (effective batch 16)
Learning rate 2 × 10⁻⁵, cosine schedule
Evaluation every 200 steps
Chat template Llama 3.1 Instruct
Step Training loss Validation loss Entropy Mean token accuracy
200 1.5638 1.5747 1.6141 67.15 %
400 1.3008 1.3514 1.3425 69.82 %
600 1.2714 1.2901 1.3146 70.83 %
800 1.2072 1.2741 1.2669 71.14 %
1,000 1.2211 1.2720 1.2690 71.15 %
1,071 1.1956 1.2720 1.2689 71.19 %

Validation loss falls by 19 %, then plateaus at 1.272; the gap between training and validation loss stays below 0.08, with no sign of overfitting.

Run logs were not archived with the weights. The values above come from the notebook outputs, as reported in the project's August deliverable. The seed, optimiser, warm-up and maximum sequence length were set in the Colab notebook and were not kept.

Evaluation

Instruction test set (964 held-out examples)

Metric Base model Fine-tuned model Comparable?
ROUGE-L F1 (primary metric) 0.1521 0.2379 Yes: same metric, same test set (+56.4 % relative)
ROUGE-1 F1 — 0.2939 —
ROUGE-2 F1 — 0.0789 —
Perplexity — 17.7211 —
Mean token accuracy — 71.20 % —

Not comparable, so not reported as a comparison:

  • BLEU: the base-model value (2.328) is SacreBLEU on a 0–100 scale; the fine-tuned value (0.0459) is BLEU on a 0–1 scale. Side by side, they suggest a regression that does not exist.
  • Token F1 (base, 0.1721) and mean token accuracy (fine-tuned, 71.20 %) are two different metrics.

Mean token accuracy is computed with teacher forcing: the model predicts each reference token given the correct previous ones. It is not a measure of response quality or of the share of acceptable answers.

Test contamination: 13 of the 964 test examples (1.3 %) have an equivalent in the training split. This does not inflate the scores noticeably.

Translation register (80 examples)

The 80 examples are all the translation-register examples of the 964-example test split. The model reaches BLEU 1.27 / chrF 29.61 (run of 21 September 2026, deterministic decoding; an independent implementation gave 1.03 / 28.2). The model knows the words but does not build the sentence. On the same 80 examples and with the same normalisation, NLLB-200 1.3B without fine-tuning reaches 22.96 BLEU / 54.08 chrF. This is why the project serves translation with a dedicated NLLB model.

Human evaluation

Not conducted. The project's human evaluation campaign was focused on the translation model (NLLB-Maloba-h1). No human rating of this model's responses exists, so the terms of reference's 85 % precision indicator is not assessed for this model.

The automatic scores above measure closeness to reference answers, not usability.

Limitations

  • Five registers learned, not instruction following. The corpus contains only 5 distinct instructions (the register tags) for 5,700 training examples. The model has learned to write in five registers, not to follow varied instructions. Retraining on the same corpus would not change this; a corpus of varied, written instructions is the condition for progress.
  • Redundant training data: 31.8 % of the training lines are redundant. The augmentation copies a line and changes it only when a French trigger word appears in it, which is rare in Lingala text. Some examples are repeated up to nine times.
  • Mislabelled register: the [INCONNU] ("unknown") tag is in fact the formal register, mislabelled by a key typo in the training notebook.
  • Unharmonised spelling: this model was trained before the corpus was orthographically harmonised. Its training data mixes spellings, including Unicode look-alikes of ɛ (Greek ε), and the model may reproduce these inconsistencies.
  • Unbalanced test set: the 964 test examples are not evenly spread across registers.
  • Weak translation performance: see above.
  • Coverage: regional varieties, and specialised or technical domains absent from the corpus, are not evaluated.
  • Human review required before any use with consequences for people.

Deployment

Component Detail
Inference server vLLM on GPU
MALOBA gateway LiteLLM, alias maloba-llm, virtual keys per application
Format OpenAI Chat Completions compatible (POST /v1/chat/completions, choices[0].message.content)
Access interface LoBAI (LibreChat-based), with Keycloak / OpenID Connect

Validation status

The gateway side is configured. The end-to-end endpoint (gateway → inference server → model) remains to be validated, with latency and error-rate measurement, once the inference server is deployed.

Traceability

  • This repository: weights uploaded on 27 July 2026 (commit c7d33fda) and unchanged since. The tokenizer configuration was fixed for compatibility with Transformers 4 and vLLM on 3, 13 and 14 August (last change cb1daa66). Later commits edit this card only. The project's evaluations and its deployment use this model.
  • QLoRA adapters: Congo-digital-service/Llama-3.1-8B-Instruct-Lingala-QLoRA-adapters, weights uploaded on 10 August 2026 (commit 21a7f7ea).
  • Correspondence between the two repositories is not demonstrated. This merged model was first published on 27 July 2026 and the adapters on 10 August 2026. The published adapters can only have produced this model if they were republished unchanged.
  • Version manifest and SHA-256 checksums: kept in the project's delivery GitHub repository.

Licence and attribution

This model is a derivative of Llama 3.1. It is distributed under the Llama 3.1 Community License, and its use must comply with the Llama 3.1 Acceptable Use Policy.

Llama 3.1 is licensed under the Llama 3.1 Community License, Copyright © Meta Platforms, Inc. All Rights Reserved.

As required by that licence:

  • this model is "Built with Llama";
  • its name begins with "Llama";
  • any redistribution must keep this notice and a copy of the licence.

Training data: the licence of the MALOBA instruction datasets used here is being aligned on the Nwulite Obodo Open Data License (NOODL-1.0), already applied to the harmonised corpus. NOODL covers the data, not this model. CDS asks users from high-income countries and commercial entities to publicly credit the MALOBA Project (UNDP Republic of Congo — language digitalisation initiative) in any publication, product, model or output derived from this model.

Environmental impact

Training ran on Google Colab. The GPU type and energy use were not recorded; emissions can be estimated with the ML Impact calculator (Lacoste et al., 2019).

Citation

@misc{cds2026lingala,
  title        = {Llama-3.1-8B-Instruct-Lingala-QLoRA},
  author       = {{Congo Digital Services}},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/Congo-digital-service/Llama-3.1-8B-Instruct-Lingala-QLoRA-merged-v2}},
  note         = {MALOBA project, UNDP Republic of Congo. Built with Llama.}
}

Contact

Card updated on 30 September 2026.

Downloads last month
865
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Congo-digital-service/Llama-3.1-8B-Instruct-Lingala-QLoRA-merged-v2

Finetuned
(3262)
this model

Dataset used to train Congo-digital-service/Llama-3.1-8B-Instruct-Lingala-QLoRA-merged-v2

Paper for Congo-digital-service/Llama-3.1-8B-Instruct-Lingala-QLoRA-merged-v2

Evaluation results

  • ROUGE-L F1 (base model 0.1521) on MALOBA Lingala instruction test set (964 examples, internal)
    self-reported
    0.238
  • ROUGE-1 F1 on MALOBA Lingala instruction test set (964 examples, internal)
    self-reported
    0.294
  • ROUGE-2 F1 on MALOBA Lingala instruction test set (964 examples, internal)
    self-reported
    0.079
  • Perplexity on MALOBA Lingala instruction test set (964 examples, internal)
    self-reported
    17.721
  • Mean token accuracy (teacher-forced) on MALOBA Lingala instruction test set (964 examples, internal)
    self-reported
    0.712
  • BLEU (0–1 scale) on MALOBA Lingala instruction test set (964 examples, internal)
    self-reported
    0.046
  • BLEU (0–100) on MALOBA translation register, test split (80 examples, internal)
    self-reported
    1.270
  • chrF on MALOBA translation register, test split (80 examples, internal)
    self-reported
    29.610