Instructions to use burkimbia/BIA-MISTRAL-7B-SACHI_merged with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use burkimbia/BIA-MISTRAL-7B-SACHI_merged with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "translation" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # 'pip install "transformers<5.0.0' from transformers import pipeline pipe = pipeline("translation", model="burkimbia/BIA-MISTRAL-7B-SACHI_merged")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("burkimbia/BIA-MISTRAL-7B-SACHI_merged") model = AutoModelForCausalLM.from_pretrained("burkimbia/BIA-MISTRAL-7B-SACHI_merged", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Unsloth Desktop
BIA-MISTRAL-7B-SACHI (merged, bf16)
French ↔ Mooré translation model from BurkimbIA,
a non-profit association in Ouagadougou building NLP tools for the languages of
Burkina Faso. Mooré (ISO 639-3 mos) is spoken by roughly 8 million people and
has very little digital text, which shapes every design choice here.
This is the merged bf16 checkpoint: the one to serve with vLLM.
Which repository do I want?
The SaChi Mistral fine-tune is published three times, from the same training run:
| Repository | Contents | Use it for |
|---|---|---|
BIA-MISTRAL-7B-SACHI |
LoRA adapter (~100 MB) | applying the fine-tune on top of your own base copy |
BIA-MISTRAL-7B-SACHI_merged (this repo) |
full weights, bf16, 3 shards, 14.5 GB | vLLM, TGI, any fp16/bf16 server |
BIA-MISTRAL-7B-SACHI_4bit |
full weights, bitsandbytes 4-bit, 7.2 GB | transformers on a small GPU |
_merged and _4bit are two saves of the same run, five hours apart: every
tokenizer file is byte-identical between them, only config.json differs by the
quantization_config block bitsandbytes adds.
Do not serve _4bit with vLLM. vLLM handles pre-quantized bitsandbytes
checkpoints poorly; use this repository instead and let the server do its own
memory management.
Access
This repository is gated: request access above, then pass an approved token
to whatever loads it (HF_TOKEN in the environment, --hf-token for
vllm serve, token= for from_pretrained). Without it the download fails
with 401 Unauthorized on config.json, and a server such as vLLM exits during
startup with no other diagnostic.
Prompt format
The model was fine-tuned on a raw completion prompt, not on a chat template.
Sending it through a /chat/completions route wraps it in a template it has
never seen and degrades the output. Use the completion route with this exact
string:
<s>
You are an expert Moore translator. Translate the provided {SRC} text to {TGT}.
The Moore alphabet is: a, ã, b, d, e, ẽ, ɛ, f, g, h, i, ĩ, ɩ, k, l, m, n, o, õ, p, r, s, t, u, ũ, ʋ, v, w, y, z.
Based on source language ({SRC}), provide the {TGT} text.
[INST]
### {SRC}:
{TEXT}
[/INST]
### {TGT}:
{SRC} and {TGT} are the English words French and Moore, in either
direction. The leading blank line and the <s> are part of the training format;
keep them.
Usage with vLLM
Serve it:
vllm serve burkimbia/BIA-MISTRAL-7B-SACHI_merged \
--served-model-name bia-mistral \
--dtype bfloat16 \
--max-model-len 2048 \
--gpu-memory-utilization 0.92 \
--enable-prefix-caching
--enable-prefix-caching is worth setting: the prompt above has a fixed
~120-token header shared by every request, so the cache pays for itself
immediately (measured hit rate around 50% on mixed traffic).
Call it through the OpenAI-compatible completions route:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
PROMPT = """
<s>
You are an expert Moore translator. Translate the provided {src} text to {tgt}.
The Moore alphabet is: a, ã, b, d, e, ẽ, ɛ, f, g, h, i, ĩ, ɩ, k, l, m, n, o, õ, p, r, s, t, u, ũ, ʋ, v, w, y, z.
Based on source language ({src}), provide the {tgt} text.
[INST]
### {src}:
{text}
[/INST]
### {tgt}:
"""
def translate(texts, src="French", tgt="Moore"):
if isinstance(texts, str):
texts = [texts]
resp = client.completions.create(
model="bia-mistral",
prompt=[PROMPT.format(src=src, tgt=tgt, text=t) for t in texts],
temperature=0, # the model was trained for greedy decoding
max_tokens=512,
)
return [c.text.strip() for c in sorted(resp.choices, key=lambda c: c.index)]
print(translate("Le marché est fermé aujourd'hui."))
# ['Raagã pagame rũndã.']
Pass a list of prompts to translate a batch in one call: vLLM runs them
together and returns choices in request order. Measured on an RTX 4090, three
sentences come back in the same wall-clock time as one.
Decoding: temperature=0. The model is a translator, not a chat assistant, and
sampling only adds noise.
Usage with transformers
If you would rather run it directly, prefer
_4bit on a
smaller GPU. With this repository:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "burkimbia/BIA-MISTRAL-7B-SACHI_merged"
tok = AutoTokenizer.from_pretrained(repo)
tok.padding_side, tok.pad_token = "left", tok.eos_token
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype=torch.bfloat16, device_map="auto")
inputs = tok(PROMPT.format(src="French", tgt="Moore", text="L'eau est froide."), return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=256, do_sample=False)
print(tok.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=True).strip())
Deployment on RunPod Serverless
Runs unchanged on the stock runpod/worker-v1-vllm image, no custom handler:
MODEL_NAME burkimbia/BIA-MISTRAL-7B-SACHI_merged
DTYPE bfloat16
MAX_MODEL_LEN 2048
GPU_MEMORY_UTILIZATION 0.92
ENABLE_PREFIX_CACHING true
OPENAI_SERVED_MODEL_NAME_OVERRIDE bia-mistral
Needs a 24 GB card or larger (14.5 GB of weights plus KV cache). Endpoint route:
https://api.runpod.ai/v2/<endpoint-id>/openai/v1.
Hardware
| Precision | Weights | Minimum GPU |
|---|---|---|
| bf16 (this repo) | 14.5 GB | 24 GB (RTX 3090/4090, A5000, L4, A40) |
4-bit (_4bit) |
7.2 GB | 12 GB |
Related models
burkimbia/BIA-NLLB-600M-11E— much smaller seq2seq translator, better latency, lower quality on long sentencesburkimbia/BIA-WHISPER-LARGE-SACHI_V2— Mooré speech recognitionburkimbia/BIA-SPARKTTS-V4— Mooré text to speech
Benchmark and leaderboard: burkimbia/mt-benchmark-public,
scored on BLEU, METEOR and chrF across five themes (administrative, daily life,
health, religious, arts and technology).
Limitations
- Low-resource. Mooré has little digital text. Output can be disfluent or wrong, and degrades on long, technical or out-of-domain sentences.
- Tone is not written in Mooré orthography, so homographs are frequent and the model can pick the wrong sense.
- The model follows the prompt above closely; changing its wording, the
alphabet line or the
### Moore:marker changes the output quality. - Outputs should be reviewed by a Mooré speaker before any consequential use, in particular for administrative, legal or medical content.
Citation
@misc{burkimbia_sachi_mistral,
title = {BIA-MISTRAL-7B-SACHI: French-Moore translation},
author = {BurkimbIA},
year = {2025},
url = {https://huggingface.co/burkimbia/BIA-MISTRAL-7B-SACHI_merged}
}
- Downloads last month
- 390
Model tree for burkimbia/BIA-MISTRAL-7B-SACHI_merged
Base model
unsloth/mistral-7b-bnb-4bit