Text Generation
MLX
PEFT
Nogai
Russian
lora
sft
nogai
turkic
translation
apple-silicon

Qwen2.5-1.5B-Nogai-SFT-Experimental (Phase 2: instruction recovery)

An mlx-lm LoRA adapter trained on Russian↔Nogai translation instructions. It sits on top of the Phase 1 model (Qwen2.5-1.5B-Nogai-LoRA fused into Qwen2.5-1.5B-Instruct). Phase 1 teaches Nogai but breaks chat-following; this adapter brings the ChatML format back. Part of NogaiLLM.

Experimental. It is trained on Bible text only, so modern-domain input (medicine, technology) often comes back with religious vocabulary instead of a translation.

Training (from adapter_config.json)

Base Qwen2.5-1.5B-Instruct with the Phase 1 adapter fused in
Data Nogai-Russian-SFT-Biblical-v1 (650 unique pairs, both directions)
LoRA rank 8, scale 20, dropout 0; q/k/v/o/gate/up/down of layers 12–27 (5.28 M parameters)
Batch / iterations / learning rate 1 / 2,400 / 2e-5, seed 0
Max sequence length 512 tokens
Hardware Apple M2 Pro, 16 GB, mlx-lm; peak memory 5.73 GB

The run crashed with a Metal Internal Error and was resumed from its last saved adapter (resume_adapter_file in the config). clear_cache_threshold was 0 and gradient checkpointing was off.

Evaluation status

The v1 dataset's validation rows also occur in its training rows. The validation/test losses from this run therefore don't measure generalisation and are not reported here. A clean re-evaluation is in progress: the v2 split, chrF++/BLEU against human translations, 3 seeds. Numbers will be added when done.

Usage (Apple Silicon / MLX)

This adapter needs the Phase 1 weights. It does not work correctly on the raw Qwen2.5 model.

pip install mlx-lm
# 1. fuse Phase 1 into the base
mlx_lm.fuse --model Qwen/Qwen2.5-1.5B-Instruct \
    --adapter-path ansarzeinulla/Qwen2.5-1.5B-Nogai-LoRA \
    --save-path local_qwen_1.5B_Nogai_Base
# 2. chat with Phase 2 on top
mlx_lm.chat --model local_qwen_1.5B_Nogai_Base \
    --adapter-path ansarzeinulla/Qwen2.5-1.5B-Nogai-SFT-Experimental --temp 0.3

Use the training prompt format, e.g. Переведи этот текст на русский язык: <Nogai text>.

For PyTorch/PEFT, see the Space code. It converts both adapters (all seven target modules, layers 12–27), merges Phase 1, then applies Phase 2.

Citation

@misc{zeinulla2026nogaillm,
  title  = {NogaiLLM: Parameter-Efficient Continued Pre-Training and Catastrophic Forgetting in Zero-Resource Turkic Languages},
  author = {Zeinulla, Ansar},
  year   = {2026},
  note   = {Manuscript under revision. Code: https://github.com/ansarzeinulla/NogaiLLM-Apple-Silicon}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ansarzeinulla/Qwen2.5-1.5B-Nogai-SFT-Experimental

Adapter
(1463)
this model

Space using ansarzeinulla/Qwen2.5-1.5B-Nogai-SFT-Experimental 1