Instructions to use centmount/llm-sft-dpo-lora-repo20260214 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Local Apps Settings
- Unsloth Desktop
llm-sft-dpo-lora-repo20260214 (SFT + DPO)
This model is fine-tuned for Structured Output (JSON/XML/TOML) using SFT + DPO.
Methods
- SFT:
- Cleaned dataset (removed markdown blocks, removed
<think>tags). - Upsampled weak formats (TOML, XML).
- Normalized TOML syntax.
- Cleaned dataset (removed markdown blocks, removed
- DPO:
- Rejected samples included
<think>tags, Markdown blocks, inline TOML, and CSV trailing commas.
- Rejected samples included
Usage
Use with Unsloth or PEFT.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Model tree for centmount/llm-sft-dpo-lora-repo20260214
Base model
Qwen/Qwen3-4B-Instruct-2507