llm-sft-dpo-lora-repo20260214 (SFT + DPO)

This model is fine-tuned for Structured Output (JSON/XML/TOML) using SFT + DPO.

Methods

  1. SFT:
    • Cleaned dataset (removed markdown blocks, removed <think> tags).
    • Upsampled weak formats (TOML, XML).
    • Normalized TOML syntax.
  2. DPO:
    • Rejected samples included <think> tags, Markdown blocks, inline TOML, and CSV trailing commas.

Usage

Use with Unsloth or PEFT.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for centmount/llm-sft-dpo-lora-repo20260214

Finetuned
(2142)
this model