dlm-jev โ€” support-ticket triage LoRA (synthetic, Korean)

A LoRA adapter for dlm-jev, a Jev-style typed decision engine on DiffusionGemma 26B-A4B (MLX 4-bit). It teaches the model one triage policy for Korean customer-support tickets, answered as four typed questions in a single decoder pass:

question type labels
category choice ๊ฒฐ์ œยท์ฒญ๊ตฌ, ๋ฒ„๊ทธ, ๊ณ„์ • ์ ‘๊ทผ, ๊ธฐ๋Šฅ ์š”์ฒญ, ์‚ฌ์šฉ ๋ฐฉ๋ฒ•, ์„œ๋น„์Šค ์žฅ์• 
urgency score ๋‚ฎ์Œ, ๋ณดํ†ต, ๋†’์Œ, ๊ธด๊ธ‰
churn noul the customer states intent to cancel / leave
engineering noul needs a code or infra fix (bugs, outages only)

Not affiliated with TypeSafe AI or Jev. "Jev" describes the interface style only.

Training

  • LoRA on decoder attention q_proj/v_proj (the encoder reuses the same layers), rank 8, scale 20, 2.8M parameters. Adam, lr 5e-5, 400 tickets, 1 epoch, batch 1 โ€” about 18 minutes on an M4 Pro (48GB), peak memory 17.5GB.
  • Loss: cross-entropy of the label-restricted distribution at each answer slot โ€” a proper scoring rule on exactly the distribution the reader serves. The encoder cache is computed outside the gradient.
  • Data: 100% synthetic (dlm_jev.synth, seed 7). Ticket attributes are sampled first and labels are derived from them, so ground truth is exact; text comes from Korean phrase banks. No real customer data.

Evaluation (raw, before temperature calibration)

Same-generator test set (300 tickets, disjoint seed):

question base + LoRA
category 0.930 1.000
urgency 0.787 0.890
churn 1.000 1.000
engineering 0.950 1.000

Because train and test share the generator, the table above can reflect template memorisation. The distribution-shift set (300 tickets, same policy, none of the training phrasing) is the better estimate:

question base + LoRA base-only / LoRA-only correct McNemar p
urgency 0.757 0.853 11 / 40 <0.001
engineering 0.953 0.997 1 / 14 0.001
category 0.953 0.920 19 / 9 0.09
churn 1.000 0.997 1 / 0 1.0

The policy parts (urgency rubric, engineering routing) transfer to unseen wording; category drifts slightly toward the templates. Outside the domain there is some forgetting (100 items each, not significance-tested): SST-2 0.93 โ†’ 0.94, AG News 0.79 โ†’ 0.75, Yelp 0.50 โ†’ 0.45.

Usage

git clone https://github.com/seongyeon1/dlm-jev && cd dlm-jev && uv sync
hf download seongyeon1/dlm-jev-tickets-lora tickets-lora.safetensors --local-dir adapters
DLM_JEV_ADAPTER=adapters/tickets-lora.safetensors uv run dlm-jev-serve
from dlm_jev.engine import DiffusionReader
from dlm_jev.synth import QUESTIONS
from dlm_jev.schema import DecideRequest

reader = DiffusionReader(adapter_path="adapters/tickets-lora.safetensors")
res = reader.decide(DecideRequest(state="๊ด€๋ฆฌ์ž ๊ณ„์ •์œผ๋กœ ๋กœ๊ทธ์ธ์ด ์•ˆ ๋ฉ๋‹ˆ๋‹ค. ์žฌ์„ค์ • ๋ฉ”์ผ๋„ ์•ˆ ์™€์š”.", questions=QUESTIONS))

Use the same four questions (dlm_jev.synth.QUESTIONS): the policy text lives in the question instructions and the adapter was trained with them.

Limitations

  • Synthetic phrasing is narrower than real tickets; validate on your own labelled tickets before relying on it.
  • Tested with mlx-vlm 0.7.4. Loading applies LoRA wrappers to the DiffusionGemma model class in that version.
  • Refit calibration temperatures on the adapter's own reads; the base model's temperatures do not carry over.

License

Apache-2.0, same as the base model.

Downloads last month

-

Downloads are not tracked for this model. How to track
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for seongyeon1/dlm-jev-tickets-lora