Instructions to use seongyeon1/dlm-jev-tickets-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use seongyeon1/dlm-jev-tickets-lora with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir dlm-jev-tickets-lora seongyeon1/dlm-jev-tickets-lora
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
dlm-jev โ support-ticket triage LoRA (synthetic, Korean)
A LoRA adapter for dlm-jev, a Jev-style typed decision engine on DiffusionGemma 26B-A4B (MLX 4-bit). It teaches the model one triage policy for Korean customer-support tickets, answered as four typed questions in a single decoder pass:
| question | type | labels |
|---|---|---|
category |
choice | ๊ฒฐ์ ยท์ฒญ๊ตฌ, ๋ฒ๊ทธ, ๊ณ์ ์ ๊ทผ, ๊ธฐ๋ฅ ์์ฒญ, ์ฌ์ฉ ๋ฐฉ๋ฒ, ์๋น์ค ์ฅ์ |
urgency |
score | ๋ฎ์, ๋ณดํต, ๋์, ๊ธด๊ธ |
churn |
noul | the customer states intent to cancel / leave |
engineering |
noul | needs a code or infra fix (bugs, outages only) |
Not affiliated with TypeSafe AI or Jev. "Jev" describes the interface style only.
Training
- LoRA on decoder attention
q_proj/v_proj(the encoder reuses the same layers), rank 8, scale 20, 2.8M parameters. Adam, lr 5e-5, 400 tickets, 1 epoch, batch 1 โ about 18 minutes on an M4 Pro (48GB), peak memory 17.5GB. - Loss: cross-entropy of the label-restricted distribution at each answer slot โ a proper scoring rule on exactly the distribution the reader serves. The encoder cache is computed outside the gradient.
- Data: 100% synthetic (
dlm_jev.synth, seed 7). Ticket attributes are sampled first and labels are derived from them, so ground truth is exact; text comes from Korean phrase banks. No real customer data.
Evaluation (raw, before temperature calibration)
Same-generator test set (300 tickets, disjoint seed):
| question | base | + LoRA |
|---|---|---|
| category | 0.930 | 1.000 |
| urgency | 0.787 | 0.890 |
| churn | 1.000 | 1.000 |
| engineering | 0.950 | 1.000 |
Because train and test share the generator, the table above can reflect template memorisation. The distribution-shift set (300 tickets, same policy, none of the training phrasing) is the better estimate:
| question | base | + LoRA | base-only / LoRA-only correct | McNemar p |
|---|---|---|---|---|
| urgency | 0.757 | 0.853 | 11 / 40 | <0.001 |
| engineering | 0.953 | 0.997 | 1 / 14 | 0.001 |
| category | 0.953 | 0.920 | 19 / 9 | 0.09 |
| churn | 1.000 | 0.997 | 1 / 0 | 1.0 |
The policy parts (urgency rubric, engineering routing) transfer to unseen wording; category drifts slightly toward the templates. Outside the domain there is some forgetting (100 items each, not significance-tested): SST-2 0.93 โ 0.94, AG News 0.79 โ 0.75, Yelp 0.50 โ 0.45.
Usage
git clone https://github.com/seongyeon1/dlm-jev && cd dlm-jev && uv sync
hf download seongyeon1/dlm-jev-tickets-lora tickets-lora.safetensors --local-dir adapters
DLM_JEV_ADAPTER=adapters/tickets-lora.safetensors uv run dlm-jev-serve
from dlm_jev.engine import DiffusionReader
from dlm_jev.synth import QUESTIONS
from dlm_jev.schema import DecideRequest
reader = DiffusionReader(adapter_path="adapters/tickets-lora.safetensors")
res = reader.decide(DecideRequest(state="๊ด๋ฆฌ์ ๊ณ์ ์ผ๋ก ๋ก๊ทธ์ธ์ด ์ ๋ฉ๋๋ค. ์ฌ์ค์ ๋ฉ์ผ๋ ์ ์์.", questions=QUESTIONS))
Use the same four questions (dlm_jev.synth.QUESTIONS): the policy text lives in the question instructions and
the adapter was trained with them.
Limitations
- Synthetic phrasing is narrower than real tickets; validate on your own labelled tickets before relying on it.
- Tested with
mlx-vlm0.7.4. Loading applies LoRA wrappers to the DiffusionGemma model class in that version. - Refit calibration temperatures on the adapter's own reads; the base model's temperatures do not carry over.
License
Apache-2.0, same as the base model.
Quantized
Model tree for seongyeon1/dlm-jev-tickets-lora
Base model
google/diffusiongemma-26B-A4B-it