klev-0.8b

klev (KV + Jev) is a 0.8B decision model: given a typed question β€” a single choice, a noul (true/false) judgement, a score, or a multi-choice item β€” it returns calibrated probabilities over the options plus an explicit rejection channel. It is built on Qwen/Qwen3.5-0.8B, fine-tuned with QLoRA, and read out through a small pointer head instead of text generation. It runs 4-bit in 0.89 GB of weights, 1.39 GB peak β€” a full 2-epoch decision-v7 run finished in 192 min on an 8 GB AMD Radeon RX 6600M (gfx1032, ROCm 7), a card the 4B arm cannot fit at all.

It was built using Unsloth: the 4-bit QLoRA fine-tune, the delimiter embedding deltas, the pointer head and the teacher-KL cache all run through Unsloth's patched Qwen 3.5 stack (FastModel). See "Built with Unsloth" below.

The name comes from KV + Jev: a Kev-lineage pointer readout serving Jev's typed decision protocol. Working name during development was M5.

Built on Kev

klev-0.8b is based on Kev and reuses part of its code. The record format, the pointer readout, the metrics and the decision-v7 data are Kev's (Apache-2.0, Jared Palmer), vendored rather than reimplemented so that klev stays byte-compatible with the Kev family β€” which is what makes the kev-0.8B column in Results a like-for-like baseline:

from Kev reused here
kev/model.py (PointerHead) model/head.py β€” the q/k readout below, plus klev's own garbage candidate
kev/api.py, kev/data.py data/format.py β€” TypeSafe record format, render(), record builder
kev/metrics.py data/metrics.py β€” the scorer used for every number in Results (verbatim)
kev/suite.py + jaredpalmer/kev-suites data/suites.py β€” frozen, sha256-checked suites
kev.calibrate scripts/calibrate.py β€” the fitted pointer temperature in head.pt
decision-v7 the 15,576 training rows (the same partitions Kev-0.8B/4B/9B trained on)
<|fim_*|> reserved tokens klev's delimiter set (see Architecture) β€” Qwen has no <unusedN> slots, so these are the <|fim_prefix|>-style slots Kev reuses

Kev is Apache-2.0 and every vendored file carries a header naming its origin. No Kev model weights are used: the base is Qwen3.5-0.8B, the head and LoRA are trained fresh. What is klev's own is the base choice, the Unsloth QLoRA training stack, the KL anchor, the delimiter embedding deltas, the rejection channel and the stitch below.

Built with Unsloth

klev was built using Unsloth (2026.9.11 with transformers 5.6.2 β€” Gemma 4 and Qwen 3.5 need transformers β‰₯5.5.0 for the architecture, and 5.6.2 is the revision that checkpoint was saved with).

stage Unsloth usage
base load unsloth.FastModel.from_pretrained(..., load_in_4bit=True)
fine-tune FastModel.get_peft_model for the rank-16 LoRA + use_gradient_checkpointing="unsloth", driven by a transformers.Trainer subclass (DistillTrainer)
trainable tokens delimiter rows via peft trainable_token_indices (5 rows, no vocabulary resize)
teacher KL cached base top-k logits, read through Unsloth's patched hidden-state/logits access
inference / evals FastModel for every decision and stitch eval

Reproducing the run needs unsloth imported before transformers / peft / trl, because it patches them at import time:

pip install unsloth            # 2026.9.11
python -c "import unsloth, transformers, peft, trl"   # unsloth first

transformers must be 5.5.0+ for the qwen3_5 architecture, and --no-deps is required when installing it alongside unsloth so the CUDA xformers dependency does not replace a ROCm torch.

Files

file what it is
adapter/ rank-16 LoRA (alpha 32) + tokenizer with the 5 delimiter special tokens
head.pt pointer head weights and the trainable delimiter-embedding deltas
train_config.json training hyperparameters (decision-v7, 2 epochs, KL anchor 0.3)
stitch/tweet-eval-lora/ external LoRA (tweet_eval hate detection) trained on the same base, alpaca format
stitch/rumoureval-lora/ external LoRA (RumourEval 2019 responding-relationship), same base and format
stitch/probe-tweet-eval.pt 32-shot nearest-class-mean probe in that LoRA's prompt space, + T, beta
stitch/probe-rumoureval.pt same, for the RumourEval LoRA (4 classes)
stitch/report-*.json full stitch reports incl. steering and fusion grids
base/ the same recipe on Qwen/Qwen3.5-0.8B-Base, the checkpoint kev-0.8B uses β€” adapter, head, config, the shared external LoRA and its fitted probe. See base/README.md.

This is not a chat model: decisions are read from the pointer head over the option/decision tokens. Reuse model/delimiters.py, model/head.py and eval/eval_decisions.py from the code repo to load and score it.

Architecture

  • Base: Qwen3.5-0.8B (dense, 1024 hidden, 24 layers, vocab 248,320, tied embeddings, no PLE embedding table), 4-bit NF4 QLoRA, loaded and trained with Unsloth (FastModel).
  • Five Qwen reserved tokens mark state / option / decision positions: <|fim_prefix|> (248060), <|fim_suffix|> (248062), <|fim_middle|> (248061), <|fim_pad|> (248063), <|repo_name|> (248064). Their embedding rows are trained as deltas, so there is no vocabulary resize. Qwen has no <bos>; the encoder asks the tokenizer for its own bos_token_id and adjusts the state-length arithmetic accordingly.
  • Pointer head (vendored from Kev, see "Built on Kev"): q/k dot product over option boundary tokens, softmax over the K options plus a learned garbage candidate for rejection.
  • No final_logit_softcapping on this base (Gemma 4 has 30.0), so the KL's softcap branch is skipped rather than approximated.
  • Trainable: 6.9M params (6.39M LoRA + 525.1K head + 5.1K deltas).

Training

  • decision-v7: 15,576 decision rows, 2 epochs, 3,894 steps, batch 1 Γ— accum 8, 192 min on one RX 6600M 8 GB.
  • Objective: pointer CE over K+1 + KL anchor 0.3 to the frozen base on decision states (top-k 32, T 1.0). Loss 3.70 β†’ 0.89; the pointer term collapses to ~0 by step 645 while the KL term holds at ~2.1–2.3, i.e. the base alignment is what moves.
  • The teacher cache is 1,413,968 content positions (~30 min to build on this card).

Two deviations from the E4B run, both forced, neither a choice

  • lr_scheduler_type is linear, not cosine. On transformers 5.6.2 the Trainer fast loop pins the learning rate at exactly 0 for every step when the scheduler is cosine and warmup_steps > 0 (verified: 6 steps at warmup 50 all report learning_rate: 0, where 1.2e-05 is correct). With cosine this recipe trains silently and does nothing.
  • Precision is fp16, not bf16. The card this was trained on is gfx1032, where Unsloth declines bf16 and the Triton bf16 path is unavailable. See scripts/rocm_env.sh in the code repo for the full set of ROCm requirements, including a hard sequence-length ceiling of ~195 tokens past which Qwen 3.5's gated-deltanet attention autotunes into a bf16 kernel gfx1030 cannot compile.

Results

Own decision suite, klev-0.8b against kev-0.8B on identical rows, each model served with a temperature fitted on decision-v7 calibration (1,148 rows) via scripts/calibrate.py.

decision-v7 development (n = 1,468, in-distribution):

klev-0.8b kev-0.8B
accuracy 0.8120 0.8120
Brier 0.2590 0.2685
ECE 0.0365 0.0274
NLL 0.4785 0.5102
coverage @ 5 % error 0.6158 0.4360
AURC 0.0592 0.0717
mean p_none (rejection) 0.016 β€”

transfer-v4 development (n = 764, six never-trained sources):

klev-0.8b kev-0.8B
accuracy 0.6636 0.6126
Brier 0.4372 0.4542
ECE 0.0207 0.0735
NLL 0.7538 0.7524
AURC 0.1719 0.1876

By question type (accuracy):

suite type klev-0.8b kev-0.8B
decision-v7 choice / noul / score 0.851 / 0.867 / 0.583 0.876 / 0.845 / 0.546
transfer-v4 choice / noul / score 0.651 / 0.736 / 0.300 0.599 / 0.671 / 0.350

Two bases are published here. The top level is Qwen/Qwen3.5-0.8B (IT); base/ is the identical recipe on Qwen/Qwen3.5-0.8B-Base, which is what kev-0.8B uses. The base matters more than it looks: it sets how strong an external task LoRA can be (tweet_eval probe 0.715 on IT vs 0.660 on -Base), and the stitch can only import what the adapter holds, so an IT-vs--Base stitch comparison measures the adapter as much as the decision model. On the shared -Base adapter the stitch gain is +7.5 pts for klev against +3.0 for kev, so about half the original +8.0 gap was the adapter's base and half is the recipe. See base/README.md and docs/m13 in the code repo.

klev-0.8b ties kev-0.8B in distribution and beats it by 5.1 points out of it, with a narrower LoRA (192 tensors / 6.4M against kev's lora_targets="all" at 372 / 11.3M) and less data (15,576 rows against kev's 22,539, which included 16,539 from a separate joint mixture).

Two caveats worth stating plainly. First, kev-0.8B's own model card reports 0.827 / 0.648, but on smaller row sets (1,264 / 656 questions); scored on these rows it measures 0.812 / 0.613, so its card is ~1.5 and ~3.5 points optimistic against this eval set β€” do not compare across row sets. Second, this is not a pure base-model ablation: the two arms use different base checkpoints (Qwen3.5-0.8B vs Qwen3.5-0.8B-Base), which matters for the stitch numbers below.

Stitch: importing an external LoRA with 32 examples

Because of the KL anchor, a LoRA trained outside this fine-tune composes with klev without wrecking its decisions. Its knowledge is linearly readable from the hidden state of its own prompt format, so fusing that readout into the pointer logits with two scalars imports it:

logits_final = logits_ptr(h_sys) + beta * log p_alp
p_alp = softmax(-d^2 / T)          # NCM in the external LoRA's prompt space

tweet_eval hate detection (200 test rows, 32 balanced shots, 2 classes, ext weight 1.0):

klev-0.8b kev-0.8B
klev/kev alone (pointer) 0.605 0.645
32-shot probe, LoRA prompt space 0.715 0.660
32-shot probe, decide state 0.565 0.705
stitched 0.715 0.675
external LoRA alone (generation) 0.730 β€”

The fused klev-0.8b equals the external LoRA's own probe, recovering 100 % of the available margin from 32 labelled rows and two scalars, while kev-0.8B gains 3.0 points. That gap is confounded, though: the two arms use different base checkpoints, and the base determines how good the external adapter can be (probe 0.715 on the IT base vs 0.660 on -Base, and 0.490 vs 0.430 on RumourEval), which caps what the stitch can import. The honest claim is that the stitch works in both, not that klev is the better stitch target.

kev-0.8B's probe, decide state of 0.705 is the more interesting single number: a linear probe on its own decide state already recovers the hate-speech signal with no external knowledge at all, so there was little left to import. The stitch's value appears to depend on the decision model not already having the capability.

RumourEval 2019 responding-relationship (351 test rows, 32 shots, 4 classes):

klev-0.8b
pointer alone 0.362
32-shot probe, LoRA prompt space 0.490
32-shot probe, decide state 0.316
stitched 0.479 (+11.7 pts)

Steering the state does not work; conditioning the readout does β€” reproducing that finding independently. Adding the LoRA's class directions to the decide state scores 0.322 against a 0.605 baseline on TweetEval, and the single global prompt direction 0.530; the gated fusion recovers everything. The knowledge is a separate linearly-readable signal in the LoRA's space, so it has to enter as a score, not as a perturbation.

Probe hyperparameters (in stitch/probe-*.pt): TweetEval T = 105.19, beta = 16.0, ext_weight = 1.0; RumourEval T = 1305.8, beta = 16.0, ext_weight = 1.0.

Install

# klev is not on PyPI: install from the repository
pip install "klev @ git+https://github.com/felipepenhorate/klev.git"

klev is the library: it carries the loader, the pointer readout, the record format and both usage paths, so a consumer does not clone this repo to use a checkpoint. Two version floors matter and pip does not know they are related to the architecture:

why
transformers >= 5.5.0 the qwen3_5 and gemma4 architectures do not exist before it. On 5.3.0 AutoConfig raises "Transformers does not recognize this architecture", which unsloth then misreports as a generic "not supported yet" ValueError β€” it reads like a missing model, not a version.
unsloth >= 2026.9.11 first release carrying gemma4 support.

Install unsloth with --no-deps if you are on ROCm, so its CUDA xformers dependency does not replace your ROCm torch build.

On AMD/ROCm also source scripts/rocm_env.sh from the repo (shipped in the wheel as klev.scripts) before importing β€” it sets the arch override, TORCH_ROCM_AOTRITON_ENABLE_BF16=0, UNSLOTH_COMPILE_DISABLE=1 and the rest. All of it is load-bearing; see docs/m11-qwen35-08b.md.

Two console scripts come with it:

klev-system-one --ckpt lumierenoir/klev-0.8b --request my_request.json   # no server
klev-serve       --ckpt lumierenoir/klev-0.8b --port 8090                # optional HTTP

load_klev() reads the base model from the checkpoint's own adapter_config.json and picks the matching delimiters, so you do not have to pass base= or preset=. Pairing a checkpoint with another model's weights would otherwise be silent nonsense rather than an error.

Usage

klev is a library first. A decision model is small and stateless, so the normal way to use one is in-process: load it once, call it as many times as you like. No server, no port, no daemon. import unsloth must come before transformers / peft (it patches them at import time), which is why it is first.

import unsloth  # noqa: F401  β€” must precede transformers / peft
from klev import load_klev, answer            # pip install "klev @ git+https://github.com/felipepenhorate/klev.git"

model, tokenizer, head = load_klev("lumierenoir/klev-0.8b")   # or a local dir with adapter/ + head.pt

result = answer(model, tokenizer, head, {
    "state": "My card was declined twice at a supermarket and the ATM refused it too.",
    "questions": {
        "is_card": {"type": "noul",   "instructions": "Is the card being declined?",
                    "criteria": {"false": "no", "true": "yes"}},
        "action":  {"type": "choice", "instructions": "What should the agent do first?",
                    "criteria": {"usage": "ask how the card was used",
                                 "atm":   "raise an ATM fault",
                                 "fraud": "run a fraud check"}},
        "urgency": {"type": "score",  "instructions": "How urgent is this?",
                    "criteria": ["no action needed", "next working day", "same day", "immediate"]},
    }},
    temperature=2.2974)

Verified output from that request:

{"answers": {
  "is_card": {"type": "noul", "noul": 0.9169, "none": 0.01314},
  "action":  {"type": "choice", "choice": "atm", "confidence": 0.4848,
              "probabilities": {"usage": 0.1311, "atm": 0.6136, "fraud": 0.0870},
              "none": 0.16828},
  "urgency": {"type": "score", "score": 1.485, "legend": {"0": "no action needed",
              "1": "next working day", "2": "same day", "3": "immediate"},
              "probabilities": {"0": 0.2936, "1": 0.2397, "2": 0.1817, "3": 0.2580, "4": 0.0270},
              "confidence": 0.6288, "none": 0.026994}}}

none is the rejection channel: the mass on "none of these apply", renormalised off the K options so the returned distribution still sums to 1.

The three question types are noul (criteria = {false, true}), choice (criteria = an option map) and score (criteria = an ordered level list). A request is validated by klev's own SystemOneRequest before any GPU work, and labels / target / _meta are ignored β€” so a labelled eval record can be passed straight through.

Temperature matters. answer(..., temperature=1.0) is the raw readout, which is what the benchmark tables above score (ECE 0.113 on decision-v7 dev). The served numbers use the per-model temperature from scripts/calibrate.py:

arm serve at
lumierenoir/klev-0.8b (IT) 2.2974
lumierenoir/klev-0.8b + /base 2.2449

head.pt stores 1.0, because it is written during training and the head only tempers in eval mode β€” so a deployment has to supply the fitted value itself. That is the one thing easy to get wrong, and it is worth 0.08 ECE.

From a shell, still no server

klev-system-one --ckpt lumierenoir/klev-0.8b                    # built-in demo
klev-system-one --ckpt lumierenoir/klev-0.8b --request r.json   # your request
echo '{"state":"...","questions":{...}}' | klev-system-one --ckpt ... --request -

which ends with a plain-language readback β€” the quickest way to sanity-check a checkpoint. The built-in demo request (a card decline, with 2 options and a 4-level severity scale):

channel_issue  noul    p(true)=0.885  -> yes   reject=0.015
next_step      choice  -> atm_fault   conf=0.371  reject=0.193
severity       score   -> 1.52        reject=0.025

If you would rather have HTTP

Worth it when several processes share one warm copy of the weights, or when the caller is not Python. Otherwise it adds a port and a failure mode to save ~1 GB of load time.

klev-serve --ckpt lumierenoir/klev-0.8b --port 8090 --temperature 2.2974
curl -s localhost:8090/v1/systemone -H 'content-type: application/json' -d @request.json
route
POST /v1/systemone {"state", "questions"} β†’ {"answers": {qid: typed answer}}
GET /v1/models loaded base, temperature, whether the rejection channel exists
GET /health {"ok": true, "model": ...}

Same routes and response shapes as kev's own server, so eval/eval_systemone_http.py scores it on the bench records unmodified β€” verified here over 120 decision-v7 development rows with zero failures. It is stdlib http.server on purpose: no fastapi/uvicorn to install. One

Identical to lumierenoir/klev-e4b apart from the base and the temperature. decision per request under a lock, since a forward is not re-entrant.

Validated locally: 703-token decision-v7 row served 200 with the correct answer, malformed bodies 400, unknown routes 404, and rows over --max-seq 400 naming the token count.

License and data

  • Base model: Apache-2.0 (Qwen/Qwen3.5-0.8B).
  • This release (adapter, head, probes): Apache-2.0.
  • Kev (github.com/jaredpalmer/kev, Apache-2.0, Jared Palmer) β€” klev is built on it and reuses its record format, PointerHead, metrics, suite loading, calibration procedure and decision-v7 data. See "Built on Kev" above.
  • Unsloth (github.com/unslothai/unsloth, Apache-2.0) β€” the training and inference stack this model was built with. See "Built with Unsloth" above.
  • Training data: decision-v7 records built from permissively licensed sets (Apache-2.0 / MIT / CC-BY-4.0), attribution manifest in the code repo.
  • stitch/*-lora are not klev's: they are external task adapters trained for this release on cardiffnlp/tweet_eval (hate; a standard public benchmark that does contain slurs) and strombergnlp/rumoureval_2019. Check those datasets' terms before redistribution.

Code

Architecture, training, evals and the stitch experiment: github.com/felipepenhorate/klev (SPEC.md, docs/m10–m12, eval/eval_lora_stitch.py). The base model is a preset (--preset qwen35-08b / e4b / e2b-qat); see docs/m12-presets.md.

Predictions do not need unsloth

Unsloth is a training dependency. The default install is torch + transformers + peft + bitsandbytes, and a published checkpoint is served through plain transformers β€” same accuracy (0.410 vs 0.410 on 195 CoSt-BR rows for klev-e4b, 0.317 vs 0.317 for klev-0.8b), 98 % identical choices. Unsloth is the train extra:

pip install "klev @ git+https://github.com/felipepenhorate/klev.git"            # predictions
pip install "klev[train] @ git+https://github.com/felipepenhorate/klev.git"     # + fine-tuning

Backend is explicit when you want the unsloth path: load_klev(ckpt, backend="unsloth"), KLEV_BACKEND=unsloth, or --backend unsloth on klev-system-one / klev-serve. See docs/m15-plain-inference.md.

Same checkpoint, two loaders

The numbers above are the unsloth path (--backend unsloth, the stack this model was trained with). klev can also be served with plain transformers + bitsandbytes and no unsloth at all β€” the default install. Measured parity for this arm: 58 CoSt-BR rows, 0.317 unsloth vs 0.317 plain, 98.3 % identical choices; the full nine-suite parity (max Β±0.4 pp on any suite) was measured on the 4B arm and is in klev-e4b's card and docs/m15-plain-inference.md.

pip install "klev @ git+https://github.com/felipepenhorate/klev.git"             # plain (default)
pip install "klev[train] @ git+https://github.com/felipepenhorate/klev.git"      # + unsloth
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for lumierenoir/klev-0.8b

Adapter
(285)
this model