Qwen3.5-4B-Jev

A locally fine-tuned, prefill-only decision model: state + typed questions → Choice, Score, and Noul.

This release contains the LoRA adapter, FP32 scalar decision head, processor, calibrated temperatures, and standalone inference code. It loads the pinned official Qwen3.5-4B base in NF4 4-bit with double quantization and BF16 computation. The base weights are downloaded separately.

jev-benchmark: 99/100 reference agreement, 10/10 context-contrast pairs passed. See the evaluation report and raw results. Results reflect this specific 100-item Chinese implicit-intent diagnostic, not general-purpose accuracy.

This is an independent Jev-like model, not an official TypeSafe model or a reproduction of its weights.

How it works

state + instructions + complete criteria + candidate
  → Qwen3.5-4B (frozen NF4 base + language LoRA)
  → last valid hidden state → FP32 scalar head
  → per-question softmax / calibrated temperature
  → typed response assembled in code
  • Choice: candidate distribution and argmax.
  • Score: ordered-level distribution and expected level, starting at 0.
  • Noul: probability of the true interpretation from false/true candidate scoring.
  • Confidence: 1 − H(p)/log(K), a concentration statistic, not a calibrated probability of correctness.
  • No response-token decoding, generated JSON, or self-reported numeric probabilities.

The complete criteria are included in each candidate input. This differs from LLM2Jev's independent yes/no scoring and normalization. The typed API design follows TypeSafe's primitives.

Quick start

Validated on Linux aarch64 / NVIDIA GB10 / CUDA 13 / Python 3.12. Use a CUDA-enabled PyTorch build compatible with your platform, then install requirements.txt. The pinned inference stack is Transformers 5.15.0, PEFT 0.20.0, bitsandbytes 0.50.2 and FLA 0.5.2. Other GPU/platform combinations have not been validated for this release.

hf download xuhaodev/Qwen3.5-4B-Jev --local-dir Qwen3.5-4B-Jev
python -m pip install -r Qwen3.5-4B-Jev/requirements.txt
import sys
sys.path.insert(0, "Qwen3.5-4B-Jev")
from qwen_jev import JevModel

model = JevModel.from_pretrained("Qwen3.5-4B-Jev")
# Optional: base_path="/path/to/the/pinned/Qwen3.5-4B"

result = model.predict(
    state="订单已经付款。仓库明确记录:尚未发货。",
    questions={
        "status": {
            "type": "choice",
            "instructions": "订单的发货状态是什么?",
            "criteria": {"pending": "尚未发货", "shipped": "已经发货"},
        },
        "paid": {"type": "noul", "instructions": "订单是否已经付款?"},
        "progress": {
            "type": "score",
            "instructions": "按订单进度判级。",
            "criteria": ["未付款", "已付款未发货", "已发货"],
        },
    },
)
print(result)

Use this loader rather than a generic AutoModelForCausalLM pipeline. The adapter is attached to the multimodal backbone and needs the scalar head and matching template to reproduce the evaluated model. No trust_remote_code auto-execution is required: the downloaded Python modules are explicitly imported. For reproducibility, pass a fixed Hub revision when downloading.

The API model identifier remains qwen35-4b-jev-v1; the public repository name is Qwen3.5-4B-Jev. String and JSON text states are supported. This release validates text only.

Training

Setting Value
Base Qwen/Qwen3.5-4B
Base revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a
Quantization NF4 4-bit, double quantization, BF16 compute
Adapter language LoRA rank 16, alpha 32, dropout 0.05
Trainable parameters 32,467,456
Training questions 6,000: 2,449 semantic references + 3,551 program-verifiable rules
Dev / calibration / internal test 600 / 600 / 600 questions, grouped source separation
Training 2 epochs, 750 optimizer updates, accumulation 16 questions
Learning rate LoRA 1e-5, scalar head 5e-5
Objective candidate CE; Score ordinal loss; verified equivalent-view consistency
Selected checkpoint epoch 2, Dev family × primitive macro NLL

Semantic data were recovered from historical reviewed synthetic training records produced using GPT-6 Luna, Grok 4.7, and GPT-6 Sol. Labels are teacher references, not new human annotations. Program-rule labels are computed from generated facts, with balanced family/label sampling. Source groups are weighted inversely to their question count. Related views stay within the same split. Raw training records are not distributed in this model repository.

The public benchmark was excluded from training, checkpoint selection and calibration. The release was hash-locked before a single 100-request final benchmark run; no weights or calibration were adjusted after seeing that result. A state-only exact-hash exclusion check and cross-split semantic near-duplicate audit were performed. Public benchmark exposure during base pretraining cannot be ruled out.

Evaluation

Benchmark: haxudev/jev-benchmark, v1.0.0, commit d6308af. Run completed 2026-10-01 21:40 UTC / 2026-10-02 local time.

Metric Result
Overall reference agreement 99/100 (99%)
Marriage subset 50/50
Girlfriend-hint subset 49/50
Same-final-utterance context pairs (both correct) 10/10
Mean / P50 / P95 full-decision latency 635 / 634 / 650 ms

Latency is from one sequential, loaded-model, loopback HTTP run on GB10, including SDK overhead; P95 uses linear interpolation. It is not a concurrent-throughput or cold-start measurement. Only love-073 differs from the reference: predicted B (0.51463), reference D (0.47288).

The independently held-out internal test achieved 99.0% agreement over 600 questions: 97.93% on 290 semantic teacher references and 100% on 310 program-rule items. Overall calibrated NLL=0.02248 and Brier=0.01135. Program families share templates across splits; these numbers do not establish unseen-template or broad task generalization.

Temperatures were fitted on the separate design-distribution calibration set: Choice=1.91865, Noul=2.37021, Score=2.02511. Calibration improved internal NLL/Brier but did not improve every metric (ECE increased from 0.00673 to 0.00963).

Scope and limits

  • Per-candidate input budget: 4,096 tokens; per-question expanded budget: 32,768 tokens.
  • Training main inputs were short (maximum candidate length 671 tokens). The 4K limit is an execution budget, not comprehensive long-context quality validation.
  • Nine ~7,800-token controlled padded probes passed; this does not establish general 8K capability.
  • Vision structure is retained, but image decision quality is not validated in this text-first release.
  • Normalized-entropy confidence is not official Jev confidence equivalence or a correctness guarantee.
  • The 100-item benchmark is synthetic, single-author and domain-specific; 99% is reference agreement.
  • Original adapter/head/calibration hashes are recorded in provenance.json; public packaging changes only portable configuration paths and inference imports.

License and references

Model adapter, decision head and included inference code: Apache-2.0. Base model: Qwen3.5-4B, Apache-2.0. The benchmark dataset is CC BY 4.0, authored by haxudev; evaluation outputs and links are included with attribution, without redistributing the full dataset.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for xuhaodev/Qwen3.5-4B-Jev

Finetuned
Qwen/Qwen3.5-4B
Adapter
(718)
this model