Instructions to use xuhaodev/Qwen3.5-4B-Jev with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use xuhaodev/Qwen3.5-4B-Jev with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Qwen3.5-4B-Jev
A locally fine-tuned, prefill-only decision model: state + typed questions → Choice, Score, and Noul.
This release contains the LoRA adapter, FP32 scalar decision head, processor, calibrated temperatures, and standalone inference code. It loads the pinned official Qwen3.5-4B base in NF4 4-bit with double quantization and BF16 computation. The base weights are downloaded separately.
jev-benchmark: 99/100 reference agreement, 10/10 context-contrast pairs passed. See the evaluation report and raw results. Results reflect this specific 100-item Chinese implicit-intent diagnostic, not general-purpose accuracy.
This is an independent Jev-like model, not an official TypeSafe model or a reproduction of its weights.
How it works
state + instructions + complete criteria + candidate
→ Qwen3.5-4B (frozen NF4 base + language LoRA)
→ last valid hidden state → FP32 scalar head
→ per-question softmax / calibrated temperature
→ typed response assembled in code
- Choice: candidate distribution and argmax.
- Score: ordered-level distribution and expected level, starting at 0.
- Noul: probability of the true interpretation from false/true candidate scoring.
- Confidence:
1 − H(p)/log(K), a concentration statistic, not a calibrated probability of correctness. - No response-token decoding, generated JSON, or self-reported numeric probabilities.
The complete criteria are included in each candidate input. This differs from LLM2Jev's independent yes/no scoring and normalization. The typed API design follows TypeSafe's primitives.
Quick start
Validated on Linux aarch64 / NVIDIA GB10 / CUDA 13 / Python 3.12.
Use a CUDA-enabled PyTorch build compatible with your platform, then install requirements.txt.
The pinned inference stack is Transformers 5.15.0, PEFT 0.20.0, bitsandbytes 0.50.2 and FLA 0.5.2.
Other GPU/platform combinations have not been validated for this release.
hf download xuhaodev/Qwen3.5-4B-Jev --local-dir Qwen3.5-4B-Jev
python -m pip install -r Qwen3.5-4B-Jev/requirements.txt
import sys
sys.path.insert(0, "Qwen3.5-4B-Jev")
from qwen_jev import JevModel
model = JevModel.from_pretrained("Qwen3.5-4B-Jev")
# Optional: base_path="/path/to/the/pinned/Qwen3.5-4B"
result = model.predict(
state="订单已经付款。仓库明确记录:尚未发货。",
questions={
"status": {
"type": "choice",
"instructions": "订单的发货状态是什么?",
"criteria": {"pending": "尚未发货", "shipped": "已经发货"},
},
"paid": {"type": "noul", "instructions": "订单是否已经付款?"},
"progress": {
"type": "score",
"instructions": "按订单进度判级。",
"criteria": ["未付款", "已付款未发货", "已发货"],
},
},
)
print(result)
Use this loader rather than a generic AutoModelForCausalLM pipeline. The adapter is attached to the
multimodal backbone and needs the scalar head and matching template to reproduce the evaluated model.
No trust_remote_code auto-execution is required: the downloaded Python modules are explicitly imported.
For reproducibility, pass a fixed Hub revision when downloading.
The API model identifier remains qwen35-4b-jev-v1; the public repository name is Qwen3.5-4B-Jev.
String and JSON text states are supported. This release validates text only.
Training
| Setting | Value |
|---|---|
| Base | Qwen/Qwen3.5-4B |
| Base revision | 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a |
| Quantization | NF4 4-bit, double quantization, BF16 compute |
| Adapter | language LoRA rank 16, alpha 32, dropout 0.05 |
| Trainable parameters | 32,467,456 |
| Training questions | 6,000: 2,449 semantic references + 3,551 program-verifiable rules |
| Dev / calibration / internal test | 600 / 600 / 600 questions, grouped source separation |
| Training | 2 epochs, 750 optimizer updates, accumulation 16 questions |
| Learning rate | LoRA 1e-5, scalar head 5e-5 |
| Objective | candidate CE; Score ordinal loss; verified equivalent-view consistency |
| Selected checkpoint | epoch 2, Dev family × primitive macro NLL |
Semantic data were recovered from historical reviewed synthetic training records produced using GPT-6 Luna, Grok 4.7, and GPT-6 Sol. Labels are teacher references, not new human annotations. Program-rule labels are computed from generated facts, with balanced family/label sampling. Source groups are weighted inversely to their question count. Related views stay within the same split. Raw training records are not distributed in this model repository.
The public benchmark was excluded from training, checkpoint selection and calibration. The release was hash-locked before a single 100-request final benchmark run; no weights or calibration were adjusted after seeing that result. A state-only exact-hash exclusion check and cross-split semantic near-duplicate audit were performed. Public benchmark exposure during base pretraining cannot be ruled out.
Evaluation
Benchmark: haxudev/jev-benchmark, v1.0.0,
commit d6308af.
Run completed 2026-10-01 21:40 UTC / 2026-10-02 local time.
| Metric | Result |
|---|---|
| Overall reference agreement | 99/100 (99%) |
| Marriage subset | 50/50 |
| Girlfriend-hint subset | 49/50 |
| Same-final-utterance context pairs (both correct) | 10/10 |
| Mean / P50 / P95 full-decision latency | 635 / 634 / 650 ms |
Latency is from one sequential, loaded-model, loopback HTTP run on GB10, including SDK overhead;
P95 uses linear interpolation. It is not a concurrent-throughput or cold-start measurement.
Only love-073 differs from the reference: predicted B (0.51463), reference D (0.47288).
The independently held-out internal test achieved 99.0% agreement over 600 questions: 97.93% on 290 semantic teacher references and 100% on 310 program-rule items. Overall calibrated NLL=0.02248 and Brier=0.01135. Program families share templates across splits; these numbers do not establish unseen-template or broad task generalization.
Temperatures were fitted on the separate design-distribution calibration set: Choice=1.91865, Noul=2.37021, Score=2.02511. Calibration improved internal NLL/Brier but did not improve every metric (ECE increased from 0.00673 to 0.00963).
Scope and limits
- Per-candidate input budget: 4,096 tokens; per-question expanded budget: 32,768 tokens.
- Training main inputs were short (maximum candidate length 671 tokens). The 4K limit is an execution budget, not comprehensive long-context quality validation.
- Nine ~7,800-token controlled padded probes passed; this does not establish general 8K capability.
- Vision structure is retained, but image decision quality is not validated in this text-first release.
- Normalized-entropy confidence is not official Jev confidence equivalence or a correctness guarantee.
- The 100-item benchmark is synthetic, single-author and domain-specific; 99% is reference agreement.
- Original adapter/head/calibration hashes are recorded in
provenance.json; public packaging changes only portable configuration paths and inference imports.
License and references
Model adapter, decision head and included inference code: Apache-2.0. Base model: Qwen3.5-4B, Apache-2.0. The benchmark dataset is CC BY 4.0, authored by haxudev; evaluation outputs and links are included with attribution, without redistributing the full dataset.
- Downloads last month
- -
Task type is invalid.