Instructions to use sankalpsthakur/qwen3-06b-typed-decisions-cloud-pilot with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use sankalpsthakur/qwen3-06b-typed-decisions-cloud-pilot with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-0.6B") model = PeftModel.from_pretrained(base_model, "sankalpsthakur/qwen3-06b-typed-decisions-cloud-pilot") - Notebooks
- Google Colab
- Kaggle
Qwen3 0.6B LoRA for grounded yes/no decisions β research pilot
This PEFT adapter was trained in private Kaggle kernel version 1 on 2026-09-30. It is a small, reproducible baseline for typed binary decisions. It is not Jev or a reproduction of Jev's undisclosed architecture. It does not establish a new small-model architecture or a general frontier-model breakthrough.
Inputs and method
- Base: Qwen/Qwen3-0.6B at commit
c1899de289a04d12100db370d81485cdf75e47ca(Apache-2.0). - Data: Praveenrajus/jev-bench at commit
18f88da81c28c2bec55edc31f63f2afdfba109ea; 1,200 BoolQ training records and 128 BoolQ validation records. The dataset card identifies BoolQ as CC-BY-SA-3.0 and grounded StrategyQA as MIT. No source passages are redistributed here. - Training: 600 steps, batch size 2, maximum sequence length 512, seed
20260930, AdamW at1e-4, LoRA rank 8, alpha 16, dropout 0.05 onq_projandv_proj, 1,146,880 trainable parameters. Kaggle assigned a Tesla T4; script runtime was 254.2 seconds. - Prompt: supply the dataset's evidence and question, request exactly
YesorNo, and disable Qwen3 thinking in the chat template. Training loss is applied only to the answer token. Evaluation scores the shared prefix's next-token logits forYesandNo, then normalizes those two logits. Thusp_yesis a conditional two-label score, not the model's unrestricted probability of answering Yes.
The adapter contains only LoRA weights. Load it with the pinned base and the prompt/readout protocol in run.py; generic free-form generation is outside the measured setting. The downloaded adapter's SHA-256 is 827fd0d08c308a89df55343ad7045d9d5b9ac22394e4fa876bdef8c2a001c845.
Evaluation
The base and adapted model were compared on the same records. The BoolQ confirmation set contains 256 test records selected after excluding all 384 local-pilot test/confirmation IDs. The transfer set contains 128 strategyqa_grounded/test records. The table reports raw scores, before calibration.
| Split | N | Base accuracy | LoRA accuracy | Paired accuracy gain, 95% bootstrap CI | Base β LoRA Brier | Base β LoRA ECE (10 bins) |
|---|---|---|---|---|---|---|
| BoolQ validation | 128 | 67.2% | 78.9% | +11.7 points [2.3, 21.1] | 0.216 β 0.193 | 0.117 β 0.189 |
| BoolQ confirmation | 256 | 71.1% | 78.9% | +7.8 points [2.0, 13.7] | 0.205 β 0.195 | 0.122 β 0.190 |
| Grounded StrategyQA transfer | 128 | 55.5% | 67.2% | +11.7 points [2.3, 21.1] | 0.289 β 0.283 | 0.181 β 0.278 |
The raw Brier reductions on both held-out sets have paired bootstrap intervals spanning zero. Raw ECE is worse after LoRA. These are narrow sample estimates, and the 128-record transfer set is one source rather than a broad out-of-distribution suite. The reported bootstrap intervals are exploratory and do not account for all experiment choices.
After inspecting those raw results, we fitted separate positive temperatures for the base and adapter by minimizing NLL on the 128 validation records only. This posthoc exploratory analysis leaves accuracy unchanged. On BoolQ confirmation, calibrated Brier is 0.190 base versus 0.158 adapter, paired reduction 0.032 [0.007, 0.057]. On grounded StrategyQA, it is 0.257 versus 0.211, paired reduction 0.045 [0.006, 0.084]. Calibrated ECE is 0.031 versus 0.035 on BoolQ and 0.086 versus 0.107 on StrategyQA: the adapter still has slightly higher ECE. Prospective calibration claims require a fresh untouched test set.
Reproduce and inspect
run.py is the complete Kaggle training/evaluation script. receipt.json records inputs, versions, settings, and scores. analyze.py checks the label-only predictions, split IDs, paired metrics, and adapter hash; calibrate.py fits validation-only temperatures and reproduces the exploratory scores. excluded-ids.json lists prior pilot IDs used for the no-overlap check. verified_summary.json, calibration_receipt.json, and the *-predictions.json files preserve the analysis inputs. All prediction files contain IDs, binary labels, and scores, with no passage or question text. Run python3 analyze.py and python3 calibrate.py from this directory to check the receipts.
reload_check.py independently loaded this public adapter revision on a fresh Kaggle T4 worker, reconstructed eight BoolQ test prompts from the pinned dataset, and compared the two-label probabilities with the original saved values. reload_receipt.json records 8/8 exact probability matches (maximum absolute difference 0.0) and the adapter hash. This verifies artifact loading and this narrow inference path, not the full evaluation set on a second run.
Scope, attribution, and release
This work tests a Qwen3 answer-token LoRA under a constrained binary decision protocol. It does not validate open-ended answers, safety use, deployment readiness, or Jev parity. Credit the Qwen team and the creators and maintainers of BoolQ, StrategyQA, and jev-bench when using these artifacts. The adapter and authored evaluation code are released under Apache-2.0. The underlying base model and source datasets retain their own licenses; the mixed-source jev-bench license does not supply a single blanket license for every source. This repository does not redistribute the source passages or questions.
Measured evaluation review β 2026-10-07
A single included CPU confirmation scores 128 unused BoolQ and 32 unused grounded StrategyQA records. Adapter/base accuracy is 83.59375%/77.34375% and 56.25%/59.375%, respectively. Both paired accuracy intervals include zero; raw confidence and transfer fail. Five of sixteen historical CPU/GPU checks fail the original 0.002 gate. No checkpoint is promoted.
Report and retained failures Β· Whole-profile evaluation scope. No external rank-one position or industrial deployment qualification is established. Original scientific objects and model weights are preserved.
- Downloads last month
- 38