--- library_name: transformers tags: - decision-model - classification - julia - open-jev - head-finetune - low-resource base_model: SupersonicLabs/Julia-1 datasets: - ZefanCai/Open-Jev pipeline_tag: text-classification --- # Qyvos **Qyvos** is a decision / classification model built from [`SupersonicLabs/Julia-1`](https://huggingface.co/SupersonicLabs/Julia-1) by head-only fine-tuning on the [`ZefanCai/Open-Jev`](https://huggingface.co/datasets/ZefanCai/Open-Jev) dataset (config `release-v2-redistributable`). It maps a **state + question + options** context to a probability distribution over the options, for three decision kinds: | qtype | kind | meaning | |-------|-------|---------| | 0 | `choice` | pick one option (categorical decision) | | 1 | `score` | grade/distribute mass over options (soft scoring) | | 2 | `noul` | binary yes/no decision (options ordered `[no, yes]`) | ## Complete backbone retention The **entire Julia-1 backbone is retained bit-exact**: - `encoder` — ModernBERT-small (22 layers, hidden 384, vocab 256k), 140,530,432 params — **frozen, untouched** - `act_head` — auxiliary action head — **frozen, untouched** Training updated only the decision head (**3,699,073 trainable params = 2.56%** of 144,292,867): - `head` — 2× `TransformerEncoderLayer(384, nhead=6, dim_ff=1536, norm_first=True)` - `type_emb` — `Embedding(3, 384)` (one embedding per decision kind) - `scorer` — `LayerNorm → Linear(384,384) → GELU → Linear(384,1)` per-option scorer Weights are **float32** end-to-end (the Julia checkpoint format has no FP16 variant; nothing was quantized or pruned). ## Training - Data: `ZefanCai/Open-Jev`, config `release-v2-redistributable`, `train` split (79,116 rows available), streamed **30,000 rows** in deterministic shuffled order - Objective: soft-target cross-entropy `-(target · log_softmax(scores)).sum(-1).mean()` — trains calibrated probability distributions, exactly matching Open-Jev's soft `target` vectors - Encoding: identical `julia.data.sequence()` marker-token serialization as the official Julia runtime (train/inference parity) - Optimizer: AdamW, lr 1e-4 (linear warmup 100 steps → linear decay to 10%), β=(0.9,0.999), weight decay 0.01 (no decay on norms/embeddings), grad clip 1.0 - Batch: **1 micro-batch × 8 accumulation** (effective batch 8), 3,750 optimizer steps - Sequence budget: `max_length=1024`, `head_length=512` - Seed: 17 ### Low-RAM training constraints (3.9 GB system RAM) No full fine-tuning was possible (or needed). The run was designed for an 8 GB-class machine: - frozen encoder ⇒ no gradients/optimizer state for 140.5M params; encoder forward runs under `no_grad` - bs=1 micro-batches, `num_workers=0`, fp32 head-only optimizer state - parquet read via **pyarrow mmap**, one row-group (~1,000 rows, a few MB) at a time; lazy per-row tokenization; **no dataset-wide RAM cache** - head-only checkpoints (~45 MB, incl. optimizer) with deterministic skip fast-forward resume - psutil RAM guard: auto-checkpoint + abort if system available RAM < 350 MB - **Measured peak RSS: 1.64 GB** ## Results (honest shuffled-mixture evaluation) Evaluation samples are drawn round-robin across **all** parquet row-groups (each region contributes equally), fixed seed 999, length-bucketed batches. Accuracy = argmax match vs the target's argmax; softCE = soft-target cross-entropy (lower is better). | model | val acc (n=900) | val softCE | test acc (n=900) | test softCE | |---|---|---|---|---| | Julia-1 base (no training) | 83.67% | 0.4542 | 83.56% | 0.4764 | | **Qyvos (30,000 rows)** | **84.11%** | **0.4282** | 83.11% | **0.4447** | Larger-sample confirmation for Qyvos on the test split: n=1500 → acc 81.67%, softCE 0.4587. Per-kind (Qyvos, n=900): | kind | val acc | test acc | |---|---|---| | choice | 66.5% | 68.2% | | score | 79.4% | 79.3% | | noul | 92.7% | 90.2% | Honest reading: Julia-1's pretrained decision head is already strong on this benchmark family, and top-1 accuracy differences between base and Qyvos are within evaluation noise (~±1.2pt stderr at n=900). What fine-tuning **did** deliver, consistently across both splits and sample sizes, is **better-calibrated probabilities (softCE −5.7% val, −6.7% test)** — which is precisely the objective that matters for Open-Jev's soft, often high-entropy targets, and for any downstream consumer of the full probability vector rather than a single argmax. The accuracy ceiling here is largely data-intrinsic: many Open-Jev targets are intentionally soft/ambiguous. ## Exact inference (official runtime) ```python import julia # from the julia/ package shipped in this repo (pip install -e julia --no-deps) engine = julia.load_model("", device="cpu", max_length=1024, head_length=512) pred = engine.predict([{ "state": "User chat: my login is blocked after the password reset.", "question": "Determine the broad category of this support ticket.", "options": ["account: Login, permissions, profile, security", "billing: Charges, invoices, refunds, subscriptions"], "type": "choice", # choice | score | noul }])[0] print(pred["index"], pred["probabilities"]) ``` Or with the bundled CLI (see `training/`): ```bash python3 scripts/infer_qyvos.py --demo # per-kind examples from the test split python3 scripts/infer_qyvos.py --eval-test 900 # honest shuffled test accuracy python3 scripts/infer_qyvos.py --row '{"state":"...","question":"...","options":["a","b"],"type":"choice"}' ``` > **Compatibility note:** on `transformers >= 5.17`, julia's *optional* fast inference path > (`julia/router/encoder.py`) references `ModernBertModel._update_attention_mask`, which was > removed upstream. The bundled scripts disable that optimization via a two-line shim > (`specialize_decision_encoder → False`); inference then uses the standard forward with > numerically equivalent results. On `transformers 5.0.x` (the version this checkpoint was > built against) no shim is needed. ## Repository layout ``` Qyvos/ ├── config.json # architecture pointer file (Julia format v1) ├── julia_config.json # Qyvos identity + head settings ├── model.safetensors # fp32 weights (backbone bit-exact + trained head) ├── encoder/config.json # ModernBERT-small config ├── tokenizer/ # tokenizer (copied bit-exact from Julia-1) ├── label_map.json # qtype id -> kind ├── qyvos_training_config.json # full training config + final metrics ├── provenance.json # base/dataset revisions + SHA-256 hashes └── training/ # all scripts used (download/inspect/train/build/infer) ``` ## Provenance - Base model: `SupersonicLabs/Julia-1` (fp32, weights SHA-256 recorded in `provenance.json`) - Dataset: `ZefanCai/Open-Jev` @ config `release-v2-redistributable` (per-shard SHA-256 in `provenance.json`) - Backbone tensors: bit-exact vs base; only `head`, `type_emb`, `scorer` differ - Final `model.safetensors` SHA-256: `1d03ac9a0c7fbc66f1bde7ecb158741b02c9f898078c0a2f0cc270fb6326f920`