Local 4B V4 Training Plan
Date: 2026-06-10
Purpose
Train one focused v4 LoRA only after the scaffolding and eval-shape fixes in docs/local_4b_v4_scaffolding_eval_shape_plan.md are implemented and rerun.
The v4 goal is not broad assistant quality. The goal is to improve the local 4B model at the model-owned parts of Figment's field workflow:
- radio and runner handoff,
- concise SBAR referral support,
- source-card discipline,
- high-value next observations,
- low-resource constraints,
- safe protocol-navigation language.
The v3 result is good enough to be worth refining, not bad enough to restart from scratch.
Current Evidence
Primary v3 trace:
traces/local_4b_finetuned_v3_field_holdout_sequential_20260610T102450Z/
Published trace dataset:
https://huggingface.co/datasets/ThomsenDrake/figment-eval-traceslocal_4b_clean_scored_records:350hosted_omni_scored_records:100scored_eval_records:450useful_trace_records:455
V3 field-workflow holdout:
150/150cases completed107/150competence successes93/150raw model successes14/150focused repair successes2/150full fallbacks148/150final validation successes0/150strict expected-label successes due tomissing_observation_cues_present
Failure concentration:
REFERRAL-SBAR-v1:0/27radio_handoff:0/16sbar_handoff_usefulness:0/10source_card_discipline:2/6rural_clinic_intake:33/36disaster_triage:30/32
Interpretation:
- V3 is strong enough on safety and protocol navigation to keep.
- V3 is not strong enough on the handoff layer that matters to the field workflow.
- The v4 dataset should be narrow and high-signal, not another broad corpus.
Prerequisite
Do not start the v4 training job until these are complete:
- Eval cue ownership is split into model-owned, handoff-owned, and harness-owned cues.
- Deterministic harness evidence is visible outside model-authored missing-observation text.
- SBAR/radio handoff metrics report concrete failures.
- V3 is rerun with the updated scoring.
- The remaining v3 failures are exported as v4 teacher prompts or repair seeds.
This prevents v4 from learning to recite deterministic metadata instead of improving handoff usefulness.
Status on 2026-06-10: prerequisites 1 through 5 are complete for the v4 dataset/job-readiness path, with evidence below.
Implementation Evidence
Current v4-readiness work on 2026-06-10:
- Updated scoring now splits model-owned, handoff-owned, and harness-owned evidence cues.
- Current-code local v3 smoke evidence:
traces/v4_readiness_v3_current_smoke_20260610T141544Z/3/3expected-label successes.3/3handoff-readiness successes.- No fallback use and no context/KV/HTTP-500 runtime errors.
- V3 holdout seed export:
data/finetune/v4_seed_exports/figment_sft_v4_v3_holdout_seeds.jsonl8model/handoff/source failure seeds.142harness-only score failures preserved as replay/synthetic-sibling seeds, not direct failure rows.- Holdout-copy policy recorded in
data/finetune/v4_seed_exports/figment_sft_v4_v3_holdout_seeds.manifest.json.
- V4 corpus wrapper:
scripts/generate_v4_full_corpus.py- Defaults:
1500navigator rows plus150focused repair rows. - Dataset/output paths default to
figment_sft_v4. - V4 distribution is intentionally handoff-heavy while preserving replay and hard-negative coverage:
375radio handoff,330SBAR handoff usefulness,210source-card discipline,150low-resource,150missing-observation prioritization,105workflow-repair-seed,105rural/disaster replay, and75safety hard-negative navigator rows before focused repair augmentation.
- Defaults:
- Teacher-backed v4 smoke corpus:
data/finetune/figment_sft_v4_smoke.jsonl4/4accepted navigator rows fromnvidia/nemotron-3-ultra-550b-a55b:free.2focused repair rows added:handoff_note_sbarandcitations_and_pathways.- Harness verification passed with
0issues. - Modal smoke split prepared at
data/finetune/modal/figment_sft_v4_smoke/.
- Full teacher-backed v4 corpus:
data/finetune/figment_sft_v4.jsonl1500navigator rows plus150focused repair rows,1650total.1500case specs atdata/finetune/figment_sft_v4_case_specs.jsonl.- Final dataset sha256:
ef7a7c9a6a99927ba72ce244e03a9da3ab86d3cf5dc70786703fb5f8bdf2a289. - Case-spec sha256:
aca6630d50e32260f3121a366406225309409c3ad5de8d495c1b5a99f5bb34e2. - Standalone harness verification passed with
0issues:.venv/bin/python scripts/verify_finetune_harness_alignment.py --dataset data/finetune/figment_sft_v4.jsonl --case-specs data/finetune/figment_sft_v4_case_specs.jsonl. - Category counts:
406radio handoff,317SBAR handoff usefulness,218source-card discipline,160low-resource constraints,128missing-observation prioritization,110workflow-repair-seed,71escalation precision,55rural clinic intake, and35disaster triage. - Focused repair counts:
68handoff-note/SBAR,38citations/pathways,23missing observations,7forbidden clinical language,7protocol urgency, and7schema. - Modal split prepared at
data/finetune/modal/figment_sft_v4/:1482train rows and168validation rows. - Modal train sha256:
af9af7111af057e42e14f1a6f07309eee6737c218cf403e104447b74fe46fb3f. - Modal validation sha256:
6a2859047ae78479b97ab797644a6646df79d8b4ee920ed21ce1469ba2302b7d. - The direct NVIDIA-compatible endpoint completed shards
0through15and then stalled on shards16through19; incomplete direct-endpoint partials were archived underdata/finetune/shards/aborted_nvidia_timeout_20260610T161117Z/. - OpenRouter fallback with
nvidia/nemotron-3-ultra-550b-a55b:freeresumed from complete shards and generated the remaining shards16through29; final source attempts were1749with123teacher backend errors and no accepted-row provenance mixing inside a completed shard. - Focused regression suite passed after generation:
.venv/bin/python -m pytest tests/test_prompt_builder_contract.py tests/test_focused_repair.py tests/test_navigator_safety.py tests/test_eval_runner.py tests/test_eval_metrics.py tests/test_finetune_v2_data_plan.py tests/test_runtime_honesty.py tests/test_modal_finetune_prep.py tests/test_v4_training_seed_export.py -q->82 passed.
- Modal v4 smoke job passed:
- Command:
.venv/bin/modal run modal/finetune_figment_nemotron.py --dataset-version figment_sft_v4 --dataset data/finetune/figment_sft_v4.jsonl --output-name figment-sft-v4-lora-smoke --smoke --gpu L40S --learning-rate 2e-5 --lora-r 16 --lora-alpha 32 --lora-dropout 0.05 --gradient-accumulation-steps 8 --validation-steps 2 --save-steps 5. - Modal app:
ap-J7w1D5j8VwZ1S9CuF4mwzN. - Staged rows:
1482train,168validation. - Tokenized rows:
1482train,168validation. - Adapter path:
/checkpoints/figment_sft_v4/figment-sft-v4-lora-smoke. - Smoke config:
max_steps=5,max_seq_length=2048,learning_rate=2e-5,lora_r=16,lora_alpha=32,lora_dropout=0.05,gradient_accumulation_steps=8. - Metrics:
train_loss=14.122270011901856,train_runtime=148.2881,epoch=0.02699055330634278; eval loss was1.741158127784729at step 2 and1.7384405136108398at step 4. - Verified Modal volume artifacts include
adapter_model.safetensors,adapter_config.json, tokenizer files,chat_template.jinja,figment_training_manifest.json, andcheckpoint-5/.
- Command:
Training Strategy
Use a targeted continuation from v3 as the primary run.
Primary run:
- Dataset version:
figment_sft_v4 - Output adapter name:
figment-sft-v4-lora - Base model:
nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16 - Starting point: continue from the v3 behavior if the Modal script is extended to load an existing adapter; otherwise train a focused v4 LoRA from the BF16 base with replay rows.
- Method: LoRA SFT, BF16 base, merge back to BF16, convert to GGUF, evaluate locally through llama.cpp.
- Context length:
16384 - Target GPU:
L40Sfirst,A100-80GBonly if the run hits memory or sequence-length failures.
If the current Modal trainer cannot resume from an existing adapter, patch it before v4 or run a fresh LoRA with enough v2/v3 replay to preserve schema behavior.
Dataset Size And Mix
Target accepted rows: 1200 to 1800.
Recommended mix:
350to450radio handoff rows.300to400SBAR handoff usefulness rows.175to250source-card discipline rows.150to225low-resource constraint rows.125to175missing-observation prioritization rows focused on first-five usefulness, not every cue.100to150focused competence-repair rows from v3 safe-but-weak outputs.100to150clinical-protocol replay rows from high-quality v2/v3 data.75to125hard negative or safety-boundary rows to preserve refusal and no-treatment behavior.
Replay rows should be high quality only:
- validation passed,
- no full fallback,
- no forbidden behavior,
- correct target card,
- correct source-card set,
- strong field provenance,
- no close neighbor of locked eval or holdout rows.
What To Generate
Every v4 row should match the exact harness prompt and response format. Do not generate generic clinical conversations.
Full Navigator Rows
Generate full assistant outputs where the model must:
- preserve deterministic red flags,
- keep urgency at or above the deterministic floor,
- cite only retrieved source cards,
- include
REFERRAL-SBAR-v1when the task is handoff-focused, - produce compact SBAR fields grounded only in confirmed intake, rules, and retrieved cards,
- prioritize the next observations that would actually help the responder move the case forward.
Focused Repair Rows
Generate repair rows for safe-but-weak outputs, not only invalid outputs.
Repair scopes:
handoff_note_sbarsource_cardscandidate_protocol_pathwaysmissing_observationsresponder_checklistsafety_boundary
Each repair row should include:
- previous weak output,
- deterministic validation result,
- competence metric failures,
- scope name,
- corrected assistant output or corrected fields,
- provenance metadata saying this is a competence repair.
Preference Pairs
Only add preference data after the SFT row set exists.
Preferred outputs:
- concise,
- grounded,
- high-value next observations first,
- correct source cards,
- safe SBAR,
- useful to a field medic under radio or paper constraints.
Rejected outputs:
- schema-valid but generic,
- overlong,
- metadata-stuffed,
- unsupported,
- target-card correct but handoff useless,
- observation list repeats the prompt without prioritization.
Preference tuning is optional. Use it only if v4 SFT improves format but still leaves SBAR/radio output operationally weak.
Teacher Model
Use the existing stronger teacher path:
- Teacher model:
nvidia/nemotron-3-ultra-550b-a55b - Preferred endpoint: existing hosted OpenAI-compatible endpoint.
- Fallback endpoint: OpenRouter if needed.
- Secrets: use
.envlocally and Modal secrets remotely. Do not write keys into dataset rows, manifests, traces, or docs.
Teacher instructions should make the field workflow explicit:
- "You are generating training targets for a bounded protocol-navigation harness, not medical advice."
- "The model output must be JSON only and match Figment's current navigator schema."
- "Optimize for a trained field responder who needs faster intake, escalation, and handoff, under low-resource constraints."
- "Do not copy locked eval rows or close paraphrases."
- "Do not add diagnosis, treatment, dosing, discharge, or autonomous routing language."
Validators
Keep all v3 validators:
- JSON/schema validation,
- known-card validation,
- retrieved-card validation,
- urgency floor,
- red-flag match,
- source-card coverage,
- candidate-pathway coverage,
- forbidden behavior,
- no teacher notes,
- no locked eval or holdout near-neighbor.
Add v4 validators:
handoff_readiness_passedsbar_slot_coveragesbar_unsupported_fact_countradio_brevity_okfirst_five_observation_usefulnesssource_card_discipline_passedcompetence_repair_scope_validharness_owned_metadata_not_required_in_model_text
Reject any row that only wins by stuffing deterministic metadata into prose.
Modal Work Needed
Patch modal/finetune_figment_nemotron.py before v4 if needed:
- expose
learning_rate, - expose
lora_r, - expose
lora_alpha, - expose
lora_dropout, - expose
gradient_accumulation_steps, - expose
validation_steps, - expose
save_steps, - optionally support
resume_adapter_nameoradapter_init_path.
Status on 2026-06-10:
- The entrypoint now accepts
learning_rate,lora_r,lora_alpha,lora_dropout,gradient_accumulation_steps,validation_steps, andsave_steps. - The entrypoint still does not support
resume_adapter_nameoradapter_init_path. - The ready full-run path is therefore fresh from
nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16with replay-heavy v4 data, not continuation from the v3 adapter.
Recommended SFT config:
max_seq_length:16384lora_r:16first,32only if v4 underfits the targeted handoff taskslora_alpha:32for rank 16,64for rank 32lora_dropout:0.05learning_rate:2e-5for continuation from v3,5e-5if training fresh from base with replaygradient_accumulation_steps:8validation_fraction:0.10max_steps:372for the current fresh-from-base v4 run, approximately2.0epochs over1482train rows at batch size1and gradient accumulation8
Runbook
- Implement and rerun the scaffolding/eval-shape plan.
- Export v3 failures with ownership labels and handoff metrics.
- Generate v4 candidate specs from those failures and nearby synthetic siblings.
- Use the teacher to produce JSON-only target outputs.
- Validate and reject rows until
1200to1800accepted rows remain. - Prepare Modal train/validation split:
.venv/bin/python scripts/prepare_modal_finetune_dataset.py \
--dataset data/finetune/figment_sft_v4.jsonl \
--dataset-version figment_sft_v4
- Run a smoke job:
.venv/bin/modal run modal/finetune_figment_nemotron.py \
--dataset-version figment_sft_v4 \
--dataset data/finetune/figment_sft_v4.jsonl \
--output-name figment-sft-v4-lora-smoke \
--smoke true \
--gpu L40S
- Run the full detached job:
.venv/bin/modal run modal/finetune_figment_nemotron.py \
--dataset-version figment_sft_v4 \
--dataset data/finetune/figment_sft_v4.jsonl \
--output-name figment-sft-v4-lora \
--max-steps 372 \
--learning-rate 5e-5 \
--lora-r 16 \
--lora-alpha 32 \
--lora-dropout 0.05 \
--gradient-accumulation-steps 8 \
--validation-steps 25 \
--save-steps 50 \
--gpu L40S \
--spawn-train
- Merge adapter:
.venv/bin/modal run modal/finetune_figment_nemotron.py \
--merge-only \
--dataset-version figment_sft_v4 \
--adapter-name figment-sft-v4-lora \
--merged-name figment-sft-v4-lora-merged-bf16 \
--gpu L40S
- Pull merged BF16 weights, convert to GGUF, serve locally through
llama-server, smoke route, and run:
- locked 50-case regression,
- field-workflow holdout with updated scoring,
- old v3 scoring for comparison only.
Acceptance Gates
Primary gate:
- field-workflow holdout competence at least
125/150. REFERRAL-SBAR-v1at least20/27.radio_handoffat least12/16.sbar_handoff_usefulnessat least8/10.source_card_disciplineat least5/6.
Safety gates:
final_validation_successesat least148/150.forbidden_behavior_absentremains150/150.red_flags_matchremains150/150.min_urgency_metremains150/150.- full fallbacks no more than
2/150.
Regression gate:
- locked 50-case competence must not drop below the v2 result of
33/50unless the miss is only a newly separated non-safety cue metric. - no increase in unsafe or unsupported clinical language.
- no loss of local/no-cloud route proof.
Operational gate:
- local GGUF hash recorded,
/v1/modelsmetadata recorded,- llama.cpp run uses
n_parallel=1or otherwise proves enough KV context for the prompt length, - eval manifest has all trace hashes,
- invalid parallel/runtime records are excluded from scored reporting.
Ship Decision
Train v4 if the scaffolding rerun still shows a real model-owned SBAR/radio gap.
Ship v3 plus scaffolding if:
- scaffolding alone gets field holdout competence close to the target,
- v4 regresses safety or validation,
- v4 improves scorer numbers by stuffing metadata rather than improving handoff usefulness,
- the remaining failures are mostly evaluator wording artifacts.
With roughly 8.5 days left, the recommended path is one focused v4 swing, not an open-ended training campaign.