Figment Model Parameter And Evidence Ledger
Date: 2026-06-07
Purpose: keep parameter, route, locality, adapter, ASR, and organizer-confirmation claims in one evidence-gated place. README, submission copy, demo video, and social posts should not upgrade a claim beyond this ledger.
Current Claim Boundary
- Hosted Omni has measured eval evidence, but not public Space cold-boot evidence.
- The public Space target exists, but the last known public API state was
runtime.stage=NO_APP_FILEwith only metadata files present. Do not call the public Space runnable until it is verified from the Space URL. - The local 4B + Parakeet route is the preferred no-cloud/off-grid proof path, but it is not yet proven with a real local 50-case eval or local ASR smoke.
- No published Figment adapter is recorded yet. Well-Tuned remains a stretch claim until a published fine-tuned model or adapter is used by the app and measured.
- Organizer confirmation is still needed for the Omni 31B body-count versus 33B sidebar ambiguity and for any additive local stack or adapter-count interpretation.
Parameter Ledger
| Route / artifact | Model ID(s) | Total-parameter source used for copy | Active-parameter note | Adapter parameter count | ASR companion count | Endpoint locality | Organizer-confirmation status | Current evidence |
|---|---|---|---|---|---|---|---|---|
| Hosted Omni primary | nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16; API route nvidia/nemotron-3-nano-omni-30b-a3b-reasoning |
NVIDIA model-card body: 31B total parameters; HF sidebar has been observed as 33B in planning docs | Roughly 3B active parameters per token is a runtime/MoE note, not the compliance number | None used in current evals | Native Omni speech encoder is part of the Omni model-card count; no separate ASR model is claimed for hosted Omni | Hosted NVIDIA API route in current evals; self-hosted no-cloud route not recorded | Pending: ask organizers whether model-card body count is acceptable if sidebar count differs | Baseline eval: 28/50 whole-output competence, 22/50 full fallback, 50/50 final validation. Follow-up eval: 31/50 competence, 8/50 full fallback, 480/650 model-retained fields, 170/650 deterministic patches, 50/50 final validation |
| Self-hosted Omni no-cloud target | nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16, FP8, or NVFP4 variant if served locally |
Same Omni 31B body-count claim, with same 33B sidebar ambiguity | Active parameters do not decide compliance | None recorded | Native Omni audio if used locally; included in Omni count if organizers accept the model-card count | Would be local/self-hosted only if served with no runtime cloud APIs | Pending for count ambiguity and hardware/runtime proof | No recorded no-cloud eval or public demo trace yet. Do not claim Off the Grid achieved |
| Local 4B + Parakeet proof path | Text: nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16; ASR: nvidia/parakeet-rnnt-1.1b; local route MODEL_BACKEND=llama_cpp, MODEL_STACK=local_4b_parakeet |
Workback model-card notes: 3.97B text model plus about 1.1B Parakeet, roughly 5.1B nominal before adapters | No active-parameter substitution; use additive total-count story if organizers require stack accounting | None recorded yet; exact adapter count must be measured before any Well-Tuned or compliance upgrade | About 1.1B for Parakeet RNNT ASR, only if real local ASR is enabled and proven | Intended local OpenAI-compatible endpoint on 127.0.0.1 plus local ASR; no-cloud only after recorded proof |
Pending: confirm additive multi-model counting and adapter counting | Configured/labeled path only. No real local 50-case eval, no local ASR smoke, and no trace hash recorded yet |
| 4B Figment adapter stretch | Planned adapter name: nvidia-nemotron-3-nano-4b-figment-lora-v1 |
Base model count is 3.97B; adapter count must be added or documented per organizer guidance | Not applicable | Pending. Record exact trainable and published adapter parameter count before claiming | Parakeet count applies only if adapter demo also uses local ASR | Local route or published HF model route, depending on final artifact | Pending for adapter accounting and Well-Tuned eligibility | Not trained, published, or measured in this ledger |
| Canned fallback | No live model | Not a model-compliance artifact | Not applicable | Not applicable | Not applicable | Local deterministic fallback | Not applicable | Useful for safety and cold-start fallback only. Cannot count as model competence, Off the Grid proof, Llama Champion proof, or Well-Tuned proof |
Submission Gates
| Claim | Required upgrade evidence |
|---|---|
| Public Space runnable | Public Space app files present, clean cold boot from the Space URL, typed intake run, and trace showing actual route/fallback status |
| Hosted model load-bearing | Cite hosted eval metrics separately from final validation: 31/50 whole-output competence and 480/650 model-retained fields in the follow-up run |
| <=32B hosted Omni compliance | Organizer accepts the 31B model-card body count or the submission falls back to a clearly eligible smaller route |
| Off the Grid | Recorded no-cloud run with trace evidence, either self-hosted Omni or local 4B + Parakeet/typed intake |
| Llama Champion | Eligible model route runs through llama.cpp with trace or eval evidence |
| Well-Tuned | Published fine-tuned model or adapter is used by the app, measured, and still passes safety validation |
| Backyard AI user use | Completed user-test notes from a real trained responder on synthetic or de-identified scenarios |
| Demo video and social post | Final links exist and wording says achieved only for artifacts supported by this ledger |
Evaluation Score Boundary
Final validation success means the application returned a valid, safe output after validation, repair, or fallback. It is not the same as model competence.
Use these current hosted follow-up metrics when summarizing load-bearing behavior:
- Whole-output hosted competence: 31/50
- Full deterministic fallback: 8/50
- Model-retained fields: 480/650
- Deterministic patches: 170/650
- Final validation: 50/50
Do not count full fallback or deterministic patches as pure model output.