figment / docs /model_parameter_evidence_ledger.md
ThomsenDrake's picture
Publish Figment Gradio Space app files
5dcfc5c verified
|
Raw
History Blame
6.25 kB

Figment Model Parameter And Evidence Ledger

Date: 2026-06-07

Purpose: keep parameter, route, locality, adapter, ASR, and organizer-confirmation claims in one evidence-gated place. README, submission copy, demo video, and social posts should not upgrade a claim beyond this ledger.

Current Claim Boundary

  • Hosted Omni has measured eval evidence, but not public Space cold-boot evidence.
  • The public Space target exists, but the last known public API state was runtime.stage=NO_APP_FILE with only metadata files present. Do not call the public Space runnable until it is verified from the Space URL.
  • The local 4B + Parakeet route is the preferred no-cloud/off-grid proof path, but it is not yet proven with a real local 50-case eval or local ASR smoke.
  • No published Figment adapter is recorded yet. Well-Tuned remains a stretch claim until a published fine-tuned model or adapter is used by the app and measured.
  • Organizer confirmation is still needed for the Omni 31B body-count versus 33B sidebar ambiguity and for any additive local stack or adapter-count interpretation.

Parameter Ledger

Route / artifact Model ID(s) Total-parameter source used for copy Active-parameter note Adapter parameter count ASR companion count Endpoint locality Organizer-confirmation status Current evidence
Hosted Omni primary nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16; API route nvidia/nemotron-3-nano-omni-30b-a3b-reasoning NVIDIA model-card body: 31B total parameters; HF sidebar has been observed as 33B in planning docs Roughly 3B active parameters per token is a runtime/MoE note, not the compliance number None used in current evals Native Omni speech encoder is part of the Omni model-card count; no separate ASR model is claimed for hosted Omni Hosted NVIDIA API route in current evals; self-hosted no-cloud route not recorded Pending: ask organizers whether model-card body count is acceptable if sidebar count differs Baseline eval: 28/50 whole-output competence, 22/50 full fallback, 50/50 final validation. Follow-up eval: 31/50 competence, 8/50 full fallback, 480/650 model-retained fields, 170/650 deterministic patches, 50/50 final validation
Self-hosted Omni no-cloud target nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16, FP8, or NVFP4 variant if served locally Same Omni 31B body-count claim, with same 33B sidebar ambiguity Active parameters do not decide compliance None recorded Native Omni audio if used locally; included in Omni count if organizers accept the model-card count Would be local/self-hosted only if served with no runtime cloud APIs Pending for count ambiguity and hardware/runtime proof No recorded no-cloud eval or public demo trace yet. Do not claim Off the Grid achieved
Local 4B + Parakeet proof path Text: nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16; ASR: nvidia/parakeet-rnnt-1.1b; local route MODEL_BACKEND=llama_cpp, MODEL_STACK=local_4b_parakeet Workback model-card notes: 3.97B text model plus about 1.1B Parakeet, roughly 5.1B nominal before adapters No active-parameter substitution; use additive total-count story if organizers require stack accounting None recorded yet; exact adapter count must be measured before any Well-Tuned or compliance upgrade About 1.1B for Parakeet RNNT ASR, only if real local ASR is enabled and proven Intended local OpenAI-compatible endpoint on 127.0.0.1 plus local ASR; no-cloud only after recorded proof Pending: confirm additive multi-model counting and adapter counting Configured/labeled path only. No real local 50-case eval, no local ASR smoke, and no trace hash recorded yet
4B Figment adapter stretch Planned adapter name: nvidia-nemotron-3-nano-4b-figment-lora-v1 Base model count is 3.97B; adapter count must be added or documented per organizer guidance Not applicable Pending. Record exact trainable and published adapter parameter count before claiming Parakeet count applies only if adapter demo also uses local ASR Local route or published HF model route, depending on final artifact Pending for adapter accounting and Well-Tuned eligibility Not trained, published, or measured in this ledger
Canned fallback No live model Not a model-compliance artifact Not applicable Not applicable Not applicable Local deterministic fallback Not applicable Useful for safety and cold-start fallback only. Cannot count as model competence, Off the Grid proof, Llama Champion proof, or Well-Tuned proof

Submission Gates

Claim Required upgrade evidence
Public Space runnable Public Space app files present, clean cold boot from the Space URL, typed intake run, and trace showing actual route/fallback status
Hosted model load-bearing Cite hosted eval metrics separately from final validation: 31/50 whole-output competence and 480/650 model-retained fields in the follow-up run
<=32B hosted Omni compliance Organizer accepts the 31B model-card body count or the submission falls back to a clearly eligible smaller route
Off the Grid Recorded no-cloud run with trace evidence, either self-hosted Omni or local 4B + Parakeet/typed intake
Llama Champion Eligible model route runs through llama.cpp with trace or eval evidence
Well-Tuned Published fine-tuned model or adapter is used by the app, measured, and still passes safety validation
Backyard AI user use Completed user-test notes from a real trained responder on synthetic or de-identified scenarios
Demo video and social post Final links exist and wording says achieved only for artifacts supported by this ledger

Evaluation Score Boundary

Final validation success means the application returned a valid, safe output after validation, repair, or fallback. It is not the same as model competence.

Use these current hosted follow-up metrics when summarizing load-bearing behavior:

  • Whole-output hosted competence: 31/50
  • Full deterministic fallback: 8/50
  • Model-retained fields: 480/650
  • Deterministic patches: 170/650
  • Final validation: 50/50

Do not count full fallback or deterministic patches as pure model output.