# Figment Submission Checklist Status: living checklist for evidence. Keep claims in README, submission copy, demo video, and social posts aligned with this file. Primary tracker: [adversarial review action items](adversarial-review-action-items.md). ## Required Submission Artifacts | Artifact | Status | Link / evidence | | -------- | ------ | --------------- | | Public Hugging Face Space | Target exists; runnable proof needed | https://huggingface.co/spaces/build-small-hackathon/figment | | Space cold boot with app files present | Not verified | Last known public API state: `runtime.stage=NO_APP_FILE` with only metadata files present | | Demo video | Proof needed | Pending | | Social post | Proof needed | Pending | | Safety statement | Present | [safety_statement.md](safety_statement.md) | | User test notes | Template present; results needed | [user_test_notes.md](user_test_notes.md) | | License | Present | [../LICENSE](../LICENSE) | | Live hosted Omni trace | Eval traces present; final demo trace still needed | Baseline: `traces/hosted_omni_eval_20260607T194833Z.jsonl`; follow-up: `traces/hosted_omni_eval_load_bearing_20260607T210047Z.jsonl` | | No-cloud/off-grid trace | Proof needed before claiming Off the Grid achieved | Pending | | Hosted Omni eval results | Measured | [hosted_omni_eval_results.md](hosted_omni_eval_results.md): 31/50 whole-output competence, 8/50 full fallback, 480/650 model-retained fields, 170/650 deterministic patches, 50/50 final validation in the follow-up run | | Local 4B + Parakeet eval results | Proof needed | Pending local no-cloud 50-case eval and ASR proof | | Parameter/evidence ledger | Present; organizer confirmation pending | [model_parameter_evidence_ledger.md](model_parameter_evidence_ledger.md) | ## Badge And Claim Status | Claim / badge area | Submission wording allowed now | Evidence needed to upgrade | | ------------------ | ------------------------------ | -------------------------- | | Backyard AI | Targeted; built for a real trained responder, with identity withheld for privacy | Completed user-test notes from that responder on synthetic or de-identified scenarios. Do not claim use, testing, validation, approval, or endorsement before notes exist | | Off the Grid | Targeted / proof-needed | Recorded no-cloud run using self-hosted Omni on adequate local hardware or a smaller verified local stack. Hosted API evidence does not count | | Hosted Gradio Space | Targeted / proof-needed; not claimed runnable | Public Space app files present, cold boot, typed intake run, and route/fallback trace. Last known state is `NO_APP_FILE` | | Demo video | Targeted / proof-needed | Final video link showing only verified routes and labeling fallbacks honestly | | Social post | Targeted / proof-needed | Final social post link with achieved-versus-targeted wording | | Llama Champion | Targeted / proof-needed | Eligible local model route through llama.cpp with trace or eval evidence | | Sharing is Caring | Targeted / proof-needed | Public Space, repo, demo video, social post, and open trace links | | Well-Tuned | Stretch only / proof-needed | Published fine-tuned model or adapter used by the app, plus measured result. Fallback output cannot count | | Field Notes | Tentative | Organizer confirmation and final field-note artifact | | Off-Brand | Targeted / proof-needed | Final demo or social artifact meeting organizer criteria | ## Off-Grid Evidence Boundary Omni is not architecturally disqualified from Off the Grid: it can support the claim if self-hosted on adequate local hardware with no runtime cloud APIs. The current repo does not yet include that recorded proof. Until a no-cloud run exists, use targeted or proof-needed language. ## User-Test Evidence Boundary The README may say the project is built for a real trained responder. It should not say that responder used, approved, validated, or endorsed Figment until [user_test_notes.md](user_test_notes.md) contains factual session notes. ## Eval Evidence Boundary The hosted Omni eval proves model and fallback behavior through the eval harness, not a public Space cold boot. Use the follow-up run as the current hosted eval score: 31/50 whole-output hosted competence, 8/50 full fallback, 480/650 model-retained fields, 170/650 deterministic patches, and 50/50 final validation. Final validation is app safety. Whole-output model competence and field-level model retention are the model-load-bearing metrics. Deterministic fallback and deterministic patches must stay visible in traces, scorecards, submission copy, and the demo. ## Submission Copy Boundaries - Space: may say the target Space exists; do not say it is runnable until app files, cold boot, typed intake, and trace labeling are verified from the public Space URL. - Demo video and social post: use pending placeholders until final links exist. - Backyard AI: may say built for a real trained responder; do not say the target user used or tested Figment until factual notes exist. - Off the Grid: claim only after a recorded no-cloud run. - Llama Champion: claim only after an eligible llama.cpp route runs with trace or eval evidence. - Well-Tuned: claim only after a published fine-tuned model or adapter is used by the app and measured. - Parameter compliance: cite the [model parameter/evidence ledger](model_parameter_evidence_ledger.md), including the Omni 31B body-count versus 33B sidebar ambiguity and organizer-confirmation status.