docs: brain final sign-off deltas — split version matrix (§10), quota rule E2, Post-v1 deferred items 4091794 unverified techfreakworm commited on Jul 21
docs: brain-ratified D8 (per-request LoRA) + final DESIGN/PLAN/README deltas 93c0b30 unverified techfreakworm commited on Jul 21
docs: reflect shipped reality — torch 2.11, per-request LoRA (D8), lazy-load (D9), adaptive residency (M2), long-form chunking (§12), sdpa-only 2d42bd8 unverified techfreakworm commited on Jul 21
Security + correctness hardening (brain code review) 2a9234b unverified techfreakworm commited on Jul 21
LoRA per-request application (ZeroGPU forks don't persist model mutations) 49b81d4 unverified techfreakworm commited on Jul 21
ZeroGPU: lazy-load models in-fork (module-level preload overran startup -> RUNTIME_ERROR) c77559f unverified techfreakworm commited on Jul 21
Fix Space health check: bind 0.0.0.0 on ZeroGPU + disable Gradio SSR e0700ef unverified techfreakworm commited on Jul 21
Fix ZeroGPU startup: skip RAM watchdog on Space (psutil misreports in container); add packages.txt (sox/ffmpeg) c8552a7 unverified techfreakworm commited on Jul 21
ZeroGPU readiness: module-level cuda preload + all-resident on 48GB card + Space metadata f357e40 unverified techfreakworm commited on Jul 21
Adaptive residency: evict other models when memory is tight (DESIGN §6) a39e055 unverified techfreakworm commited on Jul 21
LoRA live-toggle (PEFT enable/disable) + D4 UI + memory hygiene c3b59db unverified techfreakworm commited on Jul 21
Scaffold qwen-voice-studio: Qwen3-TTS core, Gradio studio app, design docs 77f7dfd unverified techfreakworm commited on Jul 21