Implementation Plan
Authored by the project brain. Ordered tasks; each has binding acceptance criteria (AC). Cross-cutting rules at the bottom apply to every task. (The lead ran ahead of this plan during design; a working core + app already exists, so several tasks land as reconciliation against the rulings rather than greenfield.)
T0 — Scaffold (do first, push immediately). Repo tree per DESIGN §4; LICENSE = Apache-2.0; .gitignore (data/, *.wav except docs/samples/, pycache, .venv, .DS_Store); requirements.txt (Space runtime, pinned exact) + requirements-dev.txt (adds pytest, playwright, psutil); README; docs/{DESIGN,PLAN,DECISIONS}.md.
AC: git push green to public GitHub; sole-operator authorship (no co-author lines, no "Generated with"); no venv/test artifacts in repo (they live under ~/Projects/tests).
T1 — config.py + device.py. DevicePolicy detection (ZeroGPU via spaces-import/env → cuda; else mps; else cpu), seed_all, memory gauge (vm_stat committed on macOS; torch.cuda stats on cuda), Watchdog (warn 72 / abort 76 GB, pageout-delta tracking, background sampling thread).
AC: policy script prints mps/bf16/sdpa/all_resident locally; gauge within ±2 GB of Activity Monitor "Memory Used"; watchdog abort path proven with an artificially low threshold; QVS_* env overrides all work.
T2 — registry.py. Sequential 3-checkpoint load per policy; per-load committed-delta logging; all_resident | single_on_demand; adaptive degrade (baseline >65 GB); codec-instance count reported.
AC: all-3 resident, peak <72 GB, per-model deltas logged; QVS_FORCE_SINGLE_RESIDENT=1 exercises degrade path; report whether codec is shared or 3×.
T3 — audio.py + engine.py. synthesize() for all three modes headless; per-request seeding; duration_estimate(); watchdog sampling active during generation.
AC: three wavs — non-silent (RMS > floor), duration >0.5 s, sr == API-returned; same seed+device ⇒ repeatable; different seed ⇒ differs; estimator sane for short/long text.
T4 — voices.py. VoicePrompt kinds ref_pair / full_prompt / xvector; save/load with meta sidecars under data/voices.
AC: roundtrip all three kinds; each usable as clone voice source; Darija speaker_embedding.pt imports as xvector{(2048,), speaker_id 3000} and generates.
T5 — lora.py. AdapterManager per DESIGN §8.
AC: DESIGN §8 (i)–(iv) all pass with the Darija adapter, including <1 s toggle and post-unload same-seed baseline match.
T6 — ui/ + app.py. Five tabs per DESIGN §5, shared Advanced builder, status strip, gr.queue(default_concurrency_limit=1), spaces no-op shim, @spaces.GPU(duration=estimate) on handlers. Use the frontend-design plugin for the visual pass (operator preference).
AC: Playwright walkthrough of EVERY tab and param group with programmatic audio asserts (file, duration, RMS, sr) + screenshots for UI state; Design→Clone bridge produces a working clone ref; memory stays <72 GB for the entire suite; one app process only.
T7 — Batch (Preset tab) + polish. Multiline → zip. AC: 3 lines → zip of 3 valid wavs; UI errors are readable (no raw tracebacks).
T8 — Pre-deploy gate. Finalize README; Space requirements.txt (torch pin ladder per D7); push GitHub.
AC: fresh-clone + README-only install reproduces the app locally (doc-follow test); repo public, Space repo private confirmed.
T9 — Space deploy + verify. Create private ZeroGPU Space techfreakworm/qwen-voice-studio; deploy; watch first boot logs; acceptance #1: two consecutive generations logging next(model.parameters()).device inside the fork = cuda both times (orphan check → else flip QVS_RESIDENCY=fork_move); if torch 2.13 rejected → pin 2.11.0 AND rerun local smoke suite on a 2.11.0 venv BEFORE redeploying; then full Playwright suite against the Space; review GPU-seconds vs 40 min/day quota; calibrate duration-estimate k.
AC: all modes + LoRA green on the Space; GitHub synced at the deploy commit; quota log reviewed; no credit top-up without operator approval.
Cross-cutting rules
Commit at least per-task, push regularly; long runs = background shells with monitoring; watchdog active in every model-touching run; never a second model-holding process; on the 2nd failed fix of any bug — stop patching, bring it to the brain for first-principles review; MPS↔CUDA outputs are never bit-compared.