Spaces:
Running on Zero
Running on Zero
File size: 17,521 Bytes
d3359f1 7e0e495 28febab d3359f1 28febab d3359f1 28febab d3359f1 28febab d3359f1 28febab d3359f1 28febab d3359f1 c3f6d17 d3359f1 f21a791 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 | # Hosting (DPLY-01)
**Decided:** 2026-08-27
**Chosen path:** `zerogpu-free`
| Field | Value |
|---|---|
| Space | WolfDavid/japanese-learning-avatar |
| Space page | https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar |
| SPACE_URL (direct, used by E2E `--space-url`) | https://wolfdavid-japanese-learning-avatar.hf.space |
| Visibility | public (`private: false`) |
| Hardware requested | `zero-a10g` (ZeroGPU) |
| Hardware applied | **`zero-a10g`** β attached on the first successful deploy (`4e3cac9`, 2026-09-05). Was `null` before that; see "Why `hardware.current` is null" below for why that was correct at the time. |
| sdk / sdk_version | gradio / **6.22.0** β corrected by plan 01-05's `README.md`, which is now the Space manifest. See "The front-matter correction, applied". |
| python_version | **3.12.12** β corrected by plan 01-05. |
| Sleep timeout | 48 h (`gcTimeout: 172800`) |
| Space repo SHA at time of writing | `b70455dcf67ff7db6b1f48f08153982d37a63dd2` (two files: `.gitattributes`, `README.md`) |
| Created | 2026-08-27T02:13:12Z |
## The decisive finding: `cpu-basic` is no longer creatable, and it fails with a payment error
Creating this Space on the default `cpu-basic` hardware was **rejected with HTTP 402 Payment
Required**. The message, verbatim:
> Static Spaces are free for everyone, but hosting Gradio and Docker Spaces on free cpu-basic
> requires a PRO subscription.
Re-issuing the identical creation request with ZeroGPU (`space_hardware="zero-a10g"`) **succeeded
immediately, with no payment prompt and no charge.**
This is empirical confirmation of the claim in `01-RESEARCH.md` and `CLAUDE.md` that Gradio Spaces
on a free personal account are PRO-gated *except* for the ZeroGPU carve-out. It is no longer a
documentation reading β it is an observed platform response on this account.
**Consequence:** ZeroGPU is now the **only** free hosting path for this project. A future phase
**cannot** silently fall back to `cpu-basic` to save a ZeroGPU slot β that fallback does not exist
for a Gradio Space on a non-PRO personal account. Any plan that proposes it is proposing a $9/mo
purchase.
## Account eligibility evidence (verified against the live public API, 2026-08-27)
Queried in this task, not quoted from RESEARCH:
- `isPro: false`; account created `2023-11-27T01:45:12.000Z` β about 2.7 years old, so the ZeroGPU
">30 days in good standing" criterion is satisfied with enormous margin.
- Email: verified (account is in good standing; ZeroGPU allocation was accepted, which requires it).
- **8 pre-existing Gradio Spaces, every one of them `requested: cpu-basic`, and zero on ZeroGPU** β
so both free ZeroGPU slots were open before this Space was created:
| Space | sdk_version | python | hardware requested | stage |
|---|---|---|---|---|
| `WolfDavid/dino` | 6.9.0 | 3.11 | cpu-basic | RUNNING |
| `WolfDavid/fea-surrogate` | 5.9.1 | 3.11 | cpu-basic | SLEEPING |
| `WolfDavid/mechspec-qwen-demo` | 5.9.1 | 3.11 | cpu-basic | SLEEPING |
| `WolfDavid/agentic-market-analyzer` | 5.9.1 | 3.11 | cpu-basic | SLEEPING |
| `WolfDavid/log-anomaly-detector` | 5.9.1 | 3.11 | cpu-basic | RUNNING |
| `WolfDavid/vision-edge` | 5.9.1 | 3.11 | cpu-basic | SLEEPING |
| `WolfDavid/nllb-translator` | 5.9.1 | 3.11 | cpu-basic | SLEEPING |
| `WolfDavid/blip-captioner` | 5.9.1 | 3.11 | cpu-basic | SLEEPING |
- `WolfDavid/dino` runs `sdk_version: 6.9.0`, so Gradio 6 is proven to work on this account.
- Note for anyone re-running this check: **`runtime.hardware.current` is `null` for every SLEEPING
Space.** `current` reports what is running right now, not what the Space is entitled to. Read
`runtime.hardware.requested` when auditing hardware tiers, or six of these eight Spaces look
unallocated.
- Note also that `GET /api/spaces?author=WolfDavid` still returned only the original 8 shortly after
creation; the list endpoint lags. `GET /api/spaces/WolfDavid/japanese-learning-avatar` returns the
new Space immediately and is the authoritative read.
## Why `hardware.current` is null, and why that is correct
```
"runtime": { "stage": "NO_APP_FILE",
"hardware": { "current": null, "requested": "zero-a10g" },
"gcTimeout": 172800 }
```
The Space repo contains only `.gitattributes` and an auto-generated `README.md`. With no `app.py`
there is nothing to schedule, so no hardware is attached. **This is the correct pre-deploy state and
is not a failure.** `requested: zero-a10g` is the binding fact β it is the tier the Space will start
on the moment plan 01-05 pushes an app.
Reachability today, measured three times consecutively with identical results:
| URL | HTTP | Meaning |
|---|---|---|
| https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar | **200** | The Space exists and is public |
| https://wolfdavid-japanese-learning-avatar.hf.space | **503** | `Your space is in error, check its status on hf.co` β no `app.py` yet |
The direct subdomain **is** provisioned (`runtime.domains[0].stage: READY`, `subdomain:
wolfdavid-japanese-learning-avatar`), it simply has nothing to serve. **Plan 01-05 turns this 503
into a 200**, and `tests/e2e/test_avatar_loop.py::test_space_reachable` β which plan 01-05 creates β
is what actually satisfies DPLY-01. DPLY-01 is **not** satisfied by this plan.
## The front-matter correction, applied (plan 01-05)
Space creation auto-generated a `README.md` whose front-matter **conflicted with the pinned
stack**. Plan 01-05 wrote a repository `README.md` that replaces it:
| Key | Auto-generated value | Value now shipped | Source of the requirement |
|---|---|---|---|
| `sdk_version` | `6.26.0` | `6.22.0` | `requirements.txt` line 1 pins `gradio==6.22.0`; `CLAUDE.md` says pin exactly |
| `python_version` | `'3.12'` | `3.12.12` | ZeroGPU provides only 3.10.13 and 3.12.12; `.python-version` is `3.12.12` |
| `emoji` | `π ` | `πΎ` | cosmetic; the auto-generated value was also mojibake in the API response |
| `colorFrom` / `colorTo` | `purple` / `gray` | `pink` / `indigo` | cosmetic |
| `license` | absent | `mit` | the repository's own code is MIT; third-party asset terms are separate |
| `short_description` | absent | present | shown on the profile card a recruiter sees first |
`hf_oauth` was deliberately **not** added: it belongs to Phase 4, and every extra consent scope is
a deterrent on the Space's first-load screen.
Left uncorrected, the Space SDK would have installed Gradio 6.26.0 while `requirements.txt` asked
for 6.22.0 β the exact "unpinned sdk_version floats you into a breaking release" failure `CLAUDE.md`
warns about, arrived at from the other direction.
## How the unrelated histories were reconciled (plan 01-05)
The Space repo was created with its own root commit β `b70455d`, "initial commit", two files
(`.gitattributes`, `README.md`) β while the local repository has a completely separate root and
~44 commits. `git merge-base master space/main` returned nothing: **the two histories share no
ancestor**, so a plain `git push space master:main` is rejected as non-fast-forward.
**Chosen: `git merge --allow-unrelated-histories space/main`, not a force-push.** Both files
conflicted and both were resolved deliberately:
- **`README.md` β resolved to ours.** The remote copy is the boilerplate quoted above and is
precisely what this plan exists to replace; there is nothing in it to preserve.
- **`.gitattributes` β resolved to the union.** Ours (plan 01-02: `*.vrm *.vvm *.wav *.onnx`)
is kept verbatim and Hugging Face's 35 default LFS patterns are appended below it. The
defaults cost nothing today β every currently tracked binary is already an LFS object under
the narrow patterns β and they arm the formats later phases will push (`.safetensors`,
`.bin`, `.pt`, `.npz`) against the 10 MiB non-LFS rejection. Dropping them to keep the file
minimal would have traded a real safety net for tidiness.
A force-push would have worked and lost nothing of substance, but the merge keeps the Space's
own creation commit in the history, which is the provenance record for the ZeroGPU allocation.
## GPU posture on the deployed Space
`DISABLE_GPU=1` **is set** as a Space variable (set by plan 01-05 via
`HfApi.add_space_variable`, 2026-09-05, description recorded on the variable itself). Phase 1
calls no GPU function, so this costs nothing and makes SC-4's demonstration honest from the
first deploy rather than retrofitted at the end. Read it back with:
```
python -c "from huggingface_hub import HfApi; print(HfApi().get_space_variables('WolfDavid/japanese-learning-avatar'))"
```
`AVATAR_TRANSPORT` is deliberately **not** set, so the Space uses the component's default,
`inline` β the transport the spike is testing. Setting it to `iframe` is the entire fallback.
### SC-4 run record β the whole loop with the GPU disabled (plan 01-09)
The platform forces one `@spaces.GPU` function to exist (`app.py::zerogpu_probe`, finding 4
below), so "no GPU function exists" is not an available proof. The honest proof is that the
whole turn loop completes without ever reaching one β and reaching it would raise, because it
checks `DISABLE_GPU` on every call.
| | |
|---|---|
| Date | 2026-09-06, 03:27β03:32 UTC |
| Space revision | `c911d74b378981e6ba0d0111eb570ed3b20927db` (plan 01-08 code, deployed by the owner) |
| Runtime at run time | `RUNNING`, hardware `zero-a10g`, read from `space_info()` in the same script |
| `DISABLE_GPU` | **`1`**, read back from `get_space_variables()` in the same script (set 2026-09-05 by plan 01-05) |
| Command | `pytest tests/e2e/test_avatar_loop.py -q --space-url https://wolfdavid-japanese-learning-avatar.hf.space` |
| Result | **15 passed, 0 failed** (14 test functions; `test_silence_rejected` runs twice) β exit 0 |
| Turns that completed | typed turn Γ5 (incl. 20 wave-5 interactions in `test_no_remount`), push-to-talk Γ2 (WASM tier), replay, slower re-synthesis; silence and cafe noise produced 0 turns |
Every row in `01-VALIDATION.md`'s Per-Task Verification Map that names a
`tests/e2e/test_avatar_loop.py` node was green in this run. The two ASR rows ran on the WASM
tier because headless Chromium has no WebGPU adapter; neither tier touches the server.
Numbers this run recorded for plan 01-10's latency harness (all on the Space's CPU, one
visitor): `synthesis_ms` 3088β6724 for γγγ«γ‘γ― across four typed turns in the session,
10194β10292 for the 36-mora long sentence (normal / speedScale 0.75); dispatchβspeech-start
4170β7252 ms for γγγ«γ‘γ―; replay 0 ms and 0 requests; pageβready 4.0β16.8 s with
`tutor.vrm` 3.0β7.4 s of that. Wake was 0.2 s every time β the Space was already running, so
true cold-from-sleep remains 01-10's manual row.
## Why not the alternatives
- **`cpu-basic` (free)** β no longer available. Attempting it returned **HTTP 402 Payment Required**
(quoted verbatim above). This option is closed, not merely discouraged.
- **`sdk: static`** β free for everyone, but VOICEVOX cannot run in a browser, so there would be no
mora timings and AVTR-02 collapses entirely. ZeroGPU is also Gradio-SDK-only, so Phase 3 would
need a rewrite. Rejected.
- **PRO ($9/mo)** β **not purchased, and not needed**: the free ZeroGPU path worked. PRO does not
improve the anonymous **visitor's** 2 min/day quota, so it would change nothing about the
free-tier design that Phase 1 is built around. No payment was made by any step of this plan.
## GPU posture
Phase 1 calls **no** GPU function. `DISABLE_GPU=1` is to be set as a Space variable and every future
GPU entry point must raise when it is set. Proof lives in `tests/test_no_gpu_on_turn_path.py` (static
AST scan) and a `DISABLE_GPU=1` deployed E2E run.
A ZeroGPU Space that never calls `@spaces.GPU` runs its web process on CPU and consumes **zero**
visitor quota, so taking the ZeroGPU tier costs nothing at runtime in Phase 1. Plan 01-04 already
established that speech synthesis is CPU-only and AST-verified free of `spaces`/`torch` imports.
## ZeroGPU slot budget
**1 of 2 free ZeroGPU slots is now consumed by this Space. 1 slot remains** for the other four
HF-profile projects. Because `cpu-basic` Gradio creation is 402-gated on this account, that final
slot is the last free Gradio Space this account can create without PRO β spend it deliberately.
## Git remote
`git remote add space https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar`
Authentication is a cached fine-grained token held by the `hf` CLI (`hf auth whoami` reports
`WolfDavid`), scoped `repo.write` / `repo.content.read` / `repo.access.read` on the WolfDavid entity.
The token value is never written to this repository, any commit message, or any log.
**No application code has been pushed to this remote.** Deployment is plan 01-05's task and is gated
behind its own spike verdict.
## First deploy record (plan 01-05, 2026-09-05)
Five pushes were needed. Four failed, and **three of the four failures were platform
constraints that no amount of local testing would have surfaced** - the full local suite
(82 quick-loop + 15 browser tests) was green before the first push and stayed green
throughout. They are recorded here in full because each one is a trap for the next Space.
| # | Commit | Stage reached | What happened |
|---|---|---|---|
| 1 | `ad85cc8` | `BUILD_ERROR` | `ResolutionImpossible`: our `spaces==0.51.1` vs the builder's injected `spaces==0.51.3` |
| 2 | `6eeb816` | `RUNTIME_ERROR` | `ModuleNotFoundError: No module named 'japanese_avatar'` - the `src/` layout is not importable on a Space |
| 3 | `e3ebc72` | `RUNTIME_ERROR` | App launched, then exited. `GRADIO_HOT_RELOAD: Launching demo not found in __main__` |
| 4 | `8c9985e` | `RUNTIME_ERROR` | SIGTERMed by the platform: `No @spaces.GPU function detected during startup` |
| 5 | `4e3cac9` | **`RUNNING`** | Live, `hardware.current: zero-a10g` |
### Finding 1: do not pin `spaces` in `requirements.txt`
The Space builder generates its own pip command and appends its own pin:
```
pip install --no-cache-dir -r /tmp/requirements.txt "torch<=2.11.0" \
gradio[oauth,mcp]==6.22.0 "uvicorn>=0.14.0" "websockets>=10.4" spaces==0.51.3
```
Two exact pins for one package is unsatisfiable. **The platform owns the `spaces` version.**
`gradio==6.22.0` in `requirements.txt` is fine because the builder injects the *same*
version - it reads `sdk_version` from the README front-matter, which is exactly why that
front-matter correction had to land in the same push.
### Finding 2: a `src/` layout package is not importable on a Space
A Space installs `requirements.txt` and nothing else; the repository is never
`pip install`ed. `app.py` now puts `src/` on `sys.path` itself rather than relying on a
`PYTHONPATH` Space variable, so a clean checkout runs with no platform configuration.
### Finding 3: the launched `Blocks` must be a module-level attribute named `demo`
Hugging Face runs the app under `gradio.utils.SpacesReloader`, whose `postrun()` calls
`getattr(watch_module, demo_name)` on every reload check. A `build_app()` factory that
keeps the `Blocks` local logs `GRADIO_HOT_RELOAD: Launching demo not found in __main__.
Using 'demo'` and then dies with **no traceback**. `demo = build_app()` at module level
is not a style preference on a Space; it is the contract.
### Finding 4 (the important one): ZeroGPU refuses to run without a `@spaces.GPU` function
```
runtime.errorMessage: "No @spaces.GPU function detected during startup"
```
The app had already bound port 7860 and printed `Running on local URL` when the platform
SIGTERMed it. **A ZeroGPU Space must declare at least one GPU entry point or it will not
run at all**, and since `cpu-basic` Gradio Spaces are 402-gated on this account (above),
there is no other free tier to move to.
`app.py::zerogpu_probe` exists solely to satisfy that startup scan. Nothing calls it, it
is not on the turn path, and it raises when `DISABLE_GPU=1` - which is set on this Space.
**This changes what SC-4 can claim.** "No `@spaces.GPU` function exists anywhere" is no
longer an available proof; the platform has taken it off the table. The honest and
equally strong assertion is "the turn loop completes without ever reaching one", which
the `DISABLE_GPU=1` deployed run demonstrates directly, since reaching it would raise.
Any later plan that asserts the absence of a GPU decorator will fail for a reason that
has nothing to do with this project's design.
### Finding 5: Gradio 6 serves an SSR shell, so do not grep the HTML for `elem_id`s
`curl -s $SPACE_URL | grep -c vrm-stage` returns **0** on a working deployment. Gradio 6
server-side-renders a shell and the component tree arrives from `GET /config`. The
equivalent check that does work:
```
curl -s https://wolfdavid-japanese-learning-avatar.hf.space/config | grep -c vrm-stage # 1
```
### Git LFS resolved correctly on the Space
Verified against the live app, not assumed - a silently unpushed pointer is the classic
failure here:
| Asset | HTTP | Bytes served |
|---|---|---|
| `/gradio_api/file=avatar/assets/tutor.vrm` | 200 | **10,776,032** (`model/vrml`) |
| `/gradio_api/file=avatar/assets/demo-konnichiwa.wav` | 200 | 50,732 |
| `/gradio_api/file=avatar/avatar.js` | 200 | 5,138 |
| `/gradio_api/file=avatar/stage.html` | 200 | 9,271 |
All 12 LFS objects (181 MB) uploaded on the first push and the Space checks them out as
real files, not 130-byte pointers.
|