WolfDavid's picture
docs(01-05): record zero-a10g as the applied hardware
7e0e495
|
Raw History Blame
15.6 kB
# Hosting (DPLY-01)
**Decided:** 2026-08-27
**Chosen path:** `zerogpu-free`
| Field | Value |
|---|---|
| Space | WolfDavid/japanese-learning-avatar |
| Space page | https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar |
| SPACE_URL (direct, used by E2E `--space-url`) | https://wolfdavid-japanese-learning-avatar.hf.space |
| Visibility | public (`private: false`) |
| Hardware requested | `zero-a10g` (ZeroGPU) |
| Hardware applied | **`zero-a10g`** — attached on the first successful deploy (`4e3cac9`, 2026-09-05). Was `null` before that; see "Why `hardware.current` is null" below for why that was correct at the time. |
| sdk / sdk_version | gradio / **6.22.0** — corrected by plan 01-05's `README.md`, which is now the Space manifest. See "The front-matter correction, applied". |
| python_version | **3.12.12** — corrected by plan 01-05. |
| Sleep timeout | 48 h (`gcTimeout: 172800`) |
| Space repo SHA at time of writing | `b70455dcf67ff7db6b1f48f08153982d37a63dd2` (two files: `.gitattributes`, `README.md`) |
| Created | 2026-08-27T02:13:12Z |
## The decisive finding: `cpu-basic` is no longer creatable, and it fails with a payment error
Creating this Space on the default `cpu-basic` hardware was **rejected with HTTP 402 Payment
Required**. The message, verbatim:
> Static Spaces are free for everyone, but hosting Gradio and Docker Spaces on free cpu-basic
> requires a PRO subscription.
Re-issuing the identical creation request with ZeroGPU (`space_hardware="zero-a10g"`) **succeeded
immediately, with no payment prompt and no charge.**
This is empirical confirmation of the claim in `01-RESEARCH.md` and `CLAUDE.md` that Gradio Spaces
on a free personal account are PRO-gated *except* for the ZeroGPU carve-out. It is no longer a
documentation reading — it is an observed platform response on this account.
**Consequence:** ZeroGPU is now the **only** free hosting path for this project. A future phase
**cannot** silently fall back to `cpu-basic` to save a ZeroGPU slot — that fallback does not exist
for a Gradio Space on a non-PRO personal account. Any plan that proposes it is proposing a $9/mo
purchase.
## Account eligibility evidence (verified against the live public API, 2026-08-27)
Queried in this task, not quoted from RESEARCH:
- `isPro: false`; account created `2023-11-27T01:45:12.000Z` — about 2.7 years old, so the ZeroGPU
">30 days in good standing" criterion is satisfied with enormous margin.
- Email: verified (account is in good standing; ZeroGPU allocation was accepted, which requires it).
- **8 pre-existing Gradio Spaces, every one of them `requested: cpu-basic`, and zero on ZeroGPU** —
so both free ZeroGPU slots were open before this Space was created:
| Space | sdk_version | python | hardware requested | stage |
|---|---|---|---|---|
| `WolfDavid/dino` | 6.9.0 | 3.11 | cpu-basic | RUNNING |
| `WolfDavid/fea-surrogate` | 5.9.1 | 3.11 | cpu-basic | SLEEPING |
| `WolfDavid/mechspec-qwen-demo` | 5.9.1 | 3.11 | cpu-basic | SLEEPING |
| `WolfDavid/agentic-market-analyzer` | 5.9.1 | 3.11 | cpu-basic | SLEEPING |
| `WolfDavid/log-anomaly-detector` | 5.9.1 | 3.11 | cpu-basic | RUNNING |
| `WolfDavid/vision-edge` | 5.9.1 | 3.11 | cpu-basic | SLEEPING |
| `WolfDavid/nllb-translator` | 5.9.1 | 3.11 | cpu-basic | SLEEPING |
| `WolfDavid/blip-captioner` | 5.9.1 | 3.11 | cpu-basic | SLEEPING |
- `WolfDavid/dino` runs `sdk_version: 6.9.0`, so Gradio 6 is proven to work on this account.
- Note for anyone re-running this check: **`runtime.hardware.current` is `null` for every SLEEPING
Space.** `current` reports what is running right now, not what the Space is entitled to. Read
`runtime.hardware.requested` when auditing hardware tiers, or six of these eight Spaces look
unallocated.
- Note also that `GET /api/spaces?author=WolfDavid` still returned only the original 8 shortly after
creation; the list endpoint lags. `GET /api/spaces/WolfDavid/japanese-learning-avatar` returns the
new Space immediately and is the authoritative read.
## Why `hardware.current` is null, and why that is correct
```
"runtime": { "stage": "NO_APP_FILE",
"hardware": { "current": null, "requested": "zero-a10g" },
"gcTimeout": 172800 }
```
The Space repo contains only `.gitattributes` and an auto-generated `README.md`. With no `app.py`
there is nothing to schedule, so no hardware is attached. **This is the correct pre-deploy state and
is not a failure.** `requested: zero-a10g` is the binding fact — it is the tier the Space will start
on the moment plan 01-05 pushes an app.
Reachability today, measured three times consecutively with identical results:
| URL | HTTP | Meaning |
|---|---|---|
| https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar | **200** | The Space exists and is public |
| https://wolfdavid-japanese-learning-avatar.hf.space | **503** | `Your space is in error, check its status on hf.co` — no `app.py` yet |
The direct subdomain **is** provisioned (`runtime.domains[0].stage: READY`, `subdomain:
wolfdavid-japanese-learning-avatar`), it simply has nothing to serve. **Plan 01-05 turns this 503
into a 200**, and `tests/e2e/test_avatar_loop.py::test_space_reachable` — which plan 01-05 creates —
is what actually satisfies DPLY-01. DPLY-01 is **not** satisfied by this plan.
## The front-matter correction, applied (plan 01-05)
Space creation auto-generated a `README.md` whose front-matter **conflicted with the pinned
stack**. Plan 01-05 wrote a repository `README.md` that replaces it:
| Key | Auto-generated value | Value now shipped | Source of the requirement |
|---|---|---|---|
| `sdk_version` | `6.26.0` | `6.22.0` | `requirements.txt` line 1 pins `gradio==6.22.0`; `CLAUDE.md` says pin exactly |
| `python_version` | `'3.12'` | `3.12.12` | ZeroGPU provides only 3.10.13 and 3.12.12; `.python-version` is `3.12.12` |
| `emoji` | `🐠` | `🗾` | cosmetic; the auto-generated value was also mojibake in the API response |
| `colorFrom` / `colorTo` | `purple` / `gray` | `pink` / `indigo` | cosmetic |
| `license` | absent | `mit` | the repository's own code is MIT; third-party asset terms are separate |
| `short_description` | absent | present | shown on the profile card a recruiter sees first |
`hf_oauth` was deliberately **not** added: it belongs to Phase 4, and every extra consent scope is
a deterrent on the Space's first-load screen.
Left uncorrected, the Space SDK would have installed Gradio 6.26.0 while `requirements.txt` asked
for 6.22.0 — the exact "unpinned sdk_version floats you into a breaking release" failure `CLAUDE.md`
warns about, arrived at from the other direction.
## How the unrelated histories were reconciled (plan 01-05)
The Space repo was created with its own root commit — `b70455d`, "initial commit", two files
(`.gitattributes`, `README.md`) — while the local repository has a completely separate root and
~44 commits. `git merge-base master space/main` returned nothing: **the two histories share no
ancestor**, so a plain `git push space master:main` is rejected as non-fast-forward.
**Chosen: `git merge --allow-unrelated-histories space/main`, not a force-push.** Both files
conflicted and both were resolved deliberately:
- **`README.md` — resolved to ours.** The remote copy is the boilerplate quoted above and is
precisely what this plan exists to replace; there is nothing in it to preserve.
- **`.gitattributes` — resolved to the union.** Ours (plan 01-02: `*.vrm *.vvm *.wav *.onnx`)
is kept verbatim and Hugging Face's 35 default LFS patterns are appended below it. The
defaults cost nothing today — every currently tracked binary is already an LFS object under
the narrow patterns — and they arm the formats later phases will push (`.safetensors`,
`.bin`, `.pt`, `.npz`) against the 10 MiB non-LFS rejection. Dropping them to keep the file
minimal would have traded a real safety net for tidiness.
A force-push would have worked and lost nothing of substance, but the merge keeps the Space's
own creation commit in the history, which is the provenance record for the ZeroGPU allocation.
## GPU posture on the deployed Space
`DISABLE_GPU=1` **is set** as a Space variable (set by plan 01-05 via
`HfApi.add_space_variable`, 2026-09-05, description recorded on the variable itself). Phase 1
calls no GPU function, so this costs nothing and makes SC-4's demonstration honest from the
first deploy rather than retrofitted at the end. Read it back with:
```
python -c "from huggingface_hub import HfApi; print(HfApi().get_space_variables('WolfDavid/japanese-learning-avatar'))"
```
`AVATAR_TRANSPORT` is deliberately **not** set, so the Space uses the component's default,
`inline` — the transport the spike is testing. Setting it to `iframe` is the entire fallback.
## Why not the alternatives
- **`cpu-basic` (free)** — no longer available. Attempting it returned **HTTP 402 Payment Required**
(quoted verbatim above). This option is closed, not merely discouraged.
- **`sdk: static`** — free for everyone, but VOICEVOX cannot run in a browser, so there would be no
mora timings and AVTR-02 collapses entirely. ZeroGPU is also Gradio-SDK-only, so Phase 3 would
need a rewrite. Rejected.
- **PRO ($9/mo)** — **not purchased, and not needed**: the free ZeroGPU path worked. PRO does not
improve the anonymous **visitor's** 2 min/day quota, so it would change nothing about the
free-tier design that Phase 1 is built around. No payment was made by any step of this plan.
## GPU posture
Phase 1 calls **no** GPU function. `DISABLE_GPU=1` is to be set as a Space variable and every future
GPU entry point must raise when it is set. Proof lives in `tests/test_no_gpu_on_turn_path.py` (static
AST scan) and a `DISABLE_GPU=1` deployed E2E run.
A ZeroGPU Space that never calls `@spaces.GPU` runs its web process on CPU and consumes **zero**
visitor quota, so taking the ZeroGPU tier costs nothing at runtime in Phase 1. Plan 01-04 already
established that speech synthesis is CPU-only and AST-verified free of `spaces`/`torch` imports.
## ZeroGPU slot budget
**1 of 2 free ZeroGPU slots is now consumed by this Space. 1 slot remains** for the other four
HF-profile projects. Because `cpu-basic` Gradio creation is 402-gated on this account, that final
slot is the last free Gradio Space this account can create without PRO — spend it deliberately.
## Git remote
`git remote add space https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar`
Authentication is a cached fine-grained token held by the `hf` CLI (`hf auth whoami` reports
`WolfDavid`), scoped `repo.write` / `repo.content.read` / `repo.access.read` on the WolfDavid entity.
The token value is never written to this repository, any commit message, or any log.
**No application code has been pushed to this remote.** Deployment is plan 01-05's task and is gated
behind its own spike verdict.
## First deploy record (plan 01-05, 2026-09-05)
Five pushes were needed. Four failed, and **three of the four failures were platform
constraints that no amount of local testing would have surfaced** - the full local suite
(82 quick-loop + 15 browser tests) was green before the first push and stayed green
throughout. They are recorded here in full because each one is a trap for the next Space.
| # | Commit | Stage reached | What happened |
|---|---|---|---|
| 1 | `ad85cc8` | `BUILD_ERROR` | `ResolutionImpossible`: our `spaces==0.51.1` vs the builder's injected `spaces==0.51.3` |
| 2 | `6eeb816` | `RUNTIME_ERROR` | `ModuleNotFoundError: No module named 'japanese_avatar'` - the `src/` layout is not importable on a Space |
| 3 | `e3ebc72` | `RUNTIME_ERROR` | App launched, then exited. `GRADIO_HOT_RELOAD: Launching demo not found in __main__` |
| 4 | `8c9985e` | `RUNTIME_ERROR` | SIGTERMed by the platform: `No @spaces.GPU function detected during startup` |
| 5 | `4e3cac9` | **`RUNNING`** | Live, `hardware.current: zero-a10g` |
### Finding 1: do not pin `spaces` in `requirements.txt`
The Space builder generates its own pip command and appends its own pin:
```
pip install --no-cache-dir -r /tmp/requirements.txt "torch<=2.11.0" \
gradio[oauth,mcp]==6.22.0 "uvicorn>=0.14.0" "websockets>=10.4" spaces==0.51.3
```
Two exact pins for one package is unsatisfiable. **The platform owns the `spaces` version.**
`gradio==6.22.0` in `requirements.txt` is fine because the builder injects the *same*
version - it reads `sdk_version` from the README front-matter, which is exactly why that
front-matter correction had to land in the same push.
### Finding 2: a `src/` layout package is not importable on a Space
A Space installs `requirements.txt` and nothing else; the repository is never
`pip install`ed. `app.py` now puts `src/` on `sys.path` itself rather than relying on a
`PYTHONPATH` Space variable, so a clean checkout runs with no platform configuration.
### Finding 3: the launched `Blocks` must be a module-level attribute named `demo`
Hugging Face runs the app under `gradio.utils.SpacesReloader`, whose `postrun()` calls
`getattr(watch_module, demo_name)` on every reload check. A `build_app()` factory that
keeps the `Blocks` local logs `GRADIO_HOT_RELOAD: Launching demo not found in __main__.
Using 'demo'` and then dies with **no traceback**. `demo = build_app()` at module level
is not a style preference on a Space; it is the contract.
### Finding 4 (the important one): ZeroGPU refuses to run without a `@spaces.GPU` function
```
runtime.errorMessage: "No @spaces.GPU function detected during startup"
```
The app had already bound port 7860 and printed `Running on local URL` when the platform
SIGTERMed it. **A ZeroGPU Space must declare at least one GPU entry point or it will not
run at all**, and since `cpu-basic` Gradio Spaces are 402-gated on this account (above),
there is no other free tier to move to.
`app.py::zerogpu_probe` exists solely to satisfy that startup scan. Nothing calls it, it
is not on the turn path, and it raises when `DISABLE_GPU=1` - which is set on this Space.
**This changes what SC-4 can claim.** "No `@spaces.GPU` function exists anywhere" is no
longer an available proof; the platform has taken it off the table. The honest and
equally strong assertion is "the turn loop completes without ever reaching one", which
the `DISABLE_GPU=1` deployed run demonstrates directly, since reaching it would raise.
Any later plan that asserts the absence of a GPU decorator will fail for a reason that
has nothing to do with this project's design.
### Finding 5: Gradio 6 serves an SSR shell, so do not grep the HTML for `elem_id`s
`curl -s $SPACE_URL | grep -c vrm-stage` returns **0** on a working deployment. Gradio 6
server-side-renders a shell and the component tree arrives from `GET /config`. The
equivalent check that does work:
```
curl -s https://wolfdavid-japanese-learning-avatar.hf.space/config | grep -c vrm-stage # 1
```
### Git LFS resolved correctly on the Space
Verified against the live app, not assumed - a silently unpushed pointer is the classic
failure here:
| Asset | HTTP | Bytes served |
|---|---|---|
| `/gradio_api/file=avatar/assets/tutor.vrm` | 200 | **10,776,032** (`model/vrml`) |
| `/gradio_api/file=avatar/assets/demo-konnichiwa.wav` | 200 | 50,732 |
| `/gradio_api/file=avatar/avatar.js` | 200 | 5,138 |
| `/gradio_api/file=avatar/stage.html` | 200 | 9,271 |
All 12 LFS objects (181 MB) uploaded on the first push and the Space checks them out as
real files, not 130-byte pointers.