Spaces:
Running on Zero
Running on Zero
|
Download docs/HOSTING.md from WolfDavid/japanese-learning-avatar: direct link, hf CLI and curl.
- Browser
- Download file 15.6 kB
-
https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar/resolve/6bacf187a16aedc18cad3b252ef933c6b3b90af2/docs/HOSTING.md
- Command line
-
hf download hf://spaces/WolfDavid/japanese-learning-avatar@6bacf187a16aedc18cad3b252ef933c6b3b90af2/docs/HOSTING.md
-
curl -L -o HOSTING.md https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar/resolve/6bacf187a16aedc18cad3b252ef933c6b3b90af2/docs/HOSTING.md
15.6 kB
| # Hosting (DPLY-01) | |
| **Decided:** 2026-08-27 | |
| **Chosen path:** `zerogpu-free` | |
| | Field | Value | | |
| |---|---| | |
| | Space | WolfDavid/japanese-learning-avatar | | |
| | Space page | https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar | | |
| | SPACE_URL (direct, used by E2E `--space-url`) | https://wolfdavid-japanese-learning-avatar.hf.space | | |
| | Visibility | public (`private: false`) | | |
| | Hardware requested | `zero-a10g` (ZeroGPU) | | |
| | Hardware applied | **`zero-a10g`** — attached on the first successful deploy (`4e3cac9`, 2026-09-05). Was `null` before that; see "Why `hardware.current` is null" below for why that was correct at the time. | | |
| | sdk / sdk_version | gradio / **6.22.0** — corrected by plan 01-05's `README.md`, which is now the Space manifest. See "The front-matter correction, applied". | | |
| | python_version | **3.12.12** — corrected by plan 01-05. | | |
| | Sleep timeout | 48 h (`gcTimeout: 172800`) | | |
| | Space repo SHA at time of writing | `b70455dcf67ff7db6b1f48f08153982d37a63dd2` (two files: `.gitattributes`, `README.md`) | | |
| | Created | 2026-08-27T02:13:12Z | | |
| ## The decisive finding: `cpu-basic` is no longer creatable, and it fails with a payment error | |
| Creating this Space on the default `cpu-basic` hardware was **rejected with HTTP 402 Payment | |
| Required**. The message, verbatim: | |
| > Static Spaces are free for everyone, but hosting Gradio and Docker Spaces on free cpu-basic | |
| > requires a PRO subscription. | |
| Re-issuing the identical creation request with ZeroGPU (`space_hardware="zero-a10g"`) **succeeded | |
| immediately, with no payment prompt and no charge.** | |
| This is empirical confirmation of the claim in `01-RESEARCH.md` and `CLAUDE.md` that Gradio Spaces | |
| on a free personal account are PRO-gated *except* for the ZeroGPU carve-out. It is no longer a | |
| documentation reading — it is an observed platform response on this account. | |
| **Consequence:** ZeroGPU is now the **only** free hosting path for this project. A future phase | |
| **cannot** silently fall back to `cpu-basic` to save a ZeroGPU slot — that fallback does not exist | |
| for a Gradio Space on a non-PRO personal account. Any plan that proposes it is proposing a $9/mo | |
| purchase. | |
| ## Account eligibility evidence (verified against the live public API, 2026-08-27) | |
| Queried in this task, not quoted from RESEARCH: | |
| - `isPro: false`; account created `2023-11-27T01:45:12.000Z` — about 2.7 years old, so the ZeroGPU | |
| ">30 days in good standing" criterion is satisfied with enormous margin. | |
| - Email: verified (account is in good standing; ZeroGPU allocation was accepted, which requires it). | |
| - **8 pre-existing Gradio Spaces, every one of them `requested: cpu-basic`, and zero on ZeroGPU** — | |
| so both free ZeroGPU slots were open before this Space was created: | |
| | Space | sdk_version | python | hardware requested | stage | | |
| |---|---|---|---|---| | |
| | `WolfDavid/dino` | 6.9.0 | 3.11 | cpu-basic | RUNNING | | |
| | `WolfDavid/fea-surrogate` | 5.9.1 | 3.11 | cpu-basic | SLEEPING | | |
| | `WolfDavid/mechspec-qwen-demo` | 5.9.1 | 3.11 | cpu-basic | SLEEPING | | |
| | `WolfDavid/agentic-market-analyzer` | 5.9.1 | 3.11 | cpu-basic | SLEEPING | | |
| | `WolfDavid/log-anomaly-detector` | 5.9.1 | 3.11 | cpu-basic | RUNNING | | |
| | `WolfDavid/vision-edge` | 5.9.1 | 3.11 | cpu-basic | SLEEPING | | |
| | `WolfDavid/nllb-translator` | 5.9.1 | 3.11 | cpu-basic | SLEEPING | | |
| | `WolfDavid/blip-captioner` | 5.9.1 | 3.11 | cpu-basic | SLEEPING | | |
| - `WolfDavid/dino` runs `sdk_version: 6.9.0`, so Gradio 6 is proven to work on this account. | |
| - Note for anyone re-running this check: **`runtime.hardware.current` is `null` for every SLEEPING | |
| Space.** `current` reports what is running right now, not what the Space is entitled to. Read | |
| `runtime.hardware.requested` when auditing hardware tiers, or six of these eight Spaces look | |
| unallocated. | |
| - Note also that `GET /api/spaces?author=WolfDavid` still returned only the original 8 shortly after | |
| creation; the list endpoint lags. `GET /api/spaces/WolfDavid/japanese-learning-avatar` returns the | |
| new Space immediately and is the authoritative read. | |
| ## Why `hardware.current` is null, and why that is correct | |
| ``` | |
| "runtime": { "stage": "NO_APP_FILE", | |
| "hardware": { "current": null, "requested": "zero-a10g" }, | |
| "gcTimeout": 172800 } | |
| ``` | |
| The Space repo contains only `.gitattributes` and an auto-generated `README.md`. With no `app.py` | |
| there is nothing to schedule, so no hardware is attached. **This is the correct pre-deploy state and | |
| is not a failure.** `requested: zero-a10g` is the binding fact — it is the tier the Space will start | |
| on the moment plan 01-05 pushes an app. | |
| Reachability today, measured three times consecutively with identical results: | |
| | URL | HTTP | Meaning | | |
| |---|---|---| | |
| | https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar | **200** | The Space exists and is public | | |
| | https://wolfdavid-japanese-learning-avatar.hf.space | **503** | `Your space is in error, check its status on hf.co` — no `app.py` yet | | |
| The direct subdomain **is** provisioned (`runtime.domains[0].stage: READY`, `subdomain: | |
| wolfdavid-japanese-learning-avatar`), it simply has nothing to serve. **Plan 01-05 turns this 503 | |
| into a 200**, and `tests/e2e/test_avatar_loop.py::test_space_reachable` — which plan 01-05 creates — | |
| is what actually satisfies DPLY-01. DPLY-01 is **not** satisfied by this plan. | |
| ## The front-matter correction, applied (plan 01-05) | |
| Space creation auto-generated a `README.md` whose front-matter **conflicted with the pinned | |
| stack**. Plan 01-05 wrote a repository `README.md` that replaces it: | |
| | Key | Auto-generated value | Value now shipped | Source of the requirement | | |
| |---|---|---|---| | |
| | `sdk_version` | `6.26.0` | `6.22.0` | `requirements.txt` line 1 pins `gradio==6.22.0`; `CLAUDE.md` says pin exactly | | |
| | `python_version` | `'3.12'` | `3.12.12` | ZeroGPU provides only 3.10.13 and 3.12.12; `.python-version` is `3.12.12` | | |
| | `emoji` | `🐠` | `🗾` | cosmetic; the auto-generated value was also mojibake in the API response | | |
| | `colorFrom` / `colorTo` | `purple` / `gray` | `pink` / `indigo` | cosmetic | | |
| | `license` | absent | `mit` | the repository's own code is MIT; third-party asset terms are separate | | |
| | `short_description` | absent | present | shown on the profile card a recruiter sees first | | |
| `hf_oauth` was deliberately **not** added: it belongs to Phase 4, and every extra consent scope is | |
| a deterrent on the Space's first-load screen. | |
| Left uncorrected, the Space SDK would have installed Gradio 6.26.0 while `requirements.txt` asked | |
| for 6.22.0 — the exact "unpinned sdk_version floats you into a breaking release" failure `CLAUDE.md` | |
| warns about, arrived at from the other direction. | |
| ## How the unrelated histories were reconciled (plan 01-05) | |
| The Space repo was created with its own root commit — `b70455d`, "initial commit", two files | |
| (`.gitattributes`, `README.md`) — while the local repository has a completely separate root and | |
| ~44 commits. `git merge-base master space/main` returned nothing: **the two histories share no | |
| ancestor**, so a plain `git push space master:main` is rejected as non-fast-forward. | |
| **Chosen: `git merge --allow-unrelated-histories space/main`, not a force-push.** Both files | |
| conflicted and both were resolved deliberately: | |
| - **`README.md` — resolved to ours.** The remote copy is the boilerplate quoted above and is | |
| precisely what this plan exists to replace; there is nothing in it to preserve. | |
| - **`.gitattributes` — resolved to the union.** Ours (plan 01-02: `*.vrm *.vvm *.wav *.onnx`) | |
| is kept verbatim and Hugging Face's 35 default LFS patterns are appended below it. The | |
| defaults cost nothing today — every currently tracked binary is already an LFS object under | |
| the narrow patterns — and they arm the formats later phases will push (`.safetensors`, | |
| `.bin`, `.pt`, `.npz`) against the 10 MiB non-LFS rejection. Dropping them to keep the file | |
| minimal would have traded a real safety net for tidiness. | |
| A force-push would have worked and lost nothing of substance, but the merge keeps the Space's | |
| own creation commit in the history, which is the provenance record for the ZeroGPU allocation. | |
| ## GPU posture on the deployed Space | |
| `DISABLE_GPU=1` **is set** as a Space variable (set by plan 01-05 via | |
| `HfApi.add_space_variable`, 2026-09-05, description recorded on the variable itself). Phase 1 | |
| calls no GPU function, so this costs nothing and makes SC-4's demonstration honest from the | |
| first deploy rather than retrofitted at the end. Read it back with: | |
| ``` | |
| python -c "from huggingface_hub import HfApi; print(HfApi().get_space_variables('WolfDavid/japanese-learning-avatar'))" | |
| ``` | |
| `AVATAR_TRANSPORT` is deliberately **not** set, so the Space uses the component's default, | |
| `inline` — the transport the spike is testing. Setting it to `iframe` is the entire fallback. | |
| ## Why not the alternatives | |
| - **`cpu-basic` (free)** — no longer available. Attempting it returned **HTTP 402 Payment Required** | |
| (quoted verbatim above). This option is closed, not merely discouraged. | |
| - **`sdk: static`** — free for everyone, but VOICEVOX cannot run in a browser, so there would be no | |
| mora timings and AVTR-02 collapses entirely. ZeroGPU is also Gradio-SDK-only, so Phase 3 would | |
| need a rewrite. Rejected. | |
| - **PRO ($9/mo)** — **not purchased, and not needed**: the free ZeroGPU path worked. PRO does not | |
| improve the anonymous **visitor's** 2 min/day quota, so it would change nothing about the | |
| free-tier design that Phase 1 is built around. No payment was made by any step of this plan. | |
| ## GPU posture | |
| Phase 1 calls **no** GPU function. `DISABLE_GPU=1` is to be set as a Space variable and every future | |
| GPU entry point must raise when it is set. Proof lives in `tests/test_no_gpu_on_turn_path.py` (static | |
| AST scan) and a `DISABLE_GPU=1` deployed E2E run. | |
| A ZeroGPU Space that never calls `@spaces.GPU` runs its web process on CPU and consumes **zero** | |
| visitor quota, so taking the ZeroGPU tier costs nothing at runtime in Phase 1. Plan 01-04 already | |
| established that speech synthesis is CPU-only and AST-verified free of `spaces`/`torch` imports. | |
| ## ZeroGPU slot budget | |
| **1 of 2 free ZeroGPU slots is now consumed by this Space. 1 slot remains** for the other four | |
| HF-profile projects. Because `cpu-basic` Gradio creation is 402-gated on this account, that final | |
| slot is the last free Gradio Space this account can create without PRO — spend it deliberately. | |
| ## Git remote | |
| `git remote add space https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar` | |
| Authentication is a cached fine-grained token held by the `hf` CLI (`hf auth whoami` reports | |
| `WolfDavid`), scoped `repo.write` / `repo.content.read` / `repo.access.read` on the WolfDavid entity. | |
| The token value is never written to this repository, any commit message, or any log. | |
| **No application code has been pushed to this remote.** Deployment is plan 01-05's task and is gated | |
| behind its own spike verdict. | |
| ## First deploy record (plan 01-05, 2026-09-05) | |
| Five pushes were needed. Four failed, and **three of the four failures were platform | |
| constraints that no amount of local testing would have surfaced** - the full local suite | |
| (82 quick-loop + 15 browser tests) was green before the first push and stayed green | |
| throughout. They are recorded here in full because each one is a trap for the next Space. | |
| | # | Commit | Stage reached | What happened | | |
| |---|---|---|---| | |
| | 1 | `ad85cc8` | `BUILD_ERROR` | `ResolutionImpossible`: our `spaces==0.51.1` vs the builder's injected `spaces==0.51.3` | | |
| | 2 | `6eeb816` | `RUNTIME_ERROR` | `ModuleNotFoundError: No module named 'japanese_avatar'` - the `src/` layout is not importable on a Space | | |
| | 3 | `e3ebc72` | `RUNTIME_ERROR` | App launched, then exited. `GRADIO_HOT_RELOAD: Launching demo not found in __main__` | | |
| | 4 | `8c9985e` | `RUNTIME_ERROR` | SIGTERMed by the platform: `No @spaces.GPU function detected during startup` | | |
| | 5 | `4e3cac9` | **`RUNNING`** | Live, `hardware.current: zero-a10g` | | |
| ### Finding 1: do not pin `spaces` in `requirements.txt` | |
| The Space builder generates its own pip command and appends its own pin: | |
| ``` | |
| pip install --no-cache-dir -r /tmp/requirements.txt "torch<=2.11.0" \ | |
| gradio[oauth,mcp]==6.22.0 "uvicorn>=0.14.0" "websockets>=10.4" spaces==0.51.3 | |
| ``` | |
| Two exact pins for one package is unsatisfiable. **The platform owns the `spaces` version.** | |
| `gradio==6.22.0` in `requirements.txt` is fine because the builder injects the *same* | |
| version - it reads `sdk_version` from the README front-matter, which is exactly why that | |
| front-matter correction had to land in the same push. | |
| ### Finding 2: a `src/` layout package is not importable on a Space | |
| A Space installs `requirements.txt` and nothing else; the repository is never | |
| `pip install`ed. `app.py` now puts `src/` on `sys.path` itself rather than relying on a | |
| `PYTHONPATH` Space variable, so a clean checkout runs with no platform configuration. | |
| ### Finding 3: the launched `Blocks` must be a module-level attribute named `demo` | |
| Hugging Face runs the app under `gradio.utils.SpacesReloader`, whose `postrun()` calls | |
| `getattr(watch_module, demo_name)` on every reload check. A `build_app()` factory that | |
| keeps the `Blocks` local logs `GRADIO_HOT_RELOAD: Launching demo not found in __main__. | |
| Using 'demo'` and then dies with **no traceback**. `demo = build_app()` at module level | |
| is not a style preference on a Space; it is the contract. | |
| ### Finding 4 (the important one): ZeroGPU refuses to run without a `@spaces.GPU` function | |
| ``` | |
| runtime.errorMessage: "No @spaces.GPU function detected during startup" | |
| ``` | |
| The app had already bound port 7860 and printed `Running on local URL` when the platform | |
| SIGTERMed it. **A ZeroGPU Space must declare at least one GPU entry point or it will not | |
| run at all**, and since `cpu-basic` Gradio Spaces are 402-gated on this account (above), | |
| there is no other free tier to move to. | |
| `app.py::zerogpu_probe` exists solely to satisfy that startup scan. Nothing calls it, it | |
| is not on the turn path, and it raises when `DISABLE_GPU=1` - which is set on this Space. | |
| **This changes what SC-4 can claim.** "No `@spaces.GPU` function exists anywhere" is no | |
| longer an available proof; the platform has taken it off the table. The honest and | |
| equally strong assertion is "the turn loop completes without ever reaching one", which | |
| the `DISABLE_GPU=1` deployed run demonstrates directly, since reaching it would raise. | |
| Any later plan that asserts the absence of a GPU decorator will fail for a reason that | |
| has nothing to do with this project's design. | |
| ### Finding 5: Gradio 6 serves an SSR shell, so do not grep the HTML for `elem_id`s | |
| `curl -s $SPACE_URL | grep -c vrm-stage` returns **0** on a working deployment. Gradio 6 | |
| server-side-renders a shell and the component tree arrives from `GET /config`. The | |
| equivalent check that does work: | |
| ``` | |
| curl -s https://wolfdavid-japanese-learning-avatar.hf.space/config | grep -c vrm-stage # 1 | |
| ``` | |
| ### Git LFS resolved correctly on the Space | |
| Verified against the live app, not assumed - a silently unpushed pointer is the classic | |
| failure here: | |
| | Asset | HTTP | Bytes served | | |
| |---|---|---| | |
| | `/gradio_api/file=avatar/assets/tutor.vrm` | 200 | **10,776,032** (`model/vrml`) | | |
| | `/gradio_api/file=avatar/assets/demo-konnichiwa.wav` | 200 | 50,732 | | |
| | `/gradio_api/file=avatar/avatar.js` | 200 | 5,138 | | |
| | `/gradio_api/file=avatar/stage.html` | 200 | 9,271 | | |
| All 12 LFS objects (181 MB) uploaded on the first push and the Space checks them out as | |
| real files, not 130-byte pointers. | |