WolfDavid's picture
docs(01-05): record zero-a10g as the applied hardware
7e0e495
|
Raw
History Blame
15.6 kB

Hosting (DPLY-01)

Decided: 2026-08-27 Chosen path: zerogpu-free

Field Value
Space WolfDavid/japanese-learning-avatar
Space page https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar
SPACE_URL (direct, used by E2E --space-url) https://wolfdavid-japanese-learning-avatar.hf.space
Visibility public (private: false)
Hardware requested zero-a10g (ZeroGPU)
Hardware applied zero-a10g β€” attached on the first successful deploy (4e3cac9, 2026-09-05). Was null before that; see "Why hardware.current is null" below for why that was correct at the time.
sdk / sdk_version gradio / 6.22.0 β€” corrected by plan 01-05's README.md, which is now the Space manifest. See "The front-matter correction, applied".
python_version 3.12.12 β€” corrected by plan 01-05.
Sleep timeout 48 h (gcTimeout: 172800)
Space repo SHA at time of writing b70455dcf67ff7db6b1f48f08153982d37a63dd2 (two files: .gitattributes, README.md)
Created 2026-08-27T02:13:12Z

The decisive finding: cpu-basic is no longer creatable, and it fails with a payment error

Creating this Space on the default cpu-basic hardware was rejected with HTTP 402 Payment Required. The message, verbatim:

Static Spaces are free for everyone, but hosting Gradio and Docker Spaces on free cpu-basic requires a PRO subscription.

Re-issuing the identical creation request with ZeroGPU (space_hardware="zero-a10g") succeeded immediately, with no payment prompt and no charge.

This is empirical confirmation of the claim in 01-RESEARCH.md and CLAUDE.md that Gradio Spaces on a free personal account are PRO-gated except for the ZeroGPU carve-out. It is no longer a documentation reading β€” it is an observed platform response on this account.

Consequence: ZeroGPU is now the only free hosting path for this project. A future phase cannot silently fall back to cpu-basic to save a ZeroGPU slot β€” that fallback does not exist for a Gradio Space on a non-PRO personal account. Any plan that proposes it is proposing a $9/mo purchase.

Account eligibility evidence (verified against the live public API, 2026-08-27)

Queried in this task, not quoted from RESEARCH:

  • isPro: false; account created 2023-11-27T01:45:12.000Z β€” about 2.7 years old, so the ZeroGPU ">30 days in good standing" criterion is satisfied with enormous margin.

  • Email: verified (account is in good standing; ZeroGPU allocation was accepted, which requires it).

  • 8 pre-existing Gradio Spaces, every one of them requested: cpu-basic, and zero on ZeroGPU β€” so both free ZeroGPU slots were open before this Space was created:

    Space sdk_version python hardware requested stage
    WolfDavid/dino 6.9.0 3.11 cpu-basic RUNNING
    WolfDavid/fea-surrogate 5.9.1 3.11 cpu-basic SLEEPING
    WolfDavid/mechspec-qwen-demo 5.9.1 3.11 cpu-basic SLEEPING
    WolfDavid/agentic-market-analyzer 5.9.1 3.11 cpu-basic SLEEPING
    WolfDavid/log-anomaly-detector 5.9.1 3.11 cpu-basic RUNNING
    WolfDavid/vision-edge 5.9.1 3.11 cpu-basic SLEEPING
    WolfDavid/nllb-translator 5.9.1 3.11 cpu-basic SLEEPING
    WolfDavid/blip-captioner 5.9.1 3.11 cpu-basic SLEEPING
  • WolfDavid/dino runs sdk_version: 6.9.0, so Gradio 6 is proven to work on this account.

  • Note for anyone re-running this check: runtime.hardware.current is null for every SLEEPING Space. current reports what is running right now, not what the Space is entitled to. Read runtime.hardware.requested when auditing hardware tiers, or six of these eight Spaces look unallocated.

  • Note also that GET /api/spaces?author=WolfDavid still returned only the original 8 shortly after creation; the list endpoint lags. GET /api/spaces/WolfDavid/japanese-learning-avatar returns the new Space immediately and is the authoritative read.

Why hardware.current is null, and why that is correct

"runtime": { "stage": "NO_APP_FILE",
             "hardware": { "current": null, "requested": "zero-a10g" },
             "gcTimeout": 172800 }

The Space repo contains only .gitattributes and an auto-generated README.md. With no app.py there is nothing to schedule, so no hardware is attached. This is the correct pre-deploy state and is not a failure. requested: zero-a10g is the binding fact β€” it is the tier the Space will start on the moment plan 01-05 pushes an app.

Reachability today, measured three times consecutively with identical results:

URL HTTP Meaning
https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar 200 The Space exists and is public
https://wolfdavid-japanese-learning-avatar.hf.space 503 Your space is in error, check its status on hf.co β€” no app.py yet

The direct subdomain is provisioned (runtime.domains[0].stage: READY, subdomain: wolfdavid-japanese-learning-avatar), it simply has nothing to serve. Plan 01-05 turns this 503 into a 200, and tests/e2e/test_avatar_loop.py::test_space_reachable β€” which plan 01-05 creates β€” is what actually satisfies DPLY-01. DPLY-01 is not satisfied by this plan.

The front-matter correction, applied (plan 01-05)

Space creation auto-generated a README.md whose front-matter conflicted with the pinned stack. Plan 01-05 wrote a repository README.md that replaces it:

Key Auto-generated value Value now shipped Source of the requirement
sdk_version 6.26.0 6.22.0 requirements.txt line 1 pins gradio==6.22.0; CLAUDE.md says pin exactly
python_version '3.12' 3.12.12 ZeroGPU provides only 3.10.13 and 3.12.12; .python-version is 3.12.12
emoji 🐠 πŸ—Ύ cosmetic; the auto-generated value was also mojibake in the API response
colorFrom / colorTo purple / gray pink / indigo cosmetic
license absent mit the repository's own code is MIT; third-party asset terms are separate
short_description absent present shown on the profile card a recruiter sees first

hf_oauth was deliberately not added: it belongs to Phase 4, and every extra consent scope is a deterrent on the Space's first-load screen.

Left uncorrected, the Space SDK would have installed Gradio 6.26.0 while requirements.txt asked for 6.22.0 β€” the exact "unpinned sdk_version floats you into a breaking release" failure CLAUDE.md warns about, arrived at from the other direction.

How the unrelated histories were reconciled (plan 01-05)

The Space repo was created with its own root commit β€” b70455d, "initial commit", two files (.gitattributes, README.md) β€” while the local repository has a completely separate root and ~44 commits. git merge-base master space/main returned nothing: the two histories share no ancestor, so a plain git push space master:main is rejected as non-fast-forward.

Chosen: git merge --allow-unrelated-histories space/main, not a force-push. Both files conflicted and both were resolved deliberately:

  • README.md β€” resolved to ours. The remote copy is the boilerplate quoted above and is precisely what this plan exists to replace; there is nothing in it to preserve.
  • .gitattributes β€” resolved to the union. Ours (plan 01-02: *.vrm *.vvm *.wav *.onnx) is kept verbatim and Hugging Face's 35 default LFS patterns are appended below it. The defaults cost nothing today β€” every currently tracked binary is already an LFS object under the narrow patterns β€” and they arm the formats later phases will push (.safetensors, .bin, .pt, .npz) against the 10 MiB non-LFS rejection. Dropping them to keep the file minimal would have traded a real safety net for tidiness.

A force-push would have worked and lost nothing of substance, but the merge keeps the Space's own creation commit in the history, which is the provenance record for the ZeroGPU allocation.

GPU posture on the deployed Space

DISABLE_GPU=1 is set as a Space variable (set by plan 01-05 via HfApi.add_space_variable, 2026-09-05, description recorded on the variable itself). Phase 1 calls no GPU function, so this costs nothing and makes SC-4's demonstration honest from the first deploy rather than retrofitted at the end. Read it back with:

python -c "from huggingface_hub import HfApi; print(HfApi().get_space_variables('WolfDavid/japanese-learning-avatar'))"

AVATAR_TRANSPORT is deliberately not set, so the Space uses the component's default, inline β€” the transport the spike is testing. Setting it to iframe is the entire fallback.

Why not the alternatives

  • cpu-basic (free) β€” no longer available. Attempting it returned HTTP 402 Payment Required (quoted verbatim above). This option is closed, not merely discouraged.
  • sdk: static β€” free for everyone, but VOICEVOX cannot run in a browser, so there would be no mora timings and AVTR-02 collapses entirely. ZeroGPU is also Gradio-SDK-only, so Phase 3 would need a rewrite. Rejected.
  • PRO ($9/mo) β€” not purchased, and not needed: the free ZeroGPU path worked. PRO does not improve the anonymous visitor's 2 min/day quota, so it would change nothing about the free-tier design that Phase 1 is built around. No payment was made by any step of this plan.

GPU posture

Phase 1 calls no GPU function. DISABLE_GPU=1 is to be set as a Space variable and every future GPU entry point must raise when it is set. Proof lives in tests/test_no_gpu_on_turn_path.py (static AST scan) and a DISABLE_GPU=1 deployed E2E run.

A ZeroGPU Space that never calls @spaces.GPU runs its web process on CPU and consumes zero visitor quota, so taking the ZeroGPU tier costs nothing at runtime in Phase 1. Plan 01-04 already established that speech synthesis is CPU-only and AST-verified free of spaces/torch imports.

ZeroGPU slot budget

1 of 2 free ZeroGPU slots is now consumed by this Space. 1 slot remains for the other four HF-profile projects. Because cpu-basic Gradio creation is 402-gated on this account, that final slot is the last free Gradio Space this account can create without PRO β€” spend it deliberately.

Git remote

git remote add space https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar

Authentication is a cached fine-grained token held by the hf CLI (hf auth whoami reports WolfDavid), scoped repo.write / repo.content.read / repo.access.read on the WolfDavid entity. The token value is never written to this repository, any commit message, or any log.

No application code has been pushed to this remote. Deployment is plan 01-05's task and is gated behind its own spike verdict.

First deploy record (plan 01-05, 2026-09-05)

Five pushes were needed. Four failed, and three of the four failures were platform constraints that no amount of local testing would have surfaced - the full local suite (82 quick-loop + 15 browser tests) was green before the first push and stayed green throughout. They are recorded here in full because each one is a trap for the next Space.

# Commit Stage reached What happened
1 ad85cc8 BUILD_ERROR ResolutionImpossible: our spaces==0.51.1 vs the builder's injected spaces==0.51.3
2 6eeb816 RUNTIME_ERROR ModuleNotFoundError: No module named 'japanese_avatar' - the src/ layout is not importable on a Space
3 e3ebc72 RUNTIME_ERROR App launched, then exited. GRADIO_HOT_RELOAD: Launching demo not found in __main__
4 8c9985e RUNTIME_ERROR SIGTERMed by the platform: No @spaces.GPU function detected during startup
5 4e3cac9 RUNNING Live, hardware.current: zero-a10g

Finding 1: do not pin spaces in requirements.txt

The Space builder generates its own pip command and appends its own pin:

pip install --no-cache-dir -r /tmp/requirements.txt "torch<=2.11.0" \
    gradio[oauth,mcp]==6.22.0 "uvicorn>=0.14.0" "websockets>=10.4" spaces==0.51.3

Two exact pins for one package is unsatisfiable. The platform owns the spaces version. gradio==6.22.0 in requirements.txt is fine because the builder injects the same version - it reads sdk_version from the README front-matter, which is exactly why that front-matter correction had to land in the same push.

Finding 2: a src/ layout package is not importable on a Space

A Space installs requirements.txt and nothing else; the repository is never pip installed. app.py now puts src/ on sys.path itself rather than relying on a PYTHONPATH Space variable, so a clean checkout runs with no platform configuration.

Finding 3: the launched Blocks must be a module-level attribute named demo

Hugging Face runs the app under gradio.utils.SpacesReloader, whose postrun() calls getattr(watch_module, demo_name) on every reload check. A build_app() factory that keeps the Blocks local logs GRADIO_HOT_RELOAD: Launching demo not found in __main__. Using 'demo' and then dies with no traceback. demo = build_app() at module level is not a style preference on a Space; it is the contract.

Finding 4 (the important one): ZeroGPU refuses to run without a @spaces.GPU function

runtime.errorMessage: "No @spaces.GPU function detected during startup"

The app had already bound port 7860 and printed Running on local URL when the platform SIGTERMed it. A ZeroGPU Space must declare at least one GPU entry point or it will not run at all, and since cpu-basic Gradio Spaces are 402-gated on this account (above), there is no other free tier to move to.

app.py::zerogpu_probe exists solely to satisfy that startup scan. Nothing calls it, it is not on the turn path, and it raises when DISABLE_GPU=1 - which is set on this Space.

This changes what SC-4 can claim. "No @spaces.GPU function exists anywhere" is no longer an available proof; the platform has taken it off the table. The honest and equally strong assertion is "the turn loop completes without ever reaching one", which the DISABLE_GPU=1 deployed run demonstrates directly, since reaching it would raise. Any later plan that asserts the absence of a GPU decorator will fail for a reason that has nothing to do with this project's design.

Finding 5: Gradio 6 serves an SSR shell, so do not grep the HTML for elem_ids

curl -s $SPACE_URL | grep -c vrm-stage returns 0 on a working deployment. Gradio 6 server-side-renders a shell and the component tree arrives from GET /config. The equivalent check that does work:

curl -s https://wolfdavid-japanese-learning-avatar.hf.space/config | grep -c vrm-stage   # 1

Git LFS resolved correctly on the Space

Verified against the live app, not assumed - a silently unpushed pointer is the classic failure here:

Asset HTTP Bytes served
/gradio_api/file=avatar/assets/tutor.vrm 200 10,776,032 (model/vrml)
/gradio_api/file=avatar/assets/demo-konnichiwa.wav 200 50,732
/gradio_api/file=avatar/avatar.js 200 5,138
/gradio_api/file=avatar/stage.html 200 9,271

All 12 LFS objects (181 MB) uploaded on the first push and the Space checks them out as real files, not 130-byte pointers.