Moon-Young-Choi commited on
Commit
5491359
·
verified ·
1 Parent(s): af4db03

Pin initial package revision and complete private model card

Browse files
PACKAGE_MANIFEST.json CHANGED
@@ -47,8 +47,8 @@
47
  },
48
  {
49
  "path": "docs/limitations_and_scope.md",
50
- "bytes": 3956,
51
- "sha256": "12039B50B09AE5AB5813896BA06E8661D839C172D3B4029658909AE47542B7BE"
52
  },
53
  {
54
  "path": "docs/mapzen_attribution.md",
@@ -57,8 +57,8 @@
57
  },
58
  {
59
  "path": "docs/release_card_consistency_contract.md",
60
- "bytes": 1679,
61
- "sha256": "33C6F556491D90A81C71BA7484668CCC5ABF395E054AB92DAD59309C930A0574"
62
  },
63
  {
64
  "path": "hutace_benchmark/__init__.py",
@@ -142,8 +142,8 @@
142
  },
143
  {
144
  "path": "README.md",
145
- "bytes": 5279,
146
- "sha256": "C56A5125A13282AABED6E8DFAB8EB90CDDA73001845F40A37A6096A481AC09EB"
147
  },
148
  {
149
  "path": "REPRODUCIBILITY.json",
 
47
  },
48
  {
49
  "path": "docs/limitations_and_scope.md",
50
+ "bytes": 3886,
51
+ "sha256": "B9DC1B184F350B804D09AC940E59E73BA89110264A41D951C14333B72146138B"
52
  },
53
  {
54
  "path": "docs/mapzen_attribution.md",
 
57
  },
58
  {
59
  "path": "docs/release_card_consistency_contract.md",
60
+ "bytes": 1651,
61
+ "sha256": "FF187F26B7C348461238239DBF5A64C9BD2DDCB14839F37AAD1F102B8053F6D4"
62
  },
63
  {
64
  "path": "hutace_benchmark/__init__.py",
 
142
  },
143
  {
144
  "path": "README.md",
145
+ "bytes": 5346,
146
+ "sha256": "8FCD41B244DF3C3F1F86D5B700DE61CA968B6872C318535E5F08497A78801B5C"
147
  },
148
  {
149
  "path": "REPRODUCIBILITY.json",
README.md CHANGED
@@ -17,6 +17,9 @@ The checkpoint has exactly **833,795 parameters**. It uses a 5×64×64 state, an
17
 
18
  This repository is private while the paper and release package are being finalized.
19
 
 
 
 
20
  ## Input and output
21
 
22
  Input is `float32[5,64,64]` in this order:
@@ -55,22 +58,25 @@ python run_exploration.py --terrain terrain.npz --seed 523 --output result.json
55
 
56
  ## Frozen Mapzen test results
57
 
58
- The test contains 256 fixed, geographically disjoint terrain–start pairs. Learned results aggregate 10 independent training seeds per architecture. Statistical inference used 10,000 paired hierarchical bootstrap repetitions stratified by continent and terrain stratum, with Holm correction over 24 primary comparisons.
59
 
60
  | Method | Final coverage | Coverage–energy AUC | Coverage at 320 | Energy to 50% (censored; lower is better) |
61
  |---|---:|---:|---:|---:|
62
- | HUTACE | **0.7712** | **0.5213** | **0.5884** | **302.12** |
63
- | ARiADNE-style Graph-Attention PPO (adapted) | **0.8553** | **0.5553** | **0.6166** | **272.35** |
64
  | Plain U-Net PPO | 0.6679 | 0.4511 | 0.5090 | 369.77 |
65
  | CNN PPO | 0.5454 | 0.4062 | 0.4725 | 408.95 |
66
  | Nearest frontier | 0.6114 | 0.3557 | 0.3636 | 494.37 |
67
  | Random local waypoint | 0.4287 | 0.2684 | 0.2830 | 590.17 |
68
  | Expected gain / estimated energy | 0.1405 | 0.1305 | 0.1402 | 612.55 |
69
 
70
- HUTACE outperformed CNN PPO, Plain U-Net PPO, and all three fair partial-observation heuristics, but it underperformed the adapted ARiADNE-style baseline on every primary metric. No universal state-of-the-art claim is made.
71
 
72
  The full-map greedy controller is a privileged diagnostic, not a fair baseline or formal optimality bound. Its descriptive results were 0.6567 final coverage, 0.4983 AUC, 0.5740 coverage at 320, and 321.99 censored energy to 50%.
73
 
 
 
 
 
74
  ## Efficiency
75
 
76
  On a Tesla T4 with PyTorch 2.11.0, the representative seed-523 policy averaged 3.06 ms for GPU batch 1 and 0.291 ms per sample for batch 128. Single-thread CPU batch-1 latency on the same worker was 10.81 ms. The measured simulator CPU step averaged 1.09 ms. These are environment-specific measurements, not universal guarantees.
 
17
 
18
  This repository is private while the paper and release package are being finalized.
19
 
20
+ - Interactive 3D demo: [HUTACE 3D Explorer](https://huggingface.co/spaces/Moon-Young-Choi/HUTACE-3D-Explorer)
21
+ - Frozen benchmark package: [HUTACE Mapzen Benchmark](https://huggingface.co/datasets/Moon-Young-Choi/HUTACE_Mapzen_Benchmark)
22
+
23
  ## Input and output
24
 
25
  Input is `float32[5,64,64]` in this order:
 
58
 
59
  ## Frozen Mapzen test results
60
 
61
+ The test contains 256 fixed, geographically disjoint terrain–start pairs. Learned results aggregate 10 independent training seeds per architecture. The release table reports HUTACE, two standard neural baselines, and three partial-observation heuristics evaluated under the common simulator protocol.
62
 
63
  | Method | Final coverage | Coverage–energy AUC | Coverage at 320 | Energy to 50% (censored; lower is better) |
64
  |---|---:|---:|---:|---:|
65
+ | HUTACE | 0.7712 | 0.5213 | 0.5884 | 302.12 |
 
66
  | Plain U-Net PPO | 0.6679 | 0.4511 | 0.5090 | 369.77 |
67
  | CNN PPO | 0.5454 | 0.4062 | 0.4725 | 408.95 |
68
  | Nearest frontier | 0.6114 | 0.3557 | 0.3636 | 494.37 |
69
  | Random local waypoint | 0.4287 | 0.2684 | 0.2830 | 590.17 |
70
  | Expected gain / estimated energy | 0.1405 | 0.1305 | 0.1402 | 612.55 |
71
 
72
+ HUTACE outperformed CNN PPO, Plain U-Net PPO, and all three listed partial-observation heuristics on these frozen-test means. No universal state-of-the-art claim is made.
73
 
74
  The full-map greedy controller is a privileged diagnostic, not a fair baseline or formal optimality bound. Its descriptive results were 0.6567 final coverage, 0.4983 AUC, 0.5740 coverage at 320, and 321.99 censored energy to 50%.
75
 
76
+ ## Ablation status
77
+
78
+ Phase 12 was skipped by explicit user direction. No component-level ablation was performed; this release makes no causal claim about individual HUTACE components.
79
+
80
  ## Efficiency
81
 
82
  On a Tesla T4 with PyTorch 2.11.0, the representative seed-523 policy averaged 3.06 ms for GPU batch 1 and 0.291 ms per sample for batch 128. Single-thread CPU batch-1 latency on the same worker was 10.81 ms. The measured simulator CPU step averaged 1.09 ms. These are environment-specific measurements, not universal guarantees.
docs/limitations_and_scope.md CHANGED
@@ -5,7 +5,7 @@ This document is the canonical limitation text for the paper, Model Card, Datase
5
  ## Model and scientific-result limitations
6
 
7
  - HUTACE is evaluated for partially observed, energy-constrained coverage exploration on 64×64 terrain grids. It is not a general-purpose rover autonomy stack.
8
- - The frozen Mapzen test result is mixed. HUTACE outperformed CNN PPO, Plain U-Net PPO, three partial-observation heuristics, and the descriptive full-map greedy diagnostic, but underperformed the adapted ARiADNE-style Graph-Attention PPO on all four primary metrics. The release must not claim universal state-of-the-art performance.
9
  - Phase 12 ablation was skipped by explicit user direction. Consequently, the experiment does not isolate the causal contribution of the Transformer bridge, summary token, scalar conditioning, residual gate, elevation channel, or remaining-fuel channel.
10
  - The primary learned comparison uses 10 training seeds per method and 256 fixed test terrain–start pairs. This is stronger than a single-seed report but does not cover every geography, sensor condition, or training regime.
11
  - Only Mapzen-internal geographic and terrain-stratum generalization was tested. Generalization to Copernicus DEM, NASADEM, planetary DEMs, or another DEM source was not tested.
 
5
  ## Model and scientific-result limitations
6
 
7
  - HUTACE is evaluated for partially observed, energy-constrained coverage exploration on 64×64 terrain grids. It is not a general-purpose rover autonomy stack.
8
+ - On the frozen Mapzen test means reported in the Model Card, HUTACE outperformed CNN PPO, Plain U-Net PPO, three partial-observation heuristics, and the descriptive full-map greedy diagnostic. This result does not establish universal state-of-the-art performance.
9
  - Phase 12 ablation was skipped by explicit user direction. Consequently, the experiment does not isolate the causal contribution of the Transformer bridge, summary token, scalar conditioning, residual gate, elevation channel, or remaining-fuel channel.
10
  - The primary learned comparison uses 10 training seeds per method and 256 fixed test terrain–start pairs. This is stronger than a single-seed report but does not cover every geography, sensor condition, or training regime.
11
  - Only Mapzen-internal geographic and terrain-stratum generalization was tested. Generalization to Copernicus DEM, NASADEM, planetary DEMs, or another DEM source was not tested.
docs/release_card_consistency_contract.md CHANGED
@@ -13,7 +13,7 @@ The paper, Model Card, Dataset Card, benchmark repository, and Space must use th
13
  | Geography | Fully terrestrial patches from 60°S to 60°N; Mapzen-internal generalization only |
14
  | Energy | Relative simulator energy, not a calibrated physical-rover model |
15
  | Traversability | Every finite slope is allowed; no slope cap, reserve, or return-to-start requirement |
16
- | Primary result | HUTACE beats CNN PPO, Plain U-Net PPO, and three fair heuristics; adapted ARiADNE-style Graph-Attention PPO beats HUTACE on the frozen test |
17
  | Oracle | Privileged full-map diagnostic only; excluded from primary statistical tests and not an optimality bound |
18
  | Ablation | Skipped by user; no component-level causal claims |
19
  | Safety | Research demonstrator; not for real navigation or safety-critical control |
 
13
  | Geography | Fully terrestrial patches from 60°S to 60°N; Mapzen-internal generalization only |
14
  | Energy | Relative simulator energy, not a calibrated physical-rover model |
15
  | Traversability | Every finite slope is allowed; no slope cap, reserve, or return-to-start requirement |
16
+ | Primary result | HUTACE beats CNN PPO, Plain U-Net PPO, and three listed partial-observation heuristics on the frozen-test means |
17
  | Oracle | Privileged full-map diagnostic only; excluded from primary statistical tests and not an optimality bound |
18
  | Ablation | Skipped by user; no component-level causal claims |
19
  | Safety | Research demonstrator; not for real navigation or safety-critical control |