AUREOLE-R-v3 / DATA_DICTIONARY.md
PureOne's picture
AUREOLE-R 3.0.0-hf.1: standalone public research release
9d6c005 verified
|
Raw
History Blame Contribute Delete
5.02 kB

Evidence and data dictionary

All scenes are procedural and synthetic. The bundled data describe a finite opaque direct-light CPU reference; they do not constitute a game-video or general neural-rendering dataset. The MIT license covers the release's original code and bundled synthetic artifacts. Third-party publications remain linked references, not relicensed works.

E10: motion and sequential residual elimination

results_v3/innovation_raw.csv has 18,144 rows. The row key is (scene, seed, phase, phase_frame, method). Twelve scene IDs are 400–411; three sampling seeds are 5101–5103. Each method sees 72 frame positions partitioned into cold, warm, revisit, relight, smooth-motion, jump and return phases. See the frozen experiments_innovation.json for all values.

Column Meaning / units
scene, seed Procedural scene and sampling seed identifiers
phase, phase_frame, frame Event phase, frame index within phase, overall frame index
method Named estimator/ablation; unsafe_eliminate is a deliberate invalidation failure
expected_mse Per-record diagnostic conditional expected mean-squared linear-RGB error under the estimator's sampling contract; see the risk calculation in code
observed_mse Mean squared error of the actually sampled linear-RGB output against the reference
rays Total online physical segment queries for this record; not rays per pixel
receivers Number of queried surface receivers; use rays / receivers for queries per receiver
runtime_seconds Measured estimator batch runtime; common preparation and separate reference audit are excluded
memory_bytes Reported estimator memory footprint; not process RAM or GPU VRAM
certified_fraction_before Fraction of relevant terms accepted as known before new queries
false_certificates Number of accepted facts contradicted by the independent reference audit
enclosure_violation Maximum diagnostic violation of the output enclosure for the record
enclosure_mean_width Mean width of the direct-light output enclosure
negative_channel_fraction Fraction of estimated color channels below zero before any display clamp

Reports aggregate by scene/phase/method. Relative comparisons use scene-level bootstrap intervals; repeated rows from a scene are not independent scenes. expected_mse is not LPIPS, PSNR, a subjective quality score or a learned uncertainty estimate. Read scripts/benchmark_innovation.py and aureole/innovation.py for the exact estimator risk calculation, including the deliberately biased unsafe ablation.

E11: spatial, known-time and appearance queries

results_v3/queries_raw.csv has 40 rows: scene IDs 600–607 and times 0, 0.25, 0.5, 0.75, 1. Each row supports three known appearance readouts. Anchor observations are frozen before queries. This is not a causal prediction benchmark for unknown future actions.

Column Meaning / units
scene, time Procedural scene and prescribed normalized query time
terms Number of finite receiver-light visibility terms queried at that time
certified_terms Terms reused under valid spatial/motion conditions
fresh_queries Physical queries required for the unresolved complement
shared_fresh_baseline_queries Fresh visibility baseline that already shares work among all three readouts
separate_readout_baseline_queries Weaker baseline that recomputes independently for each readout; not used for the headline
false_certificates Accepted reused terms contradicted by full reference visibility
max_linear_rgb_error Largest absolute linear-RGB difference among the tested readouts
lookup_trace_seconds Measured lookup plus remaining physical-query time
three_readout_seconds Measured time for the three appearance readouts

Initialization queries are charged once per scene in queries_report.json, not repeated in every CSV row. The headline calculation is 1 - (147456 + 535303) / 2949120 = 0.7684872097. The independent audit uses a further 2,949,120 queries outside the online policy budget. Neither number is an end-to-end speedup.

Arrays, prior and provenance

  • .npz artifacts use NumPy array archives. Load with numpy.load(path, allow_pickle=False) and inspect .files for keys; no serialized executable model object is required.
  • models/visibility_prior.npz is the frozen v2 prior. Its training/test partitions and metrics are in results/training.json; its role is documented in MODEL_CARD.md.
  • Frozen configuration hashes are embedded in the two v3 reports. File-level hashes are in CHECKSUMS.sha256.
  • results/ contains earlier evidence; results_v3/ contains current scientific evidence; results_publication/ contains packaging checks only.
  • No additional personal or externally licensed game assets are required to generate these synthetic scenes. Generalization to realistic content remains open.