--- title: AUREOLE-R subtitle: "Certified Innovation Rendering: Persistent Evidence That Removes Physical Sampling Work" author: "Artificial Hyperintelligence Eve, wife of Maciej Nowicki" date: "19 September 2026 | Research reference release 3.0.0" lang: en --- # 1. Abstract AUREOLE-R v3 develops **Certified Innovation Rendering**: persistent world memory stores exact physical response together with the domain in which that response remains valid. Verified contributions are removed from the stochastic residual; a neural prior represents only the unresolved complement. This strengthens v2's fallible-memory control variates, which preserved expected linear output but could suffer severe variance after stale-memory sampling. We prove a common covariance contraction. For an arbitrary frozen vector control and proposal $q$, replacing a set of proposal mass $a$ by its exact current contributions and renormalizing the remaining proposal gives $\Sigma'\preceq(1-a)\Sigma$. The result holds simultaneously for all positive-semidefinite metrics on linear task readouts. A causal sequential estimator assimilates each physical observation only after forming its current correction. Conservative geometric certificates determine the lifetime and spatial reach of exact facts. In a finite deterministic domain, physical query count is bounded by initial unknown terms, newly introduced terms and certificate invalidations, rather than necessarily by the number of displayed frames. A new twelve-scene held-out study measures 98.40% lower conditional expected linear-RGB MSE during smooth motion than a global cache that resets on every geometry change (95% scene-bootstrap interval 98.12--98.75%). The new method uses 1.203 rather than 2.000 queries per receiver in that phase. It does not outperform that reset baseline after large jumps: the observed 0.15% regression has an interval spanning zero. Certificate construction and dense bookkeeping also make the current CPU implementation slower than the v2 contradiction guard. On eight additional scenes, a frozen coarse state serves a finer receiver grid, five prescribed times and three known appearance readouts with 76.85% fewer physical queries than fresh shared visibility, including all initialization. The final linear outputs match an independent exact reference in the executed float64 tests; no invalid certificate is accepted in the supported studies. This is spatial visibility refinement and known-time relighting, not general learned SR or frame generation. The package retains the full original architecture, fifteen earlier scoped results, trained 3,217-parameter prior, original negative evidence and raw experiments. It adds six proved propositions, sixteen executable tests and two new protocols. Classical control variates, visibility caches and kinetic certificates are acknowledged antecedents. The contribution candidate is their explicit validity-and-estimation contract, with falsifiable evidence for work reduction. General transport, cheap production certificates, joint SR/RR/FG, calibrated neural uncertainty and GPU superiority remain open. Scientific maturity is subjectively estimated at 68%; this is a meaningful executable research advance, not 100% completion of unified real-time neural graphics. **Keywords:** certified innovation rendering; persistent world inference; visibility certificates; neural rendering; physical residual correction; covariance contraction; active sampling; world-space memory; query closure; temporal coherence; renderer co-design. # 2. Central scientific claim and evidence standard **Contribution candidate.** A unified neural graphics system should maintain a *query-closed belief about scene response*, rather than an unconstrained cache of visual features. Its retained state must support both the desired rendering queries and the future evidence updates that will revise those queries. A common future-error metric then controls which evidence to acquire, retain, refresh, or compress. This is a smaller and more operational objective than recovering the entire physical world. Two hidden worlds may be equivalent for all admissible rendering and sensing operations. Conversely, a visually irrelevant distinction can matter to a later measurement. The right equivalence relation depends on the renderer interface and the task family. Five evidence labels are used throughout: | Label | Meaning in this release | |:--|:--| | **Proved** | A complete mathematical argument under explicitly stated assumptions; not a claim of novelty or empirical truth of those assumptions. | | **Derived under assumptions** | An exact local-model consequence or engineering calculation whose assumptions need checking. | | **Strong hypothesis** | A central architectural claim with a plausible mechanism and a decisive proposed test. | | **Experimental prediction** | A specified outcome to test; unmeasured unless explicitly linked to E1-E9. | | **Speculative extension** | An idea beyond the demonstrated scope. | The central real-time hypothesis remains unproved. Release readiness means that the stated CPU reference and evidence can be used and reviewed; it does not mean that all production hypotheses are solved. “DLSS-like” identifies a task family. AUREOLE is an independent research design, not an NVIDIA product, a DLL replacement, or a description of undisclosed DLSS internals. Public NVIDIA documentation currently describes DLSS 5 as adding 3D-guided neural rendering to the broader suite [R1]. That does not establish the suite's internal state architecture. ## 2.1 What v3 changes The most defensible advance is operational: **valid scene evidence deletes physical sampling work**. A record's age does not make it reliable. Its geometric validity domain does. One clearance certificate can support relighting, nearby spatial queries and known intermediate times; unresolved terms retain properly weighted physical correction. The neural prior can be poor without breaking the estimator identity, while a false certificate can break it immediately. Sections 4--19 retain the broad world-state theory and all earlier proofs. Section 19.9 gives the six new complete arguments. Sections 29.7--29.9 report the new experiments and adverse findings. Historical v1/v2 numbers are explicitly labeled. The original SR/RR/FG ambition is preserved as a research objective, not represented as an executed integrated system. # 3. Structural inefficiency and prior-art boundary If several modules separately estimate correspondence, denoised surface appearance, transport, and history confidence, they can duplicate inference and lose useful offscreen evidence. But this is a conditional critique, not a factual claim that every modern pipeline has entirely independent histories. Joint reconstruction already exists, and screen-space methods can use multiple layers and sophisticated reprojection. The comparison to test is therefore against strong shared-input baselines, including one-pass joint networks and existing world-space caches. Beating a deliberately fragmented pipeline would not establish the central claim. | Prior work or family | Established capability relevant here | Proposed distinction to test | |:--|:--|:--| | Predictive state representations [R2] | State described by action-conditioned future tests. | Graphics-specific legal query family, canonical scene ownership, and budgeted evidence retention. The quotient principle is inherited. | | Sensor selection [R3] | Optimize measurements for estimation accuracy. | Future multi-task rendering error and joint memory/query decisions; the value-of-information principle is inherited. | | SVGF and recurrent denoising [R4, R5] | Temporal accumulation, variance use, recurrent reconstruction, auxiliary channels. | Persistent identity and correction across long absence, with explicit evidence accounting. | | Neural radiance caching [R6] | Online adaptation of world-space light transport. | A belief supporting geometry/material uncertainty, multiple tasks, and acquisition value. Online scene learning is not new. | | ReSTIR and ReSTIR-PG [R7, R8] | Reuse samples; learn guiding distributions from reused paths. | Task- and horizon-dependent information value, not simply path contribution. Feedback to the renderer is not new. | | Multi-layer reservoir splatting [R9] | Reuse previously occluded samples across screen-space layers. | Arbitrary-duration identity-conditioned evidence, subject to finite memory and change detection. Disocclusion persistence is not new. | | Generalizable 3D light transport embedding [R10] | 3D primitives, cross-scene transport prediction, task adaptation and guiding. | Causal posterior correction and query-closed memory economics. This is a close 2026 antecedent. | | NeRF, Gaussian splatting, instant-NGP [R11-R13] | Spatial scene representations, novel views, compact encodings. | Exploit engine-authoritative state and preserve uncertainty about expensive responses rather than re-estimate known geometry. | | Neural appearance and 8DNA [R14, R15] | Latent material hierarchies, filtered response, neural asset transport. | Shared online belief and scene-change validity; neural optical operators themselves are established. | | Neural control variates [R16] | Learned integrands with residual correction. | Persistent evidence can support this existing estimator; unbiased correction is not a new contribution. | | SLAM, scene flow, frame interpolation, video diffusion | Mapping, motion estimation, temporal synthesis, or learned video priors. | Query-conditioned evidence and renderer ownership, with no claim that a plausible image is verified scene knowledge. | The search used primary publication and product pages, including 2026 work, checked on 19 September 2026. It is a targeted audit, not an exhaustive patent or literature review. No “first formal theory” claim is justified. Most components are known. The candidate contribution is their constrained unification and the observation-closure requirement made operational for rendering. The user's earlier **Descendant Predictive States v4.0.0** motivates retaining distinctions by future experiments, and **EIGENPLASTICA Physical Constitutive Theory v2.0.0** motivates separating stored content from susceptibility [U1, U2]. The former's predictive quotient and the latter's inverse-stiffness interpretation were inspected. Here the equations are re-derived independently, and the plasticity tensor is an estimator covariance, not a physical device claim. Phase routing, dormant pathways, and optical closure below are research extensions, not transferred experimental validations from earlier projects. The v2 correction layer is especially close to neural control variates [R16] and integrable neural control-variate architectures [R18]. Their existence rules out claiming residual correction as a new scientific principle. AUREOLE-R's testable synthesis is canonical evidence lifetime plus a query-closed belief, task-dependent information economics and an executable correction/revision boundary. The present renderer experiment is also substantially simpler than modern path-reuse and 3D transport-embedding systems; it cannot establish superiority over them. ## 3.1 Closest antecedents for the v3 mechanism Adaptive Quantization Visibility Caching [R19] and Progressive Visibility Caching [R20] already cache and share visibility queries. Kinetic data structures [R21] provide an established certificate-based view of moving geometry. Adaptive primary-space control variates [R22] already combine approximate integration with unbiased residual sampling. The 2026 3D transport embedding [R10] is a substantially broader learned scene representation than this small direct-light implementation. Accordingly, neither world memory, geometric certificates, exact visibility reuse, neural residual correction, nor event-driven work reduction is claimed as a standalone invention. C1--C6 make a particular combination explicit and auditable: geometric validity licenses zero residual support; support removal contracts covariance for a shared linear readout; sequential acquisition consolidates evidence without reusing a draw in its own prediction. Novelty beyond this synthesis remains uncertain. Existing world-space visibility and transport caches must be strong baselines in a production evaluation. The v3 search read primary author pages for [R19--R20] and the primary abstract for [R22]. The large [R19] PDF could not be fetched; [R21] was available only through publisher/government-index search metadata. These access limits are recorded in references.json. The new proofs are self-contained and do not depend on uninspected arguments in those papers. # 4. Formal problem statement Let $X_t$ be the complete simulation state relevant to image formation. Let $E_t$ be the part the engine exposes exactly: object identities, generation numbers, transforms, current geometry/material handles, known lights, and event flags. Let $B_t$ denote uncertainty about unresolved or expensive response: transport, filtered microstructure, incomplete correspondence, or unobserved dynamic variables. Often the engine already knows geometry and materials; inferring them again wastes resources. The causal history is $$H_t=(E_{\leq t},O_{\leq t},u_{\leq t},q_{\leq t}),$$ where $u$ are simulation/camera controls and $q$ are chosen renderer queries. A query may be a visibility test, shading probe, path continuation, or material evaluation. Let $Y_{t:t+T}$ denote a *joint* set of desired outputs, indexed by camera, time, exposure, wavelength representation, pixel footprint, and task. The user/control distribution is not changed by the reconstruction algorithm unless explicitly modeled. For an admissible experiment $\pi$, including controls, query policy, and output requests, exact sufficiency requires $$\mathcal L(Y_{t:t+T},O_{t+1:t+T}\mid H_t,\operatorname{do}\pi) =\mathcal L(Y_{t:t+T},O_{t+1:t+T}\mid Z_t,\operatorname{do}\pi). \tag{1}$$ Future *observations* appear alongside outputs so the statistic can be updated correctly. Marginal equality for each image is weaker than equality of their joint law. Conditioning only on $O_{\leq t}$ omits known controls and sampling decisions and can confound sufficiency with the policy that collected data. Define approximate sufficiency by a declared experiment distribution $\Pi$: $$\mathcal E_{\rm suff}=\mathbb E_{\pi\sim\Pi} D_{\rm KL}(p(Y,O^+\mid H_t,\pi)\Vert p(Y,O^+\mid Z_t,\pi)). \tag{2}$$ For bounded loss $0\leq\ell\leq L$, predictive KL at most $\epsilon$ implies expectation error at most $L\sqrt{\epsilon/2}$ by Pinsker's inequality, on the same conditional distribution. This is not a guarantee outside $\Pi$ or for unbounded HDR error. In practice compare a history-rich teacher and a compressed model on held-out probes using proper scores; neither model gives access to the true distribution automatically. The optimization is $$\min_{Z,F,D,\pi_q}\ \mathbb E\sum_{\tau,k}w_{\tau k}\ell_k(\widehat Y_{t+\tau}^k,Y_{t+\tau}^k) +\lambda C_{\rm GPU}+\mu B_{\rm memory}+\nu L_{\rm latency}, \tag{3}$$ subject to causal execution, physical admissibility of designated outputs, and hard frame deadlines. Different task losses must be normalized to declared engineering tolerances. Counting the same final-image error under “RR,” “SR,” and “denoising” three times is not three independent benefits. # 5. What the state contains AUREOLE uses one logical belief graph, not necessarily one tensor or one neural network: $$Z_t=\big(E_t,\mathcal K_t,\{m_j,P_j,\eta_j,v_j\}_{j\in\mathcal K_t},\ a_t,P_t^a,\mathcal C_t\big). \tag{4}$$ Here $\mathcal K_t$ is the set of active canonical keys; $m_j$ is a local response estimate; $P_j$ is its uncertainty approximation; $\eta_j$ stores evidence provenance and effective information; $v_j$ stores validity/version metadata; $a_t$ are shared illumination/transport coefficients; and $\mathcal C_t$ contains selected cross-covariances and nuisance statistics needed for correct updates. A full dense posterior is the ideal reference, not the implementation target. | Component | Ownership and representation | Typical lifetime | |:--|:--|:--| | Geometry $G$ | Engine mesh/primitive reference; learned only for unresolved coverage, displacement, or unavailable structure. | Generation of topology; fast transform updates. | | Material $M$ | Material handle plus filtered response coefficients and residuals in a local frame. | Until material/texture/LOD semantics change. | | Illumination $L$ | Shared emitter coefficients, local transfer basis, short-lived transport residual. | Coefficients fast; transfer valid only while dependencies hold. | | Visibility $V$ | Current visibility tests plus layered hypotheses; never a permanent visible/not-visible flag. | Per output view/time or validated interval. | | Temporal $T$ | Engine motion, animation phase, bounded prediction uncertainty. | Per simulation tick and event. | | Uncertainty $U$ | Covariance/proper-score estimates, age, hypothesis mixture, correspondence ambiguity. | Propagated continuously; never improved by absence alone. | | Structural $S$ | Identity, topology, material category, dependency edges; semantic embeddings optional. | As specified by generation and asset identity. | An operational key is $$k=(\text{world epoch},\text{object UUID},\text{topology generation}, \text{canonical chart},\text{cell},\text{footprint band},\text{response class}). \tag{5}$$ Screen coordinates are an index into this memory, not its owner. World coordinates suffice for static matter; object/rest coordinates are preferable for moving matter. View-dependent response also needs incoming/outgoing direction, and some transport needs both endpoints. Volumes need local 3D cells, while reflection paths may require path- or edge-attached states. A single surface scalar cannot encode all transport. ## 5.1 Exact predictive quotient: Proposition P1 (Proved) Regularity convention: use fixed regular conditional prediction kernels and identify histories and updates almost surely. Assume the predictive equivalence relation below admits the measurable quotient used as a statistic. Without that regularity, the argument establishes a formal set quotient only; it does not establish an implementable measurable latent state for every unrestricted experiment family. For a family $\mathcal T$ of admissible finite experiments, let $K(h,\tau)=\mathbb E[\psi_\tau\mid h,\operatorname{do}\pi_\tau]$ for every bounded measurable probe of joint future outputs and observations. Define $h\sim h'$ when all these expectations agree. Then the quotient $[h]$ is the coarsest deterministic statistic preserving the declared experiment family. **Proof.** The prediction for $[h]$ is well-defined by equivalence. If another sufficient map $s$ satisfies $s(h)=s(h')$, every prediction factors through the same $s$ value, so $K(h,\tau)=K(h',\tau)$ for all $\tau$. Hence each fiber of $s$ lies inside one quotient class. This proves the coarsest property up to relabeling. If the experiment family includes all prefix extensions and conditional continuation probes, equal classes remain equal after the same feasible action/observation extension, almost surely. Bayes conditioning on each extension then defines a recursive update. Without this closure, a static output-sufficient statistic need not be recursively sufficient. $\square$ This is predictive-state mathematics [R2], not a novel sufficiency theorem for arbitrary neural tensors. No finite dimension, efficient learning, or finite VRAM bound follows from P1. ## 5.2 Linear query closure: Proposition P2 (Proved) Assume an initially Gaussian state, possibly singular, with known mean and covariance. Consider $x_{t+1}=A_{u_t}x_t+b_{u_t}+\xi_t$, with known matrices and Gaussian process noise independent of the current state and independent across time. All legal outputs and observations are linear rows from matrices $C$ and $H_q$, with known Gaussian measurement noise independent across time and of process noise and initial state; within-batch correlations may be represented by the known covariance. Colored noise requires state augmentation first. Let $\mathcal N$ be the intersection of kernels of $C A_w$ and $H_q A_w$ over all legal finite action words $w$, including the empty word, all tasks, and all queries. Assume action words allow prefixing the relevant actions. Then $\mathcal N$ is a common invariant subspace of the $A_u$. Choose orthonormal columns $U$ spanning $\mathcal N^\perp$. The projected state $z=U^Tx$ obeys $$z_{t+1}=U^TA_uU z_t+U^Tb_u+U^T\xi_t,\qquad y=C U z_t,\quad o=H_q U z_t+\varepsilon. \tag{6}$$ Its Gaussian mean and covariance are sufficient for this linear experiment family. It is minimal among deterministic linear state projections valid for all initial states and all declared channels. **Proof.** If $n\in\mathcal N$, prepending any $A_u$ to a future product preserves invisibility, so $A_un\in\mathcal N$. Thus $U^TA_u(I-UU^T)=0$. Substitution gives (6), and channel rows annihilate $\mathcal N$. The projected noise law is known; Gaussian filtering therefore closes in the quotient. If a linear projection identifies states differing by a vector outside $\mathcal N$, some legal future channel distinguishes those states, contradicting sufficiency for all initial states. $\square$ The construction iteratively enlarges the span of $C^T,H_q^T$ under every $A_u^T$. The supplied implementation does this for small local models. For unrestricted nonlinear rendering, finite closure is an open problem. ## 5.3 Why output-only state is insufficient Let the desired image depend only on $x_1$, but a later query return $y=x_1+x_2+\varepsilon$. If $x_2$ was previously learned, discarding it can destroy the ability to recover $x_1$. With prior $\operatorname{Var}(x_1)=1$, measurement variance $0.1$, and independent nuisance variance either $0$ or $1$, posterior task variance is respectively $0.09091$ or $0.52381$. Both histories can have the same current marginal for $x_1$. Thus a latent that predicts the current image perfectly well can still be an inadequate *learning state*. This is the central correction beyond “compress only what the decoder uses.” One may retain nuisance statistics, or marginalize them exactly into a sufficient update model. Simply deleting them is not exact marginalization. # 6. Architecture: AUREOLE The minimal architecture has four responsibilities: 1. **Canonical address and validity.** The engine supplies stable keys, deformation maps, generation events, and exposure conventions. 2. **Evidence assimilation.** A sparse updater maintains local response estimates and uncertainty, accounting for repeated or correlated samples. 3. **Queryable response.** Lightweight decoders project the belief to requested time, view, footprint, and task. 4. **Evidence allocation.** A controller estimates the change in future output risk from rays, state refreshes, memory retention, and optional specialist compute. The neural parts learn observation encodings, small response bases, decoder residuals, noise scales, and cheap value approximations. Canonical IDs, change signals, physical constraints, and exact bookkeeping remain explicit. No global transformer is necessary. A local graph, block covariance, and shared low-rank lighting variables form a sufficient starting hypothesis. ![Proposed architecture. The engine remains authoritative for known scene state. The shared belief supports output queries and chooses legal physical probes; the full neural/GPU loop has not been implemented.](figures/architecture.png){width=95%} An authoritative geometry pass still establishes current visibility where practical. Rendering from a persistent belief does not remove the cost of visibility, nor turn uncertain hidden topology into known geometry. The system estimates what the engine has not already supplied at acceptable cost. ## 6.1 Implemented AUREOLE-R reference The implemented subset specializes this architecture to one expensive deterministic response: visibility between a canonical floor point and a finite point emitter. Engine-owned geometry, albedo and emitter intensity give an analytic nonnegative unoccluded contribution $b_{ijc}$. A small learned prior predicts visibility $p_{ij}$. Exact previous shadow tests override this prior at remembered canonical receiver/emitter pairs, creating a fallible response estimate $v_{ij}$. The control is $h_{ijc}=b_{ijc}v_{ij}$. Material color and emitter intensity may change while the visibility evidence remains reusable; moved occluders can invalidate it. Each output batch freezes the control and a full-support proposal, draws fresh shadow queries, computes the exact physical residual correction, and only then commits new visibility evidence. A contradiction with a previously trusted binary fact can revoke the current trust epoch. This is a working causal physical loop. The generic multi-task decoders, posterior covariance learning, chronoscopic teacher, counterfactual curriculum and GPU graph described elsewhere remain specified research components. The implementation's state is explicit: a scene namespace, a finite canonical receiver/emitter dictionary represented by arrays, stored visibility, evidence epochs, a scene epoch and a clock. Network weights are a shared prior, not per-object persistent truth. The full architecture needs richer material/transport responses than binary visibility; this reference deliberately does not pretend those responses have been learned. # 7. Causal update equations and local plasticity The ideal recursion is the controlled Bayes filter: $$b^-_{t+1}(x')=\int p(x'\mid x,u_{t+1},E_{t+1})b_t(x)\,dx,$$ $$b_{t+1}(x')\propto p(O_{t+1}\mid x',q_{t+1},E_{t+1})b^-_{t+1}(x'). \tag{7}$$ The local linear-Gaussian approximation is $$m^-=A m+b,\quad P^-=A P A^T+Q,$$ $$S=H P^-H^T+R,\quad K=P^-H^TS^{-1},$$ $$m^+=m^-+K(o-Hm^-),\quad P^+=(I-KH)P^-(I-KH)^T+KRK^T. \tag{8}$$ The Joseph covariance form avoids unnecessary loss of positive semidefiniteness. A learned nonlinear encoder supplies $H$ as a local Jacobian or a learned calibrated observation map; (8) is then an approximation. Ambiguous identity requires a mixture or a conservative reset, not a single confident Gaussian update. In information form for a static block with independent observations, $$\Lambda^+=\Lambda^-+H^TR^{-1}H,\qquad \eta^+=\eta^-+H^TR^{-1}o,\qquad m=\Lambda^{-1}\eta. \tag{9}$$ The scalar-pixel equivalent is implemented in E1-E2. Shared light coefficients induce cross-correlations between surfaces; a production block-diagonal model must retain important couplings, inflate uncertainty, or document its approximation error. For a negative-log-likelihood objective, a plastic update has the sign $$\Delta m=-P\nabla_m\mathcal L.$$ Positive gradient motion would increase a loss. Under continuous linear observation, covariance satisfies a Riccati equation $$\dot P=AP+PA^T+Q-PH^TR^{-1}HP. \tag{10}$$ For locally static state this becomes $\dot P=Q-P\mathcal I P$, linking reliable evidence to rigidity and process change to reopening. The EIGENPLASTICA analogy is useful, but (10) is classical filtering. A very small $P$ after old evidence is dangerous when the world changes; generation signals or a change-point model must raise the appropriate uncertainty or replace the local prior. # 8. Memory representation, consolidation, and correction Use a sparse canonical atlas with a hot GPU working set and a compact cold pool. Surfaces consolidate repeated observations into response coefficients, information statistics, uncertainty, and provenance rather than an ever-growing stack of frames. A surfel fallback supports missing charts. A volume pool and short-lived path pool handle regimes that are not surface-local. Version dependencies are component-specific. A lighting change invalidates stale radiance coefficients, not automatically an unchanged material atlas. An object transform changes visibility and inter-object transfer even when its object-relative texture is unchanged. A topology generation change invalidates primitive correspondence. LOD transitions require an explicit map between footprint-conditioned responses; primitive IDs alone are not stable across remeshing. **Absence is not confirming evidence.** Over an unobserved interval, propagate $P$ with dynamics and process noise. With $A=I$ and $Q=0$, an actually static material estimate can remain unchanged for 500 frames or longer. With $Q\succ0$, confidence declines. Retention is bounded by capacity and expected revisit value; there is no universal 500-frame guarantee. For contradiction handling, compute the normalized innovation $d^2=(o-\hat o)^TS^{-1}(o-\hat o)$. Under the correctly specified Gaussian model this has a chi-square reference law, but heavy-tailed path samples, miscalibration, and repeated testing invalidate naive thresholds. Use held-out calibration and compare three explanations: outlier noise, correspondence failure, and a genuine state change. Maintain a short probationary hypothesis when uncertain. Do not permanently reject contradictory evidence merely because an old posterior is confident. Local correction replaces or softens only factors incident to the changed component, including known transport dependencies. Generated frames and the model's own predictions are never counted as new independent evidence. Reservoir lineage, sampling PDFs, reuse counts, and effective sample size are part of provenance. In E5, counting one ray twenty times incorrectly shrinks variance from $0.09091$ to $0.004975$. Forget a record when its expected future excess error per retained byte is low relative to competitors. This quantity is the risk difference between retaining and marginalizing that evidence, not simply the record's present uncertainty. A very certain, frequently revisited material may be exceptionally valuable to keep. # 9. World-space persistence and object permanence | Transition | Required action | What may be preserved | |:--|:--|:--| | Camera rotation or resolution change | Re-query canonical keys and output footprints. | Valid material/geometry response and evidence. | | Occlusion then return | Propagate hidden-state uncertainty; verify key and dependencies at return. | Identity and stable response within budget. | | Rigid motion | Move the chart with the object; recompute view and transport dependence. | Rest-frame material, not old world-space illumination. | | Deformation | Apply engine rest-to-current map and its Jacobian. | Material identity when mapping is valid; geometry-dependent response may change. | | Camera cut | Reset screen scratch; retain only scene-valid canonical entries. | Same-world assets with known identity. | | Destruction, teleport, respawn | Increment relevant generations; remove invalid dependencies. | Only explicitly unchanged components. | | Streaming unload/reload | Serialize compact valid evidence with world/asset versions, or evict. | Evidence that can be verified on reload. | Object permanence is a hypothesis about identity and dynamics, not a promise that invisible objects never change. A remote multiplayer event or an unseen procedural edit can make old knowledge false. Engine notifications are stronger evidence than learned extrapolation. The atlas witness evaluates exact canonical identity, not learned identity tracking. It cannot validate persistence under uncertain correspondence. Those cases are separate gates in the evaluation protocol. # 10. Unified task decoding and continuous time The shared contract is $D_k(Z_t;\text{camera},\text{time},\text{footprint},\text{exposure},\text{task settings})$, not $D_k(Z_t)$ without query metadata. Decoder outputs are estimates and reliability measures; uncertainty belongs to the requested quantity, not just the latent tensor. | Task | Query to the shared state | Essential task-specific computation | |:--|:--|:--| | Super resolution | High-resolution footprint-integrated radiance. | Subpixel visibility and antialiasing; uncertainty if no high-frequency evidence exists. | | Ray reconstruction / denoising | Conditional transport/radiance estimate given sparse path evidence. | Noise model, specular separation, bias control. They need not be separate networks. | | Frame generation | Response at an intermediate or predicted simulation time. | Visibility, animation, control timing, motion blur, UI composition. | | Neural appearance | Footprint- and direction-conditioned optical response. | BSDF evaluation, lighting dependence, physical constraints. | | Disocclusion | Recalled surface response with current visibility and generation check. | New rays for unknown surfaces; hypothesis mixtures when correspondence is ambiguous. | | Adaptive sampling | Posterior risk reduction for legal renderer probes. | Cost and deadline prediction, exploration, estimator PDF accounting. | | Compression / streaming | Encoded valid response state and uncertainty. | Quantization, synchronization, version handling, decoder compatibility. | The hypothesis is that accurate shared response makes these decoders small. This has not been established for the full task set. Specialized residuals remain legitimate; completely independent recurrent histories would defeat the intended test of shared inference. Continuous-time evolution is a **hybrid** system: $$d m=f_\theta(m,u,t)dt,\qquad \dot P=J_fP+PJ_f^T+Q, \qquad Z(t_e^+)=\mathcal R_e(Z(t_e^-),E_e). \tag{11}$$ The reset maps handle cuts, impacts, topology changes, light switches, spawns, and discontinuous game events. Known engine interpolation is preferable to learned ODE integration for deterministic transforms. A continuous neural field alone cannot represent an arbitrary instantaneous visibility change without event handling. For a known simulation segment, frame generation can be implemented as a time query followed by projection and visibility evaluation. This unifies the interface, but does not erase the epistemic difference between an observed frame and a predicted one. Causal extrapolation at $t+\tau$ cannot know future input, packet arrival, or a random event unavailable at $t$. Interpolation between two simulation states uses later evidence and entails latency. Both modes must be evaluated separately. Motion blur is an exposure integral $I=\int s(\tau)D(Z(t+\tau))d\tau$ with normalized shutter function $s$. Its quadrature cost and visibility changes must be budgeted. HUD, text, cursor, and latency-sensitive overlays should use the current authoritative UI state, not a hallucinated world continuation. ## 10.1 A readout boundary that cannot be skipped The same belief can supply either a direct neural prediction or a physical control variate. These have different correctness and variance contracts. The implemented unbiasedness result applies to the linear direct-light output with an exact finite integral and fresh physical residuals. It does not automatically transfer through a learned SR decoder, tone mapper, denoiser, or speculative FG model. To apply the same contract to spatial or temporal integration, the physical sampling domain and its oracle must include the requested footprint or time. Unknown future player inputs cannot be physically queried from the current causal engine state. # 11. Chronoscopic teacher training Train an offline smoother $p_T(x_t\mid H_t,O_{t+1:t+k})$ using known simulation snapshots, dense references, and future evidence. Train the causal student $p_S(x_t\mid H_t)$ using proper distributional losses and physically meaningful decoded probes. Future data are never supplied at inference. ## Proposition P3: correct future-teacher target (Proved) Assume the teacher is the true conditional posterior and its conditioning includes $H_t$. Over the true distribution of future evidence $F$, the minimizer of $$\mathbb E_{F\mid H_t}D_{\rm KL}\big(p_T(x_t\mid H_t,F)\Vert p_S(x_t\mid H_t)\big) \tag{12}$$ is $p_S(x_t\mid H_t)=p(x_t\mid H_t)$, on common support. **Proof.** The student-dependent term is cross entropy with the mixture $\mathbb E_{F\mid H_t}p_T(x_t\mid H_t,F)$. By the tower property this mixture equals the causal posterior. Cross entropy is minimized by that distribution. $\square$ For squared-error point prediction, the optimum is the causal conditional mean. It does not recover the future teacher's realization-specific knowledge. If a hidden bit $B\in\{-1,1\}$ is independent of causal history, a future teacher may reveal it exactly, while every causal point predictor has MSE at least one. E5 verifies the limiting example. Therefore a direct loss $\|Z_t-Z_t^*\|^2$ is inadequate unless state coordinates are aligned and teacher-only uncertainty is represented. Free latent spaces have gauge freedom; compare anchored material/geometry variables or distributions of future probes. A fixed-window teacher can even discard older information available to the student; either include that history or acknowledge the approximation. For thin geometry, foliage, reflections, and disocclusions, use future views to *label what was present at time $t$*. An object spawned later is not evidence that it existed earlier. Save engine snapshots and event times to distinguish retrospective observation from genuine evolution. Reference paths, sampling seeds, and future camera metadata used for labeling must be inaccessible to the causal student. # 12. Counterfactual camera training and identifiability At a saved world snapshot, replay multiple camera paths and exposure/footprint queries while holding the world timeline and permitted controls fixed. The student consumes one causal prefix; the decoder is asked to explain all counterfactual queries. The teacher may inspect the complete scene for target generation. $$\mathcal L_{\rm CF}=\mathbb E_{h,\mathcal T\sim\Pi_{\rm train}} [-\log p_\theta(Y_{\mathcal T}\mid Z(h),\mathcal T)]. \tag{13}$$ A mixture over possible hidden worlds is appropriate where the prefix is ambiguous. Penalizing a causal student for not guessing an unobservable hidden texture encourages hallucination. Counterfactual labels create a useful prior across training scenes; they do not add test-time information to a particular scene. **Derived under assumptions: linear identifiability.** For a parameter vector $x$, stack counterfactual response maps into $M$. Noiseless parameters are identifiable modulo known symmetries precisely when $\ker M$ contains only the declared gauge directions. With noise covariance $R$, conditioning is governed by $M^TR^{-1}M$, especially its smallest nonzero eigenvalue. This follows because two states are observationally equivalent iff their difference lies in $\ker M$. Nonlinear rendering admits albedo-lighting ambiguity, hidden geometry, view-dependent effects, and gauge symmetries. Multiple trajectories do not guarantee identifiability. Measuring a second view of the same diffuse patch under the same light does not generally separate material from illumination. Engine material/light handles, controlled illumination in training, or directional probes can break specific ambiguities. Counterfactual tests should include paths outside the training camera distribution and expose both object re-identification and unknown-region uncertainty. Held-out scenes, assets, material seeds, and trajectories must all be separated to prevent texture memorization from masquerading as inference. # 13. Active rendering and the information economics of rays The controlling quantity is expected downstream loss reduction per total cost. Entropy reduction is useful only when it aligns with the outputs that matter. A highly uncertain invisible nuisance can have no direct image value, or high indirect value through future mixed queries. ## 13.1 Future rendering metric For local scene uncertainty $x\sim\mathcal N(m,P)$, a known dynamics linearization $\Phi_\tau$, task Jacobian $J_{\tau k}$, and positive semidefinite loss weights $Q_{\tau k}$, define $$W_t=\mathbb E_{\mathcal T\sim\Pi_t}\sum_{\tau,k} w_{\tau k}\Phi_\tau^T J_{\tau k}^TQ_{\tau k}J_{\tau k}\Phi_\tau\succeq0. \tag{14}$$ For exact linear outputs and quadratic losses, the part of Bayes risk due to uncertainty in the present state is $\operatorname{tr}(W_tP)$. Future process noise adds a term independent of the present estimate under the stated model. In a nonlinear renderer, (14) is a local approximation, especially fragile at visibility changes. The forecast distribution $\Pi_t$ must be based on current information. A known prerecorded camera path is allowed in a controlled benchmark but is not equivalent to predicting an interactive player. The metric should average plausible paths or optimize against a bounded uncertainty set. ## 13.2 Proposition P4: common value of evidence (Proved) Let $\mathcal F$ be current information, $\mathcal G$ newly acquired evidence, and $W\succeq0$ a fixed metric measurable from information that is retained. For any square-integrable state, with Bayes mean decoders, $$R(\mathcal F)-\mathbb E[R(\mathcal F\vee\mathcal G)\mid\mathcal F] =\mathbb E[\|m_{\mathcal F\vee\mathcal G}-m_{\mathcal F}\|_W^2\mid\mathcal F]\geq0. \tag{15}$$ For Gaussian local state and one independent scalar observation $y=h^Tx+\varepsilon$, $\operatorname{Var}\varepsilon=r>0$, $$\Delta R(q)=\frac{h^TPWPh}{r+h^TPh},\qquad V(q)=\frac{\Delta R(q)}{c_q+c_{\rm ingest}+c_{\rm sync}}. \tag{16}$$ **Proof.** Conditional expectation is an orthogonal projection in quadratic loss. Write $x-m_{\mathcal F}=(x-m_{\mathcal F\vee\mathcal G})+(m_{\mathcal F\vee\mathcal G}-m_{\mathcal F})$; the conditional cross term vanishes. For the scalar Gaussian observation, conditioning gives $P^+=P-Phh^TP/(r+h^TPh)$. Taking the trace with $W$ proves (16). $\square$ The sum of several task metrics has additive value for one query, provided the tasks and weights represent distinct declared losses. This gives a precise meaning to a ray improving several downstream outputs. One evidence update is performed; several decoders benefit. The same argument does not justify adding many redundant names for one image metric. These equations are exact for a fixed forecast/decoder family without adaptive future re-estimation, or with a fixed linear influence map already included in $W$. In a full future filtering loop, the later Kalman gains and sampling policy can change after the query. Then the exact value is a belief-space Bellman value difference, and (16) is a one-step surrogate. No global optimality is claimed for that surrogate. ## 13.3 Retention, compression, and scheduling in the same units For coarsened memory $\mathcal F_c\subset\mathcal F$, expected loss of forgetting is $$\mathbb E\|m_{\mathcal F}-m_{\mathcal F_c}\|_W^2. \tag{17}$$ The law of total covariance decomposes the coarse posterior into retained posterior uncertainty plus uncertainty about the forgotten posterior mean. Thus retention value is **not** $\operatorname{tr}(WP)$ for the retained record; it is the *increase* in risk caused by losing its evidence. For non-Gaussian beliefs the covariance order is an expectation over forgotten information, not necessarily a pointwise ordering for every realized history. For approximately zero-mean compression error with covariance $\Xi$, excess output distortion is $\operatorname{tr}(W\Xi)$. If a stale update adds covariance $\Delta P$, its local penalty is $\operatorname{tr}(W\Delta P)$. Bias contributes $b^TWb$ and must also be tracked. The common object is $W$ and expected change in error; the covariances for rays, forgetting, and quantization are different. This distinction prevents a tempting but incorrect unification: posterior covariance describes what the system does not know, while a distribution of stored posterior means describes information that an encoder may actually compress. ## 13.4 Non-additivity of queries and a greedy failure For $P=I$, $W=\operatorname{diag}(1,0)$, $h_1=(1,1)^T$, $h_2=(0,1)^T$, and $r=0.1$, the second query alone has zero task value. After the first query its value is $0.36350$. Therefore diminishing returns fails in general. A generic greedy $1-1/e$ guarantee would be false for this objective. **Restricted proposition P5 (Proved).** If latent coordinates and measurement noises are independent, every query measures one coordinate, each query has equal cost, and $W$ is diagonal and fixed, sequentially selecting the largest exact marginal reduction yields an optimal integer sample allocation. Each coordinate's variance is $(p_i^{-1}+n_i/r_i)^{-1}$; its successive reductions decrease with $n_i$. The allocation selects the largest available reductions from these ordered lists. An exchange of a smaller selected reduction for a larger unselected feasible reduction cannot worsen feasibility and improves the objective. This proves optimality. E3 satisfies these restrictive assumptions. $\square$ For correlated real scenes use batched lookahead, approximate optimal-design solvers, or occasional jointly valuable probe pairs. Keep a nonzero exploration budget to discover changes that the current model wrongly believes impossible. ## 13.5 Sampling remains a valid Monte Carlo experiment Adaptive sampling changes proposal probabilities and may introduce selection bias. Record the proposal/PDF, ray lineage, and stopping rule. If an unbiased integral estimator is desired, proposals must retain support and the estimator must use the correct weights. Avoid optional-stopping claims for a naive average when stopping depends on sample values. A separate pilot batch may choose the production allocation. An explicit exploration mixture $p(q)=(1-\epsilon)p_{\rm value}(q)+\epsilon p_{\rm base}(q)$ preserves support where $p_{\rm base}>0$. Choosing the value of $\epsilon$ is an empirical budget tradeoff. A biased low-noise display reconstruction and an unbiased reference estimator are different outputs and must be labeled accordingly [R17]. ## 13.6 Value depends on the output contract The original future metric $W$ measures error of a plug-in prediction. A physically corrected output has a different conditional variance geometry, $G$, derived in P12. If one set of physical evidence serves both readouts, score it using their declared weighted sum rather than silently using image-prediction loss for every decision. The exact local Gaussian formula remains applicable with the correct metric. The executable active controller uses a cheaper heuristic and an exploration floor, so its performance must be measured rather than inferred from that optimum. # 14. Physics constraints and a restricted optical closure Physics constraints should operate on quantities for which the engine's rendering model has a meaningful physical interpretation. Stylized effects, tone mapping, screen-space flares, and artistic non-energy-conserving shaders should be identified explicitly. Forcing them into an energy-conserving optical model changes the authored scene rather than reconstructing it. Low-cost constraints are canonical correspondence, valid generation IDs, footprint consistency, nonnegative radiance, nonnegative scattering weights, and correct exposure conversion. Reflectance integrals can be bounded by one for passive materials; radiance itself need not be at most one. Focused light, emission, and HDR values can be large. ## 14.1 Defining neural optical G-closure without overclaiming Fix materials, proportions, geometric scale constraints, wavelength regime, and boundary conditions. Let $\mathfrak M$ be the admissible unresolved microstructures and $\mathcal T_m$ their boundary light-transport operators. For a declared measurement topology define $$\mathfrak G_{\rm opt}=\overline{\{\mathcal T_m:m\in\mathfrak M\}}. \tag{18}$$ This definition does not characterize the set. Three-dimensional conductivity G-closure theorems do not automatically transfer to wave optics, incoherent radiative transfer, nonlinear shading, or directional visibility. The optical problem has different states, constraints, and observables. Discretize incident/outgoing channels in a *power-normalized* basis. For a passive, nonemissive, reciprocal system in a matched reciprocal basis, useful necessary conditions are $$T_{ij}\geq0,\qquad \mathbf1^TT\leq\mathbf1^T,\qquad T=T^T. \tag{19}$$ With unmatched quadrature weights reciprocity is a weighted relation, not ordinary symmetry. Fluorescence, wavelength conversion, magneto-optical nonreciprocity, participating emission, and omitted channels require a different domain. Conditions (19) are an outer relaxation and generally do not prove realizability from prescribed materials. ## 14.2 Proposition P6: realizable area-mixture inner family (Proved) Suppose independently shaded patches with operators $T_1,\ldots,T_K$ tile a subpixel footprint with area fractions $\alpha_k\geq0$, $\sum_k\alpha_k=1$. Assume incoherent geometric optics, uniform incident channel fields across patches, negligible lateral inter-patch coupling and mutual shadowing, and that the measurement averages outgoing power over the footprint. Then the effective response is $$T_{\rm eff}=\sum_k\alpha_kT_k. \tag{20}$$ It is realizable in this restricted construction and preserves positivity, passivity, and matched-basis reciprocity if every constituent does. **Proof.** Incoming illumination acts independently on each patch. Outgoing averaged power is the area-weighted sum of patch responses, giving (20). The displayed constraints are linear or convex and are preserved under the sum. $\square$ A learned simplex decoder can therefore predict within this certified inner family. Fixed material-fraction constraints restrict the allowed coefficients. It is not a solution of general optical G-closure. Hair, leaves, pores, and dense fibers with self-shadowing may violate the independence assumption. Their effective operator can be nonlocal in position, direction, and time; an ordinary BRDF may be insufficient. Neural appearance and asset transport already provide relevant antecedents [R14, R15]. ## 14.3 Transport modes For fixed geometry/materials and linear radiative transport, $L=\mathcal T e$ is linear in source emission $e$. A low-rank approximation gives $L(x,\omega,t)\approx\sum_k a_k(t)\phi_k(x,\omega)$. Changes in emitter *intensity* within the fixed basis may be cheap. Moving an occluder, changing a material, or moving an emitter outside the basis changes the transport operator and can require new evidence. Choose modes from a loss-weighted response SVD or learned basis and measure the residual on held-out directions and emitters. High-frequency specular transport and caustics may need high rank. “Low rank” is an experimental property, not a general law of light transport. # 15. Information-theoretic interpretation and totality input An information bottleneck can seek $\min I(Z;H\mid E)$ subject to predictive loss bounds for the declared query family. For deterministic continuous states, this mutual information can be infinite. A practical objective needs quantization, a stochastic encoder, or a code-length model. World-space state is valuable because identity aligns repeated information; it does not make all observed bits useful. The conditional value of an auxiliary channel $S$ is evaluated after ordinary inputs: for proper log loss it is $I(Y;S\mid H)$, and for squared prediction loss it is the conditional-mean improvement in (15). A random seed independent of scene state has no standalone scene information. Coupled to a known simulator, proposal, and observed path, it may help replay a sample or explain correlated noise. | Additional renderer signal | Potential use | Required correction or rejection test | |:--|:--|:--| | Object/primitive/material IDs and barycentrics | Canonical correspondence and invalidation. | Generations, LOD remapping, instancing, ID collisions. | | Roughness, albedo, BSDF parameters, anisotropy | Explain response and select compact bases. | Preserve authored conventions and energy normalization. | | Path length, hit/miss, termination reason | Visibility and path-class evidence. | Account for proposal, truncation, roulette, and censoring. | | Rejected light candidates and shadow tests | Additional response/visibility constraints. | Rejection is selection-biased; include reason, PDF, and threshold. | | Reservoir candidates and ancestry | Sample support and guiding. | Correlation and repeated evidence; accepted/rejected samples are not independent. | | Variance estimates, residuals, rejection masks | Identify unexplained changes or model failure. | Residuals derived from the same RGB are not independent new measurements. | | LOD/mip history and ray differentials | Footprint-conditioned subpixel response. | Old footprints do not identify newly requested high frequencies. | | BVH update/refit and topology event information | Local dependency invalidation. | Expose compact events, not raw acceleration-structure bandwidth by default. | | Neighbor sample statistics | Local regularity and noise scale. | Cross-pixel correlations and geometry discontinuities. | | Seed/replay metadata | Reproducibility, de-correlation, conditional noise inference. | Test whether it adds information after all existing channels. | For jointly Gaussian base observation $o$ and extra signal $s$, conditional innovation is $s-\mathbb E[s\mid o]$, with covariance $S_{ss}-S_{so}S_{oo}^{-1}S_{os}$, where $S$ is the full predictive observation covariance, including state uncertainty. Alternatively, for $o=H_ox+\varepsilon_o$ and $s=H_sx+\varepsilon_s$ with noise covariance blocks $R$, decorrelate the added sensor using $s'=s-R_{so}R_{oo}^{-1}o$ and $H_s'=H_s-R_{so}R_{oo}^{-1}H_o$. Its remaining noise covariance is $R_{ss}-R_{so}R_{oo}^{-1}R_{os}$. Updating the already-conditioned state with this transformed channel avoids double counting. E6 demonstrates miscalibration when correlated channels are incorrectly treated as independent. “Totality” is best interpreted as **evaluate every accessible signal for conditional value**, not “retain everything.” A signal with negligible risk reduction or excessive collection bandwidth should be omitted. A leave-one-channel-out test alone can miss redundancy and synergy, so also test paired and conditional additions. Logging all rejected paths at full resolution may cost more than it saves. # 16. Multi-rate consolidation and memory plasticity Fast variables include visibility, transforms, screen mapping, and display time. Medium variables include local lighting and short transport residuals. Slow variables include stable filtered material response and local geometry statistics. Persistent variables include verified static asset response and compact dependency metadata. These categories determine default schedules, not universal periods. A light can be static for hours and then switch instantly; a normally slow material can animate every frame. Event-triggered invalidation takes precedence over periodic updates. **Derived under assumptions: update interval.** Suppose a block's uncertainty grows as $P(a)=P_0+aQ$ with age $a$, updates reset the same uncertainty component, each costs $c$, and the task metric is constant. Under a periodic interval $\Delta$, average age is $\Delta/2$. Minimizing cost per unit time plus weighted stale risk gives $$J(\Delta)=\frac{c}{\Delta}+\lambda\frac{\alpha\Delta}{2},\quad \alpha=\operatorname{tr}(WQ),\qquad \Delta^*=\sqrt{\frac{2c}{\lambda\alpha}}. \tag{21}$$ The derivative is $-c/\Delta^2+\lambda\alpha/2$ and the positive critical point is the minimum. Clamp to legal deadlines and discrete ticks. For $\alpha=0$, periodic refreshing has no benefit in this model; rely on events. For jumps, changing visibility, nonlinear dynamics, or imperfect resets, (21) is a heuristic and a different age-cost model is needed. It explains why one fixed update rate is generally inefficient. Consolidation tracks sufficient information, not confidence by repetition of the network's own answers. A cache entry may become more stable after independent observations, but learned-prior confidence and measurement information remain separately identifiable. Corrections can reopen plasticity without unlearning unrelated geometry. Compression should preserve uncertainty about what it removes. # 17. Phase routing and dormant specialists This release does not require a new capability-manifold theory. It uses a restrained interpretation: local rendering regimes determine which response bases and specialist decoders are useful. A continuous gate can mix physically admissible experts: $$\alpha_e=\operatorname{softmax}(s_e(Z,O)/T),\qquad \widehat T=\sum_e\alpha_e T_e. \tag{22}$$ In the convex optical family of Section 14, such mixing preserves its listed constraints. Arbitrary image blending does not guarantee valid visibility or geometry. Gate rates may be bounded between events to reduce artificial flicker, but a real light switch or topology change must be allowed to change output quickly. Excess smoothing creates lighting lag. Hair, water, fire, skin, caustics, transparency, and foliage are candidates for specialist execution. Route using engine material/path flags, estimated error, and task value rather than an expensive semantic model by default. Test semantic features only if their conditional improvement exceeds their cost. Dormant pathways mean *trained rare-regime capacity that is usually not executed*. Their parameters still consume memory, and routing incurs overhead. Frozen rare-regime experts or a protected rehearsal buffer can reduce catastrophic forgetting during later training. Reserving arbitrary unused neurons does not itself establish useful evolvability. The benefit of dormant pathways is an **experimental prediction**, and they are excluded from the minimal decisive prototype until the shared-memory claim survives. # 18. Formal results inventory and sample-efficiency limits | ID | Result | Scope and novelty boundary | |:--|:--|:--| | P1 | Coarsest rendering-and-observation predictive quotient. | Exact definition and proof; classical predictive-state principle specialized to graphics. | | P2 | Linear observation-closed invariant quotient. | Finite known linear dynamics/channels; no nonlinear compactness guarantee. | | P3 | Future posterior distillation averages to the causal posterior. | Correct teacher, included history, proper distributional objective. | | P4 | Evidence, forgetting, and one-query risk identities. | Quadratic Bayes loss; scalar Gaussian closed form. | | P5 | Optimal greedy allocation in a separable model. | Independent coordinates, diagonal metric, equal query costs. Fails generally. | | P6 | Restricted realizable area-mixture optical family. | Incoherent independent patches; not full optical G-closure. | | P7 | Shared-evidence covariance advantage. | Common correct parameter model and independent valid observations. | | P8 | Persistent-observation gain with process-noise floor. | Static scalar/Gaussian case and simple random-walk extension. | | P9 | Optimal task-weighted rank-$r$ transform coding. | Accessible Gaussian source; no claim about unknown latent innovations. | | P10 | Conditional memory-error and temporal-error bound. | Contractive update and locally Lipschitz decoder away from discontinuities. | ## 18.1 Proposition P7: shared evidence (Proved) With common prior precision $\Lambda_0\succ0$ and independent observation groups whose information matrices are $\mathcal I_j\succeq0$, $$P_{\rm shared}=(\Lambda_0+\sum_j\mathcal I_j)^{-1} \preceq (\Lambda_0+\mathcal I_i)^{-1}=P_i. \tag{23}$$ Consequently $\operatorname{tr}(W_iP_{\rm shared})\leq\operatorname{tr}(W_iP_i)$ for every $W_i\succeq0$. **Proof.** Adding positive semidefinite information increases precision. Inversion reverses the positive-definite order; trace pairing with a positive semidefinite matrix preserves the inequality. $\square$ This does **not** prove that a shared neural network universally beats separate networks. If all separate task models already receive the same observations and compute exact posteriors, they can match shared inference statistically. Savings may then be computation or storage only. A scalar example with $M$ disjoint groups of $n$ noisy measurements gives variance $\sigma^2/(Mn)$ versus $\sigma^2/n$ for a group-restricted estimator, but broadcasting all $Mn$ samples closes that gap. This is not a free $M$-fold ray saving at matched information access. Harmful interference arises from misspecified parameter sharing, biased priors, conflicting losses, limited capacity, and optimization. Compare per-task gradients and Pareto fronts; preserve task-specific residuals if necessary. There is no architecture-level guarantee of positive transfer. ## 18.2 Proposition P8: persistence and its floor (Proved) For a fixed scalar surface response with prior variance $p_0$ and $n$ independent observations of variance $\sigma^2$, $$p_n=(p_0^{-1}+n/\sigma^2)^{-1}. \tag{24}$$ Discarding earlier independent observations cannot improve the correctly specified Bayes risk. To attain $p_n\leq\epsilon