# Changelog All notable changes to this project will be documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). ## [Unreleased] ## [1.3.0] - 2026-09-23 ### Added - **Fused Single-Pass Kernel with Post-Step EMA (`fused_single_pass`, `ema_decay`)**: - Introduced `_lumina_v2_fused_single_pass_kernel` coupled with an asynchronous CTA micro-kernel `_lumina_v2_post_step_ema_kernel[(1,)]`. - When `fused_single_pass=True`, completely eliminates two-pass global reductions by leveraging running Exponential Moving Average (EMA) estimates of `cautious_scale` (1 / M_bar) and `bound_scale` (tanh(r) / r). - Unifies gradient filtering, momentum accumulation, variance dispersion, cautious masking, radial bounding, decoupled weight decay, stochastic rounding, and block partial reductions into a **single unified VRAM-resident pass**, slashing global GPU memory bandwidth traffic per optimizer step. - **Hybrid Master Weight Precision Architecture (`master_weights`)**: - Added multi-tier master weight flexibility via `master_weights: Union[bool, str] = False`: - `"none"` / `False` *(Default)*: Pure master-free low-precision training leveraging on-chip bitwise stochastic rounding directly on native half-precision weights. - `"semi"` / `"hybrid"` / `"half"`: 32-bit FP32 master parameters (`state["master_param"]`) coupled with 16-bit low-memory optimizer states (`exp_avg`, `exp_avg_sq` in BF16/FP16), saving ~50% optimizer state VRAM compared to traditional FP32 setups while maintaining full FP32 weight accumulation accuracy. - `"full"` / `"fp32"` / `"masterweight"`: Traditional full FP32 master parameters with full FP32 optimizer states for standard baseline comparisons. - **In-Kernel In-Flight NaN / Inf Gradient Sanitization**: - Embedded register-level defensive sanitization across all Triton GPU kernels: `is_finite_g = (g == g) & (g < 1e38) & (g > -1e38); g = tl.where(is_finite_g, g, 0.0)`. - Instantly neutralizes gradient anomalies, exploding activations, and subnormal corruptions at zero latency before they can contaminate momentum (`exp_avg`) or variance (`exp_avg_sq`) state buffers. - **Low-Precision State Buffer Serialization (`_store_state_buffer`)**: - Added a dedicated Triton state serialization routine supporting FP32, BF16, and FP16 state representations with optional bitwise stochastic rounding applied directly to momentum buffers. - **Multi-Device & Cross-Accelerator Hardware Context Isolation (`_get_device_context`)**: - Unified device stream scoping supporting NVIDIA CUDA (`torch.cuda`), Intel XPU (`torch.xpu`), and modern PyTorch 2.x unified accelerator interfaces (`torch.accelerator.device`), eliminating multi-GPU stream hijacking and cross-device race conditions. - **Strict Determinism Auto-Routing**: - Automatically queries `torch.are_deterministic_algorithms_enabled()`. When enabled under `execution="auto"`, dynamically routes execution to C++ vectorized `foreach` to guarantee 100% bitwise cross-run reproducibility, circumventing non-deterministic floating-point hardware atomic reordering. - **Granular Fault Recovery with State Guarding (`completed_indices`)**: - Implemented per-parameter completion tracking in `_dispatch_triton`. If a kernel encounters an unexpected runtime fault, completed parameters are preserved while uncompleted ones fall back cleanly to `foreach`. - Added `_single_step_v2_apply_param` and `_single_step_v1_apply_param` to safely finalize parameter updates if Pass 1 had already committed state transitions, strictly preventing double-update state corruption. ### Changed - **Per-Parameter Persistent Scratch Tensors (Elimination of Global Slicing)**: - Replaced the global dynamic slice allocator (`self._scratch_tensors[dev][:count]`) with isolated per-parameter persistent state allocations (`state["scratch"] = torch.zeros((5,), device=p.device, dtype=torch.float32)`). - Memory pointers for `[mask_sum, u_sq_sum, p_sq_sum, cautious_scale, bound_scale]` remain completely invariant throughout the entire training lifecycle, providing rock-solid stability for **CUDA Graph Capture** (`torch.accelerator.Graph`) and TorchInductor AOT compilation without memory aliasing. - **Analytically Exact Nesterov Bias Correction (beta1^(step + 1))**: - Corrected first-moment bias correction from the classical Adam formulation (`1.0 - beta1^step`) to the exact analytical expectation of the Nesterov blended momentum: `bc1_nes = 1.0 - beta1^(step + 1)`. Completely eradicates residual momentum bias during early warmup iterations. - **Adaptive CTA Tile Sizing and Dynamic Thread Warp Scaling**: - Introduced hardware-aware CTA tile scaling based on parameter volume: parameters with >= 2,097,152 elements (2M+) dynamically configure `BLOCK_SIZE = 2048` and `num_warps = 8` (versus `BLOCK_SIZE = 1024`, `num_warps = 4` for smaller tensors). - Cuts total thread-block grid counts and L2 cache atomic addition pressure by 50% on massive MLP, Attention Projection, and Vocabulary layers. - **Hardware-Native Tanh Instruction Dispatch (`tl.math.tanh` / `tl.tanh`)**: - Upgraded `_triton_tanh_fast` to dynamically inspect compiler capabilities, binding directly to native PTX Special Function Unit (SFU) instructions (`tanh.approx.f32`) on modern Triton compilers while retaining analytical sigmoid fallbacks for legacy environments. - **PRNG Decorrelation between Momentum and Parameter Updates**: - Separated RNG seed streams using bitwise alternating mask shifts: `seed_p = (seed + 0x55555555) & 0x7FFFFFFF`, mathematically eliminating cross-coupling noise correlation between state buffer rounding and parameter update rounding. - **Bumped package release version metadata to 1.3.0** across `luminav.py`. ### Fixed - **Catastrophic Floating-Point Cancellation in Radial Bounding (r -> 0)**: - Fixed precision jitter and roundoff degradation near zero by adding an asymptotic Taylor-series guard across Triton and PyTorch bounding engines: `torch.where(r < 1e-4, 1.0, tanh(r) / r)`. - **Artificial Denominator Dilation in Dispersion Scale (sigma)**: - Replaced additive epsilon inflation (`sigma + 1e-6`) with a non-distorting floor clamp (`tl.maximum(sigma, 1e-12)`), preventing artificial dampening of subtle gradient signals. - **Singular Zero-Division in Parameter Norm Clamping**: - Replaced `norm_p + 1e-6` with `torch.clamp_min(norm_p, 1.0)` / `tl.maximum(norm_p, 1.0)`, guaranteeing complete numerical stability even when parameters are initialized near zero. - **Transient Memory Pool Spikes during PyTorch Fallback Stochastic Rounding**: - Refactored `_sr_update` to stream stochastic noise generation in bounded chunks of 1,048,576 elements (~4 MB chunk size), eliminating multi-hundred megabyte temporary int32 tensor spikes during fallback execution on large parameter matrices. - **Non-Contiguous Buffer Safety & Strided Memory Copy**: - Added strict contiguous layout verification for parameters, states, and master weights with in-place synchronization (`p.copy_(p_contig)`), guarded by a one-time developer warning. ### Performance & Engine Refactoring - **~50% Global Memory Bandwidth Reduction via Fused Single-Pass Mode**: - By unifying parameter loading, gradient evaluation, state transformation, and bounding into one GPU pass under `fused_single_pass=True`, VRAM read/write cycles are halved for cautious and radially bounded steps. - **Sub-Microsecond L2 Hardware Atomic Reduction Architecture**: - Formally documented the hybrid Intra-Block Tree Reduction + Inter-Block Hardware Atomic pattern. With `BLOCK_SIZE=2048`, a 10M-parameter tensor issues only 4,883 atomic additions directly into GPU L2 cache crossbars, executing in < 0.5 µs and strictly outperforming multi-tier cascaded reduction micro-kernels by eliminating 3–5 µs driver dispatch bubbles. - **Static Buffer Address Invariance for CUDA Graphs & TorchInductor AOT**: - Persistent 5-slot state scratch buffers eliminate all dynamic slicing and memory reallocations during optimization loops, achieving 100% address invariance for CUDA Graph replay. --- ## [1.2.1] - 2026-09-17 ### Added - **In-Kernel GPU Parameter Norm Reduction (`p_sq_sum_ptr`)**: Fused parameter Euclidean norm squared `∑(p_i²)` calculation directly into Pass 1 Triton reduction kernels (`_lumina_v2_pass1_kernel` and `_lumina_v1_pass1_kernel`). Pass 2 and Pass 3 now load and resolve `norm_p` (`tl.sqrt(tl.maximum(p_sq_sum, 0.0))`) entirely inside GPU SRAM/registers without host intervention. - **Stateless CPU Integer Multiplicative PRNG Hashing**: Replaced dynamic GPU RNG tensor allocation with a pure CPU integer hash using Knuth's 32-bit golden ratio constant: `(step * 0x9E3779B9 + i * 10007) & 0x7FFFFFFF`. Generates uncorrelated 31-bit integer seeds per layer with zero GPU kernel launch overhead and zero PCIe host-device latency. - **Isolated Generator Cache for Fallback Engine (`self._generators`)**: Added an internal per-device `torch.Generator` cache, ensuring stochastic rounding in `_sr_update` operates in complete isolation from PyTorch's default CUDA generator state. ### Changed - **Consolidated Scratch Buffer Zeroing**: Eradicated redundant per-parameter slice zeroing inside parameter loops. Buffer zeroing is now batched into a single consolidated `scratch.zero_()` call per step inside `_get_scratch_buffer`. - **Scratch Buffer Stride Expansion**: - `buffer=2`: Expanded stride from 2×N to 3×N elements per parameter to host `[mask_sum, u_sq_sum, p_sq_sum]`. - `buffer=1`: Expanded stride from 3×N to 4×N elements per parameter to host `[rms_sum, mask_sum, u_sq_sum, p_sq_sum]`. - **Bumped package release version metadata to 1.2.1** across `luminav.py` and `pyproject.toml`. ### Fixed - **Host-Device Synchronization & Pipeline Stalls (`.item()` Bottleneck)**: Eradicated `norm_p = float(p_contig.float().norm().item())` and `seed = torch.randint(...).item()` from Triton step pipelines. Eliminated 200–400 synchronous PCIe roundtrips and GPU pipeline stalls per step on deep architectures (e.g., 28-layer Transformers), restoring full asynchronous GPU compute utilization. - **CUDA Graphs & `torch.compile` Graph Breaks**: Fixed fatal graph capture crashes and TorchDynamo graph breaks caused by blocking host-device synchronization during radial bounding and seed generation. The optimizer step is now 100% compliant with CUDA Graphs (`torch.accelerator.Graph`) and TorchInductor compilation. - **Transient VRAM Allocator Thrashing**: Eliminated repeated temporary FP32 parameter cloning (`p_contig.float()`) during norm evaluation, preventing memory pool fragmentation in tight VRAM regimes. - **Global RNG Pollution & Determinism Compliance**: Fixed global RNG state mutation in fallback stochastic rounding. Solved reproducibility drift in upstream Dropout, DataLoader shuffling, and RLHF reward modeling, ensuring strict adherence to PyTorch 2.10+ `torch.use_deterministic_algorithms(True)`. ### Performance & Engine Refactoring - **Batch Kernel Launch Overhead Elimination**: Removed up to 400–800 individual `cudaMemsetAsync` micro-calls per step previously triggered by per-slice `.zero_()` invocations on 1-element scratch tensors. - **Zero-Reduction Bypass for Coordinate Bounding**: Streamlined `_triton_step_v2` to immediately bypass two-pass reductions and invoke `_lumina_v2_single_pass_kernel` when `cautious=False` and coordinate bounding (`bound_mode == 2`) is selected. --- ## [1.2.0] - 2026-09-17 ### Added - **Automatic Dynamic FP16 Cliff Governor:** Built-in asymptotic boundary governor (`_triton_tanh_fast` in GPU register and PyTorch fallback) that automatically activates when `p.dtype == torch.float16`. Prevents unclipped steps from pushing parameters, activations, and residual streams into the IEEE-754 FP16 overflow cliff (> 65,504). Remains completely dormant (zero overhead) on `bfloat16` (TPU / Ampere+) and `float32`. - **Smooth Asymptotic Parameter Bounding (`bound`, `bound_type`, `bound_ratio`):** Added non-intrusive step bounding with two mathematically continuous formulations: - `"radial"` (Default): Direction-Preserving Radial Squashing (`update * tanh(r) / r`) maintaining 100.0000% gradient angular fidelity (Cosine Similarity = 1.000). - `"coordinate"`: Dual-Hyperbolic Coordinate Squashing for fine-grained per-element damping. - Fully configurable via constructor: `bound: bool = True`, `bound_type: str = "radial"`, `bound_ratio: float = 0.03`. ### Changed - **Increased Default `cautious_clamp_min` from `0.2` to `0.5`:** Enforces an exact power-of-two maximum energy scaling factor (2.00x, 2^1). Eliminates floating-point mantissa roundoff distortion in IEEE-754 hardware. - Bumped package release version metadata to `1.2.0` across `luminav.py` and `config.json`. ### Fixed * **Catastrophic FP16 Step Divergence & NaN Explosion at High Learning Rates:** Solved the step divergence and subsequent NaN crashes when training deep 28-layer Transformer architectures (e.g., Qwen3-0.6B) in pure FP16 at nominal learning rates (lr = 8e-4). * **Sparse Embedding & LM-Head Amplification:** Mitigated 5x over-amplification on sparse vocabulary gradients where active token counts are much smaller than total tensor parameters (N >> N_active). * **Verified stable 1,000-step pretraining from scratch on Qwen3-0.6B** (seq length 512, pure FP16, lr = 8e-4) on Nvidia Tesla T4 with zero gradient clipping and zero LR scheduler (loss dropped from 12.09 down to ~5.7 with zero NaNs). * **Weight Decay Numerical Parity on Fallback Engine:** Resolved a silent precision-stagnation issue in `_single_step` and `_foreach_step`. Previously, decoupled weight decay (`p.mul_(1.0 - lr * weight_decay)`) was executed directly on native FP16/BF16 tensors prior to `_sr_update`. Due to IEEE-754 Round-to-Nearest (RTN) limits, `(1 - lr * weight_decay) ≈ 0.999936` rounded back to `1.0`, causing weight decay to freeze on half-precision parameters. Weight decay is now unified into FP32 registers inside `_sr_update` alongside stochastic rounding, achieving 100% numerical parity with the Triton kernel. * **Fallback Foreach Indentation & Scope Alignment:** Corrected tensor iteration scope and indentation for `torch._foreach_add_` in non-SR fallback branches. * **Zero-Allocation GPU Clamping in `_apply_bound`:** Replaced dynamic GPU scalar allocations (`torch.tensor(1.5e-4, device=p.device)`) with native `torch.clamp_min`, eliminating transient CUDA malloc spikes per update step. ### Performance & Engine Refactoring * **Zero-Allocation GPU Bounding (`_apply_bound`):** Replaced dynamic GPU scalar tensor allocation (`torch.tensor(1.5e-4, device=p.device)`) with zero-overhead `torch.clamp_min(..., 1.5e-4)`. Eliminates transient memory allocation overhead and CUDA malloc spikes per update step. * **Unified Bound Execution Pipeline:** Added an immediate short-circuit guard (`if lr == 0.0: return update`) and streamlined the execution flow: intermediate updates are scaled in FP32 (`step_val = lr * u_float`), passed through radial/coordinate bounding, chained cleanly into the FP16 Cliff Governor, and returned in a single normalized step. --- ## [1.1.5] - 2026-09-15 ### Added - **Native Hardware Capability Detection (`_supports_native_bf16`):** Added automatic device capability detection via `torch.cuda.is_bf16_supported()` to identify Ampere SM80+ architectures and streamline native execution paths. - **Official Software Metadata & DOI:** Integrated official software DOI (`10.57967/hf/10365`) and canonical Hugging Face repository URL alongside paper archive links in module docstrings. ### Changed - **Dynamic Epsilon Clamping for FP16:** Dynamically floors effective epsilon to `max(eps, 1e-4)` when parameters or states are in `torch.float16`. Prevents bias-corrected scale factor `c2` from underflowing into IEEE-754 subnormal/zero limits on step 1. - Bumped package release version to `1.1.5`. ### Fixed - **FP16 Step-1 `NaN` from Epsilon Underflow:** Fixed catastrophic `0.0 / 0.0 = NaN` division occurring on initial training steps when parameters receive zero gradients (e.g., unselected vocabulary embeddings, padding tokens, or dropout paths). - **Second-Moment Variance Saturation (`+inf`):** Added safe upper-bound register clamping (`tl.minimum(v_new, 65000.0)` in Triton kernels and `.clamp_(0.0, 65000.0)` in PyTorch C++ foreach/single loops). Prevents unclipped initial pretraining gradient surges (`|g| > 256`) from overflowing the FP16 maximum dynamic range (65,504) into irreversible `+inf`. - **Zero-Sigma Division Guard:** Injected explicit `sigma + 1e-6` denominator guards across all Triton kernel passes and PyTorch update steps to ensure numerical safety during cold-start iterations. - Verified stable 100-step pretraining from scratch on `SmolLM-135M` using Wikimedia Wikipedia in pure FP16 on Nvidia Tesla T4 GPU with zero gradient clipping. --- ## [1.1.0] - 2026-09-11 ### Added - **Bitwise Stochastic Rounding (SR):** Integrated on-chip PRNG integer bit-manipulation hash in Triton kernels and vectorized PyTorch fallback for native FP16/BF16 training without master weights. - Unified Triton store routine introducing compile-time branching (`dtype_mode` / `use_sr: tl.constexpr`) with explicit safe casting for FP16 and BF16 pointers. - Vectorized fallback improvements utilizing C++ multi-tensor operations when stochastic rounding is disabled or tensors are FP32. - State buffer contiguity validation (`exp_avg_contig`, `exp_avg_sq_contig`) ensuring safe linear memory addressing before kernel dispatch. - Declarative configuration file (`config.json`) defining hyperparameters and optimizer presets. ### Changed - Refactored PRNG integer arithmetic to strictly stay within signed 32-bit integer limits (`< 2^31 - 1`), eliminating JIT compilation failures on newer Triton releases. - State buffers (`exp_avg`, `exp_avg_sq`) now initialize with strict contiguous memory format. ### Fixed - Fixed `ValueError: Scalar out of range for type int32` and Triton LLVM compile errors during unsigned 32-bit bitmask operations. - Resolved kernel fallback hangs and compilation overhead on CUDA devices. --- ## [1.0.2] - 2026-09-08 ### Added - Official paper reference and persistent DOI (`10.57967/hf/10270`) added to module header. - Official Apache-2.0 license notice and copyright headers. - Standardized package exports (`__all__`, `__version__ = "1.0.0"`, `__author__`, `__license__`). --- ## [1.0.1] - 2026-08-31 ### Fixed - **Memory Synchronization Fix (`423d171`):** Fixed a critical bug where non-contiguous parameter tensors (e.g., transposed linear layers or sliced weights) were updated on temporary contiguous copies without reflecting changes back to model weights. Added `p.copy_(p_contig)` after Triton execution. --- ## [1.0.0] - 2026-08-30 ### Added - Initial release of the LuminaV optimizer alongside XoneLM (`bbf30b7`). - **Zero Master-Weight Copies:** Direct parameter updates in native half-precision without allocating 4-byte FP32 master weights. - **Hyperbolic Tangent (`tanh`) Bounding Envelope:** Projects normalized momentum ratio into `(-1.0, 1.0)` to eliminate exploding gradient shocks. - **Traffic-Cop Directional Verification Gate:** Cautious update masking that zeroes out momentum steps opposing mini-batch gradient directions. - **Central Innovation Variance:** Noise dispersion tracking `(g_t - m_t)²` decoupling mini-batch turbulence from directional descent. - Dual-buffer mode (`buffer=2`, LuminaV-2B) and single-buffer extreme-poverty mode (`buffer=1`, LuminaV-1). - Custom fused OpenAI Triton kernels with automated fallback execution.