--- license: apache-2.0 language: - en tags: - pytorch - optimizer - custom-optimizer - triton - memory-efficient - stochastic-rounding - low-vram - bf16 - fp16 - cuda - xpu - llm - transformer - deep-learning - research - 3am-engineering ---

LuminaV Optimizer

We Were Too Broke for AdamW So We Trapped Gradients in a Hyperbolic Straitjacket and Hired a Traffic Cop to Slap Them

Official Upstream & Standalone Codebase | Current Version: v1.3.0 | Check `Files and Versions`

PyPI Version Changelog Benchmark Software DOI Paper DOI Config JSON License

Official Research Paper

LuminaV Paper Preview

LuminaV Optimizer Theory & Mechanics
Read LuminaV.pdf (Local Mirror)  |  Primary Paper Archive

Click the preview above to read or download the official paper PDF.

--- ## Notice: Official Upstream Repository This repository (`cloverx-id/LuminaV-Optimizer-Paper`) is the **official standalone and living development repository** for the LuminaV optimizer family. While LuminaV was originally conceived and validated as the core engine for the **[XoneLM-1.0](https://huggingface.co/cloverx-id/XoneLM-1.0-Paper)** language model series, all subsequent optimizer upgrades, low-precision Triton kernels, PyTorch standards compliance, and bug fixes are actively maintained and released directly in this repository. --- ## What's New in v1.3.0 (Major Release) The **v1.3.0** release introduces major architectural breakthroughs, featuring fused single-pass execution, multi-tier master weight flexibility, enterprise-grade fault recovery, and strict mathematical precision refinements: - **Fused Single-Pass Kernel with Post-Step EMA (`fused_single_pass=True`):** Unifies parameter loading, gradient filtering, state transformation, cautious masking, and bounding into a single VRAM pass via running Exponential Moving Average (EMA) estimates of scale factors. Reduces global GPU memory bandwidth traffic by up to ~50% per optimizer step. - **Hybrid Master Weight Precision Architecture (`master_weights`):** Added versatile precision modes: - `"none"` (Default): Pure master-free training with on-chip bitwise stochastic rounding directly on native 16-bit weights. - `"semi"` / `"half"`: 32-bit FP32 master weights with 16-bit BF16/FP16 optimizer moments, saving ~50% optimizer state VRAM while retaining full FP32 accumulation precision. - `"full"` / `"fp32"`: Traditional FP32 master weights with FP32 optimizer moments for controlled baseline benchmarks. - **In-Kernel NaN/Inf Defensive Sanitization:** Injected register-level guards (`is_finite_g`) across all Triton GPU kernels. Instantly suppresses exploding activations and poison gradients before they can corrupt running momentum or variance states. - **Per-Parameter Persistent Scratch Allocations:** Replaced shared dynamic scratch buffers with dedicated 5-slot persistent state tensors (`state["scratch"]`), guaranteeing absolute memory address invariance for flawless **CUDA Graph Capture** (`torch.accelerator.Graph`) and TorchInductor AOT compilation. - **Adaptive CTA Tiling & Hardware-Native Tanh:** Tensors with >= 2M elements automatically scale to `BLOCK_SIZE = 2048` and 8 warps, halving L2 atomic contention. Calls native PTX SFU instructions (`tanh.approx.f32`) on modern Triton compilers. - **Analytically Exact Nesterov Bias Correction:** Upgraded first-moment bias correction to `1.0 - beta1^(step + 1)`, completely eliminating early-step residual momentum bias. - **Fault-Tolerant Granular Recovery:** Per-parameter completion tracking ensures that if an unexpected fault occurs mid-step, completed parameters are preserved while uncompleted ones fall back safely without double-updating states. *(For the complete patch notes and historical version logs, see [CHANGELOG.md](https://huggingface.co/cloverx-id/LuminaV-Optimizer-Paper/blob/main/CHANGELOG.md).)* --- ## Overview **LuminaV** is a high-performance adaptive optimizer engineered specifically for deep learning workloads running directly in low precision (`FP16` / `BF16`) without maintaining redundant 4-byte FP32 master weights. By combining **Centered Innovation Variance**, **Hyperbolic Tangent (tanh) Coordinate Bounding**, a **Directional Traffic-Cop Mask**, and **On-Chip Bitwise Stochastic Rounding**, LuminaV eliminates the standard 16-byte-per-parameter memory tax imposed by classic optimizers while avoiding weight freezing, gradient shocks, and numerical underflow. --- ## Key Features 1. **Zero Master-Weight Copies (Default):** Directly mutates parameter weights in native `FP16` or `BF16`, eliminating the 4-byte FP32 master weight allocation. 2. **On-Chip Bitwise Stochastic Rounding (SR):** Implements in-register bitcast hashing in Triton to provide mathematically unbiased stochastic rounding, preventing weight stagnation during fine-grained updates or learning rate decay. 3. **Hyperbolic tanh Bounding Envelope:** Maps normalized momentum through a `(-1.0, 1.0)` transfer function, guaranteeing coordinate updates cannot explode beyond the step learning rate. 4. **The Traffic-Cop Directional Gate:** Dynamically eliminates coordinate updates whenever historical momentum conflicts with the incoming mini-batch gradient direction (`u_t · g_t <= 0`). 5. **Centered Innovation Variance:** Tracks centered innovation dispersion `(g_t - m_t)^2` rather than uncentered raw second moments, suppressing variance inflation during confident descent. 6. **Automatic FP16 Cliff Governor:** Built-in asymptotic boundary governor that dampens steps near the IEEE-754 FP16 overflow limit (> 65,504), enabling stable pure FP16 training without external schedulers or clipping. 7. **Direction-Preserving Radial Bounding:** Smooth asymptotic parameter squashing (`tanh(r)/r`) that preserves 100.000% gradient angular fidelity while capping displacement. 8. **Multi-Tier Master Weight Support:** Configurable on the fly from 100% master-free up to hybrid (`"semi"`) or full FP32 (`"full"`) modes. 9. **Dual Execution Engine:** Fully accelerated custom OpenAI Triton kernels for CUDA and Intel XPU devices, paired with vectorized C++ `torch._foreach` multi-tensor fallbacks. --- ## Installation ### From PyPI (Recommended) ```bash pip install luminav ``` For GPU acceleration via OpenAI Triton: ```bash pip install luminav[triton] ``` ### From Source (Editable Mode) ```bash git clone https://huggingface.co/cloverx-id/LuminaV-Optimizer-Paper cd LuminaV-Optimizer-Paper pip install -e . ``` ### Direct File Drop-in Alternatively, you can copy `luminav.py` directly into your working project directory without packaging overhead: ```bash wget https://huggingface.co/cloverx-id/LuminaV-Optimizer-Paper/raw/main/luminav.py ``` --- ## Quickstart ### Standard Instantiation (Master-Free Mode) ```python import torch from luminav import LuminaV # Instantiate your model in native low precision (e.g. BF16 or FP16) model = YourModel().to(device="cuda", dtype=torch.bfloat16) # Initialize LuminaV v1.3.0 optimizer = LuminaV( model.parameters(), lr=8e-4, # or 8e-5 / 8e-6 for fine-tuning betas=(0.9, 0.999), eps=1e-8, weight_decay=0.08, tau=0.8, alpha_ss=0.5, cautious=True, cautious_clamp_min=0.5, # Exact power-of-two ceiling (2.00x) buffer=2, # 2 = Dual-Buffer (Standard), 1 = Single-Buffer (Extreme Low VRAM) stochastic_rounding=True, bound=True, # Smooth asymptotic step bounding bound_type="radial", # "radial" (preserves 100% angular direction) or "coordinate" bound_ratio=0.03, master_weights="none", # "none" (Master-Free), "semi" (Hybrid FP32 Master), or "full" (Full FP32) fused_single_pass=False, # True enables ultra-fast 1-pass EMA kernel execution="auto" ) # Standard training step optimizer.zero_grad(set_to_none=True) loss = model(inputs, targets) loss.backward() optimizer.step() ``` ### Ultra-Low Memory Training (Single-Buffer Mode) To cut optimizer memory state by an additional 50% (maintaining only a single momentum buffer and collapsing variance to scalar RMS): ```python optimizer = LuminaV( model.parameters(), lr=8e-4, buffer=1, # LuminaV-1 Single-Buffer Mode alpha_ss=0.5, # Softsign presquashing factor master_weights="none" ) ``` ### Loading from `config.json` ```python import json import torch from luminav import LuminaV with open("config.json", "r") as f: config = json.load(f) # Initialize with verified default configuration optimizer = LuminaV(model.parameters(), **config["default_params"]) ``` --- ## Parameter Reference | Parameter | Type | Default | Description | | :--- | :--- | :--- | :--- | | `params` | `iterable` | *Required* | Iterable of parameters to optimize or dicts defining parameter groups. | | `lr` | `float` | `8e-4` | Learning rate (η). | | `betas` | `Tuple[float, float]` | `(0.9, 0.999)` | Coefficients (β₁, β₂) for running momentum and centered innovation variance. First-moment bias correction uses `1.0 - beta1^(step + 1)`. | | `eps` | `float` | `1e-8` | Numerical stability term (ε). Automatically floored to `1e-4` in FP16 to prevent subnormal underflow. | | `weight_decay` | `float` | `8e-2` | Decoupled weight decay coefficient (λ). | | `tau` | `float` | `0.8` | Analytical bias correction temperature parameter (τ). | | `alpha_ss` | `float` | `0.5` | Softsign dampening factor (`α_ss`) used in single-buffer mode (`buffer=1`). | | `cautious` | `bool` | `True` | If `True`, enables Traffic-Cop directional verification masking. | | `cautious_clamp_min` | `float` | `0.5` | Safety floor density clamp (`γ_min`) enforcing a power-of-two maximum energy scaling ceiling (2.00x, 2¹) and preventing division by zero. | | `buffer` | `int` | `2` | Buffer mode: `2` (Dual-buffer tracking `m_t` and `v_t`) or `1` (Single-buffer scalar RMS tracking). | | `stochastic_rounding` | `bool` | `True` | Enables bitwise stochastic rounding on native FP16/BF16 weights. | | `bound` | `bool` | `True` | If `True`, enables smooth asymptotic parameter bounding to prevent divergence in deep networks. | | `bound_type` | `str` | `"radial"` | Asymptotic bounding formulation: `"radial"` (direction-preserving squashing using `tanh(r)/r`) or `"coordinate"` (elementwise squashing). | | `bound_ratio` | `float` | `0.03` | Maximum allowed step displacement ratio relative to parameter norm or magnitude (`R = bound_ratio * max(‖p‖, 1.0)`). | | `master_weights` | `Union[bool, str]`| `"none"` | Master weight precision mode: `"none"` / `False` (Master-Free), `"semi"` / `"half"` (FP32 master with 16-bit states), or `"full"` / `"fp32"` (Full FP32). | | `fused_single_pass` | `bool` | `False` | If `True` (step > 1), fuses pass 1 and pass 2 into a single unified GPU kernel using running EMA scale estimates. | | `ema_decay` | `float` | `0.8` | Running scale decay factor (`α_ema`) used when `fused_single_pass=True`. | | `execution` | `str` | `"auto"` | Execution engine: `"auto"`, `"triton"`, `"foreach"`, or `"single"`. Automatically routes to `"foreach"` if deterministic mode is enabled. | | `seed` | `int` | `1337` | Base seed for PRNG stochastic rounding and stateless golden-ratio hashing. | --- ## Operational Modes ### LuminaV-2 (Dual-Buffer Default: `buffer=2`) Maintains first moment m_t and centered innovation variance v_t: $$ m_t = \beta_1 m_{t-1} + (1 - \beta_1) g_t $$ $$ v_t = \beta_2 v_{t-1} + (1 - \beta_2)(g_t - m_t)^2 $$ Updates are bounded through the hyperbolic tangent envelope: $$ u_t = \tanh\left(\frac{\tilde{m}_t}{\sigma_t}\right), \quad \tilde{m}_t = \beta_1 m_t + (1 - \beta_1) g_t $$ ### LuminaV-1 (Single-Buffer Extreme-Poverty Mode: `buffer=1`) Collapses variance tracking into a scalar Root-Mean-Square (RMS) across the entire tensor, saving 50% optimizer state memory by maintaining only a single state buffer (m_t): $$ \text{RMS}(\tilde{m}_t) = \sqrt{\frac{1}{N} \sum_{i=1}^N \tilde{m}_{t,i}^2 + \epsilon} $$ $$ u_t = \tanh\left(\frac{z}{1 + \alpha_{ss}|z|}\right), \quad z = \frac{\tilde{m}_t}{\tau \cdot \text{RMS}(\tilde{m}_t) + \epsilon(1 - \beta_1^{t+1})\tau} $$ --- ## Empirical Benchmarks (Qwen3.5-4B-Base) LuminaV v1.3.0 was rigorously benchmarked on a 4.0-billion parameter Large Language Model (**`Qwen/Qwen3.5-4B-Base`**) initialized from architectural config (`AutoModelForCausalLM.from_config`) under **full pretraining from scratch** conditions (all 4B parameters actively optimized, gradient checkpointing disabled) on an NVIDIA 80GB GPU.
LuminaV Qwen3.5-4B Benchmark Charts
> **Architectural Disambiguation (1P / 2P vs Buffer Count):** > The labels **1P** and **2P** refer strictly to **GPU Kernel Execution Passes** (`1P` = Fused Single-Pass Kernel via EMA, `2P` = Standard Two-Pass Kernel Reduction). **All evaluated configurations below strictly operate under the Dual-Buffer architecture (`buffer=2`)**, maintaining both the first moment (`m_t`) and centered innovation variance (`v_t`). Single-buffer mode (`buffer=1`) was not benchmarked in this suite. ### Key Benchmark Highlights | Optimizer Mode | Kernel Passes | Buffer Mode | Peak VRAM | Static State VRAM | Pure Opt Latency | Final Loss (250) | Weight Cosine Sim (vs 2P) | | :--- | :---: | :---: | :---: | :---: | :---: | :---: | :---: | | **Master-Free (1P)** | 1-Pass (EMA) | `buffer=2` (Dual) | **34.32 GB** *(-47.6%)* | **24.65 GB** *(-55.8%)* | **59.9 ms** *(1.90x faster)* | **8.8260** | 0.9445 | | **Master-Free (2P)** | 2-Pass (Exact) | `buffer=2` (Dual) | **34.32 GB** *(-47.6%)* | **24.65 GB** *(-55.8%)* | 68.4 ms | 9.1075 | 1.0000 (Exact) | | **Semi (1P)** | 1-Pass (EMA) | `buffer=2` (Dual) | 49.99 GB *(-23.6%)* | 40.32 GB *(-27.7%)* | **70.2 ms** *(1.62x faster)* | **8.8192** *(Best)* | 0.9512 | | **Full (2P - Baseline)** | 2-Pass (Exact) | `buffer=2` (Dual) | 65.47 GB | 55.80 GB | 113.8 ms | 8.9579 | 1.0000 (Exact) | - **47.6% Peak VRAM Reduction (31.15 GB Saved):** Master-Free mode cuts peak memory from 65.47 GB down to **34.32 GB**, allowing full-throughput training of a 4B parameter model on 40GB/48GB GPUs without activation checkpointing. - **1.9x Pure Optimizer Speedup via Fused Single-Pass:** Slashing redundant VRAM memory passes cuts optimizer latency from 113.8 ms (Full 2P) down to **59.9 ms (Master-Free 1P)** while retaining full dual-buffer state tracking (`buffer=2`). - **94.0% Mean Cosine Representation Fidelity:** Single-pass running EMA updates preserve 94% angular cosine similarity with exact two-pass trajectories while achieving equal or superior convergence. [Read the Full Empirical Benchmark & Reproducibility Report (BENCHMARKS.md)](https://huggingface.co/cloverx-id/LuminaV-Optimizer-Paper/blob/main/BENCHMARKS.md) [Open Interactive Reproduction Notebook in Google Colab](https://colab.research.google.com/drive/1vgNIo8O2MpH9fU-kreg08vPwaiucsXik?usp=sharing) --- ## Citation If you utilize LuminaV in your research or applications, please cite both the foundational paper and this software implementation: ```bibtex # 1. To cite the official research paper & theoretical mechanics @misc{luminamoon2026luminav_paper, author = {{Silver Moon (cloverxion)}}, organization = {Lumina Moon (cloverx-id)}, title = {{LuminaV: We Were Too Broke for AdamW So We Trapped Gradients in a Hyperbolic Straitjacket and Hired a Traffic Cop to Slap Them}}, year = {2026}, publisher = {Hugging Face}, doi = {10.57967/hf/10270}, url = {https://huggingface.co/cloverx-id/XoneLM-1.0-Paper} } # 2. To cite this software implementation & standalone codebase @software{luminamoon2026luminav_code, author = {{Silver Moon (cloverxion) and Lumina Moon Contributors}}, organization = {Lumina Moon (cloverx-id)}, title = {{LuminaV Optimizer: Official PyTorch Implementation}}, year = {2026}, publisher = {Hugging Face / PyPI}, version = {1.3.0}, doi = {10.57967/hf/10365}, url = {https://huggingface.co/cloverx-id/LuminaV-Optimizer-Paper} } ``` --- ## License All Resources are under Apache License 2.0. See [LICENSE](https://huggingface.co/cloverx-id/LuminaV-Optimizer-Paper/blob/main/LICENSE) for full terms.