---
license: apache-2.0
language:
- en
tags:
- pytorch
- optimizer
- custom-optimizer
- triton
- memory-efficient
- stochastic-rounding
- low-vram
- bf16
- fp16
- cuda
- xpu
- llm
- transformer
- deep-learning
- research
- 3am-engineering
---
---
## Notice: Official Upstream Repository
This repository (`cloverx-id/LuminaV-Optimizer-Paper`) is the **official standalone and living development repository** for the LuminaV optimizer family.
While LuminaV was originally conceived and validated as the core engine for the **[XoneLM-1.0](https://huggingface.co/cloverx-id/XoneLM-1.0-Paper)** language model series, all subsequent optimizer upgrades, low-precision Triton kernels, PyTorch standards compliance, and bug fixes are actively maintained and released directly in this repository.
---
## What's New in v1.3.0 (Major Release)
The **v1.3.0** release introduces major architectural breakthroughs, featuring fused single-pass execution, multi-tier master weight flexibility, enterprise-grade fault recovery, and strict mathematical precision refinements:
- **Fused Single-Pass Kernel with Post-Step EMA (`fused_single_pass=True`):** Unifies parameter loading, gradient filtering, state transformation, cautious masking, and bounding into a single VRAM pass via running Exponential Moving Average (EMA) estimates of scale factors. Reduces global GPU memory bandwidth traffic by up to ~50% per optimizer step.
- **Hybrid Master Weight Precision Architecture (`master_weights`):** Added versatile precision modes:
- `"none"` (Default): Pure master-free training with on-chip bitwise stochastic rounding directly on native 16-bit weights.
- `"semi"` / `"half"`: 32-bit FP32 master weights with 16-bit BF16/FP16 optimizer moments, saving ~50% optimizer state VRAM while retaining full FP32 accumulation precision.
- `"full"` / `"fp32"`: Traditional FP32 master weights with FP32 optimizer moments for controlled baseline benchmarks.
- **In-Kernel NaN/Inf Defensive Sanitization:** Injected register-level guards (`is_finite_g`) across all Triton GPU kernels. Instantly suppresses exploding activations and poison gradients before they can corrupt running momentum or variance states.
- **Per-Parameter Persistent Scratch Allocations:** Replaced shared dynamic scratch buffers with dedicated 5-slot persistent state tensors (`state["scratch"]`), guaranteeing absolute memory address invariance for flawless **CUDA Graph Capture** (`torch.accelerator.Graph`) and TorchInductor AOT compilation.
- **Adaptive CTA Tiling & Hardware-Native Tanh:** Tensors with >= 2M elements automatically scale to `BLOCK_SIZE = 2048` and 8 warps, halving L2 atomic contention. Calls native PTX SFU instructions (`tanh.approx.f32`) on modern Triton compilers.
- **Analytically Exact Nesterov Bias Correction:** Upgraded first-moment bias correction to `1.0 - beta1^(step + 1)`, completely eliminating early-step residual momentum bias.
- **Fault-Tolerant Granular Recovery:** Per-parameter completion tracking ensures that if an unexpected fault occurs mid-step, completed parameters are preserved while uncompleted ones fall back safely without double-updating states.
*(For the complete patch notes and historical version logs, see [CHANGELOG.md](https://huggingface.co/cloverx-id/LuminaV-Optimizer-Paper/blob/main/CHANGELOG.md).)*
---
## Overview
**LuminaV** is a high-performance adaptive optimizer engineered specifically for deep learning workloads running directly in low precision (`FP16` / `BF16`) without maintaining redundant 4-byte FP32 master weights.
By combining **Centered Innovation Variance**, **Hyperbolic Tangent (tanh) Coordinate Bounding**, a **Directional Traffic-Cop Mask**, and **On-Chip Bitwise Stochastic Rounding**, LuminaV eliminates the standard 16-byte-per-parameter memory tax imposed by classic optimizers while avoiding weight freezing, gradient shocks, and numerical underflow.
---
## Key Features
1. **Zero Master-Weight Copies (Default):** Directly mutates parameter weights in native `FP16` or `BF16`, eliminating the 4-byte FP32 master weight allocation.
2. **On-Chip Bitwise Stochastic Rounding (SR):** Implements in-register bitcast hashing in Triton to provide mathematically unbiased stochastic rounding, preventing weight stagnation during fine-grained updates or learning rate decay.
3. **Hyperbolic tanh Bounding Envelope:** Maps normalized momentum through a `(-1.0, 1.0)` transfer function, guaranteeing coordinate updates cannot explode beyond the step learning rate.
4. **The Traffic-Cop Directional Gate:** Dynamically eliminates coordinate updates whenever historical momentum conflicts with the incoming mini-batch gradient direction (`u_t · g_t <= 0`).
5. **Centered Innovation Variance:** Tracks centered innovation dispersion `(g_t - m_t)^2` rather than uncentered raw second moments, suppressing variance inflation during confident descent.
6. **Automatic FP16 Cliff Governor:** Built-in asymptotic boundary governor that dampens steps near the IEEE-754 FP16 overflow limit (> 65,504), enabling stable pure FP16 training without external schedulers or clipping.
7. **Direction-Preserving Radial Bounding:** Smooth asymptotic parameter squashing (`tanh(r)/r`) that preserves 100.000% gradient angular fidelity while capping displacement.
8. **Multi-Tier Master Weight Support:** Configurable on the fly from 100% master-free up to hybrid (`"semi"`) or full FP32 (`"full"`) modes.
9. **Dual Execution Engine:** Fully accelerated custom OpenAI Triton kernels for CUDA and Intel XPU devices, paired with vectorized C++ `torch._foreach` multi-tensor fallbacks.
---
## Installation
### From PyPI (Recommended)
```bash
pip install luminav
```
For GPU acceleration via OpenAI Triton:
```bash
pip install luminav[triton]
```
### From Source (Editable Mode)
```bash
git clone https://huggingface.co/cloverx-id/LuminaV-Optimizer-Paper
cd LuminaV-Optimizer-Paper
pip install -e .
```
### Direct File Drop-in
Alternatively, you can copy `luminav.py` directly into your working project directory without packaging overhead:
```bash
wget https://huggingface.co/cloverx-id/LuminaV-Optimizer-Paper/raw/main/luminav.py
```
---
## Quickstart
### Standard Instantiation (Master-Free Mode)
```python
import torch
from luminav import LuminaV
# Instantiate your model in native low precision (e.g. BF16 or FP16)
model = YourModel().to(device="cuda", dtype=torch.bfloat16)
# Initialize LuminaV v1.3.0
optimizer = LuminaV(
model.parameters(),
lr=8e-4, # or 8e-5 / 8e-6 for fine-tuning
betas=(0.9, 0.999),
eps=1e-8,
weight_decay=0.08,
tau=0.8,
alpha_ss=0.5,
cautious=True,
cautious_clamp_min=0.5, # Exact power-of-two ceiling (2.00x)
buffer=2, # 2 = Dual-Buffer (Standard), 1 = Single-Buffer (Extreme Low VRAM)
stochastic_rounding=True,
bound=True, # Smooth asymptotic step bounding
bound_type="radial", # "radial" (preserves 100% angular direction) or "coordinate"
bound_ratio=0.03,
master_weights="none", # "none" (Master-Free), "semi" (Hybrid FP32 Master), or "full" (Full FP32)
fused_single_pass=False, # True enables ultra-fast 1-pass EMA kernel
execution="auto"
)
# Standard training step
optimizer.zero_grad(set_to_none=True)
loss = model(inputs, targets)
loss.backward()
optimizer.step()
```
### Ultra-Low Memory Training (Single-Buffer Mode)
To cut optimizer memory state by an additional 50% (maintaining only a single momentum buffer and collapsing variance to scalar RMS):
```python
optimizer = LuminaV(
model.parameters(),
lr=8e-4,
buffer=1, # LuminaV-1 Single-Buffer Mode
alpha_ss=0.5, # Softsign presquashing factor
master_weights="none"
)
```
### Loading from `config.json`
```python
import json
import torch
from luminav import LuminaV
with open("config.json", "r") as f:
config = json.load(f)
# Initialize with verified default configuration
optimizer = LuminaV(model.parameters(), **config["default_params"])
```
---
## Parameter Reference
| Parameter | Type | Default | Description |
| :--- | :--- | :--- | :--- |
| `params` | `iterable` | *Required* | Iterable of parameters to optimize or dicts defining parameter groups. |
| `lr` | `float` | `8e-4` | Learning rate (η). |
| `betas` | `Tuple[float, float]` | `(0.9, 0.999)` | Coefficients (β₁, β₂) for running momentum and centered innovation variance. First-moment bias correction uses `1.0 - beta1^(step + 1)`. |
| `eps` | `float` | `1e-8` | Numerical stability term (ε). Automatically floored to `1e-4` in FP16 to prevent subnormal underflow. |
| `weight_decay` | `float` | `8e-2` | Decoupled weight decay coefficient (λ). |
| `tau` | `float` | `0.8` | Analytical bias correction temperature parameter (τ). |
| `alpha_ss` | `float` | `0.5` | Softsign dampening factor (`α_ss`) used in single-buffer mode (`buffer=1`). |
| `cautious` | `bool` | `True` | If `True`, enables Traffic-Cop directional verification masking. |
| `cautious_clamp_min` | `float` | `0.5` | Safety floor density clamp (`γ_min`) enforcing a power-of-two maximum energy scaling ceiling (2.00x, 2¹) and preventing division by zero. |
| `buffer` | `int` | `2` | Buffer mode: `2` (Dual-buffer tracking `m_t` and `v_t`) or `1` (Single-buffer scalar RMS tracking). |
| `stochastic_rounding` | `bool` | `True` | Enables bitwise stochastic rounding on native FP16/BF16 weights. |
| `bound` | `bool` | `True` | If `True`, enables smooth asymptotic parameter bounding to prevent divergence in deep networks. |
| `bound_type` | `str` | `"radial"` | Asymptotic bounding formulation: `"radial"` (direction-preserving squashing using `tanh(r)/r`) or `"coordinate"` (elementwise squashing). |
| `bound_ratio` | `float` | `0.03` | Maximum allowed step displacement ratio relative to parameter norm or magnitude (`R = bound_ratio * max(‖p‖, 1.0)`). |
| `master_weights` | `Union[bool, str]`| `"none"` | Master weight precision mode: `"none"` / `False` (Master-Free), `"semi"` / `"half"` (FP32 master with 16-bit states), or `"full"` / `"fp32"` (Full FP32). |
| `fused_single_pass` | `bool` | `False` | If `True` (step > 1), fuses pass 1 and pass 2 into a single unified GPU kernel using running EMA scale estimates. |
| `ema_decay` | `float` | `0.8` | Running scale decay factor (`α_ema`) used when `fused_single_pass=True`. |
| `execution` | `str` | `"auto"` | Execution engine: `"auto"`, `"triton"`, `"foreach"`, or `"single"`. Automatically routes to `"foreach"` if deterministic mode is enabled. |
| `seed` | `int` | `1337` | Base seed for PRNG stochastic rounding and stateless golden-ratio hashing. |
---
## Operational Modes
### LuminaV-2 (Dual-Buffer Default: `buffer=2`)
Maintains first moment m_t and centered innovation variance v_t:
$$
m_t = \beta_1 m_{t-1} + (1 - \beta_1) g_t
$$
$$
v_t = \beta_2 v_{t-1} + (1 - \beta_2)(g_t - m_t)^2
$$
Updates are bounded through the hyperbolic tangent envelope:
$$
u_t = \tanh\left(\frac{\tilde{m}_t}{\sigma_t}\right), \quad \tilde{m}_t = \beta_1 m_t + (1 - \beta_1) g_t
$$
### LuminaV-1 (Single-Buffer Extreme-Poverty Mode: `buffer=1`)
Collapses variance tracking into a scalar Root-Mean-Square (RMS) across the entire tensor, saving 50% optimizer state memory by maintaining only a single state buffer (m_t):
$$
\text{RMS}(\tilde{m}_t) = \sqrt{\frac{1}{N} \sum_{i=1}^N \tilde{m}_{t,i}^2 + \epsilon}
$$
$$
u_t = \tanh\left(\frac{z}{1 + \alpha_{ss}|z|}\right), \quad z = \frac{\tilde{m}_t}{\tau \cdot \text{RMS}(\tilde{m}_t) + \epsilon(1 - \beta_1^{t+1})\tau}
$$
---
## Empirical Benchmarks (Qwen3.5-4B-Base)
LuminaV v1.3.0 was rigorously benchmarked on a 4.0-billion parameter Large Language Model (**`Qwen/Qwen3.5-4B-Base`**) initialized from architectural config (`AutoModelForCausalLM.from_config`) under **full pretraining from scratch** conditions (all 4B parameters actively optimized, gradient checkpointing disabled) on an NVIDIA 80GB GPU.
> **Architectural Disambiguation (1P / 2P vs Buffer Count):**
> The labels **1P** and **2P** refer strictly to **GPU Kernel Execution Passes** (`1P` = Fused Single-Pass Kernel via EMA, `2P` = Standard Two-Pass Kernel Reduction). **All evaluated configurations below strictly operate under the Dual-Buffer architecture (`buffer=2`)**, maintaining both the first moment (`m_t`) and centered innovation variance (`v_t`). Single-buffer mode (`buffer=1`) was not benchmarked in this suite.
### Key Benchmark Highlights
| Optimizer Mode | Kernel Passes | Buffer Mode | Peak VRAM | Static State VRAM | Pure Opt Latency | Final Loss (250) | Weight Cosine Sim (vs 2P) |
| :--- | :---: | :---: | :---: | :---: | :---: | :---: | :---: |
| **Master-Free (1P)** | 1-Pass (EMA) | `buffer=2` (Dual) | **34.32 GB** *(-47.6%)* | **24.65 GB** *(-55.8%)* | **59.9 ms** *(1.90x faster)* | **8.8260** | 0.9445 |
| **Master-Free (2P)** | 2-Pass (Exact) | `buffer=2` (Dual) | **34.32 GB** *(-47.6%)* | **24.65 GB** *(-55.8%)* | 68.4 ms | 9.1075 | 1.0000 (Exact) |
| **Semi (1P)** | 1-Pass (EMA) | `buffer=2` (Dual) | 49.99 GB *(-23.6%)* | 40.32 GB *(-27.7%)* | **70.2 ms** *(1.62x faster)* | **8.8192** *(Best)* | 0.9512 |
| **Full (2P - Baseline)** | 2-Pass (Exact) | `buffer=2` (Dual) | 65.47 GB | 55.80 GB | 113.8 ms | 8.9579 | 1.0000 (Exact) |
- **47.6% Peak VRAM Reduction (31.15 GB Saved):** Master-Free mode cuts peak memory from 65.47 GB down to **34.32 GB**, allowing full-throughput training of a 4B parameter model on 40GB/48GB GPUs without activation checkpointing.
- **1.9x Pure Optimizer Speedup via Fused Single-Pass:** Slashing redundant VRAM memory passes cuts optimizer latency from 113.8 ms (Full 2P) down to **59.9 ms (Master-Free 1P)** while retaining full dual-buffer state tracking (`buffer=2`).
- **94.0% Mean Cosine Representation Fidelity:** Single-pass running EMA updates preserve 94% angular cosine similarity with exact two-pass trajectories while achieving equal or superior convergence.
[Read the Full Empirical Benchmark & Reproducibility Report (BENCHMARKS.md)](https://huggingface.co/cloverx-id/LuminaV-Optimizer-Paper/blob/main/BENCHMARKS.md)
[Open Interactive Reproduction Notebook in Google Colab](https://colab.research.google.com/drive/1vgNIo8O2MpH9fU-kreg08vPwaiucsXik?usp=sharing)
---
## Citation
If you utilize LuminaV in your research or applications, please cite both the foundational paper and this software implementation:
```bibtex
# 1. To cite the official research paper & theoretical mechanics
@misc{luminamoon2026luminav_paper,
author = {{Silver Moon (cloverxion)}},
organization = {Lumina Moon (cloverx-id)},
title = {{LuminaV: We Were Too Broke for AdamW So We Trapped Gradients in a Hyperbolic Straitjacket and Hired a Traffic Cop to Slap Them}},
year = {2026},
publisher = {Hugging Face},
doi = {10.57967/hf/10270},
url = {https://huggingface.co/cloverx-id/XoneLM-1.0-Paper}
}
# 2. To cite this software implementation & standalone codebase
@software{luminamoon2026luminav_code,
author = {{Silver Moon (cloverxion) and Lumina Moon Contributors}},
organization = {Lumina Moon (cloverx-id)},
title = {{LuminaV Optimizer: Official PyTorch Implementation}},
year = {2026},
publisher = {Hugging Face / PyPI},
version = {1.3.0},
doi = {10.57967/hf/10365},
url = {https://huggingface.co/cloverx-id/LuminaV-Optimizer-Paper}
}
```
---
## License
All Resources are under Apache License 2.0. See [LICENSE](https://huggingface.co/cloverx-id/LuminaV-Optimizer-Paper/blob/main/LICENSE) for full terms.