---
license: apache-2.0
language:
- en
tags:
- pytorch
- optimizer
- custom-optimizer
- triton
- memory-efficient
- stochastic-rounding
- low-vram
- bf16
- fp16
- cuda
- xpu
- llm
- transformer
- deep-learning
- research
- 3am-engineering
---
LuminaV Optimizer
We Were Too Broke for AdamW So We Trapped Gradients in a Hyperbolic Straitjacket and Hired a Traffic Cop to Slap Them
Official Upstream & Standalone Codebase | Current Version: v1.3.2 | Check `Files and Versions`
Click the preview above to read or download the official paper PDF.
Honestly, this is actually the PDF for the first version of LuminaV. So, the major updates we've made since then aren't in here. I might make a new one later. So, for now.. please look => CHANGELOG.md
---
## Notice: Official Upstream Repository
This repository (`cloverx-id/LuminaV-Optimizer-Paper`) is the **official standalone and living development repository** for the LuminaV optimizer family.
While LuminaV was originally conceived and validated as the core engine for the **[XoneLM-1.0](https://huggingface.co/cloverx-id/XoneLM-1.0-Paper)** language model series, all subsequent optimizer upgrades, low-precision Triton kernels, PyTorch standards compliance, and bug fixes are actively maintained and released directly in this repository.
---
## What's New in v1.3.0 (Major Release)
The **v1.3.0** release introduces major architectural breakthroughs, featuring fused single-pass execution, multi-tier master weight flexibility, enterprise-grade fault recovery, and strict mathematical precision refinements:
- **Fused Single-Pass Kernel with Post-Step EMA (`fused_single_pass=True`):** Unifies parameter loading, gradient filtering, state transformation, cautious masking, and bounding into a single VRAM pass via running Exponential Moving Average (EMA) estimates of scale factors. Reduces global GPU memory bandwidth traffic by up to ~50% per optimizer step.
- **Hybrid Master Weight Precision Architecture (`master_weights`):** Added versatile precision modes:
- `"none"` (Default): Pure master-free training with on-chip bitwise stochastic rounding directly on native 16-bit weights.
- `"semi"` / `"half"`: 32-bit FP32 master weights with 16-bit BF16/FP16 optimizer moments, saving ~50% optimizer state VRAM while retaining full FP32 accumulation precision.
- `"full"` / `"fp32"`: Traditional FP32 master weights with FP32 optimizer moments for controlled baseline benchmarks.
- **In-Kernel NaN/Inf Defensive Sanitization:** Injected register-level guards (`is_finite_g`) across all Triton GPU kernels. Instantly suppresses exploding activations and poison gradients before they can corrupt running momentum or variance states.
- **Per-Parameter Persistent Scratch Allocations:** Replaced shared dynamic scratch buffers with dedicated 5-slot persistent state tensors (`state["scratch"]`), guaranteeing absolute memory address invariance for flawless **CUDA Graph Capture** (`torch.accelerator.Graph`) and TorchInductor AOT compilation.
- **Adaptive CTA Tiling & Hardware-Native Tanh:** Tensors with >= 2M elements automatically scale to `BLOCK_SIZE = 2048` and 8 warps, halving L2 atomic contention. Calls native PTX SFU instructions (`tanh.approx.f32`) on modern Triton compilers.
- **Analytically Exact Nesterov Bias Correction:** Upgraded first-moment bias correction to `1.0 - beta1^(step + 1)`, completely eliminating early-step residual momentum bias.
- **Fault-Tolerant Granular Recovery:** Per-parameter completion tracking ensures that if an unexpected fault occurs mid-step, completed parameters are preserved while uncompleted ones fall back safely without double-updating states.
*(For the complete patch notes and historical version logs, see [CHANGELOG.md](https://huggingface.co/cloverx-id/LuminaV-Optimizer-Paper/blob/main/CHANGELOG.md).)*
---
## Overview
**LuminaV** is a high-performance adaptive optimizer engineered specifically for deep learning workloads running directly in low precision (`FP16` / `BF16`) without maintaining redundant 4-byte FP32 master weights.
By combining **Centered Innovation Variance**, **Hyperbolic Tangent (tanh) Coordinate Bounding**, a **Directional Traffic-Cop Mask**, and **On-Chip Bitwise Stochastic Rounding**, LuminaV eliminates the standard 16-byte-per-parameter memory tax imposed by classic optimizers while avoiding weight freezing, gradient shocks, and numerical underflow.
---
## Key Features
1. **Zero Master-Weight Copies (Default):** Directly mutates parameter weights in native `FP16` or `BF16`, eliminating the 4-byte FP32 master weight allocation.
2. **On-Chip Bitwise Stochastic Rounding (SR):** Implements in-register bitcast hashing in Triton to provide mathematically unbiased stochastic rounding, preventing weight stagnation during fine-grained updates or learning rate decay.
3. **Hyperbolic tanh Bounding Envelope:** Maps normalized momentum through a `(-1.0, 1.0)` transfer function, guaranteeing coordinate updates cannot explode beyond the step learning rate.
4. **The Traffic-Cop Directional Gate:** Dynamically eliminates coordinate updates whenever historical momentum conflicts with the incoming mini-batch gradient direction (`u_t · g_t <= 0`).
5. **Centered Innovation Variance:** Tracks centered innovation dispersion `(g_t - m_t)^2` rather than uncentered raw second moments, suppressing variance inflation during confident descent.
6. **Automatic FP16 Cliff Governor:** Built-in asymptotic boundary governor that dampens steps near the IEEE-754 FP16 overflow limit (> 65,504), enabling stable pure FP16 training without external schedulers or clipping.
7. **Direction-Preserving Radial Bounding:** Smooth asymptotic parameter squashing (`tanh(r)/r`) that preserves 100.000% gradient angular fidelity while capping displacement.
8. **Multi-Tier Master Weight Support:** Configurable on the fly from 100% master-free up to hybrid (`"semi"`) or full FP32 (`"full"`) modes.
9. **Dual Execution Engine:** Fully accelerated custom OpenAI Triton kernels for CUDA and Intel XPU devices, paired with vectorized C++ `torch._foreach` multi-tensor fallbacks.
---
## Installation
### From PyPI (Recommended)
```bash
pip install luminav
```
For GPU acceleration via OpenAI Triton:
```bash
pip install luminav[triton]
```
### From Source (Editable Mode)
```bash
git clone https://huggingface.co/cloverx-id/LuminaV-Optimizer-Paper
cd LuminaV-Optimizer-Paper
pip install -e .
```
### Direct File Drop-in
Alternatively, you can copy `luminav.py` directly into your working project directory without packaging overhead:
```bash
wget https://huggingface.co/cloverx-id/LuminaV-Optimizer-Paper/raw/main/luminav.py
```
---
## Quickstart
### Standard Instantiation (Master-Free Mode)
```python
import torch
from luminav import LuminaV
# Instantiate your model in native low precision (e.g. BF16 or FP16)
model = YourModel().to(device="cuda", dtype=torch.bfloat16)
# Initialize LuminaV v1.3.0
optimizer = LuminaV(
model.parameters(),
lr=8e-4, # or 8e-5 / 8e-6 for fine-tuning
betas=(0.9, 0.999),
eps=1e-8,
weight_decay=0.08,
tau=0.8,
alpha_ss=0.5,
cautious=True,
cautious_clamp_min=0.5, # Exact power-of-two ceiling (2.00x)
buffer=2, # 2 = Dual-Buffer (Standard), 1 = Single-Buffer (Extreme Low VRAM)
stochastic_rounding=True,
bound=True, # Smooth asymptotic step bounding
bound_type="radial", # "radial" (preserves 100% angular direction) or "coordinate"
bound_ratio=0.03,
master_weights="none", # "none" (Master-Free), "semi" (Hybrid FP32 Master), or "full" (Full FP32)
fused_single_pass=False, # True enables ultra-fast 1-pass EMA kernel
execution="auto"
)
# Standard training step
optimizer.zero_grad(set_to_none=True)
loss = model(inputs, targets)
loss.backward()
optimizer.step()
```
### Ultra-Low Memory Training (Single-Buffer Mode)
To cut optimizer memory state by an additional 50% (maintaining only a single momentum buffer and collapsing variance to scalar RMS):
```python
optimizer = LuminaV(
model.parameters(),
lr=8e-4,
buffer=1, # LuminaV-1 Single-Buffer Mode
alpha_ss=0.5, # Softsign presquashing factor
master_weights="none"
)
```
### Loading from `config.json`
```python
import json
import torch
from luminav import LuminaV
with open("config.json", "r") as f:
config = json.load(f)
# Initialize with verified default configuration
optimizer = LuminaV(model.parameters(), **config["default_params"])
```
---
## Parameter Reference
| Parameter | Type | Default | Description |
| :--- | :--- | :--- | :--- |
| `params` | `iterable` | *Required* | Iterable of parameters to optimize or dicts defining parameter groups. |
| `lr` | `float` | `8e-4` | Learning rate (η). |
| `betas` | `Tuple[float, float]` | `(0.9, 0.999)` | Coefficients (β₁, β₂) for running momentum and centered innovation variance. First-moment bias correction uses `1.0 - beta1^(step + 1)`. |
| `eps` | `float` | `1e-8` | Numerical stability term (ε). Automatically floored to `1e-4` in FP16 to prevent subnormal underflow. |
| `weight_decay` | `float` | `8e-2` | Decoupled weight decay coefficient (λ). |
| `tau` | `float` | `0.8` | Analytical bias correction temperature parameter (τ). |
| `alpha_ss` | `float` | `0.5` | Softsign dampening factor (`α_ss`) used in single-buffer mode (`buffer=1`). |
| `cautious` | `bool` | `True` | If `True`, enables Traffic-Cop directional verification masking. |
| `cautious_clamp_min` | `float` | `0.5` | Safety floor density clamp (`γ_min`) enforcing a power-of-two maximum energy scaling ceiling (2.00x, 2¹) and preventing division by zero. |
| `buffer` | `int` | `2` | Buffer mode: `2` (Dual-buffer tracking `m_t` and `v_t`) or `1` (Single-buffer scalar RMS tracking). |
| `stochastic_rounding` | `bool` | `True` | Enables bitwise stochastic rounding on native FP16/BF16 weights. |
| `bound` | `bool` | `True` | If `True`, enables smooth asymptotic parameter bounding to prevent divergence in deep networks. |
| `bound_type` | `str` | `"radial"` | Asymptotic bounding formulation: `"radial"` (direction-preserving squashing using `tanh(r)/r`) or `"coordinate"` (elementwise squashing). |
| `bound_ratio` | `float` | `0.03` | Maximum allowed step displacement ratio relative to parameter norm or magnitude (`R = bound_ratio * max(‖p‖, 1.0)`). |
| `master_weights` | `Union[bool, str]`| `"none"` | Master weight precision mode: `"none"` / `False` (Master-Free), `"semi"` / `"half"` (FP32 master with 16-bit states), or `"full"` / `"fp32"` (Full FP32). |
| `fused_single_pass` | `bool` | `False` | If `True` (step > 1), fuses pass 1 and pass 2 into a single unified GPU kernel using running EMA scale estimates. |
| `ema_decay` | `float` | `0.8` | Running scale decay factor (`α_ema`) used when `fused_single_pass=True`. |
| `execution` | `str` | `"auto"` | Execution engine: `"auto"`, `"triton"`, `"foreach"`, or `"single"`. Automatically routes to `"foreach"` if deterministic mode is enabled. |
| `seed` | `int` | `1337` | Base seed for PRNG stochastic rounding and stateless golden-ratio hashing. |
---
## Operational Modes
### LuminaV-2 (Dual-Buffer Default: `buffer=2`)
Maintains first moment m_t and centered innovation variance v_t:
$$
m_t = \beta_1 m_{t-1} + (1 - \beta_1) g_t
$$
$$
v_t = \beta_2 v_{t-1} + (1 - \beta_2)(g_t - m_t)^2
$$
Updates are bounded through the hyperbolic tangent envelope:
$$
u_t = \tanh\left(\frac{\tilde{m}_t}{\sigma_t}\right), \quad \tilde{m}_t = \beta_1 m_t + (1 - \beta_1) g_t
$$
### LuminaV-1 (Single-Buffer Extreme-Poverty Mode: `buffer=1`)
Collapses variance tracking into a scalar Root-Mean-Square (RMS) across the entire tensor, saving 50% optimizer state memory by maintaining only a single state buffer (m_t):
$$
\text{RMS}(\tilde{m}_t) = \sqrt{\frac{1}{N} \sum_{i=1}^N \tilde{m}_{t,i}^2 + \epsilon}
$$
$$
u_t = \tanh\left(\frac{z}{1 + \alpha_{ss}|z|}\right), \quad z = \frac{\tilde{m}_t}{\tau \cdot \text{RMS}(\tilde{m}_t) + \epsilon(1 - \beta_1^{t+1})\tau}
$$
---
## Empirical Benchmarks (Qwen3.5-4B-Base)
LuminaV v1.3.0 was rigorously benchmarked on a 4.0-billion parameter Large Language Model (**`Qwen/Qwen3.5-4B-Base`**) initialized from architectural config (`AutoModelForCausalLM.from_config`) under **full pretraining from scratch** conditions (all 4B parameters actively optimized, gradient checkpointing disabled) on an NVIDIA 80GB GPU.