# LuminaV Empirical Benchmark Report: Qwen3.5-4B-Base This document provides official, reproducible empirical benchmark results for **LuminaV v1.3.0**, evaluated on a 4-billion parameter Large Language Model (**Qwen/Qwen3.5-4B-Base**) initialized from architectural config (`AutoModelForCausalLM.from_config`) under full-parameter pretraining conditions from scratch, without gradient checkpointing and without parameter-efficient adapters (NOT LoRA/PEFT). [](https://colab.research.google.com/drive/1vgNIo8O2MpH9fU-kreg08vPwaiucsXik?usp=sharing) --- ## 1. Hardware & Execution Environment The benchmark was executed under an isolated compute environment with full VRAM purging (`torch.cuda.empty_cache()` and `ipc_collect`) prior to each run: | Hardware / Runtime Component | Specification | | :--- | :--- | | **GPU Accelerator** | NVIDIA A100 (80 GB VRAM) | | **Total Dedicated VRAM** | 79.3 GB | | **Execution Policy** | Max Throughput Mode (**Gradient Checkpointing DISABLED**) | | **Baseline Model VRAM** | ~7.94 GB (Model parameter footprint in native `bfloat16`) | | **PyTorch Accelerator** | Native CUDA with Custom OpenAI Triton Kernels | | **Attention Implementation** | PyTorch Scaled Dot-Product Attention (`sdpa`) | --- ## 2. Model & Dataset Configuration | Parameter | Configuration Value | | :--- | :--- | | **Target Model** | `Qwen/Qwen3.5-4B-Base` (~4.0 Billion Parameters) | | **Training Setup** | **Full Pretraining from Scratch** (Initialized via `AutoModelForCausalLM.from_config`; all 4.0B parameters actively optimized; **NOT** downstream fine-tuning, **NOT** LoRA/PEFT) | | **Weight Initialization** | Random Gaussian/Uniform architectural initialization from config (explains the theoretical initial Cross-Entropy loss of ~12.99 matching uniform vocabulary distribution: `ln(V) ≈ 12.0 - 13.0`) | | **Buffer Architecture** | **Strictly Dual-Buffer Setup (`buffer=2`)** across all tests (tracking both momentum `m_t` and centered innovation variance `v_t`). Single-buffer mode (`buffer=1`) was **NOT** evaluated in this benchmark suite. | | **Precision** | Native `torch.bfloat16` | | **Sequence Length (`seq_len`)** | 512 tokens | | **Batch Size (`batch_size`)** | 1 (Packed dense sequence) | | **Total Training Steps** | 250 steps | | **Learning Rate Schedule** | Cosine Annealing with Warmup (`WARMUP_STEPS = 25`, `eta_min = 0.0`) | | **Peak Learning Rate (`lr`)** | `8e-4` | | **Weight Decay (`weight_decay`)**| `0.08` | | **Optimizer Hyperparameters** | `betas=(0.9, 0.999)`, `eps=1e-8`, `tau=0.8`, `bound=True` (radial, ratio=0.03) | | **Streaming Multi-Corpus** | Packed bilingual stream: 40% Xone Translation Corpus, 30% Wikipedia (EN/ID), 30% CulturaX (EN/ID) | --- ## 3. Parametric Benchmark Suite > **Critical Architectural Disambiguation (1P vs 2P vs Buffer Count):** > The notations **1P** and **2P** refer strictly to **GPU Kernel Execution Passes** (`1P = Fused Single-Pass Kernel with Post-Step EMA`, `2P = Standard Two-Pass Kernel Reduction`). > **All six configurations (M1–M6) were executed exclusively under the Dual-Buffer setup (`buffer=2`)**, maintaining both the first moment (`m_t`) and centered innovation variance (`v_t`) buffers. Single-buffer scalar RMS mode (`buffer=1`) was deliberately excluded from this pretraining run. | Mode ID | Optimizer Mode | Master Weights Mode | Kernel Passes | Buffer Count | Description | | :--- | :--- | :--- | :---: | :---: | :--- | | **M1** | **Master-Free (2P)** | `"none"` (False) | **2-Pass** (`False`) | `buffer=2` (Dual) | Master-free, exact 2-pass Triton reduction with `m_t` and `v_t`. | | **M2** | **Master-Free (1P)** | `"none"` (False) | **1-Pass** (`True`) | `buffer=2` (Dual) | Master-free, 1-pass fused kernel via running EMA with `m_t` and `v_t`. | | **M3** | **Semi (2P)** | `"semi"` | **2-Pass** (`False`) | `buffer=2` (Dual) | FP32 master weights with 16-bit low-memory moments (2-pass). | | **M4** | **Semi (1P)** | `"semi"` | **1-Pass** (`True`) | `buffer=2` (Dual) | FP32 master weights with 16-bit low-memory moments (1-pass EMA). | | **M5** | **Full (2P)** | `"full"` | **2-Pass** (`False`) | `buffer=2` (Dual) | Traditional FP32 master weights and FP32 optimizer states (2-pass). | | **M6** | **Full (1P)** | `"full"` | **1-Pass** (`True`) | `buffer=2` (Dual) | Traditional FP32 master weights and FP32 optimizer states (1-pass EMA). | --- ## 4. Empirical Performance Summary All metrics are captured using hardware-synchronized CUDA events (`torch.cuda.Event`) after a 5-step warmup exclusion. ### Core Metrics Summary | Optimizer Mode | Peak VRAM | Static VRAM | Dynamic Act. | Step Time | Pure Opt Time | Initial Loss | Final Loss (Smoothed) | | :--- | :---: | :---: | :---: | :---: | :---: | :---: | :---: | | **Master-Free (2P)** | **34.32 GB** | 24.65 GB | ~9.67 GB | 1369.3 ms | 68.4 ms | 12.9938 | 9.1075 | | **Master-Free (1P)** | **34.32 GB** | 24.65 GB | ~9.67 GB | **1365.2 ms** | **59.9 ms** | 12.9938 | **8.8260** | | **Semi (2P)** | 49.99 GB | 40.32 GB | ~9.67 GB | 1399.5 ms | 87.8 ms | 12.9938 | 8.8673 | | **Semi (1P)** | 49.99 GB | 40.32 GB | ~9.67 GB | 1378.5 ms | 70.2 ms | 12.9938 | **8.8192** | | **Full (2P)** | 65.47 GB | 55.80 GB | ~9.67 GB | 1424.2 ms | 113.8 ms | 12.9938 | 8.9579 | | **Full (1P)** | 65.47 GB | 55.80 GB | ~9.67 GB | 1400.4 ms | 87.3 ms | 12.9938 | 8.9770 | --- ### Step-by-Step Loss Convergence Trajectory | Step | Master-Free (2P) | Master-Free (1P) | Semi (2P) | Semi (1P) | Full (2P) | Full (1P) | | :---: | :---: | :---: | :---: | :---: | :---: | :---: | | **1** | 12.9938 | 12.9938 | 12.9938 | 12.9938 | 12.9938 | 12.9938 | | **50** | 9.3490 | 10.0002 | 9.6643 | 9.3330 | 9.8324 | 9.6631 | | **100** | 7.4108 | 7.0516 | 7.1035 | 7.2455 | 7.3232 | 7.3519 | | **150** | 7.9167 | 7.6074 | 7.6650 | 7.8517 | 7.8090 | 7.8402 | | **200** | 7.7148 | 7.5384 | 7.5651 | 7.5854 | 7.6170 | 7.5987 | | **250** | 9.6568 | 9.3338 | 9.4042 | 9.3306 | 9.4838 | 9.5086 | --- ## 5. Visualizations & Analytical Charts The three-panel visualization below displays the pretraining loss convergence, static vs dynamic VRAM distribution, and pure optimizer kernel execution latency across all tested modes: