Theoretical Analysis of Multi-Head Double Deep Q-Network (XI-DDQN)
A convergence analysis of Q-decomposed Double Deep Q-Network (XI-DDQN) under linear function approximation, proving that decomposing the reward across interpretable "heads" (fuel, SOC, battery) preserves the convergence guarantees of monolithic DDQN, validated on a series-hybrid tractor energy-management task.
Accepted: H. Abououf, S. G. Bhatti, and Q. Ahmed, "Theoretical Analysis of Multi-Head Double Deep Q-Network with Linear Function Approximation," accepted at the 65th IEEE Conference on Decision and Control (CDC), 2026. Related: SidraBhatti/xi-ddqn-offroad-energy-management (the applied off-road energy-management paper this convergence analysis builds on).
Role & Attribution
Sidra Ghayour Bhatti β supervisory/advising role, alongside Qadeer Ahmed (PI, OSU Center for Automotive Research). Lead author and primary derivation: Hend Abououf.
Abstract
Q-decomposition Double Deep Q-Network (DDQN) enhances the transparency of the black-box DDQN by decomposing the global reward signal across subagents, improving the interpretability of the learned policy. Despite prior convergence guarantees for multi-agent SARSA and linear DQN, whether Q-decomposition preserves the convergence properties of DDQN has not been formally examined. In this paper we prove that, under standard linear function-approximation assumptions and a stability hypothesis on the double estimator, the Q-decomposed linear DDQN converges almost surely to the same projected Bellman fixed point as its monolithic counterpart, and that each head converges to its own projected fixed point under the shared policy. We further identify the conditions β realizability of the optimal Q-function in the feature span β under which this fixed point is the globally optimal Q-function. Numerical results on a hybrid tractor energy-management task are consistent with the analysis, with Q-decomposition achieving 97.4% of the fuel economy of the optimal solution obtained via dynamic programming.
Method
XI-DDQN decomposes the Q-function into one head per reward component (fuel, SOC deviation, battery-power feasibility), all sharing a common feature trunk; a shared arbitrator selects the action from the summed, weighted Q-values, and each head is trained on its own reward via a DDQN-style online/target split. The paper's key theoretical contribution is showing that because a single arbitrator (not each head individually) controls the policy, the multi-head system is algebraically equivalent to a monolithic linear DDQN β so the sum of head Q-values inherits almost-sure convergence to the projected Bellman fixed point (Theorem 1), each individual head converges to its own projected fixed point under the shared policy (Theorem 2), and under a realizability condition these fixed points coincide with the true optimal Q-functions (Theorem 3).
Fig. 1: Multi-head DDQN energy management framework β a shared trunk
with one head per reward component (fuel, SOC, battery). The online
network selects the power split, the target network evaluates it, and
the powertrain returns the decomposed reward.
Results
Evaluated on a series-hybrid tractor energy-management task (state: battery SOC + normalized power demand; action: battery power over 100 discretized levels), benchmarked against the dynamic-programming (DP) optimum:
| Metric | DP | XI-DDQN |
|---|---|---|
| Fuel consumed (gal) | 2.108 | 2.163 |
| Mean engine efficiency | 43.7% | 44.6% |
| Engine energy (kWh) | 36.554 | 36.791 |
| Battery energy (kWh) | β2.456 | β2.593 |
| Load energy (kWh) | 31.530 | 31.530 |
XI-DDQN reaches 97.4% of the DP fuel economy (0.055 gal more fuel over the full cycle) without any prior knowledge of the drive cycle or access to the engine map β consistent with the paper's almost-sure convergence result (Theorem 1). Training return increases monotonically to a stable plateau by episode 4,000, and both DP and XI-DDQN keep SOC within the [0.15, 0.85] bounds and end within 0.05 of the initial SOC, satisfying the charge-sustaining requirement.
Fig. 2β3: Left β training return (100-episode moving average) over
8,000 episodes, plateauing at 78β82 from episode 4,000. Right β battery
SOC over the drive cycle, DP (solid) vs. XI-DDQN (dashed): both end near
0.60; XI-DDQN dips further mid-cycle (0.406 vs. 0.461 for DP) but
recovers to the charge-sustaining band.
Citation
@inproceedings{abououf2026xiddqn,
author = {Abououf, Hend and Bhatti, Sidra G. and Ahmed, Qadeer},
title = {Theoretical Analysis of Multi-Head Double Deep Q-Network with Linear Function Approximation},
booktitle = {65th IEEE Conference on Decision and Control (CDC)},
year = {2026}
}