Theoretical Analysis of Multi-Head Double Deep Q-Network (XI-DDQN)

A convergence analysis of Q-decomposed Double Deep Q-Network (XI-DDQN) under linear function approximation, proving that decomposing the reward across interpretable "heads" (fuel, SOC, battery) preserves the convergence guarantees of monolithic DDQN, validated on a series-hybrid tractor energy-management task.

Accepted: H. Abououf, S. G. Bhatti, and Q. Ahmed, "Theoretical Analysis of Multi-Head Double Deep Q-Network with Linear Function Approximation," accepted at the 65th IEEE Conference on Decision and Control (CDC), 2026. Related: SidraBhatti/xi-ddqn-offroad-energy-management (the applied off-road energy-management paper this convergence analysis builds on).

Role & Attribution

Sidra Ghayour Bhatti β€” supervisory/advising role, alongside Qadeer Ahmed (PI, OSU Center for Automotive Research). Lead author and primary derivation: Hend Abououf.

Abstract

Q-decomposition Double Deep Q-Network (DDQN) enhances the transparency of the black-box DDQN by decomposing the global reward signal across subagents, improving the interpretability of the learned policy. Despite prior convergence guarantees for multi-agent SARSA and linear DQN, whether Q-decomposition preserves the convergence properties of DDQN has not been formally examined. In this paper we prove that, under standard linear function-approximation assumptions and a stability hypothesis on the double estimator, the Q-decomposed linear DDQN converges almost surely to the same projected Bellman fixed point as its monolithic counterpart, and that each head converges to its own projected fixed point under the shared policy. We further identify the conditions β€” realizability of the optimal Q-function in the feature span β€” under which this fixed point is the globally optimal Q-function. Numerical results on a hybrid tractor energy-management task are consistent with the analysis, with Q-decomposition achieving 97.4% of the fuel economy of the optimal solution obtained via dynamic programming.

Method

XI-DDQN decomposes the Q-function into one head per reward component (fuel, SOC deviation, battery-power feasibility), all sharing a common feature trunk; a shared arbitrator selects the action from the summed, weighted Q-values, and each head is trained on its own reward via a DDQN-style online/target split. The paper's key theoretical contribution is showing that because a single arbitrator (not each head individually) controls the policy, the multi-head system is algebraically equivalent to a monolithic linear DDQN β€” so the sum of head Q-values inherits almost-sure convergence to the projected Bellman fixed point (Theorem 1), each individual head converges to its own projected fixed point under the shared policy (Theorem 2), and under a realizability condition these fixed points coincide with the true optimal Q-functions (Theorem 3).

Multi-head DDQN energy management framework Fig. 1: Multi-head DDQN energy management framework β€” a shared trunk with one head per reward component (fuel, SOC, battery). The online network selects the power split, the target network evaluates it, and the powertrain returns the decomposed reward.

Results

Evaluated on a series-hybrid tractor energy-management task (state: battery SOC + normalized power demand; action: battery power over 100 discretized levels), benchmarked against the dynamic-programming (DP) optimum:

Metric DP XI-DDQN
Fuel consumed (gal) 2.108 2.163
Mean engine efficiency 43.7% 44.6%
Engine energy (kWh) 36.554 36.791
Battery energy (kWh) βˆ’2.456 βˆ’2.593
Load energy (kWh) 31.530 31.530

XI-DDQN reaches 97.4% of the DP fuel economy (0.055 gal more fuel over the full cycle) without any prior knowledge of the drive cycle or access to the engine map β€” consistent with the paper's almost-sure convergence result (Theorem 1). Training return increases monotonically to a stable plateau by episode 4,000, and both DP and XI-DDQN keep SOC within the [0.15, 0.85] bounds and end within 0.05 of the initial SOC, satisfying the charge-sustaining requirement.

Training return and SOC trajectory Fig. 2–3: Left β€” training return (100-episode moving average) over 8,000 episodes, plateauing at 78–82 from episode 4,000. Right β€” battery SOC over the drive cycle, DP (solid) vs. XI-DDQN (dashed): both end near 0.60; XI-DDQN dips further mid-cycle (0.406 vs. 0.461 for DP) but recovers to the charge-sustaining band.

Citation

@inproceedings{abououf2026xiddqn,
  author    = {Abououf, Hend and Bhatti, Sidra G. and Ahmed, Qadeer},
  title     = {Theoretical Analysis of Multi-Head Double Deep Q-Network with Linear Function Approximation},
  booktitle = {65th IEEE Conference on Decision and Control (CDC)},
  year      = {2026}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Collection including SidraBhatti/xi-ddqn-convergence-analysis-cdc2026