🛰️ Double Dueling DQN Agent - LunarLander-v3 (Stylish Precision Landing)

This repository contains a trained Double Dueling Deep Q-Network (DQN) agent for the Gymnasium LunarLander-v3 environment.

The agent was trained for 1000 episodes with an Epsilon exploration schedule decaying from 100% (1.0) to 5% (0.05), combined with custom stylish & smooth landing reward shaping (vertical attitude control $\theta \approx 0$, soft touchdown descent rate, and helipad center alignment).


🎯 Model Hyperparameters

Parameter Value
Algorithm Double Dueling DQN
Environment Gymnasium LunarLander-v3
State Dimension 8
Action Dimension 4 (Discrete: Idle, Left RCS, Main Thruster, Right RCS)
Total Episodes 1000
Epsilon Schedule 1.0 (100%) $\rightarrow$ 0.05 (5%)
Discount Factor ($\gamma$) 0.99
Target Update Soft Polyak Averaging ($\tau = 0.005$)
Learning Rate 5e-4 (Adam Optimizer)
Loss Function Smooth L1 (Huber Loss)
Replay Buffer Capacity 100,000
Batch Size 64

🚀 How to Load and Test the Model

import torch
import gymnasium as gym
from lunar_dqn_agent import DQNAgent

# 1. Initialize environment
env = gym.make("LunarLander-v3", render_mode="human")
state, _ = env.reset(seed=42)

# 2. Load trained DQN Agent
agent = DQNAgent()
agent.load("best_lunar_dqn.pth")

total_reward = 0
done = False

while not done:
    # Select best greedy action (evaluation mode)
    action, q_values, _ = agent.select_action(state, evaluate=True)
    next_state, reward, terminated, truncated, _ = env.step(action)
    done = terminated or truncated
    state = next_state
    total_reward += reward

print(f"Landing Completed! Total Reward: {total_reward:.2f}")
env.close()

🏆 Features

  • Smooth Descent: Soft vertical velocity control preventing hard impacts.
  • Upright Posture: High angular stability with minimum wobbling.
  • Precision Touchdown: Guided landing directly between the target flags.
Downloads last month
5
Video Preview
loading