🛰️ Double Dueling DQN Agent - LunarLander-v3 (Stylish Precision Landing)
This repository contains a trained Double Dueling Deep Q-Network (DQN) agent for the Gymnasium LunarLander-v3 environment.
The agent was trained for 1000 episodes with an Epsilon exploration schedule decaying from 100% (1.0) to 5% (0.05), combined with custom stylish & smooth landing reward shaping (vertical attitude control $\theta \approx 0$, soft touchdown descent rate, and helipad center alignment).
🎯 Model Hyperparameters
| Parameter | Value |
|---|---|
| Algorithm | Double Dueling DQN |
| Environment | Gymnasium LunarLander-v3 |
| State Dimension | 8 |
| Action Dimension | 4 (Discrete: Idle, Left RCS, Main Thruster, Right RCS) |
| Total Episodes | 1000 |
| Epsilon Schedule | 1.0 (100%) $\rightarrow$ 0.05 (5%) |
| Discount Factor ($\gamma$) | 0.99 |
| Target Update | Soft Polyak Averaging ($\tau = 0.005$) |
| Learning Rate | 5e-4 (Adam Optimizer) |
| Loss Function | Smooth L1 (Huber Loss) |
| Replay Buffer Capacity | 100,000 |
| Batch Size | 64 |
🚀 How to Load and Test the Model
import torch
import gymnasium as gym
from lunar_dqn_agent import DQNAgent
# 1. Initialize environment
env = gym.make("LunarLander-v3", render_mode="human")
state, _ = env.reset(seed=42)
# 2. Load trained DQN Agent
agent = DQNAgent()
agent.load("best_lunar_dqn.pth")
total_reward = 0
done = False
while not done:
# Select best greedy action (evaluation mode)
action, q_values, _ = agent.select_action(state, evaluate=True)
next_state, reward, terminated, truncated, _ = env.step(action)
done = terminated or truncated
state = next_state
total_reward += reward
print(f"Landing Completed! Total Reward: {total_reward:.2f}")
env.close()
🏆 Features
- Smooth Descent: Soft vertical velocity control preventing hard impacts.
- Upright Posture: High angular stability with minimum wobbling.
- Precision Touchdown: Guided landing directly between the target flags.
- Downloads last month
- 5