ppo-turn-heading-cessna172p

A PPO agent that flies a Cessna 172P in the JSBSim flight dynamics model. It starts on a random heading at 5,000 ft and has to turn onto a random target heading and hold altitude, using only aileron, elevator and rudder.

Simulator only. Not for flight operations. The policy was trained and tested in one simulated aircraft with no wind or turbulence. It is not certified or validated for anything else.

The agent flying the aircraft through a turn

The first 20 s of the sharpest turn in the test set, from a chase camera. The aircraft model is a low-poly stand-in drawn 5 times larger than life, but its heading, pitch and roll are the simulator's. The same flight over the full 60 s, with the instruments:

Ground track, artificial horizon, control inputs and errors

Top left is the ground track with the target heading dashed, top right an artificial horizon, then the control inputs and the heading and altitude errors.

🎯 The task

JSBSim-TurnHeadingControlTask-Cessna172P-Shaping.STANDARD-NoFG-v0:

Aircraft Cessna 172P, 120 kt cruise, 5,000 ft
Start Random heading, wings level
Goal Fly the target heading (random) and stay at 5,000 ft
Observation 17 values: altitude, attitude, body velocities, rotation rates, control positions, altitude error, sideslip, track error, steps left
Action Aileron, elevator, rudder in [-1, 1]. Throttle is fixed
Rate 5 decisions per second, 60 s episodes (299 steps)
Reward Positive and bounded per step, from altitude error and track error
Ends After 60 s, or when the aircraft is more than 1,000 ft off altitude

🚀 Usage

import gymnasium as gym
import gymnasium_jsbsim
from huggingface_hub import hf_hub_download
from stable_baselines3 import PPO
from stable_baselines3.common.vec_env import DummyVecEnv, VecNormalize

repo = "jgalego/ppo-turn-heading-cessna172p"
env = DummyVecEnv([lambda: gym.make("JSBSim-TurnHeadingControlTask-Cessna172P-Shaping.STANDARD-NoFG-v0")])
env = VecNormalize.load(hf_hub_download(repo, "vecnormalize.pkl"), env)
env.training = False
env.norm_reward = False
model = PPO.load(hf_hub_download(repo, "ppo.zip"))

obs = env.reset()
done = False
while not done:
    action, _ = model.predict(obs, deterministic=True)
    obs, reward, done, info = env.step(action)

gymnasium-jsbsim is installed with pip install git+https://github.com/JGalego/gymnasium-jsbsim. Observations must go through the saved VecNormalize statistics.

🏋️ Training

Algorithm PPO, MlpPolicy (2 x 64, tanh), Stable-Baselines3
Steps 2,000,000 in 12 parallel environments
Rollout 512 steps per environment, batch size 512
Learning rate 0.0003, discount 0.99
Normalization VecNormalize on observations and rewards
Time limit The 60 s limit counts as truncation, not as failure
Hardware cpu, 12 worker processes, 26 min

Learning curve

📊 Results

200 episodes with fresh random start and target headings. Every row flies the same episodes. Higher return is better. A run is completed if it stays within 1,000 ft for the full 60 s. Track error is averaged over the last 10 s, and the settle time is the median time until the track error stays under 5° on episodes that start more than 20° off, or "never" if fewer than half get there. Peak bank is the median over episodes.

Controller Return Completed Track error Settle time Peak bank Altitude error
PPO 290.5 ± 4.5 100% 0.1° 3.2 s 85° 1 ft
PID 174.6 ± 23.6 100% 1.1° 28.4 s 23° 117 ft
Zero input 62.1 ± 1.5 0% 87.2° never 82° 424 ft
Random 36.4 ± 15.2 2% 91.2° never 127° 319 ft

PID is a hand-tuned cascade written for comparison: it banks toward the target heading and holds altitude with pitch. Zero input leaves the controls alone, and random draws uniform commands.

Track and altitude error over time

Median and 10-90% range over the episodes, for the agent and the PID baseline.

Many turns, one frame

Each aircraft is rotated so its target heading points up. Start headings are spread over the full circle.

Aircraft converging on their target headings

What it learned

Command as a function of the error, with everything else in level flight. The PID aileron is shown for comparison.

Control law of the agent and the PID

⚠️ Limitations

  • The turns are far more aggressive than a pilot would fly. The reward does not penalize bank angle or control effort, so the agent rolls past 90° and snaps onto the heading in a few seconds. Add a penalty on bank or control rate before reading anything operational into the behavior.
  • One aircraft, one altitude, fixed throttle, calm air and a perfect state estimate.
  • Trained on whole-circle turns of a 60 s episode; longer flights or other aircraft are untested.
  • The reward is a proxy for tracking error, so a good return does not mean a good turn. Check the heading error as well.
Downloads last month
23
Video Preview
loading

Space using jgalego/ppo-turn-heading-cessna172p 1

Collection including jgalego/ppo-turn-heading-cessna172p

Evaluation results

  • Mean reward on JSBSim-TurnHeadingControlTask-Cessna172P-Shaping.STANDARD-NoFG-v0
    self-reported
    290.5 +/- 4.5