Instructions to use jgalego/ppo-turn-heading-cessna172p with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- stable-baselines3
How to use jgalego/ppo-turn-heading-cessna172p with stable-baselines3:
from huggingface_sb3 import load_from_hub checkpoint = load_from_hub( repo_id="jgalego/ppo-turn-heading-cessna172p", filename="{MODEL FILENAME}.zip", ) - Notebooks
- Google Colab
- Kaggle
ppo-turn-heading-cessna172p
A PPO agent that flies a Cessna 172P in the JSBSim flight dynamics model. It starts on a random heading at 5,000 ft and has to turn onto a random target heading and hold altitude, using only aileron, elevator and rudder.
Simulator only. Not for flight operations. The policy was trained and tested in one simulated aircraft with no wind or turbulence. It is not certified or validated for anything else.
The first 20 s of the sharpest turn in the test set, from a chase camera. The aircraft model is a low-poly stand-in drawn 5 times larger than life, but its heading, pitch and roll are the simulator's. The same flight over the full 60 s, with the instruments:
Top left is the ground track with the target heading dashed, top right an artificial horizon, then the control inputs and the heading and altitude errors.
🎯 The task
JSBSim-TurnHeadingControlTask-Cessna172P-Shaping.STANDARD-NoFG-v0:
| Aircraft | Cessna 172P, 120 kt cruise, 5,000 ft |
| Start | Random heading, wings level |
| Goal | Fly the target heading (random) and stay at 5,000 ft |
| Observation | 17 values: altitude, attitude, body velocities, rotation rates, control positions, altitude error, sideslip, track error, steps left |
| Action | Aileron, elevator, rudder in [-1, 1]. Throttle is fixed |
| Rate | 5 decisions per second, 60 s episodes (299 steps) |
| Reward | Positive and bounded per step, from altitude error and track error |
| Ends | After 60 s, or when the aircraft is more than 1,000 ft off altitude |
🚀 Usage
import gymnasium as gym
import gymnasium_jsbsim
from huggingface_hub import hf_hub_download
from stable_baselines3 import PPO
from stable_baselines3.common.vec_env import DummyVecEnv, VecNormalize
repo = "jgalego/ppo-turn-heading-cessna172p"
env = DummyVecEnv([lambda: gym.make("JSBSim-TurnHeadingControlTask-Cessna172P-Shaping.STANDARD-NoFG-v0")])
env = VecNormalize.load(hf_hub_download(repo, "vecnormalize.pkl"), env)
env.training = False
env.norm_reward = False
model = PPO.load(hf_hub_download(repo, "ppo.zip"))
obs = env.reset()
done = False
while not done:
action, _ = model.predict(obs, deterministic=True)
obs, reward, done, info = env.step(action)
gymnasium-jsbsim is installed with pip install git+https://github.com/JGalego/gymnasium-jsbsim. Observations must go through the saved VecNormalize statistics.
🏋️ Training
| Algorithm | PPO, MlpPolicy (2 x 64, tanh), Stable-Baselines3 |
| Steps | 2,000,000 in 12 parallel environments |
| Rollout | 512 steps per environment, batch size 512 |
| Learning rate | 0.0003, discount 0.99 |
| Normalization | VecNormalize on observations and rewards |
| Time limit | The 60 s limit counts as truncation, not as failure |
| Hardware | cpu, 12 worker processes, 26 min |
📊 Results
200 episodes with fresh random start and target headings. Every row flies the same episodes. Higher return is better. A run is completed if it stays within 1,000 ft for the full 60 s. Track error is averaged over the last 10 s, and the settle time is the median time until the track error stays under 5° on episodes that start more than 20° off, or "never" if fewer than half get there. Peak bank is the median over episodes.
| Controller | Return | Completed | Track error | Settle time | Peak bank | Altitude error |
|---|---|---|---|---|---|---|
| PPO | 290.5 ± 4.5 | 100% | 0.1° | 3.2 s | 85° | 1 ft |
| PID | 174.6 ± 23.6 | 100% | 1.1° | 28.4 s | 23° | 117 ft |
| Zero input | 62.1 ± 1.5 | 0% | 87.2° | never | 82° | 424 ft |
| Random | 36.4 ± 15.2 | 2% | 91.2° | never | 127° | 319 ft |
PID is a hand-tuned cascade written for comparison: it banks toward the target heading and holds altitude with pitch. Zero input leaves the controls alone, and random draws uniform commands.
Median and 10-90% range over the episodes, for the agent and the PID baseline.
Many turns, one frame
Each aircraft is rotated so its target heading points up. Start headings are spread over the full circle.
What it learned
Command as a function of the error, with everything else in level flight. The PID aileron is shown for comparison.
⚠️ Limitations
- The turns are far more aggressive than a pilot would fly. The reward does not penalize bank angle or control effort, so the agent rolls past 90° and snaps onto the heading in a few seconds. Add a penalty on bank or control rate before reading anything operational into the behavior.
- One aircraft, one altitude, fixed throttle, calm air and a perfect state estimate.
- Trained on whole-circle turns of a 60 s episode; longer flights or other aircraft are untested.
- The reward is a proxy for tracking error, so a good return does not mean a good turn. Check the heading error as well.
- Downloads last month
- 23
Space using jgalego/ppo-turn-heading-cessna172p 1
Collection including jgalego/ppo-turn-heading-cessna172p
Evaluation results
- Mean reward on JSBSim-TurnHeadingControlTask-Cessna172P-Shaping.STANDARD-NoFG-v0self-reported290.5 +/- 4.5





