LunarLander-v3 in 14 parameters

A 14-parameter neural network that solves Gymnasium LunarLander-v3.

logits = obs @ W                  # obs: (8,) float32, W: (8, 4)
action = argmax(logits)           # 0 nothing, 1 left engine, 2 main engine, 3 right engine

W is not free: it is built from 14 trainable scalars by hard-coding the exact mirror symmetry of the environment (x -> -x, vx -> -vx, angle -> -angle, angular_velocity -> -angular_velocity, the legs swap, actions left <-> right; actions 0 and 2 are self-paired). The 14 parameters live in model.safetensors as sym_theta (shape (14,)); model.py materialises the 8x4 matrix at load time, so the model really has 14 parameters.

Landing

Greedy landings recorded from the environment (three episodes, seeded 40122 / 42940 / 51305, returns 294.4 / 275.8 / 276.8):

Single episode as an animated GIF:

LunarLander-v3 landing

Evaluation

Deterministic greedy policy; "landing" and "crash" are the environment's own terminal +100 / -100 verdicts.

evaluation set mean return seeds >= 200 landings crashes
60 evaluation seeds x 100 episodes 264.06 60/60 99.55% 0.35%
60 fresh (disjoint) seeds x 100 episodes 263.83 60/60 99.48% 0.38%

The "solved" threshold for this environment is a mean return of 200; a random policy scores about -100 to -500.

Usage

import gymnasium as gym
from model import TinyLunarPolicy

policy = TinyLunarPolicy.from_pretrained("Dimitrius174/lunarlander-14params")
# or: TinyLunarPolicy.from_safetensors("model.safetensors")

print(policy.n_params)                                    # 14
action = policy.act([0.0, 1.4, 0.0, -0.6, 0.0, 0.0, 0.0, 0.0])   # int in 0..3

env = gym.make("LunarLander-v3")
obs, _ = env.reset(seed=0)
done = False
while not done:
    obs, _, term, trunc, _ = env.step(policy.act(obs))
    done = term or trunc

policy.numpy_weights() returns the materialised W of shape (8, 4).

Observation and action space

  • obs (8,): x, y, vx, vy, angle, angular_velocity, leg_left_contact, leg_right_contact
  • action (4): do nothing, fire left engine, fire main engine, fire right engine

Modes and conflicts: why 14 parameters are enough

The 14 numbers are not tied to one instance of the environment. The same weights were re-evaluated with perturbed physics — gravity g = -6, -8, -10, -12 instead of the default -10 — landing on 1.000 / 0.990 / 1.000 / 0.840 of 100 episodes each. Perturbations of this kind touch the decision only through a scale: argmax(λ · obs @ W) = argmax(obs @ W) for λ > 0, and a rescaled thrust or gravity just rescales the same bang-bang stabilisation. So one policy serves every mode: a perturbation costs zero parameters.

What does cost parameters is a conflict — two modes that demand opposite actions for the same observation. Then you need one linear piece per conflict class, and

K_needed = number of conflict classes K*, not the number of modes: success ≈ min(K / K*, 1).

family (this model unless noted) modes conflict classes K* pieces needed measured
gravity g = -6…-12 4 1 1 (the same 14 numbers) land 0.84–1.00
label swap left <-> right 2 2 1 + a fixed relabelling outside the model (0 params) 1.00 vs 0.55 (per-mode [1.00, 0.00])
label swap, single linear map only 2 2 2 (36 params) 1.00
CartPole sign-calibration gauges 5 5 5 1.00 vs 0.20 ([1,0,0,0,0] with 1)
CartPole 8 force levels, 1 gauge 8 1 1 1.00 (all 8 in one cluster)
CartPole 2 gauges × 4 force levels 8 2 2 0.97 vs 0.10 with 1
CartPole 3 gauges × 3 force levels 9 3 3 0.87 with 3 vs 0.17 with 1

The same rule answers "how many parameters do I need for this set of environments": count conflict classes, not environments. A data-driven estimator of K* (merge the modes whose optimal rules coincide on the data; the estimate is a safe upper bound) is evo_kstar_estimator.py / evo_conflict_sharpness.py; discussion in CENTROID_TO_NN.md §13–§14.

Files

file what
model.py TinyLunarPolicy: 14 parameters, symmetry built in (torch only)
model.safetensors the 14 weights (sym_theta, shape (14,))
config.json architecture and environment metadata
verify.py structural checks and live Gymnasium evaluation
record_video.py records lunarlander-14params.mp4 / .gif (headless, rgb_array)
lunarlander-14params.mp4, .gif landing videos (3 episodes / 1 episode)
requirements.txt dependencies

License

MIT. See LICENSE.

Downloads last month
37
Safetensors
Model size
14 params
Tensor type
F32
·
Video Preview
loading