LunarLander-v3 in 14 parameters
A 14-parameter neural network that solves Gymnasium
LunarLander-v3.
logits = obs @ W # obs: (8,) float32, W: (8, 4)
action = argmax(logits) # 0 nothing, 1 left engine, 2 main engine, 3 right engine
W is not free: it is built from 14 trainable scalars by hard-coding the exact mirror
symmetry of the environment (x -> -x, vx -> -vx, angle -> -angle,
angular_velocity -> -angular_velocity, the legs swap, actions left <-> right; actions 0 and 2
are self-paired). The 14 parameters live in model.safetensors as sym_theta (shape (14,));
model.py materialises the 8x4 matrix at load time, so the model really has 14 parameters.
Landing
Greedy landings recorded from the environment (three episodes, seeded 40122 / 42940 / 51305, returns 294.4 / 275.8 / 276.8):
Single episode as an animated GIF:
Evaluation
Deterministic greedy policy; "landing" and "crash" are the environment's own terminal +100 / -100
verdicts.
| evaluation set | mean return | seeds >= 200 | landings | crashes |
|---|---|---|---|---|
| 60 evaluation seeds x 100 episodes | 264.06 | 60/60 | 99.55% | 0.35% |
| 60 fresh (disjoint) seeds x 100 episodes | 263.83 | 60/60 | 99.48% | 0.38% |
The "solved" threshold for this environment is a mean return of 200; a random policy scores about -100 to -500.
Usage
import gymnasium as gym
from model import TinyLunarPolicy
policy = TinyLunarPolicy.from_pretrained("Dimitrius174/lunarlander-14params")
# or: TinyLunarPolicy.from_safetensors("model.safetensors")
print(policy.n_params) # 14
action = policy.act([0.0, 1.4, 0.0, -0.6, 0.0, 0.0, 0.0, 0.0]) # int in 0..3
env = gym.make("LunarLander-v3")
obs, _ = env.reset(seed=0)
done = False
while not done:
obs, _, term, trunc, _ = env.step(policy.act(obs))
done = term or trunc
policy.numpy_weights() returns the materialised W of shape (8, 4).
Observation and action space
obs(8,):x, y, vx, vy, angle, angular_velocity, leg_left_contact, leg_right_contactaction(4):do nothing,fire left engine,fire main engine,fire right engine
Modes and conflicts: why 14 parameters are enough
The 14 numbers are not tied to one instance of the environment. The same weights were re-evaluated with
perturbed physics — gravity g = -6, -8, -10, -12 instead of the default -10 — landing on
1.000 / 0.990 / 1.000 / 0.840 of 100 episodes each. Perturbations of this kind touch the decision
only through a scale: argmax(λ · obs @ W) = argmax(obs @ W) for λ > 0, and a rescaled thrust or
gravity just rescales the same bang-bang stabilisation. So one policy serves every mode: a
perturbation costs zero parameters.
What does cost parameters is a conflict — two modes that demand opposite actions for the same observation. Then you need one linear piece per conflict class, and
K_needed = number of conflict classes K*, not the number of modes:success ≈ min(K / K*, 1).
| family (this model unless noted) | modes | conflict classes K* | pieces needed | measured |
|---|---|---|---|---|
gravity g = -6…-12 |
4 | 1 | 1 (the same 14 numbers) | land 0.84–1.00 |
label swap left <-> right |
2 | 2 | 1 + a fixed relabelling outside the model (0 params) | 1.00 vs 0.55 (per-mode [1.00, 0.00]) |
| label swap, single linear map only | 2 | 2 | 2 (36 params) | 1.00 |
| CartPole sign-calibration gauges | 5 | 5 | 5 | 1.00 vs 0.20 ([1,0,0,0,0] with 1) |
| CartPole 8 force levels, 1 gauge | 8 | 1 | 1 | 1.00 (all 8 in one cluster) |
| CartPole 2 gauges × 4 force levels | 8 | 2 | 2 | 0.97 vs 0.10 with 1 |
| CartPole 3 gauges × 3 force levels | 9 | 3 | 3 | 0.87 with 3 vs 0.17 with 1 |
The same rule answers "how many parameters do I need for this set of environments": count conflict
classes, not environments. A data-driven estimator of K* (merge the modes whose optimal rules coincide
on the data; the estimate is a safe upper bound) is evo_kstar_estimator.py /
evo_conflict_sharpness.py; discussion in CENTROID_TO_NN.md §13–§14.
Files
| file | what |
|---|---|
model.py |
TinyLunarPolicy: 14 parameters, symmetry built in (torch only) |
model.safetensors |
the 14 weights (sym_theta, shape (14,)) |
config.json |
architecture and environment metadata |
verify.py |
structural checks and live Gymnasium evaluation |
record_video.py |
records lunarlander-14params.mp4 / .gif (headless, rgb_array) |
lunarlander-14params.mp4, .gif |
landing videos (3 episodes / 1 episode) |
requirements.txt |
dependencies |
License
MIT. See LICENSE.
- Downloads last month
- 37
