Reinforcement Learning for Orbital Transfers at the 2026 AI Winter School (Brown University)
Originally published Updated 5 min read
Training PPO policies for orbital transfers and comparing their trajectories and delta-v with a Hohmann baseline.

At the 2026 AI Winter School, hosted by the Center for the Fundamental Physics of the Universe at Brown University, I led a 2.5-hour hands-on workshop on reinforcement learning for orbital transfers.
I used a two-body transfer with a known analytic solution so we could compare learned policies with the Hohmann baseline. The notebook contains the environment, training procedure, and example outputs.
Control Problem
The notebook used nondimensional two-body dynamics: unit gravitational parameter, an initial circular orbit at radius 1, and a target circular orbit at radius 1.6. I call those radii and below. The model omitted drag, finite-duration thrust, J2 perturbations, third bodies, attitude dynamics, and mass depletion. Control was a tangential impulse applied once per simulation step.
Between impulses, the spacecraft position and velocity evolve under central gravity:
For a circular target orbit, the target specific energy and angular momentum are:
The observation vector contains normalized radius, radial velocity, tangential velocity, angular-momentum error, energy error, and the previous action. It omits the absolute orbital angle, so rotating an otherwise identical state does not change the policy input.
Hohmann Benchmark
Before training PPO, the notebook computed the Hohmann transfer. For circular, coplanar orbits with two impulsive burns, the transfer semi-major axis is:
The two burns and transfer time are:
The comparisons use total Δv, circularization error, and burn history.

The transfer ellipse touches the initial circular orbit at r1 and the target orbit at r2.

Radius history, the two impulses, and accumulated Δv for the simulated Hohmann transfer.
The second Hohmann burn is applied on the first simulation step at or after the analytic transfer time. With a timestep of 0.00050, the saved baseline ends at radius 1.5999 rather than 1.6000, and total Δv agrees with the analytic value to the displayed precision of 0.2066.
Try it: fly the transfer yourself
Watch a two-burn transfer, then try your own.
Hold Speed up or Slow down to steer toward the blue orbit.
About the model
This two-body model uses normalized units with GM = 1 and orbit radii 1 and 1.6. Browser burns follow or oppose velocity, including its radial component; the PPO notebook applies burns along the local orbital tangent. Hohmann and Greedy here are analytical and feedback controllers.
RL Formulation
The Gymnasium environment held the physics fixed and varied the control interface:
| Component | Implementation |
|---|---|
| Observation | Normalized r, vr, vt, target angular-momentum error, target energy error, previous action |
| Discrete action | coast, full prograde impulse, full retrograde impulse |
| Continuous action | throttle in [-1, 1], mapped to a signed tangential Δv impulse |
| Success criteria | tolerances on , , and |
| Failure criteria | crash/escape radius or episode timeout |
The dense reward used a combined energy/angular-momentum error:
with shaping approximately proportional to:
The reward also included fuel and ignition/switching penalties, a one-time bonus on first entry into the success region, and a reward for remaining there. PPO training used observation and reward normalization. Evaluation used frozen normalization statistics and deterministic policy actions.
Failure modes to inspect
For each policy, the notebook plots the trajectory, radius, radial velocity, thrust impulses, and cumulative Δv. Its mission report includes the burn count or active-thrust steps and closest-to-target statistics for comparison with the Hohmann transfer.
These diagnostics can reveal several failure modes:
- Repeated corrections: A discrete policy may reach the target through many small prograde and retrograde impulses, satisfying the orbit tolerances while using excessive Δv.
- Continuous small impulses: A continuous policy may keep orbital errors small by applying throttle almost continuously. The accumulated Δv reveals the cost of these corrections.
- Loose success tolerances: A policy can enter the success region while remaining on an elliptical orbit. Compare its Δv with the Hohmann transfer only after checking the final-orbit conditions.
- Radius alone: An eccentric orbit can cross the target radius without circularizing. Radial velocity, angular momentum, closest approach, and thrust history help distinguish these trajectories.
Changing the training configuration
The final section configures the following parameters through ModeConfig:
dv_mag: control authority per stepfuel_cost_penalty: cost of using Δvignition_penalty: cost of turning on or changing thrustreward_shaping_scale: strength of dense energy/angular-momentum shapingtraining_timesteps,learning_rate, andent_coef: PPO optimization and exploration behavior
For each comparison, record the configuration, seed, training budget, and success tolerances, then:
train a policy
inspect trajectory, thrust, and Δv
compare against Hohmann
explain the failure mode
change one parameter or design choice
rerun
Notebook note: The final custom experiment has a saved configuration printout that differs from the settings in its setup cell. Rerun the experiment to regenerate consistent output, and set a training seed before comparing parameter choices. The saved discrete and continuous policy reports remain available as examples.