Reinforcement Learning for Orbital Transfers at the 2026 AI Winter School (Brown University)

Training PPO policies for orbital transfers and comparing their trajectories and delta-v with a Hohmann baseline.

2026 AI Winter School banner — Brown University Department of Physics, Center for the Fundamental Physics of the Universe, January 6–9, 2026

At the 2026 AI Winter School, hosted by the Center for the Fundamental Physics of the Universe at Brown University, I led a 2.5-hour hands-on workshop on reinforcement learning for orbital transfers.

I used a two-body transfer with a known analytic solution so we could compare learned policies with the Hohmann baseline. The notebook contains the environment, training procedure, and example outputs.

Control Problem

The notebook used nondimensional two-body dynamics: unit gravitational parameter, an initial circular orbit at radius 1, and a target circular orbit at radius 1.6. I call those radii r1 and r2 below. The model omitted drag, finite-duration thrust, J2 perturbations, third bodies, attitude dynamics, and mass depletion. Control was a tangential impulse applied once per simulation step.

Between impulses, the spacecraft position and velocity evolve under central gravity:

dr→ dt = v→ dv→ dt = - μr→ r3 r = |r→| \begin{aligned}\frac{d\vec{r}}{dt} &= \vec{v} \\ \frac{d\vec{v}}{dt} &= -\frac{\mu\vec{r}}{r^3} \\ r &= |\vec{r}|\end{aligned}

For a circular target orbit, the target specific energy and angular momentum are:

E* = - μ 2r2 L* = μr2 \begin{aligned}E^* &= -\frac{\mu}{2r_2} \\ L^* &= \sqrt{\mu r_2}\end{aligned}

The observation vector contains normalized radius, radial velocity, tangential velocity, angular-momentum error, energy error, and the previous action. It omits the absolute orbital angle, so rotating an otherwise identical state does not change the policy input.

Hohmann Benchmark

Before training PPO, the notebook computed the Hohmann transfer. For circular, coplanar orbits with two impulsive burns, the transfer semi-major axis is:

aT = r1+r2 2 a_T = \frac{r_1 + r_2}{2}

The two burns and transfer time are:

Δv1 = μ ( 2r1 - 1aT ) - μr1 Δv2 = μr2 - μ ( 2r2 - 1aT ) T = π aT3 μ \begin{aligned}\Delta v_1 &= \sqrt{\mu\left(\frac{2}{r_1} - \frac{1}{a_T}\right)} - \sqrt{\frac{\mu}{r_1}} \\ \Delta v_2 &= \sqrt{\frac{\mu}{r_2}} - \sqrt{\mu\left(\frac{2}{r_2} - \frac{1}{a_T}\right)} \\ T &= \pi\sqrt{\frac{a_T^3}{\mu}}\end{aligned}

The comparisons use total Δv, circularization error, and burn history.

Hohmann transfer trajectory: the transfer ellipse touching the inner start orbit and the outer target orbit

The transfer ellipse touches the initial circular orbit at r1 and the target orbit at r2.

Verification plots for the simulated Hohmann transfer: radius versus time rising to the target, two thrust impulses showing the burn-coast-burn structure, and cumulative delta-v matching the ideal total

Radius history, the two impulses, and accumulated Δv for the simulated Hohmann transfer.

The second Hohmann burn is applied on the first simulation step at or after the analytic transfer time. With a timestep of 0.00050, the saved baseline ends at radius 1.5999 rather than 1.6000, and total Δv agrees with the analytic value to the displayed precision of 0.2066.

RL Formulation

The Gymnasium environment held the physics fixed and varied the control interface:

Component Implementation
Observation Normalized r, vr, vt, target angular-momentum error, target energy error, previous action
Discrete action coast, full prograde impulse, full retrograde impulse
Continuous action throttle in [-1, 1], mapped to a signed tangential Δv impulse
Success criteria tolerances on |r-r2|, |vr|, and |L-L*|
Failure criteria crash/escape radius or episode timeout

The dense reward used a combined energy/angular-momentum error:

err = |E-E*| |E*| + |L-L*| |L*| \mathrm{err} = \frac{|E - E^*|}{|E^*|} + \frac{|L - L^*|}{|L^*|}

with shaping approximately proportional to:

errprevious - γ errcurrent \mathrm{err}_{previous} - \gamma\,\mathrm{err}_{current}

The reward also included fuel and ignition/switching penalties, a one-time bonus on first entry into the success region, and a reward for remaining there. PPO training used observation and reward normalization. Evaluation used frozen normalization statistics and deterministic policy actions.

Failure modes to inspect

For each policy, the notebook plots the trajectory, radius, radial velocity, thrust impulses, and cumulative Δv. Its mission report includes the burn count or active-thrust steps and closest-to-target statistics for comparison with the Hohmann transfer.

These diagnostics can reveal several failure modes:

  • Repeated corrections: A discrete policy may reach the target through many small prograde and retrograde impulses, satisfying the orbit tolerances while using excessive Δv.
  • Continuous small impulses: A continuous policy may keep orbital errors small by applying throttle almost continuously. The accumulated Δv reveals the cost of these corrections.
  • Loose success tolerances: A policy can enter the success region while remaining on an elliptical orbit. Compare its Δv with the Hohmann transfer only after checking the final-orbit conditions.
  • Radius alone: An eccentric orbit can cross the target radius without circularizing. Radial velocity, angular momentum, closest approach, and thrust history help distinguish these trajectories.

Changing the training configuration

The final section configures the following parameters through ModeConfig:

  • dv_mag: control authority per step
  • fuel_cost_penalty: cost of using Δv
  • ignition_penalty: cost of turning on or changing thrust
  • reward_shaping_scale: strength of dense energy/angular-momentum shaping
  • training_timesteps, learning_rate, and ent_coef: PPO optimization and exploration behavior

For each comparison, record the configuration, seed, training budget, and success tolerances, then:

train a policy
inspect trajectory, thrust, and Δv
compare against Hohmann
explain the failure mode
change one parameter or design choice
rerun

Notebook note: The final custom experiment has a saved configuration printout that differs from the settings in its setup cell. Rerun the experiment to regenerate consistent output, and set a training seed before comparing parameter choices. The saved discrete and continuous policy reports remain available as examples.

Materials