Double Pendulum Swing-Up

← Control Systems

Two poles hang from a cart on a two-metre rail. Push the cart to swing both poles up and balance them. That's hard by hand, so a TQC reinforcement-learning policy is also available. It runs in your browser, and one network handles every target: both poles up, both down, or either one folded over the other. Pick a target and it moves the poles there from wherever they are.

Controller

Manual: drag the cart. TQC: the network drives to the target. Picking a target hands it control.

Swinging both poles up by hand is very likely impossible, but you can try Manual.

Disturbances

Delay holds back the force in both modes. Noise corrupts what the network measures (position, angles and both rates), never the force. The network was trained without either, so it is fragile: even 50 ms of delay usually breaks it, and noise above about 0.03 does too.

- Time to target (s)
Best Manual- TQC-

About this demo

The policy was trained in n-cartpole with TQC (Truncated Quantile Critics), an off-policy actor-critic that learns a whole distribution of returns rather than a single value. At every 10 ms step it sees the cart position and velocity, cos θ, sin θ and θ̇ for each pole, and the target as cos of each pole's goal angle (+1 up, −1 down). It outputs one horizontal force. Its reward peaks only when both poles sit at their targets, still, with the cart near the centre. During training the target changed every few seconds, and episodes started at every pose as well as at fully random states. That's why it can move between any two targets, and why it can catch a swing you started by hand.

Labels give one letter per pole, base pole (θ₁) first: U is up, D is down. Only DD is stable on its own. The other three have to be balanced actively, and in UD and DU one pole is folded back over the other.

"Reached" means both poles within 0.3 rad of their targets and slower than 1.5 rad/s for half a second. That's the same test the training repo's evaluation uses. The clock restarts at every reset and every new target. A time counts as TQC only if you never pushed the cart yourself, and as manual only if the network never drove. Hitting either end of the rail ends the run. The network trained on a ±0.5 m rail; this one runs to ±1 m so there's room to swing by hand, but past half a metre from centre the network is outside anything it saw in training.

The simulation uses the exact plant the network was trained on, except the rail.