This paper proposes using Signal Temporal Logic (STL) specifications to shape rewards for quadruped locomotion. The specifications encode safety bounds, gait synchronization, command tracking, and actuation limits, while smooth approximations of STL robustness provide dense rewards compatible with Proximal Policy Optimization (PPO). The authors define parameterized templates for three speed regimes: walking-trot, trot, and bound, calibrating parameters from reference rollouts. The method is instantiated on Google’s Barkour quadruped in MuJoCo XLA, with simulator parallelization and domain randomization. The abstract reports tighter velocity tracking and more stable training than handcrafted rewards.
No heat snapshots are available in the last 24 hours.