H2INT: Human-Human & Human-Robot Interaction Transformer for Dense Crowd Navigation
Original title:Human-Human & Human-Robot Interaction Transformer (H2INT) for Robot Navigation in Dense and Uncertain Crowds
Navigation in crowded environments often falters because algorithms assume uniform pedestrian reciprocity. H2INT addresses this interaction uncertainty through a reinforcement learning framework powered by a two-stage gated Transformer and recurrent policy. Instead of relying on explicit responsiveness labels, the agent infers pedestrian cooperation directly from relative positions. Validated in simulation curricula with decreasing responsiveness and deployed on physical hardware, the approach sustains safe trajectories across varying crowd densities without retraining.
Why it's worth reading
It moves past the idealized assumption of uniform pedestrian compliance, offering a practical relational reinforcement learning approach verified on physical robot hardware.