Abstract
The rigid formation control of Wheeled Mobile Robots (WMRs) in a cluttered environment is presented in this paper using a residual Reinforcement Learning (RL) framework. Unicycle robots form three shapes (a triangle, a square, and an X) and coordinate them simultaneously to different goals. A structured guidance controller is supplemented with a learned corrective policy, and the geometry of the rigid template is enforced with a Kabsch-based geometric projection, ensuring inter-agent and inter-formation collision avoidance is applied throughout. This framework is used to implement two state-of-the-art continuous-control algorithms (DDPG and TD3) and compare their performance through controlled multi-episode evaluation and component ablations under nominal and stochastic disturbance conditions. Results demonstrate that TD3 out-performs DDPG in terms of consistency and performance of the learned policies when the robot operates under nominal conditions, that the residual structure significantly improves the safety, formation accuracy and the training stability, and that the two algorithms show complementary behavior when they are disturbed. The novelty of the work is the first use of Kabsch-based rigid-template enforcement in a residual Deep RL architecture for simultaneous multi-formation navigation, where the usefulness of principled hybrid architectures that integrate classical control structure with modern RL for cooperative Multi-Robot Systems (MRS) is highlighted.