Social robot navigation in crowded environments requires balancing task efficiency with socially acceptable behavior. Reinforcement Learning (RL) has emerged as a promising approach for addressing this challenge. A critical issue in RL-based navigation techniques is the design of the reward function, which combines multiple, often conflicting, objectives such as task completion, efficiency, obstacle avoidance, and smooth human–robot interaction. This chapter presents a detailed analysis of the impact of reward function design on the performance and social compliance of RL-based navigation policies. Focusing on a social attentive RL technique, this work evaluates how individual reward terms influence navigation performance and how these interact with the usage of different human motion models. Specifically, we analyze the impact of using the Social Force Model and the Headed Social Force Model to simulate pedestrian behavior, comparing their suitability for training RL-based policies. Simulation results provide useful insights for the definition of reward functions, so as to achieve efficient and socially acceptable robot behavior in crowded environments.
Van Der Meer, T., Garulli, A., Giannitrapani, A., Quartullo, R. (2026). Reward Shaping in Learning-Based Social Navigation Systems. In Constrained Control and Machine Learning: Emerging Methodologies and Applications (pp. 271-290). Springer Science and Business Media Deutschland GmbH [10.1007/978-3-032-02709-2_12].
Reward Shaping in Learning-Based Social Navigation Systems
Tommaso Van Der Meer
;Andrea Garulli;Antonio Giannitrapani;Renato Quartullo
2026-01-01
Abstract
Social robot navigation in crowded environments requires balancing task efficiency with socially acceptable behavior. Reinforcement Learning (RL) has emerged as a promising approach for addressing this challenge. A critical issue in RL-based navigation techniques is the design of the reward function, which combines multiple, often conflicting, objectives such as task completion, efficiency, obstacle avoidance, and smooth human–robot interaction. This chapter presents a detailed analysis of the impact of reward function design on the performance and social compliance of RL-based navigation policies. Focusing on a social attentive RL technique, this work evaluates how individual reward terms influence navigation performance and how these interact with the usage of different human motion models. Specifically, we analyze the impact of using the Social Force Model and the Headed Social Force Model to simulate pedestrian behavior, comparing their suitability for training RL-based policies. Simulation results provide useful insights for the definition of reward functions, so as to achieve efficient and socially acceptable robot behavior in crowded environments.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/11365/1325834
Attenzione
Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo
