Social robot navigation in crowded environments requires balancing task efficiency with socially acceptable behavior. Reinforcement Learning (RL) has emerged as a promising approach for addressing this challenge. A critical issue in RL-based navigation techniques is the design of the reward function, which combines multiple, often conflicting, objectives such as task completion, efficiency, obstacle avoidance, and smooth human–robot interaction. This chapter presents a detailed analysis of the impact of reward function design on the performance and social compliance of RL-based navigation policies. Focusing on a social attentive RL technique, this work evaluates how individual reward terms influence navigation performance and how these interact with the usage of different human motion models. Specifically, we analyze the impact of using the Social Force Model and the Headed Social Force Model to simulate pedestrian behavior, comparing their suitability for training RL-based policies. Simulation results provide useful insights for the definition of reward functions, so as to achieve efficient and socially acceptable robot behavior in crowded environments.

Van Der Meer, T., Garulli, A., Giannitrapani, A., Quartullo, R. (2026). Reward Shaping in Learning-Based Social Navigation Systems. In G. Franzè, G. Fortino, W. Lucia, M.C. Zhou (a cura di), Constrained Control and Machine Learning: Emerging Methodologies and Applications (pp. 271-290). Springer Science and Business Media Deutschland GmbH [10.1007/978-3-032-02709-2_12].

Reward Shaping in Learning-Based Social Navigation Systems

Tommaso Van Der Meer
;
Andrea Garulli;Antonio Giannitrapani;
2026-01-01

Abstract

Social robot navigation in crowded environments requires balancing task efficiency with socially acceptable behavior. Reinforcement Learning (RL) has emerged as a promising approach for addressing this challenge. A critical issue in RL-based navigation techniques is the design of the reward function, which combines multiple, often conflicting, objectives such as task completion, efficiency, obstacle avoidance, and smooth human–robot interaction. This chapter presents a detailed analysis of the impact of reward function design on the performance and social compliance of RL-based navigation policies. Focusing on a social attentive RL technique, this work evaluates how individual reward terms influence navigation performance and how these interact with the usage of different human motion models. Specifically, we analyze the impact of using the Social Force Model and the Headed Social Force Model to simulate pedestrian behavior, comparing their suitability for training RL-based policies. Simulation results provide useful insights for the definition of reward functions, so as to achieve efficient and socially acceptable robot behavior in crowded environments.
2026
9783032027085
9783032027092
Van Der Meer, T., Garulli, A., Giannitrapani, A., Quartullo, R. (2026). Reward Shaping in Learning-Based Social Navigation Systems. In G. Franzè, G. Fortino, W. Lucia, M.C. Zhou (a cura di), Constrained Control and Machine Learning: Emerging Methodologies and Applications (pp. 271-290). Springer Science and Business Media Deutschland GmbH [10.1007/978-3-032-02709-2_12].
File in questo prodotto:
File Dimensione Formato  
Springer_book_chapter.pdf

non disponiibile

Tipologia: PDF editoriale
Licenza: NON PUBBLICO - Accesso privato/ristretto
Dimensione 339.04 kB
Formato Adobe PDF
339.04 kB Adobe PDF   Visualizza/Apri   Richiedi una copia

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11365/1325834