Emotionally Adaptive Conversational Models for Long-Term Human-Robot Interaction Using Proximal Policy Optimization
Open Access DepositedAs social robots become increasingly integrated into everyday environments, their ability to understand and respond appropriately to human emotions is essential for creating interactions that are not only functional but also emotionally meaningful. This research presents a reinforcement learning-based framework for long-term, emotion-aware dialogue generation, integrating both short- and long-term emotional state tracking to enhance the quality of human-robot interaction over time.To achieve this, the framework employs Proximal Policy Optimization (PPO), a reinforcement learning algorithm that dynamically refines conversational strategies by selecting response types based on emotional cues derived from speech-to-text input. A pre-trained emotion classifier identifies the user’s emotional state, which is used to guide response selection. This enables the robot to maintain emotionally coherent and contextually appropriate conversations across multi-turn interactions. The full system—consisting of real-time transcription, emotion classification, reinforcement learning, and language generation—was deployed on the Pepper humanoid robot. A two-week between-subjects user study (n=11) was conducted, in which participants engaged in five structured daily conversations with either a static or emotionally adaptive version of Pepper. The study evaluated the impact of emotional adaptation on three key outcomes: engagement, stress reduction, and emotional alignment. Quantitative results showed a statistically significant reduction in post-session stress levels for the adaptive group (p = 0.004) compared to the control group. Participants in the adaptive condition also reported higher engagement scores, more natural and helpful interactions, and greater mood improvements after each session. While changes in overall well-being scores were not statistically significant (p = 0.644), upward trends and positive subjective feedback suggest cumulative benefits from emotionally responsive interactions over time. These findings demonstrate the feasibility and effectiveness of reinforcement learning-based emotional adaptation in real-world social robotics. The work contributes to the advancement of emotionally intelligent conversational AI, offering promising applications in mental health support, education, and social companionship. Furthermore, the personalized PPO model and long-term emotion tracking framework provide a strong foundation for future research in adaptive, human-centered dialogue systems.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.