PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 9, 2025Frontiers in Robotics and AI3 citationsOpen Access

Exploring multimodal collaborative storytelling with Pepper: a preliminary study with zero-shot LLMs

View Full Paper
UZUnai ZabalaJEJuan EchevarríaIRIgor Rodríguez

Key Points

  • The multimodal collaborative storytelling system integrates user interaction and object recognition with Pepper.
  • Preliminary findings indicate high user acceptance of the storytelling system utilizing zero-shot large language models.
  • The robot performs storytelling using expressive gestures and speech modulation, enhancing user immersion.
  • Usability of the Llama model in interactive storytelling was assessed through user feedback collected via questionnaires.

Abstract

With the rise of large language models (LLMs), collaborative storytelling in virtual agents or chatbots has gained popularity. Despite storytelling has long been employed in social robotics as a means to educate, entertain, and persuade audiences, the integration of LLMs into such platforms remains largely unexplored. This paper presents the initial steps for a novel multimodal collaborative storytelling system in which users co-create stories with the social robot Pepper through natural language interaction and by presenting physical objects. The robot employs a YOLO-based vision system to recognize these objects and seamlessly incorporate them into the narrative. Story generation and adaptation are handled autonomously using the Llama model in a zero-shot setting, aiming to assess the usability and maturity of such models in interactive storytelling. To enhance immersion, the robot performs the final story using expressive gestures, emotional cues, and speech modulation. User feedback, collected through questionnaires and semi-structured interviews, indicates a high level of acceptance.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zabala et al. (2025) studied this question.

synapsesocial.com/papers/68e7f0af2d7e30942762c8fehttps://doi.org/10.3389/frobt.2025.1662819
Ask AI
Helpful
Bookmark
Share
View Full Paper