Esam Ghaleb receives grant to develop gesture and interaction-aware multimodal generative AI
“I’m especially excited because this is my first grant as a Principal Investigator,” Esam says. “It was awarded from a single, highly competitive pool across the natural and exact sciences. It gives us the freedom to pursue a high-risk, high-reward idea - and to build something that I hope will lead to many more projects in the future.”
Shaping understanding
The project is inspired by a simple yet powerful observation: human communication is inherently multimodal. We don’t just speak - we gesture, move, and react to each other in ways that shape understanding.
Prior research has shown that when a speaker accompanies a question with a gesture that matches their speech, listeners respond more quickly and with fewer misunderstandings. Yet, today’s interactive language models rely heavily on speech and images, often missing the embodied, dynamic aspects of human conversation. This project aims to bridge that gap by illustrating and testing the communication efficiency of a multimodal generative AI system.
“The core idea,” Esam explains, “is to build an AI system that generates gestures from speech, tailored to the objects being discussed. At the same time, the model interacts with a human partner, reads their intentions from multimodal cues - such as speech and gesture - and responds with contextually appropriate gestures that align with the partner’s behavior.”
More human-like
Unlike many models that rely solely on textual input, this system grounds its output in the full set of conversational ingredients: what is intentionally said, what is gestured, and who the interlocutor is. The result is a more human-like, embodied mode of interaction, with potential applications in assistive and interactive technologies that communicate more naturally.
Beyond its engineering goals, the project will also provide a platform for testing fundamental scientific questions. “This gives us the opportunity to examine how gesture and speech align in both production and comprehension, especially during face-to-face interaction,” says Esam. “It will help us identify which components are essential for efficient communication in multimodal systems.”
Thanks to the theoretical and computational expertise of the Multimodal Language Department at the Max Planck Institute for Psycholinguistics - combined with motion-capture suites, virtual reality labs, and a high-performance computing cluster - the team is well positioned to bring this ambitious vision to life.
Further read: Open Competition ENW | NWO
Share this page