Microsoft ends support for Internet Explorer on June 16, 2022.
We recommend using one of the browsers listed below.
Please contact your browser provider for download and installation instructions.
October 7, 2026
Information
From October 5 to 9, 2026, the 28th ACM International Conference on Multimodal Interaction (ICMI 2026), a top international conference in the field of multimodal interaction, will be held in Naples, Italy. Three papers from NTT Laboratories have been accepted to the main conference. ICMI is a leading international conference on multimodal interaction, which focuses on the integrated processing and understanding of multiple types of information (modalities), such as speech, language, facial expressions, gaze, and body movements, involved in human-human and human-AI/robot communication. ICMI 2026 will present cutting-edge research on topics including multimodal dialogue understanding and technologies for enabling natural interactions between humans and AI systems or robots.
Abbreviated names of the laboratories:
HI:Human Informatics Labs., NTT, Inc.
Ryo Ishii (HI), Chihiro Takayama (HI), Jiro Nagao (HI), Toshiki Onishi (HI), Yukiko I. Nakano (Seikei University), Junichi Sawase (HI)
For AI to understand multiparty conversations, it is important to capture not only what is being said, but also information such as vocal and facial cues, interactions among participants, and each participant’s communication behavior. In this study, we proposed a multi-context model that separately captures interaction dynamics across the entire conversation and the behavior of individual participants. We further demonstrated that the effectiveness of “self-distillation,” which uses the model’s own predictions as learning signals, depends on how dialogue information is represented, and that appropriately combining these approaches achieves high performance in dialogue understanding.
Ryo Ishii (HI), Shinichiro Eitoku (HI), Jiro Nagao (HI), Junichi Sawase (HI)
For AI agents and robots that interact with humans, it is important to generate body movements synchronized with speech without delay. In this study, we identified that conventional gesture generation models implicitly require future information during internal processing, which contributes to latency. To address this issue, we proposed a method that sequentially generates movements using only past and current information. Our evaluation confirmed that the proposed method substantially reduces processing latency while maintaining gesture quality and improves the responsiveness perceived by users.
One other paper was accepted.
Information is current as of the date of issue of the individual topics.
Please be advised that information may be outdated after that point.
WEB media that thinks about the future with NTT