Microsoft ends support for Internet Explorer on June 16, 2022.
We recommend using one of the browsers listed below.

  • Microsoft Edge(Latest version) 
  • Mozilla Firefox(Latest version) 
  • Google Chrome(Latest version) 
  • Apple Safari(Latest version) 

Please contact your browser provider for download and installation instructions.

Open search panel Close search panel Open menu Close menu

September 28, 2026

Information

NTT's 14 papers accepted for Interspeech2026, the world's largest international conference on spoken language processing

14 papers authored by NTT Laboratories have been accepted at Interspeech2026 (the 27th edition of the Interspeech Conference)Open other window , to be held in Sydney, Australia, from September 27 to October 1, 2026. Interspeech is the world’s largest and most comprehensive international conference on the science and technology of spoken language processing that supports speech communication between humans as well as between humans and machines/AI. It covers a broad range of fields, from speech recognition, speech synthesis, and spoken dialogue to phonetics. In addition to the 14 accepted papers, we will be presenting a Tutorial at Interspeech2026.

Abbreviated names of the laboratories:
CS: Communication Science Labs., NTT, Inc.
HI: Human Informatics Labs., NTT, Inc.
CMU: Carnegie Mellon University
BUT: Brno University of Technology
(Affiliations are at the time of submission.)

■Tight Boundary Prediction in Speaker Diarization Using Causal-Anticausal Consistency

Shota Horiguchi (HI), Marc Delcroix (CS), Naohiro Tawara (CS), Takanori Ashihara (HI), Atsushi Ando (HI)

Speaker diarization is the task of estimating who spoke when in an audio recording. To train speaker diarization models, a large amount of conversational data annotated with speaker-wise speech segments is required. However, the strictness of these segment labels has a substantial impact on the model outputs. In this study, we developed a model capable of estimating strict speech segments by automatically correcting less strictly annotated labels and training the model using the corrected labels. This work is expected to contribute to speech recognition and understanding in multi-speaker scenarios, such as meetings and business discussions.

■Who Spoke What When? Evaluating Spoken Language Models for Conversational ASR with Semantic and Overlap-Aware Metrics

Naohiro Tawara (CS), Samuele Cornell (CMU), Alexander Polok (CMU), Marc Delcroix (CS), Lukáš Burget (BUT), Shinji Watanabe (CMU)

Recent advances in Large Audio Language Models (Audio LLMs) have greatly improved speech recognition performance. However, their ability to accurately understand natural multi-party conversations remains unclear. In this study, we defined a semantic error metric and an analysis method for evaluating errors in overlapping speech, and used them to comprehensively evaluate the performance of Audio LLMs in conversational speech recognition. This enables a more comprehensive assessment of Audio LLMs and contributes to the realization of next-generation speech recognition technologies capable of accurately understanding real-world conversations.

■SphereVBx: Spherical Variational Bayes Clustering for Simplified EEND-VC Diarization

Petr Pálka (BUT), Jiangyu Han (BUT), Marc Delcroix (CS), Naohiro Tawara (CS), Lukáš Burget (BUT)

In multi-party conversational speech recognition, such as meeting transcription, speaker diarization plays a key role in determining who spoke and when. In this study, we proposed SphereVBx, a spherical variational Bayesian clustering method that directly clusters speaker representations on a hypersphere. The proposed method achieves comparable or better performance than conventional approaches while eliminating several complex processing steps. This work contributes to the realization of more efficient, scalable, and practical conversational speech recognition systems.

■Multi-Talker ASR Unaffected by Speaker Change Count

Naoki Makishima (HI), Suzuka Yamada (HI), Taiga Yamane (HI), Mana Ihori (HI), Tanaka Tomohiro (HI), Satoshi Suzuki (HI), Shota Orihashi (HI), Ryo Masumura (HI)

We develop an autoregressive model for multi-talker automatic speech recognition and speaker diarization robust to the number of speaker changes. Recent studies on multi-talker ASR handle multiple speakers by recursively estimating a single token sequence formed by concatenating their transcripts with speaker change tokens. However, as this technique heavily relies on context including speaker change tokens, inference performance degrades when the number of speaker changes exceeds that in the training data. To address this limitation, we introduce a novel speaker change token mask in self-attention, enabling the model to estimate token sequences, including speaker change tokens, without relying on the number of speaker changes that have occurred. This study is expected to contribute to multi-talker ASR and speaker diarization for conversational speech, including frequent backchannels and long-form audio such as meeting recordings.

■Unified Audio-Visual Modeling to Recognize Which Face Spoke When and What in Scenarios with On- and Off-Screen Participants

Naoki Makishima (HI), Suzuka Yamada (HI), Taiga Yamane (HI), Mana Ihori (HI), Tanaka Tomohiro (HI), Satoshi Suzuki (HI), Shota Orihashi (HI), Ryo Masumura (HI)

We propose an audio-visual modeling method that recognizes which face spoke when and what from videos with or without visible speakers. The conventional method represents information about which face spoke when and what was spoken as a single serialized token sequence and performs autoregressive prediction. However, it assumes that all speakers are always visible in the video, whereas in real-world environments, some speakers may be occluded or outside the camera’s view. To handle these cases, we introduce a dedicated token to indicate the absence of a speaker in videos and train the model with both on-screean and off-screen speakers. This study is expected to contribute to AI-based audio-visual communication support in situations where visual data of the target speakers is unavailable due to privacy concerns or occlusion.

■Evaluating Large Language Models Abilities for Addressee, Turn-change, and Next Speaker Prediction in Meetings

Ryo Fukuda (CS), Takatomo Kanou (CS), Siddhant Arora (CMU), Marc Delcroix (CS), Naohiro Tawara (CS), Atsunori Ogawa (CS), Yuya Chiba (CS), Atsushi Ando (HI), William Chen (CMU), Shinji Watanabe (CMU)

Advances in large language models (LLMs) have greatly improved the ability of conversational agents to understand and generate natural language, leading to growing interest in systems that can participate in multi-party conversations. In multi-party settings, agents need to understand complex turn-taking using linguistic, audio, and visual cues, but current LLM capabilities remain unclear. In this study, we compared the turn-taking abilities of LLMs and humans by testing whether they could predict who an utterance was directed to and who would speak next. The results showed that LLMs still face challenges in understanding multimodal information. These findings can support the development of AI systems that collaborate with humans in multi-party settings.

■LLM-as-Joiner: Decoupling Alignment from Language Modeling in Label-synchronous ASR

Jaeyoung Lee (HI), Masato Mimura (HI), Tatsuya Kawahara (Kyoto University), Ryo Magoshi (Kyoto University)

Recent studies have explored the use of the extensive linguistic knowledge of large language models (LLMs) for automatic speech recognition. However, conventional approaches require an LLM to process long sequences of speech data and learn how speech corresponds to text, which can increase computational cost. In this study, we propose LLM-as-Joiner, in which a speech recognition model handles the correspondence between speech and text, while the LLM focuses on understanding context and relationships between words. Experiments using English speech data and a multilingual dataset covering five languages demonstrated higher recognition accuracy and greater processing efficiency than conventional approaches, and also improved the performance of lightweight models that run on everyday devices such as smartphones. This result is expected to enable practical multilingual speech-to-text technology that delivers higher accuracy without relying on large-scale computing resources.

■Progressive Alignment Objectives for Aligner-Encoder based ASR

Jaeyoung Lee (HI), Masato Mimura (HI), Takafumi Moriya (HI)

Speech recognition systems need to correctly identify which parts of an audio signal correspond to the characters and words in the resulting text. Aligner-Encoders learn this correspondence within the model, enabling speech to be converted into text with a relatively simple architecture. However, in conventional systems, the correspondence tends to emerge suddenly near the end of processing, making training unstable and reducing accuracy, particularly for long utterances. We propose InterAligner, which also learns this correspondence at intermediate stages and gradually refines it from finer units into the final text. Experiments on English speech data substantially reduced recognition errors, with the largest improvements on extended speech such as meetings and lectures. This result is expected to improve the accuracy of speech recognition services people use every day, such as automatic meeting minutes and video captioning.

■Upcycling Pretrained Transformers into Mixture-of-Experts for Multilingual Speech Recognition

Kentaro Shinayama (HI), Masato Mimura (HI), Kohei Matsuura (HI), Jaeyoung Lee (HI)

Multilingual automatic speech recognition is a technology that enables a single model to recognize speech in multiple languages. However, because the model must learn language-specific pronunciation and writing-system characteristics within limited capacity, its recognition accuracy may be lower than that of separately trained monolingual models. In this study, we transformed a pretrained model into a Mixture-of-Experts architecture, in which multiple specialized processing components are selectively used depending on the language. This approach expands the model’s representational capacity without increasing the computational cost during inference. The results are expected to contribute to the development of highly accurate and efficient speech recognition systems capable of supporting a wide range of languages.

■Non-Autoregressive Minimum Bayes' Risk Decoding for Fast Speech Recognition

Hiroyuki Deguchi (CS), Takatomo Kano (CS), Katsuki Chousa (CS), Marc Delcroix (CS)

In long-form speech recognition, such as generating transcription from meeting audio, processing speed is just as important as recognition performance. To date, speech recognition models have broadly fallen into two types: "autoregressive" models, which are accurate but slow, and "non-autoregressive" models, which are fast but less accurate. In this work, we incorporate minimum Bayes’ risk (MBR) decoding, which improves the accuracy of non-autoregressive models without any additional training. As a result, we achieved up to a 43.1× speedup over autoregressive models while maintaining comparable accuracy. We expect this achievement to enable fast and accurate recognition of long-form audio.

■How does children's pronunciation develop? Capturing syllabic change with children's growth using unsupervised syllable discovery

Koharu Horii (CS), Naohiro Tawara (CS), Atsunori Ogawa (CS), Shoko Araki (CS)

Understanding developmental changes in pronunciation is crucial for advancing language development research and for building highly accurate child automatic speech recognition (ASR) systems. However, conventional ASR-based analyses forcibly map speech onto labels defined by adult pronunciation standards, limiting their ability to accurately capture child-specific and intermediate pronunciations. To address this, we propose a novel analysis approach using an unsupervised syllable discovery model that directly identifies syllable-level acoustic units from speech without predefined labels. By analyzing large-scale speech data from children aged 5 to 15, we automatically and objectively captured the developmental trajectory: an initial expansion of pronunciation diversity, followed by a convergence toward adult-like speech. These results provide a new data-driven framework for speech science and pave the way for highly accurate, stage-specific child ASR systems.

■Scalable Direction-Following TTS via Voice Impression-Guided Pseudo Triplet Construction

Kenichi Fujita (HI), Yusuke Ijima (HI)

We proposed a method for automatically constructing large-scale instruction-annotated speech data by generating natural language performance instructions from differences in voice impressions between speech utterances. Furthermore, by leveraging the resulting dataset, we developed a text-to-speech model capable of controlling speaking style through natural language instructions. This work is expected to contribute to speech synthesis technologies that can flexibly reflect users' expressive intentions.

■MeanVoiceFlow2: Joint Optimization of Mean Flow and Content Encoder for Fast One-Step Zero-Shot Voice Conversion

Takuhiro Kaneko (CS), Hirokazu Kameoka (CS), Kou Tanaka (CS), Yuto Kondo (CS)

In voice conversion, generative AI-based methods such as Flow Matching have attracted attention because they can reproduce high-quality speech while preserving speaker similarity. However, many of these methods require iterative computation to generate high-quality speech, resulting in long processing times. To address this issue, we previously proposed MeanVoiceFlow, a fast model that eliminates iterative computation. However, its computationally heavy process for encoding speech content has remained a bottleneck to further improving inference speed. In this study, we propose MeanVoiceFlow2, a new model that jointly optimizes a lightweight content encoder and a voice conversion model. Experimental results demonstrate that MeanVoiceFlow2 achieves more natural speech quality while maintaining speaker similarity comparable to conventional methods and accelerating inference by approximately nine times. This method is expected to be an important technology for achieving high-quality and fast voice conversion in the future.

■Latency Controllable Speech Enhancement

Hiroshi Sato (HI), Takafumi Moriya (HI), Tsubasa Ochiai (CS), Marc Delcroix (CS)

Speech enhancement reduces noise in speech to improve intelligibility. Although performance generally improves with longer processing time, conventional methods operate at a fixed latency. We propose a model that adapts its performance to the available processing time and dynamically switches latency. Experiments showed that it outperforms models trained for a single latency condition. This technology is expected to enable high-quality speech processing tailored to different applications.

Tutorial

■Conversational Speech Recognition and Analysis: Progress, Challenges, and Emerging Opportunities with Speech Language Model

Naohiro Tawara (CS), Samuele Cornell (CMU), Alexander Polok (CMU), Takafumi Moriya (MI), Marc Delcroix (HI), Lukáš Burget (BUT), Shinji Watanabe (CMU)

Recent advances in Speech Language Models (SLMs) are bringing conversational speech recognition into a new era. At the same time, how to leverage the knowledge accumulated through conventional approaches while advancing next-generation technologies has become an important research question. To address this topic, we have organized a tutorial entitled “Conversational Speech Recognition and Analysis: Progress, Challenges, and Emerging Opportunities with Speech Language Models.” The tutorial provides an overview of recent progress in conversational speech recognition and analysis, examines the opportunities and challenges brought by SLM-based approaches, and discusses promising directions for future research.

Information is current as of the date of issue of the individual topics.
Please be advised that information may be outdated after that point.