Microsoft ends support for Internet Explorer on June 16, 2022.
We recommend using one of the browsers listed below.
Please contact your browser provider for download and installation instructions.
September 11, 2026
Information
Four papers authored by NTT Laboratories have been accepted at ECCV2026 (The 19th European Conference on Computer Vision), to be held in Malmö, Sweden, from September 8 to 12, 2026. This is a flagship conference on computer vision and patter recognition where researchers seek the computational understanding, control, and generation of image and videos as well as their foundation theories.
Abbreviated names of the laboratories:
HI: Human Informatics Labs., NTT, Inc.
CD: Computer and Data Science Labs., NTT, Inc.
CS: Communication Science Labs., NTT, Inc.
(Affiliations are at the time of submission.)
Ryota Tanaka (HI), Taku Hasegawa (HI), Kyosuke Nishida (HI)
We propose CMDR, a multimodal document retrieval task that models context across multiple pages. We introduce the novel concept of Indirect Retrieval, which requires understanding document-level structure and cross-page dependencies beyond the capabilities of conventional page-level retrieval. To facilitate its evaluation, we develop CMDR-Bench, a large-scale benchmark for Indirect Retrieval. Furthermore, we propose CMDR-Embed, which jointly embeds multiple pages, and CMCL, a context-aware contrastive learning framework, achieving substantial improvements in retrieval performance. Our work advances information retrieval for long, diverse, and real-world documents.
Taiga Yamane (HI), Satoshi Suzuki (HI), Ryo Masumura (HI), Shota Orihashi (HI), Tomohiro Tanaka (HI), Mana Ihori (HI), Naoki Makishima (HI)
Multi-view pedestrian detection aims to detect pedestrians in a bird's-eye view from multi-view images. In this task, previous methods struggle to generalize to unseen camera configurations due to a lack of ability to capture 2D-3D correspondences and distortions caused by a perspective transformation when projecting image features into 3D space. To overcome this problem, we leverage a visual geometric foundation model. This model can accurately capture the 2D-3D correspondences for various camera configurations. Furthermore, this model can represent where each pixel in images is located in 3D space, enabling the projection of image features into 3D space, while avoiding distortions. By leveraging the visual geometric foundation model, our proposed MV2GF achieved significantly better generalization performance for unseen camera configurations than previous methods.
Yoko Sogabe (CD), Shiori Sugimoto (CD), Shoichiro Saito (CD), Masaki Kitahara (CD)
Diffuser-based lensless imaging uses thin optical elements instead of conventional lenses to realize compact and lightweight cameras. These cameras support attractive applications, including 3D imaging, fluorescence microscopy, high-frame-rate video imaging, and privacy-aware sensing. Because these systems do not form a direct image on the sensor, the captured measurement records the scene as a complex, multiplexed coded pattern. Computational reconstruction is therefore essential, and its accuracy depends critically on the forward model, which describes the optical process from the scene to the sensor.
In real optical systems, however, the spread of light varies between the center and the periphery of the field of view. Simplified forward models commonly used in conventional methods cannot fully capture such spatially varying aberrations and geometric distortions, leading to blur and reconstruction artifacts. In this work, we propose a method that learns the mismatch between the assumed and actual optical models as a small number of physically interpretable correction kernels and integrates them into a deep unrolled network. Experiments using real-world data demonstrate that the proposed method achieves higher reconstruction accuracy than existing methods while preserving the interpretability and computational efficiency of the physical model. These results are expected to contribute not only to diffuser-based lensless cameras but also to a broader range of computational imaging technologies using thin optical elements, including DOE lenses and metalenses.
Yusuke Oumi (Keio Univ.), Yuto Shibata (Keio Univ.), Go Irie (Keio University, Tokyo Science University), Akisato Kimura (CS), Yoshimitsu Aoki (Keio University), Mariko Isogawa (Keio University)
This paper presents the first attempt to estimate multi-person 3D poses solely from acoustic signals. Estimating the poses of multiple individuals using acoustic signals is inherently challenging due to the superposition of motion-dependent signal variations. Unlike single-person scenarios, the presence of multiple subjects leads to overlapping acoustic signatures, making it difficult to attribute specific signal changes to an individual's pose. Furthermore, the complexity is compounded by interperson reflections, which introduce intricate propagation delays that obscure the temporal motion-acoustic relationship. To address these issues, we propose a novel encoder-decoder framework consisting of two key components. First, the Acoustic Multi-scale Encoder captures diverse temporal and fine-grained frequency features to isolate subtle acoustic signatures from complex, overlapping signals. Second, the Temporal Pose Decoder employs an attention mechanism to disentangle multi-person information across successive frames. By jointly accounting for temporal dynamics and inter-person dependencies, this component precisely reconstructs frame-wise individual poses. To validate our approach, we constructed the 6-hour Acoustic Multi-person Pose (AMP) dataset consisting of 432K synchronized frames of multi-person pose and acoustic data, and demonstrated that our proposed method outperforms baseline models.
Information is current as of the date of issue of the individual topics.
Please be advised that information may be outdated after that point.
WEB media that thinks about the future with NTT