NTT Multi-Speaker ASR System for the DASR Task of CHiME-8 Challenge
Fuente:
arXiv
Saved in:
| Main Authors: | Kamo, Naoyuki, Tawara, Naohiro, Ando, Atsushi, Kano, Takatomo, Sato, Hiroshi, Ikeshita, Rintaro, Moriya, Takafumi, Horiguchi, Shota, Matsuura, Kohei, Ogawa, Atsunori, Plaquet, Alexis, Ashihara, Takanori, Ochiai, Tsubasa, Mimura, Masato, Delcroix, Marc, Nakatani, Tomohiro, Asami, Taichi, Araki, Shoko |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Microphone Array Geometry Independent Multi-Talker Distant ASR: NTT System for the DASR Task of the CHiME-8 Challenge
by: Kamo, Naoyuki, et al.
Published: (2025)
by: Kamo, Naoyuki, et al.
Published: (2025)
Applying LLMs for Rescoring N-best ASR Hypotheses of Casual Conversations: Effects of Domain Adaptation and Context Carry-over
by: Ogawa, Atsunori, et al.
Published: (2024)
by: Ogawa, Atsunori, et al.
Published: (2024)
Mamba-based Segmentation Model for Speaker Diarization
by: Plaquet, Alexis, et al.
Published: (2024)
by: Plaquet, Alexis, et al.
Published: (2024)
Interaural time difference loss for binaural target sound extraction
by: Hernandez-Olivan, Carlos, et al.
Published: (2024)
by: Hernandez-Olivan, Carlos, et al.
Published: (2024)
MOVER: Combining Multiple Meeting Recognition Systems
by: Kamo, Naoyuki, et al.
Published: (2025)
by: Kamo, Naoyuki, et al.
Published: (2025)
Sentence-wise Speech Summarization: Task, Datasets, and End-to-End Modeling with LM Knowledge Distillation
by: Matsuura, Kohei, et al.
Published: (2024)
by: Matsuura, Kohei, et al.
Published: (2024)
Dissecting the Segmentation Model of End-to-End Diarization with Vector Clustering
by: Plaquet, Alexis, et al.
Published: (2025)
by: Plaquet, Alexis, et al.
Published: (2025)
SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model
by: Hernandez-Olivan, Carlos, et al.
Published: (2024)
by: Hernandez-Olivan, Carlos, et al.
Published: (2024)
Rethinking Processing Distortions: Disentangling the Impact of Speech Enhancement Errors on Speech Recognition Performance
by: Ochiai, Tsubasa, et al.
Published: (2024)
by: Ochiai, Tsubasa, et al.
Published: (2024)
Reference Microphone Selection for Guided Source Separation based on the Normalized L-p Norm
by: Lohmann, Anselm, et al.
Published: (2025)
by: Lohmann, Anselm, et al.
Published: (2025)
Can We Really Repurpose Multi-Speaker ASR Corpus for Speaker Diarization?
by: Horiguchi, Shota, et al.
Published: (2025)
by: Horiguchi, Shota, et al.
Published: (2025)
Mitigating Non-Target Speaker Bias in Guided Speaker Embedding
by: Horiguchi, Shota, et al.
Published: (2025)
by: Horiguchi, Shota, et al.
Published: (2025)
Recursive Attentive Pooling for Extracting Speaker Embeddings from Multi-Speaker Recordings
by: Horiguchi, Shota, et al.
Published: (2024)
by: Horiguchi, Shota, et al.
Published: (2024)
Guided Speaker Embedding
by: Horiguchi, Shota, et al.
Published: (2024)
by: Horiguchi, Shota, et al.
Published: (2024)
Frontend Token Enhancement for Token-Based Speech Recognition
by: Ashihara, Takanori, et al.
Published: (2026)
by: Ashihara, Takanori, et al.
Published: (2026)
Array Geometry-Robust Attention-Based Neural Beamformer for Moving Speakers
by: Tammen, Marvin, et al.
Published: (2024)
by: Tammen, Marvin, et al.
Published: (2024)
SpeakerBeam-SS: Real-time Target Speaker Extraction with Lightweight Conv-TasNet and State Space Modeling
by: Sato, Hiroshi, et al.
Published: (2024)
by: Sato, Hiroshi, et al.
Published: (2024)
TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models
by: Peng, Junyi, et al.
Published: (2025)
by: Peng, Junyi, et al.
Published: (2025)
Probing Self-supervised Learning Models with Target Speech Extraction
by: Peng, Junyi, et al.
Published: (2024)
by: Peng, Junyi, et al.
Published: (2024)
The CHiME-8 DASR Challenge for Generalizable and Array Agnostic Distant Automatic Speech Recognition and Diarization
by: Cornell, Samuele, et al.
Published: (2024)
by: Cornell, Samuele, et al.
Published: (2024)
Investigation of Speaker Representation for Target-Speaker Speech Processing
by: Ashihara, Takanori, et al.
Published: (2024)
by: Ashihara, Takanori, et al.
Published: (2024)
Recent Trends in Distant Conversational Speech Recognition: A Review of CHiME-7 and 8 DASR Challenges
by: Cornell, Samuele, et al.
Published: (2025)
by: Cornell, Samuele, et al.
Published: (2025)
Pretraining Multi-Speaker Identification for Neural Speaker Diarization
by: Horiguchi, Shota, et al.
Published: (2025)
by: Horiguchi, Shota, et al.
Published: (2025)
Generic Speech Enhancement with Self-Supervised Representation Space Loss
by: Sato, Hiroshi, et al.
Published: (2025)
by: Sato, Hiroshi, et al.
Published: (2025)
Target Speech Extraction with Pre-trained Self-supervised Learning Models
by: Peng, Junyi, et al.
Published: (2024)
by: Peng, Junyi, et al.
Published: (2024)
Alignment-Free Training for Transducer-based Multi-Talker ASR
by: Moriya, Takafumi, et al.
Published: (2024)
by: Moriya, Takafumi, et al.
Published: (2024)
STCON System for the CHiME-8 Challenge
by: Mitrofanov, Anton, et al.
Published: (2024)
by: Mitrofanov, Anton, et al.
Published: (2024)
Microphone Array Signal Processing and Deep Learning for Speech Enhancement
by: Haeb-Umbach, Reinhold, et al.
Published: (2025)
by: Haeb-Umbach, Reinhold, et al.
Published: (2025)
What Do Self-Supervised Speech and Speaker Models Learn? New Findings From a Cross Model Layer-Wise Analysis
by: Ashihara, Takanori, et al.
Published: (2024)
by: Ashihara, Takanori, et al.
Published: (2024)
BUT System Description for CHiME-9 MCoRec Challenge
by: Klement, Dominik, et al.
Published: (2026)
by: Klement, Dominik, et al.
Published: (2026)
The USTC-NERCSLIP Systems for the CHiME-8 MMCSG Challenge
by: Jiang, Ya, et al.
Published: (2024)
by: Jiang, Ya, et al.
Published: (2024)
The USTC-NERCSLIP Systems for the CHiME-9 MCoRec Challenge
by: Jiang, Ya, et al.
Published: (2026)
by: Jiang, Ya, et al.
Published: (2026)
The USTC-NERCSLIP Systems for the CHiME-8 NOTSOFAR-1 Challenge
by: Niu, Shutong, et al.
Published: (2024)
by: Niu, Shutong, et al.
Published: (2024)
Boosting Hybrid Autoregressive Transducer-based ASR with Internal Acoustic Model Training and Dual Blank Thresholding
by: Moriya, Takafumi, et al.
Published: (2024)
by: Moriya, Takafumi, et al.
Published: (2024)
The CHiME-7 UDASE task: Unsupervised domain adaptation for conversational speech enhancement
by: Leglaive, Simon, et al.
Published: (2023)
by: Leglaive, Simon, et al.
Published: (2023)
VBx for End-to-End Neural and Clustering-based Diarization
by: Pálka, Petr, et al.
Published: (2025)
by: Pálka, Petr, et al.
Published: (2025)
Joint Enhancement and Classification using Coupled Diffusion Models of Signals and Logits
by: Nurko, Gilad, et al.
Published: (2026)
by: Nurko, Gilad, et al.
Published: (2026)
SpeechGLUE: How Well Can Self-Supervised Speech Models Capture Linguistic Knowledge?
by: Ashihara, Takanori, et al.
Published: (2023)
by: Ashihara, Takanori, et al.
Published: (2023)
Objective and subjective evaluation of speech enhancement methods in the UDASE task of the 7th CHiME challenge
by: Leglaive, Simon, et al.
Published: (2024)
by: Leglaive, Simon, et al.
Published: (2024)
Who Spoke What When? Evaluating Spoken Language Models for Conversational ASR with Semantic and Overlap-Aware Metrics
by: Tawara, Naohiro, et al.
Published: (2026)
by: Tawara, Naohiro, et al.
Published: (2026)
Similar Items
-
Microphone Array Geometry Independent Multi-Talker Distant ASR: NTT System for the DASR Task of the CHiME-8 Challenge
by: Kamo, Naoyuki, et al.
Published: (2025) -
Applying LLMs for Rescoring N-best ASR Hypotheses of Casual Conversations: Effects of Domain Adaptation and Context Carry-over
by: Ogawa, Atsunori, et al.
Published: (2024) -
Mamba-based Segmentation Model for Speaker Diarization
by: Plaquet, Alexis, et al.
Published: (2024) -
Interaural time difference loss for binaural target sound extraction
by: Hernandez-Olivan, Carlos, et al.
Published: (2024) -
MOVER: Combining Multiple Meeting Recognition Systems
by: Kamo, Naoyuki, et al.
Published: (2025)