Sentence-wise Speech Summarization: Task, Datasets, and End-to-End Modeling with LM Knowledge Distillation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Matsuura, Kohei, Ashihara, Takanori, Moriya, Takafumi, Mimura, Masato, Kano, Takatomo, Ogawa, Atsunori, Delcroix, Marc |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Applying LLMs for Rescoring N-best ASR Hypotheses of Casual Conversations: Effects of Domain Adaptation and Context Carry-over
von: Ogawa, Atsunori, et al.
Veröffentlicht: (2024)
von: Ogawa, Atsunori, et al.
Veröffentlicht: (2024)
NTT Multi-Speaker ASR System for the DASR Task of CHiME-8 Challenge
von: Kamo, Naoyuki, et al.
Veröffentlicht: (2024)
von: Kamo, Naoyuki, et al.
Veröffentlicht: (2024)
Alignment-Free Training for Transducer-based Multi-Talker ASR
von: Moriya, Takafumi, et al.
Veröffentlicht: (2024)
von: Moriya, Takafumi, et al.
Veröffentlicht: (2024)
What Do Self-Supervised Speech and Speaker Models Learn? New Findings From a Cross Model Layer-Wise Analysis
von: Ashihara, Takanori, et al.
Veröffentlicht: (2024)
von: Ashihara, Takanori, et al.
Veröffentlicht: (2024)
Microphone Array Geometry Independent Multi-Talker Distant ASR: NTT System for the DASR Task of the CHiME-8 Challenge
von: Kamo, Naoyuki, et al.
Veröffentlicht: (2025)
von: Kamo, Naoyuki, et al.
Veröffentlicht: (2025)
SpeechGLUE: How Well Can Self-Supervised Speech Models Capture Linguistic Knowledge?
von: Ashihara, Takanori, et al.
Veröffentlicht: (2023)
von: Ashihara, Takanori, et al.
Veröffentlicht: (2023)
Boosting Hybrid Autoregressive Transducer-based ASR with Internal Acoustic Model Training and Dual Blank Thresholding
von: Moriya, Takafumi, et al.
Veröffentlicht: (2024)
von: Moriya, Takafumi, et al.
Veröffentlicht: (2024)
Investigation of Speaker Representation for Target-Speaker Speech Processing
von: Ashihara, Takanori, et al.
Veröffentlicht: (2024)
von: Ashihara, Takanori, et al.
Veröffentlicht: (2024)
Frontend Token Enhancement for Token-Based Speech Recognition
von: Ashihara, Takanori, et al.
Veröffentlicht: (2026)
von: Ashihara, Takanori, et al.
Veröffentlicht: (2026)
Generic Speech Enhancement with Self-Supervised Representation Space Loss
von: Sato, Hiroshi, et al.
Veröffentlicht: (2025)
von: Sato, Hiroshi, et al.
Veröffentlicht: (2025)
All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASR
von: Moriya, Takafumi, et al.
Veröffentlicht: (2025)
von: Moriya, Takafumi, et al.
Veröffentlicht: (2025)
Chunkwise Aligners for Streaming Speech Recognition
von: Teo, Wen Shen, et al.
Veröffentlicht: (2026)
von: Teo, Wen Shen, et al.
Veröffentlicht: (2026)
SpeakerBeam-SS: Real-time Target Speaker Extraction with Lightweight Conv-TasNet and State Space Modeling
von: Sato, Hiroshi, et al.
Veröffentlicht: (2024)
von: Sato, Hiroshi, et al.
Veröffentlicht: (2024)
Recursive Attentive Pooling for Extracting Speaker Embeddings from Multi-Speaker Recordings
von: Horiguchi, Shota, et al.
Veröffentlicht: (2024)
von: Horiguchi, Shota, et al.
Veröffentlicht: (2024)
Guided Speaker Embedding
von: Horiguchi, Shota, et al.
Veröffentlicht: (2024)
von: Horiguchi, Shota, et al.
Veröffentlicht: (2024)
Lightweight Zero-shot Text-to-Speech with Mixture of Adapters
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024)
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024)
Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024)
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024)
Can We Really Repurpose Multi-Speaker ASR Corpus for Speaker Diarization?
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
Mitigating Non-Target Speaker Bias in Guided Speaker Embedding
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
Probing Self-supervised Learning Models with Target Speech Extraction
von: Peng, Junyi, et al.
Veröffentlicht: (2024)
von: Peng, Junyi, et al.
Veröffentlicht: (2024)
VBx for End-to-End Neural and Clustering-based Diarization
von: Pálka, Petr, et al.
Veröffentlicht: (2025)
von: Pálka, Petr, et al.
Veröffentlicht: (2025)
TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models
von: Peng, Junyi, et al.
Veröffentlicht: (2025)
von: Peng, Junyi, et al.
Veröffentlicht: (2025)
End-to-End Speech Recognition with Pre-trained Masked Language Model
von: Higuchi, Yosuke, et al.
Veröffentlicht: (2024)
von: Higuchi, Yosuke, et al.
Veröffentlicht: (2024)
Dissecting the Segmentation Model of End-to-End Diarization with Vector Clustering
von: Plaquet, Alexis, et al.
Veröffentlicht: (2025)
von: Plaquet, Alexis, et al.
Veröffentlicht: (2025)
An End-to-End Speech Summarization Using Large Language Model
von: Shang, Hengchao, et al.
Veröffentlicht: (2024)
von: Shang, Hengchao, et al.
Veröffentlicht: (2024)
Interpreting End-to-End Deep Learning Models for Speech Source Localization Using Layer-wise Relevance Propagation
von: Comanducci, Luca, et al.
Veröffentlicht: (2024)
von: Comanducci, Luca, et al.
Veröffentlicht: (2024)
Enhancing Fully Formatted End-to-End Speech Recognition with Knowledge Distillation via Multi-Codebook Vector Quantization
von: You, Jian, et al.
Veröffentlicht: (2025)
von: You, Jian, et al.
Veröffentlicht: (2025)
An Efficient End-to-End Approach to Noise Invariant Speech Features via Multi-Task Learning
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2024)
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2024)
Harnessing the Zero-Shot Power of Instruction-Tuned Large Language Model in End-to-End Speech Recognition
von: Higuchi, Yosuke, et al.
Veröffentlicht: (2023)
von: Higuchi, Yosuke, et al.
Veröffentlicht: (2023)
Disentangled-Transformer: An Explainable End-to-End Automatic Speech Recognition Model with Speech Content-Context Separation
von: Wang, Pu, et al.
Veröffentlicht: (2024)
von: Wang, Pu, et al.
Veröffentlicht: (2024)
End-to-End DOA-Guided Speech Extraction in Noisy Multi-Talker Scenarios
von: Jing, Kangqi, et al.
Veröffentlicht: (2025)
von: Jing, Kangqi, et al.
Veröffentlicht: (2025)
Using Adapters to Overcome Catastrophic Forgetting in End-to-End Automatic Speech Recognition
von: Eeckt, Steven Vander, et al.
Veröffentlicht: (2022)
von: Eeckt, Steven Vander, et al.
Veröffentlicht: (2022)
UniFlow: Unifying Speech Front-End Tasks via Continuous Generative Modeling
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
Lightweight and Robust Multi-Channel End-to-End Speech Recognition with Spherical Harmonic Transform
von: Kong, Xiangzhu, et al.
Veröffentlicht: (2025)
von: Kong, Xiangzhu, et al.
Veröffentlicht: (2025)
Reference Channel Selection by Multi-Channel Masking for End-to-End Multi-Channel Speech Enhancement
von: Dai, Wang, et al.
Veröffentlicht: (2024)
von: Dai, Wang, et al.
Veröffentlicht: (2024)
Breaking Walls: Pioneering Automatic Speech Recognition for Central Kurdish: End-to-End Transformer Paradigm
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2024)
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2024)
SoulX-Transcriber: A Robust End-to-End Framework for Multi-Speaker Speech Transcription
von: Dai, Yuhang, et al.
Veröffentlicht: (2026)
von: Dai, Yuhang, et al.
Veröffentlicht: (2026)
Gammatonegram Representation for End-to-End Dysarthric Speech Processing Tasks: Speech Recognition, Speaker Identification, and Intelligibility Assessment
von: Farhadipour, Aref, et al.
Veröffentlicht: (2023)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2023)
Speech-to-See: End-to-End Speech-Driven Open-Set Object Detection
von: Lu, Wenhuan, et al.
Veröffentlicht: (2025)
von: Lu, Wenhuan, et al.
Veröffentlicht: (2025)
Word Error Rate Definitions and Algorithms for Long-Form Multi-talker Speech Recognition
von: von Neumann, Thilo, et al.
Veröffentlicht: (2025)
von: von Neumann, Thilo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Applying LLMs for Rescoring N-best ASR Hypotheses of Casual Conversations: Effects of Domain Adaptation and Context Carry-over
von: Ogawa, Atsunori, et al.
Veröffentlicht: (2024) -
NTT Multi-Speaker ASR System for the DASR Task of CHiME-8 Challenge
von: Kamo, Naoyuki, et al.
Veröffentlicht: (2024) -
Alignment-Free Training for Transducer-based Multi-Talker ASR
von: Moriya, Takafumi, et al.
Veröffentlicht: (2024) -
What Do Self-Supervised Speech and Speaker Models Learn? New Findings From a Cross Model Layer-Wise Analysis
von: Ashihara, Takanori, et al.
Veröffentlicht: (2024) -
Microphone Array Geometry Independent Multi-Talker Distant ASR: NTT System for the DASR Task of the CHiME-8 Challenge
von: Kamo, Naoyuki, et al.
Veröffentlicht: (2025)