Joint Beamforming and Speaker-Attributed ASR for Real Distant-Microphone Meeting Transcription
Fuente:
arXiv
Guardado en:
| Autores principales: | Cui, Can, Sheikh, Imran Ahamad, Sadeghi, Mostafa, Vincent, Emmanuel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Improving Speaker Assignment in Speaker-Attributed ASR for Real Meeting Applications
por: Cui, Can, et al.
Publicado: (2024)
por: Cui, Can, et al.
Publicado: (2024)
End-to-end Joint Punctuated and Normalized ASR with a Limited Amount of Punctuated Training Data
por: Cui, Can, et al.
Publicado: (2023)
por: Cui, Can, et al.
Publicado: (2023)
Data-independent Beamforming for End-to-end Multichannel Multi-speaker ASR
por: Cui, Can, et al.
Publicado: (2025)
por: Cui, Can, et al.
Publicado: (2025)
The Impact of Automatic Speech Transcription on Speaker Attribution
por: Aggazzotti, Cristina, et al.
Publicado: (2025)
por: Aggazzotti, Cristina, et al.
Publicado: (2025)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
por: Nguyen, Thai-Binh, et al.
Publicado: (2024)
por: Nguyen, Thai-Binh, et al.
Publicado: (2024)
Can Authorship Attribution Models Distinguish Speakers in Speech Transcripts?
por: Aggazzotti, Cristina, et al.
Publicado: (2023)
por: Aggazzotti, Cristina, et al.
Publicado: (2023)
Speaker-Reasoner: Scaling Interaction Turns and Reasoning Patterns for Timestamped Speaker-Attributed ASR
por: Lin, Zhennan, et al.
Publicado: (2026)
por: Lin, Zhennan, et al.
Publicado: (2026)
NOTSOFAR-1 Challenge: New Datasets, Baseline, and Tasks for Distant Meeting Transcription
por: Vinnikov, Alon, et al.
Publicado: (2024)
por: Vinnikov, Alon, et al.
Publicado: (2024)
Investigating Transcription Normalization in the Faetar ASR Benchmark
por: Peckham, Leo, et al.
Publicado: (2025)
por: Peckham, Leo, et al.
Publicado: (2025)
An Exhaustive Evaluation of TTS- and VC-based Data Augmentation for ASR
por: Ogun, Sewade, et al.
Publicado: (2025)
por: Ogun, Sewade, et al.
Publicado: (2025)
CALM: Joint Contextual Acoustic-Linguistic Modeling for Personalization of Multi-Speaker ASR
por: Shakeel, Muhammad, et al.
Publicado: (2026)
por: Shakeel, Muhammad, et al.
Publicado: (2026)
Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization
por: Wang, Hsuan-Yu, et al.
Publicado: (2025)
por: Wang, Hsuan-Yu, et al.
Publicado: (2025)
Beyond Transcription: Mechanistic Interpretability in ASR
por: Glazer, Neta, et al.
Publicado: (2025)
por: Glazer, Neta, et al.
Publicado: (2025)
Recording for Eyes, Not Echoing to Ears: Contextualized Spoken-to-Written Conversion of ASR Transcripts
por: Liu, Jiaqing, et al.
Publicado: (2024)
por: Liu, Jiaqing, et al.
Publicado: (2024)
Towards Fair ASR For Second Language Speakers Using Fairness Prompted Finetuning
por: Swain, Monorama, et al.
Publicado: (2025)
por: Swain, Monorama, et al.
Publicado: (2025)
Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder
por: Shakeel, Muhammad, et al.
Publicado: (2025)
por: Shakeel, Muhammad, et al.
Publicado: (2025)
Analysis of Speech Temporal Dynamics in the Context of Speaker Verification and Voice Anonymization
por: Tomashenko, Natalia, et al.
Publicado: (2024)
por: Tomashenko, Natalia, et al.
Publicado: (2024)
ASoBO: Attentive Beamformer Selection for Distant Speaker Diarization in Meetings
por: Mariotte, Theo, et al.
Publicado: (2024)
por: Mariotte, Theo, et al.
Publicado: (2024)
Distantly-Supervised Joint Extraction with Noise-Robust Learning
por: Li, Yufei, et al.
Publicado: (2023)
por: Li, Yufei, et al.
Publicado: (2023)
Refining Transcripts With TV Subtitles by Prompt-Based Weakly Supervised Training of ASR
por: Zhao, Xinnian, et al.
Publicado: (2025)
por: Zhao, Xinnian, et al.
Publicado: (2025)
Mind the Gap: Entity-Preserved Context-Aware ASR Structured Transcriptions
por: Altinok, Duygu
Publicado: (2025)
por: Altinok, Duygu
Publicado: (2025)
Identifying Speakers in Dialogue Transcripts: A Text-based Approach Using Pretrained Language Models
por: Nguyen, Minh, et al.
Publicado: (2024)
por: Nguyen, Minh, et al.
Publicado: (2024)
Distinguishing Repetition Disfluency from Morphological Reduplication in Bangla ASR Transcripts: A Novel Corpus and Benchmarking Analysis
por: Arpa, Zaara Zabeen, et al.
Publicado: (2025)
por: Arpa, Zaara Zabeen, et al.
Publicado: (2025)
The Sound of Healthcare: Improving Medical Transcription ASR Accuracy with Large Language Models
por: Adedeji, Ayo, et al.
Publicado: (2024)
por: Adedeji, Ayo, et al.
Publicado: (2024)
Ethio-ASR: Joint Multilingual Speech Recognition and Language Identification for Ethiopian Languages
por: Abdullah, Badr M., et al.
Publicado: (2026)
por: Abdullah, Badr M., et al.
Publicado: (2026)
MeetEval: A Toolkit for Computation of Word Error Rates for Meeting Transcription Systems
por: von Neumann, Thilo, et al.
Publicado: (2023)
por: von Neumann, Thilo, et al.
Publicado: (2023)
Echoes of Phonetics: Unveiling Relevant Acoustic Cues for ASR via Feature Attribution
por: Fucci, Dennis, et al.
Publicado: (2025)
por: Fucci, Dennis, et al.
Publicado: (2025)
GETALP@AutoMin 2025: Leveraging RAG to Answer Questions based on Meeting Transcripts
por: Kang, Jeongwoo, et al.
Publicado: (2025)
por: Kang, Jeongwoo, et al.
Publicado: (2025)
TagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal Grounding
por: Huo, Mingyue, et al.
Publicado: (2026)
por: Huo, Mingyue, et al.
Publicado: (2026)
WhoSaidIt: Human-LLM Collaborative Annotation for Text-Based Multilingual Speaker-Attribute Classification
por: Gao, Lingyu, et al.
Publicado: (2026)
por: Gao, Lingyu, et al.
Publicado: (2026)
Towards Skeletal and Signer Noise Reduction in Sign Language Production via Quaternion-Based Pose Encoding and Contrastive Learning
por: Fauré, Guilhem, et al.
Publicado: (2025)
por: Fauré, Guilhem, et al.
Publicado: (2025)
Breaking the Transcription Bottleneck: Fine-tuning ASR Models for Extremely Low-Resource Fieldwork Languages
por: Liang, Siyu, et al.
Publicado: (2025)
por: Liang, Siyu, et al.
Publicado: (2025)
Action-Item-Driven Summarization of Long Meeting Transcripts
por: Golia, Logan, et al.
Publicado: (2023)
por: Golia, Logan, et al.
Publicado: (2023)
A Toolkit for Joint Speaker Diarization and Identification with Application to Speaker-Attributed ASR
por: Morrone, Giovanni, et al.
Publicado: (2024)
por: Morrone, Giovanni, et al.
Publicado: (2024)
LLaSA: A Sensor-Aware LLM for Natural Language Reasoning of Human Activity from IMU Data
por: Imran, Sheikh Asif, et al.
Publicado: (2024)
por: Imran, Sheikh Asif, et al.
Publicado: (2024)
JSPG: Dynamic Dictionary Filtering via Joint Semantic-Pinyin-Glyph Retrieval for Chinese Contextual ASR
por: Zhou, Shilin, et al.
Publicado: (2026)
por: Zhou, Shilin, et al.
Publicado: (2026)
Alzheimer Disease Classification through ASR-based Transcriptions: Exploring the Impact of Punctuation and Pauses
por: Gómez-Zaragozá, Lucía, et al.
Publicado: (2023)
por: Gómez-Zaragozá, Lucía, et al.
Publicado: (2023)
Joint Transcription of Acoustic Guitar Strumming Directions and Chords
por: Murgul, Sebastian, et al.
Publicado: (2025)
por: Murgul, Sebastian, et al.
Publicado: (2025)
NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR
por: Xie, Yuan, et al.
Publicado: (2026)
por: Xie, Yuan, et al.
Publicado: (2026)
Error Analysis in a Modular Meeting Transcription System
por: Vieting, Peter, et al.
Publicado: (2025)
por: Vieting, Peter, et al.
Publicado: (2025)
Ejemplares similares
-
Improving Speaker Assignment in Speaker-Attributed ASR for Real Meeting Applications
por: Cui, Can, et al.
Publicado: (2024) -
End-to-end Joint Punctuated and Normalized ASR with a Limited Amount of Punctuated Training Data
por: Cui, Can, et al.
Publicado: (2023) -
Data-independent Beamforming for End-to-end Multichannel Multi-speaker ASR
por: Cui, Can, et al.
Publicado: (2025) -
The Impact of Automatic Speech Transcription on Speaker Attribution
por: Aggazzotti, Cristina, et al.
Publicado: (2025) -
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
por: Nguyen, Thai-Binh, et al.
Publicado: (2024)