CALM: Joint Contextual Acoustic-Linguistic Modeling for Personalization of Multi-Speaker ASR
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shakeel, Muhammad, Fukumoto, Yosuke, Maeda, Chikara, Lin, Chyi-Jiunn, Watanabe, Shinji |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2025)
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2025)
DYNAC: Dynamic Vocabulary based Non-Autoregressive Contextualization for Speech Recognition
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
Contextualized Automatic Speech Recognition with Dynamic Vocabulary
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
Contextualized Automatic Speech Recognition with Attention-Based Bias Phrase Boosted Beam Search
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
Joint Beam Search Integrating CTC, Attention, and Transducer Decoders
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning
von: Peng, Yifan, et al.
Veröffentlicht: (2025)
von: Peng, Yifan, et al.
Veröffentlicht: (2025)
Joint Optimization of Streaming and Non-Streaming Automatic Speech Recognition with Multi-Decoder and Knowledge Distillation
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2024)
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2024)
Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2024)
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2024)
Hypothesis Clustering and Merging: Novel MultiTalker Speech Recognition with Speaker Tokens
von: Kashiwagi, Yosuke, et al.
Veröffentlicht: (2024)
von: Kashiwagi, Yosuke, et al.
Veröffentlicht: (2024)
OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
von: Guo, Pengcheng, et al.
Veröffentlicht: (2024)
von: Guo, Pengcheng, et al.
Veröffentlicht: (2024)
Revisiting Acoustic Features for Robust ASR
von: Shah, Muhammad A., et al.
Veröffentlicht: (2024)
von: Shah, Muhammad A., et al.
Veröffentlicht: (2024)
Semi-Autoregressive Streaming ASR With Label Context
von: Arora, Siddhant, et al.
Veröffentlicht: (2023)
von: Arora, Siddhant, et al.
Veröffentlicht: (2023)
Differentiable K-means for Fully-optimized Discrete Token-based ASR
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
Improving Multilingual ASR in the Wild Using Simple N-best Re-ranking
von: Yan, Brian, et al.
Veröffentlicht: (2024)
von: Yan, Brian, et al.
Veröffentlicht: (2024)
Speaker-Reasoner: Scaling Interaction Turns and Reasoning Patterns for Timestamped Speaker-Attributed ASR
von: Lin, Zhennan, et al.
Veröffentlicht: (2026)
von: Lin, Zhennan, et al.
Veröffentlicht: (2026)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2024)
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2024)
Joint ASR and Speaker Role Tagging with Serialized Output Training
von: Xu, Anfeng, et al.
Veröffentlicht: (2025)
von: Xu, Anfeng, et al.
Veröffentlicht: (2025)
Spiralformer: Low Latency Encoder for Streaming Speech Recognition with Circular Layer Skipping and Early Exiting
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2025)
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2025)
Rapid Language Adaptation for Multilingual E2E Speech Recognition Using Encoder Prompting
von: Kashiwagi, Yosuke, et al.
Veröffentlicht: (2024)
von: Kashiwagi, Yosuke, et al.
Veröffentlicht: (2024)
Speaker Targeting via Self-Speaker Adaptation for Multi-talker ASR
von: Wang, Weiqing, et al.
Veröffentlicht: (2025)
von: Wang, Weiqing, et al.
Veröffentlicht: (2025)
Enhancing the Robustness of Contextual ASR to Varying Biasing Information Volumes Through Purified Semantic Correlation Joint Modeling
von: Gu, Yue, et al.
Veröffentlicht: (2025)
von: Gu, Yue, et al.
Veröffentlicht: (2025)
ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
von: Wang, He, et al.
Veröffentlicht: (2025)
von: Wang, He, et al.
Veröffentlicht: (2025)
Can We Really Repurpose Multi-Speaker ASR Corpus for Speaker Diarization?
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
End-to-End Joint ASR and Speaker Role Diarization with Child-Adult Interactions
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
Target Speaker ASR with Whisper
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
LibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization
von: Jin, Zengrui, et al.
Veröffentlicht: (2024)
von: Jin, Zengrui, et al.
Veröffentlicht: (2024)
Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams
von: He, Xiluo, et al.
Veröffentlicht: (2025)
von: He, Xiluo, et al.
Veröffentlicht: (2025)
Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR
von: Ma, Hao, et al.
Veröffentlicht: (2025)
von: Ma, Hao, et al.
Veröffentlicht: (2025)
Segment-Level Vectorized Beam Search Based on Partially Autoregressive Inference
von: Someki, Masao, et al.
Veröffentlicht: (2023)
von: Someki, Masao, et al.
Veröffentlicht: (2023)
EFFUSE: Efficient Self-Supervised Feature Fusion for E2E ASR in Low Resource and Multilingual Scenarios
von: Srivastava, Tejes, et al.
Veröffentlicht: (2023)
von: Srivastava, Tejes, et al.
Veröffentlicht: (2023)
Minimising Biasing Word Errors for Contextual ASR with the Tree-Constrained Pointer Generator
von: Sun, Guangzhi, et al.
Veröffentlicht: (2022)
von: Sun, Guangzhi, et al.
Veröffentlicht: (2022)
Resource-Efficient Adaptation of Speech Foundation Models for Multi-Speaker ASR
von: Wang, Weiqing, et al.
Veröffentlicht: (2024)
von: Wang, Weiqing, et al.
Veröffentlicht: (2024)
Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
Speaker-Smoothed kNN Speaker Adaptation for End-to-End ASR
von: Li, Shaojun, et al.
Veröffentlicht: (2024)
von: Li, Shaojun, et al.
Veröffentlicht: (2024)
GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
von: Chen, Guoguo, et al.
Veröffentlicht: (2021)
von: Chen, Guoguo, et al.
Veröffentlicht: (2021)
Effects of Speaker Count, Duration, and Accent Diversity on Zero-Shot Accent Robustness in Low-Resource ASR
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2025)
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2025)
PromptASR for contextualized ASR with controllable style
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2023)
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2025) -
DYNAC: Dynamic Vocabulary based Non-Autoregressive Contextualization for Speech Recognition
von: Sudo, Yui, et al.
Veröffentlicht: (2025) -
Contextualized Automatic Speech Recognition with Dynamic Vocabulary
von: Sudo, Yui, et al.
Veröffentlicht: (2024) -
Contextualized Automatic Speech Recognition with Attention-Based Bias Phrase Boosted Beam Search
von: Sudo, Yui, et al.
Veröffentlicht: (2024) -
Joint Beam Search Integrating CTC, Attention, and Transducer Decoders
von: Sudo, Yui, et al.
Veröffentlicht: (2024)