PMF-CEC: Phoneme-augmented Multimodal Fusion for Context-aware ASR Error Correction with Error-specific Selective Decoding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | He, Jiajun, Toda, Tomoki |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MF-AED-AEC: Speech Emotion Recognition by Leveraging Multimodal Fusion, Asr Error Detection, and Asr Error Correction
von: He, Jiajun, et al.
Veröffentlicht: (2024)
von: He, Jiajun, et al.
Veröffentlicht: (2024)
CMT-LLM: Contextual Multi-Talker ASR Utilizing Large Language Models
von: He, Jiajun, et al.
Veröffentlicht: (2025)
von: He, Jiajun, et al.
Veröffentlicht: (2025)
FlanEC: Exploring Flan-T5 for Post-ASR Error Correction
von: La Quatra, Moreno, et al.
Veröffentlicht: (2025)
von: La Quatra, Moreno, et al.
Veröffentlicht: (2025)
A Comprehensive Study on the Effectiveness of ASR Representations for Noise-Robust Speech Emotion Recognition
von: Shi, Xiaohan, et al.
Veröffentlicht: (2023)
von: Shi, Xiaohan, et al.
Veröffentlicht: (2023)
Decoder-only Conformer with Modality-aware Sparse Mixtures of Experts for ASR
von: Lee, Jaeyoung, et al.
Veröffentlicht: (2026)
von: Lee, Jaeyoung, et al.
Veröffentlicht: (2026)
Data Augmentation for Spoken Grammatical Error Correction
von: Karanasou, Penny, et al.
Veröffentlicht: (2025)
von: Karanasou, Penny, et al.
Veröffentlicht: (2025)
Revise, Reason, and Recognize: LLM-Based Emotion Recognition via Emotion-Specific Prompts and ASR Error Correction
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
ASR Error Correction using Large Language Models
von: Ma, Rao, et al.
Veröffentlicht: (2024)
von: Ma, Rao, et al.
Veröffentlicht: (2024)
Crossmodal ASR Error Correction with Discrete Speech Units
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
ASR-EC Benchmark: Evaluating Large Language Models on Chinese ASR Error Correction
von: Wei, Victor Junqiu, et al.
Veröffentlicht: (2024)
von: Wei, Victor Junqiu, et al.
Veröffentlicht: (2024)
Revisiting ASR Error Correction with Specialized Models
von: Gu, Zijin, et al.
Veröffentlicht: (2024)
von: Gu, Zijin, et al.
Veröffentlicht: (2024)
Evolutionary Prompt Design for LLM-Based Post-ASR Error Correction
von: Sachdev, Rithik, et al.
Veröffentlicht: (2024)
von: Sachdev, Rithik, et al.
Veröffentlicht: (2024)
Articulation-Informed ASR: Integrating Articulatory Features into ASR via Auxiliary Speech Inversion and Cross-Attention Fusion
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2025)
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2025)
Omni-CLST: Error-aware Curriculum Learning with guided Selective chain-of-Thought for audio question answering
von: Zhao, Jinghua, et al.
Veröffentlicht: (2025)
von: Zhao, Jinghua, et al.
Veröffentlicht: (2025)
Large Language Model Should Understand Pinyin for Chinese ASR Error Correction
von: Li, Yuang, et al.
Veröffentlicht: (2024)
von: Li, Yuang, et al.
Veröffentlicht: (2024)
Data-Efficient ASR Personalization for Non-Normative Speech Using an Uncertainty-Based Phoneme Difficulty Score for Guided Sampling
von: Pokel, Niclas, et al.
Veröffentlicht: (2025)
von: Pokel, Niclas, et al.
Veröffentlicht: (2025)
Failing Forward: Improving Generative Error Correction for ASR with Synthetic Data and Retrieval Augmentation
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
Causal Structure Discovery for Error Diagnostics of Children's ASR
von: Singh, Vishwanath Pratap, et al.
Veröffentlicht: (2025)
von: Singh, Vishwanath Pratap, et al.
Veröffentlicht: (2025)
Advocating Character Error Rate for Multilingual ASR Evaluation
von: K, Thennal D, et al.
Veröffentlicht: (2024)
von: K, Thennal D, et al.
Veröffentlicht: (2024)
Better Pseudo-labeling with Multi-ASR Fusion and Error Correction by SpeechLLM
von: Prakash, Jeena, et al.
Veröffentlicht: (2025)
von: Prakash, Jeena, et al.
Veröffentlicht: (2025)
TurboBias: Universal ASR Context-Biasing powered by GPU-accelerated Phrase-Boosting Tree
von: Andrusenko, Andrei, et al.
Veröffentlicht: (2025)
von: Andrusenko, Andrei, et al.
Veröffentlicht: (2025)
ProKWS: Personalized Keyword Spotting via Collaborative Learning of Phonemes and Prosody
von: Pan, Jianan, et al.
Veröffentlicht: (2026)
von: Pan, Jianan, et al.
Veröffentlicht: (2026)
Benchmarking Japanese Speech Recognition on ASR-LLM Setups with Multi-Pass Augmented Generative Error Correction
von: Ko, Yuka, et al.
Veröffentlicht: (2024)
von: Ko, Yuka, et al.
Veröffentlicht: (2024)
MEBM-Phoneme: Multi-scale Enhanced BrainMagic for End-to-End MEG Phoneme Classification
von: Jinghua, Liang, et al.
Veröffentlicht: (2026)
von: Jinghua, Liang, et al.
Veröffentlicht: (2026)
Two-stage Framework for Robust Speech Emotion Recognition Using Target Speaker Extraction in Human Speech Noise Conditions
von: Mi, Jinyi, et al.
Veröffentlicht: (2024)
von: Mi, Jinyi, et al.
Veröffentlicht: (2024)
Harf-Speech: A Clinically Aligned Framework for Arabic Phoneme-Level Speech Assessment
von: Azad, Asif, et al.
Veröffentlicht: (2026)
von: Azad, Asif, et al.
Veröffentlicht: (2026)
Whispering LLaMA: A Cross-Modal Generative Error Correction Framework for Speech Recognition
von: Radhakrishnan, Srijith, et al.
Veröffentlicht: (2023)
von: Radhakrishnan, Srijith, et al.
Veröffentlicht: (2023)
LadderSym: A Multimodal Interleaved Transformer for Music Practice Error Detection
von: Chou, Benjamin Shiue-Hal, et al.
Veröffentlicht: (2025)
von: Chou, Benjamin Shiue-Hal, et al.
Veröffentlicht: (2025)
NGPU-LM: GPU-Accelerated N-Gram Language Model for Context-Biasing in Greedy ASR Decoding
von: Bataev, Vladimir, et al.
Veröffentlicht: (2025)
von: Bataev, Vladimir, et al.
Veröffentlicht: (2025)
Fun-ASR Technical Report
von: An, Keyu, et al.
Veröffentlicht: (2025)
von: An, Keyu, et al.
Veröffentlicht: (2025)
TSPC: A Two-Stage Phoneme-Centric Architecture for code-switching Vietnamese-English Speech Recognition
von: Anh, Tran Nguyen, et al.
Veröffentlicht: (2025)
von: Anh, Tran Nguyen, et al.
Veröffentlicht: (2025)
Prosody Labeling with Phoneme-BERT and Speech Foundation Models
von: Koriyama, Tomoki
Veröffentlicht: (2025)
von: Koriyama, Tomoki
Veröffentlicht: (2025)
SpeechPrune: Context-aware Token Pruning for Speech Information Retrieval
von: Lin, Yueqian, et al.
Veröffentlicht: (2024)
von: Lin, Yueqian, et al.
Veröffentlicht: (2024)
Analyzing Error Propagation in Korean Spoken QA with ASR-LLM Cascades
von: Jung, Donghyuk, et al.
Veröffentlicht: (2026)
von: Jung, Donghyuk, et al.
Veröffentlicht: (2026)
VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining
von: Zhuo, Jianheng, et al.
Veröffentlicht: (2025)
von: Zhuo, Jianheng, et al.
Veröffentlicht: (2025)
Efficient Adaptation of Multilingual Models for Japanese ASR
von: Bajo, Mark, et al.
Veröffentlicht: (2024)
von: Bajo, Mark, et al.
Veröffentlicht: (2024)
Speech Recognition on TV Series with Video-guided Post-ASR Correction
von: Yang, Haoyuan, et al.
Veröffentlicht: (2025)
von: Yang, Haoyuan, et al.
Veröffentlicht: (2025)
SEAL: Speaker Error Correction using Acoustic-conditioned Large Language Models
von: Kumar, Anurag, et al.
Veröffentlicht: (2025)
von: Kumar, Anurag, et al.
Veröffentlicht: (2025)
Detecting Music Performance Errors with Transformers
von: Chou, Benjamin Shiue-Hal, et al.
Veröffentlicht: (2025)
von: Chou, Benjamin Shiue-Hal, et al.
Veröffentlicht: (2025)
Minimising Biasing Word Errors for Contextual ASR with the Tree-Constrained Pointer Generator
von: Sun, Guangzhi, et al.
Veröffentlicht: (2022)
von: Sun, Guangzhi, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
MF-AED-AEC: Speech Emotion Recognition by Leveraging Multimodal Fusion, Asr Error Detection, and Asr Error Correction
von: He, Jiajun, et al.
Veröffentlicht: (2024) -
CMT-LLM: Contextual Multi-Talker ASR Utilizing Large Language Models
von: He, Jiajun, et al.
Veröffentlicht: (2025) -
FlanEC: Exploring Flan-T5 for Post-ASR Error Correction
von: La Quatra, Moreno, et al.
Veröffentlicht: (2025) -
A Comprehensive Study on the Effectiveness of ASR Representations for Noise-Robust Speech Emotion Recognition
von: Shi, Xiaohan, et al.
Veröffentlicht: (2023) -
Decoder-only Conformer with Modality-aware Sparse Mixtures of Experts for ASR
von: Lee, Jaeyoung, et al.
Veröffentlicht: (2026)