Recording for Eyes, Not Echoing to Ears: Contextualized Spoken-to-Written Conversion of ASR Transcripts
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Jiaqing, Deng, Chong, Zhang, Qinglin, Zhou, Shilin, Chen, Qian, Yu, Hai, Wang, Wen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Loss Masking Is Not Needed in Decoder-only Transformer for Discrete-token-based ASR
by: Chen, Qian, et al.
Published: (2023)
by: Chen, Qian, et al.
Published: (2023)
Skip-Layer Attention: Bridging Abstract and Detailed Dependencies in Transformers
by: Chen, Qian, et al.
Published: (2024)
by: Chen, Qian, et al.
Published: (2024)
OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation
by: Zhang, Qinglin, et al.
Published: (2024)
by: Zhang, Qinglin, et al.
Published: (2024)
CopyNE: Better Contextual ASR by Copying Named Entities
by: Zhou, Shilin, et al.
Published: (2023)
by: Zhou, Shilin, et al.
Published: (2023)
Improving Contextual ASR via Multi-grained Fusion with Large Language Models
by: Zhou, Shilin, et al.
Published: (2025)
by: Zhou, Shilin, et al.
Published: (2025)
JSPG: Dynamic Dictionary Filtering via Joint Semantic-Pinyin-Glyph Retrieval for Chinese Contextual ASR
by: Zhou, Shilin, et al.
Published: (2026)
by: Zhou, Shilin, et al.
Published: (2026)
DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech Representations
by: Tan, Chao-Hong, et al.
Published: (2025)
by: Tan, Chao-Hong, et al.
Published: (2025)
Multimodal Fusion and Coherence Modeling for Video Topic Segmentation
by: Yu, Hai, et al.
Published: (2024)
by: Yu, Hai, et al.
Published: (2024)
FormalASR: End-to-End Spoken Chinese to Formal Text
by: Ning, Wanyi, et al.
Published: (2026)
by: Ning, Wanyi, et al.
Published: (2026)
Spirit LM: Interleaved Spoken and Written Language Model
by: Nguyen, Tu Anh, et al.
Published: (2024)
by: Nguyen, Tu Anh, et al.
Published: (2024)
Written Term Detection Improves Spoken Term Detection
by: Yusuf, Bolaji, et al.
Published: (2024)
by: Yusuf, Bolaji, et al.
Published: (2024)
Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
by: Cui, Mingyu, et al.
Published: (2024)
by: Cui, Mingyu, et al.
Published: (2024)
Chapter C7 Spoken and Written Performatives
by: Durant, Alan, et al.
Published: (2021)
by: Durant, Alan, et al.
Published: (2021)
Contextual ASR Error Handling with LLMs Augmentation for Goal-Oriented Conversational AI
by: Asano, Yuya, et al.
Published: (2025)
by: Asano, Yuya, et al.
Published: (2025)
FGGM: Fisher-Guided Gradient Masking for Continual Learning
by: Tan, Chao-Hong, et al.
Published: (2026)
by: Tan, Chao-Hong, et al.
Published: (2026)
Investigating Transcription Normalization in the Faetar ASR Benchmark
by: Peckham, Leo, et al.
Published: (2025)
by: Peckham, Leo, et al.
Published: (2025)
Low-Resource NMT: A Case Study on the Written and Spoken Languages in Hong Kong
by: Mak, Hei Yi, et al.
Published: (2025)
by: Mak, Hei Yi, et al.
Published: (2025)
Who Spoke What When? Evaluating Spoken Language Models for Conversational ASR with Semantic and Overlap-Aware Metrics
by: Tawara, Naohiro, et al.
Published: (2026)
by: Tawara, Naohiro, et al.
Published: (2026)
UniSLU: Unified Spoken Language Understanding from Heterogeneous Cross-Task Datasets
by: Sheng, Zhichao, et al.
Published: (2025)
by: Sheng, Zhichao, et al.
Published: (2025)
How I Built ASR for Endangered Languages with a Spoken Dictionary
by: Bartley, Christopher, et al.
Published: (2025)
by: Bartley, Christopher, et al.
Published: (2025)
CTC-Assisted LLM-Based Contextual ASR
by: Yang, Guanrou, et al.
Published: (2024)
by: Yang, Guanrou, et al.
Published: (2024)
Fun-Audio-Chat Technical Report
by: Tongyi Fun Team, et al.
Published: (2025)
by: Tongyi Fun Team, et al.
Published: (2025)
Beyond Transcription: Mechanistic Interpretability in ASR
by: Glazer, Neta, et al.
Published: (2025)
by: Glazer, Neta, et al.
Published: (2025)
GenesisFunc: Multi-Agent Data Generation for Accurate and Generalizable Function-Calling
by: Xu, Hao-Xiang, et al.
Published: (2026)
by: Xu, Hao-Xiang, et al.
Published: (2026)
Lightweight Prompt Biasing for Contextualized End-to-End ASR Systems
by: Ren, Bo, et al.
Published: (2025)
by: Ren, Bo, et al.
Published: (2025)
Interventional Speech Noise Injection for ASR Generalizable Spoken Language Understanding
by: Jung, Yeonjoon, et al.
Published: (2024)
by: Jung, Yeonjoon, et al.
Published: (2024)
Analyzing Error Propagation in Korean Spoken QA with ASR-LLM Cascades
by: Jung, Donghyuk, et al.
Published: (2026)
by: Jung, Donghyuk, et al.
Published: (2026)
Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization
by: Wang, Hsuan-Yu, et al.
Published: (2025)
by: Wang, Hsuan-Yu, et al.
Published: (2025)
Enhancing the Robustness of Contextual ASR to Varying Biasing Information Volumes Through Purified Semantic Correlation Joint Modeling
by: Gu, Yue, et al.
Published: (2025)
by: Gu, Yue, et al.
Published: (2025)
LLM-as-a-Coauthor: Can Mixed Human-Written and Machine-Generated Text Be Detected?
by: Zhang, Qihui, et al.
Published: (2024)
by: Zhang, Qihui, et al.
Published: (2024)
What Do Humans Hear When Interacting? Experiments on Selective Listening for Evaluating ASR of Spoken Dialogue Systems
by: Mori, Kiyotada, et al.
Published: (2025)
by: Mori, Kiyotada, et al.
Published: (2025)
Fun-ASR Technical Report
by: An, Keyu, et al.
Published: (2025)
by: An, Keyu, et al.
Published: (2025)
EmoNews: A Spoken Dialogue System for Expressive News Conversations
by: Matsuura, Ryuki, et al.
Published: (2025)
by: Matsuura, Ryuki, et al.
Published: (2025)
Improving ASR Contextual Biasing with Guided Attention
by: Tang, Jiyang, et al.
Published: (2024)
by: Tang, Jiyang, et al.
Published: (2024)
Spoken DialogSum: An Emotion-Rich Conversational Dataset for Spoken Dialogue Summarization
by: Lu, Yen-Ju, et al.
Published: (2025)
by: Lu, Yen-Ju, et al.
Published: (2025)
Towards Conversational Medical AI with Eyes, Ears and a Voice
by: Shah, Meet, et al.
Published: (2026)
by: Shah, Meet, et al.
Published: (2026)
Joint Beamforming and Speaker-Attributed ASR for Real Distant-Microphone Meeting Transcription
by: Cui, Can, et al.
Published: (2024)
by: Cui, Can, et al.
Published: (2024)
MedSpeak: A Knowledge Graph-Aided ASR Error Correction Framework for Spoken Medical QA
by: Song, Yutong, et al.
Published: (2026)
by: Song, Yutong, et al.
Published: (2026)
Human Latency Conversational Turns for Spoken Avatar Systems
by: Jacoby, Derek, et al.
Published: (2024)
by: Jacoby, Derek, et al.
Published: (2024)
Optimal Transport Regularization for Speech Text Alignment in Spoken Language Models
by: Xu, Wenze, et al.
Published: (2025)
by: Xu, Wenze, et al.
Published: (2025)
Similar Items
-
Loss Masking Is Not Needed in Decoder-only Transformer for Discrete-token-based ASR
by: Chen, Qian, et al.
Published: (2023) -
Skip-Layer Attention: Bridging Abstract and Detailed Dependencies in Transformers
by: Chen, Qian, et al.
Published: (2024) -
OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation
by: Zhang, Qinglin, et al.
Published: (2024) -
CopyNE: Better Contextual ASR by Copying Named Entities
by: Zhou, Shilin, et al.
Published: (2023) -
Improving Contextual ASR via Multi-grained Fusion with Large Language Models
by: Zhou, Shilin, et al.
Published: (2025)