Recording for Eyes, Not Echoing to Ears: Contextualized Spoken-to-Written Conversion of ASR Transcripts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Jiaqing, Deng, Chong, Zhang, Qinglin, Zhou, Shilin, Chen, Qian, Yu, Hai, Wang, Wen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Loss Masking Is Not Needed in Decoder-only Transformer for Discrete-token-based ASR
von: Chen, Qian, et al.
Veröffentlicht: (2023)
von: Chen, Qian, et al.
Veröffentlicht: (2023)
Skip-Layer Attention: Bridging Abstract and Detailed Dependencies in Transformers
von: Chen, Qian, et al.
Veröffentlicht: (2024)
von: Chen, Qian, et al.
Veröffentlicht: (2024)
OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation
von: Zhang, Qinglin, et al.
Veröffentlicht: (2024)
von: Zhang, Qinglin, et al.
Veröffentlicht: (2024)
CopyNE: Better Contextual ASR by Copying Named Entities
von: Zhou, Shilin, et al.
Veröffentlicht: (2023)
von: Zhou, Shilin, et al.
Veröffentlicht: (2023)
Improving Contextual ASR via Multi-grained Fusion with Large Language Models
von: Zhou, Shilin, et al.
Veröffentlicht: (2025)
von: Zhou, Shilin, et al.
Veröffentlicht: (2025)
JSPG: Dynamic Dictionary Filtering via Joint Semantic-Pinyin-Glyph Retrieval for Chinese Contextual ASR
von: Zhou, Shilin, et al.
Veröffentlicht: (2026)
von: Zhou, Shilin, et al.
Veröffentlicht: (2026)
DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech Representations
von: Tan, Chao-Hong, et al.
Veröffentlicht: (2025)
von: Tan, Chao-Hong, et al.
Veröffentlicht: (2025)
Multimodal Fusion and Coherence Modeling for Video Topic Segmentation
von: Yu, Hai, et al.
Veröffentlicht: (2024)
von: Yu, Hai, et al.
Veröffentlicht: (2024)
FormalASR: End-to-End Spoken Chinese to Formal Text
von: Ning, Wanyi, et al.
Veröffentlicht: (2026)
von: Ning, Wanyi, et al.
Veröffentlicht: (2026)
Spirit LM: Interleaved Spoken and Written Language Model
von: Nguyen, Tu Anh, et al.
Veröffentlicht: (2024)
von: Nguyen, Tu Anh, et al.
Veröffentlicht: (2024)
Written Term Detection Improves Spoken Term Detection
von: Yusuf, Bolaji, et al.
Veröffentlicht: (2024)
von: Yusuf, Bolaji, et al.
Veröffentlicht: (2024)
Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
Chapter C7 Spoken and Written Performatives
von: Durant, Alan, et al.
Veröffentlicht: (2021)
von: Durant, Alan, et al.
Veröffentlicht: (2021)
Contextual ASR Error Handling with LLMs Augmentation for Goal-Oriented Conversational AI
von: Asano, Yuya, et al.
Veröffentlicht: (2025)
von: Asano, Yuya, et al.
Veröffentlicht: (2025)
FGGM: Fisher-Guided Gradient Masking for Continual Learning
von: Tan, Chao-Hong, et al.
Veröffentlicht: (2026)
von: Tan, Chao-Hong, et al.
Veröffentlicht: (2026)
Investigating Transcription Normalization in the Faetar ASR Benchmark
von: Peckham, Leo, et al.
Veröffentlicht: (2025)
von: Peckham, Leo, et al.
Veröffentlicht: (2025)
Low-Resource NMT: A Case Study on the Written and Spoken Languages in Hong Kong
von: Mak, Hei Yi, et al.
Veröffentlicht: (2025)
von: Mak, Hei Yi, et al.
Veröffentlicht: (2025)
Who Spoke What When? Evaluating Spoken Language Models for Conversational ASR with Semantic and Overlap-Aware Metrics
von: Tawara, Naohiro, et al.
Veröffentlicht: (2026)
von: Tawara, Naohiro, et al.
Veröffentlicht: (2026)
UniSLU: Unified Spoken Language Understanding from Heterogeneous Cross-Task Datasets
von: Sheng, Zhichao, et al.
Veröffentlicht: (2025)
von: Sheng, Zhichao, et al.
Veröffentlicht: (2025)
How I Built ASR for Endangered Languages with a Spoken Dictionary
von: Bartley, Christopher, et al.
Veröffentlicht: (2025)
von: Bartley, Christopher, et al.
Veröffentlicht: (2025)
CTC-Assisted LLM-Based Contextual ASR
von: Yang, Guanrou, et al.
Veröffentlicht: (2024)
von: Yang, Guanrou, et al.
Veröffentlicht: (2024)
Fun-Audio-Chat Technical Report
von: Tongyi Fun Team, et al.
Veröffentlicht: (2025)
von: Tongyi Fun Team, et al.
Veröffentlicht: (2025)
Beyond Transcription: Mechanistic Interpretability in ASR
von: Glazer, Neta, et al.
Veröffentlicht: (2025)
von: Glazer, Neta, et al.
Veröffentlicht: (2025)
GenesisFunc: Multi-Agent Data Generation for Accurate and Generalizable Function-Calling
von: Xu, Hao-Xiang, et al.
Veröffentlicht: (2026)
von: Xu, Hao-Xiang, et al.
Veröffentlicht: (2026)
Lightweight Prompt Biasing for Contextualized End-to-End ASR Systems
von: Ren, Bo, et al.
Veröffentlicht: (2025)
von: Ren, Bo, et al.
Veröffentlicht: (2025)
Interventional Speech Noise Injection for ASR Generalizable Spoken Language Understanding
von: Jung, Yeonjoon, et al.
Veröffentlicht: (2024)
von: Jung, Yeonjoon, et al.
Veröffentlicht: (2024)
Analyzing Error Propagation in Korean Spoken QA with ASR-LLM Cascades
von: Jung, Donghyuk, et al.
Veröffentlicht: (2026)
von: Jung, Donghyuk, et al.
Veröffentlicht: (2026)
Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization
von: Wang, Hsuan-Yu, et al.
Veröffentlicht: (2025)
von: Wang, Hsuan-Yu, et al.
Veröffentlicht: (2025)
Enhancing the Robustness of Contextual ASR to Varying Biasing Information Volumes Through Purified Semantic Correlation Joint Modeling
von: Gu, Yue, et al.
Veröffentlicht: (2025)
von: Gu, Yue, et al.
Veröffentlicht: (2025)
LLM-as-a-Coauthor: Can Mixed Human-Written and Machine-Generated Text Be Detected?
von: Zhang, Qihui, et al.
Veröffentlicht: (2024)
von: Zhang, Qihui, et al.
Veröffentlicht: (2024)
What Do Humans Hear When Interacting? Experiments on Selective Listening for Evaluating ASR of Spoken Dialogue Systems
von: Mori, Kiyotada, et al.
Veröffentlicht: (2025)
von: Mori, Kiyotada, et al.
Veröffentlicht: (2025)
Fun-ASR Technical Report
von: An, Keyu, et al.
Veröffentlicht: (2025)
von: An, Keyu, et al.
Veröffentlicht: (2025)
EmoNews: A Spoken Dialogue System for Expressive News Conversations
von: Matsuura, Ryuki, et al.
Veröffentlicht: (2025)
von: Matsuura, Ryuki, et al.
Veröffentlicht: (2025)
Improving ASR Contextual Biasing with Guided Attention
von: Tang, Jiyang, et al.
Veröffentlicht: (2024)
von: Tang, Jiyang, et al.
Veröffentlicht: (2024)
Spoken DialogSum: An Emotion-Rich Conversational Dataset for Spoken Dialogue Summarization
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2025)
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2025)
Towards Conversational Medical AI with Eyes, Ears and a Voice
von: Shah, Meet, et al.
Veröffentlicht: (2026)
von: Shah, Meet, et al.
Veröffentlicht: (2026)
Joint Beamforming and Speaker-Attributed ASR for Real Distant-Microphone Meeting Transcription
von: Cui, Can, et al.
Veröffentlicht: (2024)
von: Cui, Can, et al.
Veröffentlicht: (2024)
MedSpeak: A Knowledge Graph-Aided ASR Error Correction Framework for Spoken Medical QA
von: Song, Yutong, et al.
Veröffentlicht: (2026)
von: Song, Yutong, et al.
Veröffentlicht: (2026)
Human Latency Conversational Turns for Spoken Avatar Systems
von: Jacoby, Derek, et al.
Veröffentlicht: (2024)
von: Jacoby, Derek, et al.
Veröffentlicht: (2024)
Optimal Transport Regularization for Speech Text Alignment in Spoken Language Models
von: Xu, Wenze, et al.
Veröffentlicht: (2025)
von: Xu, Wenze, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Loss Masking Is Not Needed in Decoder-only Transformer for Discrete-token-based ASR
von: Chen, Qian, et al.
Veröffentlicht: (2023) -
Skip-Layer Attention: Bridging Abstract and Detailed Dependencies in Transformers
von: Chen, Qian, et al.
Veröffentlicht: (2024) -
OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation
von: Zhang, Qinglin, et al.
Veröffentlicht: (2024) -
CopyNE: Better Contextual ASR by Copying Named Entities
von: Zhou, Shilin, et al.
Veröffentlicht: (2023) -
Improving Contextual ASR via Multi-grained Fusion with Large Language Models
von: Zhou, Shilin, et al.
Veröffentlicht: (2025)