Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
Fuente:
arXiv
Guardado en:
| Autores principales: | Sharma, Roshan, Shon, Suwon, Lindsey, Mark, Dhamyal, Hira, Singh, Rita, Raj, Bhiksha |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SELM: Enhancing Speech Emotion Recognition for Out-of-Domain Scenarios
por: Bukhari, Hazim, et al.
Publicado: (2024)
por: Bukhari, Hazim, et al.
Publicado: (2024)
On the Evaluation of Speech Foundation Models for Spoken Language Understanding
por: Arora, Siddhant, et al.
Publicado: (2024)
por: Arora, Siddhant, et al.
Publicado: (2024)
Objective Measurements of Voice Quality
por: Dhamyal, Hira, et al.
Publicado: (2024)
por: Dhamyal, Hira, et al.
Publicado: (2024)
What Do Speech Foundation Models Not Learn About Speech?
por: Waheed, Abdul, et al.
Publicado: (2024)
por: Waheed, Abdul, et al.
Publicado: (2024)
Human Voice is Unique
por: Singh, Rita, et al.
Publicado: (2025)
por: Singh, Rita, et al.
Publicado: (2025)
DiscreteSLU: A Large Language Model with Self-Supervised Discrete Speech Units for Spoken Language Understanding
por: Shon, Suwon, et al.
Publicado: (2024)
por: Shon, Suwon, et al.
Publicado: (2024)
A Near-Real-Time Processing Ego Speech Filtering Pipeline Designed for Speech Interruption During Human-Robot Interaction
por: Li, Yue, et al.
Publicado: (2024)
por: Li, Yue, et al.
Publicado: (2024)
Psychophysiology-aided Perceptually Fluent Speech Analysis of Children Who Stutter
por: Xiao, Yi, et al.
Publicado: (2022)
por: Xiao, Yi, et al.
Publicado: (2022)
Recreating Neural Activity During Speech Production with Language and Speech Model Embeddings
por: Khanday, Owais Mujtaba, et al.
Publicado: (2025)
por: Khanday, Owais Mujtaba, et al.
Publicado: (2025)
USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal Synthesis
por: Yu, Luca Jiang-Tao, et al.
Publicado: (2024)
por: Yu, Luca Jiang-Tao, et al.
Publicado: (2024)
How Private is Low-Frequency Speech Audio in the Wild? An Analysis of Verbal Intelligibility by Humans and Machines
por: Liu, Ailin, et al.
Publicado: (2024)
por: Liu, Ailin, et al.
Publicado: (2024)
Cross-Lingual Speech Emotion Recognition: Humans vs. Self-Supervised Models
por: Han, Zhichen, et al.
Publicado: (2024)
por: Han, Zhichen, et al.
Publicado: (2024)
Towards Temporally Explainable Dysarthric Speech Clarity Assessment
por: Park, Seohyun, et al.
Publicado: (2025)
por: Park, Seohyun, et al.
Publicado: (2025)
Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation
por: Hu, Rui, et al.
Publicado: (2025)
por: Hu, Rui, et al.
Publicado: (2025)
SACM: SEEG-Audio Contrastive Matching for Chinese Speech Decoding
por: Wang, Hongbin, et al.
Publicado: (2025)
por: Wang, Hongbin, et al.
Publicado: (2025)
Directional Source Separation for Robust Speech Recognition on Smart Glasses
por: Feng, Tiantian, et al.
Publicado: (2023)
por: Feng, Tiantian, et al.
Publicado: (2023)
VoiceX: A Text-To-Speech Framework for Custom Voices
por: Mertes, Silvan, et al.
Publicado: (2024)
por: Mertes, Silvan, et al.
Publicado: (2024)
Personalized Speech Emotion Recognition in Human-Robot Interaction using Vision Transformers
por: Mishra, Ruchik, et al.
Publicado: (2024)
por: Mishra, Ruchik, et al.
Publicado: (2024)
Psychoacoustic Challenges Of Speech Enhancement On VoIP Platforms
por: Konan, Joseph, et al.
Publicado: (2023)
por: Konan, Joseph, et al.
Publicado: (2023)
NeuroIncept Decoder for High-Fidelity Speech Reconstruction from Neural Activity
por: Khanday, Owais Mujtaba, et al.
Publicado: (2025)
por: Khanday, Owais Mujtaba, et al.
Publicado: (2025)
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
por: Chen, Youjun, et al.
Publicado: (2025)
por: Chen, Youjun, et al.
Publicado: (2025)
Using Confidence Scores to Improve Eyes-free Detection of Speech Recognition Errors
por: Nowrin, Sadia, et al.
Publicado: (2024)
por: Nowrin, Sadia, et al.
Publicado: (2024)
Human Feedback Driven Dynamic Speech Emotion Recognition
por: Fedorov, Ilya, et al.
Publicado: (2025)
por: Fedorov, Ilya, et al.
Publicado: (2025)
Adapting Whisper for Lightweight and Efficient Automatic Speech Recognition of Children for On-device Edge Applications
por: Dutta, Satwik, et al.
Publicado: (2025)
por: Dutta, Satwik, et al.
Publicado: (2025)
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
por: Zhou, Dongliang, et al.
Publicado: (2025)
por: Zhou, Dongliang, et al.
Publicado: (2025)
Speech Robust Bench: A Robustness Benchmark For Speech Recognition
por: Shah, Muhammad A., et al.
Publicado: (2024)
por: Shah, Muhammad A., et al.
Publicado: (2024)
Domain Adaptation for Contrastive Audio-Language Models
por: Deshmukh, Soham, et al.
Publicado: (2024)
por: Deshmukh, Soham, et al.
Publicado: (2024)
Robust Dual-Modal Speech Keyword Spotting for XR Headsets
por: Cai, Zhuojiang, et al.
Publicado: (2024)
por: Cai, Zhuojiang, et al.
Publicado: (2024)
Summarizing Speech: A Comprehensive Survey
por: Retkowski, Fabian, et al.
Publicado: (2025)
por: Retkowski, Fabian, et al.
Publicado: (2025)
InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
por: Wang, Dingdong, et al.
Publicado: (2025)
por: Wang, Dingdong, et al.
Publicado: (2025)
Real-Time Word-Level Temporal Segmentation in Streaming Speech Recognition
por: Nishida, Naoto, et al.
Publicado: (2025)
por: Nishida, Naoto, et al.
Publicado: (2025)
Detecting COPD Through Speech Analysis: A Dataset of Danish Speech and Machine Learning Approach
por: Sankey-Olsen, Cuno, et al.
Publicado: (2025)
por: Sankey-Olsen, Cuno, et al.
Publicado: (2025)
Revisiting Acoustic Features for Robust ASR
por: Shah, Muhammad A., et al.
Publicado: (2024)
por: Shah, Muhammad A., et al.
Publicado: (2024)
Optimizing Multilingual Text-To-Speech with Accents & Emotions
por: Pawar, Pranav, et al.
Publicado: (2025)
por: Pawar, Pranav, et al.
Publicado: (2025)
Improving Speaker Representations Using Contrastive Losses on Multi-scale Features
por: Dixit, Satvik, et al.
Publicado: (2024)
por: Dixit, Satvik, et al.
Publicado: (2024)
Investigating the Effects of Large-Scale Pseudo-Stereo Data and Different Speech Foundation Model on Dialogue Generative Spoken Language Model
por: Fu, Yu-Kuan, et al.
Publicado: (2024)
por: Fu, Yu-Kuan, et al.
Publicado: (2024)
VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching
por: Guo, Yiwei, et al.
Publicado: (2023)
por: Guo, Yiwei, et al.
Publicado: (2023)
An End-to-End Speech Summarization Using Large Language Model
por: Shang, Hengchao, et al.
Publicado: (2024)
por: Shang, Hengchao, et al.
Publicado: (2024)
Subject Disentanglement Neural Network for Speech Envelope Reconstruction from EEG
por: Zhang, Li, et al.
Publicado: (2025)
por: Zhang, Li, et al.
Publicado: (2025)
A Cross-Modal Approach to Silent Speech with LLM-Enhanced Recognition
por: Benster, Tyler, et al.
Publicado: (2024)
por: Benster, Tyler, et al.
Publicado: (2024)
Ejemplares similares
-
SELM: Enhancing Speech Emotion Recognition for Out-of-Domain Scenarios
por: Bukhari, Hazim, et al.
Publicado: (2024) -
On the Evaluation of Speech Foundation Models for Spoken Language Understanding
por: Arora, Siddhant, et al.
Publicado: (2024) -
Objective Measurements of Voice Quality
por: Dhamyal, Hira, et al.
Publicado: (2024) -
What Do Speech Foundation Models Not Learn About Speech?
por: Waheed, Abdul, et al.
Publicado: (2024) -
Human Voice is Unique
por: Singh, Rita, et al.
Publicado: (2025)