Beyond Transcripts: Iterative Peer-Editing with Audio Unlocks High-Quality Human Summaries of Conversational Speech
Fuente:
arXiv
Guardado en:
| Autores principales: | Chaparala, Kaavya, Thebaud, Thomas, López, Jesús Villalba, Moro-Velazquez, Laureano, Viechnicki, Peter, Dehak, Najim |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Speaker Verification with Speech-Aware LLMs: Evaluation and Augmentation
por: Thebaud, Thomas, et al.
Publicado: (2026)
por: Thebaud, Thomas, et al.
Publicado: (2026)
Noise-robust Speech Separation with Fast Generative Correction
por: Wang, Helin, et al.
Publicado: (2024)
por: Wang, Helin, et al.
Publicado: (2024)
Paired by the Teacher: Turning Unpaired Data into High-Fidelity Pairs for Low-Resource Text Generation
por: Lu, Yen-Ju, et al.
Publicado: (2025)
por: Lu, Yen-Ju, et al.
Publicado: (2025)
Cross-Corpus and Cross-domain Handwriting Assessment of NeuroDegenerative Diseases via Time-Series-to-Image Conversion
por: Chavez, Gabrielle, et al.
Publicado: (2025)
por: Chavez, Gabrielle, et al.
Publicado: (2025)
Detecting Neurodegenerative Diseases using Frame-Level Handwriting Embeddings
por: Laouedj, Sarah, et al.
Publicado: (2025)
por: Laouedj, Sarah, et al.
Publicado: (2025)
MaskVCT: Masked Voice Codec Transformer for Zero-Shot Voice Conversion With Increased Controllability via Multiple Guidances
por: Lee, Junhyeok, et al.
Publicado: (2025)
por: Lee, Junhyeok, et al.
Publicado: (2025)
CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech Processing
por: Lu, Yen-Ju, et al.
Publicado: (2024)
por: Lu, Yen-Ju, et al.
Publicado: (2024)
SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline
por: Wang, Helin, et al.
Publicado: (2025)
por: Wang, Helin, et al.
Publicado: (2025)
Reconstruct! Don't Encode: Self-Supervised Representation Reconstruction Loss for High-Intelligibility and Low-Latency Streaming Neural Audio Codec
por: Lee, Junhyeok, et al.
Publicado: (2026)
por: Lee, Junhyeok, et al.
Publicado: (2026)
Spoken DialogSum: An Emotion-Rich Conversational Dataset for Spoken Dialogue Summarization
por: Lu, Yen-Ju, et al.
Publicado: (2025)
por: Lu, Yen-Ju, et al.
Publicado: (2025)
Study of Pre-processing Defenses against Adversarial Attacks on State-of-the-art Speaker Recognition Systems
por: Joshi, Sonal, et al.
Publicado: (2021)
por: Joshi, Sonal, et al.
Publicado: (2021)
Unraveling Adversarial Examples against Speaker Identification -- Techniques for Attack Detection and Victim Model Classification
por: Joshi, Sonal, et al.
Publicado: (2024)
por: Joshi, Sonal, et al.
Publicado: (2024)
Dynamics of Handwriting for Cognitive Assessment
por: Gabrielle Chavez, et al.
Publicado: (2024)
por: Gabrielle Chavez, et al.
Publicado: (2024)
Analyzing Attention Focus in the Cookie TheftPicture Description Task Using Word Alignment
por: Anna Favaro, et al.
Publicado: (2024)
por: Anna Favaro, et al.
Publicado: (2024)
Cognitive Assessment through Writing Tasks
por: Casey Chen, et al.
Publicado: (2024)
por: Casey Chen, et al.
Publicado: (2024)
Enhancing Dialogue Annotation with Speaker Characteristics Leveraging a Frozen LLM
por: Thebaud, Thomas, et al.
Publicado: (2025)
por: Thebaud, Thomas, et al.
Publicado: (2025)
Demographic Attributes Prediction from Speech Using WavLM Embeddings
por: Yang, Yuchen, et al.
Publicado: (2025)
por: Yang, Yuchen, et al.
Publicado: (2025)
DiT-Flow: Speech Enhancement Robust to Multiple Distortions based on Flow Matching in Latent Space and Diffusion Transformers
por: Cao, Tianyu, et al.
Publicado: (2026)
por: Cao, Tianyu, et al.
Publicado: (2026)
Where Do Backdoors Live? A Component-Level Analysis of Backdoor Propagation in Speech Language Models
por: Fortier, Alexandrine, et al.
Publicado: (2025)
por: Fortier, Alexandrine, et al.
Publicado: (2025)
Interpretable Features for the Assessment of Neurodegenerative Diseases through Handwriting Analysis
por: Thebaud, Thomas, et al.
Publicado: (2024)
por: Thebaud, Thomas, et al.
Publicado: (2024)
Multimodal characterization of Alzheimer's Disease using speech, eye movement, and handwriting
por: Laureano Moro‐Velazquez, et al.
Publicado: (2024)
por: Laureano Moro‐Velazquez, et al.
Publicado: (2024)
Multimodal characterization of Alzheimer’s Disease using speech, eye movement, and handwriting
por: Laureano Moro‐Velazquez, et al.
Publicado: (2024)
por: Laureano Moro‐Velazquez, et al.
Publicado: (2024)
Time Scale Network: A Shallow Neural Network For Time Series Data
por: Meyer, Trevor, et al.
Publicado: (2023)
por: Meyer, Trevor, et al.
Publicado: (2023)
CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
por: Wang, Helin, et al.
Publicado: (2025)
por: Wang, Helin, et al.
Publicado: (2025)
ReFESS-QI: Reference-Free Evaluation For Speech Separation With Joint Quality And Intelligibility Scoring
por: Frummer, Ari, et al.
Publicado: (2025)
por: Frummer, Ari, et al.
Publicado: (2025)
Multi-Target Backdoor Attacks Against Speaker Recognition
por: Fortier, Alexandrine, et al.
Publicado: (2025)
por: Fortier, Alexandrine, et al.
Publicado: (2025)
Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
por: Feng, Tiantian, et al.
Publicado: (2025)
por: Feng, Tiantian, et al.
Publicado: (2025)
Clean Label Attacks against SLU Systems
por: Xinyuan, Henry Li, et al.
Publicado: (2024)
por: Xinyuan, Henry Li, et al.
Publicado: (2024)
Multimodal Analysis of Behavior During Stroop Test for Characterization of Alzheimer’s Disease Signs
por: Trevor Meyer, et al.
Publicado: (2024)
por: Trevor Meyer, et al.
Publicado: (2024)
Adversarial Attacks and Defenses for Speech Recognition Systems
por: Żelasko, Piotr, et al.
Publicado: (2021)
por: Żelasko, Piotr, et al.
Publicado: (2021)
Mai Ho'omāuna i ka 'Ai: Language Models Improve Automatic Speech Recognition in Hawaiian
por: Chaparala, Kaavya, et al.
Publicado: (2024)
por: Chaparala, Kaavya, et al.
Publicado: (2024)
SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
por: Wang, Helin, et al.
Publicado: (2024)
por: Wang, Helin, et al.
Publicado: (2024)
GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models
por: Guan, Yaohan, et al.
Publicado: (2026)
por: Guan, Yaohan, et al.
Publicado: (2026)
SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer
por: Wang, Helin, et al.
Publicado: (2024)
por: Wang, Helin, et al.
Publicado: (2024)
Measurement of the Granularity of Vowel Production Space By Just Producible Different (JPD) Limens
por: Viechnicki, Peter
Publicado: (2025)
por: Viechnicki, Peter
Publicado: (2025)
Latent Speech-Text Transformer
por: Lu, Yen-Ju, et al.
Publicado: (2025)
por: Lu, Yen-Ju, et al.
Publicado: (2025)
Language model integration based on memory control for sequence to sequence speech recognition
por: Cho, Jaejin, et al.
Publicado: (2018)
por: Cho, Jaejin, et al.
Publicado: (2018)
Layer-Aware Early Fusion of Acoustic and Linguistic Embeddings for Cognitive Status Classification
por: Novotny, Krystof, et al.
Publicado: (2026)
por: Novotny, Krystof, et al.
Publicado: (2026)
Speech Editing -- a Summary
por: Kässmann, Tobias, et al.
Publicado: (2024)
por: Kässmann, Tobias, et al.
Publicado: (2024)
Human Detection of Political Speech Deepfakes across Transcripts, Audio, and Video
por: Groh, Matthew, et al.
Publicado: (2022)
por: Groh, Matthew, et al.
Publicado: (2022)
Ejemplares similares
-
Speaker Verification with Speech-Aware LLMs: Evaluation and Augmentation
por: Thebaud, Thomas, et al.
Publicado: (2026) -
Noise-robust Speech Separation with Fast Generative Correction
por: Wang, Helin, et al.
Publicado: (2024) -
Paired by the Teacher: Turning Unpaired Data into High-Fidelity Pairs for Low-Resource Text Generation
por: Lu, Yen-Ju, et al.
Publicado: (2025) -
Cross-Corpus and Cross-domain Handwriting Assessment of NeuroDegenerative Diseases via Time-Series-to-Image Conversion
por: Chavez, Gabrielle, et al.
Publicado: (2025) -
Detecting Neurodegenerative Diseases using Frame-Level Handwriting Embeddings
por: Laouedj, Sarah, et al.
Publicado: (2025)