Fine-Grained and Interpretable Neural Speech Editing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Morrison, Max, Churchwell, Cameron, Pruyne, Nathan, Pardo, Bryan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
High-Fidelity Neural Phonetic Posteriorgrams
von: Churchwell, Cameron, et al.
Veröffentlicht: (2024)
von: Churchwell, Cameron, et al.
Veröffentlicht: (2024)
Cross-domain Neural Pitch and Periodicity Estimation
von: Morrison, Max, et al.
Veröffentlicht: (2023)
von: Morrison, Max, et al.
Veröffentlicht: (2023)
Combolutional Neural Networks
von: Churchwell, Cameron, et al.
Veröffentlicht: (2025)
von: Churchwell, Cameron, et al.
Veröffentlicht: (2025)
Fine-Grained Quantitative Emotion Editing for Speech Generation
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
The Rhythm In Anything: Audio-Prompted Drums Generation with Masked Language Modeling
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2025)
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2025)
Emotion Neural Transducer for Fine-Grained Speech Emotion Recognition
von: Shen, Siyuan, et al.
Veröffentlicht: (2024)
von: Shen, Siyuan, et al.
Veröffentlicht: (2024)
Deep Audio Watermarks are Shallow: Limitations of Post-Hoc Watermarking Techniques for Speech
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2025)
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2025)
PartialEdit: Identifying Partial Deepfakes in the Era of Neural Speech Editing
von: Zhang, You, et al.
Veröffentlicht: (2025)
von: Zhang, You, et al.
Veröffentlicht: (2025)
Code Drift: Towards Idempotent Neural Audio Codecs
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2024)
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2024)
WildElder: A Chinese Elderly Speech Dataset from the Wild with Fine-Grained Manual Annotations
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing
von: Zhang, Hanlin, et al.
Veröffentlicht: (2026)
von: Zhang, Hanlin, et al.
Veröffentlicht: (2026)
Instance-Specific Test-Time Training for Speech Editing in the Wild
von: Kim, Taewoo, et al.
Veröffentlicht: (2025)
von: Kim, Taewoo, et al.
Veröffentlicht: (2025)
AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis
von: Qi, Tianhua, et al.
Veröffentlicht: (2026)
von: Qi, Tianhua, et al.
Veröffentlicht: (2026)
SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
von: Wang, Helin, et al.
Veröffentlicht: (2024)
von: Wang, Helin, et al.
Veröffentlicht: (2024)
CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
von: Chen, Junyang, et al.
Veröffentlicht: (2026)
von: Chen, Junyang, et al.
Veröffentlicht: (2026)
Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation
von: Guo, Haohan, et al.
Veröffentlicht: (2024)
von: Guo, Haohan, et al.
Veröffentlicht: (2024)
USM-Lite: Quantization and Sparsity Aware Fine-tuning for Speech Recognition with Universal Speech Models
von: Ding, Shaojin, et al.
Veröffentlicht: (2023)
von: Ding, Shaojin, et al.
Veröffentlicht: (2023)
Personalized Neural Speech Codec
von: Jang, Inseon, et al.
Veröffentlicht: (2024)
von: Jang, Inseon, et al.
Veröffentlicht: (2024)
Neural Vocoders as Speech Enhancers
von: Li, Andong, et al.
Veröffentlicht: (2025)
von: Li, Andong, et al.
Veröffentlicht: (2025)
Private kNN-VC: Interpretable Anonymization of Converted Speech
von: Franzreb, Carlos, et al.
Veröffentlicht: (2025)
von: Franzreb, Carlos, et al.
Veröffentlicht: (2025)
Speaker Disentanglement of Speech Pre-trained Model Based on Interpretability
von: Zhu, Xiaoxu, et al.
Veröffentlicht: (2025)
von: Zhu, Xiaoxu, et al.
Veröffentlicht: (2025)
A Neural Speech Codec for Noise Robust Speech Coding
von: Huang, Jiayi, et al.
Veröffentlicht: (2023)
von: Huang, Jiayi, et al.
Veröffentlicht: (2023)
Text2FX: Harnessing CLAP Embeddings for Text-Guided Audio Effects
von: Chu, Annie, et al.
Veröffentlicht: (2024)
von: Chu, Annie, et al.
Veröffentlicht: (2024)
Fine-Tuning Automatic Speech Recognition for People with Parkinson's: An Effective Strategy for Enhancing Speech Technology Accessibility
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2024)
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2024)
Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech Recognition
von: Wagner, Dominik, et al.
Veröffentlicht: (2025)
von: Wagner, Dominik, et al.
Veröffentlicht: (2025)
Distinguishing Neural Speech Synthesis Models Through Fingerprints in Speech Waveforms
von: Zhang, Chu Yuan, et al.
Veröffentlicht: (2023)
von: Zhang, Chu Yuan, et al.
Veröffentlicht: (2023)
Interpretable Audio Editing Evaluation via Chain-of-Thought Difference-Commonality Reasoning with Multimodal LLMs
von: Jia, Yuhang, et al.
Veröffentlicht: (2025)
von: Jia, Yuhang, et al.
Veröffentlicht: (2025)
Fine-Tuning ASR for Stuttered Speech: Personalized vs. Generalized Approaches
von: Mujtaba, Dena, et al.
Veröffentlicht: (2025)
von: Mujtaba, Dena, et al.
Veröffentlicht: (2025)
Fine-grained Preference Optimization Improves Zero-shot Text-to-Speech
von: Yao, Jixun, et al.
Veröffentlicht: (2025)
von: Yao, Jixun, et al.
Veröffentlicht: (2025)
MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model
von: Pham, The Hieu, et al.
Veröffentlicht: (2025)
von: Pham, The Hieu, et al.
Veröffentlicht: (2025)
Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
von: Li, Jiaqi, et al.
Veröffentlicht: (2024)
von: Li, Jiaqi, et al.
Veröffentlicht: (2024)
Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards
von: Ren, Yong, et al.
Veröffentlicht: (2026)
von: Ren, Yong, et al.
Veröffentlicht: (2026)
Probing the Robustness Properties of Neural Speech Codecs
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2025)
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2025)
SpatialCodec: Neural Spatial Speech Coding
von: Xu, Zhongweiyang, et al.
Veröffentlicht: (2023)
von: Xu, Zhongweiyang, et al.
Veröffentlicht: (2023)
Sketch2Sound: Controllable Audio Generation via Time-Varying Signals and Sonic Imitations
von: García, Hugo Flores, et al.
Veröffentlicht: (2024)
von: García, Hugo Flores, et al.
Veröffentlicht: (2024)
Fine-Grained Engine Fault Sound Event Detection Using Multimodal Signals
von: Fedorishin, Dennis, et al.
Veröffentlicht: (2024)
von: Fedorishin, Dennis, et al.
Veröffentlicht: (2024)
Exploring Local Interpretable Model-Agnostic Explanations for Speech Emotion Recognition with Distribution-Shift
von: Hjuler, Maja J., et al.
Veröffentlicht: (2025)
von: Hjuler, Maja J., et al.
Veröffentlicht: (2025)
Plugin Speech Enhancement: A Universal Speech Enhancement Framework Inspired by Dynamic Neural Network
von: Chen, Yanan, et al.
Veröffentlicht: (2024)
von: Chen, Yanan, et al.
Veröffentlicht: (2024)
All Neural Low-latency Directional Speech Extraction
von: Pandey, Ashutosh, et al.
Veröffentlicht: (2024)
von: Pandey, Ashutosh, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
High-Fidelity Neural Phonetic Posteriorgrams
von: Churchwell, Cameron, et al.
Veröffentlicht: (2024) -
Cross-domain Neural Pitch and Periodicity Estimation
von: Morrison, Max, et al.
Veröffentlicht: (2023) -
Combolutional Neural Networks
von: Churchwell, Cameron, et al.
Veröffentlicht: (2025) -
Fine-Grained Quantitative Emotion Editing for Speech Generation
von: Inoue, Sho, et al.
Veröffentlicht: (2024) -
The Rhythm In Anything: Audio-Prompted Drums Generation with Masked Language Modeling
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2025)