SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Hanlin, Tan, Daxin, Tao, Dehua, Chen, Xiao, Tan, Haochen, Song, Linqi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ToneUnit: A Speech Discretization Approach for Tonal Language Speech Synthesis
von: Tao, Dehua, et al.
Veröffentlicht: (2024)
von: Tao, Dehua, et al.
Veröffentlicht: (2024)
DSA-Tokenizer: Disentangled Semantic-Acoustic Tokenization via Flow Matching-based Hierarchical Fusion
von: Zhang, Hanlin, et al.
Veröffentlicht: (2026)
von: Zhang, Hanlin, et al.
Veröffentlicht: (2026)
CUEMPATHY: A Counseling Speech Dataset for Psychotherapy Research
von: Tao, Dehua, et al.
Veröffentlicht: (2024)
von: Tao, Dehua, et al.
Veröffentlicht: (2024)
MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech
von: Chen, Huakang, et al.
Veröffentlicht: (2026)
von: Chen, Huakang, et al.
Veröffentlicht: (2026)
Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards
von: Ren, Yong, et al.
Veröffentlicht: (2026)
von: Ren, Yong, et al.
Veröffentlicht: (2026)
CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
von: Chen, Junyang, et al.
Veröffentlicht: (2026)
von: Chen, Junyang, et al.
Veröffentlicht: (2026)
PartialEdit: Identifying Partial Deepfakes in the Era of Neural Speech Editing
von: Zhang, You, et al.
Veröffentlicht: (2025)
von: Zhang, You, et al.
Veröffentlicht: (2025)
Enhancing Code-switched Text-to-Speech Synthesis Capability in Large Language Models with only Monolingual Corpora
von: Xu, Jing, et al.
Veröffentlicht: (2024)
von: Xu, Jing, et al.
Veröffentlicht: (2024)
Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM
von: Cui, Wenqian, et al.
Veröffentlicht: (2026)
von: Cui, Wenqian, et al.
Veröffentlicht: (2026)
SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
von: Hou, Yixuan, et al.
Veröffentlicht: (2025)
von: Hou, Yixuan, et al.
Veröffentlicht: (2025)
MOS-Bench: Benchmarking Generalization Abilities of Subjective Speech Quality Assessment Models
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2024)
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2024)
A Large-Scale Probing Analysis of Speaker-Specific Attributes in Self-Supervised Speech Representations
von: Chiu, Aemon Yat Fei, et al.
Veröffentlicht: (2025)
von: Chiu, Aemon Yat Fei, et al.
Veröffentlicht: (2025)
SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
von: Wang, Helin, et al.
Veröffentlicht: (2024)
von: Wang, Helin, et al.
Veröffentlicht: (2024)
BENYO-S2ST-Corpus-1: A Bilingual English-to-Yoruba Direct Speech-to-Speech Translation Corpus
von: Adetiba, Emmanuel, et al.
Veröffentlicht: (2025)
von: Adetiba, Emmanuel, et al.
Veröffentlicht: (2025)
Speech-Omni-Lite: Portable Speech Interfaces for Vision-Language Models
von: Tao, Dehua, et al.
Veröffentlicht: (2026)
von: Tao, Dehua, et al.
Veröffentlicht: (2026)
MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
von: Yang, Qian, et al.
Veröffentlicht: (2024)
von: Yang, Qian, et al.
Veröffentlicht: (2024)
In This Environment, As That Speaker: A Text-Driven Framework for Multi-Attribute Speech Conversion
von: Jin, Jiawei, et al.
Veröffentlicht: (2025)
von: Jin, Jiawei, et al.
Veröffentlicht: (2025)
Fine-Grained and Interpretable Neural Speech Editing
von: Morrison, Max, et al.
Veröffentlicht: (2024)
von: Morrison, Max, et al.
Veröffentlicht: (2024)
An Attribute Interpolation Method in Speech Synthesis by Model Merging
von: Murata, Masato, et al.
Veröffentlicht: (2024)
von: Murata, Masato, et al.
Veröffentlicht: (2024)
Fine-Grained Quantitative Emotion Editing for Speech Generation
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
von: Chen, Wenxi, et al.
Veröffentlicht: (2025)
von: Chen, Wenxi, et al.
Veröffentlicht: (2025)
Sparsely Shared LoRA on Whisper for Child Speech Recognition
von: Liu, Wei, et al.
Veröffentlicht: (2023)
von: Liu, Wei, et al.
Veröffentlicht: (2023)
An Explainable Probabilistic Attribute Embedding Approach for Spoofed Speech Characterization
von: Chhibber, Manasi, et al.
Veröffentlicht: (2024)
von: Chhibber, Manasi, et al.
Veröffentlicht: (2024)
TraceableSpeech: Towards Proactively Traceable Text-to-Speech with Watermarking
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
von: Xin, Detai, et al.
Veröffentlicht: (2024)
von: Xin, Detai, et al.
Veröffentlicht: (2024)
Instance-Specific Test-Time Training for Speech Editing in the Wild
von: Kim, Taewoo, et al.
Veröffentlicht: (2025)
von: Kim, Taewoo, et al.
Veröffentlicht: (2025)
Creating Personalized Synthetic Voices from Articulation Impaired Speech Using Augmented Reconstruction Loss
von: Tian, Yusheng, et al.
Veröffentlicht: (2024)
von: Tian, Yusheng, et al.
Veröffentlicht: (2024)
Towards Ultra-Low-Power Neuromorphic Speech Enhancement with Spiking-FullSubNet
von: Hao, Xiang, et al.
Veröffentlicht: (2024)
von: Hao, Xiang, et al.
Veröffentlicht: (2024)
Scale This, Not That: Investigating Key Dataset Attributes for Efficient Speech Enhancement Scaling
von: Zhang, Leying, et al.
Veröffentlicht: (2024)
von: Zhang, Leying, et al.
Veröffentlicht: (2024)
Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
Distinguishing Neural Speech Synthesis Models Through Fingerprints in Speech Waveforms
von: Zhang, Chu Yuan, et al.
Veröffentlicht: (2023)
von: Zhang, Chu Yuan, et al.
Veröffentlicht: (2023)
NV-Bench: Benchmark of Nonverbal Vocalization Synthesis for Expressive Text-to-Speech Generation
von: Ni, Qinke, et al.
Veröffentlicht: (2026)
von: Ni, Qinke, et al.
Veröffentlicht: (2026)
DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing
von: Sahipjohn, Neha, et al.
Veröffentlicht: (2024)
von: Sahipjohn, Neha, et al.
Veröffentlicht: (2024)
Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition
von: Wang, Huimeng, et al.
Veröffentlicht: (2025)
von: Wang, Huimeng, et al.
Veröffentlicht: (2025)
Investigating Effective Speaker Property Privacy Protection in Federated Learning for Speech Emotion Recognition
von: Tan, Chao, et al.
Veröffentlicht: (2024)
von: Tan, Chao, et al.
Veröffentlicht: (2024)
WenetSpeech-Wu: Datasets, Benchmarks, and Models for a Unified Chinese Wu Dialect Speech Processing Ecosystem
von: Wang, Chengyou, et al.
Veröffentlicht: (2026)
von: Wang, Chengyou, et al.
Veröffentlicht: (2026)
Attention-Guided Adaptation for Code-Switching Speech Recognition
von: Aditya, Bobbi, et al.
Veröffentlicht: (2023)
von: Aditya, Bobbi, et al.
Veröffentlicht: (2023)
Attention-Based Beamformer For Multi-Channel Speech Enhancement
von: Bai, Jinglin, et al.
Veröffentlicht: (2024)
von: Bai, Jinglin, et al.
Veröffentlicht: (2024)
AISHELL-5: The First Open-Source In-Car Multi-Channel Multi-Speaker Speech Dataset for Automatic Speech Diarization and Recognition
von: Dai, Yuhang, et al.
Veröffentlicht: (2025)
von: Dai, Yuhang, et al.
Veröffentlicht: (2025)
The RoyalFlush Automatic Speech Diarization and Recognition System for In-Car Multi-Channel Automatic Speech Recognition Challenge
von: Tian, Jingguang, et al.
Veröffentlicht: (2024)
von: Tian, Jingguang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ToneUnit: A Speech Discretization Approach for Tonal Language Speech Synthesis
von: Tao, Dehua, et al.
Veröffentlicht: (2024) -
DSA-Tokenizer: Disentangled Semantic-Acoustic Tokenization via Flow Matching-based Hierarchical Fusion
von: Zhang, Hanlin, et al.
Veröffentlicht: (2026) -
CUEMPATHY: A Counseling Speech Dataset for Psychotherapy Research
von: Tao, Dehua, et al.
Veröffentlicht: (2024) -
MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech
von: Chen, Huakang, et al.
Veröffentlicht: (2026) -
Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards
von: Ren, Yong, et al.
Veröffentlicht: (2026)