NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xue, Liumeng, Bian, Weizhen, Pan, Jiahao, Wang, Wenxuan, Ren, Yilin, Kang, Boyi, Hu, Jingbin, Ma, Ziyang, Wang, Shuai, Qian, Xinyuan, Lee, Hung-yi, Guo, Yike |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice
von: Cheng, Sitong, et al.
Veröffentlicht: (2025)
von: Cheng, Sitong, et al.
Veröffentlicht: (2025)
Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
von: Wang, Xinsheng, et al.
Veröffentlicht: (2025)
von: Wang, Xinsheng, et al.
Veröffentlicht: (2025)
UniVocal: Unified Speech-Singing Code-Switching Synthesis
von: Shi, Yufei, et al.
Veröffentlicht: (2026)
von: Shi, Yufei, et al.
Veröffentlicht: (2026)
MeanFlowSE: One-Step Generative Speech Enhancement via MeanFlow
von: Zhu, Yike, et al.
Veröffentlicht: (2025)
von: Zhu, Yike, et al.
Veröffentlicht: (2025)
MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech
von: Chen, Huakang, et al.
Veröffentlicht: (2026)
von: Chen, Huakang, et al.
Veröffentlicht: (2026)
MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation
von: Chen, Szu-Chi, et al.
Veröffentlicht: (2026)
von: Chen, Szu-Chi, et al.
Veröffentlicht: (2026)
VocalNet-MDM: Accelerating Streaming Speech LLM via Self-Distilled Masked Diffusion Modeling
von: Cheng, Ziyang, et al.
Veröffentlicht: (2026)
von: Cheng, Ziyang, et al.
Veröffentlicht: (2026)
EmoSpeech: A Corpus of Emotionally Rich and Contextually Detailed Speech Annotations
von: Bian, Weizhen, et al.
Veröffentlicht: (2024)
von: Bian, Weizhen, et al.
Veröffentlicht: (2024)
FlashSpeech: Efficient Zero-Shot Speech Synthesis
von: Ye, Zhen, et al.
Veröffentlicht: (2024)
von: Ye, Zhen, et al.
Veröffentlicht: (2024)
TASTE-Streaming: Towards Streamable Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2026)
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2026)
SMILE: Speech Meta In-Context Learning for Low-Resource Language Automatic Speech Recognition
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2024)
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2024)
TASLA: Text-Aligned Speech Tokens with Multiple Layer-Aggregation
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2025)
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2025)
Iterate to Differentiate: Enhancing Discriminability and Reliability in Zero-Shot TTS Evaluation
von: Shen, Shengfan, et al.
Veröffentlicht: (2026)
von: Shen, Shengfan, et al.
Veröffentlicht: (2026)
Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy Speech
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
ActorMind: Emulating Human Actor Reasoning for Speech Role-Playing
von: Chen, Xi, et al.
Veröffentlicht: (2026)
von: Chen, Xi, et al.
Veröffentlicht: (2026)
voc2vec: A Foundation Model for Non-Verbal Vocalization
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
VocalNet: Speech LLM with Multi-Token Prediction for Faster and High-Quality Generation
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
NV-Bench: Benchmark of Nonverbal Vocalization Synthesis for Expressive Text-to-Speech Generation
von: Ni, Qinke, et al.
Veröffentlicht: (2026)
von: Ni, Qinke, et al.
Veröffentlicht: (2026)
An Initial Investigation of Neural Replay Simulator for Over-the-Air Adversarial Perturbations to Automatic Speaker Verification
von: Li, Jiaqi, et al.
Veröffentlicht: (2023)
von: Li, Jiaqi, et al.
Veröffentlicht: (2023)
Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages
von: Huang, Kuan-Po, et al.
Veröffentlicht: (2023)
von: Huang, Kuan-Po, et al.
Veröffentlicht: (2023)
CodecFake: Enhancing Anti-Spoofing Models Against Deepfake Audios from Codec-Based Speech Synthesis Systems
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models
von: Jiang, Feng, et al.
Veröffentlicht: (2025)
von: Jiang, Feng, et al.
Veröffentlicht: (2025)
WenetSpeech-Wu: Datasets, Benchmarks, and Models for a Unified Chinese Wu Dialect Speech Processing Ecosystem
von: Wang, Chengyou, et al.
Veröffentlicht: (2026)
von: Wang, Chengyou, et al.
Veröffentlicht: (2026)
FleSpeech: Flexibly Controllable Speech Generation with Various Prompts
von: Li, Hanzhao, et al.
Veröffentlicht: (2025)
von: Li, Hanzhao, et al.
Veröffentlicht: (2025)
Parallel Synthesis for Autoregressive Speech Generation
von: Hsu, Po-chun, et al.
Veröffentlicht: (2022)
von: Hsu, Po-chun, et al.
Veröffentlicht: (2022)
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
von: Wang, Shih-heng, et al.
Veröffentlicht: (2024)
von: Wang, Shih-heng, et al.
Veröffentlicht: (2024)
Affectron: Emotional Speech Synthesis with Affective and Contextually Aligned Nonverbal Vocalizations
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2026)
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2026)
Speech-FT: Merging Pre-trained And Fine-Tuned Speech Representation Models For Cross-Task Generalization
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2025)
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2025)
Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
von: Lin, Hsi-Che, et al.
Veröffentlicht: (2024)
von: Lin, Hsi-Che, et al.
Veröffentlicht: (2024)
SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
von: Hou, Yixuan, et al.
Veröffentlicht: (2025)
von: Hou, Yixuan, et al.
Veröffentlicht: (2025)
Analysis of Self-Supervised Speech Models on Children's Speech and Infant Vocalizations
von: Li, Jialu, et al.
Veröffentlicht: (2024)
von: Li, Jialu, et al.
Veröffentlicht: (2024)
LLM-Codec: Neural Audio Codec Meets Language Model Objectives
von: Chung, Ho-Lam, et al.
Veröffentlicht: (2026)
von: Chung, Ho-Lam, et al.
Veröffentlicht: (2026)
Transfer the linguistic representations from TTS to accent conversion with non-parallel data
von: Chen, Xi, et al.
Veröffentlicht: (2024)
von: Chen, Xi, et al.
Veröffentlicht: (2024)
Multi-level Temporal-channel Speaker Retrieval for Zero-shot Voice Conversion
von: Wang, Zhichao, et al.
Veröffentlicht: (2023)
von: Wang, Zhichao, et al.
Veröffentlicht: (2023)
DAISY: Data Adaptive Self-Supervised Early Exit for Speech Representation Models
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2024)
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2024)
Mitigating Subgroup Disparities in Multi-Label Speech Emotion Recognition: A Pseudo-Labeling and Unsupervised Learning Approach
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2025)
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2025)
LLaSE-G1: Incentivizing Generalization Capability for LLaMA-based Speech Enhancement
von: Kang, Boyi, et al.
Veröffentlicht: (2025)
von: Kang, Boyi, et al.
Veröffentlicht: (2025)
ELEGANCE: Efficient LLM Guidance for Audio-Visual Target Speech Extraction
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
Mel-RoFormer for Vocal Separation and Vocal Melody Transcription
von: Wang, Ju-Chiang, et al.
Veröffentlicht: (2024)
von: Wang, Ju-Chiang, et al.
Veröffentlicht: (2024)
Task Vector in TTS: Toward Emotionally Expressive Dialectal Speech Synthesis
von: Feng, Pengchao, et al.
Veröffentlicht: (2025)
von: Feng, Pengchao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice
von: Cheng, Sitong, et al.
Veröffentlicht: (2025) -
Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
von: Wang, Xinsheng, et al.
Veröffentlicht: (2025) -
UniVocal: Unified Speech-Singing Code-Switching Synthesis
von: Shi, Yufei, et al.
Veröffentlicht: (2026) -
MeanFlowSE: One-Step Generative Speech Enhancement via MeanFlow
von: Zhu, Yike, et al.
Veröffentlicht: (2025) -
MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech
von: Chen, Huakang, et al.
Veröffentlicht: (2026)