AnimeScore: A Preference-Based Dataset and Framework for Evaluating Anime-Like Speech Style
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Park, Joonyong, Li, Jerry |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model
von: Park, Joonyong, et al.
Veröffentlicht: (2024)
von: Park, Joonyong, et al.
Veröffentlicht: (2024)
A Pilot Study of GSLM-based Simulation of Foreign Accentuation Only Using Native Speech Corpora
von: Onda, Kentaro, et al.
Veröffentlicht: (2024)
von: Onda, Kentaro, et al.
Veröffentlicht: (2024)
SpeechAlign: Aligning Speech Generation to Human Preferences
von: Zhang, Dong, et al.
Veröffentlicht: (2024)
von: Zhang, Dong, et al.
Veröffentlicht: (2024)
Comparative Evaluation of Expressive Japanese Character Text-to-Speech with VITS and Style-BERT-VITS2
von: Rackauckas, Zackary, et al.
Veröffentlicht: (2025)
von: Rackauckas, Zackary, et al.
Veröffentlicht: (2025)
MSLM-S2ST: A Multitask Speech Language Model for Textless Speech-to-Speech Translation with Speaker Style Preservation
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
Generalized Multilingual Text-to-Speech Generation with Language-Aware Style Adaptation
von: Lou, Haowei, et al.
Veröffentlicht: (2025)
von: Lou, Haowei, et al.
Veröffentlicht: (2025)
Speech-MASSIVE: A Multilingual Speech Dataset for SLU and Beyond
von: Lee, Beomseok, et al.
Veröffentlicht: (2024)
von: Lee, Beomseok, et al.
Veröffentlicht: (2024)
Evaluating Automatic Speech Recognition Systems for Korean Meteorological Experts
von: Park, ChaeHun, et al.
Veröffentlicht: (2024)
von: Park, ChaeHun, et al.
Veröffentlicht: (2024)
OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style
von: Kang, Wonjune, et al.
Veröffentlicht: (2025)
von: Kang, Wonjune, et al.
Veröffentlicht: (2025)
YODAS: Youtube-Oriented Dataset for Audio and Speech
von: Li, Xinjian, et al.
Veröffentlicht: (2024)
von: Li, Xinjian, et al.
Veröffentlicht: (2024)
AlignCap: Aligning Speech Emotion Captioning to Human Preferences
von: Liang, Ziqi, et al.
Veröffentlicht: (2024)
von: Liang, Ziqi, et al.
Veröffentlicht: (2024)
LinTO Audio and Textual Datasets to Train and Evaluate Automatic Speech Recognition in Tunisian Arabic Dialect
von: Naouara, Hedi, et al.
Veröffentlicht: (2025)
von: Naouara, Hedi, et al.
Veröffentlicht: (2025)
CASPER: A Large Scale Spontaneous Speech Dataset
von: Xiao, Cihan, et al.
Veröffentlicht: (2025)
von: Xiao, Cihan, et al.
Veröffentlicht: (2025)
From Scores to Preferences: Redefining MOS Benchmarking for Speech Quality Reward Modeling
von: Cao, Yifei, et al.
Veröffentlicht: (2025)
von: Cao, Yifei, et al.
Veröffentlicht: (2025)
GSA-TTS : Toward Zero-Shot Speech Synthesis based on Gradual Style Adaptor
von: Lee, Seokgi, et al.
Veröffentlicht: (2025)
von: Lee, Seokgi, et al.
Veröffentlicht: (2025)
Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation
von: He, Haorui, et al.
Veröffentlicht: (2025)
von: He, Haorui, et al.
Veröffentlicht: (2025)
nEMO: Dataset of Emotional Speech in Polish
von: Christop, Iwona
Veröffentlicht: (2024)
von: Christop, Iwona
Veröffentlicht: (2024)
Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework
von: Li, Zhu, et al.
Veröffentlicht: (2025)
von: Li, Zhu, et al.
Veröffentlicht: (2025)
S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models
von: Jiang, Feng, et al.
Veröffentlicht: (2025)
von: Jiang, Feng, et al.
Veröffentlicht: (2025)
HiKE: Hierarchical Evaluation Framework for Korean-English Code-Switching Speech Recognition
von: Paik, Gio, et al.
Veröffentlicht: (2025)
von: Paik, Gio, et al.
Veröffentlicht: (2025)
EmoTale: An Enacted Speech-emotion Dataset in Danish
von: Hjuler, Maja J., et al.
Veröffentlicht: (2025)
von: Hjuler, Maja J., et al.
Veröffentlicht: (2025)
CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset
von: Yan, Brian, et al.
Veröffentlicht: (2025)
von: Yan, Brian, et al.
Veröffentlicht: (2025)
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning
von: Peng, Yifan, et al.
Veröffentlicht: (2025)
von: Peng, Yifan, et al.
Veröffentlicht: (2025)
Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
von: Li, Tianpeng, et al.
Veröffentlicht: (2025)
von: Li, Tianpeng, et al.
Veröffentlicht: (2025)
MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model
von: Park, Joonyong, et al.
Veröffentlicht: (2025)
von: Park, Joonyong, et al.
Veröffentlicht: (2025)
RoDia: A New Dataset for Romanian Dialect Identification from Speech
von: Rotaru, Codrut, et al.
Veröffentlicht: (2023)
von: Rotaru, Codrut, et al.
Veröffentlicht: (2023)
ChildGuard: A Specialized Dataset for Combatting Child-Targeted Hate Speech
von: Kashyap, Gautam Siddharth, et al.
Veröffentlicht: (2025)
von: Kashyap, Gautam Siddharth, et al.
Veröffentlicht: (2025)
LearnerVoice: A Dataset of Non-Native English Learners' Spontaneous Speech
von: Kim, Haechan, et al.
Veröffentlicht: (2024)
von: Kim, Haechan, et al.
Veröffentlicht: (2024)
StyleSinger: Style Transfer for Out-of-Domain Singing Voice Synthesis
von: Zhang, Yu, et al.
Veröffentlicht: (2023)
von: Zhang, Yu, et al.
Veröffentlicht: (2023)
AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation
von: Manakul, Potsawee, et al.
Veröffentlicht: (2025)
von: Manakul, Potsawee, et al.
Veröffentlicht: (2025)
UCorrect: An Unsupervised Framework for Automatic Speech Recognition Error Correction
von: Guo, Jiaxin, et al.
Veröffentlicht: (2024)
von: Guo, Jiaxin, et al.
Veröffentlicht: (2024)
Sagalee: an Open Source Automatic Speech Recognition Dataset for Oromo Language
von: Abu, Turi, et al.
Veröffentlicht: (2025)
von: Abu, Turi, et al.
Veröffentlicht: (2025)
HebDB: a Weakly Supervised Dataset for Hebrew Speech Processing
von: Turetzky, Arnon, et al.
Veröffentlicht: (2024)
von: Turetzky, Arnon, et al.
Veröffentlicht: (2024)
TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
Enhancing Indonesian Automatic Speech Recognition: Evaluating Multilingual Models with Diverse Speech Variabilities
von: Adila, Aulia, et al.
Veröffentlicht: (2024)
von: Adila, Aulia, et al.
Veröffentlicht: (2024)
ELF: Encoding Speaker-Specific Latent Speech Feature for Speech Synthesis
von: Kong, Jungil, et al.
Veröffentlicht: (2023)
von: Kong, Jungil, et al.
Veröffentlicht: (2023)
StoryTTS: A Highly Expressive Text-to-Speech Dataset with Rich Textual Expressiveness Annotations
von: Liu, Sen, et al.
Veröffentlicht: (2024)
von: Liu, Sen, et al.
Veröffentlicht: (2024)
Pairwise Evaluation of Accent Similarity in Speech Synthesis
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2025)
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2025)
Automatic Speech Recognition System-Independent Word Error Rate Estimation
von: Park, Chanho, et al.
Veröffentlicht: (2024)
von: Park, Chanho, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model
von: Park, Joonyong, et al.
Veröffentlicht: (2024) -
A Pilot Study of GSLM-based Simulation of Foreign Accentuation Only Using Native Speech Corpora
von: Onda, Kentaro, et al.
Veröffentlicht: (2024) -
SpeechAlign: Aligning Speech Generation to Human Preferences
von: Zhang, Dong, et al.
Veröffentlicht: (2024) -
Comparative Evaluation of Expressive Japanese Character Text-to-Speech with VITS and Style-BERT-VITS2
von: Rackauckas, Zackary, et al.
Veröffentlicht: (2025) -
MSLM-S2ST: A Multitask Speech Language Model for Textless Speech-to-Speech Translation with Speaker Style Preservation
von: Peng, Yifan, et al.
Veröffentlicht: (2024)