Making Flow-Matching-Based Zero-Shot Text-to-Speech Laugh as You Like
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kanda, Naoyuki, Wang, Xiaofei, Eskimez, Sefik Emre, Thakker, Manthan, Yang, Hemin, Zhu, Zirun, Tang, Min, Li, Canrun, Tsai, Chung-Hsien, Xiao, Zhen, Xia, Yufei, Li, Jinzhu, Liu, Yanqing, Zhao, Sheng, Zeng, Michael |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Laugh Now Cry Later: Controlling Time-Varying Emotional States of Flow-Matching-Based Zero-Shot Text-to-Speech
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
An Investigation of Noise Robustness for Flow-Matching-Based Zero-Shot TTS
von: Wang, Xiaofei, et al.
Veröffentlicht: (2024)
von: Wang, Xiaofei, et al.
Veröffentlicht: (2024)
Total-Duration-Aware Duration Modeling for Text-to-Speech Systems
von: Eskimez, Sefik Emre, et al.
Veröffentlicht: (2024)
von: Eskimez, Sefik Emre, et al.
Veröffentlicht: (2024)
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
von: Eskimez, Sefik Emre, et al.
Veröffentlicht: (2024)
von: Eskimez, Sefik Emre, et al.
Veröffentlicht: (2024)
SpeechX: Neural Codec Language Model as a Versatile Speech Transformer
von: Wang, Xiaofei, et al.
Veröffentlicht: (2023)
von: Wang, Xiaofei, et al.
Veröffentlicht: (2023)
TS3-Codec: Transformer-Based Simple Streaming Single Codec
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
Neural Speech Extraction with Human Feedback
von: Itani, Malek, et al.
Veröffentlicht: (2025)
von: Itani, Malek, et al.
Veröffentlicht: (2025)
Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
von: Li, Jiaqi, et al.
Veröffentlicht: (2024)
von: Li, Jiaqi, et al.
Veröffentlicht: (2024)
CoVoMix2: Advancing Zero-Shot Dialogue Generation with Fully Non-Autoregressive Flow Matching
von: Zhang, Leying, et al.
Veröffentlicht: (2025)
von: Zhang, Leying, et al.
Veröffentlicht: (2025)
Knowledge boosting during low-latency inference
von: Srinivas, Vidya, et al.
Veröffentlicht: (2024)
von: Srinivas, Vidya, et al.
Veröffentlicht: (2024)
Improving Practical Aspects of End-to-End Multi-Talker Speech Recognition for Online and Offline Scenarios
von: Subramanian, Aswin Shanmugam, et al.
Veröffentlicht: (2025)
von: Subramanian, Aswin Shanmugam, et al.
Veröffentlicht: (2025)
Target conversation extraction: Source separation using turn-taking dynamics
von: Chen, Tuochao, et al.
Veröffentlicht: (2024)
von: Chen, Tuochao, et al.
Veröffentlicht: (2024)
Laughing matters. You think that's funny?
Veröffentlicht: (1997)
Veröffentlicht: (1997)
DiariST: Streaming Speech Translation with Speaker Diarization
von: Yang, Mu, et al.
Veröffentlicht: (2023)
von: Yang, Mu, et al.
Veröffentlicht: (2023)
Can You Make Me Laugh? Toddlers’ and Parents’ Shared Positive Expressions in Playful Interactions
von: Anja Gampe, et al.
Veröffentlicht: (2025)
von: Anja Gampe, et al.
Veröffentlicht: (2025)
FlexiCodec: A Dynamic Neural Audio Codec for Low Frame Rates
von: Li, Jiaqi, et al.
Veröffentlicht: (2025)
von: Li, Jiaqi, et al.
Veröffentlicht: (2025)
Author! Author! Making Kids Laugh: Jon Scieszka
von: Brodie, Carolyn S.
Veröffentlicht: (2004)
von: Brodie, Carolyn S.
Veröffentlicht: (2004)
What Makes Programmers Laugh? Exploring the Submissions of the Subreddit r/ProgrammerHumor
von: Kuutila, Miikka, et al.
Veröffentlicht: (2024)
von: Kuutila, Miikka, et al.
Veröffentlicht: (2024)
Make 'Em Laugh: A Different Approach to Library Orientation.
von: Liebman, Roy
Veröffentlicht: (1980)
von: Liebman, Roy
Veröffentlicht: (1980)
Can Language Models Laugh at YouTube Short-form Videos?
von: Ko, Dayoon, et al.
Veröffentlicht: (2023)
von: Ko, Dayoon, et al.
Veröffentlicht: (2023)
The Laugh of the Tramp
von: Damian Maher
Veröffentlicht: (2025)
von: Damian Maher
Veröffentlicht: (2025)
Streaming Speaker Change Detection and Gender Classification for Transducer-Based Multi-Talker Speech Translation
von: Wang, Peidong, et al.
Veröffentlicht: (2025)
von: Wang, Peidong, et al.
Veröffentlicht: (2025)
The Set of Stable Matchings and the Core in a Matching Market with Ties and Matroid Constraints
von: Kamiyama, Naoyuki
Veröffentlicht: (2024)
von: Kamiyama, Naoyuki
Veröffentlicht: (2024)
Non-uniformly Stable Matchings
von: Kamiyama, Naoyuki
Veröffentlicht: (2024)
von: Kamiyama, Naoyuki
Veröffentlicht: (2024)
The Strongly Stable Matching Problem with Closures
von: Kamiyama, Naoyuki
Veröffentlicht: (2024)
von: Kamiyama, Naoyuki
Veröffentlicht: (2024)
Modifying an Instance of the Super-Stable Matching Problem
von: Kamiyama, Naoyuki
Veröffentlicht: (2024)
von: Kamiyama, Naoyuki
Veröffentlicht: (2024)
Spontaneous Perovskite Passivators Effectively Combined with PTAA Hole‐Transport Materials in Perovskite Solar Cells
von: Naoyuki Nishimura, et al.
Veröffentlicht: (2025)
von: Naoyuki Nishimura, et al.
Veröffentlicht: (2025)
Decoders Laugh as Loud as Encoders
von: Borodach, Eli, et al.
Veröffentlicht: (2025)
von: Borodach, Eli, et al.
Veröffentlicht: (2025)
Making the Most of the One-Shot You Got
von: Deemer, Kevin
Veröffentlicht: (2007)
von: Deemer, Kevin
Veröffentlicht: (2007)
Summary of the NOTSOFAR-1 Challenge: Highlights and Learnings
von: Abramovski, Igor, et al.
Veröffentlicht: (2025)
von: Abramovski, Igor, et al.
Veröffentlicht: (2025)
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners
von: Yuan, Ze, et al.
Veröffentlicht: (2024)
von: Yuan, Ze, et al.
Veröffentlicht: (2024)
Inform Me or Make Me Laugh: Influencer Marketing and the Role of Emotional Contagion and Information Value
von: Payal S. Kapoor, et al.
Veröffentlicht: (2024)
von: Payal S. Kapoor, et al.
Veröffentlicht: (2024)
Citizen‐Centered Public Service Design in Agile Digital Transformation: Insights From Public Mobility Services
von: Hemin Choi, et al.
Veröffentlicht: (2026)
von: Hemin Choi, et al.
Veröffentlicht: (2026)
Conversational Speech Naturalness Predictor
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
We Should Laugh So Long?
von: Nilsen, Alleen Pace
Veröffentlicht: (1986)
von: Nilsen, Alleen Pace
Veröffentlicht: (1986)
Laugh, Relate, Engage: Stylized Comment Generation for Short Videos
von: Ouyang, Xuan, et al.
Veröffentlicht: (2025)
von: Ouyang, Xuan, et al.
Veröffentlicht: (2025)
Unimodal Aggregation for CTC-based Speech Recognition
von: Fang, Ying, et al.
Veröffentlicht: (2023)
von: Fang, Ying, et al.
Veröffentlicht: (2023)
Tail Moment for Gamma‐like Risks With Arbitrary Scale Parameters
von: Bingjie Wang, et al.
Veröffentlicht: (2026)
von: Bingjie Wang, et al.
Veröffentlicht: (2026)
LightDSA: A Python-Based Hybrid Digital Signature Library and Performance Analysis of RSA, DSA, ECDSA and EdDSA in Variable Configurations, Elliptic Curve Forms and Curves
von: Serengil, Sefik, et al.
Veröffentlicht: (2025)
von: Serengil, Sefik, et al.
Veröffentlicht: (2025)
CipherFace: A Fully Homomorphic Encryption-Driven Framework for Secure Cloud-Based Facial Recognition
von: Serengil, Sefik, et al.
Veröffentlicht: (2025)
von: Serengil, Sefik, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Laugh Now Cry Later: Controlling Time-Varying Emotional States of Flow-Matching-Based Zero-Shot Text-to-Speech
von: Wu, Haibin, et al.
Veröffentlicht: (2024) -
An Investigation of Noise Robustness for Flow-Matching-Based Zero-Shot TTS
von: Wang, Xiaofei, et al.
Veröffentlicht: (2024) -
Total-Duration-Aware Duration Modeling for Text-to-Speech Systems
von: Eskimez, Sefik Emre, et al.
Veröffentlicht: (2024) -
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
von: Eskimez, Sefik Emre, et al.
Veröffentlicht: (2024) -
SpeechX: Neural Codec Language Model as a Versatile Speech Transformer
von: Wang, Xiaofei, et al.
Veröffentlicht: (2023)