Neuron-Level Emotion Control in Speech-Generative Large Audio-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Xiutian, Ulgen, Ismail Rasim, Koehn, Philipp, Schuller, Björn, Sisman, Berrak |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Discovering and Causally Validating Emotion-Sensitive Neurons in Large Audio-Language Models
von: Zhao, Xiutian, et al.
Veröffentlicht: (2026)
von: Zhao, Xiutian, et al.
Veröffentlicht: (2026)
HuLA: Prosody-Aware Anti-Spoofing with Multi-Task Learning for Expressive and Emotional Synthetic Speech
von: Mahapatra, Aurosweta, et al.
Veröffentlicht: (2025)
von: Mahapatra, Aurosweta, et al.
Veröffentlicht: (2025)
Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2025)
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2025)
DiffAnon: Diffusion-based Prosody Control for Voice Anonymization
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2026)
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2026)
Revealing Emotional Clusters in Speaker Embeddings: A Contrastive Learning Strategy for Speech Emotion Recognition
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2024)
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2024)
ProSDD: Learning Prosodic Representations for Speech Deepfake Detection against Expressive and Emotional Attacks
von: Mahapatra, Aurosweta, et al.
Veröffentlicht: (2026)
von: Mahapatra, Aurosweta, et al.
Veröffentlicht: (2026)
Rethinking Speaker Embeddings for Speech Generation: Sub-Center Modeling for Capturing Intra-Speaker Diversity
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2024)
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2024)
Text-to-Speech for Unseen Speakers via Low-Complexity Discrete Unit-Based Frame Selection
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2024)
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2024)
Discrete Unit based Masking for Improving Disentanglement in Voice Conversion
von: Lee, Philip H., et al.
Veröffentlicht: (2024)
von: Lee, Philip H., et al.
Veröffentlicht: (2024)
Can Emotion Fool Anti-spoofing?
von: Mahapatra, Aurosweta, et al.
Veröffentlicht: (2025)
von: Mahapatra, Aurosweta, et al.
Veröffentlicht: (2025)
Towards Naturalistic Voice Conversion: NaturalVoices Dataset with an Automatic Processing Pipeline
von: Salman, Ali N., et al.
Veröffentlicht: (2024)
von: Salman, Ali N., et al.
Veröffentlicht: (2024)
Enhancing Speech Emotion Recognition Through Differentiable Architecture Search
von: Rajapakshe, Thejan, et al.
Veröffentlicht: (2023)
von: Rajapakshe, Thejan, et al.
Veröffentlicht: (2023)
NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion
von: Du, Zongyang, et al.
Veröffentlicht: (2025)
von: Du, Zongyang, et al.
Veröffentlicht: (2025)
emoDARTS: Joint Optimisation of CNN & Sequential Neural Network Architectures for Superior Speech Emotion Recognition
von: Rajapakshe, Thejan, et al.
Veröffentlicht: (2024)
von: Rajapakshe, Thejan, et al.
Veröffentlicht: (2024)
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation
von: Tan, Weiting, et al.
Veröffentlicht: (2025)
von: Tan, Weiting, et al.
Veröffentlicht: (2025)
Explainable Speech Emotion Recognition: Weighted Attribute Fairness to Model Demographic Contributions to Social Bias
von: Ogunnubi, Tomisin, et al.
Veröffentlicht: (2026)
von: Ogunnubi, Tomisin, et al.
Veröffentlicht: (2026)
Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset
von: Liu, Rui, et al.
Veröffentlicht: (2025)
von: Liu, Rui, et al.
Veröffentlicht: (2025)
Style Mixture of Experts for Expressive Text-To-Speech Synthesis
von: Jawaid, Ahad, et al.
Veröffentlicht: (2024)
von: Jawaid, Ahad, et al.
Veröffentlicht: (2024)
Accent Conversion in Text-To-Speech Using Multi-Level VAE and Adversarial Training
von: Melechovsky, Jan, et al.
Veröffentlicht: (2024)
von: Melechovsky, Jan, et al.
Veröffentlicht: (2024)
Enhancing Emotional Text-to-Speech Controllability with Natural Language Guidance through Contrastive Learning and Diffusion Models
von: Jing, Xin, et al.
Veröffentlicht: (2024)
von: Jing, Xin, et al.
Veröffentlicht: (2024)
Can Large Language Models Aid in Annotating Speech Emotional Data? Uncovering New Frontiers
von: Latif, Siddique, et al.
Veröffentlicht: (2023)
von: Latif, Siddique, et al.
Veröffentlicht: (2023)
Exploring speech style spaces with language models: Emotional TTS without emotion labels
von: Chandra, Shreeram Suresh, et al.
Veröffentlicht: (2024)
von: Chandra, Shreeram Suresh, et al.
Veröffentlicht: (2024)
Prompting Large Language Models with Audio for General-Purpose Speech Summarization
von: Kang, Wonjune, et al.
Veröffentlicht: (2024)
von: Kang, Wonjune, et al.
Veröffentlicht: (2024)
ProsodyFM: Unsupervised Phrasing and Intonation Control for Intelligible Speech Synthesis
von: He, Xiangheng, et al.
Veröffentlicht: (2024)
von: He, Xiangheng, et al.
Veröffentlicht: (2024)
AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis
von: Qi, Tianhua, et al.
Veröffentlicht: (2026)
von: Qi, Tianhua, et al.
Veröffentlicht: (2026)
Who is Speaking or Who is Depressed? A Controlled Study of Speaker Leakage in Speech-Based Depression Detection
von: Yeh, Hsiang-Chen, et al.
Veröffentlicht: (2026)
von: Yeh, Hsiang-Chen, et al.
Veröffentlicht: (2026)
SNIPER Training: Single-Shot Sparse Training for Text-to-Speech
von: Lam, Perry, et al.
Veröffentlicht: (2022)
von: Lam, Perry, et al.
Veröffentlicht: (2022)
Accented Text-to-Speech Synthesis with a Conditional Variational Autoencoder
von: Melechovsky, Jan, et al.
Veröffentlicht: (2022)
von: Melechovsky, Jan, et al.
Veröffentlicht: (2022)
DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech
von: Melechovsky, Jan, et al.
Veröffentlicht: (2024)
von: Melechovsky, Jan, et al.
Veröffentlicht: (2024)
Evaluating Emotion Recognition in Spoken Language Models on Emotionally Incongruent Speech
von: Corrêa, Pedro, et al.
Veröffentlicht: (2025)
von: Corrêa, Pedro, et al.
Veröffentlicht: (2025)
Emotional Dimension Control in Language Model-Based Text-to-Speech: Spanning a Broad Spectrum of Human Emotions
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
Testing Correctness, Fairness, and Robustness of Speech Emotion Recognition Models
von: Derington, Anna, et al.
Veröffentlicht: (2023)
von: Derington, Anna, et al.
Veröffentlicht: (2023)
Audio-based Step-count Estimation for Running -- Windowing and Neural Network Baselines
von: Wagner, Philipp, et al.
Veröffentlicht: (2024)
von: Wagner, Philipp, et al.
Veröffentlicht: (2024)
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2025)
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2025)
Audio Large Language Models Can Be Descriptive Speech Quality Evaluators
von: Chen, Chen, et al.
Veröffentlicht: (2025)
von: Chen, Chen, et al.
Veröffentlicht: (2025)
SpeechVerse: A Large-scale Generalizable Audio Language Model
von: Das, Nilaksh, et al.
Veröffentlicht: (2024)
von: Das, Nilaksh, et al.
Veröffentlicht: (2024)
Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion Recognition
von: Zhao, Yan, et al.
Veröffentlicht: (2024)
von: Zhao, Yan, et al.
Veröffentlicht: (2024)
Controlling Emotion in Text-to-Speech with Natural Language Prompts
von: Bott, Thomas, et al.
Veröffentlicht: (2024)
von: Bott, Thomas, et al.
Veröffentlicht: (2024)
Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model
von: Du, Zongyang, et al.
Veröffentlicht: (2024)
von: Du, Zongyang, et al.
Veröffentlicht: (2024)
Audio Explanation Synthesis with Generative Foundation Models
von: Akman, Alican, et al.
Veröffentlicht: (2024)
von: Akman, Alican, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Discovering and Causally Validating Emotion-Sensitive Neurons in Large Audio-Language Models
von: Zhao, Xiutian, et al.
Veröffentlicht: (2026) -
HuLA: Prosody-Aware Anti-Spoofing with Multi-Task Learning for Expressive and Emotional Synthetic Speech
von: Mahapatra, Aurosweta, et al.
Veröffentlicht: (2025) -
Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2025) -
DiffAnon: Diffusion-based Prosody Control for Voice Anonymization
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2026) -
Revealing Emotional Clusters in Speaker Embeddings: A Contrastive Learning Strategy for Speech Emotion Recognition
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2024)