The Binding Effect: Analyzing How Multi-Dimensional Cues Form Gender Bias in Instruction TTS
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Kuan-Yu, Lin, Yi-Cheng, Hsieh, Po-Chung, Chou, Huang-Cheng, Hsu, Chih-Fan, Li, Jeng-Lin, Lee, Hung-yi, Ding, Jian-Jiun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SeamlessEdit: Background Noise Aware Zero-Shot Speech Editing with in-Context Enhancement
von: Chen, Kuan-Yu, et al.
Veröffentlicht: (2025)
von: Chen, Kuan-Yu, et al.
Veröffentlicht: (2025)
Bloodroot: When Watermarking Turns Poisonous For Stealthy Backdoor
von: Chen, Kuan-Yu, et al.
Veröffentlicht: (2025)
von: Chen, Kuan-Yu, et al.
Veröffentlicht: (2025)
Online Correction of Dispersion Error in 2D Waveguide Meshes
von: Fontana, Federico, et al.
Veröffentlicht: (2000)
von: Fontana, Federico, et al.
Veröffentlicht: (2000)
Auptimize: Optimal Placement of Spatial Audio Cues for Extended Reality
von: Cho, Hyunsung, et al.
Veröffentlicht: (2024)
von: Cho, Hyunsung, et al.
Veröffentlicht: (2024)
Signal-Theoretic Characterization of Waveguide Mesh Geometries for Models of Two--Dimensional Wave Propagation in Elastic Media
von: Fontana, Federico, et al.
Veröffentlicht: (2001)
von: Fontana, Federico, et al.
Veröffentlicht: (2001)
Dichotic harmony for the musical practice
von: Madgazin, Vadim R.
Veröffentlicht: (2010)
von: Madgazin, Vadim R.
Veröffentlicht: (2010)
Masked Contrastive Pre-Training Improves Music Audio Key Detection
von: Yonay, Ori, et al.
Veröffentlicht: (2026)
von: Yonay, Ori, et al.
Veröffentlicht: (2026)
SoundPlot: An Open-Source Framework for Birdsong Acoustic Analysis and Neural Synthesis with Interactive 3D Visualization
von: Mehdi, Naqcho Ali, et al.
Veröffentlicht: (2026)
von: Mehdi, Naqcho Ali, et al.
Veröffentlicht: (2026)
Spatial Audio Rendering for Real-Time Speech Translation in Virtual Meetings
von: Geleta, Margarita, et al.
Veröffentlicht: (2025)
von: Geleta, Margarita, et al.
Veröffentlicht: (2025)
MAIN-VC: Lightweight Speech Representation Disentanglement for One-shot Voice Conversion
von: Li, Pengcheng, et al.
Veröffentlicht: (2024)
von: Li, Pengcheng, et al.
Veröffentlicht: (2024)
Unsupervised Evaluation of Deep Audio Embeddings for Music Structure Analysis
von: Marmoret, Axel
Veröffentlicht: (2026)
von: Marmoret, Axel
Veröffentlicht: (2026)
MuQ-Eval: An Open-Source Per-Sample Quality Metric for AI Music Generation Evaluation
von: Zhu, Di, et al.
Veröffentlicht: (2026)
von: Zhu, Di, et al.
Veröffentlicht: (2026)
PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching
von: Vosoughi, Ali, et al.
Veröffentlicht: (2025)
von: Vosoughi, Ali, et al.
Veröffentlicht: (2025)
Compositional Phoneme Approximation for L1-Grounded L2 Pronunciation Training
von: Park, Jisang, et al.
Veröffentlicht: (2024)
von: Park, Jisang, et al.
Veröffentlicht: (2024)
Singing Timbre Popularity Assessment Based on Multimodal Large Foundation Model
von: Wang, Zihao, et al.
Veröffentlicht: (2025)
von: Wang, Zihao, et al.
Veröffentlicht: (2025)
BEAT: Tokenizing and Generating Symbolic Music by Uniform Temporal Steps
von: Qian, Lekai, et al.
Veröffentlicht: (2026)
von: Qian, Lekai, et al.
Veröffentlicht: (2026)
If You Hold Me Without Hurting Me: Pathways to Designing Game Audio for Healthy Escapism and Player Well-being
von: Nunes, Caio, et al.
Veröffentlicht: (2025)
von: Nunes, Caio, et al.
Veröffentlicht: (2025)
Matcha-TTS: A fast TTS architecture with conditional flow matching
von: Mehta, Shivam, et al.
Veröffentlicht: (2023)
von: Mehta, Shivam, et al.
Veröffentlicht: (2023)
Guitar Tone Morphing by Diffusion-based Model
von: Chen, Kuan-Yu, et al.
Veröffentlicht: (2025)
von: Chen, Kuan-Yu, et al.
Veröffentlicht: (2025)
Audio Foundation Models Outperform Symbolic Representations for Piano Performance Evaluation
von: Dhiman, Jai
Veröffentlicht: (2026)
von: Dhiman, Jai
Veröffentlicht: (2026)
Benchmarking Sub-Genre Classification For Mainstage Dance Music
von: Shu, Hongzhi, et al.
Veröffentlicht: (2024)
von: Shu, Hongzhi, et al.
Veröffentlicht: (2024)
Generation of Musical Timbres using a Text-Guided Diffusion Model
von: Yuan, Weixuan, et al.
Veröffentlicht: (2025)
von: Yuan, Weixuan, et al.
Veröffentlicht: (2025)
Self-Improvement for Audio Large Language Model using Unlabeled Speech
von: Wang, Shaowen, et al.
Veröffentlicht: (2025)
von: Wang, Shaowen, et al.
Veröffentlicht: (2025)
OBHS: An Optimized Block Huffman Scheme for Real-Time Audio Compression
von: Mahfi, Muntahi Safwan, et al.
Veröffentlicht: (2025)
von: Mahfi, Muntahi Safwan, et al.
Veröffentlicht: (2025)
A Decomposed Retrieval-Edit-Rerank Framework for Chord Generation
von: He, Qiqi, et al.
Veröffentlicht: (2026)
von: He, Qiqi, et al.
Veröffentlicht: (2026)
Quantum-Enhanced Analysis and Grading of Vocal Performance
von: Agarwal, Rohan
Veröffentlicht: (2025)
von: Agarwal, Rohan
Veröffentlicht: (2025)
ParaNoise-SV: Integrated Approach for Noise-Robust Speaker Verification with Parallel Joint Learning of Speech Enhancement and Noise Extraction
von: Kim, Minu, et al.
Veröffentlicht: (2025)
von: Kim, Minu, et al.
Veröffentlicht: (2025)
Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech
von: Mehta, Shivam, et al.
Veröffentlicht: (2024)
von: Mehta, Shivam, et al.
Veröffentlicht: (2024)
Can pre-trained Deep Learning models predict groove ratings?
von: Marmoret, Axel, et al.
Veröffentlicht: (2026)
von: Marmoret, Axel, et al.
Veröffentlicht: (2026)
Improving French Synthetic Speech Quality via SSML Prosody Control
von: Ouali, Nassima Ould, et al.
Veröffentlicht: (2025)
von: Ouali, Nassima Ould, et al.
Veröffentlicht: (2025)
Acoustic Wave Modeling Using 2D FDTD: Applications in Unreal Engine For Dynamic Sound Rendering
von: Samsurya, Bilkent
Veröffentlicht: (2025)
von: Samsurya, Bilkent
Veröffentlicht: (2025)
Refining music sample identification with a self-supervised graph neural network
von: Bhattacharjee, Aditya, et al.
Veröffentlicht: (2025)
von: Bhattacharjee, Aditya, et al.
Veröffentlicht: (2025)
Neural Proxies for Sound Synthesizers: Learning Perceptually Informed Preset Representations
von: Combes, Paolo, et al.
Veröffentlicht: (2025)
von: Combes, Paolo, et al.
Veröffentlicht: (2025)
AudioMiXR: Spatial Audio Object Manipulation with 6DoF for Sound Design in Augmented Reality
von: Woodard, Brandon, et al.
Veröffentlicht: (2025)
von: Woodard, Brandon, et al.
Veröffentlicht: (2025)
Less Stress, More Privacy: Stress Detection on Anonymized Speech of Air Traffic Controllers
von: Viswanathan, Janaki, et al.
Veröffentlicht: (2025)
von: Viswanathan, Janaki, et al.
Veröffentlicht: (2025)
SFMS-ALR: Script-First Multilingual Speech Synthesis with Adaptive Locale Resolution
von: Donepudi, Dharma Teja
Veröffentlicht: (2025)
von: Donepudi, Dharma Teja
Veröffentlicht: (2025)
REMAST: Real-time Emotion-based Music Arrangement with Soft Transition
von: Wang, Zihao, et al.
Veröffentlicht: (2023)
von: Wang, Zihao, et al.
Veröffentlicht: (2023)
GraFPrint: A GNN-Based Approach for Audio Identification
von: Bhattacharjee, Aditya, et al.
Veröffentlicht: (2024)
von: Bhattacharjee, Aditya, et al.
Veröffentlicht: (2024)
Embodied Exploration of Latent Spaces and Explainable AI
von: Wilson, Elizabeth, et al.
Veröffentlicht: (2024)
von: Wilson, Elizabeth, et al.
Veröffentlicht: (2024)
Two Sonification Methods for the MindCube
von: Liu, Fangzheng, et al.
Veröffentlicht: (2025)
von: Liu, Fangzheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SeamlessEdit: Background Noise Aware Zero-Shot Speech Editing with in-Context Enhancement
von: Chen, Kuan-Yu, et al.
Veröffentlicht: (2025) -
Bloodroot: When Watermarking Turns Poisonous For Stealthy Backdoor
von: Chen, Kuan-Yu, et al.
Veröffentlicht: (2025) -
Online Correction of Dispersion Error in 2D Waveguide Meshes
von: Fontana, Federico, et al.
Veröffentlicht: (2000) -
Auptimize: Optimal Placement of Spatial Audio Cues for Extended Reality
von: Cho, Hyunsung, et al.
Veröffentlicht: (2024) -
Signal-Theoretic Characterization of Waveguide Mesh Geometries for Models of Two--Dimensional Wave Propagation in Elastic Media
von: Fontana, Federico, et al.
Veröffentlicht: (2001)