Saved in:
| Main Authors: | Naik, Canishk, Chew, Elaine |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2412.10481 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Interactive Melody Generation System for Enhancing the Creativity of Musicians
by: Hirawata, So, et al.
Published: (2024)
by: Hirawata, So, et al.
Published: (2024)
GMM-ResNext: Combining Generative and Discriminative Models for Speaker Verification
by: Yan, Hui, et al.
Published: (2024)
by: Yan, Hui, et al.
Published: (2024)
Exploring Situated Stabilities of a Rhythm Generation System through Variational Cross-Examination
by: Kotowski, Błażej, et al.
Published: (2025)
by: Kotowski, Błażej, et al.
Published: (2025)
Revisiting Your Memory: Reconstruction of Affect-Contextualized Memory via EEG-guided Audiovisual Generation
by: Kwon, Joonwoo, et al.
Published: (2024)
by: Kwon, Joonwoo, et al.
Published: (2024)
A Cross-Modal Approach to Silent Speech with LLM-Enhanced Recognition
by: Benster, Tyler, et al.
Published: (2024)
by: Benster, Tyler, et al.
Published: (2024)
STAA-Net: A Sparse and Transferable Adversarial Attack for Speech Emotion Recognition
by: Chang, Yi, et al.
Published: (2024)
by: Chang, Yi, et al.
Published: (2024)
A Theory-Based Explainable Deep Learning Architecture for Music Emotion
by: Fong, Hortense, et al.
Published: (2024)
by: Fong, Hortense, et al.
Published: (2024)
Tidal MerzA: Combining affective modelling and autonomous code generation through Reinforcement Learning
by: Wilson, Elizabeth, et al.
Published: (2024)
by: Wilson, Elizabeth, et al.
Published: (2024)
Between the AI and Me: Analysing Listeners' Perspectives on AI- and Human-Composed Progressive Metal Music
by: Sarmento, Pedro, et al.
Published: (2024)
by: Sarmento, Pedro, et al.
Published: (2024)
DeformTune: A Deformable XAI Music Prototype for Non-Musicians
by: Xu, Ziqing, et al.
Published: (2025)
by: Xu, Ziqing, et al.
Published: (2025)
Human Perception of Audio Deepfakes
by: Müller, Nicolas M., et al.
Published: (2021)
by: Müller, Nicolas M., et al.
Published: (2021)
CabinSep: IR-Augmented Mask-Based MVDR for Real-Time In-Car Speech Separation with Distributed Heterogeneous Arrays
by: Han, Runduo, et al.
Published: (2025)
by: Han, Runduo, et al.
Published: (2025)
SynthScribe: Deep Multimodal Tools for Synthesizer Sound Retrieval and Exploration
by: Brade, Stephen, et al.
Published: (2023)
by: Brade, Stephen, et al.
Published: (2023)
VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching
by: Guo, Yiwei, et al.
Published: (2023)
by: Guo, Yiwei, et al.
Published: (2023)
EvolveCaptions: Empowering DHH Users Through Real-Time Collaborative Captioning
by: Wu, Liang-Yuan, et al.
Published: (2025)
by: Wu, Liang-Yuan, et al.
Published: (2025)
Reimagining Dance: Real-time Music Co-creation between Dancers and AI
by: Vechtomova, Olga, et al.
Published: (2025)
by: Vechtomova, Olga, et al.
Published: (2025)
LSTM-CNN Network for Audio Signature Analysis in Noisy Environments
by: Damacharla, Praveen, et al.
Published: (2023)
by: Damacharla, Praveen, et al.
Published: (2023)
MCP2OSC: Parametric Control by Natural Language
by: Fan, Yuan-Yi
Published: (2025)
by: Fan, Yuan-Yi
Published: (2025)
SounDiT: Geo-Contextual Soundscape-to-Landscape Generation
by: Wang, Junbo, et al.
Published: (2025)
by: Wang, Junbo, et al.
Published: (2025)
Freetalker: Controllable Speech and Text-Driven Gesture Generation Based on Diffusion Models for Enhanced Speaker Naturalness
by: Yang, Sicheng, et al.
Published: (2024)
by: Yang, Sicheng, et al.
Published: (2024)
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
by: Huang, Ailin, et al.
Published: (2025)
by: Huang, Ailin, et al.
Published: (2025)
More-than-Human Storytelling: Designing Longitudinal Narrative Engagements with Generative AI
by: Fabre, Émilie, et al.
Published: (2025)
by: Fabre, Émilie, et al.
Published: (2025)
MetaBGM: Dynamic Soundtrack Transformation For Continuous Multi-Scene Experiences With Ambient Awareness And Personalization
by: Liu, Haoxuan, et al.
Published: (2024)
by: Liu, Haoxuan, et al.
Published: (2024)
Proceedings of The second international workshop on eXplainable AI for the Arts (XAIxArts)
by: Bryan-Kinns, Nick, et al.
Published: (2024)
by: Bryan-Kinns, Nick, et al.
Published: (2024)
G-STAR: End-to-End Global Speaker-Tracking Attributed Recognition
by: Peng, Jing, et al.
Published: (2026)
by: Peng, Jing, et al.
Published: (2026)
Music Generation using Human-In-The-Loop Reinforcement Learning
by: Justus, Aju Ani
Published: (2025)
by: Justus, Aju Ani
Published: (2025)
Zero-Shot KWS for Children's Speech using Layer-Wise Features from SSL Models
by: Kutum, Subham, et al.
Published: (2025)
by: Kutum, Subham, et al.
Published: (2025)
DiM-Gestor: Co-Speech Gesture Generation with Adaptive Layer Normalization Mamba-2
by: Zhang, Fan, et al.
Published: (2024)
by: Zhang, Fan, et al.
Published: (2024)
Cross-Lingual Speech Emotion Recognition: Humans vs. Self-Supervised Models
by: Han, Zhichen, et al.
Published: (2024)
by: Han, Zhichen, et al.
Published: (2024)
People are poorly equipped to detect AI-powered voice clones
by: Barrington, Sarah, et al.
Published: (2024)
by: Barrington, Sarah, et al.
Published: (2024)
Language Model Can Listen While Speaking
by: Ma, Ziyang, et al.
Published: (2024)
by: Ma, Ziyang, et al.
Published: (2024)
EmoKnob: Enhance Voice Cloning with Fine-Grained Emotion Control
by: Chen, Haozhe, et al.
Published: (2024)
by: Chen, Haozhe, et al.
Published: (2024)
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
by: Chen, Qian, et al.
Published: (2025)
by: Chen, Qian, et al.
Published: (2025)
Step-Audio-EditX Technical Report
by: Yan, Chao, et al.
Published: (2025)
by: Yan, Chao, et al.
Published: (2025)
AAD-LLM: Neural Attention-Driven Auditory Scene Understanding
by: Jiang, Xilin, et al.
Published: (2025)
by: Jiang, Xilin, et al.
Published: (2025)
Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models
by: Dietrich, Juergen
Published: (2026)
by: Dietrich, Juergen
Published: (2026)
Composers' Evaluations of an AI Music Tool: Insights for Human-Centred Design
by: Row, Eleanor, et al.
Published: (2024)
by: Row, Eleanor, et al.
Published: (2024)
Homogeneous Speaker Features for On-the-Fly Dysarthric and Elderly Speaker Adaptation
by: Geng, Mengzhe, et al.
Published: (2024)
by: Geng, Mengzhe, et al.
Published: (2024)
Adaptation and Optimization of Automatic Speech Recognition (ASR) for the Maritime Domain in the Field of VHF Communication
by: Nakilcioglu, Emin Cagatay, et al.
Published: (2023)
by: Nakilcioglu, Emin Cagatay, et al.
Published: (2023)
Layer-Wise Analysis of Self-Supervised Representations for Age and Gender Classification in Children's Speech
by: Sinha, Abhijit, et al.
Published: (2025)
by: Sinha, Abhijit, et al.
Published: (2025)
Similar Items
-
Interactive Melody Generation System for Enhancing the Creativity of Musicians
by: Hirawata, So, et al.
Published: (2024) -
GMM-ResNext: Combining Generative and Discriminative Models for Speaker Verification
by: Yan, Hui, et al.
Published: (2024) -
Exploring Situated Stabilities of a Rhythm Generation System through Variational Cross-Examination
by: Kotowski, Błażej, et al.
Published: (2025) -
Revisiting Your Memory: Reconstruction of Affect-Contextualized Memory via EEG-guided Audiovisual Generation
by: Kwon, Joonwoo, et al.
Published: (2024) -
A Cross-Modal Approach to Silent Speech with LLM-Enhanced Recognition
by: Benster, Tyler, et al.
Published: (2024)