Guardado en:
| Autores principales: | Zhang, Ellie L., Liao, Duoduo, Liao, Callie C. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2511.19275 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Relationships between Keywords and Strong Beats in Lyrical Music
por: Liao, Callie C., et al.
Publicado: (2024)
por: Liao, Callie C., et al.
Publicado: (2024)
IS${}^3$ : Generic Impulsive--Stationary Sound Separation in Acoustic Scenes using Deep Filtering
por: Berger, Clémentine, et al.
Publicado: (2025)
por: Berger, Clémentine, et al.
Publicado: (2025)
Multimodal Lyrics-Rhythm Matching
por: Liao, Callie C., et al.
Publicado: (2023)
por: Liao, Callie C., et al.
Publicado: (2023)
Resounding Acoustic Fields with Reciprocity
por: Lan, Zitong, et al.
Publicado: (2025)
por: Lan, Zitong, et al.
Publicado: (2025)
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
por: Wang, Yuancheng, et al.
Publicado: (2025)
por: Wang, Yuancheng, et al.
Publicado: (2025)
Acoustic Imaging for UAV Detection: Dense Beamformed Energy Maps and U-Net SELD
por: Rodriguez, Belman Jahir, et al.
Publicado: (2025)
por: Rodriguez, Belman Jahir, et al.
Publicado: (2025)
TinyChirp: Bird Song Recognition Using TinyML Models on Low-power Wireless Acoustic Sensors
por: Huang, Zhaolan, et al.
Publicado: (2024)
por: Huang, Zhaolan, et al.
Publicado: (2024)
SWIM: Short-Window CNN Integrated with Mamba for EEG-Based Auditory Spatial Attention Decoding
por: Zhang, Ziyang, et al.
Publicado: (2024)
por: Zhang, Ziyang, et al.
Publicado: (2024)
Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration with Improved Intelligibility
por: Liu, Xiaoyu, et al.
Publicado: (2024)
por: Liu, Xiaoyu, et al.
Publicado: (2024)
Generating Moving 3D Soundscapes with Latent Diffusion Models
por: Templin, Christian, et al.
Publicado: (2025)
por: Templin, Christian, et al.
Publicado: (2025)
Soundscape Captioning using Sound Affective Quality Network and Large Language Model
por: Hou, Yuanbo, et al.
Publicado: (2024)
por: Hou, Yuanbo, et al.
Publicado: (2024)
AI-Generated Music Detection in Broadcast Monitoring
por: López-Ayala, David, et al.
Publicado: (2026)
por: López-Ayala, David, et al.
Publicado: (2026)
Classification of Heart Sounds Using Multi-Branch Deep Convolutional Network and LSTM-CNN
por: Latifi, Seyed Amir, et al.
Publicado: (2024)
por: Latifi, Seyed Amir, et al.
Publicado: (2024)
Mind the Prompt: Prompting Strategies in Audio Generations for Improving Sound Classification
por: Ronchini, Francesca, et al.
Publicado: (2025)
por: Ronchini, Francesca, et al.
Publicado: (2025)
EEG-Based Speech Decoding: A Novel Approach Using Multi-Kernel Ensemble Diffusion Models
por: Kim, Soowon, et al.
Publicado: (2024)
por: Kim, Soowon, et al.
Publicado: (2024)
CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation
por: Kim, Ji-Hoon, et al.
Publicado: (2024)
por: Kim, Ji-Hoon, et al.
Publicado: (2024)
Towards Improving Speaker Distance Estimation through Generative Impulse Response Augmentation
por: Ratnarajah, Anton, et al.
Publicado: (2026)
por: Ratnarajah, Anton, et al.
Publicado: (2026)
Robust Generative Audio Quality Assessment: Disentangling Quality from Spurious Correlations
por: Huang, Kuan-Tang, et al.
Publicado: (2026)
por: Huang, Kuan-Tang, et al.
Publicado: (2026)
CSL-L2M: Controllable Song-Level Lyric-to-Melody Generation Based on Conditional Transformer with Fine-Grained Lyric and Musical Controls
por: Chai, Li, et al.
Publicado: (2024)
por: Chai, Li, et al.
Publicado: (2024)
Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions
por: Gao, Xiaoxue, et al.
Publicado: (2025)
por: Gao, Xiaoxue, et al.
Publicado: (2025)
Construction and Evaluation of Mandarin Multimodal Emotional Speech Database
por: Ting, Zhu, et al.
Publicado: (2024)
por: Ting, Zhu, et al.
Publicado: (2024)
Expressive Acoustic Guitar Sound Synthesis with an Instrument-Specific Input Representation and Diffusion Outpainting
por: Kim, Hounsu, et al.
Publicado: (2024)
por: Kim, Hounsu, et al.
Publicado: (2024)
When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds
por: Kang, Minsu, et al.
Publicado: (2025)
por: Kang, Minsu, et al.
Publicado: (2025)
Compressing Quaternion Convolutional Neural Networks for Audio Classification
por: Singh, Arshdeep, et al.
Publicado: (2025)
por: Singh, Arshdeep, et al.
Publicado: (2025)
Toward Fully-End-to-End Listened Speech Decoding from EEG Signals
por: Lee, Jihwan, et al.
Publicado: (2024)
por: Lee, Jihwan, et al.
Publicado: (2024)
Neural Speech and Audio Coding: Modern AI Technology Meets Traditional Codecs
por: Kim, Minje, et al.
Publicado: (2024)
por: Kim, Minje, et al.
Publicado: (2024)
Differentiable Modal Synthesis for Physical Modeling of Planar String Sound and Motion Simulation
por: Lee, Jin Woo, et al.
Publicado: (2024)
por: Lee, Jin Woo, et al.
Publicado: (2024)
Towards Low-Latency Tracking of Multiple Speakers With Short-Context Speaker Embeddings
por: Iatariene, Taous, et al.
Publicado: (2025)
por: Iatariene, Taous, et al.
Publicado: (2025)
Automatic Speech Recognition using Advanced Deep Learning Approaches: A survey
por: Kheddar, Hamza, et al.
Publicado: (2024)
por: Kheddar, Hamza, et al.
Publicado: (2024)
U-DREAM: Unsupervised Dereverberation guided by a Reverberation Model
por: Bahrman, Louis, et al.
Publicado: (2025)
por: Bahrman, Louis, et al.
Publicado: (2025)
Speech Enhancement Based on Drifting Models
por: Xu, Liang, et al.
Publicado: (2026)
por: Xu, Liang, et al.
Publicado: (2026)
A Hybrid Model for Weakly-Supervised Speech Dereverberation
por: Bahrman, Louis, et al.
Publicado: (2025)
por: Bahrman, Louis, et al.
Publicado: (2025)
Perceptual Noise-Masking with Music through Deep Spectral Envelope Shaping
por: Berger, Clémentine, et al.
Publicado: (2025)
por: Berger, Clémentine, et al.
Publicado: (2025)
Wavetable Synthesis Using CVAE for Timbre Control Based on Semantic Label
por: Yutani, Tsugumasa, et al.
Publicado: (2024)
por: Yutani, Tsugumasa, et al.
Publicado: (2024)
Speech Boosting: Low-Latency Live Speech Enhancement for TWS Earbuds
por: Bae, Hanbin, et al.
Publicado: (2024)
por: Bae, Hanbin, et al.
Publicado: (2024)
Single-stage TTS with Masked Audio Token Modeling and Semantic Knowledge Distillation
por: Gállego, Gerard I., et al.
Publicado: (2024)
por: Gállego, Gerard I., et al.
Publicado: (2024)
Tool Wear Prediction in CNC Turning Operations using Ultrasonic Microphone Arrays and CNNs
por: Steckel, Jan, et al.
Publicado: (2024)
por: Steckel, Jan, et al.
Publicado: (2024)
VoicePrompter: Robust Zero-Shot Voice Conversion with Voice Prompt and Conditional Flow Matching
por: Choi, Ha-Yeong, et al.
Publicado: (2025)
por: Choi, Ha-Yeong, et al.
Publicado: (2025)
A Domain-Knowledge-Inspired Music Embedding Space and a Novel Attention Mechanism for Symbolic Music Modeling
por: Guo, Z., et al.
Publicado: (2022)
por: Guo, Z., et al.
Publicado: (2022)
Discriminating real and synthetic super-resolved audio samples using embedding-based classifiers
por: Silaev, Mikhail, et al.
Publicado: (2026)
por: Silaev, Mikhail, et al.
Publicado: (2026)
Ejemplares similares
-
Relationships between Keywords and Strong Beats in Lyrical Music
por: Liao, Callie C., et al.
Publicado: (2024) -
IS${}^3$ : Generic Impulsive--Stationary Sound Separation in Acoustic Scenes using Deep Filtering
por: Berger, Clémentine, et al.
Publicado: (2025) -
Multimodal Lyrics-Rhythm Matching
por: Liao, Callie C., et al.
Publicado: (2023) -
Resounding Acoustic Fields with Reciprocity
por: Lan, Zitong, et al.
Publicado: (2025) -
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
por: Wang, Yuancheng, et al.
Publicado: (2025)