STASE: A spatialized text-to-audio synthesis engine for music generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Chi, Tutti, Gao, Letian, Zhang, Yixiao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FxSearcher: gradient-free text-driven audio transformation
por: Ki, Hojoon, et al.
Publicado: (2025)
por: Ki, Hojoon, et al.
Publicado: (2025)
EDTC: enhance depth of text comprehension in automated audio captioning
por: Tan, Liwen, et al.
Publicado: (2024)
por: Tan, Liwen, et al.
Publicado: (2024)
Visual-based spatial audio generation system for multi-speaker environments
por: Liu, Xiaojing, et al.
Publicado: (2025)
por: Liu, Xiaojing, et al.
Publicado: (2025)
Synthetic training set generation using text-to-audio models for environmental sound classification
por: Ronchini, Francesca, et al.
Publicado: (2024)
por: Ronchini, Francesca, et al.
Publicado: (2024)
A SOUND APPROACH: Using Large Language Models to generate audio descriptions for egocentric text-audio retrieval
por: Oncescu, Andreea-Maria, et al.
Publicado: (2024)
por: Oncescu, Andreea-Maria, et al.
Publicado: (2024)
TEAdapter: Supply abundant guidance for controllable text-to-music generation
por: Zou, Jialing, et al.
Publicado: (2024)
por: Zou, Jialing, et al.
Publicado: (2024)
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio
por: Kanamori, Yusuke, et al.
Publicado: (2025)
por: Kanamori, Yusuke, et al.
Publicado: (2025)
ICGAN: An implicit conditioning method for interpretable feature control of neural audio synthesis
por: Liu, Yunyi, et al.
Publicado: (2024)
por: Liu, Yunyi, et al.
Publicado: (2024)
Deep learning based spatial aliasing reduction in beamforming for audio capture
por: Guzik, Mateusz, et al.
Publicado: (2025)
por: Guzik, Mateusz, et al.
Publicado: (2025)
DashengTokenizer: One layer is enough for unified audio understanding and generation
por: Dinkel, Heinrich, et al.
Publicado: (2026)
por: Dinkel, Heinrich, et al.
Publicado: (2026)
Expressive paragraph text-to-speech synthesis with multi-step variational autoencoder
por: Li, Xuyuan, et al.
Publicado: (2023)
por: Li, Xuyuan, et al.
Publicado: (2023)
Semi-intrusive audio evaluation: Casting non-intrusive assessment as a multi-modal text prediction task
por: Coldenhoff, Jozef, et al.
Publicado: (2024)
por: Coldenhoff, Jozef, et al.
Publicado: (2024)
Scaling up masked audio encoder learning for general audio classification
por: Dinkel, Heinrich, et al.
Publicado: (2024)
por: Dinkel, Heinrich, et al.
Publicado: (2024)
PAGURI: a user experience study of creative interaction with text-to-music models
por: Ronchini, Francesca, et al.
Publicado: (2024)
por: Ronchini, Francesca, et al.
Publicado: (2024)
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models
por: Zhang, Yixiao
Publicado: (2024)
por: Zhang, Yixiao
Publicado: (2024)
DualSpec: Text-to-spatial-audio Generation via Dual-Spectrogram Guided Diffusion Model
por: Zhao, Lei, et al.
Publicado: (2025)
por: Zhao, Lei, et al.
Publicado: (2025)
$\text{M}^3\text{PDB}$: A Multimodal, Multi-Label, Multilingual Prompt Database for Speech Generation
por: Zhu, Boyu, et al.
Publicado: (2025)
por: Zhu, Boyu, et al.
Publicado: (2025)
Audio Dialogues: Dialogues dataset for audio and music understanding
por: Goel, Arushi, et al.
Publicado: (2024)
por: Goel, Arushi, et al.
Publicado: (2024)
Exploring compressibility of transformer based text-to-music (TTM) models
por: Moschopoulos, Vasileios, et al.
Publicado: (2024)
por: Moschopoulos, Vasileios, et al.
Publicado: (2024)
GLAP: General contrastive audio-text pretraining across domains and languages
por: Dinkel, Heinrich, et al.
Publicado: (2025)
por: Dinkel, Heinrich, et al.
Publicado: (2025)
AEROMamba: An efficient architecture for audio super-resolution using generative adversarial networks and state space models
por: Abreu, Wallace, et al.
Publicado: (2024)
por: Abreu, Wallace, et al.
Publicado: (2024)
Mitigating data replication in text-to-audio generative diffusion models through anti-memorization guidance
por: Messina, Francisco, et al.
Publicado: (2025)
por: Messina, Francisco, et al.
Publicado: (2025)
MBCodec:Thorough disentangle for high-fidelity audio compression
por: Zhang, Ruonan, et al.
Publicado: (2025)
por: Zhang, Ruonan, et al.
Publicado: (2025)
Towards audio language modeling -- an overview
por: Wu, Haibin, et al.
Publicado: (2024)
por: Wu, Haibin, et al.
Publicado: (2024)
Are audio DeepFake detection models polyglots?
por: Marek, Bartłomiej, et al.
Publicado: (2024)
por: Marek, Bartłomiej, et al.
Publicado: (2024)
MusicGen-Stem: Multi-stem music generation and edition through autoregressive modeling
por: Rouard, Simon, et al.
Publicado: (2025)
por: Rouard, Simon, et al.
Publicado: (2025)
Real-time implementation of vibrato transfer as an audio effect
por: Hyrkas, Jeremy
Publicado: (2025)
por: Hyrkas, Jeremy
Publicado: (2025)
Tweaking autoregressive methods for inpainting of gaps in audio signals
por: Mokrý, Ondřej, et al.
Publicado: (2024)
por: Mokrý, Ondřej, et al.
Publicado: (2024)
TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet
por: Jeong, Jaeseok, et al.
Publicado: (2025)
por: Jeong, Jaeseok, et al.
Publicado: (2025)
ACAVCaps: Enabling large-scale training for fine-grained and diverse audio understanding
por: Niu, Yadong, et al.
Publicado: (2026)
por: Niu, Yadong, et al.
Publicado: (2026)
Speaker anonymization using neural audio codec language models
por: Panariello, Michele, et al.
Publicado: (2023)
por: Panariello, Michele, et al.
Publicado: (2023)
Regularized autoregressive modeling and its application to audio signal reconstruction
por: Mokrý, Ondřej, et al.
Publicado: (2024)
por: Mokrý, Ondřej, et al.
Publicado: (2024)
A robust audio deepfake detection system via multi-view feature
por: Yang, Yujie, et al.
Publicado: (2024)
por: Yang, Yujie, et al.
Publicado: (2024)
Testing chatbots on the creation of encoders for audio conditioned image generation
por: León, Jorge E., et al.
Publicado: (2025)
por: León, Jorge E., et al.
Publicado: (2025)
A tunable binaural audio telepresence system capable of balancing immersive and enhanced modes
por: Hsu, Yicheng, et al.
Publicado: (2024)
por: Hsu, Yicheng, et al.
Publicado: (2024)
Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models
por: Wu, Haibin, et al.
Publicado: (2024)
por: Wu, Haibin, et al.
Publicado: (2024)
Human-CLAP: Human-perception-based contrastive language-audio pretraining
por: Takano, Taisei, et al.
Publicado: (2025)
por: Takano, Taisei, et al.
Publicado: (2025)
AxLSTMs: learning self-supervised audio representations with xLSTMs
por: Yadav, Sarthak, et al.
Publicado: (2024)
por: Yadav, Sarthak, et al.
Publicado: (2024)
Enhancement by postfiltering for speech and audio coding in ad-hoc sensor networks
por: Das, Sneha, et al.
Publicado: (2020)
por: Das, Sneha, et al.
Publicado: (2020)
Modeling strategies for speech enhancement in the latent space of a neural audio codec
por: Kammoun, Sofiene, et al.
Publicado: (2025)
por: Kammoun, Sofiene, et al.
Publicado: (2025)
Ejemplares similares
-
FxSearcher: gradient-free text-driven audio transformation
por: Ki, Hojoon, et al.
Publicado: (2025) -
EDTC: enhance depth of text comprehension in automated audio captioning
por: Tan, Liwen, et al.
Publicado: (2024) -
Visual-based spatial audio generation system for multi-speaker environments
por: Liu, Xiaojing, et al.
Publicado: (2025) -
Synthetic training set generation using text-to-audio models for environmental sound classification
por: Ronchini, Francesca, et al.
Publicado: (2024) -
A SOUND APPROACH: Using Large Language Models to generate audio descriptions for egocentric text-audio retrieval
por: Oncescu, Andreea-Maria, et al.
Publicado: (2024)