PAGURI: a user experience study of creative interaction with text-to-music models
Fuente:
arXiv
Saved in:
| Main Authors: | Ronchini, Francesca, Comanducci, Luca, Perego, Gabriele, Antonacci, Fabio |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Synthetic training set generation using text-to-audio models for environmental sound classification
by: Ronchini, Francesca, et al.
Published: (2024)
by: Ronchini, Francesca, et al.
Published: (2024)
MambaFoley: Foley Sound Generation using Selective State-Space Models
by: Colombo, Marco Furio, et al.
Published: (2024)
by: Colombo, Marco Furio, et al.
Published: (2024)
Mitigating data replication in text-to-audio generative diffusion models through anti-memorization guidance
by: Messina, Francisco, et al.
Published: (2025)
by: Messina, Francisco, et al.
Published: (2025)
AI-Assisted Music Production: A User Study on Text-to-Music Models
by: Ronchini, Francesca, et al.
Published: (2025)
by: Ronchini, Francesca, et al.
Published: (2025)
Room Transfer Function Reconstruction Using Complex-valued Neural Networks and Irregularly Distributed Microphones
by: Ronchini, Francesca, et al.
Published: (2024)
by: Ronchini, Francesca, et al.
Published: (2024)
Interpreting End-to-End Deep Learning Models for Speech Source Localization Using Layer-wise Relevance Propagation
by: Comanducci, Luca, et al.
Published: (2024)
by: Comanducci, Luca, et al.
Published: (2024)
Diffused Responsibility: Analyzing the Energy Consumption of Generative Text-to-Audio Diffusion Models
by: Passoni, Riccardo, et al.
Published: (2025)
by: Passoni, Riccardo, et al.
Published: (2025)
Towards HRTF Personalization using Denoising Diffusion Models
by: Sánchez, Juan Camilo Albarracín, et al.
Published: (2025)
by: Sánchez, Juan Camilo Albarracín, et al.
Published: (2025)
Performance and energy balance: a comprehensive study of state-of-the-art sound event detection systems
by: Ronchini, Francesca, et al.
Published: (2023)
by: Ronchini, Francesca, et al.
Published: (2023)
Mind the Prompt: Prompting Strategies in Audio Generations for Improving Sound Classification
by: Ronchini, Francesca, et al.
Published: (2025)
by: Ronchini, Francesca, et al.
Published: (2025)
FakeMusicCaps: a Dataset for Detection and Attribution of Synthetic Music Generated via Text-to-Music Models
by: Comanducci, Luca, et al.
Published: (2024)
by: Comanducci, Luca, et al.
Published: (2024)
Reconstruction of Sound Field through Diffusion Models
by: Miotello, Federico, et al.
Published: (2023)
by: Miotello, Federico, et al.
Published: (2023)
Synthesis of Soundfields through Irregular Loudspeaker Arrays Based on Convolutional Neural Networks
by: Comanducci, Luca, et al.
Published: (2022)
by: Comanducci, Luca, et al.
Published: (2022)
VR-PTOLEMAIC: A Virtual Environment for the Perceptual Testing of Spatial Audio Algorithms
by: Ostan, Paolo, et al.
Published: (2025)
by: Ostan, Paolo, et al.
Published: (2025)
Sound event localization and detection based on crnn using rectangular filters and channel rotation data augmentation
by: Ronchini, Francesca, et al.
Published: (2020)
by: Ronchini, Francesca, et al.
Published: (2020)
A benchmark of state-of-the-art sound event detection systems evaluated on synthetic soundscapes
by: Ronchini, Francesca, et al.
Published: (2022)
by: Ronchini, Francesca, et al.
Published: (2022)
Low-Rank Adaptation of Deep Prior Neural Networks For Room Impulse Response Reconstruction
by: Pezzoli, Mirco, et al.
Published: (2025)
by: Pezzoli, Mirco, et al.
Published: (2025)
Acoustic source localization in the spherical harmonics domain exploiting low-rank approximations
by: Cobos, Maximo, et al.
Published: (2023)
by: Cobos, Maximo, et al.
Published: (2023)
Physics-Informed Transfer Learning for Data-Driven Sound Source Reconstruction in Near-Field Acoustic Holography
by: Luan, Xinmeng, et al.
Published: (2025)
by: Luan, Xinmeng, et al.
Published: (2025)
Exploring compressibility of transformer based text-to-music (TTM) models
by: Moschopoulos, Vasileios, et al.
Published: (2024)
by: Moschopoulos, Vasileios, et al.
Published: (2024)
STASE: A spatialized text-to-audio synthesis engine for music generation
by: Chi, Tutti, et al.
Published: (2025)
by: Chi, Tutti, et al.
Published: (2025)
TEAdapter: Supply abundant guidance for controllable text-to-music generation
by: Zou, Jialing, et al.
Published: (2024)
by: Zou, Jialing, et al.
Published: (2024)
Description and analysis of novelties introduced in DCASE Task 4 2022 on the baseline system
by: Ronchini, Francesca, et al.
Published: (2022)
by: Ronchini, Francesca, et al.
Published: (2022)
Past, Present, and Future of Spatial Audio and Room Acoustics
by: Koyama, Shoichi, et al.
Published: (2025)
by: Koyama, Shoichi, et al.
Published: (2025)
The impact of non-target events in synthetic soundscapes for sound event detection
by: Ronchini, Francesca, et al.
Published: (2021)
by: Ronchini, Francesca, et al.
Published: (2021)
Implicit neural representation with physics-informed neural networks for the reconstruction of the early part of room impulse responses
by: Pezzoli, Mirco, et al.
Published: (2023)
by: Pezzoli, Mirco, et al.
Published: (2023)
DiffusionRIR: Room Impulse Response Interpolation using Diffusion Models
by: Della Torre, Sagi, et al.
Published: (2025)
by: Della Torre, Sagi, et al.
Published: (2025)
MusicGen-Stem: Multi-stem music generation and edition through autoregressive modeling
by: Rouard, Simon, et al.
Published: (2025)
by: Rouard, Simon, et al.
Published: (2025)
$\text{M}^3\text{PDB}$: A Multimodal, Multi-Label, Multilingual Prompt Database for Speech Generation
by: Zhu, Boyu, et al.
Published: (2025)
by: Zhu, Boyu, et al.
Published: (2025)
Distilling a speech and music encoder with task arithmetic
by: Ritter-Gutierrez, Fabian, et al.
Published: (2025)
by: Ritter-Gutierrez, Fabian, et al.
Published: (2025)
EDTC: enhance depth of text comprehension in automated audio captioning
by: Tan, Liwen, et al.
Published: (2024)
by: Tan, Liwen, et al.
Published: (2024)
Text adaptation for speaker verification with speaker-text factorized embeddings
by: Yang, Yexin, et al.
Published: (2025)
by: Yang, Yexin, et al.
Published: (2025)
FxSearcher: gradient-free text-driven audio transformation
by: Ki, Hojoon, et al.
Published: (2025)
by: Ki, Hojoon, et al.
Published: (2025)
Effect of laboratory conditions on the perception of virtual stages for music
by: Accolti, Ernesto
Published: (2025)
by: Accolti, Ernesto
Published: (2025)
Expressive paragraph text-to-speech synthesis with multi-step variational autoencoder
by: Li, Xuyuan, et al.
Published: (2023)
by: Li, Xuyuan, et al.
Published: (2023)
LiveScaler: Live control of the harmony of an electronic music track
by: Rixte, Alice
Published: (2024)
by: Rixte, Alice
Published: (2024)
Investigation of perceptual music similarity focusing on each instrumental part
by: Hashizume, Yuka, et al.
Published: (2025)
by: Hashizume, Yuka, et al.
Published: (2025)
Boosting keyword spotting through on-device learnable user speech characteristics
by: Cioflan, Cristian, et al.
Published: (2024)
by: Cioflan, Cristian, et al.
Published: (2024)
A framework of text-dependent speaker verification for chinese numerical string corpus
by: Zheng, Litong, et al.
Published: (2024)
by: Zheng, Litong, et al.
Published: (2024)
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio
by: Kanamori, Yusuke, et al.
Published: (2025)
by: Kanamori, Yusuke, et al.
Published: (2025)
Similar Items
-
Synthetic training set generation using text-to-audio models for environmental sound classification
by: Ronchini, Francesca, et al.
Published: (2024) -
MambaFoley: Foley Sound Generation using Selective State-Space Models
by: Colombo, Marco Furio, et al.
Published: (2024) -
Mitigating data replication in text-to-audio generative diffusion models through anti-memorization guidance
by: Messina, Francisco, et al.
Published: (2025) -
AI-Assisted Music Production: A User Study on Text-to-Music Models
by: Ronchini, Francesca, et al.
Published: (2025) -
Room Transfer Function Reconstruction Using Complex-valued Neural Networks and Irregularly Distributed Microphones
by: Ronchini, Francesca, et al.
Published: (2024)