Diffused Responsibility: Analyzing the Energy Consumption of Generative Text-to-Audio Diffusion Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Passoni, Riccardo, Ronchini, Francesca, Comanducci, Luca, Serizel, Romain, Antonacci, Fabio |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Synthetic training set generation using text-to-audio models for environmental sound classification
von: Ronchini, Francesca, et al.
Veröffentlicht: (2024)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2024)
MambaFoley: Foley Sound Generation using Selective State-Space Models
von: Colombo, Marco Furio, et al.
Veröffentlicht: (2024)
von: Colombo, Marco Furio, et al.
Veröffentlicht: (2024)
PAGURI: a user experience study of creative interaction with text-to-music models
von: Ronchini, Francesca, et al.
Veröffentlicht: (2024)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2024)
AI-Assisted Music Production: A User Study on Text-to-Music Models
von: Ronchini, Francesca, et al.
Veröffentlicht: (2025)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2025)
Performance and energy balance: a comprehensive study of state-of-the-art sound event detection systems
von: Ronchini, Francesca, et al.
Veröffentlicht: (2023)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2023)
Mind the Prompt: Prompting Strategies in Audio Generations for Improving Sound Classification
von: Ronchini, Francesca, et al.
Veröffentlicht: (2025)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2025)
DiffusionRIR: Room Impulse Response Interpolation using Diffusion Models
von: Della Torre, Sagi, et al.
Veröffentlicht: (2025)
von: Della Torre, Sagi, et al.
Veröffentlicht: (2025)
Towards HRTF Personalization using Denoising Diffusion Models
von: Sánchez, Juan Camilo Albarracín, et al.
Veröffentlicht: (2025)
von: Sánchez, Juan Camilo Albarracín, et al.
Veröffentlicht: (2025)
A benchmark of state-of-the-art sound event detection systems evaluated on synthetic soundscapes
von: Ronchini, Francesca, et al.
Veröffentlicht: (2022)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2022)
Interpreting End-to-End Deep Learning Models for Speech Source Localization Using Layer-wise Relevance Propagation
von: Comanducci, Luca, et al.
Veröffentlicht: (2024)
von: Comanducci, Luca, et al.
Veröffentlicht: (2024)
Mitigating data replication in text-to-audio generative diffusion models through anti-memorization guidance
von: Messina, Francisco, et al.
Veröffentlicht: (2025)
von: Messina, Francisco, et al.
Veröffentlicht: (2025)
Room Transfer Function Reconstruction Using Complex-valued Neural Networks and Irregularly Distributed Microphones
von: Ronchini, Francesca, et al.
Veröffentlicht: (2024)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2024)
DreamAudio: Customized Text-to-Audio Generation with Diffusion Models
von: Yuan, Yi, et al.
Veröffentlicht: (2025)
von: Yuan, Yi, et al.
Veröffentlicht: (2025)
AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion
von: Zhao, Junqi, et al.
Veröffentlicht: (2025)
von: Zhao, Junqi, et al.
Veröffentlicht: (2025)
Energy Consumption Trends in Sound Event Detection Systems
von: Douwes, Constance, et al.
Veröffentlicht: (2024)
von: Douwes, Constance, et al.
Veröffentlicht: (2024)
Reconstruction of Sound Field through Diffusion Models
von: Miotello, Federico, et al.
Veröffentlicht: (2023)
von: Miotello, Federico, et al.
Veröffentlicht: (2023)
Frequency-Weighted Training Losses for Phoneme-Level DNN-based Speech Enhancement
von: Monir, Nasser-Eddine, et al.
Veröffentlicht: (2025)
von: Monir, Nasser-Eddine, et al.
Veröffentlicht: (2025)
The impact of non-target events in synthetic soundscapes for sound event detection
von: Ronchini, Francesca, et al.
Veröffentlicht: (2021)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2021)
Speaker Embeddings to Improve Tracking of Intermittent and Moving Speakers
von: Iatariene, Taous, et al.
Veröffentlicht: (2025)
von: Iatariene, Taous, et al.
Veröffentlicht: (2025)
Towards Low-Latency Tracking of Multiple Speakers With Short-Context Speaker Embeddings
von: Iatariene, Taous, et al.
Veröffentlicht: (2025)
von: Iatariene, Taous, et al.
Veröffentlicht: (2025)
Latent Watermarking of Audio Generative Models
von: Roman, Robin San, et al.
Veröffentlicht: (2024)
von: Roman, Robin San, et al.
Veröffentlicht: (2024)
Angular Distance Distribution Loss for Audio Classification
von: Almudévar, Antonio, et al.
Veröffentlicht: (2024)
von: Almudévar, Antonio, et al.
Veröffentlicht: (2024)
FakeMusicCaps: a Dataset for Detection and Attribution of Synthetic Music Generated via Text-to-Music Models
von: Comanducci, Luca, et al.
Veröffentlicht: (2024)
von: Comanducci, Luca, et al.
Veröffentlicht: (2024)
Description and analysis of novelties introduced in DCASE Task 4 2022 on the baseline system
von: Ronchini, Francesca, et al.
Veröffentlicht: (2022)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2022)
Diffusion-based Unsupervised Audio-visual Speech Enhancement
von: Ayilo, Jean-Eudes, et al.
Veröffentlicht: (2024)
von: Ayilo, Jean-Eudes, et al.
Veröffentlicht: (2024)
Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
ControlAudio: Tackling Text-Guided, Timing-Indicated and Intelligible Audio Generation via Progressive Diffusion Modeling
von: Jiang, Yuxuan, et al.
Veröffentlicht: (2025)
von: Jiang, Yuxuan, et al.
Veröffentlicht: (2025)
Estimating Musical Surprisal from Audio in Autoregressive Diffusion Model Noise Spaces
von: Bjare, Mathias Rose, et al.
Veröffentlicht: (2025)
von: Bjare, Mathias Rose, et al.
Veröffentlicht: (2025)
Audio-Conditioned Diffusion LLMs for ASR and Deliberation Processing
von: Wang, Mengqi, et al.
Veröffentlicht: (2025)
von: Wang, Mengqi, et al.
Veröffentlicht: (2025)
SynSonic: Augmenting Sound Event Detection through Text-to-Audio Diffusion ControlNet and Effective Sample Filtering
von: Hai, Jiarui, et al.
Veröffentlicht: (2025)
von: Hai, Jiarui, et al.
Veröffentlicht: (2025)
DualSpec: Text-to-spatial-audio Generation via Dual-Spectrogram Guided Diffusion Model
von: Zhao, Lei, et al.
Veröffentlicht: (2025)
von: Zhao, Lei, et al.
Veröffentlicht: (2025)
LiteFocus: Accelerated Diffusion Inference for Long Audio Synthesis
von: Tan, Zhenxiong, et al.
Veröffentlicht: (2024)
von: Tan, Zhenxiong, et al.
Veröffentlicht: (2024)
Inference-time Scaling for Diffusion-based Audio Super-resolution
von: Jin, Yizhu, et al.
Veröffentlicht: (2025)
von: Jin, Yizhu, et al.
Veröffentlicht: (2025)
Audio Deepfake Detection in the Age of Advanced Text-to-Speech models
von: Singh, Robin, et al.
Veröffentlicht: (2026)
von: Singh, Robin, et al.
Veröffentlicht: (2026)
TTMBA: Towards Text To Multiple Sources Binaural Audio Generation
von: He, Yuxuan, et al.
Veröffentlicht: (2025)
von: He, Yuxuan, et al.
Veröffentlicht: (2025)
Expressive Range Characterization of Open Text-to-Audio Models
von: Morse, Jonathan, et al.
Veröffentlicht: (2025)
von: Morse, Jonathan, et al.
Veröffentlicht: (2025)
Modeling Time-Variant Responses of Optical Compressors with Selective State Space Models
von: Simionato, Riccardo, et al.
Veröffentlicht: (2024)
von: Simionato, Riccardo, et al.
Veröffentlicht: (2024)
VR-PTOLEMAIC: A Virtual Environment for the Perceptual Testing of Spatial Audio Algorithms
von: Ostan, Paolo, et al.
Veröffentlicht: (2025)
von: Ostan, Paolo, et al.
Veröffentlicht: (2025)
Fine-Tuning Text-to-Speech Diffusion Models Using Reinforcement Learning with Human Feedback
von: Chen, Jingyi, et al.
Veröffentlicht: (2025)
von: Chen, Jingyi, et al.
Veröffentlicht: (2025)
ASD-Diffusion: Anomalous Sound Detection with Diffusion Models
von: Zhang, Fengrun, et al.
Veröffentlicht: (2024)
von: Zhang, Fengrun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Synthetic training set generation using text-to-audio models for environmental sound classification
von: Ronchini, Francesca, et al.
Veröffentlicht: (2024) -
MambaFoley: Foley Sound Generation using Selective State-Space Models
von: Colombo, Marco Furio, et al.
Veröffentlicht: (2024) -
PAGURI: a user experience study of creative interaction with text-to-music models
von: Ronchini, Francesca, et al.
Veröffentlicht: (2024) -
AI-Assisted Music Production: A User Study on Text-to-Music Models
von: Ronchini, Francesca, et al.
Veröffentlicht: (2025) -
Performance and energy balance: a comprehensive study of state-of-the-art sound event detection systems
von: Ronchini, Francesca, et al.
Veröffentlicht: (2023)