The Artist is Present: Traces of Artists Resigind and Spawning in Text-to-Audio AI
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Coelho, Guilherme |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Semantic and Semiotic Interplays in Text-to-Audio AI: Exploring Cognitive Dynamics and Musical Interactions
von: Coelho, Guilherme
Veröffentlicht: (2025)
von: Coelho, Guilherme
Veröffentlicht: (2025)
CAFA: a Controllable Automatic Foley Artist
von: Benita, Roi, et al.
Veröffentlicht: (2025)
von: Benita, Roi, et al.
Veröffentlicht: (2025)
AI in Music and Sound: Pedagogical Reflections, Post-Structuralist Approaches and Creative Outcomes in Seminar Practice
von: Coelho, Guilherme
Veröffentlicht: (2025)
von: Coelho, Guilherme
Veröffentlicht: (2025)
Breaking the Barriers of Text-Hungry and Audio-Deficient AI
von: Tembine, Hamidou, et al.
Veröffentlicht: (2025)
von: Tembine, Hamidou, et al.
Veröffentlicht: (2025)
Source Tracing of Audio Deepfake Systems
von: Klein, Nicholas, et al.
Veröffentlicht: (2024)
von: Klein, Nicholas, et al.
Veröffentlicht: (2024)
Past, Present, and Future of Spatial Audio and Room Acoustics
von: Koyama, Shoichi, et al.
Veröffentlicht: (2025)
von: Koyama, Shoichi, et al.
Veröffentlicht: (2025)
Open-Set Source Tracing of Audio Deepfake Systems
von: Klein, Nicholas, et al.
Veröffentlicht: (2025)
von: Klein, Nicholas, et al.
Veröffentlicht: (2025)
AudioLCM: Text-to-Audio Generation with Latent Consistency Models
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
Music Era Recognition Using Supervised Contrastive Learning and Artist Information
von: He, Qiqi, et al.
Veröffentlicht: (2024)
von: He, Qiqi, et al.
Veröffentlicht: (2024)
EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
von: Hai, Jiarui, et al.
Veröffentlicht: (2024)
von: Hai, Jiarui, et al.
Veröffentlicht: (2024)
FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio Generation
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
RPRA-ADD: Forgery Trace Enhancement-Driven Audio Deepfake Detection
von: Fu, Ruibo, et al.
Veröffentlicht: (2025)
von: Fu, Ruibo, et al.
Veröffentlicht: (2025)
AudioSpa: Spatializing Sound Events with Text
von: Feng, Linfeng, et al.
Veröffentlicht: (2025)
von: Feng, Linfeng, et al.
Veröffentlicht: (2025)
Cacophony: An Improved Contrastive Audio-Text Model
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
Enhancing Crowdsourced Audio for Text-to-Speech Models
von: Giraldo, José, et al.
Veröffentlicht: (2024)
von: Giraldo, José, et al.
Veröffentlicht: (2024)
Towards Weakly Supervised Text-to-Audio Grounding
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
AudioEval: Automatic Dual-Perspective and Multi-Dimensional Evaluation of Text-to-Audio-Generation
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
Refining Knowledge Transfer on Audio-Image Temporal Agreement for Audio-Text Cross Retrieval
von: Tsubaki, Shunsuke, et al.
Veröffentlicht: (2024)
von: Tsubaki, Shunsuke, et al.
Veröffentlicht: (2024)
PAT: Parameter-Free Audio-Text Aligner to Boost Zero-Shot Audio Classification
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback
von: Wang, Zehan, et al.
Veröffentlicht: (2025)
von: Wang, Zehan, et al.
Veröffentlicht: (2025)
Codec-Based Deepfake Source Tracing via Neural Audio Codec Taxonomy
von: Chen, Xuanjun, et al.
Veröffentlicht: (2025)
von: Chen, Xuanjun, et al.
Veröffentlicht: (2025)
Text2FX: Harnessing CLAP Embeddings for Text-Guided Audio Effects
von: Chu, Annie, et al.
Veröffentlicht: (2024)
von: Chu, Annie, et al.
Veröffentlicht: (2024)
AudioGAN: A Compact and Efficient Framework for Real-Time High-Fidelity Text-to-Audio Generation
von: Chung, HaeChun
Veröffentlicht: (2025)
von: Chung, HaeChun
Veröffentlicht: (2025)
MATS: An Audio Language Model under Text-only Supervision
von: Wang, Wen, et al.
Veröffentlicht: (2025)
von: Wang, Wen, et al.
Veröffentlicht: (2025)
PPPR: Portable Plug-in Prompt Refiner for Text to Audio Generation
von: Shi, Shuchen, et al.
Veröffentlicht: (2024)
von: Shi, Shuchen, et al.
Veröffentlicht: (2024)
From Contrast to Commonality: Audio Commonality Captioning for Enhanced Audio-Text Cross-modal Understanding in Multimodal LLMs
von: Jia, Yuhang, et al.
Veröffentlicht: (2025)
von: Jia, Yuhang, et al.
Veröffentlicht: (2025)
Improving Audio Captioning Models with Fine-grained Audio Features, Text Embedding Supervision, and LLM Mix-up Augmentation
von: Wu, Shih-Lun, et al.
Veröffentlicht: (2023)
von: Wu, Shih-Lun, et al.
Veröffentlicht: (2023)
TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
Comparative Evaluation of Text and Audio Simplification: A Methodological Replication Study
von: Barai, Prosanta, et al.
Veröffentlicht: (2025)
von: Barai, Prosanta, et al.
Veröffentlicht: (2025)
Leveraging Audio-Only Data for Text-Queried Target Sound Extraction
von: Saijo, Kohei, et al.
Veröffentlicht: (2024)
von: Saijo, Kohei, et al.
Veröffentlicht: (2024)
Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation
von: Tal, Or, et al.
Veröffentlicht: (2024)
von: Tal, Or, et al.
Veröffentlicht: (2024)
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation
von: Kumar, Sonal, et al.
Veröffentlicht: (2024)
von: Kumar, Sonal, et al.
Veröffentlicht: (2024)
LAFMA: A Latent Flow Matching Model for Text-to-Audio Generation
von: Guan, Wenhao, et al.
Veröffentlicht: (2024)
von: Guan, Wenhao, et al.
Veröffentlicht: (2024)
Exploring Text-Queried Sound Event Detection with Audio Source Separation
von: Yin, Han, et al.
Veröffentlicht: (2024)
von: Yin, Han, et al.
Veröffentlicht: (2024)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling
von: Huang, Kuan-Po, et al.
Veröffentlicht: (2025)
von: Huang, Kuan-Po, et al.
Veröffentlicht: (2025)
Improving Audio-Text Retrieval via Hierarchical Cross-Modal Interaction and Auxiliary Captions
von: Xin, Yifei, et al.
Veröffentlicht: (2023)
von: Xin, Yifei, et al.
Veröffentlicht: (2023)
Investigating Group Relative Policy Optimization for Diffusion Transformer based Text-to-Audio Generation
von: Gu, Yi, et al.
Veröffentlicht: (2026)
von: Gu, Yi, et al.
Veröffentlicht: (2026)
A Detailed Audio-Text Data Simulation Pipeline using Single-Event Sounds
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
Listen, Analyze, and Adapt to Learn New Attacks: An Exemplar-Free Class Incremental Learning Method for Audio Deepfake Source Tracing
von: Xiao, Yang, et al.
Veröffentlicht: (2025)
von: Xiao, Yang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Semantic and Semiotic Interplays in Text-to-Audio AI: Exploring Cognitive Dynamics and Musical Interactions
von: Coelho, Guilherme
Veröffentlicht: (2025) -
CAFA: a Controllable Automatic Foley Artist
von: Benita, Roi, et al.
Veröffentlicht: (2025) -
AI in Music and Sound: Pedagogical Reflections, Post-Structuralist Approaches and Creative Outcomes in Seminar Practice
von: Coelho, Guilherme
Veröffentlicht: (2025) -
Breaking the Barriers of Text-Hungry and Audio-Deficient AI
von: Tembine, Hamidou, et al.
Veröffentlicht: (2025) -
Source Tracing of Audio Deepfake Systems
von: Klein, Nicholas, et al.
Veröffentlicht: (2024)