Semantic and Semiotic Interplays in Text-to-Audio AI: Exploring Cognitive Dynamics and Musical Interactions
Fuente:
arXiv
Guardado en:
| Autor principal: | Coelho, Guilherme |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The Artist is Present: Traces of Artists Resigind and Spawning in Text-to-Audio AI
por: Coelho, Guilherme
Publicado: (2025)
por: Coelho, Guilherme
Publicado: (2025)
AI in Music and Sound: Pedagogical Reflections, Post-Structuralist Approaches and Creative Outcomes in Seminar Practice
por: Coelho, Guilherme
Publicado: (2025)
por: Coelho, Guilherme
Publicado: (2025)
Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation
por: Tal, Or, et al.
Publicado: (2024)
por: Tal, Or, et al.
Publicado: (2024)
MusicSem: A Semantically Rich Language--Audio Dataset of Natural Music Descriptions
por: Salganik, Rebecca, et al.
Publicado: (2026)
por: Salganik, Rebecca, et al.
Publicado: (2026)
Breaking the Barriers of Text-Hungry and Audio-Deficient AI
por: Tembine, Hamidou, et al.
Publicado: (2025)
por: Tembine, Hamidou, et al.
Publicado: (2025)
SemanticAudio: Audio Generation and Editing in Semantic Space
por: Dai, Zheqi, et al.
Publicado: (2026)
por: Dai, Zheqi, et al.
Publicado: (2026)
From Audio Deepfake Detection to AI-Generated Music Detection -- A Pathway and Overview
por: Li, Yupei, et al.
Publicado: (2024)
por: Li, Yupei, et al.
Publicado: (2024)
Exploring Text-Queried Sound Event Detection with Audio Source Separation
por: Yin, Han, et al.
Publicado: (2024)
por: Yin, Han, et al.
Publicado: (2024)
Enhancing Neural Audio Fingerprint Robustness to Audio Degradation for Music Identification
por: Araz, R. Oguz, et al.
Publicado: (2025)
por: Araz, R. Oguz, et al.
Publicado: (2025)
Improving Audio-Text Retrieval via Hierarchical Cross-Modal Interaction and Auxiliary Captions
por: Xin, Yifei, et al.
Publicado: (2023)
por: Xin, Yifei, et al.
Publicado: (2023)
MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation
por: Liu, Cheng, et al.
Publicado: (2025)
por: Liu, Cheng, et al.
Publicado: (2025)
Audio Conditioning for Music Generation via Discrete Bottleneck Features
por: Rouard, Simon, et al.
Publicado: (2024)
por: Rouard, Simon, et al.
Publicado: (2024)
Amphion: An Open-Source Audio, Music and Speech Generation Toolkit
por: Zhang, Xueyao, et al.
Publicado: (2023)
por: Zhang, Xueyao, et al.
Publicado: (2023)
Semantic Proximity Alignment: Towards Human Perception-consistent Audio Tagging by Aligning with Label Text Description
por: Liu, Wuyang, et al.
Publicado: (2023)
por: Liu, Wuyang, et al.
Publicado: (2023)
AudioLCM: Text-to-Audio Generation with Latent Consistency Models
por: Liu, Huadai, et al.
Publicado: (2024)
por: Liu, Huadai, et al.
Publicado: (2024)
Audio Prompt Adapter: Unleashing Music Editing Abilities for Text-to-Music with Lightweight Finetuning
por: Tsai, Fang-Duo, et al.
Publicado: (2024)
por: Tsai, Fang-Duo, et al.
Publicado: (2024)
FakeMusicCaps: a Dataset for Detection and Attribution of Synthetic Music Generated via Text-to-Music Models
por: Comanducci, Luca, et al.
Publicado: (2024)
por: Comanducci, Luca, et al.
Publicado: (2024)
Self-Supervised Multi-View Learning for Disentangled Music Audio Representations
por: Wilkins, Julia, et al.
Publicado: (2024)
por: Wilkins, Julia, et al.
Publicado: (2024)
MuseAgent-1: Interactive Grounded Multimodal Understanding of Music Scores and Performance Audio
por: Zhao, Qihao, et al.
Publicado: (2026)
por: Zhao, Qihao, et al.
Publicado: (2026)
EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
por: Hai, Jiarui, et al.
Publicado: (2024)
por: Hai, Jiarui, et al.
Publicado: (2024)
FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio Generation
por: Liu, Huadai, et al.
Publicado: (2024)
por: Liu, Huadai, et al.
Publicado: (2024)
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding
por: Guinot, Julien, et al.
Publicado: (2025)
por: Guinot, Julien, et al.
Publicado: (2025)
ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech
por: Shi, Jiatong, et al.
Publicado: (2024)
por: Shi, Jiatong, et al.
Publicado: (2024)
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning
por: Zhang, Jisi, et al.
Publicado: (2025)
por: Zhang, Jisi, et al.
Publicado: (2025)
AudioSpa: Spatializing Sound Events with Text
por: Feng, Linfeng, et al.
Publicado: (2025)
por: Feng, Linfeng, et al.
Publicado: (2025)
Cacophony: An Improved Contrastive Audio-Text Model
por: Zhu, Ge, et al.
Publicado: (2024)
por: Zhu, Ge, et al.
Publicado: (2024)
Enhancing Crowdsourced Audio for Text-to-Speech Models
por: Giraldo, José, et al.
Publicado: (2024)
por: Giraldo, José, et al.
Publicado: (2024)
Towards Weakly Supervised Text-to-Audio Grounding
por: Xu, Xuenan, et al.
Publicado: (2024)
por: Xu, Xuenan, et al.
Publicado: (2024)
AudioEval: Automatic Dual-Perspective and Multi-Dimensional Evaluation of Text-to-Audio-Generation
por: Wang, Hui, et al.
Publicado: (2025)
por: Wang, Hui, et al.
Publicado: (2025)
Refining Knowledge Transfer on Audio-Image Temporal Agreement for Audio-Text Cross Retrieval
por: Tsubaki, Shunsuke, et al.
Publicado: (2024)
por: Tsubaki, Shunsuke, et al.
Publicado: (2024)
PAT: Parameter-Free Audio-Text Aligner to Boost Zero-Shot Audio Classification
por: Seth, Ashish, et al.
Publicado: (2024)
por: Seth, Ashish, et al.
Publicado: (2024)
ITO-Master: Inference-Time Optimization for Audio Effects Modeling of Music Mastering Processors
por: Koo, Junghyun, et al.
Publicado: (2025)
por: Koo, Junghyun, et al.
Publicado: (2025)
Network Modulation Synthesis: New Algorithms for Generating Musical Audio Using Autoencoder Networks
por: Hyrkas, Jeremy
Publicado: (2021)
por: Hyrkas, Jeremy
Publicado: (2021)
Can Audio Reveal Music Performance Difficulty? Insights from the Piano Syllabus Dataset
por: Ramoneda, Pedro, et al.
Publicado: (2024)
por: Ramoneda, Pedro, et al.
Publicado: (2024)
CoPlay: Audio-agnostic Cognitive Scaling for Acoustic Sensing
por: Li, Yin, et al.
Publicado: (2024)
por: Li, Yin, et al.
Publicado: (2024)
T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback
por: Wang, Zehan, et al.
Publicado: (2025)
por: Wang, Zehan, et al.
Publicado: (2025)
Video-to-Audio Generation with Fine-grained Temporal Semantics
por: Hu, Yuchen, et al.
Publicado: (2024)
por: Hu, Yuchen, et al.
Publicado: (2024)
Text2FX: Harnessing CLAP Embeddings for Text-Guided Audio Effects
por: Chu, Annie, et al.
Publicado: (2024)
por: Chu, Annie, et al.
Publicado: (2024)
Exploring Musical Roots: Applying Audio Embeddings to Empower Influence Attribution for a Generative Music Model
por: Barnett, Julia, et al.
Publicado: (2024)
por: Barnett, Julia, et al.
Publicado: (2024)
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models
por: Zhang, Yixiao
Publicado: (2024)
por: Zhang, Yixiao
Publicado: (2024)
Ejemplares similares
-
The Artist is Present: Traces of Artists Resigind and Spawning in Text-to-Audio AI
por: Coelho, Guilherme
Publicado: (2025) -
AI in Music and Sound: Pedagogical Reflections, Post-Structuralist Approaches and Creative Outcomes in Seminar Practice
por: Coelho, Guilherme
Publicado: (2025) -
Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation
por: Tal, Or, et al.
Publicado: (2024) -
MusicSem: A Semantically Rich Language--Audio Dataset of Natural Music Descriptions
por: Salganik, Rebecca, et al.
Publicado: (2026) -
Breaking the Barriers of Text-Hungry and Audio-Deficient AI
por: Tembine, Hamidou, et al.
Publicado: (2025)