From Discord to Harmony: Decomposed Consonance-based Training for Improved Audio Chord Estimation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Poltronieri, Andrea, Serra, Xavier, Rocamora, Martín |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ChordSync: Conformer-Based Alignment of Chord Annotations to Music Audio
von: Poltronieri, Andrea, et al.
Veröffentlicht: (2024)
von: Poltronieri, Andrea, et al.
Veröffentlicht: (2024)
Leveraging Pre-Trained Autoencoders for Interpretable Prototype Learning of Music Audio
von: Alonso-Jiménez, Pablo, et al.
Veröffentlicht: (2024)
von: Alonso-Jiménez, Pablo, et al.
Veröffentlicht: (2024)
Towards Assessing Data Replication in Music Generation with Music Similarity Metrics on Raw Audio
von: Batlle-Roca, Roser, et al.
Veröffentlicht: (2024)
von: Batlle-Roca, Roser, et al.
Veröffentlicht: (2024)
FreeAudio: Training-Free Timing Planning for Controllable Long-Form Text-to-Audio Generation
von: Jiang, Yuxuan, et al.
Veröffentlicht: (2025)
von: Jiang, Yuxuan, et al.
Veröffentlicht: (2025)
Fast Text-to-Audio Generation with Adversarial Post-Training
von: Novack, Zachary, et al.
Veröffentlicht: (2025)
von: Novack, Zachary, et al.
Veröffentlicht: (2025)
Towards Generating Diverse Audio Captions via Adversarial Training
von: Mei, Xinhao, et al.
Veröffentlicht: (2022)
von: Mei, Xinhao, et al.
Veröffentlicht: (2022)
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport
von: Rho, Kyeongha, et al.
Veröffentlicht: (2025)
von: Rho, Kyeongha, et al.
Veröffentlicht: (2025)
Audio Transformers
von: Verma, Prateek, et al.
Veröffentlicht: (2021)
von: Verma, Prateek, et al.
Veröffentlicht: (2021)
AudioRole: An Audio Dataset for Character Role-Playing in Large Language Models
von: Li, Wenyu, et al.
Veröffentlicht: (2025)
von: Li, Wenyu, et al.
Veröffentlicht: (2025)
Understanding Pedestrian Movement Using Urban Sensing Technologies: The Promise of Audio-based Sensors
von: Han, Chaeyeon, et al.
Veröffentlicht: (2024)
von: Han, Chaeyeon, et al.
Veröffentlicht: (2024)
OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model
von: Li, Maomao, et al.
Veröffentlicht: (2026)
von: Li, Maomao, et al.
Veröffentlicht: (2026)
Generative AI for Music and Audio
von: Dong, Hao-Wen
Veröffentlicht: (2024)
von: Dong, Hao-Wen
Veröffentlicht: (2024)
Jailbreak-AudioBench: In-Depth Evaluation and Analysis of Jailbreak Threats for Large Audio Language Models
von: Cheng, Hao, et al.
Veröffentlicht: (2025)
von: Cheng, Hao, et al.
Veröffentlicht: (2025)
Retrieval-Augmented Text-to-Audio Generation
von: Yuan, Yi, et al.
Veröffentlicht: (2023)
von: Yuan, Yi, et al.
Veröffentlicht: (2023)
Neural Style Transfer for Audio Spectograms
von: Verma, Prateek, et al.
Veröffentlicht: (2018)
von: Verma, Prateek, et al.
Veröffentlicht: (2018)
Embedding Alignment in Code Generation for Audio
von: Kouteili, Sam, et al.
Veröffentlicht: (2025)
von: Kouteili, Sam, et al.
Veröffentlicht: (2025)
Content Adaptive Front End For Audio Classification
von: Verma, Prateek, et al.
Veröffentlicht: (2023)
von: Verma, Prateek, et al.
Veröffentlicht: (2023)
Integrating IP Broadcasting with Audio Tags: Workflow and Challenges
von: Burchett-Vass, Rhys, et al.
Veröffentlicht: (2024)
von: Burchett-Vass, Rhys, et al.
Veröffentlicht: (2024)
Unveiling Visual Biases in Audio-Visual Localization Benchmarks
von: Chen, Liangyu, et al.
Veröffentlicht: (2024)
von: Chen, Liangyu, et al.
Veröffentlicht: (2024)
PIAST: A Multimodal Piano Dataset with Audio, Symbolic and Text
von: Bang, Hayeon, et al.
Veröffentlicht: (2024)
von: Bang, Hayeon, et al.
Veröffentlicht: (2024)
LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT
von: Du, Zhihao, et al.
Veröffentlicht: (2023)
von: Du, Zhihao, et al.
Veröffentlicht: (2023)
Diverse Audio Embeddings -- Bringing Features Back Outperforms CLAP!
von: Verma, Prateek
Veröffentlicht: (2023)
von: Verma, Prateek
Veröffentlicht: (2023)
Challenge on Sound Scene Synthesis: Evaluating Text-to-Audio Generation
von: Lee, Junwon, et al.
Veröffentlicht: (2024)
von: Lee, Junwon, et al.
Veröffentlicht: (2024)
DeepFake Doctor: Diagnosing and Treating Audio-Video Fake Detection
von: Klemt, Marcel, et al.
Veröffentlicht: (2025)
von: Klemt, Marcel, et al.
Veröffentlicht: (2025)
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?
von: Li, Jia, et al.
Veröffentlicht: (2025)
von: Li, Jia, et al.
Veröffentlicht: (2025)
Leveraging Pre-trained AudioLDM for Sound Generation: A Benchmark Study
von: Yuan, Yi, et al.
Veröffentlicht: (2023)
von: Yuan, Yi, et al.
Veröffentlicht: (2023)
Listening and Seeing Again: Generative Error Correction for Audio-Visual Speech Recognition
von: Liu, Rui, et al.
Veröffentlicht: (2025)
von: Liu, Rui, et al.
Veröffentlicht: (2025)
Robust Audio Anti-Spoofing with Fusion-Reconstruction Learning on Multi-Order Spectrograms
von: Wen, Penghui, et al.
Veröffentlicht: (2023)
von: Wen, Penghui, et al.
Veröffentlicht: (2023)
SHMamba: Structured Hyperbolic State Space Model for Audio-Visual Question Answering
von: Yang, Zhe, et al.
Veröffentlicht: (2024)
von: Yang, Zhe, et al.
Veröffentlicht: (2024)
SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning
von: Nam, KiHyun, et al.
Veröffentlicht: (2026)
von: Nam, KiHyun, et al.
Veröffentlicht: (2026)
DOA-Aware Audio-Visual Self-Supervised Learning for Sound Event Localization and Detection
von: Fujita, Yoto, et al.
Veröffentlicht: (2024)
von: Fujita, Yoto, et al.
Veröffentlicht: (2024)
A Multi-Stream Fusion Approach with One-Class Learning for Audio-Visual Deepfake Detection
von: Lee, Kyungbok, et al.
Veröffentlicht: (2024)
von: Lee, Kyungbok, et al.
Veröffentlicht: (2024)
Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation
von: Zhang, Kang, et al.
Veröffentlicht: (2025)
von: Zhang, Kang, et al.
Veröffentlicht: (2025)
Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models
von: Xie, Zhifei, et al.
Veröffentlicht: (2025)
von: Xie, Zhifei, et al.
Veröffentlicht: (2025)
CHORDONOMICON: A Dataset of 666,000 Songs and their Chord Progressions
von: Kantarelis, Spyridon, et al.
Veröffentlicht: (2024)
von: Kantarelis, Spyridon, et al.
Veröffentlicht: (2024)
Purification Before Fusion: Toward Mask-Free Speech Enhancement for Robust Audio-Visual Speech Recognition
von: Wu, Linzhi, et al.
Veröffentlicht: (2026)
von: Wu, Linzhi, et al.
Veröffentlicht: (2026)
LPIPS-AttnWav2Lip: Generic Audio-Driven lip synchronization for Talking Head Generation in the Wild
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026)
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026)
DreamHead: Learning Spatial-Temporal Correspondence via Hierarchical Diffusion for Audio-driven Talking Head Synthesis
von: Hong, Fa-Ting, et al.
Veröffentlicht: (2024)
von: Hong, Fa-Ting, et al.
Veröffentlicht: (2024)
Evaluating Disentangled Representations for Controllable Music Generation
von: Ibáñez-Martínez, Laura, et al.
Veröffentlicht: (2026)
von: Ibáñez-Martínez, Laura, et al.
Veröffentlicht: (2026)
From Sound to Sight: Towards AI-authored Music Videos
von: Vitasovic, Leo, et al.
Veröffentlicht: (2025)
von: Vitasovic, Leo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ChordSync: Conformer-Based Alignment of Chord Annotations to Music Audio
von: Poltronieri, Andrea, et al.
Veröffentlicht: (2024) -
Leveraging Pre-Trained Autoencoders for Interpretable Prototype Learning of Music Audio
von: Alonso-Jiménez, Pablo, et al.
Veröffentlicht: (2024) -
Towards Assessing Data Replication in Music Generation with Music Similarity Metrics on Raw Audio
von: Batlle-Roca, Roser, et al.
Veröffentlicht: (2024) -
FreeAudio: Training-Free Timing Planning for Controllable Long-Form Text-to-Audio Generation
von: Jiang, Yuxuan, et al.
Veröffentlicht: (2025) -
Fast Text-to-Audio Generation with Adversarial Post-Training
von: Novack, Zachary, et al.
Veröffentlicht: (2025)