Dissecting Temporal Understanding in Text-to-Audio Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Oncescu, Andreea-Maria, Henriques, João F., Koepke, A. Sophia |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A SOUND APPROACH: Using Large Language Models to generate audio descriptions for egocentric text-audio retrieval
by: Oncescu, Andreea-Maria, et al.
Published: (2024)
by: Oncescu, Andreea-Maria, et al.
Published: (2024)
ASK: Adaptive Self-improving Knowledge Framework for Audio Text Retrieval
by: Fu, Siyuan, et al.
Published: (2025)
by: Fu, Siyuan, et al.
Published: (2025)
DiffATR: Diffusion-based Generative Modeling for Audio-Text Retrieval
by: Xin, Yifei, et al.
Published: (2024)
by: Xin, Yifei, et al.
Published: (2024)
Multiscale Matching Driven by Cross-Modal Similarity Consistency for Audio-Text Retrieval
by: Wang, Qian, et al.
Published: (2024)
by: Wang, Qian, et al.
Published: (2024)
EAViT: External Attention Vision Transformer for Audio Classification
by: Iqbal, Aquib, et al.
Published: (2024)
by: Iqbal, Aquib, et al.
Published: (2024)
A Novel Audio Representation for Music Genre Identification in MIR
by: Kamuni, Navin, et al.
Published: (2024)
by: Kamuni, Navin, et al.
Published: (2024)
Nested Music Transformer: Sequentially Decoding Compound Tokens in Symbolic Music and Audio Generation
by: Yoo, HaeJun, et al.
Published: (2024)
by: Yoo, HaeJun, et al.
Published: (2024)
Speaker Retrieval in the Wild: Challenges, Effectiveness and Robustness
by: Loweimi, Erfan, et al.
Published: (2025)
by: Loweimi, Erfan, et al.
Published: (2025)
Language-based Audio Retrieval with Co-Attention Networks
by: Sun, Haoran, et al.
Published: (2024)
by: Sun, Haoran, et al.
Published: (2024)
Analyzing and reducing the synthetic-to-real transfer gap in Music Information Retrieval: the task of automatic drum transcription
by: Zehren, Mickaël, et al.
Published: (2024)
by: Zehren, Mickaël, et al.
Published: (2024)
MERGE -- A Bimodal Audio-Lyrics Dataset for Static Music Emotion Recognition
by: Louro, Pedro Lima, et al.
Published: (2024)
by: Louro, Pedro Lima, et al.
Published: (2024)
On the Effect of Data-Augmentation on Local Embedding Properties in the Contrastive Learning of Music Audio Representations
by: McCallum, Matthew C., et al.
Published: (2024)
by: McCallum, Matthew C., et al.
Published: (2024)
Music Foundation Model as Generic Booster for Music Downstream Tasks
by: Liao, WeiHsiang, et al.
Published: (2024)
by: Liao, WeiHsiang, et al.
Published: (2024)
SECP: A Speech Enhancement-Based Curation Pipeline For Scalable Acquisition Of Clean Speech
by: Sabra, Adam, et al.
Published: (2024)
by: Sabra, Adam, et al.
Published: (2024)
Improving Musical Instrument Classification with Advanced Machine Learning Techniques
by: Chulev, Joanikij
Published: (2024)
by: Chulev, Joanikij
Published: (2024)
From Real to Cloned Singer Identification
by: Desblancs, Dorian, et al.
Published: (2024)
by: Desblancs, Dorian, et al.
Published: (2024)
Transfer Learning with Semi-Supervised Dataset Annotation for Birdcall Classification
by: Miyaguchi, Anthony, et al.
Published: (2023)
by: Miyaguchi, Anthony, et al.
Published: (2023)
Hybrid Losses for Hierarchical Embedding Learning
by: Tian, Haokun, et al.
Published: (2025)
by: Tian, Haokun, et al.
Published: (2025)
Deconstructing Jazz Piano Style Using Machine Learning
by: Cheston, Huw, et al.
Published: (2025)
by: Cheston, Huw, et al.
Published: (2025)
Multi-label Cross-lingual automatic music genre classification from lyrics with Sentence BERT
by: Tavares, Tiago Fernandes, et al.
Published: (2025)
by: Tavares, Tiago Fernandes, et al.
Published: (2025)
Emergent musical properties of a transformer under contrastive self-supervised learning
by: Kong, Yuexuan, et al.
Published: (2025)
by: Kong, Yuexuan, et al.
Published: (2025)
Uncertainty Estimation in the Real World: A Study on Music Emotion Recognition
by: Watcharasupat, Karn N., et al.
Published: (2025)
by: Watcharasupat, Karn N., et al.
Published: (2025)
Separate This, and All of these Things Around It: Music Source Separation via Hyperellipsoidal Queries
by: Watcharasupat, Karn N., et al.
Published: (2025)
by: Watcharasupat, Karn N., et al.
Published: (2025)
Universal Music Representations? Evaluating Foundation Models on World Music Corpora
by: Papaioannou, Charilaos, et al.
Published: (2025)
by: Papaioannou, Charilaos, et al.
Published: (2025)
Segment Length Matters: A Study of Segment Lengths on Audio Fingerprinting Performance
by: Gong, Ziling, et al.
Published: (2026)
by: Gong, Ziling, et al.
Published: (2026)
Similar but Faster: Manipulation of Tempo in Music Audio Embeddings for Tempo Prediction and Search
by: McCallum, Matthew C., et al.
Published: (2024)
by: McCallum, Matthew C., et al.
Published: (2024)
Pretrained Conformers for Audio Fingerprinting and Retrieval
by: Altwlkany, Kemal, et al.
Published: (2025)
by: Altwlkany, Kemal, et al.
Published: (2025)
Contrastive and Transfer Learning for Effective Audio Fingerprinting through a Real-World Evaluation Protocol
by: Nikou, Christos, et al.
Published: (2025)
by: Nikou, Christos, et al.
Published: (2025)
LARP: Language Audio Relational Pre-training for Cold-Start Playlist Continuation
by: Salganik, Rebecca, et al.
Published: (2024)
by: Salganik, Rebecca, et al.
Published: (2024)
Enriching Music Descriptions with a Finetuned-LLM and Metadata for Text-to-Music Retrieval
by: Doh, SeungHeon, et al.
Published: (2024)
by: Doh, SeungHeon, et al.
Published: (2024)
Exploring GPT's Ability as a Judge in Music Understanding
by: Fang, Kun, et al.
Published: (2025)
by: Fang, Kun, et al.
Published: (2025)
Expressivity-aware Music Performance Retrieval using Mid-level Perceptual Features and Emotion Word Embeddings
by: Chowdhury, Shreyan, et al.
Published: (2024)
by: Chowdhury, Shreyan, et al.
Published: (2024)
A Dataset and Baselines for Measuring and Predicting the Music Piece Memorability
by: Tseng, Li-Yang, et al.
Published: (2024)
by: Tseng, Li-Yang, et al.
Published: (2024)
Music Genre Classification: Ensemble Learning with Subcomponents-level Attention
by: Liu, Yichen, et al.
Published: (2024)
by: Liu, Yichen, et al.
Published: (2024)
Learning Normal Patterns in Musical Loops
by: Dadman, Shayan, et al.
Published: (2025)
by: Dadman, Shayan, et al.
Published: (2025)
Application of Audio Fingerprinting Techniques for Real-Time Scalable Speech Retrieval and Speech Clusterization
by: Altwlkany, Kemal, et al.
Published: (2024)
by: Altwlkany, Kemal, et al.
Published: (2024)
MUSE: Flexible Voiceprint Receptive Fields and Multi-Path Fusion Enhanced Taylor Transformer for U-Net-based Speech Enhancement
by: Lin, Zizhen, et al.
Published: (2024)
by: Lin, Zizhen, et al.
Published: (2024)
Equivariance-based self-supervised learning for audio signal recovery from clipped measurements
by: Sechaud, Victor, et al.
Published: (2024)
by: Sechaud, Victor, et al.
Published: (2024)
A Cascaded Architecture for Extractive Summarization of Multimedia Content via Audio-to-Text Alignment
by: Hossain, Tanzir, et al.
Published: (2025)
by: Hossain, Tanzir, et al.
Published: (2025)
Anchor-aware Deep Metric Learning for Audio-visual Retrieval
by: Zeng, Donghuo, et al.
Published: (2024)
by: Zeng, Donghuo, et al.
Published: (2024)
Similar Items
-
A SOUND APPROACH: Using Large Language Models to generate audio descriptions for egocentric text-audio retrieval
by: Oncescu, Andreea-Maria, et al.
Published: (2024) -
ASK: Adaptive Self-improving Knowledge Framework for Audio Text Retrieval
by: Fu, Siyuan, et al.
Published: (2025) -
DiffATR: Diffusion-based Generative Modeling for Audio-Text Retrieval
by: Xin, Yifei, et al.
Published: (2024) -
Multiscale Matching Driven by Cross-Modal Similarity Consistency for Audio-Text Retrieval
by: Wang, Qian, et al.
Published: (2024) -
EAViT: External Attention Vision Transformer for Audio Classification
by: Iqbal, Aquib, et al.
Published: (2024)