Text-based Audio Retrieval by Learning from Similarities between Audio Captions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xie, Huang, Khorrami, Khazar, Räsänen, Okko, Virtanen, Tuomas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Integrating Continuous and Binary Relevances in Audio-Text Relevance Learning
von: Xie, Huang, et al.
Veröffentlicht: (2024)
von: Xie, Huang, et al.
Veröffentlicht: (2024)
A model of early word acquisition based on realistic-scale audiovisual naming events
von: Khorrami, Khazar, et al.
Veröffentlicht: (2024)
von: Khorrami, Khazar, et al.
Veröffentlicht: (2024)
Simultaneous or Sequential Training? How Speech Representations Cooperate in a Multi-Task Self-Supervised Learning System
von: Khorrami, Khazar, et al.
Veröffentlicht: (2023)
von: Khorrami, Khazar, et al.
Veröffentlicht: (2023)
Can phones, syllables, and words emerge as side-products of cross-situational audiovisual learning? -- A computational investigation
von: Khorrami, Khazar, et al.
Veröffentlicht: (2021)
von: Khorrami, Khazar, et al.
Veröffentlicht: (2021)
Noise-to-mask Ratio Loss for Deep Neural Network based Audio Watermarking
von: Moritz, Martin, et al.
Veröffentlicht: (2024)
von: Moritz, Martin, et al.
Veröffentlicht: (2024)
Multi-label Zero-Shot Audio Classification with Temporal Attention
von: Dogan, Duygu, et al.
Veröffentlicht: (2024)
von: Dogan, Duygu, et al.
Veröffentlicht: (2024)
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025)
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025)
Speaker Distance Estimation in Enclosures from Single-Channel Audio
von: Neri, Michael, et al.
Veröffentlicht: (2024)
von: Neri, Michael, et al.
Veröffentlicht: (2024)
AudioNet: Supervised Deep Hashing for Retrieval of Similar Audio Events
von: Dutta, Sagar, et al.
Veröffentlicht: (2025)
von: Dutta, Sagar, et al.
Veröffentlicht: (2025)
Age-Dependent Analysis and Stochastic Generation of Child-Directed Speech
von: Räsänen, Okko, et al.
Veröffentlicht: (2024)
von: Räsänen, Okko, et al.
Veröffentlicht: (2024)
Automatic Contextual Audio Denoising
von: Luong, Diep, et al.
Veröffentlicht: (2026)
von: Luong, Diep, et al.
Veröffentlicht: (2026)
Automatic Live Music Song Identification Using Multi-level Deep Sequence Similarity Learning
von: Hakala, Aapo, et al.
Veröffentlicht: (2025)
von: Hakala, Aapo, et al.
Veröffentlicht: (2025)
Impact of Microphone Array Mismatches to Learning-based Replay Speech Detection
von: Neri, Michael, et al.
Veröffentlicht: (2025)
von: Neri, Michael, et al.
Veröffentlicht: (2025)
Improving Audio-Text Retrieval via Hierarchical Cross-Modal Interaction and Auxiliary Captions
von: Xin, Yifei, et al.
Veröffentlicht: (2023)
von: Xin, Yifei, et al.
Veröffentlicht: (2023)
Discrete Audio Representations for Automated Audio Captioning
von: Tian, Jingguang, et al.
Veröffentlicht: (2025)
von: Tian, Jingguang, et al.
Veröffentlicht: (2025)
CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
Inter-Speaker Relative Cues for Two-Stage Text-Guided Target Speech Extraction
von: Dai, Wang, et al.
Veröffentlicht: (2026)
von: Dai, Wang, et al.
Veröffentlicht: (2026)
MACE: Leveraging Audio for Evaluating Audio Captioning Systems
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
Evaluation of Audio-Visual Alignments in Visually Grounded Speech Models
von: Khorrami, Khazar, et al.
Veröffentlicht: (2021)
von: Khorrami, Khazar, et al.
Veröffentlicht: (2021)
Estimated Audio-Caption Correspondences Improve Language-Based Audio Retrieval
von: Primus, Paul, et al.
Veröffentlicht: (2024)
von: Primus, Paul, et al.
Veröffentlicht: (2024)
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models
von: Bai, Jisheng, et al.
Veröffentlicht: (2024)
von: Bai, Jisheng, et al.
Veröffentlicht: (2024)
Inter-Speaker Relative Cues for Text-Guided Target Speech Extraction
von: Dai, Wang, et al.
Veröffentlicht: (2025)
von: Dai, Wang, et al.
Veröffentlicht: (2025)
From Contrast to Commonality: Audio Commonality Captioning for Enhanced Audio-Text Cross-modal Understanding in Multimodal LLMs
von: Jia, Yuhang, et al.
Veröffentlicht: (2025)
von: Jia, Yuhang, et al.
Veröffentlicht: (2025)
Adversarial Representation Learning for Robust Privacy Preservation in Audio
von: Gharib, Shayan, et al.
Veröffentlicht: (2023)
von: Gharib, Shayan, et al.
Veröffentlicht: (2023)
MiDashengLM: Efficient Audio Understanding with General Audio Captions
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2025)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2025)
Enhance Temporal Relations in Audio Captioning with Sound Event Detection
von: Xie, Zeyu, et al.
Veröffentlicht: (2023)
von: Xie, Zeyu, et al.
Veröffentlicht: (2023)
Improving Audio Captioning Models with Fine-grained Audio Features, Text Embedding Supervision, and LLM Mix-up Augmentation
von: Wu, Shih-Lun, et al.
Veröffentlicht: (2023)
von: Wu, Shih-Lun, et al.
Veröffentlicht: (2023)
EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
Multi-channel Replay Speech Detection using an Adaptive Learnable Beamformer
von: Neri, Michael, et al.
Veröffentlicht: (2025)
von: Neri, Michael, et al.
Veröffentlicht: (2025)
Refining Knowledge Transfer on Audio-Image Temporal Agreement for Audio-Text Cross Retrieval
von: Tsubaki, Shunsuke, et al.
Veröffentlicht: (2024)
von: Tsubaki, Shunsuke, et al.
Veröffentlicht: (2024)
Computational modeling of early language learning from acoustic speech and audiovisual input without linguistic priors
von: Räsänen, Okko
Veröffentlicht: (2026)
von: Räsänen, Okko
Veröffentlicht: (2026)
Zimtohrli: An Efficient Psychoacoustic Audio Similarity Metric
von: Alakuijala, Jyrki, et al.
Veröffentlicht: (2025)
von: Alakuijala, Jyrki, et al.
Veröffentlicht: (2025)
Representation Learning for Audio Privacy Preservation using Source Separation and Robust Adversarial Learning
von: Luong, Diep, et al.
Veröffentlicht: (2023)
von: Luong, Diep, et al.
Veröffentlicht: (2023)
Multiscale Matching Driven by Cross-Modal Similarity Consistency for Audio-Text Retrieval
von: Wang, Qian, et al.
Veröffentlicht: (2024)
von: Wang, Qian, et al.
Veröffentlicht: (2024)
RECAP: Retrieval-Augmented Audio Captioning
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
AudioLCM: Text-to-Audio Generation with Latent Consistency Models
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
Efficient Audio Captioning with Encoder-Level Knowledge Distillation
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
Evaluating the Temporal Detection Capability of Integrated Gradients Applied on Sound Classifier
von: Dumpis, Martynas, et al.
Veröffentlicht: (2026)
von: Dumpis, Martynas, et al.
Veröffentlicht: (2026)
Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models
von: Li, Longhao, et al.
Veröffentlicht: (2026)
von: Li, Longhao, et al.
Veröffentlicht: (2026)
DRCap: Decoding CLAP Latents with Retrieval-Augmented Generation for Zero-shot Audio Captioning
von: Li, Xiquan, et al.
Veröffentlicht: (2024)
von: Li, Xiquan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Integrating Continuous and Binary Relevances in Audio-Text Relevance Learning
von: Xie, Huang, et al.
Veröffentlicht: (2024) -
A model of early word acquisition based on realistic-scale audiovisual naming events
von: Khorrami, Khazar, et al.
Veröffentlicht: (2024) -
Simultaneous or Sequential Training? How Speech Representations Cooperate in a Multi-Task Self-Supervised Learning System
von: Khorrami, Khazar, et al.
Veröffentlicht: (2023) -
Can phones, syllables, and words emerge as side-products of cross-situational audiovisual learning? -- A computational investigation
von: Khorrami, Khazar, et al.
Veröffentlicht: (2021) -
Noise-to-mask Ratio Loss for Deep Neural Network based Audio Watermarking
von: Moritz, Martin, et al.
Veröffentlicht: (2024)