Revisiting Deep Audio-Text Retrieval Through the Lens of Transportation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Luong, Manh, Nguyen, Khai, Ho, Nhat, Haf, Reza, Phung, Dinh, Qu, Lizhen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Unbiased Sliced Wasserstein Kernels for High-Quality Audio Captioning
von: Luong, Manh, et al.
Veröffentlicht: (2025)
von: Luong, Manh, et al.
Veröffentlicht: (2025)
Retrieval-Augmented Audio Deepfake Detection
von: Kang, Zuheng, et al.
Veröffentlicht: (2024)
von: Kang, Zuheng, et al.
Veröffentlicht: (2024)
Retrieval-Augmented Text-to-Audio Generation
von: Yuan, Yi, et al.
Veröffentlicht: (2023)
von: Yuan, Yi, et al.
Veröffentlicht: (2023)
AFSS: Artifact-Focused Self-Synthesis for Mitigating Bias in Audio Deepfake Detection
von: Nguyen-Le, Hai-Son, et al.
Veröffentlicht: (2026)
von: Nguyen-Le, Hai-Son, et al.
Veröffentlicht: (2026)
CoughViT: A Self-Supervised Vision Transformer for Cough Audio Representation Learning
von: Luong, Justin, et al.
Veröffentlicht: (2025)
von: Luong, Justin, et al.
Veröffentlicht: (2025)
ATRI: Mitigating Multilingual Audio Text Retrieval Inconsistencies by Reducing Data Distribution Errors
von: Yin, Yuguo, et al.
Veröffentlicht: (2025)
von: Yin, Yuguo, et al.
Veröffentlicht: (2025)
Deepfake Audio Detection Using Spectrogram-based Feature and Ensemble of Deep Learning Models
von: Pham, Lam, et al.
Veröffentlicht: (2024)
von: Pham, Lam, et al.
Veröffentlicht: (2024)
Refining Knowledge Transfer on Audio-Image Temporal Agreement for Audio-Text Cross Retrieval
von: Tsubaki, Shunsuke, et al.
Veröffentlicht: (2024)
von: Tsubaki, Shunsuke, et al.
Veröffentlicht: (2024)
AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion
von: Zhao, Junqi, et al.
Veröffentlicht: (2025)
von: Zhao, Junqi, et al.
Veröffentlicht: (2025)
DreamAudio: Customized Text-to-Audio Generation with Diffusion Models
von: Yuan, Yi, et al.
Veröffentlicht: (2025)
von: Yuan, Yi, et al.
Veröffentlicht: (2025)
AudioLens: A Closer Look at Auditory Attribute Perception of Large Audio-Language Models
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2025)
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2025)
EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
Improving Audio-Text Retrieval via Hierarchical Cross-Modal Interaction and Auxiliary Captions
von: Xin, Yifei, et al.
Veröffentlicht: (2023)
von: Xin, Yifei, et al.
Veröffentlicht: (2023)
Audio Is the Achilles' Heel: Red Teaming Audio Large Multimodal Models
von: Yang, Hao, et al.
Veröffentlicht: (2024)
von: Yang, Hao, et al.
Veröffentlicht: (2024)
In-the-wild Audio Spatialization with Flexible Text-guided Localization
von: Pan, Tianrui, et al.
Veröffentlicht: (2025)
von: Pan, Tianrui, et al.
Veröffentlicht: (2025)
Expressive Range Characterization of Open Text-to-Audio Models
von: Morse, Jonathan, et al.
Veröffentlicht: (2025)
von: Morse, Jonathan, et al.
Veröffentlicht: (2025)
TTMBA: Towards Text To Multiple Sources Binaural Audio Generation
von: He, Yuxuan, et al.
Veröffentlicht: (2025)
von: He, Yuxuan, et al.
Veröffentlicht: (2025)
Audio Deepfake Detection in the Age of Advanced Text-to-Speech models
von: Singh, Robin, et al.
Veröffentlicht: (2026)
von: Singh, Robin, et al.
Veröffentlicht: (2026)
Expanding on EnCLAP with Auxiliary Retrieval Model for Automated Audio Captioning
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
Audiobox TTA-RAG: Improving Zero-Shot and Few-Shot Text-To-Audio with Retrieval-Augmented Generation
von: Yang, Mu, et al.
Veröffentlicht: (2024)
von: Yang, Mu, et al.
Veröffentlicht: (2024)
AND: Audio Network Dissection for Interpreting Deep Acoustic Models
von: Wu, Tung-Yu, et al.
Veröffentlicht: (2024)
von: Wu, Tung-Yu, et al.
Veröffentlicht: (2024)
BrewCLIP: A Bifurcated Representation Learning Framework for Audio-Visual Retrieval
von: Lu, Zhenyu, et al.
Veröffentlicht: (2024)
von: Lu, Zhenyu, et al.
Veröffentlicht: (2024)
WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models
von: Chen, Yifu, et al.
Veröffentlicht: (2025)
von: Chen, Yifu, et al.
Veröffentlicht: (2025)
The SVASR System for Text-dependent Speaker Verification (TdSV) AAIC Challenge 2024
von: Molavi, Mohammadreza, et al.
Veröffentlicht: (2024)
von: Molavi, Mohammadreza, et al.
Veröffentlicht: (2024)
A Toolchain for Comprehensive Audio/Video Analysis Using Deep Learning Based Multimodal Approach (A use case of riot or violent context detection)
von: Pham, Lam, et al.
Veröffentlicht: (2024)
von: Pham, Lam, et al.
Veröffentlicht: (2024)
AudioRAG: A Challenging Benchmark for Audio Reasoning and Information Retrieval
von: Lin, Jingru, et al.
Veröffentlicht: (2026)
von: Lin, Jingru, et al.
Veröffentlicht: (2026)
AudioLCM: Text-to-Audio Generation with Latent Consistency Models
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
Retrieval Augmented Generation in Prompt-based Text-to-Speech Synthesis with Context-Aware Contrastive Language-Audio Pretraining
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
Dynamic Context-Aware Streaming Pretrained Language Model For Inverse Text Normalization
von: Ho, Luong, et al.
Veröffentlicht: (2025)
von: Ho, Luong, et al.
Veröffentlicht: (2025)
Enhancing Retrieval-Augmented Audio Captioning with Generation-Assisted Multimodal Querying and Progressive Learning
von: Changin, Choi, et al.
Veröffentlicht: (2024)
von: Changin, Choi, et al.
Veröffentlicht: (2024)
DRCap: Decoding CLAP Latents with Retrieval-Augmented Generation for Zero-shot Audio Captioning
von: Li, Xiquan, et al.
Veröffentlicht: (2024)
von: Li, Xiquan, et al.
Veröffentlicht: (2024)
CTC-aligned Audio-Text Embedding for Streaming Open-vocabulary Keyword Spotting
von: Jin, Sichen, et al.
Veröffentlicht: (2024)
von: Jin, Sichen, et al.
Veröffentlicht: (2024)
EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
von: Hai, Jiarui, et al.
Veröffentlicht: (2024)
von: Hai, Jiarui, et al.
Veröffentlicht: (2024)
Audio Prompt Adapter: Unleashing Music Editing Abilities for Text-to-Music with Lightweight Finetuning
von: Tsai, Fang-Duo, et al.
Veröffentlicht: (2024)
von: Tsai, Fang-Duo, et al.
Veröffentlicht: (2024)
FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio Generation
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
Bridging Language Gaps in Audio-Text Retrieval
von: Yan, Zhiyong, et al.
Veröffentlicht: (2024)
von: Yan, Zhiyong, et al.
Veröffentlicht: (2024)
Domain Adaptation Method and Modality Gap Impact in Audio-Text Models for Prototypical Sound Classification
von: Acevedo, Emiliano, et al.
Veröffentlicht: (2025)
von: Acevedo, Emiliano, et al.
Veröffentlicht: (2025)
CORD: Bridging the Audio-Text Reasoning Gap via Weighted On-policy Cross-modal Distillation
von: Hu, Jing, et al.
Veröffentlicht: (2026)
von: Hu, Jing, et al.
Veröffentlicht: (2026)
Text Prompt is Not Enough: Sound Event Enhanced Prompt Adapter for Target Style Audio Generation
von: Xiong, Chenxu, et al.
Veröffentlicht: (2024)
von: Xiong, Chenxu, et al.
Veröffentlicht: (2024)
Cacophony: An Improved Contrastive Audio-Text Model
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Unbiased Sliced Wasserstein Kernels for High-Quality Audio Captioning
von: Luong, Manh, et al.
Veröffentlicht: (2025) -
Retrieval-Augmented Audio Deepfake Detection
von: Kang, Zuheng, et al.
Veröffentlicht: (2024) -
Retrieval-Augmented Text-to-Audio Generation
von: Yuan, Yi, et al.
Veröffentlicht: (2023) -
AFSS: Artifact-Focused Self-Synthesis for Mitigating Bias in Audio Deepfake Detection
von: Nguyen-Le, Hai-Son, et al.
Veröffentlicht: (2026) -
CoughViT: A Self-Supervised Vision Transformer for Cough Audio Representation Learning
von: Luong, Justin, et al.
Veröffentlicht: (2025)