An Eye for an Ear: Zero-shot Audio Description Leveraging an Image Captioner using Audiovisual Distribution Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Malard, Hugo, Olvera, Michel, Lathuiliere, Stéphane, Essid, Slim |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization
von: Malard, Hugo, et al.
Veröffentlicht: (2024)
von: Malard, Hugo, et al.
Veröffentlicht: (2024)
SALT: Standardized Audio event Label Taxonomy
von: Stamatiadis, Paraskevas, et al.
Veröffentlicht: (2024)
von: Stamatiadis, Paraskevas, et al.
Veröffentlicht: (2024)
A sound description: Exploring prompt templates and class descriptions to enhance zero-shot audio classification
von: Olvera, Michel, et al.
Veröffentlicht: (2024)
von: Olvera, Michel, et al.
Veröffentlicht: (2024)
A Contrastive Self-Supervised Learning scheme for beat tracking amenable to few-shot learning
von: Gagnere, Antonin, et al.
Veröffentlicht: (2024)
von: Gagnere, Antonin, et al.
Veröffentlicht: (2024)
Less Forgetting for Better Generalization: Exploring Continual-learning Fine-tuning Methods for Speech Self-supervised Representations
von: Zaiem, Salah, et al.
Veröffentlicht: (2024)
von: Zaiem, Salah, et al.
Veröffentlicht: (2024)
Controlling Contrastive Self-Supervised Learning with Knowledge-Driven Multiple Hypothesis: Application to Beat Tracking
von: Gagnere, Antonin, et al.
Veröffentlicht: (2025)
von: Gagnere, Antonin, et al.
Veröffentlicht: (2025)
MACE: Leveraging Audio for Evaluating Audio Captioning Systems
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
A lightweight dual-stage framework for personalized speech enhancement based on DeepFilterNet2
von: Serre, Thomas, et al.
Veröffentlicht: (2024)
von: Serre, Thomas, et al.
Veröffentlicht: (2024)
Zero-Shot Audio Captioning Using Soft and Hard Prompts
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement
von: Serre, Thomas, et al.
Veröffentlicht: (2026)
von: Serre, Thomas, et al.
Veröffentlicht: (2026)
DRCap: Decoding CLAP Latents with Retrieval-Augmented Generation for Zero-shot Audio Captioning
von: Li, Xiquan, et al.
Veröffentlicht: (2024)
von: Li, Xiquan, et al.
Veröffentlicht: (2024)
ACES: Evaluating Automated Audio Captioning Models on the Semantics of Sounds
von: Wijngaard, Gijs, et al.
Veröffentlicht: (2024)
von: Wijngaard, Gijs, et al.
Veröffentlicht: (2024)
Discrete Audio Representations for Automated Audio Captioning
von: Tian, Jingguang, et al.
Veröffentlicht: (2025)
von: Tian, Jingguang, et al.
Veröffentlicht: (2025)
Perceptual Noise-Masking with Music through Deep Spectral Envelope Shaping
von: Berger, Clémentine, et al.
Veröffentlicht: (2025)
von: Berger, Clémentine, et al.
Veröffentlicht: (2025)
Online speaker diarization of meetings guided by speech separation
von: Gruttadauria, Elio, et al.
Veröffentlicht: (2024)
von: Gruttadauria, Elio, et al.
Veröffentlicht: (2024)
Enhancing Zero-shot Audio Classification using Sound Attribute Knowledge from Large Language Models
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
IS${}^3$ : Generic Impulsive--Stationary Sound Separation in Acoustic Scenes using Deep Filtering
von: Berger, Clémentine, et al.
Veröffentlicht: (2025)
von: Berger, Clémentine, et al.
Veröffentlicht: (2025)
MiDashengLM: Efficient Audio Understanding with General Audio Captions
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2025)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2025)
CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
Efficient Audio Captioning with Encoder-Level Knowledge Distillation
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
Multiple Choice Learning for Efficient Speech Separation with Many Speakers
von: Perera, David, et al.
Veröffentlicht: (2024)
von: Perera, David, et al.
Veröffentlicht: (2024)
LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
von: Dang, Trung, et al.
Veröffentlicht: (2024)
von: Dang, Trung, et al.
Veröffentlicht: (2024)
Enhance Temporal Relations in Audio Captioning with Sound Event Detection
von: Xie, Zeyu, et al.
Veröffentlicht: (2023)
von: Xie, Zeyu, et al.
Veröffentlicht: (2023)
CLAIR-A: Leveraging Large Language Models to Judge Audio Captions
von: Wu, Tsung-Han, et al.
Veröffentlicht: (2024)
von: Wu, Tsung-Han, et al.
Veröffentlicht: (2024)
Semantic Proximity Alignment: Towards Human Perception-consistent Audio Tagging by Aligning with Label Text Description
von: Liu, Wuyang, et al.
Veröffentlicht: (2023)
von: Liu, Wuyang, et al.
Veröffentlicht: (2023)
Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning
von: Xu, Le, et al.
Veröffentlicht: (2025)
von: Xu, Le, et al.
Veröffentlicht: (2025)
From Contrast to Commonality: Audio Commonality Captioning for Enhanced Audio-Text Cross-modal Understanding in Multimodal LLMs
von: Jia, Yuhang, et al.
Veröffentlicht: (2025)
von: Jia, Yuhang, et al.
Veröffentlicht: (2025)
Improving Audio Captioning Models with Fine-grained Audio Features, Text Embedding Supervision, and LLM Mix-up Augmentation
von: Wu, Shih-Lun, et al.
Veröffentlicht: (2023)
von: Wu, Shih-Lun, et al.
Veröffentlicht: (2023)
AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2024)
UniAudio 1.5: Large Language Model-driven Audio Codec is A Few-shot Audio Task Learner
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
Improving Audio-Text Retrieval via Hierarchical Cross-Modal Interaction and Auxiliary Captions
von: Xin, Yifei, et al.
Veröffentlicht: (2023)
von: Xin, Yifei, et al.
Veröffentlicht: (2023)
SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs
von: Chen, Wenxi, et al.
Veröffentlicht: (2024)
von: Chen, Wenxi, et al.
Veröffentlicht: (2024)
Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation
von: Wu, Yusong, et al.
Veröffentlicht: (2022)
von: Wu, Yusong, et al.
Veröffentlicht: (2022)
Speech Self-Supervised Representations Benchmarking: a Case for Larger Probing Heads
von: Zaiem, Salah, et al.
Veröffentlicht: (2023)
von: Zaiem, Salah, et al.
Veröffentlicht: (2023)
Zero- and Few-shot Sound Event Localization and Detection
von: Shimada, Kazuki, et al.
Veröffentlicht: (2023)
von: Shimada, Kazuki, et al.
Veröffentlicht: (2023)
Zero-shot Cross-lingual Voice Transfer for TTS
von: Biadsy, Fadi, et al.
Veröffentlicht: (2024)
von: Biadsy, Fadi, et al.
Veröffentlicht: (2024)
A Study on Zero-shot Non-intrusive Speech Assessment using Large Language Models
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2024)
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2024)
Resource-Efficient Reference-Free Evaluation of Audio Captions
von: Mahfuz, Rehana, et al.
Veröffentlicht: (2024)
von: Mahfuz, Rehana, et al.
Veröffentlicht: (2024)
Exploring Meta Information for Audio-based Zero-shot Bird Classification
von: Gebhard, Alexander, et al.
Veröffentlicht: (2023)
von: Gebhard, Alexander, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization
von: Malard, Hugo, et al.
Veröffentlicht: (2024) -
SALT: Standardized Audio event Label Taxonomy
von: Stamatiadis, Paraskevas, et al.
Veröffentlicht: (2024) -
A sound description: Exploring prompt templates and class descriptions to enhance zero-shot audio classification
von: Olvera, Michel, et al.
Veröffentlicht: (2024) -
A Contrastive Self-Supervised Learning scheme for beat tracking amenable to few-shot learning
von: Gagnere, Antonin, et al.
Veröffentlicht: (2024) -
Less Forgetting for Better Generalization: Exploring Continual-learning Fine-tuning Methods for Speech Self-supervised Representations
von: Zaiem, Salah, et al.
Veröffentlicht: (2024)