FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Shunian, Xie, Xinyuan, Chen, Zheshu, Zhao, Liyan, Lee, Owen, Su, Zhan, Sun, Qilin, Wang, Benyou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2024)
Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation
von: Wu, Yusong, et al.
Veröffentlicht: (2022)
von: Wu, Yusong, et al.
Veröffentlicht: (2022)
Improving Audio Captioning Models with Fine-grained Audio Features, Text Embedding Supervision, and LLM Mix-up Augmentation
von: Wu, Shih-Lun, et al.
Veröffentlicht: (2023)
von: Wu, Shih-Lun, et al.
Veröffentlicht: (2023)
Video-to-Audio Generation with Fine-grained Temporal Semantics
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
From Contrast to Commonality: Audio Commonality Captioning for Enhanced Audio-Text Cross-modal Understanding in Multimodal LLMs
von: Jia, Yuhang, et al.
Veröffentlicht: (2025)
von: Jia, Yuhang, et al.
Veröffentlicht: (2025)
Discrete Audio Representations for Automated Audio Captioning
von: Tian, Jingguang, et al.
Veröffentlicht: (2025)
von: Tian, Jingguang, et al.
Veröffentlicht: (2025)
CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
MACE: Leveraging Audio for Evaluating Audio Captioning Systems
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
MiDashengLM: Efficient Audio Understanding with General Audio Captions
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2025)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2025)
Audio-Guided Fusion Techniques for Multimodal Emotion Analysis
von: Shi, Pujin, et al.
Veröffentlicht: (2024)
von: Shi, Pujin, et al.
Veröffentlicht: (2024)
Enhance Temporal Relations in Audio Captioning with Sound Event Detection
von: Xie, Zeyu, et al.
Veröffentlicht: (2023)
von: Xie, Zeyu, et al.
Veröffentlicht: (2023)
SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs
von: Chen, Wenxi, et al.
Veröffentlicht: (2024)
von: Chen, Wenxi, et al.
Veröffentlicht: (2024)
Robust Audio-Visual Target Speaker Extraction with Emotion-Aware Multiple Enrollment Fusion
von: Jin, Zhan, et al.
Veröffentlicht: (2025)
von: Jin, Zhan, et al.
Veröffentlicht: (2025)
Audio Spatially-Guided Fusion for Audio-Visual Navigation
von: Zhou, Xinyu, et al.
Veröffentlicht: (2026)
von: Zhou, Xinyu, et al.
Veröffentlicht: (2026)
Efficient Audio Captioning with Encoder-Level Knowledge Distillation
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
UniAudio: An Audio Foundation Model Toward Universal Audio Generation
von: Yang, Dongchao, et al.
Veröffentlicht: (2023)
von: Yang, Dongchao, et al.
Veröffentlicht: (2023)
Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
Zero-Shot Audio Captioning Using Soft and Hard Prompts
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
ACES: Evaluating Automated Audio Captioning Models on the Semantics of Sounds
von: Wijngaard, Gijs, et al.
Veröffentlicht: (2024)
von: Wijngaard, Gijs, et al.
Veröffentlicht: (2024)
Why Your Tokenizer Fails in Information Fusion: A Timing-Aware Pre-Quantization Fusion for Video-Enhanced Audio Tokenization
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2026)
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2026)
Multimodal Magic Elevating Depression Detection with a Fusion of Text and Audio Intelligence
von: Gan, Lindy, et al.
Veröffentlicht: (2025)
von: Gan, Lindy, et al.
Veröffentlicht: (2025)
MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models
von: Gong, Yitian, et al.
Veröffentlicht: (2026)
von: Gong, Yitian, et al.
Veröffentlicht: (2026)
Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0
von: Wang, Zhiyong, et al.
Veröffentlicht: (2024)
von: Wang, Zhiyong, et al.
Veröffentlicht: (2024)
T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback
von: Wang, Zehan, et al.
Veröffentlicht: (2025)
von: Wang, Zehan, et al.
Veröffentlicht: (2025)
PromptSep: Generative Audio Separation via Multimodal Prompting
von: Wen, Yutong, et al.
Veröffentlicht: (2025)
von: Wen, Yutong, et al.
Veröffentlicht: (2025)
AV-SSAN: Audio-Visual Selective DoA Estimation through Explicit Multi-Band Semantic-Spatial Alignment
von: Chen, Yu, et al.
Veröffentlicht: (2025)
von: Chen, Yu, et al.
Veröffentlicht: (2025)
Estimated Audio-Caption Correspondences Improve Language-Based Audio Retrieval
von: Primus, Paul, et al.
Veröffentlicht: (2024)
von: Primus, Paul, et al.
Veröffentlicht: (2024)
Code Drift: Towards Idempotent Neural Audio Codecs
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2024)
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2024)
Spatial-Aware Conditioned Fusion for Audio-Visual Navigation
von: Wu, Shaohang, et al.
Veröffentlicht: (2026)
von: Wu, Shaohang, et al.
Veröffentlicht: (2026)
Resource-Efficient Reference-Free Evaluation of Audio Captions
von: Mahfuz, Rehana, et al.
Veröffentlicht: (2024)
von: Mahfuz, Rehana, et al.
Veröffentlicht: (2024)
Towards Fusion of Neural Audio Codec-based Representations with Spectral for Heart Murmur Classification via Bandit-based Cross-Attention Mechanism
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
DRCap: Decoding CLAP Latents with Retrieval-Augmented Generation for Zero-shot Audio Captioning
von: Li, Xiquan, et al.
Veröffentlicht: (2024)
von: Li, Xiquan, et al.
Veröffentlicht: (2024)
Aligning Audio Captions with Human Preferences
von: Hegde, Kartik, et al.
Veröffentlicht: (2025)
von: Hegde, Kartik, et al.
Veröffentlicht: (2025)
Toward Multimodal Industrial Fault Analysis: A Single-Speed Chain Conveyor Dataset with Audio and Vibration Signals
von: Chen, Zhang, et al.
Veröffentlicht: (2026)
von: Chen, Zhang, et al.
Veröffentlicht: (2026)
EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer
von: Takeuchi, Daiki, et al.
Veröffentlicht: (2025)
von: Takeuchi, Daiki, et al.
Veröffentlicht: (2025)
AudioRAG: A Challenging Benchmark for Audio Reasoning and Information Retrieval
von: Lin, Jingru, et al.
Veröffentlicht: (2026)
von: Lin, Jingru, et al.
Veröffentlicht: (2026)
AudioEditor: A Training-Free Diffusion-Based Audio Editing Framework
von: Jia, Yuhang, et al.
Veröffentlicht: (2024)
von: Jia, Yuhang, et al.
Veröffentlicht: (2024)
Automatic Contextual Audio Denoising
von: Luong, Diep, et al.
Veröffentlicht: (2026)
von: Luong, Diep, et al.
Veröffentlicht: (2026)
Improving Audio-Text Retrieval via Hierarchical Cross-Modal Interaction and Auxiliary Captions
von: Xin, Yifei, et al.
Veröffentlicht: (2023)
von: Xin, Yifei, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2024) -
Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation
von: Wu, Yusong, et al.
Veröffentlicht: (2022) -
Improving Audio Captioning Models with Fine-grained Audio Features, Text Embedding Supervision, and LLM Mix-up Augmentation
von: Wu, Shih-Lun, et al.
Veröffentlicht: (2023) -
Video-to-Audio Generation with Fine-grained Temporal Semantics
von: Hu, Yuchen, et al.
Veröffentlicht: (2024) -
From Contrast to Commonality: Audio Commonality Captioning for Enhanced Audio-Text Cross-modal Understanding in Multimodal LLMs
von: Jia, Yuhang, et al.
Veröffentlicht: (2025)