LAMB: LLM-based Audio Captioning with Modality Gap Bridging via Cauchy-Schwarz Divergence
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Hyeongkeun, Choi, Jongmin, Nam, KiHyun, Chung, Joon Son |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Diffusion-Link: Diffusion Probabilistic Model for Bridging the Audio-Text Modality Gap
von: Nam, KiHyun, et al.
Veröffentlicht: (2025)
von: Nam, KiHyun, et al.
Veröffentlicht: (2025)
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport
von: Rho, Kyeongha, et al.
Veröffentlicht: (2025)
von: Rho, Kyeongha, et al.
Veröffentlicht: (2025)
SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning
von: Nam, KiHyun, et al.
Veröffentlicht: (2026)
von: Nam, KiHyun, et al.
Veröffentlicht: (2026)
Disentangled Representation Learning for Environment-agnostic Speaker Recognition
von: Nam, KiHyun, et al.
Veröffentlicht: (2024)
von: Nam, KiHyun, et al.
Veröffentlicht: (2024)
MoLT: Mixture of Layer-Wise Tokens for Efficient Audio-Visual Learning
von: Rho, Kyeongha, et al.
Veröffentlicht: (2025)
von: Rho, Kyeongha, et al.
Veröffentlicht: (2025)
Audio Mamba: Bidirectional State Space Model for Audio Representation Learning
von: Erol, Mehmet Hamza, et al.
Veröffentlicht: (2024)
von: Erol, Mehmet Hamza, et al.
Veröffentlicht: (2024)
ElasticAST: An Audio Spectrogram Transformer for All Length and Resolutions
von: Feng, Jiu, et al.
Veröffentlicht: (2024)
von: Feng, Jiu, et al.
Veröffentlicht: (2024)
Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2024)
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2024)
Plug-and-Steer: Decoupling Separation and Selection in Audio-Visual Target Speaker Extraction
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation
von: Zhang, Kang, et al.
Veröffentlicht: (2025)
von: Zhang, Kang, et al.
Veröffentlicht: (2025)
AudioCapBench: Quick Evaluation on Audio Captioning across Sound, Music, and Speech
von: Qiu, Jielin, et al.
Veröffentlicht: (2026)
von: Qiu, Jielin, et al.
Veröffentlicht: (2026)
Cinematic Audio Source Separation Using Visual Cues
von: Zhang, Kang, et al.
Veröffentlicht: (2026)
von: Zhang, Kang, et al.
Veröffentlicht: (2026)
EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
Enhancing Retrieval-Augmented Audio Captioning with Generation-Assisted Multimodal Querying and Progressive Learning
von: Changin, Choi, et al.
Veröffentlicht: (2024)
von: Changin, Choi, et al.
Veröffentlicht: (2024)
Performance Improvement of Language-Queried Audio Source Separation Based on Caption Augmentation From Large Language Models for DCASE Challenge 2024 Task 9
von: Lee, Do Hyun, et al.
Veröffentlicht: (2024)
von: Lee, Do Hyun, et al.
Veröffentlicht: (2024)
KAD: No More FAD! An Effective and Efficient Evaluation Metric for Audio Generation
von: Chung, Yoonjin, et al.
Veröffentlicht: (2025)
von: Chung, Yoonjin, et al.
Veröffentlicht: (2025)
CORD: Bridging the Audio-Text Reasoning Gap via Weighted On-policy Cross-modal Distillation
von: Hu, Jing, et al.
Veröffentlicht: (2026)
von: Hu, Jing, et al.
Veröffentlicht: (2026)
Lightweight Audio Segmentation for Long-form Speech Translation
von: Lee, Jaesong, et al.
Veröffentlicht: (2024)
von: Lee, Jaesong, et al.
Veröffentlicht: (2024)
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024)
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024)
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
EquiAV: Leveraging Equivariance for Audio-Visual Contrastive Learning
von: Kim, Jongsuk, et al.
Veröffentlicht: (2024)
von: Kim, Jongsuk, et al.
Veröffentlicht: (2024)
CAF-Score: Calibrating CLAP with LALMs for Reference-free Audio Captioning Evaluation
von: Lee, Insung, et al.
Veröffentlicht: (2026)
von: Lee, Insung, et al.
Veröffentlicht: (2026)
LAV: Audio-Driven Dynamic Visual Generation with Neural Compression and StyleGAN2
von: Jung, Jongmin, et al.
Veröffentlicht: (2025)
von: Jung, Jongmin, et al.
Veröffentlicht: (2025)
Expanding on EnCLAP with Auxiliary Retrieval Model for Automated Audio Captioning
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
Towards Generating Diverse Audio Captions via Adversarial Training
von: Mei, Xinhao, et al.
Veröffentlicht: (2022)
von: Mei, Xinhao, et al.
Veröffentlicht: (2022)
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion
von: Chen, Shunian, et al.
Veröffentlicht: (2025)
von: Chen, Shunian, et al.
Veröffentlicht: (2025)
Cross-Modal Retrieval with Cauchy-Schwarz Divergence
von: Zhang, Jiahao, et al.
Veröffentlicht: (2025)
von: Zhang, Jiahao, et al.
Veröffentlicht: (2025)
Audio-Guided Dynamic Modality Fusion with Stereo-Aware Attention for Audio-Visual Navigation
von: Li, Jia, et al.
Veröffentlicht: (2025)
von: Li, Jia, et al.
Veröffentlicht: (2025)
EnCLAP++: Analyzing the EnCLAP Framework for Optimizing Automated Audio Captioning Performance
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
Domain Adaptation Method and Modality Gap Impact in Audio-Text Models for Prototypical Sound Classification
von: Acevedo, Emiliano, et al.
Veröffentlicht: (2025)
von: Acevedo, Emiliano, et al.
Veröffentlicht: (2025)
Let There Be Sound: Reconstructing High Quality Speech from Silent Videos
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2023)
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2023)
LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
RECAP: Retrieval-Augmented Audio Captioning
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
PIAST: A Multimodal Piano Dataset with Audio, Symbolic and Text
von: Bang, Hayeon, et al.
Veröffentlicht: (2024)
von: Bang, Hayeon, et al.
Veröffentlicht: (2024)
Classifier-Guided Captioning Across Modalities
von: Shaulov, Ariel, et al.
Veröffentlicht: (2025)
von: Shaulov, Ariel, et al.
Veröffentlicht: (2025)
Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
Audio-Maestro: Enhancing Large Audio-Language Models with Tool-Augmented Reasoning
von: Lee, Kuan-Yi, et al.
Veröffentlicht: (2025)
von: Lee, Kuan-Yi, et al.
Veröffentlicht: (2025)
TAC: Timestamped Audio Captioning
von: Kumar, Sonal, et al.
Veröffentlicht: (2026)
von: Kumar, Sonal, et al.
Veröffentlicht: (2026)
Improving Audio-Text Retrieval via Hierarchical Cross-Modal Interaction and Auxiliary Captions
von: Xin, Yifei, et al.
Veröffentlicht: (2023)
von: Xin, Yifei, et al.
Veröffentlicht: (2023)
Separate First, Fuse Later: Mitigating Cross-Modal Interference in Audio-Visual LLMs Reasoning with Modality-Specific Chain-of-Thought
von: Li, Xuanchen, et al.
Veröffentlicht: (2026)
von: Li, Xuanchen, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Diffusion-Link: Diffusion Probabilistic Model for Bridging the Audio-Text Modality Gap
von: Nam, KiHyun, et al.
Veröffentlicht: (2025) -
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport
von: Rho, Kyeongha, et al.
Veröffentlicht: (2025) -
SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning
von: Nam, KiHyun, et al.
Veröffentlicht: (2026) -
Disentangled Representation Learning for Environment-agnostic Speaker Recognition
von: Nam, KiHyun, et al.
Veröffentlicht: (2024) -
MoLT: Mixture of Layer-Wise Tokens for Efficient Audio-Visual Learning
von: Rho, Kyeongha, et al.
Veröffentlicht: (2025)