DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Ahmadian, Mona, Shirian, Amir, Guerin, Frank, Gilbert, Andrew |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FILS: Self-Supervised Video Feature Prediction In Semantic Language Space
by: Ahmadian, Mona, et al.
Published: (2024)
by: Ahmadian, Mona, et al.
Published: (2024)
Locality-aware Cross-modal Correspondence Learning for Dense Audio-Visual Events Localization
by: Xing, Ling, et al.
Published: (2024)
by: Xing, Ling, et al.
Published: (2024)
Dense Audio-Visual Event Localization under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration
by: Zhou, Ziheng, et al.
Published: (2024)
by: Zhou, Ziheng, et al.
Published: (2024)
ESG-Net: Event-Aware Semantic Guided Network for Dense Audio-Visual Event Localization
by: Li, Huilai, et al.
Published: (2025)
by: Li, Huilai, et al.
Published: (2025)
CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event Localization
by: Zhou, Jinxing, et al.
Published: (2025)
by: Zhou, Jinxing, et al.
Published: (2025)
Interpretable Action Recognition on Hard to Classify Actions
by: Anichenko, Anastasia, et al.
Published: (2024)
by: Anichenko, Anastasia, et al.
Published: (2024)
DEAR: Depth-Enhanced Action Recognition
by: Rahmaniboldaji, Sadegh, et al.
Published: (2024)
by: Rahmaniboldaji, Sadegh, et al.
Published: (2024)
Towards Open-Vocabulary Audio-Visual Event Localization
by: Zhou, Jinxing, et al.
Published: (2024)
by: Zhou, Jinxing, et al.
Published: (2024)
AVS-Mamba: Exploring Temporal and Multi-modal Mamba for Audio-Visual Segmentation
by: Gong, Sitong, et al.
Published: (2025)
by: Gong, Sitong, et al.
Published: (2025)
Object-aware Sound Source Localization via Audio-Visual Scene Understanding
by: Um, Sung Jin, et al.
Published: (2025)
by: Um, Sung Jin, et al.
Published: (2025)
InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue
by: Tong, Wenwen, et al.
Published: (2025)
by: Tong, Wenwen, et al.
Published: (2025)
CACE-Net: Co-guidance Attention and Contrastive Enhancement for Effective Audio-Visual Event Localization
by: He, Xiang, et al.
Published: (2024)
by: He, Xiang, et al.
Published: (2024)
Marten: Visual Question Answering with Mask Generation for Multi-modal Document Understanding
by: Wang, Zining, et al.
Published: (2025)
by: Wang, Zining, et al.
Published: (2025)
RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models
by: Liu, Yuqi, et al.
Published: (2025)
by: Liu, Yuqi, et al.
Published: (2025)
Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing
by: Tian, Zeyue, et al.
Published: (2026)
by: Tian, Zeyue, et al.
Published: (2026)
EgoAVU: Egocentric Audio-Visual Understanding
by: Seth, Ashish, et al.
Published: (2026)
by: Seth, Ashish, et al.
Published: (2026)
DANTE-AD: Dual-Vision Attention Network for Long-Term Audio Description
by: Deganutti, Adrienne, et al.
Published: (2025)
by: Deganutti, Adrienne, et al.
Published: (2025)
Scaling Dense Event-Stream Pretraining from Visual Foundation Models
by: Chen, Zhiwen, et al.
Published: (2026)
by: Chen, Zhiwen, et al.
Published: (2026)
Detail-Enhanced Intra- and Inter-modal Interaction for Audio-Visual Emotion Recognition
by: Shi, Tong, et al.
Published: (2024)
by: Shi, Tong, et al.
Published: (2024)
Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning
by: Ni, Minheng, et al.
Published: (2024)
by: Ni, Minheng, et al.
Published: (2024)
CAE-AV: Improving Audio-Visual Learning via Cross-modal Interactive Enrichment
by: Hu, Yunzuo, et al.
Published: (2026)
by: Hu, Yunzuo, et al.
Published: (2026)
UniAV: Unified Audio-Visual Perception for Multi-Task Video Event Localization
by: Geng, Tiantian, et al.
Published: (2024)
by: Geng, Tiantian, et al.
Published: (2024)
Multi-modal Semantic Understanding with Contrastive Cross-modal Feature Alignment
by: Zhang, Ming, et al.
Published: (2024)
by: Zhang, Ming, et al.
Published: (2024)
Privacy-Preserving Visual Localization with Event Cameras
by: Kim, Junho, et al.
Published: (2022)
by: Kim, Junho, et al.
Published: (2022)
Efficient Sparse-to-Dense Visual Localization via Compact Gaussian Scene Representation and Accelerated Dense Pose Estimation
by: Li, Zizhuo, et al.
Published: (2026)
by: Li, Zizhuo, et al.
Published: (2026)
Adapting to the Unknown: Training-Free Audio-Visual Event Perception with Dynamic Thresholds
by: Shaar, Eitan, et al.
Published: (2025)
by: Shaar, Eitan, et al.
Published: (2025)
Human-AI Divergence in Ego-centric Action Recognition under Spatial and Spatiotemporal Manipulations
by: Rahmaniboldaji, Sadegh, et al.
Published: (2026)
by: Rahmaniboldaji, Sadegh, et al.
Published: (2026)
Dense360: Dense Understanding from Omnidirectional Panoramas
by: Zhou, Yikang, et al.
Published: (2025)
by: Zhou, Yikang, et al.
Published: (2025)
Spherical World-Locking for Audio-Visual Localization in Egocentric Videos
by: Yun, Heeseung, et al.
Published: (2024)
by: Yun, Heeseung, et al.
Published: (2024)
EventLens: Leveraging Event-Aware Pretraining and Cross-modal Linking Enhances Visual Commonsense Reasoning
by: Ma, Mingjie, et al.
Published: (2024)
by: Ma, Mingjie, et al.
Published: (2024)
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
by: Li, Kunchang, et al.
Published: (2023)
by: Li, Kunchang, et al.
Published: (2023)
WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-Localization
by: Wen, Jiahao, et al.
Published: (2025)
by: Wen, Jiahao, et al.
Published: (2025)
Multi-modal Multi-task Pre-training for Improved Point Cloud Understanding
by: Liu, Liwen, et al.
Published: (2025)
by: Liu, Liwen, et al.
Published: (2025)
PreFM: Online Audio-Visual Event Parsing via Predictive Future Modeling
by: Yu, Xiao, et al.
Published: (2025)
by: Yu, Xiao, et al.
Published: (2025)
Leveraging the Video-level Semantic Consistency of Event for Audio-visual Event Localization
by: Jiang, Yuanyuan, et al.
Published: (2022)
by: Jiang, Yuanyuan, et al.
Published: (2022)
MambaFusion: Height-Fidelity Dense Global Fusion for Multi-modal 3D Object Detection
by: Wang, Hanshi, et al.
Published: (2025)
by: Wang, Hanshi, et al.
Published: (2025)
NYC-Event-VPR: A Large-Scale High-Resolution Event-Based Visual Place Recognition Dataset in Dense Urban Environments
by: Pan, Taiyi, et al.
Published: (2024)
by: Pan, Taiyi, et al.
Published: (2024)
Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech
by: Liu, Rui, et al.
Published: (2024)
by: Liu, Rui, et al.
Published: (2024)
SOUPLE: Enhancing Audio-Visual Localization and Segmentation with Learnable Prompt Contexts
by: Nguyen, Khanh Binh, et al.
Published: (2026)
by: Nguyen, Khanh Binh, et al.
Published: (2026)
Visual Grounding with Multi-modal Conditional Adaptation
by: Yao, Ruilin, et al.
Published: (2024)
by: Yao, Ruilin, et al.
Published: (2024)
Similar Items
-
FILS: Self-Supervised Video Feature Prediction In Semantic Language Space
by: Ahmadian, Mona, et al.
Published: (2024) -
Locality-aware Cross-modal Correspondence Learning for Dense Audio-Visual Events Localization
by: Xing, Ling, et al.
Published: (2024) -
Dense Audio-Visual Event Localization under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration
by: Zhou, Ziheng, et al.
Published: (2024) -
ESG-Net: Event-Aware Semantic Guided Network for Dense Audio-Visual Event Localization
by: Li, Huilai, et al.
Published: (2025) -
CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event Localization
by: Zhou, Jinxing, et al.
Published: (2025)