Locality-aware Cross-modal Correspondence Learning for Dense Audio-Visual Events Localization
Fuente:
arXiv
Saved in:
| Main Authors: | Xing, Ling, Qu, Hongyu, Yan, Rui, Shu, Xiangbo, Tang, Jinhui |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding
by: Ahmadian, Mona, et al.
Published: (2025)
by: Ahmadian, Mona, et al.
Published: (2025)
Spatio-temporal Decoupled Knowledge Compensator for Few-Shot Action Recognition
by: Qu, Hongyu, et al.
Published: (2026)
by: Qu, Hongyu, et al.
Published: (2026)
Vision-centric Token Compression in Large Language Model
by: Xing, Ling, et al.
Published: (2025)
by: Xing, Ling, et al.
Published: (2025)
EventCrab: Harnessing Frame and Point Synergy for Event-based Action Recognition and Beyond
by: Cao, Meiqi, et al.
Published: (2024)
by: Cao, Meiqi, et al.
Published: (2024)
Dense Audio-Visual Event Localization under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration
by: Zhou, Ziheng, et al.
Published: (2024)
by: Zhou, Ziheng, et al.
Published: (2024)
ESG-Net: Event-Aware Semantic Guided Network for Dense Audio-Visual Event Localization
by: Li, Huilai, et al.
Published: (2025)
by: Li, Huilai, et al.
Published: (2025)
CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event Localization
by: Zhou, Jinxing, et al.
Published: (2025)
by: Zhou, Jinxing, et al.
Published: (2025)
OmniGaze: Reward-inspired Generalizable Gaze Estimation In The Wild
by: Qu, Hongyu, et al.
Published: (2025)
by: Qu, Hongyu, et al.
Published: (2025)
Learning Clustering-based Prototypes for Compositional Zero-shot Learning
by: Qu, Hongyu, et al.
Published: (2025)
by: Qu, Hongyu, et al.
Published: (2025)
Spatiotemporal-Untrammelled Mixture of Experts for Multi-Person Motion Prediction
by: Yin, Zheng, et al.
Published: (2025)
by: Yin, Zheng, et al.
Published: (2025)
See the Text: From Tokenization to Visual Reading
by: Xing, Ling, et al.
Published: (2025)
by: Xing, Ling, et al.
Published: (2025)
MVP-Shot: Multi-Velocity Progressive-Alignment Framework for Few-Shot Action Recognition
by: Qu, Hongyu, et al.
Published: (2024)
by: Qu, Hongyu, et al.
Published: (2024)
Learning to Produce Semi-dense Correspondences for Visual Localization
by: Giang, Khang Truong, et al.
Published: (2024)
by: Giang, Khang Truong, et al.
Published: (2024)
Towards Open-Vocabulary Audio-Visual Event Localization
by: Zhou, Jinxing, et al.
Published: (2024)
by: Zhou, Jinxing, et al.
Published: (2024)
ASTRA: Let Arbitrary Subjects Transform in Video Editing
by: Shen, Fei, et al.
Published: (2025)
by: Shen, Fei, et al.
Published: (2025)
FTMoMamba: Motion Generation with Frequency and Text State Space Models
by: Li, Chengjian, et al.
Published: (2024)
by: Li, Chengjian, et al.
Published: (2024)
Object-aware Sound Source Localization via Audio-Visual Scene Understanding
by: Um, Sung Jin, et al.
Published: (2025)
by: Um, Sung Jin, et al.
Published: (2025)
MGNet: Learning Correspondences via Multiple Graphs
by: Dai, Luanyuan, et al.
Published: (2024)
by: Dai, Luanyuan, et al.
Published: (2024)
CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization
by: Xia, Rui, et al.
Published: (2025)
by: Xia, Rui, et al.
Published: (2025)
CAE-AV: Improving Audio-Visual Learning via Cross-modal Interactive Enrichment
by: Hu, Yunzuo, et al.
Published: (2026)
by: Hu, Yunzuo, et al.
Published: (2026)
Efficient Sparse-to-Dense Visual Localization via Compact Gaussian Scene Representation and Accelerated Dense Pose Estimation
by: Li, Zizhuo, et al.
Published: (2026)
by: Li, Zizhuo, et al.
Published: (2026)
AlignMamba: Enhancing Multimodal Mamba with Local and Global Cross-modal Alignment
by: Li, Yan, et al.
Published: (2024)
by: Li, Yan, et al.
Published: (2024)
CACE-Net: Co-guidance Attention and Contrastive Enhancement for Effective Audio-Visual Event Localization
by: He, Xiang, et al.
Published: (2024)
by: He, Xiang, et al.
Published: (2024)
TEST-V: TEst-time Support-set Tuning for Zero-shot Video Classification
by: Yan, Rui, et al.
Published: (2025)
by: Yan, Rui, et al.
Published: (2025)
Dynamic in Static: Hybrid Visual Correspondence for Self-Supervised Video Object Segmentation
by: Pei, Gensheng, et al.
Published: (2024)
by: Pei, Gensheng, et al.
Published: (2024)
Attack-Augmentation Mixing-Contrastive Skeletal Representation Learning
by: Xu, Binqian, et al.
Published: (2023)
by: Xu, Binqian, et al.
Published: (2023)
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer
by: Huang, Shaofei, et al.
Published: (2025)
by: Huang, Shaofei, et al.
Published: (2025)
Cross-modal Context-aware Learning for Visual Prompt Guided Multimodal Image Understanding in Remote Sensing
by: Zhang, Xu, et al.
Published: (2025)
by: Zhang, Xu, et al.
Published: (2025)
Weakly-Supervised Learning of Dense Functional Correspondences
by: Stojanov, Stefan, et al.
Published: (2025)
by: Stojanov, Stefan, et al.
Published: (2025)
Privacy-Preserving Visual Localization with Event Cameras
by: Kim, Junho, et al.
Published: (2022)
by: Kim, Junho, et al.
Published: (2022)
Cross-modal Active Complementary Learning with Self-refining Correspondence
by: Qin, Yang, et al.
Published: (2023)
by: Qin, Yang, et al.
Published: (2023)
Modeling State Shifting via Local-Global Distillation for Event-Frame Gaze Tracking
by: Li, Jiading, et al.
Published: (2024)
by: Li, Jiading, et al.
Published: (2024)
UniAV: Unified Audio-Visual Perception for Multi-Task Video Event Localization
by: Geng, Tiantian, et al.
Published: (2024)
by: Geng, Tiantian, et al.
Published: (2024)
Audio-visual Event Localization on Portrait Mode Short Videos
by: Liu, Wuyang, et al.
Published: (2025)
by: Liu, Wuyang, et al.
Published: (2025)
Pic@Point: Cross-Modal Learning by Local and Global Point-Picture Correspondence
by: Herzog, Vencia, et al.
Published: (2024)
by: Herzog, Vencia, et al.
Published: (2024)
CorrAdaptor: Adaptive Local Context Learning for Correspondence Pruning
by: Zhu, Wei, et al.
Published: (2024)
by: Zhu, Wei, et al.
Published: (2024)
SDPL: Shifting-Dense Partition Learning for UAV-View Geo-Localization
by: Chen, Quan, et al.
Published: (2024)
by: Chen, Quan, et al.
Published: (2024)
Spherical World-Locking for Audio-Visual Localization in Egocentric Videos
by: Yun, Heeseung, et al.
Published: (2024)
by: Yun, Heeseung, et al.
Published: (2024)
Global Cross-Modal Geo-Localization: A Million-Scale Dataset and a Physical Consistency Learning Framework
by: Hu, Yutong, et al.
Published: (2026)
by: Hu, Yutong, et al.
Published: (2026)
GPT4Ego: Unleashing the Potential of Pre-trained Models for Zero-Shot Egocentric Action Recognition
by: Dai, Guangzhao, et al.
Published: (2024)
by: Dai, Guangzhao, et al.
Published: (2024)
Similar Items
-
DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding
by: Ahmadian, Mona, et al.
Published: (2025) -
Spatio-temporal Decoupled Knowledge Compensator for Few-Shot Action Recognition
by: Qu, Hongyu, et al.
Published: (2026) -
Vision-centric Token Compression in Large Language Model
by: Xing, Ling, et al.
Published: (2025) -
EventCrab: Harnessing Frame and Point Synergy for Event-based Action Recognition and Beyond
by: Cao, Meiqi, et al.
Published: (2024) -
Dense Audio-Visual Event Localization under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration
by: Zhou, Ziheng, et al.
Published: (2024)