Leveraging the Video-level Semantic Consistency of Event for Audio-visual Event Localization
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Yuanyuan, Yin, Jianqin, Dang, Yonghao |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ESG-Net: Event-Aware Semantic Guided Network for Dense Audio-Visual Event Localization
by: Li, Huilai, et al.
Published: (2025)
by: Li, Huilai, et al.
Published: (2025)
Audio-visual Event Localization on Portrait Mode Short Videos
by: Liu, Wuyang, et al.
Published: (2025)
by: Liu, Wuyang, et al.
Published: (2025)
Localizing Events in Videos with Multimodal Queries
by: Zhang, Gengyuan, et al.
Published: (2024)
by: Zhang, Gengyuan, et al.
Published: (2024)
CLIP-Powered TASS: Target-Aware Single-Stream Network for Audio-Visual Question Answering
by: Jiang, Yuanyuan, et al.
Published: (2024)
by: Jiang, Yuanyuan, et al.
Published: (2024)
Event-VStream: Event-Driven Real-Time Understanding for Long Video Streams
by: Guo, Zhenghui, et al.
Published: (2026)
by: Guo, Zhenghui, et al.
Published: (2026)
MaskSem: Semantic-Guided Masking for Learning 3D Hybrid High-Order Motion Representation
by: Wei, Wei, et al.
Published: (2025)
by: Wei, Wei, et al.
Published: (2025)
A Two-stream Hybrid CNN-Transformer Network for Skeleton-based Human Interaction Recognition
by: Yin, Ruoqi, et al.
Published: (2023)
by: Yin, Ruoqi, et al.
Published: (2023)
Semantic Event Graphs for Long-Form Video Question Answering
by: Dixit, Aradhya, et al.
Published: (2026)
by: Dixit, Aradhya, et al.
Published: (2026)
Label-anticipated Event Disentanglement for Audio-Visual Video Parsing
by: Zhou, Jinxing, et al.
Published: (2024)
by: Zhou, Jinxing, et al.
Published: (2024)
A Generically Contrastive Spatiotemporal Representation Enhancement for 3D Skeleton Action Recognition
by: Zhang, Shaojie, et al.
Published: (2023)
by: Zhang, Shaojie, et al.
Published: (2023)
EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing
by: Li, Huilai, et al.
Published: (2026)
by: Li, Huilai, et al.
Published: (2026)
CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event Localization
by: Zhou, Jinxing, et al.
Published: (2025)
by: Zhou, Jinxing, et al.
Published: (2025)
When and Where do Events Switch in Multi-Event Video Generation?
by: Liao, Ruotong, et al.
Published: (2025)
by: Liao, Ruotong, et al.
Published: (2025)
LSA: Localized Semantic Alignment for Enhancing Temporal Consistency in Traffic Video Generation
by: Karimov, Mirlan, et al.
Published: (2026)
by: Karimov, Mirlan, et al.
Published: (2026)
Kinematics Modeling Network for Video-based Human Pose Estimation
by: Dang, Yonghao, et al.
Published: (2022)
by: Dang, Yonghao, et al.
Published: (2022)
VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in Videos
by: Liang, Baoyu, et al.
Published: (2025)
by: Liang, Baoyu, et al.
Published: (2025)
UVCG: Leveraging Temporal Consistency for Universal Video Protection
by: Li, KaiZhou, et al.
Published: (2024)
by: Li, KaiZhou, et al.
Published: (2024)
Event-Enhanced Blurry Video Super-Resolution
by: Kai, Dachun, et al.
Published: (2025)
by: Kai, Dachun, et al.
Published: (2025)
Enhancing Long Video Understanding via Hierarchical Event-Based Memory
by: Cheng, Dingxin, et al.
Published: (2024)
by: Cheng, Dingxin, et al.
Published: (2024)
Unleashing the Power of CNN and Transformer for Balanced RGB-Event Video Recognition
by: Wang, Xiao, et al.
Published: (2023)
by: Wang, Xiao, et al.
Published: (2023)
SiT-MLP: A Simple MLP with Point-wise Topology Feature Learning for Skeleton-based Action Recognition
by: Zhang, Shaojie, et al.
Published: (2023)
by: Zhang, Shaojie, et al.
Published: (2023)
ActivityCLIP: Enhancing Group Activity Recognition by Mining Complementary Information from Text to Supplement Image Modality
by: Xu, Guoliang, et al.
Published: (2024)
by: Xu, Guoliang, et al.
Published: (2024)
VERHallu: Evaluating and Mitigating Event Relation Hallucination in Video Large Language Models
by: Zhang, Zefan, et al.
Published: (2026)
by: Zhang, Zefan, et al.
Published: (2026)
ENTER: Event Based Interpretable Reasoning for VideoQA
by: Ayyubi, Hammad, et al.
Published: (2025)
by: Ayyubi, Hammad, et al.
Published: (2025)
V2CE: Video to Continuous Events Simulator
by: Zhang, Zhongyang, et al.
Published: (2023)
by: Zhang, Zhongyang, et al.
Published: (2023)
Uneven Event Modeling for Partially Relevant Video Retrieval
by: Zhu, Sa, et al.
Published: (2025)
by: Zhu, Sa, et al.
Published: (2025)
A Survey of Video Datasets for Grounded Event Understanding
by: Sanders, Kate, et al.
Published: (2024)
by: Sanders, Kate, et al.
Published: (2024)
Abnormal Event Detection In Videos Using Deep Embedding
by: Venkatrayappa, Darshan
Published: (2024)
by: Venkatrayappa, Darshan
Published: (2024)
Question-Answering Dense Video Events
by: Qin, Hangyu, et al.
Published: (2024)
by: Qin, Hangyu, et al.
Published: (2024)
Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios
by: Yan, Peizheng, et al.
Published: (2026)
by: Yan, Peizheng, et al.
Published: (2026)
Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events
by: Liu, Xiaolin, et al.
Published: (2026)
by: Liu, Xiaolin, et al.
Published: (2026)
Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable Events
by: Chinchure, Aditya, et al.
Published: (2024)
by: Chinchure, Aditya, et al.
Published: (2024)
What Happens When: Learning Temporal Orders of Events in Videos
by: Ahn, Daechul, et al.
Published: (2025)
by: Ahn, Daechul, et al.
Published: (2025)
Low-power, Continuous Remote Behavioral Localization with Event Cameras
by: Hamann, Friedhelm, et al.
Published: (2023)
by: Hamann, Friedhelm, et al.
Published: (2023)
F$^3$Set: Towards Analyzing Fast, Frequent, and Fine-grained Events from Videos
by: Liu, Zhaoyu, et al.
Published: (2025)
by: Liu, Zhaoyu, et al.
Published: (2025)
EventSTR: A Benchmark Dataset and Baselines for Event Stream based Scene Text Recognition
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
A Survey: Spatiotemporal Consistency in Video Generation
by: Yin, Zhiyu, et al.
Published: (2025)
by: Yin, Zhiyu, et al.
Published: (2025)
Semi-Supervised Audio-Visual Video Action Recognition with Audio Source Localization Guided Mixup
by: Kang, Seokun, et al.
Published: (2025)
by: Kang, Seokun, et al.
Published: (2025)
EvTexture: Event-driven Texture Enhancement for Video Super-Resolution
by: Kai, Dachun, et al.
Published: (2024)
by: Kai, Dachun, et al.
Published: (2024)
Sherlock: Towards Multi-scene Video Abnormal Event Extraction and Localization via a Global-local Spatial-sensitive LLM
by: Ma, Junxiao, et al.
Published: (2025)
by: Ma, Junxiao, et al.
Published: (2025)
Similar Items
-
ESG-Net: Event-Aware Semantic Guided Network for Dense Audio-Visual Event Localization
by: Li, Huilai, et al.
Published: (2025) -
Audio-visual Event Localization on Portrait Mode Short Videos
by: Liu, Wuyang, et al.
Published: (2025) -
Localizing Events in Videos with Multimodal Queries
by: Zhang, Gengyuan, et al.
Published: (2024) -
CLIP-Powered TASS: Target-Aware Single-Stream Network for Audio-Visual Question Answering
by: Jiang, Yuanyuan, et al.
Published: (2024) -
Event-VStream: Event-Driven Real-Time Understanding for Long Video Streams
by: Guo, Zhenghui, et al.
Published: (2026)