ContextDet: Temporal Action Detection with Adaptive Context Aggregation
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Ning, Xiao, Yun, Peng, Xiaopeng, Chang, Xiaojun, Wang, Xuanhong, Fang, Dingyi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Detecting Multimodal Situations with Insufficient Context and Abstaining from Baseless Predictions
by: Liu, Junzhang, et al.
Published: (2024)
by: Liu, Junzhang, et al.
Published: (2024)
TIDE : Temporal-Aware Sparse Autoencoders for Interpretable Diffusion Transformers in Image Generation
by: Huang, Victor Shea-Jay, et al.
Published: (2025)
by: Huang, Victor Shea-Jay, et al.
Published: (2025)
Multiple Contexts and Frequencies Aggregation Network forDeepfake Detection
by: Li, Zifeng, et al.
Published: (2024)
by: Li, Zifeng, et al.
Published: (2024)
Label-anticipated Event Disentanglement for Audio-Visual Video Parsing
by: Zhou, Jinxing, et al.
Published: (2024)
by: Zhou, Jinxing, et al.
Published: (2024)
VidCtx: Context-aware Video Question Answering with Image Models
by: Goulas, Andreas, et al.
Published: (2024)
by: Goulas, Andreas, et al.
Published: (2024)
ContextualStory: Consistent Visual Storytelling with Spatially-Enhanced and Storyline Context
by: Zheng, Sixiao, et al.
Published: (2024)
by: Zheng, Sixiao, et al.
Published: (2024)
VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control
by: Bian, Yuxuan, et al.
Published: (2025)
by: Bian, Yuxuan, et al.
Published: (2025)
Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence
by: Meng, Jiahao, et al.
Published: (2025)
by: Meng, Jiahao, et al.
Published: (2025)
Multimodal Long Video Modeling Based on Temporal Dynamic Context
by: Hao, Haoran, et al.
Published: (2025)
by: Hao, Haoran, et al.
Published: (2025)
IVAC-P2L: Leveraging Irregular Repetition Priors for Improving Video Action Counting
by: Wang, Hang, et al.
Published: (2024)
by: Wang, Hang, et al.
Published: (2024)
SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection
by: Qi, Peng, et al.
Published: (2024)
by: Qi, Peng, et al.
Published: (2024)
DeformTrace: A Deformable State Space Model with Relay Tokens for Temporal Forgery Localization
by: Zhu, Xiaodong, et al.
Published: (2026)
by: Zhu, Xiaodong, et al.
Published: (2026)
MSCrackMamba: Leveraging Vision Mamba for Crack Detection in Fused Multispectral Imagery
by: Zhu, Qinfeng, et al.
Published: (2024)
by: Zhu, Qinfeng, et al.
Published: (2024)
AutoAWG: Adverse Weather Generation with Adaptive Multi-Controls for Automotive Videos
by: Hu, Jiagao, et al.
Published: (2026)
by: Hu, Jiagao, et al.
Published: (2026)
Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation
by: Zhang, Zhicheng, et al.
Published: (2026)
by: Zhang, Zhicheng, et al.
Published: (2026)
LEAF-Mamba: Local Emphatic and Adaptive Fusion State Space Model for RGB-D Salient Object Detection
by: Wu, Lanhu, et al.
Published: (2025)
by: Wu, Lanhu, et al.
Published: (2025)
IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection
by: Guan, Hong, et al.
Published: (2024)
by: Guan, Hong, et al.
Published: (2024)
PanMatch: Unleashing the Potential of Large Vision Models for Unified Matching Models
by: Zhang, Yongjian, et al.
Published: (2025)
by: Zhang, Yongjian, et al.
Published: (2025)
Enhancing Visible-Infrared Person Re-identification with Modality- and Instance-aware Visual Prompt Learning
by: Wu, Ruiqi, et al.
Published: (2024)
by: Wu, Ruiqi, et al.
Published: (2024)
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models
by: Sun, Zeyi, et al.
Published: (2024)
by: Sun, Zeyi, et al.
Published: (2024)
Large-scale Multi-Modal Pre-trained Models: A Comprehensive Survey
by: Wang, Xiao, et al.
Published: (2023)
by: Wang, Xiao, et al.
Published: (2023)
FlexiAct: Towards Flexible Action Control in Heterogeneous Scenarios
by: Zhang, Shiyi, et al.
Published: (2025)
by: Zhang, Shiyi, et al.
Published: (2025)
Moiré Video Authentication: A Physical Signature Against AI Video Generation
by: Qing, Yuan, et al.
Published: (2026)
by: Qing, Yuan, et al.
Published: (2026)
Official-NV: An LLM-Generated News Video Dataset for Multimodal Fake News Detection
by: Wang, Yihao, et al.
Published: (2024)
by: Wang, Yihao, et al.
Published: (2024)
URMF: Uncertainty-aware Robust Multimodal Fusion for Multimodal Sarcasm Detection
by: Wang, Zhenyu, et al.
Published: (2026)
by: Wang, Zhenyu, et al.
Published: (2026)
LMVD: A Large-Scale Multimodal Vlog Dataset for Depression Detection in the Wild
by: He, Lang, et al.
Published: (2024)
by: He, Lang, et al.
Published: (2024)
Balancing Privacy and Action Performance: A Penalty-Driven Approach to Image Anonymization
by: Aslam, Nazia, et al.
Published: (2025)
by: Aslam, Nazia, et al.
Published: (2025)
CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models
by: Poppi, Tobia, et al.
Published: (2026)
by: Poppi, Tobia, et al.
Published: (2026)
Grounding is All You Need? Dual Temporal Grounding for Video Dialog
by: Qin, You, et al.
Published: (2024)
by: Qin, You, et al.
Published: (2024)
Exploring the latent space of diffusion models directly through singular value decomposition
by: Wang, Li, et al.
Published: (2025)
by: Wang, Li, et al.
Published: (2025)
WAVECLIP: Wavelet Tokenization for Adaptive-Resolution CLIP
by: Kimhi, Moshe, et al.
Published: (2025)
by: Kimhi, Moshe, et al.
Published: (2025)
Enhancing Self-Supervised Talking Head Forgery Detection via a Training-Free Dual-System Framework
by: Liu, Ke, et al.
Published: (2026)
by: Liu, Ke, et al.
Published: (2026)
Detecting Misinformation in Multimedia Content through Cross-Modal Entity Consistency: A Dual Learning Approach
by: Fu, Zhe, et al.
Published: (2024)
by: Fu, Zhe, et al.
Published: (2024)
Understanding Temporal Logic Consistency in Video-Language Models through Cross-Modal Attention Discriminability
by: Li, Chengzhi, et al.
Published: (2025)
by: Li, Chengzhi, et al.
Published: (2025)
VIA: Unified Spatiotemporal Video Adaptation Framework for Global and Local Video Editing
by: Gu, Jing, et al.
Published: (2024)
by: Gu, Jing, et al.
Published: (2024)
AdaCoder: Adaptive Prompt Compression for Programmatic Visual Question Answering
by: Ukai, Mahiro, et al.
Published: (2024)
by: Ukai, Mahiro, et al.
Published: (2024)
Robust Fuzzy Multi-view Learning under View Conflict
by: Duan, Siyuan, et al.
Published: (2026)
by: Duan, Siyuan, et al.
Published: (2026)
RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Racket Analysis
by: Dong, Linfeng, et al.
Published: (2025)
by: Dong, Linfeng, et al.
Published: (2025)
Boosting Facial Action Unit Detection Through Jointly Learning Facial Landmark Detection and Domain Separation and Reconstruction
by: Shang, Ziqiao, et al.
Published: (2023)
by: Shang, Ziqiao, et al.
Published: (2023)
OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation
by: Gan, Qijun, et al.
Published: (2025)
by: Gan, Qijun, et al.
Published: (2025)
Similar Items
-
Detecting Multimodal Situations with Insufficient Context and Abstaining from Baseless Predictions
by: Liu, Junzhang, et al.
Published: (2024) -
TIDE : Temporal-Aware Sparse Autoencoders for Interpretable Diffusion Transformers in Image Generation
by: Huang, Victor Shea-Jay, et al.
Published: (2025) -
Multiple Contexts and Frequencies Aggregation Network forDeepfake Detection
by: Li, Zifeng, et al.
Published: (2024) -
Label-anticipated Event Disentanglement for Audio-Visual Video Parsing
by: Zhou, Jinxing, et al.
Published: (2024) -
VidCtx: Context-aware Video Question Answering with Image Models
by: Goulas, Andreas, et al.
Published: (2024)