Uncertainty-boosted Robust Video Activity Anticipation
Fuente:
arXiv
Saved in:
| Main Authors: | Qi, Zhaobo, Wang, Shuhui, Zhang, Weigang, Huang, Qingming |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bias-Conflict Sample Synthesis and Adversarial Removal Debias Strategy for Temporal Sentence Grounding in Video
by: Qi, Zhaobo, et al.
Published: (2024)
by: Qi, Zhaobo, et al.
Published: (2024)
Masked Temporal Interpolation Diffusion for Procedure Planning in Instructional Videos
by: Zhou, Yufan, et al.
Published: (2025)
by: Zhou, Yufan, et al.
Published: (2025)
Enhancing Pre-trained Representation Classifiability can Boost its Interpretability
by: Shen, Shufan, et al.
Published: (2025)
by: Shen, Shufan, et al.
Published: (2025)
Prompting Video-Language Foundation Models with Domain-specific Fine-grained Heuristics for Video Question Answering
by: Yu, Ting, et al.
Published: (2024)
by: Yu, Ting, et al.
Published: (2024)
VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept Set
by: Shen, Shufan, et al.
Published: (2025)
by: Shen, Shufan, et al.
Published: (2025)
Kernelized Sparse Fine-Tuning with Bi-level Parameter Competition for Vision Models
by: Shen, Shufan, et al.
Published: (2025)
by: Shen, Shufan, et al.
Published: (2025)
Limb-Aware Virtual Try-On Network with Progressive Clothing Warping
by: Zhang, Shengping, et al.
Published: (2025)
by: Zhang, Shengping, et al.
Published: (2025)
A Comprehensive Survey of 3D Dense Captioning: Localizing and Describing Objects in 3D Scenes
by: Yu, Ting, et al.
Published: (2024)
by: Yu, Ting, et al.
Published: (2024)
Expanding Sparse Tuning for Low Memory Usage
by: Shen, Shufan, et al.
Published: (2024)
by: Shen, Shufan, et al.
Published: (2024)
STAND: Semantic Anchoring Constraint with Dual-Granularity Disambiguation for Remote Sensing Image Change Captioning
by: Gong, Yanpei, et al.
Published: (2026)
by: Gong, Yanpei, et al.
Published: (2026)
AIFIND: Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery Detection
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
SOVC: Subject-Oriented Video Captioning
by: Teng, Chang, et al.
Published: (2023)
by: Teng, Chang, et al.
Published: (2023)
TowerDataset: A Heterogeneous Benchmark for Transmission Corridor Segmentation with a Global-Local Fusion Framework
by: Cui, Xu, et al.
Published: (2026)
by: Cui, Xu, et al.
Published: (2026)
EVA: An Embodied World Model for Future Video Anticipation
by: Chi, Xiaowei, et al.
Published: (2024)
by: Chi, Xiaowei, et al.
Published: (2024)
Anticipating Next Active Objects for Egocentric Videos
by: Thakur, Sanket, et al.
Published: (2023)
by: Thakur, Sanket, et al.
Published: (2023)
Action-Guided Attention for Video Action Anticipation
by: Tai, Tsung-Ming, et al.
Published: (2026)
by: Tai, Tsung-Ming, et al.
Published: (2026)
AntGPT: Can Large Language Models Help Long-term Action Anticipation from Videos?
by: Zhao, Qi, et al.
Published: (2023)
by: Zhao, Qi, et al.
Published: (2023)
Locate-then-Sparsify: Attribution Guided Sparse Strategy for Visual Hallucination Mitigation
by: Dang, Tiantian, et al.
Published: (2026)
by: Dang, Tiantian, et al.
Published: (2026)
FRAME: Pre-Training Video Feature Representations via Anticipation and Memory
by: TV, Sethuraman, et al.
Published: (2025)
by: TV, Sethuraman, et al.
Published: (2025)
Anticipating Object State Changes in Long Procedural Videos
by: Manousaki, Victoria, et al.
Published: (2024)
by: Manousaki, Victoria, et al.
Published: (2024)
Uncertainty-aware Long-tailed Weights Model the Utility of Pseudo-labels for Semi-supervised Learning
by: Wu, Jiaqi, et al.
Published: (2025)
by: Wu, Jiaqi, et al.
Published: (2025)
CANeRV: Content Adaptive Neural Representation for Video Compression
by: Tang, Lv, et al.
Published: (2025)
by: Tang, Lv, et al.
Published: (2025)
RETTA: Retrieval-Enhanced Test-Time Adaptation for Zero-Shot Video Captioning
by: Ma, Yunchuan, et al.
Published: (2024)
by: Ma, Yunchuan, et al.
Published: (2024)
Collaborative Temporal Consistency Learning for Point-supervised Natural Language Video Localization
by: Tao, Zhuo, et al.
Published: (2025)
by: Tao, Zhuo, et al.
Published: (2025)
Not All Pairs are Equal: Hierarchical Learning for Average-Precision-Oriented Video Retrieval
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
From Detection to Anticipation: Online Understanding of Struggles across Various Tasks and Activities
by: Feng, Shijia, et al.
Published: (2025)
by: Feng, Shijia, et al.
Published: (2025)
HyperGLM: HyperGraph for Video Scene Graph Generation and Anticipation
by: Nguyen, Trong-Thuan, et al.
Published: (2024)
by: Nguyen, Trong-Thuan, et al.
Published: (2024)
EgoInteract: Synthetic Egocentric Videos Generation for Interaction Understanding and Anticipation
by: Leonardi, Rosario, et al.
Published: (2026)
by: Leonardi, Rosario, et al.
Published: (2026)
Bootstrapping Physics-Grounded Video Generation through VLM-Guided Iterative Self-Refinement
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Towards Unbiased and Robust Spatio-Temporal Scene Graph Generation and Anticipation
by: Peddi, Rohith, et al.
Published: (2024)
by: Peddi, Rohith, et al.
Published: (2024)
When the Future Becomes the Past: Taming Temporal Correspondence for Self-supervised Video Representation Learning
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Enhancing Sample Utilization in Noise-Robust Deep Metric Learning With Subgroup-Based Positive-Pair Selection
by: Yu, Zhipeng, et al.
Published: (2025)
by: Yu, Zhipeng, et al.
Published: (2025)
The Devil is in the Distributions: Explicit Modeling of Scene Content is Key in Zero-Shot Video Captioning
by: Tian, Mingkai, et al.
Published: (2025)
by: Tian, Mingkai, et al.
Published: (2025)
Robust Partial 3D Point Cloud Registration via Confidence Estimation under Global Context
by: Wang, Yongqiang, et al.
Published: (2025)
by: Wang, Yongqiang, et al.
Published: (2025)
COMICS: End-to-end Bi-grained Contrastive Learning for Multi-face Forgery Detection
by: Zhang, Cong, et al.
Published: (2023)
by: Zhang, Cong, et al.
Published: (2023)
Multimodal Large Models Are Effective Action Anticipators
by: Wang, Binglu, et al.
Published: (2025)
by: Wang, Binglu, et al.
Published: (2025)
From Static to Dynamic: Exploring Self-supervised Image-to-Video Representation Transfer Learning
by: Liu, Yang, et al.
Published: (2026)
by: Liu, Yang, et al.
Published: (2026)
Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?
by: Benavent-Lledo, Manuel, et al.
Published: (2025)
by: Benavent-Lledo, Manuel, et al.
Published: (2025)
Decorrelating Structure via Adapters Makes Ensemble Learning Practical for Semi-supervised Learning
by: Wu, Jiaqi, et al.
Published: (2024)
by: Wu, Jiaqi, et al.
Published: (2024)
Desensitizing for Improving Corruption Robustness in Point Cloud Classification through Adversarial Training
by: Tian, Zhiqiang, et al.
Published: (2025)
by: Tian, Zhiqiang, et al.
Published: (2025)
Similar Items
-
Bias-Conflict Sample Synthesis and Adversarial Removal Debias Strategy for Temporal Sentence Grounding in Video
by: Qi, Zhaobo, et al.
Published: (2024) -
Masked Temporal Interpolation Diffusion for Procedure Planning in Instructional Videos
by: Zhou, Yufan, et al.
Published: (2025) -
Enhancing Pre-trained Representation Classifiability can Boost its Interpretability
by: Shen, Shufan, et al.
Published: (2025) -
Prompting Video-Language Foundation Models with Domain-specific Fine-grained Heuristics for Video Question Answering
by: Yu, Ting, et al.
Published: (2024) -
VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept Set
by: Shen, Shufan, et al.
Published: (2025)