Training-Free Action Recognition and Goal Inference with Dynamic Frame Selection
Fuente:
arXiv
Saved in:
| Main Authors: | Keat, Ee Yeo, Hao, Zhang, Matyasko, Alexander, Fernando, Basura |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RCA: Region Conditioned Adaptation for Visual Abductive Reasoning
by: Zhang, Hao, et al.
Published: (2023)
by: Zhang, Hao, et al.
Published: (2023)
Improving Temporal Action Segmentation via Constraint-Aware Decoding
by: Ee, Yeo Keat, et al.
Published: (2026)
by: Ee, Yeo Keat, et al.
Published: (2026)
Neuro Symbolic Knowledge Reasoning for Procedural Video Question Answering
by: Fernando, Basura, et al.
Published: (2025)
by: Fernando, Basura, et al.
Published: (2025)
IMoRe: Implicit Program-Guided Reasoning for Human Motion Q&A
by: Li, Chen, et al.
Published: (2025)
by: Li, Chen, et al.
Published: (2025)
Predicting the Next Action by Modeling the Abstract Goal
by: Roy, Debaditya, et al.
Published: (2022)
by: Roy, Debaditya, et al.
Published: (2022)
Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion
by: Rawal, Ishaan Singh, et al.
Published: (2023)
by: Rawal, Ishaan Singh, et al.
Published: (2023)
Modelling Spatio-Temporal Interactions For Compositional Action Recognition
by: Rajendiran, Ramanathan, et al.
Published: (2023)
by: Rajendiran, Ramanathan, et al.
Published: (2023)
Inferring Past Human Actions in Homes with Abductive Reasoning
by: Tan, Clement, et al.
Published: (2022)
by: Tan, Clement, et al.
Published: (2022)
Interaction Region Visual Transformer for Egocentric Action Anticipation
by: Roy, Debaditya, et al.
Published: (2022)
by: Roy, Debaditya, et al.
Published: (2022)
TsallisPGD: Adaptive Gradient Weighting for Adversarial Attacks on Semantic Segmentation
by: Matyasko, Alexander, et al.
Published: (2026)
by: Matyasko, Alexander, et al.
Published: (2026)
Mitigating Easy Option Bias in Multiple-Choice Question Answering
by: Zhang, Hao, et al.
Published: (2025)
by: Zhang, Hao, et al.
Published: (2025)
Learning to Visually Connect Actions and their Effects
by: Parmar, Paritosh, et al.
Published: (2024)
by: Parmar, Paritosh, et al.
Published: (2024)
Effectively Leveraging CLIP for Generating Situational Summaries of Images and Videos
by: Verma, Dhruv, et al.
Published: (2024)
by: Verma, Dhruv, et al.
Published: (2024)
HMR3D: Hierarchical Multimodal Representation for 3D Scene Understanding with Large Vision-Language Model
by: Li, Chen, et al.
Published: (2025)
by: Li, Chen, et al.
Published: (2025)
Learning to Generate Long-term Future Narrations Describing Activities of Daily Living
by: Rajendiran, Ramanathan, et al.
Published: (2025)
by: Rajendiran, Ramanathan, et al.
Published: (2025)
Selective Volume Mixup for Video Action Recognition
by: Tan, Yi, et al.
Published: (2023)
by: Tan, Yi, et al.
Published: (2023)
Informative Sample Selection Model for Skeleton-based Action Recognition with Limited Training Samples
by: Tu, Zhigang, et al.
Published: (2025)
by: Tu, Zhigang, et al.
Published: (2025)
MoCrop: Training Free Motion Guided Cropping for Efficient Video Action Recognition
by: Huang, Binhua, et al.
Published: (2025)
by: Huang, Binhua, et al.
Published: (2025)
EventCrab: Harnessing Frame and Point Synergy for Event-based Action Recognition and Beyond
by: Cao, Meiqi, et al.
Published: (2024)
by: Cao, Meiqi, et al.
Published: (2024)
Situational Scene Graph for Structured Human-centric Situation Understanding
by: Sugandhika, Chinthani, et al.
Published: (2024)
by: Sugandhika, Chinthani, et al.
Published: (2024)
VOST-SGG: VLM-Aided One-Stage Spatio-Temporal Scene Graph Generation
by: Sugandhika, Chinthani, et al.
Published: (2025)
by: Sugandhika, Chinthani, et al.
Published: (2025)
Know-Show: Benchmarking Video-Language Models on Spatio-Temporal Grounded Reasoning
by: Sugandhika, Chinthani, et al.
Published: (2025)
by: Sugandhika, Chinthani, et al.
Published: (2025)
Action Selection Learning for Multi-label Multi-view Action Recognition
by: Nguyen, Trung Thanh, et al.
Published: (2024)
by: Nguyen, Trung Thanh, et al.
Published: (2024)
MaxInfo: A Training-Free Key-Frame Selection Method Using Maximum Volume for Enhanced Video Understanding
by: Li, Pengyi, et al.
Published: (2025)
by: Li, Pengyi, et al.
Published: (2025)
Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos
by: Feng, X., et al.
Published: (2026)
by: Feng, X., et al.
Published: (2026)
Frame Order Matters: A Temporal Sequence-Aware Model for Few-Shot Action Recognition
by: Li, Bozheng, et al.
Published: (2024)
by: Li, Bozheng, et al.
Published: (2024)
Spatial-Temporal Perception with Causal Inference for Naturalistic Driving Action Recognition
by: Chang, Qing, et al.
Published: (2025)
by: Chang, Qing, et al.
Published: (2025)
DiffuseSlide: Training-Free High Frame Rate Video Generation Diffusion
by: Hwang, Geunmin, et al.
Published: (2025)
by: Hwang, Geunmin, et al.
Published: (2025)
SVFormer: A Direct Training Spiking Transformer for Efficient Video Action Recognition
by: Yu, Liutao, et al.
Published: (2024)
by: Yu, Liutao, et al.
Published: (2024)
EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models
by: Yang, Yantai, et al.
Published: (2025)
by: Yang, Yantai, et al.
Published: (2025)
FIS-DiT: Breaking the Few-Step Video Inference Barrier via Training-Free Frame Interleaved Sparsity
by: Tang, Jian, et al.
Published: (2026)
by: Tang, Jian, et al.
Published: (2026)
Towards Micro-Action Recognition with Limited Annotations: An Asynchronous Pseudo Labeling and Training Approach
by: Zhang, Yan, et al.
Published: (2025)
by: Zhang, Yan, et al.
Published: (2025)
Action Unit Enhance Dynamic Facial Expression Recognition
by: Liu, Feng, et al.
Published: (2025)
by: Liu, Feng, et al.
Published: (2025)
High-Performance Inference Graph Convolutional Networks for Skeleton-Based Action Recognition
by: Wang, Junyi, et al.
Published: (2023)
by: Wang, Junyi, et al.
Published: (2023)
Frame Guidance: Training-Free Guidance for Frame-Level Control in Video Diffusion Models
by: Jang, Sangwon, et al.
Published: (2025)
by: Jang, Sangwon, et al.
Published: (2025)
GCN-DevLSTM: Path Development for Skeleton-Based Action Recognition
by: Jiang, Lei, et al.
Published: (2024)
by: Jiang, Lei, et al.
Published: (2024)
Training-Free Zero-Shot Temporal Action Detection with Vision-Language Models
by: Han, Chaolei, et al.
Published: (2025)
by: Han, Chaolei, et al.
Published: (2025)
Boosting Skeleton-based Zero-Shot Action Recognition with Training-Free Test-Time Adaptation
by: Zhu, Jingmin, et al.
Published: (2025)
by: Zhu, Jingmin, et al.
Published: (2025)
TSkel-Mamba: Temporal Dynamic Modeling via State Space Model for Human Skeleton-based Action Recognition
by: Liu, Yanan, et al.
Published: (2025)
by: Liu, Yanan, et al.
Published: (2025)
PaPr: Training-Free One-Step Patch Pruning with Lightweight ConvNets for Faster Inference
by: Mahmud, Tanvir, et al.
Published: (2024)
by: Mahmud, Tanvir, et al.
Published: (2024)
Similar Items
-
RCA: Region Conditioned Adaptation for Visual Abductive Reasoning
by: Zhang, Hao, et al.
Published: (2023) -
Improving Temporal Action Segmentation via Constraint-Aware Decoding
by: Ee, Yeo Keat, et al.
Published: (2026) -
Neuro Symbolic Knowledge Reasoning for Procedural Video Question Answering
by: Fernando, Basura, et al.
Published: (2025) -
IMoRe: Implicit Program-Guided Reasoning for Human Motion Q&A
by: Li, Chen, et al.
Published: (2025) -
Predicting the Next Action by Modeling the Abstract Goal
by: Roy, Debaditya, et al.
Published: (2022)