From Recognition to Prediction: Leveraging Sequence Reasoning for Action Anticipation
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Xin, Hao, Chao, Yu, Zitong, Yue, Huanjing, Yang, Jingyu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Adversarial Robustness in RGB-Skeleton Action Recognition: Leveraging Attention Modality Reweighter
por: Liu, Chao, et al.
Publicado: (2024)
por: Liu, Chao, et al.
Publicado: (2024)
A Simple yet Effective Network based on Vision Transformer for Camouflaged Object and Salient Object Detection
por: Hao, Chao, et al.
Publicado: (2024)
por: Hao, Chao, et al.
Publicado: (2024)
AUFormer: Vision Transformers are Parameter-Efficient Facial Action Unit Detectors
por: Yuan, Kaishen, et al.
Publicado: (2024)
por: Yuan, Kaishen, et al.
Publicado: (2024)
Distribution-Specific Learning for Joint Salient and Camouflaged Object Detection
por: Hao, Chao, et al.
Publicado: (2025)
por: Hao, Chao, et al.
Publicado: (2025)
FEALLM: Advancing Facial Emotion Analysis in Multimodal Large Language Models with Emotional Synergy and Reasoning
por: Hu, Zhuozhao, et al.
Publicado: (2025)
por: Hu, Zhuozhao, et al.
Publicado: (2025)
Zero-Shot Video Restoration and Enhancement Using Pre-Trained Image Diffusion Model
por: Cao, Cong, et al.
Publicado: (2024)
por: Cao, Cong, et al.
Publicado: (2024)
Zero-Shot Video Restoration and Enhancement with Assistance of Video Diffusion Models
por: Cao, Cong, et al.
Publicado: (2026)
por: Cao, Cong, et al.
Publicado: (2026)
EMO-LLaMA: Enhancing Facial Emotion Understanding with Instruction Tuning
por: Xing, Bohao, et al.
Publicado: (2024)
por: Xing, Bohao, et al.
Publicado: (2024)
KeDuSR: Real-World Dual-Lens Super-Resolution via Kernel-Free Matching
por: Yue, Huanjing, et al.
Publicado: (2023)
por: Yue, Huanjing, et al.
Publicado: (2023)
F2HDR: Two-Stage HDR Video Reconstruction via Flow Adapter and Physical Motion Modeling
por: Yue, Huanjing, et al.
Publicado: (2026)
por: Yue, Huanjing, et al.
Publicado: (2026)
RISAM: Referring Image Segmentation via Mutual-Aware Attention Features
por: Zhang, Mengxi, et al.
Publicado: (2023)
por: Zhang, Mengxi, et al.
Publicado: (2023)
DeeDSR: Towards Real-World Image Super-Resolution via Degradation-Aware Stable Diffusion
por: Bi, Chunyang, et al.
Publicado: (2024)
por: Bi, Chunyang, et al.
Publicado: (2024)
RViDeformer: Efficient Raw Video Denoising Transformer with a Larger Benchmark Dataset
por: Yue, Huanjing, et al.
Publicado: (2023)
por: Yue, Huanjing, et al.
Publicado: (2023)
Bidirectional Action Sequence Learning for Long-term Action Anticipation with Large Language Models
por: Sato, Yuji, et al.
Publicado: (2025)
por: Sato, Yuji, et al.
Publicado: (2025)
AULLM++: Structural Reasoning with Large Language Models for Micro-Expression Recognition
por: Liu, Zhishu, et al.
Publicado: (2026)
por: Liu, Zhishu, et al.
Publicado: (2026)
Multimodal Large Models Are Effective Action Anticipators
por: Wang, Binglu, et al.
Publicado: (2025)
por: Wang, Binglu, et al.
Publicado: (2025)
Action-Guided Attention for Video Action Anticipation
por: Tai, Tsung-Ming, et al.
Publicado: (2026)
por: Tai, Tsung-Ming, et al.
Publicado: (2026)
Efficient HDR Reconstruction from Real-World Raw Images
por: Yang, Qirui, et al.
Publicado: (2023)
por: Yang, Qirui, et al.
Publicado: (2023)
SpiralDiff: Spiral Diffusion with LoRA for RGB-to-RAW Conversion Across Cameras
por: Yue, Huanjing, et al.
Publicado: (2026)
por: Yue, Huanjing, et al.
Publicado: (2026)
Intention-Guided Cognitive Reasoning for Egocentric Long-Term Action Anticipation
por: Chu, Qiaohui, et al.
Publicado: (2025)
por: Chu, Qiaohui, et al.
Publicado: (2025)
CameraMaster: Unified Camera Semantic-Parameter Control for Photography Retouching
por: Yang, Qirui, et al.
Publicado: (2025)
por: Yang, Qirui, et al.
Publicado: (2025)
SuPRA: Surgical Phase Recognition and Anticipation for Intra-Operative Planning
por: Boels, Maxence, et al.
Publicado: (2024)
por: Boels, Maxence, et al.
Publicado: (2024)
AU-TTT: Vision Test-Time Training model for Facial Action Unit Detection
por: Xing, Bohao, et al.
Publicado: (2025)
por: Xing, Bohao, et al.
Publicado: (2025)
Domain Generalization using Action Sequences for Egocentric Action Recognition
por: Nasirimajd, Amirshayan, et al.
Publicado: (2025)
por: Nasirimajd, Amirshayan, et al.
Publicado: (2025)
Leveraging Temporal Contextualization for Video Action Recognition
por: Kim, Minji, et al.
Publicado: (2024)
por: Kim, Minji, et al.
Publicado: (2024)
Learning Adaptive Lighting via Channel-Aware Guidance
por: Yang, Qirui, et al.
Publicado: (2024)
por: Yang, Qirui, et al.
Publicado: (2024)
SUGAR: Learning Skeleton Representation with Visual-Motion Knowledge for Action Recognition
por: Ye, Qilang, et al.
Publicado: (2025)
por: Ye, Qilang, et al.
Publicado: (2025)
GCN-DevLSTM: Path Development for Skeleton-Based Action Recognition
por: Jiang, Lei, et al.
Publicado: (2024)
por: Jiang, Lei, et al.
Publicado: (2024)
CoStoDet-DDPM: Collaborative Training of Stochastic and Deterministic Models Improves Surgical Workflow Anticipation and Recognition
por: Yang, Kaixiang, et al.
Publicado: (2025)
por: Yang, Kaixiang, et al.
Publicado: (2025)
Frame Order Matters: A Temporal Sequence-Aware Model for Few-Shot Action Recognition
por: Li, Bozheng, et al.
Publicado: (2024)
por: Li, Bozheng, et al.
Publicado: (2024)
Human Action Anticipation: A Survey
por: Lai, Bolin, et al.
Publicado: (2024)
por: Lai, Bolin, et al.
Publicado: (2024)
YUV20K: A Complexity-Driven Benchmark and Trajectory-Aware Alignment Model for Video Camouflaged Object Detection
por: Liu, Yiyu, et al.
Publicado: (2026)
por: Liu, Yiyu, et al.
Publicado: (2026)
Interaction Region Visual Transformer for Egocentric Action Anticipation
por: Roy, Debaditya, et al.
Publicado: (2022)
por: Roy, Debaditya, et al.
Publicado: (2022)
Understanding Multimodal Complementarity for Single-Frame Action Anticipation
por: Benavent-Lledo, Manuel, et al.
Publicado: (2026)
por: Benavent-Lledo, Manuel, et al.
Publicado: (2026)
A Survey on Deep Learning Techniques for Action Anticipation
por: Zhong, Zeyun, et al.
Publicado: (2023)
por: Zhong, Zeyun, et al.
Publicado: (2023)
Vision and Intention Boost Large Language Model in Long-Term Action Anticipation
por: Cao, Congqi, et al.
Publicado: (2025)
por: Cao, Congqi, et al.
Publicado: (2025)
StyleQoRA: Quality-Aware Low-Rank Adaptation for Few-Shot Multi-Style Editing
por: Cao, Cong, et al.
Publicado: (2025)
por: Cao, Cong, et al.
Publicado: (2025)
Accident Anticipation via Temporal Occurrence Prediction
por: Zhao, Tianhao, et al.
Publicado: (2025)
por: Zhao, Tianhao, et al.
Publicado: (2025)
Answering Diverse Questions via Text Attached with Key Audio-Visual Clues
por: Ye, Qilang, et al.
Publicado: (2024)
por: Ye, Qilang, et al.
Publicado: (2024)
Leveraging Foundation Models for Multimodal Graph-Based Action Recognition
por: Ziaeetabar, Fatemeh, et al.
Publicado: (2025)
por: Ziaeetabar, Fatemeh, et al.
Publicado: (2025)
Ejemplares similares
-
Adversarial Robustness in RGB-Skeleton Action Recognition: Leveraging Attention Modality Reweighter
por: Liu, Chao, et al.
Publicado: (2024) -
A Simple yet Effective Network based on Vision Transformer for Camouflaged Object and Salient Object Detection
por: Hao, Chao, et al.
Publicado: (2024) -
AUFormer: Vision Transformers are Parameter-Efficient Facial Action Unit Detectors
por: Yuan, Kaishen, et al.
Publicado: (2024) -
Distribution-Specific Learning for Joint Salient and Camouflaged Object Detection
por: Hao, Chao, et al.
Publicado: (2025) -
FEALLM: Advancing Facial Emotion Analysis in Multimodal Large Language Models with Emotional Synergy and Reasoning
por: Hu, Zhuozhao, et al.
Publicado: (2025)