AntGPT: Can Large Language Models Help Long-term Action Anticipation from Videos?
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Qi, Wang, Shijie, Zhang, Ce, Fu, Changcheng, Do, Minh Quan, Agarwal, Nakul, Lee, Kwonjoon, Sun, Chen |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vamos: Versatile Action Models for Video Understanding
by: Wang, Shijie, et al.
Published: (2023)
by: Wang, Shijie, et al.
Published: (2023)
Can't make an Omelette without Breaking some Eggs: Plausible Action Anticipation using Large Video-Language Models
by: Mittal, Himangi, et al.
Published: (2024)
by: Mittal, Himangi, et al.
Published: (2024)
Bidirectional Action Sequence Learning for Long-term Action Anticipation with Large Language Models
by: Sato, Yuji, et al.
Published: (2025)
by: Sato, Yuji, et al.
Published: (2025)
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos
by: Ghoddoosian, Reza, et al.
Published: (2024)
by: Ghoddoosian, Reza, et al.
Published: (2024)
How Can Objects Help Video-Language Understanding?
by: Tang, Zitian, et al.
Published: (2025)
by: Tang, Zitian, et al.
Published: (2025)
Can Hallucination Correction Improve Video-Language Alignment?
by: Zhao, Lingjun, et al.
Published: (2025)
by: Zhao, Lingjun, et al.
Published: (2025)
Vision and Intention Boost Large Language Model in Long-Term Action Anticipation
by: Cao, Congqi, et al.
Published: (2025)
by: Cao, Congqi, et al.
Published: (2025)
M2D2M: Multi-Motion Generation from Text with Discrete Diffusion Models
by: Chi, Seunggeun, et al.
Published: (2024)
by: Chi, Seunggeun, et al.
Published: (2024)
Action-Guided Attention for Video Action Anticipation
by: Tai, Tsung-Ming, et al.
Published: (2026)
by: Tai, Tsung-Ming, et al.
Published: (2026)
Follow the Rules: Reasoning for Video Anomaly Detection with Large Language Models
by: Yang, Yuchen, et al.
Published: (2024)
by: Yang, Yuchen, et al.
Published: (2024)
Overcoming Multi-step Complexity in Multimodal Theory-of-Mind Reasoning: A Scalable Bayesian Planner
by: Zhang, Chunhui, et al.
Published: (2025)
by: Zhang, Chunhui, et al.
Published: (2025)
Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?
by: Benavent-Lledo, Manuel, et al.
Published: (2025)
by: Benavent-Lledo, Manuel, et al.
Published: (2025)
SpatialAnt: Autonomous Zero-Shot Robot Navigation via Active Scene Reconstruction and Visual Anticipation
by: Zhang, Jiwen, et al.
Published: (2026)
by: Zhang, Jiwen, et al.
Published: (2026)
Multimodal Large Models Are Effective Action Anticipators
by: Wang, Binglu, et al.
Published: (2025)
by: Wang, Binglu, et al.
Published: (2025)
ViCor: Bridging Visual Understanding and Commonsense Reasoning with Large Language Models
by: Zhou, Kaiwen, et al.
Published: (2023)
by: Zhou, Kaiwen, et al.
Published: (2023)
Action Anticipation from SoccerNet Football Video Broadcasts
by: Dalal, Mohamad, et al.
Published: (2025)
by: Dalal, Mohamad, et al.
Published: (2025)
Pose-Aware Weakly-Supervised Action Segmentation
by: Zhao, Seth Z., et al.
Published: (2025)
by: Zhao, Seth Z., et al.
Published: (2025)
Anticipating Object State Changes in Long Procedural Videos
by: Manousaki, Victoria, et al.
Published: (2024)
by: Manousaki, Victoria, et al.
Published: (2024)
Learning Robust Reasoning through Guided Adversarial Self-Play
by: Li, Shuozhe, et al.
Published: (2026)
by: Li, Shuozhe, et al.
Published: (2026)
Efficient Long-Tail Learning in Latent Space by sampling Synthetic Data
by: Sharma, Nakul
Published: (2025)
by: Sharma, Nakul
Published: (2025)
SWAG: Long-term Surgical Workflow Prediction with Generative-based Anticipation
by: Boels, Maxence, et al.
Published: (2024)
by: Boels, Maxence, et al.
Published: (2024)
ViTGAN: Training GANs with Vision Transformers
by: Lee, Kwonjoon, et al.
Published: (2021)
by: Lee, Kwonjoon, et al.
Published: (2021)
Can Uniform Meaning Representation Help GPT-4 Translate from Indigenous Languages?
by: Wein, Shira
Published: (2025)
by: Wein, Shira
Published: (2025)
Intention-Guided Cognitive Reasoning for Egocentric Long-Term Action Anticipation
by: Chu, Qiaohui, et al.
Published: (2025)
by: Chu, Qiaohui, et al.
Published: (2025)
Uncertainty-boosted Robust Video Activity Anticipation
by: Qi, Zhaobo, et al.
Published: (2024)
by: Qi, Zhaobo, et al.
Published: (2024)
Semantically Guided Action Anticipation
by: Diko, Anxhelo, et al.
Published: (2024)
by: Diko, Anxhelo, et al.
Published: (2024)
Exploring Social Desirability Response Bias in Large Language Models: Evidence from GPT-4 Simulations
by: Lee, Sanguk, et al.
Published: (2024)
by: Lee, Sanguk, et al.
Published: (2024)
Can Large Language Models Automatically Jailbreak GPT-4V?
by: Wu, Yuanwei, et al.
Published: (2024)
by: Wu, Yuanwei, et al.
Published: (2024)
Bloated Disclosures: Can ChatGPT Help Investors Process Information?
by: Kim, Alex, et al.
Published: (2023)
by: Kim, Alex, et al.
Published: (2023)
Human Action Anticipation: A Survey
by: Lai, Bolin, et al.
Published: (2024)
by: Lai, Bolin, et al.
Published: (2024)
Task-Aware Resolution Optimization for Visual Large Language Models
by: Luo, Weiqing, et al.
Published: (2025)
by: Luo, Weiqing, et al.
Published: (2025)
Technical Report for Ego4D Long-Term Action Anticipation Challenge 2025
by: Chu, Qiaohui, et al.
Published: (2025)
by: Chu, Qiaohui, et al.
Published: (2025)
Anticipating Innovation Using Large Language Models
by: Fenoaltea, Enrico Maria, et al.
Published: (2026)
by: Fenoaltea, Enrico Maria, et al.
Published: (2026)
Does the Market Anticipate? Can it? Should it?
by: Wren, Kangda Ken
Published: (2026)
by: Wren, Kangda Ken
Published: (2026)
SurgAnt-ViVQA: Learning to Anticipate Surgical Events through GRU-Driven Temporal Cross-Attention
by: Dhake, Shreyas C., et al.
Published: (2025)
by: Dhake, Shreyas C., et al.
Published: (2025)
Importance Weighting Can Help Large Language Models Self-Improve
by: Jiang, Chunyang, et al.
Published: (2024)
by: Jiang, Chunyang, et al.
Published: (2024)
Can Large Language Models Help Experimental Design for Causal Discovery?
by: Li, Junyi, et al.
Published: (2025)
by: Li, Junyi, et al.
Published: (2025)
Human-like Working Memory Interference in Large Language Models
by: Xiong, Hua-Dong, et al.
Published: (2026)
by: Xiong, Hua-Dong, et al.
Published: (2026)
Long-term Pre-training for Temporal Action Detection with Transformers
by: Kim, Jihwan, et al.
Published: (2024)
by: Kim, Jihwan, et al.
Published: (2024)
Can Small Language Models Help Large Language Models Reason Better?: LM-Guided Chain-of-Thought
by: Lee, Jooyoung, et al.
Published: (2024)
by: Lee, Jooyoung, et al.
Published: (2024)
Similar Items
-
Vamos: Versatile Action Models for Video Understanding
by: Wang, Shijie, et al.
Published: (2023) -
Can't make an Omelette without Breaking some Eggs: Plausible Action Anticipation using Large Video-Language Models
by: Mittal, Himangi, et al.
Published: (2024) -
Bidirectional Action Sequence Learning for Long-term Action Anticipation with Large Language Models
by: Sato, Yuji, et al.
Published: (2025) -
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos
by: Ghoddoosian, Reza, et al.
Published: (2024) -
How Can Objects Help Video-Language Understanding?
by: Tang, Zitian, et al.
Published: (2025)