Programmatic Video Prediction Using Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tang, Hao, Ellis, Kevin, Lohit, Suhas, Jones, Michael J., Chatterjee, Moitreya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluating Large Vision-and-Language Models on Children's Mathematical Olympiads
von: Cherian, Anoop, et al.
Veröffentlicht: (2024)
von: Cherian, Anoop, et al.
Veröffentlicht: (2024)
AssemblyBench: Physics-Aware Assembly of Complex Industrial Objects
von: Li, Danrui, et al.
Veröffentlicht: (2026)
von: Li, Danrui, et al.
Veröffentlicht: (2026)
WISE: Weighted Iterative Society-of-Experts for Robust Multimodal Multi-Agent Debate
von: Cherian, Anoop, et al.
Veröffentlicht: (2025)
von: Cherian, Anoop, et al.
Veröffentlicht: (2025)
Understanding Dynamic Compute Allocation in Recurrent Transformers
von: Moosa, Ibraheem Muhammad, et al.
Veröffentlicht: (2026)
von: Moosa, Ibraheem Muhammad, et al.
Veröffentlicht: (2026)
FreBIS: Frequency-Based Stratification for Neural Implicit Surface Representations
von: Sawada, Naoko, et al.
Veröffentlicht: (2025)
von: Sawada, Naoko, et al.
Veröffentlicht: (2025)
PAVE: Patching and Adapting Video Large Language Models
von: Liu, Zhuoming, et al.
Veröffentlicht: (2025)
von: Liu, Zhuoming, et al.
Veröffentlicht: (2025)
MORSE-500: A Programmatically Controllable Video Benchmark to Stress-Test Multimodal Reasoning
von: Cai, Zikui, et al.
Veröffentlicht: (2025)
von: Cai, Zikui, et al.
Veröffentlicht: (2025)
Undermining Image and Text Classification Algorithms Using Adversarial Attacks
von: Lunga, Langalibalele, et al.
Veröffentlicht: (2024)
von: Lunga, Langalibalele, et al.
Veröffentlicht: (2024)
TA-Prompting: Enhancing Video Large Language Models for Dense Video Captioning via Temporal Anchors
von: Cheng, Wei-Yuan, et al.
Veröffentlicht: (2026)
von: Cheng, Wei-Yuan, et al.
Veröffentlicht: (2026)
Multi-Modal Adapter for Vision-Language Models
von: Seputis, Dominykas, et al.
Veröffentlicht: (2024)
von: Seputis, Dominykas, et al.
Veröffentlicht: (2024)
To Trust Or Not To Trust Your Vision-Language Model's Prediction
von: Dong, Hao, et al.
Veröffentlicht: (2025)
von: Dong, Hao, et al.
Veröffentlicht: (2025)
Secure and Storage-Efficient Deep Learning Models for Edge AI Using Automatic Weight Generation
von: Rahaman, Habibur, et al.
Veröffentlicht: (2025)
von: Rahaman, Habibur, et al.
Veröffentlicht: (2025)
VisualPredicator: Learning Abstract World Models with Neuro-Symbolic Predicates for Robot Planning
von: Liang, Yichao, et al.
Veröffentlicht: (2024)
von: Liang, Yichao, et al.
Veröffentlicht: (2024)
Continuous Video Process: Modeling Videos as Continuous Multi-Dimensional Processes for Video Prediction
von: Shrivastava, Gaurav, et al.
Veröffentlicht: (2024)
von: Shrivastava, Gaurav, et al.
Veröffentlicht: (2024)
Rapid Motor Adaptation for Robotic Manipulator Arms
von: Liang, Yichao, et al.
Veröffentlicht: (2023)
von: Liang, Yichao, et al.
Veröffentlicht: (2023)
Nano World Models: A Minimalist Implementation of Future Video Prediction
von: Huang, Siqiao, et al.
Veröffentlicht: (2026)
von: Huang, Siqiao, et al.
Veröffentlicht: (2026)
Reasoning-Guided Grounding: Elevating Video Anomaly Detection through Multimodal Large Language Models
von: Agarwal, Sakshi, et al.
Veröffentlicht: (2026)
von: Agarwal, Sakshi, et al.
Veröffentlicht: (2026)
Causal Decoding for Hallucination-Resistant Multimodal Large Language Models
von: Tan, Shiwei, et al.
Veröffentlicht: (2026)
von: Tan, Shiwei, et al.
Veröffentlicht: (2026)
FILS: Self-Supervised Video Feature Prediction In Semantic Language Space
von: Ahmadian, Mona, et al.
Veröffentlicht: (2024)
von: Ahmadian, Mona, et al.
Veröffentlicht: (2024)
Beyond Human Vision: The Role of Large Vision Language Models in Microscope Image Analysis
von: Verma, Prateek, et al.
Veröffentlicht: (2024)
von: Verma, Prateek, et al.
Veröffentlicht: (2024)
PlanLLM: Video Procedure Planning with Refinable Large Language Models
von: Yang, Dejie, et al.
Veröffentlicht: (2024)
von: Yang, Dejie, et al.
Veröffentlicht: (2024)
Advancing Egocentric Video Question Answering with Multimodal Large Language Models
von: Patel, Alkesh, et al.
Veröffentlicht: (2025)
von: Patel, Alkesh, et al.
Veröffentlicht: (2025)
THRONE: An Object-based Hallucination Benchmark for the Free-form Generations of Large Vision-Language Models
von: Kaul, Prannay, et al.
Veröffentlicht: (2024)
von: Kaul, Prannay, et al.
Veröffentlicht: (2024)
TopoPerception: A Shortcut-Free Evaluation of Global Visual Perception in Large Vision-Language Models
von: Zhou, Wenhao, et al.
Veröffentlicht: (2025)
von: Zhou, Wenhao, et al.
Veröffentlicht: (2025)
Mitigating Hallucinations via Inter-Layer Consistency Aggregation in Large Vision-Language Models
von: Tang, Kai, et al.
Veröffentlicht: (2025)
von: Tang, Kai, et al.
Veröffentlicht: (2025)
Solving Video Inverse Problems Using Image Diffusion Models
von: Kwon, Taesung, et al.
Veröffentlicht: (2024)
von: Kwon, Taesung, et al.
Veröffentlicht: (2024)
Revisiting Feature Prediction for Learning Visual Representations from Video
von: Bardes, Adrien, et al.
Veröffentlicht: (2024)
von: Bardes, Adrien, et al.
Veröffentlicht: (2024)
Saliency-guided Emotion Modeling: Predicting Viewer Reactions from Video Stimuli
von: Yaragoppa, Akhila, et al.
Veröffentlicht: (2025)
von: Yaragoppa, Akhila, et al.
Veröffentlicht: (2025)
Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation
von: Jeong, Hyeonho, et al.
Veröffentlicht: (2024)
von: Jeong, Hyeonho, et al.
Veröffentlicht: (2024)
HourVideo: 1-Hour Video-Language Understanding
von: Chandrasegaran, Keshigeyan, et al.
Veröffentlicht: (2024)
von: Chandrasegaran, Keshigeyan, et al.
Veröffentlicht: (2024)
Bridging Vision Language Models and Symbolic Grounding for Video Question Answering
von: Ma, Haodi, et al.
Veröffentlicht: (2025)
von: Ma, Haodi, et al.
Veröffentlicht: (2025)
Automatic Mapping of Anatomical Landmarks from Free-Text Using Large Language Models: Insights from Llama-2
von: Abdi, Mohamad, et al.
Veröffentlicht: (2024)
von: Abdi, Mohamad, et al.
Veröffentlicht: (2024)
Guiding Video Prediction with Explicit Procedural Knowledge
von: Takenaka, Patrick, et al.
Veröffentlicht: (2024)
von: Takenaka, Patrick, et al.
Veröffentlicht: (2024)
Learning to Generate Rigid Body Interactions with Video Diffusion Models
von: Romero, David, et al.
Veröffentlicht: (2025)
von: Romero, David, et al.
Veröffentlicht: (2025)
Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models
von: Miao, Yanting, et al.
Veröffentlicht: (2026)
von: Miao, Yanting, et al.
Veröffentlicht: (2026)
PlayGen-MoG: Framework for Diverse Multi-Agent Play Generation via Mixture-of-Gaussians Trajectory Prediction
von: Song, Kevin
Veröffentlicht: (2026)
von: Song, Kevin
Veröffentlicht: (2026)
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement
von: Rao, Zhefan, et al.
Veröffentlicht: (2024)
von: Rao, Zhefan, et al.
Veröffentlicht: (2024)
Lightweight Unsupervised Federated Learning with Pretrained Vision Language Model
von: Yan, Hao, et al.
Veröffentlicht: (2024)
von: Yan, Hao, et al.
Veröffentlicht: (2024)
Visual Hallucinations of Multi-modal Large Language Models
von: Huang, Wen, et al.
Veröffentlicht: (2024)
von: Huang, Wen, et al.
Veröffentlicht: (2024)
Effectiveness Assessment of Recent Large Vision-Language Models
von: Jiang, Yao, et al.
Veröffentlicht: (2024)
von: Jiang, Yao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Evaluating Large Vision-and-Language Models on Children's Mathematical Olympiads
von: Cherian, Anoop, et al.
Veröffentlicht: (2024) -
AssemblyBench: Physics-Aware Assembly of Complex Industrial Objects
von: Li, Danrui, et al.
Veröffentlicht: (2026) -
WISE: Weighted Iterative Society-of-Experts for Robust Multimodal Multi-Agent Debate
von: Cherian, Anoop, et al.
Veröffentlicht: (2025) -
Understanding Dynamic Compute Allocation in Recurrent Transformers
von: Moosa, Ibraheem Muhammad, et al.
Veröffentlicht: (2026) -
FreBIS: Frequency-Based Stratification for Neural Implicit Surface Representations
von: Sawada, Naoko, et al.
Veröffentlicht: (2025)