VideoINSTA: Zero-shot Long Video Understanding via Informative Spatial-Temporal Reasoning with LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Liao, Ruotong, Erler, Max, Wang, Huiyu, Zhai, Guangyao, Zhang, Gengyuan, Ma, Yunpu, Tresp, Volker |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GenTKG: Generative Forecasting on Temporal Knowledge Graph with Large Language Models
by: Liao, Ruotong, et al.
Published: (2023)
by: Liao, Ruotong, et al.
Published: (2023)
zrLLM: Zero-Shot Relational Learning on Temporal Knowledge Graphs with Large Language Models
by: Ding, Zifeng, et al.
Published: (2023)
by: Ding, Zifeng, et al.
Published: (2023)
METok: Multi-Stage Event-based Token Compression for Efficient Long Video Understanding
by: Wang, Mengyue, et al.
Published: (2025)
by: Wang, Mengyue, et al.
Published: (2025)
Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries
by: Amoroso, Roberto, et al.
Published: (2024)
by: Amoroso, Roberto, et al.
Published: (2024)
ReEXplore: Improving MLLMs for Embodied Exploration with Contextualized Retrospective Experience Replay
by: Zhang, Gengyuan, et al.
Published: (2025)
by: Zhang, Gengyuan, et al.
Published: (2025)
Multi-event Video-Text Retrieval
by: Zhang, Gengyuan, et al.
Published: (2023)
by: Zhang, Gengyuan, et al.
Published: (2023)
Memory Helps, but Confabulation Misleads: Understanding Streaming Events in Videos with MLLMs
by: Zhang, Gengyuan, et al.
Published: (2025)
by: Zhang, Gengyuan, et al.
Published: (2025)
When and Where do Events Switch in Multi-Event Video Generation?
by: Liao, Ruotong, et al.
Published: (2025)
by: Liao, Ruotong, et al.
Published: (2025)
PyraVid: Hierarchical Multimodal Memory for Long-Horizon Video Reasoning
by: Yan, Sikuan, et al.
Published: (2026)
by: Yan, Sikuan, et al.
Published: (2026)
Can Vision-Language Models be a Good Guesser? Exploring VLMs for Times and Location Reasoning
by: Zhang, Gengyuan, et al.
Published: (2023)
by: Zhang, Gengyuan, et al.
Published: (2023)
Agentic Neural Networks: Self-Evolving Multi-Agent Systems via Textual Backpropagation
by: Ma, Xiaowen, et al.
Published: (2025)
by: Ma, Xiaowen, et al.
Published: (2025)
Differentiable Quantum Architecture Search For Job Shop Scheduling Problem
by: Sun, Yize, et al.
Published: (2024)
by: Sun, Yize, et al.
Published: (2024)
Quantum Architecture Search with Unsupervised Representation Learning
by: Sun, Yize, et al.
Published: (2024)
by: Sun, Yize, et al.
Published: (2024)
Bayes or Heisenberg: Who(se) Rules?
by: Tresp, Volker, et al.
Published: (2025)
by: Tresp, Volker, et al.
Published: (2025)
Localizing Events in Videos with Multimodal Queries
by: Zhang, Gengyuan, et al.
Published: (2024)
by: Zhang, Gengyuan, et al.
Published: (2024)
TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos
by: Fateh, Fawad Javed, et al.
Published: (2024)
by: Fateh, Fawad Javed, et al.
Published: (2024)
DyGMamba: Efficiently Modeling Long-Term Temporal Dependency on Continuous-Time Dynamic Graphs with State Space Models
by: Ding, Zifeng, et al.
Published: (2024)
by: Ding, Zifeng, et al.
Published: (2024)
Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
by: Chen, Houlun, et al.
Published: (2026)
by: Chen, Houlun, et al.
Published: (2026)
Routing-Free Mixture-of-Experts
by: Liu, Yilun, et al.
Published: (2026)
by: Liu, Yilun, et al.
Published: (2026)
SA-DQAS: Self-attention Enhanced Differentiable Quantum Architecture Search
by: Sun, Yize, et al.
Published: (2024)
by: Sun, Yize, et al.
Published: (2024)
Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents
by: Yan, Sikuan, et al.
Published: (2026)
by: Yan, Sikuan, et al.
Published: (2026)
WebPilot: A Versatile and Autonomous Multi-Agent System for Web Task Execution with Strategic Exploration
by: Zhang, Yao, et al.
Published: (2024)
by: Zhang, Yao, et al.
Published: (2024)
Parameter-Efficient Routed Fine-Tuning: Mixture-of-Experts Demands Mixture of Adaptation Modules
by: Liu, Yilun, et al.
Published: (2025)
by: Liu, Yilun, et al.
Published: (2025)
OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model
by: Zhou, Xingcheng, et al.
Published: (2025)
by: Zhou, Xingcheng, et al.
Published: (2025)
FedNano: Toward Lightweight Federated Tuning for Pretrained Multimodal Large Language Models
by: Zhang, Yao, et al.
Published: (2025)
by: Zhang, Yao, et al.
Published: (2025)
SwarmAgentic: Towards Fully Automated Agentic System Generation via Swarm Intelligence
by: Zhang, Yao, et al.
Published: (2025)
by: Zhang, Yao, et al.
Published: (2025)
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
by: Cheng, Zesen, et al.
Published: (2024)
by: Cheng, Zesen, et al.
Published: (2024)
Temporal Fact Reasoning over Hyper-Relational Knowledge Graphs
by: Ding, Zifeng, et al.
Published: (2023)
by: Ding, Zifeng, et al.
Published: (2023)
Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
by: Hu, Pengfei, et al.
Published: (2025)
by: Hu, Pengfei, et al.
Published: (2025)
INSTA-YOLO: Real-Time Instance Segmentation
by: Mohamed, Eslam, et al.
Published: (2021)
by: Mohamed, Eslam, et al.
Published: (2021)
Improving Perturbation-based Explanations by Understanding the Role of Uncertainty Calibration
by: Decker, Thomas, et al.
Published: (2025)
by: Decker, Thomas, et al.
Published: (2025)
First Experience with Real-Time Control Using Simulated VQC-Based Quantum Policies
by: Sun, Yize, et al.
Published: (2025)
by: Sun, Yize, et al.
Published: (2025)
PERFT: Parameter-Efficient Routed Fine-Tuning for Mixture-of-Expert Model
by: Liu, Yilun, et al.
Published: (2024)
by: Liu, Yilun, et al.
Published: (2024)
LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering
by: Bi, Jinhe, et al.
Published: (2024)
by: Bi, Jinhe, et al.
Published: (2024)
VideoCompressa: Data-Efficient Video Understanding via Joint Temporal Compression and Spatial Reconstruction
by: Wang, Shaobo, et al.
Published: (2025)
by: Wang, Shaobo, et al.
Published: (2025)
VideoPro: Adaptive Program Reasoning for Long Video Understanding
by: Li, Chenglin, et al.
Published: (2025)
by: Li, Chenglin, et al.
Published: (2025)
Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
by: Xu, Yifang, et al.
Published: (2025)
by: Xu, Yifang, et al.
Published: (2025)
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
by: Fu, Honghao, et al.
Published: (2026)
by: Fu, Honghao, et al.
Published: (2026)
ESVQA: Perceptual Quality Assessment of Egocentric Spatial Videos
by: Zhu, Xilei, et al.
Published: (2024)
by: Zhu, Xilei, et al.
Published: (2024)
Tool Verification for Test-Time Reinforcement Learning
by: Liao, Ruotong, et al.
Published: (2026)
by: Liao, Ruotong, et al.
Published: (2026)
Similar Items
-
GenTKG: Generative Forecasting on Temporal Knowledge Graph with Large Language Models
by: Liao, Ruotong, et al.
Published: (2023) -
zrLLM: Zero-Shot Relational Learning on Temporal Knowledge Graphs with Large Language Models
by: Ding, Zifeng, et al.
Published: (2023) -
METok: Multi-Stage Event-based Token Compression for Efficient Long Video Understanding
by: Wang, Mengyue, et al.
Published: (2025) -
Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries
by: Amoroso, Roberto, et al.
Published: (2024) -
ReEXplore: Improving MLLMs for Embodied Exploration with Contextualized Retrospective Experience Replay
by: Zhang, Gengyuan, et al.
Published: (2025)