TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shang, Chuyi, You, Amos, Subramanian, Sanjay, Darrell, Trevor, Herzig, Roei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Recursive Visual Programming
von: Ge, Jiaxin, et al.
Veröffentlicht: (2023)
von: Ge, Jiaxin, et al.
Veröffentlicht: (2023)
Latent Implicit Visual Reasoning
von: Li, Kelvin, et al.
Veröffentlicht: (2025)
von: Li, Kelvin, et al.
Veröffentlicht: (2025)
Compositional Chain-of-Thought Prompting for Large Multimodal Models
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning
von: Huang, Brandon, et al.
Veröffentlicht: (2024)
von: Huang, Brandon, et al.
Veröffentlicht: (2024)
DAVE: A VLM Vision Encoder for Document Understanding and Web Agents
von: Huang, Brandon, et al.
Veröffentlicht: (2025)
von: Huang, Brandon, et al.
Veröffentlicht: (2025)
Pose Priors from Language Models
von: Subramanian, Sanjay, et al.
Veröffentlicht: (2024)
von: Subramanian, Sanjay, et al.
Veröffentlicht: (2024)
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
von: Mitra, Chancharik, et al.
Veröffentlicht: (2025)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2025)
Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features
von: Mitra, Chancharik, et al.
Veröffentlicht: (2024)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2024)
An LMM for Efficient Video Understanding via Reinforced Compression of Video Cubes
von: Qi, Ji, et al.
Veröffentlicht: (2025)
von: Qi, Ji, et al.
Veröffentlicht: (2025)
TULIP: Towards Unified Language-Image Pretraining
von: Tang, Zineng, et al.
Veröffentlicht: (2025)
von: Tang, Zineng, et al.
Veröffentlicht: (2025)
From Generated Human Videos to Physically Plausible Robot Trajectories
von: Ni, James, et al.
Veröffentlicht: (2025)
von: Ni, James, et al.
Veröffentlicht: (2025)
Actions and Objects Pathways for Domain Adaptation in Video Question Answering
von: Mohamud, Safaa Abdullahi Moallim, et al.
Veröffentlicht: (2024)
von: Mohamud, Safaa Abdullahi Moallim, et al.
Veröffentlicht: (2024)
Meta-Adaptive Prompt Distillation for Few-Shot Visual Question Answering
von: Gupta, Akash, et al.
Veröffentlicht: (2025)
von: Gupta, Akash, et al.
Veröffentlicht: (2025)
Multi-object event graph representation learning for Video Question Answering
von: Wang, Yanan, et al.
Veröffentlicht: (2024)
von: Wang, Yanan, et al.
Veröffentlicht: (2024)
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering
von: Kugo, Noriyuki, et al.
Veröffentlicht: (2025)
von: Kugo, Noriyuki, et al.
Veröffentlicht: (2025)
LLM-grounded Video Diffusion Models
von: Lian, Long, et al.
Veröffentlicht: (2023)
von: Lian, Long, et al.
Veröffentlicht: (2023)
End-to-End Video Question Answering with Frame Scoring Mechanisms and Adaptive Sampling
von: Liang, Jianxin, et al.
Veröffentlicht: (2024)
von: Liang, Jianxin, et al.
Veröffentlicht: (2024)
STAIR: Spatial-Temporal Reasoning with Auditable Intermediate Results for Video Question Answering
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
Empowering Large Language Model for Continual Video Question Answering with Collaborative Prompting
von: Cai, Chen, et al.
Veröffentlicht: (2024)
von: Cai, Chen, et al.
Veröffentlicht: (2024)
LLMs Meet Long Video: Advancing Long Video Question Answering with An Interactive Visual Adapter in LLMs
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
Selectively Answering Visual Questions
von: Eisenschlos, Julian Martin, et al.
Veröffentlicht: (2024)
von: Eisenschlos, Julian Martin, et al.
Veröffentlicht: (2024)
AutoPresent: Designing Structured Visuals from Scratch
von: Ge, Jiaxin, et al.
Veröffentlicht: (2025)
von: Ge, Jiaxin, et al.
Veröffentlicht: (2025)
EVQAScore: A Fine-grained Metric for Video Question Answering Data Quality Evaluation
von: Liang, Hao, et al.
Veröffentlicht: (2024)
von: Liang, Hao, et al.
Veröffentlicht: (2024)
Top-down Activity Representation Learning for Video Question Answering
von: Wang, Yanan, et al.
Veröffentlicht: (2024)
von: Wang, Yanan, et al.
Veröffentlicht: (2024)
Vision-Language Models Create Cross-Modal Task Representations
von: Luo, Grace, et al.
Veröffentlicht: (2024)
von: Luo, Grace, et al.
Veröffentlicht: (2024)
VLMT: Vision-Language Multimodal Transformer for Multimodal Multi-hop Question Answering
von: Lim, Qi Zhi, et al.
Veröffentlicht: (2025)
von: Lim, Qi Zhi, et al.
Veröffentlicht: (2025)
II-MMR: Identifying and Improving Multi-modal Multi-hop Reasoning in Visual Question Answering
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
ChainReaction: Causal Chain-Guided Reasoning for Modular and Explainable Causal-Why Video Question Answering
von: Parmar, Paritosh, et al.
Veröffentlicht: (2025)
von: Parmar, Paritosh, et al.
Veröffentlicht: (2025)
Relevance-aware Multi-context Contrastive Decoding for Retrieval-augmented Visual Question Answering
von: Kim, Jongha, et al.
Veröffentlicht: (2026)
von: Kim, Jongha, et al.
Veröffentlicht: (2026)
ViQAgent: Zero-Shot Video Question Answering via Agent with Open-Vocabulary Grounding Validation
von: Montes, Tony, et al.
Veröffentlicht: (2025)
von: Montes, Tony, et al.
Veröffentlicht: (2025)
VDMA: Video Question Answering with Dynamically Generated Multi-Agents
von: Kugo, Noriyuki, et al.
Veröffentlicht: (2024)
von: Kugo, Noriyuki, et al.
Veröffentlicht: (2024)
A Gaze-grounded Visual Question Answering Dataset for Clarifying Ambiguous Japanese Questions
von: Inadumi, Shun, et al.
Veröffentlicht: (2024)
von: Inadumi, Shun, et al.
Veröffentlicht: (2024)
A Two-Stage Multitask Vision-Language Framework for Explainable Crop Disease Visual Question Answering
von: Hossain, Md. Zahid, et al.
Veröffentlicht: (2026)
von: Hossain, Md. Zahid, et al.
Veröffentlicht: (2026)
Overview of the NLPCC 2025 Shared Task 4: Multi-modal, Multilingual, and Multi-hop Medical Instructional Video Question Answering Challenge
von: Li, Bin, et al.
Veröffentlicht: (2025)
von: Li, Bin, et al.
Veröffentlicht: (2025)
In-Context Learning Enables Robot Action Prediction in LLMs
von: Yin, Yida, et al.
Veröffentlicht: (2024)
von: Yin, Yida, et al.
Veröffentlicht: (2024)
Joint Extraction Matters: Prompt-Based Visual Question Answering for Multi-Field Document Information Extraction
von: Loem, Mengsay, et al.
Veröffentlicht: (2025)
von: Loem, Mengsay, et al.
Veröffentlicht: (2025)
GeReA: Question-Aware Prompt Captions for Knowledge-based Visual Question Answering
von: Ma, Ziyu, et al.
Veröffentlicht: (2024)
von: Ma, Ziyu, et al.
Veröffentlicht: (2024)
LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning
von: Niu, Dantong, et al.
Veröffentlicht: (2024)
von: Niu, Dantong, et al.
Veröffentlicht: (2024)
GC-KBVQA: A New Four-Stage Framework for Enhancing Knowledge Based Visual Question Answering Performance
von: Moradi, Mohammad Mahdi, et al.
Veröffentlicht: (2025)
von: Moradi, Mohammad Mahdi, et al.
Veröffentlicht: (2025)
OmAgent: A Multi-modal Agent Framework for Complex Video Understanding with Task Divide-and-Conquer
von: Zhang, Lu, et al.
Veröffentlicht: (2024)
von: Zhang, Lu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Recursive Visual Programming
von: Ge, Jiaxin, et al.
Veröffentlicht: (2023) -
Latent Implicit Visual Reasoning
von: Li, Kelvin, et al.
Veröffentlicht: (2025) -
Compositional Chain-of-Thought Prompting for Large Multimodal Models
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023) -
Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning
von: Huang, Brandon, et al.
Veröffentlicht: (2024) -
DAVE: A VLM Vision Encoder for Document Understanding and Web Agents
von: Huang, Brandon, et al.
Veröffentlicht: (2025)