DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yang, Zongxin, Chen, Guikun, Li, Xiaodi, Wang, Wenguan, Yang, Yi |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Scene Graph Generation with Role-Playing Large Language Models
par: Chen, Guikun, et autres
Publié: (2024)
par: Chen, Guikun, et autres
Publié: (2024)
Hydra-SGG: Hybrid Relation Assignment for One-stage Scene Graph Generation
par: Chen, Minghan, et autres
Publié: (2024)
par: Chen, Minghan, et autres
Publié: (2024)
SinkTrack: Attention Sink based Context Anchoring for Large Language Models
par: Liu, Xu, et autres
Publié: (2026)
par: Liu, Xu, et autres
Publié: (2026)
Neural Clustering based Visual Representation Learning
par: Chen, Guikun, et autres
Publié: (2024)
par: Chen, Guikun, et autres
Publié: (2024)
A Survey on 3D Gaussian Splatting
par: Chen, Guikun, et autres
Publié: (2024)
par: Chen, Guikun, et autres
Publié: (2024)
Scalable Video Object Segmentation with Identification Mechanism
par: Yang, Zongxin, et autres
Publié: (2022)
par: Yang, Zongxin, et autres
Publié: (2022)
DIFFVSGG: Diffusion-Driven Online Video Scene Graph Generation
par: Chen, Mu, et autres
Publié: (2025)
par: Chen, Mu, et autres
Publié: (2025)
Rethinking Cross-modal Interaction from a Top-down Perspective for Referring Video Object Segmentation
par: Liang, Chen, et autres
Publié: (2021)
par: Liang, Chen, et autres
Publié: (2021)
TRACE: High-Fidelity 3D Scene Editing via Tangible Reconstruction and Geometry-Aligned Contextual Video Masking
par: Hu, Jiyuan, et autres
Publié: (2026)
par: Hu, Jiyuan, et autres
Publié: (2026)
DivScene: Towards Open-Vocabulary Object Navigation with Large Vision Language Models in Diverse Scenes
par: Wang, Zhaowei, et autres
Publié: (2024)
par: Wang, Zhaowei, et autres
Publié: (2024)
Video Understanding with Large Language Models: A Survey
par: Tang, Yolo Y., et autres
Publié: (2023)
par: Tang, Yolo Y., et autres
Publié: (2023)
Are Image-to-Video Models Good Zero-Shot Image Editors?
par: Zhang, Zechuan, et autres
Publié: (2025)
par: Zhang, Zechuan, et autres
Publié: (2025)
Zero-Shot Scene Understanding with Multimodal Large Language Models for Automated Vehicles
par: Elhenawy, Mohammed, et autres
Publié: (2025)
par: Elhenawy, Mohammed, et autres
Publié: (2025)
TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment
par: Li, Wei, et autres
Publié: (2024)
par: Li, Wei, et autres
Publié: (2024)
Navigation Instruction Generation with BEV Perception and Large Language Models
par: Fan, Sheng, et autres
Publié: (2024)
par: Fan, Sheng, et autres
Publié: (2024)
VideoAgent: Long-form Video Understanding with Large Language Model as Agent
par: Wang, Xiaohan, et autres
Publié: (2024)
par: Wang, Xiaohan, et autres
Publié: (2024)
Toward Interactive Regional Understanding in Vision-Large Language Models
par: Lee, Jungbeom, et autres
Publié: (2024)
par: Lee, Jungbeom, et autres
Publié: (2024)
GalleryGPT: Analyzing Paintings with Large Multimodal Models
par: Bin, Yi, et autres
Publié: (2024)
par: Bin, Yi, et autres
Publié: (2024)
Compositional Feature Augmentation for Unbiased Scene Graph Generation
par: Li, Lin, et autres
Publié: (2023)
par: Li, Lin, et autres
Publié: (2023)
Realizing Video Summarization from the Path of Language-based Semantic Understanding
par: Mu, Kuan-Chen, et autres
Publié: (2024)
par: Mu, Kuan-Chen, et autres
Publié: (2024)
ArtGPT-4: Towards Artistic-understanding Large Vision-Language Models with Enhanced Adapter
par: Yuan, Zhengqing, et autres
Publié: (2023)
par: Yuan, Zhengqing, et autres
Publié: (2023)
Human-Object Interaction Detection Collaborated with Large Relation-driven Diffusion Models
par: Li, Liulei, et autres
Publié: (2024)
par: Li, Liulei, et autres
Publié: (2024)
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding
par: Li, Yun, et autres
Publié: (2025)
par: Li, Yun, et autres
Publié: (2025)
Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning
par: Zhu, Yingjie, et autres
Publié: (2024)
par: Zhu, Yingjie, et autres
Publié: (2024)
Multi-Agent Visual-Language Reasoning for Comprehensive Highway Scene Understanding
par: Yang, Yunxiang, et autres
Publié: (2025)
par: Yang, Yunxiang, et autres
Publié: (2025)
Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering
par: Beliaev, Mark, et autres
Publié: (2025)
par: Beliaev, Mark, et autres
Publié: (2025)
IdealGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language Models
par: You, Haoxuan, et autres
Publié: (2023)
par: You, Haoxuan, et autres
Publié: (2023)
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
par: Maaz, Muhammad, et autres
Publié: (2023)
par: Maaz, Muhammad, et autres
Publié: (2023)
Decomposed Prototype Learning for Few-Shot Scene Graph Generation
par: Li, Xingchen, et autres
Publié: (2023)
par: Li, Xingchen, et autres
Publié: (2023)
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
par: Qi, Zhangyang, et autres
Publié: (2025)
par: Qi, Zhangyang, et autres
Publié: (2025)
GeoGPT4V: Towards Geometric Multi-modal Large Language Models with Geometric Image Generation
par: Cai, Shihao, et autres
Publié: (2024)
par: Cai, Shihao, et autres
Publié: (2024)
LSK3DNet: Towards Effective and Efficient 3D Perception with Large Sparse Kernels
par: Feng, Tuo, et autres
Publié: (2024)
par: Feng, Tuo, et autres
Publié: (2024)
Generalizable Entity Grounding via Assistance of Large Language Model
par: Qi, Lu, et autres
Publié: (2024)
par: Qi, Lu, et autres
Publié: (2024)
DreamRenderer: Taming Multi-Instance Attribute Control in Large-Scale Text-to-Image Models
par: Zhou, Dewei, et autres
Publié: (2025)
par: Zhou, Dewei, et autres
Publié: (2025)
Do as We Do, Not as You Think: the Conformity of Large Language Models
par: Weng, Zhiyuan, et autres
Publié: (2025)
par: Weng, Zhiyuan, et autres
Publié: (2025)
Volumetric Environment Representation for Vision-Language Navigation
par: Liu, Rui, et autres
Publié: (2024)
par: Liu, Rui, et autres
Publié: (2024)
Vision-Language Navigation with Energy-Based Policy
par: Liu, Rui, et autres
Publié: (2024)
par: Liu, Rui, et autres
Publié: (2024)
FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models
par: Cai, Hengxing, et autres
Publié: (2025)
par: Cai, Hengxing, et autres
Publié: (2025)
SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language Models
par: Xia, Haotian, et autres
Publié: (2024)
par: Xia, Haotian, et autres
Publié: (2024)
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
par: Zheng, Duo, et autres
Publié: (2024)
par: Zheng, Duo, et autres
Publié: (2024)
Documents similaires
-
Scene Graph Generation with Role-Playing Large Language Models
par: Chen, Guikun, et autres
Publié: (2024) -
Hydra-SGG: Hybrid Relation Assignment for One-stage Scene Graph Generation
par: Chen, Minghan, et autres
Publié: (2024) -
SinkTrack: Attention Sink based Context Anchoring for Large Language Models
par: Liu, Xu, et autres
Publié: (2026) -
Neural Clustering based Visual Representation Learning
par: Chen, Guikun, et autres
Publié: (2024) -
A Survey on 3D Gaussian Splatting
par: Chen, Guikun, et autres
Publié: (2024)