OmAgent: A Multi-modal Agent Framework for Complex Video Understanding with Task Divide-and-Conquer
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Lu, Zhao, Tiancheng, Ying, Heting, Ma, Yibo, Lee, Kyusong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OmDet: Large-scale vision-language multi-dataset pre-training with multimodal detection network
von: Zhao, Tiancheng, et al.
Veröffentlicht: (2022)
von: Zhao, Tiancheng, et al.
Veröffentlicht: (2022)
OmChat: A Recipe to Train Multimodal Language Models with Strong Long Context and Video Understanding
von: Zhao, Tiancheng, et al.
Veröffentlicht: (2024)
von: Zhao, Tiancheng, et al.
Veröffentlicht: (2024)
Real-time Transformer-based Open-Vocabulary Detection with Efficient Fusion Head
von: Zhao, Tiancheng, et al.
Veröffentlicht: (2024)
von: Zhao, Tiancheng, et al.
Veröffentlicht: (2024)
Unifying Language Agent Algorithms with Graph-based Orchestration Engine for Reproducible Agent Research
von: Zhang, Qianqian, et al.
Veröffentlicht: (2025)
von: Zhang, Qianqian, et al.
Veröffentlicht: (2025)
Preserving Knowledge in Large Language Model with Model-Agnostic Self-Decompression
von: Zhang, Zilun, et al.
Veröffentlicht: (2024)
von: Zhang, Zilun, et al.
Veröffentlicht: (2024)
Sandboxed Coding Agents are Competitive Omni-modal Task Solvers
von: Chen, Dongping, et al.
Veröffentlicht: (2026)
von: Chen, Dongping, et al.
Veröffentlicht: (2026)
MMCTAgent: Multi-modal Critical Thinking Agent Framework for Complex Visual Reasoning
von: Kumar, Somnath, et al.
Veröffentlicht: (2024)
von: Kumar, Somnath, et al.
Veröffentlicht: (2024)
DCARL: A Divide-and-Conquer Framework for Autoregressive Long-Trajectory Video Generation
von: Ouyang, Junyi, et al.
Veröffentlicht: (2026)
von: Ouyang, Junyi, et al.
Veröffentlicht: (2026)
MMOU: A Massive Multi-Task Omni Understanding and Reasoning Benchmark for Long and Complex Real-World Videos
von: Goel, Arushi, et al.
Veröffentlicht: (2026)
von: Goel, Arushi, et al.
Veröffentlicht: (2026)
DCDM: Divide-and-Conquer Diffusion Models for Consistency-Preserving Video Generation
von: Zhao, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhao, Haoyu, et al.
Veröffentlicht: (2026)
Mobile-Agent-E: Self-Evolving Mobile Assistant for Complex Tasks
von: Wang, Zhenhailong, et al.
Veröffentlicht: (2025)
von: Wang, Zhenhailong, et al.
Veröffentlicht: (2025)
Multi-modal Semantic Understanding with Contrastive Cross-modal Feature Alignment
von: Zhang, Ming, et al.
Veröffentlicht: (2024)
von: Zhang, Ming, et al.
Veröffentlicht: (2024)
DocLens : A Tool-Augmented Multi-Agent Framework for Long Visual Document Understanding
von: Zhu, Dawei, et al.
Veröffentlicht: (2025)
von: Zhu, Dawei, et al.
Veröffentlicht: (2025)
GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual Agents
von: Wang, Yunzhe, et al.
Veröffentlicht: (2026)
von: Wang, Yunzhe, et al.
Veröffentlicht: (2026)
VideoAgent: Long-form Video Understanding with Large Language Model as Agent
von: Wang, Xiaohan, et al.
Veröffentlicht: (2024)
von: Wang, Xiaohan, et al.
Veröffentlicht: (2024)
PreMind: Multi-Agent Video Understanding for Advanced Indexing of Presentation-style Videos
von: Wei, Kangda, et al.
Veröffentlicht: (2025)
von: Wei, Kangda, et al.
Veröffentlicht: (2025)
MicroCinema: A Divide-and-Conquer Approach for Text-to-Video Generation
von: Wang, Yanhui, et al.
Veröffentlicht: (2023)
von: Wang, Yanhui, et al.
Veröffentlicht: (2023)
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents
von: Wang, Xuehui, et al.
Veröffentlicht: (2025)
von: Wang, Xuehui, et al.
Veröffentlicht: (2025)
Video-Based Reward Modeling for Computer-Use Agents
von: Song, Linxin, et al.
Veröffentlicht: (2026)
von: Song, Linxin, et al.
Veröffentlicht: (2026)
DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)
von: Yang, Zongxin, et al.
Veröffentlicht: (2024)
von: Yang, Zongxin, et al.
Veröffentlicht: (2024)
RoboSVG: A Unified Framework for Interactive SVG Generation with Multi-modal Guidance
von: Wang, Jiuniu, et al.
Veröffentlicht: (2025)
von: Wang, Jiuniu, et al.
Veröffentlicht: (2025)
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
von: Yang, Rui, et al.
Veröffentlicht: (2025)
von: Yang, Rui, et al.
Veröffentlicht: (2025)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
von: Zhang, Xueqiao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueqiao, et al.
Veröffentlicht: (2025)
MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning
von: Xia, Peng, et al.
Veröffentlicht: (2025)
von: Xia, Peng, et al.
Veröffentlicht: (2025)
MuMA-ToM: Multi-modal Multi-Agent Theory of Mind
von: Shi, Haojun, et al.
Veröffentlicht: (2024)
von: Shi, Haojun, et al.
Veröffentlicht: (2024)
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
von: Wang, Junyang, et al.
Veröffentlicht: (2024)
von: Wang, Junyang, et al.
Veröffentlicht: (2024)
Overview of the NLPCC 2025 Shared Task 4: Multi-modal, Multilingual, and Multi-hop Medical Instructional Video Question Answering Challenge
von: Li, Bin, et al.
Veröffentlicht: (2025)
von: Li, Bin, et al.
Veröffentlicht: (2025)
ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding
von: Zhou, Yiyang, et al.
Veröffentlicht: (2025)
von: Zhou, Yiyang, et al.
Veröffentlicht: (2025)
Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation
von: Zhou, Yucheng, et al.
Veröffentlicht: (2025)
von: Zhou, Yucheng, et al.
Veröffentlicht: (2025)
VideoXum: Cross-modal Visual and Textural Summarization of Videos
von: Lin, Jingyang, et al.
Veröffentlicht: (2023)
von: Lin, Jingyang, et al.
Veröffentlicht: (2023)
DetectiumFire: A Comprehensive Multi-modal Dataset Bridging Vision and Language for Fire Understanding
von: Liu, Zixuan, et al.
Veröffentlicht: (2025)
von: Liu, Zixuan, et al.
Veröffentlicht: (2025)
m&m's: A Benchmark to Evaluate Tool-Use for multi-step multi-modal Tasks
von: Ma, Zixian, et al.
Veröffentlicht: (2024)
von: Ma, Zixian, et al.
Veröffentlicht: (2024)
EVA: Efficient Reinforcement Learning for End-to-End Video Agent
von: Zhang, Yaolun, et al.
Veröffentlicht: (2026)
von: Zhang, Yaolun, et al.
Veröffentlicht: (2026)
ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
von: Ma, David, et al.
Veröffentlicht: (2025)
von: Ma, David, et al.
Veröffentlicht: (2025)
Divide and Conquer: Reliable Multi-View Evidential Learning for Deepfake Detection
von: Kang, Xiaolu, et al.
Veröffentlicht: (2026)
von: Kang, Xiaolu, et al.
Veröffentlicht: (2026)
PresentAgent-2: Towards Generalist Multimodal Presentation Agents
von: Wu, Wei, et al.
Veröffentlicht: (2026)
von: Wu, Wei, et al.
Veröffentlicht: (2026)
PC-Agent: A Hierarchical Multi-Agent Collaboration Framework for Complex Task Automation on PC
von: Liu, Haowei, et al.
Veröffentlicht: (2025)
von: Liu, Haowei, et al.
Veröffentlicht: (2025)
MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos
von: Zang, Yuan, et al.
Veröffentlicht: (2025)
von: Zang, Yuan, et al.
Veröffentlicht: (2025)
Large Multi-modal Models Can Interpret Features in Large Multi-modal Models
von: Zhang, Kaichen, et al.
Veröffentlicht: (2024)
von: Zhang, Kaichen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
OmDet: Large-scale vision-language multi-dataset pre-training with multimodal detection network
von: Zhao, Tiancheng, et al.
Veröffentlicht: (2022) -
OmChat: A Recipe to Train Multimodal Language Models with Strong Long Context and Video Understanding
von: Zhao, Tiancheng, et al.
Veröffentlicht: (2024) -
Real-time Transformer-based Open-Vocabulary Detection with Efficient Fusion Head
von: Zhao, Tiancheng, et al.
Veröffentlicht: (2024) -
Unifying Language Agent Algorithms with Graph-based Orchestration Engine for Reproducible Agent Research
von: Zhang, Qianqian, et al.
Veröffentlicht: (2025) -
Preserving Knowledge in Large Language Model with Model-Agnostic Self-Decompression
von: Zhang, Zilun, et al.
Veröffentlicht: (2024)