PresentAgent-2: Towards Generalist Multimodal Presentation Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Wei, Xu, Ziyang, Zhang, Zeyu, Zhao, Yang, Tang, Hao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PresentAgent: Multimodal Agent for Presentation Video Generation
by: Shi, Jingwei, et al.
Published: (2025)
by: Shi, Jingwei, et al.
Published: (2025)
VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery
by: Ge, Jinchao, et al.
Published: (2025)
by: Ge, Jinchao, et al.
Published: (2025)
GSCo: Towards Generalizable AI in Medicine via Generalist-Specialist Collaboration
by: He, Sunan, et al.
Published: (2024)
by: He, Sunan, et al.
Published: (2024)
VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
by: Liu, Xiao, et al.
Published: (2024)
by: Liu, Xiao, et al.
Published: (2024)
MMA: Multimodal Memory Agent
by: Lu, Yihao, et al.
Published: (2026)
by: Lu, Yihao, et al.
Published: (2026)
OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
by: Wu, Zhiyong, et al.
Published: (2024)
by: Wu, Zhiyong, et al.
Published: (2024)
Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents
by: Agashe, Saaket, et al.
Published: (2025)
by: Agashe, Saaket, et al.
Published: (2025)
Summarization of Multimodal Presentations with Vision-Language Models: Study of the Effect of Modalities and Structure
by: Gigant, Théo, et al.
Published: (2025)
by: Gigant, Théo, et al.
Published: (2025)
Through the Lens of Character: Resolving Modality-Role Interference in Multimodal Role-Playing Agent
by: Tang, Yihong, et al.
Published: (2026)
by: Tang, Yihong, et al.
Published: (2026)
AutoPresent: Designing Structured Visuals from Scratch
by: Ge, Jiaxin, et al.
Published: (2025)
by: Ge, Jiaxin, et al.
Published: (2025)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
by: Zhang, Xueqiao, et al.
Published: (2025)
by: Zhang, Xueqiao, et al.
Published: (2025)
Omni-SMoLA: Boosting Generalist Multimodal Models with Soft Mixture of Low-rank Experts
by: Wu, Jialin, et al.
Published: (2023)
by: Wu, Jialin, et al.
Published: (2023)
Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning
by: LASA Team, et al.
Published: (2025)
by: LASA Team, et al.
Published: (2025)
PreMind: Multi-Agent Video Understanding for Advanced Indexing of Presentation-style Videos
by: Wei, Kangda, et al.
Published: (2025)
by: Wei, Kangda, et al.
Published: (2025)
An Embodied Generalist Agent in 3D World
by: Huang, Jiangyong, et al.
Published: (2023)
by: Huang, Jiangyong, et al.
Published: (2023)
OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
by: Kapoor, Raghav, et al.
Published: (2024)
by: Kapoor, Raghav, et al.
Published: (2024)
SurvAgent: Hierarchical CoT-Enhanced Case Banking and Dichotomy-Based Multi-Agent System for Multimodal Survival Prediction
by: Huang, Guolin, et al.
Published: (2025)
by: Huang, Guolin, et al.
Published: (2025)
EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents
by: Cheng, Zhili, et al.
Published: (2025)
by: Cheng, Zhili, et al.
Published: (2025)
Octopus v3: Technical Report for On-device Sub-billion Multimodal AI Agent
by: Chen, Wei, et al.
Published: (2024)
by: Chen, Wei, et al.
Published: (2024)
Probabilistic Concept Graph Reasoning for Multimodal Misinformation Detection
by: Yang, Ruichao, et al.
Published: (2026)
by: Yang, Ruichao, et al.
Published: (2026)
Cooperative Sentiment Agents for Multimodal Sentiment Analysis
by: Wang, Shanmin, et al.
Published: (2024)
by: Wang, Shanmin, et al.
Published: (2024)
Modality-Specialized Synergizers for Interleaved Vision-Language Generalists
by: Xu, Zhiyang, et al.
Published: (2024)
by: Xu, Zhiyang, et al.
Published: (2024)
Towards Rationality in Language and Multimodal Agents: A Survey
by: Jiang, Bowen, et al.
Published: (2024)
by: Jiang, Bowen, et al.
Published: (2024)
MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation
by: Li, Yan, et al.
Published: (2026)
by: Li, Yan, et al.
Published: (2026)
Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration
by: Wang, Junyang, et al.
Published: (2024)
by: Wang, Junyang, et al.
Published: (2024)
OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks
by: Hu, Wenbo, et al.
Published: (2026)
by: Hu, Wenbo, et al.
Published: (2026)
OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction
by: Zhang, Haonan, et al.
Published: (2025)
by: Zhang, Haonan, et al.
Published: (2025)
Large Multimodal Agents: A Survey
by: Xie, Junlin, et al.
Published: (2024)
by: Xie, Junlin, et al.
Published: (2024)
Auditing Gender Presentation Differences in Text-to-Image Models
by: Zhang, Yanzhe, et al.
Published: (2023)
by: Zhang, Yanzhe, et al.
Published: (2023)
MEXA: Towards General Multimodal Reasoning with Dynamic Multi-Expert Aggregation
by: Yu, Shoubin, et al.
Published: (2025)
by: Yu, Shoubin, et al.
Published: (2025)
OmAgent: A Multi-modal Agent Framework for Complex Video Understanding with Task Divide-and-Conquer
by: Zhang, Lu, et al.
Published: (2024)
by: Zhang, Lu, et al.
Published: (2024)
GPT-4V(ision) is a Generalist Web Agent, if Grounded
by: Zheng, Boyuan, et al.
Published: (2024)
by: Zheng, Boyuan, et al.
Published: (2024)
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
by: Wang, Junyang, et al.
Published: (2024)
by: Wang, Junyang, et al.
Published: (2024)
DeepSight: Bridging Depth Maps and Language with a Depth-Driven Multimodal Model
by: Yang, Hao, et al.
Published: (2026)
by: Yang, Hao, et al.
Published: (2026)
GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
by: Luo, Run, et al.
Published: (2025)
by: Luo, Run, et al.
Published: (2025)
DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)
by: Yang, Zongxin, et al.
Published: (2024)
by: Yang, Zongxin, et al.
Published: (2024)
MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
by: Li, Shilong, et al.
Published: (2025)
by: Li, Shilong, et al.
Published: (2025)
MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning
by: Xia, Peng, et al.
Published: (2025)
by: Xia, Peng, et al.
Published: (2025)
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction
by: Liu, Chengzhi, et al.
Published: (2026)
by: Liu, Chengzhi, et al.
Published: (2026)
Similar Items
-
PresentAgent: Multimodal Agent for Presentation Video Generation
by: Shi, Jingwei, et al.
Published: (2025) -
VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery
by: Ge, Jinchao, et al.
Published: (2025) -
GSCo: Towards Generalizable AI in Medicine via Generalist-Specialist Collaboration
by: He, Sunan, et al.
Published: (2024) -
VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
by: Liu, Xiao, et al.
Published: (2024) -
MMA: Multimodal Memory Agent
by: Lu, Yihao, et al.
Published: (2026)