PresentAgent: Multimodal Agent for Presentation Video Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shi, Jingwei, Zhang, Zeyu, Wu, Biao, Liang, Yanjie, Fang, Meng, Chen, Ling, Zhao, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PresentAgent-2: Towards Generalist Multimodal Presentation Agents
von: Wu, Wei, et al.
Veröffentlicht: (2026)
von: Wu, Wei, et al.
Veröffentlicht: (2026)
VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery
von: Ge, Jinchao, et al.
Veröffentlicht: (2025)
von: Ge, Jinchao, et al.
Veröffentlicht: (2025)
UniVid: The Open-Source Unified Video Model
von: Luo, Jiabin, et al.
Veröffentlicht: (2025)
von: Luo, Jiabin, et al.
Veröffentlicht: (2025)
Motion Avatar: Generate Human and Animal Avatars with Arbitrary Motion
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
MMA: Multimodal Memory Agent
von: Lu, Yihao, et al.
Veröffentlicht: (2026)
von: Lu, Yihao, et al.
Veröffentlicht: (2026)
VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding
von: Fan, Yue, et al.
Veröffentlicht: (2024)
von: Fan, Yue, et al.
Veröffentlicht: (2024)
Comp-Attn: Present-and-Align Attention for Compositional Video Generation
von: Zhang, Hongyu, et al.
Veröffentlicht: (2025)
von: Zhang, Hongyu, et al.
Veröffentlicht: (2025)
Refer-Agent: A Collaborative Multi-Agent System with Reasoning and Reflection for Referring Video Object Segmentation
von: Jiang, Haichao, et al.
Veröffentlicht: (2026)
von: Jiang, Haichao, et al.
Veröffentlicht: (2026)
VideoWeaver: Multimodal Multi-View Video-to-Video Transfer for Embodied Agents
von: Eskandar, George, et al.
Veröffentlicht: (2026)
von: Eskandar, George, et al.
Veröffentlicht: (2026)
Player-Centric Multimodal Prompt Generation for Large Language Model Based Identity-Aware Basketball Video Captioning
von: Xi, Zeyu, et al.
Veröffentlicht: (2025)
von: Xi, Zeyu, et al.
Veröffentlicht: (2025)
UniPPTBench: A Unified Benchmark for Presentation Generation Across Diverse Input Settings
von: Zhao, Bo, et al.
Veröffentlicht: (2026)
von: Zhao, Bo, et al.
Veröffentlicht: (2026)
PresentBench: A Fine-Grained Rubric-Based Benchmark for Slide Generation
von: Chen, Xin-Sheng, et al.
Veröffentlicht: (2026)
von: Chen, Xin-Sheng, et al.
Veröffentlicht: (2026)
Automated Movie Generation via Multi-Agent CoT Planning
von: Wu, Weijia, et al.
Veröffentlicht: (2025)
von: Wu, Weijia, et al.
Veröffentlicht: (2025)
Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
PreMind: Multi-Agent Video Understanding for Advanced Indexing of Presentation-style Videos
von: Wei, Kangda, et al.
Veröffentlicht: (2025)
von: Wei, Kangda, et al.
Veröffentlicht: (2025)
MTA-Agent: An Open Recipe for Multimodal Deep Search Agents
von: Peng, Xiangyu, et al.
Veröffentlicht: (2026)
von: Peng, Xiangyu, et al.
Veröffentlicht: (2026)
SurgVisAgent: Multimodal Agentic Model for Versatile Surgical Visual Enhancement
von: Lei, Zeyu, et al.
Veröffentlicht: (2025)
von: Lei, Zeyu, et al.
Veröffentlicht: (2025)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
von: Zhang, Xueqiao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueqiao, et al.
Veröffentlicht: (2025)
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding
von: Yin, Yufei, et al.
Veröffentlicht: (2026)
von: Yin, Yufei, et al.
Veröffentlicht: (2026)
EpiAgent: An Agent-Centric System for Ancient Inscription Restoration
von: Zhu, Shipeng, et al.
Veröffentlicht: (2026)
von: Zhu, Shipeng, et al.
Veröffentlicht: (2026)
SPIRAL: Self-Evolving Action-Conditioned Video Generation via Reflective Planning Agents
von: Yang, Yu, et al.
Veröffentlicht: (2026)
von: Yang, Yu, et al.
Veröffentlicht: (2026)
Lighting-grounded Video Generation with Renderer-based Agent Reasoning
von: Cai, Ziqi, et al.
Veröffentlicht: (2026)
von: Cai, Ziqi, et al.
Veröffentlicht: (2026)
UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist
von: Liang, Zhengyang, et al.
Veröffentlicht: (2025)
von: Liang, Zhengyang, et al.
Veröffentlicht: (2025)
EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation
von: Qiu, Zongyang, et al.
Veröffentlicht: (2025)
von: Qiu, Zongyang, et al.
Veröffentlicht: (2025)
Agent-based Video Trimming
von: Yang, Lingfeng, et al.
Veröffentlicht: (2024)
von: Yang, Lingfeng, et al.
Veröffentlicht: (2024)
Kubrick: Multimodal Agent Collaborations for Synthetic Video Generation
von: He, Liu, et al.
Veröffentlicht: (2024)
von: He, Liu, et al.
Veröffentlicht: (2024)
GEMS: Agent-Native Multimodal Generation with Memory and Skills
von: He, Zefeng, et al.
Veröffentlicht: (2026)
von: He, Zefeng, et al.
Veröffentlicht: (2026)
Multimodal Models Meet Presentation Attack Detection on ID Documents
von: Villanueva, Marina, et al.
Veröffentlicht: (2026)
von: Villanueva, Marina, et al.
Veröffentlicht: (2026)
StreamAgent: Towards Anticipatory Agents for Streaming Video Understanding
von: Yang, Haolin, et al.
Veröffentlicht: (2025)
von: Yang, Haolin, et al.
Veröffentlicht: (2025)
VideoGen-Eval: Agent-based System for Video Generation Evaluation
von: Yang, Yuhang, et al.
Veröffentlicht: (2025)
von: Yang, Yuhang, et al.
Veröffentlicht: (2025)
DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents
von: Wang, Shengqin, et al.
Veröffentlicht: (2026)
von: Wang, Shengqin, et al.
Veröffentlicht: (2026)
AffectAgent: Collaborative Multi-Agent Reasoning for Retrieval-Augmented Multimodal Emotion Recognition
von: Wang, Zeheng, et al.
Veröffentlicht: (2026)
von: Wang, Zeheng, et al.
Veröffentlicht: (2026)
MMCLIP: Cross-modal Attention Masked Modelling for Medical Language-Image Pre-Training
von: Wu, Biao, et al.
Veröffentlicht: (2024)
von: Wu, Biao, et al.
Veröffentlicht: (2024)
Spatio-Temporal Progressive Attention Model for EEG Classification in Rapid Serial Visual Presentation Task
von: Li, Yang, et al.
Veröffentlicht: (2025)
von: Li, Yang, et al.
Veröffentlicht: (2025)
Video-Based Reward Modeling for Computer-Use Agents
von: Song, Linxin, et al.
Veröffentlicht: (2026)
von: Song, Linxin, et al.
Veröffentlicht: (2026)
Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis
von: Chen, Shuang, et al.
Veröffentlicht: (2026)
von: Chen, Shuang, et al.
Veröffentlicht: (2026)
VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks
von: Jang, Lawrence, et al.
Veröffentlicht: (2024)
von: Jang, Lawrence, et al.
Veröffentlicht: (2024)
A Versatile Multimodal Agent for Multimedia Content Generation
von: Zhang, Daoan, et al.
Veröffentlicht: (2026)
von: Zhang, Daoan, et al.
Veröffentlicht: (2026)
Enhancing Video-LLM Reasoning via Agent-of-Thoughts Distillation
von: Shi, Yudi, et al.
Veröffentlicht: (2024)
von: Shi, Yudi, et al.
Veröffentlicht: (2024)
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
PresentAgent-2: Towards Generalist Multimodal Presentation Agents
von: Wu, Wei, et al.
Veröffentlicht: (2026) -
VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery
von: Ge, Jinchao, et al.
Veröffentlicht: (2025) -
UniVid: The Open-Source Unified Video Model
von: Luo, Jiabin, et al.
Veröffentlicht: (2025) -
Motion Avatar: Generate Human and Animal Avatars with Arbitrary Motion
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024) -
MMA: Multimodal Memory Agent
von: Lu, Yihao, et al.
Veröffentlicht: (2026)