Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers
Fuente:
arXiv
Saved in:
| Main Authors: | Pang, Wei, Lin, Kevin Qinghong, Jian, Xiangru, He, Xi, Torr, Philip |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Paper2Video: Automatic Video Generation from Scientific Papers
by: Zhu, Zeyu, et al.
Published: (2025)
by: Zhu, Zeyu, et al.
Published: (2025)
PosterSum: A Multimodal Benchmark for Scientific Poster Summarization
by: Saxena, Rohit, et al.
Published: (2025)
by: Saxena, Rohit, et al.
Published: (2025)
Towards Rationality in Language and Multimodal Agents: A Survey
by: Jiang, Bowen, et al.
Published: (2024)
by: Jiang, Bowen, et al.
Published: (2024)
Many Heads Are Better Than One: Improved Scientific Idea Generation by A LLM-Based Multi-Agent System
by: Su, Haoyang, et al.
Published: (2024)
by: Su, Haoyang, et al.
Published: (2024)
Enhancing Agentic Autonomous Scientific Discovery with Vision-Language Model Capabilities
by: Gandhi, Kahaan, et al.
Published: (2025)
by: Gandhi, Kahaan, et al.
Published: (2025)
Automated Vehicles Should be Connected with Natural Language
by: Gao, Xiangbo, et al.
Published: (2025)
by: Gao, Xiangbo, et al.
Published: (2025)
Judge Model for Large-scale Multimodality Benchmarks
by: Shih, Min-Han, et al.
Published: (2026)
by: Shih, Min-Han, et al.
Published: (2026)
PreMind: Multi-Agent Video Understanding for Advanced Indexing of Presentation-style Videos
by: Wei, Kangda, et al.
Published: (2025)
by: Wei, Kangda, et al.
Published: (2025)
RadAgents: Multimodal Agentic Reasoning for Chest X-ray Interpretation with Radiologist-like Workflows
by: Zhang, Kai, et al.
Published: (2025)
by: Zhang, Kai, et al.
Published: (2025)
Chain of Questions: Guiding Multimodal Curiosity in Language Models
by: Iji, Nima, et al.
Published: (2025)
by: Iji, Nima, et al.
Published: (2025)
SWIRL: A Staged Workflow for Interleaved Reinforcement Learning in Mobile GUI Control
by: Lu, Quanfeng, et al.
Published: (2025)
by: Lu, Quanfeng, et al.
Published: (2025)
Curriculum Guided Massive Multi Agent System Solving For Robust Long Horizon Tasks
by: Kar, Indrajit, et al.
Published: (2025)
by: Kar, Indrajit, et al.
Published: (2025)
MARIC: Multi-Agent Reasoning for Image Classification
by: Seo, Wonduk, et al.
Published: (2025)
by: Seo, Wonduk, et al.
Published: (2025)
EH-Benchmark Ophthalmic Hallucination Benchmark and Agent-Driven Top-Down Traceable Reasoning Workflow
by: Pan, Xiaoyu, et al.
Published: (2025)
by: Pan, Xiaoyu, et al.
Published: (2025)
An Agentic System for Rare Disease Diagnosis with Traceable Reasoning
by: Zhao, Weike, et al.
Published: (2025)
by: Zhao, Weike, et al.
Published: (2025)
AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning
by: Song, Mingyang, et al.
Published: (2026)
by: Song, Mingyang, et al.
Published: (2026)
TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in Minecraft
by: Long, Qian, et al.
Published: (2024)
by: Long, Qian, et al.
Published: (2024)
Metropolis-Hastings Captioning Game: Knowledge Fusion of Vision Language Models via Decentralized Bayesian Inference
by: Matsui, Yuta, et al.
Published: (2025)
by: Matsui, Yuta, et al.
Published: (2025)
REVECA: Adaptive Planning and Trajectory-based Validation in Cooperative Language Agents using Information Relevance and Relative Proximity
by: Seo, SeungWon, et al.
Published: (2024)
by: Seo, SeungWon, et al.
Published: (2024)
Transductive Visual Programming: Evolving Tool Libraries from Experience for Spatial Reasoning
by: Wu, Shengguang, et al.
Published: (2025)
by: Wu, Shengguang, et al.
Published: (2025)
PathFinder: A Multi-Modal Multi-Agent System for Medical Diagnostic Decision-Making Applied to Histopathology
by: Ghezloo, Fatemeh, et al.
Published: (2025)
by: Ghezloo, Fatemeh, et al.
Published: (2025)
PASK: Toward Intent-Aware Proactive Agents with Long-Term Memory
by: Xie, Zhifei, et al.
Published: (2026)
by: Xie, Zhifei, et al.
Published: (2026)
OralAgent: Integrating Reasoning, Tools, and Knowledge for Interactive Dental Image Analysis
by: Hao, Jing, et al.
Published: (2026)
by: Hao, Jing, et al.
Published: (2026)
RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
by: Chen, Tianxing, et al.
Published: (2025)
by: Chen, Tianxing, et al.
Published: (2025)
SynWorld: Virtual Scenario Synthesis for Agentic Action Knowledge Refinement
by: Fang, Runnan, et al.
Published: (2025)
by: Fang, Runnan, et al.
Published: (2025)
Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast
by: Gu, Xiangming, et al.
Published: (2024)
by: Gu, Xiangming, et al.
Published: (2024)
Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning
by: Qin, Yulei, et al.
Published: (2025)
by: Qin, Yulei, et al.
Published: (2025)
AutoGen Driven Multi Agent Framework for Iterative Crime Data Analysis and Prediction
by: Fatima, Syeda Kisaa, et al.
Published: (2025)
by: Fatima, Syeda Kisaa, et al.
Published: (2025)
SmartSnap: Proactive Evidence Seeking for Self-Verifying Agents
by: Cai, Shaofei, et al.
Published: (2025)
by: Cai, Shaofei, et al.
Published: (2025)
SkillNet: Create, Evaluate, and Connect AI Skills
by: Liang, Yuan, et al.
Published: (2026)
by: Liang, Yuan, et al.
Published: (2026)
InnoGym: Benchmarking the Innovation Potential of AI Agents
by: Zhang, Jintian, et al.
Published: (2025)
by: Zhang, Jintian, et al.
Published: (2025)
BattleAgent: Multi-modal Dynamic Emulation on Historical Battles to Complement Historical Analysis
by: Lin, Shuhang, et al.
Published: (2024)
by: Lin, Shuhang, et al.
Published: (2024)
Analyzing Persona Effects in Generated Explanations from Multimodal LLM Agents in Urban Perception
by: da Silva, Neemias, et al.
Published: (2026)
by: da Silva, Neemias, et al.
Published: (2026)
VLM-Guided Iterative Refinement for Surgical Image Segmentation with Foundation Models
by: Lou, Ange, et al.
Published: (2026)
by: Lou, Ange, et al.
Published: (2026)
SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers
by: Pramanick, Shraman, et al.
Published: (2024)
by: Pramanick, Shraman, et al.
Published: (2024)
Beyond Single Plots: A Benchmark for Question Answering on Multi-Charts
by: Efat, Azher Ahmed, et al.
Published: (2026)
by: Efat, Azher Ahmed, et al.
Published: (2026)
KnowRL: Exploring Knowledgeable Reinforcement Learning for Factuality
by: Ren, Baochang, et al.
Published: (2025)
by: Ren, Baochang, et al.
Published: (2025)
LightMem: Lightweight and Efficient Memory-Augmented Generation
by: Fang, Jizhan, et al.
Published: (2025)
by: Fang, Jizhan, et al.
Published: (2025)
Multi-Agent VQA: Exploring Multi-Agent Foundation Models in Zero-Shot Visual Question Answering
by: Jiang, Bowen, et al.
Published: (2024)
by: Jiang, Bowen, et al.
Published: (2024)
Exploring Model Kinship for Merging Large Language Models
by: Hu, Yedi, et al.
Published: (2024)
by: Hu, Yedi, et al.
Published: (2024)
Similar Items
-
Paper2Video: Automatic Video Generation from Scientific Papers
by: Zhu, Zeyu, et al.
Published: (2025) -
PosterSum: A Multimodal Benchmark for Scientific Poster Summarization
by: Saxena, Rohit, et al.
Published: (2025) -
Towards Rationality in Language and Multimodal Agents: A Survey
by: Jiang, Bowen, et al.
Published: (2024) -
Many Heads Are Better Than One: Improved Scientific Idea Generation by A LLM-Based Multi-Agent System
by: Su, Haoyang, et al.
Published: (2024) -
Enhancing Agentic Autonomous Scientific Discovery with Vision-Language Model Capabilities
by: Gandhi, Kahaan, et al.
Published: (2025)