From Pixels to Paths: A Multi-Agent Framework for Editable Scientific Illustration
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Jianwen, Zhang, Fanrui, Feng, Yukang, Li, Chuanhao, Li, Zizhen, Ai, Jiaxin, Chang, Yifan, Dai, Yu, Zhang, Kaipeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model
by: Chang, Yifan, et al.
Published: (2025)
by: Chang, Yifan, et al.
Published: (2025)
Closing the Expression Gap in LLM Instructions via Socratic Questioning
by: Sun, Jianwen, et al.
Published: (2025)
by: Sun, Jianwen, et al.
Published: (2025)
World Craft: Agentic Framework to Create Visualizable Worlds via Text
by: Sun, Jianwen, et al.
Published: (2026)
by: Sun, Jianwen, et al.
Published: (2026)
A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
by: Feng, Yukang, et al.
Published: (2025)
by: Feng, Yukang, et al.
Published: (2025)
ProSoftArena: Benchmarking Hierarchical Capabilities of Multimodal Agents in Professional Software Environments
by: Ai, Jiaxin, et al.
Published: (2025)
by: Ai, Jiaxin, et al.
Published: (2025)
ARMOR: Empowering Multimodal Understanding Model with Interleaved Multimodal Generation Capability
by: Sun, Jianwen, et al.
Published: (2025)
by: Sun, Jianwen, et al.
Published: (2025)
MeepleLM: A Virtual Playtester Simulating Diverse Subjective Experiences
by: Li, Zizhen, et al.
Published: (2026)
by: Li, Zizhen, et al.
Published: (2026)
InMind: Evaluating LLMs in Capturing and Applying Individual Human Reasoning Styles
by: Li, Zizhen, et al.
Published: (2025)
by: Li, Zizhen, et al.
Published: (2025)
AutoBG: A Board Game Design Assistant with Interactive Ideation, Iterative Rulebook Generation, and Individualized Feedback
by: Li, Zizhen, et al.
Published: (2026)
by: Li, Zizhen, et al.
Published: (2026)
Code-in-the-Loop Forensics: Agentic Tool Use for Image Forgery Detection
by: Zhang, Fanrui, et al.
Published: (2025)
by: Zhang, Fanrui, et al.
Published: (2025)
IA-T2I: Internet-Augmented Text-to-Image Generation
by: Li, Chuanhao, et al.
Published: (2025)
by: Li, Chuanhao, et al.
Published: (2025)
ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for MLLM-based Process Judges
by: Ai, Jiaxin, et al.
Published: (2025)
by: Ai, Jiaxin, et al.
Published: (2025)
MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language Models
by: Zhou, Pengfei, et al.
Published: (2025)
by: Zhou, Pengfei, et al.
Published: (2025)
LongCLI-Bench: A Preliminary Benchmark and Study for Long-horizon Agentic Programming in Command-Line Interfaces
by: Feng, Yukang, et al.
Published: (2026)
by: Feng, Yukang, et al.
Published: (2026)
LiveFigure: Generating Editable Scientific Illustration with VLM Agents
by: Shao, Chenyang, et al.
Published: (2026)
by: Shao, Chenyang, et al.
Published: (2026)
AutoFigure-Edit: Generating Editable Scientific Illustration
by: Lin, Zhen, et al.
Published: (2026)
by: Lin, Zhen, et al.
Published: (2026)
Sekai: A Video Dataset towards World Exploration
by: Li, Zhen, et al.
Published: (2025)
by: Li, Zhen, et al.
Published: (2025)
MDK12-Bench: A Comprehensive Evaluation of Multimodal Large Language Models on Multidisciplinary Exams
by: Zhou, Pengfei, et al.
Published: (2025)
by: Zhou, Pengfei, et al.
Published: (2025)
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs
by: Zhao, Haozhe, et al.
Published: (2026)
by: Zhao, Haozhe, et al.
Published: (2026)
LaGen: Towards Autoregressive LiDAR Scene Generation
by: Zhou, Sizhuo, et al.
Published: (2025)
by: Zhou, Sizhuo, et al.
Published: (2025)
AutoFigure: Generating and Refining Publication-Ready Scientific Illustrations
by: Zhu, Minjun, et al.
Published: (2026)
by: Zhu, Minjun, et al.
Published: (2026)
SearchLVLMs: A Plug-and-Play Framework for Augmenting Large Vision-Language Models by Searching Up-to-Date Internet Knowledge
by: Li, Chuanhao, et al.
Published: (2024)
by: Li, Chuanhao, et al.
Published: (2024)
STEAM: A Training-Free Congestion-Aware Enhancement Framework for Decentralized Multi-Agent Path Finding
by: Feng, Mingyang, et al.
Published: (2026)
by: Feng, Mingyang, et al.
Published: (2026)
COLE: A Hierarchical Generation Framework for Multi-Layered and Editable Graphic Design
by: Jia, Peidong, et al.
Published: (2023)
by: Jia, Peidong, et al.
Published: (2023)
Editable-DeepSC: Cross-Modal Editable Semantic Communication Systems
by: Yu, Wenbo, et al.
Published: (2023)
by: Yu, Wenbo, et al.
Published: (2023)
LawLuo: A Multi-Agent Collaborative Framework for Multi-Round Chinese Legal Consultation
by: Sun, Jingyun, et al.
Published: (2024)
by: Sun, Jingyun, et al.
Published: (2024)
SkeletonGaussian: Editable 4D Generation through Gaussian Skeletonization
by: Wu, Lifan, et al.
Published: (2026)
by: Wu, Lifan, et al.
Published: (2026)
ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Capability for Large Vision-Language Models
by: Liu, Shuo, et al.
Published: (2024)
by: Liu, Shuo, et al.
Published: (2024)
NLI4VolVis: Natural Language Interaction for Volume Visualization via LLM Multi-Agents and Editable 3D Gaussian Splatting
by: Ai, Kuangshi, et al.
Published: (2025)
by: Ai, Kuangshi, et al.
Published: (2025)
PackForcing: Short Video Training Suffices for Long Video Sampling and Long Context Inference
by: Mao, Xiaofeng, et al.
Published: (2026)
by: Mao, Xiaofeng, et al.
Published: (2026)
SVBench: Evaluation of Video Generation Models on Social Reasoning
by: Peng, Wenshuo, et al.
Published: (2025)
by: Peng, Wenshuo, et al.
Published: (2025)
MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
by: Meng, Fanqing, et al.
Published: (2024)
by: Meng, Fanqing, et al.
Published: (2024)
Point Resampling and Ray Transformation Aid to Editable NeRF Models
by: Li, Zhenyang, et al.
Published: (2024)
by: Li, Zhenyang, et al.
Published: (2024)
From Panel to Pixel: Zoom-In Vision-Language Pretraining from Biomedical Scientific Literature
by: Yuan, Kun, et al.
Published: (2025)
by: Yuan, Kun, et al.
Published: (2025)
WildWorld: A Large-Scale Dataset for Dynamic World Modeling with Actions and Explicit State toward Generative ARPG
by: Li, Zhen, et al.
Published: (2026)
by: Li, Zhen, et al.
Published: (2026)
SciNav: A General Agent Framework for Scientific Coding Tasks
by: Zhang, Tianshu, et al.
Published: (2026)
by: Zhang, Tianshu, et al.
Published: (2026)
Yume-1.5: A Text-Controlled Interactive World Generation Model
by: Mao, Xiaofeng, et al.
Published: (2025)
by: Mao, Xiaofeng, et al.
Published: (2025)
Editable Scene Simulation for Autonomous Driving via Collaborative LLM-Agents
by: Wei, Yuxi, et al.
Published: (2024)
by: Wei, Yuxi, et al.
Published: (2024)
Balancing Efficiency and Fairness: An Iterative Exchange Framework for Multi-UAV Cooperative Path Planning
by: Li, Hongzong, et al.
Published: (2025)
by: Li, Hongzong, et al.
Published: (2025)
From Pixels to Prose: Advancing Multi-Modal Language Models for Remote Sensing
by: Sun, Xintian, et al.
Published: (2024)
by: Sun, Xintian, et al.
Published: (2024)
Similar Items
-
SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model
by: Chang, Yifan, et al.
Published: (2025) -
Closing the Expression Gap in LLM Instructions via Socratic Questioning
by: Sun, Jianwen, et al.
Published: (2025) -
World Craft: Agentic Framework to Create Visualizable Worlds via Text
by: Sun, Jianwen, et al.
Published: (2026) -
A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
by: Feng, Yukang, et al.
Published: (2025) -
ProSoftArena: Benchmarking Hierarchical Capabilities of Multimodal Agents in Professional Software Environments
by: Ai, Jiaxin, et al.
Published: (2025)