Feynman: Knowledge-Infused Diagramming Agent for Scalable Visual Designs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wen, Zixin, Cai, Yifu, Lee, Kyle, Estep, Sam, Sunshine, Josh, Singh, Aarti, Chi, Yuejie, Ni, Wode
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912964552949760
author Wen, Zixin
Cai, Yifu
Lee, Kyle
Estep, Sam
Sunshine, Josh
Singh, Aarti
Chi, Yuejie
Ni, Wode
author_facet Wen, Zixin
Cai, Yifu
Lee, Kyle
Estep, Sam
Sunshine, Josh
Singh, Aarti
Chi, Yuejie
Ni, Wode
contents Visual design is an essential application of state-of-the-art multi-modal AI systems. Improving these systems requires high-quality vision-language data at scale. Despite the abundance of internet image and text data, knowledge-rich and well-aligned image-text pairs are rare. In this paper, we present a scalable diagram generation pipeline built with our agent, Feynman. To create diagrams, Feynman first enumerates domain-specific knowledge components (''ideas'') and performs code planning based on the ideas. Given the plan, Feynman translates ideas into simple declarative programs and iterates to receives feedback and visually refine diagrams. Finally, the declarative programs are rendered by the Penrose diagramming system. The optimization-based rendering of Penrose preserves the visual semantics while injecting fresh randomness into the layout, thereby producing diagrams with visual consistency and diversity. As a result, Feynman can author diagrams along with grounded captions with very little cost and time. Using Feynman, we synthesized a dataset with more than 100k well-aligned diagram-caption pairs. We also curate a visual-language benchmark, Diagramma, from freshly generated data. Diagramma can be used for evaluating the visual reasoning capabilities of vision-language models. We plan to release the dataset, benchmark, and the full agent pipeline as an open-source project.
format Preprint
id arxiv_https___arxiv_org_abs_2603_12597
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Feynman: Knowledge-Infused Diagramming Agent for Scalable Visual Designs
Wen, Zixin
Cai, Yifu
Lee, Kyle
Estep, Sam
Sunshine, Josh
Singh, Aarti
Chi, Yuejie
Ni, Wode
Machine Learning
Artificial Intelligence
Human-Computer Interaction
Multiagent Systems
Software Engineering
Visual design is an essential application of state-of-the-art multi-modal AI systems. Improving these systems requires high-quality vision-language data at scale. Despite the abundance of internet image and text data, knowledge-rich and well-aligned image-text pairs are rare. In this paper, we present a scalable diagram generation pipeline built with our agent, Feynman. To create diagrams, Feynman first enumerates domain-specific knowledge components (''ideas'') and performs code planning based on the ideas. Given the plan, Feynman translates ideas into simple declarative programs and iterates to receives feedback and visually refine diagrams. Finally, the declarative programs are rendered by the Penrose diagramming system. The optimization-based rendering of Penrose preserves the visual semantics while injecting fresh randomness into the layout, thereby producing diagrams with visual consistency and diversity. As a result, Feynman can author diagrams along with grounded captions with very little cost and time. Using Feynman, we synthesized a dataset with more than 100k well-aligned diagram-caption pairs. We also curate a visual-language benchmark, Diagramma, from freshly generated data. Diagramma can be used for evaluating the visual reasoning capabilities of vision-language models. We plan to release the dataset, benchmark, and the full agent pipeline as an open-source project.
title Feynman: Knowledge-Infused Diagramming Agent for Scalable Visual Designs
topic Machine Learning
Artificial Intelligence
Human-Computer Interaction
Multiagent Systems
Software Engineering
url https://arxiv.org/abs/2603.12597