Text2Arch: A Dataset for Generating Scientific Architecture Diagrams from Natural Language Descriptions

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Garg, Shivank, Mittal, Sankalp, Gupta, Manish
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908970374922240
author Garg, Shivank
Mittal, Sankalp
Gupta, Manish
author_facet Garg, Shivank
Mittal, Sankalp
Gupta, Manish
contents Communicating complex system designs or scientific processes through text alone is inefficient and prone to ambiguity. A system that automatically generates scientific architecture diagrams from text with high semantic fidelity can be useful in multiple applications like enterprise architecture visualization, AI-driven software design, and educational content creation. Hence, in this paper, we focus on leveraging language models to perform semantic understanding of the input text description to generate intermediate code that can be processed to generate high-fidelity architecture diagrams. Unfortunately, no clean large-scale open-access dataset exists, implying lack of any effective open models for this task. Hence, we contribute a comprehensive dataset, \system, comprising scientific architecture images, their corresponding textual descriptions, and associated DOT code representations. Leveraging this resource, we fine-tune a suite of small language models, and also perform in-context learning using GPT-4o. Through extensive experimentation, we show that \system{} models significantly outperform existing baseline models like DiagramAgent and perform at par with in-context learning-based generations from GPT-4o. We make the code, data and models publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2604_14941
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Text2Arch: A Dataset for Generating Scientific Architecture Diagrams from Natural Language Descriptions
Garg, Shivank
Mittal, Sankalp
Gupta, Manish
Computation and Language
Communicating complex system designs or scientific processes through text alone is inefficient and prone to ambiguity. A system that automatically generates scientific architecture diagrams from text with high semantic fidelity can be useful in multiple applications like enterprise architecture visualization, AI-driven software design, and educational content creation. Hence, in this paper, we focus on leveraging language models to perform semantic understanding of the input text description to generate intermediate code that can be processed to generate high-fidelity architecture diagrams. Unfortunately, no clean large-scale open-access dataset exists, implying lack of any effective open models for this task. Hence, we contribute a comprehensive dataset, \system, comprising scientific architecture images, their corresponding textual descriptions, and associated DOT code representations. Leveraging this resource, we fine-tune a suite of small language models, and also perform in-context learning using GPT-4o. Through extensive experimentation, we show that \system{} models significantly outperform existing baseline models like DiagramAgent and perform at par with in-context learning-based generations from GPT-4o. We make the code, data and models publicly available.
title Text2Arch: A Dataset for Generating Scientific Architecture Diagrams from Natural Language Descriptions
topic Computation and Language
url https://arxiv.org/abs/2604.14941