PhyDrawGen: Physically Grounded Diagram Generation from Natural Language

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Haque, Nafiul, Sakib, Syed Nazmus, Arman, Shifat E
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917545083142144
author Haque, Nafiul
Sakib, Syed Nazmus
Arman, Shifat E
author_facet Haque, Nafiul
Sakib, Syed Nazmus
Arman, Shifat E
contents Generating physics diagrams from text requires strict adherence to physical laws. While current generative models produce visually plausible outputs, they systematically hallucinate force vectors, ignore conservation laws, and violate geometric constraints. We present PhyDrawGen, a neuro-symbolic pipeline that decouples semantic scene understanding from physical constraint satisfaction. First, a large language model extracts a typed scene graph from the problem text. A deterministic solver then converts this graph into a Planar Straight-Line Graph (PSLG), encoding force balance, optical paths, and field topologies as exact geometric primitives. Finally, a fine-tuned Qwen-VL model implements a visually grounded propose-verify loop to iteratively correct any constraint violations. Evaluated on a benchmark of 1,449 problems spanning mechanics, optics, and electromagnetism, PhyDrawGen significantly outperforms GPT-5-image, Gemini 2.5 Flash, and Gemini 3 Pro, demonstrating robust physical accuracy even on unusual-object problems.
format Preprint
id arxiv_https___arxiv_org_abs_2605_30512
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PhyDrawGen: Physically Grounded Diagram Generation from Natural Language
Haque, Nafiul
Sakib, Syed Nazmus
Arman, Shifat E
Artificial Intelligence
Computer Vision and Pattern Recognition
Generating physics diagrams from text requires strict adherence to physical laws. While current generative models produce visually plausible outputs, they systematically hallucinate force vectors, ignore conservation laws, and violate geometric constraints. We present PhyDrawGen, a neuro-symbolic pipeline that decouples semantic scene understanding from physical constraint satisfaction. First, a large language model extracts a typed scene graph from the problem text. A deterministic solver then converts this graph into a Planar Straight-Line Graph (PSLG), encoding force balance, optical paths, and field topologies as exact geometric primitives. Finally, a fine-tuned Qwen-VL model implements a visually grounded propose-verify loop to iteratively correct any constraint violations. Evaluated on a benchmark of 1,449 problems spanning mechanics, optics, and electromagnetism, PhyDrawGen significantly outperforms GPT-5-image, Gemini 2.5 Flash, and Gemini 3 Pro, demonstrating robust physical accuracy even on unusual-object problems.
title PhyDrawGen: Physically Grounded Diagram Generation from Natural Language
topic Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.30512