PhyDrawGen: Physically Grounded Diagram Generation from Natural Language
Fuente:
arXiv
Saved in:
| Main Authors: | Haque, Nafiul, Sakib, Syed Nazmus, Arman, Shifat E |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PlantVillageVQA: A Visual Question Answering Dataset for Benchmarking Vision-Language Models in Plant Science
by: Sakib, Syed Nazmus, et al.
Published: (2025)
by: Sakib, Syed Nazmus, et al.
Published: (2025)
Thinking Like a Botanist: Challenging Multimodal Language Models with Intent-Driven Chain-of-Inquiry
by: Sakib, Syed Nazmus, et al.
Published: (2026)
by: Sakib, Syed Nazmus, et al.
Published: (2026)
The Surface You Test Is Not the Surface That Breaks
by: Arman, Shifat E, et al.
Published: (2026)
by: Arman, Shifat E, et al.
Published: (2026)
PATHWAYS: Evaluating Investigation and Context Discovery in AI Web Agents
by: Arman, Shifat E., et al.
Published: (2026)
by: Arman, Shifat E., et al.
Published: (2026)
PhyGround: Benchmarking Physical Reasoning in Generative World Models
by: Lin, Juyi, et al.
Published: (2026)
by: Lin, Juyi, et al.
Published: (2026)
Scene Graph-Guided Generative AI Framework for Synthesizing and Evaluating Industrial Hazard Scenarios
by: Acharjee, Sanjay, et al.
Published: (2025)
by: Acharjee, Sanjay, et al.
Published: (2025)
PhyGenHOI: Physically-Aware 4D Generation of Dynamic Human-Object Interactions
by: Benishu, Omer, et al.
Published: (2026)
by: Benishu, Omer, et al.
Published: (2026)
PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation
by: Huang, Yidong, et al.
Published: (2026)
by: Huang, Yidong, et al.
Published: (2026)
PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation
by: Xue, Qiyao, et al.
Published: (2024)
by: Xue, Qiyao, et al.
Published: (2024)
Phi-Ground Tech Report: Advancing Perception in GUI Grounding
by: Zhang, Miaosen, et al.
Published: (2025)
by: Zhang, Miaosen, et al.
Published: (2025)
PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement
by: Zhan, Yu-Wei, et al.
Published: (2025)
by: Zhan, Yu-Wei, et al.
Published: (2025)
PhyPrompt: RL-based Prompt Refinement for Physically Plausible Text-to-Video Generation
by: Wu, Shang, et al.
Published: (2026)
by: Wu, Shang, et al.
Published: (2026)
PhiP-G: Physics-Guided Text-to-3D Compositional Scene Generation
by: Li, Qixuan, et al.
Published: (2025)
by: Li, Qixuan, et al.
Published: (2025)
VideoPhy: Evaluating Physical Commonsense for Video Generation
by: Bansal, Hritik, et al.
Published: (2024)
by: Bansal, Hritik, et al.
Published: (2024)
PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation
by: Liu, Shaowei, et al.
Published: (2024)
by: Liu, Shaowei, et al.
Published: (2024)
PhyCo: Learning Controllable Physical Priors for Generative Motion
by: Narayanan, Sriram, et al.
Published: (2026)
by: Narayanan, Sriram, et al.
Published: (2026)
PhyGile: Physics-Prefix Guided Motion Generation for Agile General Humanoid Motion Tracking
by: Bao, Jiacheng, et al.
Published: (2026)
by: Bao, Jiacheng, et al.
Published: (2026)
PhyWorld: Physics-Faithful World Model for Video Generation
by: Zhao, Pu, et al.
Published: (2026)
by: Zhao, Pu, et al.
Published: (2026)
PhyTracker: An Online Tracker for Phytoplankton
by: Yu, Yang, et al.
Published: (2024)
by: Yu, Yang, et al.
Published: (2024)
VisPhyWorld: Probing Physical Reasoning via Code-Driven Video Reconstruction
by: Liang, Jiarong, et al.
Published: (2026)
by: Liang, Jiarong, et al.
Published: (2026)
"PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models
by: Gu, Jing, et al.
Published: (2025)
by: Gu, Jing, et al.
Published: (2025)
WorldAfford: Affordance Grounding based on Natural Language Instructions
by: Chen, Changmao, et al.
Published: (2024)
by: Chen, Changmao, et al.
Published: (2024)
GenSeg-R1: RL-Driven Vision-Language Grounding for Fine-Grained Referring Segmentation
by: Hegde, Sandesh, et al.
Published: (2026)
by: Hegde, Sandesh, et al.
Published: (2026)
PhyGrasp: Generalizing Robotic Grasping with Physics-informed Large Multimodal Models
by: Guo, Dingkun, et al.
Published: (2024)
by: Guo, Dingkun, et al.
Published: (2024)
The Image Reconstruction Game: Drawing Common Ground Through Iterative Multimodal Dialogue
by: Hakimov, Sherzod, et al.
Published: (2026)
by: Hakimov, Sherzod, et al.
Published: (2026)
Gen4Gen: Generative Data Pipeline for Generative Multi-Concept Composition
by: Yeh, Chun-Hsiao, et al.
Published: (2024)
by: Yeh, Chun-Hsiao, et al.
Published: (2024)
DRAGON: A Benchmark for Evidence-Grounded Visual Reasoning over Diagrams
by: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Published: (2026)
by: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Published: (2026)
3DrawAgent: Teaching LLM to Draw in 3D with Early Contrastive Experience
by: Xiao, Hongcan, et al.
Published: (2026)
by: Xiao, Hongcan, et al.
Published: (2026)
SugarcaneShuffleNet: A Very Fast, Lightweight Convolutional Neural Network for Diagnosis of 15 Sugarcane Leaf Diseases
by: Arman, Shifat E., et al.
Published: (2025)
by: Arman, Shifat E., et al.
Published: (2025)
Physical Evaluation of Naturalistic Adversarial Patches for Camera-Based Traffic-Sign Detection
by: D'Urso, Brianna, et al.
Published: (2026)
by: D'Urso, Brianna, et al.
Published: (2026)
Phi-4-reasoning-vision-15B Technical Report
by: Aneja, Jyoti, et al.
Published: (2026)
by: Aneja, Jyoti, et al.
Published: (2026)
Fine-Tuning Vision-Language Model for Automated Engineering Drawing Information Extraction
by: Khan, Muhammad Tayyab, et al.
Published: (2024)
by: Khan, Muhammad Tayyab, et al.
Published: (2024)
Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
by: Wu, Junfei, et al.
Published: (2025)
by: Wu, Junfei, et al.
Published: (2025)
GenAI-DrawIO-Creator: A Framework for Automated Diagram Generation
by: Yu, Jinze, et al.
Published: (2026)
by: Yu, Jinze, et al.
Published: (2026)
From Canopy to Ground via ForestGen3D: Learning Cross-Domain Generation of 3D Forest Structure from Aerial-to-Terrestrial LiDAR
by: Castorena, Juan, et al.
Published: (2025)
by: Castorena, Juan, et al.
Published: (2025)
VerLM: Explaining Face Verification Using Natural Language
by: Hannan, Syed Abdul, et al.
Published: (2026)
by: Hannan, Syed Abdul, et al.
Published: (2026)
StickMotion: Generating 3D Human Motions by Drawing a Stickman
by: Wang, Tao, et al.
Published: (2025)
by: Wang, Tao, et al.
Published: (2025)
SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model
by: Chang, Yifan, et al.
Published: (2025)
by: Chang, Yifan, et al.
Published: (2025)
Gen-LangSplat: Generalized Language Gaussian Splatting with Pre-Trained Feature Compression
by: Saxena, Pranav
Published: (2025)
by: Saxena, Pranav
Published: (2025)
OmniGen: Unified Image Generation
by: Xiao, Shitao, et al.
Published: (2024)
by: Xiao, Shitao, et al.
Published: (2024)
Similar Items
-
PlantVillageVQA: A Visual Question Answering Dataset for Benchmarking Vision-Language Models in Plant Science
by: Sakib, Syed Nazmus, et al.
Published: (2025) -
Thinking Like a Botanist: Challenging Multimodal Language Models with Intent-Driven Chain-of-Inquiry
by: Sakib, Syed Nazmus, et al.
Published: (2026) -
The Surface You Test Is Not the Surface That Breaks
by: Arman, Shifat E, et al.
Published: (2026) -
PATHWAYS: Evaluating Investigation and Context Discovery in AI Web Agents
by: Arman, Shifat E., et al.
Published: (2026) -
PhyGround: Benchmarking Physical Reasoning in Generative World Models
by: Lin, Juyi, et al.
Published: (2026)