From Pixels to Policies: Reinforcing Spatial Reasoning in Language Models for Content-Aware Layout Design
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Sha, Petrangeli, Stefano, Shen, Yu, Chen, Xiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLMs as Layout Designers: Enhanced Spatial Reasoning for Content-Aware Layout Generation
von: Li, Sha, et al.
Veröffentlicht: (2025)
von: Li, Sha, et al.
Veröffentlicht: (2025)
Proc3D: Procedural 3D Generation and Parametric Editing of 3D Shapes with Large Language Models
von: Raji, Fadlullah, et al.
Veröffentlicht: (2026)
von: Raji, Fadlullah, et al.
Veröffentlicht: (2026)
From Words to Worlds: Transforming One-line Prompt into Immersive Multi-modal Digital Stories with Communicative LLM Agent
von: Sohn, Samuel S., et al.
Veröffentlicht: (2024)
von: Sohn, Samuel S., et al.
Veröffentlicht: (2024)
3D-PreMise: Can Large Language Models Generate 3D Shapes with Sharp Features and Parametric Control?
von: Yuan, Zeqing, et al.
Veröffentlicht: (2024)
von: Yuan, Zeqing, et al.
Veröffentlicht: (2024)
Cutscene Agent: An LLM Agent Framework for Automated 3D Cutscene Generation
von: He, Lanshan, et al.
Veröffentlicht: (2026)
von: He, Lanshan, et al.
Veröffentlicht: (2026)
QQJ: Quantifying Qualitative Judgment for Scalable and Human-Aligned Evaluation of Generative AI
von: Veysi, Marjan, et al.
Veröffentlicht: (2026)
von: Veysi, Marjan, et al.
Veröffentlicht: (2026)
Narrative-to-Scene Generation: An LLM-Driven Pipeline for 2D Game Environments
von: Chen, Yi-Chun, et al.
Veröffentlicht: (2025)
von: Chen, Yi-Chun, et al.
Veröffentlicht: (2025)
DecoMind: A Generative AI System for Personalized Interior Design Layouts
von: Alshehri, Reema, et al.
Veröffentlicht: (2025)
von: Alshehri, Reema, et al.
Veröffentlicht: (2025)
Harnessing Adaptive Topology Representations for Zero-Shot Graph Question Answering
von: Wei, Yanbin, et al.
Veröffentlicht: (2025)
von: Wei, Yanbin, et al.
Veröffentlicht: (2025)
SuperPADL: Scaling Language-Directed Physics-Based Control with Progressive Supervised Distillation
von: Juravsky, Jordan, et al.
Veröffentlicht: (2024)
von: Juravsky, Jordan, et al.
Veröffentlicht: (2024)
PixelMan: Consistent Object Editing with Diffusion Models via Pixel Manipulation and Generation
von: Jiang, Liyao, et al.
Veröffentlicht: (2024)
von: Jiang, Liyao, et al.
Veröffentlicht: (2024)
Towards Understanding Graphical Perception in Large Multimodal Models
von: Zhang, Kai, et al.
Veröffentlicht: (2025)
von: Zhang, Kai, et al.
Veröffentlicht: (2025)
Modern Information Technologies in Scientific Research and Educational Activities
von: Malakhov, Kyrylo, et al.
Veröffentlicht: (2024)
von: Malakhov, Kyrylo, et al.
Veröffentlicht: (2024)
Layout-Corrector: Alleviating Layout Sticking Phenomenon in Discrete Diffusion Model
von: Iwai, Shoma, et al.
Veröffentlicht: (2024)
von: Iwai, Shoma, et al.
Veröffentlicht: (2024)
Grounding Language in Multi-Perspective Referential Communication
von: Tang, Zineng, et al.
Veröffentlicht: (2024)
von: Tang, Zineng, et al.
Veröffentlicht: (2024)
A Platform for Interactive AI Character Experiences
von: Wampfler, Rafael, et al.
Veröffentlicht: (2026)
von: Wampfler, Rafael, et al.
Veröffentlicht: (2026)
Co-Layout: LLM-driven Co-optimization for Interior Layout
von: Xiang, Chucheng, et al.
Veröffentlicht: (2025)
von: Xiang, Chucheng, et al.
Veröffentlicht: (2025)
DynamicGTR: Leveraging Graph Topology Representation Preferences to Boost VLM Capabilities on Graph QAs
von: Wei, Yanbin, et al.
Veröffentlicht: (2026)
von: Wei, Yanbin, et al.
Veröffentlicht: (2026)
LAYOUTDREAMER: Physics-guided Layout for Text-to-3D Compositional Scene Generation
von: Zhou, Yang, et al.
Veröffentlicht: (2025)
von: Zhou, Yang, et al.
Veröffentlicht: (2025)
StyleMotif: Multi-Modal Motion Stylization using Style-Content Cross Fusion
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
Causal Reasoning Elicits Controllable 3D Scene Generation
von: Chen, Shen, et al.
Veröffentlicht: (2025)
von: Chen, Shen, et al.
Veröffentlicht: (2025)
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation
von: Song, Lin, et al.
Veröffentlicht: (2026)
von: Song, Lin, et al.
Veröffentlicht: (2026)
Image Generation Models: A Technical History
von: Shirvani, Rouzbeh
Veröffentlicht: (2026)
von: Shirvani, Rouzbeh
Veröffentlicht: (2026)
DreamCraft: Text-Guided Generation of Functional 3D Environments in Minecraft
von: Earle, Sam, et al.
Veröffentlicht: (2024)
von: Earle, Sam, et al.
Veröffentlicht: (2024)
VisionGPT-3D: A Generalized Multimodal Agent for Enhanced 3D Vision Understanding
von: Kelly, Chris, et al.
Veröffentlicht: (2024)
von: Kelly, Chris, et al.
Veröffentlicht: (2024)
A Solver-Aided Hierarchical Language for LLM-Driven CAD Design
von: Jones, Benjamin T., et al.
Veröffentlicht: (2025)
von: Jones, Benjamin T., et al.
Veröffentlicht: (2025)
CAD-GPT: Synthesising CAD Construction Sequence with Spatial Reasoning-Enhanced Multimodal LLMs
von: Wang, Siyu, et al.
Veröffentlicht: (2024)
von: Wang, Siyu, et al.
Veröffentlicht: (2024)
PediatricsMQA: a Multi-modal Pediatrics Question Answering Benchmark
von: Bahaj, Adil, et al.
Veröffentlicht: (2025)
von: Bahaj, Adil, et al.
Veröffentlicht: (2025)
Agentic Design of Compositional Machines
von: Zhang, Wenqian, et al.
Veröffentlicht: (2025)
von: Zhang, Wenqian, et al.
Veröffentlicht: (2025)
Multi-LoRA Composition for Image Generation
von: Zhong, Ming, et al.
Veröffentlicht: (2024)
von: Zhong, Ming, et al.
Veröffentlicht: (2024)
MAPWise: Evaluating Vision-Language Models for Advanced Map Queries
von: Mukhopadhyay, Srija, et al.
Veröffentlicht: (2024)
von: Mukhopadhyay, Srija, et al.
Veröffentlicht: (2024)
CADgpt: Harnessing Natural Language Processing for 3D Modelling to Enhance Computer-Aided Design Workflows
von: Kapsalis, Timo
Veröffentlicht: (2024)
von: Kapsalis, Timo
Veröffentlicht: (2024)
NeuSDFusion: A Spatial-Aware Generative Model for 3D Shape Completion, Reconstruction, and Generation
von: Cui, Ruikai, et al.
Veröffentlicht: (2024)
von: Cui, Ruikai, et al.
Veröffentlicht: (2024)
Text-guided Controllable Mesh Refinement for Interactive 3D Modeling
von: Chen, Yun-Chun, et al.
Veröffentlicht: (2024)
von: Chen, Yun-Chun, et al.
Veröffentlicht: (2024)
SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models
von: Ray, Arijit, et al.
Veröffentlicht: (2024)
von: Ray, Arijit, et al.
Veröffentlicht: (2024)
SpringTime: Learning Simulatable Models of Cloth with Spatially-varying Constitutive Properties
von: Chen, Guanxiong, et al.
Veröffentlicht: (2025)
von: Chen, Guanxiong, et al.
Veröffentlicht: (2025)
Generating Traffic Scenarios via In-Context Learning to Learn Better Motion Planner
von: Aiersilan, Aizierjiang
Veröffentlicht: (2024)
von: Aiersilan, Aizierjiang
Veröffentlicht: (2024)
Empowering Children to Create AI-Enabled Augmented Reality Experiences
von: Zhang, Lei, et al.
Veröffentlicht: (2025)
von: Zhang, Lei, et al.
Veröffentlicht: (2025)
Multimodal Cinematic Video Synthesis Using Text-to-Image and Audio Generation Models
von: S, Sridhar, et al.
Veröffentlicht: (2025)
von: S, Sridhar, et al.
Veröffentlicht: (2025)
Stylus: Automatic Adapter Selection for Diffusion Models
von: Luo, Michael, et al.
Veröffentlicht: (2024)
von: Luo, Michael, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LLMs as Layout Designers: Enhanced Spatial Reasoning for Content-Aware Layout Generation
von: Li, Sha, et al.
Veröffentlicht: (2025) -
Proc3D: Procedural 3D Generation and Parametric Editing of 3D Shapes with Large Language Models
von: Raji, Fadlullah, et al.
Veröffentlicht: (2026) -
From Words to Worlds: Transforming One-line Prompt into Immersive Multi-modal Digital Stories with Communicative LLM Agent
von: Sohn, Samuel S., et al.
Veröffentlicht: (2024) -
3D-PreMise: Can Large Language Models Generate 3D Shapes with Sharp Features and Parametric Control?
von: Yuan, Zeqing, et al.
Veröffentlicht: (2024) -
Cutscene Agent: An LLM Agent Framework for Automated 3D Cutscene Generation
von: He, Lanshan, et al.
Veröffentlicht: (2026)