3DrawAgent: Teaching LLM to Draw in 3D with Early Contrastive Experience
Fuente:
arXiv
Guardado en:
| Autores principales: | Xiao, Hongcan, Xiao, Xinyue, Wang, Yilin, Zhang, Yue, Qi, Yonggang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
StickMotion: Generating 3D Human Motions by Drawing a Stickman
por: Wang, Tao, et al.
Publicado: (2025)
por: Wang, Tao, et al.
Publicado: (2025)
ShadowDraw: From Any Object to Shadow-Drawing Compositional Art
por: Luo, Rundong, et al.
Publicado: (2025)
por: Luo, Rundong, et al.
Publicado: (2025)
VS-LLM: Visual-Semantic Depression Assessment based on LLM for Drawing Projection Test
por: Wu, Meiqi, et al.
Publicado: (2025)
por: Wu, Meiqi, et al.
Publicado: (2025)
Text-Enhanced Panoptic Symbol Spotting in CAD Drawings
por: Liu, Xianlin, et al.
Publicado: (2025)
por: Liu, Xianlin, et al.
Publicado: (2025)
ViRED: Prediction of Visual Relations in Engineering Drawings
por: Gu, Chao, et al.
Publicado: (2024)
por: Gu, Chao, et al.
Publicado: (2024)
From Drawings to Decisions: A Hybrid Vision-Language Framework for Parsing 2D Engineering Drawings into Structured Manufacturing Knowledge
por: Khan, Muhammad Tayyab, et al.
Publicado: (2025)
por: Khan, Muhammad Tayyab, et al.
Publicado: (2025)
DrawMotion: Generating 3D Human Motions by Freehand Drawing
por: Wang, Tao, et al.
Publicado: (2026)
por: Wang, Tao, et al.
Publicado: (2026)
SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model
por: Chang, Yifan, et al.
Publicado: (2025)
por: Chang, Yifan, et al.
Publicado: (2025)
MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing Understanding
por: Kou, Qian, et al.
Publicado: (2026)
por: Kou, Qian, et al.
Publicado: (2026)
Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
por: Wu, Junfei, et al.
Publicado: (2025)
por: Wu, Junfei, et al.
Publicado: (2025)
Fine-Tuning Vision-Language Model for Automated Engineering Drawing Information Extraction
por: Khan, Muhammad Tayyab, et al.
Publicado: (2024)
por: Khan, Muhammad Tayyab, et al.
Publicado: (2024)
PhyDrawGen: Physically Grounded Diagram Generation from Natural Language
por: Haque, Nafiul, et al.
Publicado: (2026)
por: Haque, Nafiul, et al.
Publicado: (2026)
Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing
por: Zeng, Ziyun, et al.
Publicado: (2025)
por: Zeng, Ziyun, et al.
Publicado: (2025)
Few Channels Draw The Whole Picture: Revealing Massive Activations in Diffusion Transformers
por: Turri, Evelyn, et al.
Publicado: (2026)
por: Turri, Evelyn, et al.
Publicado: (2026)
The Image Reconstruction Game: Drawing Common Ground Through Iterative Multimodal Dialogue
por: Hakimov, Sherzod, et al.
Publicado: (2026)
por: Hakimov, Sherzod, et al.
Publicado: (2026)
PCEvE: Part Contribution Evaluation Based Model Explanation for Human Figure Drawing Assessment and Beyond
por: Lee, Jongseo, et al.
Publicado: (2024)
por: Lee, Jongseo, et al.
Publicado: (2024)
Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation
por: Shi, Hanlei, et al.
Publicado: (2025)
por: Shi, Hanlei, et al.
Publicado: (2025)
Generating Sketches in a Hierarchical Auto-Regressive Process for Flexible Sketch Drawing Manipulation at Stroke-Level
por: Zang, Sicong, et al.
Publicado: (2025)
por: Zang, Sicong, et al.
Publicado: (2025)
DrawVideo: Generating Long Video from Storyboard Keyframe Sketches
por: Xu, Chuanzhi, et al.
Publicado: (2026)
por: Xu, Chuanzhi, et al.
Publicado: (2026)
PanGu-Draw: Advancing Resource-Efficient Text-to-Image Synthesis with Time-Decoupled Training and Reusable Coop-Diffusion
por: Lu, Guansong, et al.
Publicado: (2023)
por: Lu, Guansong, et al.
Publicado: (2023)
Automated Parsing of Engineering Drawings for Structured Information Extraction Using a Fine-tuned Document Understanding Transformer
por: Khan, Muhammad Tayyab, et al.
Publicado: (2025)
por: Khan, Muhammad Tayyab, et al.
Publicado: (2025)
How to Enable LLM with 3D Capacity? A Survey of Spatial Reasoning in LLM
por: Zha, Jirong, et al.
Publicado: (2025)
por: Zha, Jirong, et al.
Publicado: (2025)
PyPotteryInk: One-Step Diffusion Model for Sketch to Publication-ready Archaeological Drawings
por: Cardarelli, Lorenzo
Publicado: (2025)
por: Cardarelli, Lorenzo
Publicado: (2025)
Drawing the Line: Deep Segmentation for Extracting Art from Ancient Etruscan Mirrors
por: Sterzinger, Rafael, et al.
Publicado: (2024)
por: Sterzinger, Rafael, et al.
Publicado: (2024)
Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models
por: Wang, Xiaoyan, et al.
Publicado: (2025)
por: Wang, Xiaoyan, et al.
Publicado: (2025)
3D-Agent:Tri-Modal Multi-Agent Collaboration for Scalable 3D Object Annotation
por: Zhang, Jusheng, et al.
Publicado: (2026)
por: Zhang, Jusheng, et al.
Publicado: (2026)
Draw Your Mind: Personalized Generation via Condition-Level Modeling in Text-to-Image Diffusion Models
por: Kim, Hyungjin, et al.
Publicado: (2025)
por: Kim, Hyungjin, et al.
Publicado: (2025)
Pencils to Pixels: A Systematic Study of Creative Drawings across Children, Adults and AI
por: Nath, Surabhi S, et al.
Publicado: (2025)
por: Nath, Surabhi S, et al.
Publicado: (2025)
RegionPLC: Regional Point-Language Contrastive Learning for Open-World 3D Scene Understanding
por: Yang, Jihan, et al.
Publicado: (2023)
por: Yang, Jihan, et al.
Publicado: (2023)
Is Contrastive Distillation Enough for Learning Comprehensive 3D Representations?
por: Zhang, Yifan, et al.
Publicado: (2024)
por: Zhang, Yifan, et al.
Publicado: (2024)
3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decoding
por: Ogunleye, Makanjuola, et al.
Publicado: (2026)
por: Ogunleye, Makanjuola, et al.
Publicado: (2026)
Semantic Aware Feature Extraction for Enhanced 3D Reconstruction
por: Nap, Ronald, et al.
Publicado: (2026)
por: Nap, Ronald, et al.
Publicado: (2026)
Speed3R: Sparse Feed-forward 3D Reconstruction Models
por: Ren, Weining, et al.
Publicado: (2026)
por: Ren, Weining, et al.
Publicado: (2026)
Wonder3D++: Cross-domain Diffusion for High-fidelity 3D Generation from a Single Image
por: Yang, Yuxiao, et al.
Publicado: (2025)
por: Yang, Yuxiao, et al.
Publicado: (2025)
MUSES: 3D-Controllable Image Generation via Multi-Modal Agent Collaboration
por: Ding, Yanbo, et al.
Publicado: (2024)
por: Ding, Yanbo, et al.
Publicado: (2024)
Hyperbolic Contrastive Learning for Hierarchical 3D Point Cloud Embedding
por: Liu, Yingjie, et al.
Publicado: (2025)
por: Liu, Yingjie, et al.
Publicado: (2025)
Efficient3D: A Unified Framework for Adaptive and Debiased Token Reduction in 3D MLLMs
por: Lin, Yuhui, et al.
Publicado: (2026)
por: Lin, Yuhui, et al.
Publicado: (2026)
Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning
por: Yeh, Chun-Hsiao, et al.
Publicado: (2026)
por: Yeh, Chun-Hsiao, et al.
Publicado: (2026)
Spatio-Temporal Token Pruning for Efficient High-Resolution GUI Agents
por: Xu, Zhou, et al.
Publicado: (2026)
por: Xu, Zhou, et al.
Publicado: (2026)
TGBFormer: Transformer-GraphFormer Blender Network for Video Object Detection
por: Qi, Qiang, et al.
Publicado: (2025)
por: Qi, Qiang, et al.
Publicado: (2025)
Ejemplares similares
-
StickMotion: Generating 3D Human Motions by Drawing a Stickman
por: Wang, Tao, et al.
Publicado: (2025) -
ShadowDraw: From Any Object to Shadow-Drawing Compositional Art
por: Luo, Rundong, et al.
Publicado: (2025) -
VS-LLM: Visual-Semantic Depression Assessment based on LLM for Drawing Projection Test
por: Wu, Meiqi, et al.
Publicado: (2025) -
Text-Enhanced Panoptic Symbol Spotting in CAD Drawings
por: Liu, Xianlin, et al.
Publicado: (2025) -
ViRED: Prediction of Visual Relations in Engineering Drawings
por: Gu, Chao, et al.
Publicado: (2024)