Re-Thinking Inverse Graphics With Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kulits, Peter, Feng, Haiwen, Liu, Weiyang, Abrevaya, Victoria, Black, Michael J. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reconstructing Animals and the Wild
von: Kulits, Peter, et al.
Veröffentlicht: (2024)
von: Kulits, Peter, et al.
Veröffentlicht: (2024)
Can Large Language Models Understand Symbolic Graphics Programs?
von: Qiu, Zeju, et al.
Veröffentlicht: (2024)
von: Qiu, Zeju, et al.
Veröffentlicht: (2024)
InterDyn: Controllable Interactive Dynamics with Video Diffusion Models
von: Akkerman, Rick, et al.
Veröffentlicht: (2024)
von: Akkerman, Rick, et al.
Veröffentlicht: (2024)
GenLit: Reformulating Single-Image Relighting as Video Generation
von: Bharadwaj, Shrisha, et al.
Veröffentlicht: (2024)
von: Bharadwaj, Shrisha, et al.
Veröffentlicht: (2024)
Explorative Inbetweening of Time and Space
von: Feng, Haiwen, et al.
Veröffentlicht: (2024)
von: Feng, Haiwen, et al.
Veröffentlicht: (2024)
Thinking in Blender: Staged Executable Inverse Graphics with Vision-Language Models
von: He, Guangzhao, et al.
Veröffentlicht: (2026)
von: He, Guangzhao, et al.
Veröffentlicht: (2026)
Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning
von: Yin, Shaofeng, et al.
Veröffentlicht: (2026)
von: Yin, Shaofeng, et al.
Veröffentlicht: (2026)
BrickNet: Graph-Backed Generative Brick Assembly
von: Kulits, Peter, et al.
Veröffentlicht: (2026)
von: Kulits, Peter, et al.
Veröffentlicht: (2026)
Symbolic Graphics Programming with Large Language Models
von: Chen, Yamei, et al.
Veröffentlicht: (2025)
von: Chen, Yamei, et al.
Veröffentlicht: (2025)
EgoThink: Evaluating First-Person Perspective Thinking Capability of Vision-Language Models
von: Cheng, Sijie, et al.
Veröffentlicht: (2023)
von: Cheng, Sijie, et al.
Veröffentlicht: (2023)
From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
von: Zhu, Wenxin, et al.
Veröffentlicht: (2025)
von: Zhu, Wenxin, et al.
Veröffentlicht: (2025)
Verbalized Machine Learning: Revisiting Machine Learning with Language Models
von: Xiao, Tim Z., et al.
Veröffentlicht: (2024)
von: Xiao, Tim Z., et al.
Veröffentlicht: (2024)
Graphic Design with Large Multimodal Model
von: Cheng, Yutao, et al.
Veröffentlicht: (2024)
von: Cheng, Yutao, et al.
Veröffentlicht: (2024)
Parameter-Efficient Orthogonal Finetuning via Butterfly Factorization
von: Liu, Weiyang, et al.
Veröffentlicht: (2023)
von: Liu, Weiyang, et al.
Veröffentlicht: (2023)
Think-Reflect-Revise: A Policy-Guided Reflective Framework for Safety Alignment in Large Vision Language Models
von: Weng, Fenghua, et al.
Veröffentlicht: (2025)
von: Weng, Fenghua, et al.
Veröffentlicht: (2025)
Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models
von: Li, Yunxin, et al.
Veröffentlicht: (2025)
von: Li, Yunxin, et al.
Veröffentlicht: (2025)
Toward Human Understanding with Controllable Synthesis
von: Cuevas-Velasquez, Hanz, et al.
Veröffentlicht: (2024)
von: Cuevas-Velasquez, Hanz, et al.
Veröffentlicht: (2024)
Fast-Slow Thinking GRPO for Large Vision-Language Model Reasoning
von: Xiao, Wenyi, et al.
Veröffentlicht: (2025)
von: Xiao, Wenyi, et al.
Veröffentlicht: (2025)
Visually Descriptive Language Model for Vector Graphics Reasoning
von: Wang, Zhenhailong, et al.
Veröffentlicht: (2024)
von: Wang, Zhenhailong, et al.
Veröffentlicht: (2024)
Thinking with Programming Vision: Towards a Unified View for Thinking with Images
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
ChatHuman: Chatting about 3D Humans with Tools
von: Lin, Jing, et al.
Veröffentlicht: (2024)
von: Lin, Jing, et al.
Veröffentlicht: (2024)
Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models
von: Chen, Jiaxing, et al.
Veröffentlicht: (2024)
von: Chen, Jiaxing, et al.
Veröffentlicht: (2024)
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding
von: Li, Yun, et al.
Veröffentlicht: (2025)
von: Li, Yun, et al.
Veröffentlicht: (2025)
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
von: Wu, Juncheng, et al.
Veröffentlicht: (2026)
von: Wu, Juncheng, et al.
Veröffentlicht: (2026)
Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Models
von: Jian, Pu, et al.
Veröffentlicht: (2025)
von: Jian, Pu, et al.
Veröffentlicht: (2025)
AI Based Font Pair Suggestion Modelling For Graphic Design
von: Singh, Aryan, et al.
Veröffentlicht: (2025)
von: Singh, Aryan, et al.
Veröffentlicht: (2025)
Video Understanding with Large Language Models: A Survey
von: Tang, Yolo Y., et al.
Veröffentlicht: (2023)
von: Tang, Yolo Y., et al.
Veröffentlicht: (2023)
Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models
von: Zhang, Huixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Huixuan, et al.
Veröffentlicht: (2025)
StrokeNUWA: Tokenizing Strokes for Vector Graphic Synthesis
von: Tang, Zecheng, et al.
Veröffentlicht: (2024)
von: Tang, Zecheng, et al.
Veröffentlicht: (2024)
Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts
von: Xu, Haolei, et al.
Veröffentlicht: (2026)
von: Xu, Haolei, et al.
Veröffentlicht: (2026)
Towards Understanding Graphical Perception in Large Multimodal Models
von: Zhang, Kai, et al.
Veröffentlicht: (2025)
von: Zhang, Kai, et al.
Veröffentlicht: (2025)
ReLoop: "Seeing Twice and Thinking Backwards" via Closed-loop Training to Mitigate Hallucinations in Multimodal understanding
von: Yang, Jianjiang, et al.
Veröffentlicht: (2025)
von: Yang, Jianjiang, et al.
Veröffentlicht: (2025)
Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
Leveraging Large Language Models for Scalable Vector Graphics-Driven Image Understanding
von: Cai, Mu, et al.
Veröffentlicht: (2023)
von: Cai, Mu, et al.
Veröffentlicht: (2023)
MLLMReID: Multimodal Large Language Model-based Person Re-identification
von: Yang, Shan, et al.
Veröffentlicht: (2024)
von: Yang, Shan, et al.
Veröffentlicht: (2024)
Causal Graphical Models for Vision-Language Compositional Understanding
von: Parascandolo, Fiorenzo, et al.
Veröffentlicht: (2024)
von: Parascandolo, Fiorenzo, et al.
Veröffentlicht: (2024)
Predicting 4D Hand Trajectory from Monocular Videos
von: Ye, Yufei, et al.
Veröffentlicht: (2025)
von: Ye, Yufei, et al.
Veröffentlicht: (2025)
Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward
von: Niu, Yuwei, et al.
Veröffentlicht: (2025)
von: Niu, Yuwei, et al.
Veröffentlicht: (2025)
EmotionHallucer: Evaluating Emotion Hallucinations in Multimodal Large Language Models
von: Xing, Bohao, et al.
Veröffentlicht: (2025)
von: Xing, Bohao, et al.
Veröffentlicht: (2025)
LLMVA-GEBC: Large Language Model with Video Adapter for Generic Event Boundary Captioning
von: Tang, Yolo Yunlong, et al.
Veröffentlicht: (2023)
von: Tang, Yolo Yunlong, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Reconstructing Animals and the Wild
von: Kulits, Peter, et al.
Veröffentlicht: (2024) -
Can Large Language Models Understand Symbolic Graphics Programs?
von: Qiu, Zeju, et al.
Veröffentlicht: (2024) -
InterDyn: Controllable Interactive Dynamics with Video Diffusion Models
von: Akkerman, Rick, et al.
Veröffentlicht: (2024) -
GenLit: Reformulating Single-Image Relighting as Video Generation
von: Bharadwaj, Shrisha, et al.
Veröffentlicht: (2024) -
Explorative Inbetweening of Time and Space
von: Feng, Haiwen, et al.
Veröffentlicht: (2024)