Draw2Think: Harnessing Geometry Reasoning through Constraint Engine Interaction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Juncheng, Du, Jiawei, Zhang, Xin, Zhou, Joey Tianyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Diversity-Driven Synthesis: Enhancing Dataset Distillation through Directed Weight Adjustment
von: Du, Jiawei, et al.
Veröffentlicht: (2024)
von: Du, Jiawei, et al.
Veröffentlicht: (2024)
More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
von: Liu, Chengzhi, et al.
Veröffentlicht: (2025)
von: Liu, Chengzhi, et al.
Veröffentlicht: (2025)
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
von: Wu, Juncheng, et al.
Veröffentlicht: (2026)
von: Wu, Juncheng, et al.
Veröffentlicht: (2026)
Breaking Class Barriers: Efficient Dataset Distillation via Inter-Class Feature Compensator
von: Zhang, Xin, et al.
Veröffentlicht: (2024)
von: Zhang, Xin, et al.
Veröffentlicht: (2024)
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
PVP: Polar Representation Boost for 3D Semantic Occupancy Prediction
von: Xue, Yujing, et al.
Veröffentlicht: (2024)
von: Xue, Yujing, et al.
Veröffentlicht: (2024)
Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models
von: Li, Yunxin, et al.
Veröffentlicht: (2025)
von: Li, Yunxin, et al.
Veröffentlicht: (2025)
Beyond Modality Collapse: Representations Blending for Multimodal Dataset Distillation
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
KPL: Training-Free Medical Knowledge Mining of Vision-Language Models
von: Liu, Jiaxiang, et al.
Veröffentlicht: (2025)
von: Liu, Jiaxiang, et al.
Veröffentlicht: (2025)
Spanning Training Progress: Temporal Dual-Depth Scoring (TDDS) for Enhanced Dataset Pruning
von: Zhang, Xin, et al.
Veröffentlicht: (2023)
von: Zhang, Xin, et al.
Veröffentlicht: (2023)
VTimeCoT: Thinking by Drawing for Video Temporal Grounding and Reasoning
von: Zhang, Jinglei, et al.
Veröffentlicht: (2025)
von: Zhang, Jinglei, et al.
Veröffentlicht: (2025)
Modelship Attribution: Tracing Multi-Stage Manipulations Across Generative Models
von: Tan, Zhiya, et al.
Veröffentlicht: (2025)
von: Tan, Zhiya, et al.
Veröffentlicht: (2025)
MedCoT: Medical Chain of Thought via Hierarchical Expert
von: Liu, Jiaxiang, et al.
Veröffentlicht: (2024)
von: Liu, Jiaxiang, et al.
Veröffentlicht: (2024)
Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks
von: Zhang, Wenqi, et al.
Veröffentlicht: (2025)
von: Zhang, Wenqi, et al.
Veröffentlicht: (2025)
ETCHR: Editing To Clarify and Harness Reasoning
von: Zhang, Beichen, et al.
Veröffentlicht: (2026)
von: Zhang, Beichen, et al.
Veröffentlicht: (2026)
DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding
von: Feng, Xiang, et al.
Veröffentlicht: (2026)
von: Feng, Xiang, et al.
Veröffentlicht: (2026)
MMGR: Multi-Modal Generative Reasoning
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
See, Think, Learn: A Self-Taught Multimodal Reasoner
von: Sharma, Sourabh, et al.
Veröffentlicht: (2025)
von: Sharma, Sourabh, et al.
Veröffentlicht: (2025)
ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoning
von: Tao, Xingjian, et al.
Veröffentlicht: (2026)
von: Tao, Xingjian, et al.
Veröffentlicht: (2026)
Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation
von: Zhou, Yucheng, et al.
Veröffentlicht: (2025)
von: Zhou, Yucheng, et al.
Veröffentlicht: (2025)
Evolution-aware VAriance (EVA) Coreset Selection for Medical Image Classification
von: Hong, Yuxin, et al.
Veröffentlicht: (2024)
von: Hong, Yuxin, et al.
Veröffentlicht: (2024)
Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
von: Zhu, Wenxin, et al.
Veröffentlicht: (2025)
von: Zhu, Wenxin, et al.
Veröffentlicht: (2025)
Modest-Align: Data-Efficient Alignment for Vision-Language Models
von: Liu, Jiaxiang, et al.
Veröffentlicht: (2025)
von: Liu, Jiaxiang, et al.
Veröffentlicht: (2025)
Drawing the Line: Enhancing Trustworthiness of MLLMs Through the Power of Refusal
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation
von: Zhang, Xiaofeng, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaofeng, et al.
Veröffentlicht: (2024)
Thinking with Programming Vision: Towards a Unified View for Thinking with Images
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
Thinking with Geometry: Active Geometry Integration for Spatial Reasoning
von: Li, Haoyuan, et al.
Veröffentlicht: (2026)
von: Li, Haoyuan, et al.
Veröffentlicht: (2026)
MedFrameQA: A Multi-Image Medical VQA Benchmark for Clinical Reasoning
von: Yu, Suhao, et al.
Veröffentlicht: (2025)
von: Yu, Suhao, et al.
Veröffentlicht: (2025)
AEGIS: Authenticity Evaluation Benchmark for AI-Generated Video Sequences
von: Li, Jieyu, et al.
Veröffentlicht: (2025)
von: Li, Jieyu, et al.
Veröffentlicht: (2025)
Think Visually, Reason Textually: Vision-Language Synergy in ARC
von: Zhang, Beichen, et al.
Veröffentlicht: (2025)
von: Zhang, Beichen, et al.
Veröffentlicht: (2025)
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction
von: Guo, Zichun, et al.
Veröffentlicht: (2026)
von: Guo, Zichun, et al.
Veröffentlicht: (2026)
Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attention
von: Tian, Changyuan, et al.
Veröffentlicht: (2026)
von: Tian, Changyuan, et al.
Veröffentlicht: (2026)
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
von: Du, Yifan, et al.
Veröffentlicht: (2023)
von: Du, Yifan, et al.
Veröffentlicht: (2023)
Sandboxed Coding Agents are Competitive Omni-modal Task Solvers
von: Chen, Dongping, et al.
Veröffentlicht: (2026)
von: Chen, Dongping, et al.
Veröffentlicht: (2026)
AgentThink: A Unified Framework for Tool-Augmented Chain-of-Thought Reasoning in Vision-Language Models for Autonomous Driving
von: Qian, Kangan, et al.
Veröffentlicht: (2025)
von: Qian, Kangan, et al.
Veröffentlicht: (2025)
TemMed-Bench: Evaluating Temporal Medical Image Reasoning in Vision-Language Models
von: Zhang, Junyi, et al.
Veröffentlicht: (2025)
von: Zhang, Junyi, et al.
Veröffentlicht: (2025)
Agentic Spatio-Temporal Grounding via Collaborative Reasoning
von: Zhao, Heng, et al.
Veröffentlicht: (2026)
von: Zhao, Heng, et al.
Veröffentlicht: (2026)
Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback
von: Yuan, Jiakang, et al.
Veröffentlicht: (2025)
von: Yuan, Jiakang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Diversity-Driven Synthesis: Enhancing Dataset Distillation through Directed Weight Adjustment
von: Du, Jiawei, et al.
Veröffentlicht: (2024) -
More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
von: Liu, Chengzhi, et al.
Veröffentlicht: (2025) -
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
von: Wu, Juncheng, et al.
Veröffentlicht: (2026) -
Breaking Class Barriers: Efficient Dataset Distillation via Inter-Class Feature Compensator
von: Zhang, Xin, et al.
Veröffentlicht: (2024) -
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)