CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Qi, Li, Honglin, Yu, Yingchen, Zhou, Haoyi, Yang, Lin, Bai, Song, She, Qi, Huang, Zilong, Zhao, Yunqing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ThinkGen: Generalized Thinking for Visual Generation
by: Jiao, Siyu, et al.
Published: (2025)
by: Jiao, Siyu, et al.
Published: (2025)
TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding
by: Yang, Zuhao, et al.
Published: (2025)
by: Yang, Zuhao, et al.
Published: (2025)
Let ViT Speak: Generative Language-Image Pre-training
by: Fang, Yan, et al.
Published: (2026)
by: Fang, Yan, et al.
Published: (2026)
AbductiveMLLM: Boosting Visual Abductive Reasoning Within MLLMs
by: Chang, Boyu, et al.
Published: (2026)
by: Chang, Boyu, et al.
Published: (2026)
AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning
by: Song, Mingyang, et al.
Published: (2026)
by: Song, Mingyang, et al.
Published: (2026)
GUI-Rise: Structured Reasoning and History Summarization for GUI Navigation
by: Liu, Tao, et al.
Published: (2025)
by: Liu, Tao, et al.
Published: (2025)
Text4Seg++: Advancing Image Segmentation via Generative Language Modeling
by: Lan, Mengcheng, et al.
Published: (2025)
by: Lan, Mengcheng, et al.
Published: (2025)
Monocular Normal Estimation via Shading Sequence Estimation
by: Li, Zongrui, et al.
Published: (2026)
by: Li, Zongrui, et al.
Published: (2026)
Versatile Transition Generation with Image-to-Video Diffusion
by: Yang, Zuhao, et al.
Published: (2025)
by: Yang, Zuhao, et al.
Published: (2025)
Look-Back: Implicit Visual Re-focusing in MLLM Reasoning
by: Yang, Shuo, et al.
Published: (2025)
by: Yang, Shuo, et al.
Published: (2025)
Debiasing Text-to-Image Diffusion Models
by: He, Ruifei, et al.
Published: (2024)
by: He, Ruifei, et al.
Published: (2024)
Boosting MLLM Reasoning with Text-Debiased Hint-GRPO
by: Huang, Qihan, et al.
Published: (2025)
by: Huang, Qihan, et al.
Published: (2025)
Revisiting MLLM Token Technology through the Lens of Classical Visual Coding
by: Liu, Jinming, et al.
Published: (2025)
by: Liu, Jinming, et al.
Published: (2025)
Local Happiness and Executive Compensation
by: Zilong Song, et al.
Published: (2025)
by: Zilong Song, et al.
Published: (2025)
Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression Tasks
by: Dong, Qihua, et al.
Published: (2026)
by: Dong, Qihua, et al.
Published: (2026)
Syn-GRPO: Self-Evolving Data Synthesis for MLLM Perception Reasoning
by: Huang, Qihan, et al.
Published: (2025)
by: Huang, Qihan, et al.
Published: (2025)
UniCode: Augmenting Evaluation for Code Reasoning
by: Zheng, Xinyue, et al.
Published: (2025)
by: Zheng, Xinyue, et al.
Published: (2025)
RedCode: Risky Code Execution and Generation Benchmark for Code Agents
by: Guo, Chengquan, et al.
Published: (2024)
by: Guo, Chengquan, et al.
Published: (2024)
Nonlinear Virtual Inertia Control of WTGs for Enhancing Primary Frequency Response and Suppressing Drive-Train Torsional Oscillations
by: Liu, Bi, et al.
Published: (2020)
by: Liu, Bi, et al.
Published: (2020)
Demystifying Errors in LLM Reasoning Traces: An Empirical Study of Code Execution Simulation
by: Abdollahi, Mohammad, et al.
Published: (2025)
by: Abdollahi, Mohammad, et al.
Published: (2025)
AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning
by: Zou, Jiaru, et al.
Published: (2025)
by: Zou, Jiaru, et al.
Published: (2025)
KnowPath: Knowledge-enhanced Reasoning via LLM-generated Inference Paths over Knowledge Graphs
by: Zhao, Qi, et al.
Published: (2025)
by: Zhao, Qi, et al.
Published: (2025)
Constructing Ophthalmic MLLM for Positioning-diagnosis Collaboration Through Clinical Cognitive Chain Reasoning
by: Liu, Xinyao, et al.
Published: (2025)
by: Liu, Xinyao, et al.
Published: (2025)
Code Execution as Grounded Supervision for LLM Reasoning
by: Jung, Dongwon, et al.
Published: (2025)
by: Jung, Dongwon, et al.
Published: (2025)
A Tool for In-depth Analysis of Code Execution Reasoning of Large Language Models
by: Liu, Changshu, et al.
Published: (2025)
by: Liu, Changshu, et al.
Published: (2025)
Place-it-R1: Unlocking Environment-aware Reasoning Potential of MLLM for Video Object Insertion
by: Gu, Bohai, et al.
Published: (2026)
by: Gu, Bohai, et al.
Published: (2026)
Learning Project-wise Subsequent Code Edits via Interleaving Neural-based Induction and Tool-based Deduction
by: Liu, Chenyan, et al.
Published: (2026)
by: Liu, Chenyan, et al.
Published: (2026)
DCI: A Coordinated Allocation and Filling Workload-Aware Dual-Cache Allocation GNN Inference Acceleration System
by: Luo, Yi, et al.
Published: (2025)
by: Luo, Yi, et al.
Published: (2025)
Thinking Isn't an Illusion: Overcoming the Limitations of Reasoning Models via Tool Augmentations
by: Song, Zhao, et al.
Published: (2025)
by: Song, Zhao, et al.
Published: (2025)
ChainV: Atomic Visual Hints Make Multimodal Reasoning Shorter and Better
by: Zhang, Yuan, et al.
Published: (2025)
by: Zhang, Yuan, et al.
Published: (2025)
Rethinking Visual Neglect: Steering via Context-Preference for MLLM Hallucination Mitigation
by: Wu, Jingwen, et al.
Published: (2026)
by: Wu, Jingwen, et al.
Published: (2026)
Learning to Animate Images from A Few Videos to Portray Delicate Human Actions
by: Li, Haoxin, et al.
Published: (2025)
by: Li, Haoxin, et al.
Published: (2025)
FluxEDA: A Unified Execution Infrastructure for Stateful Agentic EDA
by: Chen, Zhengrui, et al.
Published: (2026)
by: Chen, Zhengrui, et al.
Published: (2026)
Not All Noises Are Created Equally:Diffusion Noise Selection and Optimization
by: Qi, Zipeng, et al.
Published: (2024)
by: Qi, Zipeng, et al.
Published: (2024)
An Attentive Dual-Encoder Framework Leveraging Multimodal Visual and Semantic Information for Automatic OSAHS Diagnosis
by: Wei, Yingchen, et al.
Published: (2024)
by: Wei, Yingchen, et al.
Published: (2024)
Towards Long-window Anchoring in Vision-Language Model Distillation
by: Zhou, Haoyi, et al.
Published: (2025)
by: Zhou, Haoyi, et al.
Published: (2025)
AIVD: Adaptive Edge-Cloud Collaboration for Accurate and Efficient Industrial Visual Detection
by: Hu, Yunqing, et al.
Published: (2026)
by: Hu, Yunqing, et al.
Published: (2026)
Interaction-Consistent Object Removal via MLLM-Based Reasoning
by: Huang, Ching-Kai, et al.
Published: (2026)
by: Huang, Ching-Kai, et al.
Published: (2026)
What Really Improves Mathematical Reasoning: Structured Reasoning Signals Beyond Pure Code
by: Zhao, Yuze, et al.
Published: (2026)
by: Zhao, Yuze, et al.
Published: (2026)
Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning
by: Wu, Tong, et al.
Published: (2025)
by: Wu, Tong, et al.
Published: (2025)
Similar Items
-
ThinkGen: Generalized Thinking for Visual Generation
by: Jiao, Siyu, et al.
Published: (2025) -
TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding
by: Yang, Zuhao, et al.
Published: (2025) -
Let ViT Speak: Generative Language-Image Pre-training
by: Fang, Yan, et al.
Published: (2026) -
AbductiveMLLM: Boosting Visual Abductive Reasoning Within MLLMs
by: Chang, Boyu, et al.
Published: (2026) -
AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning
by: Song, Mingyang, et al.
Published: (2026)