Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Wenhao, Xia, Yu, Huang, Jinlong, Lu, Shiyin, Chen, Qing-Guo, Xu, Zhao, Luo, Weihua, Zhang, Kaifu, Wan, Yuanyu, Zhang, Lijun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization
by: Yang, Wenhao, et al.
Published: (2026)
by: Yang, Wenhao, et al.
Published: (2026)
Building Autonomous GUI Navigation via Agentic-Q Estimation and Step-Wise Policy Optimization
by: Wang, Yibo, et al.
Published: (2026)
by: Wang, Yibo, et al.
Published: (2026)
SPACE: Noise Contrastive Estimation Stabilizes Self-Play Fine-Tuning for Large Language Models
by: Wang, Yibo, et al.
Published: (2025)
by: Wang, Yibo, et al.
Published: (2025)
Ovis: Structural Embedding Alignment for Multimodal Large Language Model
by: Lu, Shiyin, et al.
Published: (2024)
by: Lu, Shiyin, et al.
Published: (2024)
MDP3: A Training-free Approach for List-wise Frame Selection in Video-LLMs
by: Sun, Hui, et al.
Published: (2025)
by: Sun, Hui, et al.
Published: (2025)
Advancing Tool-Augmented Large Language Models: Integrating Insights from Errors in Inference Trees
by: Chen, Sijia, et al.
Published: (2024)
by: Chen, Sijia, et al.
Published: (2024)
Projection-free Online Learning over Strongly Convex Sets
by: Wan, Yuanyu, et al.
Published: (2020)
by: Wan, Yuanyu, et al.
Published: (2020)
Approximate Multiplication of Sparse Matrices with Limited Space
by: Wan, Yuanyu, et al.
Published: (2020)
by: Wan, Yuanyu, et al.
Published: (2020)
Multimodal Tabular Reasoning with Privileged Structured Information
by: Jiang, Jun-Peng, et al.
Published: (2025)
by: Jiang, Jun-Peng, et al.
Published: (2025)
Triplets Better Than Pairs: Towards Stable and Effective Self-Play Fine-Tuning for LLMs
by: Wang, Yibo, et al.
Published: (2026)
by: Wang, Yibo, et al.
Published: (2026)
UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation
by: Duan, Lunhao, et al.
Published: (2024)
by: Duan, Lunhao, et al.
Published: (2024)
CHATS: Combining Human-Aligned Optimization and Test-Time Sampling for Text-to-Image Generation
by: Fu, Minghao, et al.
Published: (2025)
by: Fu, Minghao, et al.
Published: (2025)
Projection-Free Variance Reduction Methods for Stochastic Constrained Multi-Level Compositional Optimization
by: Jiang, Wei, et al.
Published: (2024)
by: Jiang, Wei, et al.
Published: (2024)
Improved Dynamic Regret for Online Frank-Wolfe
by: Wan, Yuanyu, et al.
Published: (2023)
by: Wan, Yuanyu, et al.
Published: (2023)
Revisiting Projection-Free Online Learning with Time-Varying Constraints
by: Wang, Yibo, et al.
Published: (2025)
by: Wang, Yibo, et al.
Published: (2025)
Online Nonsubmodular Optimization with Delayed Feedback in the Bandit Setting
by: Yang, Sifan, et al.
Published: (2025)
by: Yang, Sifan, et al.
Published: (2025)
Evaluating Image Caption via Cycle-consistent Text-to-Image Generation
by: Cui, Tianyu, et al.
Published: (2025)
by: Cui, Tianyu, et al.
Published: (2025)
Wings: Learning Multimodal LLMs without Text-only Forgetting
by: Zhang, Yi-Kai, et al.
Published: (2024)
by: Zhang, Yi-Kai, et al.
Published: (2024)
Sign-Based Optimizers Are Effective Under Heavy-Tailed Noise
by: Yu, Dingzhi, et al.
Published: (2026)
by: Yu, Dingzhi, et al.
Published: (2026)
Diffusion-SDPO: Safeguarded Direct Preference Optimization for Diffusion Models
by: Fu, Minghao, et al.
Published: (2025)
by: Fu, Minghao, et al.
Published: (2025)
TeEFusion: Blending Text Embeddings to Distill Classifier-Free Guidance
by: Fu, Minghao, et al.
Published: (2025)
by: Fu, Minghao, et al.
Published: (2025)
A State-Transition Framework for Efficient LLM Reasoning
by: Zhang, Liang, et al.
Published: (2026)
by: Zhang, Liang, et al.
Published: (2026)
MMCR: Advancing Visual Language Model in Multimodal Multi-Turn Contextual Reasoning
by: Yan, Dawei, et al.
Published: (2025)
by: Yan, Dawei, et al.
Published: (2025)
Improved Regret for Bandit Convex Optimization with Delayed Feedback
by: Wan, Yuanyu, et al.
Published: (2024)
by: Wan, Yuanyu, et al.
Published: (2024)
Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
by: Zhao, Shanshan, et al.
Published: (2025)
by: Zhao, Shanshan, et al.
Published: (2025)
Covert Multicast in UAV-Enabled Wireless Communication Systems With One-hop and Two-hop Strategies
by: Zhang, Wenhao, et al.
Published: (2024)
by: Zhang, Wenhao, et al.
Published: (2024)
Parrot: Multilingual Visual Instruction Tuning
by: Sun, Hai-Long, et al.
Published: (2024)
by: Sun, Hai-Long, et al.
Published: (2024)
Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images
by: Hu, JiaKui, et al.
Published: (2025)
by: Hu, JiaKui, et al.
Published: (2025)
ComfyUI-R1: Exploring Reasoning Models for Workflow Generation
by: Xu, Zhenran, et al.
Published: (2025)
by: Xu, Zhenran, et al.
Published: (2025)
Thinking by Doing: Building Efficient World Model Reasoning in LLMs via Multi-turn Interaction
by: Shu, Bao, et al.
Published: (2025)
by: Shu, Bao, et al.
Published: (2025)
Building Decision Making Models Through Language Model Regime
by: Zhang, Yu, et al.
Published: (2024)
by: Zhang, Yu, et al.
Published: (2024)
Improving Multi-turn Dialogue Consistency with Self-Recall Thinking
by: Pang, Renning, et al.
Published: (2026)
by: Pang, Renning, et al.
Published: (2026)
MANTA: Multi-turn Assessment for Nonhuman Thinking & Alignment
by: Lu, Allen, et al.
Published: (2026)
by: Lu, Allen, et al.
Published: (2026)
New Trends for Modern Machine Translation with Large Reasoning Models
by: Liu, Sinuo, et al.
Published: (2025)
by: Liu, Sinuo, et al.
Published: (2025)
Optimal and Efficient Algorithms for Decentralized Online Convex Optimization
by: Wan, Yuanyu, et al.
Published: (2024)
by: Wan, Yuanyu, et al.
Published: (2024)
Non-stationary Delayed Online Convex Optimization: From Full-information to Bandit Setting
by: Wan, Yuanyu, et al.
Published: (2023)
by: Wan, Yuanyu, et al.
Published: (2023)
Continuous Subspace Optimization for Continual Learning
by: Cheng, Quan, et al.
Published: (2025)
by: Cheng, Quan, et al.
Published: (2025)
(Perhaps) Beyond Human Translation: Harnessing Multi-Agent Collaboration for Translating Ultra-Long Literary Texts
by: Wu, Minghao, et al.
Published: (2024)
by: Wu, Minghao, et al.
Published: (2024)
DeepWideSearch: Benchmarking Depth and Width in Agentic Information Seeking
by: Lan, Tian, et al.
Published: (2025)
by: Lan, Tian, et al.
Published: (2025)
Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions
by: Zhao, Yu, et al.
Published: (2024)
by: Zhao, Yu, et al.
Published: (2024)
Similar Items
-
Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization
by: Yang, Wenhao, et al.
Published: (2026) -
Building Autonomous GUI Navigation via Agentic-Q Estimation and Step-Wise Policy Optimization
by: Wang, Yibo, et al.
Published: (2026) -
SPACE: Noise Contrastive Estimation Stabilizes Self-Play Fine-Tuning for Large Language Models
by: Wang, Yibo, et al.
Published: (2025) -
Ovis: Structural Embedding Alignment for Multimodal Large Language Model
by: Lu, Shiyin, et al.
Published: (2024) -
MDP3: A Training-free Approach for List-wise Frame Selection in Video-LLMs
by: Sun, Hui, et al.
Published: (2025)