Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Pengxiang, Gao, Zhi, Zhang, Bofei, Mi, Yapeng, Ma, Xiaojian, Shi, Chenrui, Yuan, Tao, Wu, Yuwei, Jia, Yunde, Zhu, Song-Chun, Li, Qing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage
by: Gao, Zhi, et al.
Published: (2024)
by: Gao, Zhi, et al.
Published: (2024)
FIRE: A Dataset for Feedback Integration and Refinement Evaluation of Multimodal Models
by: Li, Pengxiang, et al.
Published: (2024)
by: Li, Pengxiang, et al.
Published: (2024)
Building LLM Agents by Incorporating Insights from Computer Systems
by: Mi, Yapeng, et al.
Published: (2025)
by: Mi, Yapeng, et al.
Published: (2025)
MIRROR: Multimodal Iterative Reasoning via Reflection on Visual Regions
by: Zhang, Haoyu, et al.
Published: (2026)
by: Zhang, Haoyu, et al.
Published: (2026)
Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs
by: Zhang, Xintong, et al.
Published: (2025)
by: Zhang, Xintong, et al.
Published: (2025)
Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation
by: Li, Pengxiang, et al.
Published: (2025)
by: Li, Pengxiang, et al.
Published: (2025)
CLOVA: A Closed-Loop Visual Assistant with Tool Usage and Update
by: Gao, Zhi, et al.
Published: (2023)
by: Gao, Zhi, et al.
Published: (2023)
Multi-Step Reasoning for Embodied Question Answering via Tool Augmentation
by: Zhai, Mingliang, et al.
Published: (2025)
by: Zhai, Mingliang, et al.
Published: (2025)
TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents
by: Zhang, Bofei, et al.
Published: (2025)
by: Zhang, Bofei, et al.
Published: (2025)
Modality Alignment across Trees on Heterogeneous Hyperbolic Manifolds
by: Wu, Wei, et al.
Published: (2025)
by: Wu, Wei, et al.
Published: (2025)
GUI Knowledge Bench: Revealing the Knowledge Gap of VLMs in GUI Tasks
by: Shi, Chenrui, et al.
Published: (2025)
by: Shi, Chenrui, et al.
Published: (2025)
Geometry-aware Distance Measure for Diverse Hierarchical Structures in Hyperbolic Spaces
by: Li, Pengxiang, et al.
Published: (2025)
by: Li, Pengxiang, et al.
Published: (2025)
Memory-Centric Embodied Question Answering
by: Zhai, Mingliang, et al.
Published: (2025)
by: Zhai, Mingliang, et al.
Published: (2025)
MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
by: Mi, Yapeng, et al.
Published: (2025)
by: Mi, Yapeng, et al.
Published: (2025)
Long-Horizon Visual Imitation Learning via Plan and Code Reflection
by: Chen, Quan, et al.
Published: (2025)
by: Chen, Quan, et al.
Published: (2025)
A Set-to-Set Distance Measure in Hyperbolic Space
by: Li, Pengxiang, et al.
Published: (2025)
by: Li, Pengxiang, et al.
Published: (2025)
Facial Expression Generation Aligned with Human Preference for Natural Dyadic Interaction
by: Chen, Xu, et al.
Published: (2026)
by: Chen, Xu, et al.
Published: (2026)
VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding
by: Fan, Yue, et al.
Published: (2024)
by: Fan, Yue, et al.
Published: (2024)
Large-Scale Riemannian Meta-Optimization via Subspace Adaptation
by: Yu, Peilin, et al.
Published: (2025)
by: Yu, Peilin, et al.
Published: (2025)
Curvature Learning for Generalization of Hyperbolic Neural Networks
by: Fan, Xiaomeng, et al.
Published: (2025)
by: Fan, Xiaomeng, et al.
Published: (2025)
Multi-Sourced Compositional Generalization in Visual Question Answering
by: Li, Chuanhao, et al.
Published: (2025)
by: Li, Chuanhao, et al.
Published: (2025)
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
by: Lai, Xin, et al.
Published: (2024)
by: Lai, Xin, et al.
Published: (2024)
AdaptMMBench: Benchmarking Adaptive Multimodal Reasoning for Mode Selection and Reasoning Process
by: Zhang, Xintong, et al.
Published: (2026)
by: Zhang, Xintong, et al.
Published: (2026)
Temporally Consistent Stereo Matching
by: Zeng, Jiaxi, et al.
Published: (2024)
by: Zeng, Jiaxi, et al.
Published: (2024)
Efficient Exploration for Iterative Nash Preference Optimization
by: Nan, Tianlong, et al.
Published: (2026)
by: Nan, Tianlong, et al.
Published: (2026)
Hyperbolic Dual Feature Augmentation for Open-Environment
by: Yu, Peilin, et al.
Published: (2025)
by: Yu, Peilin, et al.
Published: (2025)
Adaptive Model Ensemble for Continual Learning
by: Mao, Yuchuan, et al.
Published: (2025)
by: Mao, Yuchuan, et al.
Published: (2025)
Beyond the Seen: Bounded Distribution Estimation for Open-Vocabulary Learning
by: Fan, Xiaomeng, et al.
Published: (2025)
by: Fan, Xiaomeng, et al.
Published: (2025)
GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation
by: Xie, Rui, et al.
Published: (2026)
by: Xie, Rui, et al.
Published: (2026)
Composition-Incremental Learning for Compositional Generalization
by: Li, Zhen, et al.
Published: (2025)
by: Li, Zhen, et al.
Published: (2025)
MMKE-Bench: A Multimodal Editing Benchmark for Diverse Visual Knowledge
by: Du, Yuntao, et al.
Published: (2025)
by: Du, Yuntao, et al.
Published: (2025)
StepTool: Enhancing Multi-Step Tool Usage in LLMs via Step-Grained Reinforcement Learning
by: Yu, Yuanqing, et al.
Published: (2024)
by: Yu, Yuanqing, et al.
Published: (2024)
Infant Agent: A Tool-Integrated, Logic-Driven Agent with Cost-Effective API Usage
by: Lei, Bin, et al.
Published: (2024)
by: Lei, Bin, et al.
Published: (2024)
Full-Step-DPO: Self-Supervised Preference Optimization with Step-wise Rewards for Mathematical Reasoning
by: Xu, Huimin, et al.
Published: (2025)
by: Xu, Huimin, et al.
Published: (2025)
DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling
by: Zhang, Zhihong, et al.
Published: (2026)
by: Zhang, Zhihong, et al.
Published: (2026)
Residual Hyperbolic Graph Convolution Networks
by: Xue, Yangkai, et al.
Published: (2024)
by: Xue, Yangkai, et al.
Published: (2024)
Multi-Label Stereo Matching for Transparent Scene Depth Estimation
by: Liu, Zhidan, et al.
Published: (2025)
by: Liu, Zhidan, et al.
Published: (2025)
SwimVG: Step-wise Multimodal Fusion and Adaption for Visual Grounding
by: Shi, Liangtao, et al.
Published: (2025)
by: Shi, Liangtao, et al.
Published: (2025)
Step-wise Distribution Alignment Guided Style Prompt Tuning for Source-free Cross-domain Few-shot Learning
by: Xu, Huali, et al.
Published: (2024)
by: Xu, Huali, et al.
Published: (2024)
CodeAgent: Enhancing Code Generation with Tool-Integrated Agent Systems for Real-World Repo-level Coding Challenges
by: Zhang, Kechi, et al.
Published: (2024)
by: Zhang, Kechi, et al.
Published: (2024)
Similar Items
-
Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage
by: Gao, Zhi, et al.
Published: (2024) -
FIRE: A Dataset for Feedback Integration and Refinement Evaluation of Multimodal Models
by: Li, Pengxiang, et al.
Published: (2024) -
Building LLM Agents by Incorporating Insights from Computer Systems
by: Mi, Yapeng, et al.
Published: (2025) -
MIRROR: Multimodal Iterative Reasoning via Reflection on Visual Regions
by: Zhang, Haoyu, et al.
Published: (2026) -
Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs
by: Zhang, Xintong, et al.
Published: (2025)