Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Zhi, Zhang, Bofei, Li, Pengxiang, Ma, Xiaojian, Yuan, Tao, Fan, Yue, Wu, Yuwei, Jia, Yunde, Zhu, Song-Chun, Li, Qing |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning
by: Li, Pengxiang, et al.
Published: (2025)
by: Li, Pengxiang, et al.
Published: (2025)
FIRE: A Dataset for Feedback Integration and Refinement Evaluation of Multimodal Models
by: Li, Pengxiang, et al.
Published: (2024)
by: Li, Pengxiang, et al.
Published: (2024)
Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs
by: Zhang, Xintong, et al.
Published: (2025)
by: Zhang, Xintong, et al.
Published: (2025)
Building LLM Agents by Incorporating Insights from Computer Systems
by: Mi, Yapeng, et al.
Published: (2025)
by: Mi, Yapeng, et al.
Published: (2025)
CLOVA: A Closed-Loop Visual Assistant with Tool Usage and Update
by: Gao, Zhi, et al.
Published: (2023)
by: Gao, Zhi, et al.
Published: (2023)
TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents
by: Zhang, Bofei, et al.
Published: (2025)
by: Zhang, Bofei, et al.
Published: (2025)
Modality Alignment across Trees on Heterogeneous Hyperbolic Manifolds
by: Wu, Wei, et al.
Published: (2025)
by: Wu, Wei, et al.
Published: (2025)
Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation
by: Li, Pengxiang, et al.
Published: (2025)
by: Li, Pengxiang, et al.
Published: (2025)
VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding
by: Fan, Yue, et al.
Published: (2024)
by: Fan, Yue, et al.
Published: (2024)
Geometry-aware Distance Measure for Diverse Hierarchical Structures in Hyperbolic Spaces
by: Li, Pengxiang, et al.
Published: (2025)
by: Li, Pengxiang, et al.
Published: (2025)
Infant Agent: A Tool-Integrated, Logic-Driven Agent with Cost-Effective API Usage
by: Lei, Bin, et al.
Published: (2024)
by: Lei, Bin, et al.
Published: (2024)
Memory-Centric Embodied Question Answering
by: Zhai, Mingliang, et al.
Published: (2025)
by: Zhai, Mingliang, et al.
Published: (2025)
Large-Scale Riemannian Meta-Optimization via Subspace Adaptation
by: Yu, Peilin, et al.
Published: (2025)
by: Yu, Peilin, et al.
Published: (2025)
Curvature Learning for Generalization of Hyperbolic Neural Networks
by: Fan, Xiaomeng, et al.
Published: (2025)
by: Fan, Xiaomeng, et al.
Published: (2025)
A Set-to-Set Distance Measure in Hyperbolic Space
by: Li, Pengxiang, et al.
Published: (2025)
by: Li, Pengxiang, et al.
Published: (2025)
MIRROR: Multimodal Iterative Reasoning via Reflection on Visual Regions
by: Zhang, Haoyu, et al.
Published: (2026)
by: Zhang, Haoyu, et al.
Published: (2026)
Multi-Step Reasoning for Embodied Question Answering via Tool Augmentation
by: Zhai, Mingliang, et al.
Published: (2025)
by: Zhai, Mingliang, et al.
Published: (2025)
Hyperbolic Dual Feature Augmentation for Open-Environment
by: Yu, Peilin, et al.
Published: (2025)
by: Yu, Peilin, et al.
Published: (2025)
Adaptive Model Ensemble for Continual Learning
by: Mao, Yuchuan, et al.
Published: (2025)
by: Mao, Yuchuan, et al.
Published: (2025)
Beyond the Seen: Bounded Distribution Estimation for Open-Vocabulary Learning
by: Fan, Xiaomeng, et al.
Published: (2025)
by: Fan, Xiaomeng, et al.
Published: (2025)
PerTouch: VLM-Driven Agent for Personalized and Semantic Image Retouching
by: Chang, Zewei, et al.
Published: (2025)
by: Chang, Zewei, et al.
Published: (2025)
ToolTok: Tool Tokenization for Efficient and Generalizable GUI Agents
by: Wang, Xiaoce, et al.
Published: (2026)
by: Wang, Xiaoce, et al.
Published: (2026)
AutoTool: Efficient Tool Selection for Large Language Model Agents
by: Jia, Jingyi, et al.
Published: (2025)
by: Jia, Jingyi, et al.
Published: (2025)
An Embodied Generalist Agent in 3D World
by: Huang, Jiangyong, et al.
Published: (2023)
by: Huang, Jiangyong, et al.
Published: (2023)
Multi-Sourced Compositional Generalization in Visual Question Answering
by: Li, Chuanhao, et al.
Published: (2025)
by: Li, Chuanhao, et al.
Published: (2025)
Mind-of-Director: Multi-modal Agent-Driven Film Previsualization via Collaborative Decision-Making
by: Nan, Shufeng, et al.
Published: (2026)
by: Nan, Shufeng, et al.
Published: (2026)
Asynchronous Tool Usage for Real-Time Agents
by: Ginart, Antonio A., et al.
Published: (2024)
by: Ginart, Antonio A., et al.
Published: (2024)
Temporally Consistent Stereo Matching
by: Zeng, Jiaxi, et al.
Published: (2024)
by: Zeng, Jiaxi, et al.
Published: (2024)
MetaToolAgent: Towards Generalizable Tool Usage in LLMs through Meta-Learning
by: Fang, Zheng, et al.
Published: (2026)
by: Fang, Zheng, et al.
Published: (2026)
MMedAgent: Learning to Use Medical Tools with Multi-modal Agent
by: Li, Binxu, et al.
Published: (2024)
by: Li, Binxu, et al.
Published: (2024)
MalTool: Malicious Tool Attacks on LLM Agents
by: Hu, Yuepeng, et al.
Published: (2026)
by: Hu, Yuepeng, et al.
Published: (2026)
Evaluating Privilege Usage of Agents with Real-World Tools
by: Zhang, Quan, et al.
Published: (2026)
by: Zhang, Quan, et al.
Published: (2026)
CACTUS: Chemistry Agent Connecting Tool-Usage to Science
by: McNaughton, Andrew D., et al.
Published: (2024)
by: McNaughton, Andrew D., et al.
Published: (2024)
Consistency of Compositional Generalization across Multiple Levels
by: Li, Chuanhao, et al.
Published: (2024)
by: Li, Chuanhao, et al.
Published: (2024)
DriveAgent-R1: Advancing VLM-based Autonomous Driving with Active Perception and Hybrid Thinking
by: Zheng, Weicheng, et al.
Published: (2025)
by: Zheng, Weicheng, et al.
Published: (2025)
Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding
by: Fan, Yue, et al.
Published: (2024)
by: Fan, Yue, et al.
Published: (2024)
From Objects to Anywhere: A Holistic Benchmark for Multi-level Visual Grounding in 3D Scenes
by: Wang, Tianxu, et al.
Published: (2025)
by: Wang, Tianxu, et al.
Published: (2025)
Towards Efficient Online Tuning of VLM Agents via Counterfactual Soft Reinforcement Learning
by: Feng, Lang, et al.
Published: (2025)
by: Feng, Lang, et al.
Published: (2025)
CodeAgent: Enhancing Code Generation with Tool-Integrated Agent Systems for Real-World Repo-level Coding Challenges
by: Zhang, Kechi, et al.
Published: (2024)
by: Zhang, Kechi, et al.
Published: (2024)
Composition-Incremental Learning for Compositional Generalization
by: Li, Zhen, et al.
Published: (2025)
by: Li, Zhen, et al.
Published: (2025)
Similar Items
-
Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning
by: Li, Pengxiang, et al.
Published: (2025) -
FIRE: A Dataset for Feedback Integration and Refinement Evaluation of Multimodal Models
by: Li, Pengxiang, et al.
Published: (2024) -
Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs
by: Zhang, Xintong, et al.
Published: (2025) -
Building LLM Agents by Incorporating Insights from Computer Systems
by: Mi, Yapeng, et al.
Published: (2025) -
CLOVA: A Closed-Loop Visual Assistant with Tool Usage and Update
by: Gao, Zhi, et al.
Published: (2023)