CoVe: Training Interactive Tool-Use Agents via Constraint-Guided Verification
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Chen, Jinpeng, Gong, Cheng, Li, Hanbo, Liu, Ziru, Tian, Zichen, Fu, Xinyu, Wu, Shi, Zhang, Chenyang, Zhang, Wu, Zhang, Suiyun, Tu, Dandan, Liu, Rui |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
GHPO: Adaptive Guidance for Stable and Efficient LLM Reinforcement Learning
par: Liu, Ziru, et autres
Publié: (2025)
par: Liu, Ziru, et autres
Publié: (2025)
Self-Distilled Trajectory-Aware Boltzmann Modeling: Bridging the Training-Inference Discrepancy in Diffusion Language Models
par: Chen, Kecheng, et autres
Publié: (2026)
par: Chen, Kecheng, et autres
Publié: (2026)
Beyond Confidence: Adaptive and Coherent Decoding for Diffusion Language Models
par: Chen, Kecheng, et autres
Publié: (2025)
par: Chen, Kecheng, et autres
Publié: (2025)
Efficient Reasoning via Reward Model
par: Wang, Yuhao, et autres
Publié: (2025)
par: Wang, Yuhao, et autres
Publié: (2025)
VisionDirector: Vision-Language Guided Closed-Loop Refinement for Generative Image Synthesis
par: Chu, Meng, et autres
Publié: (2025)
par: Chu, Meng, et autres
Publié: (2025)
Disentangling Instruction Influence in Diffusion Transformers for Parallel Multi-Instruction-Guided Image Editing
par: Liu, Hui, et autres
Publié: (2025)
par: Liu, Hui, et autres
Publié: (2025)
AutoTool: Automatic Scaling of Tool-Use Capabilities in RL via Decoupled Entropy Constraints
par: Zeng, Yirong, et autres
Publié: (2026)
par: Zeng, Yirong, et autres
Publié: (2026)
AEGIS: Exploring the Limit of World Knowledge Capabilities for Unified Mulitmodal Models
par: Lin, Jintao, et autres
Publié: (2026)
par: Lin, Jintao, et autres
Publié: (2026)
CCTU: A Benchmark for Tool Use under Complex Constraints
par: Ye, Junjie, et autres
Publié: (2026)
par: Ye, Junjie, et autres
Publié: (2026)
Training-Free Motion-Guided Video Generation with Enhanced Temporal Consistency Using Motion Consistency Loss
par: Zhang, Xinyu, et autres
Publié: (2025)
par: Zhang, Xinyu, et autres
Publié: (2025)
UI-KOBE: Knowledge-Oriented Behavior Exploration for Lightweight Graph-Guided GUI Agents
par: Chai, Yuxiang, et autres
Publié: (2026)
par: Chai, Yuxiang, et autres
Publié: (2026)
iTool: Reinforced Fine-Tuning with Dynamic Deficiency Calibration for Advanced Tool Use
par: Zeng, Yirong, et autres
Publié: (2025)
par: Zeng, Yirong, et autres
Publié: (2025)
VeCLIP: Improving CLIP Training via Visual-enriched Captions
par: Lai, Zhengfeng, et autres
Publié: (2023)
par: Lai, Zhengfeng, et autres
Publié: (2023)
VeRA: Verified Reasoning Data Augmentation at Scale
par: Cheng, Zerui, et autres
Publié: (2026)
par: Cheng, Zerui, et autres
Publié: (2026)
TopoCurate:Modeling Interaction Topology for Tool-Use Agent Training
par: Yang, Jinluan, et autres
Publié: (2026)
par: Yang, Jinluan, et autres
Publié: (2026)
The Evolution of Tool Use in LLM Agents: From Single-Tool Call to Multi-Tool Orchestration
par: Xu, Haoyuan, et autres
Publié: (2026)
par: Xu, Haoyuan, et autres
Publié: (2026)
VeCAF: Vision-language Collaborative Active Finetuning with Training Objective Awareness
par: Zhang, Rongyu, et autres
Publié: (2024)
par: Zhang, Rongyu, et autres
Publié: (2024)
ScaleEnv: Scaling Environment Synthesis from Scratch for Generalist Interactive Tool-Use Agent Training
par: Tu, Dunwei, et autres
Publié: (2026)
par: Tu, Dunwei, et autres
Publié: (2026)
VeRAPAk: Distributed Counterexample Generation via Iterative Neural Network Verification and Falsification
par: DenBleyker, Bennett, et autres
Publié: (2026)
par: DenBleyker, Bennett, et autres
Publié: (2026)
MMMamba: A Versatile Cross-Modal In Context Fusion Framework for Pan-Sharpening and Zero-Shot Image Enhancement
par: Wang, Yingying, et autres
Publié: (2025)
par: Wang, Yingying, et autres
Publié: (2025)
Hoi2Threat: An Interpretable Threat Detection Method for Human Violence Scenarios Guided by Human-Object Interaction
par: Wang, Yuhan, et autres
Publié: (2025)
par: Wang, Yuhan, et autres
Publié: (2025)
Guided by Trajectories: Repairing and Rewarding Tool-Use Trajectories for Tool-Integrated Reasoning
par: Gong, Siyu, et autres
Publié: (2026)
par: Gong, Siyu, et autres
Publié: (2026)
VeCoR -- Velocity Contrastive Regularization for Flow Matching
par: Hong, Zong-Wei, et autres
Publié: (2025)
par: Hong, Zong-Wei, et autres
Publié: (2025)
TOOLCAD: Exploring Tool-Using Large Language Models in Text-to-CAD Generation with Reinforcement Learning
par: Gong, Yifei, et autres
Publié: (2026)
par: Gong, Yifei, et autres
Publié: (2026)
Pusa V1.0: Unlocking Temporal Control in Pretrained Video Diffusion Models via Vectorized Timestep Adaptation
par: Liu, Yaofang, et autres
Publié: (2025)
par: Liu, Yaofang, et autres
Publié: (2025)
MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
par: Huang, Yue, et autres
Publié: (2023)
par: Huang, Yue, et autres
Publié: (2023)
KINDLE: Knowledge-Guided Distillation for Prior-Free Gene Regulatory Network Inference
par: Peng, Rui, et autres
Publié: (2025)
par: Peng, Rui, et autres
Publié: (2025)
LaCon: Late-Constraint Diffusion for Steerable Guided Image Synthesis
par: Liu, Chang, et autres
Publié: (2023)
par: Liu, Chang, et autres
Publié: (2023)
Process-Supervised Reinforcement Learning for Interactive Multimodal Tool-Use Agents
par: Tan, Weiting, et autres
Publié: (2025)
par: Tan, Weiting, et autres
Publié: (2025)
SInViG: A Self-Evolving Interactive Visual Agent for Human-Robot Interaction
par: Xu, Jie, et autres
Publié: (2024)
par: Xu, Jie, et autres
Publié: (2024)
Tool Zero: Training Tool-Augmented LLMs via Pure RL from Scratch
par: Zeng, Yirong, et autres
Publié: (2025)
par: Zeng, Yirong, et autres
Publié: (2025)
Schema-Aware Planning and Hybrid Knowledge Toolset for Reliable Knowledge Graph Triple Verification
par: Ma, Xinyan, et autres
Publié: (2026)
par: Ma, Xinyan, et autres
Publié: (2026)
MAS-ProVe: Understanding the Process Verification of Multi-Agent Systems
par: Venkataramani, Vishal, et autres
Publié: (2026)
par: Venkataramani, Vishal, et autres
Publié: (2026)
HyParLyVe: Hyperplane Partitioning for Neural Lyapunov Verification
par: Wayment, Jesse, et autres
Publié: (2026)
par: Wayment, Jesse, et autres
Publié: (2026)
MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents
par: Tao, Xijia, et autres
Publié: (2025)
par: Tao, Xijia, et autres
Publié: (2025)
GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces
par: Geng, Xinyu, et autres
Publié: (2026)
par: Geng, Xinyu, et autres
Publié: (2026)
A Unified Dual Consensus Approach to Distributed Optimization with Globally-Coupled Constraints
par: Liu, Zixuan, et autres
Publié: (2025)
par: Liu, Zixuan, et autres
Publié: (2025)
VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
par: Ma, Qianli, et autres
Publié: (2025)
par: Ma, Qianli, et autres
Publié: (2025)
Investigating training objective for flow matching-based speech enhancement
par: Yang, Liusha, et autres
Publié: (2025)
par: Yang, Liusha, et autres
Publié: (2025)
Causal-Aware Intelligent QoE Optimization for VR Interaction with Adaptive Keyframe Extraction
par: Zhang, Ziru, et autres
Publié: (2025)
par: Zhang, Ziru, et autres
Publié: (2025)
Documents similaires
-
GHPO: Adaptive Guidance for Stable and Efficient LLM Reinforcement Learning
par: Liu, Ziru, et autres
Publié: (2025) -
Self-Distilled Trajectory-Aware Boltzmann Modeling: Bridging the Training-Inference Discrepancy in Diffusion Language Models
par: Chen, Kecheng, et autres
Publié: (2026) -
Beyond Confidence: Adaptive and Coherent Decoding for Diffusion Language Models
par: Chen, Kecheng, et autres
Publié: (2025) -
Efficient Reasoning via Reward Model
par: Wang, Yuhao, et autres
Publié: (2025) -
VisionDirector: Vision-Language Guided Closed-Loop Refinement for Generative Image Synthesis
par: Chu, Meng, et autres
Publié: (2025)