Trajectory2Task: Training Robust Tool-Calling Agents with Synthesized Yet Verifiable Data for Complex User Intents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Ziyi, Lu, Yuxuan, Zhang, Yimeng, Chen, Pei, Dong, Ziwei, Huang, Jing, Gesi, Jiri, Tang, Xianfeng, Luo, Chen, Liu, Qun, Sang, Yisi, Lu, Hanqing, Li, Manling, Lai, Jin, Wang, Dakuo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Firefly: Illuminating Large-Scale Verified Tool-Call Data Generation from Real APIs
von: Lu, Yuxuan, et al.
Veröffentlicht: (2026)
von: Lu, Yuxuan, et al.
Veröffentlicht: (2026)
WEBSERV: A Full-Stack and RL-Ready Web Environment for Training Web Agents at Scale
von: Lu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Lu, Yuxuan, et al.
Veröffentlicht: (2025)
Customer-R1: Personalized Simulation of Human Behaviors via RL-based LLM Agent in Online Shopping
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
Can LLM Agents Simulate Multi-Turn Human Behavior? Evidence from Real Online Customer Behavior Data
von: Lu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Lu, Yuxuan, et al.
Veröffentlicht: (2025)
Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation
von: Chen, Jiaju, et al.
Veröffentlicht: (2025)
von: Chen, Jiaju, et al.
Veröffentlicht: (2025)
See, Think, Act: Online Shopper Behavior Simulation with VLM Agents
von: Zhang, Yimeng, et al.
Veröffentlicht: (2025)
von: Zhang, Yimeng, et al.
Veröffentlicht: (2025)
LLM Agent Meets Agentic AI: Can LLM Agents Simulate Customers to Evaluate Agentic-AI-based Shopping Assistants?
von: Sun, Lu, et al.
Veröffentlicht: (2025)
von: Sun, Lu, et al.
Veröffentlicht: (2025)
Adaptive Auto-Harness: Sustained Self-Improvement for Agentic System Deployment on Open-Ended Task Streams
von: Liu, Zewen, et al.
Veröffentlicht: (2026)
von: Liu, Zewen, et al.
Veröffentlicht: (2026)
UXAgent: A System for Simulating Usability Testing of Web Design with LLM Agents
von: Lu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Lu, Yuxuan, et al.
Veröffentlicht: (2025)
UXAgent: An LLM Agent-Based Usability Testing Framework for Web Design
von: Lu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Lu, Yuxuan, et al.
Veröffentlicht: (2025)
Beyond Self-learned Attention: Mitigating Attention Bias in Transformer-based Models Using Attention Guidance
von: Gesi, Jiri, et al.
Veröffentlicht: (2024)
von: Gesi, Jiri, et al.
Veröffentlicht: (2024)
More Samples or More Prompts? Exploring Effective In-Context Sampling for LLM Few-Shot Prompt Engineering
von: Yao, Bingsheng, et al.
Veröffentlicht: (2023)
von: Yao, Bingsheng, et al.
Veröffentlicht: (2023)
Shop-R1: Rewarding LLMs to Simulate Human Behavior in Online Shopping via Reinforcement Learning
von: Zhang, Yimeng, et al.
Veröffentlicht: (2025)
von: Zhang, Yimeng, et al.
Veröffentlicht: (2025)
IntentMiner: Intent Inversion Attack via Tool Call Analysis in the Model Context Protocol
von: Yao, Yunhao, et al.
Veröffentlicht: (2025)
von: Yao, Yunhao, et al.
Veröffentlicht: (2025)
Synthesizing High-Quality Visual Question Answering from Medical Documents with Generator-Verifier LMMs
von: Huang, Xiaoke, et al.
Veröffentlicht: (2025)
von: Huang, Xiaoke, et al.
Veröffentlicht: (2025)
TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use
von: He, Pengfei, et al.
Veröffentlicht: (2025)
von: He, Pengfei, et al.
Veröffentlicht: (2025)
OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
Alignment for Efficient Tool Calling of Large Language Models
von: Xu, Hongshen, et al.
Veröffentlicht: (2025)
von: Xu, Hongshen, et al.
Veröffentlicht: (2025)
DPRF: A Generalizable Dynamic Persona Refinement Framework for Optimizing Behavior Alignment Between Personalized LLM Role-Playing Agents and Humans
von: Yao, Bingsheng, et al.
Veröffentlicht: (2025)
von: Yao, Bingsheng, et al.
Veröffentlicht: (2025)
Does More Advice Help? The Effects of Second Opinions in AI-Assisted Decision Making
von: Lu, Zhuoran, et al.
Veröffentlicht: (2024)
von: Lu, Zhuoran, et al.
Veröffentlicht: (2024)
Mimic Intent, Not Just Trajectories
von: Huang, Renming, et al.
Veröffentlicht: (2026)
von: Huang, Renming, et al.
Veröffentlicht: (2026)
Emerging from Ground: Addressing Intent Deviation in Tool-Using Agents via Deriving Real Calls into Virtual Trajectories
von: Xiong, Qian, et al.
Veröffentlicht: (2026)
von: Xiong, Qian, et al.
Veröffentlicht: (2026)
Improving Generalization in Intent Detection: GRPO with Reward-Based Curriculum Sampling
von: Feng, Zihao, et al.
Veröffentlicht: (2025)
von: Feng, Zihao, et al.
Veröffentlicht: (2025)
Know Your Intent: An Autonomous Multi-Perspective LLM Agent Framework for DeFi User Transaction Intent Mining
von: Mao, Qian'ang, et al.
Veröffentlicht: (2025)
von: Mao, Qian'ang, et al.
Veröffentlicht: (2025)
Unlocking Implicit Experience: Synthesizing Tool-Use Trajectories from Text
von: Xu, Zhihao, et al.
Veröffentlicht: (2026)
von: Xu, Zhihao, et al.
Veröffentlicht: (2026)
Lean Finder: Semantic Search for Mathlib That Understands User Intents
von: Lu, Jialin, et al.
Veröffentlicht: (2025)
von: Lu, Jialin, et al.
Veröffentlicht: (2025)
Towards Robustness Analysis of E-Commerce Ranking System
von: Wang, Ningfei, et al.
Veröffentlicht: (2024)
von: Wang, Ningfei, et al.
Veröffentlicht: (2024)
ToolFlow: Boosting LLM Tool-Calling Through Natural and Coherent Dialogue Synthesis
von: Wang, Zezhong, et al.
Veröffentlicht: (2024)
von: Wang, Zezhong, et al.
Veröffentlicht: (2024)
TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories
von: Chen, Yen-Shan, et al.
Veröffentlicht: (2026)
von: Chen, Yen-Shan, et al.
Veröffentlicht: (2026)
Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
von: Lin, Minhua, et al.
Veröffentlicht: (2026)
von: Lin, Minhua, et al.
Veröffentlicht: (2026)
UniToolCall: Unifying Tool-Use Representation, Data, and Evaluation for LLM Agents
von: Liang, Yijuan, et al.
Veröffentlicht: (2026)
von: Liang, Yijuan, et al.
Veröffentlicht: (2026)
Self-Guided Defense: Adaptive Safety Alignment for Reasoning Models via Synthesized Guidelines
von: Wang, Yuhang, et al.
Veröffentlicht: (2025)
von: Wang, Yuhang, et al.
Veröffentlicht: (2025)
Who Tests the Testers? Systematic Enumeration and Coverage Audit of LLM Agent Tool Call Safety
von: Chen, Xuan, et al.
Veröffentlicht: (2026)
von: Chen, Xuan, et al.
Veröffentlicht: (2026)
ToolACE: Winning the Points of LLM Function Calling
von: Liu, Weiwen, et al.
Veröffentlicht: (2024)
von: Liu, Weiwen, et al.
Veröffentlicht: (2024)
A Population-to-individual Tuning Framework for Adapting Pretrained LM to On-device User Intent Prediction
von: Gong, Jiahui, et al.
Veröffentlicht: (2024)
von: Gong, Jiahui, et al.
Veröffentlicht: (2024)
Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers
von: Huang, Xingyue, et al.
Veröffentlicht: (2025)
von: Huang, Xingyue, et al.
Veröffentlicht: (2025)
ToolGate: Contract-Grounded and Verified Tool Execution for LLMs
von: Liu, Yanming, et al.
Veröffentlicht: (2026)
von: Liu, Yanming, et al.
Veröffentlicht: (2026)
Chain of Condition: Construct, Verify and Solve Conditions for Conditional Question Answering
von: Lin, Jiuheng, et al.
Veröffentlicht: (2024)
von: Lin, Jiuheng, et al.
Veröffentlicht: (2024)
Pragmatics of Formally Verified Yet Efficient Static Analysis, in particular for Formally Verified Compilers
von: Monniaux, David
Veröffentlicht: (2024)
von: Monniaux, David
Veröffentlicht: (2024)
The Dark Side of Human Feedback: Poisoning Large Language Models via User Inputs
von: Chen, Bocheng, et al.
Veröffentlicht: (2024)
von: Chen, Bocheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Firefly: Illuminating Large-Scale Verified Tool-Call Data Generation from Real APIs
von: Lu, Yuxuan, et al.
Veröffentlicht: (2026) -
WEBSERV: A Full-Stack and RL-Ready Web Environment for Training Web Agents at Scale
von: Lu, Yuxuan, et al.
Veröffentlicht: (2025) -
Customer-R1: Personalized Simulation of Human Behaviors via RL-based LLM Agent in Online Shopping
von: Wang, Ziyi, et al.
Veröffentlicht: (2025) -
Can LLM Agents Simulate Multi-Turn Human Behavior? Evidence from Real Online Customer Behavior Data
von: Lu, Yuxuan, et al.
Veröffentlicht: (2025) -
Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation
von: Chen, Jiaju, et al.
Veröffentlicht: (2025)