TOUCAN: Synthesizing 1.5M Tool-Agentic Data from Real-World MCP Environments
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Zhangchen, Soria, Adriana Meza, Tan, Shawn, Roy, Anurag, Agrawal, Ashish Sunil, Poovendran, Radha, Panda, Rameswar |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PRISM: Demystifying Retention and Interaction in Mid-Training
by: Runwal, Bharat, et al.
Published: (2026)
by: Runwal, Bharat, et al.
Published: (2026)
Brave: Byzantine-Resilient and Privacy-Preserving Peer-to-Peer Federated Learning
by: Xu, Zhangchen, et al.
Published: (2024)
by: Xu, Zhangchen, et al.
Published: (2024)
KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding
by: Xu, Zhangchen, et al.
Published: (2025)
by: Xu, Zhangchen, et al.
Published: (2025)
API Pack: A Massive Multi-Programming Language Dataset for API Call Generation
by: Guo, Zhen, et al.
Published: (2024)
by: Guo, Zhen, et al.
Published: (2024)
Scattered Mixture-of-Experts Implementation
by: Tan, Shawn, et al.
Published: (2024)
by: Tan, Shawn, et al.
Published: (2024)
Stronger Models are NOT Stronger Teachers for Instruction Tuning
by: Xu, Zhangchen, et al.
Published: (2024)
by: Xu, Zhangchen, et al.
Published: (2024)
ChatBug: A Common Vulnerability of Aligned LLMs Induced by Chat Templates
by: Jiang, Fengqing, et al.
Published: (2024)
by: Jiang, Fengqing, et al.
Published: (2024)
Steering Multimodal Large Language Models Decoding for Context-Aware Safety
by: Liu, Zheyuan, et al.
Published: (2025)
by: Liu, Zheyuan, et al.
Published: (2025)
Power Scheduler: A Batch Size and Token Number Agnostic Learning Rate Scheduler
by: Shen, Yikang, et al.
Published: (2024)
by: Shen, Yikang, et al.
Published: (2024)
ACE: A Model Poisoning Attack on Contribution Evaluation Methods in Federated Learning
by: Xu, Zhangchen, et al.
Published: (2024)
by: Xu, Zhangchen, et al.
Published: (2024)
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
by: Xu, Zhangchen, et al.
Published: (2024)
by: Xu, Zhangchen, et al.
Published: (2024)
Scaling Stick-Breaking Attention: An Efficient Implementation and In-depth Study
by: Tan, Shawn, et al.
Published: (2024)
by: Tan, Shawn, et al.
Published: (2024)
CleanGen: Mitigating Backdoor Attacks for Generation Tasks in Large Language Models
by: Li, Yuetai, et al.
Published: (2024)
by: Li, Yuetai, et al.
Published: (2024)
ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs
by: Jiang, Fengqing, et al.
Published: (2024)
by: Jiang, Fengqing, et al.
Published: (2024)
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning
by: Xu, Zhangchen, et al.
Published: (2025)
by: Xu, Zhangchen, et al.
Published: (2025)
Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing
by: Xu, Zhangchen, et al.
Published: (2024)
by: Xu, Zhangchen, et al.
Published: (2024)
MCP-Flow: Facilitating LLM Agents to Master Real-World, Diverse and Scaling MCP Tools
by: Wang, Wenhao, et al.
Published: (2025)
by: Wang, Wenhao, et al.
Published: (2025)
MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools
by: Guo, Zikang, et al.
Published: (2025)
by: Guo, Zikang, et al.
Published: (2025)
A cost minimization approach to fix the vocabulary size in a tokenizer for an End-to-End ASR system
by: Kopparapu, Sunil Kumar, et al.
Published: (2024)
by: Kopparapu, Sunil Kumar, et al.
Published: (2024)
SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities
by: Jiang, Fengqing, et al.
Published: (2025)
by: Jiang, Fengqing, et al.
Published: (2025)
Temporal Sampling for Forgotten Reasoning in LLMs
by: Li, Yuetai, et al.
Published: (2025)
by: Li, Yuetai, et al.
Published: (2025)
Small Models Struggle to Learn from Strong Reasoners
by: Li, Yuetai, et al.
Published: (2025)
by: Li, Yuetai, et al.
Published: (2025)
VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL
by: Feng, Yichen, et al.
Published: (2025)
by: Feng, Yichen, et al.
Published: (2025)
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
by: Wang, Zhenting, et al.
Published: (2025)
by: Wang, Zhenting, et al.
Published: (2025)
MCPTox: A Benchmark for Tool Poisoning Attack on Real-World MCP Servers
by: Wang, Zhiqiang, et al.
Published: (2025)
by: Wang, Zhiqiang, et al.
Published: (2025)
Distilling to Hybrid Attention Models via KL-Guided Layer Selection
by: Li, Yanhong, et al.
Published: (2025)
by: Li, Yanhong, et al.
Published: (2025)
Fault Tolerant Neural Control Barrier Functions for Robotic Systems under Sensor Faults and Attacks
by: Zhang, Hongchao, et al.
Published: (2024)
by: Zhang, Hongchao, et al.
Published: (2024)
Finding the Minimal Parameter Budget for Implicit Reasoning: A Data Complexity Driven Scaling Law for Language Models
by: Wang, Xinyi, et al.
Published: (2025)
by: Wang, Xinyi, et al.
Published: (2025)
MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments
by: Ganapavarapu, Giridhar, et al.
Published: (2026)
by: Ganapavarapu, Giridhar, et al.
Published: (2026)
MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers
by: Bandi, Chaithanya, et al.
Published: (2026)
by: Bandi, Chaithanya, et al.
Published: (2026)
PaTH Attention: Position Encoding via Accumulating Householder Transformations
by: Yang, Songlin, et al.
Published: (2025)
by: Yang, Songlin, et al.
Published: (2025)
MCP-SandboxScan: WASM-based Secure Execution and Runtime Analysis for MCP Tools
by: Tan, Zhuoran, et al.
Published: (2026)
by: Tan, Zhuoran, et al.
Published: (2026)
Scrutinizing Fermionic Dark Matter in Scotogenic Model with Low Reheating Temperature
by: Roy, Abhishek, et al.
Published: (2025)
by: Roy, Abhishek, et al.
Published: (2025)
MCP Pitfall Lab: Exposing Developer Pitfalls in MCP Tool Server Security under Multi-Vector Attacks
by: Hao, Run, et al.
Published: (2026)
by: Hao, Run, et al.
Published: (2026)
Spoken Grammar Assessment Using LLM
by: Kopparapu, Sunil Kumar, et al.
Published: (2024)
by: Kopparapu, Sunil Kumar, et al.
Published: (2024)
MCP-Persona: Benchmarking LLM Agents on Real-World Personal Applications via Environment Simulation
by: Wang, Wenhao, et al.
Published: (2026)
by: Wang, Wenhao, et al.
Published: (2026)
MCPVerse: An Expansive, Real-World Benchmark for Agentic Tool Use
by: Lei, Fei, et al.
Published: (2025)
by: Lei, Fei, et al.
Published: (2025)
Translation Errors Significantly Impact Low-Resource Languages in Cross-Lingual Learning
by: Agrawal, Ashish Sunil, et al.
Published: (2024)
by: Agrawal, Ashish Sunil, et al.
Published: (2024)
Boosting Zero-Shot Crosslingual Performance using LLM-Based Augmentations with Effective Data Selection
by: Fazili, Barah, et al.
Published: (2024)
by: Fazili, Barah, et al.
Published: (2024)
SoSBench: Benchmarking Safety Alignment on Six Scientific Domains
by: Jiang, Fengqing, et al.
Published: (2025)
by: Jiang, Fengqing, et al.
Published: (2025)
Similar Items
-
PRISM: Demystifying Retention and Interaction in Mid-Training
by: Runwal, Bharat, et al.
Published: (2026) -
Brave: Byzantine-Resilient and Privacy-Preserving Peer-to-Peer Federated Learning
by: Xu, Zhangchen, et al.
Published: (2024) -
KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding
by: Xu, Zhangchen, et al.
Published: (2025) -
API Pack: A Massive Multi-Programming Language Dataset for API Call Generation
by: Guo, Zhen, et al.
Published: (2024) -
Scattered Mixture-of-Experts Implementation
by: Tan, Shawn, et al.
Published: (2024)