TheMCPCompany: Creating General-purpose Agents with Task-specific Tools
Fuente:
arXiv
Salvato in:
| Autori principali: | Esfandiarpoor, Reza, Suryanarayanan, Vishwas, Bach, Stephen H., Chowdhary, Vishal, Aue, Anthony |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Follow-Up Differential Descriptions: Language Models Resolve Ambiguities for Image Classification
di: Esfandiarpoor, Reza, et al.
Pubblicazione: (2023)
di: Esfandiarpoor, Reza, et al.
Pubblicazione: (2023)
If CLIP Could Talk: Understanding Vision-Language Model Representations Through Their Preferred Concept Descriptions
di: Esfandiarpoor, Reza, et al.
Pubblicazione: (2024)
di: Esfandiarpoor, Reza, et al.
Pubblicazione: (2024)
Beyond Contrastive Learning: Synthetic Data Enables List-wise Training with Multiple Levels of Relevance
di: Esfandiarpoor, Reza, et al.
Pubblicazione: (2025)
di: Esfandiarpoor, Reza, et al.
Pubblicazione: (2025)
Local Prompt Optimization
di: Jain, Yash, et al.
Pubblicazione: (2025)
di: Jain, Yash, et al.
Pubblicazione: (2025)
VerificAgent: Domain-Specific Memory Verification for Scalable Oversight of Aligned Computer-Use Agents
di: Nguyen, Thong Q., et al.
Pubblicazione: (2025)
di: Nguyen, Thong Q., et al.
Pubblicazione: (2025)
Trove: A Flexible Toolkit for Dense Retrieval
di: Esfandiarpoor, Reza, et al.
Pubblicazione: (2025)
di: Esfandiarpoor, Reza, et al.
Pubblicazione: (2025)
GATE X-E : A Challenge Set for Gender-Fair Translations from Weakly-Gendered Languages
di: Rarrick, Spencer, et al.
Pubblicazione: (2024)
di: Rarrick, Spencer, et al.
Pubblicazione: (2024)
Learning to Generate Instruction Tuning Datasets for Zero-Shot Task Adaptation
di: Nayak, Nihal V., et al.
Pubblicazione: (2024)
di: Nayak, Nihal V., et al.
Pubblicazione: (2024)
$K$-MSHC: Unmasking Minimally Sufficient Head Circuits in Large Language Models with Experiments on Syntactic Classification Tasks
di: Chowdhary, Pratim, et al.
Pubblicazione: (2025)
di: Chowdhary, Pratim, et al.
Pubblicazione: (2025)
Multimedia Generative Script Learning for Task Planning
di: Wang, Qingyun, et al.
Pubblicazione: (2022)
di: Wang, Qingyun, et al.
Pubblicazione: (2022)
A General-purpose AI Avatar in Healthcare
di: Yan, Nicholas, et al.
Pubblicazione: (2024)
di: Yan, Nicholas, et al.
Pubblicazione: (2024)
CORE: A Conceptual Reasoning Layer for Large Language Models
di: Hegde, Vishwas, et al.
Pubblicazione: (2025)
di: Hegde, Vishwas, et al.
Pubblicazione: (2025)
Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Training
di: Yong, Zheng-Xin, et al.
Pubblicazione: (2025)
di: Yong, Zheng-Xin, et al.
Pubblicazione: (2025)
General-purpose Dataflow Model with Neuromorphic Primitives
di: Zhang, Weihao, et al.
Pubblicazione: (2024)
di: Zhang, Weihao, et al.
Pubblicazione: (2024)
AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents
di: Xie, Jingxu, et al.
Pubblicazione: (2025)
di: Xie, Jingxu, et al.
Pubblicazione: (2025)
Tool-to-Agent Retrieval: Bridging Tools and Agents for Scalable LLM Multi-Agent Systems
di: Lumer, Elias, et al.
Pubblicazione: (2025)
di: Lumer, Elias, et al.
Pubblicazione: (2025)
An Adaptive Method for Weak Supervision with Drifting Data
di: Mazzetto, Alessio, et al.
Pubblicazione: (2023)
di: Mazzetto, Alessio, et al.
Pubblicazione: (2023)
AssertBench: A Benchmark for Evaluating Self-Assertion in Large Language Models
di: Lee, Jaeho, et al.
Pubblicazione: (2025)
di: Lee, Jaeho, et al.
Pubblicazione: (2025)
LexC-Gen: Generating Data for Extremely Low-Resource Languages with Large Language Models and Bilingual Lexicons
di: Yong, Zheng-Xin, et al.
Pubblicazione: (2024)
di: Yong, Zheng-Xin, et al.
Pubblicazione: (2024)
TRACER: Verifiable Generative Provenance for Multimodal Tool-Using Agents
di: Yu, Bihui, et al.
Pubblicazione: (2026)
di: Yu, Bihui, et al.
Pubblicazione: (2026)
Doc2Agent: Scalable Generation of Tool-Using Agents from API Documentation
di: Ni, Xinyi, et al.
Pubblicazione: (2025)
di: Ni, Xinyi, et al.
Pubblicazione: (2025)
BAR: A Backward Reasoning based Agent for Complex Minecraft Tasks
di: Du, Weihong, et al.
Pubblicazione: (2025)
di: Du, Weihong, et al.
Pubblicazione: (2025)
GTA: A Benchmark for General Tool Agents
di: Wang, Jize, et al.
Pubblicazione: (2024)
di: Wang, Jize, et al.
Pubblicazione: (2024)
Development and Evaluation of a Retrieval-Augmented Generation Tool for Creating SAPPhIRE Models of Artificial Systems
di: Majumder, Anubhab, et al.
Pubblicazione: (2024)
di: Majumder, Anubhab, et al.
Pubblicazione: (2024)
Preference Tuning For Toxicity Mitigation Generalizes Across Languages
di: Li, Xiaochen, et al.
Pubblicazione: (2024)
di: Li, Xiaochen, et al.
Pubblicazione: (2024)
Revisiting Generalization Across Difficulty Levels: It's Not So Easy
di: Kordi, Yeganeh, et al.
Pubblicazione: (2025)
di: Kordi, Yeganeh, et al.
Pubblicazione: (2025)
Specializing General-purpose LLM Embeddings for Implicit Hate Speech Detection across Datasets
di: Cheremetiev, Vassiliy, et al.
Pubblicazione: (2025)
di: Cheremetiev, Vassiliy, et al.
Pubblicazione: (2025)
VeriAgent: A Tool-Integrated Multi-Agent System with Evolving Memory for PPA-Aware RTL Code Generation
di: Wang, Yaoxiang, et al.
Pubblicazione: (2026)
di: Wang, Yaoxiang, et al.
Pubblicazione: (2026)
The Tool Illusion: Rethinking Tool Use in Web Agents
di: Lou, Renze, et al.
Pubblicazione: (2026)
di: Lou, Renze, et al.
Pubblicazione: (2026)
From Failure to Mastery: Generating Hard Samples for Tool-use Agents
di: Hao, Bingguang, et al.
Pubblicazione: (2026)
di: Hao, Bingguang, et al.
Pubblicazione: (2026)
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
di: Wang, Zhenting, et al.
Pubblicazione: (2025)
di: Wang, Zhenting, et al.
Pubblicazione: (2025)
The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
di: Li, Junlong, et al.
Pubblicazione: (2025)
di: Li, Junlong, et al.
Pubblicazione: (2025)
TDAG: A Multi-Agent Framework based on Dynamic Task Decomposition and Agent Generation
di: Wang, Yaoxiang, et al.
Pubblicazione: (2024)
di: Wang, Yaoxiang, et al.
Pubblicazione: (2024)
Does It Run and Is That Enough? Revisiting Text-to-Chart Generation with a Multi-Agent Approach
di: Ford, James, et al.
Pubblicazione: (2025)
di: Ford, James, et al.
Pubblicazione: (2025)
Synergizing In-context Learning with Hints for End-to-end Task-oriented Dialog Systems
di: Saley, Vishal Vivek, et al.
Pubblicazione: (2024)
di: Saley, Vishal Vivek, et al.
Pubblicazione: (2024)
PLANTS: A Novel Problem and Dataset for Summarization of Planning-Like (PL) Tasks
di: Pallagani, Vishal, et al.
Pubblicazione: (2024)
di: Pallagani, Vishal, et al.
Pubblicazione: (2024)
Trajectory2Task: Training Robust Tool-Calling Agents with Synthesized Yet Verifiable Data for Complex User Intents
di: Wang, Ziyi, et al.
Pubblicazione: (2026)
di: Wang, Ziyi, et al.
Pubblicazione: (2026)
Using Generative Agents to Create Tip Sheets for Investigative Data Reporting
di: Veerbeek, Joris, et al.
Pubblicazione: (2024)
di: Veerbeek, Joris, et al.
Pubblicazione: (2024)
Pralekha: Cross-Lingual Document Alignment for Indic Languages
di: Suryanarayanan, Sanjay, et al.
Pubblicazione: (2024)
di: Suryanarayanan, Sanjay, et al.
Pubblicazione: (2024)
Toolshed: Scale Tool-Equipped Agents with Advanced RAG-Tool Fusion and Tool Knowledge Bases
di: Lumer, Elias, et al.
Pubblicazione: (2024)
di: Lumer, Elias, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Follow-Up Differential Descriptions: Language Models Resolve Ambiguities for Image Classification
di: Esfandiarpoor, Reza, et al.
Pubblicazione: (2023) -
If CLIP Could Talk: Understanding Vision-Language Model Representations Through Their Preferred Concept Descriptions
di: Esfandiarpoor, Reza, et al.
Pubblicazione: (2024) -
Beyond Contrastive Learning: Synthetic Data Enables List-wise Training with Multiple Levels of Relevance
di: Esfandiarpoor, Reza, et al.
Pubblicazione: (2025) -
Local Prompt Optimization
di: Jain, Yash, et al.
Pubblicazione: (2025) -
VerificAgent: Domain-Specific Memory Verification for Scalable Oversight of Aligned Computer-Use Agents
di: Nguyen, Thong Q., et al.
Pubblicazione: (2025)