Seal-Tools: Self-Instruct Tool Learning Dataset for Agent Tuning and Detailed Benchmark
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Mengsong, Zhu, Tong, Han, Han, Tan, Chuanyuan, Zhang, Xiang, Chen, Wenliang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
NesTools: A Dataset for Evaluating Nested Tool Learning Abilities of Large Language Models
von: Han, Han, et al.
Veröffentlicht: (2024)
von: Han, Han, et al.
Veröffentlicht: (2024)
Chain-of-Tools: Utilizing Massive Unseen Tools in the CoT Reasoning of Frozen Language Models
von: Wu, Mengsong, et al.
Veröffentlicht: (2025)
von: Wu, Mengsong, et al.
Veröffentlicht: (2025)
Learning to Refuse: Towards Mitigating Privacy Risks in LLMs
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024)
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024)
Is Fine-Tuning an Effective Solution? Reassessing Knowledge Editing for Unstructured Data
von: Xiong, Hao, et al.
Veröffentlicht: (2025)
von: Xiong, Hao, et al.
Veröffentlicht: (2025)
CheMatAgent: Enhancing LLMs for Chemistry and Materials Science through Tree-Search Based Tool Learning
von: Wu, Mengsong, et al.
Veröffentlicht: (2025)
von: Wu, Mengsong, et al.
Veröffentlicht: (2025)
SELT: Self-Evaluation Tree Search for LLMs with Task Decomposition
von: Wu, Mengsong, et al.
Veröffentlicht: (2025)
von: Wu, Mengsong, et al.
Veröffentlicht: (2025)
Probing Language Models for Pre-training Data Detection
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024)
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024)
UAQFact: Evaluating Factual Knowledge Utilization of LLMs on Unanswerable Questions
von: Tan, Chuanyuan, et al.
Veröffentlicht: (2025)
von: Tan, Chuanyuan, et al.
Veröffentlicht: (2025)
LiveMCPBench: Can Agents Navigate an Ocean of MCP Tools?
von: Mo, Guozhao, et al.
Veröffentlicht: (2025)
von: Mo, Guozhao, et al.
Veröffentlicht: (2025)
Learning to Instruct for Visual Instruction Tuning
von: Zhou, Zhihan, et al.
Veröffentlicht: (2025)
von: Zhou, Zhihan, et al.
Veröffentlicht: (2025)
GTA: A Benchmark for General Tool Agents
von: Wang, Jize, et al.
Veröffentlicht: (2024)
von: Wang, Jize, et al.
Veröffentlicht: (2024)
Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation
von: Qiao, Yuxuan, et al.
Veröffentlicht: (2025)
von: Qiao, Yuxuan, et al.
Veröffentlicht: (2025)
Tool Learning in the Wild: Empowering Language Models as Automatic Tool Agents
von: Shi, Zhengliang, et al.
Veröffentlicht: (2024)
von: Shi, Zhengliang, et al.
Veröffentlicht: (2024)
Semi-Instruct: Bridging Natural-Instruct and Self-Instruct for Code Large Language Models
von: Luo, Xianzhen, et al.
Veröffentlicht: (2024)
von: Luo, Xianzhen, et al.
Veröffentlicht: (2024)
AgentLongBench: A Controllable Long Benchmark For Long-Contexts Agents via Environment Rollouts
von: Fang, Shicheng, et al.
Veröffentlicht: (2026)
von: Fang, Shicheng, et al.
Veröffentlicht: (2026)
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents
von: Shen, Yujiong, et al.
Veröffentlicht: (2026)
von: Shen, Yujiong, et al.
Veröffentlicht: (2026)
Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback
von: Yan, Lecheng, et al.
Veröffentlicht: (2026)
von: Yan, Lecheng, et al.
Veröffentlicht: (2026)
Self-Tuning: Instructing LLMs to Effectively Acquire New Knowledge through Self-Teaching
von: Zhang, Xiaoying, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaoying, et al.
Veröffentlicht: (2024)
Controllable and Diverse Data Augmentation with Large Language Model for Low-Resource Open-Domain Dialogue Generation
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024)
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024)
The Tool Illusion: Rethinking Tool Use in Web Agents
von: Lou, Renze, et al.
Veröffentlicht: (2026)
von: Lou, Renze, et al.
Veröffentlicht: (2026)
FinSphere, a Real-Time Stock Analysis Agent Powered by Instruction-Tuned LLMs and Domain Tools
von: Han, Shijie, et al.
Veröffentlicht: (2025)
von: Han, Shijie, et al.
Veröffentlicht: (2025)
START: Self-taught Reasoner with Tools
von: Li, Chengpeng, et al.
Veröffentlicht: (2025)
von: Li, Chengpeng, et al.
Veröffentlicht: (2025)
GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows
von: Wang, Jize, et al.
Veröffentlicht: (2026)
von: Wang, Jize, et al.
Veröffentlicht: (2026)
ToolBridge: An Open-Source Dataset to Equip LLMs with External Tool Capabilities
von: Jin, Zhenchao, et al.
Veröffentlicht: (2024)
von: Jin, Zhenchao, et al.
Veröffentlicht: (2024)
What Affects the Stability of Tool Learning? An Empirical Study on the Robustness of Tool Learning Frameworks
von: Huang, Chengrui, et al.
Veröffentlicht: (2024)
von: Huang, Chengrui, et al.
Veröffentlicht: (2024)
R-Tuning: Instructing Large Language Models to Say `I Don't Know'
von: Zhang, Hanning, et al.
Veröffentlicht: (2023)
von: Zhang, Hanning, et al.
Veröffentlicht: (2023)
DS$^2$-Instruct: Domain-Specific Data Synthesis for Large Language Models Instruction Tuning
von: Xu, Ruiyao, et al.
Veröffentlicht: (2026)
von: Xu, Ruiyao, et al.
Veröffentlicht: (2026)
ToolGrad: Efficient Tool-use Dataset Generation with Textual "Gradients"
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
VipAct: Visual-Perception Enhancement via Specialized VLM Agent Collaboration and Tool-use
von: Zhang, Zhehao, et al.
Veröffentlicht: (2024)
von: Zhang, Zhehao, et al.
Veröffentlicht: (2024)
Dynamic Data Mixing Maximizes Instruction Tuning for Mixture-of-Experts
von: Zhu, Tong, et al.
Veröffentlicht: (2024)
von: Zhu, Tong, et al.
Veröffentlicht: (2024)
ToolACE-DEV: Self-Improving Tool Learning via Decomposition and EVolution
von: Huang, Xu, et al.
Veröffentlicht: (2025)
von: Huang, Xu, et al.
Veröffentlicht: (2025)
ToolHop: A Query-Driven Benchmark for Evaluating Large Language Models in Multi-Hop Tool Use
von: Ye, Junjie, et al.
Veröffentlicht: (2025)
von: Ye, Junjie, et al.
Veröffentlicht: (2025)
LuxInstruct: A Cross-Lingual Instruction Tuning Dataset For Luxembourgish
von: Philippy, Fred, et al.
Veröffentlicht: (2025)
von: Philippy, Fred, et al.
Veröffentlicht: (2025)
Tool-to-Agent Retrieval: Bridging Tools and Agents for Scalable LLM Multi-Agent Systems
von: Lumer, Elias, et al.
Veröffentlicht: (2025)
von: Lumer, Elias, et al.
Veröffentlicht: (2025)
AnyTool: Self-Reflective, Hierarchical Agents for Large-Scale API Calls
von: Du, Yu, et al.
Veröffentlicht: (2024)
von: Du, Yu, et al.
Veröffentlicht: (2024)
Learning to Use Tools via Cooperative and Interactive Agents
von: Shi, Zhengliang, et al.
Veröffentlicht: (2024)
von: Shi, Zhengliang, et al.
Veröffentlicht: (2024)
SciToolAgent: A Knowledge Graph-Driven Scientific Agent for Multi-Tool Integration
von: Ding, Keyan, et al.
Veröffentlicht: (2025)
von: Ding, Keyan, et al.
Veröffentlicht: (2025)
InverseCoder: Self-improving Instruction-Tuned Code LLMs with Inverse-Instruct
von: Wu, Yutong, et al.
Veröffentlicht: (2024)
von: Wu, Yutong, et al.
Veröffentlicht: (2024)
RIMRULE: Improving Tool-Using Language Agents via MDL-Guided Rule Learning
von: Gao, Xiang, et al.
Veröffentlicht: (2025)
von: Gao, Xiang, et al.
Veröffentlicht: (2025)
ToolExpander: Extending the Frontiers of Tool-Using Reinforcement Learning to Weak LLMs
von: Chen, Fu, et al.
Veröffentlicht: (2025)
von: Chen, Fu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
NesTools: A Dataset for Evaluating Nested Tool Learning Abilities of Large Language Models
von: Han, Han, et al.
Veröffentlicht: (2024) -
Chain-of-Tools: Utilizing Massive Unseen Tools in the CoT Reasoning of Frozen Language Models
von: Wu, Mengsong, et al.
Veröffentlicht: (2025) -
Learning to Refuse: Towards Mitigating Privacy Risks in LLMs
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024) -
Is Fine-Tuning an Effective Solution? Reassessing Knowledge Editing for Unstructured Data
von: Xiong, Hao, et al.
Veröffentlicht: (2025) -
CheMatAgent: Enhancing LLMs for Chemistry and Materials Science through Tree-Search Based Tool Learning
von: Wu, Mengsong, et al.
Veröffentlicht: (2025)