FamilyTool: A Multi-hop Personalized Tool Use Benchmark

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Yuxin, Guo, Yiran, Zheng, Yining, Yin, Zhangyue, Chen, Shuo, Yang, Jie, Chen, Jiajun, Li, Yuan, Huang, Xuanjing, Qiu, Xipeng
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909622929981440
author Wang, Yuxin
Guo, Yiran
Zheng, Yining
Yin, Zhangyue
Chen, Shuo
Yang, Jie
Chen, Jiajun
Li, Yuan
Huang, Xuanjing
Qiu, Xipeng
author_facet Wang, Yuxin
Guo, Yiran
Zheng, Yining
Yin, Zhangyue
Chen, Shuo
Yang, Jie
Chen, Jiajun
Li, Yuan
Huang, Xuanjing
Qiu, Xipeng
contents The integration of tool learning with Large Language Models (LLMs) has expanded their capabilities in handling complex tasks by leveraging external tools. However, existing benchmarks for tool learning inadequately address critical real-world personalized scenarios, particularly those requiring multi-hop reasoning and inductive knowledge adaptation in dynamic environments. To bridge this gap, we introduce FamilyTool, a novel benchmark grounded in a family-based knowledge graph (KG) that simulates personalized, multi-hop tool use scenarios. FamilyTool, including base and extended datasets, challenges LLMs with queries spanning from 1 to 4 relational hops (e.g., inferring familial connections and preferences) and 2 to 6 hops respectively, and incorporates an inductive KG setting where models must adapt to unseen user preferences and relationships without re-training, a common limitation in prior approaches that compromises generalization. We further propose KGETool: a simple KG-augmented evaluation pipeline to systematically assess LLMs' tool use ability in these settings. Experiments reveal significant performance gaps in state-of-the-art LLMs, with accuracy dropping sharply as hop complexity increases and inductive scenarios exposing severe generalization deficits. These findings underscore the limitations of current LLMs in handling personalized, evolving real-world contexts and highlight the urgent need for advancements in tool-learning frameworks. FamilyTool serves as a critical resource for evaluating and advancing LLM agents' reasoning, adaptability, and scalability in complex, dynamic environments. Code and dataset are available at \href{https://github.com/yxzwang/FamilyTool}{https://github.com/yxzwang/FamilyTool}.
format Preprint
id arxiv_https___arxiv_org_abs_2504_06766
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FamilyTool: A Multi-hop Personalized Tool Use Benchmark
Wang, Yuxin
Guo, Yiran
Zheng, Yining
Yin, Zhangyue
Chen, Shuo
Yang, Jie
Chen, Jiajun
Li, Yuan
Huang, Xuanjing
Qiu, Xipeng
Artificial Intelligence
Computation and Language
The integration of tool learning with Large Language Models (LLMs) has expanded their capabilities in handling complex tasks by leveraging external tools. However, existing benchmarks for tool learning inadequately address critical real-world personalized scenarios, particularly those requiring multi-hop reasoning and inductive knowledge adaptation in dynamic environments. To bridge this gap, we introduce FamilyTool, a novel benchmark grounded in a family-based knowledge graph (KG) that simulates personalized, multi-hop tool use scenarios. FamilyTool, including base and extended datasets, challenges LLMs with queries spanning from 1 to 4 relational hops (e.g., inferring familial connections and preferences) and 2 to 6 hops respectively, and incorporates an inductive KG setting where models must adapt to unseen user preferences and relationships without re-training, a common limitation in prior approaches that compromises generalization. We further propose KGETool: a simple KG-augmented evaluation pipeline to systematically assess LLMs' tool use ability in these settings. Experiments reveal significant performance gaps in state-of-the-art LLMs, with accuracy dropping sharply as hop complexity increases and inductive scenarios exposing severe generalization deficits. These findings underscore the limitations of current LLMs in handling personalized, evolving real-world contexts and highlight the urgent need for advancements in tool-learning frameworks. FamilyTool serves as a critical resource for evaluating and advancing LLM agents' reasoning, adaptability, and scalability in complex, dynamic environments. Code and dataset are available at \href{https://github.com/yxzwang/FamilyTool}{https://github.com/yxzwang/FamilyTool}.
title FamilyTool: A Multi-hop Personalized Tool Use Benchmark
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2504.06766