RiTeK: A Dataset for Large Language Models Complex Reasoning over Textual Knowledge Graphs in Medicine
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866911585385054208 |
|---|---|
| author | Huang, Jiatan Li, Mingchen Yao, Zonghai Li, Dawei Zhang, Yuxin Yang, Zhichao Xiao, Yongkang Ouyang, Feiyun Li, Xiaohan Han, Shuo Yu, Hong |
| author_facet | Huang, Jiatan Li, Mingchen Yao, Zonghai Li, Dawei Zhang, Yuxin Yang, Zhichao Xiao, Yongkang Ouyang, Feiyun Li, Xiaohan Han, Shuo Yu, Hong |
| contents | Answering complex real-world questions in the medical domain often requires accurate retrieval from medical Textual Knowledge Graphs (medical TKGs), as the relational path information from TKGs could enhance the inference ability of Large Language Models (LLMs). However, the main bottlenecks lie in the scarcity of existing medical TKGs, the limited expressiveness of their topological structures, and the lack of comprehensive evaluations of current retrievers for medical TKGs. To address these challenges, we first develop a Dataset1 for LLMs Complex Reasoning over medical Textual Knowledge Graphs (RiTeK), covering a broad range of topological structures. Specifically, we synthesize realistic user queries integrating diverse topological structures, relational information, and complex textual descriptions. We conduct a rigorous medical expert evaluation process to assess and validate the quality of our synthesized queries. RiTeK also serves as a comprehensive benchmark dataset for evaluating the capabilities of retrieval systems built upon LLMs. By assessing 11 representative retrievers on this benchmark, we observe that existing methods struggle to perform well, revealing notable limitations in current LLM-driven retrieval approaches. These findings highlight the pressing need for more effective retrieval systems tailored for semi-structured data in the medical domain. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_13987 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | RiTeK: A Dataset for Large Language Models Complex Reasoning over Textual Knowledge Graphs in Medicine Huang, Jiatan Li, Mingchen Yao, Zonghai Li, Dawei Zhang, Yuxin Yang, Zhichao Xiao, Yongkang Ouyang, Feiyun Li, Xiaohan Han, Shuo Yu, Hong Computation and Language Answering complex real-world questions in the medical domain often requires accurate retrieval from medical Textual Knowledge Graphs (medical TKGs), as the relational path information from TKGs could enhance the inference ability of Large Language Models (LLMs). However, the main bottlenecks lie in the scarcity of existing medical TKGs, the limited expressiveness of their topological structures, and the lack of comprehensive evaluations of current retrievers for medical TKGs. To address these challenges, we first develop a Dataset1 for LLMs Complex Reasoning over medical Textual Knowledge Graphs (RiTeK), covering a broad range of topological structures. Specifically, we synthesize realistic user queries integrating diverse topological structures, relational information, and complex textual descriptions. We conduct a rigorous medical expert evaluation process to assess and validate the quality of our synthesized queries. RiTeK also serves as a comprehensive benchmark dataset for evaluating the capabilities of retrieval systems built upon LLMs. By assessing 11 representative retrievers on this benchmark, we observe that existing methods struggle to perform well, revealing notable limitations in current LLM-driven retrieval approaches. These findings highlight the pressing need for more effective retrieval systems tailored for semi-structured data in the medical domain. |
| title | RiTeK: A Dataset for Large Language Models Complex Reasoning over Textual Knowledge Graphs in Medicine |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2410.13987 |