Knowledge Extraction on Semi-Structured Content: Does It Remain Relevant for Question Answering in the Era of LLMs?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Kai, Huang, Yin, Mehra, Srishti, Kachuee, Mohammad, Chen, Xilun, Tao, Renjie, Lin, Zhaojiang, Jessee, Andrea, Shah, Nirav, Betty, Alex, Liu, Yue, Kumar, Anuj, Yih, Wen-tau, Dong, Xin Luna
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911183887400960
author Sun, Kai
Huang, Yin
Mehra, Srishti
Kachuee, Mohammad
Chen, Xilun
Tao, Renjie
Lin, Zhaojiang
Jessee, Andrea
Shah, Nirav
Betty, Alex
Liu, Yue
Kumar, Anuj
Yih, Wen-tau
Dong, Xin Luna
author_facet Sun, Kai
Huang, Yin
Mehra, Srishti
Kachuee, Mohammad
Chen, Xilun
Tao, Renjie
Lin, Zhaojiang
Jessee, Andrea
Shah, Nirav
Betty, Alex
Liu, Yue
Kumar, Anuj
Yih, Wen-tau
Dong, Xin Luna
contents The advent of Large Language Models (LLMs) has significantly advanced web-based Question Answering (QA) systems over semi-structured content, raising questions about the continued utility of knowledge extraction for question answering. This paper investigates the value of triple extraction in this new paradigm by extending an existing benchmark with knowledge extraction annotations and evaluating commercial and open-source LLMs of varying sizes. Our results show that web-scale knowledge extraction remains a challenging task for LLMs. Despite achieving high QA accuracy, LLMs can still benefit from knowledge extraction, through augmentation with extracted triples and multi-task learning. These findings provide insights into the evolving role of knowledge triple extraction in web-based QA and highlight strategies for maximizing LLM effectiveness across different model sizes and resource settings.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25107
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Knowledge Extraction on Semi-Structured Content: Does It Remain Relevant for Question Answering in the Era of LLMs?
Sun, Kai
Huang, Yin
Mehra, Srishti
Kachuee, Mohammad
Chen, Xilun
Tao, Renjie
Lin, Zhaojiang
Jessee, Andrea
Shah, Nirav
Betty, Alex
Liu, Yue
Kumar, Anuj
Yih, Wen-tau
Dong, Xin Luna
Computation and Language
The advent of Large Language Models (LLMs) has significantly advanced web-based Question Answering (QA) systems over semi-structured content, raising questions about the continued utility of knowledge extraction for question answering. This paper investigates the value of triple extraction in this new paradigm by extending an existing benchmark with knowledge extraction annotations and evaluating commercial and open-source LLMs of varying sizes. Our results show that web-scale knowledge extraction remains a challenging task for LLMs. Despite achieving high QA accuracy, LLMs can still benefit from knowledge extraction, through augmentation with extracted triples and multi-task learning. These findings provide insights into the evolving role of knowledge triple extraction in web-based QA and highlight strategies for maximizing LLM effectiveness across different model sizes and resource settings.
title Knowledge Extraction on Semi-Structured Content: Does It Remain Relevant for Question Answering in the Era of LLMs?
topic Computation and Language
url https://arxiv.org/abs/2509.25107