Knowledge Extraction on Semi-Structured Content: Does It Remain Relevant for Question Answering in the Era of LLMs?
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911183887400960 |
|---|---|
| author | Sun, Kai Huang, Yin Mehra, Srishti Kachuee, Mohammad Chen, Xilun Tao, Renjie Lin, Zhaojiang Jessee, Andrea Shah, Nirav Betty, Alex Liu, Yue Kumar, Anuj Yih, Wen-tau Dong, Xin Luna |
| author_facet | Sun, Kai Huang, Yin Mehra, Srishti Kachuee, Mohammad Chen, Xilun Tao, Renjie Lin, Zhaojiang Jessee, Andrea Shah, Nirav Betty, Alex Liu, Yue Kumar, Anuj Yih, Wen-tau Dong, Xin Luna |
| contents | The advent of Large Language Models (LLMs) has significantly advanced web-based Question Answering (QA) systems over semi-structured content, raising questions about the continued utility of knowledge extraction for question answering. This paper investigates the value of triple extraction in this new paradigm by extending an existing benchmark with knowledge extraction annotations and evaluating commercial and open-source LLMs of varying sizes. Our results show that web-scale knowledge extraction remains a challenging task for LLMs. Despite achieving high QA accuracy, LLMs can still benefit from knowledge extraction, through augmentation with extracted triples and multi-task learning. These findings provide insights into the evolving role of knowledge triple extraction in web-based QA and highlight strategies for maximizing LLM effectiveness across different model sizes and resource settings. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_25107 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Knowledge Extraction on Semi-Structured Content: Does It Remain Relevant for Question Answering in the Era of LLMs? Sun, Kai Huang, Yin Mehra, Srishti Kachuee, Mohammad Chen, Xilun Tao, Renjie Lin, Zhaojiang Jessee, Andrea Shah, Nirav Betty, Alex Liu, Yue Kumar, Anuj Yih, Wen-tau Dong, Xin Luna Computation and Language The advent of Large Language Models (LLMs) has significantly advanced web-based Question Answering (QA) systems over semi-structured content, raising questions about the continued utility of knowledge extraction for question answering. This paper investigates the value of triple extraction in this new paradigm by extending an existing benchmark with knowledge extraction annotations and evaluating commercial and open-source LLMs of varying sizes. Our results show that web-scale knowledge extraction remains a challenging task for LLMs. Despite achieving high QA accuracy, LLMs can still benefit from knowledge extraction, through augmentation with extracted triples and multi-task learning. These findings provide insights into the evolving role of knowledge triple extraction in web-based QA and highlight strategies for maximizing LLM effectiveness across different model sizes and resource settings. |
| title | Knowledge Extraction on Semi-Structured Content: Does It Remain Relevant for Question Answering in the Era of LLMs? |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2509.25107 |