ODKE+: Ontology-Guided Open-Domain Knowledge Extraction with LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908520115339264 |
|---|---|
| author | Khorshidi, Samira Nikfarjam, Azadeh Shankar, Suprita Sang, Yisi Govind, Yash Jang, Hyun Kasgari, Ali McClimans, Alexis Soliman, Mohamed Konda, Vishnu Fakhry, Ahmed Qi, Xiaoguang |
| author_facet | Khorshidi, Samira Nikfarjam, Azadeh Shankar, Suprita Sang, Yisi Govind, Yash Jang, Hyun Kasgari, Ali McClimans, Alexis Soliman, Mohamed Konda, Vishnu Fakhry, Ahmed Qi, Xiaoguang |
| contents | Knowledge graphs (KGs) are foundational to many AI applications, but maintaining their freshness and completeness remains costly. We present ODKE+, a production-grade system that automatically extracts and ingests millions of open-domain facts from web sources with high precision. ODKE+ combines modular components into a scalable pipeline: (1) the Extraction Initiator detects missing or stale facts, (2) the Evidence Retriever collects supporting documents, (3) hybrid Knowledge Extractors apply both pattern-based rules and ontology-guided prompting for large language models (LLMs), (4) a lightweight Grounder validates extracted facts using a second LLM, and (5) the Corroborator ranks and normalizes candidate facts for ingestion. ODKE+ dynamically generates ontology snippets tailored to each entity type to align extractions with schema constraints, enabling scalable, type-consistent fact extraction across 195 predicates. The system supports batch and streaming modes, processing over 9 million Wikipedia pages and ingesting 19 million high-confidence facts with 98.8% precision. ODKE+ significantly improves coverage over traditional methods, achieving up to 48% overlap with third-party KGs and reducing update lag by 50 days on average. Our deployment demonstrates that LLM-based extraction, grounded in ontological structure and verification workflows, can deliver trustworthiness, production-scale knowledge ingestion with broad real-world applicability. A recording of the system demonstration is included with the submission and is also available at https://youtu.be/UcnE3_GsTWs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_04696 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | ODKE+: Ontology-Guided Open-Domain Knowledge Extraction with LLMs Khorshidi, Samira Nikfarjam, Azadeh Shankar, Suprita Sang, Yisi Govind, Yash Jang, Hyun Kasgari, Ali McClimans, Alexis Soliman, Mohamed Konda, Vishnu Fakhry, Ahmed Qi, Xiaoguang Computation and Language Artificial Intelligence Knowledge graphs (KGs) are foundational to many AI applications, but maintaining their freshness and completeness remains costly. We present ODKE+, a production-grade system that automatically extracts and ingests millions of open-domain facts from web sources with high precision. ODKE+ combines modular components into a scalable pipeline: (1) the Extraction Initiator detects missing or stale facts, (2) the Evidence Retriever collects supporting documents, (3) hybrid Knowledge Extractors apply both pattern-based rules and ontology-guided prompting for large language models (LLMs), (4) a lightweight Grounder validates extracted facts using a second LLM, and (5) the Corroborator ranks and normalizes candidate facts for ingestion. ODKE+ dynamically generates ontology snippets tailored to each entity type to align extractions with schema constraints, enabling scalable, type-consistent fact extraction across 195 predicates. The system supports batch and streaming modes, processing over 9 million Wikipedia pages and ingesting 19 million high-confidence facts with 98.8% precision. ODKE+ significantly improves coverage over traditional methods, achieving up to 48% overlap with third-party KGs and reducing update lag by 50 days on average. Our deployment demonstrates that LLM-based extraction, grounded in ontological structure and verification workflows, can deliver trustworthiness, production-scale knowledge ingestion with broad real-world applicability. A recording of the system demonstration is included with the submission and is also available at https://youtu.be/UcnE3_GsTWs. |
| title | ODKE+: Ontology-Guided Open-Domain Knowledge Extraction with LLMs |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2509.04696 |