The Hyperlexicon: A Multilingual, Sense-Centric Knowledge Graph with Hybrid Retrieval, Contextual Online Learning, and Business-Grade Interpretability

Fuente: Zenodo
Gespeichert in:
Bibliographische Detailangaben
1. Verfasser: Nelson, Benjamin
Format: Recurso digital
Sprache:Englisch
Veröffentlicht: Zenodo 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866902285354795008
author Nelson, Benjamin
author_facet Nelson, Benjamin
contents <p><strong>Introduction</strong></p> <p>The Hyperlexicon fuses symbolic and vector retrieval in a transparent, auditable framework. Built on a typed, weighted <br>lexical-semantic graph with full-text and graph-based indices, it delivers provenance-grounded answers and<br>interactive subgraphs rather than black-box outputs. It is multilingual (with cross-lingual operations), leverages<br>standard open lexical resources, runs on lightweight storage/indexes, and learns online from user feedback.<br>Every result is explainable: contributing edges, named heuristics (e.g., domain/register alignment, recency), and <br>re-ranking decisions are logged and traceable. Contextual online learning with smoothed updates adjusts weights on <br>feedback; a lightweight re-ranker balances interpretability with low latency. Expansion follows Active inference <br>(epistemic value vs. pragmatic performance).</p> <p> </p> <p><strong>What it is:</strong> A multilingual, sense-centric knowledge graph that blends full-text retrieval with a graph-based ANN vector<br>index and an interpretable reasoning layer.<br><strong>Why it matters:</strong> Delivers transparent, provenance-grounded answers with business-grade latency<br>targets (p50 ≤ 30 ms; p95 ≤ 80 ms) on commodity infrastructure.<br><strong>How it learns: </strong>Contextual online learning with smoothed updates; each user signal produces auditable ranking shifts.<br>Where it fits: Search, assistants, analytics, and compatible agent frameworks requiring explainability and low,<br>predictable latency.<br><strong>Cost advantage: </strong>Estimated 0.1%–1% of comparable LLM-only approaches; sub-cent per user at<br>internet scale (see §6).<br><strong>Risk controls: </strong>Tail-latency guardrails, named-policy logging, quarantine tags, and termination criteria for unbiased<br>stops.</p> <p><a title="Biomedical Use Case w/ BRCA1 Pathway and Node Schema" href="https://inferbenji.github.io/HyperLexicon_UseCase/">Biomedical Use Case w/ BRCA1 Pathway and Node Schema</a></p> <p><a title="Interactive Mobile Data Summary" href="https://inferbenji.github.io/HyperLexicon_UseCase/hyperlexicon_brief.html">Interactive Mobile Data Summary</a></p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_17137214
institution Zenodo
language eng
publishDate 2025
publisher Zenodo
record_format zenodo
spellingShingle The Hyperlexicon: A Multilingual, Sense-Centric Knowledge Graph with Hybrid Retrieval, Contextual Online Learning, and Business-Grade Interpretability
Nelson, Benjamin
lexical semantics, word sense disambiguation, hybrid retrieval, HNSW, FAISS, WordNet,  Wiktionary, interpretability, bandit learning, active inference, semantic search, knowledge graph.
<p><strong>Introduction</strong></p> <p>The Hyperlexicon fuses symbolic and vector retrieval in a transparent, auditable framework. Built on a typed, weighted <br>lexical-semantic graph with full-text and graph-based indices, it delivers provenance-grounded answers and<br>interactive subgraphs rather than black-box outputs. It is multilingual (with cross-lingual operations), leverages<br>standard open lexical resources, runs on lightweight storage/indexes, and learns online from user feedback.<br>Every result is explainable: contributing edges, named heuristics (e.g., domain/register alignment, recency), and <br>re-ranking decisions are logged and traceable. Contextual online learning with smoothed updates adjusts weights on <br>feedback; a lightweight re-ranker balances interpretability with low latency. Expansion follows Active inference <br>(epistemic value vs. pragmatic performance).</p> <p> </p> <p><strong>What it is:</strong> A multilingual, sense-centric knowledge graph that blends full-text retrieval with a graph-based ANN vector<br>index and an interpretable reasoning layer.<br><strong>Why it matters:</strong> Delivers transparent, provenance-grounded answers with business-grade latency<br>targets (p50 ≤ 30 ms; p95 ≤ 80 ms) on commodity infrastructure.<br><strong>How it learns: </strong>Contextual online learning with smoothed updates; each user signal produces auditable ranking shifts.<br>Where it fits: Search, assistants, analytics, and compatible agent frameworks requiring explainability and low,<br>predictable latency.<br><strong>Cost advantage: </strong>Estimated 0.1%–1% of comparable LLM-only approaches; sub-cent per user at<br>internet scale (see §6).<br><strong>Risk controls: </strong>Tail-latency guardrails, named-policy logging, quarantine tags, and termination criteria for unbiased<br>stops.</p> <p><a title="Biomedical Use Case w/ BRCA1 Pathway and Node Schema" href="https://inferbenji.github.io/HyperLexicon_UseCase/">Biomedical Use Case w/ BRCA1 Pathway and Node Schema</a></p> <p><a title="Interactive Mobile Data Summary" href="https://inferbenji.github.io/HyperLexicon_UseCase/hyperlexicon_brief.html">Interactive Mobile Data Summary</a></p>
title The Hyperlexicon: A Multilingual, Sense-Centric Knowledge Graph with Hybrid Retrieval, Contextual Online Learning, and Business-Grade Interpretability
topic lexical semantics, word sense disambiguation, hybrid retrieval, HNSW, FAISS, WordNet,  Wiktionary, interpretability, bandit learning, active inference, semantic search, knowledge graph.
url https://doi.org/10.5281/zenodo.17137214