Cross-Source Fusion Search for Heterogeneous AI Memory Systems
Fuente:
Zenodo
Salvato in:
| Autore principale: | |
|---|---|
| Natura: | Recurso digital |
| Pubblicazione: |
Zenodo
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866901623563878400 |
|---|---|
| author | Jia, Charles |
| author_facet | Jia, Charles |
| contents | <p><br> AI coding assistants maintain two distinct types of memory: semantic memory (code<br> indexes) and episodic memory (learned facts and decisions). These memories typically<br> reside in separate databases with different schemas, lifecycles, and write patterns.<br> Existing hybrid search approaches (RRF, DBSF, linear combination) assume a single data<br> source, while multi-memory systems (Mem0, MemGPT, Zep) either delegate source selection<br> to the LLM or return unranked concatenations.</p> <p> We present a cross-source fusion algorithm that exploits a shared embedding space to<br> produce unified rankings across heterogeneous databases. Our multiplicative scoring<br> formula addresses a previously undocumented vote asymmetry problem in additive<br> multi-source fusion, where independent per-source BM25 rankings inflate text-match<br> influence relative to a unified vector signal.</p> <p> Key contributions:<br> - Three-level search architecture separating per-source and cross-source fusion<br> - Multiplicative fusion formula eliminating vote asymmetry in multi-source BM25<br> - Graceful degradation to BM25-only mode when vector search is unavailable<br> - 7x latency reduction via parallel retrieval (620ms → 86ms)</p> <p> The approach requires no changes to existing per-source search pipelines.</p> <p> Project: https://github.com/gladego/index1</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_18556780 |
| institution | Zenodo |
| language | |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Cross-Source Fusion Search for Heterogeneous AI Memory Systems Jia, Charles hybrid search rank fusion <p><br> AI coding assistants maintain two distinct types of memory: semantic memory (code<br> indexes) and episodic memory (learned facts and decisions). These memories typically<br> reside in separate databases with different schemas, lifecycles, and write patterns.<br> Existing hybrid search approaches (RRF, DBSF, linear combination) assume a single data<br> source, while multi-memory systems (Mem0, MemGPT, Zep) either delegate source selection<br> to the LLM or return unranked concatenations.</p> <p> We present a cross-source fusion algorithm that exploits a shared embedding space to<br> produce unified rankings across heterogeneous databases. Our multiplicative scoring<br> formula addresses a previously undocumented vote asymmetry problem in additive<br> multi-source fusion, where independent per-source BM25 rankings inflate text-match<br> influence relative to a unified vector signal.</p> <p> Key contributions:<br> - Three-level search architecture separating per-source and cross-source fusion<br> - Multiplicative fusion formula eliminating vote asymmetry in multi-source BM25<br> - Graceful degradation to BM25-only mode when vector search is unavailable<br> - 7x latency reduction via parallel retrieval (620ms → 86ms)</p> <p> The approach requires no changes to existing per-source search pipelines.</p> <p> Project: https://github.com/gladego/index1</p> |
| title | Cross-Source Fusion Search for Heterogeneous AI Memory Systems |
| topic | hybrid search rank fusion |
| url | https://doi.org/10.5281/zenodo.18556780 |