SmartSearch: How Ranking Beats Structure for Conversational Memory Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Derehag, Jesper, Calva, Carlos, Ghiurau, Timmy
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917347807199232
author Derehag, Jesper
Calva, Carlos
Ghiurau, Timmy
author_facet Derehag, Jesper
Calva, Carlos
Ghiurau, Timmy
contents Recent conversational memory systems invest heavily in LLM-based structuring at ingestion time and learned retrieval policies at query time. We show that neither is necessary. SmartSearch retrieves from raw, unstructured conversation history using a fully deterministic pipeline: NER-weighted substring matching for recall, rule-based entity discovery for multi-hop expansion, and a CrossEncoder+ColBERT rank fusion stage -- the only learned component -- running on CPU in ~650ms. Oracle analysis on two benchmarks identifies a compilation bottleneck: retrieval recall reaches 98.6%, but without intelligent ranking only 22.5% of gold evidence survives truncation to the token budget. With score-adaptive truncation and no per-dataset tuning, SmartSearch achieves 93.5% on LoCoMo and 88.4% on LongMemEval-S, exceeding all known memory systems under the same evaluation protocol on both benchmarks while using 8.5x fewer tokens than full-context baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2603_15599
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SmartSearch: How Ranking Beats Structure for Conversational Memory Retrieval
Derehag, Jesper
Calva, Carlos
Ghiurau, Timmy
Machine Learning
Recent conversational memory systems invest heavily in LLM-based structuring at ingestion time and learned retrieval policies at query time. We show that neither is necessary. SmartSearch retrieves from raw, unstructured conversation history using a fully deterministic pipeline: NER-weighted substring matching for recall, rule-based entity discovery for multi-hop expansion, and a CrossEncoder+ColBERT rank fusion stage -- the only learned component -- running on CPU in ~650ms. Oracle analysis on two benchmarks identifies a compilation bottleneck: retrieval recall reaches 98.6%, but without intelligent ranking only 22.5% of gold evidence survives truncation to the token budget. With score-adaptive truncation and no per-dataset tuning, SmartSearch achieves 93.5% on LoCoMo and 88.4% on LongMemEval-S, exceeding all known memory systems under the same evaluation protocol on both benchmarks while using 8.5x fewer tokens than full-context baselines.
title SmartSearch: How Ranking Beats Structure for Conversational Memory Retrieval
topic Machine Learning
url https://arxiv.org/abs/2603.15599