An Ensemble Embedding Approach for Improving Semantic Caching Performance in LLM-based Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ghaffari, Shervin, Bahranifard, Zohre, Akbari, Mohammad
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912473157730304
author Ghaffari, Shervin
Bahranifard, Zohre
Akbari, Mohammad
author_facet Ghaffari, Shervin
Bahranifard, Zohre
Akbari, Mohammad
contents Semantic caching enhances the efficiency of large language model (LLM) systems by identifying semantically similar queries, storing responses once, and serving them for subsequent equivalent requests. However, existing semantic caching frameworks rely on single embedding models for query representation, which limits their ability to capture the diverse semantic relationships present in real-world query distributions. This paper presents an ensemble embedding approach that combines multiple embedding models through a trained meta-encoder to improve semantic similarity detection in LLM caching systems. We evaluate our method using the Quora Question Pairs (QQP) dataset, measuring cache hit ratios, cache miss ratios, token savings, and response times. Our ensemble approach achieves a 92\% cache hit ratio for semantically equivalent queries while maintaining an 85\% accuracy in correctly rejecting non-equivalent queries as cache misses. These results demonstrate that ensemble embedding methods significantly outperform single-model approaches in distinguishing between semantically similar and dissimilar queries, leading to more effective caching performance and reduced computational overhead in LLM-based systems.
format Preprint
id arxiv_https___arxiv_org_abs_2507_07061
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An Ensemble Embedding Approach for Improving Semantic Caching Performance in LLM-based Systems
Ghaffari, Shervin
Bahranifard, Zohre
Akbari, Mohammad
Machine Learning
68T50
I.2.7; H.3.3; I.5.1
Semantic caching enhances the efficiency of large language model (LLM) systems by identifying semantically similar queries, storing responses once, and serving them for subsequent equivalent requests. However, existing semantic caching frameworks rely on single embedding models for query representation, which limits their ability to capture the diverse semantic relationships present in real-world query distributions. This paper presents an ensemble embedding approach that combines multiple embedding models through a trained meta-encoder to improve semantic similarity detection in LLM caching systems. We evaluate our method using the Quora Question Pairs (QQP) dataset, measuring cache hit ratios, cache miss ratios, token savings, and response times. Our ensemble approach achieves a 92\% cache hit ratio for semantically equivalent queries while maintaining an 85\% accuracy in correctly rejecting non-equivalent queries as cache misses. These results demonstrate that ensemble embedding methods significantly outperform single-model approaches in distinguishing between semantically similar and dissimilar queries, leading to more effective caching performance and reduced computational overhead in LLM-based systems.
title An Ensemble Embedding Approach for Improving Semantic Caching Performance in LLM-based Systems
topic Machine Learning
68T50
I.2.7; H.3.3; I.5.1
url https://arxiv.org/abs/2507.07061