Semantic Caching for Improving Web Affordability

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Akbar, Hafsa, Athar, Danish, Rana, Muhammad Ayain Fida, Javed, Chaudhary Hammad, Uzmi, Zartash Afzal, Qazi, Ihsan Ayyub, Qazi, Zafar Ayyub
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909659955200000
author Akbar, Hafsa
Athar, Danish
Rana, Muhammad Ayain Fida
Javed, Chaudhary Hammad
Uzmi, Zartash Afzal
Qazi, Ihsan Ayyub
Qazi, Zafar Ayyub
author_facet Akbar, Hafsa
Athar, Danish
Rana, Muhammad Ayain Fida
Javed, Chaudhary Hammad
Uzmi, Zartash Afzal
Qazi, Ihsan Ayyub
Qazi, Zafar Ayyub
contents The rapid growth of web content has led to increasingly large webpages, posing significant challenges for Internet affordability, especially in developing countries where data costs remain prohibitively high. We propose semantic caching using Large Language Models (LLMs) to improve web affordability by enabling reuse of semantically similar images within webpages. Analyzing 50 leading news and media websites, encompassing 4,264 images and over 40,000 image pairs, we demonstrate potential for significant data transfer reduction, with some website categories showing up to 37% of images as replaceable. Our proof-of-concept architecture shows users can achieve approximately 10% greater byte savings compared to exact caching. We evaluate both commercial and open-source multi-modal LLMs for assessing semantic replaceability. GPT-4o performs best with a low Normalized Root Mean Square Error of 0.1735 and a weighted F1 score of 0.8374, while the open-source LLaMA 3.1 model shows comparable performance, highlighting its viability for large-scale applications. This approach offers benefits for both users and website operators, substantially reducing data transmission. We discuss ethical concerns and practical challenges, including semantic preservation, user-driven cache configuration, privacy concerns, and potential resistance from website operators
format Preprint
id arxiv_https___arxiv_org_abs_2506_20420
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Semantic Caching for Improving Web Affordability
Akbar, Hafsa
Athar, Danish
Rana, Muhammad Ayain Fida
Javed, Chaudhary Hammad
Uzmi, Zartash Afzal
Qazi, Ihsan Ayyub
Qazi, Zafar Ayyub
Networking and Internet Architecture
F.2.2, I.2.7
The rapid growth of web content has led to increasingly large webpages, posing significant challenges for Internet affordability, especially in developing countries where data costs remain prohibitively high. We propose semantic caching using Large Language Models (LLMs) to improve web affordability by enabling reuse of semantically similar images within webpages. Analyzing 50 leading news and media websites, encompassing 4,264 images and over 40,000 image pairs, we demonstrate potential for significant data transfer reduction, with some website categories showing up to 37% of images as replaceable. Our proof-of-concept architecture shows users can achieve approximately 10% greater byte savings compared to exact caching. We evaluate both commercial and open-source multi-modal LLMs for assessing semantic replaceability. GPT-4o performs best with a low Normalized Root Mean Square Error of 0.1735 and a weighted F1 score of 0.8374, while the open-source LLaMA 3.1 model shows comparable performance, highlighting its viability for large-scale applications. This approach offers benefits for both users and website operators, substantially reducing data transmission. We discuss ethical concerns and practical challenges, including semantic preservation, user-driven cache configuration, privacy concerns, and potential resistance from website operators
title Semantic Caching for Improving Web Affordability
topic Networking and Internet Architecture
F.2.2, I.2.7
url https://arxiv.org/abs/2506.20420