Information Retrieval in the Age of Generative AI: The RGB Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Garetto, Michele, Cornacchia, Alessandro, Galante, Franco, Leonardi, Emilio, Nordio, Alessandro, Tarable, Alberto
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911438855995392
author Garetto, Michele
Cornacchia, Alessandro
Galante, Franco
Leonardi, Emilio
Nordio, Alessandro
Tarable, Alberto
author_facet Garetto, Michele
Cornacchia, Alessandro
Galante, Franco
Leonardi, Emilio
Nordio, Alessandro
Tarable, Alberto
contents The advent of Large Language Models (LLMs) and generative AI is fundamentally transforming information retrieval and processing on the Internet, bringing both great potential and significant concerns regarding content authenticity and reliability. This paper presents a novel quantitative approach to shed light on the complex information dynamics arising from the growing use of generative AI tools. Despite their significant impact on the digital ecosystem, these dynamics remain largely uncharted and poorly understood. We propose a stochastic model to characterize the generation, indexing, and dissemination of information in response to new topics. This scenario particularly challenges current LLMs, which often rely on real-time Retrieval-Augmented Generation (RAG) techniques to overcome their static knowledge limitations. Our findings suggest that the rapid pace of generative AI adoption, combined with increasing user reliance, can outpace human verification, escalating the risk of inaccurate information proliferation across digital resources. An in-depth analysis of Stack Exchange data confirms that high-quality answers inevitably require substantial time and human effort to emerge. This underscores the considerable risks associated with generating persuasive text in response to new questions and highlights the critical need for responsible development and deployment of future generative AI tools.
format Preprint
id arxiv_https___arxiv_org_abs_2504_20610
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Information Retrieval in the Age of Generative AI: The RGB Model
Garetto, Michele
Cornacchia, Alessandro
Galante, Franco
Leonardi, Emilio
Nordio, Alessandro
Tarable, Alberto
Information Retrieval
Artificial Intelligence
Performance
The advent of Large Language Models (LLMs) and generative AI is fundamentally transforming information retrieval and processing on the Internet, bringing both great potential and significant concerns regarding content authenticity and reliability. This paper presents a novel quantitative approach to shed light on the complex information dynamics arising from the growing use of generative AI tools. Despite their significant impact on the digital ecosystem, these dynamics remain largely uncharted and poorly understood. We propose a stochastic model to characterize the generation, indexing, and dissemination of information in response to new topics. This scenario particularly challenges current LLMs, which often rely on real-time Retrieval-Augmented Generation (RAG) techniques to overcome their static knowledge limitations. Our findings suggest that the rapid pace of generative AI adoption, combined with increasing user reliance, can outpace human verification, escalating the risk of inaccurate information proliferation across digital resources. An in-depth analysis of Stack Exchange data confirms that high-quality answers inevitably require substantial time and human effort to emerge. This underscores the considerable risks associated with generating persuasive text in response to new questions and highlights the critical need for responsible development and deployment of future generative AI tools.
title Information Retrieval in the Age of Generative AI: The RGB Model
topic Information Retrieval
Artificial Intelligence
Performance
url https://arxiv.org/abs/2504.20610