A Library of LLM Intrinsics for Retrieval-Augmented Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Danilevsky, Marina, Greenewald, Kristjan, Gunasekara, Chulaka, Hanafi, Maeda, He, Lihong, Katsis, Yannis, Killamsetty, Krishnateja, Li, Yulong, Nandwani, Yatin, Popa, Lucian, Raghu, Dinesh, Reiss, Frederick, Shah, Vraj, Tran, Khoi-Nguyen, Zhu, Huaiyu, Lastras, Luis
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918097932255232
author Danilevsky, Marina
Greenewald, Kristjan
Gunasekara, Chulaka
Hanafi, Maeda
He, Lihong
Katsis, Yannis
Killamsetty, Krishnateja
Li, Yulong
Nandwani, Yatin
Popa, Lucian
Raghu, Dinesh
Reiss, Frederick
Shah, Vraj
Tran, Khoi-Nguyen
Zhu, Huaiyu
Lastras, Luis
author_facet Danilevsky, Marina
Greenewald, Kristjan
Gunasekara, Chulaka
Hanafi, Maeda
He, Lihong
Katsis, Yannis
Killamsetty, Krishnateja
Li, Yulong
Nandwani, Yatin
Popa, Lucian
Raghu, Dinesh
Reiss, Frederick
Shah, Vraj
Tran, Khoi-Nguyen
Zhu, Huaiyu
Lastras, Luis
contents In the developer community for large language models (LLMs), there is not yet a clean pattern analogous to a software library, to support very large scale collaboration. Even for the commonplace use case of Retrieval-Augmented Generation (RAG), it is not currently possible to write a RAG application against a well-defined set of APIs that are agreed upon by different LLM providers. Inspired by the idea of compiler intrinsics, we propose some elements of such a concept through introducing a library of LLM Intrinsics for RAG. An LLM intrinsic is defined as a capability that can be invoked through a well-defined API that is reasonably stable and independent of how the LLM intrinsic itself is implemented. The intrinsics in our library are released as LoRA adapters on HuggingFace, and through a software interface with clear structured input/output characteristics on top of vLLM as an inference platform, accompanied in both places with documentation and code. This article describes the intended usage, training details, and evaluations for each intrinsic, as well as compositions of multiple intrinsics.
format Preprint
id arxiv_https___arxiv_org_abs_2504_11704
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Library of LLM Intrinsics for Retrieval-Augmented Generation
Danilevsky, Marina
Greenewald, Kristjan
Gunasekara, Chulaka
Hanafi, Maeda
He, Lihong
Katsis, Yannis
Killamsetty, Krishnateja
Li, Yulong
Nandwani, Yatin
Popa, Lucian
Raghu, Dinesh
Reiss, Frederick
Shah, Vraj
Tran, Khoi-Nguyen
Zhu, Huaiyu
Lastras, Luis
Artificial Intelligence
I.2.7
In the developer community for large language models (LLMs), there is not yet a clean pattern analogous to a software library, to support very large scale collaboration. Even for the commonplace use case of Retrieval-Augmented Generation (RAG), it is not currently possible to write a RAG application against a well-defined set of APIs that are agreed upon by different LLM providers. Inspired by the idea of compiler intrinsics, we propose some elements of such a concept through introducing a library of LLM Intrinsics for RAG. An LLM intrinsic is defined as a capability that can be invoked through a well-defined API that is reasonably stable and independent of how the LLM intrinsic itself is implemented. The intrinsics in our library are released as LoRA adapters on HuggingFace, and through a software interface with clear structured input/output characteristics on top of vLLM as an inference platform, accompanied in both places with documentation and code. This article describes the intended usage, training details, and evaluations for each intrinsic, as well as compositions of multiple intrinsics.
title A Library of LLM Intrinsics for Retrieval-Augmented Generation
topic Artificial Intelligence
I.2.7
url https://arxiv.org/abs/2504.11704