To Memorize or to Retrieve: Scaling Laws for RAG-Considerate Pretraining
Fuente:
arXiv
Saved in:
| Main Authors: | Singh, Karan, Yu, Michael, Gangal, Varun, Tao, Zhuofu, Kumar, Sachin, Liu, Emmy, Feng, Steven Y. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models
by: Liu, Emmy, et al.
Published: (2026)
by: Liu, Emmy, et al.
Published: (2026)
A Unified Definition of Hallucination: It's The World Model, Stupid!
by: Liu, Emmy, et al.
Published: (2025)
by: Liu, Emmy, et al.
Published: (2025)
Enhancing Hallucination Detection through Perturbation-Based Synthetic Data Generation in System Responses
by: Zhang, Dongxu, et al.
Published: (2024)
by: Zhang, Dongxu, et al.
Published: (2024)
Memorization Dynamics of Fill-in-the-Middle Pretraining
by: von Arx, Tobias, et al.
Published: (2026)
by: von Arx, Tobias, et al.
Published: (2026)
MEMTRACK: Evaluating Long-Term Memory and State Tracking in Multi-Platform Dynamic Agent Environments
by: Deshpande, Darshan, et al.
Published: (2025)
by: Deshpande, Darshan, et al.
Published: (2025)
TRAIL: Trace Reasoning and Agentic Issue Localization
by: Deshpande, Darshan, et al.
Published: (2025)
by: Deshpande, Darshan, et al.
Published: (2025)
Unsupervised Memorability Modeling from Tip-of-the-Tongue Retrieval Queries
by: Bhattacharyya, Sree, et al.
Published: (2025)
by: Bhattacharyya, Sree, et al.
Published: (2025)
Not-Just-Scaling Laws: Towards a Better Understanding of the Downstream Impact of Language Model Design Decisions
by: Liu, Emmy, et al.
Published: (2025)
by: Liu, Emmy, et al.
Published: (2025)
Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG
by: Singh, Aditi, et al.
Published: (2025)
by: Singh, Aditi, et al.
Published: (2025)
Speculative RAG: Enhancing Retrieval Augmented Generation through Drafting
by: Wang, Zilong, et al.
Published: (2024)
by: Wang, Zilong, et al.
Published: (2024)
Know3-RAG: A Knowledge-aware RAG Framework with Adaptive Retrieval, Generation, and Filtering
by: Liu, Xukai, et al.
Published: (2025)
by: Liu, Xukai, et al.
Published: (2025)
Quantum-RAG and PunGPT2: Advancing Low-Resource Language Generation and Retrieval for the Punjabi Language
by: Singh, Jaskaranjeet, et al.
Published: (2025)
by: Singh, Jaskaranjeet, et al.
Published: (2025)
RAG+: Enhancing Retrieval-Augmented Generation with Application-Aware Reasoning
by: Wang, Yu, et al.
Published: (2025)
by: Wang, Yu, et al.
Published: (2025)
Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data
by: Wang, Xinyi, et al.
Published: (2024)
by: Wang, Xinyi, et al.
Published: (2024)
REIC: RAG-Enhanced Intent Classification at Scale
by: Zhang, Ziji, et al.
Published: (2025)
by: Zhang, Ziji, et al.
Published: (2025)
FB-RAG: Improving RAG with Forward and Backward Lookup
by: Chawla, Kushal, et al.
Published: (2025)
by: Chawla, Kushal, et al.
Published: (2025)
Memorization in In-Context Learning
by: Golchin, Shahriar, et al.
Published: (2024)
by: Golchin, Shahriar, et al.
Published: (2024)
RAG-Star: Enhancing Deliberative Reasoning with Retrieval Augmented Verification and Refinement
by: Jiang, Jinhao, et al.
Published: (2024)
by: Jiang, Jinhao, et al.
Published: (2024)
HM-RAG: Hierarchical Multi-Agent Multimodal Retrieval Augmented Generation
by: Liu, Pei, et al.
Published: (2025)
by: Liu, Pei, et al.
Published: (2025)
Fine-grained Claim-level RAG Benchmark for Law
by: Das, Souvick, et al.
Published: (2026)
by: Das, Souvick, et al.
Published: (2026)
Disco-RAG: Discourse-Aware Retrieval-Augmented Generation
by: Liu, Dongqi, et al.
Published: (2026)
by: Liu, Dongqi, et al.
Published: (2026)
Defenses & Enablers For Skill Injection Attacks on Terminal Based Agents
by: Fujinuma, Yoshinari, et al.
Published: (2026)
by: Fujinuma, Yoshinari, et al.
Published: (2026)
EasyRAG: Efficient Retrieval-Augmented Generation Framework for Automated Network Operations
by: Feng, Zhangchi, et al.
Published: (2024)
by: Feng, Zhangchi, et al.
Published: (2024)
MaskTab: Scalable Masked Tabular Pretraining with Scaling Laws and Distillation for Industrial Classification
by: Zheng, Bo, et al.
Published: (2026)
by: Zheng, Bo, et al.
Published: (2026)
ConflictRAG: Detecting and Resolving Knowledge Conflicts in Retrieval Augmented Generation
by: Wang, Chenyu, et al.
Published: (2026)
by: Wang, Chenyu, et al.
Published: (2026)
DuetRAG: Collaborative Retrieval-Augmented Generation
by: Jiao, Dian, et al.
Published: (2024)
by: Jiao, Dian, et al.
Published: (2024)
EfficientRAG: Efficient Retriever for Multi-Hop Question Answering
by: Zhuang, Ziyuan, et al.
Published: (2024)
by: Zhuang, Ziyuan, et al.
Published: (2024)
Scaling Retrieval Augmented Generation with RAG Fusion: Lessons from an Industry Deployment
by: Medrano, Luigi, et al.
Published: (2026)
by: Medrano, Luigi, et al.
Published: (2026)
BalanceRAG: Joint Risk Calibration for Cascaded Retrieval-Augmented Generation
by: Jia, Zijun, et al.
Published: (2026)
by: Jia, Zijun, et al.
Published: (2026)
Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies
by: Tao, Chaofan, et al.
Published: (2024)
by: Tao, Chaofan, et al.
Published: (2024)
Data Compressibility Quantifies LLM Memorization
by: Huang, Yizhan, et al.
Published: (2025)
by: Huang, Yizhan, et al.
Published: (2025)
The Scaling Laws of Skills in LLM Agent Systems
by: Chen, Charles, et al.
Published: (2026)
by: Chen, Charles, et al.
Published: (2026)
Memorization or Interpolation ? Detecting LLM Memorization through Input Perturbation Analysis
by: Djiré, Albérick Euraste, et al.
Published: (2025)
by: Djiré, Albérick Euraste, et al.
Published: (2025)
Meta-Tool: Efficient Few-Shot Tool Adaptation for Small Language Models
by: Kumar, Sachin
Published: (2026)
by: Kumar, Sachin
Published: (2026)
Bridge-RAG: An Abstract Bridge Tree Based Retrieval Augmented Generation Algorithm
by: Li, Zihang, et al.
Published: (2026)
by: Li, Zihang, et al.
Published: (2026)
From RAG to RICHES: Retrieval Interlaced with Sequence Generation
by: Jain, Palak, et al.
Published: (2024)
by: Jain, Palak, et al.
Published: (2024)
Beyond RAG for Agent Memory: Retrieval by Decoupling and Aggregation
by: Hu, Zhanghao, et al.
Published: (2026)
by: Hu, Zhanghao, et al.
Published: (2026)
AlpsBench: An LLM Personalization Benchmark for Real-Dialogue Memorization and Preference Alignment
by: Xiao, Jianfei, et al.
Published: (2026)
by: Xiao, Jianfei, et al.
Published: (2026)
Linguistic Indicators of Early Cognitive Decline in the DementiaBank Pitt Corpus: A Statistical and Machine Learning Study
by: Avetisyan, Artsvik, et al.
Published: (2026)
by: Avetisyan, Artsvik, et al.
Published: (2026)
Bringing Up a Bilingual BabyLM: Investigating Multilingual Language Acquisition Using Small-Scale Models
by: Zeng, Linda, et al.
Published: (2026)
by: Zeng, Linda, et al.
Published: (2026)
Similar Items
-
HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models
by: Liu, Emmy, et al.
Published: (2026) -
A Unified Definition of Hallucination: It's The World Model, Stupid!
by: Liu, Emmy, et al.
Published: (2025) -
Enhancing Hallucination Detection through Perturbation-Based Synthetic Data Generation in System Responses
by: Zhang, Dongxu, et al.
Published: (2024) -
Memorization Dynamics of Fill-in-the-Middle Pretraining
by: von Arx, Tobias, et al.
Published: (2026) -
MEMTRACK: Evaluating Long-Term Memory and State Tracking in Multi-Platform Dynamic Agent Environments
by: Deshpande, Darshan, et al.
Published: (2025)