Small Models, Big Insights: Leveraging Slim Proxy Models To Decide When and What to Retrieve for LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tan, Jiejun, Dou, Zhicheng, Zhu, Yutao, Guo, Peidong, Fang, Kun, Wen, Ji-Rong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MemSifter: Offloading LLM Memory Retrieval via Outcome-Driven Proxy Reasoning
von: Tan, Jiejun, et al.
Veröffentlicht: (2026)
von: Tan, Jiejun, et al.
Veröffentlicht: (2026)
HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG Systems
von: Tan, Jiejun, et al.
Veröffentlicht: (2024)
von: Tan, Jiejun, et al.
Veröffentlicht: (2024)
One Token Can Help! Learning Scalable and Pluggable Virtual Tokens for Retrieval-Augmented Large Language Models
von: Zhu, Yutao, et al.
Veröffentlicht: (2024)
von: Zhu, Yutao, et al.
Veröffentlicht: (2024)
OmniEval: An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial Domain
von: Wang, Shuting, et al.
Veröffentlicht: (2024)
von: Wang, Shuting, et al.
Veröffentlicht: (2024)
Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
Memory Matters More: Event-Centric Memory as a Logic Map for Agent Searching and Reasoning
von: Hu, Yuyang, et al.
Veröffentlicht: (2026)
von: Hu, Yuyang, et al.
Veröffentlicht: (2026)
DomainRAG: A Chinese Benchmark for Evaluating Domain-specific Retrieval-Augmented Generation
von: Wang, Shuting, et al.
Veröffentlicht: (2024)
von: Wang, Shuting, et al.
Veröffentlicht: (2024)
Progressive Multimodal Reasoning via Active Retrieval
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
Query-oriented Data Augmentation for Session Search
von: Chen, Haonan, et al.
Veröffentlicht: (2024)
von: Chen, Haonan, et al.
Veröffentlicht: (2024)
Toward General Instruction-Following Alignment for Retrieval-Augmented Generation
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
UFO: a Unified and Flexible Framework for Evaluating Factuality of Large Language Models
von: Huang, Zhaoheng, et al.
Veröffentlicht: (2024)
von: Huang, Zhaoheng, et al.
Veröffentlicht: (2024)
ChatShopBuddy: Towards Reliable Conversational Shopping Agents via Reinforcement Learning
von: Cheng, Yiruo, et al.
Veröffentlicht: (2026)
von: Cheng, Yiruo, et al.
Veröffentlicht: (2026)
BIDER: Bridging Knowledge Inconsistency for Efficient Retrieval-Augmented LLMs via Key Supporting Evidence
von: Jin, Jiajie, et al.
Veröffentlicht: (2024)
von: Jin, Jiajie, et al.
Veröffentlicht: (2024)
Single LLM, Multiple Roles: A Unified Retrieval-Augmented Generation Framework Using Role-Specific Token Optimization
von: Zhu, Yutao, et al.
Veröffentlicht: (2025)
von: Zhu, Yutao, et al.
Veröffentlicht: (2025)
R$^3$AG: Retriever Routing for Retrieval-Augmented Generation
von: Zhao, Tong, et al.
Veröffentlicht: (2026)
von: Zhao, Tong, et al.
Veröffentlicht: (2026)
DemoRank: Selecting Effective Demonstrations for Large Language Models in Ranking Task
von: Liu, Wenhan, et al.
Veröffentlicht: (2024)
von: Liu, Wenhan, et al.
Veröffentlicht: (2024)
Large Language Models for Information Retrieval: A Survey
von: Zhu, Yutao, et al.
Veröffentlicht: (2023)
von: Zhu, Yutao, et al.
Veröffentlicht: (2023)
HierSearch: A Hierarchical Enterprise Deep Search Framework Integrating Local and Web Searches
von: Tan, Jiejun, et al.
Veröffentlicht: (2025)
von: Tan, Jiejun, et al.
Veröffentlicht: (2025)
ATIR: Towards Audio-Text Interleaved Contextual Retrieval
von: Zhao, Tong, et al.
Veröffentlicht: (2026)
von: Zhao, Tong, et al.
Veröffentlicht: (2026)
EnvScaler: Scaling Tool-Interactive Environments for LLM Agent via Programmatic Synthesis
von: Song, Xiaoshuai, et al.
Veröffentlicht: (2026)
von: Song, Xiaoshuai, et al.
Veröffentlicht: (2026)
WebThinker: Empowering Large Reasoning Models with Deep Research Capability
von: Li, Xiaoxi, et al.
Veröffentlicht: (2025)
von: Li, Xiaoxi, et al.
Veröffentlicht: (2025)
INTERS: Unlocking the Power of Large Language Models in Search with Instruction Tuning
von: Zhu, Yutao, et al.
Veröffentlicht: (2024)
von: Zhu, Yutao, et al.
Veröffentlicht: (2024)
From Novice to Expert: LLM Agent Policy Optimization via Step-wise Reinforcement Learning
von: Deng, Zhirui, et al.
Veröffentlicht: (2024)
von: Deng, Zhirui, et al.
Veröffentlicht: (2024)
FlashRAG: A Modular Toolkit for Efficient Retrieval-Augmented Generation Research
von: Jin, Jiajie, et al.
Veröffentlicht: (2024)
von: Jin, Jiajie, et al.
Veröffentlicht: (2024)
DailyQA: A Benchmark to Evaluate Web Retrieval Augmented LLMs Based on Capturing Real-World Changes
von: Cheng, Jiehan, et al.
Veröffentlicht: (2025)
von: Cheng, Jiehan, et al.
Veröffentlicht: (2025)
Can Confidence Estimates Decide When Chain-of-Thought Is Necessary for LLMs?
von: Lewis-Lim, Samuel, et al.
Veröffentlicht: (2025)
von: Lewis-Lim, Samuel, et al.
Veröffentlicht: (2025)
RichRAG: Crafting Rich Responses for Multi-faceted Queries in Retrieval-Augmented Generation
von: Wang, Shuting, et al.
Veröffentlicht: (2024)
von: Wang, Shuting, et al.
Veröffentlicht: (2024)
Revisiting RAG Ensemble: A Theoretical and Mechanistic Analysis of Multi-RAG System Collaboration
von: Chen, Yifei, et al.
Veröffentlicht: (2025)
von: Chen, Yifei, et al.
Veröffentlicht: (2025)
Agentic-R: Learning to Retrieve for Agentic Search
von: Liu, Wenhan, et al.
Veröffentlicht: (2026)
von: Liu, Wenhan, et al.
Veröffentlicht: (2026)
CoRanking: Collaborative Ranking with Small and Large Ranking Agents
von: Liu, Wenhan, et al.
Veröffentlicht: (2025)
von: Liu, Wenhan, et al.
Veröffentlicht: (2025)
From Matching to Generation: A Survey on Generative Information Retrieval
von: Li, Xiaoxi, et al.
Veröffentlicht: (2024)
von: Li, Xiaoxi, et al.
Veröffentlicht: (2024)
LaSER: Internalizing Explicit Reasoning into Latent Space for Dense Retrieval
von: Jin, Jiajie, et al.
Veröffentlicht: (2026)
von: Jin, Jiajie, et al.
Veröffentlicht: (2026)
SlimLM: An Efficient Small Language Model for On-Device Document Assistance
von: Pham, Thang M., et al.
Veröffentlicht: (2024)
von: Pham, Thang M., et al.
Veröffentlicht: (2024)
LLMs + Persona-Plug = Personalized LLMs
von: Liu, Jiongnan, et al.
Veröffentlicht: (2024)
von: Liu, Jiongnan, et al.
Veröffentlicht: (2024)
PD-Loss: Proxy-Decidability for Efficient Metric Learning
von: Silva, Pedro, et al.
Veröffentlicht: (2025)
von: Silva, Pedro, et al.
Veröffentlicht: (2025)
An Analysis on Matching Mechanisms and Token Pruning for Late-interaction Models
von: Liu, Qi, et al.
Veröffentlicht: (2024)
von: Liu, Qi, et al.
Veröffentlicht: (2024)
Big Reasoning with Small Models: Instruction Retrieval at Inference Time
von: Alkiek, Kenan, et al.
Veröffentlicht: (2025)
von: Alkiek, Kenan, et al.
Veröffentlicht: (2025)
Memory-enhanced Retrieval Augmentation for Long Video Understanding
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
Majority Kernels: An Approach to Leverage Big Model Dynamics for Efficient Small Model Training
von: Mazzawi, Hanna, et al.
Veröffentlicht: (2024)
von: Mazzawi, Hanna, et al.
Veröffentlicht: (2024)
A Multi-Task Embedder For Retrieval Augmented LLMs
von: Zhang, Peitian, et al.
Veröffentlicht: (2023)
von: Zhang, Peitian, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
MemSifter: Offloading LLM Memory Retrieval via Outcome-Driven Proxy Reasoning
von: Tan, Jiejun, et al.
Veröffentlicht: (2026) -
HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG Systems
von: Tan, Jiejun, et al.
Veröffentlicht: (2024) -
One Token Can Help! Learning Scalable and Pluggable Virtual Tokens for Retrieval-Augmented Large Language Models
von: Zhu, Yutao, et al.
Veröffentlicht: (2024) -
OmniEval: An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial Domain
von: Wang, Shuting, et al.
Veröffentlicht: (2024) -
Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation
von: Dong, Guanting, et al.
Veröffentlicht: (2024)