Let your LLM generate a few tokens and you will reduce the need for retrieval
Fuente:
arXiv
Saved in:
| Main Author: | Déjean, Hervé |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
100 instances is all you need: predicting the success of a new LLM on unseen data by testing on a few instances
by: Pacchiardi, Lorenzo, et al.
Published: (2024)
by: Pacchiardi, Lorenzo, et al.
Published: (2024)
Getting the most out of your tokenizer for pre-training and domain adaptation
by: Dagan, Gautier, et al.
Published: (2024)
by: Dagan, Gautier, et al.
Published: (2024)
Retrieval-augmented generation in multilingual settings
by: Chirkova, Nadezhda, et al.
Published: (2024)
by: Chirkova, Nadezhda, et al.
Published: (2024)
Practical token pruning for foundation models in few-shot conversational virtual assistant systems
by: Qi, Haode, et al.
Published: (2024)
by: Qi, Haode, et al.
Published: (2024)
Retrieval-Augmented LLM Agents: Learning to Learn from Experience
by: Ferraz, Thomas Palmeira, et al.
Published: (2026)
by: Ferraz, Thomas Palmeira, et al.
Published: (2026)
PISCO: Pretty Simple Compression for Retrieval-Augmented Generation
by: Louis, Maxime, et al.
Published: (2025)
by: Louis, Maxime, et al.
Published: (2025)
Ambiguity is the last thing you need
by: Chivers, Emily, et al.
Published: (2024)
by: Chivers, Emily, et al.
Published: (2024)
Context Embeddings for Efficient Answer Generation in RAG
by: Rau, David, et al.
Published: (2024)
by: Rau, David, et al.
Published: (2024)
SPLADE-v3: New baselines for SPLADE
by: Lassance, Carlos, et al.
Published: (2024)
by: Lassance, Carlos, et al.
Published: (2024)
Downstream bias mitigation is all you need
by: Baksi, Arkadeep, et al.
Published: (2024)
by: Baksi, Arkadeep, et al.
Published: (2024)
Good Parenting is all you need -- Multi-agentic LLM Hallucination Mitigation
by: Kwartler, Ted, et al.
Published: (2024)
by: Kwartler, Ted, et al.
Published: (2024)
Collaboration is all you need: LLM Assisted Safe Code Translation
by: Karanjai, Rabimba, et al.
Published: (2025)
by: Karanjai, Rabimba, et al.
Published: (2025)
Reddit is all you need: Authorship profiling for Romanian
by: Ştefănescu, Ecaterina, et al.
Published: (2024)
by: Ştefănescu, Ecaterina, et al.
Published: (2024)
Unused information in token probability distribution of generative LLM: improving LLM reading comprehension through calculation of expected values
by: Zawistowski, Krystian
Published: (2024)
by: Zawistowski, Krystian
Published: (2024)
OSCAR: Online Soft Compression And Reranking
by: Louis, Maxime, et al.
Published: (2025)
by: Louis, Maxime, et al.
Published: (2025)
On multi-token prediction for efficient LLM inference
by: Mehra, Somesh, et al.
Published: (2025)
by: Mehra, Somesh, et al.
Published: (2025)
À la recherche du sens perdu: your favourite LLM might have more to say than you can understand
by: Erziev, K. O. T.
Published: (2025)
by: Erziev, K. O. T.
Published: (2025)
When retrieval outperforms generation: Dense evidence retrieval for scalable fake news detection
by: Qazi, Alamgir Munir, et al.
Published: (2025)
by: Qazi, Alamgir Munir, et al.
Published: (2025)
On the Challenges and Opportunities of Learned Sparse Retrieval for Code
by: Lupart, Simon, et al.
Published: (2026)
by: Lupart, Simon, et al.
Published: (2026)
Steps are all you need: Rethinking STEM Education with Prompt Engineering
by: Addala, Krishnasai, et al.
Published: (2024)
by: Addala, Krishnasai, et al.
Published: (2024)
Multimodal LLMs are not all you need for Pediatric Speech Language Pathology
by: Fürst, Darren, et al.
Published: (2026)
by: Fürst, Darren, et al.
Published: (2026)
Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
by: Nagarajan, Vaishnavh, et al.
Published: (2025)
by: Nagarajan, Vaishnavh, et al.
Published: (2025)
Knowledge Graphs are all you need: Leveraging KGs in Physics Question Answering
by: Addala, Krishnasai, et al.
Published: (2024)
by: Addala, Krishnasai, et al.
Published: (2024)
AUEB-Archimedes at RIRAG-2025: Is obligation concatenation really all you need?
by: Chasandras, Ioannis, et al.
Published: (2024)
by: Chasandras, Ioannis, et al.
Published: (2024)
You only need 4 extra tokens: Synergistic Test-time Adaptation for LLMs
by: Xu, Yijie, et al.
Published: (2025)
by: Xu, Yijie, et al.
Published: (2025)
Large Language Models aren't all that you need
by: Holla, Kiran Voderhobli, et al.
Published: (2024)
by: Holla, Kiran Voderhobli, et al.
Published: (2024)
Don't lie to your friends: Learning what you know from collaborative self-play
by: Eisenstein, Jacob, et al.
Published: (2025)
by: Eisenstein, Jacob, et al.
Published: (2025)
BSBench: will your LLM find the largest prime number?
by: Erziev, K. O. T.
Published: (2025)
by: Erziev, K. O. T.
Published: (2025)
Not all tokens are created equal: Perplexity Attention Weighted Networks for AI generated text detection
by: Miralles-González, Pablo, et al.
Published: (2025)
by: Miralles-González, Pablo, et al.
Published: (2025)
Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following
by: Ren, Qingyu, et al.
Published: (2025)
by: Ren, Qingyu, et al.
Published: (2025)
‛Until you're in the chair and executing your role, you don't know’: A qualitative study of the needs and perspectives of people with stroke‐related communication disabilities when returning to vocational activity
by: Lucette Lanyon, et al.
Published: (2024)
by: Lucette Lanyon, et al.
Published: (2024)
Camouflage is all you need: Evaluating and Enhancing Language Model Robustness Against Camouflage Adversarial Attacks
by: Huertas-García, Álvaro, et al.
Published: (2024)
by: Huertas-García, Álvaro, et al.
Published: (2024)
Why do LLMs attend to the first token?
by: Barbero, Federico, et al.
Published: (2025)
by: Barbero, Federico, et al.
Published: (2025)
Where is the signal in tokenization space?
by: Geh, Renato Lui, et al.
Published: (2024)
by: Geh, Renato Lui, et al.
Published: (2024)
Solving morphological analogies: from retrieval to generation
by: Marquer, Esteban, et al.
Published: (2023)
by: Marquer, Esteban, et al.
Published: (2023)
Provence: efficient and robust context pruning for retrieval-augmented generation
by: Chirkova, Nadezhda, et al.
Published: (2025)
by: Chirkova, Nadezhda, et al.
Published: (2025)
Multilingual Natural Language Processing Model for Radiology Reports -- The Summary is all you need!
by: Lindo, Mariana, et al.
Published: (2023)
by: Lindo, Mariana, et al.
Published: (2023)
Cross-Modal Safety Alignment: Is textual unlearning all you need?
by: Chakraborty, Trishna, et al.
Published: (2024)
by: Chakraborty, Trishna, et al.
Published: (2024)
BERGEN: A Benchmarking Library for Retrieval-Augmented Generation
by: Rau, David, et al.
Published: (2024)
by: Rau, David, et al.
Published: (2024)
Diffusion LLM with Native Variable Generation Lengths: Let [EOS] Lead the Way
by: Yang, Yicun, et al.
Published: (2025)
by: Yang, Yicun, et al.
Published: (2025)
Similar Items
-
100 instances is all you need: predicting the success of a new LLM on unseen data by testing on a few instances
by: Pacchiardi, Lorenzo, et al.
Published: (2024) -
Getting the most out of your tokenizer for pre-training and domain adaptation
by: Dagan, Gautier, et al.
Published: (2024) -
Retrieval-augmented generation in multilingual settings
by: Chirkova, Nadezhda, et al.
Published: (2024) -
Practical token pruning for foundation models in few-shot conversational virtual assistant systems
by: Qi, Haode, et al.
Published: (2024) -
Retrieval-Augmented LLM Agents: Learning to Learn from Experience
by: Ferraz, Thomas Palmeira, et al.
Published: (2026)