Convomem Benchmark: Why Your First 150 Conversations Don't Need RAG
Fuente:
arXiv
Salvato in:
| Autori principali: | Pakhomov, Egor, Nijkamp, Erik, Xiong, Caiming |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
xGen-small Technical Report
di: Nijkamp, Erik, et al.
Pubblicazione: (2025)
di: Nijkamp, Erik, et al.
Pubblicazione: (2025)
Don't Do RAG: When Cache-Augmented Generation is All You Need for Knowledge Tasks
di: Chan, Brian J, et al.
Pubblicazione: (2024)
di: Chan, Brian J, et al.
Pubblicazione: (2024)
Don't Throw Away Your Pretrained Model
di: Feng, Shangbin, et al.
Pubblicazione: (2025)
di: Feng, Shangbin, et al.
Pubblicazione: (2025)
You Don't Need Pre-built Graphs for RAG: Retrieval Augmented Generation with Adaptive Reasoning Structures
di: Chen, Shengyuan, et al.
Pubblicazione: (2025)
di: Chen, Shengyuan, et al.
Pubblicazione: (2025)
Why Don't Prompt-Based Fairness Metrics Correlate?
di: Zayed, Abdelrahman, et al.
Pubblicazione: (2024)
di: Zayed, Abdelrahman, et al.
Pubblicazione: (2024)
sDPO: Don't Use Your Data All at Once
di: Kim, Dahyun, et al.
Pubblicazione: (2024)
di: Kim, Dahyun, et al.
Pubblicazione: (2024)
Hatevolution: What Static Benchmarks Don't Tell Us
di: Di Bonaventura, Chiara, et al.
Pubblicazione: (2025)
di: Di Bonaventura, Chiara, et al.
Pubblicazione: (2025)
Your Students Don't Use LLMs Like You Wish They Did
di: Kobler, Sebastian, et al.
Pubblicazione: (2026)
di: Kobler, Sebastian, et al.
Pubblicazione: (2026)
UNIDOC-BENCH: A Unified Benchmark for Document-Centric Multimodal RAG
di: Peng, Xiangyu, et al.
Pubblicazione: (2025)
di: Peng, Xiangyu, et al.
Pubblicazione: (2025)
Why Don't You Know? Evaluating the Impact of Uncertainty Sources on Uncertainty Quantification in LLMs
di: Goloburda, Maiya, et al.
Pubblicazione: (2026)
di: Goloburda, Maiya, et al.
Pubblicazione: (2026)
You Don't Need Prompt Engineering Anymore: The Prompting Inversion
di: Khan, Imran
Pubblicazione: (2025)
di: Khan, Imran
Pubblicazione: (2025)
Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency
di: Wang, Chenlong, et al.
Pubblicazione: (2025)
di: Wang, Chenlong, et al.
Pubblicazione: (2025)
Don't Forget to Connect! Improving RAG with Graph-based Reranking
di: Dong, Jialin, et al.
Pubblicazione: (2024)
di: Dong, Jialin, et al.
Pubblicazione: (2024)
Attention Is All You Need But You Don't Need All Of It For Inference of Large Language Models
di: Tyukin, Georgy, et al.
Pubblicazione: (2024)
di: Tyukin, Georgy, et al.
Pubblicazione: (2024)
Fine-Tune, Don't Prompt, Your Language Model to Identify Biased Language in Clinical Notes
di: Landi, Isotta, et al.
Pubblicazione: (2026)
di: Landi, Isotta, et al.
Pubblicazione: (2026)
Don't Touch My Diacritics
di: Gorman, Kyle, et al.
Pubblicazione: (2024)
di: Gorman, Kyle, et al.
Pubblicazione: (2024)
Don't Stop the Multi-Party! On Generating Synthetic Written Multi-Party Conversations with Constraints
di: Penzo, Nicolò, et al.
Pubblicazione: (2025)
di: Penzo, Nicolò, et al.
Pubblicazione: (2025)
Don't Pay Attention
di: Hammoud, Mohammad, et al.
Pubblicazione: (2025)
di: Hammoud, Mohammad, et al.
Pubblicazione: (2025)
The Compliance Gap: Why AI Systems Promise to Follow Process Instructions but Don't
di: Shin, Kwan Soo
Pubblicazione: (2026)
di: Shin, Kwan Soo
Pubblicazione: (2026)
"Don't Teach Minerva": Guiding LLMs Through Complex Syntax for Faithful Latin Translation with RAG
di: Aguilar, Sergio Torres
Pubblicazione: (2025)
di: Aguilar, Sergio Torres
Pubblicazione: (2025)
Summary of a Haystack: A Challenge to Long-Context LLMs and RAG Systems
di: Laban, Philippe, et al.
Pubblicazione: (2024)
di: Laban, Philippe, et al.
Pubblicazione: (2024)
Don't Forget Your Reward Values: Language Model Alignment via Value-based Calibration
di: Mao, Xin, et al.
Pubblicazione: (2024)
di: Mao, Xin, et al.
Pubblicazione: (2024)
If You Don't Understand It, Don't Use It: Eliminating Trojans with Filters Between Layers
di: Hernandez, Adriano
Pubblicazione: (2024)
di: Hernandez, Adriano
Pubblicazione: (2024)
Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users
di: Balepur, Nishant, et al.
Pubblicazione: (2026)
di: Balepur, Nishant, et al.
Pubblicazione: (2026)
s3: You Don't Need That Much Data to Train a Search Agent via RL
di: Jiang, Pengcheng, et al.
Pubblicazione: (2025)
di: Jiang, Pengcheng, et al.
Pubblicazione: (2025)
Don't Say No: Jailbreaking LLM by Suppressing Refusal
di: Zhou, Yukai, et al.
Pubblicazione: (2024)
di: Zhou, Yukai, et al.
Pubblicazione: (2024)
Don't Throw Away Your Beams: Improving Consistency-based Uncertainties in LLMs via Beam Search
di: Fadeeva, Ekaterina, et al.
Pubblicazione: (2025)
di: Fadeeva, Ekaterina, et al.
Pubblicazione: (2025)
Replace, Don't Expand: Mitigating Context Dilution in Multi-Hop RAG via Fixed-Budget Evidence Assembly
di: Lahmy, Moshe, et al.
Pubblicazione: (2025)
di: Lahmy, Moshe, et al.
Pubblicazione: (2025)
Don't Retrieve, Navigate: Distilling Enterprise Knowledge into Navigable Agent Skills for QA and RAG
di: Sun, Yiqun, et al.
Pubblicazione: (2026)
di: Sun, Yiqun, et al.
Pubblicazione: (2026)
Honest AI: Fine-Tuning "Small" Language Models to Say "I Don't Know", and Reducing Hallucination in RAG
di: Chen, Xinxi, et al.
Pubblicazione: (2024)
di: Chen, Xinxi, et al.
Pubblicazione: (2024)
Don't Trust Generative Agents to Mimic Communication on Social Networks Unless You Benchmarked their Empirical Realism
di: Münker, Simon, et al.
Pubblicazione: (2025)
di: Münker, Simon, et al.
Pubblicazione: (2025)
Do RAG Systems Cover What Matters? Evaluating and Optimizing Responses with Sub-Question Coverage
di: Xie, Kaige, et al.
Pubblicazione: (2024)
di: Xie, Kaige, et al.
Pubblicazione: (2024)
Think, But Don't Overthink: Reproducing Recursive Language Models
di: Wang, Daren
Pubblicazione: (2026)
di: Wang, Daren
Pubblicazione: (2026)
Knowing You Don't Know: Learning When to Continue Search in Multi-round RAG through Self-Practicing
di: Yang, Diji, et al.
Pubblicazione: (2025)
di: Yang, Diji, et al.
Pubblicazione: (2025)
Reasoning Models Reason Well, Until They Don't
di: Rameshkumar, Revanth, et al.
Pubblicazione: (2025)
di: Rameshkumar, Revanth, et al.
Pubblicazione: (2025)
Larger Language Models Don't Care How You Think: Why Chain-of-Thought Prompting Fails in Subjective Tasks
di: Chochlakis, Georgios, et al.
Pubblicazione: (2024)
di: Chochlakis, Georgios, et al.
Pubblicazione: (2024)
Don't Throw Away Data: Better Sequence Knowledge Distillation
di: Wang, Jun, et al.
Pubblicazione: (2024)
di: Wang, Jun, et al.
Pubblicazione: (2024)
Don't Command, Cultivate: An Exploratory Study of System-2 Alignment
di: Wang, Yuhang, et al.
Pubblicazione: (2024)
di: Wang, Yuhang, et al.
Pubblicazione: (2024)
MTRAG-UN: A Benchmark for Open Challenges in Multi-Turn RAG Conversations
di: Rosenthal, Sara, et al.
Pubblicazione: (2026)
di: Rosenthal, Sara, et al.
Pubblicazione: (2026)
Don't Walk the Line: Boundary Guidance for Filtered Generation
di: Ball, Sarah, et al.
Pubblicazione: (2025)
di: Ball, Sarah, et al.
Pubblicazione: (2025)
Documenti analoghi
-
xGen-small Technical Report
di: Nijkamp, Erik, et al.
Pubblicazione: (2025) -
Don't Do RAG: When Cache-Augmented Generation is All You Need for Knowledge Tasks
di: Chan, Brian J, et al.
Pubblicazione: (2024) -
Don't Throw Away Your Pretrained Model
di: Feng, Shangbin, et al.
Pubblicazione: (2025) -
You Don't Need Pre-built Graphs for RAG: Retrieval Augmented Generation with Adaptive Reasoning Structures
di: Chen, Shengyuan, et al.
Pubblicazione: (2025) -
Why Don't Prompt-Based Fairness Metrics Correlate?
di: Zayed, Abdelrahman, et al.
Pubblicazione: (2024)