The Great Nugget Recall: Automating Fact Extraction and RAG Evaluation with Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Pradeep, Ronak, Thakur, Nandan, Upadhyay, Shivani, Campos, Daniel, Craswell, Nick, Lin, Jimmy |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Initial Nugget Evaluation Results for the TREC 2024 RAG Track with the AutoNuggetizer Framework
di: Pradeep, Ronak, et al.
Pubblicazione: (2024)
di: Pradeep, Ronak, et al.
Pubblicazione: (2024)
Support Evaluation for the TREC 2024 RAG Track: Comparing Human versus LLM Judges
di: Thakur, Nandan, et al.
Pubblicazione: (2025)
di: Thakur, Nandan, et al.
Pubblicazione: (2025)
Overview of the TREC 2025 Retrieval Augmented Generation (RAG) Track
di: Upadhyay, Shivani, et al.
Pubblicazione: (2026)
di: Upadhyay, Shivani, et al.
Pubblicazione: (2026)
A Large-Scale Study of Relevance Assessments with Large Language Models: An Initial Look
di: Upadhyay, Shivani, et al.
Pubblicazione: (2024)
di: Upadhyay, Shivani, et al.
Pubblicazione: (2024)
UMBRELA: UMbrela is the (Open-Source Reproduction of the) Bing RELevance Assessor
di: Upadhyay, Shivani, et al.
Pubblicazione: (2024)
di: Upadhyay, Shivani, et al.
Pubblicazione: (2024)
Chatbot Arena Meets Nuggets: Towards Explanations and Diagnostics in the Evaluation of LLM Responses
di: Sharifymoghaddam, Sahel, et al.
Pubblicazione: (2025)
di: Sharifymoghaddam, Sahel, et al.
Pubblicazione: (2025)
Ragnarök: A Reusable RAG Framework and Baselines for TREC 2024 Retrieval-Augmented Generation Track
di: Pradeep, Ronak, et al.
Pubblicazione: (2024)
di: Pradeep, Ronak, et al.
Pubblicazione: (2024)
An Early FIRST Reproduction and Improvements to Single-Token Decoding for Fast Listwise Reranking
di: Chen, Zijian, et al.
Pubblicazione: (2024)
di: Chen, Zijian, et al.
Pubblicazione: (2024)
UniRAG: Universal Retrieval Augmentation for Large Vision Language Models
di: Sharifymoghaddam, Sahel, et al.
Pubblicazione: (2024)
di: Sharifymoghaddam, Sahel, et al.
Pubblicazione: (2024)
Overview of the TREC 2021 deep learning track
di: Craswell, Nick, et al.
Pubblicazione: (2025)
di: Craswell, Nick, et al.
Pubblicazione: (2025)
Still Fresh? Evaluating Temporal Drift in Retrieval Benchmarks
di: Kuissi, Nathan, et al.
Pubblicazione: (2026)
di: Kuissi, Nathan, et al.
Pubblicazione: (2026)
ORBIT: Scalable and Verifiable Data Generation for Search Agents on a Tight Budget
di: Thakur, Nandan, et al.
Pubblicazione: (2026)
di: Thakur, Nandan, et al.
Pubblicazione: (2026)
Hard Negatives, Hard Lessons: Revisiting Training Data Quality for Robust Information Retrieval with LLMs
di: Thakur, Nandan, et al.
Pubblicazione: (2025)
di: Thakur, Nandan, et al.
Pubblicazione: (2025)
LLMs Can Patch Up Missing Relevance Judgments in Evaluation
di: Upadhyay, Shivani, et al.
Pubblicazione: (2024)
di: Upadhyay, Shivani, et al.
Pubblicazione: (2024)
Overview of the TREC 2022 deep learning track
di: Craswell, Nick, et al.
Pubblicazione: (2025)
di: Craswell, Nick, et al.
Pubblicazione: (2025)
Leveraging LLMs for Synthesizing Training Data Across Many Languages in Multilingual Dense Retrieval
di: Thakur, Nandan, et al.
Pubblicazione: (2023)
di: Thakur, Nandan, et al.
Pubblicazione: (2023)
FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents
di: Thakur, Nandan, et al.
Pubblicazione: (2025)
di: Thakur, Nandan, et al.
Pubblicazione: (2025)
DoGMaTiQ: Automated Generation of Question-and-Answer Nuggets for Report Evaluation
di: Li, Bryan, et al.
Pubblicazione: (2026)
di: Li, Bryan, et al.
Pubblicazione: (2026)
ConvKGYarn: Spinning Configurable and Scalable Conversational Knowledge Graph QA datasets with Large Language Models
di: Pradeep, Ronak, et al.
Pubblicazione: (2024)
di: Pradeep, Ronak, et al.
Pubblicazione: (2024)
Overview of the TREC 2023 deep learning track
di: Craswell, Nick, et al.
Pubblicazione: (2025)
di: Craswell, Nick, et al.
Pubblicazione: (2025)
Loops On Retrieval Augmented Generation (LoRAG)
di: Thakur, Ayush, et al.
Pubblicazione: (2024)
di: Thakur, Ayush, et al.
Pubblicazione: (2024)
Large language models can accurately predict searcher preferences
di: Thomas, Paul, et al.
Pubblicazione: (2023)
di: Thomas, Paul, et al.
Pubblicazione: (2023)
MedNuggetizer: Confidence-Based Information Nugget Extraction from Medical Documents
di: Donabauer, Gregor, et al.
Pubblicazione: (2025)
di: Donabauer, Gregor, et al.
Pubblicazione: (2025)
GINGER: Grounded Information Nugget-Based Generation of Responses
di: Łajewska, Weronika, et al.
Pubblicazione: (2025)
di: Łajewska, Weronika, et al.
Pubblicazione: (2025)
NuggetIndex: Governed Atomic Retrieval for Maintainable RAG
di: Zerhoudi, Saber, et al.
Pubblicazione: (2026)
di: Zerhoudi, Saber, et al.
Pubblicazione: (2026)
Recall Them All: Retrieval-Augmented Language Models for Long Object List Extraction from Long Documents
di: Singhania, Sneha, et al.
Pubblicazione: (2024)
di: Singhania, Sneha, et al.
Pubblicazione: (2024)
Towards Unification of Hallucination Detection and Fact Verification for Large Language Models
di: Su, Weihang, et al.
Pubblicazione: (2025)
di: Su, Weihang, et al.
Pubblicazione: (2025)
Evaluating RAG-Fusion with RAGElo: an Automated Elo-based Framework
di: Rackauckas, Zackary, et al.
Pubblicazione: (2024)
di: Rackauckas, Zackary, et al.
Pubblicazione: (2024)
Synthetic Test Collections for Retrieval Evaluation
di: Rahmani, Hossein A., et al.
Pubblicazione: (2024)
di: Rahmani, Hossein A., et al.
Pubblicazione: (2024)
SciDef: Automating Definition Extraction from Academic Literature with Large Language Models
di: Kučera, Filip, et al.
Pubblicazione: (2026)
di: Kučera, Filip, et al.
Pubblicazione: (2026)
"Knowing When You Don't Know": A Multilingual Relevance Assessment Dataset for Robust Retrieval-Augmented Generation
di: Thakur, Nandan, et al.
Pubblicazione: (2023)
di: Thakur, Nandan, et al.
Pubblicazione: (2023)
Fact Finder -- Enhancing Domain Expertise of Large Language Models by Incorporating Knowledge Graphs
di: Steinigen, Daniel, et al.
Pubblicazione: (2024)
di: Steinigen, Daniel, et al.
Pubblicazione: (2024)
On the Comprehensibility of Multi-structured Financial Documents using LLMs and Pre-processing Tools
di: Upadhyay, Shivani, et al.
Pubblicazione: (2025)
di: Upadhyay, Shivani, et al.
Pubblicazione: (2025)
Resolving Conflicting Evidence in Automated Fact-Checking: A Study on Retrieval-Augmented LLMs
di: Ge, Ziyu, et al.
Pubblicazione: (2025)
di: Ge, Ziyu, et al.
Pubblicazione: (2025)
Microstructures and Accuracy of Graph Recall by Large Language Models
di: Wang, Yanbang, et al.
Pubblicazione: (2024)
di: Wang, Yanbang, et al.
Pubblicazione: (2024)
Systematic Evaluation of Neural Retrieval Models on the Touché 2020 Argument Retrieval Subset of BEIR
di: Thakur, Nandan, et al.
Pubblicazione: (2024)
di: Thakur, Nandan, et al.
Pubblicazione: (2024)
DRAMA: Diverse Augmentation from Large Language Models to Smaller Dense Retrievers
di: Ma, Xueguang, et al.
Pubblicazione: (2025)
di: Ma, Xueguang, et al.
Pubblicazione: (2025)
Towards Group-aware Search Success
di: Wu, Haolun, et al.
Pubblicazione: (2024)
di: Wu, Haolun, et al.
Pubblicazione: (2024)
EnronQA: Towards Personalized RAG over Private Documents
di: Ryan, Michael J., et al.
Pubblicazione: (2025)
di: Ryan, Michael J., et al.
Pubblicazione: (2025)
Study on LLMs for Promptagator-Style Dense Retriever Training
di: Gwon, Daniel, et al.
Pubblicazione: (2025)
di: Gwon, Daniel, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Initial Nugget Evaluation Results for the TREC 2024 RAG Track with the AutoNuggetizer Framework
di: Pradeep, Ronak, et al.
Pubblicazione: (2024) -
Support Evaluation for the TREC 2024 RAG Track: Comparing Human versus LLM Judges
di: Thakur, Nandan, et al.
Pubblicazione: (2025) -
Overview of the TREC 2025 Retrieval Augmented Generation (RAG) Track
di: Upadhyay, Shivani, et al.
Pubblicazione: (2026) -
A Large-Scale Study of Relevance Assessments with Large Language Models: An Initial Look
di: Upadhyay, Shivani, et al.
Pubblicazione: (2024) -
UMBRELA: UMbrela is the (Open-Source Reproduction of the) Bing RELevance Assessor
di: Upadhyay, Shivani, et al.
Pubblicazione: (2024)