Chatbot Arena Meets Nuggets: Towards Explanations and Diagnostics in the Evaluation of LLM Responses
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Sharifymoghaddam, Sahel, Upadhyay, Shivani, Thakur, Nandan, Pradeep, Ronak, Lin, Jimmy |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Initial Nugget Evaluation Results for the TREC 2024 RAG Track with the AutoNuggetizer Framework
par: Pradeep, Ronak, et autres
Publié: (2024)
par: Pradeep, Ronak, et autres
Publié: (2024)
The Great Nugget Recall: Automating Fact Extraction and RAG Evaluation with Large Language Models
par: Pradeep, Ronak, et autres
Publié: (2025)
par: Pradeep, Ronak, et autres
Publié: (2025)
UMBRELA: UMbrela is the (Open-Source Reproduction of the) Bing RELevance Assessor
par: Upadhyay, Shivani, et autres
Publié: (2024)
par: Upadhyay, Shivani, et autres
Publié: (2024)
UniRAG: Universal Retrieval Augmentation for Large Vision Language Models
par: Sharifymoghaddam, Sahel, et autres
Publié: (2024)
par: Sharifymoghaddam, Sahel, et autres
Publié: (2024)
Support Evaluation for the TREC 2024 RAG Track: Comparing Human versus LLM Judges
par: Thakur, Nandan, et autres
Publié: (2025)
par: Thakur, Nandan, et autres
Publié: (2025)
Overview of the TREC 2025 Retrieval Augmented Generation (RAG) Track
par: Upadhyay, Shivani, et autres
Publié: (2026)
par: Upadhyay, Shivani, et autres
Publié: (2026)
Rerank Before You Reason: Analyzing Reranking Tradeoffs through Effective Token Cost in Deep Search Agents
par: Sharifymoghaddam, Sahel, et autres
Publié: (2026)
par: Sharifymoghaddam, Sahel, et autres
Publié: (2026)
Ragnarök: A Reusable RAG Framework and Baselines for TREC 2024 Retrieval-Augmented Generation Track
par: Pradeep, Ronak, et autres
Publié: (2024)
par: Pradeep, Ronak, et autres
Publié: (2024)
Lighting the Way for BRIGHT: Reproducible Baselines with Anserini, Pyserini, and RankLLM
par: Sharifymoghaddam, Sahel, et autres
Publié: (2025)
par: Sharifymoghaddam, Sahel, et autres
Publié: (2025)
A Large-Scale Study of Relevance Assessments with Large Language Models: An Initial Look
par: Upadhyay, Shivani, et autres
Publié: (2024)
par: Upadhyay, Shivani, et autres
Publié: (2024)
RankLLM: A Python Package for Reranking with LLMs
par: Sharifymoghaddam, Sahel, et autres
Publié: (2025)
par: Sharifymoghaddam, Sahel, et autres
Publié: (2025)
LLMs Can Patch Up Missing Relevance Judgments in Evaluation
par: Upadhyay, Shivani, et autres
Publié: (2024)
par: Upadhyay, Shivani, et autres
Publié: (2024)
An Early FIRST Reproduction and Improvements to Single-Token Decoding for Fast Listwise Reranking
par: Chen, Zijian, et autres
Publié: (2024)
par: Chen, Zijian, et autres
Publié: (2024)
Still Fresh? Evaluating Temporal Drift in Retrieval Benchmarks
par: Kuissi, Nathan, et autres
Publié: (2026)
par: Kuissi, Nathan, et autres
Publié: (2026)
LANCER: LLM Reranking for Nugget Coverage
par: Ju, Jia-Huei, et autres
Publié: (2026)
par: Ju, Jia-Huei, et autres
Publié: (2026)
Hard Negatives, Hard Lessons: Revisiting Training Data Quality for Robust Information Retrieval with LLMs
par: Thakur, Nandan, et autres
Publié: (2025)
par: Thakur, Nandan, et autres
Publié: (2025)
ORBIT: Scalable and Verifiable Data Generation for Search Agents on a Tight Budget
par: Thakur, Nandan, et autres
Publié: (2026)
par: Thakur, Nandan, et autres
Publié: (2026)
On the Comprehensibility of Multi-structured Financial Documents using LLMs and Pre-processing Tools
par: Upadhyay, Shivani, et autres
Publié: (2025)
par: Upadhyay, Shivani, et autres
Publié: (2025)
Systematic Evaluation of Neural Retrieval Models on the Touché 2020 Argument Retrieval Subset of BEIR
par: Thakur, Nandan, et autres
Publié: (2024)
par: Thakur, Nandan, et autres
Publié: (2024)
BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent
par: Chen, Zijian, et autres
Publié: (2025)
par: Chen, Zijian, et autres
Publié: (2025)
FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents
par: Thakur, Nandan, et autres
Publié: (2025)
par: Thakur, Nandan, et autres
Publié: (2025)
MedNuggetizer: Confidence-Based Information Nugget Extraction from Medical Documents
par: Donabauer, Gregor, et autres
Publié: (2025)
par: Donabauer, Gregor, et autres
Publié: (2025)
GINGER: Grounded Information Nugget-Based Generation of Responses
par: Łajewska, Weronika, et autres
Publié: (2025)
par: Łajewska, Weronika, et autres
Publié: (2025)
UiS-IAI@LiveRAG: Retrieval-Augmented Information Nugget-Based Generation of Responses
par: Łajewska, Weronika, et autres
Publié: (2025)
par: Łajewska, Weronika, et autres
Publié: (2025)
Conversational Gold: Evaluating Personalized Conversational Search System using Gold Nuggets
par: Abbasiantaeb, Zahra, et autres
Publié: (2025)
par: Abbasiantaeb, Zahra, et autres
Publié: (2025)
Leveraging LLMs for Synthesizing Training Data Across Many Languages in Multilingual Dense Retrieval
par: Thakur, Nandan, et autres
Publié: (2023)
par: Thakur, Nandan, et autres
Publié: (2023)
NuggetIndex: Governed Atomic Retrieval for Maintainable RAG
par: Zerhoudi, Saber, et autres
Publié: (2026)
par: Zerhoudi, Saber, et autres
Publié: (2026)
Illusions of Relevance: Arbitrary Content Injection Attacks Deceive Retrievers, Rerankers, and LLM Judges
par: Tamber, Manveer Singh, et autres
Publié: (2025)
par: Tamber, Manveer Singh, et autres
Publié: (2025)
DoGMaTiQ: Automated Generation of Question-and-Answer Nuggets for Report Evaluation
par: Li, Bryan, et autres
Publié: (2026)
par: Li, Bryan, et autres
Publié: (2026)
RankArena: A Unified Platform for Evaluating Retrieval, Reranking and RAG with Human and LLM Feedback
par: Abdallah, Abdelrahman, et autres
Publié: (2025)
par: Abdallah, Abdelrahman, et autres
Publié: (2025)
Operational Advice for Dense and Sparse Retrievers: HNSW, Flat, or Inverted Indexes?
par: Lin, Jimmy
Publié: (2024)
par: Lin, Jimmy
Publié: (2024)
NeuCLIRBench: A Modern Evaluation Collection for Monolingual, Cross-Language, and Multilingual Information Retrieval
par: Lawrie, Dawn, et autres
Publié: (2025)
par: Lawrie, Dawn, et autres
Publié: (2025)
NeuCLIRTech: Chinese Monolingual and Cross-Language Information Retrieval Evaluation in a Challenging Domain
par: Lawrie, Dawn, et autres
Publié: (2026)
par: Lawrie, Dawn, et autres
Publié: (2026)
Teaching Dense Retrieval Models to Specialize with Listwise Distillation and LLM Data Augmentation
par: Tamber, Manveer Singh, et autres
Publié: (2025)
par: Tamber, Manveer Singh, et autres
Publié: (2025)
"Knowing When You Don't Know": A Multilingual Relevance Assessment Dataset for Robust Retrieval-Augmented Generation
par: Thakur, Nandan, et autres
Publié: (2023)
par: Thakur, Nandan, et autres
Publié: (2023)
IntellBot: Retrieval Augmented LLM Chatbot for Cyber Threat Knowledge Delivery
par: Arikkat, Dincy R., et autres
Publié: (2024)
par: Arikkat, Dincy R., et autres
Publié: (2024)
Revisiting Feedback Models for HyDE
par: Jedidi, Nour, et autres
Publié: (2025)
par: Jedidi, Nour, et autres
Publié: (2025)
ConvKGYarn: Spinning Configurable and Scalable Conversational Knowledge Graph QA datasets with Large Language Models
par: Pradeep, Ronak, et autres
Publié: (2024)
par: Pradeep, Ronak, et autres
Publié: (2024)
Incorporating Q&A Nuggets into Retrieval-Augmented Generation
par: Dietz, Laura, et autres
Publié: (2026)
par: Dietz, Laura, et autres
Publié: (2026)
Walert: Putting Conversational Search Knowledge into Action by Building and Evaluating a Large Language Model-Powered Chatbot
par: Cherumanal, Sachin Pathiyan, et autres
Publié: (2024)
par: Cherumanal, Sachin Pathiyan, et autres
Publié: (2024)
Documents similaires
-
Initial Nugget Evaluation Results for the TREC 2024 RAG Track with the AutoNuggetizer Framework
par: Pradeep, Ronak, et autres
Publié: (2024) -
The Great Nugget Recall: Automating Fact Extraction and RAG Evaluation with Large Language Models
par: Pradeep, Ronak, et autres
Publié: (2025) -
UMBRELA: UMbrela is the (Open-Source Reproduction of the) Bing RELevance Assessor
par: Upadhyay, Shivani, et autres
Publié: (2024) -
UniRAG: Universal Retrieval Augmentation for Large Vision Language Models
par: Sharifymoghaddam, Sahel, et autres
Publié: (2024) -
Support Evaluation for the TREC 2024 RAG Track: Comparing Human versus LLM Judges
par: Thakur, Nandan, et autres
Publié: (2025)