Auto-ARGUE: LLM-Based Report Generation Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Walden, William, Mason, Marc, Weller, Orion, Dietz, Laura, Conroy, John, Molino, Neil, Recknor, Hannah, Li, Bryan, Liu, Gabrielle Kaili-May, Hou, Yu, Lawrie, Dawn, Mayfield, James, Yang, Eugene |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DoGMaTiQ: Automated Generation of Question-and-Answer Nuggets for Report Evaluation
by: Li, Bryan, et al.
Published: (2026)
by: Li, Bryan, et al.
Published: (2026)
HLTCOE at TREC 2024 NeuCLIR Track
by: Yang, Eugene, et al.
Published: (2025)
by: Yang, Eugene, et al.
Published: (2025)
Insider Knowledge: How Much Can RAG Systems Gain from Evaluation Secrets?
by: Dietz, Laura, et al.
Published: (2026)
by: Dietz, Laura, et al.
Published: (2026)
Incorporating Q&A Nuggets into Retrieval-Augmented Generation
by: Dietz, Laura, et al.
Published: (2026)
by: Dietz, Laura, et al.
Published: (2026)
HLTCOE at TREC 2023 NeuCLIR Track
by: Yang, Eugene, et al.
Published: (2024)
by: Yang, Eugene, et al.
Published: (2024)
Distillation for Multilingual Information Retrieval
by: Yang, Eugene, et al.
Published: (2024)
by: Yang, Eugene, et al.
Published: (2024)
HLTCOE at LiveRAG: GPT-Researcher using ColBERT retrieval
by: Duh, Kevin, et al.
Published: (2025)
by: Duh, Kevin, et al.
Published: (2025)
NevIR: Negation in Neural Information Retrieval
by: Weller, Orion, et al.
Published: (2023)
by: Weller, Orion, et al.
Published: (2023)
Investigating Retrieval-Augmented Generation Systems on Unanswerable, Uncheatable, Realistic, Multi-hop Queries
by: Liu, Gabrielle Kaili-May, et al.
Published: (2025)
by: Liu, Gabrielle Kaili-May, et al.
Published: (2025)
Language Fairness in Multilingual Information Retrieval
by: Yang, Eugene, et al.
Published: (2024)
by: Yang, Eugene, et al.
Published: (2024)
On the Evaluation of Machine-Generated Reports
by: Mayfield, James, et al.
Published: (2024)
by: Mayfield, James, et al.
Published: (2024)
mmBERT: A Modern Multilingual Encoder with Annealed Language Learning
by: Marone, Marc, et al.
Published: (2025)
by: Marone, Marc, et al.
Published: (2025)
RoutIR: Fast Serving of Retrieval Pipelines for Retrieval-Augmented Generation
by: Yang, Eugene, et al.
Published: (2026)
by: Yang, Eugene, et al.
Published: (2026)
Extending Translate-Train for ColBERT-X to African Language CLIR
by: Yang, Eugene, et al.
Published: (2024)
by: Yang, Eugene, et al.
Published: (2024)
Rank1: Test-Time Compute for Reranking in Information Retrieval
by: Weller, Orion, et al.
Published: (2025)
by: Weller, Orion, et al.
Published: (2025)
Defending Against Disinformation Attacks in Open-Domain Question Answering
by: Weller, Orion, et al.
Published: (2022)
by: Weller, Orion, et al.
Published: (2022)
Translate-Distill: Learning Cross-Language Dense Retrieval by Translation and Distillation
by: Yang, Eugene, et al.
Published: (2024)
by: Yang, Eugene, et al.
Published: (2024)
PLAID SHIRTTT for Large-Scale Streaming Dense Retrieval
by: Lawrie, Dawn, et al.
Published: (2024)
by: Lawrie, Dawn, et al.
Published: (2024)
Rank-K: Test-Time Reasoning for Listwise Reranking
by: Yang, Eugene, et al.
Published: (2025)
by: Yang, Eugene, et al.
Published: (2025)
Efficiency-Effectiveness Tradeoff of Probabilistic Structured Queries for Cross-Language Information Retrieval
by: Yang, Eugene, et al.
Published: (2024)
by: Yang, Eugene, et al.
Published: (2024)
Seq vs Seq: An Open Suite of Paired Encoders and Decoders
by: Weller, Orion, et al.
Published: (2025)
by: Weller, Orion, et al.
Published: (2025)
Overview of the TREC 2025 RAGTIME Track
by: Lawrie, Dawn, et al.
Published: (2026)
by: Lawrie, Dawn, et al.
Published: (2026)
Promptriever: Instruction-Trained Retrievers Can Be Prompted Like Language Models
by: Weller, Orion, et al.
Published: (2024)
by: Weller, Orion, et al.
Published: (2024)
Overview of the TREC 2024 NeuCLIR Track
by: Lawrie, Dawn, et al.
Published: (2025)
by: Lawrie, Dawn, et al.
Published: (2025)
Overview of the TREC 2023 NeuCLIR Track
by: Lawrie, Dawn, et al.
Published: (2024)
by: Lawrie, Dawn, et al.
Published: (2024)
MURR: Model Updating with Regularized Replay for Searching a Document Stream
by: Yang, Eugene, et al.
Published: (2025)
by: Yang, Eugene, et al.
Published: (2025)
When do Generative Query and Document Expansions Fail? A Comprehensive Study Across Methods, Retrievers, and Datasets
by: Weller, Orion, et al.
Published: (2023)
by: Weller, Orion, et al.
Published: (2023)
Beyond Relevance: On the Relationship Between Retrieval and RAG Information Coverage
by: Samuel, Saron, et al.
Published: (2026)
by: Samuel, Saron, et al.
Published: (2026)
NeuCLIRTech: Chinese Monolingual and Cross-Language Information Retrieval Evaluation in a Challenging Domain
by: Lawrie, Dawn, et al.
Published: (2026)
by: Lawrie, Dawn, et al.
Published: (2026)
NeuCLIRBench: A Modern Evaluation Collection for Monolingual, Cross-Language, and Multilingual Information Retrieval
by: Lawrie, Dawn, et al.
Published: (2025)
by: Lawrie, Dawn, et al.
Published: (2025)
mFollowIR: a Multilingual Benchmark for Instruction Following in Retrieval
by: Weller, Orion, et al.
Published: (2025)
by: Weller, Orion, et al.
Published: (2025)
FollowIR: Evaluating and Teaching Information Retrieval Models to Follow Instructions
by: Weller, Orion, et al.
Published: (2024)
by: Weller, Orion, et al.
Published: (2024)
Principled Context Engineering for RAG: Statistical Guarantees via Conformal Prediction
by: Chakraborty, Debashish, et al.
Published: (2025)
by: Chakraborty, Debashish, et al.
Published: (2025)
CoverageBench: Evaluating Information Coverage across Tasks and Domains
by: Samuel, Saron, et al.
Published: (2026)
by: Samuel, Saron, et al.
Published: (2026)
On the Theoretical Limitations of Embedding-Based Retrieval
by: Weller, Orion, et al.
Published: (2025)
by: Weller, Orion, et al.
Published: (2025)
Beyond RAG: Task-Aware KV Cache Compression for Comprehensive Knowledge Reasoning
by: Corallo, Giulio, et al.
Published: (2025)
by: Corallo, Giulio, et al.
Published: (2025)
A Workbench for Autograding Retrieve/Generate Systems
by: Dietz, Laura
Published: (2024)
by: Dietz, Laura
Published: (2024)
LLM-based relevance assessment still can't replace human relevance assessment
by: Clarke, Charles L. A., et al.
Published: (2024)
by: Clarke, Charles L. A., et al.
Published: (2024)
LLM4Rerank: LLM-based Auto-Reranking Framework for Recommendations
by: Gao, Jingtong, et al.
Published: (2024)
by: Gao, Jingtong, et al.
Published: (2024)
Criteria-Based LLM Relevance Judgments
by: Farzi, Naghmeh, et al.
Published: (2025)
by: Farzi, Naghmeh, et al.
Published: (2025)
Similar Items
-
DoGMaTiQ: Automated Generation of Question-and-Answer Nuggets for Report Evaluation
by: Li, Bryan, et al.
Published: (2026) -
HLTCOE at TREC 2024 NeuCLIR Track
by: Yang, Eugene, et al.
Published: (2025) -
Insider Knowledge: How Much Can RAG Systems Gain from Evaluation Secrets?
by: Dietz, Laura, et al.
Published: (2026) -
Incorporating Q&A Nuggets into Retrieval-Augmented Generation
by: Dietz, Laura, et al.
Published: (2026) -
HLTCOE at TREC 2023 NeuCLIR Track
by: Yang, Eugene, et al.
Published: (2024)