Systematic Evaluation of Neural Retrieval Models on the Touché 2020 Argument Retrieval Subset of BEIR
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Thakur, Nandan, Bonifacio, Luiz, Fröbe, Maik, Bondarenko, Alexander, Kamalloo, Ehsan, Potthast, Martin, Hagen, Matthias, Lin, Jimmy |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Lightning IR: Straightforward Fine-tuning and Inference of Transformer-based Language Models for Information Retrieval
von: Schlatt, Ferdinand, et al.
Veröffentlicht: (2024)
von: Schlatt, Ferdinand, et al.
Veröffentlicht: (2024)
The Viability of Crowdsourcing for RAG Evaluation
von: Gienapp, Lukas, et al.
Veröffentlicht: (2025)
von: Gienapp, Lukas, et al.
Veröffentlicht: (2025)
"Knowing When You Don't Know": A Multilingual Relevance Assessment Dataset for Robust Retrieval-Augmented Generation
von: Thakur, Nandan, et al.
Veröffentlicht: (2023)
von: Thakur, Nandan, et al.
Veröffentlicht: (2023)
Simplified Longitudinal Retrieval Experiments: A Case Study on Query Expansion and Document Boosting
von: Keller, Jüri, et al.
Veröffentlicht: (2025)
von: Keller, Jüri, et al.
Veröffentlicht: (2025)
Analyzing Adversarial Attacks on Sequence-to-Sequence Relevance Models
von: Parry, Andrew, et al.
Veröffentlicht: (2024)
von: Parry, Andrew, et al.
Veröffentlicht: (2024)
LLMs Can Patch Up Missing Relevance Judgments in Evaluation
von: Upadhyay, Shivani, et al.
Veröffentlicht: (2024)
von: Upadhyay, Shivani, et al.
Veröffentlicht: (2024)
Investigating the Effects of Sparse Attention on Cross-Encoders
von: Schlatt, Ferdinand, et al.
Veröffentlicht: (2023)
von: Schlatt, Ferdinand, et al.
Veröffentlicht: (2023)
Evaluating Generative Ad Hoc Information Retrieval
von: Gienapp, Lukas, et al.
Veröffentlicht: (2023)
von: Gienapp, Lukas, et al.
Veröffentlicht: (2023)
Counterfactual Query Rewriting to Use Historical Relevance Feedback
von: Keller, Jüri, et al.
Veröffentlicht: (2025)
von: Keller, Jüri, et al.
Veröffentlicht: (2025)
Still Fresh? Evaluating Temporal Drift in Retrieval Benchmarks
von: Kuissi, Nathan, et al.
Veröffentlicht: (2026)
von: Kuissi, Nathan, et al.
Veröffentlicht: (2026)
Hard Negatives, Hard Lessons: Revisiting Training Data Quality for Robust Information Retrieval with LLMs
von: Thakur, Nandan, et al.
Veröffentlicht: (2025)
von: Thakur, Nandan, et al.
Veröffentlicht: (2025)
Overview of the TREC 2025 Retrieval Augmented Generation (RAG) Track
von: Upadhyay, Shivani, et al.
Veröffentlicht: (2026)
von: Upadhyay, Shivani, et al.
Veröffentlicht: (2026)
FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents
von: Thakur, Nandan, et al.
Veröffentlicht: (2025)
von: Thakur, Nandan, et al.
Veröffentlicht: (2025)
Rank-DistiLLM: Closing the Effectiveness Gap Between Cross-Encoders and LLMs for Passage Re-Ranking
von: Schlatt, Ferdinand, et al.
Veröffentlicht: (2024)
von: Schlatt, Ferdinand, et al.
Veröffentlicht: (2024)
Set-Encoder: Permutation-Invariant Inter-Passage Attention for Listwise Passage Re-Ranking with Cross-Encoders
von: Schlatt, Ferdinand, et al.
Veröffentlicht: (2024)
von: Schlatt, Ferdinand, et al.
Veröffentlicht: (2024)
Hindi-BEIR : A Large Scale Retrieval Benchmark in Hindi
von: Acharya, Arkadeep, et al.
Veröffentlicht: (2024)
von: Acharya, Arkadeep, et al.
Veröffentlicht: (2024)
Wikipedia-based Datasets in Russian Information Retrieval Benchmark RusBEIR
von: Kovalev, Grigory, et al.
Veröffentlicht: (2025)
von: Kovalev, Grigory, et al.
Veröffentlicht: (2025)
Chatbot Arena Meets Nuggets: Towards Explanations and Diagnostics in the Evaluation of LLM Responses
von: Sharifymoghaddam, Sahel, et al.
Veröffentlicht: (2025)
von: Sharifymoghaddam, Sahel, et al.
Veröffentlicht: (2025)
Leveraging LLMs for Synthesizing Training Data Across Many Languages in Multilingual Dense Retrieval
von: Thakur, Nandan, et al.
Veröffentlicht: (2023)
von: Thakur, Nandan, et al.
Veröffentlicht: (2023)
BEIR-PL: Zero Shot Information Retrieval Benchmark for the Polish Language
von: Wojtasik, Konrad, et al.
Veröffentlicht: (2023)
von: Wojtasik, Konrad, et al.
Veröffentlicht: (2023)
Domain-Adaptive Dense Retrieval for Brazilian Legal Search
von: Pereira, Jayr, et al.
Veröffentlicht: (2026)
von: Pereira, Jayr, et al.
Veröffentlicht: (2026)
Ragnarök: A Reusable RAG Framework and Baselines for TREC 2024 Retrieval-Augmented Generation Track
von: Pradeep, Ronak, et al.
Veröffentlicht: (2024)
von: Pradeep, Ronak, et al.
Veröffentlicht: (2024)
Benchmarking and Building Zero-Shot Hindi Retrieval Model with Hindi-BEIR and NLLB-E5
von: Acharya, Arkadeep, et al.
Veröffentlicht: (2024)
von: Acharya, Arkadeep, et al.
Veröffentlicht: (2024)
UMBRELA: UMbrela is the (Open-Source Reproduction of the) Bing RELevance Assessor
von: Upadhyay, Shivani, et al.
Veröffentlicht: (2024)
von: Upadhyay, Shivani, et al.
Veröffentlicht: (2024)
Overview of the Plagiarism Detection Task at PAN 2025
von: Greiner-Petter, André, et al.
Veröffentlicht: (2025)
von: Greiner-Petter, André, et al.
Veröffentlicht: (2025)
ORBIT: Scalable and Verifiable Data Generation for Search Agents on a Tight Budget
von: Thakur, Nandan, et al.
Veröffentlicht: (2026)
von: Thakur, Nandan, et al.
Veröffentlicht: (2026)
Operational Advice for Dense and Sparse Retrievers: HNSW, Flat, or Inverted Indexes?
von: Lin, Jimmy
Veröffentlicht: (2024)
von: Lin, Jimmy
Veröffentlicht: (2024)
The Great Nugget Recall: Automating Fact Extraction and RAG Evaluation with Large Language Models
von: Pradeep, Ronak, et al.
Veröffentlicht: (2025)
von: Pradeep, Ronak, et al.
Veröffentlicht: (2025)
Initial Nugget Evaluation Results for the TREC 2024 RAG Track with the AutoNuggetizer Framework
von: Pradeep, Ronak, et al.
Veröffentlicht: (2024)
von: Pradeep, Ronak, et al.
Veröffentlicht: (2024)
LongRecall: A Structured Approach for Robust Recall Evaluation in Long-Form Text
von: Ardestani, MohamamdJavad, et al.
Veröffentlicht: (2025)
von: Ardestani, MohamamdJavad, et al.
Veröffentlicht: (2025)
Dimension vs. Precision: A Comparative Analysis of Autoencoders and Quantization for Efficient Vector Retrieval on BEIR SciFact
von: Pati, Satyanarayan
Veröffentlicht: (2025)
von: Pati, Satyanarayan
Veröffentlicht: (2025)
Learning Effective Representations for Retrieval Using Self-Distillation with Adaptive Relevance Margins
von: Gienapp, Lukas, et al.
Veröffentlicht: (2024)
von: Gienapp, Lukas, et al.
Veröffentlicht: (2024)
Sensitivity-Aware Retrieval-Augmented Intent Clarification
von: Larooij, Maik
Veröffentlicht: (2026)
von: Larooij, Maik
Veröffentlicht: (2026)
Support Evaluation for the TREC 2024 RAG Track: Comparing Human versus LLM Judges
von: Thakur, Nandan, et al.
Veröffentlicht: (2025)
von: Thakur, Nandan, et al.
Veröffentlicht: (2025)
Ranking Generated Answers: On the Agreement of Retrieval Models with Humans on Consumer Health Questions
von: Heineking, Sebastian, et al.
Veröffentlicht: (2024)
von: Heineking, Sebastian, et al.
Veröffentlicht: (2024)
Multilingual and Domain-Agnostic Tip-of-the-Tongue Query Generation for Simulated Evaluation
von: He, Xuhong, et al.
Veröffentlicht: (2026)
von: He, Xuhong, et al.
Veröffentlicht: (2026)
CURE: A Dataset for Clinical Understanding & Retrieval Evaluation
von: Sheikh, Nadia Athar, et al.
Veröffentlicht: (2024)
von: Sheikh, Nadia Athar, et al.
Veröffentlicht: (2024)
Illusions of Relevance: Arbitrary Content Injection Attacks Deceive Retrievers, Rerankers, and LLM Judges
von: Tamber, Manveer Singh, et al.
Veröffentlicht: (2025)
von: Tamber, Manveer Singh, et al.
Veröffentlicht: (2025)
Unifying Multimodal Retrieval via Document Screenshot Embedding
von: Ma, Xueguang, et al.
Veröffentlicht: (2024)
von: Ma, Xueguang, et al.
Veröffentlicht: (2024)
JUÁ -- A Benchmark for Information Retrieval in Brazilian Legal Text Collections
von: Pereira, Jayr, et al.
Veröffentlicht: (2026)
von: Pereira, Jayr, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Lightning IR: Straightforward Fine-tuning and Inference of Transformer-based Language Models for Information Retrieval
von: Schlatt, Ferdinand, et al.
Veröffentlicht: (2024) -
The Viability of Crowdsourcing for RAG Evaluation
von: Gienapp, Lukas, et al.
Veröffentlicht: (2025) -
"Knowing When You Don't Know": A Multilingual Relevance Assessment Dataset for Robust Retrieval-Augmented Generation
von: Thakur, Nandan, et al.
Veröffentlicht: (2023) -
Simplified Longitudinal Retrieval Experiments: A Case Study on Query Expansion and Document Boosting
von: Keller, Jüri, et al.
Veröffentlicht: (2025) -
Analyzing Adversarial Attacks on Sequence-to-Sequence Relevance Models
von: Parry, Andrew, et al.
Veröffentlicht: (2024)