Towards a rigorous evaluation of RAG systems: the challenge of due diligence
Fuente:
arXiv
Saved in:
| Main Authors: | Martinon, Grégoire, de Brionne, Alexandra Lorenzo, Bohard, Jérôme, Lojou, Antoine, Hervault, Damien, Brunel, Nicolas J-B. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Subnational Geocoding of Global Disasters Using Large Language Models
by: Ronco, Michele, et al.
Published: (2025)
by: Ronco, Michele, et al.
Published: (2025)
Space evaluation at the starting point of soccer transitions
by: Ogawa, Yohei, et al.
Published: (2025)
by: Ogawa, Yohei, et al.
Published: (2025)
Cinder: A fast and fair matchmaking system
by: Pal, Saurav
Published: (2025)
by: Pal, Saurav
Published: (2025)
BioDisco: Multi-agent hypothesis generation with dual-mode evidence, iterative feedback and temporal evaluation
by: Ke, Yujing, et al.
Published: (2025)
by: Ke, Yujing, et al.
Published: (2025)
SureMap: Simultaneous Mean Estimation for Single-Task and Multi-Task Disaggregated Evaluation
by: Khodak, Mikhail, et al.
Published: (2024)
by: Khodak, Mikhail, et al.
Published: (2024)
Bridging the Data Gap in AI Reliability Research and Establishing DR-AIR, a Comprehensive Data Repository for AI Reliability
by: Zheng, Simin, et al.
Published: (2025)
by: Zheng, Simin, et al.
Published: (2025)
Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking
by: Xu, Yang, et al.
Published: (2026)
by: Xu, Yang, et al.
Published: (2026)
Toward Reducing Unproductive Container Moves: Predicting Service Requirements and Dwell Times
by: Villalobos, Elena, et al.
Published: (2026)
by: Villalobos, Elena, et al.
Published: (2026)
Industrializing Prediction-Powered Inference: The GLIDE Library for Reliable GenAI and Agentic Systems Evaluation
by: Martinon, Grégoire, et al.
Published: (2026)
by: Martinon, Grégoire, et al.
Published: (2026)
On the Mechanistic Interpretability of Neural Networks for Causality in Bio-statistics
by: Conan, Jean-Baptiste A.
Published: (2025)
by: Conan, Jean-Baptiste A.
Published: (2025)
Decision Quality Evaluation Framework at Pinterest
by: Tian, Yuqi, et al.
Published: (2026)
by: Tian, Yuqi, et al.
Published: (2026)
Quantitative Technology Forecasting: a Review of Trend Extrapolation Methods
by: Tsai, Peng-Hung, et al.
Published: (2024)
by: Tsai, Peng-Hung, et al.
Published: (2024)
StatLLM: A Dataset for Evaluating the Performance of Large Language Models in Statistical Analysis
by: Song, Xinyi, et al.
Published: (2025)
by: Song, Xinyi, et al.
Published: (2025)
HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics
by: Luettgau, Lennart, et al.
Published: (2025)
by: Luettgau, Lennart, et al.
Published: (2025)
Data-Driven Bayesian Network Models of Hurricane Evacuation Decision Making
by: Wang, Hui Sophie, et al.
Published: (2023)
by: Wang, Hui Sophie, et al.
Published: (2023)
Decade-long Emission Forecasting with an Ensemble Model in Taiwan
by: Hung, Gordon, et al.
Published: (2025)
by: Hung, Gordon, et al.
Published: (2025)
ChatGPT and post-test probability
by: Weisenthal, Samuel J.
Published: (2023)
by: Weisenthal, Samuel J.
Published: (2023)
Process-Aware Analysis of Treatment Paths in Heart Failure Patients: A Case Study
by: Beyel, Harry H., et al.
Published: (2024)
by: Beyel, Harry H., et al.
Published: (2024)
Surrogate-Based Prevalence Measurement for Large-Scale A/B Testing
by: Xu, Zehao, et al.
Published: (2026)
by: Xu, Zehao, et al.
Published: (2026)
Rectification Difficulty and Optimal Sample Allocation in LLM-Augmented Surveys
by: Ye, Zikun, et al.
Published: (2026)
by: Ye, Zikun, et al.
Published: (2026)
Calculating Customer Lifetime Value and Churn using Beta Geometric Negative Binomial and Gamma-Gamma Distribution in a NFT based setting
by: Das, Sagarnil
Published: (2025)
by: Das, Sagarnil
Published: (2025)
Unlocking the Potential of Past Research: Using Generative AI to Reconstruct Healthcare Simulation Models
by: Monks, Thomas, et al.
Published: (2025)
by: Monks, Thomas, et al.
Published: (2025)
The Advancement of Personalized Learning Potentially Accelerated by Generative AI
by: Wei, Yuang, et al.
Published: (2024)
by: Wei, Yuang, et al.
Published: (2024)
Performance Evaluation of Large Language Models in Statistical Programming
by: Song, Xinyi, et al.
Published: (2025)
by: Song, Xinyi, et al.
Published: (2025)
TCKAN:A Novel Integrated Network Model for Predicting Mortality Risk in Sepsis Patients
by: Dong, Fanglin
Published: (2024)
by: Dong, Fanglin
Published: (2024)
Scalable Stewardship of an LLM-Assisted Clinical Benchmark with Physician Oversight
by: Ye, Junze, et al.
Published: (2025)
by: Ye, Junze, et al.
Published: (2025)
Integrating Dynamic Correlation Shifts and Weighted Benchmarking in Extreme Value Analysis
by: Panagoulias, Dimitrios P., et al.
Published: (2024)
by: Panagoulias, Dimitrios P., et al.
Published: (2024)
Classification Modeling with RNN-Based, Random Forest, and XGBoost for Imbalanced Data: A Case of Early Crash Detection in ASEAN-5 Stock Markets
by: Siswara, Deri, et al.
Published: (2024)
by: Siswara, Deri, et al.
Published: (2024)
Causal inference approach to appraise long-term effects of maintenance policy on functional performance of asphalt pavements
by: You, Lingyun, et al.
Published: (2024)
by: You, Lingyun, et al.
Published: (2024)
A survey of using EHR as real-world evidence for discovering and validating new drug indications
by: Talukdar, Nabasmita, et al.
Published: (2025)
by: Talukdar, Nabasmita, et al.
Published: (2025)
Analyzing the factors that are involved in length of inpatient stay at the hospital for diabetes patients
by: Lam, Jorden, et al.
Published: (2024)
by: Lam, Jorden, et al.
Published: (2024)
Classifying Metamorphic versus Single-Fold Proteins with Statistical Learning and AlphaFold2
by: Chen, Yongkai, et al.
Published: (2025)
by: Chen, Yongkai, et al.
Published: (2025)
A network analysis of decision strategies of human experts in steel manufacturing
by: Merten, Daniel Christopher, et al.
Published: (2021)
by: Merten, Daniel Christopher, et al.
Published: (2021)
CERES: A Probabilistic Early Warning System for Acute Food Insecurity
by: Pedersen, Tom Danny S.
Published: (2026)
by: Pedersen, Tom Danny S.
Published: (2026)
Using large language models to produce literature reviews: Usages and systematic biases of microphysics parametrizations in 2699 publications
by: Zhang, Tianhang, et al.
Published: (2025)
by: Zhang, Tianhang, et al.
Published: (2025)
Evaluating the Use of Large Language Models as Synthetic Social Agents in Social Science Research
by: Madden, Emma Rose
Published: (2025)
by: Madden, Emma Rose
Published: (2025)
A Regression Mixture Model to understand the effect of the Covid-19 pandemic on Public Transport Ridership
by: Moreau, Hugues, et al.
Published: (2024)
by: Moreau, Hugues, et al.
Published: (2024)
AI for Handball: predicting and explaining the 2024 Olympic Games tournament with Deep Learning and Large Language Models
by: Felice, Florian
Published: (2024)
by: Felice, Florian
Published: (2024)
Embracing Ambiguity: Bayesian Nonparametrics and Stakeholder Participation for Ambiguity-Aware Safety Evaluation
by: Long, Yanan
Published: (2025)
by: Long, Yanan
Published: (2025)
SEED-SET: Scalable Evolving Experimental Design for System-level Ethical Testing
by: Parashar, Anjali, et al.
Published: (2026)
by: Parashar, Anjali, et al.
Published: (2026)
Similar Items
-
Subnational Geocoding of Global Disasters Using Large Language Models
by: Ronco, Michele, et al.
Published: (2025) -
Space evaluation at the starting point of soccer transitions
by: Ogawa, Yohei, et al.
Published: (2025) -
Cinder: A fast and fair matchmaking system
by: Pal, Saurav
Published: (2025) -
BioDisco: Multi-agent hypothesis generation with dual-mode evidence, iterative feedback and temporal evaluation
by: Ke, Yujing, et al.
Published: (2025) -
SureMap: Simultaneous Mean Estimation for Single-Task and Multi-Task Disaggregated Evaluation
by: Khodak, Mikhail, et al.
Published: (2024)