RAGProbe: An Automated Approach for Evaluating RAG Applications
Fuente:
arXiv
Saved in:
| Main Authors: | Sivasothy, Shangeetha, Barnett, Scott, Kurniawan, Stefanus, Rasool, Zafaryab, Vasa, Rajesh |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Minimising changes to audit when updating decision trees
by: Simmons, Anj, et al.
Published: (2024)
by: Simmons, Anj, et al.
Published: (2024)
Large language models for generating rules, yay or nay?
by: Sivasothy, Shangeetha, et al.
Published: (2024)
by: Sivasothy, Shangeetha, et al.
Published: (2024)
The M-factor: A Novel Metric for Evaluating Neural Architecture Search in Resource-Constrained Environments
by: Thudumu, Srikanth, et al.
Published: (2025)
by: Thudumu, Srikanth, et al.
Published: (2025)
Fine-Tuning or Fine-Failing? Debunking Performance Myths in Large Language Models
by: Barnett, Scott, et al.
Published: (2024)
by: Barnett, Scott, et al.
Published: (2024)
Evaluating LLMs on Document-Based QA: Exact Answer Selection and Numerical Extraction using Cogtale dataset
by: Rasool, Zafaryab, et al.
Published: (2023)
by: Rasool, Zafaryab, et al.
Published: (2023)
LLMs for Test Input Generation for Semantic Caches
by: Rasool, Zafaryab, et al.
Published: (2024)
by: Rasool, Zafaryab, et al.
Published: (2024)
Training and Evaluating with Human Label Variation: An Empirical Study
by: Kurniawan, Kemal, et al.
Published: (2025)
by: Kurniawan, Kemal, et al.
Published: (2025)
The HalluRAG Dataset: Detecting Closed-Domain Hallucinations in RAG Applications Using an LLM's Internal States
by: Ridder, Fabian, et al.
Published: (2024)
by: Ridder, Fabian, et al.
Published: (2024)
Automating Evaluation of Diffusion Model Unlearning with (Vision-) Language Model World Knowledge
by: Yeats, Eric, et al.
Published: (2025)
by: Yeats, Eric, et al.
Published: (2025)
Topics as Entity Clusters: Entity-based Topics from Large Language Models and Graph Neural Networks
by: Loureiro, Manuel V., et al.
Published: (2023)
by: Loureiro, Manuel V., et al.
Published: (2023)
Reward-RAG: Enhancing RAG with Reward Driven Supervision
by: Nguyen, Thang, et al.
Published: (2024)
by: Nguyen, Thang, et al.
Published: (2024)
Towards Automated Safety Requirements Derivation Using Agent-based RAG
by: Balu, Balahari Vignesh, et al.
Published: (2025)
by: Balu, Balahari Vignesh, et al.
Published: (2025)
LatentRAG: Latent Reasoning and Retrieval for Efficient Agentic RAG
by: Zheng, Yijia, et al.
Published: (2026)
by: Zheng, Yijia, et al.
Published: (2026)
FutureGen: A RAG-based Approach to Generate the Future Work of Scientific Article
by: Azher, Ibrahim Al, et al.
Published: (2025)
by: Azher, Ibrahim Al, et al.
Published: (2025)
Anterior's Approach to Fairness Evaluation of Automated Prior Authorization System
by: Selvaraj, Sai P., et al.
Published: (2026)
by: Selvaraj, Sai P., et al.
Published: (2026)
TaskEval: Synthesised Evaluation for Foundation-Model Tasks
by: Widanapathiranage, Dilani, et al.
Published: (2025)
by: Widanapathiranage, Dilani, et al.
Published: (2025)
Failing to Falsify: Evaluating and Mitigating Confirmation Bias in Language Models
by: Jhaveri, Ayush Rajesh, et al.
Published: (2026)
by: Jhaveri, Ayush Rajesh, et al.
Published: (2026)
SimulRAG: Simulator-based RAG for Grounding LLMs in Long-form Scientific QA
by: Xu, Haozhou, et al.
Published: (2025)
by: Xu, Haozhou, et al.
Published: (2025)
SPICED: News Similarity Detection Dataset with Multiple Topics and Complexity Levels
by: Shushkevich, Elena, et al.
Published: (2023)
by: Shushkevich, Elena, et al.
Published: (2023)
Structured RAG for Answering Aggregative Questions
by: Koshorek, Omri, et al.
Published: (2025)
by: Koshorek, Omri, et al.
Published: (2025)
Mitigating Bias in RAG: Controlling the Embedder
by: Kim, Taeyoun, et al.
Published: (2025)
by: Kim, Taeyoun, et al.
Published: (2025)
Highlight & Summarize: RAG without the jailbreaks
by: Cherubin, Giovanni, et al.
Published: (2025)
by: Cherubin, Giovanni, et al.
Published: (2025)
P-RAG: Prompt-Enhanced Parametric RAG with LoRA and Selective CoT for Biomedical and Multi-Hop QA
by: Lyu, Xingda, et al.
Published: (2026)
by: Lyu, Xingda, et al.
Published: (2026)
Benchmarking Automated Clinical Language Simplification: Dataset, Algorithm, and Evaluation
by: Luo, Junyu, et al.
Published: (2020)
by: Luo, Junyu, et al.
Published: (2020)
Evaluating Spoken Language as a Biomarker for Automated Screening of Cognitive Impairment
by: Lima, Maria R., et al.
Published: (2025)
by: Lima, Maria R., et al.
Published: (2025)
ML-On-Rails: Safeguarding Machine Learning Models in Software Systems A Case Study
by: Abdelkader, Hala, et al.
Published: (2024)
by: Abdelkader, Hala, et al.
Published: (2024)
Preference learning made easy: Everything should be understood through win rate
by: Zhang, Lily H., et al.
Published: (2025)
by: Zhang, Lily H., et al.
Published: (2025)
Long Context RAG Performance of Large Language Models
by: Leng, Quinn, et al.
Published: (2024)
by: Leng, Quinn, et al.
Published: (2024)
Insight-RAG: Enhancing LLMs with Insight-Driven Augmentation
by: Pezeshkpour, Pouya, et al.
Published: (2025)
by: Pezeshkpour, Pouya, et al.
Published: (2025)
Operationalizing Automated Essay Scoring: A Human-Aware Approach
by: Plasencia-Calaña, Yenisel
Published: (2025)
by: Plasencia-Calaña, Yenisel
Published: (2025)
Demystifying Legalese: An Automated Approach for Summarizing and Analyzing Overlaps in Privacy Policies and Terms of Service
by: Soneji, Shikha, et al.
Published: (2024)
by: Soneji, Shikha, et al.
Published: (2024)
RAG Playground: A Framework for Systematic Evaluation of Retrieval Strategies and Prompt Engineering in RAG Systems
by: Papadimitriou, Ioannis, et al.
Published: (2024)
by: Papadimitriou, Ioannis, et al.
Published: (2024)
LLM-Independent Adaptive RAG: Let the Question Speak for Itself
by: Marina, Maria, et al.
Published: (2025)
by: Marina, Maria, et al.
Published: (2025)
PersonaBOT: Bringing Customer Personas to Life with LLMs and RAG
by: Rizwan, Muhammed, et al.
Published: (2025)
by: Rizwan, Muhammed, et al.
Published: (2025)
Diversity Enhances an LLM's Performance in RAG and Long-context Task
by: Wang, Zhichao, et al.
Published: (2025)
by: Wang, Zhichao, et al.
Published: (2025)
The Confidence Trap: Gender Bias and Predictive Certainty in LLMs
by: Sabir, Ahmed, et al.
Published: (2026)
by: Sabir, Ahmed, et al.
Published: (2026)
PrismRAG: Boosting RAG Factuality with Distractor Resilience and Strategized Reasoning
by: Kachuee, Mohammad, et al.
Published: (2025)
by: Kachuee, Mohammad, et al.
Published: (2025)
CCRS: A Zero-Shot LLM-as-a-Judge Framework for Comprehensive RAG Evaluation
by: Muhamed, Aashiq
Published: (2025)
by: Muhamed, Aashiq
Published: (2025)
Prescriptive Agents based on RAG for Automated Maintenance (PARAM)
by: Harbola, Chitranshu, et al.
Published: (2025)
by: Harbola, Chitranshu, et al.
Published: (2025)
Plan*RAG: Efficient Test-Time Planning for Retrieval Augmented Generation
by: Verma, Prakhar, et al.
Published: (2024)
by: Verma, Prakhar, et al.
Published: (2024)
Similar Items
-
Minimising changes to audit when updating decision trees
by: Simmons, Anj, et al.
Published: (2024) -
Large language models for generating rules, yay or nay?
by: Sivasothy, Shangeetha, et al.
Published: (2024) -
The M-factor: A Novel Metric for Evaluating Neural Architecture Search in Resource-Constrained Environments
by: Thudumu, Srikanth, et al.
Published: (2025) -
Fine-Tuning or Fine-Failing? Debunking Performance Myths in Large Language Models
by: Barnett, Scott, et al.
Published: (2024) -
Evaluating LLMs on Document-Based QA: Exact Answer Selection and Numerical Extraction using Cogtale dataset
by: Rasool, Zafaryab, et al.
Published: (2023)