InspectorRAGet: An Introspection Platform for RAG Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Fadnis, Kshitij, Patel, Siva Sankalp, Boni, Odellia, Katsis, Yannis, Rosenthal, Sara, Sznajder, Benjamin, Danilevsky, Marina |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Longitudinal Study on Different Annotator Feedback Loops in Complex RAG Tasks
by: Rosenthal, Sara, et al.
Published: (2025)
by: Rosenthal, Sara, et al.
Published: (2025)
RAGAPHENE: A RAG Annotation Platform with Human Enhancements and Edits
by: Fadnis, Kshitij, et al.
Published: (2025)
by: Fadnis, Kshitij, et al.
Published: (2025)
Evaluating Privacy Perceptions, Experience, and Behavior of Software Development Teams
by: Prybylo, Maxwell, et al.
Published: (2024)
by: Prybylo, Maxwell, et al.
Published: (2024)
RiskRAG: A Data-Driven Solution for Improved AI Model Risk Reporting
by: Rao, Pooja S. B., et al.
Published: (2025)
by: Rao, Pooja S. B., et al.
Published: (2025)
Leveraging Large Language Models to Enhance Domain Expert Inclusion in Data Science Workflows
by: Shih, Jasmine Y., et al.
Published: (2024)
by: Shih, Jasmine Y., et al.
Published: (2024)
A Unified, Cross-Platform Framework for Automatic GUI and Plugin Generation in Structural Bioinformatics and Beyond
by: Guo, Sikao, et al.
Published: (2026)
by: Guo, Sikao, et al.
Published: (2026)
"Always Nice and Confident, Sometimes Wrong": Developer's Experiences Engaging Large Language Models (LLMs) Versus Human-Powered Q&A Platforms for Coding Support
by: Li, Jiachen, et al.
Published: (2023)
by: Li, Jiachen, et al.
Published: (2023)
Crowdsourcing: A Framework for Usability Evaluation
by: Nasir, Muhammad
Published: (2024)
by: Nasir, Muhammad
Published: (2024)
V-SHiNE: A Virtual Smart Home Framework for Explainability Evaluation
by: Sadeghi, Mersedeh, et al.
Published: (2026)
by: Sadeghi, Mersedeh, et al.
Published: (2026)
An open-source Modular Online Psychophysics Platform (MOPP)
by: Samoilov-Kats, Yuval, et al.
Published: (2025)
by: Samoilov-Kats, Yuval, et al.
Published: (2025)
StartFlow: From Method Conception to Multi-Perspective Evaluation in UX Prototyping for Software Startups
by: Guerino, Guilherme Corredato, et al.
Published: (2026)
by: Guerino, Guilherme Corredato, et al.
Published: (2026)
GazeCopilot: Evaluating Novel Gaze-Informed Prompting for AI-Supported Code Comprehension and Readability
by: Elfares, Yasmine, et al.
Published: (2025)
by: Elfares, Yasmine, et al.
Published: (2025)
Gendered Prompting and LLM Code Review: How Gender Cues in the Prompt Shape Code Quality and Evaluation
by: Janzen, Lynn, et al.
Published: (2026)
by: Janzen, Lynn, et al.
Published: (2026)
Evaluating Citizen Satisfaction with Saudi Arabia's E-Government Services: A Standards-Based, Theory-Informed Approach
by: Alannsary, Mohammed O.
Published: (2025)
by: Alannsary, Mohammed O.
Published: (2025)
The Impact of AI-Assisted Development on Software Security: A Study of Gemini and Developer Experience
by: Jost, Nadine, et al.
Published: (2026)
by: Jost, Nadine, et al.
Published: (2026)
Elderly HealthMag: Systematic Building and Calibrating a Tool for Identifying and Evaluating Senior User Digital Health Software
by: Xiao, Yuqing, et al.
Published: (2026)
by: Xiao, Yuqing, et al.
Published: (2026)
MTRAG-UN: A Benchmark for Open Challenges in Multi-Turn RAG Conversations
by: Rosenthal, Sara, et al.
Published: (2026)
by: Rosenthal, Sara, et al.
Published: (2026)
RECAP: An End-to-End Platform for Capturing, Replaying, and Analyzing AI-Assisted Programming Interactions
by: He, Keyu, et al.
Published: (2026)
by: He, Keyu, et al.
Published: (2026)
Accodemy: AI Powered Code Learning Platform to Assist Novice Programmers in Overcoming the Fear of Coding
by: Aamina, M. A. F., et al.
Published: (2025)
by: Aamina, M. A. F., et al.
Published: (2025)
MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems
by: Katsis, Yannis, et al.
Published: (2025)
by: Katsis, Yannis, et al.
Published: (2025)
Deployment-Relevant Alignment Cannot Be Inferred from Model-Level Evaluation Alone
by: Vishwarupe, Varad, et al.
Published: (2026)
by: Vishwarupe, Varad, et al.
Published: (2026)
AI-Powered Multi-Stakeholder Ecosystems for Global Development: A Design Research Study on the GSI D-Hub Proof-of-Concept Platform
by: Mohammed, Muzakkiruddin Ahmed, et al.
Published: (2026)
by: Mohammed, Muzakkiruddin Ahmed, et al.
Published: (2026)
Sources of Underproduction in Open Source Software
by: Champion, Kaylea, et al.
Published: (2024)
by: Champion, Kaylea, et al.
Published: (2024)
Qualitative Evaluation of LLM-Designed GUI
by: Sawicki, Bartosz, et al.
Published: (2026)
by: Sawicki, Bartosz, et al.
Published: (2026)
Story Arena: A Multi-Agent Environment for Envisioning the Future of Software Engineering
by: Weisz, Justin D., et al.
Published: (2025)
by: Weisz, Justin D., et al.
Published: (2025)
To be FAIR or RIGHT? Methodological [R]esearch [I]ntegrity [G]iven [H]uman-facing [T]echnologies using the example of Learning Technologies
by: Dehne, Julian
Published: (2026)
by: Dehne, Julian
Published: (2026)
Linting Style and Substance in READMEs
by: Mynampaty, Hima, et al.
Published: (2026)
by: Mynampaty, Hima, et al.
Published: (2026)
MLLM-Based UI2Code Automation Guided by UI Layout Information
by: Wu, Fan, et al.
Published: (2025)
by: Wu, Fan, et al.
Published: (2025)
Shelter Soul: Bridging Shelters and Adopters Through Technology
by: Jagtap, Yashodip Dharmendra
Published: (2025)
by: Jagtap, Yashodip Dharmendra
Published: (2025)
Automatic Bias Detection in Source Code Review
by: Alebachew, Yoseph Berhanu, et al.
Published: (2025)
by: Alebachew, Yoseph Berhanu, et al.
Published: (2025)
CoVoL: A Cooperative Vocabulary Learning Game for Children with Autism
by: Chodkiewicz, Pawel, et al.
Published: (2025)
by: Chodkiewicz, Pawel, et al.
Published: (2025)
Systematic Literature Review of Automation and Artificial Intelligence in Usability Issue Detection
by: Kuric, Eduard, et al.
Published: (2025)
by: Kuric, Eduard, et al.
Published: (2025)
Democratizing AI Development: Local LLM Deployment for India's Developer Ecosystem in the Era of Tokenized APIs
by: Udandarao, Vikranth, et al.
Published: (2025)
by: Udandarao, Vikranth, et al.
Published: (2025)
Emotional Contagion in Code: How GitHub Emoji Reactions Shape Developer Collaboration
by: Kraishan, Obada
Published: (2025)
by: Kraishan, Obada
Published: (2025)
On the Role and Impact of GenAI Tools in Software Engineering Education
by: Qin, Qiaolin, et al.
Published: (2025)
by: Qin, Qiaolin, et al.
Published: (2025)
From Gains to Strains: Modeling Developer Burnout with GenAI Adoption
by: Feng, Zixuan, et al.
Published: (2025)
by: Feng, Zixuan, et al.
Published: (2025)
Beyond Banning AI: A First Look at GenAI Governance in Open Source Software Communities
by: Yang, Wenhao, et al.
Published: (2026)
by: Yang, Wenhao, et al.
Published: (2026)
An Empirical Investigation on the Challenges Faced by Women in the Software Industry: A Case Study
by: Trinkenreich, Bianca, et al.
Published: (2022)
by: Trinkenreich, Bianca, et al.
Published: (2022)
SmartEx: A Framework for Generating User-Centric Explanations in Smart Environments
by: Sadeghi, Mersedeh, et al.
Published: (2024)
by: Sadeghi, Mersedeh, et al.
Published: (2024)
Anteater: Interactive Visualization of Program Execution Values in Context
by: Faust, Rebecca, et al.
Published: (2019)
by: Faust, Rebecca, et al.
Published: (2019)
Similar Items
-
A Longitudinal Study on Different Annotator Feedback Loops in Complex RAG Tasks
by: Rosenthal, Sara, et al.
Published: (2025) -
RAGAPHENE: A RAG Annotation Platform with Human Enhancements and Edits
by: Fadnis, Kshitij, et al.
Published: (2025) -
Evaluating Privacy Perceptions, Experience, and Behavior of Software Development Teams
by: Prybylo, Maxwell, et al.
Published: (2024) -
RiskRAG: A Data-Driven Solution for Improved AI Model Risk Reporting
by: Rao, Pooja S. B., et al.
Published: (2025) -
Leveraging Large Language Models to Enhance Domain Expert Inclusion in Data Science Workflows
by: Shih, Jasmine Y., et al.
Published: (2024)