Scalable and Reliable Evaluation of AI Knowledge Retrieval Systems: RIKER and the Coherent Simulated Universe
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Roig, JV |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
How Much Do LLMs Hallucinate in Document Q&A Scenarios? A 172-Billion-Token Study Across Temperatures, Context Lengths, and Hardware Platforms
par: Roig, JV
Publié: (2026)
par: Roig, JV
Publié: (2026)
Towards a Standard, Enterprise-Relevant Agentic AI Benchmark: Lessons from 5.5 billion tokens' worth of agentic AI evaluations
par: Roig, JV
Publié: (2025)
par: Roig, JV
Publié: (2025)
How Do LLMs Fail In Agentic Scenarios? A Qualitative Analysis of Success and Failure Scenarios of Various LLMs in Agentic Simulations
par: Roig, JV
Publié: (2025)
par: Roig, JV
Publié: (2025)
GraphMERT: Efficient and Scalable Distillation of Reliable Knowledge Graphs from Unstructured Data
par: Belova, Margarita, et autres
Publié: (2025)
par: Belova, Margarita, et autres
Publié: (2025)
Improving Coherence and Persistence in Agentic AI for System Optimization
par: Karimi, Pantea, et autres
Publié: (2026)
par: Karimi, Pantea, et autres
Publié: (2026)
Reliable Evaluation Protocol for Low-Precision Retrieval
par: Yang, Kisu, et autres
Publié: (2025)
par: Yang, Kisu, et autres
Publié: (2025)
AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models
par: Jackson, Declan, et autres
Publié: (2025)
par: Jackson, Declan, et autres
Publié: (2025)
Automated Rubrics for Reliable Evaluation of Medical Dialogue Systems
par: Chen, Yinzhu, et autres
Publié: (2026)
par: Chen, Yinzhu, et autres
Publié: (2026)
LiveCLKTBench: Towards Reliable Evaluation of Cross-Lingual Knowledge Transfer in Multilingual LLMs
par: Guo, Pei-Fu, et autres
Publié: (2025)
par: Guo, Pei-Fu, et autres
Publié: (2025)
Knowledge Tagging System on Math Questions via LLMs with Flexible Demonstration Retriever
par: Li, Hang, et autres
Publié: (2024)
par: Li, Hang, et autres
Publié: (2024)
Bridging Legal Knowledge and AI: Retrieval-Augmented Generation with Vector Stores, Knowledge Graphs, and Hierarchical Non-negative Matrix Factorization
par: Barron, Ryan C., et autres
Publié: (2025)
par: Barron, Ryan C., et autres
Publié: (2025)
SkillFlow: Scalable and Efficient Agent Skill Retrieval System
par: Li, Fangzhou, et autres
Publié: (2025)
par: Li, Fangzhou, et autres
Publié: (2025)
CoreEval: Automatically Building Contamination-Resilient Datasets with Real-World Knowledge toward Reliable LLM Evaluation
par: Zhao, Jingqian, et autres
Publié: (2025)
par: Zhao, Jingqian, et autres
Publié: (2025)
Knowledge-Graph Based RAG System Evaluation Framework
par: Dong, Sicheng, et autres
Publié: (2025)
par: Dong, Sicheng, et autres
Publié: (2025)
Retrieval-Augmented Generation with Hierarchical Knowledge
par: Huang, Haoyu, et autres
Publié: (2025)
par: Huang, Haoyu, et autres
Publié: (2025)
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence
par: Venkit, Pranav Narayanan, et autres
Publié: (2025)
par: Venkit, Pranav Narayanan, et autres
Publié: (2025)
Uncertainty-Aware Dynamic Knowledge Graphs for Reliable Question Answering
par: Takahashi, Yu, et autres
Publié: (2025)
par: Takahashi, Yu, et autres
Publié: (2025)
Scientific Knowledge-driven Decoding Constraints Improving the Reliability of LLMs
par: Ma, Maotian, et autres
Publié: (2026)
par: Ma, Maotian, et autres
Publié: (2026)
R-Eval: A Unified Toolkit for Evaluating Domain Knowledge of Retrieval Augmented Large Language Models
par: Tu, Shangqing, et autres
Publié: (2024)
par: Tu, Shangqing, et autres
Publié: (2024)
Knowledge Graph-Guided Retrieval Augmented Generation
par: Zhu, Xiangrong, et autres
Publié: (2025)
par: Zhu, Xiangrong, et autres
Publié: (2025)
QQJ: Quantifying Qualitative Judgment for Scalable and Human-Aligned Evaluation of Generative AI
par: Veysi, Marjan, et autres
Publié: (2026)
par: Veysi, Marjan, et autres
Publié: (2026)
Retrieval-Augmented Machine Translation with Unstructured Knowledge
par: Wang, Jiaan, et autres
Publié: (2024)
par: Wang, Jiaan, et autres
Publié: (2024)
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
par: Lunardi, Riccardo, et autres
Publié: (2025)
par: Lunardi, Riccardo, et autres
Publié: (2025)
VERITAS: A Unified Approach to Reliability Evaluation
par: Ramamurthy, Rajkumar, et autres
Publié: (2024)
par: Ramamurthy, Rajkumar, et autres
Publié: (2024)
Bias-Aware Agent: Enhancing Fairness in AI-Driven Knowledge Retrieval
par: Singh, Karanbir, et autres
Publié: (2025)
par: Singh, Karanbir, et autres
Publié: (2025)
Knowledge-tuning Large Language Models with Structured Medical Knowledge Bases for Reliable Response Generation in Chinese
par: Wang, Haochun, et autres
Publié: (2023)
par: Wang, Haochun, et autres
Publié: (2023)
How Reliable are LLMs as Knowledge Bases? Re-thinking Facutality and Consistency
par: Zheng, Danna, et autres
Publié: (2024)
par: Zheng, Danna, et autres
Publié: (2024)
AI-Assisted Systematization for Evaluating GenAI Systems
par: Agarwal, Dhruv, et autres
Publié: (2026)
par: Agarwal, Dhruv, et autres
Publié: (2026)
Evaluating Reliability Asymmetries in Chinese Factual Search and AI Answers
par: Liu, Geng, et autres
Publié: (2025)
par: Liu, Geng, et autres
Publié: (2025)
Energy Landscapes Enable Reliable Abstention in Retrieval-Augmented Large Language Models for Healthcare
par: Shankar, Ravi, et autres
Publié: (2025)
par: Shankar, Ravi, et autres
Publié: (2025)
ASTRID -- An Automated and Scalable TRIaD for the Evaluation of RAG-based Clinical Question Answering Systems
par: Chowdhury, Mohita, et autres
Publié: (2025)
par: Chowdhury, Mohita, et autres
Publié: (2025)
Reliable, Adaptable, and Attributable Language Models with Retrieval
par: Asai, Akari, et autres
Publié: (2024)
par: Asai, Akari, et autres
Publié: (2024)
Evaluation Framework for AI Systems in "the Wild"
par: Jabbour, Sarah, et autres
Publié: (2025)
par: Jabbour, Sarah, et autres
Publié: (2025)
A Survey on Knowledge-Oriented Retrieval-Augmented Generation
par: Cheng, Mingyue, et autres
Publié: (2025)
par: Cheng, Mingyue, et autres
Publié: (2025)
Multilingual Retrieval-Augmented Generation for Knowledge-Intensive Task
par: Ranaldi, Leonardo, et autres
Publié: (2025)
par: Ranaldi, Leonardo, et autres
Publié: (2025)
MERIT: Memory-Enhanced Retrieval for Interpretable Knowledge Tracing
par: Li, Runze, et autres
Publié: (2026)
par: Li, Runze, et autres
Publié: (2026)
MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems
par: Katsis, Yannis, et autres
Publié: (2025)
par: Katsis, Yannis, et autres
Publié: (2025)
Do Repetitions Matter? Strengthening Reliability in LLM Evaluations
par: Gonzalez, Miguel Angel Alvarado, et autres
Publié: (2025)
par: Gonzalez, Miguel Angel Alvarado, et autres
Publié: (2025)
VERT: Reliable LLM Judges for Radiology Report Evaluation
par: Bologna, Federica, et autres
Publié: (2026)
par: Bologna, Federica, et autres
Publié: (2026)
Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
par: Pres, Itamar, et autres
Publié: (2024)
par: Pres, Itamar, et autres
Publié: (2024)
Documents similaires
-
How Much Do LLMs Hallucinate in Document Q&A Scenarios? A 172-Billion-Token Study Across Temperatures, Context Lengths, and Hardware Platforms
par: Roig, JV
Publié: (2026) -
Towards a Standard, Enterprise-Relevant Agentic AI Benchmark: Lessons from 5.5 billion tokens' worth of agentic AI evaluations
par: Roig, JV
Publié: (2025) -
How Do LLMs Fail In Agentic Scenarios? A Qualitative Analysis of Success and Failure Scenarios of Various LLMs in Agentic Simulations
par: Roig, JV
Publié: (2025) -
GraphMERT: Efficient and Scalable Distillation of Reliable Knowledge Graphs from Unstructured Data
par: Belova, Margarita, et autres
Publié: (2025) -
Improving Coherence and Persistence in Agentic AI for System Optimization
par: Karimi, Pantea, et autres
Publié: (2026)