Saved in:
| Main Author: | Roig, JV |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2601.08847 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Much Do LLMs Hallucinate in Document Q&A Scenarios? A 172-Billion-Token Study Across Temperatures, Context Lengths, and Hardware Platforms
by: Roig, JV
Published: (2026)
by: Roig, JV
Published: (2026)
Towards a Standard, Enterprise-Relevant Agentic AI Benchmark: Lessons from 5.5 billion tokens' worth of agentic AI evaluations
by: Roig, JV
Published: (2025)
by: Roig, JV
Published: (2025)
How Do LLMs Fail In Agentic Scenarios? A Qualitative Analysis of Success and Failure Scenarios of Various LLMs in Agentic Simulations
by: Roig, JV
Published: (2025)
by: Roig, JV
Published: (2025)
GraphMERT: Efficient and Scalable Distillation of Reliable Knowledge Graphs from Unstructured Data
by: Belova, Margarita, et al.
Published: (2025)
by: Belova, Margarita, et al.
Published: (2025)
Reliable Evaluation Protocol for Low-Precision Retrieval
by: Yang, Kisu, et al.
Published: (2025)
by: Yang, Kisu, et al.
Published: (2025)
Improving Coherence and Persistence in Agentic AI for System Optimization
by: Karimi, Pantea, et al.
Published: (2026)
by: Karimi, Pantea, et al.
Published: (2026)
AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models
by: Jackson, Declan, et al.
Published: (2025)
by: Jackson, Declan, et al.
Published: (2025)
Automated Rubrics for Reliable Evaluation of Medical Dialogue Systems
by: Chen, Yinzhu, et al.
Published: (2026)
by: Chen, Yinzhu, et al.
Published: (2026)
LiveCLKTBench: Towards Reliable Evaluation of Cross-Lingual Knowledge Transfer in Multilingual LLMs
by: Guo, Pei-Fu, et al.
Published: (2025)
by: Guo, Pei-Fu, et al.
Published: (2025)
SkillFlow: Scalable and Efficient Agent Skill Retrieval System
by: Li, Fangzhou, et al.
Published: (2025)
by: Li, Fangzhou, et al.
Published: (2025)
Knowledge Tagging System on Math Questions via LLMs with Flexible Demonstration Retriever
by: Li, Hang, et al.
Published: (2024)
by: Li, Hang, et al.
Published: (2024)
Bridging Legal Knowledge and AI: Retrieval-Augmented Generation with Vector Stores, Knowledge Graphs, and Hierarchical Non-negative Matrix Factorization
by: Barron, Ryan C., et al.
Published: (2025)
by: Barron, Ryan C., et al.
Published: (2025)
CoreEval: Automatically Building Contamination-Resilient Datasets with Real-World Knowledge toward Reliable LLM Evaluation
by: Zhao, Jingqian, et al.
Published: (2025)
by: Zhao, Jingqian, et al.
Published: (2025)
Retrieval-Augmented Generation with Hierarchical Knowledge
by: Huang, Haoyu, et al.
Published: (2025)
by: Huang, Haoyu, et al.
Published: (2025)
Knowledge-Graph Based RAG System Evaluation Framework
by: Dong, Sicheng, et al.
Published: (2025)
by: Dong, Sicheng, et al.
Published: (2025)
Evaluating Reliability Asymmetries in Chinese Factual Search and AI Answers
by: Liu, Geng, et al.
Published: (2025)
by: Liu, Geng, et al.
Published: (2025)
Bias-Aware Agent: Enhancing Fairness in AI-Driven Knowledge Retrieval
by: Singh, Karanbir, et al.
Published: (2025)
by: Singh, Karanbir, et al.
Published: (2025)
Uncertainty-Aware Dynamic Knowledge Graphs for Reliable Question Answering
by: Takahashi, Yu, et al.
Published: (2025)
by: Takahashi, Yu, et al.
Published: (2025)
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence
by: Venkit, Pranav Narayanan, et al.
Published: (2025)
by: Venkit, Pranav Narayanan, et al.
Published: (2025)
Scientific Knowledge-driven Decoding Constraints Improving the Reliability of LLMs
by: Ma, Maotian, et al.
Published: (2026)
by: Ma, Maotian, et al.
Published: (2026)
R-Eval: A Unified Toolkit for Evaluating Domain Knowledge of Retrieval Augmented Large Language Models
by: Tu, Shangqing, et al.
Published: (2024)
by: Tu, Shangqing, et al.
Published: (2024)
QQJ: Quantifying Qualitative Judgment for Scalable and Human-Aligned Evaluation of Generative AI
by: Veysi, Marjan, et al.
Published: (2026)
by: Veysi, Marjan, et al.
Published: (2026)
Reliable, Adaptable, and Attributable Language Models with Retrieval
by: Asai, Akari, et al.
Published: (2024)
by: Asai, Akari, et al.
Published: (2024)
Knowledge Graph-Guided Retrieval Augmented Generation
by: Zhu, Xiangrong, et al.
Published: (2025)
by: Zhu, Xiangrong, et al.
Published: (2025)
Retrieval-Augmented Machine Translation with Unstructured Knowledge
by: Wang, Jiaan, et al.
Published: (2024)
by: Wang, Jiaan, et al.
Published: (2024)
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
by: Lunardi, Riccardo, et al.
Published: (2025)
by: Lunardi, Riccardo, et al.
Published: (2025)
VERITAS: A Unified Approach to Reliability Evaluation
by: Ramamurthy, Rajkumar, et al.
Published: (2024)
by: Ramamurthy, Rajkumar, et al.
Published: (2024)
AI-Assisted Systematization for Evaluating GenAI Systems
by: Agarwal, Dhruv, et al.
Published: (2026)
by: Agarwal, Dhruv, et al.
Published: (2026)
Knowledge-tuning Large Language Models with Structured Medical Knowledge Bases for Reliable Response Generation in Chinese
by: Wang, Haochun, et al.
Published: (2023)
by: Wang, Haochun, et al.
Published: (2023)
How Reliable are LLMs as Knowledge Bases? Re-thinking Facutality and Consistency
by: Zheng, Danna, et al.
Published: (2024)
by: Zheng, Danna, et al.
Published: (2024)
Evaluation Framework for AI Systems in "the Wild"
by: Jabbour, Sarah, et al.
Published: (2025)
by: Jabbour, Sarah, et al.
Published: (2025)
Energy Landscapes Enable Reliable Abstention in Retrieval-Augmented Large Language Models for Healthcare
by: Shankar, Ravi, et al.
Published: (2025)
by: Shankar, Ravi, et al.
Published: (2025)
ASTRID -- An Automated and Scalable TRIaD for the Evaluation of RAG-based Clinical Question Answering Systems
by: Chowdhury, Mohita, et al.
Published: (2025)
by: Chowdhury, Mohita, et al.
Published: (2025)
A Survey on Knowledge-Oriented Retrieval-Augmented Generation
by: Cheng, Mingyue, et al.
Published: (2025)
by: Cheng, Mingyue, et al.
Published: (2025)
Multilingual Retrieval-Augmented Generation for Knowledge-Intensive Task
by: Ranaldi, Leonardo, et al.
Published: (2025)
by: Ranaldi, Leonardo, et al.
Published: (2025)
MERIT: Memory-Enhanced Retrieval for Interpretable Knowledge Tracing
by: Li, Runze, et al.
Published: (2026)
by: Li, Runze, et al.
Published: (2026)
Do Repetitions Matter? Strengthening Reliability in LLM Evaluations
by: Gonzalez, Miguel Angel Alvarado, et al.
Published: (2025)
by: Gonzalez, Miguel Angel Alvarado, et al.
Published: (2025)
VERT: Reliable LLM Judges for Radiology Report Evaluation
by: Bologna, Federica, et al.
Published: (2026)
by: Bologna, Federica, et al.
Published: (2026)
Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
by: Pres, Itamar, et al.
Published: (2024)
by: Pres, Itamar, et al.
Published: (2024)
Controllable and Reliable Knowledge-Intensive Task-Oriented Conversational Agents with Declarative Genie Worksheets
by: Joshi, Harshit, et al.
Published: (2024)
by: Joshi, Harshit, et al.
Published: (2024)
Similar Items
-
How Much Do LLMs Hallucinate in Document Q&A Scenarios? A 172-Billion-Token Study Across Temperatures, Context Lengths, and Hardware Platforms
by: Roig, JV
Published: (2026) -
Towards a Standard, Enterprise-Relevant Agentic AI Benchmark: Lessons from 5.5 billion tokens' worth of agentic AI evaluations
by: Roig, JV
Published: (2025) -
How Do LLMs Fail In Agentic Scenarios? A Qualitative Analysis of Success and Failure Scenarios of Various LLMs in Agentic Simulations
by: Roig, JV
Published: (2025) -
GraphMERT: Efficient and Scalable Distillation of Reliable Knowledge Graphs from Unstructured Data
by: Belova, Margarita, et al.
Published: (2025) -
Reliable Evaluation Protocol for Low-Precision Retrieval
by: Yang, Kisu, et al.
Published: (2025)