Diagnosing and Addressing Pitfalls in KG-RAG Datasets: Toward More Reliable Benchmarking
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Liangliang, Jiang, Zhuorui, Chi, Hongliang, Chen, Haoyang, Elkoumy, Mohammed, Wang, Fali, Wu, Qiong, Zhou, Zhengyi, Pan, Shirui, Wang, Suhang, Ma, Yao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Distribution Consistency based Self-Training for Graph Neural Networks with Sparse Labels
by: Wang, Fali, et al.
Published: (2024)
by: Wang, Fali, et al.
Published: (2024)
Deployment-Time Reliability of Learned Robot Policies
by: Agia, Christopher
Published: (2026)
by: Agia, Christopher
Published: (2026)
SALT-KG: A Benchmark for Semantics-Aware Learning on Enterprise Tables
by: Mulang, Isaiah Onando, et al.
Published: (2026)
by: Mulang, Isaiah Onando, et al.
Published: (2026)
Rethinking Addressing in Language Models via Contexualized Equivariant Positional Encoding
by: Zhu, Jiajun, et al.
Published: (2025)
by: Zhu, Jiajun, et al.
Published: (2025)
FATHOMS-RAG: A Framework for the Assessment of Thinking and Observation in Multimodal Systems that use Retrieval Augmented Generation
by: Hildebrand, Samuel, et al.
Published: (2025)
by: Hildebrand, Samuel, et al.
Published: (2025)
A Graph-based RAG for Energy Efficiency Question Answering
by: Campi, Riccardo, et al.
Published: (2025)
by: Campi, Riccardo, et al.
Published: (2025)
TrafficRAG: A Multimodal RAG Framework for Traffic Accident Liability Determination
by: Li, Xu, et al.
Published: (2026)
by: Li, Xu, et al.
Published: (2026)
CoMoCAVs: Cohesive Decision-Guided Motion Planning for Connected and Autonomous Vehicles with Multi-Policy Reinforcement Learning
by: Hu, Pan
Published: (2025)
by: Hu, Pan
Published: (2025)
VulScribeR: Exploring RAG-based Vulnerability Augmentation with LLMs
by: Daneshvar, Seyed Shayan, et al.
Published: (2024)
by: Daneshvar, Seyed Shayan, et al.
Published: (2024)
ARCTraj: A Dataset and Benchmark of Human Reasoning Trajectories for Abstract Problem Solving
by: Kim, Sejin, et al.
Published: (2025)
by: Kim, Sejin, et al.
Published: (2025)
ContextBench: A Benchmark for Context Retrieval in Coding Agents
by: Li, Han, et al.
Published: (2026)
by: Li, Han, et al.
Published: (2026)
From Search to Reasoning: A Five-Level RAG Capability Framework for Enterprise Data
by: Gill, Gurbinder, et al.
Published: (2025)
by: Gill, Gurbinder, et al.
Published: (2025)
Ice Cream Doesn't Cause Drowning: Benchmarking LLMs Against Statistical Pitfalls in Causal Inference
by: Du, Jin, et al.
Published: (2025)
by: Du, Jin, et al.
Published: (2025)
Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset
by: Palit, Sayon, et al.
Published: (2025)
by: Palit, Sayon, et al.
Published: (2025)
Primary Care Diagnoses as a Reliable Predictor for Orthopedic Surgical Interventions
by: Verma, Khushboo, et al.
Published: (2025)
by: Verma, Khushboo, et al.
Published: (2025)
PilotBench: A Benchmark for General Aviation Agents with Safety Constraints
by: Wu, Yalun, et al.
Published: (2026)
by: Wu, Yalun, et al.
Published: (2026)
ETOM: A Five-Level Benchmark for Evaluating Tool Orchestration within the MCP Ecosystem
by: Dong, Jia-Kai, et al.
Published: (2025)
by: Dong, Jia-Kai, et al.
Published: (2025)
eStonefish-Scenes: A Sim-to-Real Validated and Robot-Centric Event-based Optical Flow Dataset for Underwater Vehicles
by: Mansour, Jad, et al.
Published: (2025)
by: Mansour, Jad, et al.
Published: (2025)
N-Agent Ad Hoc Teamwork
by: Wang, Caroline, et al.
Published: (2024)
by: Wang, Caroline, et al.
Published: (2024)
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
by: Chehbouni, Khaoula, et al.
Published: (2025)
by: Chehbouni, Khaoula, et al.
Published: (2025)
TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation
by: Du, Bangde, et al.
Published: (2025)
by: Du, Bangde, et al.
Published: (2025)
T-HITL Effectively Addresses Problematic Associations in Image Generation and Maintains Overall Visual Quality
by: Epstein, Susan, et al.
Published: (2024)
by: Epstein, Susan, et al.
Published: (2024)
ChronoMedKG: A Temporally-Grounded Biomedical Knowledge Graph and Benchmark for Clinical Reasoning
by: Ahmed, Md Shamim, et al.
Published: (2026)
by: Ahmed, Md Shamim, et al.
Published: (2026)
Beyond Rating: A Comprehensive Evaluation and Benchmark for AI Reviews
by: Li, Bowen, et al.
Published: (2026)
by: Li, Bowen, et al.
Published: (2026)
Language Models, Graph Searching, and Supervision Adulteration: When More Supervision is Less and How to Make More More
by: Frydenlund, Arvid
Published: (2025)
by: Frydenlund, Arvid
Published: (2025)
RAPTOR-AI for Disaster OODA Loop: Hierarchical Multimodal RAG with Experience-Driven Agentic Decision-Making
by: Yasuno, Takato
Published: (2026)
by: Yasuno, Takato
Published: (2026)
Multi-Paradigm Agent Interaction in Practice:A Systematic Analysis of Generator-Evaluator, ReAct Loop,and Adversarial Evaluation in the buddyMe Framework
by: Wang, Xiaohua, et al.
Published: (2026)
by: Wang, Xiaohua, et al.
Published: (2026)
Less is More: Learning Graph Tasks with Just LLMs
by: Shirai, Sola, et al.
Published: (2025)
by: Shirai, Sola, et al.
Published: (2025)
Learning Natural Language Constraints for Safe Reinforcement Learning of Language Agents
by: Chua, Jaymari, et al.
Published: (2025)
by: Chua, Jaymari, et al.
Published: (2025)
Benchmarking Chest X-ray Diagnosis Models Across Multinational Datasets
by: Xu, Qinmei, et al.
Published: (2025)
by: Xu, Qinmei, et al.
Published: (2025)
FCBV-Net: Category-Level Robotic Garment Smoothing via Feature-Conditioned Bimanual Value Prediction
by: Daba, Mohammed, et al.
Published: (2025)
by: Daba, Mohammed, et al.
Published: (2025)
Advancing Multimodal Agent Reasoning with Long-Term Neuro-Symbolic Memory
by: Jiang, Rongjie, et al.
Published: (2026)
by: Jiang, Rongjie, et al.
Published: (2026)
Learning Can Converge Stably to the Wrong Belief under Latent Reliability
by: Zhang, Zhipeng, et al.
Published: (2026)
by: Zhang, Zhipeng, et al.
Published: (2026)
Towards Explainable and Reliable AI in Finance
by: Isufaj, Albi, et al.
Published: (2025)
by: Isufaj, Albi, et al.
Published: (2025)
The Arrival of AGI? When Expert Personas Exceed Expert Benchmarks
by: Mullens, Drake, et al.
Published: (2026)
by: Mullens, Drake, et al.
Published: (2026)
Precision at Scale: Domain-Specific Datasets On-Demand
by: Rodríguez-de-Vera, Jesús M, et al.
Published: (2024)
by: Rodríguez-de-Vera, Jesús M, et al.
Published: (2024)
Ultra-Reduced-Impact-Encased-Logging (URIEL): propose a new method for selective sustainable logging and post-harvest silvicultural treatment in tropical forest using airborne robotics systems
by: Albiero, Daniel, et al.
Published: (2026)
by: Albiero, Daniel, et al.
Published: (2026)
Bi3: A Biplatform, Bicultural, Biperson Dataset for Social Robot Navigation
by: Stratton, Andrew, et al.
Published: (2026)
by: Stratton, Andrew, et al.
Published: (2026)
WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents
by: Liu, Bingnan, et al.
Published: (2026)
by: Liu, Bingnan, et al.
Published: (2026)
VeriPlan: Integrating Formal Verification and LLMs into End-User Planning
by: Lee, Christine, et al.
Published: (2025)
by: Lee, Christine, et al.
Published: (2025)
Similar Items
-
Distribution Consistency based Self-Training for Graph Neural Networks with Sparse Labels
by: Wang, Fali, et al.
Published: (2024) -
Deployment-Time Reliability of Learned Robot Policies
by: Agia, Christopher
Published: (2026) -
SALT-KG: A Benchmark for Semantics-Aware Learning on Enterprise Tables
by: Mulang, Isaiah Onando, et al.
Published: (2026) -
Rethinking Addressing in Language Models via Contexualized Equivariant Positional Encoding
by: Zhu, Jiajun, et al.
Published: (2025) -
FATHOMS-RAG: A Framework for the Assessment of Thinking and Observation in Multimodal Systems that use Retrieval Augmented Generation
by: Hildebrand, Samuel, et al.
Published: (2025)