Semantic Layers for Reliable LLM-Powered Data Analytics: A Paired Benchmark of Accuracy and Hallucination Across Three Frontier Models
Fuente:
arXiv
Saved in:
| Main Authors: | Rumiantsau, Michael, Fokeev, Ivan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Fine-Tuning: Effective Strategies for Mitigating Hallucinations in Large Language Models for Data Analytics
by: Rumiantsau, Mikhail, et al.
Published: (2024)
by: Rumiantsau, Mikhail, et al.
Published: (2024)
Hybrid LLM/Rule-based Approaches to Business Insights Generation from Structured Data
by: Vertsel, Aliaksei, et al.
Published: (2024)
by: Vertsel, Aliaksei, et al.
Published: (2024)
Reallocating Attention Across Layers to Reduce Multimodal Hallucination
by: Lu, Haolang, et al.
Published: (2025)
by: Lu, Haolang, et al.
Published: (2025)
LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient
by: Yuan, Peiwen, et al.
Published: (2025)
by: Yuan, Peiwen, et al.
Published: (2025)
Cut Costs, Not Accuracy: LLM-Powered Data Processing with Guarantees
by: Zeighami, Sepanta, et al.
Published: (2025)
by: Zeighami, Sepanta, et al.
Published: (2025)
HalluLens: LLM Hallucination Benchmark
by: Bang, Yejin, et al.
Published: (2025)
by: Bang, Yejin, et al.
Published: (2025)
From Dispersion to Attraction: Spectral Dynamics of Hallucination Across Whisper Model Scales
by: Viakhirev, Ivan, et al.
Published: (2026)
by: Viakhirev, Ivan, et al.
Published: (2026)
LLM-Powered Swarms: A New Frontier or a Conceptual Stretch?
by: Rahman, Muhammad Atta Ur, et al.
Published: (2025)
by: Rahman, Muhammad Atta Ur, et al.
Published: (2025)
HYPE-EDIT-1: Benchmark for Measuring Reliability in Frontier Image Editing Models
by: Chan, Wing, et al.
Published: (2026)
by: Chan, Wing, et al.
Published: (2026)
AIDABench: AI Data Analytics Benchmark
by: Yang, Yibo, et al.
Published: (2026)
by: Yang, Yibo, et al.
Published: (2026)
LLM-Powered Knowledge Graphs for Enterprise Intelligence and Analytics
by: Kumar, Rajeev, et al.
Published: (2025)
by: Kumar, Rajeev, et al.
Published: (2025)
HuDEx: Integrating Hallucination Detection and Explainability for Enhancing the Reliability of LLM responses
by: Lee, Sujeong, et al.
Published: (2025)
by: Lee, Sujeong, et al.
Published: (2025)
ANAH: Analytical Annotation of Hallucinations in Large Language Models
by: Ji, Ziwei, et al.
Published: (2024)
by: Ji, Ziwei, et al.
Published: (2024)
Powering In-Database Dynamic Model Slicing for Structured Data Analytics
by: Zeng, Lingze, et al.
Published: (2024)
by: Zeng, Lingze, et al.
Published: (2024)
Ensuring Reliability of Curated EHR-Derived Data: The Validation of Accuracy for LLM/ML-Extracted Information and Data (VALID) Framework
by: Estevez, Melissa, et al.
Published: (2025)
by: Estevez, Melissa, et al.
Published: (2025)
How Does Thinking Mode Change LLM Moral Judgments? A Controlled Instant-vs-Thinking Comparison Across Five Frontier Models
by: Madur, Sai Sourabh
Published: (2026)
by: Madur, Sai Sourabh
Published: (2026)
VERSA: Verified Event Data Format for Reliable Soccer Analytics
by: Jo, Geonhee, et al.
Published: (2026)
by: Jo, Geonhee, et al.
Published: (2026)
Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment
by: Yeh, Samuel, et al.
Published: (2025)
by: Yeh, Samuel, et al.
Published: (2025)
Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges
by: Weng, Shihao, et al.
Published: (2026)
by: Weng, Shihao, et al.
Published: (2026)
Causely: A Causal Intelligence Layer for Enterprise AI A Benchmark Study on SRE and Reliability Workflows
by: Dalal, Dhairya, et al.
Published: (2026)
by: Dalal, Dhairya, et al.
Published: (2026)
CoddLLM: Empowering Large Language Models for Data Analytics
by: Zhang, Jiani, et al.
Published: (2025)
by: Zhang, Jiani, et al.
Published: (2025)
Mitigating Hallucinations in Large Language Models Via Decoder Layer Skipping
by: Li, Hanze, et al.
Published: (2026)
by: Li, Hanze, et al.
Published: (2026)
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
by: Tang, Zhenwei, et al.
Published: (2025)
by: Tang, Zhenwei, et al.
Published: (2025)
Rethinking Evaluation for LLM Hallucination Detection: A Desiderata, A New RAG-based Benchmark, New Insights
by: Chen, Wenbo, et al.
Published: (2026)
by: Chen, Wenbo, et al.
Published: (2026)
LLM-as-a-Judge for Scalable Test Coverage Evaluation: Accuracy, Operational Reliability, and Cost
by: Huang, Donghao, et al.
Published: (2025)
by: Huang, Donghao, et al.
Published: (2025)
Towards Data Governance of Frontier AI Models
by: Hausenloy, Jason, et al.
Published: (2024)
by: Hausenloy, Jason, et al.
Published: (2024)
ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding
by: Lu, Hao, et al.
Published: (2025)
by: Lu, Hao, et al.
Published: (2025)
Hallucination Detection: Robustly Discerning Reliable Answers in Large Language Models
by: Chen, Yuyan, et al.
Published: (2024)
by: Chen, Yuyan, et al.
Published: (2024)
OAEI-LLM: A Benchmark Dataset for Understanding Large Language Model Hallucinations in Ontology Matching
by: Qiang, Zhangcheng, et al.
Published: (2024)
by: Qiang, Zhangcheng, et al.
Published: (2024)
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
by: Ramesh, Shyam Sundhar, et al.
Published: (2026)
by: Ramesh, Shyam Sundhar, et al.
Published: (2026)
Surfacing Semantic Orthogonality Across Model Safety Benchmarks: A Multi-Dimensional Analysis
by: Bennion, Jonathan, et al.
Published: (2025)
by: Bennion, Jonathan, et al.
Published: (2025)
Enhancing Uncertainty Modeling with Semantic Graph for Hallucination Detection
by: Chen, Kedi, et al.
Published: (2025)
by: Chen, Kedi, et al.
Published: (2025)
Looking Beyond Accuracy: A Holistic Benchmark of ECG Foundation Models
by: Filice, Francesca, et al.
Published: (2026)
by: Filice, Francesca, et al.
Published: (2026)
HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations
by: Abdaljalil, Samir, et al.
Published: (2025)
by: Abdaljalil, Samir, et al.
Published: (2025)
Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment
by: Atasoy, I. F., et al.
Published: (2026)
by: Atasoy, I. F., et al.
Published: (2026)
Luna: An Evaluation Foundation Model to Catch Language Model Hallucinations with High Accuracy and Low Cost
by: Belyi, Masha, et al.
Published: (2024)
by: Belyi, Masha, et al.
Published: (2024)
Hallucination Detection with the Internal Layers of LLMs
by: Preiß, Martin
Published: (2025)
by: Preiß, Martin
Published: (2025)
Layers at Similar Depths Generate Similar Activations Across LLM Architectures
by: Wolfram, Christopher, et al.
Published: (2025)
by: Wolfram, Christopher, et al.
Published: (2025)
CausalFlip: A Benchmark for LLM Causal Judgment Beyond Semantic Matching
by: Wang, Yuzhe, et al.
Published: (2026)
by: Wang, Yuzhe, et al.
Published: (2026)
Beyond Accuracy: Risk-Sensitive Evaluation of Hallucinated Medical Advice
by: Doshi, Savan
Published: (2026)
by: Doshi, Savan
Published: (2026)
Similar Items
-
Beyond Fine-Tuning: Effective Strategies for Mitigating Hallucinations in Large Language Models for Data Analytics
by: Rumiantsau, Mikhail, et al.
Published: (2024) -
Hybrid LLM/Rule-based Approaches to Business Insights Generation from Structured Data
by: Vertsel, Aliaksei, et al.
Published: (2024) -
Reallocating Attention Across Layers to Reduce Multimodal Hallucination
by: Lu, Haolang, et al.
Published: (2025) -
LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient
by: Yuan, Peiwen, et al.
Published: (2025) -
Cut Costs, Not Accuracy: LLM-Powered Data Processing with Guarantees
by: Zeighami, Sepanta, et al.
Published: (2025)