Semantic Layers for Reliable LLM-Powered Data Analytics: A Paired Benchmark of Accuracy and Hallucination Across Three Frontier Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Rumiantsau, Michael, Fokeev, Ivan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Beyond Fine-Tuning: Effective Strategies for Mitigating Hallucinations in Large Language Models for Data Analytics
di: Rumiantsau, Mikhail, et al.
Pubblicazione: (2024)
di: Rumiantsau, Mikhail, et al.
Pubblicazione: (2024)
Hybrid LLM/Rule-based Approaches to Business Insights Generation from Structured Data
di: Vertsel, Aliaksei, et al.
Pubblicazione: (2024)
di: Vertsel, Aliaksei, et al.
Pubblicazione: (2024)
Reallocating Attention Across Layers to Reduce Multimodal Hallucination
di: Lu, Haolang, et al.
Pubblicazione: (2025)
di: Lu, Haolang, et al.
Pubblicazione: (2025)
LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient
di: Yuan, Peiwen, et al.
Pubblicazione: (2025)
di: Yuan, Peiwen, et al.
Pubblicazione: (2025)
Cut Costs, Not Accuracy: LLM-Powered Data Processing with Guarantees
di: Zeighami, Sepanta, et al.
Pubblicazione: (2025)
di: Zeighami, Sepanta, et al.
Pubblicazione: (2025)
HalluLens: LLM Hallucination Benchmark
di: Bang, Yejin, et al.
Pubblicazione: (2025)
di: Bang, Yejin, et al.
Pubblicazione: (2025)
From Dispersion to Attraction: Spectral Dynamics of Hallucination Across Whisper Model Scales
di: Viakhirev, Ivan, et al.
Pubblicazione: (2026)
di: Viakhirev, Ivan, et al.
Pubblicazione: (2026)
LLM-Powered Swarms: A New Frontier or a Conceptual Stretch?
di: Rahman, Muhammad Atta Ur, et al.
Pubblicazione: (2025)
di: Rahman, Muhammad Atta Ur, et al.
Pubblicazione: (2025)
HYPE-EDIT-1: Benchmark for Measuring Reliability in Frontier Image Editing Models
di: Chan, Wing, et al.
Pubblicazione: (2026)
di: Chan, Wing, et al.
Pubblicazione: (2026)
AIDABench: AI Data Analytics Benchmark
di: Yang, Yibo, et al.
Pubblicazione: (2026)
di: Yang, Yibo, et al.
Pubblicazione: (2026)
LLM-Powered Knowledge Graphs for Enterprise Intelligence and Analytics
di: Kumar, Rajeev, et al.
Pubblicazione: (2025)
di: Kumar, Rajeev, et al.
Pubblicazione: (2025)
HuDEx: Integrating Hallucination Detection and Explainability for Enhancing the Reliability of LLM responses
di: Lee, Sujeong, et al.
Pubblicazione: (2025)
di: Lee, Sujeong, et al.
Pubblicazione: (2025)
ANAH: Analytical Annotation of Hallucinations in Large Language Models
di: Ji, Ziwei, et al.
Pubblicazione: (2024)
di: Ji, Ziwei, et al.
Pubblicazione: (2024)
Powering In-Database Dynamic Model Slicing for Structured Data Analytics
di: Zeng, Lingze, et al.
Pubblicazione: (2024)
di: Zeng, Lingze, et al.
Pubblicazione: (2024)
Ensuring Reliability of Curated EHR-Derived Data: The Validation of Accuracy for LLM/ML-Extracted Information and Data (VALID) Framework
di: Estevez, Melissa, et al.
Pubblicazione: (2025)
di: Estevez, Melissa, et al.
Pubblicazione: (2025)
How Does Thinking Mode Change LLM Moral Judgments? A Controlled Instant-vs-Thinking Comparison Across Five Frontier Models
di: Madur, Sai Sourabh
Pubblicazione: (2026)
di: Madur, Sai Sourabh
Pubblicazione: (2026)
VERSA: Verified Event Data Format for Reliable Soccer Analytics
di: Jo, Geonhee, et al.
Pubblicazione: (2026)
di: Jo, Geonhee, et al.
Pubblicazione: (2026)
Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment
di: Yeh, Samuel, et al.
Pubblicazione: (2025)
di: Yeh, Samuel, et al.
Pubblicazione: (2025)
Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges
di: Weng, Shihao, et al.
Pubblicazione: (2026)
di: Weng, Shihao, et al.
Pubblicazione: (2026)
Causely: A Causal Intelligence Layer for Enterprise AI A Benchmark Study on SRE and Reliability Workflows
di: Dalal, Dhairya, et al.
Pubblicazione: (2026)
di: Dalal, Dhairya, et al.
Pubblicazione: (2026)
CoddLLM: Empowering Large Language Models for Data Analytics
di: Zhang, Jiani, et al.
Pubblicazione: (2025)
di: Zhang, Jiani, et al.
Pubblicazione: (2025)
Mitigating Hallucinations in Large Language Models Via Decoder Layer Skipping
di: Li, Hanze, et al.
Pubblicazione: (2026)
di: Li, Hanze, et al.
Pubblicazione: (2026)
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
di: Tang, Zhenwei, et al.
Pubblicazione: (2025)
di: Tang, Zhenwei, et al.
Pubblicazione: (2025)
Rethinking Evaluation for LLM Hallucination Detection: A Desiderata, A New RAG-based Benchmark, New Insights
di: Chen, Wenbo, et al.
Pubblicazione: (2026)
di: Chen, Wenbo, et al.
Pubblicazione: (2026)
LLM-as-a-Judge for Scalable Test Coverage Evaluation: Accuracy, Operational Reliability, and Cost
di: Huang, Donghao, et al.
Pubblicazione: (2025)
di: Huang, Donghao, et al.
Pubblicazione: (2025)
Towards Data Governance of Frontier AI Models
di: Hausenloy, Jason, et al.
Pubblicazione: (2024)
di: Hausenloy, Jason, et al.
Pubblicazione: (2024)
ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding
di: Lu, Hao, et al.
Pubblicazione: (2025)
di: Lu, Hao, et al.
Pubblicazione: (2025)
Hallucination Detection: Robustly Discerning Reliable Answers in Large Language Models
di: Chen, Yuyan, et al.
Pubblicazione: (2024)
di: Chen, Yuyan, et al.
Pubblicazione: (2024)
OAEI-LLM: A Benchmark Dataset for Understanding Large Language Model Hallucinations in Ontology Matching
di: Qiang, Zhangcheng, et al.
Pubblicazione: (2024)
di: Qiang, Zhangcheng, et al.
Pubblicazione: (2024)
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
di: Ramesh, Shyam Sundhar, et al.
Pubblicazione: (2026)
di: Ramesh, Shyam Sundhar, et al.
Pubblicazione: (2026)
Surfacing Semantic Orthogonality Across Model Safety Benchmarks: A Multi-Dimensional Analysis
di: Bennion, Jonathan, et al.
Pubblicazione: (2025)
di: Bennion, Jonathan, et al.
Pubblicazione: (2025)
Enhancing Uncertainty Modeling with Semantic Graph for Hallucination Detection
di: Chen, Kedi, et al.
Pubblicazione: (2025)
di: Chen, Kedi, et al.
Pubblicazione: (2025)
Looking Beyond Accuracy: A Holistic Benchmark of ECG Foundation Models
di: Filice, Francesca, et al.
Pubblicazione: (2026)
di: Filice, Francesca, et al.
Pubblicazione: (2026)
HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations
di: Abdaljalil, Samir, et al.
Pubblicazione: (2025)
di: Abdaljalil, Samir, et al.
Pubblicazione: (2025)
Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment
di: Atasoy, I. F., et al.
Pubblicazione: (2026)
di: Atasoy, I. F., et al.
Pubblicazione: (2026)
Luna: An Evaluation Foundation Model to Catch Language Model Hallucinations with High Accuracy and Low Cost
di: Belyi, Masha, et al.
Pubblicazione: (2024)
di: Belyi, Masha, et al.
Pubblicazione: (2024)
Hallucination Detection with the Internal Layers of LLMs
di: Preiß, Martin
Pubblicazione: (2025)
di: Preiß, Martin
Pubblicazione: (2025)
Layers at Similar Depths Generate Similar Activations Across LLM Architectures
di: Wolfram, Christopher, et al.
Pubblicazione: (2025)
di: Wolfram, Christopher, et al.
Pubblicazione: (2025)
CausalFlip: A Benchmark for LLM Causal Judgment Beyond Semantic Matching
di: Wang, Yuzhe, et al.
Pubblicazione: (2026)
di: Wang, Yuzhe, et al.
Pubblicazione: (2026)
Beyond Accuracy: Risk-Sensitive Evaluation of Hallucinated Medical Advice
di: Doshi, Savan
Pubblicazione: (2026)
di: Doshi, Savan
Pubblicazione: (2026)
Documenti analoghi
-
Beyond Fine-Tuning: Effective Strategies for Mitigating Hallucinations in Large Language Models for Data Analytics
di: Rumiantsau, Mikhail, et al.
Pubblicazione: (2024) -
Hybrid LLM/Rule-based Approaches to Business Insights Generation from Structured Data
di: Vertsel, Aliaksei, et al.
Pubblicazione: (2024) -
Reallocating Attention Across Layers to Reduce Multimodal Hallucination
di: Lu, Haolang, et al.
Pubblicazione: (2025) -
LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient
di: Yuan, Peiwen, et al.
Pubblicazione: (2025) -
Cut Costs, Not Accuracy: LLM-Powered Data Processing with Guarantees
di: Zeighami, Sepanta, et al.
Pubblicazione: (2025)