Reducing Hallucination in Enterprise AI Workflows via Hybrid Utility Minimum Bayes Risk (HUMBR)

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fang, Chenhao, Mola, Jordi, Harman, Mark, Nawrocki, Jason, Shrivastava, Vaibhav, Cheng, Yue, Shah, Jay Minesh, Zand, Katayoun, Tripathi, Mansi, Pudota, Arya, Becker, Matthew, Robert, Hervé, Gulati, Abhishek
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915934572118016
author Fang, Chenhao
Mola, Jordi
Harman, Mark
Nawrocki, Jason
Shrivastava, Vaibhav
Cheng, Yue
Shah, Jay Minesh
Zand, Katayoun
Tripathi, Mansi
Pudota, Arya
Becker, Matthew
Robert, Hervé
Gulati, Abhishek
author_facet Fang, Chenhao
Mola, Jordi
Harman, Mark
Nawrocki, Jason
Shrivastava, Vaibhav
Cheng, Yue
Shah, Jay Minesh
Zand, Katayoun
Tripathi, Mansi
Pudota, Arya
Becker, Matthew
Robert, Hervé
Gulati, Abhishek
contents Although LLMs drive automation, it is critical to ensure immense consideration for high-stakes enterprise workflows such as those involving legal matters, risk management, and privacy compliance. For Meta, and other organizations like ours, a single hallucinated clause in such high stakes workflows risks material consequences. We show that by framing hallucination mitigation as a Minimum Bayes Risk (MBR) problem, we can dramatically reduce this risk. Specifically, we introduce a Hybrid Utility MBR (HUMBR) framework that synthesizes semantic embedding similarity with lexical precision to identify consensus without ground-truth references, for which we derive rigorous error bounds. We complement this theoretical analysis with a comprehensive empirical evaluation on widely-used public benchmark suites (TruthfulQA and LegalBench) and also real world data from Meta production deployment. The results from our empirical study show that MBR significantly outperforms standard Universal Self-Consistency. Notably, 81% of the pipeline's suggestions were preferred over human-crafted ground truth, and critical recall failures were virtually eliminated.
format Preprint
id arxiv_https___arxiv_org_abs_2604_11141
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Reducing Hallucination in Enterprise AI Workflows via Hybrid Utility Minimum Bayes Risk (HUMBR)
Fang, Chenhao
Mola, Jordi
Harman, Mark
Nawrocki, Jason
Shrivastava, Vaibhav
Cheng, Yue
Shah, Jay Minesh
Zand, Katayoun
Tripathi, Mansi
Pudota, Arya
Becker, Matthew
Robert, Hervé
Gulati, Abhishek
Machine Learning
Cryptography and Security
Although LLMs drive automation, it is critical to ensure immense consideration for high-stakes enterprise workflows such as those involving legal matters, risk management, and privacy compliance. For Meta, and other organizations like ours, a single hallucinated clause in such high stakes workflows risks material consequences. We show that by framing hallucination mitigation as a Minimum Bayes Risk (MBR) problem, we can dramatically reduce this risk. Specifically, we introduce a Hybrid Utility MBR (HUMBR) framework that synthesizes semantic embedding similarity with lexical precision to identify consensus without ground-truth references, for which we derive rigorous error bounds. We complement this theoretical analysis with a comprehensive empirical evaluation on widely-used public benchmark suites (TruthfulQA and LegalBench) and also real world data from Meta production deployment. The results from our empirical study show that MBR significantly outperforms standard Universal Self-Consistency. Notably, 81% of the pipeline's suggestions were preferred over human-crafted ground truth, and critical recall failures were virtually eliminated.
title Reducing Hallucination in Enterprise AI Workflows via Hybrid Utility Minimum Bayes Risk (HUMBR)
topic Machine Learning
Cryptography and Security
url https://arxiv.org/abs/2604.11141