Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Maharaj, Akash V., Arbour, David, Lee, Daniel, Bhattacharya, Uttaran, Rao, Anup, Zane, Austin, Feller, Avi, Qian, Kun, Li, Yunyao
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2504.13924
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915250128814080
author Maharaj, Akash V.
Arbour, David
Lee, Daniel
Bhattacharya, Uttaran
Rao, Anup
Zane, Austin
Feller, Avi
Qian, Kun
Li, Yunyao
author_facet Maharaj, Akash V.
Arbour, David
Lee, Daniel
Bhattacharya, Uttaran
Rao, Anup
Zane, Austin
Feller, Avi
Qian, Kun
Li, Yunyao
contents Enterprise AI Assistants are increasingly deployed in domains where accuracy is paramount, making each erroneous output a potentially significant incident. This paper presents a comprehensive framework for monitoring, benchmarking, and continuously improving such complex, multi-component systems under active development by multiple teams. Our approach encompasses three key elements: (1) a hierarchical ``severity'' framework for incident detection that identifies and categorizes errors while attributing component-specific error rates, facilitating targeted improvements; (2) a scalable and principled methodology for benchmark construction, evaluation, and deployment, designed to accommodate multiple development teams, mitigate overfitting risks, and assess the downstream impact of system modifications; and (3) a continual improvement strategy leveraging multidimensional evaluation, enabling the identification and implementation of diverse enhancement opportunities. By adopting this holistic framework, organizations can systematically enhance the reliability and performance of their AI Assistants, ensuring their efficacy in critical enterprise environments. We conclude by discussing how this multifaceted evaluation approach opens avenues for various classes of enhancements, paving the way for more robust and trustworthy AI systems.
format Preprint
id arxiv_https___arxiv_org_abs_2504_13924
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluation and Incident Prevention in an Enterprise AI Assistant
Maharaj, Akash V.
Arbour, David
Lee, Daniel
Bhattacharya, Uttaran
Rao, Anup
Zane, Austin
Feller, Avi
Qian, Kun
Li, Yunyao
Artificial Intelligence
Computation and Language
Human-Computer Interaction
Enterprise AI Assistants are increasingly deployed in domains where accuracy is paramount, making each erroneous output a potentially significant incident. This paper presents a comprehensive framework for monitoring, benchmarking, and continuously improving such complex, multi-component systems under active development by multiple teams. Our approach encompasses three key elements: (1) a hierarchical ``severity'' framework for incident detection that identifies and categorizes errors while attributing component-specific error rates, facilitating targeted improvements; (2) a scalable and principled methodology for benchmark construction, evaluation, and deployment, designed to accommodate multiple development teams, mitigate overfitting risks, and assess the downstream impact of system modifications; and (3) a continual improvement strategy leveraging multidimensional evaluation, enabling the identification and implementation of diverse enhancement opportunities. By adopting this holistic framework, organizations can systematically enhance the reliability and performance of their AI Assistants, ensuring their efficacy in critical enterprise environments. We conclude by discussing how this multifaceted evaluation approach opens avenues for various classes of enhancements, paving the way for more robust and trustworthy AI systems.
title Evaluation and Incident Prevention in an Enterprise AI Assistant
topic Artificial Intelligence
Computation and Language
Human-Computer Interaction
url https://arxiv.org/abs/2504.13924