SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios
Fuente:
arXiv
Guardado en:
| Autores principales: | Clark, Jackson, Su, Yiming, Pial, Saad Mohammad Rafid, Tian, Yifang, Gniedziejko, Lily, Jacobsen, Hans-Arno, Chen, Yinfang, Xu, Tianyin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
STRATUS: A Multi-agent System for Autonomous Reliability Engineering of Modern Clouds
por: Chen, Yinfang, et al.
Publicado: (2025)
por: Chen, Yinfang, et al.
Publicado: (2025)
FAME: Failure-Aware Mixture-of-Experts for Message-Level Log Anomaly Detection
por: Wang, Huanchi, et al.
Publicado: (2026)
por: Wang, Huanchi, et al.
Publicado: (2026)
Activation of Carbon─Chlorine Bond at Chlorpromazine for Access to New Derivatives Through Nickel‐Catalyzed Kumada–Tamao–Corriu Cross‐Coupling Reaction
por: Rafid Saad Dawood
Publicado: (2025)
por: Rafid Saad Dawood
Publicado: (2025)
GALA: Can Graph-Augmented Large Language Model Agentic Workflows Elevate Root Cause Analysis?
por: Tian, Yifang, et al.
Publicado: (2025)
por: Tian, Yifang, et al.
Publicado: (2025)
Pricing and Competition for Generative AI
por: Mahmood, Rafid
Publicado: (2024)
por: Mahmood, Rafid
Publicado: (2024)
Self-attention-based Diffusion Model for Time-series Imputation in Partial Blackout Scenarios
por: Islam, Mohammad Rafid Ul, et al.
Publicado: (2025)
por: Islam, Mohammad Rafid Ul, et al.
Publicado: (2025)
Configuration Validation with Large Language Models
por: Lian, Xinyu, et al.
Publicado: (2023)
por: Lian, Xinyu, et al.
Publicado: (2023)
Deploy, Calibrate, Monitor, Heal -- No Human Required: An Autonomous AI SRE Agent for Elasticsearch
por: Mukkolakkal, Muhamed Ramees Cheriya
Publicado: (2026)
por: Mukkolakkal, Muhamed Ramees Cheriya
Publicado: (2026)
Causely: A Causal Intelligence Layer for Enterprise AI A Benchmark Study on SRE and Reliability Workflows
por: Dalal, Dhairya, et al.
Publicado: (2026)
por: Dalal, Dhairya, et al.
Publicado: (2026)
Site Reliability Engineering (SRE) and Observations on SRE Process to Make Tasks Easier
por: Puli, Balaram
Publicado: (2025)
por: Puli, Balaram
Publicado: (2025)
Beyond Performance: Measuring the Environmental Impact of Analytical Databases
por: Bachras, Michail, et al.
Publicado: (2025)
por: Bachras, Michail, et al.
Publicado: (2025)
Adversarial Robustness in Distributed Quantum Machine Learning
por: Kananian, Pouya, et al.
Publicado: (2025)
por: Kananian, Pouya, et al.
Publicado: (2025)
How Does Stake Distribution Influence Consensus? Analyzing Blockchain Decentralization
por: Motepalli, Shashank, et al.
Publicado: (2023)
por: Motepalli, Shashank, et al.
Publicado: (2023)
Decentralization in PoS Blockchain Consensus: Quantification and Advancement
por: Motepalli, Shashank, et al.
Publicado: (2025)
por: Motepalli, Shashank, et al.
Publicado: (2025)
Ksurf: Attention Kalman Filter and Principal Component Analysis for Prediction under Highly Variable Cloud Workloads
por: Dang'ana, Michael, et al.
Publicado: (2024)
por: Dang'ana, Michael, et al.
Publicado: (2024)
Adversarial Robustness of Partitioned Quantum Classifiers
por: Kananian, Pouya, et al.
Publicado: (2025)
por: Kananian, Pouya, et al.
Publicado: (2025)
Quantum simulations of quantum electrodynamics in Coulomb gauge
por: Li, Tianyin
Publicado: (2024)
por: Li, Tianyin
Publicado: (2024)
FLEdge: Benchmarking Federated Machine Learning Applications in Edge Computing Systems
por: Woisetschläger, Herbert, et al.
Publicado: (2023)
por: Woisetschläger, Herbert, et al.
Publicado: (2023)
Analyzing Common Electronic Structure Theory Algorithms for Distributed Quantum Computing
por: Jones, Grier M., et al.
Publicado: (2025)
por: Jones, Grier M., et al.
Publicado: (2025)
Distributed Quantum Computing for Chemical Applications
por: Jones, Grier M., et al.
Publicado: (2024)
por: Jones, Grier M., et al.
Publicado: (2024)
DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios
por: Meng, Jinxiang, et al.
Publicado: (2026)
por: Meng, Jinxiang, et al.
Publicado: (2026)
An Empirical Study of Production Incidents in Generative AI Cloud Services
por: Yan, Haoran, et al.
Publicado: (2025)
por: Yan, Haoran, et al.
Publicado: (2025)
Diba: A Re-configurable Stream Processor
por: Najafi, Mohammadreza, et al.
Publicado: (2023)
por: Najafi, Mohammadreza, et al.
Publicado: (2023)
AgentIF-OneDay: A Task-level Instruction-Following Benchmark for General AI Agents in Daily Scenarios
por: Chen, Kaiyuan, et al.
Publicado: (2026)
por: Chen, Kaiyuan, et al.
Publicado: (2026)
Routing, Cascades, and User Choice for LLMs
por: Mahmood, Rafid
Publicado: (2026)
por: Mahmood, Rafid
Publicado: (2026)
Bayesian Analysis for a Threshold Double Autoregressive Model With Explanatory Variables
por: Han Li, et al.
Publicado: (2025)
por: Han Li, et al.
Publicado: (2025)
Substituting Proof of Work in Blockchain with Training-Verified Collaborative Model Computation
por: Rafid, Mohammad Ishzaz Asif, et al.
Publicado: (2025)
por: Rafid, Mohammad Ishzaz Asif, et al.
Publicado: (2025)
Beyond the Benchmark: Innovative Defenses Against Prompt Injection Attacks
por: Shaheer, Safwan, et al.
Publicado: (2025)
por: Shaheer, Safwan, et al.
Publicado: (2025)
Federated Learning and AI Regulation in the European Union: Who is Responsible? -- An Interdisciplinary Analysis
por: Woisetschläger, Herbert, et al.
Publicado: (2024)
por: Woisetschläger, Herbert, et al.
Publicado: (2024)
OpenDerisk: An Industrial Framework for AI-Driven SRE, with Design, Implementation, and Case Studies
por: Di, Peng, et al.
Publicado: (2025)
por: Di, Peng, et al.
Publicado: (2025)
Faith, Truth, Fidelity
por: Jackson, Frances
Publicado: (2025)
por: Jackson, Frances
Publicado: (2025)
How Can We Train Deep Learning Models Across Clouds and Continents? An Experimental Study
por: Erben, Alexander, et al.
Publicado: (2023)
por: Erben, Alexander, et al.
Publicado: (2023)
Federated Computing -- Survey on Building Blocks, Extensions and Systems
por: Schwermer, René, et al.
Publicado: (2024)
por: Schwermer, René, et al.
Publicado: (2024)
MAR-FL: A Communication Efficient Peer-to-Peer Federated Learning System
por: Mulitze, Felix, et al.
Publicado: (2025)
por: Mulitze, Felix, et al.
Publicado: (2025)
The Americas. Living with Fidel
Publicado: (1999)
Publicado: (1999)
SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios
por: Chen, Junkai, et al.
Publicado: (2025)
por: Chen, Junkai, et al.
Publicado: (2025)
Forgetful but Faithful: A Cognitive Memory Architecture and Benchmark for Privacy-Aware Generative Agents
por: Alqithami, Saad
Publicado: (2025)
por: Alqithami, Saad
Publicado: (2025)
AIOpsLab: A Holistic Framework to Evaluate AI Agents for Enabling Autonomous Clouds
por: Chen, Yinfang, et al.
Publicado: (2025)
por: Chen, Yinfang, et al.
Publicado: (2025)
LMDG: Advancing Lateral Movement Detection Through High-Fidelity Dataset Generation
por: Mabrouk, Anas, et al.
Publicado: (2025)
por: Mabrouk, Anas, et al.
Publicado: (2025)
Adaptive Teaching in Heterogeneous Agents: Balancing Surprise in Sparse Reward Scenarios
por: Clark, Emma, et al.
Publicado: (2024)
por: Clark, Emma, et al.
Publicado: (2024)
Ejemplares similares
-
STRATUS: A Multi-agent System for Autonomous Reliability Engineering of Modern Clouds
por: Chen, Yinfang, et al.
Publicado: (2025) -
FAME: Failure-Aware Mixture-of-Experts for Message-Level Log Anomaly Detection
por: Wang, Huanchi, et al.
Publicado: (2026) -
Activation of Carbon─Chlorine Bond at Chlorpromazine for Access to New Derivatives Through Nickel‐Catalyzed Kumada–Tamao–Corriu Cross‐Coupling Reaction
por: Rafid Saad Dawood
Publicado: (2025) -
GALA: Can Graph-Augmented Large Language Model Agentic Workflows Elevate Root Cause Analysis?
por: Tian, Yifang, et al.
Publicado: (2025) -
Pricing and Competition for Generative AI
por: Mahmood, Rafid
Publicado: (2024)