ReliabilityBench: Evaluating LLM Agent Reliability Under Production-Like Stress Conditions
Fuente:
arXiv
Salvato in:
| Autore principale: | Gupta, Aayush |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Fact Grounded Attention: Eliminating Hallucination in Large Language Models Through Attention Level Knowledge Integration
di: Gupta, Aayush
Pubblicazione: (2025)
di: Gupta, Aayush
Pubblicazione: (2025)
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics
di: Peyronnet, Antoine, et al.
Pubblicazione: (2026)
di: Peyronnet, Antoine, et al.
Pubblicazione: (2026)
React-ing to Grace Hopper 200: Five Open-Weights Coding Models, One React Native App, One GH200, One Weekend
di: Potanin, Alex
Pubblicazione: (2026)
di: Potanin, Alex
Pubblicazione: (2026)
NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles
di: Jia, Xiao
Pubblicazione: (2026)
di: Jia, Xiao
Pubblicazione: (2026)
Aligning LLMs for Multilingual Consistency in Enterprise Applications
di: Agarwal, Amit, et al.
Pubblicazione: (2025)
di: Agarwal, Amit, et al.
Pubblicazione: (2025)
When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels
di: Gautam, Sushant, et al.
Pubblicazione: (2026)
di: Gautam, Sushant, et al.
Pubblicazione: (2026)
Classifier-Augmented Generation for Structured Workflow Prediction
di: Gschwind, Thomas, et al.
Pubblicazione: (2025)
di: Gschwind, Thomas, et al.
Pubblicazione: (2025)
MEDLEY-BENCH: Scale Buys Evaluation but Not Control in AI Metacognition
di: Abtahi, Farhad, et al.
Pubblicazione: (2026)
di: Abtahi, Farhad, et al.
Pubblicazione: (2026)
Bug In the Code Stack: Can LLMs Find Bugs in Large Python Code Stacks
di: Lee, Hokyung, et al.
Pubblicazione: (2024)
di: Lee, Hokyung, et al.
Pubblicazione: (2024)
SLEAN: Simple Lightweight Ensemble Analysis Network for Multi-Provider LLM Coordination: Design, Implementation, and Vibe Coding Bug Investigation Case Study
di: Vargas, Matheus J. T.
Pubblicazione: (2025)
di: Vargas, Matheus J. T.
Pubblicazione: (2025)
Generative AI and the Transformation of Software Development Practices
di: Acharya, Vivek
Pubblicazione: (2025)
di: Acharya, Vivek
Pubblicazione: (2025)
A Computational Approach to Modeling Conversational Systems: Analyzing Large-Scale Quasi-Patterned Dialogue Flows
di: Ammar, Mohamed Achref Ben, et al.
Pubblicazione: (2025)
di: Ammar, Mohamed Achref Ben, et al.
Pubblicazione: (2025)
From Noise to Diversity: Random Embedding Injection in LLM Reasoning
di: Kim, Heejun, et al.
Pubblicazione: (2026)
di: Kim, Heejun, et al.
Pubblicazione: (2026)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
di: Pather, Kaviraj, et al.
Pubblicazione: (2025)
di: Pather, Kaviraj, et al.
Pubblicazione: (2025)
BMAM: Brain-inspired Multi-Agent Memory Framework
di: Li, Yang, et al.
Pubblicazione: (2026)
di: Li, Yang, et al.
Pubblicazione: (2026)
Co-NAML-LSTUR: A Combined Model with Attentive Multi-View Learning and Long- and Short-term User Representations for News Recommendation
di: Nguyen, Minh Hoang, et al.
Pubblicazione: (2025)
di: Nguyen, Minh Hoang, et al.
Pubblicazione: (2025)
D-SMART: Enhancing LLM Dialogue Consistency via Dynamic Structured Memory And Reasoning Tree
di: Lei, Xiang, et al.
Pubblicazione: (2025)
di: Lei, Xiang, et al.
Pubblicazione: (2025)
Hallucination Detection in Large Language Models with Metamorphic Relations
di: Yang, Borui, et al.
Pubblicazione: (2025)
di: Yang, Borui, et al.
Pubblicazione: (2025)
Harnessing non-adversarial robustness in large language models
di: Zhou, Qinghua, et al.
Pubblicazione: (2026)
di: Zhou, Qinghua, et al.
Pubblicazione: (2026)
Mitigating LLM Hallucinations through Domain-Grounded Tiered Retrieval
di: Haque, Md. Asraful, et al.
Pubblicazione: (2026)
di: Haque, Md. Asraful, et al.
Pubblicazione: (2026)
Rule Extraction in Machine Learning: Chat Incremental Pattern Constructor
di: Nwokocha, Caleb Princewill
Pubblicazione: (2022)
di: Nwokocha, Caleb Princewill
Pubblicazione: (2022)
How much do LLMs learn from negative examples?
di: Hamdan, Shadi, et al.
Pubblicazione: (2025)
di: Hamdan, Shadi, et al.
Pubblicazione: (2025)
Conformal Path Reasoning: Trustworthy Knowledge Graph Question Answering via Path-Level Calibration
di: Lin, Shuhang, et al.
Pubblicazione: (2026)
di: Lin, Shuhang, et al.
Pubblicazione: (2026)
Enhancing Ultra-Low-Bit Quantization of Large Language Models Through Saliency-Aware Partial Retraining
di: Cao, Deyu, et al.
Pubblicazione: (2025)
di: Cao, Deyu, et al.
Pubblicazione: (2025)
Kodezi Chronos: A Debugging-First Language Model for Repository-Scale Code Understanding
di: Khan, Ishraq, et al.
Pubblicazione: (2025)
di: Khan, Ishraq, et al.
Pubblicazione: (2025)
MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare
di: Wang, Yihao, et al.
Pubblicazione: (2026)
di: Wang, Yihao, et al.
Pubblicazione: (2026)
Multilingual Multi-Label Emotion Classification at Scale with Synthetic Data
di: Borisov, Vadim
Pubblicazione: (2026)
di: Borisov, Vadim
Pubblicazione: (2026)
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
di: Viveiros, André G., et al.
Pubblicazione: (2025)
di: Viveiros, André G., et al.
Pubblicazione: (2025)
Thinking Longer, Not Always Smarter: Evaluating LLM Capabilities in Hierarchical Legal Reasoning
di: Zhang, Li, et al.
Pubblicazione: (2025)
di: Zhang, Li, et al.
Pubblicazione: (2025)
Approaches to Semantic Textual Similarity in Slovak Language: From Algorithms to Transformers
di: Radosky, Lukas, et al.
Pubblicazione: (2026)
di: Radosky, Lukas, et al.
Pubblicazione: (2026)
Enhancing OCR for Sino-Vietnamese Language Processing via Fine-tuned PaddleOCRv5
di: Nguyen, Minh Hoang, et al.
Pubblicazione: (2025)
di: Nguyen, Minh Hoang, et al.
Pubblicazione: (2025)
Optimizing What We Trust: Reliability-Guided QUBO Selection of Multi-Agent Weak Framing Signals for Arabic Sentiment Prediction
di: Alkhalifa, Rabab
Pubblicazione: (2026)
di: Alkhalifa, Rabab
Pubblicazione: (2026)
Tool-Genesis: A Task-Driven Tool Creation Benchmark for Self-Evolving Language Agent
di: Xia, Bowei, et al.
Pubblicazione: (2026)
di: Xia, Bowei, et al.
Pubblicazione: (2026)
LLM-supported document separation for printed reviews from zbMATH Open
di: Pluzhnikov, Ivan, et al.
Pubblicazione: (2026)
di: Pluzhnikov, Ivan, et al.
Pubblicazione: (2026)
MVTamperBench: Evaluating Robustness of Vision-Language Models
di: Agarwal, Amit, et al.
Pubblicazione: (2024)
di: Agarwal, Amit, et al.
Pubblicazione: (2024)
Smotrom tvoja pa ander drogoj verden! Resurrecting Dead Pidgin with Generative Models: Russenorsk Case Study
di: Tikhonov, Alexey, et al.
Pubblicazione: (2025)
di: Tikhonov, Alexey, et al.
Pubblicazione: (2025)
It's 2025 -- Narrative Learning is the new baseline to beat for explainable machine learning
di: Baker, Gregory D.
Pubblicazione: (2025)
di: Baker, Gregory D.
Pubblicazione: (2025)
CogniLoad: A Synthetic Natural Language Reasoning Benchmark With Tunable Length, Intrinsic Difficulty, and Distractor Density
di: Kaiser, Daniel, et al.
Pubblicazione: (2025)
di: Kaiser, Daniel, et al.
Pubblicazione: (2025)
AVEC: Bootstrapping Privacy for Local LLMs
di: Gaikwad, Madhava
Pubblicazione: (2025)
di: Gaikwad, Madhava
Pubblicazione: (2025)
Multi-Agent Synergy-Driven Iterative Visual Narrative Synthesis
di: Xi, Wang, et al.
Pubblicazione: (2025)
di: Xi, Wang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Fact Grounded Attention: Eliminating Hallucination in Large Language Models Through Attention Level Knowledge Integration
di: Gupta, Aayush
Pubblicazione: (2025) -
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics
di: Peyronnet, Antoine, et al.
Pubblicazione: (2026) -
React-ing to Grace Hopper 200: Five Open-Weights Coding Models, One React Native App, One GH200, One Weekend
di: Potanin, Alex
Pubblicazione: (2026) -
NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles
di: Jia, Xiao
Pubblicazione: (2026) -
Aligning LLMs for Multilingual Consistency in Enterprise Applications
di: Agarwal, Amit, et al.
Pubblicazione: (2025)