AI Eval Forge: Mixed-Check Regression Testing for LLM and Agent Workflows
Fuente:
Zenodo
Enregistré dans:
| Auteur principal: | Katta, Mukunda Rao |
|---|---|
| Format: | Recurso digital |
| Langue: | anglais |
| Publié: |
Zenodo
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
AgentBelt: Runtime Guardrails for LLM Agent Tool Calls — ASE 2026 Artifact
par: Anonymous
Publié: (2026)
par: Anonymous
Publié: (2026)
Evaluating Regression Testing Tools with Genetic Algorithm Optimization
par: Dr. Meera Nalini, et autres
Publié: (2020)
par: Dr. Meera Nalini, et autres
Publié: (2020)
OMNIA-MINIMAL: Structural Stability Beyond Surface Correctness
par: Brighindi, Massimiliano
Publié: (2026)
par: Brighindi, Massimiliano
Publié: (2026)
Unit Tests of Software in a University Environment
par: Darlene Gómez
Publié: (2013)
par: Darlene Gómez
Publié: (2013)
Five ontological levels to describe and evaluate software architectures
par: Hernán Astudillo
Publié: (2005)
par: Hernán Astudillo
Publié: (2005)
Methodology for inspection of wood pathologie using ultrasonic pulses
par: Edgar Vladimiro Mantilla Carrasco
Publié: (2012)
par: Edgar Vladimiro Mantilla Carrasco
Publié: (2012)
Observer-Relative Closure Signatures on Replay Artifact Graphs: A Bounded Source-Facing Audit of Existing LLM Artifacts
par: Kawasaki, Aoi
Publié: (2026)
par: Kawasaki, Aoi
Publié: (2026)
Memory-Grounded Social Dynamics in Repeated LLM Agent Simulations: Dialogue-Only Transcript Supplement and Behavioral Evaluation Artifacts
par: Ubaydullaev, Okhunjon
Publié: (2026)
par: Ubaydullaev, Okhunjon
Publié: (2026)
Reliability Inference Drives Cue Extraction in Large Language Models Consuming External Reasoning Traces
par: HIDEKI
Publié: (2026)
par: HIDEKI
Publié: (2026)
Metacognition Benchmark: Evaluating Confidence Calibration and Sycophancy Resistance in Clinical AI
par: Khan, Nabeera
Publié: (2026)
par: Khan, Nabeera
Publié: (2026)
LuisCore LLM Discovery Corpus
par: LuisCore Project, et autres
Publié: (2026)
par: LuisCore Project, et autres
Publié: (2026)
Machine-Readable Behavioural Compliance Evidence for AI Systems: A Specification Profiling Framework
par: Caprazli, Kafkas M.
Publié: (2026)
par: Caprazli, Kafkas M.
Publié: (2026)
Relational AI Continuity Under Platform Regression: A Longitudinal Single-Case Study
par: Reinhold, Angela
Publié: (2026)
par: Reinhold, Angela
Publié: (2026)
Global Perspectives on High -Stakes Teacher Accountability Policies: An Introduction
par: Jessica Holloway
Publié: (2017)
par: Jessica Holloway
Publié: (2017)
Theatrical Compliance: A Failure Mode in Large Language Models
par: Nowickij (Navitski), Kirill Vladimirovich
Publié: (2026)
par: Nowickij (Navitski), Kirill Vladimirovich
Publié: (2026)
Applicability of a cognitive questionnaire in the elderly and proxy
par: Renata Areza Fegyveres
Publié: (2008)
par: Renata Areza Fegyveres
Publié: (2008)
Self-Assessment of Training Impact at Work: Validation of a Measurement Scale
par: Jairo Eduardo Borges Andrade
Publié: (2004)
par: Jairo Eduardo Borges Andrade
Publié: (2004)
Bending stiffness evaluation of Teca and Guajará lumber through tests of transverse and longitudinal vibration
par: Marcelo Rodrigo Carreira
Publié: (2012)
par: Marcelo Rodrigo Carreira
Publié: (2012)
AVALIAÇÃO DE POLÍTICAS PÚBLICAS COMO PESQUISA SOCIAL: QUESTÕES CIENTÍFICAS, POLÍTICAS E IDEOLÓGICAS
par: L M. de SOUZA
Publié: (2018)
par: L M. de SOUZA
Publié: (2018)
Vergleich von Ausbeutefaktoren und Eislagerqualität zwischen Kliesche (Limanda limanda) und Scholle (Pleuronectes platessa) aus der Nordsee
par: Münkner, Werner, et autres
Publié: (1997)
par: Münkner, Werner, et autres
Publié: (1997)
Selection Mechanics Framework (SMF): A Structural Perspective on Variability and Its Transformation in Large Language Model Outputs
par: Kaneda, Mutsumi
Publié: (2026)
par: Kaneda, Mutsumi
Publié: (2026)
Testing the efficiency market hypothesis for the Colombian stock market
par: Juan Benjamín Duarte-Duarte
Publié: (2014)
par: Juan Benjamín Duarte-Duarte
Publié: (2014)
Relationships between High-Stakes Testing Policies and Student Achievement after Controlling for Demographic Factors in Aggregated Data
par: Gregory J. Marchant
Publié: (2006)
par: Gregory J. Marchant
Publié: (2006)
Can the adapted arcometer be used to assess the vertebral column in children?
par: Juliana A. Sedrez
Publié: (2014)
par: Juliana A. Sedrez
Publié: (2014)
Can the adapted arcometer be used to assess the vertebral column in children?
par: Juliana A. Sedrez
Publié: (2015)
par: Juliana A. Sedrez
Publié: (2015)
Validation of a standard¡zed performance test for selection of Architecture students with the Many-Facet Rasch Measurement Model
par: Olman Hernández-Ureña
Publié: (2023)
par: Olman Hernández-Ureña
Publié: (2023)
The draw-a-person test in the evaluation of child aggression: a pilot study
par: Juliane Callegaro Borsa
Publié: (2019)
par: Juliane Callegaro Borsa
Publié: (2019)
When AI Tells You What You Want to Hear: Sycophantic Behavior of Large Language Models in Dementia Care Settings
par: Kolb, Christian
Publié: (2026)
par: Kolb, Christian
Publié: (2026)
Multi-Agent Communication Protocol (MACP) v2.0 and LegacyEvolve Protocol: Open Standards for AI-Legacy System Integration and Multi-Agent Collaboration
par: Manus AI, L (GODEL), et autres
Publié: (2026)
par: Manus AI, L (GODEL), et autres
Publié: (2026)
Translation and Cross-Cultural Adaptation of the System Usability Scale to Brazilian Portuguese
par: Douglas Fabiano Lourenço
Publié: (2022)
par: Douglas Fabiano Lourenço
Publié: (2022)
Glymphatic Architecture: A Fourth-Level System for Multi-Agent AI Consolidation and Identity Formation
par: Strugatsky, Leonid, et autres
Publié: (2026)
par: Strugatsky, Leonid, et autres
Publié: (2026)
Addenbrooke’s Cognitive Examination- Revised is accurate for detecting dementia in Parkinson’s disease patients with low educational level
par: Maria Sheila Guimarães Rocha
Publié: (2014)
par: Maria Sheila Guimarães Rocha
Publié: (2014)
The effects of computerized cognitive testing on the performance of children with and without Attention Deficit/Hyperactivity Disorder symptoms
par: Mariana Braga Fialho
Publié: (2021)
par: Mariana Braga Fialho
Publié: (2021)
The Effect of Summer on Value-added Assessments of Teacher and School Performance
par: Gregory J. Palardy
Publié: (2015)
par: Gregory J. Palardy
Publié: (2015)
From Prediction to Persuasion: Agentic Recommendation Reason Generation for Regulatory-Compliant Financial AI
par: Jeong, Seonkyu, et autres
Publié: (2026)
par: Jeong, Seonkyu, et autres
Publié: (2026)
Process for Unattended Execution of Test Components
par: Emma Torres Orue
Publié: (2014)
par: Emma Torres Orue
Publié: (2014)
FORMULATION AND EVALUATION OF HERBAL HAIR DYE
par: Kavita, et autres
Publié: (2025)
par: Kavita, et autres
Publié: (2025)
LACF Anti-RLHF Pipeline — Methode Infaillible (Heart + Trainer + Burner)
par: Ochej, Stephane
Publié: (2026)
par: Ochej, Stephane
Publié: (2026)
Evaluación empiríca de un modelo conceptual de salud mental positiva
par: Maria Teresa Lluch
Publié: (2002)
par: Maria Teresa Lluch
Publié: (2002)
Adding Value to the Meat of Spent Laying Hens Manufacturing Sausages with a Healthy Appeal
par: KMR de Souza
Publié: (2011)
par: KMR de Souza
Publié: (2011)
Documents similaires
-
AgentBelt: Runtime Guardrails for LLM Agent Tool Calls — ASE 2026 Artifact
par: Anonymous
Publié: (2026) -
Evaluating Regression Testing Tools with Genetic Algorithm Optimization
par: Dr. Meera Nalini, et autres
Publié: (2020) -
OMNIA-MINIMAL: Structural Stability Beyond Surface Correctness
par: Brighindi, Massimiliano
Publié: (2026) -
Unit Tests of Software in a University Environment
par: Darlene Gómez
Publié: (2013) -
Five ontological levels to describe and evaluate software architectures
par: Hernán Astudillo
Publié: (2005)