A measurement substrate for agentic Kubernetes operations: Methodology and a case study in retrieval-compounding falsification
Fuente:
arXiv
Saved in:
| Main Authors: | Odmark, Joshua, Rubin, Gideon, van der Vyver, Deon |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Single-File Test: A Longitudinal Public-Interface Evaluation of First-Output LLM Web Generation with Social Reach Tracking
by: Palacios, Diego Cabezas
Published: (2026)
by: Palacios, Diego Cabezas
Published: (2026)
Reinforcement Learning for Dynamic Workflow Optimization in CI/CD Pipelines
by: Soni, Aniket Abhishek, et al.
Published: (2026)
by: Soni, Aniket Abhishek, et al.
Published: (2026)
Multimodal Generative AI for Story Point Estimation in Software Development
by: Islam, Mohammad Rubyet, et al.
Published: (2025)
by: Islam, Mohammad Rubyet, et al.
Published: (2025)
GitHub Copilot and Developer Productivity: An Observational Dose-Response Analysis
by: Heilman, Alex, et al.
Published: (2026)
by: Heilman, Alex, et al.
Published: (2026)
Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture
by: Iscan, Mehmet
Published: (2026)
by: Iscan, Mehmet
Published: (2026)
Predictive Analytics for Collaborators Answers, Code Quality, and Dropout on Stack Overflow
by: Zolduoarrati, Elijah, et al.
Published: (2025)
by: Zolduoarrati, Elijah, et al.
Published: (2025)
OODEval: Evaluating Large Language Models on Object-Oriented Design
by: Xiao, Bingxu, et al.
Published: (2026)
by: Xiao, Bingxu, et al.
Published: (2026)
Context Engineering for Multi-Agent LLM Code Assistants Using Elicit, NotebookLM, ChatGPT, and Claude Code
by: Haseeb, Muhammad
Published: (2025)
by: Haseeb, Muhammad
Published: (2025)
ConfProBench: A Confidence Evaluation Benchmark for MLLM-Based Process Judges
by: Zhou, Yue, et al.
Published: (2025)
by: Zhou, Yue, et al.
Published: (2025)
Foundational Requirements for Artificial General Intelligence: A Falsifiable Framework Based on Signal Prediction
by: Šprogar, Matej
Published: (2025)
by: Šprogar, Matej
Published: (2025)
Creating benchmarkable components to measure the quality ofAI-enhanced developer tools
by: Paradis, Elise, et al.
Published: (2025)
by: Paradis, Elise, et al.
Published: (2025)
How much does AI impact development speed? An enterprise-based randomized controlled trial
by: Paradis, Elise, et al.
Published: (2024)
by: Paradis, Elise, et al.
Published: (2024)
Generative AI and the Transformation of Software Development Practices
by: Acharya, Vivek
Published: (2025)
by: Acharya, Vivek
Published: (2025)
Towards Interactively Improving ML Data Preparation Code via "Shadow Pipelines"
by: Grafberger, Stefan, et al.
Published: (2024)
by: Grafberger, Stefan, et al.
Published: (2024)
Prompt Engineering Strategies for LLM-based Qualitative Coding of Psychological Safety in Software Engineering Communities: A Controlled Empirical Study
by: Alshaikh, Moaath, et al.
Published: (2026)
by: Alshaikh, Moaath, et al.
Published: (2026)
Scattered Forest Search: Smarter Code Space Exploration with LLMs
by: Light, Jonathan, et al.
Published: (2024)
by: Light, Jonathan, et al.
Published: (2024)
AI and Machine Learning Approaches for Predicting Nanoparticles Toxicity The Critical Role of Physiochemical Properties
by: Yousaf, Iqra
Published: (2024)
by: Yousaf, Iqra
Published: (2024)
Towards Generative Ray Path Sampling for Faster Point-to-Point Ray Tracing
by: Eertmans, Jérome, et al.
Published: (2024)
by: Eertmans, Jérome, et al.
Published: (2024)
Transform-Invariant Generative Ray Path Sampling for Efficient Radio Propagation Modeling
by: Eertmans, Jérome, et al.
Published: (2026)
by: Eertmans, Jérome, et al.
Published: (2026)
Stabilization Without Simplification: A Two-Dimensional Model of Software Evolution
by: Furukawa, Masaru
Published: (2026)
by: Furukawa, Masaru
Published: (2026)
From Monolith to Microservices: A Comparative Evaluation of Decomposition Frameworks
by: Weerasinghe, Mineth, et al.
Published: (2026)
by: Weerasinghe, Mineth, et al.
Published: (2026)
A Tale of Two Systems: Characterizing Architectural Complexity on Machine Learning-Enabled Systems
by: Ferreira, Renato Cordeiro
Published: (2025)
by: Ferreira, Renato Cordeiro
Published: (2025)
A Metrics-Oriented Architectural Model to Characterize Complexity on Machine Learning-Enabled Systems
by: Ferreira, Renato Cordeiro
Published: (2025)
by: Ferreira, Renato Cordeiro
Published: (2025)
BACE: LLM-based Code Generation through Bayesian Anchored Co-Evolution of Code and Test Populations
by: Silva, Kaushitha, et al.
Published: (2026)
by: Silva, Kaushitha, et al.
Published: (2026)
TML-Bench: Benchmark for Data Science Agents on Tabular ML Tasks
by: Pinchuk, Mykola
Published: (2026)
by: Pinchuk, Mykola
Published: (2026)
Making Software Metrics Useful
by: Tempero, Ewan, et al.
Published: (2026)
by: Tempero, Ewan, et al.
Published: (2026)
Illuminating Patterns of Divergence: DataDios SmartDiff for Large-Scale Data Difference Analysis
by: Poduri, Aryan, et al.
Published: (2025)
by: Poduri, Aryan, et al.
Published: (2025)
Comparing Human and LLM Generated Code: The Jury is Still Out!
by: Licorish, Sherlock A., et al.
Published: (2025)
by: Licorish, Sherlock A., et al.
Published: (2025)
MARIA: A Framework for Marginal Risk Assessment without Ground Truth in AI Systems
by: Chen, Jieshan, et al.
Published: (2025)
by: Chen, Jieshan, et al.
Published: (2025)
InterEvo-TR: Interactive Evolutionary Test Generation With Readability Assessment
by: Delgado-Pérez, Pedro, et al.
Published: (2024)
by: Delgado-Pérez, Pedro, et al.
Published: (2024)
Multi-Agent LLM Orchestration Achieves Deterministic, High-Quality Decision Support for Incident Response
by: Drammeh, Philip
Published: (2025)
by: Drammeh, Philip
Published: (2025)
How Generation Architecture Shapes Code Complexity in Multi-Agent LLM Systems: A Paired Study on HumanEval
by: Ashrafi, Nazmus
Published: (2026)
by: Ashrafi, Nazmus
Published: (2026)
Framework Matters: Energy Efficiency of UI Automation Testing Frameworks
by: Lagermann, Timmie M. R., et al.
Published: (2025)
by: Lagermann, Timmie M. R., et al.
Published: (2025)
GBM Returns the Best Prediction Performance among Regression Approaches: A Case Study of Stack Overflow Code Quality
by: Licorish, Sherlock A., et al.
Published: (2025)
by: Licorish, Sherlock A., et al.
Published: (2025)
Using LLMs to Establish Implicit User Sentiment of Software Desirability
by: Weitl-Harms, Sherri, et al.
Published: (2024)
by: Weitl-Harms, Sherri, et al.
Published: (2024)
Meta-Reinforcement Learning with Discrete World Models for Adaptive Load Balancing
by: Redovian, Cameron
Published: (2025)
by: Redovian, Cameron
Published: (2025)
Comprehensive Evaluation of Large Language Models on Software Engineering Tasks: A Multi-Task Benchmark
by: Gunawan, Go Frendi, et al.
Published: (2026)
by: Gunawan, Go Frendi, et al.
Published: (2026)
How Quickly Do Development Teams Update Their Vulnerable Dependencies?
by: Rahman, Imranur, et al.
Published: (2024)
by: Rahman, Imranur, et al.
Published: (2024)
A semantic mutation metric for metamorphic relation adequacy in scientific computing programs
by: Li, Meng, et al.
Published: (2026)
by: Li, Meng, et al.
Published: (2026)
GEML: A Grammar-based Evolutionary Machine Learning Approach for Design-Pattern Detection
by: Barbudo, Rafael, et al.
Published: (2024)
by: Barbudo, Rafael, et al.
Published: (2024)
Similar Items
-
The Single-File Test: A Longitudinal Public-Interface Evaluation of First-Output LLM Web Generation with Social Reach Tracking
by: Palacios, Diego Cabezas
Published: (2026) -
Reinforcement Learning for Dynamic Workflow Optimization in CI/CD Pipelines
by: Soni, Aniket Abhishek, et al.
Published: (2026) -
Multimodal Generative AI for Story Point Estimation in Software Development
by: Islam, Mohammad Rubyet, et al.
Published: (2025) -
GitHub Copilot and Developer Productivity: An Observational Dose-Response Analysis
by: Heilman, Alex, et al.
Published: (2026) -
Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture
by: Iscan, Mehmet
Published: (2026)