What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design
Fuente:
arXiv
Guardado en:
| Autor principal: | Bercovich, Ivan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories
por: Bercovich, Ivan, et al.
Publicado: (2026)
por: Bercovich, Ivan, et al.
Publicado: (2026)
Alljoined1 -- A dataset for EEG-to-Image decoding
por: Xu, Jonathan, et al.
Publicado: (2024)
por: Xu, Jonathan, et al.
Publicado: (2024)
AI Observability for Developer Productivity Tools: Bridging Cost Awareness and Code Quality
por: Bhati, Happy, et al.
Publicado: (2026)
por: Bhati, Happy, et al.
Publicado: (2026)
Combining Euclidean Alignment and Data Augmentation for BCI decoding
por: Rodrigues, Gustavo H., et al.
Publicado: (2024)
por: Rodrigues, Gustavo H., et al.
Publicado: (2024)
SemLoc: Structured Grounding of Free-Form LLM Reasoning for Fault Localization
por: Yang, Zhaorui, et al.
Publicado: (2026)
por: Yang, Zhaorui, et al.
Publicado: (2026)
Bayesian Hierarchical Probabilistic Forecasting of Intraday Electricity Prices
por: Nickelsen, Daniel, et al.
Publicado: (2024)
por: Nickelsen, Daniel, et al.
Publicado: (2024)
ASSERTIFY: Utilizing Large Language Models to Generate Assertions for Production Code
por: Torkamani, Mohammad Jalili, et al.
Publicado: (2024)
por: Torkamani, Mohammad Jalili, et al.
Publicado: (2024)
The Productivity-Reliability Paradox: Specification-Driven Governance for AI-Augmented Software Development
por: Farrag, Sabry E.
Publicado: (2026)
por: Farrag, Sabry E.
Publicado: (2026)
Closed-Loop Autonomous Software Development via Jira-Integrated Backlog Orchestration: A Case Study in Deterministic Control and Safety-Constrained Automation
por: Calboreanu, Elias
Publicado: (2026)
por: Calboreanu, Elias
Publicado: (2026)
A Framework for Assessing AI Agent Decisions and Outcomes in AutoML Pipelines
por: Du, Gaoyuan, et al.
Publicado: (2026)
por: Du, Gaoyuan, et al.
Publicado: (2026)
Understanding and Detecting Flaky Builds in GitHub Actions
por: Ge, Wenhao, et al.
Publicado: (2026)
por: Ge, Wenhao, et al.
Publicado: (2026)
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)
por: Young, Richard J., et al.
Publicado: (2026)
por: Young, Richard J., et al.
Publicado: (2026)
An Evaluation of a Structured Spreadsheet Development Methodology
por: Rajalingham, Kamalasen, et al.
Publicado: (2008)
por: Rajalingham, Kamalasen, et al.
Publicado: (2008)
Latent Space Representation of Electricity Market Curves: Maintaining Structural Integrity
por: Výboh, Martin, et al.
Publicado: (2025)
por: Výboh, Martin, et al.
Publicado: (2025)
Risk-Calibrated Bayesian Streaming Intrusion Detection with SRE-Aligned Decisions
por: Youssef, Michel
Publicado: (2025)
por: Youssef, Michel
Publicado: (2025)
Modeling Nonlinear Oscillator Networks Using Physics-Informed Hybrid Reservoir Computing
por: Shannon, Andrew, et al.
Publicado: (2024)
por: Shannon, Andrew, et al.
Publicado: (2024)
Perfecting Aircraft Maneuvers with Reinforcement Learning
por: Cilan, Atahan, et al.
Publicado: (2026)
por: Cilan, Atahan, et al.
Publicado: (2026)
Migration as a Probe: A Generalizable Benchmark Framework for Specialist vs. Generalist Machine-Learned Force Fields
por: Cao, Yi, et al.
Publicado: (2025)
por: Cao, Yi, et al.
Publicado: (2025)
A Systematic Evaluation of Euclidean Alignment with Deep Learning for EEG Decoding
por: Junqueira, Bruna, et al.
Publicado: (2024)
por: Junqueira, Bruna, et al.
Publicado: (2024)
Explainable Attention-Based LSTM Framework for Early Detection of AI-Assisted Ransomware via File System Behavioral Analysis
por: Nayak, Prabhudarshi, et al.
Publicado: (2026)
por: Nayak, Prabhudarshi, et al.
Publicado: (2026)
Operationalizing Cybersecurity Governance for Mitigation Planning with Attack-Path Modeling and Reinforcement Learning
por: Huff, Philip, et al.
Publicado: (2026)
por: Huff, Philip, et al.
Publicado: (2026)
RADEP: A Resilient Adaptive Defense Framework Against Model Extraction Attacks
por: Chakraborty, Amit, et al.
Publicado: (2025)
por: Chakraborty, Amit, et al.
Publicado: (2025)
AuditRepairBench: A Paired-Execution Trace Corpus for Evaluator-Channel Ranking Instability in Agent Repair
por: Hu, Yuelin, et al.
Publicado: (2026)
por: Hu, Yuelin, et al.
Publicado: (2026)
FuzzDistill: Intelligent Fuzzing Target Selection using Compile-Time Analysis and Machine Learning
por: Upadhyay, Saket
Publicado: (2024)
por: Upadhyay, Saket
Publicado: (2024)
Automating the Detection of Code Vulnerabilities by Analyzing GitHub Issues
por: Cipollone, Daniele, et al.
Publicado: (2025)
por: Cipollone, Daniele, et al.
Publicado: (2025)
Streamlining Security Vulnerability Triage with Large Language Models
por: Torkamani, Mohammad Jalili, et al.
Publicado: (2025)
por: Torkamani, Mohammad Jalili, et al.
Publicado: (2025)
Dr. Jekyll and Mr. Hyde: Two Faces of LLMs
por: Collu, Matteo Gioele, et al.
Publicado: (2023)
por: Collu, Matteo Gioele, et al.
Publicado: (2023)
Towards geological inference with process-based and deep generative modeling, part 1: training on fluvial deposits
por: Rongier, Guillaume, et al.
Publicado: (2025)
por: Rongier, Guillaume, et al.
Publicado: (2025)
Towards geological inference with process-based and deep generative modeling, part 2: inversion of fluvial deposits and latent-space disentanglement
por: Rongier, Guillaume, et al.
Publicado: (2025)
por: Rongier, Guillaume, et al.
Publicado: (2025)
CUDABeaver: Benchmarking LLM-Based Automated CUDA Debugging
por: Li, Shiyang, et al.
Publicado: (2026)
por: Li, Shiyang, et al.
Publicado: (2026)
Safe and Policy-Compliant Multi-Agent Orchestration for Enterprise AI
por: Pasupuleti, Vinil, et al.
Publicado: (2026)
por: Pasupuleti, Vinil, et al.
Publicado: (2026)
ChipCraftBrain: Validation-First RTL Generation via Multi-Agent Orchestration
por: Eryilmaz, Cagri
Publicado: (2026)
por: Eryilmaz, Cagri
Publicado: (2026)
CodeTracer: Towards Traceable Agent States
por: Li, Han, et al.
Publicado: (2026)
por: Li, Han, et al.
Publicado: (2026)
Multi-Agent Honeypot-Based Request-Response Context Dataset for Improved SQL Injection Detection Performance
por: Yu, Hao, et al.
Publicado: (2026)
por: Yu, Hao, et al.
Publicado: (2026)
Thinking is Bad: Implications of Human Error Research for Spreadsheet Research and Practice
por: Panko, Raymond R.
Publicado: (2008)
por: Panko, Raymond R.
Publicado: (2008)
On the Mistaken Assumption of Interchangeable Deep Reinforcement Learning Implementations
por: Hundal, Rajdeep Singh, et al.
Publicado: (2025)
por: Hundal, Rajdeep Singh, et al.
Publicado: (2025)
Towards Explainable Test Case Prioritisation with Learning-to-Rank Models
por: Ramírez, Aurora, et al.
Publicado: (2024)
por: Ramírez, Aurora, et al.
Publicado: (2024)
Forecasting Coccidioidomycosis (Valley Fever) in Arizona: A Graph Neural Network Approach
por: Sarabi, Ali, et al.
Publicado: (2025)
por: Sarabi, Ali, et al.
Publicado: (2025)
Learning When to Remember: Risk-Sensitive Contextual Bandits for Abstention-Aware Memory Retrieval in LLM-Based Coding Agents
por: Iscan, Mehmet
Publicado: (2026)
por: Iscan, Mehmet
Publicado: (2026)
Sim-to-reality adaptation for Deep Reinforcement Learning applied to an underwater docking application
por: Chaarani, Alaaeddine, et al.
Publicado: (2026)
por: Chaarani, Alaaeddine, et al.
Publicado: (2026)
Ejemplares similares
-
Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories
por: Bercovich, Ivan, et al.
Publicado: (2026) -
Alljoined1 -- A dataset for EEG-to-Image decoding
por: Xu, Jonathan, et al.
Publicado: (2024) -
AI Observability for Developer Productivity Tools: Bridging Cost Awareness and Code Quality
por: Bhati, Happy, et al.
Publicado: (2026) -
Combining Euclidean Alignment and Data Augmentation for BCI decoding
por: Rodrigues, Gustavo H., et al.
Publicado: (2024) -
SemLoc: Structured Grounding of Free-Form LLM Reasoning for Fault Localization
por: Yang, Zhaorui, et al.
Publicado: (2026)