SpecBench: Evaluating Specification-Level Reasoning for Software Engineering LLM Agents
Fuente:
arXiv
Salvato in:
| Autori principali: | Hamblin, Grant, Song, Kevin, Zhu, Zhanda, Jayarajan, Anand, Liu, Sihang, Vijaykumar, Nandita, Pekhimenko, Gennady |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Aegis: Taxonomy and Optimizations for Overcoming Agent-Environment Failures in LLM Agents
di: Song, Kevin, et al.
Pubblicazione: (2025)
di: Song, Kevin, et al.
Pubblicazione: (2025)
Tally: Non-Intrusive Performance Isolation for Concurrent Deep Learning Workloads
di: Zhao, Wei, et al.
Pubblicazione: (2024)
di: Zhao, Wei, et al.
Pubblicazione: (2024)
SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
di: Zhao, Bingchen, et al.
Pubblicazione: (2026)
di: Zhao, Bingchen, et al.
Pubblicazione: (2026)
SWE-WebDevBench: Evaluating Coding Agent Application Platforms as Virtual Software Agencies
di: Saxena, Siddhant, et al.
Pubblicazione: (2026)
di: Saxena, Siddhant, et al.
Pubblicazione: (2026)
Silo-Bench: A Scalable Environment for Evaluating Distributed Coordination in Multi-Agent LLM Systems
di: Zhang, Yuzhe, et al.
Pubblicazione: (2026)
di: Zhang, Yuzhe, et al.
Pubblicazione: (2026)
DPQuant: Efficient and Differentially-Private Model Training via Dynamic Quantization Scheduling
di: Gao, Yubo, et al.
Pubblicazione: (2025)
di: Gao, Yubo, et al.
Pubblicazione: (2025)
MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents
di: Zhu, Kunlun, et al.
Pubblicazione: (2025)
di: Zhu, Kunlun, et al.
Pubblicazione: (2025)
What Limits Agentic Systems Efficiency?
di: Bian, Song, et al.
Pubblicazione: (2025)
di: Bian, Song, et al.
Pubblicazione: (2025)
Facilitating Trustworthy Human-Agent Collaboration in LLM-based Multi-Agent System oriented Software Engineering
di: Ronanki, Krishna
Pubblicazione: (2025)
di: Ronanki, Krishna
Pubblicazione: (2025)
SAGE: Multi-Agent Self-Evolution for LLM Reasoning
di: Peng, Yulin, et al.
Pubblicazione: (2026)
di: Peng, Yulin, et al.
Pubblicazione: (2026)
Evaluating Collective Behaviour of Hundreds of LLM Agents
di: Willis, Richard, et al.
Pubblicazione: (2026)
di: Willis, Richard, et al.
Pubblicazione: (2026)
Spec Kit Agents: Context-Grounded Agentic Workflows
di: Taghavi, Pardis, et al.
Pubblicazione: (2026)
di: Taghavi, Pardis, et al.
Pubblicazione: (2026)
MACRO-LLM: LLM-Empowered Multi-Agent Collaborative Reasoning under Spatiotemporal Partial Observability
di: Chen, Handi, et al.
Pubblicazione: (2026)
di: Chen, Handi, et al.
Pubblicazione: (2026)
Aegis:An Advanced LLM-Based Multi-Agent for Intelligent Functional Safety Engineering
di: Shi, Lu, et al.
Pubblicazione: (2024)
di: Shi, Lu, et al.
Pubblicazione: (2024)
AgentWebBench: Benchmarking Multi-Agent Coordination in Agentic Web
di: Zhong, Shanshan, et al.
Pubblicazione: (2026)
di: Zhong, Shanshan, et al.
Pubblicazione: (2026)
Can LLM Agents Really Debate? A Controlled Study of Multi-Agent Debate in Logical Reasoning
di: Wu, Haolun, et al.
Pubblicazione: (2025)
di: Wu, Haolun, et al.
Pubblicazione: (2025)
Synergizing Logical Reasoning, Knowledge Management and Collaboration in Multi-Agent LLM System
di: Kostka, Adam, et al.
Pubblicazione: (2025)
di: Kostka, Adam, et al.
Pubblicazione: (2025)
Evaluating Multi-Agent LLM Architectures for Rare Disease Diagnosis
di: Almasoud, Ahmed
Pubblicazione: (2026)
di: Almasoud, Ahmed
Pubblicazione: (2026)
Fairness in Multi-Agent Systems for Software Engineering: An SDLC-Oriented Rapid Review
di: Yang-Smith, Corey, et al.
Pubblicazione: (2026)
di: Yang-Smith, Corey, et al.
Pubblicazione: (2026)
Adaptive Coopetition: Leveraging Coarse Verifier Signals for Resilient Multi-Agent LLM Reasoning
di: Huang, Rui Jerry, et al.
Pubblicazione: (2025)
di: Huang, Rui Jerry, et al.
Pubblicazione: (2025)
Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering
di: Zhou, Chenyu, et al.
Pubblicazione: (2026)
di: Zhou, Chenyu, et al.
Pubblicazione: (2026)
CalBench: Evaluating Coordination-Privacy Trade-offs in Multi-Agent LLMs
di: Zou, Chelsea, et al.
Pubblicazione: (2026)
di: Zou, Chelsea, et al.
Pubblicazione: (2026)
Fairness Driven Multi-Agent Path Finding Problem
di: Anand, Aditi, et al.
Pubblicazione: (2026)
di: Anand, Aditi, et al.
Pubblicazione: (2026)
LoRAFusion: Efficient LoRA Fine-Tuning for LLMs
di: Zhu, Zhanda, et al.
Pubblicazione: (2025)
di: Zhu, Zhanda, et al.
Pubblicazione: (2025)
SkillOps: Managing LLM Agent Skill Libraries as Self-Maintaining Software Ecosystems
di: Pu, Hongji, et al.
Pubblicazione: (2026)
di: Pu, Hongji, et al.
Pubblicazione: (2026)
GateLens: A Reasoning-Enhanced LLM Agent for Automotive Software Release Analytics
di: Khoee, Arsham Gholamzadeh, et al.
Pubblicazione: (2025)
di: Khoee, Arsham Gholamzadeh, et al.
Pubblicazione: (2025)
From General Relation Patterns to Task-Specific Decision-Making in Continual Multi-Agent Coordination
di: Yao, Chang, et al.
Pubblicazione: (2025)
di: Yao, Chang, et al.
Pubblicazione: (2025)
Long Live the Librarian! A Persistent Search Sub-Agent for Energy-Efficient Multi-Agent Software Engineering Systems
di: Cho, Seunghyuk, et al.
Pubblicazione: (2026)
di: Cho, Seunghyuk, et al.
Pubblicazione: (2026)
Swarm Intelligence Enhanced Reasoning: A Density-Driven Framework for LLM-Based Multi-Agent Optimization
di: Zhu, Ying, et al.
Pubblicazione: (2025)
di: Zhu, Ying, et al.
Pubblicazione: (2025)
Analyzing Code Injection Attacks on LLM-based Multi-Agent Systems in Software Development
di: Bowers, Brian, et al.
Pubblicazione: (2025)
di: Bowers, Brian, et al.
Pubblicazione: (2025)
Self-Organizing Agent Network for LLM-based Workflow Automation
di: Xiong, Yiming, et al.
Pubblicazione: (2025)
di: Xiong, Yiming, et al.
Pubblicazione: (2025)
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
di: Jiang, Yixing, et al.
Pubblicazione: (2025)
di: Jiang, Yixing, et al.
Pubblicazione: (2025)
Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems
di: Zhu, Jianing, et al.
Pubblicazione: (2026)
di: Zhu, Jianing, et al.
Pubblicazione: (2026)
Logic of Awareness in Agent's Reasoning
di: Kubono, Yudai, et al.
Pubblicazione: (2023)
di: Kubono, Yudai, et al.
Pubblicazione: (2023)
ReSo: A Reward-driven Self-organizing LLM-based Multi-Agent System for Reasoning Tasks
di: Zhou, Heng, et al.
Pubblicazione: (2025)
di: Zhou, Heng, et al.
Pubblicazione: (2025)
Multi-Agent LLM Committees for Autonomous Software Beta Testing
di: Karanam, Sumanth Bharadwaj Hachalli, et al.
Pubblicazione: (2025)
di: Karanam, Sumanth Bharadwaj Hachalli, et al.
Pubblicazione: (2025)
LLM-Enabled Multi-Agent Systems: Empirical Evaluation and Insights into Emerging Design Patterns & Paradigms
di: Renney, Harri, et al.
Pubblicazione: (2026)
di: Renney, Harri, et al.
Pubblicazione: (2026)
TrajOnco: a multi-agent framework for temporal reasoning over longitudinal EHR for multi-cancer early detection
di: Zeng, Sihang, et al.
Pubblicazione: (2026)
di: Zeng, Sihang, et al.
Pubblicazione: (2026)
DemMA: Dementia Multi-Turn Dialogue Agent with Expert-Guided Reasoning and Action Simulation
di: Song, Yutong, et al.
Pubblicazione: (2026)
di: Song, Yutong, et al.
Pubblicazione: (2026)
AgentConductor: Topology Evolution for Multi-Agent Competition-Level Code Generation
di: Wang, Siyu, et al.
Pubblicazione: (2026)
di: Wang, Siyu, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Aegis: Taxonomy and Optimizations for Overcoming Agent-Environment Failures in LLM Agents
di: Song, Kevin, et al.
Pubblicazione: (2025) -
Tally: Non-Intrusive Performance Isolation for Concurrent Deep Learning Workloads
di: Zhao, Wei, et al.
Pubblicazione: (2024) -
SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
di: Zhao, Bingchen, et al.
Pubblicazione: (2026) -
SWE-WebDevBench: Evaluating Coding Agent Application Platforms as Virtual Software Agencies
di: Saxena, Siddhant, et al.
Pubblicazione: (2026) -
Silo-Bench: A Scalable Environment for Evaluating Distributed Coordination in Multi-Agent LLM Systems
di: Zhang, Yuzhe, et al.
Pubblicazione: (2026)