Similar Items
ContractEval: A Benchmark for Evaluating Contract-Satisfying Assertions in Code Generation
by: Lim, Soohan, et al.
Published: (2025)
by: Lim, Soohan, et al.
Published: (2025)
A transfer learning approach for automatic conflicts detection in software requirement sentence pairs based on dual encoders
by: Wang, Yizheng, et al.
Published: (2025)
by: Wang, Yizheng, et al.
Published: (2025)
Tool-Genesis: A Task-Driven Tool Creation Benchmark for Self-Evolving Language Agent
by: Xia, Bowei, et al.
Published: (2026)
by: Xia, Bowei, et al.
Published: (2026)
Source Code Summarization in the Era of Large Language Models
by: Sun, Weisong, et al.
Published: (2024)
by: Sun, Weisong, et al.
Published: (2024)
Mind the Metrics: Patterns for Telemetry-Aware In-IDE AI Application Development using the Model Context Protocol (MCP)
by: Koc, Vincent, et al.
Published: (2025)
by: Koc, Vincent, et al.
Published: (2025)
TCProF: Time-Complexity Prediction SSL Framework
by: Hahn, Joonghyuk, et al.
Published: (2025)
by: Hahn, Joonghyuk, et al.
Published: (2025)
MEC$^3$O: Multi-Expert Consensus for Code Time Complexity Prediction
by: Hahn, Joonghyuk, et al.
Published: (2025)
by: Hahn, Joonghyuk, et al.
Published: (2025)
QHackBench: Benchmarking Large Language Models for Quantum Code Generation Using PennyLane Hackathon Challenges
by: Basit, Abdul, et al.
Published: (2025)
by: Basit, Abdul, et al.
Published: (2025)
SLEAN: Simple Lightweight Ensemble Analysis Network for Multi-Provider LLM Coordination: Design, Implementation, and Vibe Coding Bug Investigation Case Study
by: Vargas, Matheus J. T.
Published: (2025)
by: Vargas, Matheus J. T.
Published: (2025)
CWM: An Open-Weights LLM for Research on Code Generation with World Models
by: FAIR CodeGen team, et al.
Published: (2025)
by: FAIR CodeGen team, et al.
Published: (2025)
XPath Agent: An Efficient XPath Programming Agent Based on LLM for Web Crawler
by: Li, Yu, et al.
Published: (2024)
by: Li, Yu, et al.
Published: (2024)
From Understanding to Excelling: Template-Free Algorithm Design through Structural-Functional Co-Evolution
by: Zhao, Zhe, et al.
Published: (2025)
by: Zhao, Zhe, et al.
Published: (2025)
SATA-BENCH: Select All That Apply Benchmark for Multiple Choice Questions
by: Xu, Weijie, et al.
Published: (2025)
by: Xu, Weijie, et al.
Published: (2025)
Collaborative LLM Agents for C4 Software Architecture Design Automation
by: Szczepanik, Kamil, et al.
Published: (2025)
by: Szczepanik, Kamil, et al.
Published: (2025)
Simple and Effective Baselines for Code Summarisation Evaluation
by: Robinson, Jade, et al.
Published: (2025)
by: Robinson, Jade, et al.
Published: (2025)
Bug In the Code Stack: Can LLMs Find Bugs in Large Python Code Stacks
by: Lee, Hokyung, et al.
Published: (2024)
by: Lee, Hokyung, et al.
Published: (2024)
When LLM meets Fuzzy-TOPSIS for Personnel Selection through Automated Profile Analysis
by: Hoque, Shahria, et al.
Published: (2026)
by: Hoque, Shahria, et al.
Published: (2026)
A Framework for Testing and Adapting REST APIs as LLM Tools
by: Bandlamudi, Jayachandu, et al.
Published: (2025)
by: Bandlamudi, Jayachandu, et al.
Published: (2025)
ESALE: Enhancing Code-Summary Alignment Learning for Source Code Summarization
by: Fang, Chunrong, et al.
Published: (2024)
by: Fang, Chunrong, et al.
Published: (2024)
Mind the GAP: Text Safety Does Not Transfer to Tool-Call Safety in LLM Agents
by: Cartagena, Arnold, et al.
Published: (2026)
by: Cartagena, Arnold, et al.
Published: (2026)
Comprehensive Evaluation and Insights into the Use of Large Language Models in the Automation of Behavior-Driven Development Acceptance Test Formulation
by: Karpurapu, Shanthi, et al.
Published: (2024)
by: Karpurapu, Shanthi, et al.
Published: (2024)
Feature-Factory: Automating Software Feature Integration Using Generative AI
by: Vsevolodovna, Ruslan Idelfonso Magana
Published: (2024)
by: Vsevolodovna, Ruslan Idelfonso Magana
Published: (2024)
An Epidemiological Knowledge Graph extracted from the World Health Organization's Disease Outbreak News
by: Consoli, Sergio, et al.
Published: (2025)
by: Consoli, Sergio, et al.
Published: (2025)
AgentLeak: A Full-Stack Benchmark for Privacy Leakage in Multi-Agent LLM Systems
by: Yagoubi, Faouzi El, et al.
Published: (2026)
by: Yagoubi, Faouzi El, et al.
Published: (2026)
Controlling Long-Horizon Behavior in Language Model Agents with Explicit State Dynamics
by: Subaharan, Sukesh
Published: (2026)
by: Subaharan, Sukesh
Published: (2026)
Knowledge-Guided Multi-Agent Framework for Automated Requirements Development: A Vision
by: Huang, Jiangping, et al.
Published: (2025)
by: Huang, Jiangping, et al.
Published: (2025)
Commenting Higher-level Code Unit: Full Code, Reduced Code, or Hierarchical Code Summarization
by: Sun, Weisong, et al.
Published: (2025)
by: Sun, Weisong, et al.
Published: (2025)
LLMs taking shortcuts in test generation: A study with SAP HANA and LevelDB
by: Bekmyradov, Vekil, et al.
Published: (2026)
by: Bekmyradov, Vekil, et al.
Published: (2026)
When Many-Shot Prompting Fails: An Empirical Study of LLM Code Translation
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
Learning Software Bug Reports: A Systematic Literature Review
by: Long, Guoming, et al.
Published: (2025)
by: Long, Guoming, et al.
Published: (2025)
BEATS: Bias Evaluation and Assessment Test Suite for Large Language Models
by: Abhishek, Alok, et al.
Published: (2025)
by: Abhishek, Alok, et al.
Published: (2025)
Distilling Desired Comments for Enhanced Code Review with Large Language Models
by: Yu, Yongda, et al.
Published: (2024)
by: Yu, Yongda, et al.
Published: (2024)
The Impact of Large Language Models on Open-source Innovation: Evidence from GitHub Copilot
by: Yeverechyahu, Doron, et al.
Published: (2024)
by: Yeverechyahu, Doron, et al.
Published: (2024)
Leveraging Large Language Models for Use Case Model Generation from Software Requirements
by: Eisenreich, Tobias, et al.
Published: (2025)
by: Eisenreich, Tobias, et al.
Published: (2025)
Large Language Models are Inconsistent and Biased Evaluators
by: Stureborg, Rickard, et al.
Published: (2024)
by: Stureborg, Rickard, et al.
Published: (2024)
ContractBench: Can LLM Agents Preserve Observation Contracts?
by: Wang, Jicheng, et al.
Published: (2026)
by: Wang, Jicheng, et al.
Published: (2026)
SHARP: Social Harm Analysis via Risk Profiles for Measuring Inequities in Large Language Models
by: Abhishek, Alok, et al.
Published: (2026)
by: Abhishek, Alok, et al.
Published: (2026)
MicroRemed: Benchmarking LLMs in Microservices Remediation
by: Zhang, Lingzhe, et al.
Published: (2025)
by: Zhang, Lingzhe, et al.
Published: (2025)
Generative AI Toolkit -- a framework for increasing the quality of LLM-based applications over their whole life cycle
by: Kohl, Jens, et al.
Published: (2024)
by: Kohl, Jens, et al.
Published: (2024)
Autonomous Navigation and Collision Avoidance for Mobile Robots: Classification and Review
by: de Carvalho, Marcus Vinicius Leal, et al.
Published: (2024)
by: de Carvalho, Marcus Vinicius Leal, et al.
Published: (2024)
Similar Items
-
ContractEval: A Benchmark for Evaluating Contract-Satisfying Assertions in Code Generation
by: Lim, Soohan, et al.
Published: (2025) -
A transfer learning approach for automatic conflicts detection in software requirement sentence pairs based on dual encoders
by: Wang, Yizheng, et al.
Published: (2025) -
Tool-Genesis: A Task-Driven Tool Creation Benchmark for Self-Evolving Language Agent
by: Xia, Bowei, et al.
Published: (2026) -
Source Code Summarization in the Era of Large Language Models
by: Sun, Weisong, et al.
Published: (2024) -
Mind the Metrics: Patterns for Telemetry-Aware In-IDE AI Application Development using the Model Context Protocol (MCP)
by: Koc, Vincent, et al.
Published: (2025)