SWE Context Bench: A Benchmark for Context Learning in Coding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Jiayuan, Wu, Junde, Hu, Minhao, Zhu, Shengda, Pan, Jiazhen, Shen, Weixiang, Yang, Yijun, Liu, Fenglin, Hao, Jianye, Jin, Yueming, Ho, Qirong, Xu, Min |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Git Context Controller: Manage the Context of LLM-based Agents like Git
von: Wu, Junde, et al.
Veröffentlicht: (2025)
von: Wu, Junde, et al.
Veröffentlicht: (2025)
SWE-Bench+: Enhanced Coding Benchmark for LLMs
von: Aleithan, Reem, et al.
Veröffentlicht: (2024)
von: Aleithan, Reem, et al.
Veröffentlicht: (2024)
UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench
von: Yu, Boxi, et al.
Veröffentlicht: (2025)
von: Yu, Boxi, et al.
Veröffentlicht: (2025)
SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents
von: Wang, Yuhang, et al.
Veröffentlicht: (2026)
von: Wang, Yuhang, et al.
Veröffentlicht: (2026)
SWE-Cycle: Benchmarking Code Agents across the Complete Issue Resolution Cycle
von: Guan, Hao, et al.
Veröffentlicht: (2026)
von: Guan, Hao, et al.
Veröffentlicht: (2026)
Ask Patients with Patience: Enabling LLMs for Human-Centric Medical Dialogue with Grounded Reasoning
von: Zhu, Jiayuan, et al.
Veröffentlicht: (2025)
von: Zhu, Jiayuan, et al.
Veröffentlicht: (2025)
What's in a Benchmark? The Case of SWE-Bench in Automated Program Repair
von: Martinez, Matias, et al.
Veröffentlicht: (2026)
von: Martinez, Matias, et al.
Veröffentlicht: (2026)
LONGCODEU: Benchmarking Long-Context Language Models on Long Code Understanding
von: Li, Jia, et al.
Veröffentlicht: (2025)
von: Li, Jia, et al.
Veröffentlicht: (2025)
SWE-Bench-CL: Continual Learning for Coding Agents
von: Joshi, Thomas, et al.
Veröffentlicht: (2025)
von: Joshi, Thomas, et al.
Veröffentlicht: (2025)
SWE-Sharp-Bench: A Reproducible Benchmark for C# Software Engineering Tasks
von: Mhatre, Sanket, et al.
Veröffentlicht: (2025)
von: Mhatre, Sanket, et al.
Veröffentlicht: (2025)
SWE-TRACE: Optimizing Long-Horizon SWE Agents Through Rubric Process Reward Models and Heuristic Test-Time Scaling
von: Han, Hao, et al.
Veröffentlicht: (2026)
von: Han, Hao, et al.
Veröffentlicht: (2026)
In Line with Context: Repository-Level Code Generation via Context Inlining
von: Hu, Chao, et al.
Veröffentlicht: (2026)
von: Hu, Chao, et al.
Veröffentlicht: (2026)
Saving SWE-Bench: A Benchmark Mutation Approach for Realistic Agent Evaluation
von: Garg, Spandan, et al.
Veröffentlicht: (2025)
von: Garg, Spandan, et al.
Veröffentlicht: (2025)
SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution
von: Raghavendra, Mohit, et al.
Veröffentlicht: (2026)
von: Raghavendra, Mohit, et al.
Veröffentlicht: (2026)
SWE-QA: A Dataset and Benchmark for Complex Code Understanding
von: Elkoussy, Laïla, et al.
Veröffentlicht: (2026)
von: Elkoussy, Laïla, et al.
Veröffentlicht: (2026)
SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades
von: Lam, Man Ho, et al.
Veröffentlicht: (2026)
von: Lam, Man Ho, et al.
Veröffentlicht: (2026)
CATCODER: Repository-Level Code Generation with Relevant Code and Type Context
von: Pan, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Pan, Zhiyuan, et al.
Veröffentlicht: (2024)
One-Prompt to Segment All Medical Images
von: Wu, Junde, et al.
Veröffentlicht: (2023)
von: Wu, Junde, et al.
Veröffentlicht: (2023)
Not just Birds and Cars: Generic, Scalable and Explainable Models for Professional Visual Recognition
von: Wu, Junde, et al.
Veröffentlicht: (2024)
von: Wu, Junde, et al.
Veröffentlicht: (2024)
AACR-Bench: Evaluating Automatic Code Review with Holistic Repository-Level Context
von: Zhang, Lei, et al.
Veröffentlicht: (2026)
von: Zhang, Lei, et al.
Veröffentlicht: (2026)
Evo: Autoregressive-Diffusion Large Language Models with Evolving Balance
von: Wu, Junde, et al.
Veröffentlicht: (2026)
von: Wu, Junde, et al.
Veröffentlicht: (2026)
MRG-Bench: Evaluating and Exploring the Requirements of Context for Repository-Level Code Generation
von: Li, Haiyang
Veröffentlicht: (2025)
von: Li, Haiyang
Veröffentlicht: (2025)
SWE-WebDevBench: Evaluating Coding Agent Application Platforms as Virtual Software Agencies
von: Saxena, Siddhant, et al.
Veröffentlicht: (2026)
von: Saxena, Siddhant, et al.
Veröffentlicht: (2026)
DSCodeBench: A Realistic Benchmark for Data Science Code Generation
von: Ouyang, Shuyin, et al.
Veröffentlicht: (2025)
von: Ouyang, Shuyin, et al.
Veröffentlicht: (2025)
SWE-Bench Mobile: Can Large Language Model Agents Develop Industry-Level Mobile Applications?
von: Tian, Muxin, et al.
Veröffentlicht: (2026)
von: Tian, Muxin, et al.
Veröffentlicht: (2026)
SWE-Refactor: A Repository-Level Benchmark for Real-World LLM-Based Code Refactoring
von: Xu, Yisen, et al.
Veröffentlicht: (2026)
von: Xu, Yisen, et al.
Veröffentlicht: (2026)
Does SWE-Bench-Verified Test Agent Ability or Model Memory?
von: Prathifkumar, Thanosan, et al.
Veröffentlicht: (2025)
von: Prathifkumar, Thanosan, et al.
Veröffentlicht: (2025)
Benchmarking LLMs for Fine-Grained Code Review with Enriched Context in Practice
von: Hu, Ruida, et al.
Veröffentlicht: (2025)
von: Hu, Ruida, et al.
Veröffentlicht: (2025)
LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering
von: Qiu, Jielin, et al.
Veröffentlicht: (2025)
von: Qiu, Jielin, et al.
Veröffentlicht: (2025)
RepoMirage: Probing Repository Context Reasoning in Code Agents with Perturbations
von: Li, Hanyu, et al.
Veröffentlicht: (2026)
von: Li, Hanyu, et al.
Veröffentlicht: (2026)
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
von: Chou, Jason, et al.
Veröffentlicht: (2025)
von: Chou, Jason, et al.
Veröffentlicht: (2025)
SWE-Shepherd: Advancing PRMs for Reinforcing Code Agents
von: Dihan, Mahir Labib, et al.
Veröffentlicht: (2026)
von: Dihan, Mahir Labib, et al.
Veröffentlicht: (2026)
LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering
von: Qiu, Jielin, et al.
Veröffentlicht: (2025)
von: Qiu, Jielin, et al.
Veröffentlicht: (2025)
SWE-PRBench: Benchmarking AI Code Review Quality Against Pull Request Feedback
von: Kumar, Deepak
Veröffentlicht: (2026)
von: Kumar, Deepak
Veröffentlicht: (2026)
MedOpenClaw and MedFlowBench: Auditing Medical Agents in Full-Study Workflows
von: Shen, Weixiang, et al.
Veröffentlicht: (2026)
von: Shen, Weixiang, et al.
Veröffentlicht: (2026)
FeatureBench: Benchmarking Agentic Coding for Complex Feature Development
von: Zhou, Qixing, et al.
Veröffentlicht: (2026)
von: Zhou, Qixing, et al.
Veröffentlicht: (2026)
SWE-QA-Pro: A Representative Benchmark and Scalable Training Recipe for Repository-Level Code Understanding
von: Cai, Songcheng, et al.
Veröffentlicht: (2026)
von: Cai, Songcheng, et al.
Veröffentlicht: (2026)
MOSS: Enabling Code-Driven Evolution and Context Management for AI Agents
von: Zhu, Ming, et al.
Veröffentlicht: (2024)
von: Zhu, Ming, et al.
Veröffentlicht: (2024)
SWE-Manager: Selecting and Synthesizing Golden Proposals Before Coding
von: Tan, Boyin, et al.
Veröffentlicht: (2026)
von: Tan, Boyin, et al.
Veröffentlicht: (2026)
aiXcoder-7B-v2: Training LLMs to Fully Utilize the Long Context in Repository-level Code Completion
von: Li, Jia, et al.
Veröffentlicht: (2025)
von: Li, Jia, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Git Context Controller: Manage the Context of LLM-based Agents like Git
von: Wu, Junde, et al.
Veröffentlicht: (2025) -
SWE-Bench+: Enhanced Coding Benchmark for LLMs
von: Aleithan, Reem, et al.
Veröffentlicht: (2024) -
UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench
von: Yu, Boxi, et al.
Veröffentlicht: (2025) -
SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents
von: Wang, Yuhang, et al.
Veröffentlicht: (2026) -
SWE-Cycle: Benchmarking Code Agents across the Complete Issue Resolution Cycle
von: Guan, Hao, et al.
Veröffentlicht: (2026)