SWE-QA: A Dataset and Benchmark for Complex Code Understanding
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Elkoussy, Laïla, Perez, Julien |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
SWE-QA-Pro: A Representative Benchmark and Scalable Training Recipe for Repository-Level Code Understanding
par: Cai, Songcheng, et autres
Publié: (2026)
par: Cai, Songcheng, et autres
Publié: (2026)
SWE Context Bench: A Benchmark for Context Learning in Coding
par: Zhu, Jiayuan, et autres
Publié: (2026)
par: Zhu, Jiayuan, et autres
Publié: (2026)
SWE-PRBench: Benchmarking AI Code Review Quality Against Pull Request Feedback
par: Kumar, Deepak
Publié: (2026)
par: Kumar, Deepak
Publié: (2026)
CodeRepoQA: A Large-scale Benchmark for Software Engineering Question Answering
par: Hu, Ruida, et autres
Publié: (2024)
par: Hu, Ruida, et autres
Publié: (2024)
SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades
par: Lam, Man Ho, et autres
Publié: (2026)
par: Lam, Man Ho, et autres
Publié: (2026)
Saving SWE-Bench: A Benchmark Mutation Approach for Realistic Agent Evaluation
par: Garg, Spandan, et autres
Publié: (2025)
par: Garg, Spandan, et autres
Publié: (2025)
SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities for Large Language Models
par: Xu, Jingxuan, et autres
Publié: (2025)
par: Xu, Jingxuan, et autres
Publié: (2025)
SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks
par: Guo, Lianghong, et autres
Publié: (2025)
par: Guo, Lianghong, et autres
Publié: (2025)
SWE-Bench-CL: Continual Learning for Coding Agents
par: Joshi, Thomas, et autres
Publié: (2025)
par: Joshi, Thomas, et autres
Publié: (2025)
Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving
par: Zan, Daoguang, et autres
Publié: (2025)
par: Zan, Daoguang, et autres
Publié: (2025)
CodeSense: a Real-World Benchmark and Dataset for Code Semantic Reasoning
par: Roy, Monoshi Kumar, et autres
Publié: (2025)
par: Roy, Monoshi Kumar, et autres
Publié: (2025)
Resolving Java Code Repository Issues with iSWE Agent
par: Ganhotra, Jatin, et autres
Publié: (2026)
par: Ganhotra, Jatin, et autres
Publié: (2026)
Benchmarking Multimodal LLMs on Code Generation for Complex Interactive Webpages
par: Wu, Fan, et autres
Publié: (2026)
par: Wu, Fan, et autres
Publié: (2026)
FeatureBench: Benchmarking Agentic Coding for Complex Feature Development
par: Zhou, Qixing, et autres
Publié: (2026)
par: Zhou, Qixing, et autres
Publié: (2026)
CodeAssistBench (CAB): Dataset & Benchmarking for Multi-turn Chat-Based Code Assistance
par: Kim, Myeongsoo, et autres
Publié: (2025)
par: Kim, Myeongsoo, et autres
Publié: (2025)
APEX-SWE
par: Kottamasu, Abhi, et autres
Publié: (2026)
par: Kottamasu, Abhi, et autres
Publié: (2026)
SWE-bench-java: A GitHub Issue Resolving Benchmark for Java
par: Zan, Daoguang, et autres
Publié: (2024)
par: Zan, Daoguang, et autres
Publié: (2024)
SWE-chat: Coding Agent Interactions From Real Users in the Wild
par: Baumann, Joachim, et autres
Publié: (2026)
par: Baumann, Joachim, et autres
Publié: (2026)
BugPilot: Complex Bug Generation for Efficient Learning of SWE Skills
par: Sonwane, Atharv, et autres
Publié: (2025)
par: Sonwane, Atharv, et autres
Publié: (2025)
SWE-Universe: Scale Real-World Verifiable Environments to Millions
par: Chen, Mouxiang, et autres
Publié: (2026)
par: Chen, Mouxiang, et autres
Publié: (2026)
TOM-SWE: User Mental Modeling For Software Engineering Agents
par: Zhou, Xuhui, et autres
Publié: (2025)
par: Zhou, Xuhui, et autres
Publié: (2025)
SWE-Next: Scalable Real-World Software Engineering Tasks for Agents
par: Liang, Jiarong, et autres
Publié: (2026)
par: Liang, Jiarong, et autres
Publié: (2026)
AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation
par: Sahoo, Priyam, et autres
Publié: (2026)
par: Sahoo, Priyam, et autres
Publié: (2026)
When Agents go Astray: Course-Correcting SWE Agents with PRMs
par: Gandhi, Shubham, et autres
Publié: (2025)
par: Gandhi, Shubham, et autres
Publié: (2025)
SWE-Hub: A Unified Production System for Scalable, Executable Software Engineering Tasks
par: Zeng, Yucheng, et autres
Publié: (2026)
par: Zeng, Yucheng, et autres
Publié: (2026)
CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution
par: Gu, Alex, et autres
Publié: (2024)
par: Gu, Alex, et autres
Publié: (2024)
The SWE-Bench Illusion: When State-of-the-Art LLMs Remember Instead of Reason
par: Liang, Shanchao, et autres
Publié: (2025)
par: Liang, Shanchao, et autres
Publié: (2025)
SWE-bench Goes Live!
par: Zhang, Linghao, et autres
Publié: (2025)
par: Zhang, Linghao, et autres
Publié: (2025)
CodeWatcher: IDE Telemetry Data Extraction Tool for Understanding Coding Interactions with LLMs
par: Basha, Manaal, et autres
Publié: (2025)
par: Basha, Manaal, et autres
Publié: (2025)
SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks
par: Adamenko, Pavel, et autres
Publié: (2025)
par: Adamenko, Pavel, et autres
Publié: (2025)
SWE-AGI: Benchmarking Specification-Driven Software Construction with MoonBit in the Era of Autonomous Agents
par: Zhang, Zhirui, et autres
Publié: (2026)
par: Zhang, Zhirui, et autres
Publié: (2026)
A Benchmark for Localizing Code and Non-Code Issues in Software Projects
par: Zhang, Zejun, et autres
Publié: (2025)
par: Zhang, Zejun, et autres
Publié: (2025)
RedCode: Risky Code Execution and Generation Benchmark for Code Agents
par: Guo, Chengquan, et autres
Publié: (2024)
par: Guo, Chengquan, et autres
Publié: (2024)
Code Review Agent Benchmark
par: Zhang, Yuntong, et autres
Publié: (2026)
par: Zhang, Yuntong, et autres
Publié: (2026)
SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering?
par: Han, Tingxu, et autres
Publié: (2026)
par: Han, Tingxu, et autres
Publié: (2026)
SWE-Effi: Re-Evaluating Software AI Agent System Effectiveness Under Resource Constraints
par: Fan, Zhiyu, et autres
Publié: (2025)
par: Fan, Zhiyu, et autres
Publié: (2025)
Lingma SWE-GPT: An Open Development-Process-Centric Language Model for Automated Software Improvement
par: Ma, Yingwei, et autres
Publié: (2024)
par: Ma, Yingwei, et autres
Publié: (2024)
SPICE: An Automated SWE-Bench Labeling Pipeline for Issue Clarity, Test Coverage, and Effort Estimation
par: Oliva, Gustavo A., et autres
Publié: (2025)
par: Oliva, Gustavo A., et autres
Publié: (2025)
AKD : Adversarial Knowledge Distillation For Large Language Models Alignment on Coding tasks
par: Oulkadda, Ilyas, et autres
Publié: (2025)
par: Oulkadda, Ilyas, et autres
Publié: (2025)
FastCode: Fast and Cost-Efficient Code Understanding and Reasoning
par: Li, Zhonghang, et autres
Publié: (2026)
par: Li, Zhonghang, et autres
Publié: (2026)
Documents similaires
-
SWE-QA-Pro: A Representative Benchmark and Scalable Training Recipe for Repository-Level Code Understanding
par: Cai, Songcheng, et autres
Publié: (2026) -
SWE Context Bench: A Benchmark for Context Learning in Coding
par: Zhu, Jiayuan, et autres
Publié: (2026) -
SWE-PRBench: Benchmarking AI Code Review Quality Against Pull Request Feedback
par: Kumar, Deepak
Publié: (2026) -
CodeRepoQA: A Large-scale Benchmark for Software Engineering Question Answering
par: Hu, Ruida, et autres
Publié: (2024) -
SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades
par: Lam, Man Ho, et autres
Publié: (2026)