Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zan, Daoguang, Huang, Zhirong, Liu, Wei, Chen, Hanwu, Zhang, Linhao, Xin, Shulin, Chen, Lu, Liu, Qi, Zhong, Xiaojian, Li, Aoyan, Liu, Siyao, Xiao, Yongsheng, Chen, Liangqiang, Zhang, Yuyu, Su, Jing, Liu, Tianyu, Long, Rui, Shen, Kai, Xiang, Liang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SWE-Mirror: Scaling Issue-Resolving Datasets by Mirroring Issues Across Repositories
von: Wang, Junhao, et al.
Veröffentlicht: (2025)
von: Wang, Junhao, et al.
Veröffentlicht: (2025)
SWE-bench-java: A GitHub Issue Resolving Benchmark for Java
von: Zan, Daoguang, et al.
Veröffentlicht: (2024)
von: Zan, Daoguang, et al.
Veröffentlicht: (2024)
CodeV: Issue Resolving with Visual Data
von: Zhang, Linhao, et al.
Veröffentlicht: (2024)
von: Zhang, Linhao, et al.
Veröffentlicht: (2024)
SWE-bench Goes Live!
von: Zhang, Linghao, et al.
Veröffentlicht: (2025)
von: Zhang, Linghao, et al.
Veröffentlicht: (2025)
Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study
von: Wang, You, et al.
Veröffentlicht: (2025)
von: Wang, You, et al.
Veröffentlicht: (2025)
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
von: Jimenez, Carlos E., et al.
Veröffentlicht: (2023)
von: Jimenez, Carlos E., et al.
Veröffentlicht: (2023)
Investigating Test Overfitting on SWE-bench
von: Ahmed, Toufique, et al.
Veröffentlicht: (2025)
von: Ahmed, Toufique, et al.
Veröffentlicht: (2025)
HE-SNR: Uncovering Latent Logic via Entropy for Guiding Mid-Training on SWE-bench
von: Wang, Yueyang, et al.
Veröffentlicht: (2026)
von: Wang, Yueyang, et al.
Veröffentlicht: (2026)
FML-bench: Benchmarking Machine Learning Agents for Scientific Research
von: Zou, Qiran, et al.
Veröffentlicht: (2025)
von: Zou, Qiran, et al.
Veröffentlicht: (2025)
MMInA: Benchmarking Multihop Multimodal Internet Agents
von: Tian, Shulin, et al.
Veröffentlicht: (2024)
von: Tian, Shulin, et al.
Veröffentlicht: (2024)
A Quantum-based Database Query Scheme for Privacy Preservation in Cloud Environment
von: Liu, Wenjie, et al.
Veröffentlicht: (2020)
von: Liu, Wenjie, et al.
Veröffentlicht: (2020)
A Unitary Operator Construction Solution Based on Pauli Group for Maximal Dense Coding
von: Liu, Wenjie, et al.
Veröffentlicht: (2023)
von: Liu, Wenjie, et al.
Veröffentlicht: (2023)
The Devil is in the Neurons: Interpreting and Mitigating Social Biases in Pre-trained Language Models
von: Liu, Yan, et al.
Veröffentlicht: (2024)
von: Liu, Yan, et al.
Veröffentlicht: (2024)
Multi-dimensional reflected BSDEs driven by $G$-Brownian motion with diagonal generators
von: Li, Hanwu, et al.
Veröffentlicht: (2021)
von: Li, Hanwu, et al.
Veröffentlicht: (2021)
CNSL-bench: Benchmarking the Sign Language Understanding Capabilities of MLLMs on Chinese National Sign Language
von: Zhao, Rui, et al.
Veröffentlicht: (2026)
von: Zhao, Rui, et al.
Veröffentlicht: (2026)
SWE-QA-Pro: A Representative Benchmark and Scalable Training Recipe for Repository-Level Code Understanding
von: Cai, Songcheng, et al.
Veröffentlicht: (2026)
von: Cai, Songcheng, et al.
Veröffentlicht: (2026)
SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
von: Yang, John, et al.
Veröffentlicht: (2024)
von: Yang, John, et al.
Veröffentlicht: (2024)
Aligning CodeLLMs with Direct Preference Optimization
von: Miao, Yibo, et al.
Veröffentlicht: (2024)
von: Miao, Yibo, et al.
Veröffentlicht: (2024)
A GAN-based data poisoning framework against anomaly detection in vertical federated learning
von: Chen, Xiaolin, et al.
Veröffentlicht: (2024)
von: Chen, Xiaolin, et al.
Veröffentlicht: (2024)
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
von: Duston, Titouan, et al.
Veröffentlicht: (2025)
von: Duston, Titouan, et al.
Veröffentlicht: (2025)
RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories
von: Wang, Yanlin, et al.
Veröffentlicht: (2026)
von: Wang, Yanlin, et al.
Veröffentlicht: (2026)
CL-bench: A Benchmark for Context Learning
von: Dou, Shihan, et al.
Veröffentlicht: (2026)
von: Dou, Shihan, et al.
Veröffentlicht: (2026)
Adiabatic-Passage-Based Parameter Setting for Quantum Approximate Optimization Algorithm
von: Wu, Mingyou, et al.
Veröffentlicht: (2023)
von: Wu, Mingyou, et al.
Veröffentlicht: (2023)
Exploring the Potential of Conversational Test Suite Based Program Repair on SWE-bench
von: Cheshkov, Anton, et al.
Veröffentlicht: (2024)
von: Cheshkov, Anton, et al.
Veröffentlicht: (2024)
Can Old Tests Do New Tricks for Resolving SWE Issues?
von: Chen, Yang, et al.
Veröffentlicht: (2025)
von: Chen, Yang, et al.
Veröffentlicht: (2025)
Seed-Coder: Let the Code Model Curate Data for Itself
von: Seed, ByteDance, et al.
Veröffentlicht: (2025)
von: Seed, ByteDance, et al.
Veröffentlicht: (2025)
Kimi-Dev: Agentless Training as Skill Prior for SWE-Agents
von: Yang, Zonghan, et al.
Veröffentlicht: (2025)
von: Yang, Zonghan, et al.
Veröffentlicht: (2025)
Elephant in the Room: Unveiling the Impact of Reward Model Quality in Alignment
von: Liu, Yan, et al.
Veröffentlicht: (2024)
von: Liu, Yan, et al.
Veröffentlicht: (2024)
Microenvironment Self‐Adaptive Nanoarmor to Address Adhesion‐ and Colonization‐Related Obstacles in Impaired Intestine Promote Bacteriotherapy Against Parkinson's Disease
von: Limeng Zhu, et al.
Veröffentlicht: (2026)
von: Limeng Zhu, et al.
Veröffentlicht: (2026)
SWE-AGILE: A Software Agent Framework for Efficiently Managing Dynamic Reasoning Context
von: Lian, Shuquan, et al.
Veröffentlicht: (2026)
von: Lian, Shuquan, et al.
Veröffentlicht: (2026)
MDPBench: A Benchmark for Multilingual Document Parsing in Real-World Scenarios
von: Li, Zhang, et al.
Veröffentlicht: (2026)
von: Li, Zhang, et al.
Veröffentlicht: (2026)
Improving Natural Language Capability of Code Large Language Model
von: Li, Wei, et al.
Veröffentlicht: (2024)
von: Li, Wei, et al.
Veröffentlicht: (2024)
SWE-AGI: Benchmarking Specification-Driven Software Construction with MoonBit in the Era of Autonomous Agents
von: Zhang, Zhirui, et al.
Veröffentlicht: (2026)
von: Zhang, Zhirui, et al.
Veröffentlicht: (2026)
SWE-Cycle: Benchmarking Code Agents across the Complete Issue Resolution Cycle
von: Guan, Hao, et al.
Veröffentlicht: (2026)
von: Guan, Hao, et al.
Veröffentlicht: (2026)
Preparation investigation of photocured printing shape memory graphene/epoxy resin composites
von: Long Chen, et al.
Veröffentlicht: (2025)
von: Long Chen, et al.
Veröffentlicht: (2025)
GraphCoder: Enhancing Repository-Level Code Completion via Code Context Graph-based Retrieval and Language Model
von: Liu, Wei, et al.
Veröffentlicht: (2024)
von: Liu, Wei, et al.
Veröffentlicht: (2024)
SWE-Next: Scalable Real-World Software Engineering Tasks for Agents
von: Liang, Jiarong, et al.
Veröffentlicht: (2026)
von: Liang, Jiarong, et al.
Veröffentlicht: (2026)
Revisiting the Tertiary-induced Binary Black Hole Mergers: the Role of Superthermal Wide Tertiary Eccentricity Distributions
von: Su, Yubo, et al.
Veröffentlicht: (2024)
von: Su, Yubo, et al.
Veröffentlicht: (2024)
PredictionMarketBench: A SWE-bench-Style Framework for Backtesting Trading Agents on Prediction Markets
von: Arora, Avi, et al.
Veröffentlicht: (2026)
von: Arora, Avi, et al.
Veröffentlicht: (2026)
Conditional Expectation Backward Stochastic Differential Equations and Related Backward Stochastic Differential Equations with Conditional Reflection
von: Li, Hanwu
Veröffentlicht: (2025)
von: Li, Hanwu
Veröffentlicht: (2025)
Ähnliche Einträge
-
SWE-Mirror: Scaling Issue-Resolving Datasets by Mirroring Issues Across Repositories
von: Wang, Junhao, et al.
Veröffentlicht: (2025) -
SWE-bench-java: A GitHub Issue Resolving Benchmark for Java
von: Zan, Daoguang, et al.
Veröffentlicht: (2024) -
CodeV: Issue Resolving with Visual Data
von: Zhang, Linhao, et al.
Veröffentlicht: (2024) -
SWE-bench Goes Live!
von: Zhang, Linghao, et al.
Veröffentlicht: (2025) -
Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study
von: Wang, You, et al.
Veröffentlicht: (2025)