SWE-QA-Pro: A Representative Benchmark and Scalable Training Recipe for Repository-Level Code Understanding
Fuente:
arXiv
Salvato in:
| Autori principali: | Cai, Songcheng, Lyu, Zhiheng, Ni, Yuansheng, Chen, Xiangchao, Zhou, Baichuan, Zhu, Shenzhe, Lu, Yi, Wang, Haozhe, Ruan, Chi, Schneider, Benjamin, Zhang, Weixu, Li, Xiang, Zheng, Andy, Zhang, Yuyu, Nie, Ping, Chen, Wenhu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SWE-Next: Scalable Real-World Software Engineering Tasks for Agents
di: Liang, Jiarong, et al.
Pubblicazione: (2026)
di: Liang, Jiarong, et al.
Pubblicazione: (2026)
VisCoder2: Building Multi-Language Visualization Coding Agents
di: Ni, Yuansheng, et al.
Pubblicazione: (2025)
di: Ni, Yuansheng, et al.
Pubblicazione: (2025)
VisCoder: Fine-Tuning LLMs for Executable Python Visualization Code Generation
di: Ni, Yuansheng, et al.
Pubblicazione: (2025)
di: Ni, Yuansheng, et al.
Pubblicazione: (2025)
SWE-QA: Can Language Models Answer Repository-level Code Questions?
di: Peng, Weihan, et al.
Pubblicazione: (2025)
di: Peng, Weihan, et al.
Pubblicazione: (2025)
PixelWorld: How Far Are We from Perceiving Everything as Pixels?
di: Lyu, Zhiheng, et al.
Pubblicazione: (2025)
di: Lyu, Zhiheng, et al.
Pubblicazione: (2025)
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
di: Yue, Xiang, et al.
Pubblicazione: (2024)
di: Yue, Xiang, et al.
Pubblicazione: (2024)
ScholarCopilot: Training Large Language Models for Academic Writing with Accurate Citations
di: Wang, Yubo, et al.
Pubblicazione: (2025)
di: Wang, Yubo, et al.
Pubblicazione: (2025)
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
di: Ma, Wentao, et al.
Pubblicazione: (2025)
di: Ma, Wentao, et al.
Pubblicazione: (2025)
SWE-Spot: Building Small Repo-Experts with Repository-Centric Learning
di: Peng, Jinjun, et al.
Pubblicazione: (2026)
di: Peng, Jinjun, et al.
Pubblicazione: (2026)
ACECODER: Acing Coder RL via Automated Test-Case Synthesis
di: Zeng, Huaye, et al.
Pubblicazione: (2025)
di: Zeng, Huaye, et al.
Pubblicazione: (2025)
SWE-Bench++: A Framework for the Scalable Generation of Software Engineering Benchmarks from Open-Source Repositories
di: Wang, Lilin, et al.
Pubblicazione: (2025)
di: Wang, Lilin, et al.
Pubblicazione: (2025)
RewardHarness: Self-Evolving Agentic Post-Training
di: Zhang, Yuxuan, et al.
Pubblicazione: (2026)
di: Zhang, Yuxuan, et al.
Pubblicazione: (2026)
VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
di: Jiang, Dongfu, et al.
Pubblicazione: (2025)
di: Jiang, Dongfu, et al.
Pubblicazione: (2025)
SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories
di: Soni, Aditya Bharat, et al.
Pubblicazione: (2026)
di: Soni, Aditya Bharat, et al.
Pubblicazione: (2026)
SWE-Refactor: A Repository-Level Benchmark for Real-World LLM-Based Code Refactoring
di: Xu, Yisen, et al.
Pubblicazione: (2026)
di: Xu, Yisen, et al.
Pubblicazione: (2026)
Resolving Java Code Repository Issues with iSWE Agent
di: Ganhotra, Jatin, et al.
Pubblicazione: (2026)
di: Ganhotra, Jatin, et al.
Pubblicazione: (2026)
GenAI Arena: An Open Evaluation Platform for Generative Models
di: Jiang, Dongfu, et al.
Pubblicazione: (2024)
di: Jiang, Dongfu, et al.
Pubblicazione: (2024)
EvolveCoder: Evolving Test Cases via Adversarial Verification for Code Reinforcement Learning
di: Ruan, Chi, et al.
Pubblicazione: (2026)
di: Ruan, Chi, et al.
Pubblicazione: (2026)
SWE-QA: A Dataset and Benchmark for Complex Code Understanding
di: Elkoussy, Laïla, et al.
Pubblicazione: (2026)
di: Elkoussy, Laïla, et al.
Pubblicazione: (2026)
HarmTransform: Transforming Explicit Harmful Queries into Stealthy via Multi-Agent Debate
di: Zhu, Shenzhe
Pubblicazione: (2025)
di: Zhu, Shenzhe
Pubblicazione: (2025)
RTGen: Real-Time Generative Detection Transformer
di: Ruan, Chi, et al.
Pubblicazione: (2025)
di: Ruan, Chi, et al.
Pubblicazione: (2025)
SWE-Mirror: Scaling Issue-Resolving Datasets by Mirroring Issues Across Repositories
di: Wang, Junhao, et al.
Pubblicazione: (2025)
di: Wang, Junhao, et al.
Pubblicazione: (2025)
Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving
di: Zan, Daoguang, et al.
Pubblicazione: (2025)
di: Zan, Daoguang, et al.
Pubblicazione: (2025)
On the Creation of Representative Samples of Software Repositories
di: Gorostidi, June, et al.
Pubblicazione: (2024)
di: Gorostidi, June, et al.
Pubblicazione: (2024)
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
di: Wang, Yubo, et al.
Pubblicazione: (2024)
di: Wang, Yubo, et al.
Pubblicazione: (2024)
ABC: Achieving Better Control of Multimodal Embeddings using VLMs
di: Schneider, Benjamin, et al.
Pubblicazione: (2025)
di: Schneider, Benjamin, et al.
Pubblicazione: (2025)
OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis
di: Li, Zhuofeng, et al.
Pubblicazione: (2026)
di: Li, Zhuofeng, et al.
Pubblicazione: (2026)
SWE-fficiency: Can Language Models Optimize Real-World Repositories on Real Workloads?
di: Ma, Jeffrey Jian, et al.
Pubblicazione: (2025)
di: Ma, Jeffrey Jian, et al.
Pubblicazione: (2025)
SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories?
di: He, Xinyi, et al.
Pubblicazione: (2025)
di: He, Xinyi, et al.
Pubblicazione: (2025)
A comparison between the Avila-Gouëzel-Yoccoz norm and the Teichmüller norm
di: Su, Weixu, et al.
Pubblicazione: (2022)
di: Su, Weixu, et al.
Pubblicazione: (2022)
Volume of unit balls associated to quadratic differentials
di: Su, Weixu, et al.
Pubblicazione: (2025)
di: Su, Weixu, et al.
Pubblicazione: (2025)
CogDoc: Towards Unified thinking in Documents
di: Xu, Qixin, et al.
Pubblicazione: (2025)
di: Xu, Qixin, et al.
Pubblicazione: (2025)
Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth
di: Wu, Yuhuan, et al.
Pubblicazione: (2026)
di: Wu, Yuhuan, et al.
Pubblicazione: (2026)
Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
di: Wang, Haozhe, et al.
Pubblicazione: (2025)
di: Wang, Haozhe, et al.
Pubblicazione: (2025)
CGU-ILALab at FoodBench-QA 2026: Comparing Traditional and LLM-based Approaches for Recipe Nutrient Estimation
di: Chen, Wei-Chun, et al.
Pubblicazione: (2026)
di: Chen, Wei-Chun, et al.
Pubblicazione: (2026)
BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions
di: Yu, Tao, et al.
Pubblicazione: (2025)
di: Yu, Tao, et al.
Pubblicazione: (2025)
StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs
di: Yang, Jialin, et al.
Pubblicazione: (2025)
di: Yang, Jialin, et al.
Pubblicazione: (2025)
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
di: Deng, Xiang, et al.
Pubblicazione: (2025)
di: Deng, Xiang, et al.
Pubblicazione: (2025)
ORACLE-SWE: Quantifying the Contribution of Oracle Information Signals on SWE Agents
di: Li, Kenan, et al.
Pubblicazione: (2026)
di: Li, Kenan, et al.
Pubblicazione: (2026)
Critique-Coder: Enhancing Coder Models by Critique Reinforcement Learning
di: Ruan, Chi, et al.
Pubblicazione: (2025)
di: Ruan, Chi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
SWE-Next: Scalable Real-World Software Engineering Tasks for Agents
di: Liang, Jiarong, et al.
Pubblicazione: (2026) -
VisCoder2: Building Multi-Language Visualization Coding Agents
di: Ni, Yuansheng, et al.
Pubblicazione: (2025) -
VisCoder: Fine-Tuning LLMs for Executable Python Visualization Code Generation
di: Ni, Yuansheng, et al.
Pubblicazione: (2025) -
SWE-QA: Can Language Models Answer Repository-level Code Questions?
di: Peng, Weihan, et al.
Pubblicazione: (2025) -
PixelWorld: How Far Are We from Perceiving Everything as Pixels?
di: Lyu, Zhiheng, et al.
Pubblicazione: (2025)