SWE-chat: Coding Agent Interactions From Real Users in the Wild
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Baumann, Joachim, Padmakumar, Vishakh, Li, Xiang, Yang, John, Yang, Diyi, Koyejo, Sanmi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SWE-smith: Scaling Data for Software Engineering Agents
von: Yang, John, et al.
Veröffentlicht: (2025)
von: Yang, John, et al.
Veröffentlicht: (2025)
SparkMe: Adaptive Semi-Structured Interviewing for Qualitative Insight Discovery
von: Anugraha, David, et al.
Veröffentlicht: (2026)
von: Anugraha, David, et al.
Veröffentlicht: (2026)
TOM-SWE: User Mental Modeling For Software Engineering Agents
von: Zhou, Xuhui, et al.
Veröffentlicht: (2025)
von: Zhou, Xuhui, et al.
Veröffentlicht: (2025)
Offloading Score: Measuring AI Reliance Through Counterfactual Workflows
von: Padmakumar, Vishakh, et al.
Veröffentlicht: (2026)
von: Padmakumar, Vishakh, et al.
Veröffentlicht: (2026)
SWE-Next: Scalable Real-World Software Engineering Tasks for Agents
von: Liang, Jiarong, et al.
Veröffentlicht: (2026)
von: Liang, Jiarong, et al.
Veröffentlicht: (2026)
SWE Context Bench: A Benchmark for Context Learning in Coding
von: Zhu, Jiayuan, et al.
Veröffentlicht: (2026)
von: Zhu, Jiayuan, et al.
Veröffentlicht: (2026)
SWE-Universe: Scale Real-World Verifiable Environments to Millions
von: Chen, Mouxiang, et al.
Veröffentlicht: (2026)
von: Chen, Mouxiang, et al.
Veröffentlicht: (2026)
SWE-Bench-CL: Continual Learning for Coding Agents
von: Joshi, Thomas, et al.
Veröffentlicht: (2025)
von: Joshi, Thomas, et al.
Veröffentlicht: (2025)
TAI3: Testing Agent Integrity in Interpreting User Intent
von: Feng, Shiwei, et al.
Veröffentlicht: (2025)
von: Feng, Shiwei, et al.
Veröffentlicht: (2025)
AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation
von: Sahoo, Priyam, et al.
Veröffentlicht: (2026)
von: Sahoo, Priyam, et al.
Veröffentlicht: (2026)
SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering?
von: Han, Tingxu, et al.
Veröffentlicht: (2026)
von: Han, Tingxu, et al.
Veröffentlicht: (2026)
Resolving Java Code Repository Issues with iSWE Agent
von: Ganhotra, Jatin, et al.
Veröffentlicht: (2026)
von: Ganhotra, Jatin, et al.
Veröffentlicht: (2026)
SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
von: Yang, John, et al.
Veröffentlicht: (2024)
von: Yang, John, et al.
Veröffentlicht: (2024)
SWE-QA: A Dataset and Benchmark for Complex Code Understanding
von: Elkoussy, Laïla, et al.
Veröffentlicht: (2026)
von: Elkoussy, Laïla, et al.
Veröffentlicht: (2026)
When Agents go Astray: Course-Correcting SWE Agents with PRMs
von: Gandhi, Shubham, et al.
Veröffentlicht: (2025)
von: Gandhi, Shubham, et al.
Veröffentlicht: (2025)
CodeClash: Benchmarking Goal-Oriented Software Engineering
von: Yang, John, et al.
Veröffentlicht: (2025)
von: Yang, John, et al.
Veröffentlicht: (2025)
SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities for Large Language Models
von: Xu, Jingxuan, et al.
Veröffentlicht: (2025)
von: Xu, Jingxuan, et al.
Veröffentlicht: (2025)
SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades
von: Lam, Man Ho, et al.
Veröffentlicht: (2026)
von: Lam, Man Ho, et al.
Veröffentlicht: (2026)
Kimi-Dev: Agentless Training as Skill Prior for SWE-Agents
von: Yang, Zonghan, et al.
Veröffentlicht: (2025)
von: Yang, Zonghan, et al.
Veröffentlicht: (2025)
An Approach to Detect Abnormal Submissions for CodeWorkout Dataset
von: Hicks, Alex, et al.
Veröffentlicht: (2024)
von: Hicks, Alex, et al.
Veröffentlicht: (2024)
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
von: Jimenez, Carlos E., et al.
Veröffentlicht: (2023)
von: Jimenez, Carlos E., et al.
Veröffentlicht: (2023)
SWE-fficiency: Can Language Models Optimize Real-World Repositories on Real Workloads?
von: Ma, Jeffrey Jian, et al.
Veröffentlicht: (2025)
von: Ma, Jeffrey Jian, et al.
Veröffentlicht: (2025)
SWE-PRBench: Benchmarking AI Code Review Quality Against Pull Request Feedback
von: Kumar, Deepak
Veröffentlicht: (2026)
von: Kumar, Deepak
Veröffentlicht: (2026)
Saving SWE-Bench: A Benchmark Mutation Approach for Realistic Agent Evaluation
von: Garg, Spandan, et al.
Veröffentlicht: (2025)
von: Garg, Spandan, et al.
Veröffentlicht: (2025)
Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents
von: Huang, Jiawei, et al.
Veröffentlicht: (2026)
von: Huang, Jiawei, et al.
Veröffentlicht: (2026)
SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks
von: Guo, Lianghong, et al.
Veröffentlicht: (2025)
von: Guo, Lianghong, et al.
Veröffentlicht: (2025)
SWE-Effi: Re-Evaluating Software AI Agent System Effectiveness Under Resource Constraints
von: Fan, Zhiyu, et al.
Veröffentlicht: (2025)
von: Fan, Zhiyu, et al.
Veröffentlicht: (2025)
From User Interface to Agent Interface: Efficiency Optimization of UI Representations for LLM Agents
von: Ran, Dezhi, et al.
Veröffentlicht: (2025)
von: Ran, Dezhi, et al.
Veröffentlicht: (2025)
Code-Driven Law NO, Normware SI!
von: Sileno, Giovanni
Veröffentlicht: (2024)
von: Sileno, Giovanni
Veröffentlicht: (2024)
From Mirage to Grounding: Towards Reliable Multimodal Circuit-to-Verilog Code Generation
von: Yang, Guang, et al.
Veröffentlicht: (2026)
von: Yang, Guang, et al.
Veröffentlicht: (2026)
Automated Program Repair of Uncompilable Student Code
von: Pitts, Griffin, et al.
Veröffentlicht: (2025)
von: Pitts, Griffin, et al.
Veröffentlicht: (2025)
An Empirical Study of Proactive Coding Assistants in Real-World Software Development
von: Li, Lehui, et al.
Veröffentlicht: (2026)
von: Li, Lehui, et al.
Veröffentlicht: (2026)
Can LLMs Identify Gaps and Misconceptions in Students' Code Explanations?
von: Oli, Priti, et al.
Veröffentlicht: (2024)
von: Oli, Priti, et al.
Veröffentlicht: (2024)
$\texttt{Droid}$: A Resource Suite for AI-Generated Code Detection
von: Orel, Daniil, et al.
Veröffentlicht: (2025)
von: Orel, Daniil, et al.
Veröffentlicht: (2025)
Do Generative AI Tools Ensure Green Code? An Investigative Study
von: Sikand, Samarth, et al.
Veröffentlicht: (2025)
von: Sikand, Samarth, et al.
Veröffentlicht: (2025)
APEX-SWE
von: Kottamasu, Abhi, et al.
Veröffentlicht: (2026)
von: Kottamasu, Abhi, et al.
Veröffentlicht: (2026)
SWE-Fuse: Empowering Software Agents via Issue-free Trajectory Learning and Entropy-aware RLVR Training
von: Wen, Xin-Cheng, et al.
Veröffentlicht: (2026)
von: Wen, Xin-Cheng, et al.
Veröffentlicht: (2026)
SWE-Synth: Synthesizing Verifiable Bug-Fix Data to Enable Large Language Models in Resolving Real-World Bugs
von: Pham, Minh V. T., et al.
Veröffentlicht: (2025)
von: Pham, Minh V. T., et al.
Veröffentlicht: (2025)
ParaStudent: Generating and Evaluating Realistic Student Code by Teaching LLMs to Struggle
von: Miroyan, Mihran, et al.
Veröffentlicht: (2025)
von: Miroyan, Mihran, et al.
Veröffentlicht: (2025)
Using AI-Based Coding Assistants in Practice: State of Affairs, Perceptions, and Ways Forward
von: Sergeyuk, Agnia, et al.
Veröffentlicht: (2024)
von: Sergeyuk, Agnia, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SWE-smith: Scaling Data for Software Engineering Agents
von: Yang, John, et al.
Veröffentlicht: (2025) -
SparkMe: Adaptive Semi-Structured Interviewing for Qualitative Insight Discovery
von: Anugraha, David, et al.
Veröffentlicht: (2026) -
TOM-SWE: User Mental Modeling For Software Engineering Agents
von: Zhou, Xuhui, et al.
Veröffentlicht: (2025) -
Offloading Score: Measuring AI Reliance Through Counterfactual Workflows
von: Padmakumar, Vishakh, et al.
Veröffentlicht: (2026) -
SWE-Next: Scalable Real-World Software Engineering Tasks for Agents
von: Liang, Jiarong, et al.
Veröffentlicht: (2026)