Towards Realistic Evaluation of Commit Message Generation by Matching Online and Offline Settings
Fuente:
arXiv
Saved in:
| Main Authors: | Tsvetkov, Petr, Eliseeva, Aleksandra, Dig, Danny, Bezzubov, Alexander, Golubev, Yaroslav, Bryksin, Timofey, Zharov, Yaroslav |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On The Importance of Reasoning for Context Retrieval in Repository-Level Code Editing
by: Kovrigin, Alexander, et al.
Published: (2024)
by: Kovrigin, Alexander, et al.
Published: (2024)
EM-Assist: Safe Automated ExtractMethod Refactoring with LLMs
by: Pomian, Dorin, et al.
Published: (2024)
by: Pomian, Dorin, et al.
Published: (2024)
Untangling Knots: Leveraging LLM for Error Resolution in Computational Notebooks
by: Grotov, Konstantin, et al.
Published: (2024)
by: Grotov, Konstantin, et al.
Published: (2024)
One Step at a Time: Combining LLMs and Static Analysis to Generate Next-Step Hints for Programming Tasks
by: Birillo, Anastasiia, et al.
Published: (2024)
by: Birillo, Anastasiia, et al.
Published: (2024)
In-IDE Programming Courses: Learning Software Development in a Real-World Setting
by: Birillo, Anastasiia, et al.
Published: (2025)
by: Birillo, Anastasiia, et al.
Published: (2025)
PIPer: On-Device Environment Setup via Online Reinforcement Learning
by: Kovrigin, Alexander, et al.
Published: (2025)
by: Kovrigin, Alexander, et al.
Published: (2025)
Developer Needs and Feasible Features for AI Assistants in IDEs
by: Sergeyuk, Agnia, et al.
Published: (2024)
by: Sergeyuk, Agnia, et al.
Published: (2024)
Assessing Consensus of Developers' Views on Code Readability
by: Sergeyuk, Agnia, et al.
Published: (2024)
by: Sergeyuk, Agnia, et al.
Published: (2024)
EnvBench: A Benchmark for Automated Environment Setup
by: Eliseeva, Aleksandra, et al.
Published: (2025)
by: Eliseeva, Aleksandra, et al.
Published: (2025)
Reassessing Java Code Readability Models with a Human-Centered Approach
by: Sergeyuk, Agnia, et al.
Published: (2024)
by: Sergeyuk, Agnia, et al.
Published: (2024)
Everything You Need to Know About CS Education: Open Results from a Survey of More Than 18,000 Participants
by: Dzialets, Katsiaryna, et al.
Published: (2025)
by: Dzialets, Katsiaryna, et al.
Published: (2025)
Evolving with AI: A Longitudinal Analysis of Developer Logs
by: Sergeyuk, Agnia, et al.
Published: (2026)
by: Sergeyuk, Agnia, et al.
Published: (2026)
Unprecedented Code Change Automation: The Fusion of LLMs and Transformation by Example
by: Dilhara, Malinda, et al.
Published: (2024)
by: Dilhara, Malinda, et al.
Published: (2024)
Dynamic Retrieval-Augmented Generation
by: Shapkin, Anton, et al.
Published: (2023)
by: Shapkin, Anton, et al.
Published: (2023)
Using AI-Based Coding Assistants in Practice: State of Affairs, Perceptions, and Ways Forward
by: Sergeyuk, Agnia, et al.
Published: (2024)
by: Sergeyuk, Agnia, et al.
Published: (2024)
Long Code Arena: a Set of Benchmarks for Long-Context Code Models
by: Bogomolov, Egor, et al.
Published: (2024)
by: Bogomolov, Egor, et al.
Published: (2024)
Diff-XYZ: A Benchmark for Evaluating Diff Understanding
by: Glukhov, Evgeniy, et al.
Published: (2025)
by: Glukhov, Evgeniy, et al.
Published: (2025)
Multi-Agent Coordinated Rename Refactoring
by: Bellur, Abhiram, et al.
Published: (2026)
by: Bellur, Abhiram, et al.
Published: (2026)
Debugging Without Error Messages: How LLM Prompting Strategy Affects Programming Error Explanation Effectiveness
by: Salmon, Audrey, et al.
Published: (2025)
by: Salmon, Audrey, et al.
Published: (2025)
GitGoodBench: A Novel Benchmark For Evaluating Agentic Performance On Git
by: Lindenbauer, Tobias, et al.
Published: (2025)
by: Lindenbauer, Tobias, et al.
Published: (2025)
Together We Go Further: LLMs and IDE Static Analysis for Extract Method Refactoring
by: Pomian, Dorin, et al.
Published: (2024)
by: Pomian, Dorin, et al.
Published: (2024)
Towards LLM-Based Usability Analysis for Recommender User Interfaces
by: Lubos, Sebastian, et al.
Published: (2025)
by: Lubos, Sebastian, et al.
Published: (2025)
Qualitative Evaluation of LLM-Designed GUI
by: Sawicki, Bartosz, et al.
Published: (2026)
by: Sawicki, Bartosz, et al.
Published: (2026)
How Do Hackathons Foster Creativity? Towards AI Collaborative Evaluation of Creativity at Scale
by: Falk, Jeanette, et al.
Published: (2025)
by: Falk, Jeanette, et al.
Published: (2025)
EmoTrack: An application to Facilitate User Reflection on Their Online Behaviours
by: Zhang, Ruiyong
Published: (2026)
by: Zhang, Ruiyong
Published: (2026)
Clustering MOOC Programming Solutions to Diversify Their Presentation to Students
by: Artser, Elizaveta, et al.
Published: (2024)
by: Artser, Elizaveta, et al.
Published: (2024)
Towards the analysis of team members well-being
by: Xu, Zan, et al.
Published: (2025)
by: Xu, Zan, et al.
Published: (2025)
Towards Decoding Developer Cognition in the Age of AI Assistants
by: Haque, Ebtesam Al, et al.
Published: (2025)
by: Haque, Ebtesam Al, et al.
Published: (2025)
Towards an Understanding of Developer Experience-Driven Transparency in Software Ecosystems
by: Zacarias, Rodrigo Oliveira, et al.
Published: (2025)
by: Zacarias, Rodrigo Oliveira, et al.
Published: (2025)
Crowdsourcing: A Framework for Usability Evaluation
by: Nasir, Muhammad
Published: (2024)
by: Nasir, Muhammad
Published: (2024)
InspectorRAGet: An Introspection Platform for RAG Evaluation
by: Fadnis, Kshitij, et al.
Published: (2024)
by: Fadnis, Kshitij, et al.
Published: (2024)
Vibe Coding: Toward an AI-Native Paradigm for Semantic and Intent-Driven Programming
by: Bamil, Vinay
Published: (2025)
by: Bamil, Vinay
Published: (2025)
Towards Using Personas in Requirements Engineering: What Has Been Changed Recently?
by: Muzammel, Chowdhury Shahriar, et al.
Published: (2025)
by: Muzammel, Chowdhury Shahriar, et al.
Published: (2025)
Digital Wellbeing Redefined: Toward User-Centric Approach for Positive Social Media Engagement
by: Zhao, Yixue, et al.
Published: (2024)
by: Zhao, Yixue, et al.
Published: (2024)
Evaluating Privacy Perceptions, Experience, and Behavior of Software Development Teams
by: Prybylo, Maxwell, et al.
Published: (2024)
by: Prybylo, Maxwell, et al.
Published: (2024)
An Online A/B Testing Decision Support System for Web Usability Assessment Based on a Linguistic Decision-making Methodology: Case of Study a Virtual Learning Environment
by: Zermeño, Noe, et al.
Published: (2025)
by: Zermeño, Noe, et al.
Published: (2025)
Leveraging LLMs, IDEs, and Semantic Embeddings for Automated Move Method Refactoring
by: Bellur, Abhiram, et al.
Published: (2025)
by: Bellur, Abhiram, et al.
Published: (2025)
Commit: Online Groups with Participation Commitments
by: Popowski, Lindsay, et al.
Published: (2024)
by: Popowski, Lindsay, et al.
Published: (2024)
V-SHiNE: A Virtual Smart Home Framework for Explainability Evaluation
by: Sadeghi, Mersedeh, et al.
Published: (2026)
by: Sadeghi, Mersedeh, et al.
Published: (2026)
StartFlow: From Method Conception to Multi-Perspective Evaluation in UX Prototyping for Software Startups
by: Guerino, Guilherme Corredato, et al.
Published: (2026)
by: Guerino, Guilherme Corredato, et al.
Published: (2026)
Similar Items
-
On The Importance of Reasoning for Context Retrieval in Repository-Level Code Editing
by: Kovrigin, Alexander, et al.
Published: (2024) -
EM-Assist: Safe Automated ExtractMethod Refactoring with LLMs
by: Pomian, Dorin, et al.
Published: (2024) -
Untangling Knots: Leveraging LLM for Error Resolution in Computational Notebooks
by: Grotov, Konstantin, et al.
Published: (2024) -
One Step at a Time: Combining LLMs and Static Analysis to Generate Next-Step Hints for Programming Tasks
by: Birillo, Anastasiia, et al.
Published: (2024) -
In-IDE Programming Courses: Learning Software Development in a Real-World Setting
by: Birillo, Anastasiia, et al.
Published: (2025)