OLAF: Towards Robust LLM-Based Annotation Framework in Empirical Software Engineering
Fuente:
arXiv
Saved in:
| Main Authors: | Imran, Mia Mohammad, Zaman, Tarannum Shaila |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SECite: Analyzing and Summarizing Citations in Software Engineering Literature
by: Pyreddy, Shireesh Reddy, et al.
Published: (2026)
by: Pyreddy, Shireesh Reddy, et al.
Published: (2026)
LLPut: Investigating Large Language Models for Bug Report-Based Input Generation
by: Hasan, Alif Al, et al.
Published: (2025)
by: Hasan, Alif Al, et al.
Published: (2025)
A-ProS: Towards Reliable Autonomous Programming Through Multi-Model Feedback
by: Tabassum, Anika, et al.
Published: (2026)
by: Tabassum, Anika, et al.
Published: (2026)
DePro: Understanding the Role of LLMs in Debugging Competitive Programming Code
by: Parvez, Nabiha, et al.
Published: (2026)
by: Parvez, Nabiha, et al.
Published: (2026)
Gender Dynamics in Software Engineering: Insights from Research on Concurrency Bug Reproduction
by: Zaman, Tarannum Shaila, et al.
Published: (2025)
by: Zaman, Tarannum Shaila, et al.
Published: (2025)
Reflections on the Reproducibility of Commercial LLM Performance in Empirical Software Engineering Studies
by: Angermeir, Florian, et al.
Published: (2025)
by: Angermeir, Florian, et al.
Published: (2025)
AutoEmpirical: LLM-Based Automated Research for Empirical Software Fault Analysis
by: Yu, Jiongchi, et al.
Published: (2025)
by: Yu, Jiongchi, et al.
Published: (2025)
Where Do AI Coding Agents Fail? An Empirical Study of Failed Agentic Pull Requests in GitHub
by: Ehsani, Ramtin, et al.
Published: (2026)
by: Ehsani, Ramtin, et al.
Published: (2026)
Emotion Classification In Software Engineering Texts: A Comparative Analysis of Pre-trained Transformers Language Models
by: Imran, Mia Mohammad
Published: (2024)
by: Imran, Mia Mohammad
Published: (2024)
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
by: Wang, Ruiqi, et al.
Published: (2025)
by: Wang, Ruiqi, et al.
Published: (2025)
ALMAS: an Autonomous LLM-based Multi-Agent Software Engineering Framework
by: Tawosi, Vali, et al.
Published: (2025)
by: Tawosi, Vali, et al.
Published: (2025)
Transforming Software Development: Evaluating the Efficiency and Challenges of GitHub Copilot in Real-World Projects
by: Pandey, Ruchika, et al.
Published: (2024)
by: Pandey, Ruchika, et al.
Published: (2024)
LLM-Based Agentic Systems for Software Engineering: Challenges and Opportunities
by: Tang, Yongjian, et al.
Published: (2026)
by: Tang, Yongjian, et al.
Published: (2026)
LLM-Based Robustness Testing of Microservice Applications: An Empirical Study
by: Tigulla, Hrushitha Goud, et al.
Published: (2026)
by: Tigulla, Hrushitha Goud, et al.
Published: (2026)
Generative AI and Empirical Software Engineering: A Paradigm Shift
by: Treude, Christoph, et al.
Published: (2025)
by: Treude, Christoph, et al.
Published: (2025)
Can GPT-4 Replicate Empirical Software Engineering Research?
by: Liang, Jenny T., et al.
Published: (2023)
by: Liang, Jenny T., et al.
Published: (2023)
Evaluating LLM-Based Test Generation Under Software Evolution
by: Haroon, Sabaat, et al.
Published: (2026)
by: Haroon, Sabaat, et al.
Published: (2026)
Designing Empirical Studies on LLM-Based Code Generation: Towards a Reference Framework
by: Nascimento, Nathalia, et al.
Published: (2025)
by: Nascimento, Nathalia, et al.
Published: (2025)
Empirical Assessment of the Perception of Software Product Line Engineering by an SME before Migrating its Code Base
by: Georges, Thomas, et al.
Published: (2025)
by: Georges, Thomas, et al.
Published: (2025)
AgentForge: Execution-Grounded Multi-Agent LLM Framework for Autonomous Software Engineering
by: Kumar, Rajesh, et al.
Published: (2026)
by: Kumar, Rajesh, et al.
Published: (2026)
Rethinking the Value of Agent-Generated Tests for LLM-Based Software Engineering Agents
by: Chen, Zhi, et al.
Published: (2026)
by: Chen, Zhi, et al.
Published: (2026)
Towards Specification-Driven LLM-Based Generation of Embedded Automotive Software
by: Patil, Minal Suresh, et al.
Published: (2024)
by: Patil, Minal Suresh, et al.
Published: (2024)
Towards Comprehensive Benchmarking Infrastructure for LLMs In Software Engineering
by: Rodriguez-Cardenas, Daniel, et al.
Published: (2026)
by: Rodriguez-Cardenas, Daniel, et al.
Published: (2026)
The Rise of Agentic Testing: Multi-Agent Systems for Robust Software Quality Assurance
by: Naqvi, Saba, et al.
Published: (2026)
by: Naqvi, Saba, et al.
Published: (2026)
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
by: Fandina, Ora Nova, et al.
Published: (2025)
by: Fandina, Ora Nova, et al.
Published: (2025)
Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering
by: Zhao, Zixiao, et al.
Published: (2026)
by: Zhao, Zixiao, et al.
Published: (2026)
Unified Software Engineering Agent as AI Software Engineer
by: Applis, Leonhard, et al.
Published: (2025)
by: Applis, Leonhard, et al.
Published: (2025)
Beyond Accuracy: LLM Variability in Evidence Screening for Software Engineering SLRs
by: Hida, Gilberto Sussumu, et al.
Published: (2026)
by: Hida, Gilberto Sussumu, et al.
Published: (2026)
Software Reuse in the Generative AI Era: From Cargo Cult Towards AI Native Software Engineering
by: Mikkonen, Tommi, et al.
Published: (2025)
by: Mikkonen, Tommi, et al.
Published: (2025)
Towards a Science of Causal Interpretability in Deep Learning for Software Engineering
by: Palacio, David N.
Published: (2025)
by: Palacio, David N.
Published: (2025)
Shift-Up: A Framework for Software Engineering Guardrails in AI-native Software Development -- Initial Findings
by: Lipsanen, Petrus, et al.
Published: (2026)
by: Lipsanen, Petrus, et al.
Published: (2026)
Trae Agent: An LLM-based Agent for Software Engineering with Test-time Scaling
by: Trae Research Team, et al.
Published: (2025)
by: Trae Research Team, et al.
Published: (2025)
OrcaLoca: An LLM Agent Framework for Software Issue Localization
by: Yu, Zhongming, et al.
Published: (2025)
by: Yu, Zhongming, et al.
Published: (2025)
When Prompt Engineering Meets Software Engineering: CNL-P as Natural and Robust "APIs'' for Human-AI Interaction
by: Xing, Zhenchang, et al.
Published: (2025)
by: Xing, Zhenchang, et al.
Published: (2025)
Towards Structured, State-Aware, and Execution-Grounded Reasoning for Software Engineering Agents
by: Tse-Hsun, et al.
Published: (2026)
by: Tse-Hsun, et al.
Published: (2026)
Towards Green AI: Decoding the Energy of LLM Inference in Software Development
by: Solovyeva, Lola, et al.
Published: (2026)
by: Solovyeva, Lola, et al.
Published: (2026)
What Software Engineering Looks Like to AI Agents? -- An Empirical Study of AI-Only Technical Discourse on MoltBook
by: Huo, Junyu, et al.
Published: (2026)
by: Huo, Junyu, et al.
Published: (2026)
LLMs: A Game-Changer for Software Engineers?
by: Haque, Md Asraful
Published: (2024)
by: Haque, Md Asraful
Published: (2024)
Software Performance Engineering for Foundation Model-Powered Software
by: Zhang, Haoxiang, et al.
Published: (2024)
by: Zhang, Haoxiang, et al.
Published: (2024)
Using LLMs in Software Requirements Specifications: An Empirical Evaluation
by: Krishna, Madhava, et al.
Published: (2024)
by: Krishna, Madhava, et al.
Published: (2024)
Similar Items
-
SECite: Analyzing and Summarizing Citations in Software Engineering Literature
by: Pyreddy, Shireesh Reddy, et al.
Published: (2026) -
LLPut: Investigating Large Language Models for Bug Report-Based Input Generation
by: Hasan, Alif Al, et al.
Published: (2025) -
A-ProS: Towards Reliable Autonomous Programming Through Multi-Model Feedback
by: Tabassum, Anika, et al.
Published: (2026) -
DePro: Understanding the Role of LLMs in Debugging Competitive Programming Code
by: Parvez, Nabiha, et al.
Published: (2026) -
Gender Dynamics in Software Engineering: Insights from Research on Concurrency Bug Reproduction
by: Zaman, Tarannum Shaila, et al.
Published: (2025)