Beyond Accuracy: LLM Variability in Evidence Screening for Software Engineering SLRs
Fuente:
arXiv
Saved in:
| Main Authors: | Hida, Gilberto Sussumu, Ribeiro, Danilo Monteiro, Yahata, Erika |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding the relationships between the perceptions of burnout and instability in Software Engineering
by: Ribeiro, Danilo Monteiro
Published: (2025)
by: Ribeiro, Danilo Monteiro
Published: (2025)
Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges
by: Weng, Shihao, et al.
Published: (2026)
by: Weng, Shihao, et al.
Published: (2026)
LLM-Based Agentic Systems for Software Engineering: Challenges and Opportunities
by: Tang, Yongjian, et al.
Published: (2026)
by: Tang, Yongjian, et al.
Published: (2026)
Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering
by: Zhao, Zixiao, et al.
Published: (2026)
by: Zhao, Zixiao, et al.
Published: (2026)
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
by: Fandina, Ora Nova, et al.
Published: (2025)
by: Fandina, Ora Nova, et al.
Published: (2025)
Beyond Human-Readable: Rethinking Software Engineering Conventions for the Agentic Development Era
by: Ustynov, Dmytro
Published: (2026)
by: Ustynov, Dmytro
Published: (2026)
Unified Software Engineering Agent as AI Software Engineer
by: Applis, Leonhard, et al.
Published: (2025)
by: Applis, Leonhard, et al.
Published: (2025)
Evidence-Driven Decision Support for AI Model Selection in Research Software Engineering
by: Joonbakhsh, Alireza, et al.
Published: (2025)
by: Joonbakhsh, Alireza, et al.
Published: (2025)
It's Not About Whom You Train: An Analysis of Corporate Education in Software Engineering
by: Siqueira, Rodrigo, et al.
Published: (2026)
by: Siqueira, Rodrigo, et al.
Published: (2026)
ALMAS: an Autonomous LLM-based Multi-Agent Software Engineering Framework
by: Tawosi, Vali, et al.
Published: (2025)
by: Tawosi, Vali, et al.
Published: (2025)
Reflections on the Reproducibility of Commercial LLM Performance in Empirical Software Engineering Studies
by: Angermeir, Florian, et al.
Published: (2025)
by: Angermeir, Florian, et al.
Published: (2025)
The professional's opinion: Suggestions for improving the corporate education training process in Software Engineering
by: Siqueira, Rodrigo, et al.
Published: (2026)
by: Siqueira, Rodrigo, et al.
Published: (2026)
Rethinking the Value of Agent-Generated Tests for LLM-Based Software Engineering Agents
by: Chen, Zhi, et al.
Published: (2026)
by: Chen, Zhi, et al.
Published: (2026)
Trae Agent: An LLM-based Agent for Software Engineering with Test-time Scaling
by: Trae Research Team, et al.
Published: (2025)
by: Trae Research Team, et al.
Published: (2025)
OLAF: Towards Robust LLM-Based Annotation Framework in Empirical Software Engineering
by: Imran, Mia Mohammad, et al.
Published: (2025)
by: Imran, Mia Mohammad, et al.
Published: (2025)
AgentForge: Execution-Grounded Multi-Agent LLM Framework for Autonomous Software Engineering
by: Kumar, Rajesh, et al.
Published: (2026)
by: Kumar, Rajesh, et al.
Published: (2026)
LLMs: A Game-Changer for Software Engineers?
by: Haque, Md Asraful
Published: (2024)
by: Haque, Md Asraful
Published: (2024)
Software Performance Engineering for Foundation Model-Powered Software
by: Zhang, Haoxiang, et al.
Published: (2024)
by: Zhang, Haoxiang, et al.
Published: (2024)
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
by: Wang, Ruiqi, et al.
Published: (2025)
by: Wang, Ruiqi, et al.
Published: (2025)
LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering
by: Qiu, Jielin, et al.
Published: (2025)
by: Qiu, Jielin, et al.
Published: (2025)
LLM-based Zero-shot Triple Extraction for Automated Ontology Generation from Software Engineering Standards
by: Yue, Songhui
Published: (2025)
by: Yue, Songhui
Published: (2025)
AI-Tutoring in Software Engineering Education
by: Frankford, Eduard, et al.
Published: (2024)
by: Frankford, Eduard, et al.
Published: (2024)
Bridging Forecast Accuracy and Inventory KPIs: A Simulation-Based Software Framework
by: Fukuhara, So, et al.
Published: (2026)
by: Fukuhara, So, et al.
Published: (2026)
Designing LLM-based Multi-Agent Systems for Software Engineering Tasks: Quality Attributes, Design Patterns and Rationale
by: Cai, Yangxiao, et al.
Published: (2025)
by: Cai, Yangxiao, et al.
Published: (2025)
Causal Software Engineering: A Vision and Roadmap
by: Pietrantuono, Roberto, et al.
Published: (2026)
by: Pietrantuono, Roberto, et al.
Published: (2026)
Beyond Accuracy: Characterizing Code Comprehension Capabilities in (Large) Language Models
by: Mächtle, Felix, et al.
Published: (2026)
by: Mächtle, Felix, et al.
Published: (2026)
SEER: Sustainability Enhanced Engineering of Software Requirements
by: Roy, Mandira, et al.
Published: (2025)
by: Roy, Mandira, et al.
Published: (2025)
Agentic AI Software Engineers: Programming with Trust
by: Roychoudhury, Abhik, et al.
Published: (2025)
by: Roychoudhury, Abhik, et al.
Published: (2025)
Artificial Intelligence as a Catalyst for Innovation in Software Engineering
by: Fernández-y-Fernández, Carlos Alberto, et al.
Published: (2026)
by: Fernández-y-Fernández, Carlos Alberto, et al.
Published: (2026)
Towards Comprehensive Benchmarking Infrastructure for LLMs In Software Engineering
by: Rodriguez-Cardenas, Daniel, et al.
Published: (2026)
by: Rodriguez-Cardenas, Daniel, et al.
Published: (2026)
SECite: Analyzing and Summarizing Citations in Software Engineering Literature
by: Pyreddy, Shireesh Reddy, et al.
Published: (2026)
by: Pyreddy, Shireesh Reddy, et al.
Published: (2026)
Feedback Loops and Code Perturbations in LLM-based Software Engineering: A Case Study on a C-to-Rust Translation System
by: Weiss, Martin, et al.
Published: (2025)
by: Weiss, Martin, et al.
Published: (2025)
CodeClash: Benchmarking Goal-Oriented Software Engineering
by: Yang, John, et al.
Published: (2025)
by: Yang, John, et al.
Published: (2025)
In-Context Code-Text Learning for Bimodal Software Engineering
by: Tang, Xunzhu, et al.
Published: (2024)
by: Tang, Xunzhu, et al.
Published: (2024)
Beyond the 'Diff': Addressing Agentic Entropy in Agentic Software Development
by: Casserini, Matteo, et al.
Published: (2026)
by: Casserini, Matteo, et al.
Published: (2026)
Toward Agentic Software Engineering Beyond Code: Framing Vision, Values, and Vocabulary
by: Hoda, Rashina
Published: (2025)
by: Hoda, Rashina
Published: (2025)
Shift-Up: A Framework for Software Engineering Guardrails in AI-native Software Development -- Initial Findings
by: Lipsanen, Petrus, et al.
Published: (2026)
by: Lipsanen, Petrus, et al.
Published: (2026)
Software Reuse in the Generative AI Era: From Cargo Cult Towards AI Native Software Engineering
by: Mikkonen, Tommi, et al.
Published: (2025)
by: Mikkonen, Tommi, et al.
Published: (2025)
Beyond Accuracy: An Empirical Study on Unit Testing in Open-source Deep Learning Projects
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
Reproducible, Explainable, and Effective Evaluations of Agentic AI for Software Engineering
by: Li, Jingyue, et al.
Published: (2026)
by: Li, Jingyue, et al.
Published: (2026)
Similar Items
-
Understanding the relationships between the perceptions of burnout and instability in Software Engineering
by: Ribeiro, Danilo Monteiro
Published: (2025) -
Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges
by: Weng, Shihao, et al.
Published: (2026) -
LLM-Based Agentic Systems for Software Engineering: Challenges and Opportunities
by: Tang, Yongjian, et al.
Published: (2026) -
Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering
by: Zhao, Zixiao, et al.
Published: (2026) -
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
by: Fandina, Ora Nova, et al.
Published: (2025)