SESR-Eval: Dataset for Evaluating LLMs in the Title-Abstract Screening of Systematic Reviews
Fuente:
arXiv
Saved in:
| Main Authors: | Huotala, Aleksi, Kuutila, Miikka, Mäntylä, Mika |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Research Artifacts in Secondary Studies: A Systematic Mapping in Software Engineering
by: Huotala, Aleksi, et al.
Published: (2025)
by: Huotala, Aleksi, et al.
Published: (2025)
AISysRev -- LLM-based Tool for Title-abstract Screening
by: Huotala, Aleksi, et al.
Published: (2025)
by: Huotala, Aleksi, et al.
Published: (2025)
The Promise and Challenges of Using LLMs to Accelerate the Screening Process of Systematic Reviews
by: Huotala, Aleksi, et al.
Published: (2024)
by: Huotala, Aleksi, et al.
Published: (2024)
What Makes Programmers Laugh? Exploring the Submissions of the Subreddit r/ProgrammerHumor
by: Kuutila, Miikka, et al.
Published: (2024)
by: Kuutila, Miikka, et al.
Published: (2024)
Individual Differences Limit Predicting Well-being and Productivity Using Software Repositories: A Longitudinal Industrial Study
by: Kuutila, Miikka, et al.
Published: (2021)
by: Kuutila, Miikka, et al.
Published: (2021)
Teaching Software Metrology: The Science of Measurement for Software Engineering
by: Ralph, Paul, et al.
Published: (2024)
by: Ralph, Paul, et al.
Published: (2024)
User Personas Improve Social Sustainability by Encouraging Software Developers to Deprioritize Antisocial Features
by: Ayoola, Bimpe, et al.
Published: (2024)
by: Ayoola, Bimpe, et al.
Published: (2024)
Token Interdependency Parsing (Tipping) -- Fast and Accurate Log Parsing
by: Hashemi, Shayan, et al.
Published: (2024)
by: Hashemi, Shayan, et al.
Published: (2024)
Speed and Performance of Parserless and Unsupervised Anomaly Detection Methods on Software Logs
by: Nyyssölä, Jesse, et al.
Published: (2023)
by: Nyyssölä, Jesse, et al.
Published: (2023)
Detecting Anomalies in Software Execution Logs with Siamese Network
by: Hashemi, Shayan, et al.
Published: (2021)
by: Hashemi, Shayan, et al.
Published: (2021)
OneLog: Towards End-to-End Training in Software Log Anomaly Detection
by: Hashemi, Shayan, et al.
Published: (2021)
by: Hashemi, Shayan, et al.
Published: (2021)
Assessing REST API Test Generation Strategies with Log Coverage
by: Reinikainen, Nana, et al.
Published: (2026)
by: Reinikainen, Nana, et al.
Published: (2026)
Multi-Agent Systems for Root Cause Analysis in Microservices
by: Naakka, Alexander, et al.
Published: (2026)
by: Naakka, Alexander, et al.
Published: (2026)
AnoMod: A Dataset for Anomaly Detection and Root Cause Analysis in Microservice Systems
by: Ping, Ke, et al.
Published: (2026)
by: Ping, Ke, et al.
Published: (2026)
LogLead -- Fast and Integrated Log Loader, Enhancer, and Anomaly Detector
by: Mäntylä, Mika, et al.
Published: (2023)
by: Mäntylä, Mika, et al.
Published: (2023)
Cross-System Software Log-based Anomaly Detection Using Meta-Learning
by: Wang, Yuqing, et al.
Published: (2024)
by: Wang, Yuqing, et al.
Published: (2024)
A Comparative Study of Semantic Log Representations for Software Log-based Anomaly Detection
by: Wang, Yuqing, et al.
Published: (2026)
by: Wang, Yuqing, et al.
Published: (2026)
Show Your Title! A Scoping Review on Verbalization in Software Engineering with LLM-Assisted Screening
by: Balogh, Gergő, et al.
Published: (2025)
by: Balogh, Gergő, et al.
Published: (2025)
SR-Eval: Evaluating LLMs on Code Generation under Stepwise Requirement Refinement
by: Zhan, Zexun, et al.
Published: (2025)
by: Zhan, Zexun, et al.
Published: (2025)
ScratchEval : A Multimodal Evaluation Framework for LLMs in Block-Based Programming
by: Si, Yuan, et al.
Published: (2026)
by: Si, Yuan, et al.
Published: (2026)
Detection, Classification and Prevalence of Self-Admitted Aging Debt
by: Sridharan, Murali, et al.
Published: (2025)
by: Sridharan, Murali, et al.
Published: (2025)
TAM-Eval: Evaluating LLMs for Automated Unit Test Maintenance
by: Bruches, Elena, et al.
Published: (2026)
by: Bruches, Elena, et al.
Published: (2026)
Software Architecture Meets LLMs: A Systematic Literature Review
by: Schmid, Larissa, et al.
Published: (2025)
by: Schmid, Larissa, et al.
Published: (2025)
Hidden in Plain Sight: Where Developers Confess Self-Admitted Technical Debt
by: Sridharan, Murali, et al.
Published: (2025)
by: Sridharan, Murali, et al.
Published: (2025)
Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure
by: Yang, Zheyuan, et al.
Published: (2025)
by: Yang, Zheyuan, et al.
Published: (2025)
Cross-System Categorization of Abnormal Traces in Microservice-Based Systems via Meta-Learning
by: Wang, Yuqing, et al.
Published: (2024)
by: Wang, Yuqing, et al.
Published: (2024)
LO2: Microservice API Anomaly Dataset of Logs and Metrics
by: Bakhtin, Alexander, et al.
Published: (2025)
by: Bakhtin, Alexander, et al.
Published: (2025)
SecVulEval: Benchmarking LLMs for Real-World C/C++ Vulnerability Detection
by: Ahmed, Md Basim Uddin, et al.
Published: (2025)
by: Ahmed, Md Basim Uddin, et al.
Published: (2025)
HumanEvalComm: Benchmarking the Communication Competence of Code Generation for LLMs and LLM Agent
by: Wu, Jie JW, et al.
Published: (2024)
by: Wu, Jie JW, et al.
Published: (2024)
Staying or Leaving? How Job Satisfaction, Embeddedness and Antecedents Predict Turnover Intentions of Software Professionals
by: Kuutila, Miikka, et al.
Published: (2025)
by: Kuutila, Miikka, et al.
Published: (2025)
Can LLMs Recover Program Semantics? A Systematic Evaluation with Symbolic Execution
by: Feng, Rong, et al.
Published: (2025)
by: Feng, Rong, et al.
Published: (2025)
ScalerEval: Automated and Consistent Evaluation Testbed for Auto-scalers in Microservices
by: Xie, Shuaiyu, et al.
Published: (2025)
by: Xie, Shuaiyu, et al.
Published: (2025)
ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS
by: Xie, Bang, et al.
Published: (2026)
by: Xie, Bang, et al.
Published: (2026)
DependEval: Benchmarking LLMs for Repository Dependency Understanding
by: Du, Junjia, et al.
Published: (2025)
by: Du, Junjia, et al.
Published: (2025)
ClassEval-T: Evaluating Large Language Models in Class-Level Code Translation
by: Xue, Pengyu, et al.
Published: (2024)
by: Xue, Pengyu, et al.
Published: (2024)
LiCoEval: Evaluating LLMs on License Compliance in Code Generation
by: Xu, Weiwei, et al.
Published: (2024)
by: Xu, Weiwei, et al.
Published: (2024)
SolContractEval: A Benchmark for Evaluating Contract-Level Solidity Code Generation
by: Ye, Zhifan, et al.
Published: (2025)
by: Ye, Zhifan, et al.
Published: (2025)
ReleaseEval: A Benchmark for Evaluating Language Models in Automated Release Note Generation
by: Meng, Qianru, et al.
Published: (2025)
by: Meng, Qianru, et al.
Published: (2025)
SpecEval: Evaluating Code Comprehension in Large Language Models via Program Specifications
by: Ma, Lezhi, et al.
Published: (2024)
by: Ma, Lezhi, et al.
Published: (2024)
Towards Leveraging LLMs to Generate Abstract Penetration Test Cases from Software Architecture
by: Jafari, Mahdi, et al.
Published: (2026)
by: Jafari, Mahdi, et al.
Published: (2026)
Similar Items
-
Research Artifacts in Secondary Studies: A Systematic Mapping in Software Engineering
by: Huotala, Aleksi, et al.
Published: (2025) -
AISysRev -- LLM-based Tool for Title-abstract Screening
by: Huotala, Aleksi, et al.
Published: (2025) -
The Promise and Challenges of Using LLMs to Accelerate the Screening Process of Systematic Reviews
by: Huotala, Aleksi, et al.
Published: (2024) -
What Makes Programmers Laugh? Exploring the Submissions of the Subreddit r/ProgrammerHumor
by: Kuutila, Miikka, et al.
Published: (2024) -
Individual Differences Limit Predicting Well-being and Productivity Using Software Repositories: A Longitudinal Industrial Study
by: Kuutila, Miikka, et al.
Published: (2021)