Automating Computational Reproducibility in Social Science: Comparing Prompt-Based and Agent-Based Approaches
Fuente:
arXiv
Salvato in:
| Autori principali: | Shah, Syed Mehtab Hussain, Hopfgartner, Frank, Bleier, Arnim |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Computational Reproducibility of R Code Supplements on OSF
di: Saju, Lorraine, et al.
Pubblicazione: (2025)
di: Saju, Lorraine, et al.
Pubblicazione: (2025)
Can Coding Agents Reproduce Findings in Computational Materials Science?
di: Huang, Ziyang, et al.
Pubblicazione: (2026)
di: Huang, Ziyang, et al.
Pubblicazione: (2026)
EffiSkill: Agent Skill Based Automated Code Efficiency Optimization
di: Wang, Zimu, et al.
Pubblicazione: (2026)
di: Wang, Zimu, et al.
Pubblicazione: (2026)
ToolScope: Enhancing LLM Agent Tool Use through Tool Merging and Context-Aware Filtering
di: Liu, Marianne Menglin, et al.
Pubblicazione: (2025)
di: Liu, Marianne Menglin, et al.
Pubblicazione: (2025)
Automated Business Process Analysis: An LLM-Based Approach to Value Assessment
di: De Michele, William, et al.
Pubblicazione: (2025)
di: De Michele, William, et al.
Pubblicazione: (2025)
Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluation
di: Pysklo, Hubert M., et al.
Pubblicazione: (2026)
di: Pysklo, Hubert M., et al.
Pubblicazione: (2026)
The Prompt Alchemist: Automated LLM-Tailored Prompt Optimization for Test Case Generation
di: Gao, Shuzheng, et al.
Pubblicazione: (2025)
di: Gao, Shuzheng, et al.
Pubblicazione: (2025)
SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents
di: Badertdinov, Ibragim, et al.
Pubblicazione: (2025)
di: Badertdinov, Ibragim, et al.
Pubblicazione: (2025)
Alibaba LingmaAgent: Improving Automated Issue Resolution via Comprehensive Repository Exploration
di: Ma, Yingwei, et al.
Pubblicazione: (2024)
di: Ma, Yingwei, et al.
Pubblicazione: (2024)
ProjectEval: A Benchmark for Programming Agents Automated Evaluation on Project-Level Code Generation
di: Liu, Kaiyuan, et al.
Pubblicazione: (2025)
di: Liu, Kaiyuan, et al.
Pubblicazione: (2025)
ResearchCodeAgent: An LLM Multi-Agent System for Automated Codification of Research Methodologies
di: Gandhi, Shubham, et al.
Pubblicazione: (2025)
di: Gandhi, Shubham, et al.
Pubblicazione: (2025)
Using Large Language Models for Student-Code Guided Test Case Generation in Computer Science Education
di: Kumar, Nischal Ashok, et al.
Pubblicazione: (2024)
di: Kumar, Nischal Ashok, et al.
Pubblicazione: (2024)
What Prompts Don't Say: Understanding and Managing Underspecification in LLM Prompts
di: Yang, Chenyang, et al.
Pubblicazione: (2025)
di: Yang, Chenyang, et al.
Pubblicazione: (2025)
Terminal Agents Suffice for Enterprise Automation
di: Bechard, Patrice, et al.
Pubblicazione: (2026)
di: Bechard, Patrice, et al.
Pubblicazione: (2026)
Securing Computer-Use Agents: A Unified Architecture-Lifecycle Framework for Deployment-Grounded Reliability
di: Chen, Zejian, et al.
Pubblicazione: (2026)
di: Chen, Zejian, et al.
Pubblicazione: (2026)
Evaluating LLM-Based Goal Extraction in Requirements Engineering: Prompting Strategies and Their Limitations
di: Arnaudo, Anna, et al.
Pubblicazione: (2026)
di: Arnaudo, Anna, et al.
Pubblicazione: (2026)
Prompting and Fine-tuning Large Language Models for Automated Code Review Comment Generation
di: Haider, Md. Asif, et al.
Pubblicazione: (2024)
di: Haider, Md. Asif, et al.
Pubblicazione: (2024)
FASTRIC: Prompt Specification Language for Verifiable LLM Interactions
di: Jin, Wen-Long
Pubblicazione: (2025)
di: Jin, Wen-Long
Pubblicazione: (2025)
On Sequence-to-Sequence Models for Automated Log Parsing
di: Sorrenti, Adam, et al.
Pubblicazione: (2026)
di: Sorrenti, Adam, et al.
Pubblicazione: (2026)
SERA: Soft-Verified Efficient Repository Agents
di: Shen, Ethan, et al.
Pubblicazione: (2026)
di: Shen, Ethan, et al.
Pubblicazione: (2026)
ProbeLLM: Automating Principled Diagnosis of LLM Failures
di: Huang, Yue, et al.
Pubblicazione: (2026)
di: Huang, Yue, et al.
Pubblicazione: (2026)
ASTRAL: Automated Safety Testing of Large Language Models
di: Ugarte, Miriam, et al.
Pubblicazione: (2025)
di: Ugarte, Miriam, et al.
Pubblicazione: (2025)
GIRT-Model: Automated Generation of Issue Report Templates
di: Nikeghbal, Nafiseh, et al.
Pubblicazione: (2024)
di: Nikeghbal, Nafiseh, et al.
Pubblicazione: (2024)
Large Language Models for IT Automation Tasks: Are We There Yet?
di: Hassan, Md Mahadi, et al.
Pubblicazione: (2025)
di: Hassan, Md Mahadi, et al.
Pubblicazione: (2025)
PACE: Improving Prompt with Actor-Critic Editing for Large Language Model
di: Dong, Yihong, et al.
Pubblicazione: (2023)
di: Dong, Yihong, et al.
Pubblicazione: (2023)
Software-Based Dialogue Systems: Survey, Taxonomy and Challenges
di: Motger, Quim, et al.
Pubblicazione: (2021)
di: Motger, Quim, et al.
Pubblicazione: (2021)
PatchRecall: Patch-Driven Retrieval for Automated Program Repair
di: Dihan, Mahir Labib, et al.
Pubblicazione: (2026)
di: Dihan, Mahir Labib, et al.
Pubblicazione: (2026)
Interpretable Online Log Analysis Using Large Language Models with Prompt Strategies
di: Liu, Yilun, et al.
Pubblicazione: (2023)
di: Liu, Yilun, et al.
Pubblicazione: (2023)
RePair: Automated Program Repair with Process-based Feedback
di: Zhao, Yuze, et al.
Pubblicazione: (2024)
di: Zhao, Yuze, et al.
Pubblicazione: (2024)
Comparing Developer and LLM Biases in Code Evaluation
di: Mittal, Aditya, et al.
Pubblicazione: (2026)
di: Mittal, Aditya, et al.
Pubblicazione: (2026)
Showing LLM-Generated Code Selectively Based on Confidence of LLMs
di: Li, Jia, et al.
Pubblicazione: (2024)
di: Li, Jia, et al.
Pubblicazione: (2024)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
di: Tu, Xinming, et al.
Pubblicazione: (2026)
di: Tu, Xinming, et al.
Pubblicazione: (2026)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
di: Fu, Lingyue, et al.
Pubblicazione: (2025)
di: Fu, Lingyue, et al.
Pubblicazione: (2025)
AgentPack: A Dataset of Code Changes, Co-Authored by Agents and Humans
di: Zi, Yangtian, et al.
Pubblicazione: (2025)
di: Zi, Yangtian, et al.
Pubblicazione: (2025)
Show and Tell: Prompt Strategies for Style Control in Multi-Turn LLM Code Generation
di: Bohr, Jeremiah
Pubblicazione: (2025)
di: Bohr, Jeremiah
Pubblicazione: (2025)
(Why) Is My Prompt Getting Worse? Rethinking Regression Testing for Evolving LLM APIs
di: Ma, Wanqin, et al.
Pubblicazione: (2023)
di: Ma, Wanqin, et al.
Pubblicazione: (2023)
Exploring the Potential of Conversational AI Support for Agent-Based Social Simulation Model Design
di: Siebers, Peer-Olaf
Pubblicazione: (2024)
di: Siebers, Peer-Olaf
Pubblicazione: (2024)
Enhanced Automated Code Vulnerability Repair using Large Language Models
di: de-Fitero-Dominguez, David, et al.
Pubblicazione: (2024)
di: de-Fitero-Dominguez, David, et al.
Pubblicazione: (2024)
RGD: Multi-LLM Based Agent Debugger via Refinement and Generation Guidance
di: Jin, Haolin, et al.
Pubblicazione: (2024)
di: Jin, Haolin, et al.
Pubblicazione: (2024)
AutoMonitor-Bench: Evaluating the Reliability of LLM-Based Misbehavior Monitor
di: Yang, Shu, et al.
Pubblicazione: (2026)
di: Yang, Shu, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Computational Reproducibility of R Code Supplements on OSF
di: Saju, Lorraine, et al.
Pubblicazione: (2025) -
Can Coding Agents Reproduce Findings in Computational Materials Science?
di: Huang, Ziyang, et al.
Pubblicazione: (2026) -
EffiSkill: Agent Skill Based Automated Code Efficiency Optimization
di: Wang, Zimu, et al.
Pubblicazione: (2026) -
ToolScope: Enhancing LLM Agent Tool Use through Tool Merging and Context-Aware Filtering
di: Liu, Marianne Menglin, et al.
Pubblicazione: (2025) -
Automated Business Process Analysis: An LLM-Based Approach to Value Assessment
di: De Michele, William, et al.
Pubblicazione: (2025)