ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences
Fuente:
arXiv
Saved in:
| Main Authors: | Nguyen, Bang, Soós, Dominik, Ma, Qian, Obadage, Rochana R., Ranjan, Zack, Koneru, Sai, Szabelska, Anna, Gill, Adam, Errington, Timothy M., Nematova, Shakhlo, Rajtmajer, Sarah, Wu, Jian, Jiang, Meng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reproducibility, Replicability, and Transparency in Research: What 430 Professors Think in Universities across the USA and India
by: Chakravorti, Tatiana, et al.
Published: (2024)
by: Chakravorti, Tatiana, et al.
Published: (2024)
Perspectives from India: Opportunities and Challenges for AI Replication Prediction to Improve Confidence in Published Research
by: Chakravorti, Tatiana, et al.
Published: (2023)
by: Chakravorti, Tatiana, et al.
Published: (2023)
CC30k: A Citation Contexts Dataset for Reproducibility-Oriented Sentiment Analysis
by: Obadage, Rochana R., et al.
Published: (2025)
by: Obadage, Rochana R., et al.
Published: (2025)
Can citations tell us about a paper's reproducibility? A case study of machine learning papers
by: Obadage, Rochana R., et al.
Published: (2024)
by: Obadage, Rochana R., et al.
Published: (2024)
[Re] Network Deconvolution
by: Obadage, Rochana R., et al.
Published: (2024)
by: Obadage, Rochana R., et al.
Published: (2024)
Can Large Language Models Discern Evidence for Scientific Hypotheses? Case Studies in the Social Sciences
by: Koneru, Sai, et al.
Published: (2023)
by: Koneru, Sai, et al.
Published: (2023)
Human-AI Collaboration for Estimating Scientific Replicability
by: Chakravorti, Tatiana, et al.
Published: (2026)
by: Chakravorti, Tatiana, et al.
Published: (2026)
Context Selection for Hypothesis and Statistical Evidence Extraction from Full-Text Scientific Articles
by: Koneru, Sai, et al.
Published: (2026)
by: Koneru, Sai, et al.
Published: (2026)
The Failed Migration of Academic Twitter: A Case Study of Precocious Adopters
by: Wang, Xinyu, et al.
Published: (2024)
by: Wang, Xinyu, et al.
Published: (2024)
Many Ways to Be Fake: Benchmarking Fake News Detection Under Strategy-Driven AI Generation
by: Wang, Xinyu, et al.
Published: (2026)
by: Wang, Xinyu, et al.
Published: (2026)
Evaluating Evidence Grounding Under User Pressure in Instruction-Tuned Language Models
by: Koneru, Sai, et al.
Published: (2026)
by: Koneru, Sai, et al.
Published: (2026)
Neural Activation during Phonological Processing in Primary‐School Children with Limited Reading Experience: Insights from Rural Côte d'Ivoire
by: Kaja K. Jasińska, et al.
Published: (2024)
by: Kaja K. Jasińska, et al.
Published: (2024)
The Unappreciated Role of Intent in Algorithmic Moderation of Social Media Content
by: Wang, Xinyu, et al.
Published: (2024)
by: Wang, Xinyu, et al.
Published: (2024)
ReplicationBench: Can AI Agents Replicate Astrophysics Research Papers?
by: Ye, Christine, et al.
Published: (2025)
by: Ye, Christine, et al.
Published: (2025)
Social Scientists on the Role of AI in Research
by: Chakravorti, Tatiana, et al.
Published: (2025)
by: Chakravorti, Tatiana, et al.
Published: (2025)
Transfer Learning Approach for Railway Technical Map (RTM) Component Identification
by: Rumalshan, Obadage Rochana, et al.
Published: (2024)
by: Rumalshan, Obadage Rochana, et al.
Published: (2024)
The Cost of Replicability in Active Learning
by: Hira, Rupkatha, et al.
Published: (2024)
by: Hira, Rupkatha, et al.
Published: (2024)
Interpretable Predictability-Based AI Text Detection: A Replication Study
by: Skurla, Adam, et al.
Published: (2026)
by: Skurla, Adam, et al.
Published: (2026)
Toward Robust URL Extraction for Open Science: A Study of arXiv File Formats and Temporal Trends
by: Obadage, Rochana R., et al.
Published: (2025)
by: Obadage, Rochana R., et al.
Published: (2025)
KnowledgeGain: Evaluating and Optimizing Science News Generation for Reader Learning
by: Soós, Dominik, et al.
Published: (2026)
by: Soós, Dominik, et al.
Published: (2026)
SCENEREPLICA: Benchmarking Real-World Robot Manipulation by Creating Replicable Scenes
by: Khargonkar, Ninad, et al.
Published: (2023)
by: Khargonkar, Ninad, et al.
Published: (2023)
ZEUS: An Efficient GPU Optimization Method Integrating PSO, BFGS, and Automatic Differentiation
by: Soos, Dominik, et al.
Published: (2026)
by: Soos, Dominik, et al.
Published: (2026)
Have LLMs Reopened the Pandora's Box of AI-Generated Fake News?
by: Wang, Xinyu, et al.
Published: (2024)
by: Wang, Xinyu, et al.
Published: (2024)
Let the Results Speak: A Replication-First Paradigm for LLM Behavioral Benchmarking
by: Yuming, et al.
Published: (2026)
by: Yuming, et al.
Published: (2026)
SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers
by: Xiang, Yanzheng, et al.
Published: (2025)
by: Xiang, Yanzheng, et al.
Published: (2025)
UNIVERSAL AND NATIONAL CHARACTERISTICS SHAPING THE CONCEPTUAL MODEL OF DENDRONYMS
by: Shakhlo Mitanova
Published: (2026)
by: Shakhlo Mitanova
Published: (2026)
PaperBench: Evaluating AI's Ability to Replicate AI Research
by: Starace, Giulio, et al.
Published: (2025)
by: Starace, Giulio, et al.
Published: (2025)
Towards Using Multiple Iterated, Reproduced, and Replicated Experiments with Robots (MIRRER) for Evaluation and Benchmarking
by: Norton, Adam, et al.
Published: (2024)
by: Norton, Adam, et al.
Published: (2024)
LLM-Assisted Replication for Quantitative Social Science
by: Kubota, So, et al.
Published: (2026)
by: Kubota, So, et al.
Published: (2026)
The Difference Between "Replicable" and "Not replicable" is not Itself Scientifically Replicable
by: Devezer, Berna, et al.
Published: (2026)
by: Devezer, Berna, et al.
Published: (2026)
THE IMPORTANCE OF PROPER CAREER GUIDANCE FOR YOUNG ELEMENTARY SCHOOL STUDENTS
by: Sitora Nematova
Published: (2025)
by: Sitora Nematova
Published: (2025)
RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents
by: Black, Sid, et al.
Published: (2025)
by: Black, Sid, et al.
Published: (2025)
Replicable Clustering
by: Esfandiari, Hossein, et al.
Published: (2023)
by: Esfandiari, Hossein, et al.
Published: (2023)
Replicable Composition
by: Banihashem, Kiarash, et al.
Published: (2026)
by: Banihashem, Kiarash, et al.
Published: (2026)
Replicability of Experiment
by: John D. Norton
Published: (2015)
by: John D. Norton
Published: (2015)
Replications as Project‐Based Learning in Geographic Information Science
by: Peter Kedron, et al.
Published: (2025)
by: Peter Kedron, et al.
Published: (2025)
Stochastic Models for Replication Origin Spacings in Eukaryotic DNA Replication
by: Day, Huw, et al.
Published: (2022)
by: Day, Huw, et al.
Published: (2022)
License to Replicate: Mechanisms of Licensing Eukaryotic Origins for DNA Replication
by: Victoria Frisbie, et al.
Published: (2025)
by: Victoria Frisbie, et al.
Published: (2025)
PharmacoepidemiologicalEffects of Antibacterial Drugs for Community-Acquired Pneumonia in Children of Different Ages in a Modern Interpretation
by: Kodirova Shakhlo Salokhitdinovna
Published: (2025)
by: Kodirova Shakhlo Salokhitdinovna
Published: (2025)
Embracing Transparency: A Study of Open Science Practices Among Early Career HCI Researchers
by: Chakravorti, Tatiana, et al.
Published: (2024)
by: Chakravorti, Tatiana, et al.
Published: (2024)
Similar Items
-
Reproducibility, Replicability, and Transparency in Research: What 430 Professors Think in Universities across the USA and India
by: Chakravorti, Tatiana, et al.
Published: (2024) -
Perspectives from India: Opportunities and Challenges for AI Replication Prediction to Improve Confidence in Published Research
by: Chakravorti, Tatiana, et al.
Published: (2023) -
CC30k: A Citation Contexts Dataset for Reproducibility-Oriented Sentiment Analysis
by: Obadage, Rochana R., et al.
Published: (2025) -
Can citations tell us about a paper's reproducibility? A case study of machine learning papers
by: Obadage, Rochana R., et al.
Published: (2024) -
[Re] Network Deconvolution
by: Obadage, Rochana R., et al.
Published: (2024)