REPRO-Bench: Can Agentic AI Systems Assess the Reproducibility of Social Science Research?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Chuxuan, Zhang, Liyun, Lim, Yeji, Wadhwani, Aum, Peters, Austin, Kang, Daniel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DRAMA: Unifying Data Retrieval and Analysis for Open-Domain Analytic Queries
von: Hu, Chuxuan, et al.
Veröffentlicht: (2025)
von: Hu, Chuxuan, et al.
Veröffentlicht: (2025)
LEAP: LLM-powered End-to-end Automatic Library for Processing Social Science Queries on Unstructured Data
von: Hu, Chuxuan, et al.
Veröffentlicht: (2025)
von: Hu, Chuxuan, et al.
Veröffentlicht: (2025)
SODIUM: From Open Web Data to Queryable Databases
von: Hu, Chuxuan, et al.
Veröffentlicht: (2026)
von: Hu, Chuxuan, et al.
Veröffentlicht: (2026)
LLM-Measure: Generating Valid, Consistent, and Reproducible Text-Based Measures for Social Science Research
von: Yang, Yi, et al.
Veröffentlicht: (2024)
von: Yang, Yi, et al.
Veröffentlicht: (2024)
Can Coding Agents Reproduce Findings in Computational Materials Science?
von: Huang, Ziyang, et al.
Veröffentlicht: (2026)
von: Huang, Ziyang, et al.
Veröffentlicht: (2026)
Accelerating Social Science Research via Agentic Hypothesization and Experimentation
von: Gupta, Jishu Sen, et al.
Veröffentlicht: (2026)
von: Gupta, Jishu Sen, et al.
Veröffentlicht: (2026)
Breaking Barriers: Do Reinforcement Post Training Gains Transfer To Unseen Domains?
von: Hu, Chuxuan, et al.
Veröffentlicht: (2025)
von: Hu, Chuxuan, et al.
Veröffentlicht: (2025)
ReplicationBench: Can AI Agents Replicate Astrophysics Research Papers?
von: Ye, Christine, et al.
Veröffentlicht: (2025)
von: Ye, Christine, et al.
Veröffentlicht: (2025)
RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic Retrieval Augmented Generation Systems
von: Lin, Jingru, et al.
Veröffentlicht: (2025)
von: Lin, Jingru, et al.
Veröffentlicht: (2025)
KULTURE Bench: A Benchmark for Assessing Language Model in Korean Cultural Context
von: Wang, Xiaonan, et al.
Veröffentlicht: (2024)
von: Wang, Xiaonan, et al.
Veröffentlicht: (2024)
SysBench: Can Large Language Models Follow System Messages?
von: Qin, Yanzhao, et al.
Veröffentlicht: (2024)
von: Qin, Yanzhao, et al.
Veröffentlicht: (2024)
X-WebAgentBench: A Multilingual Interactive Web Benchmark for Evaluating Global Agentic System
von: Wang, Peng, et al.
Veröffentlicht: (2025)
von: Wang, Peng, et al.
Veröffentlicht: (2025)
SocialMemBench: Are AI Memory Systems Ready for Social Group Settings?
von: Owolabi, Olukunle
Veröffentlicht: (2026)
von: Owolabi, Olukunle
Veröffentlicht: (2026)
INTERACT: Enabling Interactive, Question-Driven Learning in Large Language Models
von: Kendapadi, Aum, et al.
Veröffentlicht: (2024)
von: Kendapadi, Aum, et al.
Veröffentlicht: (2024)
Can MLLMs Read the Room? A Multimodal Benchmark for Assessing Deception in Multi-Party Social Interactions
von: Kang, Caixin, et al.
Veröffentlicht: (2025)
von: Kang, Caixin, et al.
Veröffentlicht: (2025)
PhageBench: Can LLMs Understand Raw Bacteriophage Genomes?
von: Hou, Yusen, et al.
Veröffentlicht: (2026)
von: Hou, Yusen, et al.
Veröffentlicht: (2026)
ClawBench: Can AI Agents Complete Everyday Online Tasks?
von: Zhang, Yuxuan, et al.
Veröffentlicht: (2026)
von: Zhang, Yuxuan, et al.
Veröffentlicht: (2026)
RExBench: Can coding agents autonomously implement AI research extensions?
von: Edwards, Nicholas, et al.
Veröffentlicht: (2025)
von: Edwards, Nicholas, et al.
Veröffentlicht: (2025)
LLM Agent Meets Agentic AI: Can LLM Agents Simulate Customers to Evaluate Agentic-AI-based Shopping Assistants?
von: Sun, Lu, et al.
Veröffentlicht: (2025)
von: Sun, Lu, et al.
Veröffentlicht: (2025)
AI Idea Bench 2025: AI Research Idea Generation Benchmark
von: Qiu, Yansheng, et al.
Veröffentlicht: (2025)
von: Qiu, Yansheng, et al.
Veröffentlicht: (2025)
When Do LLM Agents Treat Surface Noise Differently from Semantic Noise? A 68-Cell Measurement Study with a Held-Out Trace-Level Validation
von: Zhang, Liyun, et al.
Veröffentlicht: (2026)
von: Zhang, Liyun, et al.
Veröffentlicht: (2026)
AI-Driven Automation Can Become the Foundation of Next-Era Science of Science Research
von: Chen, Renqi, et al.
Veröffentlicht: (2025)
von: Chen, Renqi, et al.
Veröffentlicht: (2025)
RedacBench: Can AI Erase Your Secrets?
von: Jeon, Hyunjun, et al.
Veröffentlicht: (2026)
von: Jeon, Hyunjun, et al.
Veröffentlicht: (2026)
Automating Computational Reproducibility in Social Science: Comparing Prompt-Based and Agent-Based Approaches
von: Shah, Syed Mehtab Hussain, et al.
Veröffentlicht: (2026)
von: Shah, Syed Mehtab Hussain, et al.
Veröffentlicht: (2026)
React to This (RTT): A Nonverbal Turing Test for Embodied AI
von: Zhang, Chuxuan, et al.
Veröffentlicht: (2025)
von: Zhang, Chuxuan, et al.
Veröffentlicht: (2025)
Rethinking Scale: The Efficacy of Fine-Tuned Open-Source LLMs in Large-Scale Reproducible Social Science Research
von: Carammia, Marcello, et al.
Veröffentlicht: (2024)
von: Carammia, Marcello, et al.
Veröffentlicht: (2024)
IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation
von: Schmitt, Johannes, et al.
Veröffentlicht: (2025)
von: Schmitt, Johannes, et al.
Veröffentlicht: (2025)
CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark
von: Siegel, Zachary S., et al.
Veröffentlicht: (2024)
von: Siegel, Zachary S., et al.
Veröffentlicht: (2024)
Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools
von: Magesh, Varun, et al.
Veröffentlicht: (2024)
von: Magesh, Varun, et al.
Veröffentlicht: (2024)
Can Large Language Models Transform Computational Social Science?
von: Ziems, Caleb, et al.
Veröffentlicht: (2023)
von: Ziems, Caleb, et al.
Veröffentlicht: (2023)
PaperBench: Evaluating AI's Ability to Replicate AI Research
von: Starace, Giulio, et al.
Veröffentlicht: (2025)
von: Starace, Giulio, et al.
Veröffentlicht: (2025)
AAAR-1.0: Assessing AI's Potential to Assist Research
von: Lou, Renze, et al.
Veröffentlicht: (2024)
von: Lou, Renze, et al.
Veröffentlicht: (2024)
RECAP: Reproducing Copyrighted Data from LLMs Training with an Agentic Pipeline
von: Duarte, André V., et al.
Veröffentlicht: (2025)
von: Duarte, André V., et al.
Veröffentlicht: (2025)
ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences
von: Nguyen, Bang, et al.
Veröffentlicht: (2026)
von: Nguyen, Bang, et al.
Veröffentlicht: (2026)
RareBench: Can LLMs Serve as Rare Diseases Specialists?
von: Chen, Xuanzhong, et al.
Veröffentlicht: (2024)
von: Chen, Xuanzhong, et al.
Veröffentlicht: (2024)
RefuteBench 2.0 -- Agentic Benchmark for Dynamic Evaluation of LLM Responses to Refutation Instruction
von: Yan, Jianhao, et al.
Veröffentlicht: (2025)
von: Yan, Jianhao, et al.
Veröffentlicht: (2025)
OmniGenBench: A Modular Platform for Reproducible Genomic Foundation Models Benchmarking
von: Yang, Heng, et al.
Veröffentlicht: (2025)
von: Yang, Heng, et al.
Veröffentlicht: (2025)
Salsa as a Nonverbal Embodied Language -- The CoMPAS3D Dataset and Benchmarks
von: Burkanova, Bermet, et al.
Veröffentlicht: (2025)
von: Burkanova, Bermet, et al.
Veröffentlicht: (2025)
Situating AI Agents in their World: Aspective Agentic AI for Dynamic Partially Observable Information Systems
von: Bentley, Peter J., et al.
Veröffentlicht: (2025)
von: Bentley, Peter J., et al.
Veröffentlicht: (2025)
Enriching Social Science Research via Survey Item Linking
von: Tsereteli, Tornike, et al.
Veröffentlicht: (2024)
von: Tsereteli, Tornike, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
DRAMA: Unifying Data Retrieval and Analysis for Open-Domain Analytic Queries
von: Hu, Chuxuan, et al.
Veröffentlicht: (2025) -
LEAP: LLM-powered End-to-end Automatic Library for Processing Social Science Queries on Unstructured Data
von: Hu, Chuxuan, et al.
Veröffentlicht: (2025) -
SODIUM: From Open Web Data to Queryable Databases
von: Hu, Chuxuan, et al.
Veröffentlicht: (2026) -
LLM-Measure: Generating Valid, Consistent, and Reproducible Text-Based Measures for Social Science Research
von: Yang, Yi, et al.
Veröffentlicht: (2024) -
Can Coding Agents Reproduce Findings in Computational Materials Science?
von: Huang, Ziyang, et al.
Veröffentlicht: (2026)