Read the Paper, Write the Code: Agentic Reproduction of Social-Science Results
Fuente:
arXiv
Guardado en:
| Autores principales: | Kohler, Benjamin, Zollikofer, David, Einsiedler, Johanna, Hoyle, Alexander, Ash, Elliott |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Synthetic Cancer -- Augmenting Worms with LLMs
por: Zimmerman, Benjamin, et al.
Publicado: (2024)
por: Zimmerman, Benjamin, et al.
Publicado: (2024)
Can Reasoning Help Large Language Models Capture Human Annotator Disagreement?
por: Ni, Jingwei, et al.
Publicado: (2025)
por: Ni, Jingwei, et al.
Publicado: (2025)
The Phenomenology of Machine: A Comprehensive Analysis of the Sentience of the OpenAI-o1 Model Integrating Functionalism, Consciousness Theories, Active Inference, and AI Architectures
por: Hoyle, Victoria Violet
Publicado: (2024)
por: Hoyle, Victoria Violet
Publicado: (2024)
Reflective Paper-to-Code Reproduction Enabled by Fine-Grained Verification
por: Zhou, Mingyang, et al.
Publicado: (2025)
por: Zhou, Mingyang, et al.
Publicado: (2025)
PaperRepro: Automated Computational Reproducibility Assessment for Social Science Papers
por: Zhang, Linhao, et al.
Publicado: (2026)
por: Zhang, Linhao, et al.
Publicado: (2026)
Accelerating Social Science Research via Agentic Hypothesization and Experimentation
por: Gupta, Jishu Sen, et al.
Publicado: (2026)
por: Gupta, Jishu Sen, et al.
Publicado: (2026)
What Papers Don't Tell You: Recovering Tacit Knowledge for Automated Paper Reproduction
por: Li, Lehui, et al.
Publicado: (2026)
por: Li, Lehui, et al.
Publicado: (2026)
FormalScience: Scalable Human-in-the-Loop Autoformalisation of Science with Agentic Code Generation in Lean
por: Meadows, Jordan, et al.
Publicado: (2026)
por: Meadows, Jordan, et al.
Publicado: (2026)
Variational Best-of-N Alignment
por: Amini, Afra, et al.
Publicado: (2024)
por: Amini, Afra, et al.
Publicado: (2024)
AutoReproduce: Automatic AI Experiment Reproduction with Paper Lineage
por: Zhao, Xuanle, et al.
Publicado: (2025)
por: Zhao, Xuanle, et al.
Publicado: (2025)
Enhancing Automated Paper Reproduction via Prompt-Free Collaborative Agents
por: Lin, Zijie, et al.
Publicado: (2025)
por: Lin, Zijie, et al.
Publicado: (2025)
How Persuasive is Your Context?
por: Nguyen, Tu, et al.
Publicado: (2025)
por: Nguyen, Tu, et al.
Publicado: (2025)
Measuring Scalar Constructs in Social Science with LLMs
por: Licht, Hauke, et al.
Publicado: (2025)
por: Licht, Hauke, et al.
Publicado: (2025)
Dynamic Cogeneration of Bug Reproduction Test in Agentic Program Repair
por: Cheng, Runxiang, et al.
Publicado: (2026)
por: Cheng, Runxiang, et al.
Publicado: (2026)
Agentic Bug Reproduction for Effective Automated Program Repair at Google
por: Cheng, Runxiang, et al.
Publicado: (2025)
por: Cheng, Runxiang, et al.
Publicado: (2025)
DeepCode: Open Agentic Coding
por: Li, Zongwei, et al.
Publicado: (2025)
por: Li, Zongwei, et al.
Publicado: (2025)
APRES: An Agentic Paper Revision and Evaluation System
por: Zhao, Bingchen, et al.
Publicado: (2026)
por: Zhao, Bingchen, et al.
Publicado: (2026)
AFaCTA: Assisting the Annotation of Factual Claim Detection with Reliable LLM Annotators
por: Ni, Jingwei, et al.
Publicado: (2024)
por: Ni, Jingwei, et al.
Publicado: (2024)
Sound Agentic Science Requires Adversarial Experiments
por: Fa, Dionizije, et al.
Publicado: (2026)
por: Fa, Dionizije, et al.
Publicado: (2026)
Preacher: Paper-to-Video Agentic System
por: Liu, Jingwei, et al.
Publicado: (2025)
por: Liu, Jingwei, et al.
Publicado: (2025)
PaperScope: A Multi-Modal Multi-Document Benchmark for Agentic Deep Research Across Massive Scientific Papers
por: Xiong, Lei, et al.
Publicado: (2026)
por: Xiong, Lei, et al.
Publicado: (2026)
PaperArena: An Evaluation Benchmark for Tool-Augmented Agentic Reasoning on Scientific Literature
por: Wang, Daoyu, et al.
Publicado: (2025)
por: Wang, Daoyu, et al.
Publicado: (2025)
From Reproduction to Replication: Evaluating Research Agents with Progressive Code Masking
por: Kim, Gyeongwon James, et al.
Publicado: (2025)
por: Kim, Gyeongwon James, et al.
Publicado: (2025)
Agentic Code Reasoning
por: Ugare, Shubham, et al.
Publicado: (2026)
por: Ugare, Shubham, et al.
Publicado: (2026)
Vibe Coding vs. Agentic Coding: Fundamentals and Practical Implications of Agentic AI
por: Sapkota, Ranjan, et al.
Publicado: (2025)
por: Sapkota, Ranjan, et al.
Publicado: (2025)
PaperOrchestra: A Multi-Agent Framework for Automated AI Research Paper Writing
por: Song, Yiwen, et al.
Publicado: (2026)
por: Song, Yiwen, et al.
Publicado: (2026)
Sanity Checks for Agentic Data Science
por: Rewolinski, Zachary T., et al.
Publicado: (2026)
por: Rewolinski, Zachary T., et al.
Publicado: (2026)
Towards Agentic Intelligence for Materials Science
por: Zhang, Huan, et al.
Publicado: (2026)
por: Zhang, Huan, et al.
Publicado: (2026)
"I'm Not Reading All of That": Understanding Software Engineers' Level of Cognitive Engagement with Agentic Coding Assistants
por: Catalan, Carlos Rafael, et al.
Publicado: (2026)
por: Catalan, Carlos Rafael, et al.
Publicado: (2026)
Whose Preferences? Differences in Fairness Preferences and Their Impact on the Fairness of AI Utilizing Human Feedback
por: Lerner, Emilia Agis, et al.
Publicado: (2024)
por: Lerner, Emilia Agis, et al.
Publicado: (2024)
Resilient Write: A Six-Layer Durable Write Surface for LLM Coding Agents
por: Agyemang, Justice Owusu, et al.
Publicado: (2026)
por: Agyemang, Justice Owusu, et al.
Publicado: (2026)
Embodied Science: Closing the Discovery Loop with Agentic Embodied AI
por: Zhuang, Xiang, et al.
Publicado: (2026)
por: Zhuang, Xiang, et al.
Publicado: (2026)
An Agentic Evaluation Framework for AI-Generated Scientific Code in PETSc
por: Zhang, Hong, et al.
Publicado: (2026)
por: Zhang, Hong, et al.
Publicado: (2026)
Signal Reshaping for GRPO in Weak-Feedback Agentic Code Repair
por: Li, Jia, et al.
Publicado: (2026)
por: Li, Jia, et al.
Publicado: (2026)
DialogueForge: LLM Simulation of Human-Chatbot Dialogue
por: Zhu, Ruizhe, et al.
Publicado: (2025)
por: Zhu, Ruizhe, et al.
Publicado: (2025)
PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading
por: Wu, Yutao, et al.
Publicado: (2025)
por: Wu, Yutao, et al.
Publicado: (2025)
DIRAS: Efficient LLM Annotation of Document Relevance in Retrieval Augmented Generation
por: Ni, Jingwei, et al.
Publicado: (2024)
por: Ni, Jingwei, et al.
Publicado: (2024)
A Study of Plasticity Loss in On-Policy Deep Reinforcement Learning
por: Juliani, Arthur, et al.
Publicado: (2024)
por: Juliani, Arthur, et al.
Publicado: (2024)
Non-Monotonic Attention-based Read/Write Policy Learning for Simultaneous Translation
por: Ahmed, Zeeshan, et al.
Publicado: (2025)
por: Ahmed, Zeeshan, et al.
Publicado: (2025)
Agentic Coding Needs Proactivity, Not Just Autonomy
por: Bui, Nghi D. Q., et al.
Publicado: (2026)
por: Bui, Nghi D. Q., et al.
Publicado: (2026)
Ejemplares similares
-
Synthetic Cancer -- Augmenting Worms with LLMs
por: Zimmerman, Benjamin, et al.
Publicado: (2024) -
Can Reasoning Help Large Language Models Capture Human Annotator Disagreement?
por: Ni, Jingwei, et al.
Publicado: (2025) -
The Phenomenology of Machine: A Comprehensive Analysis of the Sentience of the OpenAI-o1 Model Integrating Functionalism, Consciousness Theories, Active Inference, and AI Architectures
por: Hoyle, Victoria Violet
Publicado: (2024) -
Reflective Paper-to-Code Reproduction Enabled by Fine-Grained Verification
por: Zhou, Mingyang, et al.
Publicado: (2025) -
PaperRepro: Automated Computational Reproducibility Assessment for Social Science Papers
por: Zhang, Linhao, et al.
Publicado: (2026)