ProgramBench: Can Language Models Rebuild Programs From Scratch?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, John, Lieret, Kilian, Ma, Jeffrey, Thakkar, Parth, Pedchenko, Dmitrii, Sootla, Sten, McMilin, Emily, Yin, Pengcheng, Hou, Rui, Synnaeve, Gabriel, Yang, Diyi, Press, Ofir |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Underspecification in Language Modeling Tasks: A Causality-Informed Study of Gendered Pronoun Resolution
von: McMilin, Emily
Veröffentlicht: (2022)
von: McMilin, Emily
Veröffentlicht: (2022)
CodeClash: Benchmarking Goal-Oriented Software Engineering
von: Yang, John, et al.
Veröffentlicht: (2025)
von: Yang, John, et al.
Veröffentlicht: (2025)
SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
von: Yang, John, et al.
Veröffentlicht: (2024)
von: Yang, John, et al.
Veröffentlicht: (2024)
SWE-smith: Scaling Data for Software Engineering Agents
von: Yang, John, et al.
Veröffentlicht: (2025)
von: Yang, John, et al.
Veröffentlicht: (2025)
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
von: Yang, John, et al.
Veröffentlicht: (2024)
von: Yang, John, et al.
Veröffentlicht: (2024)
Toward Training Superintelligent Software Agents through Self-Play SWE-RL
von: Wei, Yuxiang, et al.
Veröffentlicht: (2025)
von: Wei, Yuxiang, et al.
Veröffentlicht: (2025)
AlgoTune: Can Language Models Speed Up General-Purpose Numerical Programs?
von: Press, Ori, et al.
Veröffentlicht: (2025)
von: Press, Ori, et al.
Veröffentlicht: (2025)
VideoGameBench: Can Vision-Language Models complete popular video games?
von: Zhang, Alex L., et al.
Veröffentlicht: (2025)
von: Zhang, Alex L., et al.
Veröffentlicht: (2025)
Properties of Eventually Positive Linear Input-Output Systems
von: Sootla, Aivar
Veröffentlicht: (2015)
von: Sootla, Aivar
Veröffentlicht: (2015)
AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?
von: Yoran, Ori, et al.
Veröffentlicht: (2024)
von: Yoran, Ori, et al.
Veröffentlicht: (2024)
CiteME: Can Language Models Accurately Cite Scientific Claims?
von: Press, Ori, et al.
Veröffentlicht: (2024)
von: Press, Ori, et al.
Veröffentlicht: (2024)
BigO(Bench) -- Can LLMs Generate Code with Controlled Time and Space Complexity?
von: Chambon, Pierre, et al.
Veröffentlicht: (2025)
von: Chambon, Pierre, et al.
Veröffentlicht: (2025)
Automatically Generating Questions About Scratch Programs
von: Obermüller, Florian, et al.
Veröffentlicht: (2025)
von: Obermüller, Florian, et al.
Veröffentlicht: (2025)
Detecting Gender Stereotypes in Scratch Programming Tutorials
von: Graßl, Isabella, et al.
Veröffentlicht: (2025)
von: Graßl, Isabella, et al.
Veröffentlicht: (2025)
Sicheres autonomes Flugrobotersystem für den Einsatz im Produktions- und Logistikumfeld
von: Lieret, Markus
Veröffentlicht: (2025)
von: Lieret, Markus
Veröffentlicht: (2025)
Finding Needles in Images: Can Multimodal LLMs Locate Fine Details?
von: Thakkar, Parth, et al.
Veröffentlicht: (2025)
von: Thakkar, Parth, et al.
Veröffentlicht: (2025)
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
von: Jimenez, Carlos E., et al.
Veröffentlicht: (2023)
von: Jimenez, Carlos E., et al.
Veröffentlicht: (2023)
WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch
von: Lu, Zimu, et al.
Veröffentlicht: (2025)
von: Lu, Zimu, et al.
Veröffentlicht: (2025)
SWE-fficiency: Can Language Models Optimize Real-World Repositories on Real Workloads?
von: Ma, Jeffrey Jian, et al.
Veröffentlicht: (2025)
von: Ma, Jeffrey Jian, et al.
Veröffentlicht: (2025)
EquiBench: Benchmarking Large Language Models' Reasoning about Program Semantics via Equivalence Checking
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
von: Si, Chenglei, et al.
Veröffentlicht: (2024)
von: Si, Chenglei, et al.
Veröffentlicht: (2024)
Differentiating Through a Quadratic Cone Program
von: Healey, Quill, et al.
Veröffentlicht: (2025)
von: Healey, Quill, et al.
Veröffentlicht: (2025)
Disproving Program Equivalence with LLMs
von: Allamanis, Miltiadis, et al.
Veröffentlicht: (2025)
von: Allamanis, Miltiadis, et al.
Veröffentlicht: (2025)
EnIGMA: Interactive Tools Substantially Assist LM Agents in Finding Security Vulnerabilities
von: Abramovich, Talor, et al.
Veröffentlicht: (2024)
von: Abramovich, Talor, et al.
Veröffentlicht: (2024)
Can LLMs Reason in the Wild with Programs?
von: Yang, Yuan, et al.
Veröffentlicht: (2024)
von: Yang, Yuan, et al.
Veröffentlicht: (2024)
Chapter 13 Rebuilding Relationships between Data, Method, and Theories
von: Cucina, Jeffrey M., et al.
Veröffentlicht: (2022)
von: Cucina, Jeffrey M., et al.
Veröffentlicht: (2022)
La inmigración hispana en Italia: hacia una variedad de contacto entre español e italiano
von: Milin Bonomi
Veröffentlicht: (2010)
von: Milin Bonomi
Veröffentlicht: (2010)
Entre divergencia y acomodación: el caso de los inmigrantes hispanos en Barcelona y Milán
von: Milin Bonomi
Veröffentlicht: (2010)
von: Milin Bonomi
Veröffentlicht: (2010)
Alonso, José Antonio y Rodolfo Gutiérrez (dir.). 2010. Emigración y lengua. El papel del español en las migraciones internacionales. Madrid: Ariel y Fundación Telefónica, 304 pp.
von: Milin Bonomi
Veröffentlicht: (2011)
von: Milin Bonomi
Veröffentlicht: (2011)
Can prompt cusps of WIMP dark matter be detected as individual gamma-ray sources?
von: Delos, M. Sten
Veröffentlicht: (2023)
von: Delos, M. Sten
Veröffentlicht: (2023)
ScratchEval : A Multimodal Evaluation Framework for LLMs in Block-Based Programming
von: Si, Yuan, et al.
Veröffentlicht: (2026)
von: Si, Yuan, et al.
Veröffentlicht: (2026)
Raven: Rethinking Automated Assessment for Scratch Programs via Video-Grounded Evaluation
von: Li, Donglin, et al.
Veröffentlicht: (2026)
von: Li, Donglin, et al.
Veröffentlicht: (2026)
Rebuilding Syria
Veröffentlicht: (2019)
Veröffentlicht: (2019)
ProBench: Benchmarking Large Language Models in Competitive Programming
von: Yang, Lei, et al.
Veröffentlicht: (2025)
von: Yang, Lei, et al.
Veröffentlicht: (2025)
How Rules Represent Causal Knowledge: Causal Modeling with Abductive Logic Programs
von: Rückschloß, Kilian, et al.
Veröffentlicht: (2025)
von: Rückschloß, Kilian, et al.
Veröffentlicht: (2025)
Collection Development from Scratch: Supporting a New Degree Program in Construction Management
von: Reycraft, Kimberly
Veröffentlicht: (2020)
von: Reycraft, Kimberly
Veröffentlicht: (2020)
GeoGramBench: Benchmarking the Geometric Program Reasoning in Modern LLMs
von: Luo, Shixian, et al.
Veröffentlicht: (2025)
von: Luo, Shixian, et al.
Veröffentlicht: (2025)
ReplicationBench: Can AI Agents Replicate Astrophysics Research Papers?
von: Ye, Christine, et al.
Veröffentlicht: (2025)
von: Ye, Christine, et al.
Veröffentlicht: (2025)
Dynamic Skill Adaptation for Large Language Models
von: Chen, Jiaao, et al.
Veröffentlicht: (2024)
von: Chen, Jiaao, et al.
Veröffentlicht: (2024)
Searching for Privacy Risks in LLM Agents via Simulation
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2025)
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Underspecification in Language Modeling Tasks: A Causality-Informed Study of Gendered Pronoun Resolution
von: McMilin, Emily
Veröffentlicht: (2022) -
CodeClash: Benchmarking Goal-Oriented Software Engineering
von: Yang, John, et al.
Veröffentlicht: (2025) -
SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
von: Yang, John, et al.
Veröffentlicht: (2024) -
SWE-smith: Scaling Data for Software Engineering Agents
von: Yang, John, et al.
Veröffentlicht: (2025) -
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
von: Yang, John, et al.
Veröffentlicht: (2024)