BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Zehua, Bajaj, Ati Priya, Handa, Divij, Liu, Siyu, Raj, Arvind S, Chen, Hongkai, Wang, Hulin, Liu, Yibo, Basque, Zion Leonahenahe, Nath, Souradip, Juneja, Vishal, Chapre, Nikhil, Shoshitaishvili, Yan, Doupé, Adam, Baral, Chitta, Wang, Ruoyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Root-Cause-Driven Automated Vulnerability Repair
by: Wang, Hulin, et al.
Published: (2026)
by: Wang, Hulin, et al.
Published: (2026)
Pushan: Trace-Free Deobfuscation of Virtualization-Obfuscated Binaries
by: Sudhir, Ashwin, et al.
Published: (2026)
by: Sudhir, Ashwin, et al.
Published: (2026)
When "Competency" in Reasoning Opens the Door to Vulnerability: Jailbreaking LLMs via Novel Complex Ciphers
by: Handa, Divij, et al.
Published: (2024)
by: Handa, Divij, et al.
Published: (2024)
ActionReasoningBench: Reasoning about Actions with and without Ramification Constraints
by: Handa, Divij, et al.
Published: (2024)
by: Handa, Divij, et al.
Published: (2024)
Hypothesis Generation for Materials Discovery and Design Using Goal-Driven and Constraint-Guided LLM Agents
by: Kumbhar, Shrinidhi, et al.
Published: (2025)
by: Kumbhar, Shrinidhi, et al.
Published: (2025)
GuidedSampling: Steering LLMs Towards Diverse Candidate Solutions at Inference-Time
by: Handa, Divij, et al.
Published: (2025)
by: Handa, Divij, et al.
Published: (2025)
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models
by: RRV, Aswin, et al.
Published: (2026)
by: RRV, Aswin, et al.
Published: (2026)
ThinkTuning: Instilling Cognitive Reflections without Distillation
by: RRV, Aswin, et al.
Published: (2025)
by: RRV, Aswin, et al.
Published: (2025)
D-LiFT: Improving LLM-based Decompiler Backend via Code Quality-driven Fine-tuning
by: Zou, Muqi, et al.
Published: (2025)
by: Zou, Muqi, et al.
Published: (2025)
Can ChatGPT Perform Image Splicing Detection? A Preliminary Study
by: Nath, Souradip
Published: (2025)
by: Nath, Souradip
Published: (2025)
UnSeenTimeQA: Time-Sensitive Question-Answering Beyond LLMs' Memorization
by: Uddin, Md Nayem, et al.
Published: (2024)
by: Uddin, Md Nayem, et al.
Published: (2024)
Take a Step Further: Understanding Page Spray in Linux Kernel Exploitation
by: Guo, Ziyi, et al.
Published: (2024)
by: Guo, Ziyi, et al.
Published: (2024)
OptAgent: Optimizing Query Rewriting for E-commerce via Multi-Agent Simulation
by: Handa, Divij, et al.
Published: (2025)
by: Handa, Divij, et al.
Published: (2025)
Evaluating Gender Bias of LLMs in Making Morality Judgements
by: Bajaj, Divij, et al.
Published: (2024)
by: Bajaj, Divij, et al.
Published: (2024)
Hepatoprotective Potential of Allium cepa Bulb Extract in Mitigating Lead Toxicity in Rattus norvegicus
by: Priya Bajaj and Anjani Rani
Published: (2025)
by: Priya Bajaj and Anjani Rani
Published: (2025)
Managing translocations of aquatic species / R. G. Doupé
by: Doupé, R. G
Published: (2000)
by: Doupé, R. G
Published: (2000)
Genetic variation in the growth traits of straight-bred and crossbred black bream (Acanthopagrus butcheri Munro) at 90 days of age. / R G Doupé
by: Doupé R. G
Published: (2003)
by: Doupé R. G
Published: (2003)
Short communication Visible implant fluorescent elastomer tags as pedigree markers for applied aquaculture : n evaluation using black bream acanthopagrus butcheri / R. G. Doupé
by: Doupé, R. G
Published: (2003)
by: Doupé, R. G
Published: (2003)
ActionCOMET: A Zero-shot Approach to Learn Image-specific Commonsense Concepts about Actions
by: Sampat, Shailaja Keyur, et al.
Published: (2024)
by: Sampat, Shailaja Keyur, et al.
Published: (2024)
ARVO: Atlas of Reproducible Vulnerabilities for Open Source Software
by: Mei, Xiang, et al.
Published: (2024)
by: Mei, Xiang, et al.
Published: (2024)
ASE NSW 246th general meeting and the last lecture of 2023
by: Atieh (Ati) Sadr
Published: (2024)
by: Atieh (Ati) Sadr
Published: (2024)
AEJ report for ASE NSW 62nd AGM and 247th General Meeting
by: Atieh (Ati) Sadr
Published: (2024)
by: Atieh (Ati) Sadr
Published: (2024)
Don't Blame the Annotator: Bias Already Starts in the Annotation Instructions
by: Parmar, Mihir, et al.
Published: (2022)
by: Parmar, Mihir, et al.
Published: (2022)
$λ$-ECLIPSE: Multi-Concept Personalized Text-to-Image Diffusion Models by Leveraging CLIP Latent Space
by: Patel, Maitreya, et al.
Published: (2024)
by: Patel, Maitreya, et al.
Published: (2024)
On the Robustness of Language Guidance for Low-Level Vision Tasks: Findings from Depth Estimation
by: Chatterjee, Agneet, et al.
Published: (2024)
by: Chatterjee, Agneet, et al.
Published: (2024)
MMTABREAL: Real-World Benchmark for Multimodal Table Understanding
by: Titiya, Prasham, et al.
Published: (2025)
by: Titiya, Prasham, et al.
Published: (2025)
GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning
by: Siingh, Shikhhar, et al.
Published: (2025)
by: Siingh, Shikhhar, et al.
Published: (2025)
ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion Models
by: Patel, Maitreya, et al.
Published: (2023)
by: Patel, Maitreya, et al.
Published: (2023)
Map&Make: Schema Guided Text to Table Generation
by: Ahuja, Naman, et al.
Published: (2025)
by: Ahuja, Naman, et al.
Published: (2025)
The Art of Defending: A Systematic Evaluation and Analysis of LLM Defense Strategies on Safety and Over-Defensiveness
by: Varshney, Neeraj, et al.
Published: (2023)
by: Varshney, Neeraj, et al.
Published: (2023)
Invariant Measure for Linear Stochastic PDEs in the space of Tempered distributions
by: Nath, Arvind Kumar
Published: (2024)
by: Nath, Arvind Kumar
Published: (2024)
Evaporating universes
by: Gupta, Divij
Published: (2025)
by: Gupta, Divij
Published: (2025)
MENINGKATKAN KEMAMPUAN MENULIS TEKS ULASAN CERPEN MENGGUNAKAN TEKNIK BRAINWRITING PADA SISWA KELAS VIII SMP NEGERI 8 WASILE
by: Senen, Kristin, et al.
Published: (2024)
by: Senen, Kristin, et al.
Published: (2024)
Evaluation of School Libraries in Terms of Quantity and Quality
by: Di?lekçi?, Ati?lla
Published: (2022)
by: Di?lekçi?, Ati?lla
Published: (2022)
Rethinking Information Synthesis in Multimodal Question Answering A Multi-Agent Perspective
by: Rajput, Krishna Singh, et al.
Published: (2025)
by: Rajput, Krishna Singh, et al.
Published: (2025)
Insights into Alignment: Evaluating DPO and its Variants Across Multiple Tasks
by: Saeidi, Amir, et al.
Published: (2024)
by: Saeidi, Amir, et al.
Published: (2024)
Help Me Identify: Is an LLM+VQA System All We Need to Identify Visual Concepts?
by: Sampat, Shailaja Keyur, et al.
Published: (2024)
by: Sampat, Shailaja Keyur, et al.
Published: (2024)
The value of macroinvertebrate assemblages for determining priorities in wetland rehabilitation: A case study from Lake Toolibin, Western Australia
by: Doupe, R G, et al.
Published: (1995)
by: Doupe, R G, et al.
Published: (1995)
Do Hackers Dream of Electric Teachers?: A Large-Scale, In-Situ Evaluation of Cybersecurity Student Behaviors and Performance with AI Tutors
by: Tompkins, Michael, et al.
Published: (2026)
by: Tompkins, Michael, et al.
Published: (2026)
AcT2I: Evaluating and Improving Action Depiction in Text-to-Image Models
by: Malaviya, Vatsal, et al.
Published: (2025)
by: Malaviya, Vatsal, et al.
Published: (2025)
Similar Items
-
Root-Cause-Driven Automated Vulnerability Repair
by: Wang, Hulin, et al.
Published: (2026) -
Pushan: Trace-Free Deobfuscation of Virtualization-Obfuscated Binaries
by: Sudhir, Ashwin, et al.
Published: (2026) -
When "Competency" in Reasoning Opens the Door to Vulnerability: Jailbreaking LLMs via Novel Complex Ciphers
by: Handa, Divij, et al.
Published: (2024) -
ActionReasoningBench: Reasoning about Actions with and without Ramification Constraints
by: Handa, Divij, et al.
Published: (2024) -
Hypothesis Generation for Materials Discovery and Design Using Goal-Driven and Constraint-Guided LLM Agents
by: Kumbhar, Shrinidhi, et al.
Published: (2025)