The Heap: A Contamination-Free Multilingual Code Dataset for Evaluating Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Katzy, Jonathan, Popescu, Razvan Mihai, van Deursen, Arie, Izadi, Maliheh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Automated Attention Pattern Discovery at Scale in Large Language Models
by: Katzy, Jonathan, et al.
Published: (2026)
by: Katzy, Jonathan, et al.
Published: (2026)
An Exploratory Investigation into Code License Infringements in Large Language Model Training Datasets
by: Katzy, Jonathan, et al.
Published: (2024)
by: Katzy, Jonathan, et al.
Published: (2024)
Language Models for Code Completion: A Practical Evaluation
by: Izadi, Maliheh, et al.
Published: (2024)
by: Izadi, Maliheh, et al.
Published: (2024)
Traces of Memorisation in Large Language Models for Code
by: Al-Kaswan, Ali, et al.
Published: (2023)
by: Al-Kaswan, Ali, et al.
Published: (2023)
A Qualitative Investigation into LLM-Generated Multilingual Code Comments and Automatic Evaluation Metrics
by: Katzy, Jonathan, et al.
Published: (2025)
by: Katzy, Jonathan, et al.
Published: (2025)
A Transformer-Based Approach for Smart Invocation of Automatic Code Completion
by: de Moor, Aral, et al.
Published: (2024)
by: de Moor, Aral, et al.
Published: (2024)
Code Red! On the Harmfulness of Applying Off-the-shelf Large Language Models to Programming Tasks
by: Al-Kaswan, Ali, et al.
Published: (2025)
by: Al-Kaswan, Ali, et al.
Published: (2025)
TreeRanker: Fast and Model-agnostic Ranking System for Code Suggestions in IDEs
by: Cipollone, Daniele, et al.
Published: (2025)
by: Cipollone, Daniele, et al.
Published: (2025)
Model See, Model Do? Exposure-Aware Evaluation of Bug-vs-Fix Preference in Code LLMs
by: Al-Kaswan, Ali, et al.
Published: (2026)
by: Al-Kaswan, Ali, et al.
Published: (2026)
AST-PAC: AST-guided Membership Inference for Code
by: Koohestani, Roham, et al.
Published: (2026)
by: Koohestani, Roham, et al.
Published: (2026)
Data Contamination Quiz: A Tool to Detect and Estimate Contamination in Large Language Models
by: Golchin, Shahriar, et al.
Published: (2023)
by: Golchin, Shahriar, et al.
Published: (2023)
Investigating Autonomous Agent Contributions in the Wild: Activity Patterns and Code Change over Time
by: Popescu, Razvan Mihai, et al.
Published: (2026)
by: Popescu, Razvan Mihai, et al.
Published: (2026)
Evaluating Large Language Models for Functional and Maintainable Code in Industrial Settings: A Case Study at ASML
by: Mundhra, Yash, et al.
Published: (2025)
by: Mundhra, Yash, et al.
Published: (2025)
Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture the Flag Challenges
by: Al-Kaswan, Ali, et al.
Published: (2026)
by: Al-Kaswan, Ali, et al.
Published: (2026)
Time Travel in LLMs: Tracing Data Contamination in Large Language Models
by: Golchin, Shahriar, et al.
Published: (2023)
by: Golchin, Shahriar, et al.
Published: (2023)
Is It Good Data for Multilingual Instruction Tuning or Just Bad Multilingual Evaluation for Large Language Models?
by: Chen, Pinzhen, et al.
Published: (2024)
by: Chen, Pinzhen, et al.
Published: (2024)
SwitchLingua: The First Large-Scale Multilingual and Multi-Ethnic Code-Switching Dataset
by: Xie, Peng, et al.
Published: (2025)
by: Xie, Peng, et al.
Published: (2025)
Evaluation Methodology for Large Language Models for Multilingual Document Question and Answer
by: Kahana, Adar, et al.
Published: (2024)
by: Kahana, Adar, et al.
Published: (2024)
Evaluating the Elementary Multilingual Capabilities of Large Language Models with MultiQ
by: Holtermann, Carolin, et al.
Published: (2024)
by: Holtermann, Carolin, et al.
Published: (2024)
Large Language Model Critics for Execution-Free Evaluation of Code Changes
by: Yadavally, Aashish, et al.
Published: (2025)
by: Yadavally, Aashish, et al.
Published: (2025)
AdEval: Alignment-based Dynamic Evaluation to Mitigate Data Contamination in Large Language Models
by: Fan, Yang
Published: (2025)
by: Fan, Yang
Published: (2025)
MoZIP: A Multilingual Benchmark to Evaluate Large Language Models in Intellectual Property
by: Ni, Shiwen, et al.
Published: (2024)
by: Ni, Shiwen, et al.
Published: (2024)
Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination
by: Chen, Simin, et al.
Published: (2025)
by: Chen, Simin, et al.
Published: (2025)
An Open Source Data Contamination Report for Large Language Models
by: Li, Yucheng, et al.
Published: (2023)
by: Li, Yucheng, et al.
Published: (2023)
Investigating Data Contamination in Modern Benchmarks for Large Language Models
by: Deng, Chunyuan, et al.
Published: (2023)
by: Deng, Chunyuan, et al.
Published: (2023)
A Cross-Lingual Analysis of Bias in Large Language Models Using Romanian History
by: Cocu, Matei-Iulian, et al.
Published: (2025)
by: Cocu, Matei-Iulian, et al.
Published: (2025)
EuropeMedQA Study Protocol: A Multilingual, Multimodal Medical Examination Dataset for Language Model Evaluation
by: Causio, Francesco Andrea, et al.
Published: (2026)
by: Causio, Francesco Andrea, et al.
Published: (2026)
The Roles of English in Evaluating Multilingual Language Models
by: Poelman, Wessel, et al.
Published: (2024)
by: Poelman, Wessel, et al.
Published: (2024)
Multilingual Collaborative Defense for Large Language Models
by: Li, Hongliang, et al.
Published: (2025)
by: Li, Hongliang, et al.
Published: (2025)
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence
by: Choi, Hyeong Kyu, et al.
Published: (2025)
by: Choi, Hyeong Kyu, et al.
Published: (2025)
Evaluating the Limits of Large Language Models in Multilingual Legal Reasoning
by: Ioannou, Antreas, et al.
Published: (2025)
by: Ioannou, Antreas, et al.
Published: (2025)
DynaCode: A Dynamic Complexity-Aware Code Benchmark for Evaluating Large Language Models in Code Generation
by: Hu, Wenhao, et al.
Published: (2025)
by: Hu, Wenhao, et al.
Published: (2025)
Evaluating Multilingual and Code-Switched Alignment in LLMs via Synthetic Natural Language Inference
by: Abdaljalil, Samir, et al.
Published: (2025)
by: Abdaljalil, Samir, et al.
Published: (2025)
Long Code Arena: a Set of Benchmarks for Long-Context Code Models
by: Bogomolov, Egor, et al.
Published: (2024)
by: Bogomolov, Egor, et al.
Published: (2024)
All Languages Matter: On the Multilingual Safety of Large Language Models
by: Wang, Wenxuan, et al.
Published: (2023)
by: Wang, Wenxuan, et al.
Published: (2023)
Multilingual Training and Evaluation Resources for Vision-Language Models
by: Baiamonte, Daniela, et al.
Published: (2026)
by: Baiamonte, Daniela, et al.
Published: (2026)
Evaluating Monolingual and Multilingual Large Language Models for Greek Question Answering: The DemosQA Benchmark
by: Mastrokostas, Charalampos, et al.
Published: (2026)
by: Mastrokostas, Charalampos, et al.
Published: (2026)
Do Large Language Models Have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMs
by: Guo, Yanzhu, et al.
Published: (2024)
by: Guo, Yanzhu, et al.
Published: (2024)
CodeMind: Evaluating Large Language Models for Code Reasoning
by: Liu, Changshu, et al.
Published: (2024)
by: Liu, Changshu, et al.
Published: (2024)
Multilingual Performance Biases of Large Language Models in Education
by: Gupta, Vansh, et al.
Published: (2025)
by: Gupta, Vansh, et al.
Published: (2025)
Similar Items
-
Automated Attention Pattern Discovery at Scale in Large Language Models
by: Katzy, Jonathan, et al.
Published: (2026) -
An Exploratory Investigation into Code License Infringements in Large Language Model Training Datasets
by: Katzy, Jonathan, et al.
Published: (2024) -
Language Models for Code Completion: A Practical Evaluation
by: Izadi, Maliheh, et al.
Published: (2024) -
Traces of Memorisation in Large Language Models for Code
by: Al-Kaswan, Ali, et al.
Published: (2023) -
A Qualitative Investigation into LLM-Generated Multilingual Code Comments and Automatic Evaluation Metrics
by: Katzy, Jonathan, et al.
Published: (2025)