Language model developers should report train-test overlap
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Andy K, Klyman, Kevin, Mai, Yifan, Levine, Yoav, Zhang, Yian, Bommasani, Rishi, Liang, Percy |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Robustness tests for biomedical foundation models should tailor to specifications
by: Xian, R. Patrick, et al.
Published: (2025)
by: Xian, R. Patrick, et al.
Published: (2025)
Reliable agent engineering should integrate machine-compatible organizational principles
by: Xian, R. Patrick, et al.
Published: (2025)
by: Xian, R. Patrick, et al.
Published: (2025)
Do AI Companies Make Good on Voluntary Commitments to the White House?
by: Wang, Jennifer, et al.
Published: (2025)
by: Wang, Jennifer, et al.
Published: (2025)
Ten simple rules for training scientists to make better software
by: Gallagher, Kit, et al.
Published: (2024)
by: Gallagher, Kit, et al.
Published: (2024)
Overwhelmed software developers: An Interpretative Phenomenological Analysis
by: Michels, Lisa-Marie, et al.
Published: (2024)
by: Michels, Lisa-Marie, et al.
Published: (2024)
Innovating the software engineering class through multi-team development
by: Brockenbrough, Allan
Published: (2025)
by: Brockenbrough, Allan
Published: (2025)
The 2024 Foundation Model Transparency Index
by: Bommasani, Rishi, et al.
Published: (2024)
by: Bommasani, Rishi, et al.
Published: (2024)
Thoughts on Learning Human and Programming Languages
by: Katz, Daniel S., et al.
Published: (2024)
by: Katz, Daniel S., et al.
Published: (2024)
Large Language Models are overconfident and amplify human bias
by: Sun, Fengfei, et al.
Published: (2025)
by: Sun, Fengfei, et al.
Published: (2025)
Context-Specific Instruction: A Longitudinal Study on Debugging Skill Acquisition and Retention for Novice Programmers
by: Zhang, Ziyi, et al.
Published: (2025)
by: Zhang, Ziyi, et al.
Published: (2025)
Open, Small, Rigmarole -- Evaluating Llama 3.2 3B's Feedback for Programming Exercises
by: Azaiz, Imen, et al.
Published: (2025)
by: Azaiz, Imen, et al.
Published: (2025)
The First Issue Matters: Linking Task-Level Characteristics to Long-Term Newcomer Retention in OSS
by: Hao, Yichen, et al.
Published: (2026)
by: Hao, Yichen, et al.
Published: (2026)
A First Look at the General Data Protection Regulation (GDPR) in Open-Source Software
by: Franke, Lucas, et al.
Published: (2024)
by: Franke, Lucas, et al.
Published: (2024)
Reusable MLOps: Reusable Deployment, Reusable Infrastructure and Hot-Swappable Machine Learning models and services
by: Panchal, D, et al.
Published: (2024)
by: Panchal, D, et al.
Published: (2024)
"Write in English, Nobody Understands Your Language Here": A Study of Non-English Trends in Open-Source Repositories
by: Bhuiyan, Masudul Hasan Masud, et al.
Published: (2026)
by: Bhuiyan, Masudul Hasan Masud, et al.
Published: (2026)
Directional Diffusion-Style Code Editing Pre-training
by: Liang, Qingyuan, et al.
Published: (2025)
by: Liang, Qingyuan, et al.
Published: (2025)
FastFixer: An Efficient and Effective Approach for Repairing Programming Assignments
by: Liu, Fang, et al.
Published: (2024)
by: Liu, Fang, et al.
Published: (2024)
CodeGrad: Integrating Multi-Step Verification with Gradient-Based LLM Refinement
by: Zhang, Yueke, et al.
Published: (2025)
by: Zhang, Yueke, et al.
Published: (2025)
An instrument to measure factors that constitute the socio-technical context of testing experience
by: Swillus, Mark, et al.
Published: (2025)
by: Swillus, Mark, et al.
Published: (2025)
Foundation Model Transparency Reports
by: Bommasani, Rishi, et al.
Published: (2024)
by: Bommasani, Rishi, et al.
Published: (2024)
The 2025 Foundation Model Transparency Index
by: Wan, Alexander, et al.
Published: (2025)
by: Wan, Alexander, et al.
Published: (2025)
The role of slicing in test-driven development
by: Dieste, Oscar, et al.
Published: (2024)
by: Dieste, Oscar, et al.
Published: (2024)
DPO-F+: Aligning Code Repair Feedback with Developers' Preferences
by: Fang, Zihan, et al.
Published: (2025)
by: Fang, Zihan, et al.
Published: (2025)
Coverage measurement in model-based testing of web applications: Tool support and an industrial experience report
by: Garousi, Vahid, et al.
Published: (2024)
by: Garousi, Vahid, et al.
Published: (2024)
Rapid Mobile App Development for Generative AI Agents on MIT App Inventor
by: Gao, Jaida, et al.
Published: (2024)
by: Gao, Jaida, et al.
Published: (2024)
Practical Program Repair in the Era of Large Pre-trained Language Models
by: Xia, Chunqiu Steven, et al.
Published: (2022)
by: Xia, Chunqiu Steven, et al.
Published: (2022)
Example-driven development: bridging tests and documentation
by: Nierstrasz, Oscar, et al.
Published: (2024)
by: Nierstrasz, Oscar, et al.
Published: (2024)
LLM-CompDroid: Repairing Configuration Compatibility Bugs in Android Apps with Pre-trained Large Language Models
by: Liu, Zhijie, et al.
Published: (2024)
by: Liu, Zhijie, et al.
Published: (2024)
LLM Use, Cheating, and Academic Integrity in Software Engineering Education
by: Santos, Ronnie de Souza, et al.
Published: (2026)
by: Santos, Ronnie de Souza, et al.
Published: (2026)
A Scenario Analysis of Ethical Issues in Dark Patterns and Their Research
by: Ruohonen, Jukka, et al.
Published: (2025)
by: Ruohonen, Jukka, et al.
Published: (2025)
Towards a Knowledge Base of Common Sustainability Weaknesses in Green Software Development
by: Pathania, Priyavanshi, et al.
Published: (2025)
by: Pathania, Priyavanshi, et al.
Published: (2025)
From Generation to Adaptation: Comparing AI-Assisted Strategies in High School Programming Education
by: Hu, Tong, et al.
Published: (2025)
by: Hu, Tong, et al.
Published: (2025)
Overcoming Obstacles: Challenges of Gender Inequality in Undergraduate ICT Programs
by: Souza, Angelica Pereira, et al.
Published: (2025)
by: Souza, Angelica Pereira, et al.
Published: (2025)
The State of Computational Science in Fission and Fusion Energy
by: Coto, Andrea Morales, et al.
Published: (2025)
by: Coto, Andrea Morales, et al.
Published: (2025)
Understanding: reframing automation and assurance
by: Bloomfield, Robin
Published: (2026)
by: Bloomfield, Robin
Published: (2026)
Bringing AI into the Classroom: A Structured Approach for Integrating AI into Software Engineering Education
by: Groher, Iris, et al.
Published: (2026)
by: Groher, Iris, et al.
Published: (2026)
Evaluation of Systems Programming Exercises through Tailored Static Analysis
by: Natella, Roberto
Published: (2024)
by: Natella, Roberto
Published: (2024)
How frontier AI companies could implement an internal audit function
by: Gomez, Francesca, et al.
Published: (2025)
by: Gomez, Francesca, et al.
Published: (2025)
From Pre-labeling to Production: Engineering Lessons from a Machine Learning Pipeline in the Public Sector
by: Ferreira, Ronivaldo, et al.
Published: (2025)
by: Ferreira, Ronivaldo, et al.
Published: (2025)
The Need for a Green ICT Reference Framework
by: Aiello, Marco, et al.
Published: (2026)
by: Aiello, Marco, et al.
Published: (2026)
Similar Items
-
Robustness tests for biomedical foundation models should tailor to specifications
by: Xian, R. Patrick, et al.
Published: (2025) -
Reliable agent engineering should integrate machine-compatible organizational principles
by: Xian, R. Patrick, et al.
Published: (2025) -
Do AI Companies Make Good on Voluntary Commitments to the White House?
by: Wang, Jennifer, et al.
Published: (2025) -
Ten simple rules for training scientists to make better software
by: Gallagher, Kit, et al.
Published: (2024) -
Overwhelmed software developers: An Interpretative Phenomenological Analysis
by: Michels, Lisa-Marie, et al.
Published: (2024)