Position: The Pre/Post-Training Boundary Should Govern IP in Industry-Academia ML Collaborations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bergemann, Dirk, Ghili, Soheil, Mekel-Bobrov, Nitzan
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916037037916160
author Bergemann, Dirk
Ghili, Soheil
Mekel-Bobrov, Nitzan
author_facet Bergemann, Dirk
Ghili, Soheil
Mekel-Bobrov, Nitzan
contents Industry-academia ML collaborations routinely fail to launch -- not for scientific reasons, but because academics must publish while companies must protect models trained on proprietary data, and no standard contract framework resolves this tension. Because contracts are negotiated by legal departments alone, many apparent legal disputes are incentive misalignment problems that only scientists at the table can correctly diagnose. We propose PBOS (Protect-the-Business / Open-Source-the-Science), a community-adoptable contract template anchored to a single technically-grounded boundary: pre-training artifacts (architectures, training code, benchmarks, untrained weights) are open science; post-training artifacts (weights trained on proprietary data) are business IP. This boundary is technically meaningful, legally clean, and auditable -- and could not have been drawn correctly without scientists at the negotiating table. We argue the ML community should adopt PBOS as its default contract for such collaborations.
format Preprint
id arxiv_https___arxiv_org_abs_2605_22632
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Position: The Pre/Post-Training Boundary Should Govern IP in Industry-Academia ML Collaborations
Bergemann, Dirk
Ghili, Soheil
Mekel-Bobrov, Nitzan
General Economics
Economics
Industry-academia ML collaborations routinely fail to launch -- not for scientific reasons, but because academics must publish while companies must protect models trained on proprietary data, and no standard contract framework resolves this tension. Because contracts are negotiated by legal departments alone, many apparent legal disputes are incentive misalignment problems that only scientists at the table can correctly diagnose. We propose PBOS (Protect-the-Business / Open-Source-the-Science), a community-adoptable contract template anchored to a single technically-grounded boundary: pre-training artifacts (architectures, training code, benchmarks, untrained weights) are open science; post-training artifacts (weights trained on proprietary data) are business IP. This boundary is technically meaningful, legally clean, and auditable -- and could not have been drawn correctly without scientists at the negotiating table. We argue the ML community should adopt PBOS as its default contract for such collaborations.
title Position: The Pre/Post-Training Boundary Should Govern IP in Industry-Academia ML Collaborations
topic General Economics
Economics
url https://arxiv.org/abs/2605.22632