Learning-to-Defer with Expert-Conditional Advice

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Montreuil, Yannis, Montreuil, Leïna, Carlier, Axel, Ng, Lai Xing, Ooi, Wei Tsang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914615234920448
author Montreuil, Yannis
Montreuil, Leïna
Carlier, Axel
Ng, Lai Xing
Ooi, Wei Tsang
author_facet Montreuil, Yannis
Montreuil, Leïna
Carlier, Axel
Ng, Lai Xing
Ooi, Wei Tsang
contents Learning-to-Defer routes each input to the expert that minimizes expected cost, but it assumes that the information available to every expert is fixed at decision time. Many modern systems violate this assumption: after selecting an expert, one may also choose what additional information that expert should receive, such as retrieved documents, tool outputs, or escalation context. We study this problem and call it Learning-to-Defer with advice. We show that a broad family of natural separated surrogates, which learn routing and advice with distinct heads, is inconsistent even in the smallest non-trivial setting. We then introduce an augmented surrogate that operates on the composite expert--advice action space and prove an $\mathcal{H}$-consistency guarantee together with an excess-risk transfer bound, yielding recovery of the Bayes-optimal policy in the limit. Experiments on tabular, language, and multi-modal tasks show that the resulting method improves over standard Learning-to-Defer while adapting its advice-acquisition behavior to the cost regime; a synthetic benchmark confirms the failure mode predicted for separated surrogates.
format Preprint
id arxiv_https___arxiv_org_abs_2603_14324
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learning-to-Defer with Expert-Conditional Advice
Montreuil, Yannis
Montreuil, Leïna
Carlier, Axel
Ng, Lai Xing
Ooi, Wei Tsang
Machine Learning
Learning-to-Defer routes each input to the expert that minimizes expected cost, but it assumes that the information available to every expert is fixed at decision time. Many modern systems violate this assumption: after selecting an expert, one may also choose what additional information that expert should receive, such as retrieved documents, tool outputs, or escalation context. We study this problem and call it Learning-to-Defer with advice. We show that a broad family of natural separated surrogates, which learn routing and advice with distinct heads, is inconsistent even in the smallest non-trivial setting. We then introduce an augmented surrogate that operates on the composite expert--advice action space and prove an $\mathcal{H}$-consistency guarantee together with an excess-risk transfer bound, yielding recovery of the Bayes-optimal policy in the limit. Experiments on tabular, language, and multi-modal tasks show that the resulting method improves over standard Learning-to-Defer while adapting its advice-acquisition behavior to the cost regime; a synthetic benchmark confirms the failure mode predicted for separated surrogates.
title Learning-to-Defer with Expert-Conditional Advice
topic Machine Learning
url https://arxiv.org/abs/2603.14324