Adversarial Robustness in Two-Stage Learning-to-Defer: Algorithms and Guarantees

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Montreuil, Yannis, Carlier, Axel, Ng, Lai Xing, Ooi, Wei Tsang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912552418541568
author Montreuil, Yannis
Carlier, Axel
Ng, Lai Xing
Ooi, Wei Tsang
author_facet Montreuil, Yannis
Carlier, Axel
Ng, Lai Xing
Ooi, Wei Tsang
contents Two-stage Learning-to-Defer (L2D) enables optimal task delegation by assigning each input to either a fixed main model or one of several offline experts, supporting reliable decision-making in complex, multi-agent environments. However, existing L2D frameworks assume clean inputs and are vulnerable to adversarial perturbations that can manipulate query allocation--causing costly misrouting or expert overload. We present the first comprehensive study of adversarial robustness in two-stage L2D systems. We introduce two novel attack strategie--untargeted and targeted--which respectively disrupt optimal allocations or force queries to specific agents. To defend against such threats, we propose SARD, a convex learning algorithm built on a family of surrogate losses that are provably Bayes-consistent and $(\mathcal{R}, \mathcal{G})$-consistent. These guarantees hold across classification, regression, and multi-task settings. Empirical results demonstrate that SARD significantly improves robustness under adversarial attacks while maintaining strong clean performance, marking a critical step toward secure and trustworthy L2D deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2502_01027
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Adversarial Robustness in Two-Stage Learning-to-Defer: Algorithms and Guarantees
Montreuil, Yannis
Carlier, Axel
Ng, Lai Xing
Ooi, Wei Tsang
Machine Learning
Two-stage Learning-to-Defer (L2D) enables optimal task delegation by assigning each input to either a fixed main model or one of several offline experts, supporting reliable decision-making in complex, multi-agent environments. However, existing L2D frameworks assume clean inputs and are vulnerable to adversarial perturbations that can manipulate query allocation--causing costly misrouting or expert overload. We present the first comprehensive study of adversarial robustness in two-stage L2D systems. We introduce two novel attack strategie--untargeted and targeted--which respectively disrupt optimal allocations or force queries to specific agents. To defend against such threats, we propose SARD, a convex learning algorithm built on a family of surrogate losses that are provably Bayes-consistent and $(\mathcal{R}, \mathcal{G})$-consistent. These guarantees hold across classification, regression, and multi-task settings. Empirical results demonstrate that SARD significantly improves robustness under adversarial attacks while maintaining strong clean performance, marking a critical step toward secure and trustworthy L2D deployment.
title Adversarial Robustness in Two-Stage Learning-to-Defer: Algorithms and Guarantees
topic Machine Learning
url https://arxiv.org/abs/2502.01027