Wait, that's not an option: LLMs Robustness with Incorrect Multiple-Choice Options

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Góral, Gracjan, Wiśnios, Emilia, Sankowski, Piotr, Budzianowski, Paweł
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908388350230528
author Góral, Gracjan
Wiśnios, Emilia
Sankowski, Piotr
Budzianowski, Paweł
author_facet Góral, Gracjan
Wiśnios, Emilia
Sankowski, Piotr
Budzianowski, Paweł
contents This work introduces a novel framework for evaluating LLMs' capacity to balance instruction-following with critical reasoning when presented with multiple-choice questions containing no valid answers. Through systematic evaluation across arithmetic, domain-specific knowledge, and high-stakes medical decision tasks, we demonstrate that post-training aligned models often default to selecting invalid options, while base models exhibit improved refusal capabilities that scale with model size. Our analysis reveals that alignment techniques, though intended to enhance helpfulness, can inadvertently impair models' reflective judgment--the ability to override default behaviors when faced with invalid options. We additionally conduct a parallel human study showing similar instruction-following biases, with implications for how these biases may propagate through human feedback datasets used in alignment. We provide extensive ablation studies examining the impact of model size, training techniques, and prompt engineering. Our findings highlight fundamental tensions between alignment optimization and preservation of critical reasoning capabilities, with important implications for developing more robust AI systems for real-world deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2409_00113
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Wait, that's not an option: LLMs Robustness with Incorrect Multiple-Choice Options
Góral, Gracjan
Wiśnios, Emilia
Sankowski, Piotr
Budzianowski, Paweł
Computation and Language
Artificial Intelligence
This work introduces a novel framework for evaluating LLMs' capacity to balance instruction-following with critical reasoning when presented with multiple-choice questions containing no valid answers. Through systematic evaluation across arithmetic, domain-specific knowledge, and high-stakes medical decision tasks, we demonstrate that post-training aligned models often default to selecting invalid options, while base models exhibit improved refusal capabilities that scale with model size. Our analysis reveals that alignment techniques, though intended to enhance helpfulness, can inadvertently impair models' reflective judgment--the ability to override default behaviors when faced with invalid options. We additionally conduct a parallel human study showing similar instruction-following biases, with implications for how these biases may propagate through human feedback datasets used in alignment. We provide extensive ablation studies examining the impact of model size, training techniques, and prompt engineering. Our findings highlight fundamental tensions between alignment optimization and preservation of critical reasoning capabilities, with important implications for developing more robust AI systems for real-world deployment.
title Wait, that's not an option: LLMs Robustness with Incorrect Multiple-Choice Options
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2409.00113