Understanding Performance Collapse in Layer-Pruned Large Language Models via Decision Representation Transitions

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shi, Boyu, Liu, Chang, Gao, ChuanBao, Yang, Xu, Geng, Xin
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910200591548416
author Shi, Boyu
Liu, Chang
Gao, ChuanBao
Yang, Xu
Geng, Xin
author_facet Shi, Boyu
Liu, Chang
Gao, ChuanBao
Yang, Xu
Geng, Xin
contents Layer pruning efficiently reduces Large Language Model (LLM) computational costs but often triggers sudden performance collapse. Existing representation-based analyses struggle to explain this mechanism. We propose studying pruning through decision representation. Focusing on multiple-choice tasks, we introduce two metrics, Decision Margin and Option Frequency, and an Iterative Pruning method to analyze layer-wise decision dynamics. Our findings reveal a sharp decision transition that partitions the network into two stages: a Silent Phase, where the model cannot yet predict the correct answer, and a Decisive Phase, where the correct prediction emerges. We also find that pruning the Decisive Phase has minimal impact, whereas pruning the Silent Phase triggers immediate performance collapse, highlighting its extreme sensitivity to structural changes. Therefore, we conclude that pruning-induced collapse stems from disrupting the Silent Phase, which prevents the critical decision transition from occurring.
format Preprint
id arxiv_https___arxiv_org_abs_2605_07271
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Understanding Performance Collapse in Layer-Pruned Large Language Models via Decision Representation Transitions
Shi, Boyu
Liu, Chang
Gao, ChuanBao
Yang, Xu
Geng, Xin
Computation and Language
Artificial Intelligence
Layer pruning efficiently reduces Large Language Model (LLM) computational costs but often triggers sudden performance collapse. Existing representation-based analyses struggle to explain this mechanism. We propose studying pruning through decision representation. Focusing on multiple-choice tasks, we introduce two metrics, Decision Margin and Option Frequency, and an Iterative Pruning method to analyze layer-wise decision dynamics. Our findings reveal a sharp decision transition that partitions the network into two stages: a Silent Phase, where the model cannot yet predict the correct answer, and a Decisive Phase, where the correct prediction emerges. We also find that pruning the Decisive Phase has minimal impact, whereas pruning the Silent Phase triggers immediate performance collapse, highlighting its extreme sensitivity to structural changes. Therefore, we conclude that pruning-induced collapse stems from disrupting the Silent Phase, which prevents the critical decision transition from occurring.
title Understanding Performance Collapse in Layer-Pruned Large Language Models via Decision Representation Transitions
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2605.07271