TELL-TALE: Task Efficient LLMs with Task Aware Layer Elimination

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Naim, Omar, Sharma, Krish, Barman, Niyar R, Asher, Nicholas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914550155051008
author Naim, Omar
Sharma, Krish
Barman, Niyar R
Asher, Nicholas
author_facet Naim, Omar
Sharma, Krish
Barman, Niyar R
Asher, Nicholas
contents Large Language Models (LLMs) typically come with a fixed architecture, despite growing evidence that not all layers contribute equally to every downstream task. We introduce TALE (Task-Aware Layer Elimination), an inference-time method that improves task performance by selectively removing layers that are irrelevant or detrimental for a given task. TALE optimizes task-specific performance, yielding a task-optimized architecture without retraining. Across 9 tasks and 5 model families, under both zero-shot and few-shot settings, TALE consistently matches or surpasses baseline performance while simultaneously reducing computational costs. TALE also synergizes with fine-tuning, leading to further performance improvements. Computing TALE for a new task requires modest resources, making it a practical and deployable solution for task-specialized LLM inference.
format Preprint
id arxiv_https___arxiv_org_abs_2510_22767
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TELL-TALE: Task Efficient LLMs with Task Aware Layer Elimination
Naim, Omar
Sharma, Krish
Barman, Niyar R
Asher, Nicholas
Machine Learning
Computation and Language
Large Language Models (LLMs) typically come with a fixed architecture, despite growing evidence that not all layers contribute equally to every downstream task. We introduce TALE (Task-Aware Layer Elimination), an inference-time method that improves task performance by selectively removing layers that are irrelevant or detrimental for a given task. TALE optimizes task-specific performance, yielding a task-optimized architecture without retraining. Across 9 tasks and 5 model families, under both zero-shot and few-shot settings, TALE consistently matches or surpasses baseline performance while simultaneously reducing computational costs. TALE also synergizes with fine-tuning, leading to further performance improvements. Computing TALE for a new task requires modest resources, making it a practical and deployable solution for task-specialized LLM inference.
title TELL-TALE: Task Efficient LLMs with Task Aware Layer Elimination
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2510.22767