Saved in:
Bibliographic Details
Main Authors: Kadasi, Pritam, Upperwal, Abhishek, Singh, Mayank
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.03103
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918320919281664
author Kadasi, Pritam
Upperwal, Abhishek
Singh, Mayank
author_facet Kadasi, Pritam
Upperwal, Abhishek
Singh, Mayank
contents Instruction tuning is now the default way to train and adapt large language models, but many instruction--input--output pairs are only weakly specified: for a given input, the same output can remain plausible under several alternative instructions. This raises a simple question: \emph{does the instruction uniquely determine the target output?} We propose the \textbf{Task--Specificity Score (TSS)} to quantify how much an instruction matters for predicting its output, by contrasting the true instruction against plausible alternatives for the same input. We further introduce \textbf{TSS++}, which uses hard alternatives and a small quality term to mitigate easy-negative effects. Across three instruction datasets (\textsc{Alpaca}, \textsc{Dolly-15k}, \textsc{NI-20}) and three open LLMs (Gemma, Llama, Qwen), we show that selecting task-specific examples improves downstream performance under tight token budgets and complements quality-based filters such as perplexity and IFD.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03103
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Task--Specificity Score: Measuring How Much Instructions Really Matter for Supervision
Kadasi, Pritam
Upperwal, Abhishek
Singh, Mayank
Computation and Language
Artificial Intelligence
Instruction tuning is now the default way to train and adapt large language models, but many instruction--input--output pairs are only weakly specified: for a given input, the same output can remain plausible under several alternative instructions. This raises a simple question: \emph{does the instruction uniquely determine the target output?} We propose the \textbf{Task--Specificity Score (TSS)} to quantify how much an instruction matters for predicting its output, by contrasting the true instruction against plausible alternatives for the same input. We further introduce \textbf{TSS++}, which uses hard alternatives and a small quality term to mitigate easy-negative effects. Across three instruction datasets (\textsc{Alpaca}, \textsc{Dolly-15k}, \textsc{NI-20}) and three open LLMs (Gemma, Llama, Qwen), we show that selecting task-specific examples improves downstream performance under tight token budgets and complements quality-based filters such as perplexity and IFD.
title Task--Specificity Score: Measuring How Much Instructions Really Matter for Supervision
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2602.03103