TRIM: Token-wise Attention-Derived Saliency for Data-Efficient Instruction Tuning
Fuente:
arXiv
Saved in:
| Main Authors: | Nagaraj, Manish, Choudhary, Sakshi, Saxena, Utkarsh, Ravikumar, Deepak, Roy, Kaushik |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Eigen Attention: Attention in Low-Rank Space for KV Cache Compression
by: Saxena, Utkarsh, et al.
Published: (2024)
by: Saxena, Utkarsh, et al.
Published: (2024)
Coresets from Trajectories: Selecting Data via Correlation of Loss Differences
by: Nagaraj, Manish, et al.
Published: (2025)
by: Nagaraj, Manish, et al.
Published: (2025)
OASIS: Online Activation Subspace Learning for Memory-Efficient Training
by: Choudhary, Sakshi, et al.
Published: (2026)
by: Choudhary, Sakshi, et al.
Published: (2026)
Finding the Muses: Identifying Coresets through Loss Trajectories
by: Nagaraj, Manish, et al.
Published: (2025)
by: Nagaraj, Manish, et al.
Published: (2025)
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals
by: Saxena, Utkarsh, et al.
Published: (2024)
by: Saxena, Utkarsh, et al.
Published: (2024)
TraceNAS: Zero-shot LLM Pruning via Gradient Trace Correlation
by: Malettira, Prajna G., et al.
Published: (2026)
by: Malettira, Prajna G., et al.
Published: (2026)
GradientSpace: Unsupervised Data Clustering for Improved Instruction Tuning
by: Sridharan, Shrihari, et al.
Published: (2025)
by: Sridharan, Shrihari, et al.
Published: (2025)
Averaging Rate Scheduler for Decentralized Learning on Heterogeneous Data
by: Aketi, Sai Aparna, et al.
Published: (2024)
by: Aketi, Sai Aparna, et al.
Published: (2024)
KVLinC : KV Cache Quantization with Hadamard Rotation and Linear Correction
by: Saxena, Utkarsh, et al.
Published: (2025)
by: Saxena, Utkarsh, et al.
Published: (2025)
TOFU: Towards Obfuscated Federated Updates by Encoding Weight Updates into Gradients from Proxy Data
by: Garg, Isha, et al.
Published: (2022)
by: Garg, Isha, et al.
Published: (2022)
SADDLe: Sharpness-Aware Decentralized Deep Learning with Heterogeneous Data
by: Choudhary, Sakshi, et al.
Published: (2024)
by: Choudhary, Sakshi, et al.
Published: (2024)
Advancing Compressed Video Action Recognition through Progressive Knowledge Distillation
by: Soufleri, Efstathia, et al.
Published: (2024)
by: Soufleri, Efstathia, et al.
Published: (2024)
Mixture of Weight-shared Heterogeneous Group Attention Experts for Dynamic Token-wise KV Optimization
by: Song, Guanghui, et al.
Published: (2025)
by: Song, Guanghui, et al.
Published: (2025)
Beyond Higher Rank: Token-wise Input-Output Projections for Efficient Low-Rank Adaptation
by: Li, Shiwei, et al.
Published: (2025)
by: Li, Shiwei, et al.
Published: (2025)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
by: Jo, Dongwon, et al.
Published: (2026)
by: Jo, Dongwon, et al.
Published: (2026)
Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
by: Alwis, Praditha, et al.
Published: (2026)
by: Alwis, Praditha, et al.
Published: (2026)
Data Diversity Matters for Robust Instruction Tuning
by: Bukharin, Alexander, et al.
Published: (2023)
by: Bukharin, Alexander, et al.
Published: (2023)
STS: Efficient Sparse Attention with Speculative Token Sparsity
by: Xu, Ceyu, et al.
Published: (2026)
by: Xu, Ceyu, et al.
Published: (2026)
Federated Data-Efficient Instruction Tuning for Large Language Models
by: Qin, Zhen, et al.
Published: (2024)
by: Qin, Zhen, et al.
Published: (2024)
Mortgage Language Model: Domain-Adaptive Pretraining with Residual Instruction, Alignment Tuning, and Task-Specific Routing
by: Jain, Manish, et al.
Published: (2025)
by: Jain, Manish, et al.
Published: (2025)
Memory-Efficient Fine-Tuning of Transformers via Token Selection
by: Simoulin, Antoine, et al.
Published: (2025)
by: Simoulin, Antoine, et al.
Published: (2025)
TopoPrune: Robust Data Pruning via Unified Latent Space Topology
by: Roy, Arjun, et al.
Published: (2026)
by: Roy, Arjun, et al.
Published: (2026)
Energy-Gated Attention: Spectral Salience as an Inductive Bias for Transformer Attention
by: Zeris, Athanasios
Published: (2026)
by: Zeris, Athanasios
Published: (2026)
The Easy Path to Robustness: Coreset Selection using Sample Hardness
by: Ramesh, Pranav, et al.
Published: (2025)
by: Ramesh, Pranav, et al.
Published: (2025)
Curvature Clues: Decoding Deep Learning Privacy with Input Loss Curvature
by: Ravikumar, Deepak, et al.
Published: (2024)
by: Ravikumar, Deepak, et al.
Published: (2024)
CODE-CL: Conceptor-Based Gradient Projection for Deep Continual Learning
by: Apolinario, Marco Paul E., et al.
Published: (2024)
by: Apolinario, Marco Paul E., et al.
Published: (2024)
Saliency Attention and Semantic Similarity-Driven Adversarial Perturbation
by: Waghela, Hetvi, et al.
Published: (2024)
by: Waghela, Hetvi, et al.
Published: (2024)
Homogenizing Non-IID datasets via In-Distribution Knowledge Distillation for Decentralized Learning
by: Ravikumar, Deepak, et al.
Published: (2023)
by: Ravikumar, Deepak, et al.
Published: (2023)
Parameter Efficient Instruction Tuning: An Empirical Study
by: He, Pengfei
Published: (2024)
by: He, Pengfei
Published: (2024)
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
by: Zarch, Hossein Entezari, et al.
Published: (2025)
by: Zarch, Hossein Entezari, et al.
Published: (2025)
Token Distillation: Attention-aware Input Embeddings For New Tokens
by: Dobler, Konstantin, et al.
Published: (2025)
by: Dobler, Konstantin, et al.
Published: (2025)
Instruction Mining: Instruction Data Selection for Tuning Large Language Models
by: Cao, Yihan, et al.
Published: (2023)
by: Cao, Yihan, et al.
Published: (2023)
The Best Instruction-Tuning Data are Those That Fit
by: Zhang, Dylan, et al.
Published: (2025)
by: Zhang, Dylan, et al.
Published: (2025)
Bridging the Dimensional Chasm: Uncover Layer-wise Dimensional Reduction in Transformers through Token Correlation
by: Song, Zhuo-Yang, et al.
Published: (2025)
by: Song, Zhuo-Yang, et al.
Published: (2025)
How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
Multiplex Thinking: Reasoning via Token-wise Branch-and-Merge
by: Tang, Yao, et al.
Published: (2026)
by: Tang, Yao, et al.
Published: (2026)
Learning When to Attend: Conditional Memory Access for Long-Context LLMs
by: Choudhary, Sakshi, et al.
Published: (2026)
by: Choudhary, Sakshi, et al.
Published: (2026)
Prompt-Based Bias Calibration for Better Zero/Few-Shot Learning of Language Models
by: He, Kang, et al.
Published: (2024)
by: He, Kang, et al.
Published: (2024)
Selective Reflection-Tuning: Student-Selected Data Recycling for LLM Instruction-Tuning
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
SMART: Submodular Data Mixture Strategy for Instruction Tuning
by: Renduchintala, H S V N S Kowndinya, et al.
Published: (2024)
by: Renduchintala, H S V N S Kowndinya, et al.
Published: (2024)
Similar Items
-
Eigen Attention: Attention in Low-Rank Space for KV Cache Compression
by: Saxena, Utkarsh, et al.
Published: (2024) -
Coresets from Trajectories: Selecting Data via Correlation of Loss Differences
by: Nagaraj, Manish, et al.
Published: (2025) -
OASIS: Online Activation Subspace Learning for Memory-Efficient Training
by: Choudhary, Sakshi, et al.
Published: (2026) -
Finding the Muses: Identifying Coresets through Loss Trajectories
by: Nagaraj, Manish, et al.
Published: (2025) -
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals
by: Saxena, Utkarsh, et al.
Published: (2024)