Identifying and Evaluating Inactive Heads in Pretrained LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Sandoval-Segura, Pedro, Wang, Xijun, Panda, Ashwinee, Goldblum, Micah, Basri, Ronen, Goldstein, Tom, Jacobs, David |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Token Prediction via Self-Distillation
by: Kirchenbauer, John, et al.
Published: (2026)
by: Kirchenbauer, John, et al.
Published: (2026)
Gemstones: A Model Suite for Multi-Faceted Scaling Laws
by: McLeish, Sean, et al.
Published: (2025)
by: McLeish, Sean, et al.
Published: (2025)
Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models
by: Jain, Neel, et al.
Published: (2024)
by: Jain, Neel, et al.
Published: (2024)
LoRI: Reducing Cross-Task Interference in Multi-Task Low-Rank Adaptation
by: Zhang, Juzheng, et al.
Published: (2025)
by: Zhang, Juzheng, et al.
Published: (2025)
FineGRAIN: Evaluating Failure Modes of Text-to-Image Models with Vision Language Model Judges
by: Hayes, Kevin David, et al.
Published: (2025)
by: Hayes, Kevin David, et al.
Published: (2025)
CinePile: A Long Video Question Answering Dataset and Benchmark
by: Rawal, Ruchit, et al.
Published: (2024)
by: Rawal, Ruchit, et al.
Published: (2024)
Speculating Experts Accelerates Inference for Mixture-of-Experts
by: Madan, Vivan, et al.
Published: (2026)
by: Madan, Vivan, et al.
Published: (2026)
Commercial LLM Agents Are Already Vulnerable to Simple Yet Dangerous Attacks
by: Li, Ang, et al.
Published: (2025)
by: Li, Ang, et al.
Published: (2025)
Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text
by: Hans, Abhimanyu, et al.
Published: (2024)
by: Hans, Abhimanyu, et al.
Published: (2024)
When Is Compositional Reasoning Learnable from Verifiable Rewards?
by: Barzilai, Daniel, et al.
Published: (2026)
by: Barzilai, Daniel, et al.
Published: (2026)
Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
by: McLeish, Sean, et al.
Published: (2025)
by: McLeish, Sean, et al.
Published: (2025)
vTune: Verifiable Fine-Tuning for LLMs Through Backdooring
by: Zhang, Eva, et al.
Published: (2024)
by: Zhang, Eva, et al.
Published: (2024)
Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
by: Panda, Ashwinee, et al.
Published: (2025)
by: Panda, Ashwinee, et al.
Published: (2025)
Likelihood Training of Cascaded Diffusion Models via Hierarchical Volume-preserving Maps
by: Li, Henry, et al.
Published: (2025)
by: Li, Henry, et al.
Published: (2025)
Privacy-Preserving Mechanisms Enable Cheap Verifiable Inference of LLMs
by: Pal, Arka, et al.
Published: (2026)
by: Pal, Arka, et al.
Published: (2026)
Controlling the Inductive Bias of Wide Neural Networks by Modifying the Kernel's Spectrum
by: Geifman, Amnon, et al.
Published: (2023)
by: Geifman, Amnon, et al.
Published: (2023)
Alignment-Constrained Dynamic Pruning for LLMs: Identifying and Preserving Alignment-Critical Circuits
by: Patel, Dev, et al.
Published: (2025)
by: Patel, Dev, et al.
Published: (2025)
Dynamic Delayed Tree Expansion For Improved Multi-Path Speculative Decoding
by: Thomas, Rahul, et al.
Published: (2026)
by: Thomas, Rahul, et al.
Published: (2026)
Measuring Style Similarity in Diffusion Models
by: Somepalli, Gowthami, et al.
Published: (2024)
by: Somepalli, Gowthami, et al.
Published: (2024)
An Attack to Break Permutation-Based Private Third-Party Inference Schemes for LLMs
by: Thomas, Rahul, et al.
Published: (2025)
by: Thomas, Rahul, et al.
Published: (2025)
The No Free Lunch Theorem, Kolmogorov Complexity, and the Role of Inductive Biases in Machine Learning
by: Goldblum, Micah, et al.
Published: (2023)
by: Goldblum, Micah, et al.
Published: (2023)
A Simple Baseline for Predicting Events with Auto-Regressive Tabular Transformers
by: Stein, Alex, et al.
Published: (2024)
by: Stein, Alex, et al.
Published: (2024)
Knowing What You Know Is Not Enough: Large Language Model Confidences Don't Align With Their Actions
by: Pal, Arka, et al.
Published: (2025)
by: Pal, Arka, et al.
Published: (2025)
On Training in Imagination
by: Timor, Nadav, et al.
Published: (2026)
by: Timor, Nadav, et al.
Published: (2026)
Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful
by: Marek, Martin, et al.
Published: (2025)
by: Marek, Martin, et al.
Published: (2025)
Compute Better Spent: Replacing Dense Layers with Structured Matrices
by: Qiu, Shikai, et al.
Published: (2024)
by: Qiu, Shikai, et al.
Published: (2024)
On the Reliability of Watermarks for Large Language Models
by: Kirchenbauer, John, et al.
Published: (2023)
by: Kirchenbauer, John, et al.
Published: (2023)
Querying Kernel Methods Suffices for Reconstructing their Training Data
by: Barzilai, Daniel, et al.
Published: (2025)
by: Barzilai, Daniel, et al.
Published: (2025)
SpectralNet: Spectral Clustering using Deep Neural Networks
by: Shaham, Uri, et al.
Published: (2018)
by: Shaham, Uri, et al.
Published: (2018)
Adaptive Retention & Correction: Test-Time Training for Continual Learning
by: Chen, Haoran, et al.
Published: (2024)
by: Chen, Haoran, et al.
Published: (2024)
A Capacity-Based Rationale for Multi-Head Attention
by: Adler, Micah
Published: (2025)
by: Adler, Micah
Published: (2025)
Teach LLMs to Phish: Stealing Private Information from Language Models
by: Panda, Ashwinee, et al.
Published: (2024)
by: Panda, Ashwinee, et al.
Published: (2024)
DynaGuard: A Dynamic Guardian Model With User-Defined Policies
by: Hoover, Monte, et al.
Published: (2025)
by: Hoover, Monte, et al.
Published: (2025)
Amortized Latent Steering: Low-Cost Alternative to Test-Time Optimization
by: Egbuna, Nathan, et al.
Published: (2025)
by: Egbuna, Nathan, et al.
Published: (2025)
The Lie Derivative for Measuring Learned Equivariance
by: Gruver, Nate, et al.
Published: (2022)
by: Gruver, Nate, et al.
Published: (2022)
Unlocking Tokens as Data Points for Generalization Bounds on Larger Language Models
by: Lotfi, Sanae, et al.
Published: (2024)
by: Lotfi, Sanae, et al.
Published: (2024)
Transformers Boost the Performance of Decision Trees on Tabular Data across Sample Sizes
by: Jayawardhana, Mayuka, et al.
Published: (2025)
by: Jayawardhana, Mayuka, et al.
Published: (2025)
Generating Potent Poisons and Backdoors from Scratch with Guided Diffusion
by: Souri, Hossein, et al.
Published: (2024)
by: Souri, Hossein, et al.
Published: (2024)
Cascade: Token-Sharded Private LLM Inference
by: Thomas, Rahul, et al.
Published: (2025)
by: Thomas, Rahul, et al.
Published: (2025)
A New Linear Scaling Rule for Private Adaptive Hyperparameter Optimization
by: Panda, Ashwinee, et al.
Published: (2022)
by: Panda, Ashwinee, et al.
Published: (2022)
Similar Items
-
Multi-Token Prediction via Self-Distillation
by: Kirchenbauer, John, et al.
Published: (2026) -
Gemstones: A Model Suite for Multi-Faceted Scaling Laws
by: McLeish, Sean, et al.
Published: (2025) -
Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models
by: Jain, Neel, et al.
Published: (2024) -
LoRI: Reducing Cross-Task Interference in Multi-Task Low-Rank Adaptation
by: Zhang, Juzheng, et al.
Published: (2025) -
FineGRAIN: Evaluating Failure Modes of Text-to-Image Models with Vision Language Model Judges
by: Hayes, Kevin David, et al.
Published: (2025)