Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Salfati, Samuel |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ACE: Exploring Activation Cosine Similarity and Variance for Accurate and Calibration-Efficient LLM Pruning
par: Mi, Zhendong, et autres
Publié: (2025)
par: Mi, Zhendong, et autres
Publié: (2025)
Transformer Scalability Crisis: The First Comprehensive Empirical Analysis of Performance Walls in Modern Language Models
par: Moghadasi, Mahdi Naser, et autres
Publié: (2026)
par: Moghadasi, Mahdi Naser, et autres
Publié: (2026)
TensorLens: End-to-End Transformer Analysis via High-Order Attention Tensors
par: Atad, Ido Andrew, et autres
Publié: (2026)
par: Atad, Ido Andrew, et autres
Publié: (2026)
Quantization Undoes Alignment: Bias Emergence in Compressed LLMs Across Models and Precision Levels
par: Rath, Plawan Kumar, et autres
Publié: (2026)
par: Rath, Plawan Kumar, et autres
Publié: (2026)
On the Limits of Learned Importance Scoring for KV Cache Compression
par: Steele, Brady
Publié: (2026)
par: Steele, Brady
Publié: (2026)
The Geometry of Thought: How Scale Restructures Reasoning In Large Language Models
par: Anderson, Samuel Cyrenius
Publié: (2026)
par: Anderson, Samuel Cyrenius
Publié: (2026)
Revisiting LRP: Positional Attribution as the Missing Ingredient for Transformer Explainability
par: Bakish, Yarden, et autres
Publié: (2025)
par: Bakish, Yarden, et autres
Publié: (2025)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
par: Fadli, Samih
Publié: (2025)
par: Fadli, Samih
Publié: (2025)
Merge-Bench: Resolve Merge Conflicts with Large Language Models
par: Schesch, Benedikt, et autres
Publié: (2026)
par: Schesch, Benedikt, et autres
Publié: (2026)
Neural Activation Patterns Across Language Model Architectures: A Comprehensive Analysis of Cognitive Task Performance
par: Naser-Moghadasi, Mahdi, et autres
Publié: (2026)
par: Naser-Moghadasi, Mahdi, et autres
Publié: (2026)
Learned Relay Representations for Forward-Thinking Discrete Diffusion Models
par: Rozonoyer, Benjamin, et autres
Publié: (2026)
par: Rozonoyer, Benjamin, et autres
Publié: (2026)
Stratified Hazard Sampling: Minimal-Variance Event Scheduling for CTMC/DTMC Discrete Diffusion and Flow Models
par: Jang, Seunghwan, et autres
Publié: (2026)
par: Jang, Seunghwan, et autres
Publié: (2026)
Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring
par: Heyman, Alex, et autres
Publié: (2025)
par: Heyman, Alex, et autres
Publié: (2025)
Memory Bank Compression for Continual Adaptation of Large Language Models
par: Katraouras, Thomas, et autres
Publié: (2026)
par: Katraouras, Thomas, et autres
Publié: (2026)
When Models Can't Follow: Testing Instruction Adherence Across 256 LLMs
par: Young, Richard J., et autres
Publié: (2025)
par: Young, Richard J., et autres
Publié: (2025)
Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs
par: Easley, Eric, et autres
Publié: (2026)
par: Easley, Eric, et autres
Publié: (2026)
Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features
par: McCann, Jordan F.
Publié: (2026)
par: McCann, Jordan F.
Publié: (2026)
Super Apriel: One Checkpoint, Many Speeds
par: Labs, SLAM, et autres
Publié: (2026)
par: Labs, SLAM, et autres
Publié: (2026)
Why LoRA Resists Label Noise: A Theoretical Framework for Noise-Robust Parameter-Efficient Fine-Tuning
par: Steele, Brady
Publié: (2026)
par: Steele, Brady
Publié: (2026)
Scalable GPU-Accelerated Euler Characteristic Curves: Optimization and Differentiable Learning for PyTorch
par: Saxena, Udit
Publié: (2025)
par: Saxena, Udit
Publié: (2025)
Synergy over Discrepancy: A Partition-Based Approach to Multi-Domain LLM Fine-Tuning
par: Ye, Hua, et autres
Publié: (2025)
par: Ye, Hua, et autres
Publié: (2025)
KerZOO: Kernel Function Informed Zeroth-Order Optimization for Accurate and Accelerated LLM Fine-Tuning
par: Mi, Zhendong, et autres
Publié: (2025)
par: Mi, Zhendong, et autres
Publié: (2025)
QuAnTS: Question Answering on Time Series
par: Divo, Felix, et autres
Publié: (2025)
par: Divo, Felix, et autres
Publié: (2025)
LLM Vocabulary Compression for Low-Compute Environments
par: Vennam, Sreeram, et autres
Publié: (2024)
par: Vennam, Sreeram, et autres
Publié: (2024)
Latent Cache Flow: Model-to-Model Communication Without Text
par: Rossi, Maximillian, et autres
Publié: (2026)
par: Rossi, Maximillian, et autres
Publié: (2026)
Recurrent Memory-Augmented Transformers with Chunked Attention for Long-Context Language Modeling
par: Kashyap, Ankit
Publié: (2025)
par: Kashyap, Ankit
Publié: (2025)
The Data Efficiency Frontier of Financial Foundation Models: Scaling Laws from Continued Pretraining
par: Ponnock, Jesse
Publié: (2025)
par: Ponnock, Jesse
Publié: (2025)
Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models
par: Shravan, Rohan
Publié: (2026)
par: Shravan, Rohan
Publié: (2026)
Closing the Curvature Gap: Full Transformer Hessians and Their Implications for Scaling Laws
par: Petrov, Egor, et autres
Publié: (2025)
par: Petrov, Egor, et autres
Publié: (2025)
Continuous-Depth Transformers with Learned Control Dynamics
par: Jemley, Peter
Publié: (2026)
par: Jemley, Peter
Publié: (2026)
Scaling Trends for Multi-Hop Contextual Reasoning in Mid-Scale Language Models
par: Steele, Brady, et autres
Publié: (2026)
par: Steele, Brady, et autres
Publié: (2026)
Shattered Compositionality: Counterintuitive Learning Dynamics of Transformers for Arithmetic
par: Zhao, Xingyu, et autres
Publié: (2026)
par: Zhao, Xingyu, et autres
Publié: (2026)
Task-Conditioned Routing Signatures in Sparse Mixture-of-Experts Transformers
par: Avinash, Mynampati Sri Ranganadha
Publié: (2026)
par: Avinash, Mynampati Sri Ranganadha
Publié: (2026)
Thread Detection and Response Generation using Transformers with Prompt Optimisation
par: T, Kevin Joshua, et autres
Publié: (2024)
par: T, Kevin Joshua, et autres
Publié: (2024)
Semantic Convergence: Investigating Shared Representations Across Scaled LLMs
par: Son, Daniel, et autres
Publié: (2025)
par: Son, Daniel, et autres
Publié: (2025)
Discovering Transformer Circuits via a Hybrid Attribution and Pruning Framework
par: Gu, Hao, et autres
Publié: (2025)
par: Gu, Hao, et autres
Publié: (2025)
Transformers Boost the Performance of Decision Trees on Tabular Data across Sample Sizes
par: Jayawardhana, Mayuka, et autres
Publié: (2025)
par: Jayawardhana, Mayuka, et autres
Publié: (2025)
CircuitProbe: Predicting Reasoning Circuits in Transformers via Stability Zone Detection
par: Panuganti, Rajkiran
Publié: (2026)
par: Panuganti, Rajkiran
Publié: (2026)
The Right Answer, the Wrong Direction: Why Transformers Fail at Counting and How to Fix It
par: Garcia, Gabriel
Publié: (2026)
par: Garcia, Gabriel
Publié: (2026)
From Syntax to Semantics: Unveiling the Emergence of Chirality in SMILES Translation Models
par: Li, Zehao, et autres
Publié: (2026)
par: Li, Zehao, et autres
Publié: (2026)
Documents similaires
-
ACE: Exploring Activation Cosine Similarity and Variance for Accurate and Calibration-Efficient LLM Pruning
par: Mi, Zhendong, et autres
Publié: (2025) -
Transformer Scalability Crisis: The First Comprehensive Empirical Analysis of Performance Walls in Modern Language Models
par: Moghadasi, Mahdi Naser, et autres
Publié: (2026) -
TensorLens: End-to-End Transformer Analysis via High-Order Attention Tensors
par: Atad, Ido Andrew, et autres
Publié: (2026) -
Quantization Undoes Alignment: Bias Emergence in Compressed LLMs Across Models and Precision Levels
par: Rath, Plawan Kumar, et autres
Publié: (2026) -
On the Limits of Learned Importance Scoring for KV Cache Compression
par: Steele, Brady
Publié: (2026)