Supernodes and Halos: Loss-Critical Hubs in LLM Feed-Forward Layers
Fuente:
arXiv
Salvato in:
| Autori principali: | Cherilyn, Audrey, Safaai, Houman |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Amortized Vine Copulas for High-Dimensional Density and Information Estimation
di: Safaai, Houman
Pubblicazione: (2026)
di: Safaai, Houman
Pubblicazione: (2026)
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
di: Bozic, Vukasin, et al.
Pubblicazione: (2023)
di: Bozic, Vukasin, et al.
Pubblicazione: (2023)
Merging Feed-Forward Sublayers for Compressed Transformers
di: Verma, Neha, et al.
Pubblicazione: (2025)
di: Verma, Neha, et al.
Pubblicazione: (2025)
Dynamic Vine Copulas: Detecting and Quantifying Time-Varying Higher-Order Interactions
di: Safaai, Houman, et al.
Pubblicazione: (2026)
di: Safaai, Houman, et al.
Pubblicazione: (2026)
FFN-SkipLLM: A Hidden Gem for Autoregressive Decoding with Adaptive Feed Forward Skipping
di: Jaiswal, Ajay, et al.
Pubblicazione: (2024)
di: Jaiswal, Ajay, et al.
Pubblicazione: (2024)
HaloScope: Harnessing Unlabeled LLM Generations for Hallucination Detection
di: Du, Xuefeng, et al.
Pubblicazione: (2024)
di: Du, Xuefeng, et al.
Pubblicazione: (2024)
Flash Multi-Head Feed-Forward Network
di: Zhang, Minshen, et al.
Pubblicazione: (2025)
di: Zhang, Minshen, et al.
Pubblicazione: (2025)
Task Relevance Is Not Local Replaceability: A Two-Axis View of Channel Information
di: Safaai, Houman, et al.
Pubblicazione: (2026)
di: Safaai, Houman, et al.
Pubblicazione: (2026)
Sparse Layers are Critical to Scaling Looped Language Models
di: Lee, Ryan, et al.
Pubblicazione: (2026)
di: Lee, Ryan, et al.
Pubblicazione: (2026)
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
di: Kim, Eunsu, et al.
Pubblicazione: (2025)
di: Kim, Eunsu, et al.
Pubblicazione: (2025)
Spectral Insights into Data-Oblivious Critical Layers in Large Language Models
di: Liu, Xuyuan, et al.
Pubblicazione: (2025)
di: Liu, Xuyuan, et al.
Pubblicazione: (2025)
When Chain-of-Thought Fails, the Solution Hides in the Hidden States
di: Mehrafarin, Houman, et al.
Pubblicazione: (2026)
di: Mehrafarin, Houman, et al.
Pubblicazione: (2026)
Improving LLM Final Representations with Inter-Layer Geometry
di: Ulanovski, Tom, et al.
Pubblicazione: (2026)
di: Ulanovski, Tom, et al.
Pubblicazione: (2026)
Distributional Associations vs In-Context Reasoning: A Study of Feed-forward and Attention Layers
di: Chen, Lei, et al.
Pubblicazione: (2024)
di: Chen, Lei, et al.
Pubblicazione: (2024)
CLAA: Cross-Layer Attention Aggregation for Accelerating LLM Prefill
di: McDanel, Bradley, et al.
Pubblicazione: (2026)
di: McDanel, Bradley, et al.
Pubblicazione: (2026)
Fast Forwarding Low-Rank Training
di: Rahamim, Adir, et al.
Pubblicazione: (2024)
di: Rahamim, Adir, et al.
Pubblicazione: (2024)
LIDS: LLM Summary Inference Under the Layered Lens
di: Park, Dylan, et al.
Pubblicazione: (2026)
di: Park, Dylan, et al.
Pubblicazione: (2026)
Toward LLM-Supported Automated Assessment of Critical Thinking Subskills
di: Peczuh, Marisa C., et al.
Pubblicazione: (2025)
di: Peczuh, Marisa C., et al.
Pubblicazione: (2025)
Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference
di: Taniguchi, Rei, et al.
Pubblicazione: (2026)
di: Taniguchi, Rei, et al.
Pubblicazione: (2026)
Fine-Tuning Language Models with Just Forward Passes
di: Malladi, Sadhika, et al.
Pubblicazione: (2023)
di: Malladi, Sadhika, et al.
Pubblicazione: (2023)
nanoLM: an Affordable LLM Pre-training Benchmark via Accurate Loss Prediction across Scales
di: Yao, Yiqun, et al.
Pubblicazione: (2023)
di: Yao, Yiqun, et al.
Pubblicazione: (2023)
LCSB: Layer-Cyclic Selective Backpropagation for Memory-Efficient On-Device LLM Fine-Tuning
di: Park, Juneyoung, et al.
Pubblicazione: (2026)
di: Park, Juneyoung, et al.
Pubblicazione: (2026)
FBQuant: FeedBack Quantization for Large Language Models
di: Liu, Yijiang, et al.
Pubblicazione: (2025)
di: Liu, Yijiang, et al.
Pubblicazione: (2025)
ReLook: Vision-Grounded RL with a Multimodal LLM Critic for Agentic Web Coding
di: Li, Yuhang, et al.
Pubblicazione: (2025)
di: Li, Yuhang, et al.
Pubblicazione: (2025)
DASH: Input-Aware Dynamic Layer Skipping for Efficient LLM Inference with Markov Decision Policies
di: Yang, Ning, et al.
Pubblicazione: (2025)
di: Yang, Ning, et al.
Pubblicazione: (2025)
Dr.LLM: Dynamic Layer Routing in LLMs
di: Heakl, Ahmed, et al.
Pubblicazione: (2025)
di: Heakl, Ahmed, et al.
Pubblicazione: (2025)
LESA: Learnable LLM Layer Scaling-Up
di: Yang, Yifei, et al.
Pubblicazione: (2025)
di: Yang, Yifei, et al.
Pubblicazione: (2025)
LLM Unlearning via Loss Adjustment with Only Forget Data
di: Wang, Yaxuan, et al.
Pubblicazione: (2024)
di: Wang, Yaxuan, et al.
Pubblicazione: (2024)
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache
di: Sharma, Akshat, et al.
Pubblicazione: (2024)
di: Sharma, Akshat, et al.
Pubblicazione: (2024)
Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes
di: Kolawole, Steven, et al.
Pubblicazione: (2024)
di: Kolawole, Steven, et al.
Pubblicazione: (2024)
SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal Budget
di: Wang, Zihao, et al.
Pubblicazione: (2024)
di: Wang, Zihao, et al.
Pubblicazione: (2024)
ChatGPT in Linear Algebra: Strides Forward, Steps to Go
di: Bagno, Eli, et al.
Pubblicazione: (2024)
di: Bagno, Eli, et al.
Pubblicazione: (2024)
TrimLLM: Progressive Layer Dropping for Domain-Specific LLMs
di: Hu, Lanxiang, et al.
Pubblicazione: (2024)
di: Hu, Lanxiang, et al.
Pubblicazione: (2024)
Concept Layers: Enhancing Interpretability and Intervenability via LLM Conceptualization
di: Bidusa, Or Raphael, et al.
Pubblicazione: (2025)
di: Bidusa, Or Raphael, et al.
Pubblicazione: (2025)
LayerBoost: Layer-Aware Attention Reduction for Efficient LLMs
di: Souibgui, Mohamed Ali, et al.
Pubblicazione: (2026)
di: Souibgui, Mohamed Ali, et al.
Pubblicazione: (2026)
ConfLayers: Adaptive Confidence-based Layer Skipping for Self-Speculative Decoding
di: Amer, Walaa, et al.
Pubblicazione: (2026)
di: Amer, Walaa, et al.
Pubblicazione: (2026)
From Memorization to Reasoning in the Spectrum of Loss Curvature
di: Merullo, Jack, et al.
Pubblicazione: (2025)
di: Merullo, Jack, et al.
Pubblicazione: (2025)
Hyperparameter Loss Surfaces Are Simple Near their Optima
di: Lourie, Nicholas, et al.
Pubblicazione: (2025)
di: Lourie, Nicholas, et al.
Pubblicazione: (2025)
FFCL: Forward-Forward Net with Cortical Loops, Training and Inference on Edge Without Backpropagation
di: Karkehabadi, Ali, et al.
Pubblicazione: (2024)
di: Karkehabadi, Ali, et al.
Pubblicazione: (2024)
FAAST: Forward-Only Associative Learning via Closed-Form Fast Weights for Test-Time Supervised Adaptation
di: Bao, Guangsheng, et al.
Pubblicazione: (2026)
di: Bao, Guangsheng, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Amortized Vine Copulas for High-Dimensional Density and Information Estimation
di: Safaai, Houman
Pubblicazione: (2026) -
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
di: Bozic, Vukasin, et al.
Pubblicazione: (2023) -
Merging Feed-Forward Sublayers for Compressed Transformers
di: Verma, Neha, et al.
Pubblicazione: (2025) -
Dynamic Vine Copulas: Detecting and Quantifying Time-Varying Higher-Order Interactions
di: Safaai, Houman, et al.
Pubblicazione: (2026) -
FFN-SkipLLM: A Hidden Gem for Autoregressive Decoding with Adaptive Feed Forward Skipping
di: Jaiswal, Ajay, et al.
Pubblicazione: (2024)