End-to-End Training Induces Information Bottleneck through Layer-Role Differentiation: A Comparative Analysis with Layer-wise Training
Fuente:
arXiv
Saved in:
| Main Authors: | Sakamoto, Keitaro, Sato, Issei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Explaining Grokking and Information Bottleneck through Neural Collapse Emergence
by: Sakamoto, Keitaro, et al.
Published: (2025)
by: Sakamoto, Keitaro, et al.
Published: (2025)
Benign Overfitting in Token Selection of Attention Mechanism
by: Sakamoto, Keitaro, et al.
Published: (2024)
by: Sakamoto, Keitaro, et al.
Published: (2024)
LOTOS: Layer-wise Orthogonalization for Training Robust Ensembles
by: Ebrahimpour-Boroojeny, Ali, et al.
Published: (2024)
by: Ebrahimpour-Boroojeny, Ali, et al.
Published: (2024)
Information-Theoretic Greedy Layer-wise Training for Traffic Sign Recognition
by: Lyu, Shuyan, et al.
Published: (2025)
by: Lyu, Shuyan, et al.
Published: (2025)
Rationalizing Transformer Predictions via End-To-End Differentiable Self-Training
by: Brinner, Marc, et al.
Published: (2025)
by: Brinner, Marc, et al.
Published: (2025)
Are Transformers with One Layer Self-Attention Using Low-Rank Weight Matrices Universal Approximators?
by: Kajitsuka, Tokio, et al.
Published: (2023)
by: Kajitsuka, Tokio, et al.
Published: (2023)
Exploring Layer-wise Information Effectiveness for Post-Training Quantization in Small Language Models
by: Xiao, He, et al.
Published: (2025)
by: Xiao, He, et al.
Published: (2025)
End-to-End Test-Time Training for Long Context
by: Tandon, Arnuv, et al.
Published: (2025)
by: Tandon, Arnuv, et al.
Published: (2025)
Efficient Knowledge Deletion from Trained Models through Layer-wise Partial Machine Unlearning
by: Gogineni, Vinay Chakravarthi, et al.
Published: (2024)
by: Gogineni, Vinay Chakravarthi, et al.
Published: (2024)
Resource-Efficient Federated Multimodal Learning via Layer-wise and Progressive Training
by: Tun, Ye Lin, et al.
Published: (2024)
by: Tun, Ye Lin, et al.
Published: (2024)
Sobolev Training of End-to-End Optimization Proxies
by: Rosemberg, Andrew W., et al.
Published: (2025)
by: Rosemberg, Andrew W., et al.
Published: (2025)
LaRA: Layer-wise Representation Analysis for Detecting Data Contamination in RL Post-Training
by: Gwak, Minju, et al.
Published: (2026)
by: Gwak, Minju, et al.
Published: (2026)
Efficient and Flexible Neural Network Training through Layer-wise Feedback Propagation
by: Weber, Leander, et al.
Published: (2023)
by: Weber, Leander, et al.
Published: (2023)
Conformal Risk Training: End-to-End Optimization of Conformal Risk Control
by: Yeh, Christopher, et al.
Published: (2025)
by: Yeh, Christopher, et al.
Published: (2025)
End-to-End Training for Back-Translation with Categorical Reparameterization Trick
by: Heo, DongNyeong, et al.
Published: (2022)
by: Heo, DongNyeong, et al.
Published: (2022)
Time-Varying Audio Effect Modeling by End-to-End Adversarial Training
by: Bourdin, Yann, et al.
Published: (2025)
by: Bourdin, Yann, et al.
Published: (2025)
On Expressive Power of Looped Transformers: Theoretical Analysis and Enhancement via Timestep Encoding
by: Xu, Kevin, et al.
Published: (2024)
by: Xu, Kevin, et al.
Published: (2024)
End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost
by: Tan, Qitao, et al.
Published: (2025)
by: Tan, Qitao, et al.
Published: (2025)
Revisiting End-To-End Sparse Autoencoder Training: A Short Finetune Is All You Need
by: Karvonen, Adam
Published: (2025)
by: Karvonen, Adam
Published: (2025)
Max-pooling Network Revisited: Analyzing the Role of Semantic Probability in Multiple Instance Learning for Hallucination Detection
by: Fujikawa, Shota, et al.
Published: (2026)
by: Fujikawa, Shota, et al.
Published: (2026)
Constraint-Aware Flow Matching: Decision Aligned End-to-End Training for Constrained Sampling
by: Christopher, Jacob K., et al.
Published: (2026)
by: Christopher, Jacob K., et al.
Published: (2026)
ResearchGPT: Benchmarking and Training LLMs for End-to-End Computer Science Research Workflows
by: Wang, Penghao, et al.
Published: (2025)
by: Wang, Penghao, et al.
Published: (2025)
Understanding Generalization in Physics Informed Models through Affine Variety Dimensions
by: Koshizuka, Takeshi, et al.
Published: (2025)
by: Koshizuka, Takeshi, et al.
Published: (2025)
How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
Rethinking Layer-wise Model Merging through Chain of Merges
by: Buzzega, Pietro, et al.
Published: (2025)
by: Buzzega, Pietro, et al.
Published: (2025)
Can Test-time Computation Mitigate Reproduction Bias in Neural Symbolic Regression?
by: Sato, Shun, et al.
Published: (2025)
by: Sato, Shun, et al.
Published: (2025)
DOGE-Train: Discrete Optimization on GPU with End-to-end Training
by: Abbas, Ahmed, et al.
Published: (2022)
by: Abbas, Ahmed, et al.
Published: (2022)
A Layer-wise Analysis of Supervised Fine-Tuning
by: Zhao, Qinghua, et al.
Published: (2026)
by: Zhao, Qinghua, et al.
Published: (2026)
OneLog: Towards End-to-End Training in Software Log Anomaly Detection
by: Hashemi, Shayan, et al.
Published: (2021)
by: Hashemi, Shayan, et al.
Published: (2021)
Learning to Remember: End-to-End Training of Memory Agents for Long-Context Reasoning
by: Zhang, Kehao, et al.
Published: (2026)
by: Zhang, Kehao, et al.
Published: (2026)
Unsupervised End-to-End Training with a Self-Defined Target
by: Liu, Dongshu, et al.
Published: (2024)
by: Liu, Dongshu, et al.
Published: (2024)
End-to-End Training for Unified Tokenization and Latent Denoising
by: Duggal, Shivam, et al.
Published: (2026)
by: Duggal, Shivam, et al.
Published: (2026)
PETRA: Parallel End-to-end Training with Reversible Architectures
by: Rivaud, Stéphane, et al.
Published: (2024)
by: Rivaud, Stéphane, et al.
Published: (2024)
Predictive Concept Decoders: Training Scalable End-to-End Interpretability Assistants
by: Huang, Vincent, et al.
Published: (2025)
by: Huang, Vincent, et al.
Published: (2025)
DAF: An Efficient End-to-End Dynamic Activation Framework for on-Device DNN Training
by: Liu, Renyuan, et al.
Published: (2025)
by: Liu, Renyuan, et al.
Published: (2025)
Layer-wise Linear Mode Connectivity
by: Adilova, Linara, et al.
Published: (2023)
by: Adilova, Linara, et al.
Published: (2023)
Layer-wise Derivative Controlled Networks
by: Martnishn, Rowan, et al.
Published: (2026)
by: Martnishn, Rowan, et al.
Published: (2026)
End-to-end Planner Training for Language Modeling
by: Cornille, Nathan, et al.
Published: (2024)
by: Cornille, Nathan, et al.
Published: (2024)
Disentangling Recall and Reasoning in Transformer Models through Layer-wise Attention and Activation Analysis
by: Fartale, Harshwardhan, et al.
Published: (2025)
by: Fartale, Harshwardhan, et al.
Published: (2025)
A Comparative analysis of Layer-wise Representational Capacity in AR and Diffusion LLMs
by: Goel, Raghavv, et al.
Published: (2026)
by: Goel, Raghavv, et al.
Published: (2026)
Similar Items
-
Explaining Grokking and Information Bottleneck through Neural Collapse Emergence
by: Sakamoto, Keitaro, et al.
Published: (2025) -
Benign Overfitting in Token Selection of Attention Mechanism
by: Sakamoto, Keitaro, et al.
Published: (2024) -
LOTOS: Layer-wise Orthogonalization for Training Robust Ensembles
by: Ebrahimpour-Boroojeny, Ali, et al.
Published: (2024) -
Information-Theoretic Greedy Layer-wise Training for Traffic Sign Recognition
by: Lyu, Shuyan, et al.
Published: (2025) -
Rationalizing Transformer Predictions via End-To-End Differentiable Self-Training
by: Brinner, Marc, et al.
Published: (2025)