Progressive Feedforward Collapse of ResNet Training
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Sicong, Gai, Kuo, Zhang, Shihua |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Correcting Stochastic Update Bias in Preconditioned Language Model Optimizers
by: Nayak, Nikhil, et al.
Published: (2026)
by: Nayak, Nikhil, et al.
Published: (2026)
Interpretability Can Be Actionable
by: Orgad, Hadas, et al.
Published: (2026)
by: Orgad, Hadas, et al.
Published: (2026)
Improving Time Series Classification with Representation Soft Label Smoothing
by: Ma, Hengyi, et al.
Published: (2024)
by: Ma, Hengyi, et al.
Published: (2024)
Energy-Efficient Deep Learning Without Backpropagation: A Rigorous Evaluation of Forward-Only Algorithms
by: Spyra, Przemysław, et al.
Published: (2025)
by: Spyra, Przemysław, et al.
Published: (2025)
Graph Neural Networks Need Cluster-Normalize-Activate Modules
by: Skryagin, Arseny, et al.
Published: (2024)
by: Skryagin, Arseny, et al.
Published: (2024)
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference
by: Jørgensen, Tollef Emil
Published: (2025)
by: Jørgensen, Tollef Emil
Published: (2025)
Matryoshka Policy Gradient for Entropy-Regularized RL: Convergence and Global Optimality
by: Ged, François, et al.
Published: (2023)
by: Ged, François, et al.
Published: (2023)
On Privacy Leakage in Tabular Diffusion Models: Influential Factors, Attacker Knowledge, and Metrics
by: Shafieinejad, Masoumeh, et al.
Published: (2026)
by: Shafieinejad, Masoumeh, et al.
Published: (2026)
Deceptive Diffusion: Generating Synthetic Adversarial Examples
by: Beerens, Lucas, et al.
Published: (2024)
by: Beerens, Lucas, et al.
Published: (2024)
Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure
by: Sarkar, Nilesh, et al.
Published: (2026)
by: Sarkar, Nilesh, et al.
Published: (2026)
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
The Curious Case of In-Training Compression of State Space Models
by: Chahine, Makram, et al.
Published: (2025)
by: Chahine, Makram, et al.
Published: (2025)
ElegansNet: a brief scientific report and initial experiments
by: Bardozzo, Francesco, et al.
Published: (2023)
by: Bardozzo, Francesco, et al.
Published: (2023)
Position: Mechanistic Interpretability Must Disclose Identification Assumptions for Causal Claims
by: Lin, Zezheng, et al.
Published: (2026)
by: Lin, Zezheng, et al.
Published: (2026)
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
by: Badshah, Sher, et al.
Published: (2024)
by: Badshah, Sher, et al.
Published: (2024)
Large Language Models Report Subjective Experience Under Self-Referential Processing
by: Berg, Cameron, et al.
Published: (2025)
by: Berg, Cameron, et al.
Published: (2025)
Are Transformers with One Layer Self-Attention Using Low-Rank Weight Matrices Universal Approximators?
by: Kajitsuka, Tokio, et al.
Published: (2023)
by: Kajitsuka, Tokio, et al.
Published: (2023)
Safe Continual Reinforcement Learning Methods for Nonstationary Environments. Towards a Survey of the State of the Art
by: Tomashevskiy, Timofey
Published: (2026)
by: Tomashevskiy, Timofey
Published: (2026)
UR4NNV: Neural Network Verification, Under-approximation Reachability Works!
by: Liang, Zhen, et al.
Published: (2024)
by: Liang, Zhen, et al.
Published: (2024)
The First MPDD Challenge: Multimodal Personality-aware Depression Detection
by: Fu, Changzeng, et al.
Published: (2025)
by: Fu, Changzeng, et al.
Published: (2025)
Union of Experts: Adapting Hierarchical Routing to Equivalently Decomposed Transformer
by: Yang, Yujiao, et al.
Published: (2025)
by: Yang, Yujiao, et al.
Published: (2025)
The Matching Principle: A Geometric Theory of Loss Functions for Nuisance-Robust Representation Learning
by: Rajput, Vishal
Published: (2026)
by: Rajput, Vishal
Published: (2026)
Stealth edits to large language models
by: Sutton, Oliver J., et al.
Published: (2024)
by: Sutton, Oliver J., et al.
Published: (2024)
Efficient compression of neural networks and datasets
by: Barth, Lukas Silvester, et al.
Published: (2025)
by: Barth, Lukas Silvester, et al.
Published: (2025)
Regime Change Hypothesis: Foundations for Decoupled Dynamics in Neural Network Training
by: Pérez-Corral, Cristian, et al.
Published: (2026)
by: Pérez-Corral, Cristian, et al.
Published: (2026)
OTAD: An Optimal Transport-Induced Robust Model for Agnostic Adversarial Attack
by: Gai, Kuo, et al.
Published: (2024)
by: Gai, Kuo, et al.
Published: (2024)
JAM: Controllable and Responsible Text Generation via Causal Reasoning and Latent Vector Manipulation
by: Huang, Yingbing, et al.
Published: (2025)
by: Huang, Yingbing, et al.
Published: (2025)
A ZeNN architecture to avoid the Gaussian trap
by: Carvalho, Luís, et al.
Published: (2025)
by: Carvalho, Luís, et al.
Published: (2025)
Robust DDoS-Attack Classification with 3D CNNs Against Adversarial Methods
by: Bragg, Landon, et al.
Published: (2025)
by: Bragg, Landon, et al.
Published: (2025)
Variance Reduced Policy Gradient Method for Multi-Objective Reinforcement Learning
by: Guidobene, Davide, et al.
Published: (2025)
by: Guidobene, Davide, et al.
Published: (2025)
Customizing Graph Neural Networks using Path Reweighting
by: Chen, Jianpeng, et al.
Published: (2021)
by: Chen, Jianpeng, et al.
Published: (2021)
Adaptive Latent-Space Constraints in Personalized Federated Learning
by: Ayromlou, Sana, et al.
Published: (2025)
by: Ayromlou, Sana, et al.
Published: (2025)
Sharp higher order convergence rates for the Adam optimizer
by: Dereich, Steffen, et al.
Published: (2025)
by: Dereich, Steffen, et al.
Published: (2025)
JacNet: Learning Functions with Structured Jacobians
by: Lorraine, Jonathan, et al.
Published: (2024)
by: Lorraine, Jonathan, et al.
Published: (2024)
Correcting Gradient-Based Circuit Localization via Interaction-Aware Backpropagation
by: Edin, Joakim, et al.
Published: (2025)
by: Edin, Joakim, et al.
Published: (2025)
Visual Categorization Across Minds and Models: Cognitive Analysis of Human Labeling and Neuro-Symbolic Integration
by: Kabgere, Chethana Prasad
Published: (2025)
by: Kabgere, Chethana Prasad
Published: (2025)
Enhancing Feature Selection and Interpretability in AI Regression Tasks Through Feature Attribution
by: Hinterleitner, Alexander, et al.
Published: (2024)
by: Hinterleitner, Alexander, et al.
Published: (2024)
Scalable Heterogeneous Graph Foundation Models for Data-Driven Optimal Power Flow in Smart Grids
by: Pasini, Massimiliano Lupo, et al.
Published: (2026)
by: Pasini, Massimiliano Lupo, et al.
Published: (2026)
DeepPersona: A Generative Engine for Scaling Deep Synthetic Personas
by: Wang, Zhen, et al.
Published: (2025)
by: Wang, Zhen, et al.
Published: (2025)
TraXion: Rethinking Pre-training Frameworks for Mobility and Beyond
by: Hsu, Shang-Ling, et al.
Published: (2026)
by: Hsu, Shang-Ling, et al.
Published: (2026)
Similar Items
-
Correcting Stochastic Update Bias in Preconditioned Language Model Optimizers
by: Nayak, Nikhil, et al.
Published: (2026) -
Interpretability Can Be Actionable
by: Orgad, Hadas, et al.
Published: (2026) -
Improving Time Series Classification with Representation Soft Label Smoothing
by: Ma, Hengyi, et al.
Published: (2024) -
Energy-Efficient Deep Learning Without Backpropagation: A Rigorous Evaluation of Forward-Only Algorithms
by: Spyra, Przemysław, et al.
Published: (2025) -
Graph Neural Networks Need Cluster-Normalize-Activate Modules
by: Skryagin, Arseny, et al.
Published: (2024)