Training Memory in Deep Neural Networks: Mechanisms, Evidence, and Measurement Gaps
Fuente:
arXiv
Saved in:
| Main Authors: | Sevetlidis, Vasileios, Pavlidis, George |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Process-Tensor Tomography of SGD: Measuring Non-Markovian Memory via Back-Flow of Distinguishability
by: Sevetlidis, Vasileios, et al.
Published: (2026)
by: Sevetlidis, Vasileios, et al.
Published: (2026)
Gauge-invariant representation holonomy
by: Sevetlidis, Vasileios, et al.
Published: (2026)
by: Sevetlidis, Vasileios, et al.
Published: (2026)
Angular Regularization for Positive-Unlabeled Learning on the Hypersphere
by: Sevetlidis, Vasileios, et al.
Published: (2025)
by: Sevetlidis, Vasileios, et al.
Published: (2025)
Dens-PU: PU Learning with Density-Based Positive Labeled Augmentation
by: Sevetlidis, Vasileios, et al.
Published: (2023)
by: Sevetlidis, Vasileios, et al.
Published: (2023)
A Fiber Criterion for Representation Identifiability in Supervised Learning
by: Sevetlidis, Vasileios
Published: (2026)
by: Sevetlidis, Vasileios
Published: (2026)
Defect detection using weakly supervised learning
by: Sevetlidis, Vasileios, et al.
Published: (2023)
by: Sevetlidis, Vasileios, et al.
Published: (2023)
Deep Neural Networks with 3D Point Clouds for Empirical Friction Measurements in Hydrodynamic Flood Models
by: Haces-Garcia, Francisco, et al.
Published: (2024)
by: Haces-Garcia, Francisco, et al.
Published: (2024)
Armada: Memory-Efficient Distributed Training of Large-Scale Graph Neural Networks
by: Waleffe, Roger, et al.
Published: (2025)
by: Waleffe, Roger, et al.
Published: (2025)
Sonnet: Spectral Operator Neural Network for Multivariable Time Series Forecasting
by: Shu, Yuxuan, et al.
Published: (2025)
by: Shu, Yuxuan, et al.
Published: (2025)
Evaluation of Bio-Inspired Models under Different Learning Settings For Energy Efficiency in Network Traffic Prediction
by: Tsiolakis, Theodoros, et al.
Published: (2024)
by: Tsiolakis, Theodoros, et al.
Published: (2024)
On the Hardness of Training Deep Neural Networks Discretely
by: Doron-Arad, Ilan
Published: (2024)
by: Doron-Arad, Ilan
Published: (2024)
Inverted Activations: Reducing Memory Footprint in Neural Network Training
by: Novikov, Georgii, et al.
Published: (2024)
by: Novikov, Georgii, et al.
Published: (2024)
Memory Faults in Activation-sparse Quantized Deep Neural Networks: Analysis and Mitigation using Sharpness-aware Training
by: Malhotra, Akul, et al.
Published: (2024)
by: Malhotra, Akul, et al.
Published: (2024)
On Measuring Intrinsic Causal Attributions in Deep Neural Networks
by: Saha, Saptarshi, et al.
Published: (2025)
by: Saha, Saptarshi, et al.
Published: (2025)
Training Deep Morphological Neural Networks as Universal Approximators
by: Fotopoulos, Konstantinos, et al.
Published: (2025)
by: Fotopoulos, Konstantinos, et al.
Published: (2025)
Concurrent Training and Layer Pruning of Deep Neural Networks
by: Guenter, Valentin Frank Ingmar, et al.
Published: (2024)
by: Guenter, Valentin Frank Ingmar, et al.
Published: (2024)
Bulk-Switching Memristor-based Compute-In-Memory Module for Deep Neural Network Training
by: Wu, Yuting, et al.
Published: (2023)
by: Wu, Yuting, et al.
Published: (2023)
Advancing Physics Data Analysis through Machine Learning and Physics-Informed Neural Networks
by: Vatellis, Vasileios
Published: (2024)
by: Vatellis, Vasileios
Published: (2024)
Exact Gauss-Newton Optimization for Training Deep Neural Networks
by: Korbit, Mikalai, et al.
Published: (2024)
by: Korbit, Mikalai, et al.
Published: (2024)
Deep Neural Network Training as Random Effects: An Optimization-Inference Duality
by: Yao, Minhao, et al.
Published: (2026)
by: Yao, Minhao, et al.
Published: (2026)
Solving Inverse Problems with Deep Linear Neural Networks: Global Convergence Guarantees for Gradient Descent with Weight Decay
by: Laus, Hannah, et al.
Published: (2025)
by: Laus, Hannah, et al.
Published: (2025)
CPT: Efficient Deep Neural Network Training via Cyclic Precision
by: Fu, Yonggan, et al.
Published: (2021)
by: Fu, Yonggan, et al.
Published: (2021)
Complexity-Aware Training of Deep Neural Networks for Optimal Structure Discovery
by: Guenter, Valentin Frank Ingmar, et al.
Published: (2024)
by: Guenter, Valentin Frank Ingmar, et al.
Published: (2024)
TensorGRaD: Tensor Gradient Robust Decomposition for Memory-Efficient Neural Operator Training
by: Loeschcke, Sebastian, et al.
Published: (2025)
by: Loeschcke, Sebastian, et al.
Published: (2025)
GPU Memory Usage Optimization for Backward Propagation in Deep Network Training
by: Hong, Ding-Yong, et al.
Published: (2025)
by: Hong, Ding-Yong, et al.
Published: (2025)
NeuZip: Memory-Efficient Training and Inference with Dynamic Compression of Neural Networks
by: Hao, Yongchang, et al.
Published: (2024)
by: Hao, Yongchang, et al.
Published: (2024)
Benchmarking Neural Network Training Algorithms
by: Dahl, George E., et al.
Published: (2023)
by: Dahl, George E., et al.
Published: (2023)
Convolutional Neural Networks for Accurate Measurement of Train Speed
by: Tian, Haitao, et al.
Published: (2025)
by: Tian, Haitao, et al.
Published: (2025)
Chordal Sparsity for Lipschitz Constant Estimation of Deep Neural Networks
by: Xue, Anton, et al.
Published: (2022)
by: Xue, Anton, et al.
Published: (2022)
FreshGNN: Reducing Memory Access via Stable Historical Embeddings for Graph Neural Network Training
by: Huang, Kezhao, et al.
Published: (2023)
by: Huang, Kezhao, et al.
Published: (2023)
Theory-to-Practice Gap for Neural Networks and Neural Operators
by: Grohs, Philipp, et al.
Published: (2025)
by: Grohs, Philipp, et al.
Published: (2025)
Layerwise Progressive Freezing Enables STE-Free Training of Deep Binary Neural Networks
by: Smith, Evan Gibson, et al.
Published: (2026)
by: Smith, Evan Gibson, et al.
Published: (2026)
Memory Analysis on the Training Course of DeepSeek Models
by: Zhang, Ping, et al.
Published: (2025)
by: Zhang, Ping, et al.
Published: (2025)
PreNeT: Leveraging Computational Features to Predict Deep Neural Network Training Time
by: Pourali, Alireza, et al.
Published: (2024)
by: Pourali, Alireza, et al.
Published: (2024)
SortedNet: A Scalable and Generalized Framework for Training Modular Deep Neural Networks
by: Valipour, Mojtaba, et al.
Published: (2023)
by: Valipour, Mojtaba, et al.
Published: (2023)
AdaBet: Gradient-free Layer Selection for Efficient Training of Deep Neural Networks
by: Tenison, Irene, et al.
Published: (2025)
by: Tenison, Irene, et al.
Published: (2025)
Towards Deep Encrypted Training: Low-Latency, Memory-Efficient, and High-Throughput Inference for Privacy-Preserving Neural Networks
by: Njungle, Nges Brian, et al.
Published: (2026)
by: Njungle, Nges Brian, et al.
Published: (2026)
On the Dataless Training of Neural Networks
by: Velasquez, Alvaro, et al.
Published: (2025)
by: Velasquez, Alvaro, et al.
Published: (2025)
Using the IBM Analog In-Memory Hardware Acceleration Kit for Neural Network Training and Inference
by: Gallo, Manuel Le, et al.
Published: (2023)
by: Gallo, Manuel Le, et al.
Published: (2023)
Self-Abstraction Learning for Effective and Stable Training of Deep Neural Networks
by: Cho, Wonyong, et al.
Published: (2026)
by: Cho, Wonyong, et al.
Published: (2026)
Similar Items
-
Process-Tensor Tomography of SGD: Measuring Non-Markovian Memory via Back-Flow of Distinguishability
by: Sevetlidis, Vasileios, et al.
Published: (2026) -
Gauge-invariant representation holonomy
by: Sevetlidis, Vasileios, et al.
Published: (2026) -
Angular Regularization for Positive-Unlabeled Learning on the Hypersphere
by: Sevetlidis, Vasileios, et al.
Published: (2025) -
Dens-PU: PU Learning with Density-Based Positive Labeled Augmentation
by: Sevetlidis, Vasileios, et al.
Published: (2023) -
A Fiber Criterion for Representation Identifiability in Supervised Learning
by: Sevetlidis, Vasileios
Published: (2026)