Understanding Generalization, Robustness, and Interpretability in Low-Capacity Neural Networks
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Kumar, Yash |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards understanding how attention mechanism works in deep learning
von: Ruan, Tianyu, et al.
Veröffentlicht: (2024)
von: Ruan, Tianyu, et al.
Veröffentlicht: (2024)
CVCM Track Circuits Pre-emptive Failure Diagnostics for Predictive Maintenance Using Deep Neural Networks
von: Mukherjee, Debdeep, et al.
Veröffentlicht: (2025)
von: Mukherjee, Debdeep, et al.
Veröffentlicht: (2025)
multivariateGPT: a decoder-only transformer for multivariate categorical and numeric data
von: Loza, Andrew J., et al.
Veröffentlicht: (2025)
von: Loza, Andrew J., et al.
Veröffentlicht: (2025)
DecompKAN: Decomposed Patch-KAN for Long-Term Time Series Forecasting
von: Mysore, Naveen
Veröffentlicht: (2026)
von: Mysore, Naveen
Veröffentlicht: (2026)
Scaling Laws in the Tiny Regime: How Small Models Change Their Mistakes
von: Alnemari, Mohammed, et al.
Veröffentlicht: (2026)
von: Alnemari, Mohammed, et al.
Veröffentlicht: (2026)
TGraphX: Tensor-Aware Graph Neural Network for Multi-Dimensional Feature Learning
von: Sajjadi, Arash, et al.
Veröffentlicht: (2025)
von: Sajjadi, Arash, et al.
Veröffentlicht: (2025)
Cross-Architecture Knowledge Distillation (KD) for Retinal Fundus Image Anomaly Detection on NVIDIA Jetson Nano
von: Yilmaz, Berk, et al.
Veröffentlicht: (2025)
von: Yilmaz, Berk, et al.
Veröffentlicht: (2025)
The Normalized Difference Layer: A Differentiable Spectral Index Formulation for Deep Learning
von: Lotfi, Ali, et al.
Veröffentlicht: (2026)
von: Lotfi, Ali, et al.
Veröffentlicht: (2026)
Efficient Morphology-Control Co-Design via Stackelberg Proximal Policy Optimization
von: Dai, Yanning, et al.
Veröffentlicht: (2026)
von: Dai, Yanning, et al.
Veröffentlicht: (2026)
Optimized Gradient Clipping for Noisy Label Learning
von: Ye, Xichen, et al.
Veröffentlicht: (2024)
von: Ye, Xichen, et al.
Veröffentlicht: (2024)
Pulse-Driven Neural Architecture: Learnable Oscillatory Dynamics for Robust Continuous-Time Sequence Processing
von: Sharma, Paras
Veröffentlicht: (2026)
von: Sharma, Paras
Veröffentlicht: (2026)
PolyGLU: State-Conditional Activation Routing in Transformer Feed-Forward Networks
von: Medeiros, Daniel Nobrega
Veröffentlicht: (2026)
von: Medeiros, Daniel Nobrega
Veröffentlicht: (2026)
Neural Encoding for Image Recall: Human-Like Memory
von: Foussereau, Virgile, et al.
Veröffentlicht: (2024)
von: Foussereau, Virgile, et al.
Veröffentlicht: (2024)
GraphNNK -- Graph Classification and Interpretability
von: Bolevic, Zeljko, et al.
Veröffentlicht: (2026)
von: Bolevic, Zeljko, et al.
Veröffentlicht: (2026)
Explicit Dropout: Deterministic Regularization for Transformer Architectures
von: Agrawal, Vidhi, et al.
Veröffentlicht: (2026)
von: Agrawal, Vidhi, et al.
Veröffentlicht: (2026)
Massive Redundancy in Gradient Transport Enables Sparse Online Learning
von: Merin, Aur Shalev
Veröffentlicht: (2026)
von: Merin, Aur Shalev
Veröffentlicht: (2026)
Optimizing Inference in Transformer-Based Models: A Multi-Method Benchmark
von: Ho, Siu Hang, et al.
Veröffentlicht: (2025)
von: Ho, Siu Hang, et al.
Veröffentlicht: (2025)
Rethinking Visual Intelligence: Insights from Video Pretraining
von: Acuaviva, Pablo, et al.
Veröffentlicht: (2025)
von: Acuaviva, Pablo, et al.
Veröffentlicht: (2025)
Revisiting GAN with Bayes-Optimal Discrimination
von: Naeini, Mohammadreza Tavasoli, et al.
Veröffentlicht: (2025)
von: Naeini, Mohammadreza Tavasoli, et al.
Veröffentlicht: (2025)
A Hybrid Inductive-Transductive Network for Traffic Flow Imputation on Unsampled Locations
von: Rahimiasl, Mohammadmahdi, et al.
Veröffentlicht: (2025)
von: Rahimiasl, Mohammadmahdi, et al.
Veröffentlicht: (2025)
JacNet: Learning Functions with Structured Jacobians
von: Lorraine, Jonathan, et al.
Veröffentlicht: (2024)
von: Lorraine, Jonathan, et al.
Veröffentlicht: (2024)
Machine Unlearning for Class Removal through SISA-based Deep Neural Network Architectures
von: Mahi, Ishrak Hamim, et al.
Veröffentlicht: (2026)
von: Mahi, Ishrak Hamim, et al.
Veröffentlicht: (2026)
H-Model: Dynamic Neural Architectures for Adaptive Processing
von: Hospodarchuk, Dmytro
Veröffentlicht: (2025)
von: Hospodarchuk, Dmytro
Veröffentlicht: (2025)
Predictive Modeling of Maritime Radar Data Using Transformer Architecture
von: Qesaraku, Bjorna, et al.
Veröffentlicht: (2025)
von: Qesaraku, Bjorna, et al.
Veröffentlicht: (2025)
Revisiting Non-separable Binary Classification and its Applications in Anomaly Detection
von: Lau, Matthew, et al.
Veröffentlicht: (2023)
von: Lau, Matthew, et al.
Veröffentlicht: (2023)
Predicting Traffic Accident Severity with Deep Neural Networks
von: Bibb, Meghan, et al.
Veröffentlicht: (2025)
von: Bibb, Meghan, et al.
Veröffentlicht: (2025)
Kolmogorov-Arnold Attention: Is Learnable Attention Better For Vision Transformers?
von: Maity, Subhajit, et al.
Veröffentlicht: (2025)
von: Maity, Subhajit, et al.
Veröffentlicht: (2025)
Scalable, Technology-Agnostic Diagnosis and Predictive Maintenance for Point Machine using Deep Learning
von: Di Santi, Eduardo, et al.
Veröffentlicht: (2025)
von: Di Santi, Eduardo, et al.
Veröffentlicht: (2025)
Closing the Theory-Practice Gap in Spiking Transformers via Effective Dimension
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
Temporal Functional Circuits: From Spline Plots to Faithful Explanations in KAN Forecasting
von: Mysore, Naveen
Veröffentlicht: (2026)
von: Mysore, Naveen
Veröffentlicht: (2026)
EchoLSTM: A Self-Reflective Recurrent Network for Stabilizing Long-Range Memory
von: K, Prasanth K, et al.
Veröffentlicht: (2025)
von: K, Prasanth K, et al.
Veröffentlicht: (2025)
CellARC: Measuring Intelligence with Cellular Automata
von: Lžičař, Miroslav
Veröffentlicht: (2025)
von: Lžičař, Miroslav
Veröffentlicht: (2025)
Neural Reasoning Networks: Efficient Interpretable Neural Networks With Automatic Textual Explanations
von: Carrow, Stephen, et al.
Veröffentlicht: (2024)
von: Carrow, Stephen, et al.
Veröffentlicht: (2024)
Is Cambodia the World's Largest Cashew Producer?
von: Chaya, Veasna, et al.
Veröffentlicht: (2024)
von: Chaya, Veasna, et al.
Veröffentlicht: (2024)
A generalised novel loss function for computational fluid dynamics
von: Cooper-Baldock, Zachary, et al.
Veröffentlicht: (2024)
von: Cooper-Baldock, Zachary, et al.
Veröffentlicht: (2024)
SigGate-GT: Taming Over-Smoothing in Graph Transformers via Sigmoid-Gated Attention
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
VDW-GNNs: Vector diffusion wavelets for geometric graph neural networks
von: Johnson, David R., et al.
Veröffentlicht: (2025)
von: Johnson, David R., et al.
Veröffentlicht: (2025)
torchsom: The Reference PyTorch Library for Self-Organizing Maps
von: Berthier, Louis, et al.
Veröffentlicht: (2025)
von: Berthier, Louis, et al.
Veröffentlicht: (2025)
Neural Velocity for hyperparameter tuning
von: Dalmasso, Gianluca, et al.
Veröffentlicht: (2025)
von: Dalmasso, Gianluca, et al.
Veröffentlicht: (2025)
On the Equivalence of Regression and Classification
von: Jayadeva, et al.
Veröffentlicht: (2025)
von: Jayadeva, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Towards understanding how attention mechanism works in deep learning
von: Ruan, Tianyu, et al.
Veröffentlicht: (2024) -
CVCM Track Circuits Pre-emptive Failure Diagnostics for Predictive Maintenance Using Deep Neural Networks
von: Mukherjee, Debdeep, et al.
Veröffentlicht: (2025) -
multivariateGPT: a decoder-only transformer for multivariate categorical and numeric data
von: Loza, Andrew J., et al.
Veröffentlicht: (2025) -
DecompKAN: Decomposed Patch-KAN for Long-Term Time Series Forecasting
von: Mysore, Naveen
Veröffentlicht: (2026) -
Scaling Laws in the Tiny Regime: How Small Models Change Their Mistakes
von: Alnemari, Mohammed, et al.
Veröffentlicht: (2026)