Understanding Generalization, Robustness, and Interpretability in Low-Capacity Neural Networks
Fuente:
arXiv
Guardado en:
| Autor principal: | Kumar, Yash |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Towards understanding how attention mechanism works in deep learning
por: Ruan, Tianyu, et al.
Publicado: (2024)
por: Ruan, Tianyu, et al.
Publicado: (2024)
CVCM Track Circuits Pre-emptive Failure Diagnostics for Predictive Maintenance Using Deep Neural Networks
por: Mukherjee, Debdeep, et al.
Publicado: (2025)
por: Mukherjee, Debdeep, et al.
Publicado: (2025)
multivariateGPT: a decoder-only transformer for multivariate categorical and numeric data
por: Loza, Andrew J., et al.
Publicado: (2025)
por: Loza, Andrew J., et al.
Publicado: (2025)
DecompKAN: Decomposed Patch-KAN for Long-Term Time Series Forecasting
por: Mysore, Naveen
Publicado: (2026)
por: Mysore, Naveen
Publicado: (2026)
Scaling Laws in the Tiny Regime: How Small Models Change Their Mistakes
por: Alnemari, Mohammed, et al.
Publicado: (2026)
por: Alnemari, Mohammed, et al.
Publicado: (2026)
TGraphX: Tensor-Aware Graph Neural Network for Multi-Dimensional Feature Learning
por: Sajjadi, Arash, et al.
Publicado: (2025)
por: Sajjadi, Arash, et al.
Publicado: (2025)
Cross-Architecture Knowledge Distillation (KD) for Retinal Fundus Image Anomaly Detection on NVIDIA Jetson Nano
por: Yilmaz, Berk, et al.
Publicado: (2025)
por: Yilmaz, Berk, et al.
Publicado: (2025)
The Normalized Difference Layer: A Differentiable Spectral Index Formulation for Deep Learning
por: Lotfi, Ali, et al.
Publicado: (2026)
por: Lotfi, Ali, et al.
Publicado: (2026)
Efficient Morphology-Control Co-Design via Stackelberg Proximal Policy Optimization
por: Dai, Yanning, et al.
Publicado: (2026)
por: Dai, Yanning, et al.
Publicado: (2026)
Optimized Gradient Clipping for Noisy Label Learning
por: Ye, Xichen, et al.
Publicado: (2024)
por: Ye, Xichen, et al.
Publicado: (2024)
Pulse-Driven Neural Architecture: Learnable Oscillatory Dynamics for Robust Continuous-Time Sequence Processing
por: Sharma, Paras
Publicado: (2026)
por: Sharma, Paras
Publicado: (2026)
PolyGLU: State-Conditional Activation Routing in Transformer Feed-Forward Networks
por: Medeiros, Daniel Nobrega
Publicado: (2026)
por: Medeiros, Daniel Nobrega
Publicado: (2026)
Neural Encoding for Image Recall: Human-Like Memory
por: Foussereau, Virgile, et al.
Publicado: (2024)
por: Foussereau, Virgile, et al.
Publicado: (2024)
GraphNNK -- Graph Classification and Interpretability
por: Bolevic, Zeljko, et al.
Publicado: (2026)
por: Bolevic, Zeljko, et al.
Publicado: (2026)
Explicit Dropout: Deterministic Regularization for Transformer Architectures
por: Agrawal, Vidhi, et al.
Publicado: (2026)
por: Agrawal, Vidhi, et al.
Publicado: (2026)
Massive Redundancy in Gradient Transport Enables Sparse Online Learning
por: Merin, Aur Shalev
Publicado: (2026)
por: Merin, Aur Shalev
Publicado: (2026)
Optimizing Inference in Transformer-Based Models: A Multi-Method Benchmark
por: Ho, Siu Hang, et al.
Publicado: (2025)
por: Ho, Siu Hang, et al.
Publicado: (2025)
Rethinking Visual Intelligence: Insights from Video Pretraining
por: Acuaviva, Pablo, et al.
Publicado: (2025)
por: Acuaviva, Pablo, et al.
Publicado: (2025)
Revisiting GAN with Bayes-Optimal Discrimination
por: Naeini, Mohammadreza Tavasoli, et al.
Publicado: (2025)
por: Naeini, Mohammadreza Tavasoli, et al.
Publicado: (2025)
A Hybrid Inductive-Transductive Network for Traffic Flow Imputation on Unsampled Locations
por: Rahimiasl, Mohammadmahdi, et al.
Publicado: (2025)
por: Rahimiasl, Mohammadmahdi, et al.
Publicado: (2025)
JacNet: Learning Functions with Structured Jacobians
por: Lorraine, Jonathan, et al.
Publicado: (2024)
por: Lorraine, Jonathan, et al.
Publicado: (2024)
Machine Unlearning for Class Removal through SISA-based Deep Neural Network Architectures
por: Mahi, Ishrak Hamim, et al.
Publicado: (2026)
por: Mahi, Ishrak Hamim, et al.
Publicado: (2026)
H-Model: Dynamic Neural Architectures for Adaptive Processing
por: Hospodarchuk, Dmytro
Publicado: (2025)
por: Hospodarchuk, Dmytro
Publicado: (2025)
Predictive Modeling of Maritime Radar Data Using Transformer Architecture
por: Qesaraku, Bjorna, et al.
Publicado: (2025)
por: Qesaraku, Bjorna, et al.
Publicado: (2025)
Revisiting Non-separable Binary Classification and its Applications in Anomaly Detection
por: Lau, Matthew, et al.
Publicado: (2023)
por: Lau, Matthew, et al.
Publicado: (2023)
Predicting Traffic Accident Severity with Deep Neural Networks
por: Bibb, Meghan, et al.
Publicado: (2025)
por: Bibb, Meghan, et al.
Publicado: (2025)
Kolmogorov-Arnold Attention: Is Learnable Attention Better For Vision Transformers?
por: Maity, Subhajit, et al.
Publicado: (2025)
por: Maity, Subhajit, et al.
Publicado: (2025)
Scalable, Technology-Agnostic Diagnosis and Predictive Maintenance for Point Machine using Deep Learning
por: Di Santi, Eduardo, et al.
Publicado: (2025)
por: Di Santi, Eduardo, et al.
Publicado: (2025)
Closing the Theory-Practice Gap in Spiking Transformers via Effective Dimension
por: Guo, Dongxin, et al.
Publicado: (2026)
por: Guo, Dongxin, et al.
Publicado: (2026)
Temporal Functional Circuits: From Spline Plots to Faithful Explanations in KAN Forecasting
por: Mysore, Naveen
Publicado: (2026)
por: Mysore, Naveen
Publicado: (2026)
EchoLSTM: A Self-Reflective Recurrent Network for Stabilizing Long-Range Memory
por: K, Prasanth K, et al.
Publicado: (2025)
por: K, Prasanth K, et al.
Publicado: (2025)
CellARC: Measuring Intelligence with Cellular Automata
por: Lžičař, Miroslav
Publicado: (2025)
por: Lžičař, Miroslav
Publicado: (2025)
Neural Reasoning Networks: Efficient Interpretable Neural Networks With Automatic Textual Explanations
por: Carrow, Stephen, et al.
Publicado: (2024)
por: Carrow, Stephen, et al.
Publicado: (2024)
Is Cambodia the World's Largest Cashew Producer?
por: Chaya, Veasna, et al.
Publicado: (2024)
por: Chaya, Veasna, et al.
Publicado: (2024)
A generalised novel loss function for computational fluid dynamics
por: Cooper-Baldock, Zachary, et al.
Publicado: (2024)
por: Cooper-Baldock, Zachary, et al.
Publicado: (2024)
SigGate-GT: Taming Over-Smoothing in Graph Transformers via Sigmoid-Gated Attention
por: Guo, Dongxin, et al.
Publicado: (2026)
por: Guo, Dongxin, et al.
Publicado: (2026)
VDW-GNNs: Vector diffusion wavelets for geometric graph neural networks
por: Johnson, David R., et al.
Publicado: (2025)
por: Johnson, David R., et al.
Publicado: (2025)
torchsom: The Reference PyTorch Library for Self-Organizing Maps
por: Berthier, Louis, et al.
Publicado: (2025)
por: Berthier, Louis, et al.
Publicado: (2025)
Neural Velocity for hyperparameter tuning
por: Dalmasso, Gianluca, et al.
Publicado: (2025)
por: Dalmasso, Gianluca, et al.
Publicado: (2025)
On the Equivalence of Regression and Classification
por: Jayadeva, et al.
Publicado: (2025)
por: Jayadeva, et al.
Publicado: (2025)
Ejemplares similares
-
Towards understanding how attention mechanism works in deep learning
por: Ruan, Tianyu, et al.
Publicado: (2024) -
CVCM Track Circuits Pre-emptive Failure Diagnostics for Predictive Maintenance Using Deep Neural Networks
por: Mukherjee, Debdeep, et al.
Publicado: (2025) -
multivariateGPT: a decoder-only transformer for multivariate categorical and numeric data
por: Loza, Andrew J., et al.
Publicado: (2025) -
DecompKAN: Decomposed Patch-KAN for Long-Term Time Series Forecasting
por: Mysore, Naveen
Publicado: (2026) -
Scaling Laws in the Tiny Regime: How Small Models Change Their Mistakes
por: Alnemari, Mohammed, et al.
Publicado: (2026)