Grokking Beyond the Euclidean Norm of Model Parameters
Fuente:
arXiv
Saved in:
| Main Authors: | Notsawo, Pascal Jr Tikeng, Dumas, Guillaume, Rabusseau, Guillaume |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Grokking Finite-Dimensional Algebra
by: Notsawo, Pascal Jr Tikeng, et al.
Published: (2026)
by: Notsawo, Pascal Jr Tikeng, et al.
Published: (2026)
Benefits and Limitations of Communication in Multi-Agent Reasoning
by: Rizvi-Martel, Michael, et al.
Published: (2025)
by: Rizvi-Martel, Michael, et al.
Published: (2025)
The Geometric Inductive Bias of Grokking: Bypassing Phase Transitions via Architectural Topology
by: Yıldırım, Alper
Published: (2026)
by: Yıldırım, Alper
Published: (2026)
Deep Memory Search: A Metaheuristic Approach for Optimizing Heuristic Search
by: Hedar, Abdel-Rahman, et al.
Published: (2024)
by: Hedar, Abdel-Rahman, et al.
Published: (2024)
Why Online Reinforcement Learning is Causal
by: Schulte, Oliver, et al.
Published: (2024)
by: Schulte, Oliver, et al.
Published: (2024)
Learning Agents With Prioritization and Parameter Noise in Continuous State and Action Space
by: Mangannavar, Rajesh, et al.
Published: (2024)
by: Mangannavar, Rajesh, et al.
Published: (2024)
A Machine Learning Framework for Turbofan Health Estimation via Inverse Problem Formulation
by: Leyli-Abadi, Milad, et al.
Published: (2026)
by: Leyli-Abadi, Milad, et al.
Published: (2026)
Grokking in the Wild: Data Augmentation for Real-World Multi-Hop Reasoning with Transformers
by: Abramov, Roman, et al.
Published: (2025)
by: Abramov, Roman, et al.
Published: (2025)
Large Language Models as Attribution Regularizers for Efficient Model Training
by: Vukadin, Davor, et al.
Published: (2025)
by: Vukadin, Davor, et al.
Published: (2025)
Approximate Domain Unlearning for Vision-Language Models
by: Kawamura, Kodai, et al.
Published: (2025)
by: Kawamura, Kodai, et al.
Published: (2025)
Evaluating Model Explanations without Ground Truth
by: Rawal, Kaivalya, et al.
Published: (2025)
by: Rawal, Kaivalya, et al.
Published: (2025)
Scaling Offline RL via Efficient and Expressive Shortcut Models
by: Espinosa-Dice, Nicolas, et al.
Published: (2025)
by: Espinosa-Dice, Nicolas, et al.
Published: (2025)
Architectural Proprioception in State Space Models: Thermodynamic Training Induces Anticipatory Halt Detection
by: Noon, Jay
Published: (2026)
by: Noon, Jay
Published: (2026)
Black Box Model Explanations and the Human Interpretability Expectations -- An Analysis in the Context of Homicide Prediction
by: Ribeiro, José, et al.
Published: (2022)
by: Ribeiro, José, et al.
Published: (2022)
A Practical Approach to using Supervised Machine Learning Models to Classify Aviation Safety Occurrences
by: Siow, Bryan Y.
Published: (2025)
by: Siow, Bryan Y.
Published: (2025)
A Systematic Evaluation of Euclidean Alignment with Deep Learning for EEG Decoding
by: Junqueira, Bruna, et al.
Published: (2024)
by: Junqueira, Bruna, et al.
Published: (2024)
Spectral Compact Training: Pre-Training Large Language Models via Permanent Truncated SVD and Stiefel QR Retraction
by: Kohlberger, Björn Roman
Published: (2026)
by: Kohlberger, Björn Roman
Published: (2026)
Model Fusion via Retrofitting
by: Luenam, Phoomraphee, et al.
Published: (2025)
by: Luenam, Phoomraphee, et al.
Published: (2025)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
by: Fadli, Samih
Published: (2025)
by: Fadli, Samih
Published: (2025)
AGWM: Affordance-Grounded World Models for Environments with Compositional Prerequisites
by: Zhang, Qinshi, et al.
Published: (2026)
by: Zhang, Qinshi, et al.
Published: (2026)
The Geometry of Thought: How Scale Restructures Reasoning In Large Language Models
by: Anderson, Samuel Cyrenius
Published: (2026)
by: Anderson, Samuel Cyrenius
Published: (2026)
On Semantic Loss Fine-Tuning Approach for Preventing Model Collapse in Causal Reasoning
by: Deshmukh, Pratik, et al.
Published: (2026)
by: Deshmukh, Pratik, et al.
Published: (2026)
Reasoning Large Language Model Errors Arise from Hallucinating Critical Problem Features
by: Heyman, Alex, et al.
Published: (2025)
by: Heyman, Alex, et al.
Published: (2025)
HEHRGNN: A Unified Embedding Model for Knowledge Graphs with Hyperedges and Hyper-Relational Edges
by: Rajagopalamenon, Rajesh, et al.
Published: (2026)
by: Rajagopalamenon, Rajesh, et al.
Published: (2026)
Quantization Undoes Alignment: Bias Emergence in Compressed LLMs Across Models and Precision Levels
by: Rath, Plawan Kumar, et al.
Published: (2026)
by: Rath, Plawan Kumar, et al.
Published: (2026)
Beyond Hallucinations: A Composite Score for Measuring Reliability in Open-Source Large Language Models
by: Salla, Rohit Kumar, et al.
Published: (2025)
by: Salla, Rohit Kumar, et al.
Published: (2025)
Fusion-Based Neural Generalization for Predicting Temperature Fields in Industrial PET Preform Heating
by: Alsheikh, Ahmad, et al.
Published: (2025)
by: Alsheikh, Ahmad, et al.
Published: (2025)
The Anti-Ouroboros Effect: Emergent Resilience in Large Language Models from Recursive Selective Feedback
by: Adapala, Sai Teja Reddy
Published: (2025)
by: Adapala, Sai Teja Reddy
Published: (2025)
What Do World Models Learn in RL? Probing Latent Representations in Learned Environment Simulators
by: Zhang, Xinyu
Published: (2026)
by: Zhang, Xinyu
Published: (2026)
I-GLIDE: Input Groups for Latent Health Indicators in Degradation Estimation
by: Thil, Lucas, et al.
Published: (2025)
by: Thil, Lucas, et al.
Published: (2025)
TACO: Tackling Over-correction in Federated Learning with Tailored Adaptive Correction
by: Liu, Weijie, et al.
Published: (2025)
by: Liu, Weijie, et al.
Published: (2025)
DataRater: Meta-Learned Dataset Curation
by: Calian, Dan A., et al.
Published: (2025)
by: Calian, Dan A., et al.
Published: (2025)
The Lattice Geometry of Neural Network Quantization -- A Short Equivalence Proof of GPTQ and Babai's Algorithm
by: Birnick, Johann
Published: (2025)
by: Birnick, Johann
Published: (2025)
Unlearning's Blind Spots: Over-Unlearning and Prototypical Relearning Attack
by: Ha, SeungBum, et al.
Published: (2025)
by: Ha, SeungBum, et al.
Published: (2025)
Residual Reservoir Memory Networks
by: Pinna, Matteo, et al.
Published: (2025)
by: Pinna, Matteo, et al.
Published: (2025)
FreRA: A Frequency-Refined Augmentation for Contrastive Learning on Time Series Classification
by: Tian, Tian, et al.
Published: (2025)
by: Tian, Tian, et al.
Published: (2025)
Deep Residual Echo State Networks: exploring residual orthogonal connections in untrained Recurrent Neural Networks
by: Pinna, Matteo, et al.
Published: (2025)
by: Pinna, Matteo, et al.
Published: (2025)
1 bit is all we need: binary normalized neural networks
by: Cabral, Eduardo Lobo Lustoda, et al.
Published: (2025)
by: Cabral, Eduardo Lobo Lustoda, et al.
Published: (2025)
Multi-Task Reinforcement Learning with Language-Encoded Gated Policy Networks
by: Arora, Rushiv
Published: (2025)
by: Arora, Rushiv
Published: (2025)
Global-Order GFlowNets
by: Pastor-Pérez, Lluís, et al.
Published: (2025)
by: Pastor-Pérez, Lluís, et al.
Published: (2025)
Similar Items
-
Grokking Finite-Dimensional Algebra
by: Notsawo, Pascal Jr Tikeng, et al.
Published: (2026) -
Benefits and Limitations of Communication in Multi-Agent Reasoning
by: Rizvi-Martel, Michael, et al.
Published: (2025) -
The Geometric Inductive Bias of Grokking: Bypassing Phase Transitions via Architectural Topology
by: Yıldırım, Alper
Published: (2026) -
Deep Memory Search: A Metaheuristic Approach for Optimizing Heuristic Search
by: Hedar, Abdel-Rahman, et al.
Published: (2024) -
Why Online Reinforcement Learning is Causal
by: Schulte, Oliver, et al.
Published: (2024)