Collapse-Free Prototype Readout Layer for Transformer Encoders
Fuente:
arXiv
Saved in:
| Main Authors: | Cirrincione, Giansalvo, Kumar, Rahul Ranjeev |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DDCL-INCRT: A Self-Organising Transformer with Hierarchical Prototype Structure (Theoretical Foundations)
by: Cirrincione, Giansalvo
Published: (2026)
by: Cirrincione, Giansalvo
Published: (2026)
DDCL: Deep Dual Competitive Learning: A Differentiable End-to-End Framework for Unsupervised Prototype-Based Representation Learning
by: Cirrincione, Giansalvo
Published: (2026)
by: Cirrincione, Giansalvo
Published: (2026)
INCRT: An Incremental Transformer That Determines Its Own Architecture
by: Cirrincione, Giansalvo
Published: (2026)
by: Cirrincione, Giansalvo
Published: (2026)
Hierarchical Kernel Transformer: Multi-Scale Attention with an Information-Theoretic Approximation Analysis
by: Cirrincione, Giansalvo
Published: (2026)
by: Cirrincione, Giansalvo
Published: (2026)
Efficient Deep Spiking Multi-Layer Perceptrons with Multiplication-Free Inference
by: Li, Boyan, et al.
Published: (2023)
by: Li, Boyan, et al.
Published: (2023)
Vertical Federated Continual Learning via Evolving Prototype Knowledge
by: Wang, Shuo, et al.
Published: (2025)
by: Wang, Shuo, et al.
Published: (2025)
Modeling the Human Visual System: Comparative Insights from Response-Optimized and Task-Optimized Vision Models, Language Models, and different Readout Mechanisms
by: Saha, Shreya, et al.
Published: (2024)
by: Saha, Shreya, et al.
Published: (2024)
Rank, Head-Channel Non-Identifiability, and Symmetry Breaking: A Precise Analysis of Representational Collapse in Transformers
by: Cirrincione, Giansalvo
Published: (2026)
by: Cirrincione, Giansalvo
Published: (2026)
Representation Learning in a Decomposed Encoder Design for Bio-inspired Hebbian Learning
by: Jaziri, Achref, et al.
Published: (2023)
by: Jaziri, Achref, et al.
Published: (2023)
Prototype-based interpretation of the functionality of neurons in winner-take-all neural networks
by: Sabzevar, Ramin Zarei, et al.
Published: (2020)
by: Sabzevar, Ramin Zarei, et al.
Published: (2020)
STAL: Spike Threshold Adaptive Learning Encoder for Classification of Pain-Related Biosignal Data
by: Hens, Freek, et al.
Published: (2024)
by: Hens, Freek, et al.
Published: (2024)
Mixture of Experts with Soft Nearest Neighbor Loss: Resolving Expert Collapse via Representation Disentanglement
by: Agarap, Abien Fred, et al.
Published: (2026)
by: Agarap, Abien Fred, et al.
Published: (2026)
Approximating Matrix Functions with Deep Neural Networks and Transformers
by: Padmanabhan, Rahul, et al.
Published: (2026)
by: Padmanabhan, Rahul, et al.
Published: (2026)
Rethinking Recurrent Neural Networks for Time Series Forecasting: A Reinforced Recurrent Encoder with Prediction-Oriented Proximal Policy Optimization
by: Lai, Xin, et al.
Published: (2026)
by: Lai, Xin, et al.
Published: (2026)
Univariate Radial Basis Function Layers: Brain-inspired Deep Neural Layers for Low-Dimensional Inputs
by: Jost, Daniel, et al.
Published: (2023)
by: Jost, Daniel, et al.
Published: (2023)
Empirical Analysis of Nature-Inspired Algorithms for Autism Spectrum Disorder Detection Using 3D Video Dataset
by: Panchal, Aneesh, et al.
Published: (2025)
by: Panchal, Aneesh, et al.
Published: (2025)
A Hormone-inspired Emotion Layer for Transformer language models (HELT)
by: Reda, Eslam, et al.
Published: (2026)
by: Reda, Eslam, et al.
Published: (2026)
Prototype Analysis in Hopfield Networks with Hebbian Learning
by: McAlister, Hayden, et al.
Published: (2024)
by: McAlister, Hayden, et al.
Published: (2024)
FGATT: A Robust Framework for Wireless Data Imputation Using Fuzzy Graph Attention Networks and Transformer Encoders
by: Xing, Jinming, et al.
Published: (2024)
by: Xing, Jinming, et al.
Published: (2024)
Deep Neural Regression Collapse
by: Rangamani, Akshay, et al.
Published: (2026)
by: Rangamani, Akshay, et al.
Published: (2026)
Affinity Is Not Enough: Recovering the Free Energy Principle in Mixture-of-Experts
by: Wong, Man Yung
Published: (2026)
by: Wong, Man Yung
Published: (2026)
Looped Transformers are Better at Learning Learning Algorithms
by: Yang, Liu, et al.
Published: (2023)
by: Yang, Liu, et al.
Published: (2023)
A Backpropagation-Free Feedback-Hebbian Network for Continual Learning Dynamics
by: Li, Josh, et al.
Published: (2026)
by: Li, Josh, et al.
Published: (2026)
Reward-Modulated Local Learning in Spiking Encoders: Controlled Benchmarks with STDP and Hybrid Rate Readouts
by: Chakraborty, Debjyoti
Published: (2026)
by: Chakraborty, Debjyoti
Published: (2026)
Utilizing Evolution Strategies to Train Transformers in Reinforcement Learning
by: Lorenc, Matyáš, et al.
Published: (2025)
by: Lorenc, Matyáš, et al.
Published: (2025)
LeanKAN: A Parameter-Lean Kolmogorov-Arnold Network Layer with Improved Memory Efficiency and Convergence Behavior
by: Koenig, Benjamin C., et al.
Published: (2025)
by: Koenig, Benjamin C., et al.
Published: (2025)
Genetic Programming with Transformer-Based Mutation for Approximate Circuit Design
by: Galeta, Ondrej, et al.
Published: (2026)
by: Galeta, Ondrej, et al.
Published: (2026)
Utilizing Novelty-based Evolution Strategies to Train Transformers in Reinforcement Learning
by: Lorenc, Matyáš, et al.
Published: (2025)
by: Lorenc, Matyáš, et al.
Published: (2025)
Transformer Semantic Genetic Programming for d-dimensional Symbolic Regression Problems
by: Anthes, Philipp, et al.
Published: (2025)
by: Anthes, Philipp, et al.
Published: (2025)
Prototype-Grounded Concept Models for Verifiable Concept Alignment
by: Colamonaco, Stefano, et al.
Published: (2026)
by: Colamonaco, Stefano, et al.
Published: (2026)
Transforming Datasets to Requested Complexity with Projection-based Many-Objective Genetic Algorithm
by: Komorniczak, Joanna
Published: (2025)
by: Komorniczak, Joanna
Published: (2025)
SpikeGraphormer: A High-Performance Graph Transformer with Spiking Graph Attention
by: Sun, Yundong, et al.
Published: (2024)
by: Sun, Yundong, et al.
Published: (2024)
Indian Wedding System Optimization (IWSO): A Novel Socially Inspired Metaheuristic with Operational Design and Analysis
by: Saxena, Deepika, et al.
Published: (2026)
by: Saxena, Deepika, et al.
Published: (2026)
Sample-based Dynamic Hierarchical Transformer with Layer and Head Flexibility via Contextual Bandit
by: Meng, Fanfei, et al.
Published: (2023)
by: Meng, Fanfei, et al.
Published: (2023)
SpikingGamma: Surrogate-Gradient Free and Temporally Precise Online Training of Spiking Neural Networks with Smoothed Delays
by: Koopman, Roel, et al.
Published: (2026)
by: Koopman, Roel, et al.
Published: (2026)
GridPE: Unifying Positional Encoding in Transformers with a Grid Cell-Inspired Framework
by: Li, Boyang, et al.
Published: (2024)
by: Li, Boyang, et al.
Published: (2024)
GRASP: GRouped Activation Shared Parameterization for Parameter-Efficient Fine-Tuning and Robust Inference of Transformers
by: Bal, Malyaban, et al.
Published: (2025)
by: Bal, Malyaban, et al.
Published: (2025)
Spatiotemporal Forecasting of Traffic Flow using Wavelet-based Temporal Attention
by: Jakhmola, Yash, et al.
Published: (2024)
by: Jakhmola, Yash, et al.
Published: (2024)
GT-SNT: A Linear-Time Transformer for Large-Scale Graphs via Spiking Node Tokenization
by: Zhang, Huizhe, et al.
Published: (2025)
by: Zhang, Huizhe, et al.
Published: (2025)
Differential Evolution Algorithm based Hyper-Parameters Selection of Transformer Neural Network Model for Load Forecasting
by: Sen, Anuvab, et al.
Published: (2023)
by: Sen, Anuvab, et al.
Published: (2023)
Similar Items
-
DDCL-INCRT: A Self-Organising Transformer with Hierarchical Prototype Structure (Theoretical Foundations)
by: Cirrincione, Giansalvo
Published: (2026) -
DDCL: Deep Dual Competitive Learning: A Differentiable End-to-End Framework for Unsupervised Prototype-Based Representation Learning
by: Cirrincione, Giansalvo
Published: (2026) -
INCRT: An Incremental Transformer That Determines Its Own Architecture
by: Cirrincione, Giansalvo
Published: (2026) -
Hierarchical Kernel Transformer: Multi-Scale Attention with an Information-Theoretic Approximation Analysis
by: Cirrincione, Giansalvo
Published: (2026) -
Efficient Deep Spiking Multi-Layer Perceptrons with Multiplication-Free Inference
by: Li, Boyan, et al.
Published: (2023)