Kan Extension Transformers: A Categorical Unification of Attention, Diffusion, and Predict-Detach Self-Conditioning
Fuente:
arXiv
Saved in:
| Main Author: | Mahadevan, Sridhar |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GAIA: Categorical Foundations of Generative AI
by: Mahadevan, Sridhar
Published: (2024)
by: Mahadevan, Sridhar
Published: (2024)
Universal Decision Learners
by: Mahadevan, Sridhar
Published: (2026)
by: Mahadevan, Sridhar
Published: (2026)
Learning Is a Kan Extension
by: Pugh, Matthew, et al.
Published: (2025)
by: Pugh, Matthew, et al.
Published: (2025)
Universal Reinforcement Learning in Coalgebras: Asynchronous Stochastic Computation via Conduction
by: Mahadevan, Sridhar
Published: (2025)
by: Mahadevan, Sridhar
Published: (2025)
Consciousness as a Functor
by: Mahadevan, Sridhar
Published: (2025)
by: Mahadevan, Sridhar
Published: (2025)
Universal Imitation Games
by: Mahadevan, Sridhar
Published: (2024)
by: Mahadevan, Sridhar
Published: (2024)
CSQL: Mapping Documents into Causal Databases
by: Mahadevan, Sridhar
Published: (2026)
by: Mahadevan, Sridhar
Published: (2026)
Self-Attention as a Parametric Endofunctor: A Categorical Framework for Transformer Architectures
by: O'Neill, Charles
Published: (2025)
by: O'Neill, Charles
Published: (2025)
A Unification of Discrete, Gaussian, and Simplicial Diffusion
by: Chandra, Nuria Alina, et al.
Published: (2025)
by: Chandra, Nuria Alina, et al.
Published: (2025)
A Rose by Any Other Name Would Smell as Sweet: Categorical Homotopy Theory for Large Language Models
by: Mahadevan, Sridhar
Published: (2025)
by: Mahadevan, Sridhar
Published: (2025)
Crossfusor: A Cross-Attention Transformer Enhanced Conditional Diffusion Model for Car-Following Trajectory Prediction
by: You, Junwei, et al.
Published: (2024)
by: You, Junwei, et al.
Published: (2024)
Conditional Diffusion Modeling with Attention for Probabilistic Battery Capacity Prediction under Real-World Condition
by: Jiang, Chunlin, et al.
Published: (2025)
by: Jiang, Chunlin, et al.
Published: (2025)
Spectral Conditioning of Attention Improves Transformer Performance
by: Saratchandran, Hemanth, et al.
Published: (2026)
by: Saratchandran, Hemanth, et al.
Published: (2026)
Categorical Reparameterization with Denoising Diffusion models
by: Gourevitch, Samson, et al.
Published: (2026)
by: Gourevitch, Samson, et al.
Published: (2026)
Unified Discrete Diffusion for Categorical Data
by: Zhao, Lingxiao, et al.
Published: (2024)
by: Zhao, Lingxiao, et al.
Published: (2024)
On Understanding Attention-Based In-Context Learning for Categorical Data
by: Wang, Aaron T., et al.
Published: (2024)
by: Wang, Aaron T., et al.
Published: (2024)
DiffER: Categorical Diffusion for Chemical Retrosynthesis
by: Current, Sean, et al.
Published: (2025)
by: Current, Sean, et al.
Published: (2025)
Can Kans (re)discover predictive models for Direct-Drive Laser Fusion?
by: Ejaz, Rahman, et al.
Published: (2024)
by: Ejaz, Rahman, et al.
Published: (2024)
DGTN: Graph-Enhanced Transformer with Diffusive Attention Gating Mechanism for Enzyme DDG Prediction
by: Lin, Abigail
Published: (2025)
by: Lin, Abigail
Published: (2025)
Understanding Differential Transformer Unchains Pretrained Self-Attentions
by: Kong, Chaerin, et al.
Published: (2025)
by: Kong, Chaerin, et al.
Published: (2025)
Self-Attention as Distributional Projection: A Unified Interpretation of Transformer Architecture
by: Mehta, Nihal
Published: (2025)
by: Mehta, Nihal
Published: (2025)
Continuously Augmented Discrete Diffusion model for Categorical Generative Modeling
by: Zheng, Huangjie, et al.
Published: (2025)
by: Zheng, Huangjie, et al.
Published: (2025)
Graph Convolutions Enrich the Self-Attention in Transformers!
by: Choi, Jeongwhan, et al.
Published: (2023)
by: Choi, Jeongwhan, et al.
Published: (2023)
Graph Diffusion Transformers for Multi-Conditional Molecular Generation
by: Liu, Gang, et al.
Published: (2024)
by: Liu, Gang, et al.
Published: (2024)
Categorical Distributions are Effective Neural Network Outputs for Event Prediction
by: Doran, Kevin, et al.
Published: (2025)
by: Doran, Kevin, et al.
Published: (2025)
Attention-guided Spectrogram Sequence Modeling with CNNs for Music Genre Classification
by: Sridhar, Aditya
Published: (2024)
by: Sridhar, Aditya
Published: (2024)
CSAI: Conditional Self-Attention Imputation for Healthcare Time-series
by: Qian, Linglong, et al.
Published: (2023)
by: Qian, Linglong, et al.
Published: (2023)
NoiseFormer -- Noise Diffused Symmetric Attention Transformer
by: Kumar, Phani, et al.
Published: (2026)
by: Kumar, Phani, et al.
Published: (2026)
CAST: Clustering Self-Attention using Surrogate Tokens for Efficient Transformers
by: van Engelenhoven, Adjorn, et al.
Published: (2024)
by: van Engelenhoven, Adjorn, et al.
Published: (2024)
Attention Sinks Are Provably Necessary in Softmax Transformers: Evidence from Trigger-Conditional Tasks
by: Ran-Milo, Yuval
Published: (2026)
by: Ran-Milo, Yuval
Published: (2026)
Multistability of Self-Attention Dynamics in Transformers
by: Altafini, Claudio
Published: (2025)
by: Altafini, Claudio
Published: (2025)
Transformers for Tabular Data: A Training Perspective of Self-Attention via Optimal Transport
by: Quadrio, Alessandro, et al.
Published: (2025)
by: Quadrio, Alessandro, et al.
Published: (2025)
Encoding Predictability and Legibility for Style-Conditioned Diffusion Policy
by: Crétides, Adrien Jacquet, et al.
Published: (2026)
by: Crétides, Adrien Jacquet, et al.
Published: (2026)
Unification and Optimization of Robust Supervised Learning
by: Hanselle, Jonas, et al.
Published: (2026)
by: Hanselle, Jonas, et al.
Published: (2026)
Quantum Adaptive Self-Attention for Quantum Transformer Models
by: Chen, Chi-Sheng, et al.
Published: (2025)
by: Chen, Chi-Sheng, et al.
Published: (2025)
Uncertainty Estimation of Transformers' Predictions via Topological Analysis of the Attention Matrices
by: Kostenok, Elizaveta, et al.
Published: (2023)
by: Kostenok, Elizaveta, et al.
Published: (2023)
Budgeted Attention Allocation: Cost-Conditioned Compute Control for Efficient Transformers
by: Nidhi, Amrit
Published: (2026)
by: Nidhi, Amrit
Published: (2026)
FlashMask: Efficient and Rich Mask Extension of FlashAttention
by: Wang, Guoxia, et al.
Published: (2024)
by: Wang, Guoxia, et al.
Published: (2024)
Label Attention Network for Temporal Sets Prediction: You Were Looking at a Wrong Self-Attention
by: Kovtun, Elizaveta, et al.
Published: (2023)
by: Kovtun, Elizaveta, et al.
Published: (2023)
Simple Self-Conditioning Adaptation for Masked Diffusion Models
by: Cardei, Michael, et al.
Published: (2026)
by: Cardei, Michael, et al.
Published: (2026)
Similar Items
-
GAIA: Categorical Foundations of Generative AI
by: Mahadevan, Sridhar
Published: (2024) -
Universal Decision Learners
by: Mahadevan, Sridhar
Published: (2026) -
Learning Is a Kan Extension
by: Pugh, Matthew, et al.
Published: (2025) -
Universal Reinforcement Learning in Coalgebras: Asynchronous Stochastic Computation via Conduction
by: Mahadevan, Sridhar
Published: (2025) -
Consciousness as a Functor
by: Mahadevan, Sridhar
Published: (2025)