Outlier-Efficient Hopfield Layers for Large Transformer-Based Models
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Jerry Yao-Chieh, Chang, Pei-Hsuan, Luo, Robin, Chen, Hong-Yu, Li, Weijian, Wang, Wei-Po, Liu, Han |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
STanHop: Sparse Tandem Hopfield Model for Memory-Enhanced Time Series Prediction
by: Wu, Dennis, et al.
Published: (2023)
by: Wu, Dennis, et al.
Published: (2023)
Nonparametric Modern Hopfield Models
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
Provably Optimal Memory Capacity for Modern Hopfield Models: Transformer-Compatible Dense Associative Memories as Spherical Codes
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
Uniform Memory Retrieval with Larger Capacity for Modern Hopfield Models
by: Wu, Dennis, et al.
Published: (2024)
by: Wu, Dennis, et al.
Published: (2024)
On Computational Limits of Modern Hopfield Models: A Fine-Grained Complexity Analysis
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
BiSHop: Bi-Directional Cellular Learning for Tabular Data with Generalized Sparse Modern Hopfield Model
by: Xu, Chenwei, et al.
Published: (2024)
by: Xu, Chenwei, et al.
Published: (2024)
Fast and Low-Cost Genomic Foundation Models via Outlier Removal
by: Luo, Haozheng, et al.
Published: (2025)
by: Luo, Haozheng, et al.
Published: (2025)
On Statistical Rates and Provably Efficient Criteria of Latent Diffusion Transformers (DiTs)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
Fundamental Limits of Prompt Tuning Transformers: Universality, Capacity and Efficiency
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
Transformer Approximations from ReLUs
by: Hu, Jerry Yao-Chieh, et al.
Published: (2026)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2026)
In-Context Deep Learning via Transformer Models
by: Wu, Weimin, et al.
Published: (2024)
by: Wu, Weimin, et al.
Published: (2024)
In-Context Algorithm Emulation in Fixed-Weight Transformers
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
VESTA: A Versatile SNN-Based Transformer Accelerator with Unified PEs for Multiple Computational Layers
by: Chen, Ching-Yao, et al.
Published: (2025)
by: Chen, Ching-Yao, et al.
Published: (2025)
Taming Outlier Tokens in Diffusion Transformers
by: Wu, Xiaoyu, et al.
Published: (2026)
by: Wu, Xiaoyu, et al.
Published: (2026)
Universal Approximation with Softmax Attention
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
Computational Limits of Low-Rank Adaptation (LoRA) Fine-Tuning for Transformer Models
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
On Statistical Rates of Conditional Diffusion Transformers: Approximation, Estimation and Minimax Optimality
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
An Exploratory Framework for Future SETI Applications: Detecting Generative Reactivity via Language Models
by: Yu, Po-Chieh
Published: (2025)
by: Yu, Po-Chieh
Published: (2025)
Attention Mechanism, Max-Affine Partition, and Universal Approximation
by: Liu, Hude, et al.
Published: (2025)
by: Liu, Hude, et al.
Published: (2025)
Error Exponents for Quantum Packing Problems via An Operator Layer Cake Theorem
by: Cheng, Hao-Chung, et al.
Published: (2025)
by: Cheng, Hao-Chung, et al.
Published: (2025)
Who Brings the Frisbee: Probing Hidden Hallucination Factors in Large Vision-Language Model via Causality Analysis
by: Huang, Po-Hsuan, et al.
Published: (2024)
by: Huang, Po-Hsuan, et al.
Published: (2024)
Mind the Inconspicuous: Revealing the Hidden Weakness in Aligned LLMs' Refusal Boundaries
by: Yu, Jiahao, et al.
Published: (2024)
by: Yu, Jiahao, et al.
Published: (2024)
Discrete Flow Matching Policy Optimization
by: Su, Maojiang, et al.
Published: (2026)
by: Su, Maojiang, et al.
Published: (2026)
On Flow Matching KL Divergence
by: Su, Maojiang, et al.
Published: (2025)
by: Su, Maojiang, et al.
Published: (2025)
Unveiling the Latent Directions of Reflection in Large Language Models
by: Chang, Fu-Chieh, et al.
Published: (2025)
by: Chang, Fu-Chieh, et al.
Published: (2025)
Residual 1D CNN for Low SFR Surface Density Regression: A Design Note
by: Yu, Po-Chieh
Published: (2025)
by: Yu, Po-Chieh
Published: (2025)
Layer Cake Representations for Quantum Divergences
by: Liu, Po-Chieh, et al.
Published: (2025)
by: Liu, Po-Chieh, et al.
Published: (2025)
A Theoretical Framework for OOD Robustness in Transformers using Gevrey Classes
by: Wang, Yu, et al.
Published: (2025)
by: Wang, Yu, et al.
Published: (2025)
Unraveling Arithmetic in Large Language Models: The Role of Algebraic Structures
by: Chang, Fu-Chieh, et al.
Published: (2024)
by: Chang, Fu-Chieh, et al.
Published: (2024)
Dense Hopfield Networks in the Teacher-Student Setting
by: Thériault, Robin, et al.
Published: (2024)
by: Thériault, Robin, et al.
Published: (2024)
Differentially Private Kernel Density Estimation
by: Liu, Erzhi, et al.
Published: (2024)
by: Liu, Erzhi, et al.
Published: (2024)
On Differentially Private String Distances
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
Decoupled Alignment for Robust Plug-and-Play Adaptation
by: Luo, Haozheng, et al.
Published: (2024)
by: Luo, Haozheng, et al.
Published: (2024)
Anti-Length Shift: Dynamic Outlier Truncation for Training Efficient Reasoning Models
by: Wu, Wei, et al.
Published: (2026)
by: Wu, Wei, et al.
Published: (2026)
Self‐knotted feeding jejunostomy tube in an esophageal cancer patient: A case report and review of the literature
by: Po‐Hsuan Wu, et al.
Published: (2024)
by: Po‐Hsuan Wu, et al.
Published: (2024)
ReCross: Efficient Embedding Reduction Scheme for In-Memory Computing using ReRAM-Based Crossbar
by: Lai, Yu-Hong, et al.
Published: (2025)
by: Lai, Yu-Hong, et al.
Published: (2025)
QuantTune: Optimizing Model Quantization with Adaptive Outlier-Driven Fine Tuning
by: Chen, Jiun-Man, et al.
Published: (2024)
by: Chen, Jiun-Man, et al.
Published: (2024)
Pareto-Optimal Energy Alignment for Designing Nature-Like Antibodies
by: Wen, Yibo, et al.
Published: (2024)
by: Wen, Yibo, et al.
Published: (2024)
Minimalist Softmax Attention Provably Learns Constrained Boolean Functions
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
Self-Attention with State-Object Weighted Combination for Compositional Zero Shot Learning
by: Chang, Cheng-Hong, et al.
Published: (2025)
by: Chang, Cheng-Hong, et al.
Published: (2025)
Similar Items
-
STanHop: Sparse Tandem Hopfield Model for Memory-Enhanced Time Series Prediction
by: Wu, Dennis, et al.
Published: (2023) -
Nonparametric Modern Hopfield Models
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024) -
Provably Optimal Memory Capacity for Modern Hopfield Models: Transformer-Compatible Dense Associative Memories as Spherical Codes
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024) -
Uniform Memory Retrieval with Larger Capacity for Modern Hopfield Models
by: Wu, Dennis, et al.
Published: (2024) -
On Computational Limits of Modern Hopfield Models: A Fine-Grained Complexity Analysis
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)