Attention-Only Transformers via Unrolled Subspace Denoising
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Peng, Lu, Yifu, Yu, Yaodong, Pai, Druv, Qu, Qing, Ma, Yi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction
by: Wu, Ziyang, et al.
Published: (2024)
by: Wu, Ziyang, et al.
Published: (2024)
Masked Completion via Structured Diffusion with White-Box Transformers
by: Pai, Druv, et al.
Published: (2024)
by: Pai, Druv, et al.
Published: (2024)
A Global Geometric Analysis of Maximal Coding Rate Reduction
by: Wang, Peng, et al.
Published: (2024)
by: Wang, Peng, et al.
Published: (2024)
On the Edge of Memorization in Diffusion Models
by: Buchanan, Sam, et al.
Published: (2025)
by: Buchanan, Sam, et al.
Published: (2025)
White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is?
by: Yu, Yaodong, et al.
Published: (2023)
by: Yu, Yaodong, et al.
Published: (2023)
Exploring Low-Dimensional Subspaces in Diffusion Models for Controllable Image Editing
by: Chen, Siyi, et al.
Published: (2024)
by: Chen, Siyi, et al.
Published: (2024)
Active-Dormant Attention Heads: Mechanistically Demystifying Extreme-Token Phenomena in LLMs
by: Guo, Tianyu, et al.
Published: (2024)
by: Guo, Tianyu, et al.
Published: (2024)
Diffusion Models Learn Low-Dimensional Distributions via Subspace Clustering
by: Wang, Peng, et al.
Published: (2024)
by: Wang, Peng, et al.
Published: (2024)
Lightweight Transformer for EEG Classification via Balanced Signed Graph Algorithm Unrolling
by: Yao, Junyi, et al.
Published: (2025)
by: Yao, Junyi, et al.
Published: (2025)
Multi-Omics Analysis for Cancer Subtype Inference via Unrolling Graph Smoothness Priors
by: Lu, Jielong, et al.
Published: (2025)
by: Lu, Jielong, et al.
Published: (2025)
A Constrained Optimization Perspective of Unrolled Transformers
by: Porras-Valenzuela, Javier, et al.
Published: (2026)
by: Porras-Valenzuela, Javier, et al.
Published: (2026)
Contextual Text Denoising with Masked Language Models
by: Sun, Yifu, et al.
Published: (2019)
by: Sun, Yifu, et al.
Published: (2019)
Scaling White-Box Transformers for Vision
by: Yang, Jinrui, et al.
Published: (2024)
by: Yang, Jinrui, et al.
Published: (2024)
Attention-Refined Unrolling for Sparse Sequential micro-Doppler Reconstruction
by: Mazzieri, Riccardo, et al.
Published: (2023)
by: Mazzieri, Riccardo, et al.
Published: (2023)
Cross-Validated Cross-Channel Self-Attention and Denoising for Automatic Modulation Classification
by: Suman, Prakash, et al.
Published: (2026)
by: Suman, Prakash, et al.
Published: (2026)
Transformers as Unrolled Inference in Probabilistic Laplacian Eigenmaps: An Interpretation and Potential Improvements
by: Ravuri, Aditya, et al.
Published: (2025)
by: Ravuri, Aditya, et al.
Published: (2025)
Training Data Attribution via Approximate Unrolled Differentiation
by: Bae, Juhan, et al.
Published: (2024)
by: Bae, Juhan, et al.
Published: (2024)
Lightweight and Interpretable Transformer via Mixed Graph Algorithm Unrolling for Traffic Forecast
by: Qi, Ji, et al.
Published: (2025)
by: Qi, Ji, et al.
Published: (2025)
The Emergence of Reproducibility and Generalizability in Diffusion Models
by: Zhang, Huijie, et al.
Published: (2023)
by: Zhang, Huijie, et al.
Published: (2023)
Learning Expressive Random Feature Models via Parametrized Activations
by: Ma, Zailin, et al.
Published: (2024)
by: Ma, Zailin, et al.
Published: (2024)
Precipitation Nowcasting Using Diffusion Transformer with Causal Attention
by: Li, ChaoRong, et al.
Published: (2024)
by: Li, ChaoRong, et al.
Published: (2024)
Understanding the Curse of Unrolling
by: Mehmood, Sheheryar, et al.
Published: (2026)
by: Mehmood, Sheheryar, et al.
Published: (2026)
Continual Learning with Query-Only Attention
by: Bekal, Gautham, et al.
Published: (2025)
by: Bekal, Gautham, et al.
Published: (2025)
Unrolled Neural Networks for Constrained Optimization
by: Hadou, Samar, et al.
Published: (2026)
by: Hadou, Samar, et al.
Published: (2026)
An Efficient Unsupervised Framework for Convex Quadratic Programs via Deep Unrolling
by: Yang, Linxin, et al.
Published: (2024)
by: Yang, Linxin, et al.
Published: (2024)
Efficient Resource-Constrained Training of Transformers via Subspace Optimization
by: Nguyen, Le-Trung, et al.
Published: (2025)
by: Nguyen, Le-Trung, et al.
Published: (2025)
Independent and Decentralized Learning in Markov Potential Games
by: Maheshwari, Chinmay, et al.
Published: (2022)
by: Maheshwari, Chinmay, et al.
Published: (2022)
Stochastic Unrolled Federated Learning
by: Hadou, Samar, et al.
Published: (2023)
by: Hadou, Samar, et al.
Published: (2023)
EEGDnet: Fusing Non-Local and Local Self-Similarity for 1-D EEG Signal Denoising with 2-D Transformer
by: Yi, Peng, et al.
Published: (2021)
by: Yi, Peng, et al.
Published: (2021)
Interpretable Lightweight Transformer via Unrolling of Learned Graph Smoothness Priors
by: Do, Tam Thuc, et al.
Published: (2024)
by: Do, Tam Thuc, et al.
Published: (2024)
Unrolled Graph Neural Networks for Constrained Optimization
by: Hadou, Samar, et al.
Published: (2025)
by: Hadou, Samar, et al.
Published: (2025)
DS2TA: Denoising Spiking Transformer with Attenuated Spatiotemporal Attention
by: Xu, Boxun, et al.
Published: (2024)
by: Xu, Boxun, et al.
Published: (2024)
Emergence of Minimal Circuits for Indirect Object Identification in Attention-Only Transformers
by: Adhikari, Rabin
Published: (2025)
by: Adhikari, Rabin
Published: (2025)
Subspace Control: Turning Constrained Model Steering into Controllable Spectral Optimization
by: Huang, Yancheng, et al.
Published: (2026)
by: Huang, Yancheng, et al.
Published: (2026)
Out-of-Distribution Generalization of In-Context Learning: A Low-Dimensional Subspace Perspective
by: Kwon, Soo Min, et al.
Published: (2025)
by: Kwon, Soo Min, et al.
Published: (2025)
PDHG-Unrolled Learning-to-Optimize Method for Large-Scale Linear Programming
by: Li, Bingheng, et al.
Published: (2024)
by: Li, Bingheng, et al.
Published: (2024)
Shallow Diffuse: Robust and Invisible Watermarking through Low-Dimensional Subspaces in Diffusion Models
by: Li, Wenda, et al.
Published: (2024)
by: Li, Wenda, et al.
Published: (2024)
Robust Stochastically-Descending Unrolled Networks
by: Hadou, Samar, et al.
Published: (2023)
by: Hadou, Samar, et al.
Published: (2023)
Recurrent Transformer U-Net Surrogate for Flow Modeling and Data Assimilation in Subsurface Formations with Faults
by: Han, Yifu, et al.
Published: (2025)
by: Han, Yifu, et al.
Published: (2025)
Joint Data Inpainting and Graph Learning via Unrolled Neural Networks
by: Batreddy, Subbareddy, et al.
Published: (2024)
by: Batreddy, Subbareddy, et al.
Published: (2024)
Similar Items
-
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction
by: Wu, Ziyang, et al.
Published: (2024) -
Masked Completion via Structured Diffusion with White-Box Transformers
by: Pai, Druv, et al.
Published: (2024) -
A Global Geometric Analysis of Maximal Coding Rate Reduction
by: Wang, Peng, et al.
Published: (2024) -
On the Edge of Memorization in Diffusion Models
by: Buchanan, Sam, et al.
Published: (2025) -
White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is?
by: Yu, Yaodong, et al.
Published: (2023)