Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Ziyang, Ding, Tianjiao, Lu, Yifu, Pai, Druv, Zhang, Jingyuan, Wang, Weida, Yu, Yaodong, Ma, Yi, Haeffele, Benjamin D. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is?
di: Yu, Yaodong, et al.
Pubblicazione: (2023)
di: Yu, Yaodong, et al.
Pubblicazione: (2023)
Attention-Only Transformers via Unrolled Subspace Denoising
di: Wang, Peng, et al.
Pubblicazione: (2025)
di: Wang, Peng, et al.
Pubblicazione: (2025)
Masked Completion via Structured Diffusion with White-Box Transformers
di: Pai, Druv, et al.
Pubblicazione: (2024)
di: Pai, Druv, et al.
Pubblicazione: (2024)
Image Clustering via the Principle of Rate Reduction in the Age of Pretrained Models
di: Chu, Tianzhe, et al.
Pubblicazione: (2023)
di: Chu, Tianzhe, et al.
Pubblicazione: (2023)
A Global Geometric Analysis of Maximal Coding Rate Reduction
di: Wang, Peng, et al.
Pubblicazione: (2024)
di: Wang, Peng, et al.
Pubblicazione: (2024)
Simplifying DINO via Coding Rate Regularization
di: Wu, Ziyang, et al.
Pubblicazione: (2025)
di: Wu, Ziyang, et al.
Pubblicazione: (2025)
Scaling White-Box Transformers for Vision
di: Yang, Jinrui, et al.
Pubblicazione: (2024)
di: Yang, Jinrui, et al.
Pubblicazione: (2024)
Active-Dormant Attention Heads: Mechanistically Demystifying Extreme-Token Phenomena in LLMs
di: Guo, Tianyu, et al.
Pubblicazione: (2024)
di: Guo, Tianyu, et al.
Pubblicazione: (2024)
On the Edge of Memorization in Diffusion Models
di: Buchanan, Sam, et al.
Pubblicazione: (2025)
di: Buchanan, Sam, et al.
Pubblicazione: (2025)
Independent and Decentralized Learning in Markov Potential Games
di: Maheshwari, Chinmay, et al.
Pubblicazione: (2022)
di: Maheshwari, Chinmay, et al.
Pubblicazione: (2022)
BLISS: Global Blind Identification of Linear Systems with Sparse Inputs
di: Poe, Kyle, et al.
Pubblicazione: (2026)
di: Poe, Kyle, et al.
Pubblicazione: (2026)
Measurement-Efficient Variational Quantum Linear Solver for Carleman-Linearized Nonlinear Dynamics
di: Liu, Yunya, et al.
Pubblicazione: (2026)
di: Liu, Yunya, et al.
Pubblicazione: (2026)
An Overview of Low-Rank Structures in the Training and Adaptation of Large Models
di: Balzano, Laura, et al.
Pubblicazione: (2025)
di: Balzano, Laura, et al.
Pubblicazione: (2025)
Congestion Pricing for Efficiency and Equity: Theory and Applications to the San Francisco Bay Area
di: Maheshwari, Chinmay, et al.
Pubblicazione: (2024)
di: Maheshwari, Chinmay, et al.
Pubblicazione: (2024)
Nonconvex Linear System Identification with Minimal State Representation
di: Tadipatri, Uday Kiran Reddy, et al.
Pubblicazione: (2025)
di: Tadipatri, Uday Kiran Reddy, et al.
Pubblicazione: (2025)
Consistent Time-of-Flight Depth Denoising via Graph-Informed Geometric Attention
di: Wang, Weida, et al.
Pubblicazione: (2025)
di: Wang, Weida, et al.
Pubblicazione: (2025)
Joint Chance-constrained Game for Coordinating Renewable Microgrids with Service Delivery Risk: A Bayesian Optimization Approach
di: Ding, Yifu, et al.
Pubblicazione: (2023)
di: Ding, Yifu, et al.
Pubblicazione: (2023)
Variational Linear Attention: Stable Associative Memory for Long-Context Transformers
di: Pandey, Vishal, et al.
Pubblicazione: (2026)
di: Pandey, Vishal, et al.
Pubblicazione: (2026)
Wave Physics-informed Matrix Factorizations
di: Tetali, Harsha Vardhan, et al.
Pubblicazione: (2023)
di: Tetali, Harsha Vardhan, et al.
Pubblicazione: (2023)
Diagonal-Tiled Mixed-Precision Attention for Efficient Low-Bit MXFP Inference
di: Ding, Yifu, et al.
Pubblicazione: (2026)
di: Ding, Yifu, et al.
Pubblicazione: (2026)
Ctrl123: Consistent Novel View Synthesis via Closed-Loop Transcription
di: Zhao, Hongxiang, et al.
Pubblicazione: (2024)
di: Zhao, Hongxiang, et al.
Pubblicazione: (2024)
Resi-VidTok: An Efficient and Decomposed Progressive Tokenization Framework for Ultra-Low-Rate and Lightweight Video Transmission
di: Liu, Zhenyu, et al.
Pubblicazione: (2025)
di: Liu, Zhenyu, et al.
Pubblicazione: (2025)
Dimension Reduction for Large-Scale Federated Data: Statistical Rate and Asymptotic Inference
di: Shen, Shuting, et al.
Pubblicazione: (2023)
di: Shen, Shuting, et al.
Pubblicazione: (2023)
Efectos de la musicoterapia integrativa en niños y niñas con autismo: un estudio de caso múltiple
di: Tianjiao Ma
Pubblicazione: (2025)
di: Tianjiao Ma
Pubblicazione: (2025)
On Statistical Rates of Conditional Diffusion Transformers: Approximation, Estimation and Minimax Optimality
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2024)
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2024)
Source data for article "High-throughput single-cell density measurements enable dynamic profiling of immune cell and drug response from patient samples"
di: Wu, Weida
Pubblicazione: (2025)
di: Wu, Weida
Pubblicazione: (2025)
GT-SNT: A Linear-Time Transformer for Large-Scale Graphs via Spiking Node Tokenization
di: Zhang, Huizhe, et al.
Pubblicazione: (2025)
di: Zhang, Huizhe, et al.
Pubblicazione: (2025)
Neighbor-Aware Token Reduction via Hilbert Curve for Vision Transformers
di: Li, Yunge, et al.
Pubblicazione: (2025)
di: Li, Yunge, et al.
Pubblicazione: (2025)
MBTSAD: Mitigating Backdoors in Language Models Based on Token Splitting and Attention Distillation
di: Ding, Yidong, et al.
Pubblicazione: (2025)
di: Ding, Yidong, et al.
Pubblicazione: (2025)
Efficient Real-Time Aircraft ETA Prediction via Feature Tokenization Transformer
di: Huang, Liping, et al.
Pubblicazione: (2025)
di: Huang, Liping, et al.
Pubblicazione: (2025)
SALAD: Achieve High-Sparsity Attention via Efficient Linear Attention Tuning for Video Diffusion Transformer
di: Fang, Tongcheng, et al.
Pubblicazione: (2026)
di: Fang, Tongcheng, et al.
Pubblicazione: (2026)
MHLA: Restoring Expressivity of Linear Attention via Token-Level Multi-Head
di: Zhang, Kewei, et al.
Pubblicazione: (2026)
di: Zhang, Kewei, et al.
Pubblicazione: (2026)
BWTA: Accurate and Efficient Binarized Transformer by Algorithm-Hardware Co-design
di: Ding, Yifu, et al.
Pubblicazione: (2026)
di: Ding, Yifu, et al.
Pubblicazione: (2026)
SPOT: Sparsification with Attention Dynamics via Token Relevance in Vision Transformers
di: Schlesinger, Oded, et al.
Pubblicazione: (2025)
di: Schlesinger, Oded, et al.
Pubblicazione: (2025)
Adaptive Stain Normalization for Cross-Domain Medical Histology
di: Xu, Tianyue, et al.
Pubblicazione: (2025)
di: Xu, Tianyue, et al.
Pubblicazione: (2025)
Purify Unlearnable Examples via Rate-Constrained Variational Autoencoders
di: Yu, Yi, et al.
Pubblicazione: (2024)
di: Yu, Yi, et al.
Pubblicazione: (2024)
Parallax: Parameterized Local Linear Attention for Language Modeling
di: Zuo, Yifei, et al.
Pubblicazione: (2026)
di: Zuo, Yifei, et al.
Pubblicazione: (2026)
A Convex Relaxation Approach to Generalization Analysis for Parallel Positively Homogeneous Networks
di: Tadipatri, Uday Kiran Reddy, et al.
Pubblicazione: (2024)
di: Tadipatri, Uday Kiran Reddy, et al.
Pubblicazione: (2024)
Attention Mechanism for LLM-based Agents Dynamic Diffusion under Information Asymmetry
di: Zhang, Yiwen, et al.
Pubblicazione: (2025)
di: Zhang, Yiwen, et al.
Pubblicazione: (2025)
Cascade Token Selection for Transformer Attention Acceleration
di: Thomas, Stephen J.
Pubblicazione: (2026)
di: Thomas, Stephen J.
Pubblicazione: (2026)
Documenti analoghi
-
White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is?
di: Yu, Yaodong, et al.
Pubblicazione: (2023) -
Attention-Only Transformers via Unrolled Subspace Denoising
di: Wang, Peng, et al.
Pubblicazione: (2025) -
Masked Completion via Structured Diffusion with White-Box Transformers
di: Pai, Druv, et al.
Pubblicazione: (2024) -
Image Clustering via the Principle of Rate Reduction in the Age of Pretrained Models
di: Chu, Tianzhe, et al.
Pubblicazione: (2023) -
A Global Geometric Analysis of Maximal Coding Rate Reduction
di: Wang, Peng, et al.
Pubblicazione: (2024)