When Can We Solve the Weighted Low Rank Approximation Problem in Truly Subquadratic Time?
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Chenyang, Liang, Yingyu, Shi, Zhenmei, Song, Zhao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RoPE Attention Can Be Trained in Almost Linear Time
di: Cao, Yang, et al.
Pubblicazione: (2024)
di: Cao, Yang, et al.
Pubblicazione: (2024)
Modern Hopfield Networks Require Chain-of-Thought to Solve $\mathsf{NC}^1$-Hard Problems
di: Cao, Yang, et al.
Pubblicazione: (2024)
di: Cao, Yang, et al.
Pubblicazione: (2024)
Looped ReLU MLPs May Be All You Need as Practical Programmable Computers
di: Liang, Yingyu, et al.
Pubblicazione: (2024)
di: Liang, Yingyu, et al.
Pubblicazione: (2024)
On the Computational Capability of Graph Neural Networks: A Circuit Complexity Bound Perspective
di: Li, Xiaoyu, et al.
Pubblicazione: (2025)
di: Li, Xiaoyu, et al.
Pubblicazione: (2025)
Time and Memory Trade-off of KV-Cache Compression in Tensor Transformer Decoding
di: Chen, Yifang, et al.
Pubblicazione: (2025)
di: Chen, Yifang, et al.
Pubblicazione: (2025)
The Computational Limits of State-Space Models and Mamba via the Lens of Circuit Complexity
di: Chen, Yifang, et al.
Pubblicazione: (2024)
di: Chen, Yifang, et al.
Pubblicazione: (2024)
Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers
di: Li, Xiaoyu, et al.
Pubblicazione: (2024)
di: Li, Xiaoyu, et al.
Pubblicazione: (2024)
Circuit Complexity Bounds for Visual Autoregressive Model
di: Ke, Yekun, et al.
Pubblicazione: (2025)
di: Ke, Yekun, et al.
Pubblicazione: (2025)
Neural Algorithmic Reasoning for Hypergraphs with Looped Transformers
di: Huang, Zekai, et al.
Pubblicazione: (2025)
di: Huang, Zekai, et al.
Pubblicazione: (2025)
On Fine-Grained I/O Complexity of Attention Backward Passes
di: Li, Xiaoyu, et al.
Pubblicazione: (2024)
di: Li, Xiaoyu, et al.
Pubblicazione: (2024)
Circuit Complexity Bounds for RoPE-based Transformer Architecture
di: Chen, Bo, et al.
Pubblicazione: (2024)
di: Chen, Bo, et al.
Pubblicazione: (2024)
On Computational Limits and Provably Efficient Criteria of Visual Autoregressive Models: A Fine-Grained Complexity Analysis
di: Ke, Yekun, et al.
Pubblicazione: (2025)
di: Ke, Yekun, et al.
Pubblicazione: (2025)
Multi-Layer Transformers Gradient Can be Approximated in Almost Linear Time
di: Liang, Yingyu, et al.
Pubblicazione: (2024)
di: Liang, Yingyu, et al.
Pubblicazione: (2024)
Provable Failure of Language Models in Learning Majority Boolean Logic via Gradient Descent
di: Chen, Bo, et al.
Pubblicazione: (2025)
di: Chen, Bo, et al.
Pubblicazione: (2025)
On Computational Limits of FlowAR Models: Expressivity and Efficiency
di: Cao, Yang, et al.
Pubblicazione: (2025)
di: Cao, Yang, et al.
Pubblicazione: (2025)
Training Tensor Attention Efficiently: From Cubic to Almost Linear Time
di: Cao, Yang, et al.
Pubblicazione: (2024)
di: Cao, Yang, et al.
Pubblicazione: (2024)
Computational Limits of Low-Rank Adaptation (LoRA) Fine-Tuning for Transformer Models
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2024)
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2024)
Exploring the Frontiers of Softmax: Provable Optimization, Applications in Diffusion Model, and Beyond
di: Cao, Yang, et al.
Pubblicazione: (2024)
di: Cao, Yang, et al.
Pubblicazione: (2024)
Beyond Linear Approximations: A Novel Pruning Approach for Attention Matrix
di: Liang, Yingyu, et al.
Pubblicazione: (2024)
di: Liang, Yingyu, et al.
Pubblicazione: (2024)
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond
di: Ke, Yekun, et al.
Pubblicazione: (2024)
di: Ke, Yekun, et al.
Pubblicazione: (2024)
Universal Approximation of Visual Autoregressive Transformers
di: Chen, Yifang, et al.
Pubblicazione: (2025)
di: Chen, Yifang, et al.
Pubblicazione: (2025)
Bypassing the Exponential Dependency: Looped Transformers Efficiently Learn In-context by Multi-step Gradient Descent
di: Chen, Bo, et al.
Pubblicazione: (2024)
di: Chen, Bo, et al.
Pubblicazione: (2024)
Towards Infinite-Long Prefix in Transformer
di: Liang, Yingyu, et al.
Pubblicazione: (2024)
di: Liang, Yingyu, et al.
Pubblicazione: (2024)
Lossless Model Compression via Joint Low-Rank Factorization Optimization
di: Zhang, Boyang, et al.
Pubblicazione: (2024)
di: Zhang, Boyang, et al.
Pubblicazione: (2024)
A Tighter Complexity Analysis of SparseGPT
di: Li, Xiaoyu, et al.
Pubblicazione: (2024)
di: Li, Xiaoyu, et al.
Pubblicazione: (2024)
Do Large Language Models Have Compositional Ability? An Investigation into Limitations and Scalability
di: Xu, Zhuoyan, et al.
Pubblicazione: (2024)
di: Xu, Zhuoyan, et al.
Pubblicazione: (2024)
Unraveling the Smoothness Properties of Diffusion Models: A Gaussian Mixture Perspective
di: Liang, Yingyu, et al.
Pubblicazione: (2024)
di: Liang, Yingyu, et al.
Pubblicazione: (2024)
Advancing the Understanding of Fixed Point Iterations in Deep Neural Networks: A Detailed Analytical Study
di: Ke, Yekun, et al.
Pubblicazione: (2024)
di: Ke, Yekun, et al.
Pubblicazione: (2024)
Scaling Law Phenomena Across Regression Paradigms: Multiple and Kernel Approaches
di: Chen, Yifang, et al.
Pubblicazione: (2025)
di: Chen, Yifang, et al.
Pubblicazione: (2025)
HSR-Enhanced Sparse Attention Acceleration
di: Chen, Bo, et al.
Pubblicazione: (2024)
di: Chen, Bo, et al.
Pubblicazione: (2024)
Provable Differentially Private Computation of the Cross-Attention Mechanism
di: Ke, Yekun, et al.
Pubblicazione: (2024)
di: Ke, Yekun, et al.
Pubblicazione: (2024)
Differential Privacy Mechanisms in Neural Tangent Kernel Regression
di: Gu, Jiuxiang, et al.
Pubblicazione: (2024)
di: Gu, Jiuxiang, et al.
Pubblicazione: (2024)
Conv-Basis: A New Paradigm for Efficient Attention Inference and Gradient Computation in Transformers
di: Liang, Yingyu, et al.
Pubblicazione: (2024)
di: Liang, Yingyu, et al.
Pubblicazione: (2024)
Why Larger Language Models Do In-context Learning Differently?
di: Shi, Zhenmei, et al.
Pubblicazione: (2024)
di: Shi, Zhenmei, et al.
Pubblicazione: (2024)
HOFAR: High-Order Augmentation of Flow Autoregressive Transformers
di: Liang, Yingyu, et al.
Pubblicazione: (2025)
di: Liang, Yingyu, et al.
Pubblicazione: (2025)
Kernel Regression in Structured Non-IID Settings: Theory and Implications for Denoising Score Learning
di: Zhang, Dechen, et al.
Pubblicazione: (2025)
di: Zhang, Dechen, et al.
Pubblicazione: (2025)
Can You Count to Nine? A Human Evaluation Benchmark for Counting Limits in Modern Text-to-Video Models
di: Guo, Xuyang, et al.
Pubblicazione: (2025)
di: Guo, Xuyang, et al.
Pubblicazione: (2025)
Reinforced Generation of Combinatorial Structures: Hardness of Approximation
di: Nagda, Ansh, et al.
Pubblicazione: (2025)
di: Nagda, Ansh, et al.
Pubblicazione: (2025)
Visual Autoregressive Transformers Must Use $Ω(n^2 d)$ Memory
di: Cao, Yang, et al.
Pubblicazione: (2025)
di: Cao, Yang, et al.
Pubblicazione: (2025)
Low-Rank Matrix Approximation for Neural Network Compression
di: Cherukuri, Kalyan, et al.
Pubblicazione: (2025)
di: Cherukuri, Kalyan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
RoPE Attention Can Be Trained in Almost Linear Time
di: Cao, Yang, et al.
Pubblicazione: (2024) -
Modern Hopfield Networks Require Chain-of-Thought to Solve $\mathsf{NC}^1$-Hard Problems
di: Cao, Yang, et al.
Pubblicazione: (2024) -
Looped ReLU MLPs May Be All You Need as Practical Programmable Computers
di: Liang, Yingyu, et al.
Pubblicazione: (2024) -
On the Computational Capability of Graph Neural Networks: A Circuit Complexity Bound Perspective
di: Li, Xiaoyu, et al.
Pubblicazione: (2025) -
Time and Memory Trade-off of KV-Cache Compression in Tensor Transformer Decoding
di: Chen, Yifang, et al.
Pubblicazione: (2025)