Enregistré dans:
| Auteurs principaux: | Zhang, Michael, Bhatia, Kush, Kumbong, Hermann, Ré, Christopher |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2402.04347 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Cookbook: A framework for improving LLM generative abilities via programmatic data generating templates
par: Narayan, Avanika, et autres
Publié: (2024)
par: Narayan, Avanika, et autres
Publié: (2024)
LowRA: Accurate and Efficient LoRA Fine-Tuning of LLMs under 2 Bits
par: Zhou, Zikai, et autres
Publié: (2025)
par: Zhou, Zikai, et autres
Publié: (2025)
Why Softmax Attention Outperforms Linear Attention
par: Deng, Yichuan, et autres
Publié: (2023)
par: Deng, Yichuan, et autres
Publié: (2023)
Automated Rewards via LLM-Generated Progress Functions
par: Sarukkai, Vishnu, et autres
Publié: (2024)
par: Sarukkai, Vishnu, et autres
Publié: (2024)
Kimi Linear: An Expressive, Efficient Attention Architecture
par: Kimi Team, et autres
Publié: (2025)
par: Kimi Team, et autres
Publié: (2025)
The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought
par: Brösamle, Moritz, et autres
Publié: (2026)
par: Brösamle, Moritz, et autres
Publié: (2026)
Stochastic Attention: Connectome-Inspired Randomized Routing for Expressive Linear-Time Attention
par: Jin, Zehao, et autres
Publié: (2026)
par: Jin, Zehao, et autres
Publié: (2026)
RAM-Net: Expressive Linear Attention with Selectively Addressable Memory
par: Xiao, Kaicheng, et autres
Publié: (2026)
par: Xiao, Kaicheng, et autres
Publié: (2026)
Scalable-Softmax Is Superior for Attention
par: Nakanishi, Ken M.
Publié: (2025)
par: Nakanishi, Ken M.
Publié: (2025)
Softmax Attention with Constant Cost per Token
par: Heinsen, Franz A.
Publié: (2024)
par: Heinsen, Franz A.
Publié: (2024)
Forgetting Transformer: Softmax Attention with a Forget Gate
par: Lin, Zhixuan, et autres
Publié: (2025)
par: Lin, Zhixuan, et autres
Publié: (2025)
Evaluating the fairness of task-adaptive pretraining on unlabeled test data before few-shot text classification
par: Dubey, Kush
Publié: (2024)
par: Dubey, Kush
Publié: (2024)
In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
par: Collins, Liam, et autres
Publié: (2024)
par: Collins, Liam, et autres
Publié: (2024)
LoLCATs: On Low-Rank Linearizing of Large Language Models
par: Zhang, Michael, et autres
Publié: (2024)
par: Zhang, Michael, et autres
Publié: (2024)
More Expressive Attention with Negative Weights
par: Lv, Ang, et autres
Publié: (2024)
par: Lv, Ang, et autres
Publié: (2024)
To Softmax, or not to Softmax: that is the question when applying Active Learning for Transformer Models
par: Gonsior, Julius, et autres
Publié: (2022)
par: Gonsior, Julius, et autres
Publié: (2022)
Softmax Transformers are Turing-Complete
par: Jiang, Hongjian, et autres
Publié: (2025)
par: Jiang, Hongjian, et autres
Publié: (2025)
Taipan: Efficient and Expressive State Space Language Models with Selective Attention
par: Van Nguyen, Chien, et autres
Publié: (2024)
par: Van Nguyen, Chien, et autres
Publié: (2024)
The Information Geometry of Softmax: Probing and Steering
par: Park, Kiho, et autres
Publié: (2026)
par: Park, Kiho, et autres
Publié: (2026)
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling
par: Dadgarnia, Alireza, et autres
Publié: (2026)
par: Dadgarnia, Alireza, et autres
Publié: (2026)
Beyond Mimicry to Contextual Guidance: Knowledge Distillation for Interactive AI
par: Wang, Tong, et autres
Publié: (2024)
par: Wang, Tong, et autres
Publié: (2024)
SEA: Sparse Linear Attention with Estimated Attention Mask
par: Lee, Heejun, et autres
Publié: (2023)
par: Lee, Heejun, et autres
Publié: (2023)
Linear Attention Sequence Parallelism
par: Sun, Weigao, et autres
Publié: (2024)
par: Sun, Weigao, et autres
Publié: (2024)
Higher-order Linear Attention
par: Zhang, Yifan, et autres
Publié: (2025)
par: Zhang, Yifan, et autres
Publié: (2025)
On the Expressiveness of Softmax Attention: A Recurrent Neural Network Perspective
par: Mongaras, Gabriel, et autres
Publié: (2025)
par: Mongaras, Gabriel, et autres
Publié: (2025)
Trading Complexity for Expressivity Through Structured Generalized Linear Token Mixing
par: Fagnou, Erwan, et autres
Publié: (2026)
par: Fagnou, Erwan, et autres
Publié: (2026)
Scaling Linear Attention with Sparse State Expansion
par: Pan, Yuqi, et autres
Publié: (2025)
par: Pan, Yuqi, et autres
Publié: (2025)
HMAR: Efficient Hierarchical Masked Auto-Regressive Image Generation
par: Kumbong, Hermann, et autres
Publié: (2025)
par: Kumbong, Hermann, et autres
Publié: (2025)
Simple linear attention language models balance the recall-throughput tradeoff
par: Arora, Simran, et autres
Publié: (2024)
par: Arora, Simran, et autres
Publié: (2024)
Aioli: A Unified Optimization Framework for Language Model Data Mixing
par: Chen, Mayee F., et autres
Publié: (2024)
par: Chen, Mayee F., et autres
Publié: (2024)
Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers
par: Li, Xiaoyu, et autres
Publié: (2024)
par: Li, Xiaoyu, et autres
Publié: (2024)
Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
par: He, Mutian, et autres
Publié: (2025)
par: He, Mutian, et autres
Publié: (2025)
Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention Models
par: Deng, Difan, et autres
Publié: (2026)
par: Deng, Difan, et autres
Publié: (2026)
Gated Linear Attention Transformers with Hardware-Efficient Training
par: Yang, Songlin, et autres
Publié: (2023)
par: Yang, Songlin, et autres
Publié: (2023)
Counting Like Transformers: Compiling Temporal Counting Logic Into Softmax Transformers
par: Yang, Andy, et autres
Publié: (2024)
par: Yang, Andy, et autres
Publié: (2024)
The Expressive Capacity of State Space Models: A Formal Language Perspective
par: Sarrof, Yash, et autres
Publié: (2024)
par: Sarrof, Yash, et autres
Publié: (2024)
Unifying Linear-Time Attention via Latent Probabilistic Modelling
par: Dolga, Rares, et autres
Publié: (2024)
par: Dolga, Rares, et autres
Publié: (2024)
Transformer Based Linear Attention with Optimized GPU Kernel Implementation
par: Gerami, Armin, et autres
Publié: (2025)
par: Gerami, Armin, et autres
Publié: (2025)
LoLA: Low-Rank Linear Attention With Sparse Caching
par: McDermott, Luke, et autres
Publié: (2025)
par: McDermott, Luke, et autres
Publié: (2025)
Probing Geometry of Next Token Prediction Using Cumulant Expansion of the Softmax Entropy
par: Viswanathan, Karthik, et autres
Publié: (2025)
par: Viswanathan, Karthik, et autres
Publié: (2025)
Documents similaires
-
Cookbook: A framework for improving LLM generative abilities via programmatic data generating templates
par: Narayan, Avanika, et autres
Publié: (2024) -
LowRA: Accurate and Efficient LoRA Fine-Tuning of LLMs under 2 Bits
par: Zhou, Zikai, et autres
Publié: (2025) -
Why Softmax Attention Outperforms Linear Attention
par: Deng, Yichuan, et autres
Publié: (2023) -
Automated Rewards via LLM-Generated Progress Functions
par: Sarukkai, Vishnu, et autres
Publié: (2024) -
Kimi Linear: An Expressive, Efficient Attention Architecture
par: Kimi Team, et autres
Publié: (2025)