Gespeichert in:
| Hauptverfasser: | Graef, Nils, Wasielewski, Andrew |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2503.05840 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
KV-weights are all you need for skipless transformers
von: Graef, Nils
Veröffentlicht: (2024)
von: Graef, Nils
Veröffentlicht: (2024)
FlashNorm: Fast Normalization for Transformers
von: Graef, Nils, et al.
Veröffentlicht: (2024)
von: Graef, Nils, et al.
Veröffentlicht: (2024)
Linear attention is (maybe) all you need (to understand transformer optimization)
von: Ahn, Kwangjun, et al.
Veröffentlicht: (2023)
von: Ahn, Kwangjun, et al.
Veröffentlicht: (2023)
Transformer tricks: Precomputing the first layer
von: Graef, Nils
Veröffentlicht: (2024)
von: Graef, Nils
Veröffentlicht: (2024)
Is attention all you need in medical image analysis? A review
von: Papanastasiou, Giorgos, et al.
Veröffentlicht: (2023)
von: Papanastasiou, Giorgos, et al.
Veröffentlicht: (2023)
One protein is all you need
von: Bushuiev, Anton, et al.
Veröffentlicht: (2024)
von: Bushuiev, Anton, et al.
Veröffentlicht: (2024)
Graph is all you need? Lightweight data-agnostic neural architecture search without training
von: Huang, Zhenhan, et al.
Veröffentlicht: (2024)
von: Huang, Zhenhan, et al.
Veröffentlicht: (2024)
Kolmogorov GAM Networks are all you need!
von: Polson, Sarah, et al.
Veröffentlicht: (2025)
von: Polson, Sarah, et al.
Veröffentlicht: (2025)
Tabular Data: Is Deep Learning all you need?
von: Zabërgja, Guri, et al.
Veröffentlicht: (2024)
von: Zabërgja, Guri, et al.
Veröffentlicht: (2024)
Image compositing is all you need for data augmentation
von: Shermaine, Ang Jia Ning, et al.
Veröffentlicht: (2025)
von: Shermaine, Ang Jia Ning, et al.
Veröffentlicht: (2025)
Large Language Models aren't all that you need
von: Holla, Kiran Voderhobli, et al.
Veröffentlicht: (2024)
von: Holla, Kiran Voderhobli, et al.
Veröffentlicht: (2024)
Attention and Compression is all you need for Controllably Efficient Language Models
von: Prakash, Jatin, et al.
Veröffentlicht: (2025)
von: Prakash, Jatin, et al.
Veröffentlicht: (2025)
Experts are all you need: A Composable Framework for Large Language Model Inference
von: Sridharan, Shrihari, et al.
Veröffentlicht: (2025)
von: Sridharan, Shrihari, et al.
Veröffentlicht: (2025)
Cross-Modal Safety Alignment: Is textual unlearning all you need?
von: Chakraborty, Trishna, et al.
Veröffentlicht: (2024)
von: Chakraborty, Trishna, et al.
Veröffentlicht: (2024)
Attention is all you need for boosting graph convolutional neural network
von: Wu, Yinwei
Veröffentlicht: (2024)
von: Wu, Yinwei
Veröffentlicht: (2024)
Addition is almost all you need: Compressing large language models with double binary factorization
von: Boža, Vladimír, et al.
Veröffentlicht: (2025)
von: Boža, Vladimír, et al.
Veröffentlicht: (2025)
BaKlaVa -- Budgeted Allocation of KV cache for Long-context Inference
von: Gulhan, Ahmed Burak, et al.
Veröffentlicht: (2025)
von: Gulhan, Ahmed Burak, et al.
Veröffentlicht: (2025)
Self-attention as an attractor network: transient memories without backpropagation
von: D'Amico, Francesco, et al.
Veröffentlicht: (2024)
von: D'Amico, Francesco, et al.
Veröffentlicht: (2024)
Simulation-based inference with scattering representations: scattering is all you need
von: Lin, Kiyam, et al.
Veröffentlicht: (2024)
von: Lin, Kiyam, et al.
Veröffentlicht: (2024)
DC is all you need: describing ReLU from a signal processing standpoint
von: Kechris, Christodoulos, et al.
Veröffentlicht: (2024)
von: Kechris, Christodoulos, et al.
Veröffentlicht: (2024)
Few Labels are all you need: A Weakly Supervised Framework for Appliance Localization in Smart-Meter Series
von: Petralia, Adrien, et al.
Veröffentlicht: (2025)
von: Petralia, Adrien, et al.
Veröffentlicht: (2025)
In-context denoising with one-layer transformers: connections between attention and associative memory retrieval
von: Smart, Matthew, et al.
Veröffentlicht: (2025)
von: Smart, Matthew, et al.
Veröffentlicht: (2025)
Signformer is all you need: Towards Edge AI for Sign Language
von: Yang, Eta
Veröffentlicht: (2024)
von: Yang, Eta
Veröffentlicht: (2024)
An experimental study of KV cache reuse strategies in chunk-level caching systems
von: Cestola, Samuel, et al.
Veröffentlicht: (2026)
von: Cestola, Samuel, et al.
Veröffentlicht: (2026)
An extension of linear self-attention for in-context learning
von: Hagiwara, Katsuyuki
Veröffentlicht: (2025)
von: Hagiwara, Katsuyuki
Veröffentlicht: (2025)
Why you don't overfit, and don't need Bayes if you only train for one epoch
von: Aitchison, Laurence
Veröffentlicht: (2024)
von: Aitchison, Laurence
Veröffentlicht: (2024)
KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent Systems
von: Ye, Hancheng, et al.
Veröffentlicht: (2025)
von: Ye, Hancheng, et al.
Veröffentlicht: (2025)
Neural Operator: Is data all you need to model the world? An insight into the paradigm of data-driven scientific ML
von: Viswanath, Hrishikesh, et al.
Veröffentlicht: (2023)
von: Viswanath, Hrishikesh, et al.
Veröffentlicht: (2023)
Attention is all you need for an improved CNN-based flash flood susceptibility modeling. The case of the ungauged Rheraya watershed, Morocco
von: Elghouat, Akram, et al.
Veröffentlicht: (2024)
von: Elghouat, Akram, et al.
Veröffentlicht: (2024)
On Surjectivity of Neural Networks: Can you elicit any behavior from your model?
von: Jiang, Haozhe, et al.
Veröffentlicht: (2025)
von: Jiang, Haozhe, et al.
Veröffentlicht: (2025)
What are you sinking? A geometric approach on attention sink
von: Ruscio, Valeria, et al.
Veröffentlicht: (2025)
von: Ruscio, Valeria, et al.
Veröffentlicht: (2025)
100 instances is all you need: predicting the success of a new LLM on unseen data by testing on a few instances
von: Pacchiardi, Lorenzo, et al.
Veröffentlicht: (2024)
von: Pacchiardi, Lorenzo, et al.
Veröffentlicht: (2024)
SALS: Sparse Attention in Latent Space for KV cache Compression
von: Mu, Junlin, et al.
Veröffentlicht: (2025)
von: Mu, Junlin, et al.
Veröffentlicht: (2025)
SlimMoE: Structured Compression of Large MoE Models via Expert Slimming and Distillation
von: Li, Zichong, et al.
Veröffentlicht: (2025)
von: Li, Zichong, et al.
Veröffentlicht: (2025)
SlimDiff: Training-Free, Activation-Guided Hands-free Slimming of Diffusion Models
von: Roy, Arani, et al.
Veröffentlicht: (2025)
von: Roy, Arani, et al.
Veröffentlicht: (2025)
TransformerFAM: Feedback attention is working memory
von: Hwang, Dongseong, et al.
Veröffentlicht: (2024)
von: Hwang, Dongseong, et al.
Veröffentlicht: (2024)
Asymptotic theory of in-context learning by linear attention
von: Lu, Yue M., et al.
Veröffentlicht: (2024)
von: Lu, Yue M., et al.
Veröffentlicht: (2024)
An Innovative CGL-MHA Model for Sarcasm Sentiment Recognition Using the MindSpore Framework
von: Qin, Zhenkai, et al.
Veröffentlicht: (2024)
von: Qin, Zhenkai, et al.
Veröffentlicht: (2024)
Are nuclear masks all you need for improved out-of-domain generalisation? A closer look at cancer classification in histopathology
von: Tomar, Dhananjay, et al.
Veröffentlicht: (2024)
von: Tomar, Dhananjay, et al.
Veröffentlicht: (2024)
Is attention all you need to solve the correlated electron problem?
von: Geier, Max, et al.
Veröffentlicht: (2025)
von: Geier, Max, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
KV-weights are all you need for skipless transformers
von: Graef, Nils
Veröffentlicht: (2024) -
FlashNorm: Fast Normalization for Transformers
von: Graef, Nils, et al.
Veröffentlicht: (2024) -
Linear attention is (maybe) all you need (to understand transformer optimization)
von: Ahn, Kwangjun, et al.
Veröffentlicht: (2023) -
Transformer tricks: Precomputing the first layer
von: Graef, Nils
Veröffentlicht: (2024) -
Is attention all you need in medical image analysis? A review
von: Papanastasiou, Giorgos, et al.
Veröffentlicht: (2023)