On Difficulties of Attention Factorization through Shared Memory
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yorsh, Uladzislau, Holeňa, Martin, Bojar, Ondřej, Herel, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Adversarial Testing as a Tool for Interpretability: Length-based Overfitting of Elementary Functions in Transformers
von: Zavoral, Patrik, et al.
Veröffentlicht: (2024)
von: Zavoral, Patrik, et al.
Veröffentlicht: (2024)
Alquist 5.0: Dialogue Trees Meet Generative Models. A Novel Approach for Enhancing SocialBot Conversations
von: Kobza, Ondřej, et al.
Veröffentlicht: (2023)
von: Kobza, Ondřej, et al.
Veröffentlicht: (2023)
Geometric Reasoning in the Embedding Space
von: Hůla, Jan, et al.
Veröffentlicht: (2025)
von: Hůla, Jan, et al.
Veröffentlicht: (2025)
Higher-Order Message Passing for Glycan Representation Learning
von: Joeres, Roman, et al.
Veröffentlicht: (2024)
von: Joeres, Roman, et al.
Veröffentlicht: (2024)
Rethinking Thinking Tokens: Understanding Why They Underperform in Practice
von: Vennam, Sreeram, et al.
Veröffentlicht: (2024)
von: Vennam, Sreeram, et al.
Veröffentlicht: (2024)
Dataset Difficulty and the Role of Inductive Bias
von: Kwok, Devin, et al.
Veröffentlicht: (2024)
von: Kwok, Devin, et al.
Veröffentlicht: (2024)
Optimizing Reasoning Efficiency through Prompt Difficulty Prediction
von: Zhao, Bo, et al.
Veröffentlicht: (2025)
von: Zhao, Bo, et al.
Veröffentlicht: (2025)
Quaternion Self-Attention with Shared Scores
von: Yamauchi, Shogo, et al.
Veröffentlicht: (2026)
von: Yamauchi, Shogo, et al.
Veröffentlicht: (2026)
Reducing Fine-Tuning Memory Overhead by Approximate and Memory-Sharing Backpropagation
von: Yang, Yuchen, et al.
Veröffentlicht: (2024)
von: Yang, Yuchen, et al.
Veröffentlicht: (2024)
Linear Memory SE(2) Invariant Attention
von: Pronovost, Ethan, et al.
Veröffentlicht: (2025)
von: Pronovost, Ethan, et al.
Veröffentlicht: (2025)
DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation
von: Zhou, Yang, et al.
Veröffentlicht: (2026)
von: Zhou, Yang, et al.
Veröffentlicht: (2026)
Predicting Chess Puzzle Difficulty with Transformers
von: Miłosz, Szymon, et al.
Veröffentlicht: (2024)
von: Miłosz, Szymon, et al.
Veröffentlicht: (2024)
UMoE: Unifying Attention and FFN with Shared Experts
von: Yang, Yuanhang, et al.
Veröffentlicht: (2025)
von: Yang, Yuanhang, et al.
Veröffentlicht: (2025)
Difficulties with Evaluating a Deception Detector for AIs
von: Smith, Lewis, et al.
Veröffentlicht: (2025)
von: Smith, Lewis, et al.
Veröffentlicht: (2025)
Understanding the role of FFNs in driving multilingual behaviour in LLMs
von: Bhattacharya, Sunit, et al.
Veröffentlicht: (2024)
von: Bhattacharya, Sunit, et al.
Veröffentlicht: (2024)
Quality and Quantity of Machine Translation References for Automatic Metrics
von: Zouhar, Vilém, et al.
Veröffentlicht: (2024)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2024)
Finetuning LLMs for EvaCun 2025 token prediction shared task
von: Jon, Josef, et al.
Veröffentlicht: (2025)
von: Jon, Josef, et al.
Veröffentlicht: (2025)
Intrinsic vs. Extrinsic Evaluation of Czech Sentence Embeddings: Semantic Relevance Doesn't Help with MT Evaluation
von: Barančíková, Petra, et al.
Veröffentlicht: (2025)
von: Barančíková, Petra, et al.
Veröffentlicht: (2025)
Overview of the Sensemaking Task at the ELOQUENT 2025 Lab: LLMs as Teachers, Students and Evaluators
von: Šindelář, Pavel, et al.
Veröffentlicht: (2025)
von: Šindelář, Pavel, et al.
Veröffentlicht: (2025)
Long-Form End-to-End Speech Translation via Latent Alignment Segmentation
von: Polák, Peter, et al.
Veröffentlicht: (2023)
von: Polák, Peter, et al.
Veröffentlicht: (2023)
End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs
von: Luu, Nam, et al.
Veröffentlicht: (2025)
von: Luu, Nam, et al.
Veröffentlicht: (2025)
Learning to Forget Attention: Memory Consolidation for Adaptive Compute Reduction
von: Shihab, Ibne Farabi, et al.
Veröffentlicht: (2026)
von: Shihab, Ibne Farabi, et al.
Veröffentlicht: (2026)
Interdomain Attention: Beyond Token-Level Key-Value Memory
von: Kiyohara, Naoki, et al.
Veröffentlicht: (2026)
von: Kiyohara, Naoki, et al.
Veröffentlicht: (2026)
Dynamic Expert Sharing: Decoupling Memory from Parallelism in Mixture-of-Experts Diffusion LLMs
von: Chen, Hao Mark, et al.
Veröffentlicht: (2026)
von: Chen, Hao Mark, et al.
Veröffentlicht: (2026)
Optimizing Coverage and Difficulty in Reinforcement Learning for Quiz Composition
von: Silva, Ricardo Pedro Querido Andrade, et al.
Veröffentlicht: (2026)
von: Silva, Ricardo Pedro Querido Andrade, et al.
Veröffentlicht: (2026)
Memory-Efficient Optimization with Factorized Hamiltonian Descent
von: Nguyen, Son, et al.
Veröffentlicht: (2024)
von: Nguyen, Son, et al.
Veröffentlicht: (2024)
GanitLLM: Difficulty-Aware Bengali Mathematical Reasoning through Curriculum-GRPO
von: Dipta, Shubhashis Roy, et al.
Veröffentlicht: (2026)
von: Dipta, Shubhashis Roy, et al.
Veröffentlicht: (2026)
Navigating Noise: A Study of How Noise Influences Generalisation and Calibration of Neural Networks
von: Ferianc, Martin, et al.
Veröffentlicht: (2023)
von: Ferianc, Martin, et al.
Veröffentlicht: (2023)
Cross-Domain Latent Factors Sharing via Implicit Matrix Factorization
von: Samra, Abdulaziz, et al.
Veröffentlicht: (2024)
von: Samra, Abdulaziz, et al.
Veröffentlicht: (2024)
LoMA: Lossless Compressed Memory Attention
von: Wang, Yumeng, et al.
Veröffentlicht: (2024)
von: Wang, Yumeng, et al.
Veröffentlicht: (2024)
Adaptive Memory Decay for Log-Linear Attention
von: Amin, Yaxita, et al.
Veröffentlicht: (2026)
von: Amin, Yaxita, et al.
Veröffentlicht: (2026)
vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention
von: Prabhu, Ramya, et al.
Veröffentlicht: (2024)
von: Prabhu, Ramya, et al.
Veröffentlicht: (2024)
Share Your Attention: Transformer Weight Sharing via Matrix-based Dictionary Learning
von: Zhussip, Magauiya, et al.
Veröffentlicht: (2025)
von: Zhussip, Magauiya, et al.
Veröffentlicht: (2025)
GatedFWA: Linear Flash Windowed Attention with Gated Associative Memory
von: Liu, Jiaxu, et al.
Veröffentlicht: (2025)
von: Liu, Jiaxu, et al.
Veröffentlicht: (2025)
An Attention-based Feature Memory Design for Energy-Efficient Continual Learning
von: Wang, Yuandou, et al.
Veröffentlicht: (2025)
von: Wang, Yuandou, et al.
Veröffentlicht: (2025)
Temporal Attention for Adaptive Control of Euler-Lagrange Systems with Unobservable Memory
von: Cirrincione, Giansalvo, et al.
Veröffentlicht: (2026)
von: Cirrincione, Giansalvo, et al.
Veröffentlicht: (2026)
Variational Linear Attention: Stable Associative Memory for Long-Context Transformers
von: Pandey, Vishal, et al.
Veröffentlicht: (2026)
von: Pandey, Vishal, et al.
Veröffentlicht: (2026)
VeriDispatcher: Multi-Model Dispatching through Pre-Inference Difficulty Prediction for RTL Generation Optimization
von: Wang, Zeng, et al.
Veröffentlicht: (2025)
von: Wang, Zeng, et al.
Veröffentlicht: (2025)
The Course Difficulty Analysis Cookbook
von: Baucks, Frederik, et al.
Veröffentlicht: (2025)
von: Baucks, Frederik, et al.
Veröffentlicht: (2025)
Which LLMs are Difficult to Detect? A Detailed Analysis of Potential Factors Contributing to Difficulties in LLM Text Detection
von: Thorat, Shantanu, et al.
Veröffentlicht: (2024)
von: Thorat, Shantanu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Adversarial Testing as a Tool for Interpretability: Length-based Overfitting of Elementary Functions in Transformers
von: Zavoral, Patrik, et al.
Veröffentlicht: (2024) -
Alquist 5.0: Dialogue Trees Meet Generative Models. A Novel Approach for Enhancing SocialBot Conversations
von: Kobza, Ondřej, et al.
Veröffentlicht: (2023) -
Geometric Reasoning in the Embedding Space
von: Hůla, Jan, et al.
Veröffentlicht: (2025) -
Higher-Order Message Passing for Glycan Representation Learning
von: Joeres, Roman, et al.
Veröffentlicht: (2024) -
Rethinking Thinking Tokens: Understanding Why They Underperform in Practice
von: Vennam, Sreeram, et al.
Veröffentlicht: (2024)