Empirical Capacity Model for Self-Attention Neural Networks
Fuente:
arXiv
Guardado en:
| Autores principales: | Härmä, Aki, Pietrasik, Marcin, Wilbik, Anna |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Hierarchical Blockmodelling for Knowledge Graphs
por: Pietrasik, Marcin, et al.
Publicado: (2024)
por: Pietrasik, Marcin, et al.
Publicado: (2024)
When Attention Sink Emerges in Language Models: An Empirical View
por: Gu, Xiangming, et al.
Publicado: (2024)
por: Gu, Xiangming, et al.
Publicado: (2024)
Fully Autonomous Programming using Iterative Multi-Agent Debugging with Large Language Models
por: Grishina, Anastasiia, et al.
Publicado: (2025)
por: Grishina, Anastasiia, et al.
Publicado: (2025)
Deep Learning Methods for Detecting Thermal Runaway Events in Battery Production Lines
por: Athanasopoulos, Athanasios, et al.
Publicado: (2025)
por: Athanasopoulos, Athanasios, et al.
Publicado: (2025)
Scaling Fine-Grained MoE Beyond 50B Parameters: Empirical Evaluation and Practical Insights
por: Krajewski, Jakub, et al.
Publicado: (2025)
por: Krajewski, Jakub, et al.
Publicado: (2025)
From Self-Attention to Markov Models: Unveiling the Dynamics of Generative Transformers
por: Ildiz, M. Emrullah, et al.
Publicado: (2024)
por: Ildiz, M. Emrullah, et al.
Publicado: (2024)
Direct-Inverse Prompting: Analyzing LLMs' Discriminative Capacity in Self-Improving Generation
por: Ahn, Jihyun Janice, et al.
Publicado: (2024)
por: Ahn, Jihyun Janice, et al.
Publicado: (2024)
Capacity Matters: a Proof-of-Concept for Transformer Memorization on Real-World Data
por: Changalidis, Anton, et al.
Publicado: (2025)
por: Changalidis, Anton, et al.
Publicado: (2025)
Towards the Law of Capacity Gap in Distilling Language Models
por: Zhang, Chen, et al.
Publicado: (2023)
por: Zhang, Chen, et al.
Publicado: (2023)
RecurFormer: Not All Transformer Heads Need Self-Attention
por: Yan, Ruiqing, et al.
Publicado: (2024)
por: Yan, Ruiqing, et al.
Publicado: (2024)
Dynamics of Spontaneous Topic Changes in Next Token Prediction with Self-Attention
por: Jia, Mumin, et al.
Publicado: (2025)
por: Jia, Mumin, et al.
Publicado: (2025)
Self-Improving World Modelling with Latent Actions
por: Qiu, Yifu, et al.
Publicado: (2026)
por: Qiu, Yifu, et al.
Publicado: (2026)
Rare but Severe Neural Machine Translation Errors Induced by Minimal Deletion: An Empirical Study on Chinese and English
por: Shi, Ruikang, et al.
Publicado: (2022)
por: Shi, Ruikang, et al.
Publicado: (2022)
Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws
por: Allen-Zhu, Zeyuan, et al.
Publicado: (2024)
por: Allen-Zhu, Zeyuan, et al.
Publicado: (2024)
GMLM: Bridging Graph Neural Networks and Language Models for Heterophilic Node Classification
por: Sinha, Aarush
Publicado: (2025)
por: Sinha, Aarush
Publicado: (2025)
Increasing Model Capacity for Free: A Simple Strategy for Parameter Efficient Fine-tuning
por: Song, Haobo, et al.
Publicado: (2024)
por: Song, Haobo, et al.
Publicado: (2024)
Neighborhood-Order Learning Graph Attention Network for Fake News Detection
por: Lakzaei, Batool, et al.
Publicado: (2025)
por: Lakzaei, Batool, et al.
Publicado: (2025)
Machine Learning for Pattern Detection in Printhead Nozzle Logging
por: Prianikov, Nikola, et al.
Publicado: (2025)
por: Prianikov, Nikola, et al.
Publicado: (2025)
Predicting the Lifespan of Industrial Printheads with Survival Analysis
por: Parii, Dan, et al.
Publicado: (2025)
por: Parii, Dan, et al.
Publicado: (2025)
Native Hybrid Attention for Efficient Sequence Modeling
por: Du, Jusen, et al.
Publicado: (2025)
por: Du, Jusen, et al.
Publicado: (2025)
Self-generated Replay Memories for Continual Neural Machine Translation
por: Resta, Michele, et al.
Publicado: (2024)
por: Resta, Michele, et al.
Publicado: (2024)
Parallax: Parameterized Local Linear Attention for Language Modeling
por: Zuo, Yifei, et al.
Publicado: (2026)
por: Zuo, Yifei, et al.
Publicado: (2026)
Switchable Decision: Dynamic Neural Generation Networks
por: Zhang, Shujian, et al.
Publicado: (2024)
por: Zhang, Shujian, et al.
Publicado: (2024)
An Empirical Comparison of Text Summarization: A Multi-Dimensional Evaluation of Large Language Models
por: Janakiraman, Anantharaman, et al.
Publicado: (2025)
por: Janakiraman, Anantharaman, et al.
Publicado: (2025)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
por: Wang, Guangtao, et al.
Publicado: (2025)
por: Wang, Guangtao, et al.
Publicado: (2025)
CrisisKAN: Knowledge-infused and Explainable Multimodal Attention Network for Crisis Event Classification
por: Gupta, Shubham, et al.
Publicado: (2024)
por: Gupta, Shubham, et al.
Publicado: (2024)
Relation-Aware Network with Attention-Based Loss for Few-Shot Knowledge Graph Completion
por: Qiao, Qiao, et al.
Publicado: (2023)
por: Qiao, Qiao, et al.
Publicado: (2023)
CAISSON: Concept-Augmented Inference Suite of Self-Organizing Neural Networks
por: Halperin, Igor
Publicado: (2024)
por: Halperin, Igor
Publicado: (2024)
Prediction is not Explanation: Revisiting the Explanatory Capacity of Mapping Embeddings
por: Herasimchyk, Hanna, et al.
Publicado: (2025)
por: Herasimchyk, Hanna, et al.
Publicado: (2025)
Attention Speaks Volumes: Localizing and Mitigating Bias in Language Models
por: Adiga, Rishabh, et al.
Publicado: (2024)
por: Adiga, Rishabh, et al.
Publicado: (2024)
Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry
por: Li, Zhuochun, et al.
Publicado: (2026)
por: Li, Zhuochun, et al.
Publicado: (2026)
Understanding and Improving Length Generalization in Hierarchical Sparse Attention Models
por: Leng, Jiaqi, et al.
Publicado: (2025)
por: Leng, Jiaqi, et al.
Publicado: (2025)
Sliding Window Attention Training for Efficient Large Language Models
por: Fu, Zichuan, et al.
Publicado: (2025)
por: Fu, Zichuan, et al.
Publicado: (2025)
Cost-Optimal Grouped-Query Attention for Long-Context Modeling
por: Chen, Yingfa, et al.
Publicado: (2025)
por: Chen, Yingfa, et al.
Publicado: (2025)
Instruction Following by Principled Boosting Attention of Large Language Models
por: Guardieiro, Vitoria, et al.
Publicado: (2025)
por: Guardieiro, Vitoria, et al.
Publicado: (2025)
SelfIE: Self-Interpretation of Large Language Model Embeddings
por: Chen, Haozhe, et al.
Publicado: (2024)
por: Chen, Haozhe, et al.
Publicado: (2024)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
por: Deng, Yichuan, et al.
Publicado: (2024)
por: Deng, Yichuan, et al.
Publicado: (2024)
Is Bigger Edit Batch Size Always Better? -- An Empirical Study on Model Editing with Llama-3
por: Yoon, Junsang, et al.
Publicado: (2024)
por: Yoon, Junsang, et al.
Publicado: (2024)
Do Large Language Models Truly Grasp Mathematics? An Empirical Exploration From Cognitive Psychology
por: Xie, Wei, et al.
Publicado: (2024)
por: Xie, Wei, et al.
Publicado: (2024)
Interactive Training: Feedback-Driven Neural Network Optimization
por: Zhang, Wentao, et al.
Publicado: (2025)
por: Zhang, Wentao, et al.
Publicado: (2025)
Ejemplares similares
-
Hierarchical Blockmodelling for Knowledge Graphs
por: Pietrasik, Marcin, et al.
Publicado: (2024) -
When Attention Sink Emerges in Language Models: An Empirical View
por: Gu, Xiangming, et al.
Publicado: (2024) -
Fully Autonomous Programming using Iterative Multi-Agent Debugging with Large Language Models
por: Grishina, Anastasiia, et al.
Publicado: (2025) -
Deep Learning Methods for Detecting Thermal Runaway Events in Battery Production Lines
por: Athanasopoulos, Athanasios, et al.
Publicado: (2025) -
Scaling Fine-Grained MoE Beyond 50B Parameters: Empirical Evaluation and Practical Insights
por: Krajewski, Jakub, et al.
Publicado: (2025)