Breaking the Attention Bottleneck
Fuente:
arXiv
Salvato in:
| Autore principale: | Hilsenbek, Kalle |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Breaking Training Bottlenecks: Effective and Stable Reinforcement Learning for Coding Models
di: Li, Zongqian, et al.
Pubblicazione: (2026)
di: Li, Zongqian, et al.
Pubblicazione: (2026)
MUDDFormer: Breaking Residual Bottlenecks in Transformers via Multiway Dynamic Dense Connections
di: Xiao, Da, et al.
Pubblicazione: (2025)
di: Xiao, Da, et al.
Pubblicazione: (2025)
Breaking Quadratic Barriers: A Non-Attention LLM for Ultra-Long Context Horizons
di: Kiruluta, Andrew, et al.
Pubblicazione: (2025)
di: Kiruluta, Andrew, et al.
Pubblicazione: (2025)
Efficient Knowledge Injection in LLMs via Self-Distillation
di: Kujanpää, Kalle, et al.
Pubblicazione: (2024)
di: Kujanpää, Kalle, et al.
Pubblicazione: (2024)
Investigating Affect Mining Techniques for Annotation Sample Selection in the Creation of Finnish Affective Speech Corpus
di: Lahtinen, Kalle, et al.
Pubblicazione: (2025)
di: Lahtinen, Kalle, et al.
Pubblicazione: (2025)
Thinking with Patterns: Breaking the Perceptual Bottleneck in Visual Planning via Pattern Induction
di: Jian, Yichang, et al.
Pubblicazione: (2026)
di: Jian, Yichang, et al.
Pubblicazione: (2026)
Concept Bottleneck Large Language Models
di: Sun, Chung-En, et al.
Pubblicazione: (2024)
di: Sun, Chung-En, et al.
Pubblicazione: (2024)
Robust Training of Vector Quantized Bottleneck Models
di: Łańcucki, Adrian, et al.
Pubblicazione: (2020)
di: Łańcucki, Adrian, et al.
Pubblicazione: (2020)
Scaling Stick-Breaking Attention: An Efficient Implementation and In-depth Study
di: Tan, Shawn, et al.
Pubblicazione: (2024)
di: Tan, Shawn, et al.
Pubblicazione: (2024)
RNNs are not Transformers (Yet): The Key Bottleneck on In-context Retrieval
di: Wen, Kaiyue, et al.
Pubblicazione: (2024)
di: Wen, Kaiyue, et al.
Pubblicazione: (2024)
ShifaMind: A Multiplicative Concept Bottleneck for Interpretable ICD-10 Coding
di: Syed, Mohammed Sameer, et al.
Pubblicazione: (2026)
di: Syed, Mohammed Sameer, et al.
Pubblicazione: (2026)
Policy Learning with a Language Bottleneck
di: Srivastava, Megha, et al.
Pubblicazione: (2024)
di: Srivastava, Megha, et al.
Pubblicazione: (2024)
Breaking Symmetry When Training Transformers
di: Zuo, Chunsheng, et al.
Pubblicazione: (2024)
di: Zuo, Chunsheng, et al.
Pubblicazione: (2024)
Single Character Perturbations Break LLM Alignment
di: Lin, Leon, et al.
Pubblicazione: (2024)
di: Lin, Leon, et al.
Pubblicazione: (2024)
Language Bottleneck Models for Qualitative Knowledge State Modeling
di: Berthon, Antonin, et al.
Pubblicazione: (2025)
di: Berthon, Antonin, et al.
Pubblicazione: (2025)
Circuit Breaking: Removing Model Behaviors with Targeted Ablation
di: Li, Maximilian, et al.
Pubblicazione: (2023)
di: Li, Maximilian, et al.
Pubblicazione: (2023)
Why Softmax Attention Outperforms Linear Attention
di: Deng, Yichuan, et al.
Pubblicazione: (2023)
di: Deng, Yichuan, et al.
Pubblicazione: (2023)
Does Alignment Tuning Really Break LLMs' Internal Confidence?
di: Oh, Hongseok, et al.
Pubblicazione: (2024)
di: Oh, Hongseok, et al.
Pubblicazione: (2024)
Break the Sequential Dependency of LLM Inference Using Lookahead Decoding
di: Fu, Yichao, et al.
Pubblicazione: (2024)
di: Fu, Yichao, et al.
Pubblicazione: (2024)
Breaking MLPerf Training: A Case Study on Optimizing BERT
di: Kim, Yongdeok, et al.
Pubblicazione: (2024)
di: Kim, Yongdeok, et al.
Pubblicazione: (2024)
STBLLM: Breaking the 1-Bit Barrier with Structured Binary LLMs
di: Dong, Peijie, et al.
Pubblicazione: (2024)
di: Dong, Peijie, et al.
Pubblicazione: (2024)
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity
di: Lochab, Anamika, et al.
Pubblicazione: (2026)
di: Lochab, Anamika, et al.
Pubblicazione: (2026)
SEA: Sparse Linear Attention with Estimated Attention Mask
di: Lee, Heejun, et al.
Pubblicazione: (2023)
di: Lee, Heejun, et al.
Pubblicazione: (2023)
Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers
di: Wong, Liang Ze
Pubblicazione: (2025)
di: Wong, Liang Ze
Pubblicazione: (2025)
What Makes and Breaks Safety Fine-tuning? A Mechanistic Study
di: Jain, Samyak, et al.
Pubblicazione: (2024)
di: Jain, Samyak, et al.
Pubblicazione: (2024)
Breaking the Benchmark: Revealing LLM Bias via Minimal Contextual Augmentation
di: Miandoab, Kaveh Eskandari, et al.
Pubblicazione: (2025)
di: Miandoab, Kaveh Eskandari, et al.
Pubblicazione: (2025)
SPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive Learning
di: Yu, Qifan, et al.
Pubblicazione: (2026)
di: Yu, Qifan, et al.
Pubblicazione: (2026)
Breaking BERT: Gradient Attack on Twitter Sentiment Analysis for Targeted Misclassification
di: Subedi, Akil Raj, et al.
Pubblicazione: (2025)
di: Subedi, Akil Raj, et al.
Pubblicazione: (2025)
Don't Pay Attention, PLANT It: Pretraining Attention via Learning-to-Rank
di: Roy, Debjyoti Saha, et al.
Pubblicazione: (2024)
di: Roy, Debjyoti Saha, et al.
Pubblicazione: (2024)
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
di: Qiu, Quantong, et al.
Pubblicazione: (2026)
di: Qiu, Quantong, et al.
Pubblicazione: (2026)
PowerAttention: Exponentially Scaling of Receptive Fields for Effective Sparse Attention
di: Chen, Lida, et al.
Pubblicazione: (2025)
di: Chen, Lida, et al.
Pubblicazione: (2025)
Heterogeneity in Formal Linguistic Competence of Language Models: Is Data the Real Bottleneck?
di: Renduchintala, H S V N S Kowndinya, et al.
Pubblicazione: (2026)
di: Renduchintala, H S V N S Kowndinya, et al.
Pubblicazione: (2026)
LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety
di: Yang, Junxiao, et al.
Pubblicazione: (2026)
di: Yang, Junxiao, et al.
Pubblicazione: (2026)
ABBEL: LLM Agents Acting through Belief Bottlenecks Expressed in Language
di: Lidayan, Aly, et al.
Pubblicazione: (2025)
di: Lidayan, Aly, et al.
Pubblicazione: (2025)
Learning to Attribute with Attention
di: Cohen-Wang, Benjamin, et al.
Pubblicazione: (2025)
di: Cohen-Wang, Benjamin, et al.
Pubblicazione: (2025)
Exclusive Self Attention
di: Zhai, Shuangfei
Pubblicazione: (2026)
di: Zhai, Shuangfei
Pubblicazione: (2026)
Scale-invariant Attention
di: Anson, Ben, et al.
Pubblicazione: (2025)
di: Anson, Ben, et al.
Pubblicazione: (2025)
The Million-Label NER: Breaking Scale Barriers with GLiNER bi-encoder
di: Stepanov, Ihor, et al.
Pubblicazione: (2026)
di: Stepanov, Ihor, et al.
Pubblicazione: (2026)
Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters
di: Shyam, Vasudev, et al.
Pubblicazione: (2024)
di: Shyam, Vasudev, et al.
Pubblicazione: (2024)
Stochastic Attention: Connectome-Inspired Randomized Routing for Expressive Linear-Time Attention
di: Jin, Zehao, et al.
Pubblicazione: (2026)
di: Jin, Zehao, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Breaking Training Bottlenecks: Effective and Stable Reinforcement Learning for Coding Models
di: Li, Zongqian, et al.
Pubblicazione: (2026) -
MUDDFormer: Breaking Residual Bottlenecks in Transformers via Multiway Dynamic Dense Connections
di: Xiao, Da, et al.
Pubblicazione: (2025) -
Breaking Quadratic Barriers: A Non-Attention LLM for Ultra-Long Context Horizons
di: Kiruluta, Andrew, et al.
Pubblicazione: (2025) -
Efficient Knowledge Injection in LLMs via Self-Distillation
di: Kujanpää, Kalle, et al.
Pubblicazione: (2024) -
Investigating Affect Mining Techniques for Annotation Sample Selection in the Creation of Finnish Affective Speech Corpus
di: Lahtinen, Kalle, et al.
Pubblicazione: (2025)