The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
Fuente:
arXiv
Saved in:
| Main Authors: | Fichtl, Alexander M., Bohn, Jeremias, Kelber, Josefin, Mosca, Edoardo, Groh, Georg |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adapter-based Approaches to Knowledge-enhanced Language Models -- A Survey
by: Fichtl, Alexander, et al.
Published: (2024)
by: Fichtl, Alexander, et al.
Published: (2024)
Simpler becomes Harder: Do LLMs Exhibit a Coherent Behavior on Simplified Corpora?
by: Anschütz, Miriam, et al.
Published: (2024)
by: Anschütz, Miriam, et al.
Published: (2024)
Granuscore: A Reference-Free Measure of Granularity for Text Analysis and Question Answering
by: Ellinger, Lukas, et al.
Published: (2026)
by: Ellinger, Lukas, et al.
Published: (2026)
Depth-Recurrent Attention Mixtures: Giving Latent Reasoning the Attention it Deserves
by: Knupp, Jonas, et al.
Published: (2026)
by: Knupp, Jonas, et al.
Published: (2026)
CAT: Circular-Convolutional Attention for Sub-Quadratic Transformers
by: Yamada, Yoshihiro
Published: (2025)
by: Yamada, Yoshihiro
Published: (2025)
It Depends: Resolving Referential Ambiguity in Minimal Contexts with Commonsense Knowledge
by: Ellinger, Lukas, et al.
Published: (2025)
by: Ellinger, Lukas, et al.
Published: (2025)
Simplifications are Absolutists: How Simplified Language Reduces Word Sense Awareness in LLM-Generated Definitions
by: Ellinger, Lukas, et al.
Published: (2025)
by: Ellinger, Lukas, et al.
Published: (2025)
Images Speak Volumes: User-Centric Assessment of Image Generation for Accessible Communication
by: Anschütz, Miriam, et al.
Published: (2024)
by: Anschütz, Miriam, et al.
Published: (2024)
Cross-lingual Text Classification Transfer: The Case of Ukrainian
by: Dementieva, Daryna, et al.
Published: (2024)
by: Dementieva, Daryna, et al.
Published: (2024)
To Bias or Not to Bias: Detecting bias in News with bias-detector
by: Ghosh, Himel, et al.
Published: (2025)
by: Ghosh, Himel, et al.
Published: (2025)
Toxicity Classification in Ukrainian
by: Dementieva, Daryna, et al.
Published: (2024)
by: Dementieva, Daryna, et al.
Published: (2024)
Profiling Bias in LLMs: Stereotype Dimensions in Contextual Word Embeddings
by: Schuster, Carolin M., et al.
Published: (2024)
by: Schuster, Carolin M., et al.
Published: (2024)
Semantic Component Analysis: Introducing Multi-Topic Distributions to Clustering-Based Topic Modeling
by: Eichin, Florian, et al.
Published: (2024)
by: Eichin, Florian, et al.
Published: (2024)
Beyond Hate: Differentiating Uncivil and Intolerant Speech in Multimodal Content Moderation
by: Herrmann, Nils A., et al.
Published: (2026)
by: Herrmann, Nils A., et al.
Published: (2026)
A Comprehensive Evaluation of Cognitive Biases in LLMs
by: Malberg, Simon, et al.
Published: (2024)
by: Malberg, Simon, et al.
Published: (2024)
Tuning Into Bias: A Computational Study of Gender Bias in Song Lyrics
by: Chen, Danqing, et al.
Published: (2024)
by: Chen, Danqing, et al.
Published: (2024)
TUM-MiKaNi at SemEval-2025 Task 3: Towards Multilingual and Knowledge-Aware Non-factual Hallucination Identification
by: Anschütz, Miriam, et al.
Published: (2025)
by: Anschütz, Miriam, et al.
Published: (2025)
Safer Reasoning Traces: Measuring and Mitigating Chain-of-Thought Leakage in LLMs
by: Ahrend, Patrick, et al.
Published: (2026)
by: Ahrend, Patrick, et al.
Published: (2026)
The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs
by: Nawrot, Piotr, et al.
Published: (2025)
by: Nawrot, Piotr, et al.
Published: (2025)
Crafting Tomorrow's Headlines: Neural News Generation and Detection in English, Turkish, Hungarian, and Persian
by: Üyük, Cem, et al.
Published: (2024)
by: Üyük, Cem, et al.
Published: (2024)
Beat and Downbeat Tracking in Performance MIDI Using an End-to-End Transformer Architecture
by: Murgul, Sebastian, et al.
Published: (2025)
by: Murgul, Sebastian, et al.
Published: (2025)
Extra Global Attention Designation Using Keyword Detection in Sparse Transformer Architectures
by: Lucas, Evan, et al.
Published: (2024)
by: Lucas, Evan, et al.
Published: (2024)
Do We Need Bigger Models for Science? Task-Aware Retrieval with Small Language Models
by: Kelber, Florian, et al.
Published: (2026)
by: Kelber, Florian, et al.
Published: (2026)
German4All -- A Dataset and Model for Readability-Controlled Paraphrasing in German
by: Anschütz, Miriam, et al.
Published: (2025)
by: Anschütz, Miriam, et al.
Published: (2025)
DIALECTIC: A Multi-Agent System for Startup Evaluation
by: Bae, Jae Yoon, et al.
Published: (2026)
by: Bae, Jae Yoon, et al.
Published: (2026)
Multilingual European Language Models: Benchmarking Approaches and Challenges
by: Barth, Fabio, et al.
Published: (2025)
by: Barth, Fabio, et al.
Published: (2025)
Pose-Based Sign Language Spotting via an End-to-End Encoder Architecture
by: Johnny, Samuel Ebimobowei, et al.
Published: (2025)
by: Johnny, Samuel Ebimobowei, et al.
Published: (2025)
Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning
by: Wang, Junxuan, et al.
Published: (2025)
by: Wang, Junxuan, et al.
Published: (2025)
The Counting Power of Transformers
by: Sälzer, Marco, et al.
Published: (2025)
by: Sälzer, Marco, et al.
Published: (2025)
VeriLocc: End-to-End Cross-Architecture Register Allocation via LLM
by: Jin, Lesheng, et al.
Published: (2025)
by: Jin, Lesheng, et al.
Published: (2025)
Assessment of RAG and Fine-Tuning for Industrial Question-Answering-Applications
by: Sturm, Jakob, et al.
Published: (2026)
by: Sturm, Jakob, et al.
Published: (2026)
Extracting Rule-based Descriptions of Attention Features in Transformers
by: Friedman, Dan, et al.
Published: (2025)
by: Friedman, Dan, et al.
Published: (2025)
A Transformer with Stack Attention
by: Li, Jiaoda, et al.
Published: (2024)
by: Li, Jiaoda, et al.
Published: (2024)
Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention
by: Guo, Zhenyu, et al.
Published: (2025)
by: Guo, Zhenyu, et al.
Published: (2025)
Rationalizing Transformer Predictions via End-To-End Differentiable Self-Training
by: Brinner, Marc, et al.
Published: (2025)
by: Brinner, Marc, et al.
Published: (2025)
Symmetric Dot-Product Attention for Efficient Training of BERT Language Models
by: Courtois, Martin, et al.
Published: (2024)
by: Courtois, Martin, et al.
Published: (2024)
Characterizing the Expressivity of Local Attention in Transformers
by: Li, Jiaoda, et al.
Published: (2026)
by: Li, Jiaoda, et al.
Published: (2026)
Anisotropy Is Inherent to Self-Attention in Transformers
by: Godey, Nathan, et al.
Published: (2024)
by: Godey, Nathan, et al.
Published: (2024)
A Training-free Sub-quadratic Cost Transformer Model Serving Framework With Hierarchically Pruned Attention
by: Lee, Heejun, et al.
Published: (2024)
by: Lee, Heejun, et al.
Published: (2024)
Canvas: End-to-End Kernel Architecture Search in Neural Networks
by: Zhao, Chenggang, et al.
Published: (2023)
by: Zhao, Chenggang, et al.
Published: (2023)
Similar Items
-
Adapter-based Approaches to Knowledge-enhanced Language Models -- A Survey
by: Fichtl, Alexander, et al.
Published: (2024) -
Simpler becomes Harder: Do LLMs Exhibit a Coherent Behavior on Simplified Corpora?
by: Anschütz, Miriam, et al.
Published: (2024) -
Granuscore: A Reference-Free Measure of Granularity for Text Analysis and Question Answering
by: Ellinger, Lukas, et al.
Published: (2026) -
Depth-Recurrent Attention Mixtures: Giving Latent Reasoning the Attention it Deserves
by: Knupp, Jonas, et al.
Published: (2026) -
CAT: Circular-Convolutional Attention for Sub-Quadratic Transformers
by: Yamada, Yoshihiro
Published: (2025)