Transformer Scalability Crisis: The First Comprehensive Empirical Analysis of Performance Walls in Modern Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Moghadasi, Mahdi Naser, Ghaderi, Faezeh |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Neural Activation Patterns Across Language Model Architectures: A Comprehensive Analysis of Cognitive Task Performance
par: Naser-Moghadasi, Mahdi, et autres
Publié: (2026)
par: Naser-Moghadasi, Mahdi, et autres
Publié: (2026)
Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales
par: Salfati, Samuel
Publié: (2026)
par: Salfati, Samuel
Publié: (2026)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
par: Fadli, Samih
Publié: (2025)
par: Fadli, Samih
Publié: (2025)
TensorLens: End-to-End Transformer Analysis via High-Order Attention Tensors
par: Atad, Ido Andrew, et autres
Publié: (2026)
par: Atad, Ido Andrew, et autres
Publié: (2026)
Merge-Bench: Resolve Merge Conflicts with Large Language Models
par: Schesch, Benedikt, et autres
Publié: (2026)
par: Schesch, Benedikt, et autres
Publié: (2026)
Scalable GPU-Accelerated Euler Characteristic Curves: Optimization and Differentiable Learning for PyTorch
par: Saxena, Udit
Publié: (2025)
par: Saxena, Udit
Publié: (2025)
Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring
par: Heyman, Alex, et autres
Publié: (2025)
par: Heyman, Alex, et autres
Publié: (2025)
Revisiting LRP: Positional Attribution as the Missing Ingredient for Transformer Explainability
par: Bakish, Yarden, et autres
Publié: (2025)
par: Bakish, Yarden, et autres
Publié: (2025)
Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive Survey
par: Huang, Yunpeng, et autres
Publié: (2023)
par: Huang, Yunpeng, et autres
Publié: (2023)
Learned Relay Representations for Forward-Thinking Discrete Diffusion Models
par: Rozonoyer, Benjamin, et autres
Publié: (2026)
par: Rozonoyer, Benjamin, et autres
Publié: (2026)
Recurrent Memory-Augmented Transformers with Chunked Attention for Long-Context Language Modeling
par: Kashyap, Ankit
Publié: (2025)
par: Kashyap, Ankit
Publié: (2025)
ACE: Exploring Activation Cosine Similarity and Variance for Accurate and Calibration-Efficient LLM Pruning
par: Mi, Zhendong, et autres
Publié: (2025)
par: Mi, Zhendong, et autres
Publié: (2025)
Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs
par: Easley, Eric, et autres
Publié: (2026)
par: Easley, Eric, et autres
Publié: (2026)
Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features
par: McCann, Jordan F.
Publié: (2026)
par: McCann, Jordan F.
Publié: (2026)
Super Apriel: One Checkpoint, Many Speeds
par: Labs, SLAM, et autres
Publié: (2026)
par: Labs, SLAM, et autres
Publié: (2026)
Synergy over Discrepancy: A Partition-Based Approach to Multi-Domain LLM Fine-Tuning
par: Ye, Hua, et autres
Publié: (2025)
par: Ye, Hua, et autres
Publié: (2025)
KerZOO: Kernel Function Informed Zeroth-Order Optimization for Accurate and Accelerated LLM Fine-Tuning
par: Mi, Zhendong, et autres
Publié: (2025)
par: Mi, Zhendong, et autres
Publié: (2025)
QuAnTS: Question Answering on Time Series
par: Divo, Felix, et autres
Publié: (2025)
par: Divo, Felix, et autres
Publié: (2025)
Why LoRA Resists Label Noise: A Theoretical Framework for Noise-Robust Parameter-Efficient Fine-Tuning
par: Steele, Brady
Publié: (2026)
par: Steele, Brady
Publié: (2026)
Extreme AutoML: Analysis of Classification, Regression, and NLP Performance
par: Ratner, Edward, et autres
Publié: (2024)
par: Ratner, Edward, et autres
Publié: (2024)
Language Models Are Implicitly Continuous
par: Marro, Samuele, et autres
Publié: (2025)
par: Marro, Samuele, et autres
Publié: (2025)
Latent Cache Flow: Model-to-Model Communication Without Text
par: Rossi, Maximillian, et autres
Publié: (2026)
par: Rossi, Maximillian, et autres
Publié: (2026)
Pre-trained Models Perform the Best When Token Distributions Follow Zipf's Law
par: He, Yanjin, et autres
Publié: (2025)
par: He, Yanjin, et autres
Publié: (2025)
Combining Language and Topic Models for Hierarchical Text Classification
par: Toit, Jaco du, et autres
Publié: (2025)
par: Toit, Jaco du, et autres
Publié: (2025)
Memory Bank Compression for Continual Adaptation of Large Language Models
par: Katraouras, Thomas, et autres
Publié: (2026)
par: Katraouras, Thomas, et autres
Publié: (2026)
Transformers Boost the Performance of Decision Trees on Tabular Data across Sample Sizes
par: Jayawardhana, Mayuka, et autres
Publié: (2025)
par: Jayawardhana, Mayuka, et autres
Publié: (2025)
Rethinking Addressing in Language Models via Contexualized Equivariant Positional Encoding
par: Zhu, Jiajun, et autres
Publié: (2025)
par: Zhu, Jiajun, et autres
Publié: (2025)
The Geometry of Thought: How Scale Restructures Reasoning In Large Language Models
par: Anderson, Samuel Cyrenius
Publié: (2026)
par: Anderson, Samuel Cyrenius
Publié: (2026)
Continuous-Depth Transformers with Learned Control Dynamics
par: Jemley, Peter
Publié: (2026)
par: Jemley, Peter
Publié: (2026)
PoTS: Proof-of-Training-Steps for Backdoor Detection in Large Language Models
par: Seddik, Issam, et autres
Publié: (2025)
par: Seddik, Issam, et autres
Publié: (2025)
Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models
par: Shravan, Rohan
Publié: (2026)
par: Shravan, Rohan
Publié: (2026)
How Language Models Process Out-of-Distribution Inputs: A Two-Pathway Framework
par: Saghir, Hamidreza
Publié: (2026)
par: Saghir, Hamidreza
Publié: (2026)
Reasoning Large Language Model Errors Arise from Hallucinating Critical Problem Features
par: Heyman, Alex, et autres
Publié: (2025)
par: Heyman, Alex, et autres
Publié: (2025)
Shattered Compositionality: Counterintuitive Learning Dynamics of Transformers for Arithmetic
par: Zhao, Xingyu, et autres
Publié: (2026)
par: Zhao, Xingyu, et autres
Publié: (2026)
The Anti-Ouroboros Effect: Emergent Resilience in Large Language Models from Recursive Selective Feedback
par: Adapala, Sai Teja Reddy
Publié: (2025)
par: Adapala, Sai Teja Reddy
Publié: (2025)
Thread Detection and Response Generation using Transformers with Prompt Optimisation
par: T, Kevin Joshua, et autres
Publié: (2024)
par: T, Kevin Joshua, et autres
Publié: (2024)
Task-Conditioned Routing Signatures in Sparse Mixture-of-Experts Transformers
par: Avinash, Mynampati Sri Ranganadha
Publié: (2026)
par: Avinash, Mynampati Sri Ranganadha
Publié: (2026)
Discovering Transformer Circuits via a Hybrid Attribution and Pruning Framework
par: Gu, Hao, et autres
Publié: (2025)
par: Gu, Hao, et autres
Publié: (2025)
Future Token Prediction -- Causal Language Modelling with Per-Token Semantic State Vector for Multi-Token Prediction
par: Walker, Nicholas
Publié: (2024)
par: Walker, Nicholas
Publié: (2024)
CircuitProbe: Predicting Reasoning Circuits in Transformers via Stability Zone Detection
par: Panuganti, Rajkiran
Publié: (2026)
par: Panuganti, Rajkiran
Publié: (2026)
Documents similaires
-
Neural Activation Patterns Across Language Model Architectures: A Comprehensive Analysis of Cognitive Task Performance
par: Naser-Moghadasi, Mahdi, et autres
Publié: (2026) -
Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales
par: Salfati, Samuel
Publié: (2026) -
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
par: Fadli, Samih
Publié: (2025) -
TensorLens: End-to-End Transformer Analysis via High-Order Attention Tensors
par: Atad, Ido Andrew, et autres
Publié: (2026) -
Merge-Bench: Resolve Merge Conflicts with Large Language Models
par: Schesch, Benedikt, et autres
Publié: (2026)