Reducing the Transformer Architecture to a Minimum
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bermeitinger, Bernhard, Hrycej, Tomas, Pavone, Massimo, Kath, Julianus, Handschuh, Siegfried |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Make Deep Networks Shallow Again
von: Bermeitinger, Bernhard, et al.
Veröffentlicht: (2023)
von: Bermeitinger, Bernhard, et al.
Veröffentlicht: (2023)
A Convexity-dependent Two-Phase Training Algorithm for Deep Neural Networks
von: Hrycej, Tomas, et al.
Veröffentlicht: (2025)
von: Hrycej, Tomas, et al.
Veröffentlicht: (2025)
Efficient Neural Network Training via Subset Pretraining
von: Spörer, Jan, et al.
Veröffentlicht: (2024)
von: Spörer, Jan, et al.
Veröffentlicht: (2024)
Is More Data Worth the Cost? Dataset Scaling Laws in a Tiny Attention-Only Decoder
von: Wiegand, Götz-Henrik, et al.
Veröffentlicht: (2026)
von: Wiegand, Götz-Henrik, et al.
Veröffentlicht: (2026)
Scalable Reinforcement Learning-based Neural Architecture Search
von: Cassimon, Amber, et al.
Veröffentlicht: (2024)
von: Cassimon, Amber, et al.
Veröffentlicht: (2024)
What Is the Minimum Architecture for Prolepsis? Early Irrevocable Commitment Across Tasks in Small Transformers
von: Jacopin, Éric
Veröffentlicht: (2026)
von: Jacopin, Éric
Veröffentlicht: (2026)
Designing a Classifier for Active Fire Detection from Multispectral Satellite Imagery Using Neural Architecture Search
von: Cassimon, Amber, et al.
Veröffentlicht: (2024)
von: Cassimon, Amber, et al.
Veröffentlicht: (2024)
The Speed-up Factor: A Quantitative Multi-Iteration Active Learning Performance Metric
von: Kath, Hannes, et al.
Veröffentlicht: (2026)
von: Kath, Hannes, et al.
Veröffentlicht: (2026)
Minimum Reduced-Order Models via Causal Inference
von: Chen, Nan, et al.
Veröffentlicht: (2024)
von: Chen, Nan, et al.
Veröffentlicht: (2024)
Reducing Hallucination in Enterprise AI Workflows via Hybrid Utility Minimum Bayes Risk (HUMBR)
von: Fang, Chenhao, et al.
Veröffentlicht: (2026)
von: Fang, Chenhao, et al.
Veröffentlicht: (2026)
Hybrid Dual-Path Linear Transformations for Efficient Transformer Architectures
von: Khasia, Vladimer
Veröffentlicht: (2026)
von: Khasia, Vladimer
Veröffentlicht: (2026)
On Limitations of the Transformer Architecture
von: Peng, Binghui, et al.
Veröffentlicht: (2024)
von: Peng, Binghui, et al.
Veröffentlicht: (2024)
In Transformer We Trust? A Perspective on Transformer Architecture Failure Modes
von: Mondal, Trishit, et al.
Veröffentlicht: (2026)
von: Mondal, Trishit, et al.
Veröffentlicht: (2026)
Separations in the Representational Capabilities of Transformers and Recurrent Architectures
von: Bhattamishra, Satwik, et al.
Veröffentlicht: (2024)
von: Bhattamishra, Satwik, et al.
Veröffentlicht: (2024)
Approximation Rate of the Transformer Architecture for Sequence Modeling
von: Jiang, Haotian, et al.
Veröffentlicht: (2023)
von: Jiang, Haotian, et al.
Veröffentlicht: (2023)
The Power of Architecture: Deep Dive into Transformer Architectures for Long-Term Time Series Forecasting
von: Shen, Lefei, et al.
Veröffentlicht: (2025)
von: Shen, Lefei, et al.
Veröffentlicht: (2025)
Minimum-Excess-Work Guidance
von: Kolloff, Christopher, et al.
Veröffentlicht: (2025)
von: Kolloff, Christopher, et al.
Veröffentlicht: (2025)
Architecture Determines Observability of Transformers
von: Carmichael, Thomas
Veröffentlicht: (2026)
von: Carmichael, Thomas
Veröffentlicht: (2026)
Born a Transformer -- Always a Transformer? On the Effect of Pretraining on Architectural Abilities
von: Jobanputra, Mayank, et al.
Veröffentlicht: (2025)
von: Jobanputra, Mayank, et al.
Veröffentlicht: (2025)
Disentangling and Integrating Relational and Sensory Information in Transformer Architectures
von: Altabaa, Awni, et al.
Veröffentlicht: (2024)
von: Altabaa, Awni, et al.
Veröffentlicht: (2024)
Rethinking Graph Transformer Architecture Design for Node Classification
von: Zhou, Jiajun, et al.
Veröffentlicht: (2024)
von: Zhou, Jiajun, et al.
Veröffentlicht: (2024)
SPARTAN: A Sparse Transformer World Model Attending to What Matters
von: Lei, Anson, et al.
Veröffentlicht: (2024)
von: Lei, Anson, et al.
Veröffentlicht: (2024)
On the Universality of Transformer Architectures; How Much Attention Is Enough?
von: Abbasi, Amirreza, et al.
Veröffentlicht: (2025)
von: Abbasi, Amirreza, et al.
Veröffentlicht: (2025)
Self-Attention as a Parametric Endofunctor: A Categorical Framework for Transformer Architectures
von: O'Neill, Charles
Veröffentlicht: (2025)
von: O'Neill, Charles
Veröffentlicht: (2025)
Deep Learning Warm Starts for Trajectory Optimization on the International Space Station
von: Banerjee, Somrita, et al.
Veröffentlicht: (2025)
von: Banerjee, Somrita, et al.
Veröffentlicht: (2025)
Adaptive Meta-Learning for Identification of Rover-Terrain Dynamics
von: Banerjee, S., et al.
Veröffentlicht: (2020)
von: Banerjee, S., et al.
Veröffentlicht: (2020)
Finding Clustering Algorithms in the Transformer Architecture
von: Clarkson, Kenneth L., et al.
Veröffentlicht: (2025)
von: Clarkson, Kenneth L., et al.
Veröffentlicht: (2025)
GNN-Transformer Cooperative Architecture for Trustworthy Graph Contrastive Learning
von: Liang, Jianqing, et al.
Veröffentlicht: (2024)
von: Liang, Jianqing, et al.
Veröffentlicht: (2024)
Surprise Potential as a Measure of Interactivity in Driving Scenarios
von: Ding, Wenhao, et al.
Veröffentlicht: (2025)
von: Ding, Wenhao, et al.
Veröffentlicht: (2025)
NASH: Neural Architecture and Accelerator Search for Multiplication-Reduced Hybrid Models
von: Xu, Yang, et al.
Veröffentlicht: (2024)
von: Xu, Yang, et al.
Veröffentlicht: (2024)
From PEFT to DEFT: Parameter Efficient Finetuning for Reducing Activation Density in Transformers
von: Runwal, Bharat, et al.
Veröffentlicht: (2024)
von: Runwal, Bharat, et al.
Veröffentlicht: (2024)
Quantum Re-Uploading for Calorimetry: Optimized Architectures with Extended Expressivity
von: Cassé, Léa, et al.
Veröffentlicht: (2024)
von: Cassé, Léa, et al.
Veröffentlicht: (2024)
Longitudinal Targeted Minimum Loss-based Estimation with Temporal-Difference Heterogeneous Transformer
von: Shirakawa, Toru, et al.
Veröffentlicht: (2024)
von: Shirakawa, Toru, et al.
Veröffentlicht: (2024)
Minimum-Norm Interpolation Under Covariate Shift
von: Mallinar, Neil, et al.
Veröffentlicht: (2024)
von: Mallinar, Neil, et al.
Veröffentlicht: (2024)
Minimum distance classification for nonlinear dynamical systems
von: Martinez, Dominique
Veröffentlicht: (2026)
von: Martinez, Dominique
Veröffentlicht: (2026)
AnchorGT: Efficient and Flexible Attention Architecture for Scalable Graph Transformers
von: Zhu, Wenhao, et al.
Veröffentlicht: (2024)
von: Zhu, Wenhao, et al.
Veröffentlicht: (2024)
Self-Attention as Distributional Projection: A Unified Interpretation of Transformer Architecture
von: Mehta, Nihal
Veröffentlicht: (2025)
von: Mehta, Nihal
Veröffentlicht: (2025)
A Transformer-based Autoregressive Decoder Architecture for Hierarchical Text Classification
von: Yousef, Younes, et al.
Veröffentlicht: (2025)
von: Yousef, Younes, et al.
Veröffentlicht: (2025)
Self-Supervised Transformer Architecture for Change Detection in Radio Access Networks
von: Kozlov, Igor, et al.
Veröffentlicht: (2023)
von: Kozlov, Igor, et al.
Veröffentlicht: (2023)
FairJob: A Real-World Dataset for Fairness in Online Systems
von: Vladimirova, Mariia, et al.
Veröffentlicht: (2024)
von: Vladimirova, Mariia, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Make Deep Networks Shallow Again
von: Bermeitinger, Bernhard, et al.
Veröffentlicht: (2023) -
A Convexity-dependent Two-Phase Training Algorithm for Deep Neural Networks
von: Hrycej, Tomas, et al.
Veröffentlicht: (2025) -
Efficient Neural Network Training via Subset Pretraining
von: Spörer, Jan, et al.
Veröffentlicht: (2024) -
Is More Data Worth the Cost? Dataset Scaling Laws in a Tiny Attention-Only Decoder
von: Wiegand, Götz-Henrik, et al.
Veröffentlicht: (2026) -
Scalable Reinforcement Learning-based Neural Architecture Search
von: Cassimon, Amber, et al.
Veröffentlicht: (2024)