Optimizing Inference in Transformer-Based Models: A Multi-Method Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Ho, Siu Hang, Ganesan, Prasad, Duong, Nguyen, Schlabig, Daniel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Explicit Dropout: Deterministic Regularization for Transformer Architectures
by: Agrawal, Vidhi, et al.
Published: (2026)
by: Agrawal, Vidhi, et al.
Published: (2026)
Massive Redundancy in Gradient Transport Enables Sparse Online Learning
by: Merin, Aur Shalev
Published: (2026)
by: Merin, Aur Shalev
Published: (2026)
Is Cambodia the World's Largest Cashew Producer?
by: Chaya, Veasna, et al.
Published: (2024)
by: Chaya, Veasna, et al.
Published: (2024)
Closing the Theory-Practice Gap in Spiking Transformers via Effective Dimension
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
H-Model: Dynamic Neural Architectures for Adaptive Processing
by: Hospodarchuk, Dmytro
Published: (2025)
by: Hospodarchuk, Dmytro
Published: (2025)
Scaling Laws in the Tiny Regime: How Small Models Change Their Mistakes
by: Alnemari, Mohammed, et al.
Published: (2026)
by: Alnemari, Mohammed, et al.
Published: (2026)
torchsom: The Reference PyTorch Library for Self-Organizing Maps
by: Berthier, Louis, et al.
Published: (2025)
by: Berthier, Louis, et al.
Published: (2025)
Tricks and Plug-ins for Gradient Boosting with Transformers
by: Fang, Biyi, et al.
Published: (2025)
by: Fang, Biyi, et al.
Published: (2025)
PolyGLU: State-Conditional Activation Routing in Transformer Feed-Forward Networks
by: Medeiros, Daniel Nobrega
Published: (2026)
by: Medeiros, Daniel Nobrega
Published: (2026)
CellARC: Measuring Intelligence with Cellular Automata
by: Lžičař, Miroslav
Published: (2025)
by: Lžičař, Miroslav
Published: (2025)
multivariateGPT: a decoder-only transformer for multivariate categorical and numeric data
by: Loza, Andrew J., et al.
Published: (2025)
by: Loza, Andrew J., et al.
Published: (2025)
VDW-GNNs: Vector diffusion wavelets for geometric graph neural networks
by: Johnson, David R., et al.
Published: (2025)
by: Johnson, David R., et al.
Published: (2025)
DecompKAN: Decomposed Patch-KAN for Long-Term Time Series Forecasting
by: Mysore, Naveen
Published: (2026)
by: Mysore, Naveen
Published: (2026)
Efficient Morphology-Control Co-Design via Stackelberg Proximal Policy Optimization
by: Dai, Yanning, et al.
Published: (2026)
by: Dai, Yanning, et al.
Published: (2026)
SigGate-GT: Taming Over-Smoothing in Graph Transformers via Sigmoid-Gated Attention
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
CVCM Track Circuits Pre-emptive Failure Diagnostics for Predictive Maintenance Using Deep Neural Networks
by: Mukherjee, Debdeep, et al.
Published: (2025)
by: Mukherjee, Debdeep, et al.
Published: (2025)
EchoLSTM: A Self-Reflective Recurrent Network for Stabilizing Long-Range Memory
by: K, Prasanth K, et al.
Published: (2025)
by: K, Prasanth K, et al.
Published: (2025)
Predictive Modeling of Maritime Radar Data Using Transformer Architecture
by: Qesaraku, Bjorna, et al.
Published: (2025)
by: Qesaraku, Bjorna, et al.
Published: (2025)
Predicting Traffic Accident Severity with Deep Neural Networks
by: Bibb, Meghan, et al.
Published: (2025)
by: Bibb, Meghan, et al.
Published: (2025)
An in-depth look at approximation via deep and narrow neural networks
by: Dommel, Joris, et al.
Published: (2025)
by: Dommel, Joris, et al.
Published: (2025)
Adaptive Negative Scheduling for Graph Contrastive Learning
by: Ali, Adnan, et al.
Published: (2026)
by: Ali, Adnan, et al.
Published: (2026)
The Impact of Data Characteristics on GNN Evaluation for Detecting Fake News
by: Karn, Isha, et al.
Published: (2025)
by: Karn, Isha, et al.
Published: (2025)
The Normalized Difference Layer: A Differentiable Spectral Index Formulation for Deep Learning
by: Lotfi, Ali, et al.
Published: (2026)
by: Lotfi, Ali, et al.
Published: (2026)
Complex-Valued Phase-Coherent Transformer
by: Hioki, Leona
Published: (2026)
by: Hioki, Leona
Published: (2026)
Contract-Driven QoE Auditing for Speech and Singing Services: From MOS Regression to Service Graphs
by: Du, Wenzhang
Published: (2025)
by: Du, Wenzhang
Published: (2025)
Optimized Gradient Clipping for Noisy Label Learning
by: Ye, Xichen, et al.
Published: (2024)
by: Ye, Xichen, et al.
Published: (2024)
Scalable, Technology-Agnostic Diagnosis and Predictive Maintenance for Point Machine using Deep Learning
by: Di Santi, Eduardo, et al.
Published: (2025)
by: Di Santi, Eduardo, et al.
Published: (2025)
Revisiting GAN with Bayes-Optimal Discrimination
by: Naeini, Mohammadreza Tavasoli, et al.
Published: (2025)
by: Naeini, Mohammadreza Tavasoli, et al.
Published: (2025)
Kolmogorov-Arnold Attention: Is Learnable Attention Better For Vision Transformers?
by: Maity, Subhajit, et al.
Published: (2025)
by: Maity, Subhajit, et al.
Published: (2025)
JacNet: Learning Functions with Structured Jacobians
by: Lorraine, Jonathan, et al.
Published: (2024)
by: Lorraine, Jonathan, et al.
Published: (2024)
Revisiting Non-separable Binary Classification and its Applications in Anomaly Detection
by: Lau, Matthew, et al.
Published: (2023)
by: Lau, Matthew, et al.
Published: (2023)
RCUKF: Data-Driven Modeling Meets Bayesian Estimation
by: Anurag, Kumar, et al.
Published: (2025)
by: Anurag, Kumar, et al.
Published: (2025)
Neural Encoding for Image Recall: Human-Like Memory
by: Foussereau, Virgile, et al.
Published: (2024)
by: Foussereau, Virgile, et al.
Published: (2024)
Benchmarking Catastrophic Forgetting Mitigation Methods in Federated Time Series Forecasting
by: Hallak, Khaled, et al.
Published: (2025)
by: Hallak, Khaled, et al.
Published: (2025)
A Hybrid Inductive-Transductive Network for Traffic Flow Imputation on Unsampled Locations
by: Rahimiasl, Mohammadmahdi, et al.
Published: (2025)
by: Rahimiasl, Mohammadmahdi, et al.
Published: (2025)
Pulse-Driven Neural Architecture: Learnable Oscillatory Dynamics for Robust Continuous-Time Sequence Processing
by: Sharma, Paras
Published: (2026)
by: Sharma, Paras
Published: (2026)
PH-VAE: A Polynomial Hierarchical Variational Autoencoder Towards Disentangled Representation Learning
by: Chen, Xi, et al.
Published: (2025)
by: Chen, Xi, et al.
Published: (2025)
GraphNNK -- Graph Classification and Interpretability
by: Bolevic, Zeljko, et al.
Published: (2026)
by: Bolevic, Zeljko, et al.
Published: (2026)
Estimating the Event-Related Potential from Few EEG Trials
by: Nørskov, Anders Vestergaard, et al.
Published: (2025)
by: Nørskov, Anders Vestergaard, et al.
Published: (2025)
Temporal Functional Circuits: From Spline Plots to Faithful Explanations in KAN Forecasting
by: Mysore, Naveen
Published: (2026)
by: Mysore, Naveen
Published: (2026)
Similar Items
-
Explicit Dropout: Deterministic Regularization for Transformer Architectures
by: Agrawal, Vidhi, et al.
Published: (2026) -
Massive Redundancy in Gradient Transport Enables Sparse Online Learning
by: Merin, Aur Shalev
Published: (2026) -
Is Cambodia the World's Largest Cashew Producer?
by: Chaya, Veasna, et al.
Published: (2024) -
Closing the Theory-Practice Gap in Spiking Transformers via Effective Dimension
by: Guo, Dongxin, et al.
Published: (2026) -
H-Model: Dynamic Neural Architectures for Adaptive Processing
by: Hospodarchuk, Dmytro
Published: (2025)