You Only Need Your Transformer 25% of the Time: Meaning-First Execution for Eliminating Unnecessary Inference
Fuente:
arXiv
Saved in:
| Main Author: | Shamim, Ryan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
You Only Debias Once: Towards Flexible Accuracy-Fairness Trade-offs at Inference Time
by: Han, Xiaotian, et al.
Published: (2025)
by: Han, Xiaotian, et al.
Published: (2025)
Attention Once Is All You Need: Efficient Streaming Inference with Stateful Transformers
by: Norgren, Victor
Published: (2026)
by: Norgren, Victor
Published: (2026)
Adaptive Test-Time Training for Predicting Need for Invasive Mechanical Ventilation in Multi-Center Cohorts
by: Lu, Xiaolei, et al.
Published: (2025)
by: Lu, Xiaolei, et al.
Published: (2025)
The Residual Stream Is All You Need: On the Redundancy of the KV Cache in Transformer Inference
by: Qasim, Kaleem Ullah, et al.
Published: (2026)
by: Qasim, Kaleem Ullah, et al.
Published: (2026)
Attention Is All You Need But You Don't Need All Of It For Inference of Large Language Models
by: Tyukin, Georgy, et al.
Published: (2024)
by: Tyukin, Georgy, et al.
Published: (2024)
Probabilities Are All You Need: A Probability-Only Approach to Uncertainty Estimation in Large Language Models
by: Nguyen, Manh, et al.
Published: (2025)
by: Nguyen, Manh, et al.
Published: (2025)
Context is All You Need
by: Delanois, Jean Erik, et al.
Published: (2026)
by: Delanois, Jean Erik, et al.
Published: (2026)
You Only Accept Samples Once: Fast, Self-Correcting Stochastic Variational Inference
by: Dayta, Dominic B.
Published: (2024)
by: Dayta, Dominic B.
Published: (2024)
Catch-Up Distillation: You Only Need to Train Once for Accelerating Sampling
by: Shao, Shitong, et al.
Published: (2023)
by: Shao, Shitong, et al.
Published: (2023)
You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories
by: Wei, Zhepei, et al.
Published: (2026)
by: Wei, Zhepei, et al.
Published: (2026)
Occam's Razor is Only as Sharp as Your ELBO
by: Harvey, Ethan, et al.
Published: (2026)
by: Harvey, Ethan, et al.
Published: (2026)
Scalable Data Attribution via Forward-Only Test-Time Inference
by: Ma, Sibo, et al.
Published: (2025)
by: Ma, Sibo, et al.
Published: (2025)
Mamba or Transformer for Time Series Forecasting? Mixture of Universals (MoU) Is All You Need
by: Peng, Sijia, et al.
Published: (2024)
by: Peng, Sijia, et al.
Published: (2024)
SmartMem: Layout Transformation Elimination and Adaptation for Efficient DNN Execution on Mobile
by: Niu, Wei, et al.
Published: (2024)
by: Niu, Wei, et al.
Published: (2024)
You Only Train Once
by: Sakaridis, Christos
Published: (2025)
by: Sakaridis, Christos
Published: (2025)
Attention is All You Need Until You Need Retention
by: Yaslioglu, M. Murat
Published: (2025)
by: Yaslioglu, M. Murat
Published: (2025)
Harnesses for Inference-Time Alignment over Execution Trajectories
by: Wang, Boyuan, et al.
Published: (2026)
by: Wang, Boyuan, et al.
Published: (2026)
FUNU: Boosting Machine Unlearning Efficiency by Filtering Unnecessary Unlearning
by: Li, Zitong, et al.
Published: (2025)
by: Li, Zitong, et al.
Published: (2025)
Accuracy is Not All You Need
by: Dutta, Abhinav, et al.
Published: (2024)
by: Dutta, Abhinav, et al.
Published: (2024)
Attention Is Not All You Need: The Importance of Feedforward Networks in Transformer Models
by: Gerber, Isaac
Published: (2025)
by: Gerber, Isaac
Published: (2025)
Are High-Degree Representations Really Unnecessary in Equivariant Graph Neural Networks?
by: Cen, Jiacheng, et al.
Published: (2024)
by: Cen, Jiacheng, et al.
Published: (2024)
Embedding Is (Almost) All You Need: Retrieval-Augmented Inference for Generalizable Genomic Prediction Tasks
by: Datta, Nirjhor, et al.
Published: (2025)
by: Datta, Nirjhor, et al.
Published: (2025)
Triplet Interaction Improves Graph Transformers: Accurate Molecular Graph Learning with Triplet Graph Transformers
by: Hussain, Md Shamim, et al.
Published: (2024)
by: Hussain, Md Shamim, et al.
Published: (2024)
Masked Generative Transformer Is What You Need for Image Editing
by: Chow, Wei, et al.
Published: (2026)
by: Chow, Wei, et al.
Published: (2026)
Simple Feedfoward Neural Networks are Almost All You Need for Time Series Forecasting
by: Sun, Fan-Keng, et al.
Published: (2025)
by: Sun, Fan-Keng, et al.
Published: (2025)
Relevance Isn't All You Need: Scaling RAG Systems With Inference-Time Compute Via Multi-Criteria Reranking
by: LeVine, Will, et al.
Published: (2025)
by: LeVine, Will, et al.
Published: (2025)
Grokking as Structural Inference: Transformers Need Bayesian Lottery Tickets
by: Hidajat, Kai, et al.
Published: (2026)
by: Hidajat, Kai, et al.
Published: (2026)
Optimisation Is Not What You Need
by: Ibias, Alfredo
Published: (2025)
by: Ibias, Alfredo
Published: (2025)
Some Attention is All You Need for Retrieval
by: Michalak, Felix, et al.
Published: (2025)
by: Michalak, Felix, et al.
Published: (2025)
Half Search Space is All You Need
by: Rumiantsev, Pavel, et al.
Published: (2025)
by: Rumiantsev, Pavel, et al.
Published: (2025)
Top-$nσ$: Not All Logits Are You Need
by: Tang, Chenxia, et al.
Published: (2024)
by: Tang, Chenxia, et al.
Published: (2024)
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
by: Baroni, Luca, et al.
Published: (2025)
by: Baroni, Luca, et al.
Published: (2025)
Positional Knowledge is All You Need: Position-induced Transformer (PiT) for Operator Learning
by: Chen, Junfeng, et al.
Published: (2024)
by: Chen, Junfeng, et al.
Published: (2024)
Improving Prediction of Need for Mechanical Ventilation using Cross-Attention
by: Mohanty, Anwesh, et al.
Published: (2024)
by: Mohanty, Anwesh, et al.
Published: (2024)
You Only Spike Once: Improving Energy-Efficient Neuromorphic Inference to ANN-Level Accuracy
by: P, Srivatsa, et al.
Published: (2020)
by: P, Srivatsa, et al.
Published: (2020)
Deterministic Inference across Tensor Parallel Sizes That Eliminates Training-Inference Mismatch
by: Zhang, Ziyang, et al.
Published: (2025)
by: Zhang, Ziyang, et al.
Published: (2025)
Memory Is All You Need: An Overview of Compute-in-Memory Architectures for Accelerating Large Language Model Inference
by: Wolters, Christopher, et al.
Published: (2024)
by: Wolters, Christopher, et al.
Published: (2024)
Ask Your Distribution Shift if Pre-Training is Right for You
by: Cohen-Wang, Benjamin, et al.
Published: (2024)
by: Cohen-Wang, Benjamin, et al.
Published: (2024)
You Only Train Once: Differentiable Subset Selection for Omics Data
by: Chopard, Daphné, et al.
Published: (2025)
by: Chopard, Daphné, et al.
Published: (2025)
Attention-Only Transformers via Unrolled Subspace Denoising
by: Wang, Peng, et al.
Published: (2025)
by: Wang, Peng, et al.
Published: (2025)
Similar Items
-
You Only Debias Once: Towards Flexible Accuracy-Fairness Trade-offs at Inference Time
by: Han, Xiaotian, et al.
Published: (2025) -
Attention Once Is All You Need: Efficient Streaming Inference with Stateful Transformers
by: Norgren, Victor
Published: (2026) -
Adaptive Test-Time Training for Predicting Need for Invasive Mechanical Ventilation in Multi-Center Cohorts
by: Lu, Xiaolei, et al.
Published: (2025) -
The Residual Stream Is All You Need: On the Redundancy of the KV Cache in Transformer Inference
by: Qasim, Kaleem Ullah, et al.
Published: (2026) -
Attention Is All You Need But You Don't Need All Of It For Inference of Large Language Models
by: Tyukin, Georgy, et al.
Published: (2024)