Do Efficient Transformers Really Save Computation?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Kai, Ackermann, Jan, He, Zhenyu, Feng, Guhao, Zhang, Bohang, Feng, Yunzhen, Ye, Qiwei, He, Di, Wang, Liwei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length Extrapolation
von: He, Zhenyu, et al.
Veröffentlicht: (2024)
von: He, Zhenyu, et al.
Veröffentlicht: (2024)
Theoretical Benefit and Limitation of Diffusion Language Model
von: Feng, Guhao, et al.
Veröffentlicht: (2025)
von: Feng, Guhao, et al.
Veröffentlicht: (2025)
Efficient Reasoning for Large Reasoning Language Models via Certainty-Guided Reflection Suppression
von: Huang, Jiameng, et al.
Veröffentlicht: (2025)
von: Huang, Jiameng, et al.
Veröffentlicht: (2025)
How Numerical Precision Affects Arithmetical Reasoning Capabilities of LLMs
von: Feng, Guhao, et al.
Veröffentlicht: (2024)
von: Feng, Guhao, et al.
Veröffentlicht: (2024)
In-Place Test-Time Training
von: Feng, Guhao, et al.
Veröffentlicht: (2026)
von: Feng, Guhao, et al.
Veröffentlicht: (2026)
DPO Meets PPO: Reinforced Token Optimization for RLHF
von: Zhong, Han, et al.
Veröffentlicht: (2024)
von: Zhong, Han, et al.
Veröffentlicht: (2024)
A Tale of Tails: Model Collapse as a Change of Scaling Laws
von: Dohmatob, Elvis, et al.
Veröffentlicht: (2024)
von: Dohmatob, Elvis, et al.
Veröffentlicht: (2024)
Rethinking Model-based, Policy-based, and Value-based Reinforcement Learning via the Lens of Representation Complexity
von: Feng, Guhao, et al.
Veröffentlicht: (2023)
von: Feng, Guhao, et al.
Veröffentlicht: (2023)
From Sequence to Structure: Uncovering Substructure Reasoning in Transformers
von: Dai, Xinnan, et al.
Veröffentlicht: (2025)
von: Dai, Xinnan, et al.
Veröffentlicht: (2025)
Efficient Agent Training for Computer Use
von: He, Yanheng, et al.
Veröffentlicht: (2025)
von: He, Yanheng, et al.
Veröffentlicht: (2025)
Enhancing Auto-regressive Chain-of-Thought through Loop-Aligned Reasoning
von: Yu, Qifan, et al.
Veröffentlicht: (2025)
von: Yu, Qifan, et al.
Veröffentlicht: (2025)
Adaptive Computation Pruning for the Forgetting Transformer
von: Lin, Zhixuan, et al.
Veröffentlicht: (2025)
von: Lin, Zhixuan, et al.
Veröffentlicht: (2025)
Let the Code LLM Edit Itself When You Edit the Code
von: He, Zhenyu, et al.
Veröffentlicht: (2024)
von: He, Zhenyu, et al.
Veröffentlicht: (2024)
Model Collapse Demystified: The Case of Regression
von: Dohmatob, Elvis, et al.
Veröffentlicht: (2024)
von: Dohmatob, Elvis, et al.
Veröffentlicht: (2024)
REST: Retrieval-Based Speculative Decoding
von: He, Zhenyu, et al.
Veröffentlicht: (2023)
von: He, Zhenyu, et al.
Veröffentlicht: (2023)
Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models?
von: Wang, Yanbo, et al.
Veröffentlicht: (2025)
von: Wang, Yanbo, et al.
Veröffentlicht: (2025)
SABER: Switchable and Balanced Training for Efficient LLM Reasoning
von: Zhao, Kai, et al.
Veröffentlicht: (2025)
von: Zhao, Kai, et al.
Veröffentlicht: (2025)
Understanding Dynamic Compute Allocation in Recurrent Transformers
von: Moosa, Ibraheem Muhammad, et al.
Veröffentlicht: (2026)
von: Moosa, Ibraheem Muhammad, et al.
Veröffentlicht: (2026)
Unleashing Diverse Thinking Modes in LLMs through Multi-Agent Collaboration
von: He, Zhixuan, et al.
Veröffentlicht: (2025)
von: He, Zhixuan, et al.
Veröffentlicht: (2025)
Parameter Efficient Instruction Tuning: An Empirical Study
von: He, Pengfei
Veröffentlicht: (2024)
von: He, Pengfei
Veröffentlicht: (2024)
Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention
von: Guo, Zhenyu, et al.
Veröffentlicht: (2025)
von: Guo, Zhenyu, et al.
Veröffentlicht: (2025)
Does LLM Alignment Really Need Diversity? An Empirical Study of Adapting RLVR Methods for Moral Reasoning
von: Zhang, Zhaowei, et al.
Veröffentlicht: (2026)
von: Zhang, Zhaowei, et al.
Veröffentlicht: (2026)
Multimodal large language model for wheat breeding: a new exploration of smart breeding
von: Yang, Guofeng, et al.
Veröffentlicht: (2024)
von: Yang, Guofeng, et al.
Veröffentlicht: (2024)
Mixture of Universal Experts: Scaling Virtual Width via Depth-Width Transformation
von: Chen, Yilong, et al.
Veröffentlicht: (2026)
von: Chen, Yilong, et al.
Veröffentlicht: (2026)
Projected Compression: Trainable Projection for Efficient Transformer Compression
von: Stefaniak, Maciej, et al.
Veröffentlicht: (2025)
von: Stefaniak, Maciej, et al.
Veröffentlicht: (2025)
A Survey of Quantized Graph Representation Learning: Connecting Graph Structures with Large Language Models
von: Lin, Qika, et al.
Veröffentlicht: (2025)
von: Lin, Qika, et al.
Veröffentlicht: (2025)
Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback
von: Ackermann, Johannes, et al.
Veröffentlicht: (2025)
von: Ackermann, Johannes, et al.
Veröffentlicht: (2025)
Athena: Efficient Block-Wise Post-Training Quantization for Large Language Models Using Second-Order Matrix Derivative Information
von: Wang, Yanshu, et al.
Veröffentlicht: (2024)
von: Wang, Yanshu, et al.
Veröffentlicht: (2024)
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
von: Yang, Zhihe, et al.
Veröffentlicht: (2025)
von: Yang, Zhihe, et al.
Veröffentlicht: (2025)
JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention
von: Tian, Yuandong, et al.
Veröffentlicht: (2023)
von: Tian, Yuandong, et al.
Veröffentlicht: (2023)
Benchmarking the Energy Savings with Speculative Decoding Strategies
von: Dutta, Rohit, et al.
Veröffentlicht: (2026)
von: Dutta, Rohit, et al.
Veröffentlicht: (2026)
Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models
von: Vendrell, Victor Conchello, et al.
Veröffentlicht: (2026)
von: Vendrell, Victor Conchello, et al.
Veröffentlicht: (2026)
Self-Improvement as Coherence Optimization: A Theoretical Account
von: Qiu, Tianyi, et al.
Veröffentlicht: (2026)
von: Qiu, Tianyi, et al.
Veröffentlicht: (2026)
A Survey of Test-Time Compute: From Intuitive Inference to Deliberate Reasoning
von: Ji, Yixin, et al.
Veröffentlicht: (2025)
von: Ji, Yixin, et al.
Veröffentlicht: (2025)
LeanK: Learnable K Cache Channel Pruning for Efficient Decoding
von: Zhang, Yike, et al.
Veröffentlicht: (2025)
von: Zhang, Yike, et al.
Veröffentlicht: (2025)
Beyond Model Collapse: Scaling Up with Synthesized Data Requires Verification
von: Feng, Yunzhen, et al.
Veröffentlicht: (2024)
von: Feng, Yunzhen, et al.
Veröffentlicht: (2024)
Understanding Subword Compositionality of Large Language Models
von: Peng, Qiwei, et al.
Veröffentlicht: (2025)
von: Peng, Qiwei, et al.
Veröffentlicht: (2025)
Transformers Struggle to Learn to Search
von: Saparov, Abulhair, et al.
Veröffentlicht: (2024)
von: Saparov, Abulhair, et al.
Veröffentlicht: (2024)
Position: The Turing-Completeness of Autoregressive Transformers Relies Heavily on Context Management
von: Cui, Guanyu, et al.
Veröffentlicht: (2026)
von: Cui, Guanyu, et al.
Veröffentlicht: (2026)
What Matters in Transformers? Not All Attention is Needed
von: He, Shwai, et al.
Veröffentlicht: (2024)
von: He, Shwai, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length Extrapolation
von: He, Zhenyu, et al.
Veröffentlicht: (2024) -
Theoretical Benefit and Limitation of Diffusion Language Model
von: Feng, Guhao, et al.
Veröffentlicht: (2025) -
Efficient Reasoning for Large Reasoning Language Models via Certainty-Guided Reflection Suppression
von: Huang, Jiameng, et al.
Veröffentlicht: (2025) -
How Numerical Precision Affects Arithmetical Reasoning Capabilities of LLMs
von: Feng, Guhao, et al.
Veröffentlicht: (2024) -
In-Place Test-Time Training
von: Feng, Guhao, et al.
Veröffentlicht: (2026)