Predictor-Corrector Enhanced Transformers with Exponential Moving Average Coefficient Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Bei, Zheng, Tong, Wang, Rui, Liu, Jiahao, Guo, Qingyan, Guo, Junliang, Tan, Xu, Xiao, Tong, Zhu, Jingbo, Wang, Jingang, Cai, Xunliang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EIT: Enhanced Interactive Transformer
by: Zheng, Tong, et al.
Published: (2022)
by: Zheng, Tong, et al.
Published: (2022)
IIET: Efficient Numerical Transformer via Implicit Iterative Euler Method
by: Liu, Xinyu, et al.
Published: (2025)
by: Liu, Xinyu, et al.
Published: (2025)
Speculative Decoding via Early-exiting for Faster LLM Inference with Thompson Sampling Control Mechanism
by: Liu, Jiahao, et al.
Published: (2024)
by: Liu, Jiahao, et al.
Published: (2024)
Mitigating Reversal Curse in Large Language Models via Semantic-aware Permutation Training
by: Guo, Qingyan, et al.
Published: (2024)
by: Guo, Qingyan, et al.
Published: (2024)
MTR-Suite: A Framework for Evaluating and Synthesizing Conversational Retrieval Benchmarks
by: Ruan, Junhao, et al.
Published: (2026)
by: Ruan, Junhao, et al.
Published: (2026)
EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers
by: Guo, Qingyan, et al.
Published: (2023)
by: Guo, Qingyan, et al.
Published: (2023)
SpanNorm: Reconciling Training Stability and Performance in Deep Transformers
by: Wang, Chao, et al.
Published: (2026)
by: Wang, Chao, et al.
Published: (2026)
LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance
by: Fan, Yuchun, et al.
Published: (2026)
by: Fan, Yuchun, et al.
Published: (2026)
ReMamba: Equip Mamba with Effective Long-Sequence Modeling
by: Yuan, Danlong, et al.
Published: (2024)
by: Yuan, Danlong, et al.
Published: (2024)
Forgetting Curve: A Reliable Method for Evaluating Memorization Capability for Long-context Models
by: Liu, Xinyu, et al.
Published: (2024)
by: Liu, Xinyu, et al.
Published: (2024)
What Makes Quantization for Large Language Models Hard? An Empirical Study from the Lens of Perturbation
by: Gong, Zhuocheng, et al.
Published: (2024)
by: Gong, Zhuocheng, et al.
Published: (2024)
Early Exit Is a Natural Capability in Transformer-based Models: An Empirical Study on Early Exit without Joint Optimization
by: Shan, Weiqiao, et al.
Published: (2024)
by: Shan, Weiqiao, et al.
Published: (2024)
Causal Autoregressive Diffusion Language Model
by: Ruan, Junhao, et al.
Published: (2026)
by: Ruan, Junhao, et al.
Published: (2026)
Monitoring the Coefficient of Variation Using a Synthetic Exponentially Weighted Moving Average Chart
by: Bryan Chek Hui Thien, et al.
Published: (2025)
by: Bryan Chek Hui Thien, et al.
Published: (2025)
PartialFormer: Modeling Part Instead of Whole for Machine Translation
by: Zheng, Tong, et al.
Published: (2023)
by: Zheng, Tong, et al.
Published: (2023)
Dynamic Fisher-weighted Model Merging via Bayesian Optimization
by: Lee, Sanwoo, et al.
Published: (2025)
by: Lee, Sanwoo, et al.
Published: (2025)
Libra: Assessing and Improving Reward Model by Learning to Think
by: Zhou, Meng, et al.
Published: (2025)
by: Zhou, Meng, et al.
Published: (2025)
Foundations of Large Language Models
by: Xiao, Tong, et al.
Published: (2025)
by: Xiao, Tong, et al.
Published: (2025)
Graph-Structured Speculative Decoding
by: Gong, Zhuocheng, et al.
Published: (2024)
by: Gong, Zhuocheng, et al.
Published: (2024)
Comparisons of Optimal Generally Weighted Moving Average and Exponentially Weighted Moving Average Charts
by: Steven E. Rigdon
Published: (2025)
by: Steven E. Rigdon
Published: (2025)
Break a Lag: Triple Exponential Moving Average for Enhanced Optimization
by: Peleg, Roi, et al.
Published: (2023)
by: Peleg, Roi, et al.
Published: (2023)
FedEMA: Federated Exponential Moving Averaging with Negative Entropy Regularizer in Autonomous Driving
by: Kou, Wei-Bin, et al.
Published: (2025)
by: Kou, Wei-Bin, et al.
Published: (2025)
Classifier-Free Guidance is a Predictor-Corrector
by: Bradley, Arwen, et al.
Published: (2024)
by: Bradley, Arwen, et al.
Published: (2024)
Parallel Decoding via Hidden Transfer for Lossless Large Language Model Acceleration
by: Wu, Pengfei, et al.
Published: (2024)
by: Wu, Pengfei, et al.
Published: (2024)
FIRP: Faster LLM inference via future intermediate representation prediction
by: Wu, Pengfei, et al.
Published: (2024)
by: Wu, Pengfei, et al.
Published: (2024)
TEAM-SimHRA: A Team-Based Simulation Framework for Human Reliability Analysis Using Multi-Agent Large Language Models
by: Xiao, Xingyu, et al.
Published: (2026)
by: Xiao, Xingyu, et al.
Published: (2026)
Exponential Moving Average of Weights in Deep Learning: Dynamics and Benefits
by: Morales-Brotons, Daniel, et al.
Published: (2024)
by: Morales-Brotons, Daniel, et al.
Published: (2024)
A Comprehensive Review of Human Error in Risk-Informed Decision Making: Integrating Human Reliability Assessment, Artificial Intelligence, and Human Performance Models
by: Xiao, Xingyu, et al.
Published: (2025)
by: Xiao, Xingyu, et al.
Published: (2025)
Earlier Tokens Contribute More: Learning Direct Preference Optimization From Temporal Decay Perspective
by: Shao, Ruichen, et al.
Published: (2025)
by: Shao, Ruichen, et al.
Published: (2025)
A Predictor‐Corrector Linearized High‐Order FEM for Nonlinear Time‐Fractional Parabolic Equations With Distributed Delay and Variable Coefficients
by: Ujwal Warbhe
Published: (2026)
by: Ujwal Warbhe
Published: (2026)
Rethinking the Sampling Criteria in Reinforcement Learning for LLM Reasoning: A Competence-Difficulty Alignment Perspective
by: Kong, Deyang, et al.
Published: (2025)
by: Kong, Deyang, et al.
Published: (2025)
DC-Solver: Improving Predictor-Corrector Diffusion Sampler via Dynamic Compensation
by: Zhao, Wenliang, et al.
Published: (2024)
by: Zhao, Wenliang, et al.
Published: (2024)
One Size Does Not Fit All: A Distribution-Aware Sparsification for More Precise Model Merging
by: Luo, Yingfeng, et al.
Published: (2025)
by: Luo, Yingfeng, et al.
Published: (2025)
The Prevalence of Misreporting and Misinterpreting Correlation Coefficients in Biomedical Literature
by: Xu, Jiayang, et al.
Published: (2025)
by: Xu, Jiayang, et al.
Published: (2025)
NeedleInATable: Exploring Long-Context Capability of Large Language Models towards Long-Structured Tables
by: Wang, Lanrui, et al.
Published: (2025)
by: Wang, Lanrui, et al.
Published: (2025)
DemaFormer: Damped Exponential Moving Average Transformer with Energy-Based Modeling for Temporal Language Grounding
by: Nguyen, Thong, et al.
Published: (2023)
by: Nguyen, Thong, et al.
Published: (2023)
Hybrid Explicit-Implicit Predictor-Corrector Exponential Time-Differencing Multistep Padé Schemes for Semilinear Parabolic Equations with Time-Delay
by: Dai, Haishen, et al.
Published: (2025)
by: Dai, Haishen, et al.
Published: (2025)
Higher-Dimensional Moving Averages and Submanifold Genericity
by: Cheng, Jiajun, et al.
Published: (2025)
by: Cheng, Jiajun, et al.
Published: (2025)
Ltri-LLM: Streaming Long Context Inference for LLMs with Training-Free Dynamic Triangular Attention Pattern
by: Tang, Hongyin, et al.
Published: (2024)
by: Tang, Hongyin, et al.
Published: (2024)
EMAG: Self-Rectifying Diffusion Sampling with Exponential Moving Average Guidance
by: Yadav, Ankit, et al.
Published: (2025)
by: Yadav, Ankit, et al.
Published: (2025)
Similar Items
-
EIT: Enhanced Interactive Transformer
by: Zheng, Tong, et al.
Published: (2022) -
IIET: Efficient Numerical Transformer via Implicit Iterative Euler Method
by: Liu, Xinyu, et al.
Published: (2025) -
Speculative Decoding via Early-exiting for Faster LLM Inference with Thompson Sampling Control Mechanism
by: Liu, Jiahao, et al.
Published: (2024) -
Mitigating Reversal Curse in Large Language Models via Semantic-aware Permutation Training
by: Guo, Qingyan, et al.
Published: (2024) -
MTR-Suite: A Framework for Evaluating and Synthesizing Conversational Retrieval Benchmarks
by: Ruan, Junhao, et al.
Published: (2026)