Transformers as Intrinsic Optimizers: Forward Inference through the Energy Principle
Fuente:
arXiv
Guardado en:
| Autores principales: | Ren, Ruifeng, Ouyang, Sheng, Tang, Huayi, Liu, Yong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity
por: Ren, Ruifeng, et al.
Publicado: (2025)
por: Ren, Ruifeng, et al.
Publicado: (2025)
Towards Understanding How Transformers Learn In-context Through a Representation Learning Lens
por: Ren, Ruifeng, et al.
Publicado: (2023)
por: Ren, Ruifeng, et al.
Publicado: (2023)
PAC-Bayesian Generalization Bounds for Graph Convolutional Networks on Inductive Node Classification
por: Tang, Huayi, et al.
Publicado: (2025)
por: Tang, Huayi, et al.
Publicado: (2025)
Information-Theoretic Generalization Bounds for Transductive Learning and its Applications
por: Tang, Huayi, et al.
Publicado: (2023)
por: Tang, Huayi, et al.
Publicado: (2023)
Perfect Alignment May be Poisonous to Graph Contrastive Learning
por: Liu, Jingyu, et al.
Publicado: (2023)
por: Liu, Jingyu, et al.
Publicado: (2023)
Sparsity is Combinatorial Depth: Quantifying MoE Expressivity via Tropical Geometry
por: Su, Ye, et al.
Publicado: (2026)
por: Su, Ye, et al.
Publicado: (2026)
Kaczmarz Linear Attention
por: Zou, Jiaxuan, et al.
Publicado: (2026)
por: Zou, Jiaxuan, et al.
Publicado: (2026)
Towards Auto-Regressive Next-Token Prediction: In-Context Learning Emerges from Generalization
por: Gong, Zixuan, et al.
Publicado: (2025)
por: Gong, Zixuan, et al.
Publicado: (2025)
Understanding Model Ensemble in Transferable Adversarial Attack
por: Yao, Wei, et al.
Publicado: (2024)
por: Yao, Wei, et al.
Publicado: (2024)
Exploring the Limitations of Mamba in COPY and CoT Reasoning
por: Ren, Ruifeng, et al.
Publicado: (2024)
por: Ren, Ruifeng, et al.
Publicado: (2024)
LightFF: Lightweight Inference for Forward-Forward Algorithm
por: Aminifar, Amin, et al.
Publicado: (2024)
por: Aminifar, Amin, et al.
Publicado: (2024)
Mono-Forward: Revisiting Forward-Forward through Objective-Locality Decomposition
por: Gong, James, et al.
Publicado: (2025)
por: Gong, James, et al.
Publicado: (2025)
Effective Frontiers: A Unification of Neural Scaling Laws
por: Zou, Jiaxuan, et al.
Publicado: (2026)
por: Zou, Jiaxuan, et al.
Publicado: (2026)
On Weak-to-Strong Generalization and f-Divergence
por: Yao, Wei, et al.
Publicado: (2025)
por: Yao, Wei, et al.
Publicado: (2025)
FFCL: Forward-Forward Net with Cortical Loops, Training and Inference on Edge Without Backpropagation
por: Karkehabadi, Ali, et al.
Publicado: (2024)
por: Karkehabadi, Ali, et al.
Publicado: (2024)
Energy-Efficient Transformer Inference: Optimization Strategies for Time Series Classification
por: Kermani, Arshia, et al.
Publicado: (2025)
por: Kermani, Arshia, et al.
Publicado: (2025)
Prefix Grouper: Efficient GRPO Training through Shared-Prefix Forward
por: Liu, Zikang, et al.
Publicado: (2025)
por: Liu, Zikang, et al.
Publicado: (2025)
Compositional Generalization from Learned Skills via CoT Training: A Theoretical and Structural Analysis for Reasoning
por: Yao, Xinhao, et al.
Publicado: (2025)
por: Yao, Xinhao, et al.
Publicado: (2025)
Forward-Forward Autoencoder Architectures for Energy-Efficient Wireless Communications
por: Seifert, Daniel, et al.
Publicado: (2025)
por: Seifert, Daniel, et al.
Publicado: (2025)
Refining Latent Representations: A Generative SSL Approach for Heterogeneous Graph Learning
por: Hu, Yulan, et al.
Publicado: (2023)
por: Hu, Yulan, et al.
Publicado: (2023)
VIGraph: Generative Self-supervised Learning for Class-Imbalanced Node Classification
por: Hu, Yulan, et al.
Publicado: (2023)
por: Hu, Yulan, et al.
Publicado: (2023)
Energy-Efficient Vision Transformer Inference for Edge-AI Deployment
por: Amanzhol, Nursultan, et al.
Publicado: (2025)
por: Amanzhol, Nursultan, et al.
Publicado: (2025)
Shrinking the Giant : Quasi-Weightless Transformers for Low Energy Inference
por: Nag, Shashank, et al.
Publicado: (2024)
por: Nag, Shashank, et al.
Publicado: (2024)
Rate-Distortion Optimization for Transformer Inference
por: de Andrade, Anderson, et al.
Publicado: (2026)
por: de Andrade, Anderson, et al.
Publicado: (2026)
Efficient and Principled Scientific Discovery through Bayesian Optimization: A Tutorial
por: Yu, Zhongwei, et al.
Publicado: (2026)
por: Yu, Zhongwei, et al.
Publicado: (2026)
Selective Attention: Enhancing Transformer through Principled Context Control
por: Zhang, Xuechen, et al.
Publicado: (2024)
por: Zhang, Xuechen, et al.
Publicado: (2024)
Intrinsic Rewards for Exploration without Harm from Observational Noise: A Simulation Study Based on the Free Energy Principle
por: Tinker, Theodore Jerome, et al.
Publicado: (2024)
por: Tinker, Theodore Jerome, et al.
Publicado: (2024)
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization
por: Pang, Bowen, et al.
Publicado: (2025)
por: Pang, Bowen, et al.
Publicado: (2025)
Transformer-like Inference from Optimal Control
por: Kudre, Aditya, et al.
Publicado: (2026)
por: Kudre, Aditya, et al.
Publicado: (2026)
EnergyLens: Predictive Energy-Aware Exploration for Multi-GPU LLM Inference Optimization
por: Song, Zhiye, et al.
Publicado: (2026)
por: Song, Zhiye, et al.
Publicado: (2026)
Scalable Data Attribution via Forward-Only Test-Time Inference
por: Ma, Sibo, et al.
Publicado: (2025)
por: Ma, Sibo, et al.
Publicado: (2025)
Addressing Class Imbalance with Probabilistic Graphical Models and Variational Inference
por: Lou, Yujia, et al.
Publicado: (2025)
por: Lou, Yujia, et al.
Publicado: (2025)
Stochastic Forward-Forward Learning through Representational Dimensionality Compression
por: Zhu, Zhichao, et al.
Publicado: (2025)
por: Zhu, Zhichao, et al.
Publicado: (2025)
Contrastive Forward-Forward: A Training Algorithm of Vision Transformer
por: Aghagolzadeh, Hossein, et al.
Publicado: (2025)
por: Aghagolzadeh, Hossein, et al.
Publicado: (2025)
GUNDAM: Aligning Large Language Models with Graph Understanding
por: Ouyang, Sheng, et al.
Publicado: (2024)
por: Ouyang, Sheng, et al.
Publicado: (2024)
POSEIDON: Physics-Optimized Seismic Energy Inference and Detection Operating Network
por: Kriuk, Boris, et al.
Publicado: (2026)
por: Kriuk, Boris, et al.
Publicado: (2026)
R$^2$Energy: A Large-Scale Benchmark for Robust Renewable Energy Forecasting under Diverse and Extreme Conditions
por: Sheng, Zhi, et al.
Publicado: (2026)
por: Sheng, Zhi, et al.
Publicado: (2026)
Energy-Guided Diffusion Sampling for Offline-to-Online Reinforcement Learning
por: Liu, Xu-Hui, et al.
Publicado: (2024)
por: Liu, Xu-Hui, et al.
Publicado: (2024)
Energy-Efficient Supervised Learning with a Binary Stochastic Forward-Forward Algorithm
por: Jaiswal, Risi, et al.
Publicado: (2025)
por: Jaiswal, Risi, et al.
Publicado: (2025)
Second-Order Forward-Mode Automatic Differentiation for Optimization
por: Cobb, Adam D., et al.
Publicado: (2024)
por: Cobb, Adam D., et al.
Publicado: (2024)
Ejemplares similares
-
Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity
por: Ren, Ruifeng, et al.
Publicado: (2025) -
Towards Understanding How Transformers Learn In-context Through a Representation Learning Lens
por: Ren, Ruifeng, et al.
Publicado: (2023) -
PAC-Bayesian Generalization Bounds for Graph Convolutional Networks on Inductive Node Classification
por: Tang, Huayi, et al.
Publicado: (2025) -
Information-Theoretic Generalization Bounds for Transductive Learning and its Applications
por: Tang, Huayi, et al.
Publicado: (2023) -
Perfect Alignment May be Poisonous to Graph Contrastive Learning
por: Liu, Jingyu, et al.
Publicado: (2023)