Transformers as Intrinsic Optimizers: Forward Inference through the Energy Principle
Fuente:
arXiv
Salvato in:
| Autori principali: | Ren, Ruifeng, Ouyang, Sheng, Tang, Huayi, Liu, Yong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity
di: Ren, Ruifeng, et al.
Pubblicazione: (2025)
di: Ren, Ruifeng, et al.
Pubblicazione: (2025)
Towards Understanding How Transformers Learn In-context Through a Representation Learning Lens
di: Ren, Ruifeng, et al.
Pubblicazione: (2023)
di: Ren, Ruifeng, et al.
Pubblicazione: (2023)
PAC-Bayesian Generalization Bounds for Graph Convolutional Networks on Inductive Node Classification
di: Tang, Huayi, et al.
Pubblicazione: (2025)
di: Tang, Huayi, et al.
Pubblicazione: (2025)
Information-Theoretic Generalization Bounds for Transductive Learning and its Applications
di: Tang, Huayi, et al.
Pubblicazione: (2023)
di: Tang, Huayi, et al.
Pubblicazione: (2023)
Perfect Alignment May be Poisonous to Graph Contrastive Learning
di: Liu, Jingyu, et al.
Pubblicazione: (2023)
di: Liu, Jingyu, et al.
Pubblicazione: (2023)
Sparsity is Combinatorial Depth: Quantifying MoE Expressivity via Tropical Geometry
di: Su, Ye, et al.
Pubblicazione: (2026)
di: Su, Ye, et al.
Pubblicazione: (2026)
Kaczmarz Linear Attention
di: Zou, Jiaxuan, et al.
Pubblicazione: (2026)
di: Zou, Jiaxuan, et al.
Pubblicazione: (2026)
Towards Auto-Regressive Next-Token Prediction: In-Context Learning Emerges from Generalization
di: Gong, Zixuan, et al.
Pubblicazione: (2025)
di: Gong, Zixuan, et al.
Pubblicazione: (2025)
Understanding Model Ensemble in Transferable Adversarial Attack
di: Yao, Wei, et al.
Pubblicazione: (2024)
di: Yao, Wei, et al.
Pubblicazione: (2024)
Exploring the Limitations of Mamba in COPY and CoT Reasoning
di: Ren, Ruifeng, et al.
Pubblicazione: (2024)
di: Ren, Ruifeng, et al.
Pubblicazione: (2024)
LightFF: Lightweight Inference for Forward-Forward Algorithm
di: Aminifar, Amin, et al.
Pubblicazione: (2024)
di: Aminifar, Amin, et al.
Pubblicazione: (2024)
Mono-Forward: Revisiting Forward-Forward through Objective-Locality Decomposition
di: Gong, James, et al.
Pubblicazione: (2025)
di: Gong, James, et al.
Pubblicazione: (2025)
Effective Frontiers: A Unification of Neural Scaling Laws
di: Zou, Jiaxuan, et al.
Pubblicazione: (2026)
di: Zou, Jiaxuan, et al.
Pubblicazione: (2026)
On Weak-to-Strong Generalization and f-Divergence
di: Yao, Wei, et al.
Pubblicazione: (2025)
di: Yao, Wei, et al.
Pubblicazione: (2025)
FFCL: Forward-Forward Net with Cortical Loops, Training and Inference on Edge Without Backpropagation
di: Karkehabadi, Ali, et al.
Pubblicazione: (2024)
di: Karkehabadi, Ali, et al.
Pubblicazione: (2024)
Energy-Efficient Transformer Inference: Optimization Strategies for Time Series Classification
di: Kermani, Arshia, et al.
Pubblicazione: (2025)
di: Kermani, Arshia, et al.
Pubblicazione: (2025)
Prefix Grouper: Efficient GRPO Training through Shared-Prefix Forward
di: Liu, Zikang, et al.
Pubblicazione: (2025)
di: Liu, Zikang, et al.
Pubblicazione: (2025)
Compositional Generalization from Learned Skills via CoT Training: A Theoretical and Structural Analysis for Reasoning
di: Yao, Xinhao, et al.
Pubblicazione: (2025)
di: Yao, Xinhao, et al.
Pubblicazione: (2025)
Forward-Forward Autoencoder Architectures for Energy-Efficient Wireless Communications
di: Seifert, Daniel, et al.
Pubblicazione: (2025)
di: Seifert, Daniel, et al.
Pubblicazione: (2025)
Refining Latent Representations: A Generative SSL Approach for Heterogeneous Graph Learning
di: Hu, Yulan, et al.
Pubblicazione: (2023)
di: Hu, Yulan, et al.
Pubblicazione: (2023)
VIGraph: Generative Self-supervised Learning for Class-Imbalanced Node Classification
di: Hu, Yulan, et al.
Pubblicazione: (2023)
di: Hu, Yulan, et al.
Pubblicazione: (2023)
Energy-Efficient Vision Transformer Inference for Edge-AI Deployment
di: Amanzhol, Nursultan, et al.
Pubblicazione: (2025)
di: Amanzhol, Nursultan, et al.
Pubblicazione: (2025)
Shrinking the Giant : Quasi-Weightless Transformers for Low Energy Inference
di: Nag, Shashank, et al.
Pubblicazione: (2024)
di: Nag, Shashank, et al.
Pubblicazione: (2024)
Rate-Distortion Optimization for Transformer Inference
di: de Andrade, Anderson, et al.
Pubblicazione: (2026)
di: de Andrade, Anderson, et al.
Pubblicazione: (2026)
Efficient and Principled Scientific Discovery through Bayesian Optimization: A Tutorial
di: Yu, Zhongwei, et al.
Pubblicazione: (2026)
di: Yu, Zhongwei, et al.
Pubblicazione: (2026)
Selective Attention: Enhancing Transformer through Principled Context Control
di: Zhang, Xuechen, et al.
Pubblicazione: (2024)
di: Zhang, Xuechen, et al.
Pubblicazione: (2024)
Intrinsic Rewards for Exploration without Harm from Observational Noise: A Simulation Study Based on the Free Energy Principle
di: Tinker, Theodore Jerome, et al.
Pubblicazione: (2024)
di: Tinker, Theodore Jerome, et al.
Pubblicazione: (2024)
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization
di: Pang, Bowen, et al.
Pubblicazione: (2025)
di: Pang, Bowen, et al.
Pubblicazione: (2025)
Transformer-like Inference from Optimal Control
di: Kudre, Aditya, et al.
Pubblicazione: (2026)
di: Kudre, Aditya, et al.
Pubblicazione: (2026)
EnergyLens: Predictive Energy-Aware Exploration for Multi-GPU LLM Inference Optimization
di: Song, Zhiye, et al.
Pubblicazione: (2026)
di: Song, Zhiye, et al.
Pubblicazione: (2026)
Scalable Data Attribution via Forward-Only Test-Time Inference
di: Ma, Sibo, et al.
Pubblicazione: (2025)
di: Ma, Sibo, et al.
Pubblicazione: (2025)
Addressing Class Imbalance with Probabilistic Graphical Models and Variational Inference
di: Lou, Yujia, et al.
Pubblicazione: (2025)
di: Lou, Yujia, et al.
Pubblicazione: (2025)
Stochastic Forward-Forward Learning through Representational Dimensionality Compression
di: Zhu, Zhichao, et al.
Pubblicazione: (2025)
di: Zhu, Zhichao, et al.
Pubblicazione: (2025)
Contrastive Forward-Forward: A Training Algorithm of Vision Transformer
di: Aghagolzadeh, Hossein, et al.
Pubblicazione: (2025)
di: Aghagolzadeh, Hossein, et al.
Pubblicazione: (2025)
GUNDAM: Aligning Large Language Models with Graph Understanding
di: Ouyang, Sheng, et al.
Pubblicazione: (2024)
di: Ouyang, Sheng, et al.
Pubblicazione: (2024)
POSEIDON: Physics-Optimized Seismic Energy Inference and Detection Operating Network
di: Kriuk, Boris, et al.
Pubblicazione: (2026)
di: Kriuk, Boris, et al.
Pubblicazione: (2026)
R$^2$Energy: A Large-Scale Benchmark for Robust Renewable Energy Forecasting under Diverse and Extreme Conditions
di: Sheng, Zhi, et al.
Pubblicazione: (2026)
di: Sheng, Zhi, et al.
Pubblicazione: (2026)
Energy-Guided Diffusion Sampling for Offline-to-Online Reinforcement Learning
di: Liu, Xu-Hui, et al.
Pubblicazione: (2024)
di: Liu, Xu-Hui, et al.
Pubblicazione: (2024)
Energy-Efficient Supervised Learning with a Binary Stochastic Forward-Forward Algorithm
di: Jaiswal, Risi, et al.
Pubblicazione: (2025)
di: Jaiswal, Risi, et al.
Pubblicazione: (2025)
Second-Order Forward-Mode Automatic Differentiation for Optimization
di: Cobb, Adam D., et al.
Pubblicazione: (2024)
di: Cobb, Adam D., et al.
Pubblicazione: (2024)
Documenti analoghi
-
Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity
di: Ren, Ruifeng, et al.
Pubblicazione: (2025) -
Towards Understanding How Transformers Learn In-context Through a Representation Learning Lens
di: Ren, Ruifeng, et al.
Pubblicazione: (2023) -
PAC-Bayesian Generalization Bounds for Graph Convolutional Networks on Inductive Node Classification
di: Tang, Huayi, et al.
Pubblicazione: (2025) -
Information-Theoretic Generalization Bounds for Transductive Learning and its Applications
di: Tang, Huayi, et al.
Pubblicazione: (2023) -
Perfect Alignment May be Poisonous to Graph Contrastive Learning
di: Liu, Jingyu, et al.
Pubblicazione: (2023)