Revisiting Transformer Layer Parameterization Through Causal Energy Minimization
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Jin, Couturier, Camille, Rühle, Victor, Rajmohan, Saravan, Hensman, James |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TurboAttention: Efficient Attention Approximation For High Throughputs LLMs
by: Kang, Hao, et al.
Published: (2024)
by: Kang, Hao, et al.
Published: (2024)
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models
by: Couturier, Camille, et al.
Published: (2025)
by: Couturier, Camille, et al.
Published: (2025)
LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation
by: Han, Dongge, et al.
Published: (2025)
by: Han, Dongge, et al.
Published: (2025)
Hybrid-RACA: Hybrid Retrieval-Augmented Composition Assistance for Real-time Text Prediction
by: Xia, Menglin, et al.
Published: (2023)
by: Xia, Menglin, et al.
Published: (2023)
Enhancing Reasoning Capabilities of Small Language Models with Blueprints and Prompt Template Search
by: Han, Dongge, et al.
Published: (2025)
by: Han, Dongge, et al.
Published: (2025)
Exploring How LLMs Capture and Represent Domain-Specific Knowledge
by: Garcia, Mirian Hipolito, et al.
Published: (2025)
by: Garcia, Mirian Hipolito, et al.
Published: (2025)
DiSK: A Diffusion Model for Structured Knowledge
by: Kitouni, Ouail, et al.
Published: (2023)
by: Kitouni, Ouail, et al.
Published: (2023)
OptRot: Mitigating Weight Outliers via Data-Free Rotations for Post-Training Quantization
by: Gadhikar, Advait, et al.
Published: (2025)
by: Gadhikar, Advait, et al.
Published: (2025)
AMPO: Active Multi-Preference Optimization for Self-play Preference Selection
by: Gupta, Taneesh, et al.
Published: (2025)
by: Gupta, Taneesh, et al.
Published: (2025)
REFA: Reference Free Alignment for multi-preference optimization
by: Gupta, Taneesh, et al.
Published: (2024)
by: Gupta, Taneesh, et al.
Published: (2024)
Lean Attention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers
by: Sanovar, Rya, et al.
Published: (2024)
by: Sanovar, Rya, et al.
Published: (2024)
LLM Reasoning as Trajectories: Step-Specific Representation Geometry and Correctness Signals
by: Sun, Lihao, et al.
Published: (2026)
by: Sun, Lihao, et al.
Published: (2026)
Multi-Preference Optimization: Generalizing DPO via Set-Level Contrasts
by: Gupta, Taneesh, et al.
Published: (2024)
by: Gupta, Taneesh, et al.
Published: (2024)
Everything of Thoughts: Defying the Law of Penrose Triangle for Thought Generation
by: Ding, Ruomeng, et al.
Published: (2023)
by: Ding, Ruomeng, et al.
Published: (2023)
AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task Generation
by: Hu, Mengkang, et al.
Published: (2024)
by: Hu, Mengkang, et al.
Published: (2024)
Pretrain Value, Not Reward: Decoupled Value Policy Optimization
by: Huang, Chenghua, et al.
Published: (2025)
by: Huang, Chenghua, et al.
Published: (2025)
Revisiting Few-Shot Learning from a Causal Perspective
by: Lin, Guoliang, et al.
Published: (2022)
by: Lin, Guoliang, et al.
Published: (2022)
Peri-LN: Revisiting Normalization Layer in the Transformer Architecture
by: Kim, Jeonghoon, et al.
Published: (2025)
by: Kim, Jeonghoon, et al.
Published: (2025)
Simulating Environments with Reasoning Models for Agent Training
by: Li, Yuetai, et al.
Published: (2025)
by: Li, Yuetai, et al.
Published: (2025)
The Minimal Search Space for Conditional Causal Bandits
by: Simoes, Francisco N. F. Q., et al.
Published: (2025)
by: Simoes, Francisco N. F. Q., et al.
Published: (2025)
Importance Sampling for Nonlinear Models
by: Rajmohan, Prakash Palanivelu, et al.
Published: (2025)
by: Rajmohan, Prakash Palanivelu, et al.
Published: (2025)
ARO: A New Lens On Matrix Optimization For Large Models
by: Gong, Wenbo, et al.
Published: (2026)
by: Gong, Wenbo, et al.
Published: (2026)
Revisiting Sharpness-Aware Minimization: A More Faithful and Effective Implementation
by: Chen, Jianlong, et al.
Published: (2026)
by: Chen, Jianlong, et al.
Published: (2026)
Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs
by: Wang, Qibin, et al.
Published: (2025)
by: Wang, Qibin, et al.
Published: (2025)
COIN: Chance-Constrained Imitation Learning for Uncertainty-aware Adaptive Resource Oversubscription Policy
by: Wang, Lu, et al.
Published: (2024)
by: Wang, Lu, et al.
Published: (2024)
Causal Layering via Conditional Entropy
by: Feigenbaum, Itai, et al.
Published: (2024)
by: Feigenbaum, Itai, et al.
Published: (2024)
Generalizing Knowledge Graph Embedding with Universal Orthogonal Parameterization
by: Li, Rui, et al.
Published: (2024)
by: Li, Rui, et al.
Published: (2024)
Transformer Is Inherently a Causal Learner
by: Wang, Xinyue, et al.
Published: (2026)
by: Wang, Xinyue, et al.
Published: (2026)
Generalizable Reasoning through Compositional Energy Minimization
by: Oarga, Alexandru, et al.
Published: (2025)
by: Oarga, Alexandru, et al.
Published: (2025)
A Recipe for Causal Graph Regression: Confounding Effects Revisited
by: Yin, Yujia, et al.
Published: (2025)
by: Yin, Yujia, et al.
Published: (2025)
Parameterized Projected Bellman Operator
by: Vincent, Théo, et al.
Published: (2023)
by: Vincent, Théo, et al.
Published: (2023)
Estimating Causal Effects from Learned Causal Networks
by: Raichev, Anna, et al.
Published: (2024)
by: Raichev, Anna, et al.
Published: (2024)
Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity
by: Xia, Menglin, et al.
Published: (2026)
by: Xia, Menglin, et al.
Published: (2026)
Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity
by: Ren, Ruifeng, et al.
Published: (2025)
by: Ren, Ruifeng, et al.
Published: (2025)
Stabilizing Sharpness-aware Minimization Through A Simple Renormalization Strategy
by: Tan, Chengli, et al.
Published: (2024)
by: Tan, Chengli, et al.
Published: (2024)
Diffusion Controller: Framework, Algorithms and Parameterization
by: Yang, Tong, et al.
Published: (2026)
by: Yang, Tong, et al.
Published: (2026)
Ensuring Fair LLM Serving Amid Diverse Applications
by: Khan, Redwan Ibne Seraj, et al.
Published: (2024)
by: Khan, Redwan Ibne Seraj, et al.
Published: (2024)
Probing the Embedding Space of Transformers via Minimal Token Perturbations
by: Conti, Eddie, et al.
Published: (2025)
by: Conti, Eddie, et al.
Published: (2025)
Transformer Normalisation Layers and the Independence of Semantic Subspaces
by: Menary, Stephen, et al.
Published: (2024)
by: Menary, Stephen, et al.
Published: (2024)
Dynamic Layer Tying for Parameter-Efficient Transformers
by: Hay, Tamir David, et al.
Published: (2024)
by: Hay, Tamir David, et al.
Published: (2024)
Similar Items
-
TurboAttention: Efficient Attention Approximation For High Throughputs LLMs
by: Kang, Hao, et al.
Published: (2024) -
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models
by: Couturier, Camille, et al.
Published: (2025) -
LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation
by: Han, Dongge, et al.
Published: (2025) -
Hybrid-RACA: Hybrid Retrieval-Augmented Composition Assistance for Real-time Text Prediction
by: Xia, Menglin, et al.
Published: (2023) -
Enhancing Reasoning Capabilities of Small Language Models with Blueprints and Prompt Template Search
by: Han, Dongge, et al.
Published: (2025)