Allocation of Parameters in Transformers
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yu, Ruoxi, Jiang, Haotian, Cheng, Jingpu, Yu, Penghao, Li, Qianxiao, Li, Zhong |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
The Effect of Attention Head Count on Transformer Approximation
par: Yu, Penghao, et autres
Publié: (2025)
par: Yu, Penghao, et autres
Publié: (2025)
InfoFlow: A Framework for Multi-Layer Transformer Analysis
par: Yu, Penghao, et autres
Publié: (2026)
par: Yu, Penghao, et autres
Publié: (2026)
The Effect of Depth on the Expressivity of Deep Linear State-Space Models
par: Bao, Zeyu, et autres
Publié: (2025)
par: Bao, Zeyu, et autres
Publié: (2025)
Approximation Rate of the Transformer Architecture for Sequence Modeling
par: Jiang, Haotian, et autres
Publié: (2023)
par: Jiang, Haotian, et autres
Publié: (2023)
Machine Unlearning under Retain-Forget Entanglement
par: Cheng, Jingpu, et autres
Publié: (2026)
par: Cheng, Jingpu, et autres
Publié: (2026)
A unified framework for establishing the universal approximation of transformer-type architectures
par: Cheng, Jingpu, et autres
Publié: (2025)
par: Cheng, Jingpu, et autres
Publié: (2025)
Deep learning and the rate of approximation by flows
par: Cheng, Jingpu, et autres
Publié: (2026)
par: Cheng, Jingpu, et autres
Publié: (2026)
Numerical Investigation of Sequence Modeling Theory using Controllable Memory Functions
par: Jiang, Haotian, et autres
Publié: (2025)
par: Jiang, Haotian, et autres
Publié: (2025)
Inverse Approximation Theory for Nonlinear Recurrent Neural Networks
par: Wang, Shida, et autres
Publié: (2023)
par: Wang, Shida, et autres
Publié: (2023)
Learning task-specific predictive models for scientific computing
par: Yin, Jianyuan, et autres
Publié: (2025)
par: Yin, Jianyuan, et autres
Publié: (2025)
From Generalization Analysis to Optimization Designs for State Space Models
par: Liu, Fusheng, et autres
Publié: (2024)
par: Liu, Fusheng, et autres
Publié: (2024)
Autocorrelation Matters: Understanding the Role of Initialization Schemes for State Space Models
par: Liu, Fusheng, et autres
Publié: (2024)
par: Liu, Fusheng, et autres
Publié: (2024)
Accelerating Legacy Numerical Solvers by Non-intrusive Gradient-based Meta-solving
par: Arisaka, Sohei, et autres
Publié: (2024)
par: Arisaka, Sohei, et autres
Publié: (2024)
Unifying back-propagation and forward-forward algorithms through model predictive control
par: Ren, Lianhai, et autres
Publié: (2024)
par: Ren, Lianhai, et autres
Publié: (2024)
DynGMA: a robust approach for learning stochastic differential equations from data
par: Zhu, Aiqing, et autres
Publié: (2024)
par: Zhu, Aiqing, et autres
Publié: (2024)
Learning Macroscopic Dynamics from Partial Microscopic Observations
par: Chen, Mengyi, et autres
Publié: (2024)
par: Chen, Mengyi, et autres
Publié: (2024)
StableSSM: Alleviating the Curse of Memory in State-space Models through Stable Reparameterization
par: Wang, Shida, et autres
Publié: (2023)
par: Wang, Shida, et autres
Publié: (2023)
Learning Permutation-invariant Macroscopic Dynamics
par: Han, Zhichao, et autres
Publié: (2026)
par: Han, Zhichao, et autres
Publié: (2026)
How Transformers Get Rich: Approximation and Dynamics Analysis
par: Wang, Mingze, et autres
Publié: (2024)
par: Wang, Mingze, et autres
Publié: (2024)
Continuity-Preserving Convolutional Autoencoders for Learning Continuous Latent Dynamical Models from Images
par: Zhu, Aiqing, et autres
Publié: (2025)
par: Zhu, Aiqing, et autres
Publié: (2025)
MSign: An Optimizer Preventing Training Instability in Large Language Models via Stable Rank Restoration
par: Ren, Lianhai, et autres
Publié: (2026)
par: Ren, Lianhai, et autres
Publié: (2026)
Closed-Form Concept Erasure via Double Projections
par: Zhang, Chi, et autres
Publié: (2026)
par: Zhang, Chi, et autres
Publié: (2026)
Mitigating distribution shift in machine learning-augmented hybrid simulation
par: Zhao, Jiaxi, et autres
Publié: (2024)
par: Zhao, Jiaxi, et autres
Publié: (2024)
On the Curse of Memory in Recurrent Neural Networks: Approximation and Optimization Analysis
par: Li, Zhong, et autres
Publié: (2020)
par: Li, Zhong, et autres
Publié: (2020)
TUNI: A Textual Unimodal Detector for Identity Inference in CLIP Models
par: Li, Songze, et autres
Publié: (2024)
par: Li, Songze, et autres
Publié: (2024)
Scalable learning of macroscopic stochastic dynamics
par: Chen, Mengyi, et autres
Publié: (2025)
par: Chen, Mengyi, et autres
Publié: (2025)
Identifiable learning of dissipative dynamics
par: Zhu, Aiqing, et autres
Publié: (2025)
par: Zhu, Aiqing, et autres
Publié: (2025)
Hidden Representation Clustering with Multi-Task Representation Learning towards Robust Online Budget Allocation
par: Wang, Xiaohan, et autres
Publié: (2025)
par: Wang, Xiaohan, et autres
Publié: (2025)
Terminally constrained flow-based generative models from an optimal control perspective
par: Gao, Weiguo, et autres
Publié: (2026)
par: Gao, Weiguo, et autres
Publié: (2026)
Parameter Optimization with Conscious Allocation (POCA)
par: Inman, Joshua, et autres
Publié: (2023)
par: Inman, Joshua, et autres
Publié: (2023)
DMRL: Data- and Model-aware Reward Learning for Data Extraction
par: Wang, Zhiqiang, et autres
Publié: (2025)
par: Wang, Zhiqiang, et autres
Publié: (2025)
Shuttle Between the Instructions and the Parameters of Large Language Models
par: Sun, Wangtao, et autres
Publié: (2025)
par: Sun, Wangtao, et autres
Publié: (2025)
From PEFT to DEFT: Parameter Efficient Finetuning for Reducing Activation Density in Transformers
par: Runwal, Bharat, et autres
Publié: (2024)
par: Runwal, Bharat, et autres
Publié: (2024)
Tiny-Critic RAG: Empowering Agentic Fallback with Parameter-Efficient Small Language Models
par: Wu, Yichao, et autres
Publié: (2026)
par: Wu, Yichao, et autres
Publié: (2026)
LoRA-SP: Streamlined Partial Parameter Adaptation for Resource-Efficient Fine-Tuning of Large Language Models
par: Wu, Yichao, et autres
Publié: (2024)
par: Wu, Yichao, et autres
Publié: (2024)
Data Efficient Any Transformer-to-Mamba Distillation via Attention Bridge
par: Wang, Penghao, et autres
Publié: (2025)
par: Wang, Penghao, et autres
Publié: (2025)
Embed and Emulate: Contrastive representations for simulation-based inference
par: Jiang, Ruoxi, et autres
Publié: (2024)
par: Jiang, Ruoxi, et autres
Publié: (2024)
Parameter-Efficient Fine-Tuning via Circular Convolution
par: Chen, Aochuan, et autres
Publié: (2024)
par: Chen, Aochuan, et autres
Publié: (2024)
SGNO: Spectral Generator Neural Operators for Stable Long Horizon PDE Rollouts
par: Li, Jiayi, et autres
Publié: (2026)
par: Li, Jiayi, et autres
Publié: (2026)
Heterogeneity-Informed Meta-Parameter Learning for Spatiotemporal Time Series Forecasting
par: Dong, Zheng, et autres
Publié: (2024)
par: Dong, Zheng, et autres
Publié: (2024)
Documents similaires
-
The Effect of Attention Head Count on Transformer Approximation
par: Yu, Penghao, et autres
Publié: (2025) -
InfoFlow: A Framework for Multi-Layer Transformer Analysis
par: Yu, Penghao, et autres
Publié: (2026) -
The Effect of Depth on the Expressivity of Deep Linear State-Space Models
par: Bao, Zeyu, et autres
Publié: (2025) -
Approximation Rate of the Transformer Architecture for Sequence Modeling
par: Jiang, Haotian, et autres
Publié: (2023) -
Machine Unlearning under Retain-Forget Entanglement
par: Cheng, Jingpu, et autres
Publié: (2026)