Dynamic Layer Tying for Parameter-Efficient Transformers
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Hay, Tamir David, Wolf, Lior |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Accelerating Error Correction Code Transformers
par: Levy, Matan, et autres
Publié: (2024)
par: Levy, Matan, et autres
Publié: (2024)
Factor Graph Optimization of Error-Correcting Codes for Belief Propagation Decoding
par: Choukroun, Yoni, et autres
Publié: (2024)
par: Choukroun, Yoni, et autres
Publié: (2024)
Module-Aware Parameter-Efficient Machine Unlearning on Transformers
par: Bao, Wenjie, et autres
Publié: (2025)
par: Bao, Wenjie, et autres
Publié: (2025)
Outlier-Efficient Hopfield Layers for Large Transformer-Based Models
par: Hu, Jerry Yao-Chieh, et autres
Publié: (2024)
par: Hu, Jerry Yao-Chieh, et autres
Publié: (2024)
Understanding and Guiding Layer Placement in Parameter-Efficient Fine-Tuning of Large Language Models
par: Xu, Yichen, et autres
Publié: (2026)
par: Xu, Yichen, et autres
Publié: (2026)
Diversity of Transformer Layers: One Aspect of Parameter Scaling Laws
par: Kamigaito, Hidetaka, et autres
Publié: (2025)
par: Kamigaito, Hidetaka, et autres
Publié: (2025)
TWISTED-RL: Hierarchical Skilled Agents for Knot-Tying without Human Demonstrations
par: Freund, Guy, et autres
Publié: (2026)
par: Freund, Guy, et autres
Publié: (2026)
On The Statistical Representation Properties Of The Perturb-Softmax And The Perturb-Argmax Probability Distributions
par: Indelman, Hedda Cohen, et autres
Publié: (2024)
par: Indelman, Hedda Cohen, et autres
Publié: (2024)
Echo-LoRA: Parameter-Efficient Fine-Tuning via Cross-Layer Representation Injection
par: Peng, Yihang, et autres
Publié: (2026)
par: Peng, Yihang, et autres
Publié: (2026)
Evaluating Neural Networks for Early Maritime Threat Detection
par: Tella, Dhanush, et autres
Publié: (2024)
par: Tella, Dhanush, et autres
Publié: (2024)
Finding Clustering Algorithms in the Transformer Architecture
par: Clarkson, Kenneth L., et autres
Publié: (2025)
par: Clarkson, Kenneth L., et autres
Publié: (2025)
Backward Lens: Projecting Language Model Gradients into the Vocabulary Space
par: Katz, Shahar, et autres
Publié: (2024)
par: Katz, Shahar, et autres
Publié: (2024)
Partial Parameter Updates for Efficient Distributed Training
par: Filippova, Anastasiia, et autres
Publié: (2025)
par: Filippova, Anastasiia, et autres
Publié: (2025)
Communication Dynamics Neural Networks: FFT-Diagonalized Layers for Improved Hessian Conditioning at Reduced Parameter Count
par: Pan, Lurong
Publié: (2026)
par: Pan, Lurong
Publié: (2026)
PETAH: Parameter Efficient Task Adaptation for Hybrid Transformers in a resource-limited Context
par: Augustin, Maximilian, et autres
Publié: (2024)
par: Augustin, Maximilian, et autres
Publié: (2024)
Parameter Efficient Mamba Tuning via Projector-targeted Diagonal-centric Linear Transformation
par: Ham, Seokil, et autres
Publié: (2024)
par: Ham, Seokil, et autres
Publié: (2024)
Parameter-Efficient Fine-Tuning with Discrete Fourier Transform
par: Gao, Ziqi, et autres
Publié: (2024)
par: Gao, Ziqi, et autres
Publié: (2024)
Set-based Neural Network Encoding Without Weight Tying
par: Andreis, Bruno, et autres
Publié: (2023)
par: Andreis, Bruno, et autres
Publié: (2023)
Parameter-Efficient Fine-Tuning for HAR: Integrating LoRA and QLoRA into Transformer Models
par: Seregina, Irina, et autres
Publié: (2025)
par: Seregina, Irina, et autres
Publié: (2025)
Simulus: Combining Improvements in Sample-Efficient World Model Agents
par: Cohen, Lior, et autres
Publié: (2025)
par: Cohen, Lior, et autres
Publié: (2025)
Non-Uniform Parameter-Wise Model Merging
par: Camacho, Albert Manuel Orozco, et autres
Publié: (2024)
par: Camacho, Albert Manuel Orozco, et autres
Publié: (2024)
Is there Value in Reinforcement Learning?
par: Fox, Lior, et autres
Publié: (2025)
par: Fox, Lior, et autres
Publié: (2025)
Hybrid Dynamic Pruning: A Pathway to Efficient Transformer Inference
par: Jaradat, Ghadeer, et autres
Publié: (2024)
par: Jaradat, Ghadeer, et autres
Publié: (2024)
SGFormer: Single-Layer Graph Transformers with Approximation-Free Linear Complexity
par: Wu, Qitian, et autres
Publié: (2024)
par: Wu, Qitian, et autres
Publié: (2024)
Anomaly Detection with Variance Stabilized Density Estimation
par: Rozner, Amit, et autres
Publié: (2023)
par: Rozner, Amit, et autres
Publié: (2023)
Dynamic Mixture of Experts: An Auto-Tuning Approach for Efficient Transformer Models
par: Guo, Yongxin, et autres
Publié: (2024)
par: Guo, Yongxin, et autres
Publié: (2024)
PAINET: A Principled Efficient Transformer for 3D Dynamics Modeling
par: Yang, Kai, et autres
Publié: (2025)
par: Yang, Kai, et autres
Publié: (2025)
Enhancing GNNs with Architecture-Agnostic Graph Transformations: A Systematic Analysis
par: Li, Zhifei, et autres
Publié: (2024)
par: Li, Zhifei, et autres
Publié: (2024)
Solo Connection: A Parameter Efficient Fine-Tuning Technique for Transformers
par: Pathak, Harsh Nilesh, et autres
Publié: (2025)
par: Pathak, Harsh Nilesh, et autres
Publié: (2025)
Unified Parameter-Efficient Unlearning for LLMs
par: Ding, Chenlu, et autres
Publié: (2024)
par: Ding, Chenlu, et autres
Publié: (2024)
Transformer Normalisation Layers and the Independence of Semantic Subspaces
par: Menary, Stephen, et autres
Publié: (2024)
par: Menary, Stephen, et autres
Publié: (2024)
Layer Specialization Underlying Compositional Reasoning in Transformers
par: Liu, Jing
Publié: (2025)
par: Liu, Jing
Publié: (2025)
DeciMamba: Exploring the Length Extrapolation Potential of Mamba
par: Ben-Kish, Assaf, et autres
Publié: (2024)
par: Ben-Kish, Assaf, et autres
Publié: (2024)
Overflow Prevention Enhances Long-Context Recurrent LLMs
par: Ben-Kish, Assaf, et autres
Publié: (2025)
par: Ben-Kish, Assaf, et autres
Publié: (2025)
LEAP: Layer-wise Exit-Aware Pretraining for Efficient Transformer Inference
par: Kapadia, Shashank, et autres
Publié: (2026)
par: Kapadia, Shashank, et autres
Publié: (2026)
Deep Active Speech Cancellation with Mamba-Masking Network
par: Mishaly, Yehuda, et autres
Publié: (2025)
par: Mishaly, Yehuda, et autres
Publié: (2025)
Overclocking LLM Reasoning: Monitoring and Controlling Thinking Path Lengths in LLMs
par: Eisenstadt, Roy, et autres
Publié: (2025)
par: Eisenstadt, Roy, et autres
Publié: (2025)
Universal Approximation Theorem for a Single-Layer Transformer
par: Gumaan, Esmail
Publié: (2025)
par: Gumaan, Esmail
Publié: (2025)
WARP: Guaranteed Inner-Layer Repair of NLP Transformers
par: Hsu, Hsin-Ling, et autres
Publié: (2026)
par: Hsu, Hsin-Ling, et autres
Publié: (2026)
Efficiently Aligning Draft Models via Parameter- and Data-Efficient Adaptation
par: Lin, Luxi, et autres
Publié: (2026)
par: Lin, Luxi, et autres
Publié: (2026)
Documents similaires
-
Accelerating Error Correction Code Transformers
par: Levy, Matan, et autres
Publié: (2024) -
Factor Graph Optimization of Error-Correcting Codes for Belief Propagation Decoding
par: Choukroun, Yoni, et autres
Publié: (2024) -
Module-Aware Parameter-Efficient Machine Unlearning on Transformers
par: Bao, Wenjie, et autres
Publié: (2025) -
Outlier-Efficient Hopfield Layers for Large Transformer-Based Models
par: Hu, Jerry Yao-Chieh, et autres
Publié: (2024) -
Understanding and Guiding Layer Placement in Parameter-Efficient Fine-Tuning of Large Language Models
par: Xu, Yichen, et autres
Publié: (2026)