Growing Transformers: Modular Composition and Layer-wise Expansion on a Frozen Substrate
Fuente:
arXiv
Saved in:
| Main Author: | Bochkov, A. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Emergent Semantics Beyond Token Embeddings: Transformer LMs with Frozen Visual Unicode Representations
by: Bochkov, A.
Published: (2025)
by: Bochkov, A.
Published: (2025)
Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA
by: Bae, Sangmin, et al.
Published: (2024)
by: Bae, Sangmin, et al.
Published: (2024)
Bridging the Dimensional Chasm: Uncover Layer-wise Dimensional Reduction in Transformers through Token Correlation
by: Song, Zhuo-Yang, et al.
Published: (2025)
by: Song, Zhuo-Yang, et al.
Published: (2025)
On the Effect of Uncertainty on Layer-wise Inference Dynamics
by: Kim, Sunwoo, et al.
Published: (2025)
by: Kim, Sunwoo, et al.
Published: (2025)
Frozen Transformers in Language Models Are Effective Visual Encoder Layers
by: Pang, Ziqi, et al.
Published: (2023)
by: Pang, Ziqi, et al.
Published: (2023)
LEAP: Layer-wise Exit-Aware Pretraining for Efficient Transformer Inference
by: Kapadia, Shashank, et al.
Published: (2026)
by: Kapadia, Shashank, et al.
Published: (2026)
Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models
by: Mondorf, Philipp, et al.
Published: (2024)
by: Mondorf, Philipp, et al.
Published: (2024)
Iterative Layer-wise Distillation for Efficient Compression of Large Language Models
by: Kovalev, Grigory, et al.
Published: (2025)
by: Kovalev, Grigory, et al.
Published: (2025)
Internal Chain-of-Thought: Empirical Evidence for Layer-wise Subtask Scheduling in LLMs
by: Yang, Zhipeng, et al.
Published: (2025)
by: Yang, Zhipeng, et al.
Published: (2025)
Layer-wise Representation Dynamics: An Empirical Investigation Across Embedders and Base LLMs
by: Jiang, Jingzhou, et al.
Published: (2026)
by: Jiang, Jingzhou, et al.
Published: (2026)
A Comparative analysis of Layer-wise Representational Capacity in AR and Diffusion LLMs
by: Goel, Raghavv, et al.
Published: (2026)
by: Goel, Raghavv, et al.
Published: (2026)
Language Models Without a Trainable Input Embedding Table: Learning from Fixed Minimal Binary Token Codes
by: Bochkov, A.
Published: (2026)
by: Bochkov, A.
Published: (2026)
Rehearsal-Free Modular and Compositional Continual Learning for Language Models
by: Wang, Mingyang, et al.
Published: (2024)
by: Wang, Mingyang, et al.
Published: (2024)
MoFE: Mixture of Frozen Experts Architecture
by: Seo, Jean, et al.
Published: (2025)
by: Seo, Jean, et al.
Published: (2025)
SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal Budget
by: Wang, Zihao, et al.
Published: (2024)
by: Wang, Zihao, et al.
Published: (2024)
GRASS: Gradient-based Adaptive Layer-wise Importance Sampling for Memory-efficient Large Language Model Fine-tuning
by: Tian, Kaiyuan, et al.
Published: (2026)
by: Tian, Kaiyuan, et al.
Published: (2026)
Layer-wise Importance Matters: Less Memory for Better Performance in Parameter-efficient Fine-tuning of Large Language Models
by: Yao, Kai, et al.
Published: (2024)
by: Yao, Kai, et al.
Published: (2024)
Lost in State Space: Probing Frozen Mamba Representations
by: Wagh, Bhagyashree, et al.
Published: (2026)
by: Wagh, Bhagyashree, et al.
Published: (2026)
MaskPrune: Mask-based LLM Pruning for Layer-wise Uniform Structures
by: Qin, Jiayu, et al.
Published: (2025)
by: Qin, Jiayu, et al.
Published: (2025)
SpecBound: Adaptive Bounded Self-Speculation with Layer-wise Confidence Calibration
by: Wen, Zhuofan, et al.
Published: (2026)
by: Wen, Zhuofan, et al.
Published: (2026)
CAST: Compositional Analysis via Spectral Tracking for Understanding Transformer Layer Functions
by: Fu, Zihao, et al.
Published: (2025)
by: Fu, Zihao, et al.
Published: (2025)
Learning to Skip the Middle Layers of Transformers
by: Lawson, Tim, et al.
Published: (2025)
by: Lawson, Tim, et al.
Published: (2025)
Beyond Outliers: A Data-Free Layer-wise Mixed-Precision Quantization Approach Driven by Numerical and Structural Dual-Sensitivity
by: Zhang, Hengyuan, et al.
Published: (2026)
by: Zhang, Hengyuan, et al.
Published: (2026)
Landscape-Aware Growing: The Power of a Little LAG
by: Karp, Stefani, et al.
Published: (2024)
by: Karp, Stefani, et al.
Published: (2024)
LayerNorm Induces Recency Bias in Transformer Decoders
by: Kim, Junu, et al.
Published: (2025)
by: Kim, Junu, et al.
Published: (2025)
Provable Knowledge Acquisition and Extraction in One-Layer Transformers
by: Xu, Ruichen, et al.
Published: (2025)
by: Xu, Ruichen, et al.
Published: (2025)
ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings
by: Hao, Shibo, et al.
Published: (2023)
by: Hao, Shibo, et al.
Published: (2023)
METHOD: Modular Efficient Transformer for Health Outcome Discovery
by: Qian, Linglong, et al.
Published: (2025)
by: Qian, Linglong, et al.
Published: (2025)
Dynamic Encoder Size Based on Data-Driven Layer-wise Pruning for Speech Recognition
by: Xu, Jingjing, et al.
Published: (2024)
by: Xu, Jingjing, et al.
Published: (2024)
Out-of-Distribution Detection by Leveraging Between-Layer Transformation Smoothness
by: Jelenić, Fran, et al.
Published: (2023)
by: Jelenić, Fran, et al.
Published: (2023)
Mechanism and Emergence of Stacked Attention Heads in Multi-Layer Transformers
by: Musat, Tiberiu
Published: (2024)
by: Musat, Tiberiu
Published: (2024)
How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
by: Brandon, William, et al.
Published: (2024)
by: Brandon, William, et al.
Published: (2024)
Complexity Control Facilitates Reasoning-Based Compositional Generalization in Transformers
by: Zhang, Zhongwang, et al.
Published: (2025)
by: Zhang, Zhongwang, et al.
Published: (2025)
Borrowed Geometry: Cross-Distribution Head-Importance Fingerprints of Frozen Pretrained Gemma 4 31B
by: Bektursun, Abay
Published: (2026)
by: Bektursun, Abay
Published: (2026)
CharED: Character-wise Ensemble Decoding for Large Language Models
by: Gu, Kevin, et al.
Published: (2024)
by: Gu, Kevin, et al.
Published: (2024)
MeTHanol: Modularized Thinking Language Models with Intermediate Layer Thinking, Decoding and Bootstrapping Reasoning
by: Xi, Ningyuan, et al.
Published: (2024)
by: Xi, Ningyuan, et al.
Published: (2024)
TRIM: Token-wise Attention-Derived Saliency for Data-Efficient Instruction Tuning
by: Nagaraj, Manish, et al.
Published: (2025)
by: Nagaraj, Manish, et al.
Published: (2025)
CoT-UQ: Improving Response-wise Uncertainty Quantification in LLMs with Chain-of-Thought
by: Zhang, Boxuan, et al.
Published: (2025)
by: Zhang, Boxuan, et al.
Published: (2025)
ULMA: Unified Language Model Alignment with Human Demonstration and Point-wise Preference
by: Cai, Tianchi, et al.
Published: (2023)
by: Cai, Tianchi, et al.
Published: (2023)
Similar Items
-
Emergent Semantics Beyond Token Embeddings: Transformer LMs with Frozen Visual Unicode Representations
by: Bochkov, A.
Published: (2025) -
Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA
by: Bae, Sangmin, et al.
Published: (2024) -
Bridging the Dimensional Chasm: Uncover Layer-wise Dimensional Reduction in Transformers through Token Correlation
by: Song, Zhuo-Yang, et al.
Published: (2025) -
On the Effect of Uncertainty on Layer-wise Inference Dynamics
by: Kim, Sunwoo, et al.
Published: (2025) -
Frozen Transformers in Language Models Are Effective Visual Encoder Layers
by: Pang, Ziqi, et al.
Published: (2023)