Early Transformers: A study on Efficient Training of Transformer Models through Early-Bird Lottery Tickets
Fuente:
arXiv
Saved in:
| Main Author: | Cheekati, Shravan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TextAge: A Curated and Diverse Text Dataset for Age Classification
by: Cheekati, Shravan, et al.
Published: (2024)
by: Cheekati, Shravan, et al.
Published: (2024)
KS-Lottery: Finding Certified Lottery Tickets for Multilingual Language Models
by: Yuan, Fei, et al.
Published: (2024)
by: Yuan, Fei, et al.
Published: (2024)
Early-Bird GCNs: Graph-Network Co-Optimization Towards More Efficient GCN Training and Inference via Drawing Early-Bird Lottery Tickets
by: You, Haoran, et al.
Published: (2021)
by: You, Haoran, et al.
Published: (2021)
Transformers with Selective Access to Early Representations
by: Gunasekaran, Skye, et al.
Published: (2026)
by: Gunasekaran, Skye, et al.
Published: (2026)
Toy Combinatorial Interpretability Models Reveal Lottery Tickets in Early Feature Space
by: Bebchuk, Alon, et al.
Published: (2026)
by: Bebchuk, Alon, et al.
Published: (2026)
Drawing Early-Bird Tickets: Towards More Efficient Training of Deep Networks
by: You, Haoran, et al.
Published: (2019)
by: You, Haoran, et al.
Published: (2019)
A Survey of Lottery Ticket Hypothesis
by: Liu, Bohan, et al.
Published: (2024)
by: Liu, Bohan, et al.
Published: (2024)
LOTUS: Improving Transformer Efficiency with Sparsity Pruning and Data Lottery Tickets
by: Upadhyay, Ojasw
Published: (2024)
by: Upadhyay, Ojasw
Published: (2024)
Bayesian Lottery Ticket Hypothesis
by: Kuhn, Nicholas, et al.
Published: (2026)
by: Kuhn, Nicholas, et al.
Published: (2026)
Grokking as Structural Inference: Transformers Need Bayesian Lottery Tickets
by: Hidajat, Kai, et al.
Published: (2026)
by: Hidajat, Kai, et al.
Published: (2026)
Escaping the Mode Lottery: Multi-Response Training Improves Language Model Generalization
by: Amin, Hasan, et al.
Published: (2026)
by: Amin, Hasan, et al.
Published: (2026)
Gated Linear Attention Transformers with Hardware-Efficient Training
by: Yang, Songlin, et al.
Published: (2023)
by: Yang, Songlin, et al.
Published: (2023)
FLOP-Efficient Training: Early Stopping Based on Test-Time Compute Awareness
by: Amer, Hossam, et al.
Published: (2026)
by: Amer, Hossam, et al.
Published: (2026)
Winning the Lottery by Preserving Network Training Dynamics with Concrete Ticket Search
by: Arora, Tanay, et al.
Published: (2025)
by: Arora, Tanay, et al.
Published: (2025)
Pruning Unsafe Tickets: A Resource-Efficient Framework for Safer and More Robust LLMs
by: Si, Wai Man, et al.
Published: (2026)
by: Si, Wai Man, et al.
Published: (2026)
Towards Scalable Lottery Ticket Networks using Genetic Algorithms
by: Schönberger, Julian, et al.
Published: (2025)
by: Schönberger, Julian, et al.
Published: (2025)
A Multi-Level Framework for Accelerating Training Transformer Models
by: Zou, Longwei, et al.
Published: (2024)
by: Zou, Longwei, et al.
Published: (2024)
Teaching Transformers Causal Reasoning through Axiomatic Training
by: Vashishtha, Aniket, et al.
Published: (2024)
by: Vashishtha, Aniket, et al.
Published: (2024)
Development of Pre-Trained Transformer-based Models for the Nepali Language
by: Thapa, Prajwal, et al.
Published: (2024)
by: Thapa, Prajwal, et al.
Published: (2024)
HELIOS: Adaptive Model And Early-Exit Selection for Efficient LLM Inference Serving
by: Kumar, Avinash, et al.
Published: (2025)
by: Kumar, Avinash, et al.
Published: (2025)
Breaking Symmetry When Training Transformers
by: Zuo, Chunsheng, et al.
Published: (2024)
by: Zuo, Chunsheng, et al.
Published: (2024)
On the Sparsity of the Strong Lottery Ticket Hypothesis
by: Natale, Emanuele, et al.
Published: (2024)
by: Natale, Emanuele, et al.
Published: (2024)
Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models
by: Shravan, Rohan
Published: (2026)
by: Shravan, Rohan
Published: (2026)
Random Masking Finds Winning Tickets for Parameter Efficient Fine-tuning
by: Xu, Jing, et al.
Published: (2024)
by: Xu, Jing, et al.
Published: (2024)
What Is the Minimum Architecture for Prolepsis? Early Irrevocable Commitment Across Tasks in Small Transformers
by: Jacopin, Éric
Published: (2026)
by: Jacopin, Éric
Published: (2026)
Rethinking Attention Output Projection: Structured Hadamard Transforms for Efficient Transformers
by: Aggarwal, Shubham, et al.
Published: (2026)
by: Aggarwal, Shubham, et al.
Published: (2026)
Thinking into the Future: Latent Lookahead Training for Transformers
by: Noci, Lorenzo, et al.
Published: (2026)
by: Noci, Lorenzo, et al.
Published: (2026)
Late-to-Early Training: LET LLMs Learn Earlier, So Faster and Better
by: Zhao, Ji, et al.
Published: (2026)
by: Zhao, Ji, et al.
Published: (2026)
SuperTickets: Drawing Task-Agnostic Lottery Tickets from Supernets via Jointly Architecture Searching and Parameter Pruning
by: You, Haoran, et al.
Published: (2022)
by: You, Haoran, et al.
Published: (2022)
An Efficient Inference Framework for Early-exit Large Language Models
by: Miao, Ruijie, et al.
Published: (2024)
by: Miao, Ruijie, et al.
Published: (2024)
On Mesa-Optimization in Autoregressively Trained Transformers: Emergence and Capability
by: Zheng, Chenyu, et al.
Published: (2024)
by: Zheng, Chenyu, et al.
Published: (2024)
Insights into the Lottery Ticket Hypothesis and Iterative Magnitude Pruning
by: Saleem, Tausifa Jan, et al.
Published: (2024)
by: Saleem, Tausifa Jan, et al.
Published: (2024)
Dual Path Attribution: Efficient Attribution for SwiGLU-Transformers through Layer-Wise Target Propagation
by: Jantsch, Lasse Marten, et al.
Published: (2026)
by: Jantsch, Lasse Marten, et al.
Published: (2026)
Diffusion Language Models Generation Can Be Halted Early
by: Vaina, Sofia Maria Lo Cicero, et al.
Published: (2023)
by: Vaina, Sofia Maria Lo Cicero, et al.
Published: (2023)
Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel Reasoning
by: Bi, Jiaxi, et al.
Published: (2026)
by: Bi, Jiaxi, et al.
Published: (2026)
Ultra Memory-Efficient On-FPGA Training of Transformers via Tensor-Compressed Optimization
by: Tian, Jiayi, et al.
Published: (2025)
by: Tian, Jiayi, et al.
Published: (2025)
HybridNorm: Towards Stable and Efficient Transformer Training via Hybrid Normalization
by: Zhuo, Zhijian, et al.
Published: (2025)
by: Zhuo, Zhijian, et al.
Published: (2025)
Trainable Transformer in Transformer
by: Panigrahi, Abhishek, et al.
Published: (2023)
by: Panigrahi, Abhishek, et al.
Published: (2023)
The Open-World Lottery Ticket Hypothesis for OOD Intent Classification
by: Zhou, Yunhua, et al.
Published: (2022)
by: Zhou, Yunhua, et al.
Published: (2022)
A General and Efficient Training for Transformer via Token Expansion
by: Huang, Wenxuan, et al.
Published: (2024)
by: Huang, Wenxuan, et al.
Published: (2024)
Similar Items
-
TextAge: A Curated and Diverse Text Dataset for Age Classification
by: Cheekati, Shravan, et al.
Published: (2024) -
KS-Lottery: Finding Certified Lottery Tickets for Multilingual Language Models
by: Yuan, Fei, et al.
Published: (2024) -
Early-Bird GCNs: Graph-Network Co-Optimization Towards More Efficient GCN Training and Inference via Drawing Early-Bird Lottery Tickets
by: You, Haoran, et al.
Published: (2021) -
Transformers with Selective Access to Early Representations
by: Gunasekaran, Skye, et al.
Published: (2026) -
Toy Combinatorial Interpretability Models Reveal Lottery Tickets in Early Feature Space
by: Bebchuk, Alon, et al.
Published: (2026)