The Strong Lottery Ticket Hypothesis for Multi-Head Attention Mechanisms
Fuente:
arXiv
Saved in:
| Main Authors: | Otsuka, Hikari, Chijiwa, Daiki, Okoshi, Yasuyuki, Fujiki, Daichi, Takeuchi, Susumu, Motomura, Masato |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Partially Frozen Random Networks Contain Compact Strong Lottery Tickets
by: Otsuka, Hikari, et al.
Published: (2024)
by: Otsuka, Hikari, et al.
Published: (2024)
AQPIM: Breaking the PIM Capacity Wall for LLMs with In-Memory Activation Quantization
by: Matsushima, Kosuke, et al.
Published: (2026)
by: Matsushima, Kosuke, et al.
Published: (2026)
On the Sparsity of the Strong Lottery Ticket Hypothesis
by: Natale, Emanuele, et al.
Published: (2024)
by: Natale, Emanuele, et al.
Published: (2024)
Portable Reward Tuning: Towards Reusable Fine-Tuning across Different Pretrained Models
by: Chijiwa, Daiki, et al.
Published: (2025)
by: Chijiwa, Daiki, et al.
Published: (2025)
Investigating the Lottery Ticket Hypothesis for Variational Quantum Circuits
by: Kölle, Michael, et al.
Published: (2025)
by: Kölle, Michael, et al.
Published: (2025)
Context Memorization for Efficient Long Context Generation
by: Okoshi, Yasuyuki, et al.
Published: (2026)
by: Okoshi, Yasuyuki, et al.
Published: (2026)
Rethinking Optimal Verification Granularity for Compute-Efficient Test-Time Scaling
by: Chen, Hao Mark, et al.
Published: (2025)
by: Chen, Hao Mark, et al.
Published: (2025)
Binary Quadratic Quantization: Beyond First-Order Quantization for Real-Valued Matrix Compression
by: Kuroki, Kyo, et al.
Published: (2025)
by: Kuroki, Kyo, et al.
Published: (2025)
KS-Lottery: Finding Certified Lottery Tickets for Multilingual Language Models
by: Yuan, Fei, et al.
Published: (2024)
by: Yuan, Fei, et al.
Published: (2024)
Grokking as Structural Inference: Transformers Need Bayesian Lottery Tickets
by: Hidajat, Kai, et al.
Published: (2026)
by: Hidajat, Kai, et al.
Published: (2026)
Uncovering a Winning Lottery Ticket with Continuously Relaxed Bernoulli Gates
by: Tsayag, Itamar, et al.
Published: (2026)
by: Tsayag, Itamar, et al.
Published: (2026)
Lossless Vocabulary Reduction for Auto-Regressive Language Models
by: Chijiwa, Daiki, et al.
Published: (2025)
by: Chijiwa, Daiki, et al.
Published: (2025)
Bayesian Lottery Ticket Hypothesis
by: Kuhn, Nicholas, et al.
Published: (2026)
by: Kuhn, Nicholas, et al.
Published: (2026)
Quantization vs Pruning: Insights from the Strong Lottery Ticket Hypothesis
by: Kumar, Aakash, et al.
Published: (2025)
by: Kumar, Aakash, et al.
Published: (2025)
Rationale-Enhanced Decoding for Multi-modal Chain-of-Thought
by: Yamaguchi, Shin'ya, et al.
Published: (2025)
by: Yamaguchi, Shin'ya, et al.
Published: (2025)
A Neural Scaling Law from Lottery Ticket Ensembling
by: Liu, Ziming, et al.
Published: (2023)
by: Liu, Ziming, et al.
Published: (2023)
A Survey of Lottery Ticket Hypothesis
by: Liu, Bohan, et al.
Published: (2024)
by: Liu, Bohan, et al.
Published: (2024)
The Multiple Ticket Hypothesis: Random Sparse Subnetworks Suffice for RLVR
by: Adewuyi, Israel, et al.
Published: (2026)
by: Adewuyi, Israel, et al.
Published: (2026)
LOTUS: Improving Transformer Efficiency with Sparsity Pruning and Data Lottery Tickets
by: Upadhyay, Ojasw
Published: (2024)
by: Upadhyay, Ojasw
Published: (2024)
Toward Data Efficient Model Merging between Different Datasets without Performance Degradation
by: Yamada, Masanori, et al.
Published: (2023)
by: Yamada, Masanori, et al.
Published: (2023)
Insights into the Lottery Ticket Hypothesis and Iterative Magnitude Pruning
by: Saleem, Tausifa Jan, et al.
Published: (2024)
by: Saleem, Tausifa Jan, et al.
Published: (2024)
Transfer Learning with Pre-trained Conditional Generative Models
by: Yamaguchi, Shin'ya, et al.
Published: (2022)
by: Yamaguchi, Shin'ya, et al.
Published: (2022)
Winning the Lottery by Preserving Network Training Dynamics with Concrete Ticket Search
by: Arora, Tanay, et al.
Published: (2025)
by: Arora, Tanay, et al.
Published: (2025)
Zero-shot Concept Bottleneck Models
by: Yamaguchi, Shin'ya, et al.
Published: (2025)
by: Yamaguchi, Shin'ya, et al.
Published: (2025)
Adaptive Random Feature Regularization on Fine-tuning Deep Neural Networks
by: Yamaguchi, Shin'ya, et al.
Published: (2024)
by: Yamaguchi, Shin'ya, et al.
Published: (2024)
Parallel In-context Learning for Large Vision Language Models
by: Yamaguchi, Shin'ya, et al.
Published: (2026)
by: Yamaguchi, Shin'ya, et al.
Published: (2026)
Uncovering Critical Features for Deepfake Detection through the Lottery Ticket Hypothesis
by: Amin, Lisan Al, et al.
Published: (2025)
by: Amin, Lisan Al, et al.
Published: (2025)
AdaBlock-dLLM: Semantic-Aware Diffusion LLM Inference via Adaptive Block Size
by: Lu, Guanxi, et al.
Published: (2025)
by: Lu, Guanxi, et al.
Published: (2025)
XicorAttention: Time Series Transformer Using Attention with Nonlinear Correlation
by: Kimura, Daichi, et al.
Published: (2025)
by: Kimura, Daichi, et al.
Published: (2025)
Geometric Analysis of Token Selection in Multi-Head Attention
by: Mudarisov, Timur, et al.
Published: (2026)
by: Mudarisov, Timur, et al.
Published: (2026)
Superiority of Multi-Head Attention in In-Context Linear Regression
by: Cui, Yingqian, et al.
Published: (2024)
by: Cui, Yingqian, et al.
Published: (2024)
Post-pre-training for Modality Alignment in Vision-Language Foundation Models
by: Yamaguchi, Shin'ya, et al.
Published: (2025)
by: Yamaguchi, Shin'ya, et al.
Published: (2025)
MoH: Multi-Head Attention as Mixture-of-Head Attention
by: Jin, Peng, et al.
Published: (2024)
by: Jin, Peng, et al.
Published: (2024)
The Lottery LLM Hypothesis, Rethinking What Abilities Should LLM Compression Preserve?
by: Tang, Zhenheng, et al.
Published: (2025)
by: Tang, Zhenheng, et al.
Published: (2025)
TransMLA: Multi-Head Latent Attention Is All You Need
by: Meng, Fanxu, et al.
Published: (2025)
by: Meng, Fanxu, et al.
Published: (2025)
Feature Lottery? A Bifurcation Theory of Concept Emergence
by: Yang, Fuming
Published: (2026)
by: Yang, Fuming
Published: (2026)
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
by: Li, Cheng, et al.
Published: (2025)
by: Li, Cheng, et al.
Published: (2025)
Playing the Lottery With Concave Regularizers for Sparse Trainable Neural Networks
by: Fracastoro, Giulia, et al.
Published: (2025)
by: Fracastoro, Giulia, et al.
Published: (2025)
Jackpot! Alignment as a Maximal Lottery
by: Maura-Rivero, Roberto-Rafael, et al.
Published: (2025)
by: Maura-Rivero, Roberto-Rafael, et al.
Published: (2025)
CARE: Covariance-Aware and Rank-Enhanced Decomposition for Enabling Multi-Head Latent Attention
by: Zhou, Zhongzhu, et al.
Published: (2026)
by: Zhou, Zhongzhu, et al.
Published: (2026)
Similar Items
-
Partially Frozen Random Networks Contain Compact Strong Lottery Tickets
by: Otsuka, Hikari, et al.
Published: (2024) -
AQPIM: Breaking the PIM Capacity Wall for LLMs with In-Memory Activation Quantization
by: Matsushima, Kosuke, et al.
Published: (2026) -
On the Sparsity of the Strong Lottery Ticket Hypothesis
by: Natale, Emanuele, et al.
Published: (2024) -
Portable Reward Tuning: Towards Reusable Fine-Tuning across Different Pretrained Models
by: Chijiwa, Daiki, et al.
Published: (2025) -
Investigating the Lottery Ticket Hypothesis for Variational Quantum Circuits
by: Kölle, Michael, et al.
Published: (2025)