First Attentions Last: Better Exploiting First Attentions for Efficient Transformer Training
Fuente:
arXiv
Guardado en:
| Autores principales: | Kim, Gyudong, Na, Hyukju, Kim, Jin Hyeon, Jang, Hyunsung, Park, Jaemin, Hwang, Jaegi, Ha, Namkoo, Kim, Seungryong, Kim, Young Geun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CorGi: Contribution-Guided Block-Wise Interval Caching for Training-Free Acceleration of Diffusion Transformers
por: Son, Yonglak, et al.
Publicado: (2025)
por: Son, Yonglak, et al.
Publicado: (2025)
HeteroSwitch: Characterizing and Taming System-Induced Data Heterogeneity in Federated Learning
por: Kim, Gyudong, et al.
Publicado: (2024)
por: Kim, Gyudong, et al.
Publicado: (2024)
Improving Cone-Beam CT Image Quality with Knowledge Distillation-Enhanced Diffusion Model in Imbalanced Data Settings
por: Hwang, Joonil, et al.
Publicado: (2024)
por: Hwang, Joonil, et al.
Publicado: (2024)
Self-Rectifying Diffusion Sampling with Perturbed-Attention Guidance
por: Ahn, Donghoon, et al.
Publicado: (2024)
por: Ahn, Donghoon, et al.
Publicado: (2024)
When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models
por: Choi, Jiho, et al.
Publicado: (2026)
por: Choi, Jiho, et al.
Publicado: (2026)
CAMEO: Correspondence-Attention Alignment for Multi-View Diffusion Models
por: Kwon, Minkyung, et al.
Publicado: (2025)
por: Kwon, Minkyung, et al.
Publicado: (2025)
Aligned Novel View Image and Geometry Synthesis via Cross-modal Attention Instillation
por: Kwak, Min-Seop, et al.
Publicado: (2025)
por: Kwak, Min-Seop, et al.
Publicado: (2025)
First Report of Fig Leaf Rust Caused by Cerotolium fici in South Korea
por: Hyo‐Jeong Kim, et al.
Publicado: (2025)
por: Hyo‐Jeong Kim, et al.
Publicado: (2025)
Unified Diffusion Transformer for High-fidelity Text-Aware Image Restoration
por: Kim, Jin Hyeon, et al.
Publicado: (2025)
por: Kim, Jin Hyeon, et al.
Publicado: (2025)
PAtt: A Pattern Attention Network for ETA Prediction Using Historical Speed Profiles
por: Kim, ByeoungDo, et al.
Publicado: (2026)
por: Kim, ByeoungDo, et al.
Publicado: (2026)
Entropy Meets Importance: A Unified Head Importance-Entropy Score for Stable and Efficient Transformer Pruning
por: Choi, Minsik, et al.
Publicado: (2025)
por: Choi, Minsik, et al.
Publicado: (2025)
Magnetic Field Control Using an Electromagnetic Actuation System with Combined Air‐Core and Metal‐Core Coils
por: Nader Latifi Gharamaleki, et al.
Publicado: (2024)
por: Nader Latifi Gharamaleki, et al.
Publicado: (2024)
DreamMatcher: Appearance Matching Self-Attention for Semantically-Consistent Text-to-Image Personalization
por: Nam, Jisu, et al.
Publicado: (2024)
por: Nam, Jisu, et al.
Publicado: (2024)
Gated Linear Attention Transformers with Hardware-Efficient Training
por: Yang, Songlin, et al.
Publicado: (2023)
por: Yang, Songlin, et al.
Publicado: (2023)
Electrolyte‐Controlled Stripping Behavior of Electroplated Lithium Toward Efficient Lithium Metal Anodes
por: Jaemin Hwang, et al.
Publicado: (2026)
por: Jaemin Hwang, et al.
Publicado: (2026)
First record of the family Maxillipiidae (Crustacea: Malacostraca: Amphipoda) from Korea, including a new species and a new record species of the genus Maxillipius
por: Lee, Jeong-Hyeon, et al.
Publicado: (2025)
por: Lee, Jeong-Hyeon, et al.
Publicado: (2025)
KoACD: The First Korean Adolescent Dataset for Cognitive Distortion Analysis via Role-Switching Multi-LLM Negotiation
por: Kim, JunSeo, et al.
Publicado: (2025)
por: Kim, JunSeo, et al.
Publicado: (2025)
Attention-Aided MMSE for OFDM Channel Estimation: Learning Linear Filters with Attention
por: Ha, TaeJun, et al.
Publicado: (2025)
por: Ha, TaeJun, et al.
Publicado: (2025)
Self Attention with Temporal Prior: Can We Learn More from Arrow of Time?
por: Kim, Kyung Geun, et al.
Publicado: (2023)
por: Kim, Kyung Geun, et al.
Publicado: (2023)
FEAST: Fully Connected Expressive Attention for Spatial Transcriptomics
por: Jeong, Taejin, et al.
Publicado: (2026)
por: Jeong, Taejin, et al.
Publicado: (2026)
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers
por: Bae, Jongseong, et al.
Publicado: (2024)
por: Bae, Jongseong, et al.
Publicado: (2024)
Adaptive Graph Rewiring to Mitigate Over-Squashing in Mesh-Based GNNs for Fluid Dynamics Simulations
por: Seo, Sangwoo, et al.
Publicado: (2025)
por: Seo, Sangwoo, et al.
Publicado: (2025)
On a lower bound of Hausdorff dimension of weighted singular vectors
por: Kim, Taehyeong, et al.
Publicado: (2022)
por: Kim, Taehyeong, et al.
Publicado: (2022)
DiffuseSlide: Training-Free High Frame Rate Video Generation Diffusion
por: Hwang, Geunmin, et al.
Publicado: (2025)
por: Hwang, Geunmin, et al.
Publicado: (2025)
A Training-free Sub-quadratic Cost Transformer Model Serving Framework With Hierarchically Pruned Attention
por: Lee, Heejun, et al.
Publicado: (2024)
por: Lee, Heejun, et al.
Publicado: (2024)
Rice E3 ligase OsRFPH2‐16 acts as a negative regulator to mediate the degradation of OsPIP1;1 under salt stress
por: Jong Ho Kim, et al.
Publicado: (2025)
por: Jong Ho Kim, et al.
Publicado: (2025)
Learning Video Temporal Dynamics with Cross-Modal Attention for Robust Audio-Visual Speech Recognition
por: Kim, Sungnyun, et al.
Publicado: (2024)
por: Kim, Sungnyun, et al.
Publicado: (2024)
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
por: Kim, Jinyeong, et al.
Publicado: (2025)
por: Kim, Jinyeong, et al.
Publicado: (2025)
Thermodynamic Isomorphism of Transformers: A Lagrangian Approach to Attention Dynamics
por: Kim, Gunn
Publicado: (2026)
por: Kim, Gunn
Publicado: (2026)
Text-Aware Image Restoration with Diffusion Models
por: Min, Jaewon, et al.
Publicado: (2025)
por: Min, Jaewon, et al.
Publicado: (2025)
Not All Adapters Matter: Selective Adapter Freezing for Memory-Efficient Fine-Tuning of Language Models
por: Son, Hyegang, et al.
Publicado: (2024)
por: Son, Hyegang, et al.
Publicado: (2024)
Spatial Transport Optimization by Repositioning Attention Map for Training-Free Text-to-Image Synthesis
por: Han, Woojung, et al.
Publicado: (2025)
por: Han, Woojung, et al.
Publicado: (2025)
MoDiTalker: Motion-Disentangled Diffusion Model for High-Fidelity Talking Head Generation
por: Kim, Seyeon, et al.
Publicado: (2024)
por: Kim, Seyeon, et al.
Publicado: (2024)
Are Self-Attentions Effective for Time Series Forecasting?
por: Kim, Dongbin, et al.
Publicado: (2024)
por: Kim, Dongbin, et al.
Publicado: (2024)
Far‐infrared irradiation attenuates platelet‐derived growth factor‐stimulated vascular smooth muscle cell migration through protein phosphatase 2A‐mediated Akt inhibition
por: Na‐Young Lee, et al.
Publicado: (2025)
por: Na‐Young Lee, et al.
Publicado: (2025)
Recent Advances in Nanotechnology‐Mediated Noninvasive Transdermal and Topical Delivery of Proteins
por: Junghyeon Ko, et al.
Publicado: (2024)
por: Junghyeon Ko, et al.
Publicado: (2024)
Rethinking the Use of Vision Transformers for AI-Generated Image Detection
por: Park, NaHyeon, et al.
Publicado: (2025)
por: Park, NaHyeon, et al.
Publicado: (2025)
SPARTA: Advancing Sparse Attention in Spiking Neural Networks via Spike-Timing-Based Prioritization
por: Jang, Minsuk, et al.
Publicado: (2025)
por: Jang, Minsuk, et al.
Publicado: (2025)
Atomic-Scale Mechanisms of SiO$_2$ Plasma-Enhanced Chemical Vapor Deposition Revealed by Molecular Dynamics with a Machine-Learning Interatomic Potential
por: Kim, Jaehoon, et al.
Publicado: (2026)
por: Kim, Jaehoon, et al.
Publicado: (2026)
SEA: Sparse Linear Attention with Estimated Attention Mask
por: Lee, Heejun, et al.
Publicado: (2023)
por: Lee, Heejun, et al.
Publicado: (2023)
Ejemplares similares
-
CorGi: Contribution-Guided Block-Wise Interval Caching for Training-Free Acceleration of Diffusion Transformers
por: Son, Yonglak, et al.
Publicado: (2025) -
HeteroSwitch: Characterizing and Taming System-Induced Data Heterogeneity in Federated Learning
por: Kim, Gyudong, et al.
Publicado: (2024) -
Improving Cone-Beam CT Image Quality with Knowledge Distillation-Enhanced Diffusion Model in Imbalanced Data Settings
por: Hwang, Joonil, et al.
Publicado: (2024) -
Self-Rectifying Diffusion Sampling with Perturbed-Attention Guidance
por: Ahn, Donghoon, et al.
Publicado: (2024) -
When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models
por: Choi, Jiho, et al.
Publicado: (2026)