EPAS: Efficient Training with Progressive Activation Sharing
Fuente:
arXiv
Saved in:
| Main Authors: | Karim, Rezaul, Dialameh, Maryam, Liu, Yang, Chen, Boxing, Ahmed, Walid |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Distributed Hybrid Parallelism for Large Language Models: Comparative Study and System Design Guide
by: Amer, Hossam, et al.
Published: (2026)
by: Amer, Hossam, et al.
Published: (2026)
ECHO-LLaMA: Efficient Caching for High-Performance LLaMA Training
by: Dialameh, Maryam, et al.
Published: (2025)
by: Dialameh, Maryam, et al.
Published: (2025)
CommFuse: Hiding Tail Latency via Communication Decomposition and Fusion for Distributed LLM Training
by: Karim, Rezaul, et al.
Published: (2026)
by: Karim, Rezaul, et al.
Published: (2026)
Automated Road Extraction from Satellite Imagery Integrating Dense Depthwise Dilated Separable Spatial Pyramid Pooling with DeepLabV3+
by: Mahara, Arpan, et al.
Published: (2024)
by: Mahara, Arpan, et al.
Published: (2024)
Improving Resnet-9 Generalization Trained on Small Datasets
by: Awad, Omar Mohamed, et al.
Published: (2023)
by: Awad, Omar Mohamed, et al.
Published: (2023)
FLOP-Efficient Training: Early Stopping Based on Test-Time Compute Awareness
by: Amer, Hossam, et al.
Published: (2026)
by: Amer, Hossam, et al.
Published: (2026)
Multi-branch Spatio-Temporal Graph Neural Network For Efficient Ice Layer Thickness Prediction
by: Liu, Zesheng, et al.
Published: (2024)
by: Liu, Zesheng, et al.
Published: (2024)
K-STEMIT: Knowledge-Informed Spatio-Temporal Efficient Multi-Branch Graph Neural Network for Subsurface Stratigraphy Thickness Estimation from Radar Data
by: Liu, Zesheng, et al.
Published: (2026)
by: Liu, Zesheng, et al.
Published: (2026)
Progressive Data Dropout: An Embarrassingly Simple Approach to Faster Training
by: Sathiyanarayanan, Shriram M, et al.
Published: (2025)
by: Sathiyanarayanan, Shriram M, et al.
Published: (2025)
Post-training Quantization for Text-to-Image Diffusion Models with Progressive Calibration and Activation Relaxing
by: Tang, Siao, et al.
Published: (2023)
by: Tang, Siao, et al.
Published: (2023)
Progressive Monitoring of Generative Model Training Evolution
by: Prasad, Vidya, et al.
Published: (2024)
by: Prasad, Vidya, et al.
Published: (2024)
Efficient Training of Diffusion Mixture-of-Experts Models: A Practical Recipe
by: Liu, Yahui, et al.
Published: (2025)
by: Liu, Yahui, et al.
Published: (2025)
Leveraging Hierarchical Feature Sharing for Efficient Dataset Condensation
by: Zheng, Haizhong, et al.
Published: (2023)
by: Zheng, Haizhong, et al.
Published: (2023)
Explaining the Impact of Training on Vision Models via Activation Clustering
by: Boubekki, Ahcène, et al.
Published: (2024)
by: Boubekki, Ahcène, et al.
Published: (2024)
CLIP-UP: A Simple and Efficient Mixture-of-Experts CLIP Training Recipe with Sparse Upcycling
by: Wang, Xinze, et al.
Published: (2025)
by: Wang, Xinze, et al.
Published: (2025)
Efficient Lifelong Model Evaluation in an Era of Rapid Progress
by: Prabhu, Ameya, et al.
Published: (2024)
by: Prabhu, Ameya, et al.
Published: (2024)
Parameter-Efficient Fine-Tuning for Pre-Trained Vision Models: A Survey and Benchmark
by: Xin, Yi, et al.
Published: (2024)
by: Xin, Yi, et al.
Published: (2024)
Progressive Autoregressive Video Diffusion Models
by: Xie, Desai, et al.
Published: (2024)
by: Xie, Desai, et al.
Published: (2024)
GRIT-LP: Graph Transformer with Long-Range Skip Connection and Partitioned Spatial Graphs for Accurate Ice Layer Thickness Prediction
by: Liu, Zesheng, et al.
Published: (2025)
by: Liu, Zesheng, et al.
Published: (2025)
Learning Spatio-Temporal Patterns of Polar Ice Layers With Physics-Informed Graph Neural Network
by: Liu, Zesheng, et al.
Published: (2024)
by: Liu, Zesheng, et al.
Published: (2024)
ST-GRIT: Spatio-Temporal Graph Transformer For Internal Ice Layer Thickness Prediction
by: Liu, Zesheng, et al.
Published: (2025)
by: Liu, Zesheng, et al.
Published: (2025)
NeuCLIP: Efficient Large-Scale CLIP Training with Neural Normalizer Optimization
by: Wei, Xiyuan, et al.
Published: (2025)
by: Wei, Xiyuan, et al.
Published: (2025)
Exploring the Role of Convolutional Neural Networks (CNN) in Dental Radiography Segmentation: A Comprehensive Systematic Literature Review
by: Brahmi, Walid, et al.
Published: (2024)
by: Brahmi, Walid, et al.
Published: (2024)
Spanning Training Progress: Temporal Dual-Depth Scoring (TDDS) for Enhanced Dataset Pruning
by: Zhang, Xin, et al.
Published: (2023)
by: Zhang, Xin, et al.
Published: (2023)
DANCE: DAta-Network Co-optimization for Efficient Segmentation Model Training and Inference
by: Li, Chaojian, et al.
Published: (2021)
by: Li, Chaojian, et al.
Published: (2021)
SlimDiff: Training-Free, Activation-Guided Hands-free Slimming of Diffusion Models
by: Roy, Arani, et al.
Published: (2025)
by: Roy, Arani, et al.
Published: (2025)
EEG Foundation Models: Progresses, Benchmarking, and Open Problems
by: Liu, Dingkun, et al.
Published: (2026)
by: Liu, Dingkun, et al.
Published: (2026)
Towards Accurate and Efficient Sub-8-Bit Integer Training
by: Guo, Wenjin, et al.
Published: (2024)
by: Guo, Wenjin, et al.
Published: (2024)
LightCache: Memory-Efficient, Training-Free Acceleration for Video Generation
by: Xiao, Yang, et al.
Published: (2025)
by: Xiao, Yang, et al.
Published: (2025)
Visual prompting reimagined: The power of the Activation Prompts
by: Zhang, Yihua, et al.
Published: (2026)
by: Zhang, Yihua, et al.
Published: (2026)
DQA: An Efficient Method for Deep Quantization of Deep Neural Network Activations
by: Hu, Wenhao, et al.
Published: (2024)
by: Hu, Wenhao, et al.
Published: (2024)
Beyond Top Activations: Efficient and Reliable Crowdsourced Evaluation of Automated Interpretability
by: Oikarinen, Tuomas, et al.
Published: (2025)
by: Oikarinen, Tuomas, et al.
Published: (2025)
TPCL: Task Progressive Curriculum Learning for Robust Visual Question Answering
by: Akl, Ahmed, et al.
Published: (2024)
by: Akl, Ahmed, et al.
Published: (2024)
Deep Nets with Subsampling Layers Unwittingly Discard Useful Activations at Test-Time
by: Yang, Chiao-An, et al.
Published: (2024)
by: Yang, Chiao-An, et al.
Published: (2024)
Quantize-then-Rectify: Efficient VQ-VAE Training
by: Zhang, Borui, et al.
Published: (2025)
by: Zhang, Borui, et al.
Published: (2025)
GQKVA: Efficient Pre-training of Transformers by Grouping Queries, Keys, and Values
by: Javadi, Farnoosh, et al.
Published: (2023)
by: Javadi, Farnoosh, et al.
Published: (2023)
EfficientTrain++: Generalized Curriculum Learning for Efficient Visual Backbone Training
by: Wang, Yulin, et al.
Published: (2024)
by: Wang, Yulin, et al.
Published: (2024)
Enhancing Post-Training Quantization via Future Activation Awareness
by: Lv, Zheqi, et al.
Published: (2026)
by: Lv, Zheqi, et al.
Published: (2026)
Progressive Compression with Universally Quantized Diffusion Models
by: Yang, Yibo, et al.
Published: (2024)
by: Yang, Yibo, et al.
Published: (2024)
LPT++: Efficient Training on Mixture of Long-tailed Experts
by: Dong, Bowen, et al.
Published: (2024)
by: Dong, Bowen, et al.
Published: (2024)
Similar Items
-
Distributed Hybrid Parallelism for Large Language Models: Comparative Study and System Design Guide
by: Amer, Hossam, et al.
Published: (2026) -
ECHO-LLaMA: Efficient Caching for High-Performance LLaMA Training
by: Dialameh, Maryam, et al.
Published: (2025) -
CommFuse: Hiding Tail Latency via Communication Decomposition and Fusion for Distributed LLM Training
by: Karim, Rezaul, et al.
Published: (2026) -
Automated Road Extraction from Satellite Imagery Integrating Dense Depthwise Dilated Separable Spatial Pyramid Pooling with DeepLabV3+
by: Mahara, Arpan, et al.
Published: (2024) -
Improving Resnet-9 Generalization Trained on Small Datasets
by: Awad, Omar Mohamed, et al.
Published: (2023)