Saved in:
| Main Authors: | Hu, Pihe, Li, Shaolong, Huang, Longbo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2408.11746 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Value-Based Deep Multi-Agent Reinforcement Learning with Dynamic Sparse Training
by: Hu, Pihe, et al.
Published: (2024)
by: Hu, Pihe, et al.
Published: (2024)
Sparse-IFT: Sparse Iso-FLOP Transformations for Maximizing Training Efficiency
by: Thangarasa, Vithursan, et al.
Published: (2023)
by: Thangarasa, Vithursan, et al.
Published: (2023)
Finite-time Convergence Analysis of Actor-Critic with Evolving Reward
by: Hu, Rui, et al.
Published: (2025)
by: Hu, Rui, et al.
Published: (2025)
Beyond Squared Error: Exploring Loss Design for Enhanced Training of Generative Flow Networks
by: Hu, Rui, et al.
Published: (2024)
by: Hu, Rui, et al.
Published: (2024)
Beyond the Proxy: Trajectory-Distilled Guidance for Offline GFlowNet Training
by: Chen, Ruishuo, et al.
Published: (2025)
by: Chen, Ruishuo, et al.
Published: (2025)
FLOP-Efficient Training: Early Stopping Based on Test-Time Compute Awareness
by: Amer, Hossam, et al.
Published: (2026)
by: Amer, Hossam, et al.
Published: (2026)
Accelerating Transformer Pre-training with 2:4 Sparsity
by: Hu, Yuezhou, et al.
Published: (2024)
by: Hu, Yuezhou, et al.
Published: (2024)
FALCON: FLOP-Aware Combinatorial Optimization for Neural Network Pruning
by: Meng, Xiang, et al.
Published: (2024)
by: Meng, Xiang, et al.
Published: (2024)
Real-Time Parallel Counterfactual Regret Minimization
by: Li, Boning, et al.
Published: (2026)
by: Li, Boning, et al.
Published: (2026)
Finite-Time Convergence Analysis of ODE-based Generative Models for Stochastic Interpolants
by: Liu, Yuhao, et al.
Published: (2025)
by: Liu, Yuhao, et al.
Published: (2025)
Finite-Time Analysis of Discrete-Time Stochastic Interpolants
by: Liu, Yuhao, et al.
Published: (2025)
by: Liu, Yuhao, et al.
Published: (2025)
EcoSpa: Efficient Transformer Training with Coupled Sparsity
by: Xiao, Jinqi, et al.
Published: (2025)
by: Xiao, Jinqi, et al.
Published: (2025)
Accelerating Transformer Inference and Training with 2:4 Activation Sparsity
by: Haziza, Daniel, et al.
Published: (2025)
by: Haziza, Daniel, et al.
Published: (2025)
Adversarial Network Optimization under Bandit Feedback: Maximizing Utility in Non-Stationary Multi-Hop Networks
by: Dai, Yan, et al.
Published: (2024)
by: Dai, Yan, et al.
Published: (2024)
Provably Efficient Partially Observable Risk-Sensitive Reinforcement Learning with Hindsight Observation
by: Zhang, Tonghe, et al.
Published: (2024)
by: Zhang, Tonghe, et al.
Published: (2024)
RL-CFR: Improving Action Abstraction for Imperfect Information Extensive-Form Games with Reinforcement Learning
by: Li, Boning, et al.
Published: (2024)
by: Li, Boning, et al.
Published: (2024)
Problems with Chinchilla Approach 2: Systematic Biases in IsoFLOP Parabola Fits
by: Czech, Eric, et al.
Published: (2026)
by: Czech, Eric, et al.
Published: (2026)
uniINF: Best-of-Both-Worlds Algorithm for Parameter-Free Heavy-Tailed MABs
by: Chen, Yu, et al.
Published: (2024)
by: Chen, Yu, et al.
Published: (2024)
Layer-Aware Influence for Online Data Valuation Estimation
by: Yang, Ziao, et al.
Published: (2025)
by: Yang, Ziao, et al.
Published: (2025)
LatentMoE: Toward Optimal Accuracy per FLOP and Parameter in Mixture of Experts
by: Elango, Venmugil, et al.
Published: (2026)
by: Elango, Venmugil, et al.
Published: (2026)
Reparameterization Proximal Policy Optimization
by: Zhong, Hai, et al.
Published: (2025)
by: Zhong, Hai, et al.
Published: (2025)
Reparameterization Flow Policy Optimization
by: Zhong, Hai, et al.
Published: (2026)
by: Zhong, Hai, et al.
Published: (2026)
Provable Risk-Sensitive Distributional Reinforcement Learning with General Function Approximation
by: Chen, Yu, et al.
Published: (2024)
by: Chen, Yu, et al.
Published: (2024)
Continuous K-Max Bandits
by: Chen, Yu, et al.
Published: (2025)
by: Chen, Yu, et al.
Published: (2025)
ProxSparse: Regularized Learning of Semi-Structured Sparsity Masks for Pretrained LLMs
by: Liu, Hongyi, et al.
Published: (2025)
by: Liu, Hongyi, et al.
Published: (2025)
Adaptive Sparsity Level during Training for Efficient Time Series Forecasting with Transformers
by: Atashgahi, Zahra, et al.
Published: (2023)
by: Atashgahi, Zahra, et al.
Published: (2023)
Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs
by: Ling Team, et al.
Published: (2025)
by: Ling Team, et al.
Published: (2025)
CAST: Continuous and Differentiable Semi-Structured Sparsity-Aware Training for Large Language Models
by: Huang, Weiyu, et al.
Published: (2025)
by: Huang, Weiyu, et al.
Published: (2025)
Understanding the Training and Generalization of Pretrained Transformer for Sequential Decision Making
by: Wang, Hanzhao, et al.
Published: (2024)
by: Wang, Hanzhao, et al.
Published: (2024)
Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers
by: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Published: (2024)
by: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Published: (2024)
Homeostasis and Sparsity in Transformer
by: Kotyuzanskiy, Leonid, et al.
Published: (2024)
by: Kotyuzanskiy, Leonid, et al.
Published: (2024)
Beyond Shallow Behavior: Task-Efficient Value-Based Multi-Task Offline MARL via Skill Discovery
by: Wang, Xun, et al.
Published: (2025)
by: Wang, Xun, et al.
Published: (2025)
PowerFlow: Unlocking the Dual Nature of LLMs via Principled Distribution Matching
by: Chen, Ruishuo, et al.
Published: (2026)
by: Chen, Ruishuo, et al.
Published: (2026)
Neural networks can be FLOP-efficient integrators of 1D oscillatory integrands
by: Sinha, Anshuman, et al.
Published: (2024)
by: Sinha, Anshuman, et al.
Published: (2024)
Best-of-Both-Worlds for Heavy-Tailed Markov Decision Processes
by: Chen, Yu, et al.
Published: (2026)
by: Chen, Yu, et al.
Published: (2026)
Learnable Permutation for Structured Sparsity on Transformer Models
by: Li, Zekai, et al.
Published: (2026)
by: Li, Zekai, et al.
Published: (2026)
From Solo to Symphony: Orchestrating Multi-Agent Collaboration with Single-Agent Demos
by: Wang, Xun, et al.
Published: (2025)
by: Wang, Xun, et al.
Published: (2025)
A Quadratic Synchronization Rule for Distributed Deep Learning
by: Gu, Xinran, et al.
Published: (2023)
by: Gu, Xinran, et al.
Published: (2023)
Spark Transformer: Reactivating Sparsity in FFN and Attention
by: You, Chong, et al.
Published: (2025)
by: You, Chong, et al.
Published: (2025)
Self-Ablating Transformers: More Interpretability, Less Sparsity
by: Ferrao, Jeremias, et al.
Published: (2025)
by: Ferrao, Jeremias, et al.
Published: (2025)
Similar Items
-
Value-Based Deep Multi-Agent Reinforcement Learning with Dynamic Sparse Training
by: Hu, Pihe, et al.
Published: (2024) -
Sparse-IFT: Sparse Iso-FLOP Transformations for Maximizing Training Efficiency
by: Thangarasa, Vithursan, et al.
Published: (2023) -
Finite-time Convergence Analysis of Actor-Critic with Evolving Reward
by: Hu, Rui, et al.
Published: (2025) -
Beyond Squared Error: Exploring Loss Design for Enhanced Training of Generative Flow Networks
by: Hu, Rui, et al.
Published: (2024) -
Beyond the Proxy: Trajectory-Distilled Guidance for Offline GFlowNet Training
by: Chen, Ruishuo, et al.
Published: (2025)