PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zeman, Deng, Yuan, Zhong, Peilin, Razaviyayn, Meisam, Mirrokni, Vahab |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Addax: Utilizing Zeroth-Order Gradients to Improve Memory Efficiency and Performance of SGD for Fine-Tuning Language Models
by: Li, Zeman, et al.
Published: (2024)
by: Li, Zeman, et al.
Published: (2024)
Memory Caching: RNNs with Growing Memory
by: Behrouz, Ali, et al.
Published: (2026)
by: Behrouz, Ali, et al.
Published: (2026)
Nested Learning: The Illusion of Deep Learning Architectures
by: Behrouz, Ali, et al.
Published: (2025)
by: Behrouz, Ali, et al.
Published: (2025)
It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization
by: Behrouz, Ali, et al.
Published: (2025)
by: Behrouz, Ali, et al.
Published: (2025)
Synthetic Text Generation for Training Large Language Models via Gradient Matching
by: Nguyen, Dang, et al.
Published: (2025)
by: Nguyen, Dang, et al.
Published: (2025)
TNT: Improving Chunkwise Training for Test-Time Memorization
by: Li, Zeman, et al.
Published: (2025)
by: Li, Zeman, et al.
Published: (2025)
Sampling and Loss Weights in Multi-Domain Training
by: Salmani, Mahdi, et al.
Published: (2025)
by: Salmani, Mahdi, et al.
Published: (2025)
Less is More: Convergence Benefits of Fewer Data Weight Updates over Longer Horizon
by: Das, Rudrajit, et al.
Published: (2026)
by: Das, Rudrajit, et al.
Published: (2026)
ATLAS: Learning to Optimally Memorize the Context at Test Time
by: Behrouz, Ali, et al.
Published: (2025)
by: Behrouz, Ali, et al.
Published: (2025)
Titans: Learning to Memorize at Test Time
by: Behrouz, Ali, et al.
Published: (2024)
by: Behrouz, Ali, et al.
Published: (2024)
PolySketchFormer: Fast Transformers via Sketching Polynomial Kernels
by: Kacham, Praneeth, et al.
Published: (2023)
by: Kacham, Praneeth, et al.
Published: (2023)
DiSK: Differentially Private Optimizer with Simplified Kalman Filter for Noise Reduction
by: Zhang, Xinwei, et al.
Published: (2024)
by: Zhang, Xinwei, et al.
Published: (2024)
Optimal Differentially Private Model Training with Public Data
by: Lowy, Andrew, et al.
Published: (2023)
by: Lowy, Andrew, et al.
Published: (2023)
MS-SSM: A Multi-Scale State Space Model for Efficient Sequence Modeling
by: Karami, Mahdi, et al.
Published: (2025)
by: Karami, Mahdi, et al.
Published: (2025)
Adaptively Private Next-Token Prediction of Large Language Models
by: Flemings, James, et al.
Published: (2024)
by: Flemings, James, et al.
Published: (2024)
Private Stochastic Optimization With Large Worst-Case Lipschitz Parameter
by: Lowy, Andrew, et al.
Published: (2022)
by: Lowy, Andrew, et al.
Published: (2022)
On the Inherent Privacy of Zeroth Order Projected Gradient Descent
by: Gupta, Devansh, et al.
Published: (2025)
by: Gupta, Devansh, et al.
Published: (2025)
Differentially Private Graph Learning via Sensitivity-Bounded Personalized PageRank
by: Epasto, Alessandro, et al.
Published: (2022)
by: Epasto, Alessandro, et al.
Published: (2022)
Understanding the Role of Training Data in Test-Time Scaling
by: Javanmard, Adel, et al.
Published: (2025)
by: Javanmard, Adel, et al.
Published: (2025)
Private Federated Learning Without a Trusted Server: Optimal Algorithms for Convex Losses
by: Lowy, Andrew, et al.
Published: (2021)
by: Lowy, Andrew, et al.
Published: (2021)
Policy Gradient Converges to the Globally Optimal Policy for Nearly Linear-Quadratic Regulators
by: Han, Yinbin, et al.
Published: (2023)
by: Han, Yinbin, et al.
Published: (2023)
PriorBoost: An Adaptive Algorithm for Learning from Aggregate Responses
by: Javanmard, Adel, et al.
Published: (2024)
by: Javanmard, Adel, et al.
Published: (2024)
Output Perturbation for Differentially Private Convex Optimization: Faster and More General
by: Lowy, Andrew, et al.
Published: (2021)
by: Lowy, Andrew, et al.
Published: (2021)
Differentially Private Next-Token Prediction of Large Language Models
by: Flemings, James, et al.
Published: (2024)
by: Flemings, James, et al.
Published: (2024)
A Stochastic Optimization Framework for Private and Fair Learning From Decentralized Data
by: Gupta, Devansh, et al.
Published: (2024)
by: Gupta, Devansh, et al.
Published: (2024)
High-Dimensional Geometric Streaming for Nearly Low Rank Data
by: Esfandiari, Hossein, et al.
Published: (2024)
by: Esfandiari, Hossein, et al.
Published: (2024)
Differentially Private In-context Learning via Sampling Few-shot Mixed with Zero-shot Outputs
by: Flemings, James, et al.
Published: (2025)
by: Flemings, James, et al.
Published: (2025)
Efficient Data Selection at Scale via Influence Distillation
by: Nikdan, Mahdi, et al.
Published: (2025)
by: Nikdan, Mahdi, et al.
Published: (2025)
Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning Models
by: Javanmard, Adel, et al.
Published: (2026)
by: Javanmard, Adel, et al.
Published: (2026)
Lattice: Learning to Efficiently Compress the Memory
by: Karami, Mahdi, et al.
Published: (2025)
by: Karami, Mahdi, et al.
Published: (2025)
DOPPLER: Differentially Private Optimizers with Low-pass Filter for Privacy Noise Reduction
by: Zhang, Xinwei, et al.
Published: (2024)
by: Zhang, Xinwei, et al.
Published: (2024)
f-FERM: A Scalable Framework for Robust Fair Empirical Risk Minimization
by: Baharlouei, Sina, et al.
Published: (2023)
by: Baharlouei, Sina, et al.
Published: (2023)
Neural Network-Based Score Estimation in Diffusion Models: Optimization and Generalization
by: Han, Yinbin, et al.
Published: (2024)
by: Han, Yinbin, et al.
Published: (2024)
Stochastic Control for Fine-tuning Diffusion Models: Optimality, Regularity, and Convergence
by: Han, Yinbin, et al.
Published: (2024)
by: Han, Yinbin, et al.
Published: (2024)
Perturb-and-Project: Differentially Private Similarities and Marginals
by: Cohen-Addad, Vincent, et al.
Published: (2024)
by: Cohen-Addad, Vincent, et al.
Published: (2024)
Differentially Private Synthetic Data Release for Topics API Outputs
by: Dick, Travis, et al.
Published: (2025)
by: Dick, Travis, et al.
Published: (2025)
Learning Rate Schedules in the Presence of Distribution Shift
by: Fahrbach, Matthew, et al.
Published: (2023)
by: Fahrbach, Matthew, et al.
Published: (2023)
Optimistic Rates for Learning from Label Proportions
by: Li, Gene, et al.
Published: (2024)
by: Li, Gene, et al.
Published: (2024)
Trellis: Learning to Compress Key-Value Memory in Attention Models
by: Karami, Mahdi, et al.
Published: (2025)
by: Karami, Mahdi, et al.
Published: (2025)
Tradeoffs between convergence rate and noise amplification for momentum-based accelerated optimization algorithms
by: Mohammadi, Hesameddin, et al.
Published: (2022)
by: Mohammadi, Hesameddin, et al.
Published: (2022)
Similar Items
-
Addax: Utilizing Zeroth-Order Gradients to Improve Memory Efficiency and Performance of SGD for Fine-Tuning Language Models
by: Li, Zeman, et al.
Published: (2024) -
Memory Caching: RNNs with Growing Memory
by: Behrouz, Ali, et al.
Published: (2026) -
Nested Learning: The Illusion of Deep Learning Architectures
by: Behrouz, Ali, et al.
Published: (2025) -
It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization
by: Behrouz, Ali, et al.
Published: (2025) -
Synthetic Text Generation for Training Large Language Models via Gradient Matching
by: Nguyen, Dang, et al.
Published: (2025)