Sampling and Loss Weights in Multi-Domain Training
Fuente:
arXiv
Saved in:
| Main Authors: | Salmani, Mahdi, Worah, Pratik, Razaviyayn, Meisam, Mirrokni, Vahab |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization
by: Behrouz, Ali, et al.
Published: (2025)
by: Behrouz, Ali, et al.
Published: (2025)
Nested Learning: The Illusion of Deep Learning Architectures
by: Behrouz, Ali, et al.
Published: (2025)
by: Behrouz, Ali, et al.
Published: (2025)
TNT: Improving Chunkwise Training for Test-Time Memorization
by: Li, Zeman, et al.
Published: (2025)
by: Li, Zeman, et al.
Published: (2025)
Memory Caching: RNNs with Growing Memory
by: Behrouz, Ali, et al.
Published: (2026)
by: Behrouz, Ali, et al.
Published: (2026)
ECO: Quantized Training without Full-Precision Master Weights
by: Nikdan, Mahdi, et al.
Published: (2026)
by: Nikdan, Mahdi, et al.
Published: (2026)
Early Stopping for Large Reasoning Models via Confidence Dynamics
by: Hosseini, Parsa, et al.
Published: (2026)
by: Hosseini, Parsa, et al.
Published: (2026)
Less is More: Convergence Benefits of Fewer Data Weight Updates over Longer Horizon
by: Das, Rudrajit, et al.
Published: (2026)
by: Das, Rudrajit, et al.
Published: (2026)
Lattice: Learning to Efficiently Compress the Memory
by: Karami, Mahdi, et al.
Published: (2025)
by: Karami, Mahdi, et al.
Published: (2025)
Learning Rate Schedules in the Presence of Distribution Shift
by: Fahrbach, Matthew, et al.
Published: (2023)
by: Fahrbach, Matthew, et al.
Published: (2023)
Understanding the Role of Training Data in Test-Time Scaling
by: Javanmard, Adel, et al.
Published: (2025)
by: Javanmard, Adel, et al.
Published: (2025)
PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts
by: Li, Zeman, et al.
Published: (2025)
by: Li, Zeman, et al.
Published: (2025)
Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning Models
by: Javanmard, Adel, et al.
Published: (2026)
by: Javanmard, Adel, et al.
Published: (2026)
Synthetic Text Generation for Training Large Language Models via Gradient Matching
by: Nguyen, Dang, et al.
Published: (2025)
by: Nguyen, Dang, et al.
Published: (2025)
Learning from Aggregate responses: Instance Level versus Bag Level Loss Functions
by: Javanmard, Adel, et al.
Published: (2024)
by: Javanmard, Adel, et al.
Published: (2024)
Titans: Learning to Memorize at Test Time
by: Behrouz, Ali, et al.
Published: (2024)
by: Behrouz, Ali, et al.
Published: (2024)
Optimistic Rates for Learning from Label Proportions
by: Li, Gene, et al.
Published: (2024)
by: Li, Gene, et al.
Published: (2024)
DiSK: Differentially Private Optimizer with Simplified Kalman Filter for Noise Reduction
by: Zhang, Xinwei, et al.
Published: (2024)
by: Zhang, Xinwei, et al.
Published: (2024)
Addax: Utilizing Zeroth-Order Gradients to Improve Memory Efficiency and Performance of SGD for Fine-Tuning Language Models
by: Li, Zeman, et al.
Published: (2024)
by: Li, Zeman, et al.
Published: (2024)
PolarQuant: Quantizing KV Caches with Polar Transformation
by: Han, Insu, et al.
Published: (2025)
by: Han, Insu, et al.
Published: (2025)
SubGen: Token Generation in Sublinear Time and Memory
by: Zandieh, Amir, et al.
Published: (2024)
by: Zandieh, Amir, et al.
Published: (2024)
TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate
by: Zandieh, Amir, et al.
Published: (2025)
by: Zandieh, Amir, et al.
Published: (2025)
A Goemans-Williamson type algorithm for identifying subcohorts in clinical trials
by: Worah, Pratik
Published: (2025)
by: Worah, Pratik
Published: (2025)
Enhancing selectivity using Wasserstein distance based reweighing
by: Worah, Pratik
Published: (2024)
by: Worah, Pratik
Published: (2024)
Private Federated Learning Without a Trusted Server: Optimal Algorithms for Convex Losses
by: Lowy, Andrew, et al.
Published: (2021)
by: Lowy, Andrew, et al.
Published: (2021)
ATLAS: Learning to Optimally Memorize the Context at Test Time
by: Behrouz, Ali, et al.
Published: (2025)
by: Behrouz, Ali, et al.
Published: (2025)
Adaptive Weighted Loss for Sequential Recommendations on Sparse Domains
by: Mittal, Akshay, et al.
Published: (2025)
by: Mittal, Akshay, et al.
Published: (2025)
Understanding Transformer Reasoning Capabilities via Graph Algorithms
by: Sanford, Clayton, et al.
Published: (2024)
by: Sanford, Clayton, et al.
Published: (2024)
MemLoss: Enhancing Adversarial Training with Recycling Adversarial Examples
by: Mahdi, Soroush, et al.
Published: (2025)
by: Mahdi, Soroush, et al.
Published: (2025)
Meta-GCN: A Dynamically Weighted Loss Minimization Method for Dealing with the Data Imbalance in Graph Neural Networks
by: Mohammadizadeh, Mahdi, et al.
Published: (2024)
by: Mohammadizadeh, Mahdi, et al.
Published: (2024)
Output Perturbation for Differentially Private Convex Optimization: Faster and More General
by: Lowy, Andrew, et al.
Published: (2021)
by: Lowy, Andrew, et al.
Published: (2021)
Differentially Private Synthetic Data Release for Topics API Outputs
by: Dick, Travis, et al.
Published: (2025)
by: Dick, Travis, et al.
Published: (2025)
Private Stochastic Optimization With Large Worst-Case Lipschitz Parameter
by: Lowy, Andrew, et al.
Published: (2022)
by: Lowy, Andrew, et al.
Published: (2022)
MS-SSM: A Multi-Scale State Space Model for Efficient Sequence Modeling
by: Karami, Mahdi, et al.
Published: (2025)
by: Karami, Mahdi, et al.
Published: (2025)
When Losses Align: Gradient-Based Composite Loss Weighting for Efficient Pretraining
by: Karpukhin, Ivan, et al.
Published: (2026)
by: Karpukhin, Ivan, et al.
Published: (2026)
Trellis: Learning to Compress Key-Value Memory in Attention Models
by: Karami, Mahdi, et al.
Published: (2025)
by: Karami, Mahdi, et al.
Published: (2025)
f-FERM: A Scalable Framework for Robust Fair Empirical Risk Minimization
by: Baharlouei, Sina, et al.
Published: (2023)
by: Baharlouei, Sina, et al.
Published: (2023)
Neural Network-Based Score Estimation in Diffusion Models: Optimization and Generalization
by: Han, Yinbin, et al.
Published: (2024)
by: Han, Yinbin, et al.
Published: (2024)
Rewriting the Budget: A General Framework for Black-Box Attacks Under Cost Asymmetry
by: Salmani, Mahdi, et al.
Published: (2025)
by: Salmani, Mahdi, et al.
Published: (2025)
Debate as Reward: A Multi-Agent Reward System for Scientific Ideation via RL Post-Training
by: Salimi, Moein, et al.
Published: (2026)
by: Salimi, Moein, et al.
Published: (2026)
MDGMIX: Boundary-Aware Subgraph Mixing for Multi-Domain Graph Pre-Training
by: Zheng, Ziyu, et al.
Published: (2026)
by: Zheng, Ziyu, et al.
Published: (2026)
Similar Items
-
It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization
by: Behrouz, Ali, et al.
Published: (2025) -
Nested Learning: The Illusion of Deep Learning Architectures
by: Behrouz, Ali, et al.
Published: (2025) -
TNT: Improving Chunkwise Training for Test-Time Memorization
by: Li, Zeman, et al.
Published: (2025) -
Memory Caching: RNNs with Growing Memory
by: Behrouz, Ali, et al.
Published: (2026) -
ECO: Quantized Training without Full-Precision Master Weights
by: Nikdan, Mahdi, et al.
Published: (2026)