Understanding the Training Speedup from Sampling with Approximate Losses
Fuente:
arXiv
Saved in:
| Main Authors: | Das, Rudrajit, Chen, Xi, Ieong, Bertram, Bansal, Parikshit, Sanghavi, Sujay |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enabling Approximate Joint Sampling in Diffusion LMs
by: Bansal, Parikshit, et al.
Published: (2025)
by: Bansal, Parikshit, et al.
Published: (2025)
Context-Free Synthetic Data Mitigates Forgetting
by: Bansal, Parikshit, et al.
Published: (2025)
by: Bansal, Parikshit, et al.
Published: (2025)
Understanding Self-Supervised Learning via Gaussian Mixture Models
by: Bansal, Parikshit, et al.
Published: (2024)
by: Bansal, Parikshit, et al.
Published: (2024)
Upweighting Easy Samples in Fine-Tuning Mitigates Forgetting
by: Sanyal, Sunny, et al.
Published: (2025)
by: Sanyal, Sunny, et al.
Published: (2025)
Towards Quantifying the Preconditioning Effect of Adam
by: Das, Rudrajit, et al.
Published: (2024)
by: Das, Rudrajit, et al.
Published: (2024)
Test-Time Speculation
by: Kumar, Avinash, et al.
Published: (2026)
by: Kumar, Avinash, et al.
Published: (2026)
Retraining with Predicted Hard Labels Provably Increases Model Accuracy
by: Das, Rudrajit, et al.
Published: (2024)
by: Das, Rudrajit, et al.
Published: (2024)
HiSpec: Hierarchical Speculative Decoding for LLMs
by: Kumar, Avinash, et al.
Published: (2025)
by: Kumar, Avinash, et al.
Published: (2025)
Learning Mixtures of Experts with EM: A Mirror Descent Perspective
by: Fruytier, Quentin, et al.
Published: (2024)
by: Fruytier, Quentin, et al.
Published: (2024)
Geometric Median (GM) Matching for Robust Data Pruning
by: Acharya, Anish, et al.
Published: (2024)
by: Acharya, Anish, et al.
Published: (2024)
Self-Boost via Optimal Retraining: An Analysis via Approximate Message Passing
by: Javanmard, Adel, et al.
Published: (2025)
by: Javanmard, Adel, et al.
Published: (2025)
Blocking Bandits
by: Basu, Soumya, et al.
Published: (2019)
by: Basu, Soumya, et al.
Published: (2019)
Geometric Median Matching for Robust k-Subset Selection from Noisy Data
by: Acharya, Anish, et al.
Published: (2025)
by: Acharya, Anish, et al.
Published: (2025)
Asymptotically-Optimal Gaussian Bandits with Side Observations
by: Atsidakou, Alexia, et al.
Published: (2025)
by: Atsidakou, Alexia, et al.
Published: (2025)
Finite-Time Logarithmic Bayes Regret Upper Bounds
by: Atsidakou, Alexia, et al.
Published: (2023)
by: Atsidakou, Alexia, et al.
Published: (2023)
Small LLMs with Expert Blocks Are Good Enough for Hyperparamter Tuning
by: Naphade, Om, et al.
Published: (2025)
by: Naphade, Om, et al.
Published: (2025)
Understanding Contrastive Representation Learning from Positive Unlabeled (PU) Data
by: Acharya, Anish, et al.
Published: (2024)
by: Acharya, Anish, et al.
Published: (2024)
Adaptive and Optimal Second-order Optimistic Methods for Minimax Optimization
by: Jiang, Ruichen, et al.
Published: (2024)
by: Jiang, Ruichen, et al.
Published: (2024)
Less is More: Convergence Benefits of Fewer Data Weight Updates over Longer Horizon
by: Das, Rudrajit, et al.
Published: (2026)
by: Das, Rudrajit, et al.
Published: (2026)
Time Weaver: A Conditional Time Series Generation Model
by: Narasimhan, Sai Shankar, et al.
Published: (2024)
by: Narasimhan, Sai Shankar, et al.
Published: (2024)
When Attention Collapses: How Degenerate Layers in LLMs Enable Smaller, Stronger Models
by: Sanyal, Sunny, et al.
Published: (2024)
by: Sanyal, Sunny, et al.
Published: (2024)
In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
by: Collins, Liam, et al.
Published: (2024)
by: Collins, Liam, et al.
Published: (2024)
Entropy Aware Reward Guidance for Diffusion Language Model Alignment
by: Tejaswi, Atula, et al.
Published: (2026)
by: Tejaswi, Atula, et al.
Published: (2026)
Understanding Uncertainty Sampling via Equivalent Loss
by: Liu, Shang, et al.
Published: (2023)
by: Liu, Shang, et al.
Published: (2023)
Beyond Speedup -- Utilizing KV Cache for Sampling and Reasoning
by: Xing, Zeyu, et al.
Published: (2026)
by: Xing, Zeyu, et al.
Published: (2026)
Pretrained deep models outperform GBDTs in Learning-To-Rank under label scarcity
by: Hou, Charlie, et al.
Published: (2023)
by: Hou, Charlie, et al.
Published: (2023)
Sculpting Latent Spaces With MMD: Disentanglement With Programmable Priors
by: Fruytier, Quentin, et al.
Published: (2025)
by: Fruytier, Quentin, et al.
Published: (2025)
Quantum Speedup for Spectral Approximation of Kronecker Products
by: Gao, Yeqi, et al.
Published: (2024)
by: Gao, Yeqi, et al.
Published: (2024)
Positive Unlabeled Contrastive Learning
by: Acharya, Anish, et al.
Published: (2022)
by: Acharya, Anish, et al.
Published: (2022)
From Edges to Depth: Probing the Spatial Hierarchy in Vision Transformers
by: Sanghavi, Jainum
Published: (2026)
by: Sanghavi, Jainum
Published: (2026)
Efficiently Training Neural Networks for Imperfect Information Games by Sampling Information Sets
by: Bertram, Timo, et al.
Published: (2024)
by: Bertram, Timo, et al.
Published: (2024)
Power Flow Approximations for Multiphase Distribution Networks using Gaussian Processes
by: Glover, Daniel, et al.
Published: (2025)
by: Glover, Daniel, et al.
Published: (2025)
EPiC: Towards Lossless Speedup for Reasoning Training through Edge-Preserving CoT Condensation
by: Jia, Jinghan, et al.
Published: (2025)
by: Jia, Jinghan, et al.
Published: (2025)
Sampling and Loss Weights in Multi-Domain Training
by: Salmani, Mahdi, et al.
Published: (2025)
by: Salmani, Mahdi, et al.
Published: (2025)
InfoPO: On Mutual Information Maximization for Large Language Model Alignment
by: Xiao, Teng, et al.
Published: (2025)
by: Xiao, Teng, et al.
Published: (2025)
MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation
by: Ekbote, Chanakya, et al.
Published: (2025)
by: Ekbote, Chanakya, et al.
Published: (2025)
Unified Graph Networks (UGN): A Deep Neural Framework for Solving Graph Problems
by: Dawn, Rudrajit, et al.
Published: (2025)
by: Dawn, Rudrajit, et al.
Published: (2025)
Approximate Equivariance in Reinforcement Learning
by: Park, Jung Yeon, et al.
Published: (2024)
by: Park, Jung Yeon, et al.
Published: (2024)
Sample Complexity of Causal Identification with Temporal Heterogeneity
by: Rathod, Ameya, et al.
Published: (2026)
by: Rathod, Ameya, et al.
Published: (2026)
Towards Understanding the Influence of Training Samples on Explanations
by: Artelt, André, et al.
Published: (2024)
by: Artelt, André, et al.
Published: (2024)
Similar Items
-
Enabling Approximate Joint Sampling in Diffusion LMs
by: Bansal, Parikshit, et al.
Published: (2025) -
Context-Free Synthetic Data Mitigates Forgetting
by: Bansal, Parikshit, et al.
Published: (2025) -
Understanding Self-Supervised Learning via Gaussian Mixture Models
by: Bansal, Parikshit, et al.
Published: (2024) -
Upweighting Easy Samples in Fine-Tuning Mitigates Forgetting
by: Sanyal, Sunny, et al.
Published: (2025) -
Towards Quantifying the Preconditioning Effect of Adam
by: Das, Rudrajit, et al.
Published: (2024)