When Losses Align: Gradient-Based Composite Loss Weighting for Efficient Pretraining
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Karpukhin, Ivan, Savchenko, Andrey |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Detecting the Future: All-at-Once Event Sequence Forecasting with Horizon Matching
von: Karpukhin, Ivan, et al.
Veröffentlicht: (2024)
von: Karpukhin, Ivan, et al.
Veröffentlicht: (2024)
HT-Transformer: Event Sequences Classification by Accumulating Prefix Information with History Tokens
von: Karpukhin, Ivan, et al.
Veröffentlicht: (2025)
von: Karpukhin, Ivan, et al.
Veröffentlicht: (2025)
Multimodal Banking Dataset: Understanding Client Needs through Event Sequences
von: Mollaev, Dzhambulat, et al.
Veröffentlicht: (2024)
von: Mollaev, Dzhambulat, et al.
Veröffentlicht: (2024)
Gradient Aligned Regression via Pairwise Losses
von: Zhu, Dixian, et al.
Veröffentlicht: (2024)
von: Zhu, Dixian, et al.
Veröffentlicht: (2024)
HoTPP Benchmark: Are We Good at the Long Horizon Events Forecasting?
von: Karpukhin, Ivan, et al.
Veröffentlicht: (2024)
von: Karpukhin, Ivan, et al.
Veröffentlicht: (2024)
Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining
von: Sow, Daouda, et al.
Veröffentlicht: (2025)
von: Sow, Daouda, et al.
Veröffentlicht: (2025)
HN-MVTS: HyperNetwork-based Multivariate Time Series Forecasting
von: Savchenko, Andrey, et al.
Veröffentlicht: (2025)
von: Savchenko, Andrey, et al.
Veröffentlicht: (2025)
Sampling and Loss Weights in Multi-Domain Training
von: Salmani, Mahdi, et al.
Veröffentlicht: (2025)
von: Salmani, Mahdi, et al.
Veröffentlicht: (2025)
Adaptive Weighted Loss for Sequential Recommendations on Sparse Domains
von: Mittal, Akshay, et al.
Veröffentlicht: (2025)
von: Mittal, Akshay, et al.
Veröffentlicht: (2025)
Loss- and Reward-Weighting for Efficient Distributed Reinforcement Learning
von: Holen, Martin, et al.
Veröffentlicht: (2023)
von: Holen, Martin, et al.
Veröffentlicht: (2023)
Simplicial SMOTE: Oversampling Solution to the Imbalanced Learning Problem
von: Kachan, Oleg, et al.
Veröffentlicht: (2025)
von: Kachan, Oleg, et al.
Veröffentlicht: (2025)
Beyond Isolated Clients: Integrating Graph-Based Embeddings into Event Sequence Models
von: Proshian, Harry, et al.
Veröffentlicht: (2026)
von: Proshian, Harry, et al.
Veröffentlicht: (2026)
Mitigating Long-Tailed Anomaly Score Distributions with Importance-Weighted Loss
von: Lee, Jungi, et al.
Veröffentlicht: (2026)
von: Lee, Jungi, et al.
Veröffentlicht: (2026)
Latent Factorization of Tensors with Threshold Distance Weighted Loss for Traffic Data Estimation
von: Yang, Lei
Veröffentlicht: (2025)
von: Yang, Lei
Veröffentlicht: (2025)
Analytical Uncertainty-Based Loss Weighting in Multi-Task Learning
von: Kirchdorfer, Lukas, et al.
Veröffentlicht: (2024)
von: Kirchdorfer, Lukas, et al.
Veröffentlicht: (2024)
AdaKD: Dynamic Knowledge Distillation of ASR models using Adaptive Loss Weighting
von: Ganguly, Shreyan, et al.
Veröffentlicht: (2024)
von: Ganguly, Shreyan, et al.
Veröffentlicht: (2024)
CP Loss: Channel-wise Perceptual Loss for Time Series Forecasting
von: Zha, Yaohua, et al.
Veröffentlicht: (2026)
von: Zha, Yaohua, et al.
Veröffentlicht: (2026)
Elliptic Loss Regularization
von: Hasan, Ali, et al.
Veröffentlicht: (2025)
von: Hasan, Ali, et al.
Veröffentlicht: (2025)
A new membership inference attack that spots memorization in generative and predictive models: Loss-Based with Reference Model algorithm (LBRM)
von: Taleb, Faiz, et al.
Veröffentlicht: (2025)
von: Taleb, Faiz, et al.
Veröffentlicht: (2025)
Diversifying Deep Ensembles: A Saliency Map Approach for Enhanced OOD Detection, Calibration, and Accuracy
von: Dereka, Stanislav, et al.
Veröffentlicht: (2023)
von: Dereka, Stanislav, et al.
Veröffentlicht: (2023)
XGrad: Boosting Gradient-Based Optimizers With Weight Prediction
von: Guan, Lei, et al.
Veröffentlicht: (2023)
von: Guan, Lei, et al.
Veröffentlicht: (2023)
Investigating the Histogram Loss in Regression
von: Imani, Ehsan, et al.
Veröffentlicht: (2024)
von: Imani, Ehsan, et al.
Veröffentlicht: (2024)
There is a Singularity in the Loss Landscape
von: Lowell, Mark
Veröffentlicht: (2022)
von: Lowell, Mark
Veröffentlicht: (2022)
When to Stop Reusing: Dynamic Gradient Gating for Sample-Efficient RLVR
von: Miao, Yuchun, et al.
Veröffentlicht: (2026)
von: Miao, Yuchun, et al.
Veröffentlicht: (2026)
Meta-GCN: A Dynamically Weighted Loss Minimization Method for Dealing with the Data Imbalance in Graph Neural Networks
von: Mohammadizadeh, Mahdi, et al.
Veröffentlicht: (2024)
von: Mohammadizadeh, Mahdi, et al.
Veröffentlicht: (2024)
Gradient Weight-normalized Low-rank Projection for Efficient LLM Training
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024)
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024)
When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models
von: Xie, Tong, et al.
Veröffentlicht: (2025)
von: Xie, Tong, et al.
Veröffentlicht: (2025)
Loss-to-Loss Prediction: Scaling Laws for All Datasets
von: Brandfonbrener, David, et al.
Veröffentlicht: (2024)
von: Brandfonbrener, David, et al.
Veröffentlicht: (2024)
Weight-Entanglement Meets Gradient-Based Neural Architecture Search
von: Sukthanker, Rhea Sanjay, et al.
Veröffentlicht: (2023)
von: Sukthanker, Rhea Sanjay, et al.
Veröffentlicht: (2023)
Barycentric Neural Networks and Length-Weighted Persistent Entropy Loss: A Green Geometric and Topological Framework for Function Approximation
von: Toscano-Duran, Victor, et al.
Veröffentlicht: (2025)
von: Toscano-Duran, Victor, et al.
Veröffentlicht: (2025)
Premier-TACO is a Few-Shot Policy Learner: Pretraining Multitask Representation via Temporal Action-Driven Contrastive Loss
von: Zheng, Ruijie, et al.
Veröffentlicht: (2024)
von: Zheng, Ruijie, et al.
Veröffentlicht: (2024)
Rethinking Losses for Diffusion Bridge Samplers
von: Sanokowski, Sebastian, et al.
Veröffentlicht: (2025)
von: Sanokowski, Sebastian, et al.
Veröffentlicht: (2025)
Global Minimizers of Sigmoid Contrastive Loss
von: Bangachev, Kiril, et al.
Veröffentlicht: (2025)
von: Bangachev, Kiril, et al.
Veröffentlicht: (2025)
Deep Metric Loss for Multimodal Learning
von: Moon, Sehwan, et al.
Veröffentlicht: (2023)
von: Moon, Sehwan, et al.
Veröffentlicht: (2023)
Meta-Learning Adaptive Loss Functions
von: Raymond, Christian, et al.
Veröffentlicht: (2023)
von: Raymond, Christian, et al.
Veröffentlicht: (2023)
Neural Network Plasticity and Loss Sharpness
von: Koster, Max, et al.
Veröffentlicht: (2024)
von: Koster, Max, et al.
Veröffentlicht: (2024)
LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws
von: Mayilvahanan, Prasanna, et al.
Veröffentlicht: (2025)
von: Mayilvahanan, Prasanna, et al.
Veröffentlicht: (2025)
Fundamental Limits of Deep Learning-Based Binary Classifiers Trained with Hinge Loss
von: Getu, Tilahun M., et al.
Veröffentlicht: (2023)
von: Getu, Tilahun M., et al.
Veröffentlicht: (2023)
Holdout-Loss-Based Data Selection for LLM Finetuning via In-Context Learning
von: Zhang, Ling, et al.
Veröffentlicht: (2025)
von: Zhang, Ling, et al.
Veröffentlicht: (2025)
Decoupled Split Learning via Auxiliary Loss
von: Zihad, Anower, et al.
Veröffentlicht: (2026)
von: Zihad, Anower, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Detecting the Future: All-at-Once Event Sequence Forecasting with Horizon Matching
von: Karpukhin, Ivan, et al.
Veröffentlicht: (2024) -
HT-Transformer: Event Sequences Classification by Accumulating Prefix Information with History Tokens
von: Karpukhin, Ivan, et al.
Veröffentlicht: (2025) -
Multimodal Banking Dataset: Understanding Client Needs through Event Sequences
von: Mollaev, Dzhambulat, et al.
Veröffentlicht: (2024) -
Gradient Aligned Regression via Pairwise Losses
von: Zhu, Dixian, et al.
Veröffentlicht: (2024) -
HoTPP Benchmark: Are We Good at the Long Horizon Events Forecasting?
von: Karpukhin, Ivan, et al.
Veröffentlicht: (2024)