MiLe Loss: a New Entropy-Weighed Loss for Mitigating the Bias of Learning Difficulties in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Su, Zhenpeng, Wu, Xing, Bai, Xue, Lin, Zijia, Chen, Hui, Ding, Guiguang, Zhou, Wei, Hu, Songlin |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MaskMoE: Boosting Token-Level Learning via Routing Mask in Mixture-of-Experts
by: Su, Zhenpeng, et al.
Published: (2024)
by: Su, Zhenpeng, et al.
Published: (2024)
CartesianMoE: Boosting Knowledge Sharing among Experts via Cartesian Product Routing in Mixture-of-Experts
by: Su, Zhenpeng, et al.
Published: (2024)
by: Su, Zhenpeng, et al.
Published: (2024)
Temporal Scaling Law for Large Language Models
by: Xiong, Yizhe, et al.
Published: (2024)
by: Xiong, Yizhe, et al.
Published: (2024)
Scaffold-BPE: Enhancing Byte Pair Encoding for Large Language Models with Simple and Effective Scaffold Token Removal
by: Lian, Haoran, et al.
Published: (2024)
by: Lian, Haoran, et al.
Published: (2024)
LSNet: See Large, Focus Small
by: Wang, Ao, et al.
Published: (2025)
by: Wang, Ao, et al.
Published: (2025)
Task-level Distributionally Robust Optimization for Large Language Model-based Dense Retrieval
by: Ma, Guangyuan, et al.
Published: (2024)
by: Ma, Guangyuan, et al.
Published: (2024)
DSMoE: Matrix-Partitioned Experts with Dynamic Routing for Computation-Efficient Dense LLMs
by: Lv, Minxuan, et al.
Published: (2025)
by: Lv, Minxuan, et al.
Published: (2025)
UniAttn: Reducing Inference Costs via Softmax Unification for Post-Training LLMs
by: Xiong, Yizhe, et al.
Published: (2025)
by: Xiong, Yizhe, et al.
Published: (2025)
RepViT-SAM: Towards Real-Time Segmenting Anything
by: Wang, Ao, et al.
Published: (2023)
by: Wang, Ao, et al.
Published: (2023)
RepViT: Revisiting Mobile CNN From ViT Perspective
by: Wang, Ao, et al.
Published: (2023)
by: Wang, Ao, et al.
Published: (2023)
Dial-MAE: ConTextual Masked Auto-Encoder for Retrieval-based Dialogue Systems
by: Su, Zhenpeng, et al.
Published: (2023)
by: Su, Zhenpeng, et al.
Published: (2023)
HC3 Plus: A Semantic-Invariant Human ChatGPT Comparison Corpus
by: Su, Zhenpeng, et al.
Published: (2023)
by: Su, Zhenpeng, et al.
Published: (2023)
Finedeep: Mitigating Sparse Activation in Dense LLMs via Multi-Layer Fine-Grained Experts
by: Pan, Leiyu, et al.
Published: (2025)
by: Pan, Leiyu, et al.
Published: (2025)
LBPE: Long-token-first Tokenization to Improve Large Language Models
by: Lian, Haoran, et al.
Published: (2024)
by: Lian, Haoran, et al.
Published: (2024)
LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context Instructions
by: Gao, Chaochen, et al.
Published: (2025)
by: Gao, Chaochen, et al.
Published: (2025)
Drop your Decoder: Pre-training with Bag-of-Word Prediction for Dense Passage Retrieval
by: Ma, Guangyuan, et al.
Published: (2024)
by: Ma, Guangyuan, et al.
Published: (2024)
CAIT: Triple-Win Compression towards High Accuracy, Fast Inference, and Favorable Transferability For ViTs
by: Wang, Ao, et al.
Published: (2023)
by: Wang, Ao, et al.
Published: (2023)
[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs
by: Wang, Ao, et al.
Published: (2024)
by: Wang, Ao, et al.
Published: (2024)
YOLOE: Real-Time Seeing Anything
by: Wang, Ao, et al.
Published: (2025)
by: Wang, Ao, et al.
Published: (2025)
Entropy Gain and Information Loss by Measurements
by: Wang, Xing M.
Published: (2019)
by: Wang, Xing M.
Published: (2019)
EntropyLong: Effective Long-Context Training via Predictive Uncertainty
by: Jia, Junlong, et al.
Published: (2025)
by: Jia, Junlong, et al.
Published: (2025)
NExtLong: Toward Effective Long-Context Training without Long Documents
by: Gao, Chaochen, et al.
Published: (2025)
by: Gao, Chaochen, et al.
Published: (2025)
Epistemic Uncertainty-Weighted Loss for Visual Bias Mitigation
by: Stone, Rebecca S, et al.
Published: (2022)
by: Stone, Rebecca S, et al.
Published: (2022)
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning
by: Su, Zhenpeng, et al.
Published: (2025)
by: Su, Zhenpeng, et al.
Published: (2025)
Breaking the Stage Barrier: A Novel Single-Stage Approach to Long Context Extension for Large Language Models
by: Lian, Haoran, et al.
Published: (2024)
by: Lian, Haoran, et al.
Published: (2024)
Fast Quiet-STaR: Thinking Without Thought Tokens
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
Learn from the Learnt: Source-Free Active Domain Adaptation via Contrastive Sampling and Visual Persistence
by: Lyu, Mengyao, et al.
Published: (2024)
by: Lyu, Mengyao, et al.
Published: (2024)
YOLOv10: Real-Time End-to-End Object Detection
by: Wang, Ao, et al.
Published: (2024)
by: Wang, Ao, et al.
Published: (2024)
Mitigating the Bias of Large Language Model Evaluation
by: Zhou, Hongli, et al.
Published: (2024)
by: Zhou, Hongli, et al.
Published: (2024)
Multi-Task Learning Using Uncertainty to Weigh Losses for Heterogeneous Face Attribute Estimation
by: Yuan, Huaqing, et al.
Published: (2024)
by: Yuan, Huaqing, et al.
Published: (2024)
Advancing Reliable Test-Time Adaptation of Vision-Language Models under Visual Variations
by: Liang, Yiwen, et al.
Published: (2025)
by: Liang, Yiwen, et al.
Published: (2025)
Mitigating Bias in Facial Recognition Systems: Centroid Fairness Loss Optimization
by: Conti, Jean-Rémy, et al.
Published: (2025)
by: Conti, Jean-Rémy, et al.
Published: (2025)
LightRetriever: A LLM-based Text Retrieval Architecture with Extremely Faster Query Inference
by: Ma, Guangyuan, et al.
Published: (2025)
by: Ma, Guangyuan, et al.
Published: (2025)
Context Enhancement with Reconstruction as Sequence for Unified Unsupervised Anomaly Detection
by: Yang, Hui-Yue, et al.
Published: (2024)
by: Yang, Hui-Yue, et al.
Published: (2024)
Initialization of Large Language Models via Reparameterization to Mitigate Loss Spikes
by: Nishida, Kosuke, et al.
Published: (2024)
by: Nishida, Kosuke, et al.
Published: (2024)
Bias in Large Language Models: Origin, Evaluation, and Mitigation
by: Guo, Yufei, et al.
Published: (2024)
by: Guo, Yufei, et al.
Published: (2024)
The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking
by: Miao, Yuchun, et al.
Published: (2025)
by: Miao, Yuchun, et al.
Published: (2025)
Structure-aware Propagation Generation with Large Language Models for Fake News Detection
by: Chen, Mengyang, et al.
Published: (2025)
by: Chen, Mengyang, et al.
Published: (2025)
Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models
by: Fu, Mingyu, et al.
Published: (2025)
by: Fu, Mingyu, et al.
Published: (2025)
Neutralizing Token Aggregation via Information Augmentation for Efficient Test-Time Adaptation
by: Xiong, Yizhe, et al.
Published: (2025)
by: Xiong, Yizhe, et al.
Published: (2025)
Similar Items
-
MaskMoE: Boosting Token-Level Learning via Routing Mask in Mixture-of-Experts
by: Su, Zhenpeng, et al.
Published: (2024) -
CartesianMoE: Boosting Knowledge Sharing among Experts via Cartesian Product Routing in Mixture-of-Experts
by: Su, Zhenpeng, et al.
Published: (2024) -
Temporal Scaling Law for Large Language Models
by: Xiong, Yizhe, et al.
Published: (2024) -
Scaffold-BPE: Enhancing Byte Pair Encoding for Large Language Models with Simple and Effective Scaffold Token Removal
by: Lian, Haoran, et al.
Published: (2024) -
LSNet: See Large, Focus Small
by: Wang, Ao, et al.
Published: (2025)