Gespeichert in:
| Hauptverfasser: | Li, Tianjian, Xu, Haoran, Tan, Weiting, Murray, Kenton, Khashabi, Daniel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2410.04579 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Error Norm Truncation: Robust Training in the Presence of Data Noise for Text Generation Models
von: Li, Tianjian, et al.
Veröffentlicht: (2023)
von: Li, Tianjian, et al.
Veröffentlicht: (2023)
SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning
von: Li, Tianjian, et al.
Veröffentlicht: (2025)
von: Li, Tianjian, et al.
Veröffentlicht: (2025)
Jointly Reinforcing Diversity and Quality in Language Model Generations
von: Li, Tianjian, et al.
Veröffentlicht: (2025)
von: Li, Tianjian, et al.
Veröffentlicht: (2025)
The Translation Barrier Hypothesis: Multilingual Generation with Large Language Models Suffers from Implicit Translation Failure
von: Bafna, Niyati, et al.
Veröffentlicht: (2025)
von: Bafna, Niyati, et al.
Veröffentlicht: (2025)
DOTResize: Reducing LLM Width via Discrete Optimal Transport-based Neuron Merging
von: Verma, Neha, et al.
Veröffentlicht: (2025)
von: Verma, Neha, et al.
Veröffentlicht: (2025)
Merging Feed-Forward Sublayers for Compressed Transformers
von: Verma, Neha, et al.
Veröffentlicht: (2025)
von: Verma, Neha, et al.
Veröffentlicht: (2025)
Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher
von: Uzunoglu, Arda, et al.
Veröffentlicht: (2026)
von: Uzunoglu, Arda, et al.
Veröffentlicht: (2026)
The Flaw of Averages: Quantifying Uniformity of Performance on Benchmarks
von: Uzunoglu, Arda, et al.
Veröffentlicht: (2025)
von: Uzunoglu, Arda, et al.
Veröffentlicht: (2025)
IA2: Alignment with ICL Activations Improves Supervised Fine-Tuning
von: Mishra, Aayush, et al.
Veröffentlicht: (2025)
von: Mishra, Aayush, et al.
Veröffentlicht: (2025)
Do pretrained Transformers Learn In-Context by Gradient Descent?
von: Shen, Lingfeng, et al.
Veröffentlicht: (2023)
von: Shen, Lingfeng, et al.
Veröffentlicht: (2023)
Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data
von: Zhang, Jingyu, et al.
Veröffentlicht: (2024)
von: Zhang, Jingyu, et al.
Veröffentlicht: (2024)
SELAUR: Self Evolving LLM Agent via Uncertainty-aware Rewards
von: Zhang, Dengjia, et al.
Veröffentlicht: (2026)
von: Zhang, Dengjia, et al.
Veröffentlicht: (2026)
DISCO Balances the Scales: Adaptive Domain- and Difficulty-Aware Reinforcement Learning on Imbalanced Data
von: Zhou, Yuhang, et al.
Veröffentlicht: (2025)
von: Zhou, Yuhang, et al.
Veröffentlicht: (2025)
Position: The Turing-Completeness of Autoregressive Transformers Relies Heavily on Context Management
von: Cui, Guanyu, et al.
Veröffentlicht: (2026)
von: Cui, Guanyu, et al.
Veröffentlicht: (2026)
DiffNorm: Self-Supervised Normalization for Non-autoregressive Speech-to-speech Translation
von: Tan, Weiting, et al.
Veröffentlicht: (2024)
von: Tan, Weiting, et al.
Veröffentlicht: (2024)
AnchorAL: Computationally Efficient Active Learning for Large and Imbalanced Datasets
von: Lesci, Pietro, et al.
Veröffentlicht: (2024)
von: Lesci, Pietro, et al.
Veröffentlicht: (2024)
Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation
von: Xu, Haoran, et al.
Veröffentlicht: (2024)
von: Xu, Haoran, et al.
Veröffentlicht: (2024)
k-SemStamp: A Clustering-Based Semantic Watermark for Detection of Machine-Generated Text
von: Hou, Abe Bohan, et al.
Veröffentlicht: (2024)
von: Hou, Abe Bohan, et al.
Veröffentlicht: (2024)
Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations
von: Liu, Wei, et al.
Veröffentlicht: (2026)
von: Liu, Wei, et al.
Veröffentlicht: (2026)
It Takes Two: On the Seamlessness between Reward and Policy Model in RLHF
von: Lu, Taiming, et al.
Veröffentlicht: (2024)
von: Lu, Taiming, et al.
Veröffentlicht: (2024)
SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses
von: Jiang, Dongwei, et al.
Veröffentlicht: (2024)
von: Jiang, Dongwei, et al.
Veröffentlicht: (2024)
The Language Barrier: Dissecting Safety Challenges of LLMs in Multilingual Contexts
von: Shen, Lingfeng, et al.
Veröffentlicht: (2024)
von: Shen, Lingfeng, et al.
Veröffentlicht: (2024)
Mitigating Training Imbalance in LLM Fine-Tuning via Selective Parameter Merging
von: Ju, Yiming, et al.
Veröffentlicht: (2024)
von: Ju, Yiming, et al.
Veröffentlicht: (2024)
SABER: Switchable and Balanced Training for Efficient LLM Reasoning
von: Zhao, Kai, et al.
Veröffentlicht: (2025)
von: Zhao, Kai, et al.
Veröffentlicht: (2025)
RedPajama: an Open Dataset for Training Large Language Models
von: Weber, Maurice, et al.
Veröffentlicht: (2024)
von: Weber, Maurice, et al.
Veröffentlicht: (2024)
Training Superior Sparse Autoencoders for Instruct Models
von: Li, Jiaming, et al.
Veröffentlicht: (2025)
von: Li, Jiaming, et al.
Veröffentlicht: (2025)
Training a Huggingface Model on AWS Sagemaker (Without Tears)
von: Tan, Liling
Veröffentlicht: (2025)
von: Tan, Liling
Veröffentlicht: (2025)
Effective Reasoning Chains Reduce Intrinsic Dimensionality
von: Prasad, Archiki, et al.
Veröffentlicht: (2026)
von: Prasad, Archiki, et al.
Veröffentlicht: (2026)
Exploring Imbalanced Annotations for Effective In-Context Learning
von: Gao, Hongfu, et al.
Veröffentlicht: (2025)
von: Gao, Hongfu, et al.
Veröffentlicht: (2025)
Dataset Decomposition: Faster LLM Training with Variable Sequence Length Curriculum
von: Pouransari, Hadi, et al.
Veröffentlicht: (2024)
von: Pouransari, Hadi, et al.
Veröffentlicht: (2024)
Beyond Mode-Seeking RL: Trajectory-Balance Post-Training for Diffusion Language Models
von: Ahmadi, Saba, et al.
Veröffentlicht: (2026)
von: Ahmadi, Saba, et al.
Veröffentlicht: (2026)
Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models
von: Qiu, Zihan, et al.
Veröffentlicht: (2025)
von: Qiu, Zihan, et al.
Veröffentlicht: (2025)
Improving Diffusion Language Model Decoding through Joint Search in Generation Order and Token Space
von: Shen, Yangyi, et al.
Veröffentlicht: (2026)
von: Shen, Yangyi, et al.
Veröffentlicht: (2026)
MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training
von: Li, Jiacheng, et al.
Veröffentlicht: (2026)
von: Li, Jiacheng, et al.
Veröffentlicht: (2026)
LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs
von: Xia, Yunhui, et al.
Veröffentlicht: (2025)
von: Xia, Yunhui, et al.
Veröffentlicht: (2025)
Online Bayesian Imbalanced Learning with Bregman-Calibrated Deep Networks
von: Alsulaimawi, Zahir
Veröffentlicht: (2026)
von: Alsulaimawi, Zahir
Veröffentlicht: (2026)
Exploiting Vocabulary Frequency Imbalance in Language Model Pre-training
von: Chung, Woojin, et al.
Veröffentlicht: (2025)
von: Chung, Woojin, et al.
Veröffentlicht: (2025)
Data Augmentation for Classification of Negative Pregnancy Outcomes in Imbalanced Data
von: Biswas, Md Badsha
Veröffentlicht: (2025)
von: Biswas, Md Badsha
Veröffentlicht: (2025)
ALTA: Compiler-Based Analysis of Transformers
von: Shaw, Peter, et al.
Veröffentlicht: (2024)
von: Shaw, Peter, et al.
Veröffentlicht: (2024)
BigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks
von: Rodriguez, Juan, et al.
Veröffentlicht: (2024)
von: Rodriguez, Juan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Error Norm Truncation: Robust Training in the Presence of Data Noise for Text Generation Models
von: Li, Tianjian, et al.
Veröffentlicht: (2023) -
SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning
von: Li, Tianjian, et al.
Veröffentlicht: (2025) -
Jointly Reinforcing Diversity and Quality in Language Model Generations
von: Li, Tianjian, et al.
Veröffentlicht: (2025) -
The Translation Barrier Hypothesis: Multilingual Generation with Large Language Models Suffers from Implicit Translation Failure
von: Bafna, Niyati, et al.
Veröffentlicht: (2025) -
DOTResize: Reducing LLM Width via Discrete Optimal Transport-based Neuron Merging
von: Verma, Neha, et al.
Veröffentlicht: (2025)