Balancing the Budget: Understanding Trade-offs Between Supervised and Preference-Based Finetuning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Raghavendra, Mohit, Kang, Junmo, Ritter, Alan |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMs
par: Park, Jungsoo, et autres
Publié: (2025)
par: Park, Jungsoo, et autres
Publié: (2025)
No-Regret Learning in Bilateral Trade via Global Budget Balance
par: Bernasconi, Martino, et autres
Publié: (2023)
par: Bernasconi, Martino, et autres
Publié: (2023)
ParallelBench: Understanding the Trade-offs of Parallel Decoding in Diffusion LLMs
par: Kang, Wonjun, et autres
Publié: (2025)
par: Kang, Wonjun, et autres
Publié: (2025)
Architecture Selection via the Trade-off Between Accuracy and Robustness
par: Deng, Zhun, et autres
Publié: (2019)
par: Deng, Zhun, et autres
Publié: (2019)
Meta-Tuning LLMs to Leverage Lexical Knowledge for Generalizable Language Style Understanding
par: Guo, Ruohao, et autres
Publié: (2023)
par: Guo, Ruohao, et autres
Publié: (2023)
Agentic Rubrics as Contextual Verifiers for SWE Agents
par: Raghavendra, Mohit, et autres
Publié: (2026)
par: Raghavendra, Mohit, et autres
Publié: (2026)
Tabular Data Understanding with LLMs: A Survey of Recent Advances and Challenges
par: Wu, Xiaofeng, et autres
Publié: (2025)
par: Wu, Xiaofeng, et autres
Publié: (2025)
Orthogonal Finetuning for Direct Preference Optimization
par: Yang, Chenxu, et autres
Publié: (2024)
par: Yang, Chenxu, et autres
Publié: (2024)
Revisiting the Superficial Alignment Hypothesis
par: Raghavendra, Mohit, et autres
Publié: (2024)
par: Raghavendra, Mohit, et autres
Publié: (2024)
Logits-Based Finetuning
par: Li, Jingyao, et autres
Publié: (2025)
par: Li, Jingyao, et autres
Publié: (2025)
Balancing Speed and Stability: The Trade-offs of FP8 vs. BF16 Training in LLMs
par: Fujii, Kazuki, et autres
Publié: (2024)
par: Fujii, Kazuki, et autres
Publié: (2024)
Understanding Finetuning for Factual Knowledge Extraction
par: Ghosal, Gaurav, et autres
Publié: (2024)
par: Ghosal, Gaurav, et autres
Publié: (2024)
To Distill or Decide? Understanding the Algorithmic Trade-off in Partially Observable Reinforcement Learning
par: Song, Yuda, et autres
Publié: (2025)
par: Song, Yuda, et autres
Publié: (2025)
Understanding Generalization of Federated Learning: the Trade-off between Model Stability and Optimization
par: Zeng, Dun, et autres
Publié: (2024)
par: Zeng, Dun, et autres
Publié: (2024)
Self-MoE: Towards Compositional Large Language Models with Self-Specialized Experts
par: Kang, Junmo, et autres
Publié: (2024)
par: Kang, Junmo, et autres
Publié: (2024)
Do Heavy Tails Help Diffusion? On the Subtle Trade-off Between Initialization and Training
par: Cherkaoui, Hamza, et autres
Publié: (2026)
par: Cherkaoui, Hamza, et autres
Publié: (2026)
Connecting Thompson Sampling and UCB: Towards More Efficient Trade-offs Between Privacy and Regret
par: Hu, Bingshan, et autres
Publié: (2025)
par: Hu, Bingshan, et autres
Publié: (2025)
PEFT-Arena: Understanding Parameter-Efficient Finetuning from a Stability-Plasticity Perspective
par: Huang, Yangyi, et autres
Publié: (2026)
par: Huang, Yangyi, et autres
Publié: (2026)
Pareto Continual Learning: Preference-Conditioned Learning and Adaption for Dynamic Stability-Plasticity Trade-off
par: Lai, Song, et autres
Publié: (2025)
par: Lai, Song, et autres
Publié: (2025)
Differential Privacy for Anomaly Detection: Analyzing the Trade-off Between Privacy and Explainability
par: Ezzeddine, Fatima, et autres
Publié: (2024)
par: Ezzeddine, Fatima, et autres
Publié: (2024)
A Discrepancy-Based Perspective on Dataset Condensation
par: Chen, Tong, et autres
Publié: (2025)
par: Chen, Tong, et autres
Publié: (2025)
Understanding the Quality-Diversity Trade-off in Diffusion Language Models
par: Buzzard, Zak
Publié: (2025)
par: Buzzard, Zak
Publié: (2025)
Stable-GFlowNet: Toward Diverse and Robust LLM Red-Teaming via Contrastive Trajectory Balance
par: Kwon, Minchan, et autres
Publié: (2026)
par: Kwon, Minchan, et autres
Publié: (2026)
CoBa: Convergence Balancer for Multitask Finetuning of Large Language Models
par: Gong, Zi, et autres
Publié: (2024)
par: Gong, Zi, et autres
Publié: (2024)
Trading Convergence Rate with Computational Budget in High Dimensional Bayesian Optimization
par: Tran-The, Hung, et autres
Publié: (2019)
par: Tran-The, Hung, et autres
Publié: (2019)
Understanding and Improving Noisy Embedding Techniques in Instruction Finetuning
par: Yadav, Abhay
Publié: (2026)
par: Yadav, Abhay
Publié: (2026)
Trade-offs Between Individual and Group Fairness in Machine Learning: A Comprehensive Review
par: Benítez-Peña, Sandra, et autres
Publié: (2026)
par: Benítez-Peña, Sandra, et autres
Publié: (2026)
Improving the Trade-off Between Watermark Strength and Speculative Sampling Efficiency for Language Models
par: He, Weiqing, et autres
Publié: (2026)
par: He, Weiqing, et autres
Publié: (2026)
Splats under Pressure: Exploring Performance-Energy Trade-offs in Real-Time 3D Gaussian Splatting under Constrained GPU Budgets
par: Tajwar, Muhammad Fahim, et autres
Publié: (2026)
par: Tajwar, Muhammad Fahim, et autres
Publié: (2026)
Mitigating Accuracy-Robustness Trade-off via Balanced Multi-Teacher Adversarial Distillation
par: Zhao, Shiji, et autres
Publié: (2023)
par: Zhao, Shiji, et autres
Publié: (2023)
Understanding the Trade-offs in Accuracy and Uncertainty Quantification: Architecture and Inference Choices in Bayesian Neural Networks
par: Sheinkman, Alisa, et autres
Publié: (2025)
par: Sheinkman, Alisa, et autres
Publié: (2025)
The Finetuner's Fallacy: When to Pretrain with Your Finetuning Data
par: Baek, Christina, et autres
Publié: (2026)
par: Baek, Christina, et autres
Publié: (2026)
DARA: Few-shot Budget Allocation in Online Advertising via In-Context Decision Making with RL-Finetuned LLMs
par: Song, Mingxuan, et autres
Publié: (2026)
par: Song, Mingxuan, et autres
Publié: (2026)
CATTO: Balancing Preferences and Confidence in Language Models
par: Parikh, Nisarg, et autres
Publié: (2026)
par: Parikh, Nisarg, et autres
Publié: (2026)
Reconstruction of Personally Identifiable Information from Supervised Finetuned Models
par: Furukawa, Sae, et autres
Publié: (2026)
par: Furukawa, Sae, et autres
Publié: (2026)
Language Models can Self-Improve at State-Value Estimation for Better Search
par: Mendes, Ethan, et autres
Publié: (2025)
par: Mendes, Ethan, et autres
Publié: (2025)
Transparent Trade-offs between Properties of Explanations
par: Tadesse, Hiwot Belay, et autres
Publié: (2024)
par: Tadesse, Hiwot Belay, et autres
Publié: (2024)
Reasoning-Finetuning Repurposes Latent Representations in Base Models
par: Ward, Jake, et autres
Publié: (2025)
par: Ward, Jake, et autres
Publié: (2025)
Towards Understanding Dual BN In Hybrid Adversarial Training
par: Zhang, Chenshuang, et autres
Publié: (2024)
par: Zhang, Chenshuang, et autres
Publié: (2024)
Emissions and Performance Trade-off Between Small and Large Language Models
par: Garg, Anandita, et autres
Publié: (2025)
par: Garg, Anandita, et autres
Publié: (2025)
Documents similaires
-
Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMs
par: Park, Jungsoo, et autres
Publié: (2025) -
No-Regret Learning in Bilateral Trade via Global Budget Balance
par: Bernasconi, Martino, et autres
Publié: (2023) -
ParallelBench: Understanding the Trade-offs of Parallel Decoding in Diffusion LLMs
par: Kang, Wonjun, et autres
Publié: (2025) -
Architecture Selection via the Trade-off Between Accuracy and Robustness
par: Deng, Zhun, et autres
Publié: (2019) -
Meta-Tuning LLMs to Leverage Lexical Knowledge for Generalizable Language Style Understanding
par: Guo, Ruohao, et autres
Publié: (2023)