Iterative Label Refinement Matters More than Preference Optimization under Weak Supervision
Fuente:
arXiv
Guardado en:
| Autores principales: | Ye, Yaowen, Laidlaw, Cassidy, Steinhardt, Jacob |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Which Attention Heads Matter for In-Context Learning?
por: Yin, Kayo, et al.
Publicado: (2025)
por: Yin, Kayo, et al.
Publicado: (2025)
Discovering Latent Knowledge in Language Models Without Supervision
por: Burns, Collin, et al.
Publicado: (2022)
por: Burns, Collin, et al.
Publicado: (2022)
How do Language Models Bind Entities in Context?
por: Feng, Jiahai, et al.
Publicado: (2023)
por: Feng, Jiahai, et al.
Publicado: (2023)
Geometric-Averaged Preference Optimization for Soft Preference Labels
por: Furuta, Hiroki, et al.
Publicado: (2024)
por: Furuta, Hiroki, et al.
Publicado: (2024)
Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF
por: Siththaranjan, Anand, et al.
Publicado: (2023)
por: Siththaranjan, Anand, et al.
Publicado: (2023)
Overthinking the Truth: Understanding how Language Models Process False Demonstrations
por: Halawi, Danny, et al.
Publicado: (2023)
por: Halawi, Danny, et al.
Publicado: (2023)
Feedback Loops With Language Models Drive In-Context Reward Hacking
por: Pan, Alexander, et al.
Publicado: (2024)
por: Pan, Alexander, et al.
Publicado: (2024)
Explaining Datasets in Words: Statistical Models with Natural Language Parameters
por: Zhong, Ruiqi, et al.
Publicado: (2024)
por: Zhong, Ruiqi, et al.
Publicado: (2024)
Modulated Intervention Preference Optimization (MIPO): Keep the Easy, Refine the Difficult
por: Jang, Cheolhun
Publicado: (2024)
por: Jang, Cheolhun
Publicado: (2024)
Training Language Models to Explain Their Own Computations
por: Li, Belinda Z., et al.
Publicado: (2025)
por: Li, Belinda Z., et al.
Publicado: (2025)
Learning a Generative Meta-Model of LLM Activations
por: Luo, Grace, et al.
Publicado: (2026)
por: Luo, Grace, et al.
Publicado: (2026)
Evolutionary Guided Decoding: Iterative Value Refinement for LLMs
por: Liu, Zhenhua, et al.
Publicado: (2025)
por: Liu, Zhenhua, et al.
Publicado: (2025)
Understanding In-context Learning of Addition via Activation Subspaces
por: Hu, Xinyan, et al.
Publicado: (2025)
por: Hu, Xinyan, et al.
Publicado: (2025)
Predictive Concept Decoders: Training Scalable End-to-End Interpretability Assistants
por: Huang, Vincent, et al.
Publicado: (2025)
por: Huang, Vincent, et al.
Publicado: (2025)
Stable Preference Optimization: A Bilevel Approach to Catastrophic Preference Shift
por: Jian, Chengtao, et al.
Publicado: (2025)
por: Jian, Chengtao, et al.
Publicado: (2025)
On the Role of Preference Variance in Preference Optimization
por: Guo, Jiacheng, et al.
Publicado: (2025)
por: Guo, Jiacheng, et al.
Publicado: (2025)
Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
por: Zhang, Yuheng, et al.
Publicado: (2024)
por: Zhang, Yuheng, et al.
Publicado: (2024)
Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
por: Laidlaw, Cassidy, et al.
Publicado: (2024)
por: Laidlaw, Cassidy, et al.
Publicado: (2024)
De Jure: Iterative LLM Self-Refinement for Structured Extraction of Regulatory Rules
por: Guliani, Keerat, et al.
Publicado: (2026)
por: Guliani, Keerat, et al.
Publicado: (2026)
Approaching Human-Level Forecasting with Language Models
por: Halawi, Danny, et al.
Publicado: (2024)
por: Halawi, Danny, et al.
Publicado: (2024)
Less is More: Improving LLM Alignment via Preference Data Selection
por: Deng, Xun, et al.
Publicado: (2025)
por: Deng, Xun, et al.
Publicado: (2025)
ROPO: Robust Preference Optimization for Large Language Models
por: Liang, Xize, et al.
Publicado: (2024)
por: Liang, Xize, et al.
Publicado: (2024)
Beyond the Last Answer: Your Reasoning Trace Uncovers More than You Think
por: Hammoud, Hasan Abed Al Kader, et al.
Publicado: (2025)
por: Hammoud, Hasan Abed Al Kader, et al.
Publicado: (2025)
Iterative Length-Regularized Direct Preference Optimization: A Case Study on Improving 7B Language Models to GPT-4 Level
por: Liu, Jie, et al.
Publicado: (2024)
por: Liu, Jie, et al.
Publicado: (2024)
ToolACE-R: Model-aware Iterative Training and Adaptive Refinement for Tool Learning
por: Zeng, Xingshan, et al.
Publicado: (2025)
por: Zeng, Xingshan, et al.
Publicado: (2025)
Weakly Supervised Veracity Classification with LLM-Predicted Credibility Signals
por: Leite, João A., et al.
Publicado: (2023)
por: Leite, João A., et al.
Publicado: (2023)
Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement
por: Xiong, Weimin, et al.
Publicado: (2024)
por: Xiong, Weimin, et al.
Publicado: (2024)
Instances and Labels: Hierarchy-aware Joint Supervised Contrastive Learning for Hierarchical Multi-Label Text Classification
por: Yu, Simon, et al.
Publicado: (2023)
por: Yu, Simon, et al.
Publicado: (2023)
Filtered Direct Preference Optimization
por: Morimura, Tetsuro, et al.
Publicado: (2024)
por: Morimura, Tetsuro, et al.
Publicado: (2024)
Direct Preference Optimization with an Offset
por: Amini, Afra, et al.
Publicado: (2024)
por: Amini, Afra, et al.
Publicado: (2024)
Self-Consistency Preference Optimization
por: Prasad, Archiki, et al.
Publicado: (2024)
por: Prasad, Archiki, et al.
Publicado: (2024)
AMPO: Active Multi-Preference Optimization for Self-play Preference Selection
por: Gupta, Taneesh, et al.
Publicado: (2025)
por: Gupta, Taneesh, et al.
Publicado: (2025)
Eliciting Language Model Behaviors with Investigator Agents
por: Li, Xiang Lisa, et al.
Publicado: (2025)
por: Li, Xiang Lisa, et al.
Publicado: (2025)
Entropy Controllable Direct Preference Optimization
por: Omura, Motoki, et al.
Publicado: (2024)
por: Omura, Motoki, et al.
Publicado: (2024)
Orthogonal Finetuning for Direct Preference Optimization
por: Yang, Chenxu, et al.
Publicado: (2024)
por: Yang, Chenxu, et al.
Publicado: (2024)
How Do Latent Reasoning Methods Perform Under Weak and Strong Supervision?
por: Cui, Yingqian, et al.
Publicado: (2026)
por: Cui, Yingqian, et al.
Publicado: (2026)
Large Language Models Assume People are More Rational than We Really are
por: Liu, Ryan, et al.
Publicado: (2024)
por: Liu, Ryan, et al.
Publicado: (2024)
Self-Supervised Prompt Optimization
por: Xiang, Jinyu, et al.
Publicado: (2025)
por: Xiang, Jinyu, et al.
Publicado: (2025)
Packing Analysis: Packing Is More Appropriate for Large Models or Datasets in Supervised Fine-tuning
por: Wang, Shuhe, et al.
Publicado: (2024)
por: Wang, Shuhe, et al.
Publicado: (2024)
An Experimental Design Framework for Label-Efficient Supervised Finetuning of Large Language Models
por: Bhatt, Gantavya, et al.
Publicado: (2024)
por: Bhatt, Gantavya, et al.
Publicado: (2024)
Ejemplares similares
-
Which Attention Heads Matter for In-Context Learning?
por: Yin, Kayo, et al.
Publicado: (2025) -
Discovering Latent Knowledge in Language Models Without Supervision
por: Burns, Collin, et al.
Publicado: (2022) -
How do Language Models Bind Entities in Context?
por: Feng, Jiahai, et al.
Publicado: (2023) -
Geometric-Averaged Preference Optimization for Soft Preference Labels
por: Furuta, Hiroki, et al.
Publicado: (2024) -
Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF
por: Siththaranjan, Anand, et al.
Publicado: (2023)