Test-Time Alignment via Hypothesis Reweighting
Fuente:
arXiv
Guardado en:
| Autores principales: | Lee, Yoonho, Williams, Jonathan, Marklund, Henrik, Sharma, Archit, Mitchell, Eric, Singh, Anikait, Finn, Chelsea |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Feedback Descent: Open-Ended Text Optimization via Pairwise Comparison
por: Lee, Yoonho, et al.
Publicado: (2025)
por: Lee, Yoonho, et al.
Publicado: (2025)
FSPO: Few-Shot Optimization of Synthetic Preferences Personalizes to Real Users
por: Singh, Anikait, et al.
Publicado: (2025)
por: Singh, Anikait, et al.
Publicado: (2025)
RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems
por: Qu, Yuxiao, et al.
Publicado: (2025)
por: Qu, Yuxiao, et al.
Publicado: (2025)
Calibrating Language Models with Adaptive Temperature Scaling
por: Xie, Johnathan, et al.
Publicado: (2024)
por: Xie, Johnathan, et al.
Publicado: (2024)
Bidirectional Decoding: Improving Action Chunking via Guided Test-Time Sampling
por: Liu, Yuejiang, et al.
Publicado: (2024)
por: Liu, Yuejiang, et al.
Publicado: (2024)
Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data
por: Tajwar, Fahim, et al.
Publicado: (2024)
por: Tajwar, Fahim, et al.
Publicado: (2024)
Self-Guided Masked Autoencoders for Domain-Agnostic Self-Supervised Learning
por: Xie, Johnathan, et al.
Publicado: (2024)
por: Xie, Johnathan, et al.
Publicado: (2024)
Grounding by Trying: LLMs with Reinforcement Learning-Enhanced Retrieval
por: Hsu, Sheryl, et al.
Publicado: (2024)
por: Hsu, Sheryl, et al.
Publicado: (2024)
A Critical Evaluation of AI Feedback for Aligning Large Language Models
por: Sharma, Archit, et al.
Publicado: (2024)
por: Sharma, Archit, et al.
Publicado: (2024)
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
por: Rafailov, Rafael, et al.
Publicado: (2023)
por: Rafailov, Rafael, et al.
Publicado: (2023)
Conservative Prediction via Data-Driven Confidence Minimization
por: Choi, Caroline, et al.
Publicado: (2023)
por: Choi, Caroline, et al.
Publicado: (2023)
RLVF: Learning from Verbal Feedback without Overgeneralization
por: Stephan, Moritz, et al.
Publicado: (2024)
por: Stephan, Moritz, et al.
Publicado: (2024)
Clarify: Improving Model Robustness With Natural Language Corrections
por: Lee, Yoonho, et al.
Publicado: (2024)
por: Lee, Yoonho, et al.
Publicado: (2024)
AutoFT: Learning an Objective for Robust Fine-Tuning
por: Choi, Caroline, et al.
Publicado: (2024)
por: Choi, Caroline, et al.
Publicado: (2024)
Adapt On-the-Go: Behavior Modulation for Single-Life Robot Deployment
por: Chen, Annie S., et al.
Publicado: (2023)
por: Chen, Annie S., et al.
Publicado: (2023)
Choice Between Partial Trajectories: Disentangling Goals from Beliefs
por: Marklund, Henrik, et al.
Publicado: (2024)
por: Marklund, Henrik, et al.
Publicado: (2024)
Maintaining Plasticity in Continual Learning via Regenerative Regularization
por: Kumar, Saurabh, et al.
Publicado: (2023)
por: Kumar, Saurabh, et al.
Publicado: (2023)
Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-Tuning
por: Nakamoto, Mitsuhiko, et al.
Publicado: (2023)
por: Nakamoto, Mitsuhiko, et al.
Publicado: (2023)
Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation
por: Manvi, Rohin, et al.
Publicado: (2024)
por: Manvi, Rohin, et al.
Publicado: (2024)
MemER: Scaling Up Memory for Robot Control via Experience Retrieval
por: Sridhar, Ajay, et al.
Publicado: (2025)
por: Sridhar, Ajay, et al.
Publicado: (2025)
Consequentialist Objectives and Catastrophe
por: Marklund, Henrik, et al.
Publicado: (2026)
por: Marklund, Henrik, et al.
Publicado: (2026)
Misalignment from Treating Means as Ends
por: Marklund, Henrik, et al.
Publicado: (2025)
por: Marklund, Henrik, et al.
Publicado: (2025)
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
por: Mark, Max Sobol, et al.
Publicado: (2024)
por: Mark, Max Sobol, et al.
Publicado: (2024)
DAJ: Data-Reweighted LLM Judge for Test-Time Scaling in Code Generation
por: Qin, Peijia, et al.
Publicado: (2026)
por: Qin, Peijia, et al.
Publicado: (2026)
Yell At Your Robot: Improving On-the-Fly from Language Corrections
por: Shi, Lucy Xiaoyang, et al.
Publicado: (2024)
por: Shi, Lucy Xiaoyang, et al.
Publicado: (2024)
Hypothesis Testing for Generalized Thurstone Models
por: Makur, Anuran, et al.
Publicado: (2025)
por: Makur, Anuran, et al.
Publicado: (2025)
D5RL: Diverse Datasets for Data-Driven Deep Reinforcement Learning
por: Rafailov, Rafael, et al.
Publicado: (2024)
por: Rafailov, Rafael, et al.
Publicado: (2024)
Operationalising the Superficial Alignment Hypothesis via Task Complexity
por: Vergara-Browne, Tomás, et al.
Publicado: (2026)
por: Vergara-Browne, Tomás, et al.
Publicado: (2026)
Granular feedback merits sophisticated aggregation
por: Kagrecha, Anmol, et al.
Publicado: (2025)
por: Kagrecha, Anmol, et al.
Publicado: (2025)
AutoCompress: Critical Layer Isolation for Efficient Transformer Compression
por: Thorat, Archit
Publicado: (2026)
por: Thorat, Archit
Publicado: (2026)
Robust Reward Alignment via Hypothesis Space Batch Cutting
por: Xie, Zhixian, et al.
Publicado: (2025)
por: Xie, Zhixian, et al.
Publicado: (2025)
Affordance-Guided Reinforcement Learning via Visual Prompting
por: Lee, Olivia Y., et al.
Publicado: (2024)
por: Lee, Olivia Y., et al.
Publicado: (2024)
Strategic Hypothesis Testing
por: Hossain, Safwan, et al.
Publicado: (2025)
por: Hossain, Safwan, et al.
Publicado: (2025)
Reinforcement Learning via Implicit Imitation Guidance
por: Dong, Perry, et al.
Publicado: (2025)
por: Dong, Perry, et al.
Publicado: (2025)
Universal Neural Functionals
por: Zhou, Allan, et al.
Publicado: (2024)
por: Zhou, Allan, et al.
Publicado: (2024)
Learning Long-Context Diffusion Policies via Past-Token Prediction
por: Torne, Marcel, et al.
Publicado: (2025)
por: Torne, Marcel, et al.
Publicado: (2025)
Hypothesis Testing the Circuit Hypothesis in LLMs
por: Shi, Claudia, et al.
Publicado: (2024)
por: Shi, Claudia, et al.
Publicado: (2024)
Minimax Hypothesis Testing for the Bradley-Terry-Luce Model
por: Makur, Anuran, et al.
Publicado: (2024)
por: Makur, Anuran, et al.
Publicado: (2024)
The Master Key Hypothesis: Unlocking Cross-Model Capability Transfer via Linear Subspace Alignment
por: Balasubramanian, Rishab, et al.
Publicado: (2026)
por: Balasubramanian, Rishab, et al.
Publicado: (2026)
From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function
por: Rafailov, Rafael, et al.
Publicado: (2024)
por: Rafailov, Rafael, et al.
Publicado: (2024)
Ejemplares similares
-
Feedback Descent: Open-Ended Text Optimization via Pairwise Comparison
por: Lee, Yoonho, et al.
Publicado: (2025) -
FSPO: Few-Shot Optimization of Synthetic Preferences Personalizes to Real Users
por: Singh, Anikait, et al.
Publicado: (2025) -
RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems
por: Qu, Yuxiao, et al.
Publicado: (2025) -
Calibrating Language Models with Adaptive Temperature Scaling
por: Xie, Johnathan, et al.
Publicado: (2024) -
Bidirectional Decoding: Improving Action Chunking via Guided Test-Time Sampling
por: Liu, Yuejiang, et al.
Publicado: (2024)