Test-Time Alignment via Hypothesis Reweighting
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Yoonho, Williams, Jonathan, Marklund, Henrik, Sharma, Archit, Mitchell, Eric, Singh, Anikait, Finn, Chelsea |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Feedback Descent: Open-Ended Text Optimization via Pairwise Comparison
by: Lee, Yoonho, et al.
Published: (2025)
by: Lee, Yoonho, et al.
Published: (2025)
FSPO: Few-Shot Optimization of Synthetic Preferences Personalizes to Real Users
by: Singh, Anikait, et al.
Published: (2025)
by: Singh, Anikait, et al.
Published: (2025)
RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems
by: Qu, Yuxiao, et al.
Published: (2025)
by: Qu, Yuxiao, et al.
Published: (2025)
Calibrating Language Models with Adaptive Temperature Scaling
by: Xie, Johnathan, et al.
Published: (2024)
by: Xie, Johnathan, et al.
Published: (2024)
Bidirectional Decoding: Improving Action Chunking via Guided Test-Time Sampling
by: Liu, Yuejiang, et al.
Published: (2024)
by: Liu, Yuejiang, et al.
Published: (2024)
Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data
by: Tajwar, Fahim, et al.
Published: (2024)
by: Tajwar, Fahim, et al.
Published: (2024)
Self-Guided Masked Autoencoders for Domain-Agnostic Self-Supervised Learning
by: Xie, Johnathan, et al.
Published: (2024)
by: Xie, Johnathan, et al.
Published: (2024)
Grounding by Trying: LLMs with Reinforcement Learning-Enhanced Retrieval
by: Hsu, Sheryl, et al.
Published: (2024)
by: Hsu, Sheryl, et al.
Published: (2024)
A Critical Evaluation of AI Feedback for Aligning Large Language Models
by: Sharma, Archit, et al.
Published: (2024)
by: Sharma, Archit, et al.
Published: (2024)
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
by: Rafailov, Rafael, et al.
Published: (2023)
by: Rafailov, Rafael, et al.
Published: (2023)
Conservative Prediction via Data-Driven Confidence Minimization
by: Choi, Caroline, et al.
Published: (2023)
by: Choi, Caroline, et al.
Published: (2023)
RLVF: Learning from Verbal Feedback without Overgeneralization
by: Stephan, Moritz, et al.
Published: (2024)
by: Stephan, Moritz, et al.
Published: (2024)
Clarify: Improving Model Robustness With Natural Language Corrections
by: Lee, Yoonho, et al.
Published: (2024)
by: Lee, Yoonho, et al.
Published: (2024)
AutoFT: Learning an Objective for Robust Fine-Tuning
by: Choi, Caroline, et al.
Published: (2024)
by: Choi, Caroline, et al.
Published: (2024)
Adapt On-the-Go: Behavior Modulation for Single-Life Robot Deployment
by: Chen, Annie S., et al.
Published: (2023)
by: Chen, Annie S., et al.
Published: (2023)
Choice Between Partial Trajectories: Disentangling Goals from Beliefs
by: Marklund, Henrik, et al.
Published: (2024)
by: Marklund, Henrik, et al.
Published: (2024)
Maintaining Plasticity in Continual Learning via Regenerative Regularization
by: Kumar, Saurabh, et al.
Published: (2023)
by: Kumar, Saurabh, et al.
Published: (2023)
Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-Tuning
by: Nakamoto, Mitsuhiko, et al.
Published: (2023)
by: Nakamoto, Mitsuhiko, et al.
Published: (2023)
Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation
by: Manvi, Rohin, et al.
Published: (2024)
by: Manvi, Rohin, et al.
Published: (2024)
MemER: Scaling Up Memory for Robot Control via Experience Retrieval
by: Sridhar, Ajay, et al.
Published: (2025)
by: Sridhar, Ajay, et al.
Published: (2025)
Consequentialist Objectives and Catastrophe
by: Marklund, Henrik, et al.
Published: (2026)
by: Marklund, Henrik, et al.
Published: (2026)
Misalignment from Treating Means as Ends
by: Marklund, Henrik, et al.
Published: (2025)
by: Marklund, Henrik, et al.
Published: (2025)
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
by: Mark, Max Sobol, et al.
Published: (2024)
by: Mark, Max Sobol, et al.
Published: (2024)
DAJ: Data-Reweighted LLM Judge for Test-Time Scaling in Code Generation
by: Qin, Peijia, et al.
Published: (2026)
by: Qin, Peijia, et al.
Published: (2026)
Yell At Your Robot: Improving On-the-Fly from Language Corrections
by: Shi, Lucy Xiaoyang, et al.
Published: (2024)
by: Shi, Lucy Xiaoyang, et al.
Published: (2024)
Hypothesis Testing for Generalized Thurstone Models
by: Makur, Anuran, et al.
Published: (2025)
by: Makur, Anuran, et al.
Published: (2025)
D5RL: Diverse Datasets for Data-Driven Deep Reinforcement Learning
by: Rafailov, Rafael, et al.
Published: (2024)
by: Rafailov, Rafael, et al.
Published: (2024)
Operationalising the Superficial Alignment Hypothesis via Task Complexity
by: Vergara-Browne, Tomás, et al.
Published: (2026)
by: Vergara-Browne, Tomás, et al.
Published: (2026)
Granular feedback merits sophisticated aggregation
by: Kagrecha, Anmol, et al.
Published: (2025)
by: Kagrecha, Anmol, et al.
Published: (2025)
AutoCompress: Critical Layer Isolation for Efficient Transformer Compression
by: Thorat, Archit
Published: (2026)
by: Thorat, Archit
Published: (2026)
Robust Reward Alignment via Hypothesis Space Batch Cutting
by: Xie, Zhixian, et al.
Published: (2025)
by: Xie, Zhixian, et al.
Published: (2025)
Affordance-Guided Reinforcement Learning via Visual Prompting
by: Lee, Olivia Y., et al.
Published: (2024)
by: Lee, Olivia Y., et al.
Published: (2024)
Strategic Hypothesis Testing
by: Hossain, Safwan, et al.
Published: (2025)
by: Hossain, Safwan, et al.
Published: (2025)
Reinforcement Learning via Implicit Imitation Guidance
by: Dong, Perry, et al.
Published: (2025)
by: Dong, Perry, et al.
Published: (2025)
Universal Neural Functionals
by: Zhou, Allan, et al.
Published: (2024)
by: Zhou, Allan, et al.
Published: (2024)
Learning Long-Context Diffusion Policies via Past-Token Prediction
by: Torne, Marcel, et al.
Published: (2025)
by: Torne, Marcel, et al.
Published: (2025)
Hypothesis Testing the Circuit Hypothesis in LLMs
by: Shi, Claudia, et al.
Published: (2024)
by: Shi, Claudia, et al.
Published: (2024)
Minimax Hypothesis Testing for the Bradley-Terry-Luce Model
by: Makur, Anuran, et al.
Published: (2024)
by: Makur, Anuran, et al.
Published: (2024)
The Master Key Hypothesis: Unlocking Cross-Model Capability Transfer via Linear Subspace Alignment
by: Balasubramanian, Rishab, et al.
Published: (2026)
by: Balasubramanian, Rishab, et al.
Published: (2026)
From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function
by: Rafailov, Rafael, et al.
Published: (2024)
by: Rafailov, Rafael, et al.
Published: (2024)
Similar Items
-
Feedback Descent: Open-Ended Text Optimization via Pairwise Comparison
by: Lee, Yoonho, et al.
Published: (2025) -
FSPO: Few-Shot Optimization of Synthetic Preferences Personalizes to Real Users
by: Singh, Anikait, et al.
Published: (2025) -
RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems
by: Qu, Yuxiao, et al.
Published: (2025) -
Calibrating Language Models with Adaptive Temperature Scaling
by: Xie, Johnathan, et al.
Published: (2024) -
Bidirectional Decoding: Improving Action Chunking via Guided Test-Time Sampling
by: Liu, Yuejiang, et al.
Published: (2024)