Learning to Choose or Choosing to Learn: Best-of-N vs. Supervised Fine-Tuning for Bit String Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Somerstep, Seamus, Raman, Vinod, Subedi, Unique, Sun, Yuekai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Combinatorial Characterization of Supervised Online Learnability
by: Raman, Vinod, et al.
Published: (2023)
by: Raman, Vinod, et al.
Published: (2023)
Online Learning with Set-Valued Feedback
by: Raman, Vinod, et al.
Published: (2023)
by: Raman, Vinod, et al.
Published: (2023)
Online Infinite-Dimensional Regression: Learning Linear Operators
by: Raman, Vinod, et al.
Published: (2023)
by: Raman, Vinod, et al.
Published: (2023)
Algorithmic Fairness in Performative Policy Learning: Escaping the Impossibility of Group Fairness
by: Somerstep, Seamus, et al.
Published: (2024)
by: Somerstep, Seamus, et al.
Published: (2024)
Learning In Reverse Causal Strategic Environments With Ramifications on Two Sided Markets
by: Somerstep, Seamus, et al.
Published: (2024)
by: Somerstep, Seamus, et al.
Published: (2024)
Multiclass Transductive Online Learning
by: Hanneke, Steve, et al.
Published: (2024)
by: Hanneke, Steve, et al.
Published: (2024)
The Complexity of Sequential Prediction in Dynamical Systems
by: Raman, Vinod, et al.
Published: (2024)
by: Raman, Vinod, et al.
Published: (2024)
A Characterization of Multioutput Learnability
by: Raman, Vinod, et al.
Published: (2023)
by: Raman, Vinod, et al.
Published: (2023)
Smoothed Online Classification can be Harder than Batch Classification
by: Raman, Vinod, et al.
Published: (2024)
by: Raman, Vinod, et al.
Published: (2024)
Apple Tasting: Combinatorial Dimensions and Minimax Rates
by: Raman, Vinod, et al.
Published: (2023)
by: Raman, Vinod, et al.
Published: (2023)
Multiclass Online Learnability under Bandit Feedback
by: Raman, Ananth, et al.
Published: (2023)
by: Raman, Ananth, et al.
Published: (2023)
Operator Learning: A Statistical Perspective
by: Subedi, Unique, et al.
Published: (2025)
by: Subedi, Unique, et al.
Published: (2025)
On the Benefits of Active Data Collection in Operator Learning
by: Subedi, Unique, et al.
Published: (2024)
by: Subedi, Unique, et al.
Published: (2024)
Limitations of refinement methods for weak to strong generalization
by: Somerstep, Seamus, et al.
Published: (2025)
by: Somerstep, Seamus, et al.
Published: (2025)
Is Zero-Shot Super-Resolution Possible in Operator Learning?
by: Subedi, Unique, et al.
Published: (2026)
by: Subedi, Unique, et al.
Published: (2026)
Sloth: scaling laws for LLM skills to predict multi-benchmark performance across families
by: Polo, Felipe Maia, et al.
Published: (2024)
by: Polo, Felipe Maia, et al.
Published: (2024)
Controlling Statistical, Discretization, and Truncation Errors in Learning Fourier Linear Operators
by: Subedi, Unique, et al.
Published: (2024)
by: Subedi, Unique, et al.
Published: (2024)
Operator Learning for Schrödinger Equation: Unitarity, Error Bounds, and Time Generalization
by: Patel, Yash, et al.
Published: (2025)
by: Patel, Yash, et al.
Published: (2025)
Optimal Stopping vs Best-of-$N$ for Inference Time Optimization
by: Kalayci, Yusuf, et al.
Published: (2025)
by: Kalayci, Yusuf, et al.
Published: (2025)
Microfoundation Inference for Strategic Prediction
by: Bracale, Daniele, et al.
Published: (2024)
by: Bracale, Daniele, et al.
Published: (2024)
A transfer learning framework for weak-to-strong generalization
by: Somerstep, Seamus, et al.
Published: (2024)
by: Somerstep, Seamus, et al.
Published: (2024)
AdaBoN: Adaptive Best-of-N Alignment
by: Raman, Vinod, et al.
Published: (2025)
by: Raman, Vinod, et al.
Published: (2025)
Tracking the Best Expert Privately
by: Saha, Aadirupa, et al.
Published: (2025)
by: Saha, Aadirupa, et al.
Published: (2025)
Learning from Streaming Data when Users Choose
by: Su, Jinyan, et al.
Published: (2024)
by: Su, Jinyan, et al.
Published: (2024)
Generation from Noisy Examples
by: Raman, Ananth, et al.
Published: (2025)
by: Raman, Ananth, et al.
Published: (2025)
Supervised Fine-Tuning as Inverse Reinforcement Learning
by: Sun, Hao
Published: (2024)
by: Sun, Hao
Published: (2024)
A Latent Variable Framework for Scaling Laws in Large Language Models
by: Cai, Peiyao, et al.
Published: (2025)
by: Cai, Peiyao, et al.
Published: (2025)
Learning ON Large Datasets Using Bit-String Trees
by: Gupta, Prashant
Published: (2025)
by: Gupta, Prashant
Published: (2025)
Choosing a Proxy Metric from Past Experiments
by: Tripuraneni, Nilesh, et al.
Published: (2023)
by: Tripuraneni, Nilesh, et al.
Published: (2023)
CYCle: Choosing Your Collaborators Wisely to Enhance Collaborative Fairness in Decentralized Learning
by: Tastan, Nurbek, et al.
Published: (2025)
by: Tastan, Nurbek, et al.
Published: (2025)
Choosing How to Remember: Adaptive Memory Structures for LLM Agents
by: Lu, Mingfei, et al.
Published: (2026)
by: Lu, Mingfei, et al.
Published: (2026)
Let Quantum Neural Networks Choose Their Own Frequencies
by: Jaderberg, Ben, et al.
Published: (2023)
by: Jaderberg, Ben, et al.
Published: (2023)
Choosing a Classical Planner with Graph Neural Networks
by: Vatter, Jana, et al.
Published: (2024)
by: Vatter, Jana, et al.
Published: (2024)
How to Choose a Reinforcement-Learning Algorithm
by: Bongratz, Fabian, et al.
Published: (2024)
by: Bongratz, Fabian, et al.
Published: (2024)
Choosing the Right Weights: Balancing Value, Strategy, and Noise in Recommender Systems
by: Milli, Smitha, et al.
Published: (2023)
by: Milli, Smitha, et al.
Published: (2023)
Transductive and Learning-Augmented Online Regression
by: Raman, Vinod, et al.
Published: (2025)
by: Raman, Vinod, et al.
Published: (2025)
On Generation in Metric Spaces
by: Li, Jiaxun, et al.
Published: (2026)
by: Li, Jiaxun, et al.
Published: (2026)
Goal-Conditioned Supervised Learning for LLM Fine-Tuning
by: Li, Shijun, et al.
Published: (2026)
by: Li, Shijun, et al.
Published: (2026)
Generation through the lens of learning theory
by: Li, Jiaxun, et al.
Published: (2024)
by: Li, Jiaxun, et al.
Published: (2024)
Erasing the Bias: Fine-Tuning Foundation Models for Semi-Supervised Learning
by: Gan, Kai, et al.
Published: (2024)
by: Gan, Kai, et al.
Published: (2024)
Similar Items
-
A Combinatorial Characterization of Supervised Online Learnability
by: Raman, Vinod, et al.
Published: (2023) -
Online Learning with Set-Valued Feedback
by: Raman, Vinod, et al.
Published: (2023) -
Online Infinite-Dimensional Regression: Learning Linear Operators
by: Raman, Vinod, et al.
Published: (2023) -
Algorithmic Fairness in Performative Policy Learning: Escaping the Impossibility of Group Fairness
by: Somerstep, Seamus, et al.
Published: (2024) -
Learning In Reverse Causal Strategic Environments With Ramifications on Two Sided Markets
by: Somerstep, Seamus, et al.
Published: (2024)