Labels or Preferences? Budget-Constrained Learning with Human Judgments over AI-Generated Outputs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Dong, Zihan, Hou, Xiaotian, Wu, Ruijia, Zhang, Linjun |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Evaluating LLMs When They Do Not Know the Answer: Statistical Evaluation of Mathematical Reasoning via Comparative Signals
par: Dong, Zihan, et autres
Publié: (2026)
par: Dong, Zihan, et autres
Publié: (2026)
Unified Inference Framework for Single and Multi-Player Performative Prediction: Method and Asymptotic Optimality
par: Zhang, Zhixian, et autres
Publié: (2026)
par: Zhang, Zhixian, et autres
Publié: (2026)
Residual Feature Integration is Sufficient to Prevent Negative Transfer
par: Xu, Yichen, et autres
Publié: (2025)
par: Xu, Yichen, et autres
Publié: (2025)
Provable Reward-Agnostic Preference-Based Reinforcement Learning
par: Zhan, Wenhao, et autres
Publié: (2023)
par: Zhan, Wenhao, et autres
Publié: (2023)
Fixed-Budget Differentially Private Best Arm Identification
par: Chen, Zhirui, et autres
Publié: (2024)
par: Chen, Zhirui, et autres
Publié: (2024)
A Unified Pair-GRPO Family: From Implicit to Explicit Preference Constraints for Stable and General RL Alignment
par: Yu, Hao
Publié: (2026)
par: Yu, Hao
Publié: (2026)
Unified Algorithms for RL with Decision-Estimation Coefficients: PAC, Reward-Free, Preference-Based Learning, and Beyond
par: Chen, Fan, et autres
Publié: (2022)
par: Chen, Fan, et autres
Publié: (2022)
A Statistical Hypothesis Testing Framework for Data Misappropriation Detection in Large Language Models
par: Cai, Yinpeng, et autres
Publié: (2025)
par: Cai, Yinpeng, et autres
Publié: (2025)
Compression, Generalization and Learning
par: Campi, Marco C., et autres
Publié: (2023)
par: Campi, Marco C., et autres
Publié: (2023)
Label Noise Robustness of Conformal Prediction
par: Einbinder, Bat-Sheva, et autres
Publié: (2022)
par: Einbinder, Bat-Sheva, et autres
Publié: (2022)
Counterfactual Generative Modeling with Variational Causal Inference
par: Wu, Yulun, et autres
Publié: (2024)
par: Wu, Yulun, et autres
Publié: (2024)
A Theory of the Mechanics of Information: Generalization Through Measurement of Uncertainty (Learning is Measuring)
par: Hazard, Christopher J., et autres
Publié: (2025)
par: Hazard, Christopher J., et autres
Publié: (2025)
Neural Networks Learn Generic Multi-Index Models Near Information-Theoretic Limit
par: Zhang, Bohan, et autres
Publié: (2025)
par: Zhang, Bohan, et autres
Publié: (2025)
Psychometric Tests for AI Agents and Their Moduli Space
par: Chojecki, Przemyslaw
Publié: (2025)
par: Chojecki, Przemyslaw
Publié: (2025)
Adaptive auditing of AI systems with anytime-valid guarantees
par: Zhou, Siyu, et autres
Publié: (2026)
par: Zhou, Siyu, et autres
Publié: (2026)
Solving a Research Problem in Mathematical Statistics with AI Assistance
par: Dobriban, Edgar
Publié: (2025)
par: Dobriban, Edgar
Publié: (2025)
Towards a Sharp Analysis of Offline Policy Learning for $f$-Divergence-Regularized Contextual Bandits
par: Zhao, Qingyue, et autres
Publié: (2025)
par: Zhao, Qingyue, et autres
Publié: (2025)
On the Statistical Capacity of Deep Generative Models
par: Tam, Edric, et autres
Publié: (2025)
par: Tam, Edric, et autres
Publié: (2025)
Learning Interpretable Concepts: Unifying Causal Representation Learning and Foundation Models
par: Rajendran, Goutham, et autres
Publié: (2024)
par: Rajendran, Goutham, et autres
Publié: (2024)
Characteristic Learning for Provable One Step Generation
par: Ding, Zhao, et autres
Publié: (2024)
par: Ding, Zhao, et autres
Publié: (2024)
Neural Networks Generalize on Low Complexity Data
par: Chatterjee, Sourav, et autres
Publié: (2024)
par: Chatterjee, Sourav, et autres
Publié: (2024)
Generalization and Scaling Laws for Mixture-of-Experts Transformers
par: Mayaki, Mansour Zoubeirou a
Publié: (2026)
par: Mayaki, Mansour Zoubeirou a
Publié: (2026)
FraPPE: Fast and Efficient Preference-based Pure Exploration
par: Das, Udvas, et autres
Publié: (2025)
par: Das, Udvas, et autres
Publié: (2025)
Towards Efficient Online Exploration for Reinforcement Learning with Human Feedback
par: Li, Gen, et autres
Publié: (2025)
par: Li, Gen, et autres
Publié: (2025)
Online Learning with Unknown Constraints
par: Sridharan, Karthik, et autres
Publié: (2024)
par: Sridharan, Karthik, et autres
Publié: (2024)
Training Implicit Generative Models via an Invariant Statistical Loss
par: de Frutos, José Manuel, et autres
Publié: (2024)
par: de Frutos, José Manuel, et autres
Publié: (2024)
Adaptive Sample Aggregation In Transfer Learning
par: Hanneke, Steve, et autres
Publié: (2024)
par: Hanneke, Steve, et autres
Publié: (2024)
On the Statistical Properties of Generative Adversarial Models for Low Intrinsic Data Dimension
par: Chakraborty, Saptarshi, et autres
Publié: (2024)
par: Chakraborty, Saptarshi, et autres
Publié: (2024)
Conformal Prediction for Privacy-Preserving Machine Learning
par: Balinsky, Alexander David, et autres
Publié: (2025)
par: Balinsky, Alexander David, et autres
Publié: (2025)
Learning with Differentially Private (Sliced) Wasserstein Gradients
par: Rodríguez-Vítores, David, et autres
Publié: (2025)
par: Rodríguez-Vítores, David, et autres
Publié: (2025)
Generalization Properties of Score-matching Diffusion Models for Intrinsically Low-dimensional Data
par: Chakraborty, Saptarshi, et autres
Publié: (2026)
par: Chakraborty, Saptarshi, et autres
Publié: (2026)
U-Nets as Belief Propagation: Efficient Classification, Denoising, and Diffusion in Generative Hierarchical Models
par: Mei, Song
Publié: (2024)
par: Mei, Song
Publié: (2024)
Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits
par: Chen, Fan, et autres
Publié: (2025)
par: Chen, Fan, et autres
Publié: (2025)
Tail-Aware Information-Theoretic Generalization for RLHF and SGLD
par: Zhang, Huiming, et autres
Publié: (2026)
par: Zhang, Huiming, et autres
Publié: (2026)
Understanding In-Context Learning on Structured Manifolds: Bridging Attention to Kernel Methods
par: Shen, Zhaiming, et autres
Publié: (2025)
par: Shen, Zhaiming, et autres
Publié: (2025)
Learning Hierarchical Polynomials of Multiple Nonlinear Features with Three-Layer Networks
par: Fu, Hengyu, et autres
Publié: (2024)
par: Fu, Hengyu, et autres
Publié: (2024)
Chemical Reaction Networks Learn Better than Spiking Neural Networks
par: Jaffard, Sophie, et autres
Publié: (2026)
par: Jaffard, Sophie, et autres
Publié: (2026)
Is Behavior Cloning All You Need? Understanding Horizon in Imitation Learning
par: Foster, Dylan J., et autres
Publié: (2024)
par: Foster, Dylan J., et autres
Publié: (2024)
Beyond identifiability: Learning causal representations with few environments and finite samples
par: Lee, Inbeom, et autres
Publié: (2026)
par: Lee, Inbeom, et autres
Publié: (2026)
Can Generative Artificial Intelligence Survive Data Contamination? Theoretical Guarantees under Contaminated Recursive Training
par: Wang, Kevin, et autres
Publié: (2026)
par: Wang, Kevin, et autres
Publié: (2026)
Documents similaires
-
Evaluating LLMs When They Do Not Know the Answer: Statistical Evaluation of Mathematical Reasoning via Comparative Signals
par: Dong, Zihan, et autres
Publié: (2026) -
Unified Inference Framework for Single and Multi-Player Performative Prediction: Method and Asymptotic Optimality
par: Zhang, Zhixian, et autres
Publié: (2026) -
Residual Feature Integration is Sufficient to Prevent Negative Transfer
par: Xu, Yichen, et autres
Publié: (2025) -
Provable Reward-Agnostic Preference-Based Reinforcement Learning
par: Zhan, Wenhao, et autres
Publié: (2023) -
Fixed-Budget Differentially Private Best Arm Identification
par: Chen, Zhirui, et autres
Publié: (2024)