The Limits of Preference Data for Post-Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Eric, Dai, Jessica, Awasthi, Pranjal |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Nash CoT: Multi-Path Inference with Preference Equilibrium
von: Zhang, Ziqi, et al.
Veröffentlicht: (2024)
von: Zhang, Ziqi, et al.
Veröffentlicht: (2024)
COMAL: A Convergent Meta-Algorithm for Aligning LLMs with General Preferences
von: Liu, Yixin, et al.
Veröffentlicht: (2024)
von: Liu, Yixin, et al.
Veröffentlicht: (2024)
Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
von: Zhang, Yuheng, et al.
Veröffentlicht: (2024)
von: Zhang, Yuheng, et al.
Veröffentlicht: (2024)
PrefBench: Evaluating Zero-Shot LLM Agents in Hidden-Preference Personalized Pricing Negotiations
von: Lei, Yingjie
Veröffentlicht: (2026)
von: Lei, Yingjie
Veröffentlicht: (2026)
Truthfulness Despite Weak Supervision: Evaluating and Training LLMs Using Peer Prediction
von: Qiu, Tianyi Alex, et al.
Veröffentlicht: (2026)
von: Qiu, Tianyi Alex, et al.
Veröffentlicht: (2026)
GameTalk: Training LLMs for Strategic Conversation
von: Vendrell, Victor Conchello, et al.
Veröffentlicht: (2026)
von: Vendrell, Victor Conchello, et al.
Veröffentlicht: (2026)
On the Fundamental Impossibility of Hallucination Control in Large Language Models
von: Karpowicz, Michał P.
Veröffentlicht: (2025)
von: Karpowicz, Michał P.
Veröffentlicht: (2025)
Multi-Head Attention Is a Multi-Player Game
von: Chakrabarti, Kushal, et al.
Veröffentlicht: (2026)
von: Chakrabarti, Kushal, et al.
Veröffentlicht: (2026)
Ad Auctions for LLMs via Retrieval Augmented Generation
von: Hajiaghayi, MohammadTaghi, et al.
Veröffentlicht: (2024)
von: Hajiaghayi, MohammadTaghi, et al.
Veröffentlicht: (2024)
Asymptotic Universal Alignment: A New Alignment Framework via Test-Time Scaling
von: Cai, Yang, et al.
Veröffentlicht: (2026)
von: Cai, Yang, et al.
Veröffentlicht: (2026)
On The Truthfulness of 'Surprisingly Likely' Responses of Large Language Models
von: Goel, Naman
Veröffentlicht: (2023)
von: Goel, Naman
Veröffentlicht: (2023)
Language Alignment via Nash-learning and Adaptive feedback
von: Azarafrooz, Ari, et al.
Veröffentlicht: (2024)
von: Azarafrooz, Ari, et al.
Veröffentlicht: (2024)
Used Car Salesbots? Honesty and Credulity of LLMs as Bargaining Agents under Partial Information
von: Miceli-Barone, Antonio Valerio, et al.
Veröffentlicht: (2026)
von: Miceli-Barone, Antonio Valerio, et al.
Veröffentlicht: (2026)
LLM-Powered Preference Elicitation in Combinatorial Assignment
von: Soumalias, Ermis, et al.
Veröffentlicht: (2025)
von: Soumalias, Ermis, et al.
Veröffentlicht: (2025)
Bandits with Preference Feedback: A Stackelberg Game Perspective
von: Pásztor, Barna, et al.
Veröffentlicht: (2024)
von: Pásztor, Barna, et al.
Veröffentlicht: (2024)
Language Self-Play For Data-Free Training
von: Kuba, Jakub Grudzien, et al.
Veröffentlicht: (2025)
von: Kuba, Jakub Grudzien, et al.
Veröffentlicht: (2025)
Preferences Evolve And So Should Your Bandits: Bandits with Evolving States for Online Platforms
von: Khosravi, Khashayar, et al.
Veröffentlicht: (2023)
von: Khosravi, Khashayar, et al.
Veröffentlicht: (2023)
Efficiently Training Neural Networks for Imperfect Information Games by Sampling Information Sets
von: Bertram, Timo, et al.
Veröffentlicht: (2024)
von: Bertram, Timo, et al.
Veröffentlicht: (2024)
Representative Social Choice: From Learning Theory to AI Alignment
von: Qiu, Tianyi
Veröffentlicht: (2024)
von: Qiu, Tianyi
Veröffentlicht: (2024)
GLEE: A Unified Framework and Benchmark for Language-based Economic Environments
von: Shapira, Eilam, et al.
Veröffentlicht: (2024)
von: Shapira, Eilam, et al.
Veröffentlicht: (2024)
Preference-Based Multi-Agent Reinforcement Learning: Data Coverage and Algorithmic Techniques
von: Zhang, Natalia, et al.
Veröffentlicht: (2024)
von: Zhang, Natalia, et al.
Veröffentlicht: (2024)
On the Impact of the Utility in Semivalue-based Data Valuation
von: Tamine, Mélissa, et al.
Veröffentlicht: (2025)
von: Tamine, Mélissa, et al.
Veröffentlicht: (2025)
Incentivizing Truthful Language Models via Peer Elicitation Games
von: Chen, Baiting, et al.
Veröffentlicht: (2025)
von: Chen, Baiting, et al.
Veröffentlicht: (2025)
Ranking Abuse via Strategic Pairwise Data Perturbations
von: Yao, Junyi, et al.
Veröffentlicht: (2026)
von: Yao, Junyi, et al.
Veröffentlicht: (2026)
Trustworthy Machine Learning under Social and Adversarial Data Sources
von: Shao, Han
Veröffentlicht: (2024)
von: Shao, Han
Veröffentlicht: (2024)
VickreyFeedback: Cost-efficient Data Construction for Reinforcement Learning from Human Feedback
von: Zhang, Guoxi, et al.
Veröffentlicht: (2024)
von: Zhang, Guoxi, et al.
Veröffentlicht: (2024)
Federated Learning for Data Market: Shapley-UCB for Seller Selection and Incentives
von: Chen, Kongyang, et al.
Veröffentlicht: (2024)
von: Chen, Kongyang, et al.
Veröffentlicht: (2024)
Peer-Predictive Self-Training for Language Model Reasoning
von: Feng, Shi, et al.
Veröffentlicht: (2026)
von: Feng, Shi, et al.
Veröffentlicht: (2026)
Agent-oriented Joint Decision Support for Data Owners in Auction-based Federated Learning
von: Tang, Xiaoli, et al.
Veröffentlicht: (2024)
von: Tang, Xiaoli, et al.
Veröffentlicht: (2024)
How Can Incentives and Cut Layer Selection Influence Data Contribution in Split Federated Learning?
von: Lee, Joohyung, et al.
Veröffentlicht: (2024)
von: Lee, Joohyung, et al.
Veröffentlicht: (2024)
Reasoning Like an Economist: Post-Training on Economic Problems Induces Strategic Generalization in LLMs
von: Zhou, Yufa, et al.
Veröffentlicht: (2025)
von: Zhou, Yufa, et al.
Veröffentlicht: (2025)
LUDOBENCH: Evaluating LLM Behavioural Decision-Making Through Spot-Based Board Game Scenarios in Ludo
von: Jain, Ojas, et al.
Veröffentlicht: (2026)
von: Jain, Ojas, et al.
Veröffentlicht: (2026)
The Price Reversal Phenomenon: When Cheaper Reasoning Models Cost More
von: Chen, Lingjiao, et al.
Veröffentlicht: (2026)
von: Chen, Lingjiao, et al.
Veröffentlicht: (2026)
Game-theoretic LLM: Agent Workflow for Negotiation Games
von: Hua, Wenyue, et al.
Veröffentlicht: (2024)
von: Hua, Wenyue, et al.
Veröffentlicht: (2024)
PokerBench: Training Large Language Models to become Professional Poker Players
von: Zhuang, Richard, et al.
Veröffentlicht: (2025)
von: Zhuang, Richard, et al.
Veröffentlicht: (2025)
Can LLMs Replace Economic Choice Prediction Labs? The Case of Language-based Persuasion Games
von: Shapira, Eilam, et al.
Veröffentlicht: (2024)
von: Shapira, Eilam, et al.
Veröffentlicht: (2024)
LLM-Auction: Generative Auction towards LLM-Native Advertising
von: Zhao, Chujie, et al.
Veröffentlicht: (2025)
von: Zhao, Chujie, et al.
Veröffentlicht: (2025)
Meta-Computing Enhanced Federated Learning in IIoT: Satisfaction-Aware Incentive Scheme via DRL-Based Stackelberg Game
von: Li, Xiaohuan, et al.
Veröffentlicht: (2025)
von: Li, Xiaohuan, et al.
Veröffentlicht: (2025)
Stackelberg Learning from Human Feedback: Preference Optimization as a Sequential Game
von: Pásztor, Barna, et al.
Veröffentlicht: (2025)
von: Pásztor, Barna, et al.
Veröffentlicht: (2025)
Learning Aggregation Rules in Participatory Budgeting: A Data-Driven Approach
von: Fairstein, Roy, et al.
Veröffentlicht: (2024)
von: Fairstein, Roy, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Nash CoT: Multi-Path Inference with Preference Equilibrium
von: Zhang, Ziqi, et al.
Veröffentlicht: (2024) -
COMAL: A Convergent Meta-Algorithm for Aligning LLMs with General Preferences
von: Liu, Yixin, et al.
Veröffentlicht: (2024) -
Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
von: Zhang, Yuheng, et al.
Veröffentlicht: (2024) -
PrefBench: Evaluating Zero-Shot LLM Agents in Hidden-Preference Personalized Pricing Negotiations
von: Lei, Yingjie
Veröffentlicht: (2026) -
Truthfulness Despite Weak Supervision: Evaluating and Training LLMs Using Peer Prediction
von: Qiu, Tianyi Alex, et al.
Veröffentlicht: (2026)