Nash CoT: Multi-Path Inference with Preference Equilibrium
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Ziqi, Wang, Cunxiang, Xiao, Xiong, Zhang, Yue, Wang, Donglin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
by: Zhang, Yuheng, et al.
Published: (2024)
by: Zhang, Yuheng, et al.
Published: (2024)
Large-Scale Auto-bidding with Nash Equilibrium Constraints
by: Mou, Zhiyu, et al.
Published: (2025)
by: Mou, Zhiyu, et al.
Published: (2025)
Paths to Equilibrium in Games
by: Yongacoglu, Bora, et al.
Published: (2024)
by: Yongacoglu, Bora, et al.
Published: (2024)
Language Alignment via Nash-learning and Adaptive feedback
by: Azarafrooz, Ari, et al.
Published: (2024)
by: Azarafrooz, Ari, et al.
Published: (2024)
Accelerating Nash Equilibrium Convergence in Monte Carlo Settings Through Counterfactual Value Based Fictitious Play
by: Qi, Ju, et al.
Published: (2023)
by: Qi, Ju, et al.
Published: (2023)
The Limits of Preference Data for Post-Training
by: Zhao, Eric, et al.
Published: (2025)
by: Zhao, Eric, et al.
Published: (2025)
COMAL: A Convergent Meta-Algorithm for Aligning LLMs with General Preferences
by: Liu, Yixin, et al.
Published: (2024)
by: Liu, Yixin, et al.
Published: (2024)
PrefBench: Evaluating Zero-Shot LLM Agents in Hidden-Preference Personalized Pricing Negotiations
by: Lei, Yingjie
Published: (2026)
by: Lei, Yingjie
Published: (2026)
Policy Optimization finds Nash Equilibrium in Regularized General-Sum LQ Games
by: Zaman, Muhammad Aneeq uz, et al.
Published: (2024)
by: Zaman, Muhammad Aneeq uz, et al.
Published: (2024)
Multi-Head Attention Is a Multi-Player Game
by: Chakrabarti, Kushal, et al.
Published: (2026)
by: Chakrabarti, Kushal, et al.
Published: (2026)
How Likely Do LLMs with CoT Mimic Human Reasoning?
by: Bao, Guangsheng, et al.
Published: (2024)
by: Bao, Guangsheng, et al.
Published: (2024)
Convergence to Nash Equilibrium and No-regret Guarantee in (Markov) Potential Games
by: Dong, Jing, et al.
Published: (2024)
by: Dong, Jing, et al.
Published: (2024)
Beyond Nash Equilibrium: Bounded Rationality of LLMs and humans in Strategic Decision-making
by: Zheng, Kehan, et al.
Published: (2025)
by: Zheng, Kehan, et al.
Published: (2025)
Data Poisoning to Fake a Nash Equilibrium in Markov Games
by: Wu, Young, et al.
Published: (2023)
by: Wu, Young, et al.
Published: (2023)
Explore Reinforced: Equilibrium Approximation with Reinforcement Learning
by: Yu, Ryan, et al.
Published: (2024)
by: Yu, Ryan, et al.
Published: (2024)
Preference-Based Multi-Agent Reinforcement Learning: Data Coverage and Algorithmic Techniques
by: Zhang, Natalia, et al.
Published: (2024)
by: Zhang, Natalia, et al.
Published: (2024)
Online Learning and Equilibrium Computation with Ranking Feedback
by: Liu, Mingyang, et al.
Published: (2026)
by: Liu, Mingyang, et al.
Published: (2026)
Large Language Models Playing Mixed Strategy Nash Equilibrium Games
by: Silva, Alonso
Published: (2024)
by: Silva, Alonso
Published: (2024)
Ad Auctions for LLMs via Retrieval Augmented Generation
by: Hajiaghayi, MohammadTaghi, et al.
Published: (2024)
by: Hajiaghayi, MohammadTaghi, et al.
Published: (2024)
On the Fundamental Impossibility of Hallucination Control in Large Language Models
by: Karpowicz, Michał P.
Published: (2025)
by: Karpowicz, Michał P.
Published: (2025)
Asymptotic Universal Alignment: A New Alignment Framework via Test-Time Scaling
by: Cai, Yang, et al.
Published: (2026)
by: Cai, Yang, et al.
Published: (2026)
On The Truthfulness of 'Surprisingly Likely' Responses of Large Language Models
by: Goel, Naman
Published: (2023)
by: Goel, Naman
Published: (2023)
Truthfulness Despite Weak Supervision: Evaluating and Training LLMs Using Peer Prediction
by: Qiu, Tianyi Alex, et al.
Published: (2026)
by: Qiu, Tianyi Alex, et al.
Published: (2026)
Used Car Salesbots? Honesty and Credulity of LLMs as Bargaining Agents under Partial Information
by: Miceli-Barone, Antonio Valerio, et al.
Published: (2026)
by: Miceli-Barone, Antonio Valerio, et al.
Published: (2026)
LLM-Powered Preference Elicitation in Combinatorial Assignment
by: Soumalias, Ermis, et al.
Published: (2025)
by: Soumalias, Ermis, et al.
Published: (2025)
Bandits with Preference Feedback: A Stackelberg Game Perspective
by: Pásztor, Barna, et al.
Published: (2024)
by: Pásztor, Barna, et al.
Published: (2024)
Minimally Modifying a Markov Game to Achieve Any Nash Equilibrium and Value
by: Wu, Young, et al.
Published: (2023)
by: Wu, Young, et al.
Published: (2023)
Approximate Nash Equilibrium Learning for n-Player Markov Games in Dynamic Pricing
by: Liu, Larkin
Published: (2022)
by: Liu, Larkin
Published: (2022)
Fusion-PSRO: Nash Policy Fusion for Policy Space Response Oracles
by: Lian, Jiesong, et al.
Published: (2024)
by: Lian, Jiesong, et al.
Published: (2024)
Nash Learning from Human Feedback
by: Munos, Rémi, et al.
Published: (2023)
by: Munos, Rémi, et al.
Published: (2023)
Statistical Impossibility and Possibility of Aligning LLMs with Human Preferences: From Condorcet Paradox to Nash Equilibrium
by: Liu, Kaizhao, et al.
Published: (2025)
by: Liu, Kaizhao, et al.
Published: (2025)
Game-theoretic LLM: Agent Workflow for Negotiation Games
by: Hua, Wenyue, et al.
Published: (2024)
by: Hua, Wenyue, et al.
Published: (2024)
Mixed Strategy Nash Equilibrium for Crowd Navigation
by: Sun, Max Muchen, et al.
Published: (2024)
by: Sun, Max Muchen, et al.
Published: (2024)
Verbalized Bayesian Persuasion
by: Li, Wenhao, et al.
Published: (2025)
by: Li, Wenhao, et al.
Published: (2025)
Preferences Evolve And So Should Your Bandits: Bandits with Evolving States for Online Platforms
by: Khosravi, Khashayar, et al.
Published: (2023)
by: Khosravi, Khashayar, et al.
Published: (2023)
Discovering Expert-Level Nash Equilibrium Algorithms with Large Language Models
by: Li, Hanyu, et al.
Published: (2025)
by: Li, Hanyu, et al.
Published: (2025)
Gnothi Seauton: Empowering Faithful Self-Interpretability in Black-Box Transformers
by: Wang, Shaobo, et al.
Published: (2024)
by: Wang, Shaobo, et al.
Published: (2024)
Puzzle it Out: Local-to-Global World Model for Offline Multi-Agent Reinforcement Learning
by: Li, Sijia, et al.
Published: (2026)
by: Li, Sijia, et al.
Published: (2026)
Multi-Agent Training beyond Zero-Sum with Correlated Equilibrium Meta-Solvers
by: Marris, Luke, et al.
Published: (2021)
by: Marris, Luke, et al.
Published: (2021)
GemNet: Menu-Based, Strategy-Proof Multi-Bidder Auctions Through Deep Learning
by: Wang, Tonghan, et al.
Published: (2024)
by: Wang, Tonghan, et al.
Published: (2024)
Similar Items
-
Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
by: Zhang, Yuheng, et al.
Published: (2024) -
Large-Scale Auto-bidding with Nash Equilibrium Constraints
by: Mou, Zhiyu, et al.
Published: (2025) -
Paths to Equilibrium in Games
by: Yongacoglu, Bora, et al.
Published: (2024) -
Language Alignment via Nash-learning and Adaptive feedback
by: Azarafrooz, Ari, et al.
Published: (2024) -
Accelerating Nash Equilibrium Convergence in Monte Carlo Settings Through Counterfactual Value Based Fictitious Play
by: Qi, Ju, et al.
Published: (2023)