Reasoning, Memorization, and Fine-Tuning Language Models for Non-Cooperative Games
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Yunhao, Berthellemy, Leonard, Topcu, Ufuk |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Human-Agent Cooperation in Games under Incomplete Information through Natural Language Communication
von: Chen, Shenghui, et al.
Veröffentlicht: (2024)
von: Chen, Shenghui, et al.
Veröffentlicht: (2024)
Fine-Tuning Language Models Using Formal Methods Feedback
von: Yang, Yunhao, et al.
Veröffentlicht: (2023)
von: Yang, Yunhao, et al.
Veröffentlicht: (2023)
Multimodal Pretrained Models for Verifiable Sequential Decision-Making: Planning, Grounding, and Perception
von: Yang, Yunhao, et al.
Veröffentlicht: (2023)
von: Yang, Yunhao, et al.
Veröffentlicht: (2023)
Memorization in Fine-Tuned Large Language Models
von: Savine, Danil
Veröffentlicht: (2025)
von: Savine, Danil
Veröffentlicht: (2025)
Joint Verification and Refinement of Language Models for Safety-Constrained Planning
von: Yang, Yunhao, et al.
Veröffentlicht: (2024)
von: Yang, Yunhao, et al.
Veröffentlicht: (2024)
IG-MCTS: Human-in-the-Loop Cooperative Navigation under Incomplete Information
von: Chen, Shenghui, et al.
Veröffentlicht: (2025)
von: Chen, Shenghui, et al.
Veröffentlicht: (2025)
Impact of Fine-Tuning Methods on Memorization in Large Language Models
von: Hou, Jie, et al.
Veröffentlicht: (2025)
von: Hou, Jie, et al.
Veröffentlicht: (2025)
Unintended Memorization of Sensitive Information in Fine-Tuned Language Models
von: Szep, Marton, et al.
Veröffentlicht: (2026)
von: Szep, Marton, et al.
Veröffentlicht: (2026)
Human-Agent Coordination in Games under Incomplete Information via Multi-Step Intent
von: Chen, Shenghui, et al.
Veröffentlicht: (2024)
von: Chen, Shenghui, et al.
Veröffentlicht: (2024)
Evaluating Human Trust in LLM-Based Planners: A Preliminary Study
von: Chen, Shenghui, et al.
Veröffentlicht: (2025)
von: Chen, Shenghui, et al.
Veröffentlicht: (2025)
Assessing and Mitigating Data Memorization Risks in Fine-Tuned Large Language Models
von: Ramakrishnan, Badrinath, et al.
Veröffentlicht: (2025)
von: Ramakrishnan, Badrinath, et al.
Veröffentlicht: (2025)
Do LLMs Strategically Reveal, Conceal, and Infer Information? A Theoretical and Empirical Analysis in The Chameleon Game
von: Karabag, Mustafa O., et al.
Veröffentlicht: (2025)
von: Karabag, Mustafa O., et al.
Veröffentlicht: (2025)
Know Where You're Uncertain When Planning with Multimodal Foundation Models: A Formal Framework
von: Bhatt, Neel P., et al.
Veröffentlicht: (2024)
von: Bhatt, Neel P., et al.
Veröffentlicht: (2024)
Foundation Models for Logistics: Toward Certifiable, Conversational Planning Interfaces
von: Yang, Yunhao, et al.
Veröffentlicht: (2025)
von: Yang, Yunhao, et al.
Veröffentlicht: (2025)
Zero-Shot Reinforcement Learning via Function Encoders
von: Ingebrand, Tyler, et al.
Veröffentlicht: (2024)
von: Ingebrand, Tyler, et al.
Veröffentlicht: (2024)
Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning
von: Tian, Chang, et al.
Veröffentlicht: (2025)
von: Tian, Chang, et al.
Veröffentlicht: (2025)
Exploring Memorization in Fine-tuned Language Models
von: Zeng, Shenglai, et al.
Veröffentlicht: (2023)
von: Zeng, Shenglai, et al.
Veröffentlicht: (2023)
Neural Port-Hamiltonian Differential Algebraic Equations for Compositional Learning of Electrical Networks
von: Neary, Cyrus, et al.
Veröffentlicht: (2024)
von: Neary, Cyrus, et al.
Veröffentlicht: (2024)
RepV: Safety-Separable Latent Spaces for Scalable Neurosymbolic Plan Verification
von: Yang, Yunhao, et al.
Veröffentlicht: (2025)
von: Yang, Yunhao, et al.
Veröffentlicht: (2025)
VLN-Zero: Rapid Exploration and Cache-Enabled Neurosymbolic Vision-Language Planning for Zero-Shot Transfer in Robot Navigation
von: Bhatt, Neel P., et al.
Veröffentlicht: (2025)
von: Bhatt, Neel P., et al.
Veröffentlicht: (2025)
Multi-Environment POMDPs: Discrete Model Uncertainty Under Partial Observability
von: Bovy, Eline M., et al.
Veröffentlicht: (2025)
von: Bovy, Eline M., et al.
Veröffentlicht: (2025)
When Should a Leader Act Suboptimally? The Role of Inferability in Repeated Stackelberg Games
von: Karabag, Mustafa O., et al.
Veröffentlicht: (2023)
von: Karabag, Mustafa O., et al.
Veröffentlicht: (2023)
Learning to Coordinate without Communication under Incomplete Information
von: Chen, Shenghui, et al.
Veröffentlicht: (2024)
von: Chen, Shenghui, et al.
Veröffentlicht: (2024)
Adaptive Shielding for Safe Reinforcement Learning under Hidden-Parameter Dynamics Shifts
von: Kwon, Minjae, et al.
Veröffentlicht: (2025)
von: Kwon, Minjae, et al.
Veröffentlicht: (2025)
Sequential Resource Trading Using Comparison-Based Gradient Estimation
von: Murthy, Surya, et al.
Veröffentlicht: (2024)
von: Murthy, Surya, et al.
Veröffentlicht: (2024)
Using Large Language Models to Automate and Expedite Reinforcement Learning with Reward Machine
von: Alsadat, Shayan Meshkat, et al.
Veröffentlicht: (2024)
von: Alsadat, Shayan Meshkat, et al.
Veröffentlicht: (2024)
Unveiling the Impact of Coding Data Instruction Fine-Tuning on Large Language Models Reasoning
von: Zhang, Xinlu, et al.
Veröffentlicht: (2024)
von: Zhang, Xinlu, et al.
Veröffentlicht: (2024)
Categorical semantics of compositional reinforcement learning
von: Bakirtzis, Georgios, et al.
Veröffentlicht: (2022)
von: Bakirtzis, Georgios, et al.
Veröffentlicht: (2022)
Fine-Tuning Large Language Models for Cooperative Tactical Deconfliction of Small Unmanned Aerial Systems
von: Sharifi, Iman, et al.
Veröffentlicht: (2026)
von: Sharifi, Iman, et al.
Veröffentlicht: (2026)
The Reasoning-Memorization Interplay in Language Models Is Mediated by a Single Direction
von: Hong, Yihuai, et al.
Veröffentlicht: (2025)
von: Hong, Yihuai, et al.
Veröffentlicht: (2025)
Why Do LLMs Struggle in Strategic Play? Broken Links Between Observations, Beliefs, and Actions
von: Sobotka, Jan, et al.
Veröffentlicht: (2026)
von: Sobotka, Jan, et al.
Veröffentlicht: (2026)
Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning of Vision Language Models
von: Tan, Huajie, et al.
Veröffentlicht: (2025)
von: Tan, Huajie, et al.
Veröffentlicht: (2025)
GRIP: In-Parameter Graph Reasoning through Fine-Tuning Large Language Models
von: Feng, Jiarui, et al.
Veröffentlicht: (2025)
von: Feng, Jiarui, et al.
Veröffentlicht: (2025)
Robust Multi-Agent Reinforcement Learning for Small UAS Separation Assurance under GPS Degradation and Spoofing
von: Zongo, Alex, et al.
Veröffentlicht: (2026)
von: Zongo, Alex, et al.
Veröffentlicht: (2026)
Mitigating Memorization In Language Models
von: Sakarvadia, Mansi, et al.
Veröffentlicht: (2024)
von: Sakarvadia, Mansi, et al.
Veröffentlicht: (2024)
Online Foundation Model Selection in Robotics
von: Li, Po-han, et al.
Veröffentlicht: (2024)
von: Li, Po-han, et al.
Veröffentlicht: (2024)
On-Policy Supervised Fine-Tuning for Efficient Reasoning
von: Zhao, Anhao, et al.
Veröffentlicht: (2026)
von: Zhao, Anhao, et al.
Veröffentlicht: (2026)
Pramana: Fine-Tuning Large Language Models for Epistemic Reasoning through Navya-Nyaya
von: Sathish, Sharath
Veröffentlicht: (2026)
von: Sathish, Sharath
Veröffentlicht: (2026)
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models
von: Liao, Yi, et al.
Veröffentlicht: (2025)
von: Liao, Yi, et al.
Veröffentlicht: (2025)
Enhance Reasoning for Large Language Models in the Game Werewolf
von: Wu, Shuang, et al.
Veröffentlicht: (2024)
von: Wu, Shuang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Human-Agent Cooperation in Games under Incomplete Information through Natural Language Communication
von: Chen, Shenghui, et al.
Veröffentlicht: (2024) -
Fine-Tuning Language Models Using Formal Methods Feedback
von: Yang, Yunhao, et al.
Veröffentlicht: (2023) -
Multimodal Pretrained Models for Verifiable Sequential Decision-Making: Planning, Grounding, and Perception
von: Yang, Yunhao, et al.
Veröffentlicht: (2023) -
Memorization in Fine-Tuned Large Language Models
von: Savine, Danil
Veröffentlicht: (2025) -
Joint Verification and Refinement of Language Models for Safety-Constrained Planning
von: Yang, Yunhao, et al.
Veröffentlicht: (2024)