Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Maura-Rivero, Roberto-Rafael, Nagpal, Chirag, Patel, Roma, Visin, Francesco |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Jackpot! Alignment as a Maximal Lottery
von: Maura-Rivero, Roberto-Rafael, et al.
Veröffentlicht: (2025)
von: Maura-Rivero, Roberto-Rafael, et al.
Veröffentlicht: (2025)
Language Models Trained to do Arithmetic Predict Human Risky and Intertemporal Choice
von: Zhu, Jian-Qiao, et al.
Veröffentlicht: (2024)
von: Zhu, Jian-Qiao, et al.
Veröffentlicht: (2024)
Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards
von: Liu, Shuze Daniel, et al.
Veröffentlicht: (2026)
von: Liu, Shuze Daniel, et al.
Veröffentlicht: (2026)
Double Jeopardy and Climate Impact in the Use of Large Language Models: Socio-economic Disparities and Reduced Utility for Non-English Speakers
von: Solatorio, Aivin V., et al.
Veröffentlicht: (2024)
von: Solatorio, Aivin V., et al.
Veröffentlicht: (2024)
Can Large Language Models Replace Human Subjects? A Large-Scale Replication of Scenario-Based Experiments in Psychology and Management
von: Cui, Ziyan, et al.
Veröffentlicht: (2024)
von: Cui, Ziyan, et al.
Veröffentlicht: (2024)
A Unified Framework to Classify Business Activities into International Standard Industrial Classification through Large Language Models for Circular Economy
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
Transforming the Voice of the Customer: Large Language Models for Identifying Customer Needs
von: Timoshenko, Artem, et al.
Veröffentlicht: (2025)
von: Timoshenko, Artem, et al.
Veröffentlicht: (2025)
Integrating Natural Language Processing Techniques of Text Mining Into Financial System: Applications and Limitations
von: Millo, Denisa, et al.
Veröffentlicht: (2024)
von: Millo, Denisa, et al.
Veröffentlicht: (2024)
Think, Speak, Decide: Language-Augmented Multi-Agent Reinforcement Learning for Economic Decision-Making
von: Ma, Heyang, et al.
Veröffentlicht: (2025)
von: Ma, Heyang, et al.
Veröffentlicht: (2025)
LLMs Model Non-WEIRD Populations: Experiments with Synthetic Cultural Agents
von: Gonzalez-Bonorino, Augusto, et al.
Veröffentlicht: (2025)
von: Gonzalez-Bonorino, Augusto, et al.
Veröffentlicht: (2025)
Approximating Auction Equilibria with Reinforcement Learning
von: Rawat, Pranjal
Veröffentlicht: (2024)
von: Rawat, Pranjal
Veröffentlicht: (2024)
Dynamic Reinsurance Treaty Bidding via Multi-Agent Reinforcement Learning
von: Dong, Stella C., et al.
Veröffentlicht: (2025)
von: Dong, Stella C., et al.
Veröffentlicht: (2025)
Leveraging AI and NLP for Bank Marketing: A Systematic Review and Gap Analysis
von: Gerling, Christopher, et al.
Veröffentlicht: (2024)
von: Gerling, Christopher, et al.
Veröffentlicht: (2024)
From Transcripts to Insights: Uncovering Corporate Risks Using Generative AI
von: Kim, Alex, et al.
Veröffentlicht: (2023)
von: Kim, Alex, et al.
Veröffentlicht: (2023)
AI evaluation may bias perceptions: The importance of context in interpreting academic writing
von: Wu, Shang, et al.
Veröffentlicht: (2026)
von: Wu, Shang, et al.
Veröffentlicht: (2026)
Divergent LLM Adoption and Heterogeneous Convergence Paths in Research Writing
von: Lin, Cong William, et al.
Veröffentlicht: (2025)
von: Lin, Cong William, et al.
Veröffentlicht: (2025)
Anchors in the Machine: Behavioral and Attributional Evidence of Anchoring Bias in LLMs
von: Valencia-Clavijo, Felipe
Veröffentlicht: (2025)
von: Valencia-Clavijo, Felipe
Veröffentlicht: (2025)
Using Artificial Intelligence to Unlock Crowdfunding Success for Small Businesses
von: Ye, Teng, et al.
Veröffentlicht: (2024)
von: Ye, Teng, et al.
Veröffentlicht: (2024)
AI Patents in the United States and China: Measurement, Organization, and Knowledge Flows
von: Fang, Hanming, et al.
Veröffentlicht: (2026)
von: Fang, Hanming, et al.
Veröffentlicht: (2026)
Calibrating Behavioral Parameters with Large Language Models
von: Yee, Brandon, et al.
Veröffentlicht: (2026)
von: Yee, Brandon, et al.
Veröffentlicht: (2026)
Reproducing and Extending Experiments in Behavioral Strategy with Large Language Models
von: Albert, Daniel, et al.
Veröffentlicht: (2024)
von: Albert, Daniel, et al.
Veröffentlicht: (2024)
Training for Obsolescence? The AI-Driven Education Trap
von: Peterson, Andrew J.
Veröffentlicht: (2025)
von: Peterson, Andrew J.
Veröffentlicht: (2025)
Beyond Code: The Multidimensional Impacts of Large Language Models in Software Development
von: Bonabi, Sardar, et al.
Veröffentlicht: (2025)
von: Bonabi, Sardar, et al.
Veröffentlicht: (2025)
Reinforcement Learning and Consumption-Savings Behavior
von: Kaplowitz, Brandon
Veröffentlicht: (2025)
von: Kaplowitz, Brandon
Veröffentlicht: (2025)
EconAgentic in DePIN Markets: A Large Language Model Approach to the Sharing Economy of Decentralized Physical Infrastructure
von: Liu, Yulin, et al.
Veröffentlicht: (2025)
von: Liu, Yulin, et al.
Veröffentlicht: (2025)
AI Skills Improve Job Prospects: Causal Evidence from a Hiring Experiment
von: Stephany, Fabian, et al.
Veröffentlicht: (2026)
von: Stephany, Fabian, et al.
Veröffentlicht: (2026)
The Economics of p(doom): Scenarios of Existential Risk and Economic Growth in the Age of Transformative AI
von: Growiec, Jakub, et al.
Veröffentlicht: (2025)
von: Growiec, Jakub, et al.
Veröffentlicht: (2025)
Orchestrating Rewards in the Era of Intelligence-Driven Commerce
von: Oamen, Paul Osemudiame, et al.
Veröffentlicht: (2025)
von: Oamen, Paul Osemudiame, et al.
Veröffentlicht: (2025)
Carrot, stick, or both? Price incentives for sustainable food choice in competitive environments
von: Salvi, Francesco, et al.
Veröffentlicht: (2025)
von: Salvi, Francesco, et al.
Veröffentlicht: (2025)
From Individual Learning to Market Equilibrium: Correcting Structural and Parametric Biases in RL Simulations of Economic Models
von: Chen, Ruxin, et al.
Veröffentlicht: (2025)
von: Chen, Ruxin, et al.
Veröffentlicht: (2025)
The AI Productivity Index (APEX)
von: Vidgen, Bertie, et al.
Veröffentlicht: (2025)
von: Vidgen, Bertie, et al.
Veröffentlicht: (2025)
AI Realtor: Towards Grounded Persuasive Language Generation for Automated Copywriting
von: Wu, Jibang, et al.
Veröffentlicht: (2025)
von: Wu, Jibang, et al.
Veröffentlicht: (2025)
Quantifying Systemic Vulnerability in the Foundation Model Industry
von: Pirrone, Claudio, et al.
Veröffentlicht: (2025)
von: Pirrone, Claudio, et al.
Veröffentlicht: (2025)
AI Safety Should Prioritize the Future of Work
von: Hazra, Sanchaita, et al.
Veröffentlicht: (2025)
von: Hazra, Sanchaita, et al.
Veröffentlicht: (2025)
Modelling of Economic Implications of Bias in AI-Powered Health Emergency Response Systems
von: Bahamazava, Katsiaryna
Veröffentlicht: (2024)
von: Bahamazava, Katsiaryna
Veröffentlicht: (2024)
AI-Driven Scenarios for Urban Mobility: Quantifying the Role of ODE Models and Scenario Planning in Reducing Traffic Congestion
von: Bahamazava, Katsiaryna
Veröffentlicht: (2024)
von: Bahamazava, Katsiaryna
Veröffentlicht: (2024)
Unveiling the Role of Artificial Intelligence and Stock Market Growth in Achieving Carbon Neutrality in the United States: An ARDL Model Analysis
von: Rafi, Azizul Hakim, et al.
Veröffentlicht: (2024)
von: Rafi, Azizul Hakim, et al.
Veröffentlicht: (2024)
Is there "Secret Sauce'' in Large Language Model Development?
von: Mertens, Matthias, et al.
Veröffentlicht: (2026)
von: Mertens, Matthias, et al.
Veröffentlicht: (2026)
Large Language Models at Work in China's Labor Market
von: Chen, Qin, et al.
Veröffentlicht: (2023)
von: Chen, Qin, et al.
Veröffentlicht: (2023)
Improving Task Instructions for Data Annotators: How Clear Rules and Higher Pay Increase Performance in Data Annotation in the AI Economy
von: Laux, Johann, et al.
Veröffentlicht: (2023)
von: Laux, Johann, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Jackpot! Alignment as a Maximal Lottery
von: Maura-Rivero, Roberto-Rafael, et al.
Veröffentlicht: (2025) -
Language Models Trained to do Arithmetic Predict Human Risky and Intertemporal Choice
von: Zhu, Jian-Qiao, et al.
Veröffentlicht: (2024) -
Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards
von: Liu, Shuze Daniel, et al.
Veröffentlicht: (2026) -
Double Jeopardy and Climate Impact in the Use of Large Language Models: Socio-economic Disparities and Reduced Utility for Non-English Speakers
von: Solatorio, Aivin V., et al.
Veröffentlicht: (2024) -
Can Large Language Models Replace Human Subjects? A Large-Scale Replication of Scenario-Based Experiments in Psychology and Management
von: Cui, Ziyan, et al.
Veröffentlicht: (2024)