Rewarding Creativity: A Human-Aligned Generative Reward Model for Reinforcement Learning in Storytelling
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zhaoyan, Lei, Hang, Wang, Yujia, Liu, Lanbo, Liu, Hao, Yu, Liang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Direct Generation: A Decomposed Approach to Well-Crafted Screenwriting with LLMs
by: Lei, Hang, et al.
Published: (2025)
by: Lei, Hang, et al.
Published: (2025)
RTTC: Reward-Guided Collaborative Test-Time Compute
by: Muñoz, J. Pablo, et al.
Published: (2025)
by: Muñoz, J. Pablo, et al.
Published: (2025)
Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking
by: Singha, Disha
Published: (2026)
by: Singha, Disha
Published: (2026)
Intervention Complexity as a Canonical Reward and a Measure of Intelligence
by: McCane, Brendan
Published: (2026)
by: McCane, Brendan
Published: (2026)
Reward is not enough: can we liberate AI from the reinforcement learning paradigm?
by: Glukhov, Vacslav
Published: (2022)
by: Glukhov, Vacslav
Published: (2022)
In-Context Learning May Not Elicit Trustworthy Reasoning: A-Not-B Errors in Pretrained Language Models
by: Han, Pengrui, et al.
Published: (2024)
by: Han, Pengrui, et al.
Published: (2024)
Fanar: An Arabic-Centric Multimodal Generative AI Platform
by: Fanar Team, et al.
Published: (2025)
by: Fanar Team, et al.
Published: (2025)
Next Token Prediction Is a Dead End for Creativity
by: Olatunji, Ibukun, et al.
Published: (2025)
by: Olatunji, Ibukun, et al.
Published: (2025)
ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models
by: Dumitru, Razvan-Gabriel, et al.
Published: (2025)
by: Dumitru, Razvan-Gabriel, et al.
Published: (2025)
OG-RAG: Ontology-Grounded Retrieval-Augmented Generation For Large Language Models
by: Sharma, Kartik, et al.
Published: (2024)
by: Sharma, Kartik, et al.
Published: (2024)
Generating Causal Explanations of Vehicular Agent Behavioural Interactions with Learnt Reward Profiles
by: Howard, Rhys, et al.
Published: (2025)
by: Howard, Rhys, et al.
Published: (2025)
From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments
by: Luo, Lijing, et al.
Published: (2026)
by: Luo, Lijing, et al.
Published: (2026)
Modeling Emotions and Ethics with Large Language Models
by: Chang, Edward Y.
Published: (2024)
by: Chang, Edward Y.
Published: (2024)
ConCISE: A Reference-Free Conciseness Evaluation Metric for LLM-Generated Answers
by: Ghafari, Seyed Mohssen, et al.
Published: (2025)
by: Ghafari, Seyed Mohssen, et al.
Published: (2025)
Reducing Selection Bias in Large Language Models
by: Eicher, J. E., et al.
Published: (2024)
by: Eicher, J. E., et al.
Published: (2024)
Quo Vadis ChatGPT? From Large Language Models to Large Knowledge Models
by: Venkatasubramanian, Venkat, et al.
Published: (2024)
by: Venkatasubramanian, Venkat, et al.
Published: (2024)
Reward Machines for Deep RL in Noisy and Uncertain Environments
by: Li, Andrew C., et al.
Published: (2024)
by: Li, Andrew C., et al.
Published: (2024)
Training Language Models to Win Debates with Self-Play Improves Judge Accuracy
by: Arnesen, Samuel, et al.
Published: (2024)
by: Arnesen, Samuel, et al.
Published: (2024)
Revisiting Parameter-Based Knowledge Editing in Large Language Models: Theoretical Limits and Empirical Evidence
by: Ren, Wanying, et al.
Published: (2026)
by: Ren, Wanying, et al.
Published: (2026)
How much can change in a year? Revisiting Evaluation in Multi-Agent Reinforcement Learning
by: Singh, Siddarth, et al.
Published: (2023)
by: Singh, Siddarth, et al.
Published: (2023)
Box Maze: A Process-Control Architecture for Reliable LLM Reasoning
by: Qiang, Zou
Published: (2026)
by: Qiang, Zou
Published: (2026)
Cognitively Inspired Components for Social Conversational Agents
by: Clay, Alex, et al.
Published: (2023)
by: Clay, Alex, et al.
Published: (2023)
CuentosIE: can a chatbot about "tales with a message" help to teach emotional intelligence?
by: Ferrández, Antonio, et al.
Published: (2024)
by: Ferrández, Antonio, et al.
Published: (2024)
FedLoGe: Joint Local and Generic Federated Learning under Long-tailed Data
by: Xiao, Zikai, et al.
Published: (2024)
by: Xiao, Zikai, et al.
Published: (2024)
BabyReasoningBench: Generating Developmentally-Inspired Reasoning Tasks for Evaluating Baby Language Models
by: Dhole, Kaustubh D.
Published: (2026)
by: Dhole, Kaustubh D.
Published: (2026)
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization
by: Khanna, Danush, et al.
Published: (2025)
by: Khanna, Danush, et al.
Published: (2025)
OpenAI Cribbed Our Tax Example, But Can GPT-4 Really Do Tax?
by: Blair-Stanek, Andrew, et al.
Published: (2023)
by: Blair-Stanek, Andrew, et al.
Published: (2023)
HCAST: Human-Calibrated Autonomy Software Tasks
by: Rein, David, et al.
Published: (2025)
by: Rein, David, et al.
Published: (2025)
Noisy Data is Destructive to Reinforcement Learning with Verifiable Rewards
by: Zhu, Yuxuan, et al.
Published: (2026)
by: Zhu, Yuxuan, et al.
Published: (2026)
Generative AI for Synthetic Data Generation: Methods, Challenges and the Future
by: Guo, Xu, et al.
Published: (2024)
by: Guo, Xu, et al.
Published: (2024)
Sensemaking in Novel Environments: How Human Cognition Can Inform Artificial Agents
by: Patterson, Robert E., et al.
Published: (2025)
by: Patterson, Robert E., et al.
Published: (2025)
Evaluating Large Language Models for Causal Modeling
by: Razouk, Houssam, et al.
Published: (2024)
by: Razouk, Houssam, et al.
Published: (2024)
Deep Reinforcement Learning for Adverse Garage Scenario Generation
by: Li, Kai
Published: (2024)
by: Li, Kai
Published: (2024)
ARCTraj: A Dataset and Benchmark of Human Reasoning Trajectories for Abstract Problem Solving
by: Kim, Sejin, et al.
Published: (2025)
by: Kim, Sejin, et al.
Published: (2025)
Learning from a Generative AI Predecessor -- The Many Motivations for Interacting with Conversational Agents
by: Brinkman, Donald, et al.
Published: (2023)
by: Brinkman, Donald, et al.
Published: (2023)
Machine Learning and Theory Ladenness -- A Phenomenological Account
by: Termine, Alberto, et al.
Published: (2024)
by: Termine, Alberto, et al.
Published: (2024)
Interpretability-Guided Bi-objective Optimization: Aligning Accuracy and Explainability
by: Fouladi, Kasra, et al.
Published: (2026)
by: Fouladi, Kasra, et al.
Published: (2026)
Mamba-Shedder: Post-Transformer Compression for Efficient Selective Structured State Space Models
by: Muñoz, J. Pablo, et al.
Published: (2025)
by: Muñoz, J. Pablo, et al.
Published: (2025)
Mutagenesis screen to map the functions of parameters of Large Language Models
by: Hu, Yue, et al.
Published: (2024)
by: Hu, Yue, et al.
Published: (2024)
RWKV-7 "Goose" with Expressive Dynamic State Evolution
by: Peng, Bo, et al.
Published: (2025)
by: Peng, Bo, et al.
Published: (2025)
Similar Items
-
Beyond Direct Generation: A Decomposed Approach to Well-Crafted Screenwriting with LLMs
by: Lei, Hang, et al.
Published: (2025) -
RTTC: Reward-Guided Collaborative Test-Time Compute
by: Muñoz, J. Pablo, et al.
Published: (2025) -
Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking
by: Singha, Disha
Published: (2026) -
Intervention Complexity as a Canonical Reward and a Measure of Intelligence
by: McCane, Brendan
Published: (2026) -
Reward is not enough: can we liberate AI from the reinforcement learning paradigm?
by: Glukhov, Vacslav
Published: (2022)