Automated Rewards via LLM-Generated Progress Functions
Fuente:
arXiv
Salvato in:
| Autori principali: | Sarukkai, Vishnu, Shacklett, Brennan, Majercik, Zander, Bhatia, Kush, Ré, Christopher, Fatahalian, Kayvon |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Self-Generated In-Context Examples Improve LLM Agents for Sequential Decision-Making Tasks
di: Sarukkai, Vishnu, et al.
Pubblicazione: (2025)
di: Sarukkai, Vishnu, et al.
Pubblicazione: (2025)
Learning to Move Like Professional Counter-Strike Players
di: Durst, David, et al.
Pubblicazione: (2024)
di: Durst, David, et al.
Pubblicazione: (2024)
Learning to Ball: Composing Policies for Long-Horizon Basketball Moves
di: Xu, Pei, et al.
Pubblicazione: (2025)
di: Xu, Pei, et al.
Pubblicazione: (2025)
Inference-Time Distillation: Cost-Efficient Agents Without Fine-Tuning or Manual Prompt Engineering
di: Sarukkai, Vishnu, et al.
Pubblicazione: (2025)
di: Sarukkai, Vishnu, et al.
Pubblicazione: (2025)
Cookbook: A framework for improving LLM generative abilities via programmatic data generating templates
di: Narayan, Avanika, et al.
Pubblicazione: (2024)
di: Narayan, Avanika, et al.
Pubblicazione: (2024)
The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry
di: Zhang, Michael, et al.
Pubblicazione: (2024)
di: Zhang, Michael, et al.
Pubblicazione: (2024)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
di: Alam, Firoj, et al.
Pubblicazione: (2026)
di: Alam, Firoj, et al.
Pubblicazione: (2026)
Consistency Is the Key: Detecting Hallucinations in LLM Generated Text By Checking Inconsistencies About Key Facts
di: Gupta, Raavi, et al.
Pubblicazione: (2025)
di: Gupta, Raavi, et al.
Pubblicazione: (2025)
DreamPRM-Code: Function-as-Step Process Reward Model with Label Correction for LLM Coding
di: Zhang, Ruiyi, et al.
Pubblicazione: (2025)
di: Zhang, Ruiyi, et al.
Pubblicazione: (2025)
Selective Preference Optimization via Token-Level Reward Function Estimation
di: Yang, Kailai, et al.
Pubblicazione: (2024)
di: Yang, Kailai, et al.
Pubblicazione: (2024)
ProgRM: Build Better GUI Agents with Progress Rewards
di: Zhang, Danyang, et al.
Pubblicazione: (2025)
di: Zhang, Danyang, et al.
Pubblicazione: (2025)
Block and Detail: Scaffolding Sketch-to-Image Generation
di: Sarukkai, Vishnu, et al.
Pubblicazione: (2024)
di: Sarukkai, Vishnu, et al.
Pubblicazione: (2024)
Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences
di: Karlekar, Sweta, et al.
Pubblicazione: (2026)
di: Karlekar, Sweta, et al.
Pubblicazione: (2026)
Multi-LLM QA with Embodied Exploration
di: Patel, Bhrij, et al.
Pubblicazione: (2024)
di: Patel, Bhrij, et al.
Pubblicazione: (2024)
Automated Text Scoring in the Age of Generative AI for the GPU-poor
di: Ormerod, Christopher Michael, et al.
Pubblicazione: (2024)
di: Ormerod, Christopher Michael, et al.
Pubblicazione: (2024)
FlowRL: Matching Reward Distributions for LLM Reasoning
di: Zhu, Xuekai, et al.
Pubblicazione: (2025)
di: Zhu, Xuekai, et al.
Pubblicazione: (2025)
Using Large Language Models to Automate and Expedite Reinforcement Learning with Reward Machine
di: Alsadat, Shayan Meshkat, et al.
Pubblicazione: (2024)
di: Alsadat, Shayan Meshkat, et al.
Pubblicazione: (2024)
Functional Entropy: Predicting Functional Correctness in LLM-Generated Code with Uncertainty Quantification
di: Bouchard, Dylan, et al.
Pubblicazione: (2026)
di: Bouchard, Dylan, et al.
Pubblicazione: (2026)
On Designing Effective RL Reward at Training Time for LLM Reasoning
di: Gao, Jiaxuan, et al.
Pubblicazione: (2024)
di: Gao, Jiaxuan, et al.
Pubblicazione: (2024)
ReDit: Reward Dithering for Improved LLM Policy Optimization
di: Wei, Chenxing, et al.
Pubblicazione: (2025)
di: Wei, Chenxing, et al.
Pubblicazione: (2025)
Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM Alignment
di: Cheng, Ruoxi, et al.
Pubblicazione: (2025)
di: Cheng, Ruoxi, et al.
Pubblicazione: (2025)
Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents
di: Xia, Fanzeng, et al.
Pubblicazione: (2024)
di: Xia, Fanzeng, et al.
Pubblicazione: (2024)
AI Metropolis: Scaling Large Language Model-based Multi-Agent Simulation with Out-of-order Execution
di: Xie, Zhiqiang, et al.
Pubblicazione: (2024)
di: Xie, Zhiqiang, et al.
Pubblicazione: (2024)
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training
di: Liang, Yu, et al.
Pubblicazione: (2026)
di: Liang, Yu, et al.
Pubblicazione: (2026)
Mitigating Spurious Correlations in NLI via LLM-Synthesized Counterfactuals and Dynamic Balanced Sampling
di: Jaimes, Christopher Román
Pubblicazione: (2025)
di: Jaimes, Christopher Román
Pubblicazione: (2025)
On the Effect of Instruction Tuning Loss on Generalization
di: Chatterjee, Anwoy, et al.
Pubblicazione: (2025)
di: Chatterjee, Anwoy, et al.
Pubblicazione: (2025)
Multilinguality in LLM-Designed Reward Functions for Restless Bandits: Effects on Task Performance and Fairness
di: Parthasarathy, Ambreesh, et al.
Pubblicazione: (2025)
di: Parthasarathy, Ambreesh, et al.
Pubblicazione: (2025)
Distribution-Aware Reward: Reinforcement Learning over Predictive Distributions for LLM Regression
di: Park, Jungsoo, et al.
Pubblicazione: (2026)
di: Park, Jungsoo, et al.
Pubblicazione: (2026)
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment
di: Zhang, Jiazheng, et al.
Pubblicazione: (2025)
di: Zhang, Jiazheng, et al.
Pubblicazione: (2025)
Exploring RL-based LLM Training for Formal Language Tasks with Programmed Rewards
di: Padula, Alexander G., et al.
Pubblicazione: (2024)
di: Padula, Alexander G., et al.
Pubblicazione: (2024)
Aioli: A Unified Optimization Framework for Language Model Data Mixing
di: Chen, Mayee F., et al.
Pubblicazione: (2024)
di: Chen, Mayee F., et al.
Pubblicazione: (2024)
AgentRM: Enhancing Agent Generalization with Reward Modeling
di: Xia, Yu, et al.
Pubblicazione: (2025)
di: Xia, Yu, et al.
Pubblicazione: (2025)
TrimLLM: Progressive Layer Dropping for Domain-Specific LLMs
di: Hu, Lanxiang, et al.
Pubblicazione: (2024)
di: Hu, Lanxiang, et al.
Pubblicazione: (2024)
Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization
di: Li, Chenliang, et al.
Pubblicazione: (2025)
di: Li, Chenliang, et al.
Pubblicazione: (2025)
Mining Intrinsic Rewards from LLM Hidden States for Efficient Best-of-N Sampling
di: Guo, Jizhou, et al.
Pubblicazione: (2025)
di: Guo, Jizhou, et al.
Pubblicazione: (2025)
Reward Shaping to Mitigate Reward Hacking in RLHF
di: Fu, Jiayi, et al.
Pubblicazione: (2025)
di: Fu, Jiayi, et al.
Pubblicazione: (2025)
COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation
di: Zhou, Tianyi, et al.
Pubblicazione: (2026)
di: Zhou, Tianyi, et al.
Pubblicazione: (2026)
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
di: Yang, Wenkai, et al.
Pubblicazione: (2026)
di: Yang, Wenkai, et al.
Pubblicazione: (2026)
Multi-Turn Code Generation Through Single-Step Rewards
di: Jain, Arnav Kumar, et al.
Pubblicazione: (2025)
di: Jain, Arnav Kumar, et al.
Pubblicazione: (2025)
From Faithfulness to Correctness: Generative Reward Models that Think Critically
di: Ma, Qiyao, et al.
Pubblicazione: (2025)
di: Ma, Qiyao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Self-Generated In-Context Examples Improve LLM Agents for Sequential Decision-Making Tasks
di: Sarukkai, Vishnu, et al.
Pubblicazione: (2025) -
Learning to Move Like Professional Counter-Strike Players
di: Durst, David, et al.
Pubblicazione: (2024) -
Learning to Ball: Composing Policies for Long-Horizon Basketball Moves
di: Xu, Pei, et al.
Pubblicazione: (2025) -
Inference-Time Distillation: Cost-Efficient Agents Without Fine-Tuning or Manual Prompt Engineering
di: Sarukkai, Vishnu, et al.
Pubblicazione: (2025) -
Cookbook: A framework for improving LLM generative abilities via programmatic data generating templates
di: Narayan, Avanika, et al.
Pubblicazione: (2024)