Efficient RL for optimizing conversation level outcomes with an LLM-based tutor
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nam, Hyunji, Gottesman, Omer, Zhang, Amy, Foster, Dean, Brunskill, Emma, Ungar, Lyle |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mitigating LLM biases toward spurious social contexts using direct preference optimization
von: Nam, Hyunji, et al.
Veröffentlicht: (2026)
von: Nam, Hyunji, et al.
Veröffentlicht: (2026)
Closing the Confidence-Faithfulness Gap in Large Language Models
von: Miao, Miranda Muqing, et al.
Veröffentlicht: (2026)
von: Miao, Miranda Muqing, et al.
Veröffentlicht: (2026)
Maximizing Mutual Information Between Prompt and Response Improves LLM Performance With No Additional Data
von: Nam, Hyunji, et al.
Veröffentlicht: (2026)
von: Nam, Hyunji, et al.
Veröffentlicht: (2026)
Evaluating and Optimizing Educational Content with Large Language Model Judgments
von: He-Yueya, Joy, et al.
Veröffentlicht: (2024)
von: He-Yueya, Joy, et al.
Veröffentlicht: (2024)
The Impact of Language Mixing on Bilingual LLM Reasoning
von: Li, Yihao, et al.
Veröffentlicht: (2025)
von: Li, Yihao, et al.
Veröffentlicht: (2025)
Predicting Long Term Sequential Policy Value Using Softer Surrogates
von: Nam, Hyunji, et al.
Veröffentlicht: (2024)
von: Nam, Hyunji, et al.
Veröffentlicht: (2024)
Dynamic benchmarking framework for LLM-based conversational data capture
von: Aluffi, Pietro Alessandro, et al.
Veröffentlicht: (2025)
von: Aluffi, Pietro Alessandro, et al.
Veröffentlicht: (2025)
ZeroTuning: Unlocking the Initial Token's Power to Enhance Large Language Models Without Training
von: Han, Feijiang, et al.
Veröffentlicht: (2025)
von: Han, Feijiang, et al.
Veröffentlicht: (2025)
Psychometric Alignment: Capturing Human Knowledge Distributions via Language Models
von: He-Yueya, Joy, et al.
Veröffentlicht: (2024)
von: He-Yueya, Joy, et al.
Veröffentlicht: (2024)
When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks?
von: Hatgis-Kessell, Stephane, et al.
Veröffentlicht: (2026)
von: Hatgis-Kessell, Stephane, et al.
Veröffentlicht: (2026)
Exploring the generalization of LLM truth directions on conversational formats
von: Ichmoukhamedov, Timour, et al.
Veröffentlicht: (2025)
von: Ichmoukhamedov, Timour, et al.
Veröffentlicht: (2025)
GIANTS: Generative Insight Anticipation from Scientific Literature
von: He-Yueya, Joy, et al.
Veröffentlicht: (2026)
von: He-Yueya, Joy, et al.
Veröffentlicht: (2026)
PsychAdapter: Adapting LLM Transformers to Reflect Traits, Personality and Mental Health
von: Vu, Huy, et al.
Veröffentlicht: (2024)
von: Vu, Huy, et al.
Veröffentlicht: (2024)
CORG: Generating Answers from Complex, Interrelated Contexts
von: Lee, Hyunji, et al.
Veröffentlicht: (2025)
von: Lee, Hyunji, et al.
Veröffentlicht: (2025)
Math Neurosurgery: Isolating Language Models' Math Reasoning Abilities Using Only Forward Passes
von: Christ, Bryan R., et al.
Veröffentlicht: (2024)
von: Christ, Bryan R., et al.
Veröffentlicht: (2024)
SafeCtrl-RL: Inference-Time Adaptive Behaviour Control for LLM Dialogue via RL-Driven Prompt Optimisation
von: Orme, Michael, et al.
Veröffentlicht: (2026)
von: Orme, Michael, et al.
Veröffentlicht: (2026)
Miner:Mining Intrinsic Mastery for Data-Efficient RL in Large Reasoning Models
von: Jiang, Shuyang, et al.
Veröffentlicht: (2026)
von: Jiang, Shuyang, et al.
Veröffentlicht: (2026)
Reasoning Core: A Scalable RL Environment for LLM Symbolic Reasoning
von: Lacombe, Valentin, et al.
Veröffentlicht: (2025)
von: Lacombe, Valentin, et al.
Veröffentlicht: (2025)
Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning
von: Xie, Tian, et al.
Veröffentlicht: (2025)
von: Xie, Tian, et al.
Veröffentlicht: (2025)
Improving LLM Reliability with RAG in Religious Question-Answering: MufassirQAS
von: Alan, Ahmet Yusuf, et al.
Veröffentlicht: (2024)
von: Alan, Ahmet Yusuf, et al.
Veröffentlicht: (2024)
Dopamin: Transformer-based Comment Classifiers through Domain Post-Training and Multi-level Layer Aggregation
von: Hai, Nam Le, et al.
Veröffentlicht: (2024)
von: Hai, Nam Le, et al.
Veröffentlicht: (2024)
ReviewRL: Towards Automated Scientific Review with RL
von: Zeng, Sihang, et al.
Veröffentlicht: (2025)
von: Zeng, Sihang, et al.
Veröffentlicht: (2025)
Exploring RL-based LLM Training for Formal Language Tasks with Programmed Rewards
von: Padula, Alexander G., et al.
Veröffentlicht: (2024)
von: Padula, Alexander G., et al.
Veröffentlicht: (2024)
FlowRL: Matching Reward Distributions for LLM Reasoning
von: Zhu, Xuekai, et al.
Veröffentlicht: (2025)
von: Zhu, Xuekai, et al.
Veröffentlicht: (2025)
DetectRL: Benchmarking LLM-Generated Text Detection in Real-World Scenarios
von: Wu, Junchao, et al.
Veröffentlicht: (2024)
von: Wu, Junchao, et al.
Veröffentlicht: (2024)
Efficient Prompting for LLM-based Generative Internet of Things
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
Discerning minds or generic tutors? Evaluating instructional guidance capabilities in Socratic LLMs
von: Liu, Ying, et al.
Veröffentlicht: (2025)
von: Liu, Ying, et al.
Veröffentlicht: (2025)
MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent
von: Yu, Hongli, et al.
Veröffentlicht: (2025)
von: Yu, Hongli, et al.
Veröffentlicht: (2025)
A review on the use of large language models as virtual tutors
von: García-Méndez, Silvia, et al.
Veröffentlicht: (2024)
von: García-Méndez, Silvia, et al.
Veröffentlicht: (2024)
The GPT Surprise: Offering Large Language Model Chat in a Massive Coding Class Reduced Engagement but Increased Adopters Exam Performances
von: Nie, Allen, et al.
Veröffentlicht: (2024)
von: Nie, Allen, et al.
Veröffentlicht: (2024)
Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL
von: Liu, Che, et al.
Veröffentlicht: (2025)
von: Liu, Che, et al.
Veröffentlicht: (2025)
Scaling LLM Multi-turn RL with End-to-end Summarization-based Context Management
von: Lu, Miao, et al.
Veröffentlicht: (2025)
von: Lu, Miao, et al.
Veröffentlicht: (2025)
Clinical Note Bloat Reduction for Efficient LLM Use
von: Cahoon, Jordan L., et al.
Veröffentlicht: (2026)
von: Cahoon, Jordan L., et al.
Veröffentlicht: (2026)
Generalized Correctness Models: Learning Calibrated and Model-Agnostic Correctness Predictors from Historical Patterns
von: Xiao, Hanqi, et al.
Veröffentlicht: (2025)
von: Xiao, Hanqi, et al.
Veröffentlicht: (2025)
Large Language Models Show Human-like Social Desirability Biases in Survey Responses
von: Salecha, Aadesh, et al.
Veröffentlicht: (2024)
von: Salecha, Aadesh, et al.
Veröffentlicht: (2024)
Efficient and Accurate Memorable Conversation Model using DPO based on sLLM
von: Seo, Youngkyung, et al.
Veröffentlicht: (2024)
von: Seo, Youngkyung, et al.
Veröffentlicht: (2024)
Balancing Faithfulness and Performance in Reasoning via Multi-Listener Soft Execution
von: Sivakumaran, Nithin, et al.
Veröffentlicht: (2026)
von: Sivakumaran, Nithin, et al.
Veröffentlicht: (2026)
Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels
von: Cen, Zhepeng, et al.
Veröffentlicht: (2025)
von: Cen, Zhepeng, et al.
Veröffentlicht: (2025)
Beyond Markovian: Reflective Exploration via Bayes-Adaptive RL for LLM Reasoning
von: Zhang, Shenao, et al.
Veröffentlicht: (2025)
von: Zhang, Shenao, et al.
Veröffentlicht: (2025)
Sparse-RL: Breaking the Memory Wall in LLM Reinforcement Learning via Stable Sparse Rollouts
von: Luo, Sijia, et al.
Veröffentlicht: (2026)
von: Luo, Sijia, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Mitigating LLM biases toward spurious social contexts using direct preference optimization
von: Nam, Hyunji, et al.
Veröffentlicht: (2026) -
Closing the Confidence-Faithfulness Gap in Large Language Models
von: Miao, Miranda Muqing, et al.
Veröffentlicht: (2026) -
Maximizing Mutual Information Between Prompt and Response Improves LLM Performance With No Additional Data
von: Nam, Hyunji, et al.
Veröffentlicht: (2026) -
Evaluating and Optimizing Educational Content with Large Language Model Judgments
von: He-Yueya, Joy, et al.
Veröffentlicht: (2024) -
The Impact of Language Mixing on Bilingual LLM Reasoning
von: Li, Yihao, et al.
Veröffentlicht: (2025)