PERSA: Reinforcement Learning for Professor-Style Personalized Feedback with LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Ranjan, Ravi, Grover, Utkarsh, Lin, Xiaomin, Polyzou, Agoritsa |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
G-Drift MIA: Membership Inference via Gradient-Induced Feature Drift in LLMs
by: Ranjan, Ravi, et al.
Published: (2026)
by: Ranjan, Ravi, et al.
Published: (2026)
RAZOR: Ratio-Aware Layer Editing for Targeted Unlearning in Vision Transformers and Diffusion Models
by: Ranjan, Ravi, et al.
Published: (2026)
by: Ranjan, Ravi, et al.
Published: (2026)
CatRAG: Functor-Guided Structural Debiasing with Retrieval Augmentation for Fair LLMs
by: Ranjan, Ravi, et al.
Published: (2026)
by: Ranjan, Ravi, et al.
Published: (2026)
VLA-Forget: Vision-Language-Action Unlearning for Embodied Foundation Models
by: Ranjan, Ravi, et al.
Published: (2026)
by: Ranjan, Ravi, et al.
Published: (2026)
Position: LLMs Must Use Functor-Based and RAG-Driven Bias Mitigation for Fairness
by: Ranjan, Ravi, et al.
Published: (2026)
by: Ranjan, Ravi, et al.
Published: (2026)
Embodied Foundation Models at the Edge: A Survey of Deployment Constraints and Mitigation Strategies
by: Grover, Utkarsh, et al.
Published: (2026)
by: Grover, Utkarsh, et al.
Published: (2026)
How Good Are Large Language Models for Course Recommendation in MOOCs?
by: Ma, Boxuan, et al.
Published: (2025)
by: Ma, Boxuan, et al.
Published: (2025)
FAM-Bench: A Multimodal Benchmark for Condition-Aware Food-as-Medicine Reasoning
by: Mao, Mingyang, et al.
Published: (2026)
by: Mao, Mingyang, et al.
Published: (2026)
Swap-guided Preference Learning for Personalized Reinforcement Learning from Human Feedback
by: Kim, Gihoon, et al.
Published: (2026)
by: Kim, Gihoon, et al.
Published: (2026)
Reinforcement Learning from Multi-role Debates as Feedback for Bias Mitigation in LLMs
by: Cheng, Ruoxi, et al.
Published: (2024)
by: Cheng, Ruoxi, et al.
Published: (2024)
LANPO: Bootstrapping Language and Numerical Feedback for Reinforcement Learning in LLMs
by: Li, Ang, et al.
Published: (2025)
by: Li, Ang, et al.
Published: (2025)
RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning
by: Gehring, Jonas, et al.
Published: (2024)
by: Gehring, Jonas, et al.
Published: (2024)
RLPF: Reinforcement Learning from Prediction Feedback for User Summarization with LLMs
by: Wu, Jiaxing, et al.
Published: (2024)
by: Wu, Jiaxing, et al.
Published: (2024)
Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning
by: Poddar, Sriyash, et al.
Published: (2024)
by: Poddar, Sriyash, et al.
Published: (2024)
Aurora: Neuro-Symbolic AI Driven Advising Agent
by: Lugones, Lorena Amanda Quincoso, et al.
Published: (2026)
by: Lugones, Lorena Amanda Quincoso, et al.
Published: (2026)
Reinforcement Learning from User Feedback
by: Han, Eric, et al.
Published: (2025)
by: Han, Eric, et al.
Published: (2025)
Training LLMs with Reinforcement Learning for Intent-Aware Personalized Question Answering
by: Amirizaniani, Maryam, et al.
Published: (2026)
by: Amirizaniani, Maryam, et al.
Published: (2026)
Privacy Preserving Reinforcement Learning with One-Sided Feedback
by: Cong, Lin William, et al.
Published: (2026)
by: Cong, Lin William, et al.
Published: (2026)
Wisdom of the Crowd: Reinforcement Learning from Coevolutionary Collective Feedback
by: Yuan, Wenzhen, et al.
Published: (2025)
by: Yuan, Wenzhen, et al.
Published: (2025)
Enhancing LLMs for Physics Problem-Solving using Reinforcement Learning with Human-AI Feedback
by: Anand, Avinash, et al.
Published: (2024)
by: Anand, Avinash, et al.
Published: (2024)
Curriculum-RLAIF: Curriculum Alignment with Reinforcement Learning from AI Feedback
by: Lin, Jiaye, et al.
Published: (2025)
by: Lin, Jiaye, et al.
Published: (2025)
LLMSense: Harnessing LLMs for High-level Reasoning Over Spatiotemporal Sensor Traces
by: Ouyang, Xiaomin, et al.
Published: (2024)
by: Ouyang, Xiaomin, et al.
Published: (2024)
Personalizing LLMs with Binary Feedback: A Preference-Corrected Optimization Framework
by: Ma, Xilai, et al.
Published: (2026)
by: Ma, Xilai, et al.
Published: (2026)
Prompted Policy Search: Reinforcement Learning through Linguistic and Numerical Reasoning in LLMs
by: Zhou, Yifan, et al.
Published: (2025)
by: Zhou, Yifan, et al.
Published: (2025)
Formalizing Style in Personal Narratives
by: Cortal, Gustave, et al.
Published: (2025)
by: Cortal, Gustave, et al.
Published: (2025)
Styles + Persona-plug = Customized LLMs
by: Song, Yutong, et al.
Published: (2026)
by: Song, Yutong, et al.
Published: (2026)
Orcust: Stepwise-Feedback Reinforcement Learning for GUI Agent
by: Lu, Junyu, et al.
Published: (2025)
by: Lu, Junyu, et al.
Published: (2025)
Navigating Noisy Feedback: Enhancing Reinforcement Learning with Error-Prone Language Models
by: Lin, Muhan, et al.
Published: (2024)
by: Lin, Muhan, et al.
Published: (2024)
RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs
by: Chaudhari, Shreyas, et al.
Published: (2024)
by: Chaudhari, Shreyas, et al.
Published: (2024)
Peering Through Preferences: Unraveling Feedback Acquisition for Aligning Large Language Models
by: Bansal, Hritik, et al.
Published: (2023)
by: Bansal, Hritik, et al.
Published: (2023)
Missing Data Imputation With Granular Semantics and AI-driven Pipeline for Bankruptcy Prediction
by: Chakraborty, Debarati, et al.
Published: (2024)
by: Chakraborty, Debarati, et al.
Published: (2024)
Reinforcement Learning with Backtracking Feedback
by: Sel, Bilgehan, et al.
Published: (2026)
by: Sel, Bilgehan, et al.
Published: (2026)
RLAE: Reinforcement Learning-Assisted Ensemble for LLMs
by: Fu, Yuqian, et al.
Published: (2025)
by: Fu, Yuqian, et al.
Published: (2025)
Reinforcement Learning from Reflective Feedback (RLRF): Aligning and Improving LLMs via Fine-Grained Self-Reflection
by: Lee, Kyungjae, et al.
Published: (2024)
by: Lee, Kyungjae, et al.
Published: (2024)
Text2Grad: Reinforcement Learning from Natural Language Feedback
by: Wang, Hanyang, et al.
Published: (2025)
by: Wang, Hanyang, et al.
Published: (2025)
RLAF: Reinforcement Learning from Automaton Feedback
by: Alinejad, Mahyar, et al.
Published: (2025)
by: Alinejad, Mahyar, et al.
Published: (2025)
BitRL-Light: 1-bit LLM Agents with Deep Reinforcement Learning for Energy-Efficient Smart Home Lighting Optimization
by: Gupta, Ravi, et al.
Published: (2025)
by: Gupta, Ravi, et al.
Published: (2025)
ICRL: Learning to Internalize Self-Critique with Reinforcement Learning
by: Lin, Jianbo, et al.
Published: (2026)
by: Lin, Jianbo, et al.
Published: (2026)
Learning Personalized Agents from Human Feedback
by: Liang, Kaiqu, et al.
Published: (2026)
by: Liang, Kaiqu, et al.
Published: (2026)
Emergent Hierarchical Reasoning in LLMs through Reinforcement Learning
by: Wang, Haozhe, et al.
Published: (2025)
by: Wang, Haozhe, et al.
Published: (2025)
Similar Items
-
G-Drift MIA: Membership Inference via Gradient-Induced Feature Drift in LLMs
by: Ranjan, Ravi, et al.
Published: (2026) -
RAZOR: Ratio-Aware Layer Editing for Targeted Unlearning in Vision Transformers and Diffusion Models
by: Ranjan, Ravi, et al.
Published: (2026) -
CatRAG: Functor-Guided Structural Debiasing with Retrieval Augmentation for Fair LLMs
by: Ranjan, Ravi, et al.
Published: (2026) -
VLA-Forget: Vision-Language-Action Unlearning for Embodied Foundation Models
by: Ranjan, Ravi, et al.
Published: (2026) -
Position: LLMs Must Use Functor-Based and RAG-Driven Bias Mitigation for Fairness
by: Ranjan, Ravi, et al.
Published: (2026)