Inference-Time Personalized Alignment with a Few User Preference Queries
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pădurean, Victor-Alexandru, Kamalaruban, Parameswaran, Kotalwar, Nachiket, Gotovos, Alkis, Singla, Adish |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hints-In-Browser: Benchmarking Language Models for Programming Feedback Generation
von: Kotalwar, Nachiket, et al.
Veröffentlicht: (2024)
von: Kotalwar, Nachiket, et al.
Veröffentlicht: (2024)
Prompt Programming: A Platform for Dialogue-based Computational Problem Solving with Generative AI Models
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2025)
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2025)
Curriculum Design for Trajectory-Constrained Agent: Compressing Chain-of-Thought Tokens in LLMs
von: Tzannetos, Georgios, et al.
Veröffentlicht: (2025)
von: Tzannetos, Georgios, et al.
Veröffentlicht: (2025)
Informativeness of Reward Functions in Reinforcement Learning
von: Devidze, Rati, et al.
Veröffentlicht: (2024)
von: Devidze, Rati, et al.
Veröffentlicht: (2024)
Proximal Curriculum with Task Correlations for Deep Reinforcement Learning
von: Tzannetos, Georgios, et al.
Veröffentlicht: (2024)
von: Tzannetos, Georgios, et al.
Veröffentlicht: (2024)
Neural Task Synthesis for Visual Programming
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2023)
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2023)
Synthesizing High-Quality Programming Tasks with LLM-based Expert and Student Agents
von: Nguyen, Manh Hung, et al.
Veröffentlicht: (2025)
von: Nguyen, Manh Hung, et al.
Veröffentlicht: (2025)
Policy Teaching via Data Poisoning in Learning from Human Preferences
von: Nika, Andi, et al.
Veröffentlicht: (2025)
von: Nika, Andi, et al.
Veröffentlicht: (2025)
Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences
von: Nika, Andi, et al.
Veröffentlicht: (2024)
von: Nika, Andi, et al.
Veröffentlicht: (2024)
Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback
von: Nika, Andi, et al.
Veröffentlicht: (2026)
von: Nika, Andi, et al.
Veröffentlicht: (2026)
Corruption Robust Offline Reinforcement Learning with Human Feedback
von: Mandal, Debmalya, et al.
Veröffentlicht: (2024)
von: Mandal, Debmalya, et al.
Veröffentlicht: (2024)
Benchmarking Generative Models on Computational Thinking Tests in Elementary Visual Programming
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2024)
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2024)
Humanizing Automated Programming Feedback: Fine-Tuning Generative Models with Student-Written Feedback
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2025)
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2025)
BugSpotter: Automated Generation of Code Debugging Exercises
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2024)
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2024)
Interleaving Natural Language Prompting with Code Editing for Solving Programming Tasks with Generative AI Models
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2025)
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2025)
The Right Kind of Help: Evaluating the Effectiveness of Intervention Methods in Elementary-Level Visual Programming
von: Ghosh, Ahana, et al.
Veröffentlicht: (2025)
von: Ghosh, Ahana, et al.
Veröffentlicht: (2025)
Exploring the Impact of Quizzes Interleaved with Write-Code Tasks in Elementary-Level Visual Programming
von: Ghosh, Ahana, et al.
Veröffentlicht: (2024)
von: Ghosh, Ahana, et al.
Veröffentlicht: (2024)
Hints Help Finding and Fixing Bugs Differently in Python and Text-based Program Representations
von: Rawal, Ruchit, et al.
Veröffentlicht: (2024)
von: Rawal, Ruchit, et al.
Veröffentlicht: (2024)
Benchmarking the Robustness of Agentic Systems to Adversarially-Induced Harms
von: Nöther, Jonathan, et al.
Veröffentlicht: (2025)
von: Nöther, Jonathan, et al.
Veröffentlicht: (2025)
Text-Diffusion Red-Teaming of Large Language Models: Unveiling Harmful Behaviors with Proximity Constraints
von: Nöther, Jonathan, et al.
Veröffentlicht: (2025)
von: Nöther, Jonathan, et al.
Veröffentlicht: (2025)
Adversarially Robust Decision Transformer
von: Tang, Xiaohang, et al.
Veröffentlicht: (2024)
von: Tang, Xiaohang, et al.
Veröffentlicht: (2024)
MaMa: A Game-Theoretic Approach for Designing Safe Agentic Systems
von: Nöther, Jonathan, et al.
Veröffentlicht: (2026)
von: Nöther, Jonathan, et al.
Veröffentlicht: (2026)
Towards Generalizable Agents in Text-Based Educational Environments: A Study of Integrating RL with LLMs
von: Radmehr, Bahar, et al.
Veröffentlicht: (2024)
von: Radmehr, Bahar, et al.
Veröffentlicht: (2024)
Learning Personalized Decision Support Policies
von: Bhatt, Umang, et al.
Veröffentlicht: (2023)
von: Bhatt, Umang, et al.
Veröffentlicht: (2023)
Learning Embeddings for Sequential Tasks Using Population of Agents
von: Mahajan, Mridul, et al.
Veröffentlicht: (2023)
von: Mahajan, Mridul, et al.
Veröffentlicht: (2023)
Emergent Bias and Fairness in Multi-Agent Decision Systems
von: Madigan, Maeve, et al.
Veröffentlicht: (2025)
von: Madigan, Maeve, et al.
Veröffentlicht: (2025)
Corruption-Robust Offline Two-Player Zero-Sum Markov Games
von: Nika, Andi, et al.
Veröffentlicht: (2024)
von: Nika, Andi, et al.
Veröffentlicht: (2024)
Federated Variational Preference Alignment with Gumbel-Softmax Prior for Personalized User Preferences
von: Koo, Jabin, et al.
Veröffentlicht: (2026)
von: Koo, Jabin, et al.
Veröffentlicht: (2026)
Evaluating Fairness in Transaction Fraud Models: Fairness Metrics, Bias Audits, and Challenges
von: Kamalaruban, Parameswaran, et al.
Veröffentlicht: (2024)
von: Kamalaruban, Parameswaran, et al.
Veröffentlicht: (2024)
FSPO: Few-Shot Optimization of Synthetic Preferences Personalizes to Real Users
von: Singh, Anikait, et al.
Veröffentlicht: (2025)
von: Singh, Anikait, et al.
Veröffentlicht: (2025)
Formal Models of Active Learning from Contrastive Examples
von: Mansouri, Farnam, et al.
Veröffentlicht: (2025)
von: Mansouri, Farnam, et al.
Veröffentlicht: (2025)
Fairness-Aware Low-Rank Adaptation Under Demographic Privacy Constraints
von: Kamalaruban, Parameswaran, et al.
Veröffentlicht: (2025)
von: Kamalaruban, Parameswaran, et al.
Veröffentlicht: (2025)
PITA: Preference-Guided Inference-Time Alignment for LLM Post-Training
von: Bobbili, Sarat Chandra, et al.
Veröffentlicht: (2025)
von: Bobbili, Sarat Chandra, et al.
Veröffentlicht: (2025)
PersonalizedRouter: Personalized LLM Routing via Graph-based User Preference Modeling
von: Dai, Zhongjie, et al.
Veröffentlicht: (2025)
von: Dai, Zhongjie, et al.
Veröffentlicht: (2025)
Learning Half-Spaces from Perturbed Contrastive Examples
von: Ravari, Aryan Alavi Razavi, et al.
Veröffentlicht: (2026)
von: Ravari, Aryan Alavi Razavi, et al.
Veröffentlicht: (2026)
Hybrid-Gym: Training Coding Agents to Generalize Across Tasks
von: Xie, Yiqing, et al.
Veröffentlicht: (2026)
von: Xie, Yiqing, et al.
Veröffentlicht: (2026)
Optimal Decision Making Under Strategic Behavior
von: Tsirtsis, Stratis, et al.
Veröffentlicht: (2019)
von: Tsirtsis, Stratis, et al.
Veröffentlicht: (2019)
DMAP: A Distribution Map for Text
von: Kempton, Tom, et al.
Veröffentlicht: (2026)
von: Kempton, Tom, et al.
Veröffentlicht: (2026)
Amulet: ReAlignment During Test Time for Personalized Preference Adaptation of LLMs
von: Zhang, Zhaowei, et al.
Veröffentlicht: (2025)
von: Zhang, Zhaowei, et al.
Veröffentlicht: (2025)
GEM: Generative Entropy-Guided Preference Modeling for Few-shot Alignment of LLMs
von: Zhao, Yiyang, et al.
Veröffentlicht: (2025)
von: Zhao, Yiyang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Hints-In-Browser: Benchmarking Language Models for Programming Feedback Generation
von: Kotalwar, Nachiket, et al.
Veröffentlicht: (2024) -
Prompt Programming: A Platform for Dialogue-based Computational Problem Solving with Generative AI Models
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2025) -
Curriculum Design for Trajectory-Constrained Agent: Compressing Chain-of-Thought Tokens in LLMs
von: Tzannetos, Georgios, et al.
Veröffentlicht: (2025) -
Informativeness of Reward Functions in Reinforcement Learning
von: Devidze, Rati, et al.
Veröffentlicht: (2024) -
Proximal Curriculum with Task Correlations for Deep Reinforcement Learning
von: Tzannetos, Georgios, et al.
Veröffentlicht: (2024)