Enhancing RLHF with Human Gaze Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Galliamov, Karim, Titov, Ivan, Pershin, Ilya |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Knowledge Offloading: Decomposing LLMs into Sparse Backbones and Memory Modules
by: Galliamov, Karim, et al.
Published: (2026)
by: Galliamov, Karim, et al.
Published: (2026)
Refining Joint Text and Source Code Embeddings for Retrieval Task with Parameter-Efficient Fine-Tuning
by: Galliamov, Karim, et al.
Published: (2024)
by: Galliamov, Karim, et al.
Published: (2024)
[Re] FairDICE: A Fair Tradeoff in Multi-objective Offline RL
by: Adema, Peter, et al.
Published: (2026)
by: Adema, Peter, et al.
Published: (2026)
Concepts' Information Bottleneck Models
by: Galliamov, Karim, et al.
Published: (2026)
by: Galliamov, Karim, et al.
Published: (2026)
Influencing Humans to Conform to Preference Models for RLHF
by: Hatgis-Kessell, Stephane, et al.
Published: (2025)
by: Hatgis-Kessell, Stephane, et al.
Published: (2025)
Calibration-Gated LLM Pseudo-Observations for Online Contextual Bandits
by: Pershin, Maksim, et al.
Published: (2026)
by: Pershin, Maksim, et al.
Published: (2026)
Awareness of uncertainty in classification using a multivariate model and multi-views
by: Kornaev, Alexey, et al.
Published: (2024)
by: Kornaev, Alexey, et al.
Published: (2024)
Optimal Design for Reward Modeling in RLHF
by: Scheid, Antoine, et al.
Published: (2024)
by: Scheid, Antoine, et al.
Published: (2024)
RLHF Workflow: From Reward Modeling to Online RLHF
by: Dong, Hanze, et al.
Published: (2024)
by: Dong, Hanze, et al.
Published: (2024)
What's New in My Data? Novelty Exploration via Contrastive Generation
by: Isonuma, Masaru, et al.
Published: (2024)
by: Isonuma, Masaru, et al.
Published: (2024)
A Descriptive and Normative Theory of Human Beliefs in RLHF
by: Dandekar, Sylee, et al.
Published: (2025)
by: Dandekar, Sylee, et al.
Published: (2025)
Shared Doubt: Zero-shot Cross-Lingual Confidence Estimation for Language Models
by: Kyriakou, Athina, et al.
Published: (2026)
by: Kyriakou, Athina, et al.
Published: (2026)
Learning a Pessimistic Reward Model in RLHF
by: Xu, Yinglun, et al.
Published: (2025)
by: Xu, Yinglun, et al.
Published: (2025)
SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences
by: Mukherjee, Arpan, et al.
Published: (2025)
by: Mukherjee, Arpan, et al.
Published: (2025)
Detecting and Pruning Prominent but Detrimental Neurons in Large Language Models
by: Ali, Ameen, et al.
Published: (2025)
by: Ali, Ameen, et al.
Published: (2025)
Cache & Distil: Optimising API Calls to Large Language Models
by: Ramírez, Guillem, et al.
Published: (2023)
by: Ramírez, Guillem, et al.
Published: (2023)
Zero-Shot Gaze-based Volumetric Medical Image Segmentation
by: Shmykova, Tatyana, et al.
Published: (2025)
by: Shmykova, Tatyana, et al.
Published: (2025)
Autoencoding Conditional Neural Processes for Representation Learning
by: Prokhorov, Victor, et al.
Published: (2023)
by: Prokhorov, Victor, et al.
Published: (2023)
Mitigating Copy Bias in In-Context Learning through Neuron Pruning
by: Ali, Ameen, et al.
Published: (2024)
by: Ali, Ameen, et al.
Published: (2024)
WPO: Enhancing RLHF with Weighted Preference Optimization
by: Zhou, Wenxuan, et al.
Published: (2024)
by: Zhou, Wenxuan, et al.
Published: (2024)
Mitigating the Alignment Tax of RLHF
by: Lin, Yong, et al.
Published: (2023)
by: Lin, Yong, et al.
Published: (2023)
Provably Efficient Online RLHF with One-Pass Reward Modeling
by: Li, Long-Fei, et al.
Published: (2025)
by: Li, Long-Fei, et al.
Published: (2025)
Factored Causal Representation Learning for Robust Reward Modeling in RLHF
by: Yang, Yupei, et al.
Published: (2026)
by: Yang, Yupei, et al.
Published: (2026)
Finding Culture-Sensitive Neurons in Vision-Language Models
by: Zhao, Xiutian, et al.
Published: (2025)
by: Zhao, Xiutian, et al.
Published: (2025)
Gaze-Assisted Medical Image Segmentation
by: Khaertdinova, Leila, et al.
Published: (2024)
by: Khaertdinova, Leila, et al.
Published: (2024)
MaxMin-RLHF: Alignment with Diverse Human Preferences
by: Chakraborty, Souradip, et al.
Published: (2024)
by: Chakraborty, Souradip, et al.
Published: (2024)
Joint Localization and Activation Editing for Low-Resource Fine-Tuning
by: Lai, Wen, et al.
Published: (2025)
by: Lai, Wen, et al.
Published: (2025)
Value Imprint: A Technique for Auditing the Human Values Embedded in RLHF Datasets
by: Obi, Ike, et al.
Published: (2024)
by: Obi, Ike, et al.
Published: (2024)
Towards true discovery of the differential equations
by: Hvatov, Alexander, et al.
Published: (2023)
by: Hvatov, Alexander, et al.
Published: (2023)
Reward Model Overoptimisation in Iterated RLHF
by: Wolf, Lorenz, et al.
Published: (2025)
by: Wolf, Lorenz, et al.
Published: (2025)
How to Evaluate Reward Models for RLHF
by: Frick, Evan, et al.
Published: (2024)
by: Frick, Evan, et al.
Published: (2024)
Policy Filtration for RLHF to Mitigate Noise in Reward Models
by: Zhang, Chuheng, et al.
Published: (2024)
by: Zhang, Chuheng, et al.
Published: (2024)
ACE-RLHF: Automated Code Evaluation and Socratic Feedback Generation Tool using Large Language Models and Reinforcement Learning with Human Feedback
by: Rahman, Tasnia, et al.
Published: (2025)
by: Rahman, Tasnia, et al.
Published: (2025)
On a Connection Between Imitation Learning and RLHF
by: Xiao, Teng, et al.
Published: (2025)
by: Xiao, Teng, et al.
Published: (2025)
Explaining and Preventing Alignment Collapse in Iterative RLHF
by: Gauthier, Etienne, et al.
Published: (2026)
by: Gauthier, Etienne, et al.
Published: (2026)
Offline Constrained RLHF with Multiple Preference Oracles
by: Latham, Brenden, et al.
Published: (2026)
by: Latham, Brenden, et al.
Published: (2026)
Understanding and Alleviating Memory Consumption in RLHF for LLMs
by: Zhou, Jin, et al.
Published: (2024)
by: Zhou, Jin, et al.
Published: (2024)
On The Global Convergence Of Online RLHF With Neural Parametrization
by: Gaur, Mudit, et al.
Published: (2024)
by: Gaur, Mudit, et al.
Published: (2024)
Unifying Stable Optimization and Reference Regularization in RLHF
by: He, Li, et al.
Published: (2026)
by: He, Li, et al.
Published: (2026)
Towards a Theoretical Understanding to the Generalization of RLHF
by: Li, Zhaochun, et al.
Published: (2026)
by: Li, Zhaochun, et al.
Published: (2026)
Similar Items
-
Knowledge Offloading: Decomposing LLMs into Sparse Backbones and Memory Modules
by: Galliamov, Karim, et al.
Published: (2026) -
Refining Joint Text and Source Code Embeddings for Retrieval Task with Parameter-Efficient Fine-Tuning
by: Galliamov, Karim, et al.
Published: (2024) -
[Re] FairDICE: A Fair Tradeoff in Multi-objective Offline RL
by: Adema, Peter, et al.
Published: (2026) -
Concepts' Information Bottleneck Models
by: Galliamov, Karim, et al.
Published: (2026) -
Influencing Humans to Conform to Preference Models for RLHF
by: Hatgis-Kessell, Stephane, et al.
Published: (2025)