Elephant in the Room: Unveiling the Impact of Reward Model Quality in Alignment
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Yan, Yi, Xiaoyuan, Chen, Xiaokang, Yao, Jing, Yi, Jingwei, Zan, Daoguang, Liu, Zheng, Xie, Xing, Ho, Tsung-Yi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Devil is in the Neurons: Interpreting and Mitigating Social Biases in Pre-trained Language Models
di: Liu, Yan, et al.
Pubblicazione: (2024)
di: Liu, Yan, et al.
Pubblicazione: (2024)
CLAVE: An Adaptive Framework for Evaluating Values of LLM Generated Responses
di: Yao, Jing, et al.
Pubblicazione: (2024)
di: Yao, Jing, et al.
Pubblicazione: (2024)
CAReDiO: Cultural Alignment via Representativeness and Distinctiveness Guided Data Optimization
di: Yao, Jing, et al.
Pubblicazione: (2025)
di: Yao, Jing, et al.
Pubblicazione: (2025)
Negating Negatives: Alignment with Human Negative Samples via Distributional Dispreference Optimization
di: Duan, Shitong, et al.
Pubblicazione: (2024)
di: Duan, Shitong, et al.
Pubblicazione: (2024)
Beyond Human Norms: Unveiling Unique Values of Large Language Models through Interdisciplinary Approaches
di: Biedma, Pablo, et al.
Pubblicazione: (2024)
di: Biedma, Pablo, et al.
Pubblicazione: (2024)
On the Essence and Prospect: An Investigation of Alignment Approaches for Big Models
di: Wang, Xinpeng, et al.
Pubblicazione: (2024)
di: Wang, Xinpeng, et al.
Pubblicazione: (2024)
SR-GRPO: Stable Rank as an Intrinsic Geometric Reward for Large Language Model Alignment
di: Tang, Yixuan, et al.
Pubblicazione: (2025)
di: Tang, Yixuan, et al.
Pubblicazione: (2025)
PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
di: Jiang, Han, et al.
Pubblicazione: (2025)
di: Jiang, Han, et al.
Pubblicazione: (2025)
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
di: Lee, Jaehyeok, et al.
Pubblicazione: (2026)
di: Lee, Jaehyeok, et al.
Pubblicazione: (2026)
Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language Models
di: Guo, Hanze, et al.
Pubblicazione: (2025)
di: Guo, Hanze, et al.
Pubblicazione: (2025)
ToolNet: Connecting Large Language Models with Massive Tools via Tool Graph
di: Liu, Xukun, et al.
Pubblicazione: (2024)
di: Liu, Xukun, et al.
Pubblicazione: (2024)
MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models?
di: Yong, Xixian, et al.
Pubblicazione: (2025)
di: Yong, Xixian, et al.
Pubblicazione: (2025)
Improving Natural Language Capability of Code Large Language Model
di: Li, Wei, et al.
Pubblicazione: (2024)
di: Li, Wei, et al.
Pubblicazione: (2024)
Hey, That's My Data! Token-Only Dataset Inference in Large Language Models
di: Xiong, Chen, et al.
Pubblicazione: (2025)
di: Xiong, Chen, et al.
Pubblicazione: (2025)
Negation: A Pink Elephant in the Large Language Models' Room?
di: Vrabcová, Tereza, et al.
Pubblicazione: (2025)
di: Vrabcová, Tereza, et al.
Pubblicazione: (2025)
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
di: Choi, Sooyung, et al.
Pubblicazione: (2025)
di: Choi, Sooyung, et al.
Pubblicazione: (2025)
Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes
di: Hu, Xiaomeng, et al.
Pubblicazione: (2024)
di: Hu, Xiaomeng, et al.
Pubblicazione: (2024)
CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention
di: Hu, Xiaomeng, et al.
Pubblicazione: (2025)
di: Hu, Xiaomeng, et al.
Pubblicazione: (2025)
Leveraging Implicit Sentiments: Enhancing Reliability and Validity in Psychological Trait Evaluation of LLMs
di: Ma, Huanhuan, et al.
Pubblicazione: (2025)
di: Ma, Huanhuan, et al.
Pubblicazione: (2025)
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
di: Hu, Xiaomeng, et al.
Pubblicazione: (2025)
di: Hu, Xiaomeng, et al.
Pubblicazione: (2025)
Modelling Intertextuality with N-gram Embeddings
di: Xing, Yi
Pubblicazione: (2025)
di: Xing, Yi
Pubblicazione: (2025)
MoVa: Towards Generalizable Classification of Human Morals and Values
di: Chen, Ziyu, et al.
Pubblicazione: (2025)
di: Chen, Ziyu, et al.
Pubblicazione: (2025)
IROTE: Human-like Traits Elicitation of Large Language Model via In-Context Self-Reflective Optimization
di: Bai, Yuzhuo, et al.
Pubblicazione: (2025)
di: Bai, Yuzhuo, et al.
Pubblicazione: (2025)
The Incomplete Bridge: How AI Research (Mis)Engages with Psychology
di: Jiang, Han, et al.
Pubblicazione: (2025)
di: Jiang, Han, et al.
Pubblicazione: (2025)
Raising the Bar: Investigating the Values of Large Language Models via Generative Evolving Testing
di: Jiang, Han, et al.
Pubblicazione: (2024)
di: Jiang, Han, et al.
Pubblicazione: (2024)
Human Values Matter: Investigating How Misalignment Shapes Collective Behaviors in LLM Agent Communities
di: Zhang, Xiangxu, et al.
Pubblicazione: (2026)
di: Zhang, Xiangxu, et al.
Pubblicazione: (2026)
Denevil: Towards Deciphering and Navigating the Ethical Values of Large Language Models via Instruction Learning
di: Duan, Shitong, et al.
Pubblicazione: (2023)
di: Duan, Shitong, et al.
Pubblicazione: (2023)
Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models
di: Yi, Jingwei, et al.
Pubblicazione: (2023)
di: Yi, Jingwei, et al.
Pubblicazione: (2023)
Does LLM Alignment Really Need Diversity? An Empirical Study of Adapting RLVR Methods for Moral Reasoning
di: Zhang, Zhaowei, et al.
Pubblicazione: (2026)
di: Zhang, Zhaowei, et al.
Pubblicazione: (2026)
PERM: Psychology-grounded Empathetic Reward Modeling for Large Language Models
di: Wang, Chengbing, et al.
Pubblicazione: (2026)
di: Wang, Chengbing, et al.
Pubblicazione: (2026)
CDEval: A Benchmark for Measuring the Cultural Dimensions of Large Language Models
di: Wang, Yuhang, et al.
Pubblicazione: (2023)
di: Wang, Yuhang, et al.
Pubblicazione: (2023)
Unveiling Environmental Impacts of Large Language Model Serving: A Functional Unit View
di: Wu, Yanran, et al.
Pubblicazione: (2025)
di: Wu, Yanran, et al.
Pubblicazione: (2025)
Improving Long Text Understanding with Knowledge Distilled from Summarization Model
di: Liu, Yan, et al.
Pubblicazione: (2024)
di: Liu, Yan, et al.
Pubblicazione: (2024)
Dynamic Rewarding with Prompt Optimization Enables Tuning-free Self-Alignment of Language Models
di: Singla, Somanshu, et al.
Pubblicazione: (2024)
di: Singla, Somanshu, et al.
Pubblicazione: (2024)
The Elephant in the Room: Analyzing the Presence of Big Tech in Natural Language Processing Research
di: Abdalla, Mohamed, et al.
Pubblicazione: (2023)
di: Abdalla, Mohamed, et al.
Pubblicazione: (2023)
The Elephant in the Coreference Room: Resolving Coreference in Full-Length French Fiction Works
di: Bourgois, Antoine, et al.
Pubblicazione: (2025)
di: Bourgois, Antoine, et al.
Pubblicazione: (2025)
AdAEM: An Adaptively and Automated Extensible Measurement of LLMs' Value Difference
di: Yao, Jing, et al.
Pubblicazione: (2025)
di: Yao, Jing, et al.
Pubblicazione: (2025)
LaF-GRPO: In-Situ Navigation Instruction Generation for the Visually Impaired via GRPO with LLM-as-Follower Reward
di: Zhao, Yi, et al.
Pubblicazione: (2025)
di: Zhao, Yi, et al.
Pubblicazione: (2025)
Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets
di: Hsiung, Lei, et al.
Pubblicazione: (2025)
di: Hsiung, Lei, et al.
Pubblicazione: (2025)
Contextualized Privacy Defense for LLM Agents
di: Wen, Yule, et al.
Pubblicazione: (2026)
di: Wen, Yule, et al.
Pubblicazione: (2026)
Documenti analoghi
-
The Devil is in the Neurons: Interpreting and Mitigating Social Biases in Pre-trained Language Models
di: Liu, Yan, et al.
Pubblicazione: (2024) -
CLAVE: An Adaptive Framework for Evaluating Values of LLM Generated Responses
di: Yao, Jing, et al.
Pubblicazione: (2024) -
CAReDiO: Cultural Alignment via Representativeness and Distinctiveness Guided Data Optimization
di: Yao, Jing, et al.
Pubblicazione: (2025) -
Negating Negatives: Alignment with Human Negative Samples via Distributional Dispreference Optimization
di: Duan, Shitong, et al.
Pubblicazione: (2024) -
Beyond Human Norms: Unveiling Unique Values of Large Language Models through Interdisciplinary Approaches
di: Biedma, Pablo, et al.
Pubblicazione: (2024)