Dishonesty in Helpful and Harmless Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Youcheng, Tang, Jingkun, Feng, Duanyu, Zhang, Zheng, Lei, Wenqiang, Lv, Jiancheng, Cohn, Anthony G. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Cross-model Transferability among Large Language Models on the Platonic Representations of Concepts
von: Huang, Youcheng, et al.
Veröffentlicht: (2025)
von: Huang, Youcheng, et al.
Veröffentlicht: (2025)
DREditor: An Time-efficient Approach for Building a Domain-specific Dense Retrieval Model
von: Huang, Chen, et al.
Veröffentlicht: (2024)
von: Huang, Chen, et al.
Veröffentlicht: (2024)
See the Unseen: Better Context-Consistent Knowledge-Editing by Noises
von: Huang, Youcheng, et al.
Veröffentlicht: (2024)
von: Huang, Youcheng, et al.
Veröffentlicht: (2024)
Legend: Leveraging Representation Engineering to Annotate Safety Margin for Preference Datasets
von: Feng, Duanyu, et al.
Veröffentlicht: (2024)
von: Feng, Duanyu, et al.
Veröffentlicht: (2024)
AMaPO: Adaptive Margin-attached Preference Optimization for Language Model Alignment
von: Deng, Ruibo, et al.
Veröffentlicht: (2025)
von: Deng, Ruibo, et al.
Veröffentlicht: (2025)
Effective and Efficient Adversarial Detection for Vision-Language Models via A Single Vector
von: Huang, Youcheng, et al.
Veröffentlicht: (2024)
von: Huang, Youcheng, et al.
Veröffentlicht: (2024)
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information
von: Huang, Youcheng, et al.
Veröffentlicht: (2025)
von: Huang, Youcheng, et al.
Veröffentlicht: (2025)
Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective
von: Feng, Duanyu, et al.
Veröffentlicht: (2024)
von: Feng, Duanyu, et al.
Veröffentlicht: (2024)
Selective Annotation via Data Allocation: These Data Should Be Triaged to Experts for Annotation Rather Than the Model
von: Huang, Chen, et al.
Veröffentlicht: (2024)
von: Huang, Chen, et al.
Veröffentlicht: (2024)
Can Large Language Models Understand Internet Buzzwords Through User-Generated Content
von: Huang, Chen, et al.
Veröffentlicht: (2025)
von: Huang, Chen, et al.
Veröffentlicht: (2025)
Beyond the Safety Bundle: Auditing the Helpful and Harmless Dataset
von: Chehbouni, Khaoula, et al.
Veröffentlicht: (2024)
von: Chehbouni, Khaoula, et al.
Veröffentlicht: (2024)
ARAIDA: Analogical Reasoning-Augmented Interactive Data Annotation
von: Huang, Chen, et al.
Veröffentlicht: (2024)
von: Huang, Chen, et al.
Veröffentlicht: (2024)
Too Helpful, Too Harmless, Too Honest or Just Right?
von: Kashyap, Gautam Siddharth, et al.
Veröffentlicht: (2025)
von: Kashyap, Gautam Siddharth, et al.
Veröffentlicht: (2025)
BAR: A Backward Reasoning based Agent for Complex Minecraft Tasks
von: Du, Weihong, et al.
Veröffentlicht: (2025)
von: Du, Weihong, et al.
Veröffentlicht: (2025)
H3Fusion: Helpful, Harmless, Honest Fusion of Aligned LLMs
von: Tekin, Selim Furkan, et al.
Veröffentlicht: (2024)
von: Tekin, Selim Furkan, et al.
Veröffentlicht: (2024)
AdvisorQA: Towards Helpful and Harmless Advice-seeking Question Answering with Collective Intelligence
von: Kim, Minbeom, et al.
Veröffentlicht: (2024)
von: Kim, Minbeom, et al.
Veröffentlicht: (2024)
InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance
von: Wang, Pengyu, et al.
Veröffentlicht: (2024)
von: Wang, Pengyu, et al.
Veröffentlicht: (2024)
Concept -- An Evaluation Protocol on Conversational Recommender Systems with System-centric and User-centric Factors
von: Huang, Chen, et al.
Veröffentlicht: (2024)
von: Huang, Chen, et al.
Veröffentlicht: (2024)
We Think, Therefore We Align LLMs to Helpful, Harmless and Honest Before They Go Wrong
von: Kashyap, Gautam Siddharth, et al.
Veröffentlicht: (2025)
von: Kashyap, Gautam Siddharth, et al.
Veröffentlicht: (2025)
How to Enable Effective Cooperation Between Humans and NLP Models: A Survey of Principles, Formalizations, and Beyond
von: Huang, Chen, et al.
Veröffentlicht: (2025)
von: Huang, Chen, et al.
Veröffentlicht: (2025)
Mix Data or Merge Models? Balancing the Helpfulness, Honesty, and Harmlessness of Large Language Model via Model Merging
von: Yang, Jinluan, et al.
Veröffentlicht: (2025)
von: Yang, Jinluan, et al.
Veröffentlicht: (2025)
Towards Understanding the Influence of Reward Margin on Preference Model Performance
von: Qin, Bowen, et al.
Veröffentlicht: (2024)
von: Qin, Bowen, et al.
Veröffentlicht: (2024)
Can Large Language Models Reason about the Region Connection Calculus?
von: Cohn, Anthony G, et al.
Veröffentlicht: (2024)
von: Cohn, Anthony G, et al.
Veröffentlicht: (2024)
Evaluating the Ability of Large Language Models to Reason about Cardinal Directions
von: Cohn, Anthony G, et al.
Veröffentlicht: (2024)
von: Cohn, Anthony G, et al.
Veröffentlicht: (2024)
Evaluating the Ability of Large Language Models to Reason about Cardinal Directions, Revisited
von: Cohn, Anthony G, et al.
Veröffentlicht: (2025)
von: Cohn, Anthony G, et al.
Veröffentlicht: (2025)
Balancing Enhancement, Harmlessness, and General Capabilities: Enhancing Conversational LLMs with Direct RLHF
von: Zheng, Chen, et al.
Veröffentlicht: (2024)
von: Zheng, Chen, et al.
Veröffentlicht: (2024)
Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores
von: Blackwell, Robert E., et al.
Veröffentlicht: (2024)
von: Blackwell, Robert E., et al.
Veröffentlicht: (2024)
Exploring Spatial Representations in the Historical Lake District Texts with LLM-based Relation Extraction
von: Haris, Erum, et al.
Veröffentlicht: (2024)
von: Haris, Erum, et al.
Veröffentlicht: (2024)
Reinforcement Learning Amplifies Emergent Misalignment from Harmless Rewards
von: Jørgenvåg, Magnus, et al.
Veröffentlicht: (2026)
von: Jørgenvåg, Magnus, et al.
Veröffentlicht: (2026)
Adaptive Helpfulness-Harmlessness Alignment with Preference Vectors
von: Liang, Ren-Wei, et al.
Veröffentlicht: (2025)
von: Liang, Ren-Wei, et al.
Veröffentlicht: (2025)
FFT: Towards Harmlessness Evaluation and Analysis for LLMs with Factuality, Fairness, Toxicity
von: Cui, Shiyao, et al.
Veröffentlicht: (2023)
von: Cui, Shiyao, et al.
Veröffentlicht: (2023)
Towards Harmless Multimodal Assistants with Blind Preference Optimization
von: Li, Yongqi, et al.
Veröffentlicht: (2025)
von: Li, Yongqi, et al.
Veröffentlicht: (2025)
Alignment Helps Make the Most of Multimodal Data
von: Arnold, Christian, et al.
Veröffentlicht: (2024)
von: Arnold, Christian, et al.
Veröffentlicht: (2024)
Empirical Study on Updating Key-Value Memories in Transformer Feed-forward Layers
von: Qiu, Zihan, et al.
Veröffentlicht: (2024)
von: Qiu, Zihan, et al.
Veröffentlicht: (2024)
LLMs Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions
von: Hu, Xuhao, et al.
Veröffentlicht: (2025)
von: Hu, Xuhao, et al.
Veröffentlicht: (2025)
When Harmless Words Harm: A New Threat to LLM Safety via Conceptual Triggers
von: Zhang, Zhaoxin, et al.
Veröffentlicht: (2025)
von: Zhang, Zhaoxin, et al.
Veröffentlicht: (2025)
Compromising Honesty and Harmlessness in Language Models via Deception Attacks
von: Vaugrante, Laurène, et al.
Veröffentlicht: (2025)
von: Vaugrante, Laurène, et al.
Veröffentlicht: (2025)
Reframing Spatial Reasoning Evaluation in Language Models: A Real-World Simulation Benchmark for Qualitative Reasoning
von: Li, Fangjun, et al.
Veröffentlicht: (2024)
von: Li, Fangjun, et al.
Veröffentlicht: (2024)
"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth
von: Li, Yaqiong, et al.
Veröffentlicht: (2025)
von: Li, Yaqiong, et al.
Veröffentlicht: (2025)
The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment
von: Brach, William, et al.
Veröffentlicht: (2026)
von: Brach, William, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Cross-model Transferability among Large Language Models on the Platonic Representations of Concepts
von: Huang, Youcheng, et al.
Veröffentlicht: (2025) -
DREditor: An Time-efficient Approach for Building a Domain-specific Dense Retrieval Model
von: Huang, Chen, et al.
Veröffentlicht: (2024) -
See the Unseen: Better Context-Consistent Knowledge-Editing by Noises
von: Huang, Youcheng, et al.
Veröffentlicht: (2024) -
Legend: Leveraging Representation Engineering to Annotate Safety Margin for Preference Datasets
von: Feng, Duanyu, et al.
Veröffentlicht: (2024) -
AMaPO: Adaptive Margin-attached Preference Optimization for Language Model Alignment
von: Deng, Ruibo, et al.
Veröffentlicht: (2025)