A Survey on Progress in LLM Alignment from the Perspective of Reward Design
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ji, Miaomiao, Wu, Yanqiu, Wu, Zhibin, Wang, Shoujin, Yang, Jian, Dras, Mark, Naseem, Usman |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Should LLM Safety Be More Than Refusing Harmful Instructions?
von: Maskey, Utsav, et al.
Veröffentlicht: (2025)
von: Maskey, Utsav, et al.
Veröffentlicht: (2025)
Over-Refusal and Representation Subspaces: A Mechanistic Analysis of Task-Conditioned Refusal in Aligned LLMs
von: Maskey, Utsav, et al.
Veröffentlicht: (2026)
von: Maskey, Utsav, et al.
Veröffentlicht: (2026)
SHIELD: Classifier-Guided Prompting for Robust and Safer LVLMs
von: Ren, Juan, et al.
Veröffentlicht: (2025)
von: Ren, Juan, et al.
Veröffentlicht: (2025)
Seeing the Threat: Vulnerabilities in Vision-Language Models to Adversarial Attack
von: Ren, Juan, et al.
Veröffentlicht: (2025)
von: Ren, Juan, et al.
Veröffentlicht: (2025)
Steering Over-refusals Towards Safety in Retrieval Augmented Generation
von: Maskey, Utsav, et al.
Veröffentlicht: (2025)
von: Maskey, Utsav, et al.
Veröffentlicht: (2025)
VITAL: A New Dataset for Benchmarking Pluralistic Alignment in Healthcare
von: Shetty, Anudeex, et al.
Veröffentlicht: (2025)
von: Shetty, Anudeex, et al.
Veröffentlicht: (2025)
Beyond the Black Box: Demystifying Multi-Turn LLM Reasoning with VISTA
von: Zhang, Yiran, et al.
Veröffentlicht: (2025)
von: Zhang, Yiran, et al.
Veröffentlicht: (2025)
Steering Towards Fairness: Mitigating Political Bias in LLMs
von: Nadeem, Afrozah, et al.
Veröffentlicht: (2025)
von: Nadeem, Afrozah, et al.
Veröffentlicht: (2025)
Fairness Evaluation and Inference Level Mitigation in LLMs
von: Nadeem, Afrozah, et al.
Veröffentlicht: (2025)
von: Nadeem, Afrozah, et al.
Veröffentlicht: (2025)
Framing Political Bias in Multilingual LLMs Across Pakistani Languages
von: Nadeem, Afrozah, et al.
Veröffentlicht: (2025)
von: Nadeem, Afrozah, et al.
Veröffentlicht: (2025)
We Think, Therefore We Align LLMs to Helpful, Harmless and Honest Before They Go Wrong
von: Kashyap, Gautam Siddharth, et al.
Veröffentlicht: (2025)
von: Kashyap, Gautam Siddharth, et al.
Veröffentlicht: (2025)
Too Helpful, Too Harmless, Too Honest or Just Right?
von: Kashyap, Gautam Siddharth, et al.
Veröffentlicht: (2025)
von: Kashyap, Gautam Siddharth, et al.
Veröffentlicht: (2025)
AlignCultura: Towards Culturally Aligned Large Language Models?
von: Kashyap, Gautam Siddharth, et al.
Veröffentlicht: (2026)
von: Kashyap, Gautam Siddharth, et al.
Veröffentlicht: (2026)
When the Model Said 'No Comment', We Knew Helpfulness Was Dead, Honesty Was Alive, and Safety Was Terrified
von: Kashyap, Gautam Siddharth, et al.
Veröffentlicht: (2026)
von: Kashyap, Gautam Siddharth, et al.
Veröffentlicht: (2026)
Mechanistic Interpretability for Large Language Model Alignment: Progress, Challenges, and Future Directions
von: Naseem, Usman
Veröffentlicht: (2026)
von: Naseem, Usman
Veröffentlicht: (2026)
CogMem: A Cognitive Memory Architecture for Sustained Multi-Turn Reasoning in Large Language Models
von: Zhang, Yiran, et al.
Veröffentlicht: (2025)
von: Zhang, Yiran, et al.
Veröffentlicht: (2025)
SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation Steering
von: Maskey, Utsav, et al.
Veröffentlicht: (2025)
von: Maskey, Utsav, et al.
Veröffentlicht: (2025)
VaxGuard: A Multi-Generator, Multi-Type, and Multi-Role Dataset for Detecting LLM-Generated Vaccine Misinformation
von: Ahmad, Syed Talal, et al.
Veröffentlicht: (2025)
von: Ahmad, Syed Talal, et al.
Veröffentlicht: (2025)
Medical Question Summarization with Entity-driven Contrastive Learning
von: Lu, Wenpeng, et al.
Veröffentlicht: (2023)
von: Lu, Wenpeng, et al.
Veröffentlicht: (2023)
MSynFD: Multi-hop Syntax aware Fake News Detection
von: Xiao, Liang, et al.
Veröffentlicht: (2024)
von: Xiao, Liang, et al.
Veröffentlicht: (2024)
Agentic Moderation: Multi-Agent Design for Safer Vision-Language Models
von: Ren, Juan, et al.
Veröffentlicht: (2025)
von: Ren, Juan, et al.
Veröffentlicht: (2025)
Cultural Palette: Pluralising Culture Alignment via Multi-agent Palette
von: Yuan, Jiahao, et al.
Veröffentlicht: (2024)
von: Yuan, Jiahao, et al.
Veröffentlicht: (2024)
Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards
von: Wu, Xiaobao
Veröffentlicht: (2025)
von: Wu, Xiaobao
Veröffentlicht: (2025)
VISPA: Pluralistic Alignment via Automatic Value Selection and Activation
von: Zheng, Shenyan, et al.
Veröffentlicht: (2026)
von: Zheng, Shenyan, et al.
Veröffentlicht: (2026)
Can Pruning Improve Reasoning? Revisiting Long-CoT Compression with Capability in Mind for Better Reasoning
von: Zhao, Shangziqi, et al.
Veröffentlicht: (2025)
von: Zhao, Shangziqi, et al.
Veröffentlicht: (2025)
Unifying Tree Search Algorithm and Reward Design for LLM Reasoning: A Survey
von: Wei, Jiaqi, et al.
Veröffentlicht: (2025)
von: Wei, Jiaqi, et al.
Veröffentlicht: (2025)
Pluralistic Alignment for Healthcare: A Role-Driven Framework
von: Zhong, Jiayou, et al.
Veröffentlicht: (2025)
von: Zhong, Jiayou, et al.
Veröffentlicht: (2025)
Myanmar XNLI: Building a Dataset and Exploring Low-resource Approaches to Natural Language Inference with Myanmar
von: Htet, Aung Kyaw, et al.
Veröffentlicht: (2025)
von: Htet, Aung Kyaw, et al.
Veröffentlicht: (2025)
DUAL-Bench: Measuring Over-Refusal and Robustness in Vision-Language Models
von: Ren, Kaixuan, et al.
Veröffentlicht: (2025)
von: Ren, Kaixuan, et al.
Veröffentlicht: (2025)
Benchmarking Large Language Models for Cryptanalysis and Side-Channel Vulnerabilities
von: Maskey, Utsav, et al.
Veröffentlicht: (2025)
von: Maskey, Utsav, et al.
Veröffentlicht: (2025)
Enhancing ESG Impact Type Identification through Early Fusion and Multilingual Models
von: Veeramani, Hariram, et al.
Veröffentlicht: (2024)
von: Veeramani, Hariram, et al.
Veröffentlicht: (2024)
PersoDPO: Scalable Preference Optimization for Instruction-Adherent, Persona-Grounded Dialogue via Multi-LLM Evaluation
von: Afzoon, Saleh, et al.
Veröffentlicht: (2026)
von: Afzoon, Saleh, et al.
Veröffentlicht: (2026)
AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025)
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025)
PersoPilot: An Adaptive AI-Copilot for Transparent Contextualized Persona Classification and Personalized Response Generation
von: Afzoon, Saleh, et al.
Veröffentlicht: (2026)
von: Afzoon, Saleh, et al.
Veröffentlicht: (2026)
PACR: Progressively Ascending Confidence Reward for LLM Reasoning
von: Yoon, Eunseop, et al.
Veröffentlicht: (2025)
von: Yoon, Eunseop, et al.
Veröffentlicht: (2025)
Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy
von: Afzoon, Saleh, et al.
Veröffentlicht: (2025)
von: Afzoon, Saleh, et al.
Veröffentlicht: (2025)
Teaching LLM to be Persuasive: Reward-Enhanced Policy Optimization for Alignment from Heterogeneous Rewards
von: Zeng, Xia, et al.
Veröffentlicht: (2025)
von: Zeng, Xia, et al.
Veröffentlicht: (2025)
RMB: Comprehensively Benchmarking Reward Models in LLM Alignment
von: Zhou, Enyu, et al.
Veröffentlicht: (2024)
von: Zhou, Enyu, et al.
Veröffentlicht: (2024)
A Survey on Personalized Alignment -- The Missing Piece for Large Language Models in Real-World Applications
von: Guan, Jian, et al.
Veröffentlicht: (2025)
von: Guan, Jian, et al.
Veröffentlicht: (2025)
XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content
von: Abishethvarman, Vadivel, et al.
Veröffentlicht: (2025)
von: Abishethvarman, Vadivel, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Should LLM Safety Be More Than Refusing Harmful Instructions?
von: Maskey, Utsav, et al.
Veröffentlicht: (2025) -
Over-Refusal and Representation Subspaces: A Mechanistic Analysis of Task-Conditioned Refusal in Aligned LLMs
von: Maskey, Utsav, et al.
Veröffentlicht: (2026) -
SHIELD: Classifier-Guided Prompting for Robust and Safer LVLMs
von: Ren, Juan, et al.
Veröffentlicht: (2025) -
Seeing the Threat: Vulnerabilities in Vision-Language Models to Adversarial Attack
von: Ren, Juan, et al.
Veröffentlicht: (2025) -
Steering Over-refusals Towards Safety in Retrieval Augmented Generation
von: Maskey, Utsav, et al.
Veröffentlicht: (2025)