Legend: Leveraging Representation Engineering to Annotate Safety Margin for Preference Datasets
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Feng, Duanyu, Qin, Bowen, Huang, Chen, Huang, Youcheng, Zhang, Zheng, Lei, Wenqiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Cross-model Transferability among Large Language Models on the Platonic Representations of Concepts
von: Huang, Youcheng, et al.
Veröffentlicht: (2025)
von: Huang, Youcheng, et al.
Veröffentlicht: (2025)
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information
von: Huang, Youcheng, et al.
Veröffentlicht: (2025)
von: Huang, Youcheng, et al.
Veröffentlicht: (2025)
Towards Understanding the Influence of Reward Margin on Preference Model Performance
von: Qin, Bowen, et al.
Veröffentlicht: (2024)
von: Qin, Bowen, et al.
Veröffentlicht: (2024)
Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective
von: Feng, Duanyu, et al.
Veröffentlicht: (2024)
von: Feng, Duanyu, et al.
Veröffentlicht: (2024)
AMaPO: Adaptive Margin-attached Preference Optimization for Language Model Alignment
von: Deng, Ruibo, et al.
Veröffentlicht: (2025)
von: Deng, Ruibo, et al.
Veröffentlicht: (2025)
ZEBRA: Leveraging Model-Behavioral Knowledge for Zero-Annotation Preference Dataset Construction
von: Jung, Jeesu, et al.
Veröffentlicht: (2025)
von: Jung, Jeesu, et al.
Veröffentlicht: (2025)
Dishonesty in Helpful and Harmless Alignment
von: Huang, Youcheng, et al.
Veröffentlicht: (2024)
von: Huang, Youcheng, et al.
Veröffentlicht: (2024)
SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model
von: Zhang, Yongting, et al.
Veröffentlicht: (2024)
von: Zhang, Yongting, et al.
Veröffentlicht: (2024)
Beyond Persuasion: Towards Conversational Recommender System with Credible Explanations
von: Qin, Peixin, et al.
Veröffentlicht: (2024)
von: Qin, Peixin, et al.
Veröffentlicht: (2024)
DREditor: An Time-efficient Approach for Building a Domain-specific Dense Retrieval Model
von: Huang, Chen, et al.
Veröffentlicht: (2024)
von: Huang, Chen, et al.
Veröffentlicht: (2024)
Concept -- An Evaluation Protocol on Conversational Recommender Systems with System-centric and User-centric Factors
von: Huang, Chen, et al.
Veröffentlicht: (2024)
von: Huang, Chen, et al.
Veröffentlicht: (2024)
Larger or Smaller Reward Margins to Select Preferences for Alignment?
von: Huang, Kexin, et al.
Veröffentlicht: (2025)
von: Huang, Kexin, et al.
Veröffentlicht: (2025)
Towards Proactive Information Probing: Customer Service Chatbots Harvesting Value from Conversation
von: Huang, Chen, et al.
Veröffentlicht: (2026)
von: Huang, Chen, et al.
Veröffentlicht: (2026)
Beyond Prompt: Fine-grained Simulation of Cognitively Impaired Standardized Patients via Stochastic Steering
von: Zhang, Weikang, et al.
Veröffentlicht: (2026)
von: Zhang, Weikang, et al.
Veröffentlicht: (2026)
Preference Ranking Optimization for Human Alignment
von: Song, Feifan, et al.
Veröffentlicht: (2023)
von: Song, Feifan, et al.
Veröffentlicht: (2023)
Adaptive Margin RLHF via Preference over Preferences
von: Chittepu, Yaswanth, et al.
Veröffentlicht: (2025)
von: Chittepu, Yaswanth, et al.
Veröffentlicht: (2025)
Reasoning with OmniThought: A Large CoT Dataset with Verbosity and Cognitive Difficulty Annotations
von: Cai, Wenrui, et al.
Veröffentlicht: (2025)
von: Cai, Wenrui, et al.
Veröffentlicht: (2025)
CMDAG: A Chinese Metaphor Dataset with Annotated Grounds as CoT for Boosting Metaphor Generation
von: Shao, Yujie, et al.
Veröffentlicht: (2024)
von: Shao, Yujie, et al.
Veröffentlicht: (2024)
Lying with Truths: Open-Channel Multi-Agent Collusion for Belief Manipulation via Generative Montage
von: Hu, Jinwei, et al.
Veröffentlicht: (2026)
von: Hu, Jinwei, et al.
Veröffentlicht: (2026)
LRHP: Learning Representations for Human Preferences via Preference Pairs
von: Wang, Chenglong, et al.
Veröffentlicht: (2024)
von: Wang, Chenglong, et al.
Veröffentlicht: (2024)
METRO: Towards Strategy Induction from Expert Dialogue Transcripts for Non-collaborative Dialogues
von: Yang, Haofu, et al.
Veröffentlicht: (2026)
von: Yang, Haofu, et al.
Veröffentlicht: (2026)
Disentangling Preference Representation and Text Generation for Efficient Individual Preference Alignment
von: Zhang, Jianfei, et al.
Veröffentlicht: (2024)
von: Zhang, Jianfei, et al.
Veröffentlicht: (2024)
Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective
von: Li, Tianlong, et al.
Veröffentlicht: (2024)
von: Li, Tianlong, et al.
Veröffentlicht: (2024)
E2Edev: Benchmarking Large Language Models in End-to-End Software Development Task
von: Liu, Jingyao, et al.
Veröffentlicht: (2025)
von: Liu, Jingyao, et al.
Veröffentlicht: (2025)
PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
von: Ji, Jiaming, et al.
Veröffentlicht: (2024)
von: Ji, Jiaming, et al.
Veröffentlicht: (2024)
Users as Annotators: LLM Preference Learning from Comparison Mode
von: Cai, Zhongze, et al.
Veröffentlicht: (2025)
von: Cai, Zhongze, et al.
Veröffentlicht: (2025)
AISafetyLab: A Comprehensive Framework for AI Safety Evaluation and Improvement
von: Zhang, Zhexin, et al.
Veröffentlicht: (2025)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2025)
Memory-Augmented Knowledge Fusion with Safety-Aware Decoding for Domain-Adaptive Question Answering
von: Fu, Lei, et al.
Veröffentlicht: (2025)
von: Fu, Lei, et al.
Veröffentlicht: (2025)
STAR: Semantic-Tuned and Tail-Adaptive Retriever for Graph-Augmented Generation
von: Li, Shuai, et al.
Veröffentlicht: (2026)
von: Li, Shuai, et al.
Veröffentlicht: (2026)
PLaD: Preference-based Large Language Model Distillation with Pseudo-Preference Pairs
von: Zhang, Rongzhi, et al.
Veröffentlicht: (2024)
von: Zhang, Rongzhi, et al.
Veröffentlicht: (2024)
DynamicGTR: Leveraging Graph Topology Representation Preferences to Boost VLM Capabilities on Graph QAs
von: Wei, Yanbin, et al.
Veröffentlicht: (2026)
von: Wei, Yanbin, et al.
Veröffentlicht: (2026)
How to Enable Effective Cooperation Between Humans and NLP Models: A Survey of Principles, Formalizations, and Beyond
von: Huang, Chen, et al.
Veröffentlicht: (2025)
von: Huang, Chen, et al.
Veröffentlicht: (2025)
Entity Alignment with Noisy Annotations from Large Language Models
von: Chen, Shengyuan, et al.
Veröffentlicht: (2024)
von: Chen, Shengyuan, et al.
Veröffentlicht: (2024)
METER: Evaluating Multi-Level Contextual Causal Reasoning in Large Language Models
von: Li, Pengfeng, et al.
Veröffentlicht: (2026)
von: Li, Pengfeng, et al.
Veröffentlicht: (2026)
Responsible Agentic AI Requires Explicit Provenance
von: Hu, Jinwei, et al.
Veröffentlicht: (2026)
von: Hu, Jinwei, et al.
Veröffentlicht: (2026)
Self-supervised Preference Optimization: Enhance Your Language Model with Preference Degree Awareness
von: Li, Jian, et al.
Veröffentlicht: (2024)
von: Li, Jian, et al.
Veröffentlicht: (2024)
Empowering Many, Biasing a Few: Generalist Credit Scoring through Large Language Models
von: Feng, Duanyu, et al.
Veröffentlicht: (2023)
von: Feng, Duanyu, et al.
Veröffentlicht: (2023)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
von: Kim, Dongyoung, et al.
Veröffentlicht: (2024)
von: Kim, Dongyoung, et al.
Veröffentlicht: (2024)
Towards Data-Centric RLHF: Simple Metrics for Preference Dataset Comparison
von: Shen, Judy Hanwen, et al.
Veröffentlicht: (2024)
von: Shen, Judy Hanwen, et al.
Veröffentlicht: (2024)
See the Unseen: Better Context-Consistent Knowledge-Editing by Noises
von: Huang, Youcheng, et al.
Veröffentlicht: (2024)
von: Huang, Youcheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Cross-model Transferability among Large Language Models on the Platonic Representations of Concepts
von: Huang, Youcheng, et al.
Veröffentlicht: (2025) -
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information
von: Huang, Youcheng, et al.
Veröffentlicht: (2025) -
Towards Understanding the Influence of Reward Margin on Preference Model Performance
von: Qin, Bowen, et al.
Veröffentlicht: (2024) -
Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective
von: Feng, Duanyu, et al.
Veröffentlicht: (2024) -
AMaPO: Adaptive Margin-attached Preference Optimization for Language Model Alignment
von: Deng, Ruibo, et al.
Veröffentlicht: (2025)