Negating Negatives: Alignment with Human Negative Samples via Distributional Dispreference Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Duan, Shitong, Yi, Xiaoyuan, Zhang, Peng, Liu, Yan, Liu, Zheng, Lu, Tun, Xie, Xing, Gu, Ning |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Denevil: Towards Deciphering and Navigating the Ethical Values of Large Language Models via Instruction Learning
von: Duan, Shitong, et al.
Veröffentlicht: (2023)
von: Duan, Shitong, et al.
Veröffentlicht: (2023)
IROTE: Human-like Traits Elicitation of Large Language Model via In-Context Self-Reflective Optimization
von: Bai, Yuzhuo, et al.
Veröffentlicht: (2025)
von: Bai, Yuzhuo, et al.
Veröffentlicht: (2025)
AdAEM: An Adaptively and Automated Extensible Measurement of LLMs' Value Difference
von: Yao, Jing, et al.
Veröffentlicht: (2025)
von: Yao, Jing, et al.
Veröffentlicht: (2025)
Correcting Negative Bias in Large Language Models through Negative Attention Score Alignment
von: Yu, Sangwon, et al.
Veröffentlicht: (2024)
von: Yu, Sangwon, et al.
Veröffentlicht: (2024)
Enhancing Semantics in Multimodal Chain of Thought via Soft Negative Sampling
von: Zheng, Guangmin, et al.
Veröffentlicht: (2024)
von: Zheng, Guangmin, et al.
Veröffentlicht: (2024)
MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples
von: Xie, Shuo, et al.
Veröffentlicht: (2024)
von: Xie, Shuo, et al.
Veröffentlicht: (2024)
On the Essence and Prospect: An Investigation of Alignment Approaches for Big Models
von: Wang, Xinpeng, et al.
Veröffentlicht: (2024)
von: Wang, Xinpeng, et al.
Veröffentlicht: (2024)
More Expressive Attention with Negative Weights
von: Lv, Ang, et al.
Veröffentlicht: (2024)
von: Lv, Ang, et al.
Veröffentlicht: (2024)
Simplicity Prevails: Rethinking Negative Preference Optimization for LLM Unlearning
von: Fan, Chongyu, et al.
Veröffentlicht: (2024)
von: Fan, Chongyu, et al.
Veröffentlicht: (2024)
Elephant in the Room: Unveiling the Impact of Reward Model Quality in Alignment
von: Liu, Yan, et al.
Veröffentlicht: (2024)
von: Liu, Yan, et al.
Veröffentlicht: (2024)
HNCSE: Advancing Sentence Embeddings via Hybrid Contrastive Learning with Hard Negatives
von: Liu, Wenxiao, et al.
Veröffentlicht: (2024)
von: Liu, Wenxiao, et al.
Veröffentlicht: (2024)
Evaluating Negative Sampling Approaches for Neural Topic Models
von: Adhya, Suman, et al.
Veröffentlicht: (2025)
von: Adhya, Suman, et al.
Veröffentlicht: (2025)
Boosting Protein Language Models with Negative Sample Mining
von: Xu, Yaoyao, et al.
Veröffentlicht: (2024)
von: Xu, Yaoyao, et al.
Veröffentlicht: (2024)
DRAGON: Guard LLM Unlearning in Context via Negative Detection and Reasoning
von: Wang, Yaxuan, et al.
Veröffentlicht: (2025)
von: Wang, Yaxuan, et al.
Veröffentlicht: (2025)
Negation Triplet Extraction with Syntactic Dependency and Semantic Consistency
von: Shi, Yuchen, et al.
Veröffentlicht: (2024)
von: Shi, Yuchen, et al.
Veröffentlicht: (2024)
PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
von: Jiang, Han, et al.
Veröffentlicht: (2025)
von: Jiang, Han, et al.
Veröffentlicht: (2025)
Teaching with Lies: Curriculum DPO on Synthetic Negatives for Hallucination Detection
von: Pandit, Shrey, et al.
Veröffentlicht: (2025)
von: Pandit, Shrey, et al.
Veröffentlicht: (2025)
Generating Diverse Negations from Affirmative Sentences
von: Vasquez, Darian Rodriguez, et al.
Veröffentlicht: (2024)
von: Vasquez, Darian Rodriguez, et al.
Veröffentlicht: (2024)
Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference Optimization
von: Duan, Shaohua, et al.
Veröffentlicht: (2025)
von: Duan, Shaohua, et al.
Veröffentlicht: (2025)
Negative Preference Optimization: From Catastrophic Collapse to Effective Unlearning
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2024)
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2024)
Diversified and Adaptive Negative Sampling on Knowledge Graphs
von: Liu, Ran, et al.
Veröffentlicht: (2024)
von: Liu, Ran, et al.
Veröffentlicht: (2024)
ToolNet: Connecting Large Language Models with Massive Tools via Tool Graph
von: Liu, Xukun, et al.
Veröffentlicht: (2024)
von: Liu, Xukun, et al.
Veröffentlicht: (2024)
Diffusion-based Hierarchical Negative Sampling for Multimodal Knowledge Graph Completion
von: Niu, Guanglin, et al.
Veröffentlicht: (2025)
von: Niu, Guanglin, et al.
Veröffentlicht: (2025)
Semantic Gravity Wells: Why Negative Constraints Backfire
von: Rana, Shailesh
Veröffentlicht: (2026)
von: Rana, Shailesh
Veröffentlicht: (2026)
The Impact of Negated Text on Hallucination with Large Language Models
von: Seo, Jaehyung, et al.
Veröffentlicht: (2025)
von: Seo, Jaehyung, et al.
Veröffentlicht: (2025)
Taming LLMs with Negative Samples: A Reference-Free Framework to Evaluate Presentation Content with Actionable Feedback
von: Muppidi, Ananth, et al.
Veröffentlicht: (2025)
von: Muppidi, Ananth, et al.
Veröffentlicht: (2025)
Carrot and Stick: Inducing Self-Motivation with Positive & Negative Feedback
von: Sohn, Jimin, et al.
Veröffentlicht: (2024)
von: Sohn, Jimin, et al.
Veröffentlicht: (2024)
Mitigating the Negative Impact of Over-association for Conversational Query Production
von: Wang, Ante, et al.
Veröffentlicht: (2024)
von: Wang, Ante, et al.
Veröffentlicht: (2024)
NCO: A Versatile Plug-in for Handling Negative Constraints in Decoding
von: Jin, Hyundong, et al.
Veröffentlicht: (2026)
von: Jin, Hyundong, et al.
Veröffentlicht: (2026)
Negative Advantages Is a Double-Edged Sword: Calibrating advantages in GRPO for Search Agents
von: Wu, Jiayi, et al.
Veröffentlicht: (2026)
von: Wu, Jiayi, et al.
Veröffentlicht: (2026)
CLAVE: An Adaptive Framework for Evaluating Values of LLM Generated Responses
von: Yao, Jing, et al.
Veröffentlicht: (2024)
von: Yao, Jing, et al.
Veröffentlicht: (2024)
Momentum Contrastive Learning with Enhanced Negative Sampling and Hard Negative Filtering
von: Hoang, Duy, et al.
Veröffentlicht: (2025)
von: Hoang, Duy, et al.
Veröffentlicht: (2025)
RIFT: Repurposing Negative Samples via Reward-Informed Fine-Tuning
von: Liu, Zehua, et al.
Veröffentlicht: (2026)
von: Liu, Zehua, et al.
Veröffentlicht: (2026)
Boosting Knowledge Graph Foundation Models via Enhanced Negative Sampling
von: Liu, Yinan, et al.
Veröffentlicht: (2026)
von: Liu, Yinan, et al.
Veröffentlicht: (2026)
Beyond Human Norms: Unveiling Unique Values of Large Language Models through Interdisciplinary Approaches
von: Biedma, Pablo, et al.
Veröffentlicht: (2024)
von: Biedma, Pablo, et al.
Veröffentlicht: (2024)
InceptionXML: A Lightweight Framework with Synchronized Negative Sampling for Short Text Extreme Classification
von: Kharbanda, Siddhant, et al.
Veröffentlicht: (2021)
von: Kharbanda, Siddhant, et al.
Veröffentlicht: (2021)
Towards Minimal Targeted Updates of Language Models with Targeted Negative Training
von: Zhang, Lily H., et al.
Veröffentlicht: (2024)
von: Zhang, Lily H., et al.
Veröffentlicht: (2024)
MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models?
von: Yong, Xixian, et al.
Veröffentlicht: (2025)
von: Yong, Xixian, et al.
Veröffentlicht: (2025)
Causal Negative Sampling via Diffusion Model for Out-of-Distribution Recommendation
von: Zhao, Chu, et al.
Veröffentlicht: (2025)
von: Zhao, Chu, et al.
Veröffentlicht: (2025)
MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual Questions
von: Zhu, Yanxu, et al.
Veröffentlicht: (2025)
von: Zhu, Yanxu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Denevil: Towards Deciphering and Navigating the Ethical Values of Large Language Models via Instruction Learning
von: Duan, Shitong, et al.
Veröffentlicht: (2023) -
IROTE: Human-like Traits Elicitation of Large Language Model via In-Context Self-Reflective Optimization
von: Bai, Yuzhuo, et al.
Veröffentlicht: (2025) -
AdAEM: An Adaptively and Automated Extensible Measurement of LLMs' Value Difference
von: Yao, Jing, et al.
Veröffentlicht: (2025) -
Correcting Negative Bias in Large Language Models through Negative Attention Score Alignment
von: Yu, Sangwon, et al.
Veröffentlicht: (2024) -
Enhancing Semantics in Multimodal Chain of Thought via Soft Negative Sampling
von: Zheng, Guangmin, et al.
Veröffentlicht: (2024)