Density-Guided Response Optimization: Community-Grounded Alignment via Implicit Acceptance Signals
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Gerard, Patrick, Volkova, Svitlana |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Community-Aligned Behavior Under Uncertainty: Evidence of Epistemic Stance Transfer in LLMs
par: Gerard, Patrick, et autres
Publié: (2025)
par: Gerard, Patrick, et autres
Publié: (2025)
Proactive Defense: Compound AI for Detecting Persuasion Attacks and Measuring Inoculation Effectiveness
par: Volkova, Svitlana, et autres
Publié: (2025)
par: Volkova, Svitlana, et autres
Publié: (2025)
OASST-ETC Dataset: Alignment Signals from Eye-tracking Analysis of LLM Responses
par: Lopez-Cardona, Angela, et autres
Publié: (2025)
par: Lopez-Cardona, Angela, et autres
Publié: (2025)
On-the-fly Preference Alignment via Principle-Guided Decoding
par: Zhu, Mingye, et autres
Publié: (2025)
par: Zhu, Mingye, et autres
Publié: (2025)
MuPHI: Learning Implicit Multimodal Harm Reasoning via Semantically Grounded Reward Optimization
par: Saha, Anisha, et autres
Publié: (2026)
par: Saha, Anisha, et autres
Publié: (2026)
Implicit Cross-Lingual Rewarding for Efficient Multilingual Preference Alignment
par: Yang, Wen, et autres
Publié: (2025)
par: Yang, Wen, et autres
Publié: (2025)
KPC-cF: Aspect-Based Sentiment Analysis via Implicit-Feature Alignment with Corpus Filtering
par: Nam, Kibeom
Publié: (2024)
par: Nam, Kibeom
Publié: (2024)
On the Limitations of Steering in Language Model Alignment
par: Niranjan, Chebrolu, et autres
Publié: (2025)
par: Niranjan, Chebrolu, et autres
Publié: (2025)
Exploring Big Five Personality and AI Capability Effects in LLM-Simulated Negotiation Dialogues
par: Cohen, Myke C., et autres
Publié: (2025)
par: Cohen, Myke C., et autres
Publié: (2025)
WISTERIA: Weak Implicit Signal-based Temporal Relation Extraction with Attention
par: Do, Duy Dao, et autres
Publié: (2026)
par: Do, Duy Dao, et autres
Publié: (2026)
InFact: Informativeness Alignment for Improved LLM Factuality
par: Cohen, Roi, et autres
Publié: (2025)
par: Cohen, Roi, et autres
Publié: (2025)
Empathy Level Alignment via Reinforcement Learning for Empathetic Response Generation
par: Ma, Hui, et autres
Publié: (2024)
par: Ma, Hui, et autres
Publié: (2024)
Graph Alignment Topology as an Inductive Bias for Grounding Detection
par: Landes, Paul, et autres
Publié: (2026)
par: Landes, Paul, et autres
Publié: (2026)
Improving Implicit Hate Speech Detection via a Community-Driven Multi-Agent Framework
par: Gajewska, Ewelina, et autres
Publié: (2026)
par: Gajewska, Ewelina, et autres
Publié: (2026)
Relative Density Ratio Optimization for Stable and Statistically Consistent Model Alignment
par: Takahashi, Hiroshi, et autres
Publié: (2026)
par: Takahashi, Hiroshi, et autres
Publié: (2026)
Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race
par: Sun, Lihao, et autres
Publié: (2025)
par: Sun, Lihao, et autres
Publié: (2025)
Implicit Behavioral Alignment of Language Agents in High-Stakes Crowd Simulations
par: Wang, Yunzhe, et autres
Publié: (2025)
par: Wang, Yunzhe, et autres
Publié: (2025)
Knowledge Graphs are Implicit Reward Models: Path-Derived Signals Enable Compositional Reasoning
par: Kansal, Yuval, et autres
Publié: (2026)
par: Kansal, Yuval, et autres
Publié: (2026)
Self-Guided Defense: Adaptive Safety Alignment for Reasoning Models via Synthesized Guidelines
par: Wang, Yuhang, et autres
Publié: (2025)
par: Wang, Yuhang, et autres
Publié: (2025)
How to Understand "Support"? An Implicit-enhanced Causal Inference Approach for Weakly-supervised Phrase Grounding
par: Luo, Jiamin, et autres
Publié: (2024)
par: Luo, Jiamin, et autres
Publié: (2024)
PA-RAG: RAG Alignment via Multi-Perspective Preference Optimization
par: Wu, Jiayi, et autres
Publié: (2024)
par: Wu, Jiayi, et autres
Publié: (2024)
REAL: Response Embedding-based Alignment for LLMs
par: Zhang, Honggen, et autres
Publié: (2024)
par: Zhang, Honggen, et autres
Publié: (2024)
MACPO: Weak-to-Strong Alignment via Multi-Agent Contrastive Preference Optimization
par: Lyu, Yougang, et autres
Publié: (2024)
par: Lyu, Yougang, et autres
Publié: (2024)
Negating Negatives: Alignment with Human Negative Samples via Distributional Dispreference Optimization
par: Duan, Shitong, et autres
Publié: (2024)
par: Duan, Shitong, et autres
Publié: (2024)
ProFit: Leveraging High-Value Signals in SFT via Probability-Guided Token Selection
par: Liu, Tao, et autres
Publié: (2026)
par: Liu, Tao, et autres
Publié: (2026)
Polypersona: Persona-Grounded LLM for Synthetic Survey Responses
par: Dash, Tejaswani, et autres
Publié: (2025)
par: Dash, Tejaswani, et autres
Publié: (2025)
MMA-ASIA: A Multilingual and Multimodal Alignment Framework for Culturally-Grounded Evaluation
par: Zheng, Weihua, et autres
Publié: (2025)
par: Zheng, Weihua, et autres
Publié: (2025)
Acceptance Dynamics Across Cognitive Domains in Speculative Decoding
par: Mahmoud, Saif
Publié: (2026)
par: Mahmoud, Saif
Publié: (2026)
Preference Ranking Optimization for Human Alignment
par: Song, Feifan, et autres
Publié: (2023)
par: Song, Feifan, et autres
Publié: (2023)
ICDPO: Effectively Borrowing Alignment Capability of Others via In-context Direct Preference Optimization
par: Song, Feifan, et autres
Publié: (2024)
par: Song, Feifan, et autres
Publié: (2024)
Towards Interpretable Time Series Foundation Models
par: Boileau, Matthieu, et autres
Publié: (2025)
par: Boileau, Matthieu, et autres
Publié: (2025)
Nudging: Inference-time Alignment of LLMs via Guided Decoding
par: Fei, Yu, et autres
Publié: (2024)
par: Fei, Yu, et autres
Publié: (2024)
Modeling Information Narrative Detection and Evolution on Telegram during the Russia-Ukraine War
par: Gerard, Patrick, et autres
Publié: (2024)
par: Gerard, Patrick, et autres
Publié: (2024)
Factors That Support Grounded Responses in LLM Conversations: A Rapid Review
par: Iwashima, Gabriele Cesar, et autres
Publié: (2025)
par: Iwashima, Gabriele Cesar, et autres
Publié: (2025)
PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
par: Jiang, Han, et autres
Publié: (2025)
par: Jiang, Han, et autres
Publié: (2025)
HumanLM: Simulating Users with State Alignment Beats Response Imitation
par: Wu, Shirley, et autres
Publié: (2026)
par: Wu, Shirley, et autres
Publié: (2026)
The Thinking Therapist: Training Large Language Models to Deliver Acceptance and Commitment Therapy using Supervised Fine-Tuning and Odds Ratio Policy Optimization
par: Tahir, Talha
Publié: (2025)
par: Tahir, Talha
Publié: (2025)
Imperfectly Cooperative Human-AI Interactions: Comparing the Impacts of Human and AI Attributes in Simulated and User Studies
par: Cohen, Myke C., et autres
Publié: (2026)
par: Cohen, Myke C., et autres
Publié: (2026)
RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization
par: Bao, Qiming, et autres
Publié: (2026)
par: Bao, Qiming, et autres
Publié: (2026)
PrinciplismQA: A Philosophy-Grounded Approach to Assessing LLM-Human Clinical Medical Ethics Alignment
par: Hong, Chang, et autres
Publié: (2025)
par: Hong, Chang, et autres
Publié: (2025)
Documents similaires
-
Community-Aligned Behavior Under Uncertainty: Evidence of Epistemic Stance Transfer in LLMs
par: Gerard, Patrick, et autres
Publié: (2025) -
Proactive Defense: Compound AI for Detecting Persuasion Attacks and Measuring Inoculation Effectiveness
par: Volkova, Svitlana, et autres
Publié: (2025) -
OASST-ETC Dataset: Alignment Signals from Eye-tracking Analysis of LLM Responses
par: Lopez-Cardona, Angela, et autres
Publié: (2025) -
On-the-fly Preference Alignment via Principle-Guided Decoding
par: Zhu, Mingye, et autres
Publié: (2025) -
MuPHI: Learning Implicit Multimodal Harm Reasoning via Semantically Grounded Reward Optimization
par: Saha, Anisha, et autres
Publié: (2026)