Anchored Preference Optimization and Contrastive Revisions: Addressing Underspecification in Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | D'Oosterlinck, Karel, Xu, Winnie, Develder, Chris, Demeester, Thomas, Singh, Amanpreet, Potts, Christopher, Kiela, Douwe, Mehri, Shikib |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
In-Context Learning for Extreme Multi-Label Classification
von: D'Oosterlinck, Karel, et al.
Veröffentlicht: (2024)
von: D'Oosterlinck, Karel, et al.
Veröffentlicht: (2024)
Updating CLIP to Prefer Descriptions Over Captions
von: Zur, Amir, et al.
Veröffentlicht: (2024)
von: Zur, Amir, et al.
Veröffentlicht: (2024)
Reflective Context Learning: Studying the Optimization Primitives of Context Space
von: Vassilyev, Nikita, et al.
Veröffentlicht: (2026)
von: Vassilyev, Nikita, et al.
Veröffentlicht: (2026)
KTO: Model Alignment as Prospect Theoretic Optimization
von: Ethayarajh, Kawin, et al.
Veröffentlicht: (2024)
von: Ethayarajh, Kawin, et al.
Veröffentlicht: (2024)
LMUnit: Fine-grained Evaluation with Natural Language Unit Tests
von: Saad-Falcon, Jon, et al.
Veröffentlicht: (2024)
von: Saad-Falcon, Jon, et al.
Veröffentlicht: (2024)
Document Optimization for Black-Box Retrieval via Reinforcement Learning
von: Uzan, Omri, et al.
Veröffentlicht: (2026)
von: Uzan, Omri, et al.
Veröffentlicht: (2026)
HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
GAMBIT: A Three-Mode Benchmark for Adversarial Robustness in Multi-Agent LLM Collectives
von: Mercier, Alexandre Le, et al.
Veröffentlicht: (2026)
von: Mercier, Alexandre Le, et al.
Veröffentlicht: (2026)
CLASP: Defending Hybrid Large Language Models Against Hidden State Poisoning Attacks
von: Mercier, Alexandre Le, et al.
Veröffentlicht: (2026)
von: Mercier, Alexandre Le, et al.
Veröffentlicht: (2026)
Hidden State Poisoning Attacks against Mamba-based Language Models
von: Mercier, Alexandre Le, et al.
Veröffentlicht: (2026)
von: Mercier, Alexandre Le, et al.
Veröffentlicht: (2026)
Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets
von: Tran, Dat, et al.
Veröffentlicht: (2026)
von: Tran, Dat, et al.
Veröffentlicht: (2026)
I am a Strange Dataset: Metalinguistic Tests for Language Models
von: Thrush, Tristan, et al.
Veröffentlicht: (2024)
von: Thrush, Tristan, et al.
Veröffentlicht: (2024)
Goal Alignment in LLM-Based User Simulators for Conversational AI
von: Mehri, Shuhaib, et al.
Veröffentlicht: (2025)
von: Mehri, Shuhaib, et al.
Veröffentlicht: (2025)
Single- vs. Dual-Prompt Dialogue Generation with LLMs for Job Interviews in Human Resources
von: De Baer, Joachim, et al.
Veröffentlicht: (2025)
von: De Baer, Joachim, et al.
Veröffentlicht: (2025)
SkillMatch: Evaluating Self-supervised Learning of Skill Relatedness
von: Decorte, Jens-Joris, et al.
Veröffentlicht: (2024)
von: Decorte, Jens-Joris, et al.
Veröffentlicht: (2024)
Efficient Text Encoders for Labor Market Analysis
von: Decorte, Jens-Joris, et al.
Veröffentlicht: (2025)
von: Decorte, Jens-Joris, et al.
Veröffentlicht: (2025)
On the Biased Assessment of Expert Finding Systems
von: Decorte, Jens-Joris, et al.
Veröffentlicht: (2024)
von: Decorte, Jens-Joris, et al.
Veröffentlicht: (2024)
Generative Representational Instruction Tuning
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2024)
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2024)
BlitzRank: Principled Zero-shot Ranking Agents with Tournament Graphs
von: Agrawal, Sheshansh, et al.
Veröffentlicht: (2026)
von: Agrawal, Sheshansh, et al.
Veröffentlicht: (2026)
Composing Policy Gradients and Prompt Optimization for Language Model Programs
von: Ziems, Noah, et al.
Veröffentlicht: (2025)
von: Ziems, Noah, et al.
Veröffentlicht: (2025)
Anchor Points: Benchmarking Models with Much Fewer Examples
von: Vivek, Rajan, et al.
Veröffentlicht: (2023)
von: Vivek, Rajan, et al.
Veröffentlicht: (2023)
Leveraging Diffusion Perturbations for Measuring Fairness in Computer Vision
von: Lui, Nicholas, et al.
Veröffentlicht: (2023)
von: Lui, Nicholas, et al.
Veröffentlicht: (2023)
Lynx: An Open Source Hallucination Evaluation Model
von: Ravi, Selvan Sunitha, et al.
Veröffentlicht: (2024)
von: Ravi, Selvan Sunitha, et al.
Veröffentlicht: (2024)
Accounting for Underspecification in Statistical Claims of Model Superiority
von: Sanchez, Thomas, et al.
Veröffentlicht: (2025)
von: Sanchez, Thomas, et al.
Veröffentlicht: (2025)
Nearest Neighbor Normalization Improves Multimodal Retrieval
von: Chowdhury, Neil, et al.
Veröffentlicht: (2024)
von: Chowdhury, Neil, et al.
Veröffentlicht: (2024)
Transcendental Optimization in Geometric Programming via Power Series Approximations
von: singh, Amanpreet
Veröffentlicht: (2025)
von: singh, Amanpreet
Veröffentlicht: (2025)
Enhancing Software Vulnerability Detection Using Code Property Graphs and Convolutional Neural Networks
von: Saimbhi, Amanpreet Singh
Veröffentlicht: (2025)
von: Saimbhi, Amanpreet Singh
Veröffentlicht: (2025)
Querying Databases with Function Calling
von: Shorten, Connor, et al.
Veröffentlicht: (2025)
von: Shorten, Connor, et al.
Veröffentlicht: (2025)
ADPO: Anchored Direct Preference Optimization
von: Zixian, Wang
Veröffentlicht: (2025)
von: Zixian, Wang
Veröffentlicht: (2025)
LLM Theory of Mind and Alignment: Opportunities and Risks
von: Street, Winnie
Veröffentlicht: (2024)
von: Street, Winnie
Veröffentlicht: (2024)
Multi-market value-stacking: Battery control for combined imbalance participation and non-uniform FCR bidding
von: Hendrickx, Celle, et al.
Veröffentlicht: (2026)
von: Hendrickx, Celle, et al.
Veröffentlicht: (2026)
Explainable Reinforcement Learning-based Home Energy Management Systems using Differentiable Decision Trees
von: Gokhale, Gargya, et al.
Veröffentlicht: (2024)
von: Gokhale, Gargya, et al.
Veröffentlicht: (2024)
Robustness of Spatio-temporal Graph Neural Networks for Fault Location in Partially Observable Distribution Grids
von: Karabulut, Burak, et al.
Veröffentlicht: (2026)
von: Karabulut, Burak, et al.
Veröffentlicht: (2026)
Generalization of Graph Neural Network Models for Distribution Grid Fault Detection
von: Karabulut, Burak, et al.
Veröffentlicht: (2025)
von: Karabulut, Burak, et al.
Veröffentlicht: (2025)
FoldA: Computing Partial-Order Alignments Using Directed Net Unfoldings
von: Geurtjens, Douwe, et al.
Veröffentlicht: (2025)
von: Geurtjens, Douwe, et al.
Veröffentlicht: (2025)
LHAW: Controllable Underspecification for Long-Horizon Tasks
von: Pu, George, et al.
Veröffentlicht: (2026)
von: Pu, George, et al.
Veröffentlicht: (2026)
Learning Where It Matters: Geometric Anchoring for Robust Preference Alignment
von: Cho, Youngjae, et al.
Veröffentlicht: (2026)
von: Cho, Youngjae, et al.
Veröffentlicht: (2026)
Dynamic Expert-Guided Model Averaging for Causal Discovery
von: Tench, Adrick, et al.
Veröffentlicht: (2026)
von: Tench, Adrick, et al.
Veröffentlicht: (2026)
Annotation-Assisted Learning of Treatment Policies From Multimodal Electronic Health Records
von: Arno, Henri, et al.
Veröffentlicht: (2025)
von: Arno, Henri, et al.
Veröffentlicht: (2025)
ExtractBench: A Benchmark and Evaluation Methodology for Complex Structured Extraction
von: Ferguson, Nick, et al.
Veröffentlicht: (2026)
von: Ferguson, Nick, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
In-Context Learning for Extreme Multi-Label Classification
von: D'Oosterlinck, Karel, et al.
Veröffentlicht: (2024) -
Updating CLIP to Prefer Descriptions Over Captions
von: Zur, Amir, et al.
Veröffentlicht: (2024) -
Reflective Context Learning: Studying the Optimization Primitives of Context Space
von: Vassilyev, Nikita, et al.
Veröffentlicht: (2026) -
KTO: Model Alignment as Prospect Theoretic Optimization
von: Ethayarajh, Kawin, et al.
Veröffentlicht: (2024) -
LMUnit: Fine-grained Evaluation with Natural Language Unit Tests
von: Saad-Falcon, Jon, et al.
Veröffentlicht: (2024)