The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bou, Matthieu, Patel, Nyal, Jagota, Arjun, Krishna, Satyapriya, Parbhoo, Sonali |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning from Failures: Understanding LLM Alignment through Failure-Aware Inverse RL
von: Patel, Nyal, et al.
Veröffentlicht: (2025)
von: Patel, Nyal, et al.
Veröffentlicht: (2025)
Insights from the Inverse: Reconstructing LLM Training Goals Through Inverse Reinforcement Learning
von: Joselowitz, Jared, et al.
Veröffentlicht: (2024)
von: Joselowitz, Jared, et al.
Veröffentlicht: (2024)
On Almost Surely Safe Alignment of Large Language Models at Inference-Time
von: Ji, Xiaotong, et al.
Veröffentlicht: (2025)
von: Ji, Xiaotong, et al.
Veröffentlicht: (2025)
References Improve LLM Alignment in Non-Verifiable Domains
von: Shi, Kejian, et al.
Veröffentlicht: (2026)
von: Shi, Kejian, et al.
Veröffentlicht: (2026)
Tree-Based Leakage Inspection and Control in Concept Bottleneck Models
von: Ragkousis, Angelos, et al.
Veröffentlicht: (2024)
von: Ragkousis, Angelos, et al.
Veröffentlicht: (2024)
Uncovering Cross-Objective Interference in Multi-Objective Alignment
von: Lu, Yining, et al.
Veröffentlicht: (2026)
von: Lu, Yining, et al.
Veröffentlicht: (2026)
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
von: Kroeger, Nicholas, et al.
Veröffentlicht: (2023)
von: Kroeger, Nicholas, et al.
Veröffentlicht: (2023)
Improving LLM Safety Alignment with Dual-Objective Optimization
von: Zhao, Xuandong, et al.
Veröffentlicht: (2025)
von: Zhao, Xuandong, et al.
Veröffentlicht: (2025)
Why LLMs Fail at Causal Discovery and How Interventional Agents Escape
von: Roy, Amartya, et al.
Veröffentlicht: (2026)
von: Roy, Amartya, et al.
Veröffentlicht: (2026)
Causal Bayesian Optimization with Unknown Graphs
von: Durand, Jean, et al.
Veröffentlicht: (2025)
von: Durand, Jean, et al.
Veröffentlicht: (2025)
Drifting Objectives for Refining Discrete Diffusion Language Models
von: Oba, Daisuke, et al.
Veröffentlicht: (2026)
von: Oba, Daisuke, et al.
Veröffentlicht: (2026)
Importing Phantoms: Measuring LLM Package Hallucination Vulnerabilities
von: Krishna, Arjun, et al.
Veröffentlicht: (2025)
von: Krishna, Arjun, et al.
Veröffentlicht: (2025)
Compress the Context, Keep the Commitments: A Formal Framework for Verifiable LLM Context Compression
von: Trukhina, Natalia, et al.
Veröffentlicht: (2026)
von: Trukhina, Natalia, et al.
Veröffentlicht: (2026)
Robust Multi-Objective Preference Alignment with Online DPO
von: Gupta, Raghav, et al.
Veröffentlicht: (2025)
von: Gupta, Raghav, et al.
Veröffentlicht: (2025)
Multi-Objective Alignment of Language Models for Personalized Psychotherapy
von: Beikzadeh, Mehrab, et al.
Veröffentlicht: (2026)
von: Beikzadeh, Mehrab, et al.
Veröffentlicht: (2026)
Discovering Implicit Large Language Model Alignment Objectives
von: Chen, Edward, et al.
Veröffentlicht: (2026)
von: Chen, Edward, et al.
Veröffentlicht: (2026)
VerifierQ: Enhancing LLM Test Time Compute with Q-Learning-based Verifiers
von: Qi, Jianing, et al.
Veröffentlicht: (2024)
von: Qi, Jianing, et al.
Veröffentlicht: (2024)
Solving the Inverse Alignment Problem for Efficient RLHF
von: Krishna, Shambhavi, et al.
Veröffentlicht: (2024)
von: Krishna, Shambhavi, et al.
Veröffentlicht: (2024)
DualDiffusion: A Speculative Decoding Strategy for Masked Diffusion Models
von: Goyal, Satyam, et al.
Veröffentlicht: (2026)
von: Goyal, Satyam, et al.
Veröffentlicht: (2026)
Blockwise Advantage Estimation for Multi-Objective RL with Verifiable Rewards
von: Pavlenko, Kirill, et al.
Veröffentlicht: (2026)
von: Pavlenko, Kirill, et al.
Veröffentlicht: (2026)
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
Reward-free Alignment for Conflicting Objectives
von: Chen, Peter, et al.
Veröffentlicht: (2026)
von: Chen, Peter, et al.
Veröffentlicht: (2026)
De Jure: Iterative LLM Self-Refinement for Structured Extraction of Regulatory Rules
von: Guliani, Keerat, et al.
Veröffentlicht: (2026)
von: Guliani, Keerat, et al.
Veröffentlicht: (2026)
Learning to Optimize Multi-Objective Alignment Through Dynamic Reward Weighting
von: Lu, Yining, et al.
Veröffentlicht: (2025)
von: Lu, Yining, et al.
Veröffentlicht: (2025)
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2026)
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2026)
Do regularization methods for shortcut mitigation work as intended?
von: Hong, Haoyang, et al.
Veröffentlicht: (2025)
von: Hong, Haoyang, et al.
Veröffentlicht: (2025)
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
von: Setlur, Amrith, et al.
Veröffentlicht: (2024)
von: Setlur, Amrith, et al.
Veröffentlicht: (2024)
Refined Direct Preference Optimization with Synthetic Data for Behavioral Alignment of LLMs
von: Gallego, Víctor
Veröffentlicht: (2024)
von: Gallego, Víctor
Veröffentlicht: (2024)
OrthAlign: Orthogonal Subspace Decomposition for Non-Interfering Multi-Objective Alignment
von: Lin, Liang, et al.
Veröffentlicht: (2025)
von: Lin, Liang, et al.
Veröffentlicht: (2025)
More Than a Score: Probing the Impact of Prompt Specificity on LLM Code Generation
von: Zi, Yangtian, et al.
Veröffentlicht: (2025)
von: Zi, Yangtian, et al.
Veröffentlicht: (2025)
Pareto Multi-Objective Alignment for Language Models
von: He, Qiang, et al.
Veröffentlicht: (2025)
von: He, Qiang, et al.
Veröffentlicht: (2025)
Bayesian Inverse Transition Learning: Learning Dynamics From Near-Optimal Trajectories
von: Benac, Leo, et al.
Veröffentlicht: (2024)
von: Benac, Leo, et al.
Veröffentlicht: (2024)
Mixture of Attentions For Speculative Decoding
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)
Self-Improvement Towards Pareto Optimality: Mitigating Preference Conflicts in Multi-Objective Alignment
von: Li, Moxin, et al.
Veröffentlicht: (2025)
von: Li, Moxin, et al.
Veröffentlicht: (2025)
Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs
von: Shukor, Mustafa, et al.
Veröffentlicht: (2024)
von: Shukor, Mustafa, et al.
Veröffentlicht: (2024)
Refining Input Guardrails: Enhancing LLM-as-a-Judge Efficiency Through Chain-of-Thought Fine-Tuning and Alignment
von: Rad, Melissa Kazemi, et al.
Veröffentlicht: (2025)
von: Rad, Melissa Kazemi, et al.
Veröffentlicht: (2025)
SparsePO: Controlling Preference Alignment of LLMs via Sparse Token Masks
von: Christopoulou, Fenia, et al.
Veröffentlicht: (2024)
von: Christopoulou, Fenia, et al.
Veröffentlicht: (2024)
Concept-driven Off Policy Evaluation
von: Majumdar, Ritam, et al.
Veröffentlicht: (2024)
von: Majumdar, Ritam, et al.
Veröffentlicht: (2024)
Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards
von: Shen, Yiran, et al.
Veröffentlicht: (2025)
von: Shen, Yiran, et al.
Veröffentlicht: (2025)
Guarantee Regions for Local Explanations
von: Havasi, Marton, et al.
Veröffentlicht: (2024)
von: Havasi, Marton, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Learning from Failures: Understanding LLM Alignment through Failure-Aware Inverse RL
von: Patel, Nyal, et al.
Veröffentlicht: (2025) -
Insights from the Inverse: Reconstructing LLM Training Goals Through Inverse Reinforcement Learning
von: Joselowitz, Jared, et al.
Veröffentlicht: (2024) -
On Almost Surely Safe Alignment of Large Language Models at Inference-Time
von: Ji, Xiaotong, et al.
Veröffentlicht: (2025) -
References Improve LLM Alignment in Non-Verifiable Domains
von: Shi, Kejian, et al.
Veröffentlicht: (2026) -
Tree-Based Leakage Inspection and Control in Concept Bottleneck Models
von: Ragkousis, Angelos, et al.
Veröffentlicht: (2024)