On Almost Surely Safe Alignment of Large Language Models at Inference-Time
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ji, Xiaotong, Ramesh, Shyam Sundhar, Zimmer, Matthieu, Bogunovic, Ilija, Wang, Jun, Ammar, Haitham Bou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2026)
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2026)
Group Robust Preference Optimization in Reward-free RLHF
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2024)
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2024)
Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2025)
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2025)
Decoding as Optimisation on the Probability Simplex: From Top-K to Top-P (Nucleus) to Best-of-K Samplers
von: Ji, Xiaotong, et al.
Veröffentlicht: (2026)
von: Ji, Xiaotong, et al.
Veröffentlicht: (2026)
Scalable Power Sampling: Unlocking Efficient, Training-Free Reasoning for LLMs via Distribution Sharpening
von: Ji, Xiaotong, et al.
Veröffentlicht: (2026)
von: Ji, Xiaotong, et al.
Veröffentlicht: (2026)
Robust Multi-Objective Controlled Decoding of Large Language Models
von: Son, Seongho, et al.
Veröffentlicht: (2025)
von: Son, Seongho, et al.
Veröffentlicht: (2025)
Distributionally Robust Model-based Reinforcement Learning with Large State Spaces
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2023)
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2023)
Mixture of Attentions For Speculative Decoding
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)
The Model Knows, the Decoder Finds: Future Value Guided Particle Power Sampling
von: Nguyen, Tu, et al.
Veröffentlicht: (2026)
von: Nguyen, Tu, et al.
Veröffentlicht: (2026)
The $\mathbf{Y}$-Combinator for LLMs: Solving Long-Context Rot with $λ$-Calculus
von: Roy, Amartya, et al.
Veröffentlicht: (2026)
von: Roy, Amartya, et al.
Veröffentlicht: (2026)
Bourbaki: Self-Generated and Goal-Conditioned MDPs for Theorem Proving
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2025)
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2025)
Risk-Controlled Lean-as-Judge for Natural-Language Mathematical Reasoning
von: Bourigault, Pauline, et al.
Veröffentlicht: (2026)
von: Bourigault, Pauline, et al.
Veröffentlicht: (2026)
LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs?
von: Ziomek, Juliusz, et al.
Veröffentlicht: (2026)
von: Ziomek, Juliusz, et al.
Veröffentlicht: (2026)
SparsePO: Controlling Preference Alignment of LLMs via Sparse Token Masks
von: Christopoulou, Fenia, et al.
Veröffentlicht: (2024)
von: Christopoulou, Fenia, et al.
Veröffentlicht: (2024)
Many of Your DPOs are Secretly One: Attempting Unification Through Mutual Information
von: Tutnov, Rasul, et al.
Veröffentlicht: (2025)
von: Tutnov, Rasul, et al.
Veröffentlicht: (2025)
Al-Khwarizmi: Discovering Physical Laws with Foundation Models
von: Mower, Christopher E., et al.
Veröffentlicht: (2025)
von: Mower, Christopher E., et al.
Veröffentlicht: (2025)
Safe Reinforcement Learning on the Constraint Manifold: Theory and Applications
von: Liu, Puze, et al.
Veröffentlicht: (2024)
von: Liu, Puze, et al.
Veröffentlicht: (2024)
Subjective Depth and Timescale Transformers: Learning Where and When to Compute
von: Wieser, Frederico, et al.
Veröffentlicht: (2025)
von: Wieser, Frederico, et al.
Veröffentlicht: (2025)
RSPO: Regularized Self-Play Alignment of Large Language Models
von: Tang, Xiaohang, et al.
Veröffentlicht: (2025)
von: Tang, Xiaohang, et al.
Veröffentlicht: (2025)
Untangling Component Imbalance in Hybrid Linear Attention Conversion Methods
von: Benfeghoul, Martin, et al.
Veröffentlicht: (2025)
von: Benfeghoul, Martin, et al.
Veröffentlicht: (2025)
SuRe: Surprise-Driven Prioritised Replay for Continual LLM Learning
von: Hazard, Hugo, et al.
Veröffentlicht: (2025)
von: Hazard, Hugo, et al.
Veröffentlicht: (2025)
Efficient Reinforcement Learning with Large Language Model Priors
von: Yan, Xue, et al.
Veröffentlicht: (2024)
von: Yan, Xue, et al.
Veröffentlicht: (2024)
Bayesian Reward Models for LLM Alignment
von: Yang, Adam X., et al.
Veröffentlicht: (2024)
von: Yang, Adam X., et al.
Veröffentlicht: (2024)
Overton Pluralistic Reinforcement Learning for Large Language Models
von: Fu, Yu, et al.
Veröffentlicht: (2026)
von: Fu, Yu, et al.
Veröffentlicht: (2026)
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives
von: Bou, Matthieu, et al.
Veröffentlicht: (2025)
von: Bou, Matthieu, et al.
Veröffentlicht: (2025)
Learning from Failures: Understanding LLM Alignment through Failure-Aware Inverse RL
von: Patel, Nyal, et al.
Veröffentlicht: (2025)
von: Patel, Nyal, et al.
Veröffentlicht: (2025)
Why Can Large Language Models Generate Correct Chain-of-Thoughts?
von: Tutunov, Rasul, et al.
Veröffentlicht: (2023)
von: Tutunov, Rasul, et al.
Veröffentlicht: (2023)
Human-inspired Episodic Memory for Infinite Context LLMs
von: Fountas, Zafeirios, et al.
Veröffentlicht: (2024)
von: Fountas, Zafeirios, et al.
Veröffentlicht: (2024)
wd1: Weighted Policy Optimization for Reasoning in Diffusion Language Models
von: Tang, Xiaohang, et al.
Veröffentlicht: (2025)
von: Tang, Xiaohang, et al.
Veröffentlicht: (2025)
Sample-efficient Bayesian Optimisation Using Known Invariances
von: Brown, Theodore, et al.
Veröffentlicht: (2024)
von: Brown, Theodore, et al.
Veröffentlicht: (2024)
Robust Bayesian Optimisation with Unbounded Corruptions
von: Ezzerg, Abdelhamid, et al.
Veröffentlicht: (2025)
von: Ezzerg, Abdelhamid, et al.
Veröffentlicht: (2025)
Bottlenecked Transformers: Periodic KV Cache Consolidation for Generalised Reasoning
von: Oomerjee, Adnan, et al.
Veröffentlicht: (2025)
von: Oomerjee, Adnan, et al.
Veröffentlicht: (2025)
Contextual Causal Bayesian Optimisation
von: Arsenyan, Vahan, et al.
Veröffentlicht: (2023)
von: Arsenyan, Vahan, et al.
Veröffentlicht: (2023)
PROSAC: Provably Safe Certification for Machine Learning Models under Adversarial Attacks
von: Feng, Chen, et al.
Veröffentlicht: (2024)
von: Feng, Chen, et al.
Veröffentlicht: (2024)
The Mosaic Memory of Large Language Models
von: Shilov, Igor, et al.
Veröffentlicht: (2024)
von: Shilov, Igor, et al.
Veröffentlicht: (2024)
Did the Neurons Read your Book? Document-level Membership Inference for Large Language Models
von: Meeus, Matthieu, et al.
Veröffentlicht: (2023)
von: Meeus, Matthieu, et al.
Veröffentlicht: (2023)
FlashDecoding++: Faster Large Language Model Inference on GPUs
von: Hong, Ke, et al.
Veröffentlicht: (2023)
von: Hong, Ke, et al.
Veröffentlicht: (2023)
LongAlign: A Recipe for Long Context Alignment of Large Language Models
von: Bai, Yushi, et al.
Veröffentlicht: (2024)
von: Bai, Yushi, et al.
Veröffentlicht: (2024)
Entropic-Time Inference: Self-Organizing Large Language Model Decoding Beyond Attention
von: Kiruluta, Andrew
Veröffentlicht: (2026)
von: Kiruluta, Andrew
Veröffentlicht: (2026)
Tree-OPO: Off-policy Monte Carlo Tree-Guided Advantage Optimization for Multistep Reasoning
von: Huang, Bingning, et al.
Veröffentlicht: (2025)
von: Huang, Bingning, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2026) -
Group Robust Preference Optimization in Reward-free RLHF
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2024) -
Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2025) -
Decoding as Optimisation on the Probability Simplex: From Top-K to Top-P (Nucleus) to Best-of-K Samplers
von: Ji, Xiaotong, et al.
Veröffentlicht: (2026) -
Scalable Power Sampling: Unlocking Efficient, Training-Free Reasoning for LLMs via Distribution Sharpening
von: Ji, Xiaotong, et al.
Veröffentlicht: (2026)