Selective Safety Steering via Value-Filtered Decoding
Fuente:
arXiv
Salvato in:
| Autori principali: | Einbinder, Bat-Sheva, Davidov, Hen, Teh, Yee Whye, Gal, Yarin, Romano, Yaniv |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Semi-Supervised Risk Control via Prediction-Powered Inference
di: Einbinder, Bat-Sheva, et al.
Pubblicazione: (2024)
di: Einbinder, Bat-Sheva, et al.
Pubblicazione: (2024)
Label Noise Robustness of Conformal Prediction
di: Einbinder, Bat-Sheva, et al.
Pubblicazione: (2022)
di: Einbinder, Bat-Sheva, et al.
Pubblicazione: (2022)
Calibrated Predictive Lower Bounds on Time-to-Unsafe-Sampling in LLMs
di: Davidov, Hen, et al.
Pubblicazione: (2025)
di: Davidov, Hen, et al.
Pubblicazione: (2025)
Protected Test-Time Adaptation via Online Entropy Matching: A Betting Approach
di: Bar, Yarin, et al.
Pubblicazione: (2024)
di: Bar, Yarin, et al.
Pubblicazione: (2024)
NoProp: Training Neural Networks without Full Back-propagation or Full Forward-propagation
di: Li, Qinyu, et al.
Pubblicazione: (2025)
di: Li, Qinyu, et al.
Pubblicazione: (2025)
SymDiff: Equivariant Diffusion via Stochastic Symmetrisation
di: Zhang, Leo, et al.
Pubblicazione: (2024)
di: Zhang, Leo, et al.
Pubblicazione: (2024)
Testing For Distribution Shifts with Conditional Conformal Test Martingales
di: Shaer, Shalev, et al.
Pubblicazione: (2026)
di: Shaer, Shalev, et al.
Pubblicazione: (2026)
Kalman Filter for Online Classification of Non-Stationary Data
di: Titsias, Michalis K., et al.
Pubblicazione: (2023)
di: Titsias, Michalis K., et al.
Pubblicazione: (2023)
The Edge-of-Reach Problem in Offline Model-Based Reinforcement Learning
di: Sims, Anya, et al.
Pubblicazione: (2024)
di: Sims, Anya, et al.
Pubblicazione: (2024)
Incorporating Unlabelled Data into Bayesian Neural Networks
di: Sharma, Mrinank, et al.
Pubblicazione: (2023)
di: Sharma, Mrinank, et al.
Pubblicazione: (2023)
Verifier-Backed Hard Problem Generation for Mathematical Reasoning
di: Lai, Yuhang, et al.
Pubblicazione: (2026)
di: Lai, Yuhang, et al.
Pubblicazione: (2026)
L3Ms -- Lagrange Large Language Models
di: Dhillon, Guneet S., et al.
Pubblicazione: (2024)
di: Dhillon, Guneet S., et al.
Pubblicazione: (2024)
Rao-Blackwellised Reparameterisation Gradients
di: Lam, Kevin H., et al.
Pubblicazione: (2025)
di: Lam, Kevin H., et al.
Pubblicazione: (2025)
Manifold Aware Denoising Score Matching (MAD)
di: Levy-Jurgenson, Alona, et al.
Pubblicazione: (2026)
di: Levy-Jurgenson, Alona, et al.
Pubblicazione: (2026)
SofT-GRPO: Surpassing Discrete-Token LLM Reinforcement Learning via Gumbel-Reparameterized Soft-Thinking Policy Optimization
di: Zheng, Zhi, et al.
Pubblicazione: (2025)
di: Zheng, Zhi, et al.
Pubblicazione: (2025)
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization
di: Tolochinsky, Elad, et al.
Pubblicazione: (2026)
di: Tolochinsky, Elad, et al.
Pubblicazione: (2026)
Meta Flow Maps enable scalable reward alignment
di: Potaptchik, Peter, et al.
Pubblicazione: (2026)
di: Potaptchik, Peter, et al.
Pubblicazione: (2026)
Extending Epistemic Uncertainty Beyond Parameters Would Assist in Designing Reliable LLMs
di: Nguyen-Hien, T. Duy, et al.
Pubblicazione: (2025)
di: Nguyen-Hien, T. Duy, et al.
Pubblicazione: (2025)
Meta-Learning Objectives for Preference Optimization
di: Alfano, Carlo, et al.
Pubblicazione: (2024)
di: Alfano, Carlo, et al.
Pubblicazione: (2024)
Semi-Supervised Hypothesis Testing by Betting on Predictions
di: Tenzer, Yaniv, et al.
Pubblicazione: (2026)
di: Tenzer, Yaniv, et al.
Pubblicazione: (2026)
EvIL: Evolution Strategies for Generalisable Imitation Learning
di: Sapora, Silvia, et al.
Pubblicazione: (2024)
di: Sapora, Silvia, et al.
Pubblicazione: (2024)
SigmaDock: Untwisting Molecular Docking With Fragment-Based SE(3) Diffusion
di: Prat, Alvaro, et al.
Pubblicazione: (2025)
di: Prat, Alvaro, et al.
Pubblicazione: (2025)
Metropolis-Adjusted Diffusion Models
di: Lam, Kevin H., et al.
Pubblicazione: (2026)
di: Lam, Kevin H., et al.
Pubblicazione: (2026)
Amortized Probabilistic Detection of Communities in Graphs
di: Wang, Yueqi, et al.
Pubblicazione: (2020)
di: Wang, Yueqi, et al.
Pubblicazione: (2020)
How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation
di: Feldman, Shai, et al.
Pubblicazione: (2026)
di: Feldman, Shai, et al.
Pubblicazione: (2026)
Robust Conformal Prediction Using Privileged Information
di: Feldman, Shai, et al.
Pubblicazione: (2024)
di: Feldman, Shai, et al.
Pubblicazione: (2024)
Online Adaptation of Language Models with a Memory of Amortized Contexts
di: Tack, Jihoon, et al.
Pubblicazione: (2024)
di: Tack, Jihoon, et al.
Pubblicazione: (2024)
INNOCENT III (1198-1216) ET L’ANCIEN TESTAMENT: POLITIQUE ET EXÉGÈSE DANS LA DELIBERATIO DOMINI PAPAE INNOCENTII SUPER FACTO IMPERII DE TRIBUS ELECTIS
di: Albert Bat Sheva
Pubblicazione: (2016)
di: Albert Bat Sheva
Pubblicazione: (2016)
Prompting Strategies for Enabling Large Language Models to Infer Causation from Correlation
di: Sgouritsa, Eleni, et al.
Pubblicazione: (2024)
di: Sgouritsa, Eleni, et al.
Pubblicazione: (2024)
Context-Guided Diffusion for Out-of-Distribution Molecular and Protein Design
di: Klarner, Leo, et al.
Pubblicazione: (2024)
di: Klarner, Leo, et al.
Pubblicazione: (2024)
Simple Baselines are Competitive with Code Evolution
di: Gideoni, Yonatan, et al.
Pubblicazione: (2026)
di: Gideoni, Yonatan, et al.
Pubblicazione: (2026)
The Benefits and Risks of Transductive Approaches for AI Fairness
di: Razzak, Muhammed, et al.
Pubblicazione: (2024)
di: Razzak, Muhammed, et al.
Pubblicazione: (2024)
Pivotal Auto-Encoder via Self-Normalizing ReLU
di: Goldenstein, Nelson, et al.
Pubblicazione: (2024)
di: Goldenstein, Nelson, et al.
Pubblicazione: (2024)
Unleashing the Power of Meta-tuning for Few-shot Generalization Through Sparse Interpolated Experts
di: Chen, Shengzhuang, et al.
Pubblicazione: (2024)
di: Chen, Shengzhuang, et al.
Pubblicazione: (2024)
Do Multilingual LLMs Think In English?
di: Schut, Lisa, et al.
Pubblicazione: (2025)
di: Schut, Lisa, et al.
Pubblicazione: (2025)
In-Context Learning Learns Label Relationships but Is Not Conventional Learning
di: Kossen, Jannik, et al.
Pubblicazione: (2023)
di: Kossen, Jannik, et al.
Pubblicazione: (2023)
Non-Stationary Learning of Neural Networks with Automatic Soft Parameter Reset
di: Galashov, Alexandre, et al.
Pubblicazione: (2024)
di: Galashov, Alexandre, et al.
Pubblicazione: (2024)
Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning
di: Melo, Luckeciano C., et al.
Pubblicazione: (2025)
di: Melo, Luckeciano C., et al.
Pubblicazione: (2025)
Temporal-Difference Variational Continual Learning
di: Melo, Luckeciano C., et al.
Pubblicazione: (2024)
di: Melo, Luckeciano C., et al.
Pubblicazione: (2024)
Variational Flow Maps: Make Some Noise for One-Step Conditional Generation
di: Mammadov, Abbas, et al.
Pubblicazione: (2026)
di: Mammadov, Abbas, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Semi-Supervised Risk Control via Prediction-Powered Inference
di: Einbinder, Bat-Sheva, et al.
Pubblicazione: (2024) -
Label Noise Robustness of Conformal Prediction
di: Einbinder, Bat-Sheva, et al.
Pubblicazione: (2022) -
Calibrated Predictive Lower Bounds on Time-to-Unsafe-Sampling in LLMs
di: Davidov, Hen, et al.
Pubblicazione: (2025) -
Protected Test-Time Adaptation via Online Entropy Matching: A Betting Approach
di: Bar, Yarin, et al.
Pubblicazione: (2024) -
NoProp: Training Neural Networks without Full Back-propagation or Full Forward-propagation
di: Li, Qinyu, et al.
Pubblicazione: (2025)