Hair-Trigger Alignment: Black-Box Evaluation Cannot Guarantee Post-Update Alignment
Fuente:
arXiv
Guardado en:
| Autores principales: | Bakman, Yavuz, Yaldiz, Duygu Nur, Avestimehr, Salman, Karimireddy, Sai Praneeth |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Reconsidering LLM Uncertainty Estimation Methods in the Wild
por: Bakman, Yavuz, et al.
Publicado: (2025)
por: Bakman, Yavuz, et al.
Publicado: (2025)
Reject Only Critical Tokens: Pivot-Aware Speculative Decoding
por: Ziashahabi, Amir, et al.
Publicado: (2025)
por: Ziashahabi, Amir, et al.
Publicado: (2025)
Conformal Prediction Adaptive to Unknown Subpopulation Shifts
por: Wang, Nien-Shao, et al.
Publicado: (2025)
por: Wang, Nien-Shao, et al.
Publicado: (2025)
MARS: Meaning-Aware Response Scoring for Uncertainty Estimation in Generative LLMs
por: Bakman, Yavuz Faruk, et al.
Publicado: (2024)
por: Bakman, Yavuz Faruk, et al.
Publicado: (2024)
Uncertainty as Feature Gaps: Epistemic Uncertainty Quantification of LLMs in Contextual Question-Answering
por: Bakman, Yavuz, et al.
Publicado: (2025)
por: Bakman, Yavuz, et al.
Publicado: (2025)
Uncertainty Quantification for Hallucination Detection in Large Language Models: Foundations, Methodology, and Future Directions
por: Kang, Sungmin, et al.
Publicado: (2025)
por: Kang, Sungmin, et al.
Publicado: (2025)
Un-considering Contextual Information: Assessing LLMs' Understanding of Indexical Elements
por: Oguz, Metehan, et al.
Publicado: (2025)
por: Oguz, Metehan, et al.
Publicado: (2025)
Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs
por: Yun, Vincent-Daniel, et al.
Publicado: (2026)
por: Yun, Vincent-Daniel, et al.
Publicado: (2026)
CroMo-Mixup: Augmenting Cross-Model Representations for Continual Self-Supervised Learning
por: Mushtaq, Erum, et al.
Publicado: (2024)
por: Mushtaq, Erum, et al.
Publicado: (2024)
Optimization with Access to Auxiliary Information
por: Chayti, El Mahdi, et al.
Publicado: (2022)
por: Chayti, El Mahdi, et al.
Publicado: (2022)
TruthTorchLM: A Comprehensive Library for Predicting Truthfulness in LLM Outputs
por: Yaldiz, Duygu Nur, et al.
Publicado: (2025)
por: Yaldiz, Duygu Nur, et al.
Publicado: (2025)
On the Limits of Momentum in Decentralized and Federated Optimization
por: Zaccone, Riccardo, et al.
Publicado: (2025)
por: Zaccone, Riccardo, et al.
Publicado: (2025)
Do Not Design, Learn: A Trainable Scoring Function for Uncertainty Estimation in Generative LLMs
por: Yaldiz, Duygu Nur, et al.
Publicado: (2024)
por: Yaldiz, Duygu Nur, et al.
Publicado: (2024)
LIA: Privacy-Preserving Data Quality Evaluation in Federated Learning Using a Lazy Influence Approximation
por: Rokvic, Ljubomir, et al.
Publicado: (2022)
por: Rokvic, Ljubomir, et al.
Publicado: (2022)
Collaborative Heterogeneous Causal Inference Beyond Meta-analysis
por: Guo, Tianyu, et al.
Publicado: (2024)
por: Guo, Tianyu, et al.
Publicado: (2024)
DAVED: Data Acquisition via Experimental Design for Data Markets
por: Lu, Charles, et al.
Publicado: (2024)
por: Lu, Charles, et al.
Publicado: (2024)
Defection-Free Collaboration between Competitors in a Learning System
por: Werner, Mariel, et al.
Publicado: (2024)
por: Werner, Mariel, et al.
Publicado: (2024)
Do Data Valuations Make Good Data Prices?
por: Fan, Dongyang, et al.
Publicado: (2025)
por: Fan, Dongyang, et al.
Publicado: (2025)
VoxGuard: Evaluating User and Attribute Privacy in Speech via Membership Inference Attacks
por: Tsaprazlis, Efthymios, et al.
Publicado: (2025)
por: Tsaprazlis, Efthymios, et al.
Publicado: (2025)
A Differentially Private Kaplan-Meier Estimator for Privacy-Preserving Survival Analysis
por: Veeraragavan, Narasimha Raghavan, et al.
Publicado: (2024)
por: Veeraragavan, Narasimha Raghavan, et al.
Publicado: (2024)
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
por: Fan, Dongyang, et al.
Publicado: (2025)
por: Fan, Dongyang, et al.
Publicado: (2025)
Entropy-driven Fair and Effective Federated Learning
por: Wang, Lin, et al.
Publicado: (2023)
por: Wang, Lin, et al.
Publicado: (2023)
Communication-Efficient Heterogeneous Federated Learning with Generalized Heavy-Ball Momentum
por: Zaccone, Riccardo, et al.
Publicado: (2023)
por: Zaccone, Riccardo, et al.
Publicado: (2023)
Balancing Classification and Calibration Performance in Decision-Making LLMs via Calibration Aware Reinforcement Learning
por: Yaldiz, Duygu Nur, et al.
Publicado: (2026)
por: Yaldiz, Duygu Nur, et al.
Publicado: (2026)
f-INE: A Hypothesis Testing Framework for Estimating Influence under Training Randomness
por: Panda, Subhodip, et al.
Publicado: (2025)
por: Panda, Subhodip, et al.
Publicado: (2025)
Black-Box Behavioral Distillation Breaks Safety Alignment in Medical LLMs
por: Jahan, Sohely, et al.
Publicado: (2025)
por: Jahan, Sohely, et al.
Publicado: (2025)
Robust Multi-Agent LLMs under Byzantine Faults
por: Lee, Haejoon, et al.
Publicado: (2026)
por: Lee, Haejoon, et al.
Publicado: (2026)
Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization
por: Nguyen, Tuan, et al.
Publicado: (2025)
por: Nguyen, Tuan, et al.
Publicado: (2025)
Deployment-Relevant Alignment Cannot Be Inferred from Model-Level Evaluation Alone
por: Vishwarupe, Varad, et al.
Publicado: (2026)
por: Vishwarupe, Varad, et al.
Publicado: (2026)
Differentially Private Federated Learning without Noise Addition: When is it Possible?
por: Zhang, Jiang, et al.
Publicado: (2024)
por: Zhang, Jiang, et al.
Publicado: (2024)
ATP: Enabling Fast LLM Serving via Attention on Top Principal Keys
por: Niu, Yue, et al.
Publicado: (2024)
por: Niu, Yue, et al.
Publicado: (2024)
A Closer Look at Personalized Fine-Tuning in Heterogeneous Federated Learning
por: Chen, Minghui, et al.
Publicado: (2025)
por: Chen, Minghui, et al.
Publicado: (2025)
GEM: A Scale-Aware and Distribution-Sensitive Sparse Fine-Tuning Framework for Effective Downstream Adaptation
por: Kang, Sungmin, et al.
Publicado: (2025)
por: Kang, Sungmin, et al.
Publicado: (2025)
Towards Distillation Guarantees under Algorithmic Alignment for Combinatorial Optimization
por: Le, Thien, et al.
Publicado: (2026)
por: Le, Thien, et al.
Publicado: (2026)
Trust-Region Noise Search for Black-Box Alignment of Diffusion and Flow Models
por: Schweiger, Niklas, et al.
Publicado: (2026)
por: Schweiger, Niklas, et al.
Publicado: (2026)
Hawk: Accurate and Fast Privacy-Preserving Machine Learning Using Secure Lookup Table Computation
por: Saleem, Hamza, et al.
Publicado: (2024)
por: Saleem, Hamza, et al.
Publicado: (2024)
Conformal Alignment: Knowing When to Trust Foundation Models with Guarantees
por: Gui, Yu, et al.
Publicado: (2024)
por: Gui, Yu, et al.
Publicado: (2024)
Information Theoretic Guarantees For Policy Alignment In Large Language Models
por: Mroueh, Youssef
Publicado: (2024)
por: Mroueh, Youssef
Publicado: (2024)
Asymptotics of Language Model Alignment
por: Yang, Joy Qiping, et al.
Publicado: (2024)
por: Yang, Joy Qiping, et al.
Publicado: (2024)
No Certificate for Alignment: Two Independent Impossibilities and the Pareto Frontier of Achievable Safety Guarantees
por: Agarwal, Ayushi
Publicado: (2026)
por: Agarwal, Ayushi
Publicado: (2026)
Ejemplares similares
-
Reconsidering LLM Uncertainty Estimation Methods in the Wild
por: Bakman, Yavuz, et al.
Publicado: (2025) -
Reject Only Critical Tokens: Pivot-Aware Speculative Decoding
por: Ziashahabi, Amir, et al.
Publicado: (2025) -
Conformal Prediction Adaptive to Unknown Subpopulation Shifts
por: Wang, Nien-Shao, et al.
Publicado: (2025) -
MARS: Meaning-Aware Response Scoring for Uncertainty Estimation in Generative LLMs
por: Bakman, Yavuz Faruk, et al.
Publicado: (2024) -
Uncertainty as Feature Gaps: Epistemic Uncertainty Quantification of LLMs in Contextual Question-Answering
por: Bakman, Yavuz, et al.
Publicado: (2025)