Rethinking the Evaluation of Alignment Methods: Insights into Diversity, Generalisation, and Safety
Fuente:
arXiv
Saved in:
| Main Authors: | Janiak, Denis, Moska, Julia, Motyka, Dawid, Seweryn, Karolina, Walkowiak, Paweł, Żuk, Bartosz, Janz, Arkadiusz |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unpacking Robustness in Inflectional Languages: Adversarial Evaluation and Mechanistic Insights
by: Walkowiak, Paweł, et al.
Published: (2025)
by: Walkowiak, Paweł, et al.
Published: (2025)
Hallucination Detection in LLMs Using Spectral Features of Attention Maps
by: Binkowski, Jakub, et al.
Published: (2025)
by: Binkowski, Jakub, et al.
Published: (2025)
FactSelfCheck: Fact-Level Black-Box Hallucination Detection for LLMs
by: Sawczyn, Albert, et al.
Published: (2025)
by: Sawczyn, Albert, et al.
Published: (2025)
Understanding the Effects of RLHF on LLM Generalisation and Diversity
by: Kirk, Robert, et al.
Published: (2023)
by: Kirk, Robert, et al.
Published: (2023)
Do Generalisation Results Generalise?
by: Boglioni, Matteo, et al.
Published: (2025)
by: Boglioni, Matteo, et al.
Published: (2025)
Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents
by: Tang, Zhengyang, et al.
Published: (2026)
by: Tang, Zhengyang, et al.
Published: (2026)
A Symbolic Framework for Evaluating Mathematical Reasoning and Generalisation with Transformers
by: Meadows, Jordan, et al.
Published: (2023)
by: Meadows, Jordan, et al.
Published: (2023)
The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs
by: Janiak, Denis, et al.
Published: (2025)
by: Janiak, Denis, et al.
Published: (2025)
Isotonic Quantile Regression Averaging for uncertainty quantification of electricity price forecasts
by: Lipiecki, Arkadiusz, et al.
Published: (2025)
by: Lipiecki, Arkadiusz, et al.
Published: (2025)
Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs
by: Chen, Yen-Shan, et al.
Published: (2026)
by: Chen, Yen-Shan, et al.
Published: (2026)
StylOch at PAN: Gradient-Boosted Trees with Frequency-Based Stylometric Features
by: Ochab, Jeremi K., et al.
Published: (2025)
by: Ochab, Jeremi K., et al.
Published: (2025)
SafeArena: Evaluating the Safety of Autonomous Web Agents
by: Tur, Ada Defne, et al.
Published: (2025)
by: Tur, Ada Defne, et al.
Published: (2025)
Sample, Don't Search: Rethinking Test-Time Alignment for Language Models
by: Faria, Gonçalo, et al.
Published: (2025)
by: Faria, Gonçalo, et al.
Published: (2025)
Integrating gender inclusivity into large language models via instruction tuning
by: Wróblewska, Alina, et al.
Published: (2025)
by: Wróblewska, Alina, et al.
Published: (2025)
Rethinking Regularization Methods for Knowledge Graph Completion
by: Li, Linyu, et al.
Published: (2025)
by: Li, Linyu, et al.
Published: (2025)
Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection
by: Sun, Guanglong, et al.
Published: (2026)
by: Sun, Guanglong, et al.
Published: (2026)
Evaluating LLMs Robustness in Less Resourced Languages with Proxy Models
by: Chrabąszcz, Maciej, et al.
Published: (2025)
by: Chrabąszcz, Maciej, et al.
Published: (2025)
Advancing LLM Safe Alignment with Safety Representation Ranking
by: Du, Tianqi, et al.
Published: (2025)
by: Du, Tianqi, et al.
Published: (2025)
Test-Time Safety Alignment
by: Saglam, Baturay, et al.
Published: (2026)
by: Saglam, Baturay, et al.
Published: (2026)
Rethinking the Role of Proxy Rewards in Language Model Alignment
by: Kim, Sungdong, et al.
Published: (2024)
by: Kim, Sungdong, et al.
Published: (2024)
Few Tokens, Big Leverage: Preserving Safety Alignment by Constraining Safety Tokens during Fine-tuning
by: Wang, Guoli, et al.
Published: (2026)
by: Wang, Guoli, et al.
Published: (2026)
Towards Generalising Neural Topical Representations
by: Yang, Xiaohao, et al.
Published: (2023)
by: Yang, Xiaohao, et al.
Published: (2023)
Survey of Action Recognition, Spotting and Spatio-Temporal Localization in Soccer -- Current Trends and Research Perspectives
by: Seweryn, Karolina, et al.
Published: (2023)
by: Seweryn, Karolina, et al.
Published: (2023)
MULTIVERSE: Exposing Large Language Model Alignment Problems in Diverse Worlds
by: Jin, Xiaolong, et al.
Published: (2024)
by: Jin, Xiaolong, et al.
Published: (2024)
Rethinking of Encoder-based Warm-start Methods in Hyperparameter Optimization
by: Płudowski, Dawid, et al.
Published: (2024)
by: Płudowski, Dawid, et al.
Published: (2024)
SABER: Uncovering Vulnerabilities in Safety Alignment via Cross-Layer Residual Connection
by: Joshi, Maithili, et al.
Published: (2025)
by: Joshi, Maithili, et al.
Published: (2025)
Multilingual Safety Alignment via Self-Distillation
by: Qin, Ruiyang, et al.
Published: (2026)
by: Qin, Ruiyang, et al.
Published: (2026)
PL-Guard: Benchmarking Language Model Safety for Polish
by: Krasnodębska, Aleksandra, et al.
Published: (2025)
by: Krasnodębska, Aleksandra, et al.
Published: (2025)
Rethinking LLM Evaluation: Can We Evaluate LLMs with 200x Less Data?
by: Wang, Shaobo, et al.
Published: (2025)
by: Wang, Shaobo, et al.
Published: (2025)
Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons
by: Chen, Jianhui, et al.
Published: (2024)
by: Chen, Jianhui, et al.
Published: (2024)
The Role of Language Imbalance in Cross-lingual Generalisation: Insights from Cloned Language Experiments
by: Schäfer, Anton, et al.
Published: (2024)
by: Schäfer, Anton, et al.
Published: (2024)
Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection
by: Deshpande, Vijeta, et al.
Published: (2026)
by: Deshpande, Vijeta, et al.
Published: (2026)
TRA: Better Length Generalisation with Threshold Relative Attention
by: Opper, Mattia, et al.
Published: (2025)
by: Opper, Mattia, et al.
Published: (2025)
MobileSafetyBench: Evaluating Safety of Autonomous Agents in Mobile Device Control
by: Lee, Juyong, et al.
Published: (2024)
by: Lee, Juyong, et al.
Published: (2024)
Rethinking Perplexity: Revealing the Impact of Input Length on Perplexity Evaluation in LLMs
by: Cheng, Letian, et al.
Published: (2026)
by: Cheng, Letian, et al.
Published: (2026)
Safety Alignment Can Be Not Superficial With Explicit Safety Signals
by: Li, Jianwei, et al.
Published: (2025)
by: Li, Jianwei, et al.
Published: (2025)
ProMIL: Probabilistic Multiple Instance Learning for Medical Imaging
by: Struski, Łukasz, et al.
Published: (2023)
by: Struski, Łukasz, et al.
Published: (2023)
Improving LLM Safety Alignment with Dual-Objective Optimization
by: Zhao, Xuandong, et al.
Published: (2025)
by: Zhao, Xuandong, et al.
Published: (2025)
Course-Correction: Safety Alignment Using Synthetic Preferences
by: Xu, Rongwu, et al.
Published: (2024)
by: Xu, Rongwu, et al.
Published: (2024)
Rethinking Layer Relevance in Large Language Models Beyond Cosine Similarity
by: Hinostroza, Cristian, et al.
Published: (2026)
by: Hinostroza, Cristian, et al.
Published: (2026)
Similar Items
-
Unpacking Robustness in Inflectional Languages: Adversarial Evaluation and Mechanistic Insights
by: Walkowiak, Paweł, et al.
Published: (2025) -
Hallucination Detection in LLMs Using Spectral Features of Attention Maps
by: Binkowski, Jakub, et al.
Published: (2025) -
FactSelfCheck: Fact-Level Black-Box Hallucination Detection for LLMs
by: Sawczyn, Albert, et al.
Published: (2025) -
Understanding the Effects of RLHF on LLM Generalisation and Diversity
by: Kirk, Robert, et al.
Published: (2023) -
Do Generalisation Results Generalise?
by: Boglioni, Matteo, et al.
Published: (2025)