Alignment Quality Index (AQI) : Beyond Refusals: AQI as an Intrinsic Alignment Diagnostic via Latent Geometry, Cluster Divergence, and Layer wise Pooled Representations

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Borah, Abhilekh, Sharma, Chhavi, Khanna, Danush, Bhatt, Utkarsh, Singh, Gurpreet, Abdullah, Hasnat Md, Ravi, Raghav Kaushik, Jain, Vinija, Patel, Jyoti, Singh, Shubham, Sharma, Vasu, Vats, Arpita, Raja, Rahul, Chadha, Aman, Das, Amitava
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913896460189696
author Borah, Abhilekh
Sharma, Chhavi
Khanna, Danush
Bhatt, Utkarsh
Singh, Gurpreet
Abdullah, Hasnat Md
Ravi, Raghav Kaushik
Jain, Vinija
Patel, Jyoti
Singh, Shubham
Sharma, Vasu
Vats, Arpita
Raja, Rahul
Chadha, Aman
Das, Amitava
author_facet Borah, Abhilekh
Sharma, Chhavi
Khanna, Danush
Bhatt, Utkarsh
Singh, Gurpreet
Abdullah, Hasnat Md
Ravi, Raghav Kaushik
Jain, Vinija
Patel, Jyoti
Singh, Shubham
Sharma, Vasu
Vats, Arpita
Raja, Rahul
Chadha, Aman
Das, Amitava
contents Alignment is no longer a luxury, it is a necessity. As large language models (LLMs) enter high-stakes domains like education, healthcare, governance, and law, their behavior must reliably reflect human-aligned values and safety constraints. Yet current evaluations rely heavily on behavioral proxies such as refusal rates, G-Eval scores, and toxicity classifiers, all of which have critical blind spots. Aligned models are often vulnerable to jailbreaking, stochasticity of generation, and alignment faking. To address this issue, we introduce the Alignment Quality Index (AQI). This novel geometric and prompt-invariant metric empirically assesses LLM alignment by analyzing the separation of safe and unsafe activations in latent space. By combining measures such as the Davies-Bouldin Score (DBS), Dunn Index (DI), Xie-Beni Index (XBI), and Calinski-Harabasz Index (CHI) across various formulations, AQI captures clustering quality to detect hidden misalignments and jailbreak risks, even when outputs appear compliant. AQI also serves as an early warning signal for alignment faking, offering a robust, decoding invariant tool for behavior agnostic safety auditing. Additionally, we propose the LITMUS dataset to facilitate robust evaluation under these challenging conditions. Empirical tests on LITMUS across different models trained under DPO, GRPO, and RLHF conditions demonstrate AQI's correlation with external judges and ability to reveal vulnerabilities missed by refusal metrics. We make our implementation publicly available to foster future research in this area.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13901
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Alignment Quality Index (AQI) : Beyond Refusals: AQI as an Intrinsic Alignment Diagnostic via Latent Geometry, Cluster Divergence, and Layer wise Pooled Representations
Borah, Abhilekh
Sharma, Chhavi
Khanna, Danush
Bhatt, Utkarsh
Singh, Gurpreet
Abdullah, Hasnat Md
Ravi, Raghav Kaushik
Jain, Vinija
Patel, Jyoti
Singh, Shubham
Sharma, Vasu
Vats, Arpita
Raja, Rahul
Chadha, Aman
Das, Amitava
Computation and Language
Artificial Intelligence
Alignment is no longer a luxury, it is a necessity. As large language models (LLMs) enter high-stakes domains like education, healthcare, governance, and law, their behavior must reliably reflect human-aligned values and safety constraints. Yet current evaluations rely heavily on behavioral proxies such as refusal rates, G-Eval scores, and toxicity classifiers, all of which have critical blind spots. Aligned models are often vulnerable to jailbreaking, stochasticity of generation, and alignment faking. To address this issue, we introduce the Alignment Quality Index (AQI). This novel geometric and prompt-invariant metric empirically assesses LLM alignment by analyzing the separation of safe and unsafe activations in latent space. By combining measures such as the Davies-Bouldin Score (DBS), Dunn Index (DI), Xie-Beni Index (XBI), and Calinski-Harabasz Index (CHI) across various formulations, AQI captures clustering quality to detect hidden misalignments and jailbreak risks, even when outputs appear compliant. AQI also serves as an early warning signal for alignment faking, offering a robust, decoding invariant tool for behavior agnostic safety auditing. Additionally, we propose the LITMUS dataset to facilitate robust evaluation under these challenging conditions. Empirical tests on LITMUS across different models trained under DPO, GRPO, and RLHF conditions demonstrate AQI's correlation with external judges and ability to reveal vulnerabilities missed by refusal metrics. We make our implementation publicly available to foster future research in this area.
title Alignment Quality Index (AQI) : Beyond Refusals: AQI as an Intrinsic Alignment Diagnostic via Latent Geometry, Cluster Divergence, and Layer wise Pooled Representations
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.13901