Alignment Quality Index (AQI) : Beyond Refusals: AQI as an Intrinsic Alignment Diagnostic via Latent Geometry, Cluster Divergence, and Layer wise Pooled Representations
Fuente:
arXiv
Saved in:
| Main Authors: | Borah, Abhilekh, Sharma, Chhavi, Khanna, Danush, Bhatt, Utkarsh, Singh, Gurpreet, Abdullah, Hasnat Md, Ravi, Raghav Kaushik, Jain, Vinija, Patel, Jyoti, Singh, Shubham, Sharma, Vasu, Vats, Arpita, Raja, Rahul, Chadha, Aman, Das, Amitava |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AlignGuard-LoRA: Alignment-Preserving Fine-Tuning via Fisher-Guided Decomposition and Riemannian-Geodesic Collision Regularization
by: Das, Amitava, et al.
Published: (2025)
by: Das, Amitava, et al.
Published: (2025)
YINYANG-ALIGN: Benchmarking Contradictory Objectives and Proposing Multi-Objective Optimization based DPO for Text-to-Image Alignment
by: Das, Amitava, et al.
Published: (2025)
by: Das, Amitava, et al.
Published: (2025)
DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization
by: Das, Amitava, et al.
Published: (2025)
by: Das, Amitava, et al.
Published: (2025)
Air in Your Neighborhood: Fine-Grained AQI Forecasting Using Mobile Sensor Data
by: Sharma, Aaryam
Published: (2025)
by: Sharma, Aaryam
Published: (2025)
AdversariaL attacK sAfety aLIgnment(ALKALI): Safeguarding LLMs through GRACE: Geometric Representation-Aware Contrastive Enhancement- Introducing Adversarial Vulnerability Quality Index (AVQI)
by: Khanna, Danush, et al.
Published: (2025)
by: Khanna, Danush, et al.
Published: (2025)
CliME: Evaluating Multimodal Climate Discourse on Social Media and the Climate Alignment Quotient (CAQ)
by: Borah, Abhilekh, et al.
Published: (2025)
by: Borah, Abhilekh, et al.
Published: (2025)
TRACEALIGN -- Tracing the Drift: Attributing Alignment Failures to Training-Time Belief Sources in LLMs
by: Das, Amitava, et al.
Published: (2025)
by: Das, Amitava, et al.
Published: (2025)
DETONATE: A Benchmark for Text-to-Image Alignment and Kernelized Direct Preference Optimization
by: Prasad, Renjith, et al.
Published: (2025)
by: Prasad, Renjith, et al.
Published: (2025)
AlignMerge - Alignment-Preserving Large Language Model Merging via Fisher-Guided Geometric Constraints
by: Roy, Aniruddha, et al.
Published: (2025)
by: Roy, Aniruddha, et al.
Published: (2025)
ECLIPTICA -- A Framework for Switchable LLM Alignment via CITA - Contrastive Instruction-Tuned Alignment
by: Wanaskar, Kapil, et al.
Published: (2026)
by: Wanaskar, Kapil, et al.
Published: (2026)
Exploring the Impact of Large Language Models on Recommender Systems: An Extensive Review
by: Vats, Arpita, et al.
Published: (2024)
by: Vats, Arpita, et al.
Published: (2024)
Predictive Modelling of Air Quality Index (AQI) Across Diverse Cities and States of India using Machine Learning: Investigating the Influence of Punjab's Stubble Burning on AQI Variability
by: Sidhu, Kamaljeet Kaur, et al.
Published: (2024)
by: Sidhu, Kamaljeet Kaur, et al.
Published: (2024)
Peccavi: Visual Paraphrase Attack Safe and Distortion Free Image Watermarking Technique for AI-Generated Images
by: Dixit, Shreyas, et al.
Published: (2025)
by: Dixit, Shreyas, et al.
Published: (2025)
Predicting Solar Energy Generation with Machine Learning based on AQI and Weather Features
by: Shah, Arjun, et al.
Published: (2024)
by: Shah, Arjun, et al.
Published: (2024)
Multilingual State Space Models for Structured Question Answering in Indic Languages
by: Vats, Arpita, et al.
Published: (2025)
by: Vats, Arpita, et al.
Published: (2025)
D-STEER - Preference Alignment Techniques Learn to Behave, not to Believe -- Beneath the Surface, DPO as Steering Vector Perturbation in Activation Space
by: Raina, Samarth, et al.
Published: (2025)
by: Raina, Samarth, et al.
Published: (2025)
SPINAL -- Scaling-law and Preference Integration in Neural Alignment Layers
by: Das, Arion, et al.
Published: (2026)
by: Das, Arion, et al.
Published: (2026)
Alignment Faking - the Train -> Deploy Asymmetry: Through a Game-Theoretic Lens with Bayesian-Stackelberg Equilibria
by: Garg, Kartik, et al.
Published: (2025)
by: Garg, Kartik, et al.
Published: (2025)
Predicting Lung Disease Severity via Image-Based AQI Analysis using Deep Learning Techniques
by: Mahajan, Anvita, et al.
Published: (2024)
by: Mahajan, Anvita, et al.
Published: (2024)
M$^2$FedAQI: Multimodal Federated Learning for Air Quality Prediction on Heterogeneous Edge Devices
by: Nepal, Manjil, et al.
Published: (2026)
by: Nepal, Manjil, et al.
Published: (2026)
Forecasting of Multiple Seasonal Categorical Time Series Using Fourier Series with Application to AQI Data of Kolkata
by: Ghosh, Anirban, et al.
Published: (2026)
by: Ghosh, Anirban, et al.
Published: (2026)
RADIANT: Retrieval AugmenteD entIty-context AligNmenT -- Introducing RAG-ability and Entity-Context Divergence
by: Rawte, Vipula, et al.
Published: (2025)
by: Rawte, Vipula, et al.
Published: (2025)
The Brittleness of AI-Generated Image Watermarking Techniques: Examining Their Robustness Against Visual Paraphrasing Attacks
by: Barman, Niyar R, et al.
Published: (2024)
by: Barman, Niyar R, et al.
Published: (2024)
Position: The Complexity of Perfect AI Alignment -- Formalizing the RLHF Trilemma
by: Sahoo, Subramanyam, et al.
Published: (2025)
by: Sahoo, Subramanyam, et al.
Published: (2025)
MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models
by: Saha, Partha Pratim, et al.
Published: (2026)
by: Saha, Partha Pratim, et al.
Published: (2026)
Alignment For Performance Improvement in Conversation Bots
by: Garg, Raghav, et al.
Published: (2024)
by: Garg, Raghav, et al.
Published: (2024)
SAHOO: Safeguarded Alignment for High-Order Optimization Objectives in Recursive Self-Improvement
by: Sahoo, Subramanyam, et al.
Published: (2026)
by: Sahoo, Subramanyam, et al.
Published: (2026)
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization
by: Khanna, Danush, et al.
Published: (2025)
by: Khanna, Danush, et al.
Published: (2025)
Alignment Monitoring
by: Henzinger, Thomas A., et al.
Published: (2025)
by: Henzinger, Thomas A., et al.
Published: (2025)
The Visual Counter Turing Test (VCT2): A Benchmark for Evaluating AI-Generated Image Detection and the Visual AI Index (VAI)
by: Imanpour, Nasrin, et al.
Published: (2024)
by: Imanpour, Nasrin, et al.
Published: (2024)
AMBEDKAR-A Multi-level Bias Elimination through a Decoding Approach with Knowledge Augmentation for Robust Constitutional Alignment of Language Models
by: Mukhopadhyay, Snehasis, et al.
Published: (2025)
by: Mukhopadhyay, Snehasis, et al.
Published: (2025)
Neural FOXP2 -- Language Specific Neuron Steering for Targeted Language Improvement in LLMs
by: Saha, Anusa, et al.
Published: (2026)
by: Saha, Anusa, et al.
Published: (2026)
RLHF: A comprehensive Survey for Cultural, Multimodal and Low Latency Alignment Methods
by: Sharma, Raghav, et al.
Published: (2025)
by: Sharma, Raghav, et al.
Published: (2025)
Reasoning or Rhetoric? An Empirical Analysis of Moral Reasoning Explanations in Large Language Models
by: Kasat, Aryan, et al.
Published: (2026)
by: Kasat, Aryan, et al.
Published: (2026)
Born With a Silver Spoon? Investigating Socioeconomic Bias in Large Language Models
by: Singh, Smriti, et al.
Published: (2024)
by: Singh, Smriti, et al.
Published: (2024)
Curriculum Learning for Safety Alignment
by: Kumar, Sandeep, et al.
Published: (2026)
by: Kumar, Sandeep, et al.
Published: (2026)
Data-driven Policies For Two-stage Stochastic Linear Programs
by: Sharma, Chhavi, et al.
Published: (2026)
by: Sharma, Chhavi, et al.
Published: (2026)
Existence results for singular p-biharmonic problem with Hardy potential and critical Hardy-Sobolev exponent
by: Singh, Gurpreet
Published: (2024)
by: Singh, Gurpreet
Published: (2024)
LLMsAgainstHate @ NLU of Devanagari Script Languages 2025: Hate Speech Detection and Target Identification in Devanagari Languages via Parameter Efficient Fine-Tuning of LLMs
by: Sidibomma, Rushendra, et al.
Published: (2024)
by: Sidibomma, Rushendra, et al.
Published: (2024)
MAAT: Multi-phase Adapter-Aware Targeted Unlearning
by: Yagnik, Suryash, et al.
Published: (2026)
by: Yagnik, Suryash, et al.
Published: (2026)
Similar Items
-
AlignGuard-LoRA: Alignment-Preserving Fine-Tuning via Fisher-Guided Decomposition and Riemannian-Geodesic Collision Regularization
by: Das, Amitava, et al.
Published: (2025) -
YINYANG-ALIGN: Benchmarking Contradictory Objectives and Proposing Multi-Objective Optimization based DPO for Text-to-Image Alignment
by: Das, Amitava, et al.
Published: (2025) -
DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization
by: Das, Amitava, et al.
Published: (2025) -
Air in Your Neighborhood: Fine-Grained AQI Forecasting Using Mobile Sensor Data
by: Sharma, Aaryam
Published: (2025) -
AdversariaL attacK sAfety aLIgnment(ALKALI): Safeguarding LLMs through GRACE: Geometric Representation-Aware Contrastive Enhancement- Introducing Adversarial Vulnerability Quality Index (AVQI)
by: Khanna, Danush, et al.
Published: (2025)