Position: The Complexity of Perfect AI Alignment -- Formalizing the RLHF Trilemma
Fuente:
arXiv
Saved in:
| Main Authors: | Sahoo, Subramanyam, Chadha, Aman, Jain, Vinija, Chaudhary, Divya |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SAHOO: Safeguarded Alignment for High-Order Optimization Objectives in Recursive Self-Improvement
by: Sahoo, Subramanyam, et al.
Published: (2026)
by: Sahoo, Subramanyam, et al.
Published: (2026)
I Can't Believe It's Not Robust: Catastrophic Collapse of Safety Classifiers under Embedding Drift
by: Sahoo, Subramanyam, et al.
Published: (2026)
by: Sahoo, Subramanyam, et al.
Published: (2026)
When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning
by: Sahoo, Subramanyam, et al.
Published: (2026)
by: Sahoo, Subramanyam, et al.
Published: (2026)
The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness
by: Sahoo, Subramanyam, et al.
Published: (2026)
by: Sahoo, Subramanyam, et al.
Published: (2026)
Dial E for Ethical Enforcement: institutional VETO power as a governance primitive
by: Sahoo, Subramanyam, et al.
Published: (2026)
by: Sahoo, Subramanyam, et al.
Published: (2026)
AlignGuard-LoRA: Alignment-Preserving Fine-Tuning via Fisher-Guided Decomposition and Riemannian-Geodesic Collision Regularization
by: Das, Amitava, et al.
Published: (2025)
by: Das, Amitava, et al.
Published: (2025)
ECLIPTICA -- A Framework for Switchable LLM Alignment via CITA - Contrastive Instruction-Tuned Alignment
by: Wanaskar, Kapil, et al.
Published: (2026)
by: Wanaskar, Kapil, et al.
Published: (2026)
Are Small Language Models Ready to Compete with Large Language Models for Practical Applications?
by: Sinha, Neelabh, et al.
Published: (2024)
by: Sinha, Neelabh, et al.
Published: (2024)
D-STEER - Preference Alignment Techniques Learn to Behave, not to Believe -- Beneath the Surface, DPO as Steering Vector Perturbation in Activation Space
by: Raina, Samarth, et al.
Published: (2025)
by: Raina, Samarth, et al.
Published: (2025)
Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types
by: Sinha, Neelabh, et al.
Published: (2024)
by: Sinha, Neelabh, et al.
Published: (2024)
Catch Me If You Can: How Smaller Reasoning Models Pretend to Reason with Mathematical Fidelity
by: Sahoo, Subramanyam, et al.
Published: (2025)
by: Sahoo, Subramanyam, et al.
Published: (2025)
Decoding the Diversity: A Review of the Indic AI Research Landscape
by: KJ, Sankalp, et al.
Published: (2024)
by: KJ, Sankalp, et al.
Published: (2024)
SPINAL -- Scaling-law and Preference Integration in Neural Alignment Layers
by: Das, Arion, et al.
Published: (2026)
by: Das, Arion, et al.
Published: (2026)
How Culturally Aware are Vision-Language Models?
by: Burda-Lassen, Olena, et al.
Published: (2024)
by: Burda-Lassen, Olena, et al.
Published: (2024)
A Comprehensive Survey of Hallucination in Large Language, Image, Video and Audio Foundation Models
by: Sahoo, Pranab, et al.
Published: (2024)
by: Sahoo, Pranab, et al.
Published: (2024)
MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models
by: Saha, Partha Pratim, et al.
Published: (2026)
by: Saha, Partha Pratim, et al.
Published: (2026)
Moral Sensitivity in LLMs: A Tiered Evaluation of Contextual Bias via Behavioral Profiling and Mechanistic Interpretability
by: Aggarwal, Yash, et al.
Published: (2026)
by: Aggarwal, Yash, et al.
Published: (2026)
MAAT: Multi-phase Adapter-Aware Targeted Unlearning
by: Yagnik, Suryash, et al.
Published: (2026)
by: Yagnik, Suryash, et al.
Published: (2026)
The Double Life of Code World Models: Provably Unmasking Malicious Behavior Through Execution Traces
by: Sahoo, Subramanyam
Published: (2025)
by: Sahoo, Subramanyam
Published: (2025)
The Good, The Bad, and The Hybrid: A Reward Structure Showdown in Reasoning Models Training
by: Sahoo, Subramanyam
Published: (2025)
by: Sahoo, Subramanyam
Published: (2025)
The Horcrux: Mechanistically Interpretable Task Decomposition for Detecting and Mitigating Reward Hacking in Embodied AI Systems
by: Sahoo, Subramanyam, et al.
Published: (2025)
by: Sahoo, Subramanyam, et al.
Published: (2025)
Solving the Inverse Alignment Problem for Efficient RLHF
by: Krishna, Shambhavi, et al.
Published: (2024)
by: Krishna, Shambhavi, et al.
Published: (2024)
TRACEALIGN -- Tracing the Drift: Attributing Alignment Failures to Training-Time Belief Sources in LLMs
by: Das, Amitava, et al.
Published: (2025)
by: Das, Amitava, et al.
Published: (2025)
Parameter Efficient Fine Tuning: A Comprehensive Analysis Across Applications
by: Balne, Charith Chandra Sai, et al.
Published: (2024)
by: Balne, Charith Chandra Sai, et al.
Published: (2024)
SKETCH: Structured Knowledge Enhanced Text Comprehension for Holistic Retrieval
by: Mahalingam, Aakash, et al.
Published: (2024)
by: Mahalingam, Aakash, et al.
Published: (2024)
Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
by: Sahoo, Subramanyam
Published: (2026)
by: Sahoo, Subramanyam
Published: (2026)
PermaFrost-Attack: Stealth Pretraining Seeding(SPS) for planting Logic Landmines During LLM Training
by: Kumar, Harsh, et al.
Published: (2026)
by: Kumar, Harsh, et al.
Published: (2026)
Boardwalk Empire: How Generative AI is Revolutionizing Economic Paradigms
by: Sahoo, Subramanyam, et al.
Published: (2024)
by: Sahoo, Subramanyam, et al.
Published: (2024)
AdversariaL attacK sAfety aLIgnment(ALKALI): Safeguarding LLMs through GRACE: Geometric Representation-Aware Contrastive Enhancement- Introducing Adversarial Vulnerability Quality Index (AVQI)
by: Khanna, Danush, et al.
Published: (2025)
by: Khanna, Danush, et al.
Published: (2025)
A Novel XAI-Enhanced Quantum Adversarial Networks for Velocity Dispersion Modeling in MaNGA Galaxies
by: Narkedimilli, Sathwik, et al.
Published: (2025)
by: Narkedimilli, Sathwik, et al.
Published: (2025)
The Deepfake Detective: Interpreting Neural Forensics Through Sparse Features and Manifolds
by: Sahoo, Subramanyam, et al.
Published: (2025)
by: Sahoo, Subramanyam, et al.
Published: (2025)
Mitigating the Alignment Tax of RLHF
by: Lin, Yong, et al.
Published: (2023)
by: Lin, Yong, et al.
Published: (2023)
The Perfect Blend: Redefining RLHF with Mixture of Judges
by: Xu, Tengyu, et al.
Published: (2024)
by: Xu, Tengyu, et al.
Published: (2024)
Overview of Factify5WQA: Fact Verification through 5W Question-Answering
by: Suresh, Suryavardan, et al.
Published: (2024)
by: Suresh, Suryavardan, et al.
Published: (2024)
SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference
by: Chauhan, Anay, et al.
Published: (2026)
by: Chauhan, Anay, et al.
Published: (2026)
Leveraging Geolocation in Clinical Records to Improve Alzheimer's Disease Diagnosis Using DMV Framework
by: Zhang, Peng, et al.
Published: (2025)
by: Zhang, Peng, et al.
Published: (2025)
Explaining and Preventing Alignment Collapse in Iterative RLHF
by: Gauthier, Etienne, et al.
Published: (2026)
by: Gauthier, Etienne, et al.
Published: (2026)
SQL-of-Thought: Multi-agentic Text-to-SQL with Guided Error Correction
by: Chaturvedi, Saumya, et al.
Published: (2025)
by: Chaturvedi, Saumya, et al.
Published: (2025)
AlignMerge - Alignment-Preserving Large Language Model Merging via Fisher-Guided Geometric Constraints
by: Roy, Aniruddha, et al.
Published: (2025)
by: Roy, Aniruddha, et al.
Published: (2025)
Beyond RLHF: A Unified Theoretical Framework of Alignment
by: Yun, Jihun, et al.
Published: (2025)
by: Yun, Jihun, et al.
Published: (2025)
Similar Items
-
SAHOO: Safeguarded Alignment for High-Order Optimization Objectives in Recursive Self-Improvement
by: Sahoo, Subramanyam, et al.
Published: (2026) -
I Can't Believe It's Not Robust: Catastrophic Collapse of Safety Classifiers under Embedding Drift
by: Sahoo, Subramanyam, et al.
Published: (2026) -
When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning
by: Sahoo, Subramanyam, et al.
Published: (2026) -
The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness
by: Sahoo, Subramanyam, et al.
Published: (2026) -
Dial E for Ethical Enforcement: institutional VETO power as a governance primitive
by: Sahoo, Subramanyam, et al.
Published: (2026)