Alignment Faking - the Train -> Deploy Asymmetry: Through a Game-Theoretic Lens with Bayesian-Stackelberg Equilibria
Fuente:
arXiv
Guardado en:
| Autores principales: | Garg, Kartik, Mishra, Shourya, Sinha, Kartikeya, Singh, Ojaswi Pratap, Chopra, Ayush, Rai, Kanishk, Sheikh, Ammar, Maheshwari, Raghav, Chadha, Aman, Jain, Vinija, Das, Amitava |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
TRACEALIGN -- Tracing the Drift: Attributing Alignment Failures to Training-Time Belief Sources in LLMs
por: Das, Amitava, et al.
Publicado: (2025)
por: Das, Amitava, et al.
Publicado: (2025)
Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types
por: Sinha, Neelabh, et al.
Publicado: (2024)
por: Sinha, Neelabh, et al.
Publicado: (2024)
Are Small Language Models Ready to Compete with Large Language Models for Practical Applications?
por: Sinha, Neelabh, et al.
Publicado: (2024)
por: Sinha, Neelabh, et al.
Publicado: (2024)
AlignGuard-LoRA: Alignment-Preserving Fine-Tuning via Fisher-Guided Decomposition and Riemannian-Geodesic Collision Regularization
por: Das, Amitava, et al.
Publicado: (2025)
por: Das, Amitava, et al.
Publicado: (2025)
Stochastic CHAOS: Why Deterministic Inference Kills, and Distributional Variability Is the Heartbeat of Artifical Cognition
por: Joshi, Tanmay, et al.
Publicado: (2026)
por: Joshi, Tanmay, et al.
Publicado: (2026)
D-STEER - Preference Alignment Techniques Learn to Behave, not to Believe -- Beneath the Surface, DPO as Steering Vector Perturbation in Activation Space
por: Raina, Samarth, et al.
Publicado: (2025)
por: Raina, Samarth, et al.
Publicado: (2025)
AlignMerge - Alignment-Preserving Large Language Model Merging via Fisher-Guided Geometric Constraints
por: Roy, Aniruddha, et al.
Publicado: (2025)
por: Roy, Aniruddha, et al.
Publicado: (2025)
Neural FOXP2 -- Language Specific Neuron Steering for Targeted Language Improvement in LLMs
por: Saha, Anusa, et al.
Publicado: (2026)
por: Saha, Anusa, et al.
Publicado: (2026)
ECLIPTICA -- A Framework for Switchable LLM Alignment via CITA - Contrastive Instruction-Tuned Alignment
por: Wanaskar, Kapil, et al.
Publicado: (2026)
por: Wanaskar, Kapil, et al.
Publicado: (2026)
Refining Text-to-Image Generation: Towards Accurate Training-Free Glyph-Enhanced Image Generation
por: Lakhanpal, Sanyam, et al.
Publicado: (2024)
por: Lakhanpal, Sanyam, et al.
Publicado: (2024)
LLMsAgainstHate @ NLU of Devanagari Script Languages 2025: Hate Speech Detection and Target Identification in Devanagari Languages via Parameter Efficient Fine-Tuning of LLMs
por: Sidibomma, Rushendra, et al.
Publicado: (2024)
por: Sidibomma, Rushendra, et al.
Publicado: (2024)
MAAT: Multi-phase Adapter-Aware Targeted Unlearning
por: Yagnik, Suryash, et al.
Publicado: (2026)
por: Yagnik, Suryash, et al.
Publicado: (2026)
SPINAL -- Scaling-law and Preference Integration in Neural Alignment Layers
por: Das, Arion, et al.
Publicado: (2026)
por: Das, Arion, et al.
Publicado: (2026)
Peccavi: Visual Paraphrase Attack Safe and Distortion Free Image Watermarking Technique for AI-Generated Images
por: Dixit, Shreyas, et al.
Publicado: (2025)
por: Dixit, Shreyas, et al.
Publicado: (2025)
PermaFrost-Attack: Stealth Pretraining Seeding(SPS) for planting Logic Landmines During LLM Training
por: Kumar, Harsh, et al.
Publicado: (2026)
por: Kumar, Harsh, et al.
Publicado: (2026)
A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications
por: Sahoo, Pranab, et al.
Publicado: (2024)
por: Sahoo, Pranab, et al.
Publicado: (2024)
When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning
por: Sahoo, Subramanyam, et al.
Publicado: (2026)
por: Sahoo, Subramanyam, et al.
Publicado: (2026)
Reasoning or Rhetoric? An Empirical Analysis of Moral Reasoning Explanations in Large Language Models
por: Kasat, Aryan, et al.
Publicado: (2026)
por: Kasat, Aryan, et al.
Publicado: (2026)
Dial E for Ethical Enforcement: institutional VETO power as a governance primitive
por: Sahoo, Subramanyam, et al.
Publicado: (2026)
por: Sahoo, Subramanyam, et al.
Publicado: (2026)
The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness
por: Sahoo, Subramanyam, et al.
Publicado: (2026)
por: Sahoo, Subramanyam, et al.
Publicado: (2026)
Born With a Silver Spoon? Investigating Socioeconomic Bias in Large Language Models
por: Singh, Smriti, et al.
Publicado: (2024)
por: Singh, Smriti, et al.
Publicado: (2024)
SAHOO: Safeguarded Alignment for High-Order Optimization Objectives in Recursive Self-Improvement
por: Sahoo, Subramanyam, et al.
Publicado: (2026)
por: Sahoo, Subramanyam, et al.
Publicado: (2026)
I Can't Believe It's Not Robust: Catastrophic Collapse of Safety Classifiers under Embedding Drift
por: Sahoo, Subramanyam, et al.
Publicado: (2026)
por: Sahoo, Subramanyam, et al.
Publicado: (2026)
Position: The Complexity of Perfect AI Alignment -- Formalizing the RLHF Trilemma
por: Sahoo, Subramanyam, et al.
Publicado: (2025)
por: Sahoo, Subramanyam, et al.
Publicado: (2025)
Exploring the Impact of Large Language Models on Recommender Systems: An Extensive Review
por: Vats, Arpita, et al.
Publicado: (2024)
por: Vats, Arpita, et al.
Publicado: (2024)
On the Relationship between Sentence Analogy Identification and Sentence Structure Encoding in Large Language Models
por: Wijesiriwardene, Thilini, et al.
Publicado: (2023)
por: Wijesiriwardene, Thilini, et al.
Publicado: (2023)
The What, Why, and How of Context Length Extension Techniques in Large Language Models -- A Detailed Survey
por: Pawar, Saurav, et al.
Publicado: (2024)
por: Pawar, Saurav, et al.
Publicado: (2024)
MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models
por: Saha, Partha Pratim, et al.
Publicado: (2026)
por: Saha, Partha Pratim, et al.
Publicado: (2026)
Language Models Entangle Language and Culture
por: Jain, Shourya, et al.
Publicado: (2026)
por: Jain, Shourya, et al.
Publicado: (2026)
Robust Stackelberg Equilibria
por: Gan, Jiarui, et al.
Publicado: (2023)
por: Gan, Jiarui, et al.
Publicado: (2023)
AMBEDKAR-A Multi-level Bias Elimination through a Decoding Approach with Knowledge Augmentation for Robust Constitutional Alignment of Language Models
por: Mukhopadhyay, Snehasis, et al.
Publicado: (2025)
por: Mukhopadhyay, Snehasis, et al.
Publicado: (2025)
KnowledgePrompts: Exploring the Abilities of Large Language Models to Solve Proportional Analogies via Knowledge-Enhanced Prompting
por: Wijesiriwardene, Thilini, et al.
Publicado: (2024)
por: Wijesiriwardene, Thilini, et al.
Publicado: (2024)
YINYANG-ALIGN: Benchmarking Contradictory Objectives and Proposing Multi-Objective Optimization based DPO for Text-to-Image Alignment
por: Das, Amitava, et al.
Publicado: (2025)
por: Das, Amitava, et al.
Publicado: (2025)
SEPSIS: I Can Catch Your Lies -- A New Paradigm for Deception Detection
por: Rani, Anku, et al.
Publicado: (2023)
por: Rani, Anku, et al.
Publicado: (2023)
SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference
por: Chauhan, Anay, et al.
Publicado: (2026)
por: Chauhan, Anay, et al.
Publicado: (2026)
How Culturally Aware are Vision-Language Models?
por: Burda-Lassen, Olena, et al.
Publicado: (2024)
por: Burda-Lassen, Olena, et al.
Publicado: (2024)
CONSCIENTIA: Can LLM Agents Learn to Strategize? Emergent Deception and Trust in a Multi-Agent NYC Simulation
por: Sinha, Aarush, et al.
Publicado: (2026)
por: Sinha, Aarush, et al.
Publicado: (2026)
SleepWalk: A Three-Tier Benchmark for Stress-Testing Instruction-Guided Vision-Language Navigation
por: Rawal, Niyati, et al.
Publicado: (2026)
por: Rawal, Niyati, et al.
Publicado: (2026)
A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models
por: Tonmoy, S. M Towhidul Islam, et al.
Publicado: (2024)
por: Tonmoy, S. M Towhidul Islam, et al.
Publicado: (2024)
Source-Free Domain Adaptation with Diffusion-Guided Source Data Generation
por: Chopra, Shivang, et al.
Publicado: (2024)
por: Chopra, Shivang, et al.
Publicado: (2024)
Ejemplares similares
-
TRACEALIGN -- Tracing the Drift: Attributing Alignment Failures to Training-Time Belief Sources in LLMs
por: Das, Amitava, et al.
Publicado: (2025) -
Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types
por: Sinha, Neelabh, et al.
Publicado: (2024) -
Are Small Language Models Ready to Compete with Large Language Models for Practical Applications?
por: Sinha, Neelabh, et al.
Publicado: (2024) -
AlignGuard-LoRA: Alignment-Preserving Fine-Tuning via Fisher-Guided Decomposition and Riemannian-Geodesic Collision Regularization
por: Das, Amitava, et al.
Publicado: (2025) -
Stochastic CHAOS: Why Deterministic Inference Kills, and Distributional Variability Is the Heartbeat of Artifical Cognition
por: Joshi, Tanmay, et al.
Publicado: (2026)