The Realignment Problem: When Right becomes Wrong in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Sharma, Aakash Sen, Sanyal, Debdeep, Ray, Manodeep, Srivastava, Vivek, Karande, Shirish, Mandal, Murari |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs
by: Sanyal, Debdeep, et al.
Published: (2025)
by: Sanyal, Debdeep, et al.
Published: (2025)
Nine Ways to Break Copyright Law and Why Our LLM Won't: A Fair Use Aligned Generation Framework
by: Sharma, Aakash Sen, et al.
Published: (2025)
by: Sharma, Aakash Sen, et al.
Published: (2025)
Policy Optimization Prefers The Path of Least Resistance
by: Sanyal, Debdeep, et al.
Published: (2025)
by: Sanyal, Debdeep, et al.
Published: (2025)
Agents Are All You Need for LLM Unlearning
by: Sanyal, Debdeep, et al.
Published: (2025)
by: Sanyal, Debdeep, et al.
Published: (2025)
Forgetting is Competition: Rethinking Unlearning as Representation Interference in Diffusion Models
by: Ranjan, Ashutosh, et al.
Published: (2026)
by: Ranjan, Ashutosh, et al.
Published: (2026)
OrgAccess: A Benchmark for Role Based Access Control in Organization Scale LLMs
by: Sanyal, Debdeep, et al.
Published: (2025)
by: Sanyal, Debdeep, et al.
Published: (2025)
Guardians of Generation: Dynamic Inference-Time Copyright Shielding with Adaptive Guidance for AI Image Generation
by: Roy, Soham, et al.
Published: (2025)
by: Roy, Soham, et al.
Published: (2025)
Confidence is Not Competence
by: Sanyal, Debdeep, et al.
Published: (2025)
by: Sanyal, Debdeep, et al.
Published: (2025)
The Silent Brush: Evaluating Artistic Style Leakage in AI Art Generation
by: Joshi, Ninad, et al.
Published: (2026)
by: Joshi, Ninad, et al.
Published: (2026)
Polaris: A Gödel Agent Framework for Small Language Models through Experience-Abstracted Policy Repair
by: Kakade, Aditya, et al.
Published: (2026)
by: Kakade, Aditya, et al.
Published: (2026)
Persuasion Games using Large Language Models
by: Ramani, Ganesh Prasath, et al.
Published: (2024)
by: Ramani, Ganesh Prasath, et al.
Published: (2024)
UnStar: Unlearning with Self-Taught Anti-Sample Reasoning for LLMs
by: Sinha, Yash, et al.
Published: (2024)
by: Sinha, Yash, et al.
Published: (2024)
Beyond Accuracy: Diagnosing Algebraic Reasoning Failures in LLMs Across Nine Complexity Dimensions
by: Patil, Parth, et al.
Published: (2026)
by: Patil, Parth, et al.
Published: (2026)
Easy Problems That LLMs Get Wrong
by: Williams, Sean, et al.
Published: (2024)
by: Williams, Sean, et al.
Published: (2024)
NagaNLP: Bootstrapping NLP for Low-Resource Nagamese Creole with Human-in-the-Loop Synthetic Data
by: Maiti, Agniva, et al.
Published: (2025)
by: Maiti, Agniva, et al.
Published: (2025)
Varying Shades of Wrong: Aligning LLMs with Wrong Answers Only
by: Yao, Jihan, et al.
Published: (2024)
by: Yao, Jihan, et al.
Published: (2024)
Flexible Realignment of Language Models
by: Zhu, Wenhong, et al.
Published: (2025)
by: Zhu, Wenhong, et al.
Published: (2025)
time2time: Causal Intervention in Hidden States to Simulate Rare Events in Time Series Foundation Models
by: Sanyal, Debdeep, et al.
Published: (2025)
by: Sanyal, Debdeep, et al.
Published: (2025)
Distinguishing Right from Wrong in Debates: Attribution Analysis of Chinese Harmful Memes
by: Wang, Weiming, et al.
Published: (2026)
by: Wang, Weiming, et al.
Published: (2026)
Perplexity Cannot Always Tell Right from Wrong
by: Veličković, Petar, et al.
Published: (2026)
by: Veličković, Petar, et al.
Published: (2026)
LLM-as-a-Judge for Time Series Explanations
by: Sivalingam, Preetham, et al.
Published: (2026)
by: Sivalingam, Preetham, et al.
Published: (2026)
Evaluating LLMs' Mathematical Reasoning in Financial Document Question Answering
by: Srivastava, Pragya, et al.
Published: (2024)
by: Srivastava, Pragya, et al.
Published: (2024)
When Flores Bloomz Wrong: Cross-Direction Contamination in Machine Translation Evaluation
by: Tan, David, et al.
Published: (2026)
by: Tan, David, et al.
Published: (2026)
When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks
by: Fleisig, Eve, et al.
Published: (2023)
by: Fleisig, Eve, et al.
Published: (2023)
When Attention Collapses: How Degenerate Layers in LLMs Enable Smaller, Stronger Models
by: Sanyal, Sunny, et al.
Published: (2024)
by: Sanyal, Sunny, et al.
Published: (2024)
The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training
by: Zhang, Rui, et al.
Published: (2026)
by: Zhang, Rui, et al.
Published: (2026)
Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong
by: Fu, Tairan, et al.
Published: (2025)
by: Fu, Tairan, et al.
Published: (2025)
Do LLMs Signal When They're Right? Evidence from Neuron Agreement
by: Chen, Kang, et al.
Published: (2025)
by: Chen, Kang, et al.
Published: (2025)
Decoding-time Realignment of Language Models
by: Liu, Tianlin, et al.
Published: (2024)
by: Liu, Tianlin, et al.
Published: (2024)
Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering
by: Molfese, Francesco Maria, et al.
Published: (2025)
by: Molfese, Francesco Maria, et al.
Published: (2025)
Constructions Are So Difficult That Even Large Language Models Get Them Right for the Wrong Reasons
by: Zhou, Shijia, et al.
Published: (2024)
by: Zhou, Shijia, et al.
Published: (2024)
IndicParam: Benchmark to evaluate LLMs on low-resource Indic Languages
by: Maheshwari, Ayush, et al.
Published: (2025)
by: Maheshwari, Ayush, et al.
Published: (2025)
When AI companions become witty: Can human brain recognize AI-generated irony?
by: Rao, Xiaohui, et al.
Published: (2025)
by: Rao, Xiaohui, et al.
Published: (2025)
Ask Early, Ask Late, Ask Right: When Does Clarification Timing Matter for Long-Horizon Agents?
by: Gulati, Anmol, et al.
Published: (2026)
by: Gulati, Anmol, et al.
Published: (2026)
When Reject Turns into Accept: Quantifying the Vulnerability of LLM-Based Scientific Reviewers to Indirect Prompt Injection
by: Sahoo, Devanshu, et al.
Published: (2025)
by: Sahoo, Devanshu, et al.
Published: (2025)
GeoAlign: Geometric Feature Realignment for MLLM Spatial Reasoning
by: Liu, Zhaochen, et al.
Published: (2026)
by: Liu, Zhaochen, et al.
Published: (2026)
Large Language Models Help Humans Verify Truthfulness -- Except When They Are Convincingly Wrong
by: Si, Chenglei, et al.
Published: (2023)
by: Si, Chenglei, et al.
Published: (2023)
Simpler becomes Harder: Do LLMs Exhibit a Coherent Behavior on Simplified Corpora?
by: Anschütz, Miriam, et al.
Published: (2024)
by: Anschütz, Miriam, et al.
Published: (2024)
Structured Context Recomposition for Large Language Models Using Probabilistic Layer Realignment
by: Teel, Jonathan, et al.
Published: (2025)
by: Teel, Jonathan, et al.
Published: (2025)
Measuring Representation Robustness in Large Language Models for Geometry
by: Jawandhia, Vedant, et al.
Published: (2026)
by: Jawandhia, Vedant, et al.
Published: (2026)
Similar Items
-
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs
by: Sanyal, Debdeep, et al.
Published: (2025) -
Nine Ways to Break Copyright Law and Why Our LLM Won't: A Fair Use Aligned Generation Framework
by: Sharma, Aakash Sen, et al.
Published: (2025) -
Policy Optimization Prefers The Path of Least Resistance
by: Sanyal, Debdeep, et al.
Published: (2025) -
Agents Are All You Need for LLM Unlearning
by: Sanyal, Debdeep, et al.
Published: (2025) -
Forgetting is Competition: Rethinking Unlearning as Representation Interference in Diffusion Models
by: Ranjan, Ashutosh, et al.
Published: (2026)