SAHOO: Safeguarded Alignment for High-Order Optimization Objectives in Recursive Self-Improvement
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sahoo, Subramanyam, Chadha, Aman, Jain, Vinija, Chaudhary, Divya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
I Can't Believe It's Not Robust: Catastrophic Collapse of Safety Classifiers under Embedding Drift
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
Position: The Complexity of Perfect AI Alignment -- Formalizing the RLHF Trilemma
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2025)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2025)
Are Small Language Models Ready to Compete with Large Language Models for Practical Applications?
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024)
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024)
Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024)
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024)
SPINAL -- Scaling-law and Preference Integration in Neural Alignment Layers
von: Das, Arion, et al.
Veröffentlicht: (2026)
von: Das, Arion, et al.
Veröffentlicht: (2026)
Decoding the Diversity: A Review of the Indic AI Research Landscape
von: KJ, Sankalp, et al.
Veröffentlicht: (2024)
von: KJ, Sankalp, et al.
Veröffentlicht: (2024)
MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models
von: Saha, Partha Pratim, et al.
Veröffentlicht: (2026)
von: Saha, Partha Pratim, et al.
Veröffentlicht: (2026)
How Culturally Aware are Vision-Language Models?
von: Burda-Lassen, Olena, et al.
Veröffentlicht: (2024)
von: Burda-Lassen, Olena, et al.
Veröffentlicht: (2024)
Parameter Efficient Fine Tuning: A Comprehensive Analysis Across Applications
von: Balne, Charith Chandra Sai, et al.
Veröffentlicht: (2024)
von: Balne, Charith Chandra Sai, et al.
Veröffentlicht: (2024)
Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
von: Sahoo, Subramanyam
Veröffentlicht: (2026)
von: Sahoo, Subramanyam
Veröffentlicht: (2026)
Dial E for Ethical Enforcement: institutional VETO power as a governance primitive
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
PermaFrost-Attack: Stealth Pretraining Seeding(SPS) for planting Logic Landmines During LLM Training
von: Kumar, Harsh, et al.
Veröffentlicht: (2026)
von: Kumar, Harsh, et al.
Veröffentlicht: (2026)
A Comprehensive Survey of Hallucination in Large Language, Image, Video and Audio Foundation Models
von: Sahoo, Pranab, et al.
Veröffentlicht: (2024)
von: Sahoo, Pranab, et al.
Veröffentlicht: (2024)
AlignGuard-LoRA: Alignment-Preserving Fine-Tuning via Fisher-Guided Decomposition and Riemannian-Geodesic Collision Regularization
von: Das, Amitava, et al.
Veröffentlicht: (2025)
von: Das, Amitava, et al.
Veröffentlicht: (2025)
Neural FOXP2 -- Language Specific Neuron Steering for Targeted Language Improvement in LLMs
von: Saha, Anusa, et al.
Veröffentlicht: (2026)
von: Saha, Anusa, et al.
Veröffentlicht: (2026)
ECLIPTICA -- A Framework for Switchable LLM Alignment via CITA - Contrastive Instruction-Tuned Alignment
von: Wanaskar, Kapil, et al.
Veröffentlicht: (2026)
von: Wanaskar, Kapil, et al.
Veröffentlicht: (2026)
Overview of Factify5WQA: Fact Verification through 5W Question-Answering
von: Suresh, Suryavardan, et al.
Veröffentlicht: (2024)
von: Suresh, Suryavardan, et al.
Veröffentlicht: (2024)
Catch Me If You Can: How Smaller Reasoning Models Pretend to Reason with Mathematical Fidelity
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2025)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2025)
Born With a Silver Spoon? Investigating Socioeconomic Bias in Large Language Models
von: Singh, Smriti, et al.
Veröffentlicht: (2024)
von: Singh, Smriti, et al.
Veröffentlicht: (2024)
DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization
von: Das, Amitava, et al.
Veröffentlicht: (2025)
von: Das, Amitava, et al.
Veröffentlicht: (2025)
Assessing LLM Reliability on Temporally Recent Open-Domain Questions
von: Krishnappa, Pushwitha, et al.
Veröffentlicht: (2026)
von: Krishnappa, Pushwitha, et al.
Veröffentlicht: (2026)
A Comprehensive Survey of Accelerated Generation Techniques in Large Language Models
von: Khoshnoodi, Mahsa, et al.
Veröffentlicht: (2024)
von: Khoshnoodi, Mahsa, et al.
Veröffentlicht: (2024)
Multilingual State Space Models for Structured Question Answering in Indic Languages
von: Vats, Arpita, et al.
Veröffentlicht: (2025)
von: Vats, Arpita, et al.
Veröffentlicht: (2025)
A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications
von: Sahoo, Pranab, et al.
Veröffentlicht: (2024)
von: Sahoo, Pranab, et al.
Veröffentlicht: (2024)
Can Large Language Models Infer Causal Relationships from Real-World Text?
von: Saklad, Ryan, et al.
Veröffentlicht: (2025)
von: Saklad, Ryan, et al.
Veröffentlicht: (2025)
Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation
von: Zelikman, Eric, et al.
Veröffentlicht: (2023)
von: Zelikman, Eric, et al.
Veröffentlicht: (2023)
Investigating Annotator Bias in Large Language Models for Hate Speech Detection
von: Das, Amit, et al.
Veröffentlicht: (2024)
von: Das, Amit, et al.
Veröffentlicht: (2024)
Exploring the Frontier of Vision-Language Models: A Survey of Current Methodologies and Future Directions
von: Ghosh, Akash, et al.
Veröffentlicht: (2024)
von: Ghosh, Akash, et al.
Veröffentlicht: (2024)
TRACEALIGN -- Tracing the Drift: Attributing Alignment Failures to Training-Time Belief Sources in LLMs
von: Das, Amitava, et al.
Veröffentlicht: (2025)
von: Das, Amitava, et al.
Veröffentlicht: (2025)
Self-Improvement as Coherence Optimization: A Theoretical Account
von: Qiu, Tianyi, et al.
Veröffentlicht: (2026)
von: Qiu, Tianyi, et al.
Veröffentlicht: (2026)
AdversariaL attacK sAfety aLIgnment(ALKALI): Safeguarding LLMs through GRACE: Geometric Representation-Aware Contrastive Enhancement- Introducing Adversarial Vulnerability Quality Index (AVQI)
von: Khanna, Danush, et al.
Veröffentlicht: (2025)
von: Khanna, Danush, et al.
Veröffentlicht: (2025)
Reward-free Alignment for Conflicting Objectives
von: Chen, Peter, et al.
Veröffentlicht: (2026)
von: Chen, Peter, et al.
Veröffentlicht: (2026)
Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language Models
von: Li, Chengao, et al.
Veröffentlicht: (2025)
von: Li, Chengao, et al.
Veröffentlicht: (2025)
MAAT: Multi-phase Adapter-Aware Targeted Unlearning
von: Yagnik, Suryash, et al.
Veröffentlicht: (2026)
von: Yagnik, Suryash, et al.
Veröffentlicht: (2026)
Pareto Multi-Objective Alignment for Language Models
von: He, Qiang, et al.
Veröffentlicht: (2025)
von: He, Qiang, et al.
Veröffentlicht: (2025)
Certifying Knowledge Comprehension in LLMs
von: Chaudhary, Isha, et al.
Veröffentlicht: (2024)
von: Chaudhary, Isha, et al.
Veröffentlicht: (2024)
Self-Play Preference Optimization for Language Model Alignment
von: Wu, Yue, et al.
Veröffentlicht: (2024)
von: Wu, Yue, et al.
Veröffentlicht: (2024)
SKETCH: Structured Knowledge Enhanced Text Comprehension for Holistic Retrieval
von: Mahalingam, Aakash, et al.
Veröffentlicht: (2024)
von: Mahalingam, Aakash, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026) -
The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026) -
I Can't Believe It's Not Robust: Catastrophic Collapse of Safety Classifiers under Embedding Drift
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026) -
Position: The Complexity of Perfect AI Alignment -- Formalizing the RLHF Trilemma
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2025) -
Are Small Language Models Ready to Compete with Large Language Models for Practical Applications?
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024)