DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Das, Amitava, Trivedy, Suranjana, Khanna, Danush, Roy, Rajarshi, Singh, Gurpreet, Ghosh, Basab, Narsupalli, Yaswanth, Jain, Vinija, Sharma, Vasu, Reganti, Aishwarya Naresh, Chadha, Aman |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
YINYANG-ALIGN: Benchmarking Contradictory Objectives and Proposing Multi-Objective Optimization based DPO for Text-to-Image Alignment
von: Das, Amitava, et al.
Veröffentlicht: (2025)
von: Das, Amitava, et al.
Veröffentlicht: (2025)
AdversariaL attacK sAfety aLIgnment(ALKALI): Safeguarding LLMs through GRACE: Geometric Representation-Aware Contrastive Enhancement- Introducing Adversarial Vulnerability Quality Index (AVQI)
von: Khanna, Danush, et al.
Veröffentlicht: (2025)
von: Khanna, Danush, et al.
Veröffentlicht: (2025)
RADIANT: Retrieval AugmenteD entIty-context AligNmenT -- Introducing RAG-ability and Entity-Context Divergence
von: Rawte, Vipula, et al.
Veröffentlicht: (2025)
von: Rawte, Vipula, et al.
Veröffentlicht: (2025)
PermaFrost-Attack: Stealth Pretraining Seeding(SPS) for planting Logic Landmines During LLM Training
von: Kumar, Harsh, et al.
Veröffentlicht: (2026)
von: Kumar, Harsh, et al.
Veröffentlicht: (2026)
On the Relationship between Sentence Analogy Identification and Sentence Structure Encoding in Large Language Models
von: Wijesiriwardene, Thilini, et al.
Veröffentlicht: (2023)
von: Wijesiriwardene, Thilini, et al.
Veröffentlicht: (2023)
D-STEER - Preference Alignment Techniques Learn to Behave, not to Believe -- Beneath the Surface, DPO as Steering Vector Perturbation in Activation Space
von: Raina, Samarth, et al.
Veröffentlicht: (2025)
von: Raina, Samarth, et al.
Veröffentlicht: (2025)
SleepWalk: A Three-Tier Benchmark for Stress-Testing Instruction-Guided Vision-Language Navigation
von: Rawal, Niyati, et al.
Veröffentlicht: (2026)
von: Rawal, Niyati, et al.
Veröffentlicht: (2026)
DETONATE: A Benchmark for Text-to-Image Alignment and Kernelized Direct Preference Optimization
von: Prasad, Renjith, et al.
Veröffentlicht: (2025)
von: Prasad, Renjith, et al.
Veröffentlicht: (2025)
TRACEALIGN -- Tracing the Drift: Attributing Alignment Failures to Training-Time Belief Sources in LLMs
von: Das, Amitava, et al.
Veröffentlicht: (2025)
von: Das, Amitava, et al.
Veröffentlicht: (2025)
Overview of Factify5WQA: Fact Verification through 5W Question-Answering
von: Suresh, Suryavardan, et al.
Veröffentlicht: (2024)
von: Suresh, Suryavardan, et al.
Veröffentlicht: (2024)
CONSCIENTIA: Can LLM Agents Learn to Strategize? Emergent Deception and Trust in a Multi-Agent NYC Simulation
von: Sinha, Aarush, et al.
Veröffentlicht: (2026)
von: Sinha, Aarush, et al.
Veröffentlicht: (2026)
AlignGuard-LoRA: Alignment-Preserving Fine-Tuning via Fisher-Guided Decomposition and Riemannian-Geodesic Collision Regularization
von: Das, Amitava, et al.
Veröffentlicht: (2025)
von: Das, Amitava, et al.
Veröffentlicht: (2025)
Alignment Quality Index (AQI) : Beyond Refusals: AQI as an Intrinsic Alignment Diagnostic via Latent Geometry, Cluster Divergence, and Layer wise Pooled Representations
von: Borah, Abhilekh, et al.
Veröffentlicht: (2025)
von: Borah, Abhilekh, et al.
Veröffentlicht: (2025)
ViBe: A Text-to-Video Benchmark for Evaluating Hallucination in Large Multimodal Models
von: Rawte, Vipula, et al.
Veröffentlicht: (2024)
von: Rawte, Vipula, et al.
Veröffentlicht: (2024)
Findings of the Counter Turing Test: AI-Generated Image Detection
von: Roy, Rajarshi, et al.
Veröffentlicht: (2026)
von: Roy, Rajarshi, et al.
Veröffentlicht: (2026)
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization
von: Khanna, Danush, et al.
Veröffentlicht: (2025)
von: Khanna, Danush, et al.
Veröffentlicht: (2025)
Peccavi: Visual Paraphrase Attack Safe and Distortion Free Image Watermarking Technique for AI-Generated Images
von: Dixit, Shreyas, et al.
Veröffentlicht: (2025)
von: Dixit, Shreyas, et al.
Veröffentlicht: (2025)
SPINAL -- Scaling-law and Preference Integration in Neural Alignment Layers
von: Das, Arion, et al.
Veröffentlicht: (2026)
von: Das, Arion, et al.
Veröffentlicht: (2026)
A Comprehensive Dataset for Human vs. AI Generated Text Detection
von: Roy, Rajarshi, et al.
Veröffentlicht: (2025)
von: Roy, Rajarshi, et al.
Veröffentlicht: (2025)
A Comprehensive Dataset for Human vs. AI Generated Image Detection
von: Roy, Rajarshi, et al.
Veröffentlicht: (2026)
von: Roy, Rajarshi, et al.
Veröffentlicht: (2026)
AlignMerge - Alignment-Preserving Large Language Model Merging via Fisher-Guided Geometric Constraints
von: Roy, Aniruddha, et al.
Veröffentlicht: (2025)
von: Roy, Aniruddha, et al.
Veröffentlicht: (2025)
Neural FOXP2 -- Language Specific Neuron Steering for Targeted Language Improvement in LLMs
von: Saha, Anusa, et al.
Veröffentlicht: (2026)
von: Saha, Anusa, et al.
Veröffentlicht: (2026)
ECLIPTICA -- A Framework for Switchable LLM Alignment via CITA - Contrastive Instruction-Tuned Alignment
von: Wanaskar, Kapil, et al.
Veröffentlicht: (2026)
von: Wanaskar, Kapil, et al.
Veröffentlicht: (2026)
Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024)
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024)
Are Small Language Models Ready to Compete with Large Language Models for Practical Applications?
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024)
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024)
SEPSIS: I Can Catch Your Lies -- A New Paradigm for Deception Detection
von: Rani, Anku, et al.
Veröffentlicht: (2023)
von: Rani, Anku, et al.
Veröffentlicht: (2023)
LLMsAgainstHate @ NLU of Devanagari Script Languages 2025: Hate Speech Detection and Target Identification in Devanagari Languages via Parameter Efficient Fine-Tuning of LLMs
von: Sidibomma, Rushendra, et al.
Veröffentlicht: (2024)
von: Sidibomma, Rushendra, et al.
Veröffentlicht: (2024)
MAAT: Multi-phase Adapter-Aware Targeted Unlearning
von: Yagnik, Suryash, et al.
Veröffentlicht: (2026)
von: Yagnik, Suryash, et al.
Veröffentlicht: (2026)
Smoothness of Function in Terms of Kernel and Its Relation With K ‐Derivative
von: Suranjana Deb
Veröffentlicht: (2025)
von: Suranjana Deb
Veröffentlicht: (2025)
Exploring the Frontier of Vision-Language Models: A Survey of Current Methodologies and Future Directions
von: Ghosh, Akash, et al.
Veröffentlicht: (2024)
von: Ghosh, Akash, et al.
Veröffentlicht: (2024)
When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
Reasoning or Rhetoric? An Empirical Analysis of Moral Reasoning Explanations in Large Language Models
von: Kasat, Aryan, et al.
Veröffentlicht: (2026)
von: Kasat, Aryan, et al.
Veröffentlicht: (2026)
Dial E for Ethical Enforcement: institutional VETO power as a governance primitive
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
Born With a Silver Spoon? Investigating Socioeconomic Bias in Large Language Models
von: Singh, Smriti, et al.
Veröffentlicht: (2024)
von: Singh, Smriti, et al.
Veröffentlicht: (2024)
SAHOO: Safeguarded Alignment for High-Order Optimization Objectives in Recursive Self-Improvement
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
I Can't Believe It's Not Robust: Catastrophic Collapse of Safety Classifiers under Embedding Drift
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
Position: The Complexity of Perfect AI Alignment -- Formalizing the RLHF Trilemma
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2025)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2025)
Exploring the Impact of Large Language Models on Recommender Systems: An Extensive Review
von: Vats, Arpita, et al.
Veröffentlicht: (2024)
von: Vats, Arpita, et al.
Veröffentlicht: (2024)
DeHate: A Stable Diffusion-based Multimodal Approach to Mitigate Hate Speech in Images
von: Dalal, Dwip, et al.
Veröffentlicht: (2025)
von: Dalal, Dwip, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
YINYANG-ALIGN: Benchmarking Contradictory Objectives and Proposing Multi-Objective Optimization based DPO for Text-to-Image Alignment
von: Das, Amitava, et al.
Veröffentlicht: (2025) -
AdversariaL attacK sAfety aLIgnment(ALKALI): Safeguarding LLMs through GRACE: Geometric Representation-Aware Contrastive Enhancement- Introducing Adversarial Vulnerability Quality Index (AVQI)
von: Khanna, Danush, et al.
Veröffentlicht: (2025) -
RADIANT: Retrieval AugmenteD entIty-context AligNmenT -- Introducing RAG-ability and Entity-Context Divergence
von: Rawte, Vipula, et al.
Veröffentlicht: (2025) -
PermaFrost-Attack: Stealth Pretraining Seeding(SPS) for planting Logic Landmines During LLM Training
von: Kumar, Harsh, et al.
Veröffentlicht: (2026) -
On the Relationship between Sentence Analogy Identification and Sentence Structure Encoding in Large Language Models
von: Wijesiriwardene, Thilini, et al.
Veröffentlicht: (2023)