YINYANG-ALIGN: Benchmarking Contradictory Objectives and Proposing Multi-Objective Optimization based DPO for Text-to-Image Alignment
Fuente:
arXiv
Salvato in:
| Autori principali: | Das, Amitava, Narsupalli, Yaswanth, Singh, Gurpreet, Jain, Vinija, Sharma, Vasu, Trivedy, Suranjana, Chadha, Aman, Sheth, Amit |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization
di: Das, Amitava, et al.
Pubblicazione: (2025)
di: Das, Amitava, et al.
Pubblicazione: (2025)
AdversariaL attacK sAfety aLIgnment(ALKALI): Safeguarding LLMs through GRACE: Geometric Representation-Aware Contrastive Enhancement- Introducing Adversarial Vulnerability Quality Index (AVQI)
di: Khanna, Danush, et al.
Pubblicazione: (2025)
di: Khanna, Danush, et al.
Pubblicazione: (2025)
PermaFrost-Attack: Stealth Pretraining Seeding(SPS) for planting Logic Landmines During LLM Training
di: Kumar, Harsh, et al.
Pubblicazione: (2026)
di: Kumar, Harsh, et al.
Pubblicazione: (2026)
SleepWalk: A Three-Tier Benchmark for Stress-Testing Instruction-Guided Vision-Language Navigation
di: Rawal, Niyati, et al.
Pubblicazione: (2026)
di: Rawal, Niyati, et al.
Pubblicazione: (2026)
TRACEALIGN -- Tracing the Drift: Attributing Alignment Failures to Training-Time Belief Sources in LLMs
di: Das, Amitava, et al.
Pubblicazione: (2025)
di: Das, Amitava, et al.
Pubblicazione: (2025)
D-STEER - Preference Alignment Techniques Learn to Behave, not to Believe -- Beneath the Surface, DPO as Steering Vector Perturbation in Activation Space
di: Raina, Samarth, et al.
Pubblicazione: (2025)
di: Raina, Samarth, et al.
Pubblicazione: (2025)
DETONATE: A Benchmark for Text-to-Image Alignment and Kernelized Direct Preference Optimization
di: Prasad, Renjith, et al.
Pubblicazione: (2025)
di: Prasad, Renjith, et al.
Pubblicazione: (2025)
AlignGuard-LoRA: Alignment-Preserving Fine-Tuning via Fisher-Guided Decomposition and Riemannian-Geodesic Collision Regularization
di: Das, Amitava, et al.
Pubblicazione: (2025)
di: Das, Amitava, et al.
Pubblicazione: (2025)
SAHOO: Safeguarded Alignment for High-Order Optimization Objectives in Recursive Self-Improvement
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
ECLIPTICA -- A Framework for Switchable LLM Alignment via CITA - Contrastive Instruction-Tuned Alignment
di: Wanaskar, Kapil, et al.
Pubblicazione: (2026)
di: Wanaskar, Kapil, et al.
Pubblicazione: (2026)
RADIANT: Retrieval AugmenteD entIty-context AligNmenT -- Introducing RAG-ability and Entity-Context Divergence
di: Rawte, Vipula, et al.
Pubblicazione: (2025)
di: Rawte, Vipula, et al.
Pubblicazione: (2025)
SPINAL -- Scaling-law and Preference Integration in Neural Alignment Layers
di: Das, Arion, et al.
Pubblicazione: (2026)
di: Das, Arion, et al.
Pubblicazione: (2026)
CONSCIENTIA: Can LLM Agents Learn to Strategize? Emergent Deception and Trust in a Multi-Agent NYC Simulation
di: Sinha, Aarush, et al.
Pubblicazione: (2026)
di: Sinha, Aarush, et al.
Pubblicazione: (2026)
AlignMerge - Alignment-Preserving Large Language Model Merging via Fisher-Guided Geometric Constraints
di: Roy, Aniruddha, et al.
Pubblicazione: (2025)
di: Roy, Aniruddha, et al.
Pubblicazione: (2025)
Peccavi: Visual Paraphrase Attack Safe and Distortion Free Image Watermarking Technique for AI-Generated Images
di: Dixit, Shreyas, et al.
Pubblicazione: (2025)
di: Dixit, Shreyas, et al.
Pubblicazione: (2025)
The Brittleness of AI-Generated Image Watermarking Techniques: Examining Their Robustness Against Visual Paraphrasing Attacks
di: Barman, Niyar R, et al.
Pubblicazione: (2024)
di: Barman, Niyar R, et al.
Pubblicazione: (2024)
On the Relationship between Sentence Analogy Identification and Sentence Structure Encoding in Large Language Models
di: Wijesiriwardene, Thilini, et al.
Pubblicazione: (2023)
di: Wijesiriwardene, Thilini, et al.
Pubblicazione: (2023)
KnowledgePrompts: Exploring the Abilities of Large Language Models to Solve Proportional Analogies via Knowledge-Enhanced Prompting
di: Wijesiriwardene, Thilini, et al.
Pubblicazione: (2024)
di: Wijesiriwardene, Thilini, et al.
Pubblicazione: (2024)
SEPSIS: I Can Catch Your Lies -- A New Paradigm for Deception Detection
di: Rani, Anku, et al.
Pubblicazione: (2023)
di: Rani, Anku, et al.
Pubblicazione: (2023)
ViBe: A Text-to-Video Benchmark for Evaluating Hallucination in Large Multimodal Models
di: Rawte, Vipula, et al.
Pubblicazione: (2024)
di: Rawte, Vipula, et al.
Pubblicazione: (2024)
Neural FOXP2 -- Language Specific Neuron Steering for Targeted Language Improvement in LLMs
di: Saha, Anusa, et al.
Pubblicazione: (2026)
di: Saha, Anusa, et al.
Pubblicazione: (2026)
MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models
di: Saha, Partha Pratim, et al.
Pubblicazione: (2026)
di: Saha, Partha Pratim, et al.
Pubblicazione: (2026)
Alignment Quality Index (AQI) : Beyond Refusals: AQI as an Intrinsic Alignment Diagnostic via Latent Geometry, Cluster Divergence, and Layer wise Pooled Representations
di: Borah, Abhilekh, et al.
Pubblicazione: (2025)
di: Borah, Abhilekh, et al.
Pubblicazione: (2025)
Overview of Factify5WQA: Fact Verification through 5W Question-Answering
di: Suresh, Suryavardan, et al.
Pubblicazione: (2024)
di: Suresh, Suryavardan, et al.
Pubblicazione: (2024)
LLMsAgainstHate @ NLU of Devanagari Script Languages 2025: Hate Speech Detection and Target Identification in Devanagari Languages via Parameter Efficient Fine-Tuning of LLMs
di: Sidibomma, Rushendra, et al.
Pubblicazione: (2024)
di: Sidibomma, Rushendra, et al.
Pubblicazione: (2024)
MAAT: Multi-phase Adapter-Aware Targeted Unlearning
di: Yagnik, Suryash, et al.
Pubblicazione: (2026)
di: Yagnik, Suryash, et al.
Pubblicazione: (2026)
SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference
di: Chauhan, Anay, et al.
Pubblicazione: (2026)
di: Chauhan, Anay, et al.
Pubblicazione: (2026)
Assessing LLM Reliability on Temporally Recent Open-Domain Questions
di: Krishnappa, Pushwitha, et al.
Pubblicazione: (2026)
di: Krishnappa, Pushwitha, et al.
Pubblicazione: (2026)
Position: The Complexity of Perfect AI Alignment -- Formalizing the RLHF Trilemma
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2025)
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2025)
AMBEDKAR-A Multi-level Bias Elimination through a Decoding Approach with Knowledge Augmentation for Robust Constitutional Alignment of Language Models
di: Mukhopadhyay, Snehasis, et al.
Pubblicazione: (2025)
di: Mukhopadhyay, Snehasis, et al.
Pubblicazione: (2025)
Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types
di: Sinha, Neelabh, et al.
Pubblicazione: (2024)
di: Sinha, Neelabh, et al.
Pubblicazione: (2024)
Are Small Language Models Ready to Compete with Large Language Models for Practical Applications?
di: Sinha, Neelabh, et al.
Pubblicazione: (2024)
di: Sinha, Neelabh, et al.
Pubblicazione: (2024)
Reasoning or Rhetoric? An Empirical Analysis of Moral Reasoning Explanations in Large Language Models
di: Kasat, Aryan, et al.
Pubblicazione: (2026)
di: Kasat, Aryan, et al.
Pubblicazione: (2026)
Born With a Silver Spoon? Investigating Socioeconomic Bias in Large Language Models
di: Singh, Smriti, et al.
Pubblicazione: (2024)
di: Singh, Smriti, et al.
Pubblicazione: (2024)
A Comprehensive Dataset for Human vs. AI Generated Text Detection
di: Roy, Rajarshi, et al.
Pubblicazione: (2025)
di: Roy, Rajarshi, et al.
Pubblicazione: (2025)
On the Robustness of Lexicase Selection to Contradictory Objectives
di: Shahbandegan, Shakiba, et al.
Pubblicazione: (2024)
di: Shahbandegan, Shakiba, et al.
Pubblicazione: (2024)
The What, Why, and How of Context Length Extension Techniques in Large Language Models -- A Detailed Survey
di: Pawar, Saurav, et al.
Pubblicazione: (2024)
di: Pawar, Saurav, et al.
Pubblicazione: (2024)
Refining Text-to-Image Generation: Towards Accurate Training-Free Glyph-Enhanced Image Generation
di: Lakhanpal, Sanyam, et al.
Pubblicazione: (2024)
di: Lakhanpal, Sanyam, et al.
Pubblicazione: (2024)
Findings of the Counter Turing Test: AI-Generated Image Detection
di: Roy, Rajarshi, et al.
Pubblicazione: (2026)
di: Roy, Rajarshi, et al.
Pubblicazione: (2026)
When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
Documenti analoghi
-
DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization
di: Das, Amitava, et al.
Pubblicazione: (2025) -
AdversariaL attacK sAfety aLIgnment(ALKALI): Safeguarding LLMs through GRACE: Geometric Representation-Aware Contrastive Enhancement- Introducing Adversarial Vulnerability Quality Index (AVQI)
di: Khanna, Danush, et al.
Pubblicazione: (2025) -
PermaFrost-Attack: Stealth Pretraining Seeding(SPS) for planting Logic Landmines During LLM Training
di: Kumar, Harsh, et al.
Pubblicazione: (2026) -
SleepWalk: A Three-Tier Benchmark for Stress-Testing Instruction-Guided Vision-Language Navigation
di: Rawal, Niyati, et al.
Pubblicazione: (2026) -
TRACEALIGN -- Tracing the Drift: Attributing Alignment Failures to Training-Time Belief Sources in LLMs
di: Das, Amitava, et al.
Pubblicazione: (2025)