One Goal, Many Challenges: Robust Preference Optimization Amid Content-Aware and Multi-Source Noise
Fuente:
arXiv
Saved in:
| Main Authors: | Afzali, Amirabbas, Afsharrad, Amirhossein, Mousavi, Seyed Shahabeddin, Lall, Sanjay |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LORE: Lagrangian-Optimized Robust Embeddings for Visual Encoders
by: Khodabandeh, Borna, et al.
Published: (2025)
by: Khodabandeh, Borna, et al.
Published: (2025)
Multi-Agent Stage-wise Conservative Linear Bandits
by: Afsharrad, Amirhossein, et al.
Published: (2025)
by: Afsharrad, Amirhossein, et al.
Published: (2025)
Cooperative Multi-Agent Constrained Stochastic Linear Bandits
by: Afsharrad, Amirhossein, et al.
Published: (2024)
by: Afsharrad, Amirhossein, et al.
Published: (2024)
Beyond Binary Preferences: A Principled Framework for Reward Modeling with Ordinal Feedback
by: Afsharrad, Amirhossein, et al.
Published: (2026)
by: Afsharrad, Amirhossein, et al.
Published: (2026)
Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems
by: Allbert, Rumi, et al.
Published: (2025)
by: Allbert, Rumi, et al.
Published: (2025)
The Gray Area: Characterizing Moderator Disagreement on Reddit
by: Alipour, Shayan, et al.
Published: (2026)
by: Alipour, Shayan, et al.
Published: (2026)
When Weak LLMs Speak with Confidence, Preference Alignment Gets Stronger
by: Afzali, Amirabbas, et al.
Published: (2026)
by: Afzali, Amirabbas, et al.
Published: (2026)
Aligning Visual Contrastive learning models via Preference Optimization
by: Afzali, Amirabbas, et al.
Published: (2024)
by: Afzali, Amirabbas, et al.
Published: (2024)
Controlling Gender Bias in Retrieval via a Backpack Architecture
by: Afzali, Amirabbas, et al.
Published: (2025)
by: Afzali, Amirabbas, et al.
Published: (2025)
Geometry-Aware Decoding with Wasserstein-Regularized Truncation and Mass Penalties for Large Language Models
by: Davoodi, Arash Gholami, et al.
Published: (2026)
by: Davoodi, Arash Gholami, et al.
Published: (2026)
Clustering Time Series Data with Gaussian Mixture Embeddings in a Graph Autoencoder Framework
by: Afzali, Amirabbas, et al.
Published: (2024)
by: Afzali, Amirabbas, et al.
Published: (2024)
On-Policy Distillation of Language Models for Autonomous Vehicle Motion Planning
by: Afsharrad, Amirhossein, et al.
Published: (2026)
by: Afsharrad, Amirhossein, et al.
Published: (2026)
Preference Optimization with Multi-Sample Comparisons
by: Wang, Chaoqi, et al.
Published: (2024)
by: Wang, Chaoqi, et al.
Published: (2024)
Group Robust Preference Optimization in Reward-free RLHF
by: Ramesh, Shyam Sundhar, et al.
Published: (2024)
by: Ramesh, Shyam Sundhar, et al.
Published: (2024)
Robust Preference Optimization through Reward Model Distillation
by: Fisch, Adam, et al.
Published: (2024)
by: Fisch, Adam, et al.
Published: (2024)
RLHF Can Speak Many Languages: Unlocking Multilingual Preference Optimization for LLMs
by: Dang, John, et al.
Published: (2024)
by: Dang, John, et al.
Published: (2024)
Robust Multi-Objective Preference Alignment with Online DPO
by: Gupta, Raghav, et al.
Published: (2025)
by: Gupta, Raghav, et al.
Published: (2025)
Flextron: Many-in-One Flexible Large Language Model
by: Cai, Ruisi, et al.
Published: (2024)
by: Cai, Ruisi, et al.
Published: (2024)
Direct Multi-Turn Preference Optimization for Language Agents
by: Shi, Wentao, et al.
Published: (2024)
by: Shi, Wentao, et al.
Published: (2024)
Multi-Reference Preference Optimization for Large Language Models
by: Le, Hung, et al.
Published: (2024)
by: Le, Hung, et al.
Published: (2024)
Multi-Response Preference Optimization with Augmented Ranking Dataset
by: Gwon, Hansle, et al.
Published: (2024)
by: Gwon, Hansle, et al.
Published: (2024)
How Many Human Judgments Are Enough? Feasibility Limits of Human Preference Evaluation
by: Lee, Wilson Y.
Published: (2026)
by: Lee, Wilson Y.
Published: (2026)
ROPO: Robust Preference Optimization for Large Language Models
by: Liang, Xize, et al.
Published: (2024)
by: Liang, Xize, et al.
Published: (2024)
AMPO: Active Multi-Preference Optimization for Self-play Preference Selection
by: Gupta, Taneesh, et al.
Published: (2025)
by: Gupta, Taneesh, et al.
Published: (2025)
Out of One, Many: Using Language Models to Simulate Human Samples
by: Argyle, Lisa P., et al.
Published: (2022)
by: Argyle, Lisa P., et al.
Published: (2022)
The Impact of Quantization on the Robustness of Transformer-based Text Classifiers
by: Neshaei, Seyed Parsa, et al.
Published: (2024)
by: Neshaei, Seyed Parsa, et al.
Published: (2024)
Adversarial Training of Two-Layer Polynomial and ReLU Activation Networks via Convex Optimization
by: Kuelbs, Daniel, et al.
Published: (2024)
by: Kuelbs, Daniel, et al.
Published: (2024)
LLMs Are Not Intelligent Thinkers: Introducing Mathematical Topic Tree Benchmark for Comprehensive Evaluation of LLMs
by: Davoodi, Arash Gholami, et al.
Published: (2024)
by: Davoodi, Arash Gholami, et al.
Published: (2024)
Many of Your DPOs are Secretly One: Attempting Unification Through Mutual Information
by: Tutnov, Rasul, et al.
Published: (2025)
by: Tutnov, Rasul, et al.
Published: (2025)
Many Minds from One Model: Bayesian-Inspired Transformers for Population Diversity
by: Yang, Diji, et al.
Published: (2025)
by: Yang, Diji, et al.
Published: (2025)
VERI-DPO: Evidence-Aware Alignment for Clinical Summarization via Claim Verification and Direct Preference Optimization
by: Liu, Weixin, et al.
Published: (2026)
by: Liu, Weixin, et al.
Published: (2026)
Speculative Decoding Scaling Laws (SDSL): Throughput Optimization Made Simple
by: Bozorgkhoo, Amirhossein, et al.
Published: (2026)
by: Bozorgkhoo, Amirhossein, et al.
Published: (2026)
Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm
by: Shashidhar, Sarvesh, et al.
Published: (2025)
by: Shashidhar, Sarvesh, et al.
Published: (2025)
On the Noise Robustness of In-Context Learning for Text Generation
by: Gao, Hongfu, et al.
Published: (2024)
by: Gao, Hongfu, et al.
Published: (2024)
RoseRAG: Robust Retrieval-augmented Generation with Small-scale LLMs via Margin-aware Preference Optimization
by: Liu, Tianci, et al.
Published: (2025)
by: Liu, Tianci, et al.
Published: (2025)
On the Role of Preference Variance in Preference Optimization
by: Guo, Jiacheng, et al.
Published: (2025)
by: Guo, Jiacheng, et al.
Published: (2025)
PORT: Preference Optimization on Reasoning Traces
by: Lahlou, Salem, et al.
Published: (2024)
by: Lahlou, Salem, et al.
Published: (2024)
Length Desensitization in Direct Preference Optimization
by: Liu, Wei, et al.
Published: (2024)
by: Liu, Wei, et al.
Published: (2024)
Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference Optimization
by: Wu, Junkang, et al.
Published: (2024)
by: Wu, Junkang, et al.
Published: (2024)
ULTra: Unveiling Latent Token Interpretability in Transformer-Based Understanding and Segmentation
by: Hosseini, Hesam, et al.
Published: (2024)
by: Hosseini, Hesam, et al.
Published: (2024)
Similar Items
-
LORE: Lagrangian-Optimized Robust Embeddings for Visual Encoders
by: Khodabandeh, Borna, et al.
Published: (2025) -
Multi-Agent Stage-wise Conservative Linear Bandits
by: Afsharrad, Amirhossein, et al.
Published: (2025) -
Cooperative Multi-Agent Constrained Stochastic Linear Bandits
by: Afsharrad, Amirhossein, et al.
Published: (2024) -
Beyond Binary Preferences: A Principled Framework for Reward Modeling with Ordinal Feedback
by: Afsharrad, Amirhossein, et al.
Published: (2026) -
Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems
by: Allbert, Rumi, et al.
Published: (2025)