Alignment-Aware Model Adaptation via Feedback-Guided Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Bhatt, Gaurav, Chinchure, Aditya, Zhou, Jiawei, Sigal, Leonid |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Preventing Catastrophic Forgetting through Memory Networks in Continuous Detection
by: Bhatt, Gaurav, et al.
Published: (2024)
by: Bhatt, Gaurav, et al.
Published: (2024)
RewardRank: Optimizing True Learning-to-Rank Utility
by: Bhatt, Gaurav, et al.
Published: (2025)
by: Bhatt, Gaurav, et al.
Published: (2025)
TIBET: Identifying and Evaluating Biases in Text-to-Image Generative Models
by: Chinchure, Aditya, et al.
Published: (2023)
by: Chinchure, Aditya, et al.
Published: (2023)
BiasConnect: Investigating Bias Interactions in Text-to-Image Models
by: Shukla, Pushkar, et al.
Published: (2025)
by: Shukla, Pushkar, et al.
Published: (2025)
SPIKE-RL: Video-LLMs meet Bayesian Surprise
by: Ravi, Sahithya, et al.
Published: (2025)
by: Ravi, Sahithya, et al.
Published: (2025)
Spotlight: Identifying and Localizing Video Generation Errors Using VLMs
by: Chinchure, Aditya, et al.
Published: (2025)
by: Chinchure, Aditya, et al.
Published: (2025)
ORCE: Order-Aware Alignment of Verbalized Confidence in Large Language Models
by: Li, Chen, et al.
Published: (2026)
by: Li, Chen, et al.
Published: (2026)
Towards Reliable, Uncertainty-Aware Alignment
by: Banerjee, Debangshu, et al.
Published: (2025)
by: Banerjee, Debangshu, et al.
Published: (2025)
Property-Guided Molecular Generation and Optimization via Latent Flows
by: Lobo, Alexander Arjun, et al.
Published: (2026)
by: Lobo, Alexander Arjun, et al.
Published: (2026)
All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding
by: Rahman, Tanzila, et al.
Published: (2026)
by: Rahman, Tanzila, et al.
Published: (2026)
Improving the Throughput of Diffusion-based Large Language Models via a Training-Free Confidence-Aware Calibration
by: Shen, Jucheng, et al.
Published: (2025)
by: Shen, Jucheng, et al.
Published: (2025)
ADiff4TPP: Asynchronous Diffusion Models for Temporal Point Processes
by: Mukherjee, Amartya, et al.
Published: (2025)
by: Mukherjee, Amartya, et al.
Published: (2025)
Efficient Algorithms for Logistic Contextual Slate Bandits with Bandit Feedback
by: Goyal, Tanmay, et al.
Published: (2025)
by: Goyal, Tanmay, et al.
Published: (2025)
Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable Events
by: Chinchure, Aditya, et al.
Published: (2024)
by: Chinchure, Aditya, et al.
Published: (2024)
Comparing Bad Apples to Good Oranges: Aligning Large Language Models via Joint Preference Optimization
by: Bansal, Hritik, et al.
Published: (2024)
by: Bansal, Hritik, et al.
Published: (2024)
Group Preference Optimization: Few-Shot Alignment of Large Language Models
by: Zhao, Siyan, et al.
Published: (2023)
by: Zhao, Siyan, et al.
Published: (2023)
What Has Been Overlooked in Contrastive Source-Free Domain Adaptation: Leveraging Source-Informed Latent Augmentation within Neighborhood Context
by: Wang, Jing, et al.
Published: (2024)
by: Wang, Jing, et al.
Published: (2024)
Do LLMs Benefit from User and Item Embeddings in Recommendation Tasks?
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2026)
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2026)
Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities
by: Chandhok, Shivam, et al.
Published: (2025)
by: Chandhok, Shivam, et al.
Published: (2025)
Uncertainty-Guided Alignment for Unsupervised Domain Adaptation in Regression
by: Nejjar, Ismail, et al.
Published: (2024)
by: Nejjar, Ismail, et al.
Published: (2024)
Prototype-Guided Pseudo-Labeling with Neighborhood-Aware Consistency for Unsupervised Adaptation
by: Ali, Eman, et al.
Published: (2025)
by: Ali, Eman, et al.
Published: (2025)
IterInject: Indirect Prompt Injection Against LLM Agents via Feedback-Guided Iterative Optimization
by: Chen, Zixuan, et al.
Published: (2026)
by: Chen, Zixuan, et al.
Published: (2026)
Geometric Moment Alignment for Domain Adaptation via Siegel Embeddings
by: Gharib, Shayan, et al.
Published: (2025)
by: Gharib, Shayan, et al.
Published: (2025)
Learning ECG Image Representations via Dual Physiological-Aware Alignments
by: Pham, Hung Manh, et al.
Published: (2026)
by: Pham, Hung Manh, et al.
Published: (2026)
AI-Guided Design and Optimization of Graphite-Based Anodes via Iterative Experimental Feedback
by: Du, Qian, et al.
Published: (2026)
by: Du, Qian, et al.
Published: (2026)
Experience-Guided Adaptation of Inference-Time Reasoning Strategies
by: Stein, Adam, et al.
Published: (2025)
by: Stein, Adam, et al.
Published: (2025)
Reasoning Elicitation in Language Models via Counterfactual Feedback
by: Hüyük, Alihan, et al.
Published: (2024)
by: Hüyük, Alihan, et al.
Published: (2024)
Testing the Feasibility of Linear Programs with Bandit Feedback
by: Gangrade, Aditya, et al.
Published: (2024)
by: Gangrade, Aditya, et al.
Published: (2024)
Framework-agnostic Semantically-aware Global Reasoning for Segmentation
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2022)
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2022)
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks
by: Fan, Wan-Cyuan, et al.
Published: (2024)
by: Fan, Wan-Cyuan, et al.
Published: (2024)
Accelerated Predictive Coding Networks via Direct Kolen-Pollack Feedback Alignment
by: Casnici, Davide, et al.
Published: (2026)
by: Casnici, Davide, et al.
Published: (2026)
Revealing Weaknesses in Text Watermarking Through Self-Information Rewrite Attacks
by: Cheng, Yixin, et al.
Published: (2025)
by: Cheng, Yixin, et al.
Published: (2025)
Inpainting-Guided Policy Optimization for Diffusion Large Language Models
by: Zhao, Siyan, et al.
Published: (2025)
by: Zhao, Siyan, et al.
Published: (2025)
Smoothing the Score Function for Generalization in Diffusion Models: An Optimization-based Explanation Framework
by: Zhou, Xinyu, et al.
Published: (2026)
by: Zhou, Xinyu, et al.
Published: (2026)
Beyond Sharp Minima: Robust LLM Unlearning via Feedback-Guided Multi-Point Optimization
by: Wu, Wenhan, et al.
Published: (2025)
by: Wu, Wenhan, et al.
Published: (2025)
A Judge-Aware Ranking Framework for Evaluating Large Language Models without Ground Truth
by: Xu, Mingyuan, et al.
Published: (2026)
by: Xu, Mingyuan, et al.
Published: (2026)
Reflective Preference Optimization (RPO): Enhancing On-Policy Alignment via Hint-Guided Reflection
by: Zhao, Zihui, et al.
Published: (2025)
by: Zhao, Zihui, et al.
Published: (2025)
Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization
by: Peng, Xiyue, et al.
Published: (2024)
by: Peng, Xiyue, et al.
Published: (2024)
SAFE: Stable Alignment Finetuning with Entropy-Aware Predictive Control for Reinforcement Learning from Human Feedback (RLHF)
by: Maity, Dipan
Published: (2026)
by: Maity, Dipan
Published: (2026)
GUIDE: Guided Initialization and Distillation of Embeddings
by: Trinh, Khoa, et al.
Published: (2025)
by: Trinh, Khoa, et al.
Published: (2025)
Similar Items
-
Preventing Catastrophic Forgetting through Memory Networks in Continuous Detection
by: Bhatt, Gaurav, et al.
Published: (2024) -
RewardRank: Optimizing True Learning-to-Rank Utility
by: Bhatt, Gaurav, et al.
Published: (2025) -
TIBET: Identifying and Evaluating Biases in Text-to-Image Generative Models
by: Chinchure, Aditya, et al.
Published: (2023) -
BiasConnect: Investigating Bias Interactions in Text-to-Image Models
by: Shukla, Pushkar, et al.
Published: (2025) -
SPIKE-RL: Video-LLMs meet Bayesian Surprise
by: Ravi, Sahithya, et al.
Published: (2025)