Entropy Aware Reward Guidance for Diffusion Language Model Alignment
Fuente:
arXiv
Salvato in:
| Autori principali: | Tejaswi, Atula, Rout, Litu, Caramanis, Constantine, Shakkottai, Sanjay, Sanghavi, Sujay |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Anchored Diffusion Language Model
di: Rout, Litu, et al.
Pubblicazione: (2025)
di: Rout, Litu, et al.
Pubblicazione: (2025)
AnCoder: Anchored Code Generation via Discrete Diffusion Models
di: Xue, Anton, et al.
Pubblicazione: (2026)
di: Xue, Anton, et al.
Pubblicazione: (2026)
In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
di: Collins, Liam, et al.
Pubblicazione: (2024)
di: Collins, Liam, et al.
Pubblicazione: (2024)
RARe: Retrieval Augmented Retrieval with In-Context Examples
di: Tejaswi, Atula, et al.
Pubblicazione: (2024)
di: Tejaswi, Atula, et al.
Pubblicazione: (2024)
Exploring Design Choices for Building Language-Specific LLMs
di: Tejaswi, Atula, et al.
Pubblicazione: (2024)
di: Tejaswi, Atula, et al.
Pubblicazione: (2024)
Asymptotically-Optimal Gaussian Bandits with Side Observations
di: Atsidakou, Alexia, et al.
Pubblicazione: (2025)
di: Atsidakou, Alexia, et al.
Pubblicazione: (2025)
Efficient Approximate Posterior Sampling with Annealed Langevin Monte Carlo
di: Parulekar, Advait, et al.
Pubblicazione: (2025)
di: Parulekar, Advait, et al.
Pubblicazione: (2025)
SVFT: Parameter-Efficient Fine-Tuning with Singular Vectors
di: Lingam, Vijay, et al.
Pubblicazione: (2024)
di: Lingam, Vijay, et al.
Pubblicazione: (2024)
RB-Modulation: Training-Free Personalization of Diffusion Models using Stochastic Optimal Control
di: Rout, Litu, et al.
Pubblicazione: (2024)
di: Rout, Litu, et al.
Pubblicazione: (2024)
Semantic Image Inversion and Editing using Rectified Stochastic Differential Equations
di: Rout, Litu, et al.
Pubblicazione: (2024)
di: Rout, Litu, et al.
Pubblicazione: (2024)
Test-Time Anchoring for Discrete Diffusion Posterior Sampling
di: Rout, Litu, et al.
Pubblicazione: (2025)
di: Rout, Litu, et al.
Pubblicazione: (2025)
HiSpec: Hierarchical Speculative Decoding for LLMs
di: Kumar, Avinash, et al.
Pubblicazione: (2025)
di: Kumar, Avinash, et al.
Pubblicazione: (2025)
When Attention Collapses: How Degenerate Layers in LLMs Enable Smaller, Stronger Models
di: Sanyal, Sunny, et al.
Pubblicazione: (2024)
di: Sanyal, Sunny, et al.
Pubblicazione: (2024)
Constrained Posterior Sampling: Time Series Generation with Hard Constraints
di: Narasimhan, Sai Shankar, et al.
Pubblicazione: (2024)
di: Narasimhan, Sai Shankar, et al.
Pubblicazione: (2024)
On the Robustness of Reward Models for Language Model Alignment
di: Hong, Jiwoo, et al.
Pubblicazione: (2025)
di: Hong, Jiwoo, et al.
Pubblicazione: (2025)
Enabling Approximate Joint Sampling in Diffusion LMs
di: Bansal, Parikshit, et al.
Pubblicazione: (2025)
di: Bansal, Parikshit, et al.
Pubblicazione: (2025)
M2Lingual: Enhancing Multilingual, Multi-Turn Instruction Alignment in Large Language Models
di: Maheshwary, Rishabh, et al.
Pubblicazione: (2024)
di: Maheshwary, Rishabh, et al.
Pubblicazione: (2024)
PITA: Preference-Guided Inference-Time Alignment for LLM Post-Training
di: Bobbili, Sarat Chandra, et al.
Pubblicazione: (2025)
di: Bobbili, Sarat Chandra, et al.
Pubblicazione: (2025)
Rethinking the Role of Proxy Rewards in Language Model Alignment
di: Kim, Sungdong, et al.
Pubblicazione: (2024)
di: Kim, Sungdong, et al.
Pubblicazione: (2024)
Uncertainty-Aware Gradient Signal-to-Noise Data Selection for Instruction Tuning
di: Yuan, Zhihang, et al.
Pubblicazione: (2026)
di: Yuan, Zhihang, et al.
Pubblicazione: (2026)
Reviving The Classics: Active Reward Modeling in Large Language Model Alignment
di: Shen, Yunyi, et al.
Pubblicazione: (2025)
di: Shen, Yunyi, et al.
Pubblicazione: (2025)
TABES: Trajectory-Aware Backward-on-Entropy Steering for Masked Diffusion Models
di: Saini, Shreshth, et al.
Pubblicazione: (2026)
di: Saini, Shreshth, et al.
Pubblicazione: (2026)
SALMON: Self-Alignment with Instructable Reward Models
di: Sun, Zhiqing, et al.
Pubblicazione: (2023)
di: Sun, Zhiqing, et al.
Pubblicazione: (2023)
RewardFlow: Topology-Aware Reward Propagation on State Graphs for Agentic RL with Large Language Models
di: Feng, Xiao, et al.
Pubblicazione: (2026)
di: Feng, Xiao, et al.
Pubblicazione: (2026)
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs
di: Sun, Hao, et al.
Pubblicazione: (2025)
di: Sun, Hao, et al.
Pubblicazione: (2025)
Curry-DPO: Enhancing Alignment using Curriculum Learning & Ranked Preferences
di: Pattnaik, Pulkit, et al.
Pubblicazione: (2024)
di: Pattnaik, Pulkit, et al.
Pubblicazione: (2024)
RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models
di: Yang, Daniel, et al.
Pubblicazione: (2026)
di: Yang, Daniel, et al.
Pubblicazione: (2026)
Succeeding at Scale: Automated Dataset Construction and Query-Side Adaptation for Multi-Tenant Search
di: Jain, Prateek, et al.
Pubblicazione: (2026)
di: Jain, Prateek, et al.
Pubblicazione: (2026)
Sink-Aware Pruning for Diffusion Language Models
di: Myrzakhan, Aidar, et al.
Pubblicazione: (2026)
di: Myrzakhan, Aidar, et al.
Pubblicazione: (2026)
More Bang for the Buck: Process Reward Modeling with Entropy-Driven Uncertainty
di: Cao, Lang, et al.
Pubblicazione: (2025)
di: Cao, Lang, et al.
Pubblicazione: (2025)
Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
di: Rafailov, Rafael, et al.
Pubblicazione: (2024)
di: Rafailov, Rafael, et al.
Pubblicazione: (2024)
Reward-free Alignment for Conflicting Objectives
di: Chen, Peter, et al.
Pubblicazione: (2026)
di: Chen, Peter, et al.
Pubblicazione: (2026)
ARGS: Alignment as Reward-Guided Search
di: Khanov, Maxim, et al.
Pubblicazione: (2024)
di: Khanov, Maxim, et al.
Pubblicazione: (2024)
Adaptive Guidance for Retrieval-Augmented Masked Diffusion Models
di: Kim, Jaemin, et al.
Pubblicazione: (2026)
di: Kim, Jaemin, et al.
Pubblicazione: (2026)
Entropy Centroids as Intrinsic Rewards for Test-Time Scaling
di: Zhao, Wenshuo, et al.
Pubblicazione: (2026)
di: Zhao, Wenshuo, et al.
Pubblicazione: (2026)
Finite-Time Logarithmic Bayes Regret Upper Bounds
di: Atsidakou, Alexia, et al.
Pubblicazione: (2023)
di: Atsidakou, Alexia, et al.
Pubblicazione: (2023)
ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models
di: Baek, Jinheon, et al.
Pubblicazione: (2024)
di: Baek, Jinheon, et al.
Pubblicazione: (2024)
Adaptive Segment-level Reward: Bridging the Gap Between Action and Reward Space in Alignment
di: Li, Yanshi, et al.
Pubblicazione: (2024)
di: Li, Yanshi, et al.
Pubblicazione: (2024)
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment
di: Zhang, Jiazheng, et al.
Pubblicazione: (2025)
di: Zhang, Jiazheng, et al.
Pubblicazione: (2025)
Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment
di: Yang, Rui, et al.
Pubblicazione: (2024)
di: Yang, Rui, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Anchored Diffusion Language Model
di: Rout, Litu, et al.
Pubblicazione: (2025) -
AnCoder: Anchored Code Generation via Discrete Diffusion Models
di: Xue, Anton, et al.
Pubblicazione: (2026) -
In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
di: Collins, Liam, et al.
Pubblicazione: (2024) -
RARe: Retrieval Augmented Retrieval with In-Context Examples
di: Tejaswi, Atula, et al.
Pubblicazione: (2024) -
Exploring Design Choices for Building Language-Specific LLMs
di: Tejaswi, Atula, et al.
Pubblicazione: (2024)