Test-Time Scaling in Diffusion LLMs via Hidden Semi-Autoregressive Experts
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Jihoon, Moon, Hoyeon, Zhai, Kevin, Chithanar, Arun Kumar, Sahu, Anit Kumar, Kar, Soummya, Lee, Chul, Chakraborty, Souradip, Bedi, Amrit Singh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MIRA: Towards Mitigating Reward Hacking in Inference-Time Alignment of T2I Diffusion Models
by: Zhai, Kevin, et al.
Published: (2025)
by: Zhai, Kevin, et al.
Published: (2025)
Towards Realistic Mechanisms That Incentivize Federated Participation and Contribution
by: Bornstein, Marco, et al.
Published: (2023)
by: Bornstein, Marco, et al.
Published: (2023)
RL with Learnable Textual Feedback: A Bilevel Approach
by: Singh, Utsav, et al.
Published: (2026)
by: Singh, Utsav, et al.
Published: (2026)
Beyond Text: Utilizing Vocal Cues to Improve Decision Making in LLMs for Robot Navigation Tasks
by: Sun, Xingpeng, et al.
Published: (2024)
by: Sun, Xingpeng, et al.
Published: (2024)
Why Pass@k Optimization Can Degrade Pass@1: Prompt Interference in LLM Post-training
by: Barakat, Anas, et al.
Published: (2026)
by: Barakat, Anas, et al.
Published: (2026)
Direct Preference Optimization for Primitive-Enabled Hierarchical RL: A Bilevel Approach
by: Singh, Utsav, et al.
Published: (2024)
by: Singh, Utsav, et al.
Published: (2024)
LIAR: Leveraging Inference Time Alignment (Best-of-N) to Jailbreak LLMs in Seconds
by: Beetham, James, et al.
Published: (2024)
by: Beetham, James, et al.
Published: (2024)
On the Global Optimality of Policy Gradient Methods in General Utility Reinforcement Learning
by: Barakat, Anas, et al.
Published: (2024)
by: Barakat, Anas, et al.
Published: (2024)
Code Comprehension then Auditing for Unsupervised LLM Evaluation
by: Patel, Bhrij, et al.
Published: (2024)
by: Patel, Bhrij, et al.
Published: (2024)
REBEL: Reward Regularization-Based Approach for Robotic Reinforcement Learning from Human Feedback
by: Chakraborty, Souradip, et al.
Published: (2023)
by: Chakraborty, Souradip, et al.
Published: (2023)
Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time
by: Chehade, Mohamad, et al.
Published: (2025)
by: Chehade, Mohamad, et al.
Published: (2025)
BalancedDPO: Adaptive Multi-Metric Alignment
by: Tamboli, Dipesh, et al.
Published: (2025)
by: Tamboli, Dipesh, et al.
Published: (2025)
Uncertainty-Aware Answer Selection for Improved Reasoning in Multi-LLM Systems
by: Agrawal, Aakriti, et al.
Published: (2025)
by: Agrawal, Aakriti, et al.
Published: (2025)
PROPS: Progressively Private Self-alignment of Large Language Models
by: Teku, Noel, et al.
Published: (2025)
by: Teku, Noel, et al.
Published: (2025)
Safety Recovery in Reasoning Models Is Only a Few Early Steering Steps Away
by: Ghosal, Soumya Suvra, et al.
Published: (2026)
by: Ghosal, Soumya Suvra, et al.
Published: (2026)
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment
by: Trivedi, Prashant, et al.
Published: (2025)
by: Trivedi, Prashant, et al.
Published: (2025)
VARP: Reinforcement Learning from Vision-Language Model Feedback with Agent Regularized Preferences
by: Singh, Anukriti, et al.
Published: (2025)
by: Singh, Anukriti, et al.
Published: (2025)
DIPPER: Direct Preference Optimization to Accelerate Primitive-Enabled Hierarchical Reinforcement Learning
by: Singh, Utsav, et al.
Published: (2024)
by: Singh, Utsav, et al.
Published: (2024)
PARL: A Unified Framework for Policy Alignment in Reinforcement Learning from Human Feedback
by: Chakraborty, Souradip, et al.
Published: (2023)
by: Chakraborty, Souradip, et al.
Published: (2023)
Does Thinking More always Help? Mirage of Test-Time Scaling in Reasoning Models
by: Ghosal, Soumya Suvra, et al.
Published: (2025)
by: Ghosal, Soumya Suvra, et al.
Published: (2025)
Transfer Q Star: Principled Decoding for LLM Alignment
by: Chakraborty, Souradip, et al.
Published: (2024)
by: Chakraborty, Souradip, et al.
Published: (2024)
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
AI Cap-and-Trade: Efficiency Incentives for Accessibility and Sustainability
by: Bornstein, Marco, et al.
Published: (2026)
by: Bornstein, Marco, et al.
Published: (2026)
MaxMin-RLHF: Alignment with Diverse Human Preferences
by: Chakraborty, Souradip, et al.
Published: (2024)
by: Chakraborty, Souradip, et al.
Published: (2024)
On the Role of Feedback in Test-Time Scaling of Agentic AI Workflows
by: Chakraborty, Souradip, et al.
Published: (2025)
by: Chakraborty, Souradip, et al.
Published: (2025)
Draft-Conditioned Constrained Decoding for Structured Generation in LLMs
by: Reddy, Avinash, et al.
Published: (2026)
by: Reddy, Avinash, et al.
Published: (2026)
Stability Analysis of a B-Spline Deep Neural Operator for Nonlinear Systems
by: Romagnoli, Raffaele, et al.
Published: (2025)
by: Romagnoli, Raffaele, et al.
Published: (2025)
SAIL: Self-Improving Efficient Online Alignment of Large Language Models
by: Ding, Mucong, et al.
Published: (2024)
by: Ding, Mucong, et al.
Published: (2024)
On the Vulnerability of LLM/VLM-Controlled Robotics
by: Wu, Xiyang, et al.
Published: (2024)
by: Wu, Xiyang, et al.
Published: (2024)
Achieving Zero Constraint Violation for Constrained Reinforcement Learning via Conservative Natural Policy Gradient Primal-Dual Algorithm
by: Bai, Qinbo, et al.
Published: (2022)
by: Bai, Qinbo, et al.
Published: (2022)
Computational Imaging for Long-Term Prediction of Solar Irradiance
by: Julian, Leron, et al.
Published: (2024)
by: Julian, Leron, et al.
Published: (2024)
TRAM: Test-Time Risk Adaptation with Mixture of Agents
by: Chehade, Mohamad Fares El Hajj, et al.
Published: (2024)
by: Chehade, Mohamad Fares El Hajj, et al.
Published: (2024)
DMCA: Dense Multi-agent Navigation using Attention and Communication
by: Arul, Senthil Hariharan, et al.
Published: (2022)
by: Arul, Senthil Hariharan, et al.
Published: (2022)
Sample-Optimal Zero-Violation Safety For Continuous Control
by: Ray, Ritabrata, et al.
Published: (2024)
by: Ray, Ritabrata, et al.
Published: (2024)
Distributed Truncated Predictive Control for Networked Systems under Uncertainty: Stability and Near-Optimality Guarantee
by: Xu, Eric, et al.
Published: (2023)
by: Xu, Eric, et al.
Published: (2023)
Smoothed Gradient Clipping and Error Feedback for Decentralized Optimization under Symmetric Heavy-Tailed Noise
by: Yu, Shuhua, et al.
Published: (2023)
by: Yu, Shuhua, et al.
Published: (2023)
Model-Free Learning and Optimal Policy Design in Multi-Agent MDPs Under Probabilistic Agent Dropout
by: Fiscko, Carmel, et al.
Published: (2023)
by: Fiscko, Carmel, et al.
Published: (2023)
Decentralized Nonconvex Optimization under Heavy-Tailed Noise: Normalization and Optimal Convergence
by: Yu, Shuhua, et al.
Published: (2025)
by: Yu, Shuhua, et al.
Published: (2025)
Learning When to Trust Which Teacher for Weakly Supervised ASR
by: Agrawal, Aakriti, et al.
Published: (2023)
by: Agrawal, Aakriti, et al.
Published: (2023)
The effectiveness of RANS ‐based turbulence models predicting the PCM ‐connected air conditioning unit performance
by: Arun Kumar Sao, et al.
Published: (2026)
by: Arun Kumar Sao, et al.
Published: (2026)
Similar Items
-
MIRA: Towards Mitigating Reward Hacking in Inference-Time Alignment of T2I Diffusion Models
by: Zhai, Kevin, et al.
Published: (2025) -
Towards Realistic Mechanisms That Incentivize Federated Participation and Contribution
by: Bornstein, Marco, et al.
Published: (2023) -
RL with Learnable Textual Feedback: A Bilevel Approach
by: Singh, Utsav, et al.
Published: (2026) -
Beyond Text: Utilizing Vocal Cues to Improve Decision Making in LLMs for Robot Navigation Tasks
by: Sun, Xingpeng, et al.
Published: (2024) -
Why Pass@k Optimization Can Degrade Pass@1: Prompt Interference in LLM Post-training
by: Barakat, Anas, et al.
Published: (2026)