Latent Reward Steering: An Adaptive Inference-Time Framework that Implicitly Promotes Cognitive Behaviors in Reasoning LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Jiakang, Zhu, Guanyu, Jin, Can, Huang, Chenxi, Yu, Dexu, Chen, Ronghao, Zhou, Yang, Peng, Hongwu, Lan, Xuanqi, Metaxas, Dimitris N., Li, Youhua |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight
by: Jin, Can, et al.
Published: (2026)
by: Jin, Can, et al.
Published: (2026)
Two Heads are Better Than One: Test-time Scaling of Multi-agent Collaborative Reasoning
by: Jin, Can, et al.
Published: (2025)
by: Jin, Can, et al.
Published: (2025)
Test-Time Spectrum-Aware Latent Steering for Zero-Shot Generalization in Vision-Language Models
by: Dafnis, Konstantinos M., et al.
Published: (2025)
by: Dafnis, Konstantinos M., et al.
Published: (2025)
Learning from Teaching Regularization: Generalizable Correlations Should be Easy to Imitate
by: Jin, Can, et al.
Published: (2024)
by: Jin, Can, et al.
Published: (2024)
GSS: Gated Subspace Steering for Selective Memorization Mitigation in LLMs
by: Zhang, Xuanqi, et al.
Published: (2026)
by: Zhang, Xuanqi, et al.
Published: (2026)
Chain of Mindset: Reasoning with Adaptive Cognitive Modes
by: Jiang, Tianyi, et al.
Published: (2026)
by: Jiang, Tianyi, et al.
Published: (2026)
Thinking in Latents: Adaptive Anchor Refinement for Implicit Reasoning in LLMs
by: Sheshanarayana, Disha, et al.
Published: (2026)
by: Sheshanarayana, Disha, et al.
Published: (2026)
Your Reward Function for RL is Your Best PRM for Search: Unifying RL and Search-Based TTS
by: Jin, Can, et al.
Published: (2025)
by: Jin, Can, et al.
Published: (2025)
Beyond Explicit Edges: Robust Reasoning over Noisy and Sparse Knowledge Graphs
by: Gao, Hang, et al.
Published: (2026)
by: Gao, Hang, et al.
Published: (2026)
On the Role of Language Representations in Auto-Bidding: Findings and Implications
by: Zhu, Guanyu, et al.
Published: (2026)
by: Zhu, Guanyu, et al.
Published: (2026)
Steering Rectified Flow Models in the Vector Field for Controlled Image Generation
by: Patel, Maitreya, et al.
Published: (2024)
by: Patel, Maitreya, et al.
Published: (2024)
RISER: Orchestrating Latent Reasoning Skills for Adaptive Activation Steering
by: Ye, Wencheng, et al.
Published: (2026)
by: Ye, Wencheng, et al.
Published: (2026)
ATLAS: Adaptive Test-Time Latent Steering with External Verifiers for Enhancing LLMs Reasoning
by: Nguyen, Tuc, et al.
Published: (2026)
by: Nguyen, Tuc, et al.
Published: (2026)
Beyond Interpretability: When, Why, and How Sparse Autoencoders Enable Label-Free Visual Steering
by: Chatzoudis, Gerasimos, et al.
Published: (2025)
by: Chatzoudis, Gerasimos, et al.
Published: (2025)
Distribution-Aware Reward Estimation for Test-Time Reinforcement Learning
by: Du, Bodong, et al.
Published: (2026)
by: Du, Bodong, et al.
Published: (2026)
MPDiT: Multi-Patch Global-to-Local Transformer Architecture For Efficient Flow Matching and Diffusion Model
by: Dao, Quan, et al.
Published: (2026)
by: Dao, Quan, et al.
Published: (2026)
Latent Implicit Visual Reasoning
by: Li, Kelvin, et al.
Published: (2025)
by: Li, Kelvin, et al.
Published: (2025)
Implicit In-context Learning
by: Li, Zhuowei, et al.
Published: (2024)
by: Li, Zhuowei, et al.
Published: (2024)
Improved Training Technique for Latent Consistency Models
by: Dao, Quan, et al.
Published: (2025)
by: Dao, Quan, et al.
Published: (2025)
Understanding and Steering the Cognitive Behaviors of Reasoning Models at Test-Time
by: Zhang, Zhenyu, et al.
Published: (2025)
by: Zhang, Zhenyu, et al.
Published: (2025)
Reasoning over Precedents Alongside Statutes: Case-Augmented Deliberative Alignment for LLM Safety
by: Jin, Can, et al.
Published: (2026)
by: Jin, Can, et al.
Published: (2026)
CogDual: Enhancing Dual Cognition of LLMs via Reinforcement Learning with Implicit Rule-Based Rewards
by: Liu, Cheng, et al.
Published: (2025)
by: Liu, Cheng, et al.
Published: (2025)
Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
by: Wen, Xumeng, et al.
Published: (2025)
by: Wen, Xumeng, et al.
Published: (2025)
RankFlow: A Multi-Role Collaborative Reranking Workflow Utilizing Large Language Models
by: Jin, Can, et al.
Published: (2025)
by: Jin, Can, et al.
Published: (2025)
DTop-p MoE: Sparsity-Controlled Dynamic Top-p MoE for Foundation Model Pre-training
by: Jin, Can, et al.
Published: (2025)
by: Jin, Can, et al.
Published: (2025)
APEER: Automatic Prompt Engineering Enhances Large Language Model Reranking
by: Jin, Can, et al.
Published: (2024)
by: Jin, Can, et al.
Published: (2024)
Mitigating Cognitive Inertia in Large Reasoning Models via Latent Spike Steering
by: Lee, Seojin, et al.
Published: (2026)
by: Lee, Seojin, et al.
Published: (2026)
Token-Controlled Re-ranking for Sequential Recommendation via LLMs
by: Dai, Wenxi, et al.
Published: (2025)
by: Dai, Wenxi, et al.
Published: (2025)
SteerConf: Steering LLMs for Confidence Elicitation
by: Zhou, Ziang, et al.
Published: (2025)
by: Zhou, Ziang, et al.
Published: (2025)
Fractional Reasoning via Latent Steering Vectors Improves Inference Time Compute
by: Liu, Sheng, et al.
Published: (2025)
by: Liu, Sheng, et al.
Published: (2025)
FairSteer: Inference Time Debiasing for LLMs with Dynamic Activation Steering
by: Li, Yichen, et al.
Published: (2025)
by: Li, Yichen, et al.
Published: (2025)
DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation
by: Zhou, Yang, et al.
Published: (2026)
by: Zhou, Yang, et al.
Published: (2026)
iCLP: Large Language Model Reasoning with Implicit Cognition Latent Planning
by: Chen, Sijia, et al.
Published: (2025)
by: Chen, Sijia, et al.
Published: (2025)
Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR
by: Wang, Jiakang, et al.
Published: (2025)
by: Wang, Jiakang, et al.
Published: (2025)
LED: LLM Enhanced Open-Vocabulary Object Detection without Human Curated Data Generation
by: Zhou, Yang, et al.
Published: (2025)
by: Zhou, Yang, et al.
Published: (2025)
Exploring the Personality Traits of LLMs through Latent Features Steering
by: Yang, Shu, et al.
Published: (2024)
by: Yang, Shu, et al.
Published: (2024)
Cognitive Aesthetic Evaluation Model - Implementation Code
by: Chenxi, Jin
Published: (2026)
by: Chenxi, Jin
Published: (2026)
Improving Visual Reasoning with Iterative Evidence Refinement
by: Shi, Zeru, et al.
Published: (2026)
by: Shi, Zeru, et al.
Published: (2026)
Score-Guided Diffusion for 3D Human Recovery
by: Stathopoulos, Anastasis, et al.
Published: (2024)
by: Stathopoulos, Anastasis, et al.
Published: (2024)
Adaptive Fuzzy‐Based Event‐Triggered Consensus of Switched Nonlinear Multiagent Systems With Communication Faults and State‐Dependent Switchings
by: Ronghao Zhang, et al.
Published: (2025)
by: Ronghao Zhang, et al.
Published: (2025)
Similar Items
-
Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight
by: Jin, Can, et al.
Published: (2026) -
Two Heads are Better Than One: Test-time Scaling of Multi-agent Collaborative Reasoning
by: Jin, Can, et al.
Published: (2025) -
Test-Time Spectrum-Aware Latent Steering for Zero-Shot Generalization in Vision-Language Models
by: Dafnis, Konstantinos M., et al.
Published: (2025) -
Learning from Teaching Regularization: Generalizable Correlations Should be Easy to Imitate
by: Jin, Can, et al.
Published: (2024) -
GSS: Gated Subspace Steering for Selective Memorization Mitigation in LLMs
by: Zhang, Xuanqi, et al.
Published: (2026)