Seek in the Dark: Reasoning via Test-Time Instance-Level Policy Gradient in Latent Space
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Hengli, Li, Chenxi, Wu, Tong, Zhu, Xuekai, Wang, Yuxuan, Yu, Zhaoxin, Jiang, Eric Hanchen, Zhu, Song-Chun, Jia, Zixia, Wu, Ying Nian, Zheng, Zilong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
by: Mi, Yapeng, et al.
Published: (2025)
by: Mi, Yapeng, et al.
Published: (2025)
Discrete Markov Bridge
by: Li, Hengli, et al.
Published: (2025)
by: Li, Hengli, et al.
Published: (2025)
BEDA: Belief Estimation as Probabilistic Constraints for Performing Strategic Dialogue Acts
by: Li, Hengli, et al.
Published: (2025)
by: Li, Hengli, et al.
Published: (2025)
TokenSwift: Lossless Acceleration of Ultra Long Sequence Generation
by: Wu, Tong, et al.
Published: (2025)
by: Wu, Tong, et al.
Published: (2025)
Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning
by: Wu, Tong, et al.
Published: (2025)
by: Wu, Tong, et al.
Published: (2025)
MindDial: Belief Dynamics Tracking with Theory-of-Mind Modeling for Situated Neural Dialogue Generation
by: Qiu, Shuwen, et al.
Published: (2023)
by: Qiu, Shuwen, et al.
Published: (2023)
LangSuitE: Planning, Controlling and Interacting with Large Language Models in Embodied Text Environments
by: Jia, Zixia, et al.
Published: (2024)
by: Jia, Zixia, et al.
Published: (2024)
$\nabla$-Reasoner: LLM Reasoning via Test-Time Gradient Descent in Latent Space
by: Wang, Peihao, et al.
Published: (2026)
by: Wang, Peihao, et al.
Published: (2026)
How to Synthesize Text Data without Model Collapse?
by: Zhu, Xuekai, et al.
Published: (2024)
by: Zhu, Xuekai, et al.
Published: (2024)
Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space
by: Chen, Chao, et al.
Published: (2025)
by: Chen, Chao, et al.
Published: (2025)
TongSearch-QR: Reinforced Query Reasoning for Retrieval
by: Qin, Xubo, et al.
Published: (2025)
by: Qin, Xubo, et al.
Published: (2025)
The Thinking Pixel: Recursive Sparse Reasoning in Multimodal Diffusion Latents
by: Sun, Yuwei, et al.
Published: (2026)
by: Sun, Yuwei, et al.
Published: (2026)
ManCAR: Manifold-Constrained Latent Reasoning with Adaptive Test-Time Computation for Sequential Recommendation
by: Yang, Kun, et al.
Published: (2026)
by: Yang, Kun, et al.
Published: (2026)
Multiple Latent Space Mapping for Compressed Dark Image Enhancement
by: Zeng, Yi, et al.
Published: (2024)
by: Zeng, Yi, et al.
Published: (2024)
Agent Q-Mix: Selecting the Right Action for LLM Multi-Agent Systems through Reinforcement Learning
by: Jiang, Eric Hanchen, et al.
Published: (2026)
by: Jiang, Eric Hanchen, et al.
Published: (2026)
MorphSeek: Fine-grained Latent Representation-Level Policy Optimization for Deformable Image Registration
by: Zhang, Runxun, et al.
Published: (2025)
by: Zhang, Runxun, et al.
Published: (2025)
Skew-adjoint linear relatioins between Banach spaces
by: Li, Hanchen, et al.
Published: (2024)
by: Li, Hanchen, et al.
Published: (2024)
Designing Instance-Level Sampling Schedules via REINFORCE with James-Stein Shrinkage
by: Yu, Peiyu, et al.
Published: (2025)
by: Yu, Peiyu, et al.
Published: (2025)
Understanding and Leveraging the Expert Specialization of Context Faithfulness in Mixture-of-Experts LLMs
by: Bai, Jun, et al.
Published: (2025)
by: Bai, Jun, et al.
Published: (2025)
Anomalous bulk-edge correspondence of nonlinear Rice-Mele model
by: Bai, Chenxi, et al.
Published: (2025)
by: Bai, Chenxi, et al.
Published: (2025)
Fractional Thouless pumping of solitons: a unique manifestation of bulk-edge correspondence of nonlinear eigenvalue problems
by: Bai, Chenxi, et al.
Published: (2025)
by: Bai, Chenxi, et al.
Published: (2025)
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
by: Zuo, Yuxin, et al.
Published: (2025)
by: Zuo, Yuxin, et al.
Published: (2025)
Statistical Inference on Latent Space Models for Network Data
by: Li, Jinming, et al.
Published: (2023)
by: Li, Jinming, et al.
Published: (2023)
Learning to Rank Chain-of-Thought: Using a Small Model
by: Jiang, Eric Hanchen, et al.
Published: (2025)
by: Jiang, Eric Hanchen, et al.
Published: (2025)
The AI Hippocampus: How Far are We From Human Memory?
by: Jia, Zixia, et al.
Published: (2026)
by: Jia, Zixia, et al.
Published: (2026)
AsynFusion: Towards Asynchronous Latent Consistency Models for Decoupled Whole-Body Audio-Driven Avatars
by: Zhang, Tianbao, et al.
Published: (2025)
by: Zhang, Tianbao, et al.
Published: (2025)
Policy Gradient Guidance Enables Test Time Control
by: Qi, Jianing, et al.
Published: (2025)
by: Qi, Jianing, et al.
Published: (2025)
Q-Tacit: Image Quality Assessment via Latent Visual Reasoning
by: Jiang, Yuxuan, et al.
Published: (2026)
by: Jiang, Yuxuan, et al.
Published: (2026)
Latent Space Energy-based Neural ODEs
by: Cheng, Sheng, et al.
Published: (2024)
by: Cheng, Sheng, et al.
Published: (2024)
Combining Supervised Learning and Reinforcement Learning for Multi-Label Classification Tasks with Partial Labels
by: Jia, Zixia, et al.
Published: (2024)
by: Jia, Zixia, et al.
Published: (2024)
Adaptive Preference Optimization with Uncertainty-aware Utility Anchor
by: Wang, Xiaobo, et al.
Published: (2025)
by: Wang, Xiaobo, et al.
Published: (2025)
A Study on the Ancient theater of official house in The Taihang mountain area of North Henan Province in China
by: Hengli Peng
Published: (2023)
by: Hengli Peng
Published: (2023)
InterPreT: Interactive Predicate Learning from Language Feedback for Generalizable Task Planning
by: Han, Muzhi, et al.
Published: (2024)
by: Han, Muzhi, et al.
Published: (2024)
DisenReason: Behavior Disentanglement and Latent Reasoning for Shared-Account Sequential Recommendation
by: Cheng, Jiawei, et al.
Published: (2026)
by: Cheng, Jiawei, et al.
Published: (2026)
ReflectEvo: Improving Meta Introspection of Small LLMs by Learning Self-Reflection
by: Li, Jiaqi, et al.
Published: (2025)
by: Li, Jiaqi, et al.
Published: (2025)
Latent-GRPO: Group Relative Policy Optimization for Latent Reasoning
by: Deng, Jingcheng, et al.
Published: (2026)
by: Deng, Jingcheng, et al.
Published: (2026)
Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
by: Lou, Chao, et al.
Published: (2024)
by: Lou, Chao, et al.
Published: (2024)
LaSER: Internalizing Explicit Reasoning into Latent Space for Dense Retrieval
by: Jin, Jiajie, et al.
Published: (2026)
by: Jin, Jiajie, et al.
Published: (2026)
ExoViP: Step-by-step Verification and Exploration with Exoskeleton Modules for Compositional Visual Reasoning
by: Wang, Yuxuan, et al.
Published: (2024)
by: Wang, Yuxuan, et al.
Published: (2024)
Wasserstein Proximal Policy Gradient
by: Zhu, Zhaoyu, et al.
Published: (2026)
by: Zhu, Zhaoyu, et al.
Published: (2026)
Similar Items
-
MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
by: Mi, Yapeng, et al.
Published: (2025) -
Discrete Markov Bridge
by: Li, Hengli, et al.
Published: (2025) -
BEDA: Belief Estimation as Probabilistic Constraints for Performing Strategic Dialogue Acts
by: Li, Hengli, et al.
Published: (2025) -
TokenSwift: Lossless Acceleration of Ultra Long Sequence Generation
by: Wu, Tong, et al.
Published: (2025) -
Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning
by: Wu, Tong, et al.
Published: (2025)