Restoring Exploration after Post-Training: Latent Exploration Decoding for Large Reasoning Models
Fuente:
arXiv
Saved in:
| Main Authors: | Tan, Wenhui, Parascandolo, Fiorenzo, Sangineto, Enver, Ju, Jianzhong, Luo, Zhenbo, Cao, Qian, Cucchiara, Rita, Song, Ruihua, Luan, Jian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BFS-PO: Best-First Search for Large Reasoning Models
by: Parascandolo, Fiorenzo, et al.
Published: (2026)
by: Parascandolo, Fiorenzo, et al.
Published: (2026)
Causal Graphical Models for Vision-Language Compositional Understanding
by: Parascandolo, Fiorenzo, et al.
Published: (2024)
by: Parascandolo, Fiorenzo, et al.
Published: (2024)
Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning Chains
by: Tan, Wenhui, et al.
Published: (2025)
by: Tan, Wenhui, et al.
Published: (2025)
One Transformer for All Time Series: Representing and Training with Time-Dependent Heterogeneous Tabular Data
by: Luetto, Simone, et al.
Published: (2023)
by: Luetto, Simone, et al.
Published: (2023)
Think-Clip-Sample: Slow-Fast Frame Selection for Video Understanding
by: Tan, Wenhui, et al.
Published: (2026)
by: Tan, Wenhui, et al.
Published: (2026)
Diffusion Transformers for Tabular Data Time Series Generation
by: Garuti, Fabrizio, et al.
Published: (2025)
by: Garuti, Fabrizio, et al.
Published: (2025)
MSJoE: Jointly Evolving MLLM and Sampler for Efficient Long-Form Video Understanding
by: Tan, Wenhui, et al.
Published: (2026)
by: Tan, Wenhui, et al.
Published: (2026)
Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation
by: Li, Jiaze, et al.
Published: (2026)
by: Li, Jiaze, et al.
Published: (2026)
ThinkOmni: Lifting Textual Reasoning to Omni-modal Scenarios via Guidance Decoding
by: Guan, Yiran, et al.
Published: (2026)
by: Guan, Yiran, et al.
Published: (2026)
REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding
by: Li, Jiaze, et al.
Published: (2025)
by: Li, Jiaze, et al.
Published: (2025)
Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension
by: Xu, Haoran, et al.
Published: (2026)
by: Xu, Haoran, et al.
Published: (2026)
TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding
by: Xu, Boshen, et al.
Published: (2025)
by: Xu, Boshen, et al.
Published: (2025)
Embodied Agents for Efficient Exploration and Smart Scene Description
by: Bigazzi, Roberto, et al.
Published: (2023)
by: Bigazzi, Roberto, et al.
Published: (2023)
Federated Joint Learning for Domain and Class Generalization
by: Xu, Haoran, et al.
Published: (2026)
by: Xu, Haoran, et al.
Published: (2026)
Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
by: Guan, Yiran, et al.
Published: (2026)
by: Guan, Yiran, et al.
Published: (2026)
Xiaomi MiMo-VL-Miloco Technical Report
by: Li, Jiaze, et al.
Published: (2025)
by: Li, Jiaze, et al.
Published: (2025)
RoLD: Robot Latent Diffusion for Multi-task Policy Modeling
by: Tan, Wenhui, et al.
Published: (2024)
by: Tan, Wenhui, et al.
Published: (2024)
Chapter Alberto Magnaghi, ovvero gli imprescindibili riferimenti territoriali
by: Fabio, Parascandolo
Published: (2026)
by: Fabio, Parascandolo
Published: (2026)
Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle
by: Zhu, Linghao, et al.
Published: (2025)
by: Zhu, Linghao, et al.
Published: (2025)
Direction-Aware Diagonal Autoregressive Image Generation
by: Xu, Yijia, et al.
Published: (2025)
by: Xu, Yijia, et al.
Published: (2025)
LLaVA-SG: Leveraging Scene Graphs as Visual Semantic Expression in Vision-Language Models
by: Wang, Jingyi, et al.
Published: (2024)
by: Wang, Jingyi, et al.
Published: (2024)
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
by: Wang, Ye, et al.
Published: (2025)
by: Wang, Ye, et al.
Published: (2025)
TIR-Agent: Training an Explorative and Efficient Agent for Image Restoration
by: Zhang, Yisheng, et al.
Published: (2026)
by: Zhang, Yisheng, et al.
Published: (2026)
ABRA: Teleporting Fine-Tuned Knowledge Across Domains for Open-Vocabulary Object Detection
by: Bernardi, Mattia, et al.
Published: (2026)
by: Bernardi, Mattia, et al.
Published: (2026)
ReLaX: Reasoning with Latent Exploration for Large Reasoning Models
by: Zhang, Shimin, et al.
Published: (2025)
by: Zhang, Shimin, et al.
Published: (2025)
Cook and Clean Together: Teaching Embodied Agents for Parallel Task Execution
by: Liang, Dingkang, et al.
Published: (2025)
by: Liang, Dingkang, et al.
Published: (2025)
Decoding Cortical Microcircuits: A Generative Model for Latent Space Exploration and Controlled Synthesis
by: Liu, Xingyu, et al.
Published: (2025)
by: Liu, Xingyu, et al.
Published: (2025)
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior
by: Huang, Zeyi, et al.
Published: (2026)
by: Huang, Zeyi, et al.
Published: (2026)
Reasoning Palette: Modulating Reasoning via Latent Contextualization for Controllable Exploration for (V)LMs
by: Long, Rujiao, et al.
Published: (2025)
by: Long, Rujiao, et al.
Published: (2025)
Guiding the Experts: Semantic Priors for Efficient and Focused MoE Routing
by: Min, Chengxi, et al.
Published: (2025)
by: Min, Chengxi, et al.
Published: (2025)
Semantic Residual Prompts for Continual Learning
by: Menabue, Martin, et al.
Published: (2024)
by: Menabue, Martin, et al.
Published: (2024)
JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
by: Chao, Jianghan, et al.
Published: (2025)
by: Chao, Jianghan, et al.
Published: (2025)
Unrewarded Exploration in Large Language Models Reveals Latent Learning from Psychology
by: Xiong, Jian, et al.
Published: (2026)
by: Xiong, Jian, et al.
Published: (2026)
AutoLink: Autonomous Schema Exploration and Expansion for Scalable Schema Linking in Text-to-SQL at Scale
by: Wang, Ziyang, et al.
Published: (2025)
by: Wang, Ziyang, et al.
Published: (2025)
Multimodal-Conditioned Latent Diffusion Models for Fashion Image Editing
by: Baldrati, Alberto, et al.
Published: (2024)
by: Baldrati, Alberto, et al.
Published: (2024)
Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition
by: Zhang, Yiru, et al.
Published: (2025)
by: Zhang, Yiru, et al.
Published: (2025)
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs
by: Zhang, Shaojie, et al.
Published: (2025)
by: Zhang, Shaojie, et al.
Published: (2025)
Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization
by: Hua, Xingyuan, et al.
Published: (2026)
by: Hua, Xingyuan, et al.
Published: (2026)
Fluent and Accurate Image Captioning with a Self-Trained Reward Model
by: Moratelli, Nicholas, et al.
Published: (2024)
by: Moratelli, Nicholas, et al.
Published: (2024)
Trajectory Forecasting through Low-Rank Adaptation of Discrete Latent Codes
by: Benaglia, Riccardo, et al.
Published: (2024)
by: Benaglia, Riccardo, et al.
Published: (2024)
Similar Items
-
BFS-PO: Best-First Search for Large Reasoning Models
by: Parascandolo, Fiorenzo, et al.
Published: (2026) -
Causal Graphical Models for Vision-Language Compositional Understanding
by: Parascandolo, Fiorenzo, et al.
Published: (2024) -
Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning Chains
by: Tan, Wenhui, et al.
Published: (2025) -
One Transformer for All Time Series: Representing and Training with Time-Dependent Heterogeneous Tabular Data
by: Luetto, Simone, et al.
Published: (2023) -
Think-Clip-Sample: Slow-Fast Frame Selection for Video Understanding
by: Tan, Wenhui, et al.
Published: (2026)