Navigate the Unknown: Enhancing LLM Reasoning with Intrinsic Motivation Guided Exploration
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gao, Jingtong, Pan, Ling, Wang, Yejing, Zhong, Rui, Lu, Chi, Wang, Maolin, Cai, Qingpeng, Jiang, Peng, Zhao, Xiangyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Generative Auto-Bidding with Value-Guided Explorations
von: Gao, Jingtong, et al.
Veröffentlicht: (2025)
von: Gao, Jingtong, et al.
Veröffentlicht: (2025)
Reinforced Preference Optimization for Reasoning-Augmented Recommendations
von: Gao, Jingtong, et al.
Veröffentlicht: (2026)
von: Gao, Jingtong, et al.
Veröffentlicht: (2026)
Job Skill Extraction via LLM-Centric Multi-Module Framework
von: Li, Guojing, et al.
Veröffentlicht: (2026)
von: Li, Guojing, et al.
Veröffentlicht: (2026)
Generative Auto-Bidding in Large-Scale Competitive Auctions via Diffusion Completer-Aligner
von: Li, Yewen, et al.
Veröffentlicht: (2025)
von: Li, Yewen, et al.
Veröffentlicht: (2025)
From Principles to Applications: A Comprehensive Survey of Discrete Tokenizers in Generation, Comprehension, Recommendation, and Information Retrieval
von: Jia, Jian, et al.
Veröffentlicht: (2025)
von: Jia, Jian, et al.
Veröffentlicht: (2025)
Count Counts: Motivating Exploration in LLM Reasoning with Count-based Intrinsic Rewards
von: Zhang, Xuan, et al.
Veröffentlicht: (2025)
von: Zhang, Xuan, et al.
Veröffentlicht: (2025)
Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards
von: He, Haoran, et al.
Veröffentlicht: (2025)
von: He, Haoran, et al.
Veröffentlicht: (2025)
SEARCH-R: Structured Entity-Aware Retrieval with Chain-of-Reasoning Navigator for Multi-hop Question Answering
von: Fu, Yuqing, et al.
Veröffentlicht: (2026)
von: Fu, Yuqing, et al.
Veröffentlicht: (2026)
ThoughtProbe: Classifier-Guided Thought Space Exploration Leveraging LLM Intrinsic Reasoning
von: Wang, Zijian, et al.
Veröffentlicht: (2025)
von: Wang, Zijian, et al.
Veröffentlicht: (2025)
GAS: Generative Auto-bidding with Post-training Search
von: Li, Yewen, et al.
Veröffentlicht: (2024)
von: Li, Yewen, et al.
Veröffentlicht: (2024)
Learning Off-policy with Model-based Intrinsic Motivation For Active Online Exploration
von: Wang, Yibo, et al.
Veröffentlicht: (2024)
von: Wang, Yibo, et al.
Veröffentlicht: (2024)
MetaLoRA: Tensor-Enhanced Adaptive Low-Rank Fine-tuning
von: Wang, Maolin, et al.
Veröffentlicht: (2025)
von: Wang, Maolin, et al.
Veröffentlicht: (2025)
LLM-Powered User Simulator for Recommender System
von: Zhang, Zijian, et al.
Veröffentlicht: (2024)
von: Zhang, Zijian, et al.
Veröffentlicht: (2024)
Intrinsic-Motivation Multi-Robot Social Formation Navigation with Coordinated Exploration
von: Fu, Hao, et al.
Veröffentlicht: (2025)
von: Fu, Hao, et al.
Veröffentlicht: (2025)
Behavior Modeling Space Reconstruction for E-Commerce Search
von: Wang, Yejing, et al.
Veröffentlicht: (2025)
von: Wang, Yejing, et al.
Veröffentlicht: (2025)
LBM: Hierarchical Large Auto-Bidding Model via Reasoning and Acting
von: Li, Yewen, et al.
Veröffentlicht: (2026)
von: Li, Yewen, et al.
Veröffentlicht: (2026)
TrackRec: Iterative Alternating Feedback with Chain-of-Thought via Preference Alignment for Recommendation
von: Xia, Yu, et al.
Veröffentlicht: (2025)
von: Xia, Yu, et al.
Veröffentlicht: (2025)
Scenario-Wise Rec: A Multi-Scenario Recommendation Benchmark
von: Li, Xiaopeng, et al.
Veröffentlicht: (2024)
von: Li, Xiaopeng, et al.
Veröffentlicht: (2024)
AdaSwitch: Balancing Exploration and Guidance in Knowledge Distillation via Adaptive Switching
von: Peng, Jingyu, et al.
Veröffentlicht: (2025)
von: Peng, Jingyu, et al.
Veröffentlicht: (2025)
From Laws to Motivation: Guiding Exploration through Law-Based Reasoning and Rewards
von: Chen, Ziyu, et al.
Veröffentlicht: (2024)
von: Chen, Ziyu, et al.
Veröffentlicht: (2024)
DLCRec: A Novel Approach for Managing Diversity in LLM-Based Recommender Systems
von: Chen, Jiaju, et al.
Veröffentlicht: (2024)
von: Chen, Jiaju, et al.
Veröffentlicht: (2024)
Random Policy Evaluation Uncovers Policies of Generative Flow Networks
von: He, Haoran, et al.
Veröffentlicht: (2024)
von: He, Haoran, et al.
Veröffentlicht: (2024)
Enhancing Surface Neural Implicits with Curvature-Guided Sampling and Uncertainty-Augmented Representations
von: Sang, Lu, et al.
Veröffentlicht: (2023)
von: Sang, Lu, et al.
Veröffentlicht: (2023)
Intrinsically-Motivated Humans and Agents in Open-World Exploration
von: Lidayan, Aly, et al.
Veröffentlicht: (2025)
von: Lidayan, Aly, et al.
Veröffentlicht: (2025)
Structured Spectral Reasoning for Frequency-Adaptive Multimodal Recommendation
von: Yang, Wei, et al.
Veröffentlicht: (2025)
von: Yang, Wei, et al.
Veröffentlicht: (2025)
Conditional Memory Enhanced Item Representation for Generative Recommendation
von: Liu, Ziwei, et al.
Veröffentlicht: (2026)
von: Liu, Ziwei, et al.
Veröffentlicht: (2026)
GLINT-RU: Gated Lightweight Intelligent Recurrent Units for Sequential Recommender Systems
von: Zhang, Sheng, et al.
Veröffentlicht: (2024)
von: Zhang, Sheng, et al.
Veröffentlicht: (2024)
Reasoning through Exploration: A Reinforcement Learning Framework for Robust Function Calling
von: Hao, Bingguang, et al.
Veröffentlicht: (2025)
von: Hao, Bingguang, et al.
Veröffentlicht: (2025)
Exploring Spatial Representation to Enhance LLM Reasoning in Aerial Vision-Language Navigation
von: Gao, Yunpeng, et al.
Veröffentlicht: (2024)
von: Gao, Yunpeng, et al.
Veröffentlicht: (2024)
SIGMA: Selective Gated Mamba for Sequential Recommendation
von: Liu, Ziwei, et al.
Veröffentlicht: (2024)
von: Liu, Ziwei, et al.
Veröffentlicht: (2024)
Empowering Denoising Sequential Recommendation with Large Language Model Embeddings
von: Wu, Tongzhou, et al.
Veröffentlicht: (2025)
von: Wu, Tongzhou, et al.
Veröffentlicht: (2025)
Position-Aware Drafting for Inference Acceleration in LLM-Based Generative List-Wise Recommendation
von: Chen, Jiaju, et al.
Veröffentlicht: (2026)
von: Chen, Jiaju, et al.
Veröffentlicht: (2026)
Curriculum Learning for Efficient Chain-of-Thought Distillation via Structure-Aware Masking and GRPO
von: Yu, Bowen, et al.
Veröffentlicht: (2026)
von: Yu, Bowen, et al.
Veröffentlicht: (2026)
Large Language Model Enhanced Recommender Systems: A Survey
von: Liu, Qidong, et al.
Veröffentlicht: (2024)
von: Liu, Qidong, et al.
Veröffentlicht: (2024)
Enhancing Conversational Agents via Task-Oriented Adversarial Memory Adaptation
von: Deng, Yimin, et al.
Veröffentlicht: (2026)
von: Deng, Yimin, et al.
Veröffentlicht: (2026)
LASER: LLM Agent with State-Space Exploration for Web Navigation
von: Ma, Kaixin, et al.
Veröffentlicht: (2023)
von: Ma, Kaixin, et al.
Veröffentlicht: (2023)
LLM4Rerank: LLM-based Auto-Reranking Framework for Recommendations
von: Gao, Jingtong, et al.
Veröffentlicht: (2024)
von: Gao, Jingtong, et al.
Veröffentlicht: (2024)
LLM-EDT: Large Language Model Enhanced Cross-domain Sequential Recommendation with Dual-phase Training
von: Liu, Ziwei, et al.
Veröffentlicht: (2025)
von: Liu, Ziwei, et al.
Veröffentlicht: (2025)
MCTuner: Spatial Decomposition-Enhanced Database Tuning via LLM-Guided Exploration
von: Yan, Zihan, et al.
Veröffentlicht: (2025)
von: Yan, Zihan, et al.
Veröffentlicht: (2025)
Can LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM Reasoning
von: Liang, Zhenwen, et al.
Veröffentlicht: (2025)
von: Liang, Zhenwen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Generative Auto-Bidding with Value-Guided Explorations
von: Gao, Jingtong, et al.
Veröffentlicht: (2025) -
Reinforced Preference Optimization for Reasoning-Augmented Recommendations
von: Gao, Jingtong, et al.
Veröffentlicht: (2026) -
Job Skill Extraction via LLM-Centric Multi-Module Framework
von: Li, Guojing, et al.
Veröffentlicht: (2026) -
Generative Auto-Bidding in Large-Scale Competitive Auctions via Diffusion Completer-Aligner
von: Li, Yewen, et al.
Veröffentlicht: (2025) -
From Principles to Applications: A Comprehensive Survey of Discrete Tokenizers in Generation, Comprehension, Recommendation, and Information Retrieval
von: Jia, Jian, et al.
Veröffentlicht: (2025)