Agent Explorative Policy Optimization for Multimodal Agentic Reasoning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Kang, Minki, Diao, Shizhe, Hachiuma, Ryo, Hwang, Sung Ju, Molchanov, Pavlo, Wang, Yu-Chiang Frank, Lee, Byung-Kwan |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs
par: Lee, Chanuk, et autres
Publié: (2026)
par: Lee, Chanuk, et autres
Publié: (2026)
Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR
par: Lee, Chanuk, et autres
Publié: (2026)
par: Lee, Chanuk, et autres
Publié: (2026)
GenRecal: Generation after Recalibration from Large to Small Vision-Language Models
par: Lee, Byung-Kwan, et autres
Publié: (2025)
par: Lee, Byung-Kwan, et autres
Publié: (2025)
Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
par: Lee, Byung-Kwan, et autres
Publié: (2025)
par: Lee, Byung-Kwan, et autres
Publié: (2025)
Distilling LLM Agent into Small Models with Retrieval and Code Tools
par: Kang, Minki, et autres
Publié: (2025)
par: Kang, Minki, et autres
Publié: (2025)
Latent Paraphrasing: Perturbation on Layers Improves Knowledge Injection in Language Models
par: Kang, Minki, et autres
Publié: (2024)
par: Kang, Minki, et autres
Publié: (2024)
PREPING: Building Agent Memory without Tasks
par: Choi, Yumin, et autres
Publié: (2026)
par: Choi, Yumin, et autres
Publié: (2026)
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
par: Liu, Shih-Yang, et autres
Publié: (2026)
par: Liu, Shih-Yang, et autres
Publié: (2026)
BroRL: Scaling Reinforcement Learning via Broadened Exploration
par: Hu, Jian, et autres
Publié: (2025)
par: Hu, Jian, et autres
Publié: (2025)
Memory Transfer Learning: How Memories are Transferred Across Domains in Coding Agents
par: Kim, Kangsan, et autres
Publié: (2026)
par: Kim, Kangsan, et autres
Publié: (2026)
VLsI: Verbalized Layers-to-Interactions from Large to Small Vision Language Models
par: Lee, Byung-Kwan, et autres
Publié: (2024)
par: Lee, Byung-Kwan, et autres
Publié: (2024)
Unified Reinforcement and Imitation Learning for Vision-Language Models
par: Lee, Byung-Kwan, et autres
Publié: (2025)
par: Lee, Byung-Kwan, et autres
Publié: (2025)
Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception
par: Wan, Zhen, et autres
Publié: (2026)
par: Wan, Zhen, et autres
Publié: (2026)
Entropy-Regularized Process Reward Model
par: Zhang, Hanning, et autres
Publié: (2024)
par: Zhang, Hanning, et autres
Publié: (2024)
WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning
par: Yeo, Woongyeong, et autres
Publié: (2025)
par: Yeo, Woongyeong, et autres
Publié: (2025)
Fast-dLLM v2: Efficient Block-Diffusion LLM
par: Wu, Chengyue, et autres
Publié: (2025)
par: Wu, Chengyue, et autres
Publié: (2025)
SafeRoute: Adaptive Model Selection for Efficient and Accurate Safety Guardrails in Large Language Models
par: Lee, Seanie, et autres
Publié: (2025)
par: Lee, Seanie, et autres
Publié: (2025)
Towards Predicting Any Human Trajectory In Context
par: Fujii, Ryo, et autres
Publié: (2025)
par: Fujii, Ryo, et autres
Publié: (2025)
DLER: Doing Length pEnalty Right - Incentivizing More Intelligence per Token via Reinforcement Learning
par: Liu, Shih-Yang, et autres
Publié: (2025)
par: Liu, Shih-Yang, et autres
Publié: (2025)
ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and Judge
par: Wang, Zhilin, et autres
Publié: (2025)
par: Wang, Zhilin, et autres
Publié: (2025)
LongMamba: Enhancing Mamba's Long Context Capabilities via Training-Free Receptive Field Enlargement
par: Ye, Zhifan, et autres
Publié: (2025)
par: Ye, Zhifan, et autres
Publié: (2025)
$\textbf{PLUM}$: Improving Code LMs with Execution-Guided On-Policy Preference Learning Driven By Synthetic Test Cases
par: Zhang, Dylan, et autres
Publié: (2024)
par: Zhang, Dylan, et autres
Publié: (2024)
Small Language Models are the Future of Agentic AI
par: Belcak, Peter, et autres
Publié: (2025)
par: Belcak, Peter, et autres
Publié: (2025)
Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMs
par: Choi, Yumin, et autres
Publié: (2025)
par: Choi, Yumin, et autres
Publié: (2025)
HarmAug: Effective Data Augmentation for Knowledge Distillation of Safety Guard Models
par: Lee, Seanie, et autres
Publié: (2024)
par: Lee, Seanie, et autres
Publié: (2024)
OmniRetrieval: Unified Retrieval across Heterogeneous Knowledge Sources
par: Baek, Jinheon, et autres
Publié: (2026)
par: Baek, Jinheon, et autres
Publié: (2026)
Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation
par: Yu, Seonghoon, et autres
Publié: (2026)
par: Yu, Seonghoon, et autres
Publié: (2026)
CodeGraph: Enhancing Graph Reasoning of LLMs with Code
par: Cai, Qiaolong, et autres
Publié: (2024)
par: Cai, Qiaolong, et autres
Publié: (2024)
Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training
par: Diao, Shizhe, et autres
Publié: (2025)
par: Diao, Shizhe, et autres
Publié: (2025)
CCR 2.0: High-level Reasoning for Conditional Refinements
par: Song, Youngju, et autres
Publié: (2025)
par: Song, Youngju, et autres
Publié: (2025)
Automatic Prompt Augmentation and Selection with Chain-of-Thought from Labeled Data
par: Shum, KaShun, et autres
Publié: (2023)
par: Shum, KaShun, et autres
Publié: (2023)
DSPO: Stable and Efficient Policy Optimization for Agentic Search and Reasoning
par: Gu, Chenyang, et autres
Publié: (2025)
par: Gu, Chenyang, et autres
Publié: (2025)
Unified Multimodal Interleaved Document Representation for Retrieval
par: Lee, Jaewoo, et autres
Publié: (2024)
par: Lee, Jaewoo, et autres
Publié: (2024)
AutoML-Agent: A Multi-Agent LLM Framework for Full-Pipeline AutoML
par: Trirat, Patara, et autres
Publié: (2024)
par: Trirat, Patara, et autres
Publié: (2024)
System Prompt Optimization with Meta-Learning
par: Choi, Yumin, et autres
Publié: (2025)
par: Choi, Yumin, et autres
Publié: (2025)
DPEPO: Diverse Parallel Exploration Policy Optimization for LLM-based Agents
par: Zhang, Junshuo, et autres
Publié: (2026)
par: Zhang, Junshuo, et autres
Publié: (2026)
ConstraintChecker: A Plugin for Large Language Models to Reason on Commonsense Knowledge Bases
par: Do, Quyet V., et autres
Publié: (2024)
par: Do, Quyet V., et autres
Publié: (2024)
Learn from Weaknesses: Automated Domain Specialization for Small Computer-Use Agents
par: Kim, Suji, et autres
Publié: (2026)
par: Kim, Suji, et autres
Publié: (2026)
Perception-Aware Policy Optimization for Multimodal Reasoning
par: Wang, Zhenhailong, et autres
Publié: (2025)
par: Wang, Zhenhailong, et autres
Publié: (2025)
When Thoughts Meet Facts: Reusable Reasoning for Long-Context LMs
par: Jeong, Soyeong, et autres
Publié: (2025)
par: Jeong, Soyeong, et autres
Publié: (2025)
Documents similaires
-
SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs
par: Lee, Chanuk, et autres
Publié: (2026) -
Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR
par: Lee, Chanuk, et autres
Publié: (2026) -
GenRecal: Generation after Recalibration from Large to Small Vision-Language Models
par: Lee, Byung-Kwan, et autres
Publié: (2025) -
Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
par: Lee, Byung-Kwan, et autres
Publié: (2025) -
Distilling LLM Agent into Small Models with Retrieval and Code Tools
par: Kang, Minki, et autres
Publié: (2025)