Maximizing Rollout Informativeness under a Fixed Budget: A Submodular View of Tree Search for Tool-Use Agentic Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Yuelin, Yu, Zhenbo, Cheng, Zhengxue, Liu, Wei, Song, Li |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
by: Estevanell-Valladares, Ernesto L., et al.
Published: (2025)
by: Estevanell-Valladares, Ernesto L., et al.
Published: (2025)
Murphys Laws of AI Alignment: Why the Gap Always Wins
by: Gaikwad, Madhava
Published: (2025)
by: Gaikwad, Madhava
Published: (2025)
Emotion-Inspired Learning Signals (EILS): A Homeostatic Framework for Adaptive Autonomous Agents
by: Tiwari, Dhruv
Published: (2025)
by: Tiwari, Dhruv
Published: (2025)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
by: Pather, Kaviraj, et al.
Published: (2025)
by: Pather, Kaviraj, et al.
Published: (2025)
RoboMoRe: LLM-based Robot Co-design via Joint Optimization of Morphology and Reward
by: Fang, Jiawei, et al.
Published: (2025)
by: Fang, Jiawei, et al.
Published: (2025)
Approaches to Semantic Textual Similarity in Slovak Language: From Algorithms to Transformers
by: Radosky, Lukas, et al.
Published: (2026)
by: Radosky, Lukas, et al.
Published: (2026)
UniPROT: Uniform Prototype Selection via Partial Optimal Transport with Submodular Guarantees
by: Chanda, Prateek, et al.
Published: (2026)
by: Chanda, Prateek, et al.
Published: (2026)
Adaptive Minds: Empowering Agents with LoRA-as-Tools
by: Shekar, Pavan C, et al.
Published: (2025)
by: Shekar, Pavan C, et al.
Published: (2025)
Generative AI and the Transformation of Software Development Practices
by: Acharya, Vivek
Published: (2025)
by: Acharya, Vivek
Published: (2025)
Streaming Continual Learning for Unified Adaptive Intelligence in Dynamic Environments
by: Giannini, Federico, et al.
Published: (2026)
by: Giannini, Federico, et al.
Published: (2026)
Empirical Evaluation of QAOA with Zero Noise Extrapolation on NISQ Hardware for Carbon Credit Portfolio Optimization in the Brazilian Cerrado
by: Ribeiro, Hugo José
Published: (2026)
by: Ribeiro, Hugo José
Published: (2026)
CogniLoad: A Synthetic Natural Language Reasoning Benchmark With Tunable Length, Intrinsic Difficulty, and Distractor Density
by: Kaiser, Daniel, et al.
Published: (2025)
by: Kaiser, Daniel, et al.
Published: (2025)
Intersymbolic AI: Interlinking Symbolic AI and Subsymbolic AI
by: Platzer, André
Published: (2024)
by: Platzer, André
Published: (2024)
Diffusion-MPC in Discrete Domains: Feasibility Constraints, Horizon Effects, and Critic Alignment: Case study with Tetris
by: Wang, Haochuan Kevin
Published: (2026)
by: Wang, Haochuan Kevin
Published: (2026)
Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models
by: Fu, Tianyu, et al.
Published: (2025)
by: Fu, Tianyu, et al.
Published: (2025)
optimize_anything: A Universal API for Optimizing any Text Parameter
by: Agrawal, Lakshya A, et al.
Published: (2026)
by: Agrawal, Lakshya A, et al.
Published: (2026)
Enhancing Diversity in Multi-objective Feature Selection
by: Miyandoab, Sevil Zanjani, et al.
Published: (2024)
by: Miyandoab, Sevil Zanjani, et al.
Published: (2024)
Improving LLM Agent Planning with In-Context Learning via Atomic Fact Augmentation and Lookahead Search
by: Holt, Samuel, et al.
Published: (2025)
by: Holt, Samuel, et al.
Published: (2025)
An Improved Adaptive PID Optimizer with Enhanced Convergence and Stability for Deep Learning
by: Saini, Saurabh, et al.
Published: (2026)
by: Saini, Saurabh, et al.
Published: (2026)
SALE-Based Offline Reinforcement Learning with Ensemble Q-Networks
by: Chun, Zheng
Published: (2025)
by: Chun, Zheng
Published: (2025)
AI Agents: Evolution, Architecture, and Real-World Applications
by: Krishnan, Naveen
Published: (2025)
by: Krishnan, Naveen
Published: (2025)
CAPE: Corrective Actions from Precondition Errors using Large Language Models
by: Raman, Shreyas Sundara, et al.
Published: (2022)
by: Raman, Shreyas Sundara, et al.
Published: (2022)
Discovering Algorithms with Computational Language Processing
by: Bourdais, Theo, et al.
Published: (2025)
by: Bourdais, Theo, et al.
Published: (2025)
PLUGH: A Benchmark for Spatial Understanding and Reasoning in Large Language Models
by: Tikhonov, Alexey
Published: (2024)
by: Tikhonov, Alexey
Published: (2024)
EvoPref: Multi-Objective Evolutionary Optimization Discovers Diverse LLM Alignments Beyond Gradient Descent
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Parameter-Efficient Neuroevolution for Diverse LLM Generation: Quality-Diversity Optimization via Prompt Embedding Evolution
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Dual-Channel Feature Fusion for Joint Prediction in Dynamic Signed Weighted Networks
by: Zhang, Gaoxin, et al.
Published: (2026)
by: Zhang, Gaoxin, et al.
Published: (2026)
DISC: Dynamic Decomposition Improves LLM Inference Scaling
by: Light, Jonathan, et al.
Published: (2025)
by: Light, Jonathan, et al.
Published: (2025)
A Comparison Between Decision Transformers and Traditional Offline Reinforcement Learning Algorithms
by: Caunhye, Ali Murtaza, et al.
Published: (2025)
by: Caunhye, Ali Murtaza, et al.
Published: (2025)
SMOSE: Sparse Mixture of Shallow Experts for Interpretable Reinforcement Learning in Continuous Control Tasks
by: Vincze, Mátyás, et al.
Published: (2024)
by: Vincze, Mátyás, et al.
Published: (2024)
A Structural Threshold in Decision Capacity Governs Collapse in Self-Play Reinforcement Learning
by: Kujur, Arahan
Published: (2026)
by: Kujur, Arahan
Published: (2026)
Evolving machine learning workflows through interactive AutoML
by: Barbudo, Rafael, et al.
Published: (2024)
by: Barbudo, Rafael, et al.
Published: (2024)
Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation
by: Osman, Asim, et al.
Published: (2026)
by: Osman, Asim, et al.
Published: (2026)
Computational Hardness of Reinforcement Learning with Partial $q^π$-Realizability
by: Karimi, Shayan, et al.
Published: (2025)
by: Karimi, Shayan, et al.
Published: (2025)
STRIDE: A Self-Reflective Agent Framework for Reliable Automatic Equation Discovery
by: Su, Jiarui, et al.
Published: (2026)
by: Su, Jiarui, et al.
Published: (2026)
The Geometry of Thought: Disclosing the Transformer as a Tropical Polynomial Circuit
by: Alpay, Faruk, et al.
Published: (2026)
by: Alpay, Faruk, et al.
Published: (2026)
BudgetMLAgent: A Cost-Effective LLM Multi-Agent system for Automating Machine Learning Tasks
by: Gandhi, Shubham, et al.
Published: (2024)
by: Gandhi, Shubham, et al.
Published: (2024)
Generative AI in Transportation Planning: A Survey
by: Da, Longchao, et al.
Published: (2025)
by: Da, Longchao, et al.
Published: (2025)
CellARC: Measuring Intelligence with Cellular Automata
by: Lžičař, Miroslav
Published: (2025)
by: Lžičař, Miroslav
Published: (2025)
Context Engineering for Multi-Agent LLM Code Assistants Using Elicit, NotebookLM, ChatGPT, and Claude Code
by: Haseeb, Muhammad
Published: (2025)
by: Haseeb, Muhammad
Published: (2025)
Similar Items
-
XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
by: Estevanell-Valladares, Ernesto L., et al.
Published: (2025) -
Murphys Laws of AI Alignment: Why the Gap Always Wins
by: Gaikwad, Madhava
Published: (2025) -
Emotion-Inspired Learning Signals (EILS): A Homeostatic Framework for Adaptive Autonomous Agents
by: Tiwari, Dhruv
Published: (2025) -
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
by: Pather, Kaviraj, et al.
Published: (2025) -
RoboMoRe: LLM-based Robot Co-design via Joint Optimization of Morphology and Reward
by: Fang, Jiawei, et al.
Published: (2025)