Enhanced LLM Reasoning by Optimizing Reward Functions with Search-Driven Reinforcement Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Ahmadi, Arash, Sharif, Sarah, Yaser, Banad |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Improving Aviation Safety Analysis: Automated HFACS Classification Using Reinforcement Learning with Group Relative Policy Optimization
por: Ahmadi, Arash, et al.
Publicado: (2025)
por: Ahmadi, Arash, et al.
Publicado: (2025)
A Comparative Study of Sampling Methods with Cross-Validation in the FedHome Framework
por: Ahmadi, Arash, et al.
Publicado: (2024)
por: Ahmadi, Arash, et al.
Publicado: (2024)
MCP Bridge: A Lightweight, LLM-Agnostic RESTful Proxy for Model Context Protocol Servers
por: Ahmadi, Arash, et al.
Publicado: (2025)
por: Ahmadi, Arash, et al.
Publicado: (2025)
A Cloud-Edge Framework for Energy-Efficient Event-Driven Control: An Integration of Online Supervised Learning, Spiking Neural Networks and Local Plasticity Rules
por: Ahmadvand, Reza, et al.
Publicado: (2024)
por: Ahmadvand, Reza, et al.
Publicado: (2024)
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning
por: Zhao, Qingfei, et al.
Publicado: (2025)
por: Zhao, Qingfei, et al.
Publicado: (2025)
Enhancing LLM Reasoning with Reward-guided Tree Search
por: Jiang, Jinhao, et al.
Publicado: (2024)
por: Jiang, Jinhao, et al.
Publicado: (2024)
Neuromorphic Digital-Twin-based Controller for Indoor Multi-UAV Systems Deployment
por: Ahmadvand, Reza, et al.
Publicado: (2025)
por: Ahmadvand, Reza, et al.
Publicado: (2025)
No Dense Tensors Needed: Fully Sparse Object Detection on Event-Camera Voxel Grids
por: Sadoun, Mohamad Yazan, et al.
Publicado: (2026)
por: Sadoun, Mohamad Yazan, et al.
Publicado: (2026)
Brain-inspired spike-timing plasticity for reliable label-efficient event-camera vision
por: Sadoun, Mohamad Yazan, et al.
Publicado: (2026)
por: Sadoun, Mohamad Yazan, et al.
Publicado: (2026)
Event-based Heterogeneous Information Processing for Online Vision-based Obstacle Detection and Localization
por: Ahmadvand, Reza, et al.
Publicado: (2026)
por: Ahmadvand, Reza, et al.
Publicado: (2026)
Neuromorphic Robust Estimation of Nonlinear Dynamical Systems Applied to Satellite Rendezvous
por: Ahmadvand, Reza, et al.
Publicado: (2024)
por: Ahmadvand, Reza, et al.
Publicado: (2024)
Swarm Intelligence in Collision-free Formation Control for Multi-UAV Systems with 3D Obstacle Avoidance Maneuvers
por: Ahmadvand, Reza, et al.
Publicado: (2024)
por: Ahmadvand, Reza, et al.
Publicado: (2024)
Design and Performance Analysis of an Ultra-Low Power Integrate-and-Fire Neuron Circuit Using Nanoscale Side-contacted Field Effect Diode Technology
por: Motaman, Seyedmohamadjavad, et al.
Publicado: (2024)
por: Motaman, Seyedmohamadjavad, et al.
Publicado: (2024)
Ultra-Low-Power Spiking Neurons in 7 nm FinFET Technology: A Comparative Analysis of Leaky Integrate-and-Fire, Morris-Lecar, and Axon-Hillock Architectures
por: Larsh, Logan, et al.
Publicado: (2025)
por: Larsh, Logan, et al.
Publicado: (2025)
Group Distributionally Robust Optimization-Driven Reinforcement Learning for LLM Reasoning
por: Panaganti, Kishan, et al.
Publicado: (2026)
por: Panaganti, Kishan, et al.
Publicado: (2026)
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
por: Zhang, Kongcheng, et al.
Publicado: (2025)
por: Zhang, Kongcheng, et al.
Publicado: (2025)
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search
por: Shen, Maohao, et al.
Publicado: (2025)
por: Shen, Maohao, et al.
Publicado: (2025)
Unifying Tree Search Algorithm and Reward Design for LLM Reasoning: A Survey
por: Wei, Jiaqi, et al.
Publicado: (2025)
por: Wei, Jiaqi, et al.
Publicado: (2025)
A Unified Phase-native Computational Principle Governs Hippocampal Spike Timing and Neural Coding
por: Ahmadvand, Reza, et al.
Publicado: (2026)
por: Ahmadvand, Reza, et al.
Publicado: (2026)
Investigating the Effect of Electrical and Thermal Transport Properties on Oxide-Based Memristors Performance and Reliability
por: Gooran-Shoorakchaly, Armin, et al.
Publicado: (2024)
por: Gooran-Shoorakchaly, Armin, et al.
Publicado: (2024)
Dual Micro-Ring Resonators with Angular GST Modulation: Enabling Ultra-Fast Nonlinear Activation for Neuromorphic Photonics
por: Karimkhani, Hossein, et al.
Publicado: (2025)
por: Karimkhani, Hossein, et al.
Publicado: (2025)
Accelerated Discovery of Vanadium Oxide Compositions: A WGAN-VAE Framework for Materials Design
por: Ebrahimzadeh, Danial, et al.
Publicado: (2025)
por: Ebrahimzadeh, Danial, et al.
Publicado: (2025)
Design of Tunable Perfect Absorbers in the Mid-IR Spectrum Using Graphene-Based Multilayer Structures: Emerging Applications in Atmospheric Window Matching
por: Nazari, Masoumeh, et al.
Publicado: (2024)
por: Nazari, Masoumeh, et al.
Publicado: (2024)
Synthesizability Prediction of Crystalline Structures with a Hierarchical Transformer and Uncertainty Quantification
por: Ebrahimzadeh, Danial, et al.
Publicado: (2025)
por: Ebrahimzadeh, Danial, et al.
Publicado: (2025)
Teaching LLM to be Persuasive: Reward-Enhanced Policy Optimization for Alignment from Heterogeneous Rewards
por: Zeng, Xia, et al.
Publicado: (2025)
por: Zeng, Xia, et al.
Publicado: (2025)
InfoFlow: Reinforcing Search Agent Via Reward Density Optimization
por: Luo, Kun, et al.
Publicado: (2025)
por: Luo, Kun, et al.
Publicado: (2025)
Metacognition as Reward: Reinforcing LLM Reasoning via Knowledge and Regulation Signals
por: Chen, Sirui, et al.
Publicado: (2026)
por: Chen, Sirui, et al.
Publicado: (2026)
From Sight to Insight: Improving Visual Reasoning Capabilities of Multimodal Models via Reinforcement Learning
por: Sharif, Omar, et al.
Publicado: (2026)
por: Sharif, Omar, et al.
Publicado: (2026)
On-Chip Learning with Memristor-Based Neural Networks: Assessing Accuracy and Efficiency Under Device Variations, Conductance Errors, and Input Noise
por: Eslami, M. Reza, et al.
Publicado: (2024)
por: Eslami, M. Reza, et al.
Publicado: (2024)
SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning
por: Wan, Zhongwei, et al.
Publicado: (2025)
por: Wan, Zhongwei, et al.
Publicado: (2025)
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
por: Jin, Bowen, et al.
Publicado: (2025)
por: Jin, Bowen, et al.
Publicado: (2025)
Countering Reward Over-optimization in LLM with Demonstration-Guided Reinforcement Learning
por: Rita, Mathieu, et al.
Publicado: (2024)
por: Rita, Mathieu, et al.
Publicado: (2024)
Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning
por: Liu, Qihao, et al.
Publicado: (2025)
por: Liu, Qihao, et al.
Publicado: (2025)
MRO: Enhancing Reasoning in Diffusion Language Models via Multi-Reward Optimization
por: Wang, Chenglong, et al.
Publicado: (2025)
por: Wang, Chenglong, et al.
Publicado: (2025)
From Roots to Rewards: Dynamic Tree Reasoning with Reinforcement Learning
por: Bahloul, Ahmed, et al.
Publicado: (2025)
por: Bahloul, Ahmed, et al.
Publicado: (2025)
ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning
por: Chen, Mingyang, et al.
Publicado: (2025)
por: Chen, Mingyang, et al.
Publicado: (2025)
ConvSearch-R1: Enhancing Query Reformulation for Conversational Search with Reasoning via Reinforcement Learning
por: Zhu, Changtai, et al.
Publicado: (2025)
por: Zhu, Changtai, et al.
Publicado: (2025)
Toward Digital Twins in 3D IC Packaging: A Critical Review of Physics, Data, and Hybrid Architectures
por: Datta, Gourab, et al.
Publicado: (2026)
por: Datta, Gourab, et al.
Publicado: (2026)
Novel Pigeon-inspired 3D Obstacle Detection and Avoidance Maneuver for Multi-UAV Systems
por: Ahmadvand, Reza, et al.
Publicado: (2025)
por: Ahmadvand, Reza, et al.
Publicado: (2025)
TransMatch: A Transfer-Learning Framework for Defect Detection in Laser Powder Bed Fusion Additive Manufacturing
por: Ilani, Mohsen Asghari, et al.
Publicado: (2025)
por: Ilani, Mohsen Asghari, et al.
Publicado: (2025)
Ejemplares similares
-
Improving Aviation Safety Analysis: Automated HFACS Classification Using Reinforcement Learning with Group Relative Policy Optimization
por: Ahmadi, Arash, et al.
Publicado: (2025) -
A Comparative Study of Sampling Methods with Cross-Validation in the FedHome Framework
por: Ahmadi, Arash, et al.
Publicado: (2024) -
MCP Bridge: A Lightweight, LLM-Agnostic RESTful Proxy for Model Context Protocol Servers
por: Ahmadi, Arash, et al.
Publicado: (2025) -
A Cloud-Edge Framework for Energy-Efficient Event-Driven Control: An Integration of Online Supervised Learning, Spiking Neural Networks and Local Plasticity Rules
por: Ahmadvand, Reza, et al.
Publicado: (2024) -
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning
por: Zhao, Qingfei, et al.
Publicado: (2025)