Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Akgül, Ömer Faruk, Kannan, Rajgopal, Neiswanger, Willie, Prasanna, Viktor |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LYNX: Learning Dynamic Exits for Confidence-Controlled Reasoning
von: Akgül, Ömer Faruk, et al.
Veröffentlicht: (2025)
von: Akgül, Ömer Faruk, et al.
Veröffentlicht: (2025)
Conformal Prediction for Federated Graph Neural Networks with Missing Neighbor Information
von: Akgül, Ömer Faruk, et al.
Veröffentlicht: (2024)
von: Akgül, Ömer Faruk, et al.
Veröffentlicht: (2024)
RECIPE-TKG: From Sparse History to Structured Reasoning for LLM-based Temporal Knowledge Graph Completion
von: Akgül, Ömer Faruk, et al.
Veröffentlicht: (2025)
von: Akgül, Ömer Faruk, et al.
Veröffentlicht: (2025)
Tina: Tiny Reasoning Models via LoRA
von: Wang, Shangshang, et al.
Veröffentlicht: (2025)
von: Wang, Shangshang, et al.
Veröffentlicht: (2025)
Resa: Transparent Reasoning Models via SAEs
von: Wang, Shangshang, et al.
Veröffentlicht: (2025)
von: Wang, Shangshang, et al.
Veröffentlicht: (2025)
HierRouter: Coordinated Routing of Specialized Large Language Models via Reinforcement Learning
von: Gupta, Nikunj, et al.
Veröffentlicht: (2025)
von: Gupta, Nikunj, et al.
Veröffentlicht: (2025)
Accelerating Sparse MTTKRP for Small Tensor Decomposition on GPU
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
Sparse MTTKRP Acceleration for Tensor Decomposition on GPU
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2024)
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2024)
AMPED: Accelerating MTTKRP for Billion-Scale Sparse Tensor Decomposition on Multiple GPUs
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
Training Diverse Graph Experts for Ensembles: A Systematic Empirical Study
von: Deng, Gangda, et al.
Veröffentlicht: (2025)
von: Deng, Gangda, et al.
Veröffentlicht: (2025)
Adversarial Training in Low-Label Regimes with Margin-Based Interpolation
von: Ye, Tian, et al.
Veröffentlicht: (2024)
von: Ye, Tian, et al.
Veröffentlicht: (2024)
From Calibration to Collaboration: LLM Uncertainty Quantification Should Be More Human-Centered
von: Devic, Siddartha, et al.
Veröffentlicht: (2025)
von: Devic, Siddartha, et al.
Veröffentlicht: (2025)
Do LLM-derived graph priors improve multi-agent coordination?
von: Gupta, Nikunj, et al.
Veröffentlicht: (2026)
von: Gupta, Nikunj, et al.
Veröffentlicht: (2026)
FACTUAL: A Novel Framework for Contrastive Learning Based Robust SAR Image Classification
von: Wang, Xu, et al.
Veröffentlicht: (2024)
von: Wang, Xu, et al.
Veröffentlicht: (2024)
Studying the Effects of Self-Attention on SAR Automatic Target Recognition
von: Fein-Ashley, Jacob, et al.
Veröffentlicht: (2024)
von: Fein-Ashley, Jacob, et al.
Veröffentlicht: (2024)
LLM Unlearning Without an Expert Curated Dataset
von: Zhu, Xiaoyuan, et al.
Veröffentlicht: (2025)
von: Zhu, Xiaoyuan, et al.
Veröffentlicht: (2025)
GCV-Turbo: End-to-end Acceleration of GNN-based Computer Vision Tasks on FPGA
von: Zhang, Bingyi, et al.
Veröffentlicht: (2024)
von: Zhang, Bingyi, et al.
Veröffentlicht: (2024)
Contextual Feedback Loops: Amplifying Deep Reasoning with Iterative Top-Down Feedback
von: Fein-Ashley, Jacob, et al.
Veröffentlicht: (2024)
von: Fein-Ashley, Jacob, et al.
Veröffentlicht: (2024)
ConsensusDrop: Fusing Visual and Cross-Modal Saliency for Efficient Vision Language Models
von: Parikh, Dhruv, et al.
Veröffentlicht: (2026)
von: Parikh, Dhruv, et al.
Veröffentlicht: (2026)
Uncertainty-Aware SAR ATR: Defending Against Adversarial Attacks via Bayesian Neural Networks
von: Ye, Tian, et al.
Veröffentlicht: (2024)
von: Ye, Tian, et al.
Veröffentlicht: (2024)
DeLLMa: Decision Making Under Uncertainty with Large Language Models
von: Liu, Ollie, et al.
Veröffentlicht: (2024)
von: Liu, Ollie, et al.
Veröffentlicht: (2024)
SPARC-RAG: Adaptive Sequential-Parallel Scaling with Context Management for Retrieval-Augmented Generation
von: Yang, Yuxin, et al.
Veröffentlicht: (2026)
von: Yang, Yuxin, et al.
Veröffentlicht: (2026)
Latent Denoising Improves Visual Alignment in Large Multimodal Models
von: Parikh, Dhruv, et al.
Veröffentlicht: (2026)
von: Parikh, Dhruv, et al.
Veröffentlicht: (2026)
Adversarial Knapsack for Sequential Competitive Resource Allocation
von: Thakoor, Omkar, et al.
Veröffentlicht: (2025)
von: Thakoor, Omkar, et al.
Veröffentlicht: (2025)
Euclid: Supercharging Multimodal LLMs with Synthetic High-Fidelity Visual Descriptions
von: Zhang, Jiarui, et al.
Veröffentlicht: (2024)
von: Zhang, Jiarui, et al.
Veröffentlicht: (2024)
Attention, Distillation, and Tabularization: Towards Practical Neural Network-Based Prefetching
von: Zhang, Pengmiao, et al.
Veröffentlicht: (2023)
von: Zhang, Pengmiao, et al.
Veröffentlicht: (2023)
Estimating LLM Consistency: A User Baseline vs Surrogate Metrics
von: Wu, Xiaoyuan, et al.
Veröffentlicht: (2025)
von: Wu, Xiaoyuan, et al.
Veröffentlicht: (2025)
Benchmarking Deep Learning Classifiers for SAR Automatic Target Recognition
von: Fein-Ashley, Jacob, et al.
Veröffentlicht: (2023)
von: Fein-Ashley, Jacob, et al.
Veröffentlicht: (2023)
Action-Graph Policies: Learning Action Co-dependencies in Multi-Agent Reinforcement Learning
von: Gupta, Nikunj, et al.
Veröffentlicht: (2026)
von: Gupta, Nikunj, et al.
Veröffentlicht: (2026)
TypeBandit: Type-Level Context Allocation and Reweighting for Effective Attribute Completion in Heterogeneous Graph Neural Networks
von: Wang, Ta-Yang, et al.
Veröffentlicht: (2026)
von: Wang, Ta-Yang, et al.
Veröffentlicht: (2026)
FAME: FPGA Acceleration of Secure Matrix Multiplication with Homomorphic Encryption
von: Xu, Zhihan, et al.
Veröffentlicht: (2025)
von: Xu, Zhihan, et al.
Veröffentlicht: (2025)
ClusterViG: Efficient Globally Aware Vision GNNs via Image Partitioning
von: Parikh, Dhruv, et al.
Veröffentlicht: (2025)
von: Parikh, Dhruv, et al.
Veröffentlicht: (2025)
PAHD: Perception-Action based Human Decision Making using Explainable Graph Neural Networks on SAR Images
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2024)
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2024)
SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning
von: Limozin, Alexis, et al.
Veröffentlicht: (2026)
von: Limozin, Alexis, et al.
Veröffentlicht: (2026)
CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
AI-University: An LLM-based platform for instructional alignment to scientific classrooms
von: Shojaei, Mostafa Faghih, et al.
Veröffentlicht: (2025)
von: Shojaei, Mostafa Faghih, et al.
Veröffentlicht: (2025)
A Single Graph Convolution Is All You Need: Efficient Grayscale Image Classification
von: Fein-Ashley, Jacob, et al.
Veröffentlicht: (2024)
von: Fein-Ashley, Jacob, et al.
Veröffentlicht: (2024)
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search
von: Hou, Zhenyu, et al.
Veröffentlicht: (2025)
von: Hou, Zhenyu, et al.
Veröffentlicht: (2025)
Mixture of Thoughts: Learning to Aggregate What Experts Think, Not Just What They Say
von: Fein-Ashley, Jacob, et al.
Veröffentlicht: (2025)
von: Fein-Ashley, Jacob, et al.
Veröffentlicht: (2025)
Accelerating ViT Inference on FPGA through Static and Dynamic Pruning
von: Parikh, Dhruv, et al.
Veröffentlicht: (2024)
von: Parikh, Dhruv, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LYNX: Learning Dynamic Exits for Confidence-Controlled Reasoning
von: Akgül, Ömer Faruk, et al.
Veröffentlicht: (2025) -
Conformal Prediction for Federated Graph Neural Networks with Missing Neighbor Information
von: Akgül, Ömer Faruk, et al.
Veröffentlicht: (2024) -
RECIPE-TKG: From Sparse History to Structured Reasoning for LLM-based Temporal Knowledge Graph Completion
von: Akgül, Ömer Faruk, et al.
Veröffentlicht: (2025) -
Tina: Tiny Reasoning Models via LoRA
von: Wang, Shangshang, et al.
Veröffentlicht: (2025) -
Resa: Transparent Reasoning Models via SAEs
von: Wang, Shangshang, et al.
Veröffentlicht: (2025)