Expanding the Action Space of LLMs to Reason Beyond Language
Fuente:
arXiv
Salvato in:
| Autori principali: | Yue, Zhongqi, Wang, Weishi, Zhan, Yundaichuan, Li, Juncheng, Dahlmeier, Daniel, Johansson, Fredrik D. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
IncomeSCM: From tabular data set to time-series simulator and causal estimation benchmark
di: Johansson, Fredrik D.
Pubblicazione: (2024)
di: Johansson, Fredrik D.
Pubblicazione: (2024)
Latent Preference Bandits
di: Mwai, Newton, et al.
Pubblicazione: (2025)
di: Mwai, Newton, et al.
Pubblicazione: (2025)
Unsupervised domain adaptation by learning using privileged information
di: Breitholtz, Adam, et al.
Pubblicazione: (2023)
di: Breitholtz, Adam, et al.
Pubblicazione: (2023)
Latent Order Bandits
di: Carlsson, Emil, et al.
Pubblicazione: (2026)
di: Carlsson, Emil, et al.
Pubblicazione: (2026)
DynaAct: Large Language Model Reasoning with Dynamic Action Spaces
di: Zhao, Xueliang, et al.
Pubblicazione: (2025)
di: Zhao, Xueliang, et al.
Pubblicazione: (2025)
Beyond Alignment: Expanding Reasoning Capacity via Manifold-Reshaping Policy Optimization
di: Wang, Dayu, et al.
Pubblicazione: (2026)
di: Wang, Dayu, et al.
Pubblicazione: (2026)
CARE: Turning LLMs Into Causal Reasoning Expert
di: Dong, Juncheng, et al.
Pubblicazione: (2025)
di: Dong, Juncheng, et al.
Pubblicazione: (2025)
Federated Learning with Heterogeneous and Private Label Sets
di: Breitholtz, Adam, et al.
Pubblicazione: (2025)
di: Breitholtz, Adam, et al.
Pubblicazione: (2025)
Can LLMs Learn to Reason Robustly under Noisy Supervision?
di: Yang, Shenzhi, et al.
Pubblicazione: (2026)
di: Yang, Shenzhi, et al.
Pubblicazione: (2026)
Knowledge Boundary Discovery for Large Language Models
di: Wang, Ziquan, et al.
Pubblicazione: (2026)
di: Wang, Ziquan, et al.
Pubblicazione: (2026)
Prediction Models That Learn to Avoid Missing Values
di: Stempfle, Lena, et al.
Pubblicazione: (2025)
di: Stempfle, Lena, et al.
Pubblicazione: (2025)
Pure Exploration in Bandits with Linear Constraints
di: Carlsson, Emil, et al.
Pubblicazione: (2023)
di: Carlsson, Emil, et al.
Pubblicazione: (2023)
Active Preference Learning for Ordering Items In- and Out-of-sample
di: Bergström, Herman, et al.
Pubblicazione: (2024)
di: Bergström, Herman, et al.
Pubblicazione: (2024)
Overcoming label shift with target-aware federated learning
di: Zec, Edvin Listo, et al.
Pubblicazione: (2024)
di: Zec, Edvin Listo, et al.
Pubblicazione: (2024)
Beyond Test-Time Memory: State-Space Optimal Control for LLM Reasoning
di: Wang, Peihao, et al.
Pubblicazione: (2026)
di: Wang, Peihao, et al.
Pubblicazione: (2026)
Steerable Neural ODEs on Homogeneous Spaces
di: Andersdotter, Emma, et al.
Pubblicazione: (2026)
di: Andersdotter, Emma, et al.
Pubblicazione: (2026)
Reasoning in Transformers -- Mitigating Spurious Correlations and Reasoning Shortcuts
di: Enström, Daniel, et al.
Pubblicazione: (2024)
di: Enström, Daniel, et al.
Pubblicazione: (2024)
Beyond the Answer: Decoding the Behavior of LLMs as Scientific Reasoners
di: Pandey, Rohan, et al.
Pubblicazione: (2026)
di: Pandey, Rohan, et al.
Pubblicazione: (2026)
Identifiable Latent Bandits: Leveraging observational data for personalized decision-making
di: Balcıoğlu, Ahmet Zahid, et al.
Pubblicazione: (2024)
di: Balcıoğlu, Ahmet Zahid, et al.
Pubblicazione: (2024)
BioBridge: Bridging Proteins and Language for Enhanced Biological Reasoning with LLMs
di: Wang, Yujia, et al.
Pubblicazione: (2026)
di: Wang, Yujia, et al.
Pubblicazione: (2026)
OLR-WAA: Adaptive and Drift-Resilient Online Regression with Dynamic Weighted Averaging
di: Abu-Shaira, Mohammad, et al.
Pubblicazione: (2025)
di: Abu-Shaira, Mohammad, et al.
Pubblicazione: (2025)
Reasoning through Verifiable Forecast Actions: Consistency-Grounded RL for Financial LLMs
di: Chen, Jialin, et al.
Pubblicazione: (2026)
di: Chen, Jialin, et al.
Pubblicazione: (2026)
When are radiology reports useful for training medical image classifiers?
di: Bergström, Herman, et al.
Pubblicazione: (2025)
di: Bergström, Herman, et al.
Pubblicazione: (2025)
Handling missing values in clinical machine learning: Insights from an expert study
di: Stempfle, Lena, et al.
Pubblicazione: (2024)
di: Stempfle, Lena, et al.
Pubblicazione: (2024)
Fine Tuning Methods for Low-resource Languages
di: Bakkenes, Tim, et al.
Pubblicazione: (2025)
di: Bakkenes, Tim, et al.
Pubblicazione: (2025)
Pragmatic Policy Development via Interpretable Behavior Cloning
di: Matsson, Anton, et al.
Pubblicazione: (2025)
di: Matsson, Anton, et al.
Pubblicazione: (2025)
Reasoning Beyond Limits: Advances and Open Problems for LLMs
di: Ferrag, Mohamed Amine, et al.
Pubblicazione: (2025)
di: Ferrag, Mohamed Amine, et al.
Pubblicazione: (2025)
Document Intelligence in the Era of Large Language Models: A Survey
di: Wang, Weishi, et al.
Pubblicazione: (2025)
di: Wang, Weishi, et al.
Pubblicazione: (2025)
GENSR: Symbolic Regression Based in Equation Generative Space
di: Li, Qian, et al.
Pubblicazione: (2026)
di: Li, Qian, et al.
Pubblicazione: (2026)
FlashEvaluator: Expanding Search Space with Parallel Evaluation
di: Feng, Chao, et al.
Pubblicazione: (2026)
di: Feng, Chao, et al.
Pubblicazione: (2026)
Teleportation With Null Space Gradient Projection for Optimization Acceleration
di: Wu, Zihao, et al.
Pubblicazione: (2025)
di: Wu, Zihao, et al.
Pubblicazione: (2025)
Unveiling Statistical Significance of Online Regression over Multiple Datasets
di: Abu-Shaira, Mohammad, et al.
Pubblicazione: (2025)
di: Abu-Shaira, Mohammad, et al.
Pubblicazione: (2025)
Online Domain-aware LLM Decoding for Continual Domain Evolution
di: Abu-Shaira, Mohammad, et al.
Pubblicazione: (2026)
di: Abu-Shaira, Mohammad, et al.
Pubblicazione: (2026)
Diagnosing Shortcut-Induced Rigidity in Continual Learning: The Einstellung Rigidity Index (ERI)
di: Gu, Kai, et al.
Pubblicazione: (2025)
di: Gu, Kai, et al.
Pubblicazione: (2025)
Learning plug-in surrogate endpoints for randomized experiments
di: Margueritte, Alessandro-Umberto, et al.
Pubblicazione: (2026)
di: Margueritte, Alessandro-Umberto, et al.
Pubblicazione: (2026)
Continuous Reasoning for Vision-Language-Action
di: Wu, Yueh-Hua, et al.
Pubblicazione: (2026)
di: Wu, Yueh-Hua, et al.
Pubblicazione: (2026)
ConfClip: Confidence-Weighted and Clipped Reward for Reinforcement Learning in LLMs
di: Zhang, Bonan, et al.
Pubblicazione: (2025)
di: Zhang, Bonan, et al.
Pubblicazione: (2025)
Graph-Augmented LLMs for Personalized Health Insights: A Case Study in Sleep Analysis
di: Subramanian, Ajan, et al.
Pubblicazione: (2024)
di: Subramanian, Ajan, et al.
Pubblicazione: (2024)
Simple Policy Gradients for Reasoning with Diffusion Language Models
di: Zhan, Anthony
Pubblicazione: (2025)
di: Zhan, Anthony
Pubblicazione: (2025)
Co-Evolving Latent Action World Models
di: Wang, Yucen, et al.
Pubblicazione: (2025)
di: Wang, Yucen, et al.
Pubblicazione: (2025)
Documenti analoghi
-
IncomeSCM: From tabular data set to time-series simulator and causal estimation benchmark
di: Johansson, Fredrik D.
Pubblicazione: (2024) -
Latent Preference Bandits
di: Mwai, Newton, et al.
Pubblicazione: (2025) -
Unsupervised domain adaptation by learning using privileged information
di: Breitholtz, Adam, et al.
Pubblicazione: (2023) -
Latent Order Bandits
di: Carlsson, Emil, et al.
Pubblicazione: (2026) -
DynaAct: Large Language Model Reasoning with Dynamic Action Spaces
di: Zhao, Xueliang, et al.
Pubblicazione: (2025)