Efficient Computation of Blackwell Optimal Policies using Rational Functions
Fuente:
arXiv
Salvato in:
| Autori principali: | Mukherjee, Dibyangshu, Kalyanakrishnan, Shivaram |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Howard's Policy Iteration is Subexponential for Deterministic Markov Decision Problems with Rewards of Fixed Bit-size and Arbitrary Discount Factor
di: Mukherjee, Dibyangshu, et al.
Pubblicazione: (2025)
di: Mukherjee, Dibyangshu, et al.
Pubblicazione: (2025)
On-line Learning in Tree MDPs by Treating Policies as Bandit Arms
di: Shah, Anvay, et al.
Pubblicazione: (2026)
di: Shah, Anvay, et al.
Pubblicazione: (2026)
Scaling Inference-Efficient Language Models
di: Bian, Song, et al.
Pubblicazione: (2025)
di: Bian, Song, et al.
Pubblicazione: (2025)
Rao-Blackwellized POMDP Planning
di: Lee, Jiho, et al.
Pubblicazione: (2024)
di: Lee, Jiho, et al.
Pubblicazione: (2024)
Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs
di: Bian, Song, et al.
Pubblicazione: (2025)
di: Bian, Song, et al.
Pubblicazione: (2025)
Tesserae: Scalable Placement Policies for Deep Learning Workloads
di: Bian, Song, et al.
Pubblicazione: (2025)
di: Bian, Song, et al.
Pubblicazione: (2025)
A View of the Certainty-Equivalence Method for PAC RL as an Application of the Trajectory Tree Method
di: Kalyanakrishnan, Shivaram, et al.
Pubblicazione: (2025)
di: Kalyanakrishnan, Shivaram, et al.
Pubblicazione: (2025)
Performative Policy Gradient: Optimality in Performative Reinforcement Learning
di: Basu, Debabrota, et al.
Pubblicazione: (2025)
di: Basu, Debabrota, et al.
Pubblicazione: (2025)
Learning Optimal and Sample-Efficient Decision Policies with Guarantees
di: Shao, Daqian
Pubblicazione: (2026)
di: Shao, Daqian
Pubblicazione: (2026)
Learnable Game-theoretic Policy Optimization for Data-centric Self-explanation Rationalization
di: Zhao, Yunxiao, et al.
Pubblicazione: (2025)
di: Zhao, Yunxiao, et al.
Pubblicazione: (2025)
Evaluating CUDA Tile for AI Workloads on Hopper and Blackwell GPUs
di: Yadav, Divakar Kumar, et al.
Pubblicazione: (2026)
di: Yadav, Divakar Kumar, et al.
Pubblicazione: (2026)
EcoAlign: An Economically Rational Framework for Efficient LVLM Alignment
di: Cheng, Ruoxi, et al.
Pubblicazione: (2025)
di: Cheng, Ruoxi, et al.
Pubblicazione: (2025)
Implementing Rational Choice Functions with LLMs and Measuring their Alignment with User Preferences
di: Karnysheva, Anna, et al.
Pubblicazione: (2025)
di: Karnysheva, Anna, et al.
Pubblicazione: (2025)
ROI-Reasoning: Rational Optimization for Inference via Pre-Computation Meta-Cognition
di: Zhao, Muyang, et al.
Pubblicazione: (2026)
di: Zhao, Muyang, et al.
Pubblicazione: (2026)
Step-level Optimization for Efficient Computer-use Agents
di: Wei, Jinbiao, et al.
Pubblicazione: (2026)
di: Wei, Jinbiao, et al.
Pubblicazione: (2026)
Rethinking Optimal Verification Granularity for Compute-Efficient Test-Time Scaling
di: Chen, Hao Mark, et al.
Pubblicazione: (2025)
di: Chen, Hao Mark, et al.
Pubblicazione: (2025)
Rationality Check! Benchmarking the Rationality of Large Language Models
di: Zhou, Zhilun, et al.
Pubblicazione: (2025)
di: Zhou, Zhilun, et al.
Pubblicazione: (2025)
Test-time Verification via Optimal Transport: Coverage, ROC, & Sub-optimality
di: Mukherjee, Arpan, et al.
Pubblicazione: (2025)
di: Mukherjee, Arpan, et al.
Pubblicazione: (2025)
Using Common Random Numbers for Simulation-based Planning with Rollouts
di: Yadav, Sandarbh, et al.
Pubblicazione: (2026)
di: Yadav, Sandarbh, et al.
Pubblicazione: (2026)
Latent Modulated Function for Computational Optimal Continuous Image Representation
di: He, Zongyao, et al.
Pubblicazione: (2024)
di: He, Zongyao, et al.
Pubblicazione: (2024)
Uncovering the Computational Ingredients of Human-Like Representations in LLMs
di: Studdiford, Zach, et al.
Pubblicazione: (2025)
di: Studdiford, Zach, et al.
Pubblicazione: (2025)
Minimax Optimal and Computationally Efficient Algorithms for Distributionally Robust Offline Reinforcement Learning
di: Liu, Zhishuai, et al.
Pubblicazione: (2024)
di: Liu, Zhishuai, et al.
Pubblicazione: (2024)
POETS: Uncertainty-Aware LLM Optimization via Compute-Efficient Policy Ensembles
di: Menet, Nicolas, et al.
Pubblicazione: (2026)
di: Menet, Nicolas, et al.
Pubblicazione: (2026)
RV-Syn: Rational and Verifiable Mathematical Reasoning Data Synthesis based on Structured Function Library
di: Wang, Jiapeng, et al.
Pubblicazione: (2025)
di: Wang, Jiapeng, et al.
Pubblicazione: (2025)
What Limits Agentic Systems Efficiency?
di: Bian, Song, et al.
Pubblicazione: (2025)
di: Bian, Song, et al.
Pubblicazione: (2025)
Partial Policy Gradients for RL in LLMs
di: Mathur, Puneet, et al.
Pubblicazione: (2026)
di: Mathur, Puneet, et al.
Pubblicazione: (2026)
Optimal Policy Minimum Bayesian Risk
di: Astudillo, Ramón Fernandez, et al.
Pubblicazione: (2025)
di: Astudillo, Ramón Fernandez, et al.
Pubblicazione: (2025)
The Function-Representation Model of Computation
di: Ibias, Alfredo, et al.
Pubblicazione: (2024)
di: Ibias, Alfredo, et al.
Pubblicazione: (2024)
The AI Policy Module: Developing Computer Science Student Competency in AI Ethics and Policy
di: Weichert, James, et al.
Pubblicazione: (2025)
di: Weichert, James, et al.
Pubblicazione: (2025)
LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models
di: Chang, Tzu-Tao, et al.
Pubblicazione: (2025)
di: Chang, Tzu-Tao, et al.
Pubblicazione: (2025)
Are Protein Language Models Compute Optimal?
di: Serrano, Yaiza, et al.
Pubblicazione: (2024)
di: Serrano, Yaiza, et al.
Pubblicazione: (2024)
Functional Critics Are Essential for Actor-Critic: From Off-Policy Stability to Efficient Exploration
di: Bai, Qinxun, et al.
Pubblicazione: (2025)
di: Bai, Qinxun, et al.
Pubblicazione: (2025)
RationAnomaly: Log Anomaly Detection with Rationality via Chain-of-Thought and Reinforcement Learning
di: Xu, Song, et al.
Pubblicazione: (2025)
di: Xu, Song, et al.
Pubblicazione: (2025)
Private LLM Inference on Consumer Blackwell GPUs: A Practical Guide for Cost-Effective Local Deployment in SMEs
di: Knoop, Jonathan, et al.
Pubblicazione: (2026)
di: Knoop, Jonathan, et al.
Pubblicazione: (2026)
Sample and Computationally Efficient Continuous-Time Reinforcement Learning with General Function Approximation
di: Zhao, Runze, et al.
Pubblicazione: (2025)
di: Zhao, Runze, et al.
Pubblicazione: (2025)
Tabula: Efficiently Computing Nonlinear Activation Functions for Secure Neural Network Inference
di: Lam, Maximilian, et al.
Pubblicazione: (2022)
di: Lam, Maximilian, et al.
Pubblicazione: (2022)
Proximal Policy Optimization with Graph Neural Networks for Optimal Power Flow
di: López-Cardona, Ángela, et al.
Pubblicazione: (2022)
di: López-Cardona, Ángela, et al.
Pubblicazione: (2022)
PGT-I: Scaling Spatiotemporal GNNs with Memory-Efficient Distributed Training
di: Ockerman, Seth, et al.
Pubblicazione: (2025)
di: Ockerman, Seth, et al.
Pubblicazione: (2025)
On-Policy Supervised Fine-Tuning for Efficient Reasoning
di: Zhao, Anhao, et al.
Pubblicazione: (2026)
di: Zhao, Anhao, et al.
Pubblicazione: (2026)
Rational Inverse Reasoning
di: Zandonati, Ben, et al.
Pubblicazione: (2025)
di: Zandonati, Ben, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Howard's Policy Iteration is Subexponential for Deterministic Markov Decision Problems with Rewards of Fixed Bit-size and Arbitrary Discount Factor
di: Mukherjee, Dibyangshu, et al.
Pubblicazione: (2025) -
On-line Learning in Tree MDPs by Treating Policies as Bandit Arms
di: Shah, Anvay, et al.
Pubblicazione: (2026) -
Scaling Inference-Efficient Language Models
di: Bian, Song, et al.
Pubblicazione: (2025) -
Rao-Blackwellized POMDP Planning
di: Lee, Jiho, et al.
Pubblicazione: (2024) -
Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs
di: Bian, Song, et al.
Pubblicazione: (2025)