Most Likely Sequence Generation for $n$-Grams, Transformers, HMMs, and Markov Chains, by Using Rollout Algorithms
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Yuchao, Bertsekas, Dimitri |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Superior Computer Chess with Model Predictive Control, Reinforcement Learning, and Rollout
di: Gundawar, Atharva, et al.
Pubblicazione: (2024)
di: Gundawar, Atharva, et al.
Pubblicazione: (2024)
Model Predictive Control and Reinforcement Learning: A Unified Framework Based on Dynamic Programming
di: Bertsekas, Dimitri P.
Pubblicazione: (2024)
di: Bertsekas, Dimitri P.
Pubblicazione: (2024)
Feature-Based Belief Aggregation for Partially Observable Markov Decision Problems
di: Li, Yuchao, et al.
Pubblicazione: (2025)
di: Li, Yuchao, et al.
Pubblicazione: (2025)
An Error Bound for Aggregation in Approximate Dynamic Programming
di: Li, Yuchao, et al.
Pubblicazione: (2025)
di: Li, Yuchao, et al.
Pubblicazione: (2025)
Adaptive Network Security Policies via Belief Aggregation and Rollout
di: Hammar, Kim, et al.
Pubblicazione: (2025)
di: Hammar, Kim, et al.
Pubblicazione: (2025)
Fine-tuning Smaller Language Models for Question Answering over Financial Documents
di: Phogat, Karmvir Singh, et al.
Pubblicazione: (2024)
di: Phogat, Karmvir Singh, et al.
Pubblicazione: (2024)
Are More LLM Calls All You Need? Towards Scaling Laws of Compound Inference Systems
di: Chen, Lingjiao, et al.
Pubblicazione: (2024)
di: Chen, Lingjiao, et al.
Pubblicazione: (2024)
LogicGuard: Improving Embodied LLM agents through Temporal Logic based Critics
di: Gokhale, Anand, et al.
Pubblicazione: (2025)
di: Gokhale, Anand, et al.
Pubblicazione: (2025)
Language Models as Efficient Reward Function Searchers for Custom-Environment Multi-Objective Reinforcement
di: Xie, Guanwen, et al.
Pubblicazione: (2024)
di: Xie, Guanwen, et al.
Pubblicazione: (2024)
PreFT: Prefill-only finetuning for efficient inference
di: Lanpouthakoun, Andrew, et al.
Pubblicazione: (2026)
di: Lanpouthakoun, Andrew, et al.
Pubblicazione: (2026)
A Survey on Large Language Model-empowered Autonomous Driving
di: Zhu, Yuxuan, et al.
Pubblicazione: (2024)
di: Zhu, Yuxuan, et al.
Pubblicazione: (2024)
ACING: Actor-Critic for Instruction Learning in Black-Box LLMs
di: Kharrat, Salma, et al.
Pubblicazione: (2024)
di: Kharrat, Salma, et al.
Pubblicazione: (2024)
ControlAgent: Automating Control System Design via Novel Integration of LLM Agents and Domain Expertise
di: Guo, Xingang, et al.
Pubblicazione: (2024)
di: Guo, Xingang, et al.
Pubblicazione: (2024)
Context, Reasoning, and Hierarchy: A Cost-Performance Study of Compound LLM Agent Design in an Adversarial POMDP
di: Bogdanov, Igor, et al.
Pubblicazione: (2026)
di: Bogdanov, Igor, et al.
Pubblicazione: (2026)
FORGE: Self-Evolving Agent Memory With No Weight Updates via Population Broadcast
di: Bogdanov, Igor, et al.
Pubblicazione: (2026)
di: Bogdanov, Igor, et al.
Pubblicazione: (2026)
FALCON: Autonomous Cyber Threat Intelligence Mining with LLMs for IDS Rule Generation
di: Mitra, Shaswata, et al.
Pubblicazione: (2025)
di: Mitra, Shaswata, et al.
Pubblicazione: (2025)
A Unified Generative-AI Framework for Smart Energy Infrastructure: Intelligent Gas Distribution, Utility Billing, Carbon Analytics, and Quantum-Inspired Optimisation
di: Manjunath, Pavan, et al.
Pubblicazione: (2026)
di: Manjunath, Pavan, et al.
Pubblicazione: (2026)
Preference Adaptive and Sequential Text-to-Image Generation
di: Nabati, Ofir, et al.
Pubblicazione: (2024)
di: Nabati, Ofir, et al.
Pubblicazione: (2024)
On Limitation of Transformer for Learning HMMs
di: Hu, Jiachen, et al.
Pubblicazione: (2024)
di: Hu, Jiachen, et al.
Pubblicazione: (2024)
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
di: Xu, Yixuan Even, et al.
Pubblicazione: (2025)
di: Xu, Yixuan Even, et al.
Pubblicazione: (2025)
Fake-Mamba: Real-Time Speech Deepfake Detection Using Bidirectional Mamba as Self-Attention's Alternative
di: Xuan, Xi, et al.
Pubblicazione: (2025)
di: Xuan, Xi, et al.
Pubblicazione: (2025)
Generating Chain-of-Thoughts with a Pairwise-Comparison Approach to Searching for the Most Promising Intermediate Thought
di: Zhang, Zhen-Yu, et al.
Pubblicazione: (2024)
di: Zhang, Zhen-Yu, et al.
Pubblicazione: (2024)
SciNav: A General Agent Framework for Scientific Coding Tasks
di: Zhang, Tianshu, et al.
Pubblicazione: (2026)
di: Zhang, Tianshu, et al.
Pubblicazione: (2026)
Semilinear Dynamic Programming: Analysis, Algorithms, and Certainty Equivalence Properties
di: Li, Yuchao, et al.
Pubblicazione: (2025)
di: Li, Yuchao, et al.
Pubblicazione: (2025)
Near Optimal Convergence to Coarse Correlated Equilibrium in General-Sum Markov Games
di: Yorulmaz, Asrin Efe, et al.
Pubblicazione: (2025)
di: Yorulmaz, Asrin Efe, et al.
Pubblicazione: (2025)
Transformers as Support Vector Machines
di: Tarzanagh, Davoud Ataee, et al.
Pubblicazione: (2023)
di: Tarzanagh, Davoud Ataee, et al.
Pubblicazione: (2023)
OCMDP: Observation-Constrained Markov Decision Process
di: Wang, Taiyi, et al.
Pubblicazione: (2024)
di: Wang, Taiyi, et al.
Pubblicazione: (2024)
1-2-3-Go! Policy Synthesis for Parameterized Markov Decision Processes via Decision-Tree Learning and Generalization
di: Azeem, Muqsit, et al.
Pubblicazione: (2024)
di: Azeem, Muqsit, et al.
Pubblicazione: (2024)
MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs
di: Mitra, Purbesh, et al.
Pubblicazione: (2025)
di: Mitra, Purbesh, et al.
Pubblicazione: (2025)
Mitigating Backdoor Threats to Large Language Models: Advancement and Challenges
di: Liu, Qin, et al.
Pubblicazione: (2024)
di: Liu, Qin, et al.
Pubblicazione: (2024)
Conformal Off-Policy Evaluation in Markov Decision Processes
di: Foffano, Daniele, et al.
Pubblicazione: (2023)
di: Foffano, Daniele, et al.
Pubblicazione: (2023)
Large Language Models as Markov Chains
di: Zekri, Oussama, et al.
Pubblicazione: (2024)
di: Zekri, Oussama, et al.
Pubblicazione: (2024)
From Self-Attention to Markov Models: Unveiling the Dynamics of Generative Transformers
di: Ildiz, M. Emrullah, et al.
Pubblicazione: (2024)
di: Ildiz, M. Emrullah, et al.
Pubblicazione: (2024)
On-Line Policy Iteration with Trajectory-Driven Policy Generation
di: Li, Yuchao, et al.
Pubblicazione: (2026)
di: Li, Yuchao, et al.
Pubblicazione: (2026)
GenSafe: A Generalizable Safety Enhancer for Safe Reinforcement Learning Algorithms Based on Reduced Order Markov Decision Process Model
di: Zhou, Zhehua, et al.
Pubblicazione: (2024)
di: Zhou, Zhehua, et al.
Pubblicazione: (2024)
Multi-Modal Drift Forecasting of Leeway Objects via Navier-Stokes-Guided CNN and Sequence-to-Sequence Attention-Based Models
di: Adesunkanmi, Rahmat K., et al.
Pubblicazione: (2025)
di: Adesunkanmi, Rahmat K., et al.
Pubblicazione: (2025)
Unsupervised Detection of Spatiotemporal Anomalies in PMU Data Using Transformer-Based BiGAN
di: Hossain, Muhammad Imran, et al.
Pubblicazione: (2025)
di: Hossain, Muhammad Imran, et al.
Pubblicazione: (2025)
Hierarchical Neuro-Symbolic Decision Transformer
di: Baheri, Ali, et al.
Pubblicazione: (2025)
di: Baheri, Ali, et al.
Pubblicazione: (2025)
An Offline Risk-aware Policy Selection Method for Bayesian Markov Decision Processes
di: Angelotti, Giorgio, et al.
Pubblicazione: (2021)
di: Angelotti, Giorgio, et al.
Pubblicazione: (2021)
On the Convergence of Modified Policy Iteration in Risk Sensitive Exponential Cost Markov Decision Processes
di: Murthy, Yashaswini, et al.
Pubblicazione: (2023)
di: Murthy, Yashaswini, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Superior Computer Chess with Model Predictive Control, Reinforcement Learning, and Rollout
di: Gundawar, Atharva, et al.
Pubblicazione: (2024) -
Model Predictive Control and Reinforcement Learning: A Unified Framework Based on Dynamic Programming
di: Bertsekas, Dimitri P.
Pubblicazione: (2024) -
Feature-Based Belief Aggregation for Partially Observable Markov Decision Problems
di: Li, Yuchao, et al.
Pubblicazione: (2025) -
An Error Bound for Aggregation in Approximate Dynamic Programming
di: Li, Yuchao, et al.
Pubblicazione: (2025) -
Adaptive Network Security Policies via Belief Aggregation and Rollout
di: Hammar, Kim, et al.
Pubblicazione: (2025)