Synthetic POMDPs to Challenge Memory-Augmented RL: Memory Demand Structure Modeling
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Yongyi, Li, Lingfeng, Chen, Bozhou, Li, Ang, Liu, Hanyu, Zheng, Qirui, Yang, Xionghui, Li, Wenxin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Decoupling Return-to-Go for Efficient Decision Transformer
di: Wang, Yongyi, et al.
Pubblicazione: (2026)
di: Wang, Yongyi, et al.
Pubblicazione: (2026)
Beyond Autoregressive RTG: Conditioning via Injection Outside Sequential Modeling in Decision Transformer
di: Wang, Yongyi, et al.
Pubblicazione: (2026)
di: Wang, Yongyi, et al.
Pubblicazione: (2026)
ShuttleEnv: An Interactive Data-Driven RL Environment for Badminton Strategy Modeling
di: Li, Ang, et al.
Pubblicazione: (2026)
di: Li, Ang, et al.
Pubblicazione: (2026)
BotzoneBench: Scalable LLM Evaluation via Graded AI Anchors
di: Li, Lingfeng, et al.
Pubblicazione: (2026)
di: Li, Lingfeng, et al.
Pubblicazione: (2026)
Style-Preserving Policy Optimization for Game Agents
di: Li, Lingfeng, et al.
Pubblicazione: (2025)
di: Li, Lingfeng, et al.
Pubblicazione: (2025)
Mxplainer: Explain and Learn Insights by Imitating Mahjong Agents
di: Li, Lingfeng, et al.
Pubblicazione: (2025)
di: Li, Lingfeng, et al.
Pubblicazione: (2025)
Constructing Non-Markovian Decision Process via History Aggregator
di: Wang, Yongyi, et al.
Pubblicazione: (2025)
di: Wang, Yongyi, et al.
Pubblicazione: (2025)
Robust Finite-Memory Policy Gradients for Hidden-Model POMDPs
di: Galesloot, Maris F. L., et al.
Pubblicazione: (2025)
di: Galesloot, Maris F. L., et al.
Pubblicazione: (2025)
Adapting Rules of Official International Mahjong for Online Players
di: Wang, Chucai, et al.
Pubblicazione: (2026)
di: Wang, Chucai, et al.
Pubblicazione: (2026)
HAGE: Harnessing Agentic Memory via RL-Driven Weighted Graph Evolution
di: Jiang, Dongming, et al.
Pubblicazione: (2026)
di: Jiang, Dongming, et al.
Pubblicazione: (2026)
Generalizing Multi-Step Inverse Models for Representation Learning to Finite-Memory POMDPs
di: Wu, Lili, et al.
Pubblicazione: (2024)
di: Wu, Lili, et al.
Pubblicazione: (2024)
MindBridge: Scalable and Cross-Model Knowledge Editing via Memory-Augmented Modality
di: Li, Shuaike, et al.
Pubblicazione: (2025)
di: Li, Shuaike, et al.
Pubblicazione: (2025)
MemLong: Memory-Augmented Retrieval for Long Text Modeling
di: Liu, Weijie, et al.
Pubblicazione: (2024)
di: Liu, Weijie, et al.
Pubblicazione: (2024)
Scalable Policy-Based RL Algorithms for POMDPs
di: Anjarlekar, Ameya, et al.
Pubblicazione: (2025)
di: Anjarlekar, Ameya, et al.
Pubblicazione: (2025)
R-Debater: Retrieval-Augmented Debate Generation through Argumentative Memory
di: Li, Maoyuan, et al.
Pubblicazione: (2025)
di: Li, Maoyuan, et al.
Pubblicazione: (2025)
Retrieval-Augmented Decision Transformer: External Memory for In-context RL
di: Schmied, Thomas, et al.
Pubblicazione: (2024)
di: Schmied, Thomas, et al.
Pubblicazione: (2024)
Explainable Representation of Finite-Memory Policies for POMDPs using Decision Trees
di: Azeem, Muqsit, et al.
Pubblicazione: (2024)
di: Azeem, Muqsit, et al.
Pubblicazione: (2024)
SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety
di: Liu, Zhe, et al.
Pubblicazione: (2026)
di: Liu, Zhe, et al.
Pubblicazione: (2026)
Investigating Memory in Model-Free RL with POPGym Arcade
di: Wang, Zekang, et al.
Pubblicazione: (2025)
di: Wang, Zekang, et al.
Pubblicazione: (2025)
GRAVITY: Architecture-Agnostic Structured Anchoring for Long-Horizon Conversational Memory
di: Sun, Yushi, et al.
Pubblicazione: (2026)
di: Sun, Yushi, et al.
Pubblicazione: (2026)
Memory in Large Language Models: Mechanisms, Evaluation and Evolution
di: Zhang, Dianxing, et al.
Pubblicazione: (2025)
di: Zhang, Dianxing, et al.
Pubblicazione: (2025)
MemPO: Self-Memory Policy Optimization for Long-Horizon Agents
di: Li, Ruoran, et al.
Pubblicazione: (2026)
di: Li, Ruoran, et al.
Pubblicazione: (2026)
Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization
di: Liu, Zeyuan, et al.
Pubblicazione: (2026)
di: Liu, Zeyuan, et al.
Pubblicazione: (2026)
Scalable Solution Methods for Dec-POMDPs with Deterministic Dynamics
di: You, Yang, et al.
Pubblicazione: (2025)
di: You, Yang, et al.
Pubblicazione: (2025)
Contrastive Augmented Graph2Graph Memory Interaction for Few Shot Continual Learning
di: Qi, Biqing, et al.
Pubblicazione: (2024)
di: Qi, Biqing, et al.
Pubblicazione: (2024)
Pareto-guided Pipeline for Distilling Featherweight AI Agents in Mobile MOBA Games
di: Yang, Xionghui, et al.
Pubblicazione: (2026)
di: Yang, Xionghui, et al.
Pubblicazione: (2026)
Boundary-Guided Policy Optimization for Memory-efficient RL of Diffusion Large Language Models
di: Lin, Nianyi, et al.
Pubblicazione: (2025)
di: Lin, Nianyi, et al.
Pubblicazione: (2025)
MemoryMamba: Memory-Augmented State Space Model for Defect Recognition
di: Wang, Qianning, et al.
Pubblicazione: (2024)
di: Wang, Qianning, et al.
Pubblicazione: (2024)
ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators
di: Zou, Guoqiang, et al.
Pubblicazione: (2025)
di: Zou, Guoqiang, et al.
Pubblicazione: (2025)
LifeAlign: Lifelong Alignment for Large Language Models with Memory-Augmented Focalized Preference Optimization
di: Li, Junsong, et al.
Pubblicazione: (2025)
di: Li, Junsong, et al.
Pubblicazione: (2025)
Generalizable End-to-End Tool-Use RL with Synthetic CodeGym
di: Du, Weihua, et al.
Pubblicazione: (2025)
di: Du, Weihua, et al.
Pubblicazione: (2025)
Automated Reformulation of Robust Optimization via Memory-Augmented Large Language Models
di: Chen, Jinbiao, et al.
Pubblicazione: (2026)
di: Chen, Jinbiao, et al.
Pubblicazione: (2026)
MemoryKT: An Integrative Memory-and-Forgetting Method for Knowledge Tracing
di: Lin, Mingrong, et al.
Pubblicazione: (2025)
di: Lin, Mingrong, et al.
Pubblicazione: (2025)
State Contamination in Memory-Augmented LLM Agents
di: Wang, Yian, et al.
Pubblicazione: (2026)
di: Wang, Yian, et al.
Pubblicazione: (2026)
MemRec: Collaborative Memory-Augmented Agentic Recommender System
di: Chen, Weixin, et al.
Pubblicazione: (2026)
di: Chen, Weixin, et al.
Pubblicazione: (2026)
Symmetry-Guided Memory Augmentation for Efficient Locomotion Learning
di: Bao, Kaixi, et al.
Pubblicazione: (2025)
di: Bao, Kaixi, et al.
Pubblicazione: (2025)
Memory, Consciousness and Large Language Model
di: Li, Jitang, et al.
Pubblicazione: (2024)
di: Li, Jitang, et al.
Pubblicazione: (2024)
MetaState: Persistent Working Memory Enhances Reasoning in Discrete Diffusion Language Models
di: Xia, Kejing, et al.
Pubblicazione: (2026)
di: Xia, Kejing, et al.
Pubblicazione: (2026)
$δ$-mem: Efficient Online Memory for Large Language Models
di: Lei, Jingdi, et al.
Pubblicazione: (2026)
di: Lei, Jingdi, et al.
Pubblicazione: (2026)
Coinvisor: An RL-Enhanced Chatbot Agent for Interactive Cryptocurrency Investment Analysis
di: Chen, Chong, et al.
Pubblicazione: (2025)
di: Chen, Chong, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Decoupling Return-to-Go for Efficient Decision Transformer
di: Wang, Yongyi, et al.
Pubblicazione: (2026) -
Beyond Autoregressive RTG: Conditioning via Injection Outside Sequential Modeling in Decision Transformer
di: Wang, Yongyi, et al.
Pubblicazione: (2026) -
ShuttleEnv: An Interactive Data-Driven RL Environment for Badminton Strategy Modeling
di: Li, Ang, et al.
Pubblicazione: (2026) -
BotzoneBench: Scalable LLM Evaluation via Graded AI Anchors
di: Li, Lingfeng, et al.
Pubblicazione: (2026) -
Style-Preserving Policy Optimization for Game Agents
di: Li, Lingfeng, et al.
Pubblicazione: (2025)