Generalized Policy Gradient with History-Aware Decision Transformer for Reliable Routing over Graph Signals
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wei, Xing, Wang, Yuanhang, Zhao, Duoxiang, Zhang, Zezhou, Qin, Hao, Ouyang, Yuqi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Graph-Aware Diffusion for Signal Generation
von: Rozada, Sergio, et al.
Veröffentlicht: (2025)
von: Rozada, Sergio, et al.
Veröffentlicht: (2025)
Policy Gradient for Robust Markov Decision Processes
von: Wang, Qiuhao, et al.
Veröffentlicht: (2024)
von: Wang, Qiuhao, et al.
Veröffentlicht: (2024)
GPG: Generalized Policy Gradient Theorem for Transformer-based Policies
von: Mao, Hangyu, et al.
Veröffentlicht: (2025)
von: Mao, Hangyu, et al.
Veröffentlicht: (2025)
Adjusting the Output of Decision Transformer with Action Gradient
von: Lin, Rui, et al.
Veröffentlicht: (2025)
von: Lin, Rui, et al.
Veröffentlicht: (2025)
MASPO: Unifying Gradient Utilization, Probability Mass, and Signal Reliability for Robust and Sample-Efficient LLM Reasoning
von: Fu, Xiaoliang, et al.
Veröffentlicht: (2026)
von: Fu, Xiaoliang, et al.
Veröffentlicht: (2026)
Uncertainty-Aware Decision Transformer for Stochastic Driving Environments
von: Li, Zenan, et al.
Veröffentlicht: (2023)
von: Li, Zenan, et al.
Veröffentlicht: (2023)
Is Your Explanation Reliable: Confidence-Aware Explanation on Graph Neural Networks
von: Zhang, Jiaxing, et al.
Veröffentlicht: (2025)
von: Zhang, Jiaxing, et al.
Veröffentlicht: (2025)
Online Finetuning Decision Transformers with Pure RL Gradients
von: Luo, Junkai, et al.
Veröffentlicht: (2026)
von: Luo, Junkai, et al.
Veröffentlicht: (2026)
Learning General Policies with Policy Gradient Methods
von: Ståhlberg, Simon, et al.
Veröffentlicht: (2025)
von: Ståhlberg, Simon, et al.
Veröffentlicht: (2025)
Reinforcement Learning Gradients as Vitamin for Online Finetuning Decision Transformers
von: Yan, Kai, et al.
Veröffentlicht: (2024)
von: Yan, Kai, et al.
Veröffentlicht: (2024)
RACER: Risk-Aware Calibrated Efficient Routing for Large Language Models
von: Hao, Sai, et al.
Veröffentlicht: (2026)
von: Hao, Sai, et al.
Veröffentlicht: (2026)
TRACE: Distilling Where It Matters via Token-Routed Self On-Policy Alignment
von: Wang, Jiaxuan, et al.
Veröffentlicht: (2026)
von: Wang, Jiaxuan, et al.
Veröffentlicht: (2026)
Signal-Adaptive Trust Regions for Gradient-Free Optimization of Recurrent Spiking Neural Networks
von: Li, Jinhao, et al.
Veröffentlicht: (2026)
von: Li, Jinhao, et al.
Veröffentlicht: (2026)
Learning Spatio-Temporal Dynamics for Trajectory Recovery via Time-Aware Transformer
von: Sun, Tian, et al.
Veröffentlicht: (2025)
von: Sun, Tian, et al.
Veröffentlicht: (2025)
Do Transformer World Models Give Better Policy Gradients?
von: Ma, Michel, et al.
Veröffentlicht: (2024)
von: Ma, Michel, et al.
Veröffentlicht: (2024)
Solving Robust Markov Decision Processes: Generic, Reliable, Efficient
von: Meggendorfer, Tobias, et al.
Veröffentlicht: (2024)
von: Meggendorfer, Tobias, et al.
Veröffentlicht: (2024)
Policy Gradient Algorithms with Monte Carlo Tree Learning for Non-Markov Decision Processes
von: Morimura, Tetsuro, et al.
Veröffentlicht: (2022)
von: Morimura, Tetsuro, et al.
Veröffentlicht: (2022)
SeqRoute: Global Budget-Aware Sequential LLM Routing via Offline Reinforcement Learning
von: Xu, Zhongling, et al.
Veröffentlicht: (2026)
von: Xu, Zhongling, et al.
Veröffentlicht: (2026)
Gradient Routing: Masking Gradients to Localize Computation in Neural Networks
von: Cloud, Alex, et al.
Veröffentlicht: (2024)
von: Cloud, Alex, et al.
Veröffentlicht: (2024)
Data Fusion-Enhanced Decision Transformer for Stable Cross-Domain Generalization
von: Wang, Guojian, et al.
Veröffentlicht: (2025)
von: Wang, Guojian, et al.
Veröffentlicht: (2025)
Input Domain Aware MoE: Decoupling Routing Decisions from Task Optimization in Mixture of Experts
von: Hua, Yongxiang, et al.
Veröffentlicht: (2025)
von: Hua, Yongxiang, et al.
Veröffentlicht: (2025)
GraphGPT: Generative Pre-trained Graph Eulerian Transformer
von: Zhao, Qifang, et al.
Veröffentlicht: (2023)
von: Zhao, Qifang, et al.
Veröffentlicht: (2023)
Validated Intent Compilation for Constrained Routing in LEO Mega-Constellations
von: Li, Yuanhang
Veröffentlicht: (2026)
von: Li, Yuanhang
Veröffentlicht: (2026)
Regret Analysis of Policy Gradient Algorithm for Infinite Horizon Average Reward Markov Decision Processes
von: Bai, Qinbo, et al.
Veröffentlicht: (2023)
von: Bai, Qinbo, et al.
Veröffentlicht: (2023)
Physics-Guided Tiny-Mamba Transformer for Reliability-Aware Early Fault Warning
von: Li, Changyu, et al.
Veröffentlicht: (2026)
von: Li, Changyu, et al.
Veröffentlicht: (2026)
Directional Routing in Transformers
von: Taylor, Kevin
Veröffentlicht: (2026)
von: Taylor, Kevin
Veröffentlicht: (2026)
Complexity-Aware Deep Symbolic Regression with Robust Risk-Seeking Policy Gradients
von: Bastiani, Zachary, et al.
Veröffentlicht: (2024)
von: Bastiani, Zachary, et al.
Veröffentlicht: (2024)
Improved Sample Complexity Analysis of Natural Policy Gradient Algorithm with General Parameterization for Infinite Horizon Discounted Reward Markov Decision Processes
von: Mondal, Washim Uddin, et al.
Veröffentlicht: (2023)
von: Mondal, Washim Uddin, et al.
Veröffentlicht: (2023)
Retrieval-Aware Distillation for Transformer-SSM Hybrids
von: Bick, Aviv, et al.
Veröffentlicht: (2026)
von: Bick, Aviv, et al.
Veröffentlicht: (2026)
RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs
von: Xu, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Xu, Zhiyuan, et al.
Veröffentlicht: (2026)
Distance-Misaligned Training in Graph Transformers and Adaptive Graph-Aware Control
von: Hou, Qinhan, et al.
Veröffentlicht: (2026)
von: Hou, Qinhan, et al.
Veröffentlicht: (2026)
PR-CapsNet: Pseudo-Riemannian Capsule Network with Adaptive Curvature Routing for Graph Learning
von: Qin, Ye, et al.
Veröffentlicht: (2025)
von: Qin, Ye, et al.
Veröffentlicht: (2025)
The Devil Is in Gradient Entanglement: Energy-Aware Gradient Coordinator for Robust Generalized Category Discovery
von: Zheng, Haiyang, et al.
Veröffentlicht: (2026)
von: Zheng, Haiyang, et al.
Veröffentlicht: (2026)
Discrepancy-Aware Graph Mask Auto-Encoder
von: Zheng, Ziyu, et al.
Veröffentlicht: (2025)
von: Zheng, Ziyu, et al.
Veröffentlicht: (2025)
Uncertainty-Aware Gradient Signal-to-Noise Data Selection for Instruction Tuning
von: Yuan, Zhihang, et al.
Veröffentlicht: (2026)
von: Yuan, Zhihang, et al.
Veröffentlicht: (2026)
ImaginationPolicy: Towards Generalizable, Precise and Reliable End-to-End Policy for Robotic Manipulation
von: Lu, Dekun, et al.
Veröffentlicht: (2025)
von: Lu, Dekun, et al.
Veröffentlicht: (2025)
Task-Aware Harmony Multi-Task Decision Transformer for Offline Reinforcement Learning
von: Fan, Ziqing, et al.
Veröffentlicht: (2024)
von: Fan, Ziqing, et al.
Veröffentlicht: (2024)
Randomized Advantage Transformation (RAT): Computing Natural Policy Gradients via Direct Backpropagation
von: Sun, Mingfei
Veröffentlicht: (2026)
von: Sun, Mingfei
Veröffentlicht: (2026)
From Imperfect Signals to Trustworthy Structure: Confidence-Aware Inference from Heterogeneous and Reliability-Varying Utility Data
von: Li, Haoran, et al.
Veröffentlicht: (2025)
von: Li, Haoran, et al.
Veröffentlicht: (2025)
One Router to Route Them All: Homogeneous Expert Routing for Heterogeneous Graph Transformers
von: Shakirov, Georgiy, et al.
Veröffentlicht: (2025)
von: Shakirov, Georgiy, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Graph-Aware Diffusion for Signal Generation
von: Rozada, Sergio, et al.
Veröffentlicht: (2025) -
Policy Gradient for Robust Markov Decision Processes
von: Wang, Qiuhao, et al.
Veröffentlicht: (2024) -
GPG: Generalized Policy Gradient Theorem for Transformer-based Policies
von: Mao, Hangyu, et al.
Veröffentlicht: (2025) -
Adjusting the Output of Decision Transformer with Action Gradient
von: Lin, Rui, et al.
Veröffentlicht: (2025) -
MASPO: Unifying Gradient Utilization, Probability Mass, and Signal Reliability for Robust and Sample-Efficient LLM Reasoning
von: Fu, Xiaoliang, et al.
Veröffentlicht: (2026)