LLMs for High-Frequency Decision-Making: Normalized Action Reward-Guided Consistency Policy Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Yang, Li, Zihao, Jiang, Zhiyu, Ma, Dandan, Liu, Ganchao, Zhao, Wenzhe |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SAGE-LLM: Towards Safe and Generalizable LLM Controller with Fuzzy-CBF Verification and Graph-Structured Knowledge Retrieval for UAV Decision
by: Zhao, Wenzhe, et al.
Published: (2026)
by: Zhao, Wenzhe, et al.
Published: (2026)
Stream-level flow matching with Gaussian processes
by: Wei, Ganchao, et al.
Published: (2024)
by: Wei, Ganchao, et al.
Published: (2024)
VLASCD: A Visual Language Action Model for Simultaneous Chatting and Decision Making
by: Tang, Zuojin, et al.
Published: (2024)
by: Tang, Zuojin, et al.
Published: (2024)
HALO: Hallucination Analysis and Learning Optimization to Empower LLMs with Retrieval-Augmented Context for Guided Clinical Decision Making
by: Anjum, Sumera, et al.
Published: (2024)
by: Anjum, Sumera, et al.
Published: (2024)
FreqPolicy: Efficient Flow-based Visuomotor Policy via Frequency Consistency
by: Su, Yifei, et al.
Published: (2025)
by: Su, Yifei, et al.
Published: (2025)
Transformer-Enhanced Motion Planner: Attention-Guided Sampling for State-Specific Decision Making
by: Zhuang, Lei, et al.
Published: (2024)
by: Zhuang, Lei, et al.
Published: (2024)
Driving with Regulation: Trustworthy and Interpretable Decision-Making for Autonomous Driving with Retrieval-Augmented Reasoning
by: Cai, Tianhui, et al.
Published: (2024)
by: Cai, Tianhui, et al.
Published: (2024)
Meta-Aligner: Bidirectional Preference-Policy Optimization for Multi-Objective LLMs Alignment
by: Xu, Wenzhe, et al.
Published: (2026)
by: Xu, Wenzhe, et al.
Published: (2026)
Reward Guided Latent Consistency Distillation
by: Li, Jiachen, et al.
Published: (2024)
by: Li, Jiachen, et al.
Published: (2024)
SOLID: a Framework of Synergizing Optimization and LLMs for Intelligent Decision-Making
by: Wang, Yinsheng, et al.
Published: (2025)
by: Wang, Yinsheng, et al.
Published: (2025)
Pretrain Value, Not Reward: Decoupled Value Policy Optimization
by: Huang, Chenghua, et al.
Published: (2025)
by: Huang, Chenghua, et al.
Published: (2025)
Reward Design for Justifiable Sequential Decision-Making
by: Sukovic, Aleksa, et al.
Published: (2024)
by: Sukovic, Aleksa, et al.
Published: (2024)
Cognitive Bias in Decision-Making with LLMs
by: Echterhoff, Jessica, et al.
Published: (2024)
by: Echterhoff, Jessica, et al.
Published: (2024)
Evaluating LLMs for Police Decision-Making: A Framework Based on Police Action Scenarios
by: Lee, Sangyub, et al.
Published: (2026)
by: Lee, Sangyub, et al.
Published: (2026)
The State-Action-Reward-State-Action Algorithm in Spatial Prisoner's Dilemma Game
by: Yang, Lanyu, et al.
Published: (2024)
by: Yang, Lanyu, et al.
Published: (2024)
WARBENCH: A Comprehensive Benchmark for Evaluating LLMs in Military Decision-Making
by: Li, Zongjie, et al.
Published: (2026)
by: Li, Zongjie, et al.
Published: (2026)
Reflective Preference Optimization (RPO): Enhancing On-Policy Alignment via Hint-Guided Reflection
by: Zhao, Zihui, et al.
Published: (2025)
by: Zhao, Zihui, et al.
Published: (2025)
The Consistency-Acceptability Divergence of LLMs in Judicial Decision-Making: Task and Stakeholder Dimensions
by: MingDa, Zhang, et al.
Published: (2025)
by: MingDa, Zhang, et al.
Published: (2025)
Frequency-Enhanced Diffusion Models: Curriculum-Guided Semantic Alignment for Zero-Shot Skeleton Action Recognition
by: Zhou, Yuxi, et al.
Published: (2026)
by: Zhou, Yuxi, et al.
Published: (2026)
RRPO: Robust Reward Policy Optimization for LLM-based Emotional TTS
by: Wang, Cong, et al.
Published: (2025)
by: Wang, Cong, et al.
Published: (2025)
Policy Filtration for RLHF to Mitigate Noise in Reward Models
by: Zhang, Chuheng, et al.
Published: (2024)
by: Zhang, Chuheng, et al.
Published: (2024)
Decision Flow Policy Optimization
by: Hu, Jifeng, et al.
Published: (2025)
by: Hu, Jifeng, et al.
Published: (2025)
Difficulty-Estimated Policy Optimization
by: Zhao, Yu, et al.
Published: (2026)
by: Zhao, Yu, et al.
Published: (2026)
CROP: Conservative Reward for Model-based Offline Policy Optimization
by: Li, Hao, et al.
Published: (2023)
by: Li, Hao, et al.
Published: (2023)
ActionCodec: What Makes for Good Action Tokenizers
by: Dong, Zibin, et al.
Published: (2026)
by: Dong, Zibin, et al.
Published: (2026)
Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
by: Li, Manling, et al.
Published: (2024)
by: Li, Manling, et al.
Published: (2024)
One-Way Policy Optimization for Self-Evolving LLMs
by: Yang, Shuo, et al.
Published: (2026)
by: Yang, Shuo, et al.
Published: (2026)
DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization
by: Li, Gang, et al.
Published: (2025)
by: Li, Gang, et al.
Published: (2025)
RiskPO: Risk-based Policy Optimization via Verifiable Reward for LLM Post-Training
by: Ren, Tao, et al.
Published: (2025)
by: Ren, Tao, et al.
Published: (2025)
Adversarial Testing in LLMs: Insights into Decision-Making Vulnerabilities
by: Zhang, Lili, et al.
Published: (2025)
by: Zhang, Lili, et al.
Published: (2025)
NOTAM-Evolve: A Knowledge-Guided Self-Evolving Optimization Framework with LLMs for NOTAM Interpretation
by: Liu, Maoqi, et al.
Published: (2025)
by: Liu, Maoqi, et al.
Published: (2025)
SMAC-R1: The Emergence of Intelligence in Decision-Making Tasks
by: Deng, Yue, et al.
Published: (2024)
by: Deng, Yue, et al.
Published: (2024)
Intrinsic Reward Policy Optimization for Sparse-Reward Environments
by: Cho, Minjae, et al.
Published: (2026)
by: Cho, Minjae, et al.
Published: (2026)
Leveraging Information Consistency in Frequency and Spatial Domain for Adversarial Attacks
by: Jin, Zhibo, et al.
Published: (2024)
by: Jin, Zhibo, et al.
Published: (2024)
Preferred-Action-Optimized Diffusion Policies for Offline Reinforcement Learning
by: Zhang, Tianle, et al.
Published: (2024)
by: Zhang, Tianle, et al.
Published: (2024)
Mediator-Guided Multi-Agent Collaboration among Open-Source Models for Medical Decision-Making
by: Chen, Kaitao, et al.
Published: (2025)
by: Chen, Kaitao, et al.
Published: (2025)
Strategy-Aware Optimization Modeling with Reasoning LLMs
by: Zhao, Ruiqing, et al.
Published: (2026)
by: Zhao, Ruiqing, et al.
Published: (2026)
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models
by: Zhou, Xin, et al.
Published: (2025)
by: Zhou, Xin, et al.
Published: (2025)
IDEA: An Interpretable and Editable Decision-Making Framework for LLMs via Verbal-to-Numeric Calibration
by: He, Yanji, et al.
Published: (2026)
by: He, Yanji, et al.
Published: (2026)
Unifying and Optimizing Data Values for Selection via Sequential Decision-Making
by: Chi, Hongliang, et al.
Published: (2025)
by: Chi, Hongliang, et al.
Published: (2025)
Similar Items
-
SAGE-LLM: Towards Safe and Generalizable LLM Controller with Fuzzy-CBF Verification and Graph-Structured Knowledge Retrieval for UAV Decision
by: Zhao, Wenzhe, et al.
Published: (2026) -
Stream-level flow matching with Gaussian processes
by: Wei, Ganchao, et al.
Published: (2024) -
VLASCD: A Visual Language Action Model for Simultaneous Chatting and Decision Making
by: Tang, Zuojin, et al.
Published: (2024) -
HALO: Hallucination Analysis and Learning Optimization to Empower LLMs with Retrieval-Augmented Context for Guided Clinical Decision Making
by: Anjum, Sumera, et al.
Published: (2024) -
FreqPolicy: Efficient Flow-based Visuomotor Policy via Frequency Consistency
by: Su, Yifei, et al.
Published: (2025)