Online Causal Kalman Filtering for Stable and Effective Policy Optimization
Fuente:
arXiv
Guardado en:
| Autores principales: | He, Shuo, Feng, Lang, Cheng, Xin, Feng, Lei, An, Bo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Hierarchy-of-Groups Policy Optimization for Long-Horizon Agentic Tasks
por: He, Shuo, et al.
Publicado: (2026)
por: He, Shuo, et al.
Publicado: (2026)
Is Crowdsourcing Breaking Your Bank? Cost-Effective Fine-Tuning of Pre-trained Language Models with Proximal Policy Optimization
por: Yang, Shuo, et al.
Publicado: (2024)
por: Yang, Shuo, et al.
Publicado: (2024)
COPO: Causal-Oriented Policy Optimization for Hallucinations of MLLMs
por: Guo, Peizheng, et al.
Publicado: (2025)
por: Guo, Peizheng, et al.
Publicado: (2025)
An Information Bottleneck Perspective for Effective Noise Filtering on Retrieval-Augmented Generation
por: Zhu, Kun, et al.
Publicado: (2024)
por: Zhu, Kun, et al.
Publicado: (2024)
Causally-Enhanced Reinforcement Policy Optimization
por: Wang, Xiangqi, et al.
Publicado: (2025)
por: Wang, Xiangqi, et al.
Publicado: (2025)
StableMask: Refining Causal Masking in Decoder-only Transformer
por: Yin, Qingyu, et al.
Publicado: (2024)
por: Yin, Qingyu, et al.
Publicado: (2024)
Online Difficulty Filtering for Reasoning Oriented Reinforcement Learning
por: Bae, Sanghwan, et al.
Publicado: (2025)
por: Bae, Sanghwan, et al.
Publicado: (2025)
Optimizing Instruction Synthesis: Effective Exploration of Evolutionary Space with Tree Search
por: Li, Chenglin, et al.
Publicado: (2024)
por: Li, Chenglin, et al.
Publicado: (2024)
Causal Discovery and Counterfactual Reasoning to Optimize Persuasive Dialogue Policies
por: Zeng, Donghuo, et al.
Publicado: (2025)
por: Zeng, Donghuo, et al.
Publicado: (2025)
TrustDataFilter:Leveraging Trusted Knowledge Base Data for More Effective Filtering of Unknown Information
por: Zhang, Jinghong, et al.
Publicado: (2025)
por: Zhang, Jinghong, et al.
Publicado: (2025)
Hallucinate Less by Thinking More: Aspect-Based Causal Abstention for Large Language Models
por: Nguyen, Vy, et al.
Publicado: (2025)
por: Nguyen, Vy, et al.
Publicado: (2025)
Beyond Uniform Credit: Causal Credit Assignment for Policy Optimization
por: Khandoga, Mykola, et al.
Publicado: (2026)
por: Khandoga, Mykola, et al.
Publicado: (2026)
Unveiling Causal Reasoning in Large Language Models: Reality or Mirage?
por: Chi, Haoang, et al.
Publicado: (2025)
por: Chi, Haoang, et al.
Publicado: (2025)
Aligning Large Language Models to Follow Instructions and Hallucinate Less via Effective Data Filtering
por: Si, Shuzheng, et al.
Publicado: (2025)
por: Si, Shuzheng, et al.
Publicado: (2025)
AT$^2$PO: Agentic Turn-based Policy Optimization via Tree Search
por: Zong, Zefang, et al.
Publicado: (2026)
por: Zong, Zefang, et al.
Publicado: (2026)
Nuance Matters: Probing Epistemic Consistency in Causal Reasoning
por: Cui, Shaobo, et al.
Publicado: (2024)
por: Cui, Shaobo, et al.
Publicado: (2024)
VEPO: Variable Entropy Policy Optimization for Low-Resource Language Foundation Models
por: Liu, Chonghan, et al.
Publicado: (2026)
por: Liu, Chonghan, et al.
Publicado: (2026)
Filter-then-Weight: Online Data Selection and Reweighting for LLM Fine-Tuning
por: Wang, Fangxin, et al.
Publicado: (2026)
por: Wang, Fangxin, et al.
Publicado: (2026)
Retroformer: Retrospective Large Language Agents with Policy Gradient Optimization
por: Yao, Weiran, et al.
Publicado: (2023)
por: Yao, Weiran, et al.
Publicado: (2023)
Fibration Policy Optimization
por: Li, Chang, et al.
Publicado: (2026)
por: Li, Chang, et al.
Publicado: (2026)
Optima: Optimizing Effectiveness and Efficiency for LLM-Based Multi-Agent System
por: Chen, Weize, et al.
Publicado: (2024)
por: Chen, Weize, et al.
Publicado: (2024)
LLMs Are Prone to Fallacies in Causal Inference
por: Joshi, Nitish, et al.
Publicado: (2024)
por: Joshi, Nitish, et al.
Publicado: (2024)
ICDPO: Effectively Borrowing Alignment Capability of Others via In-context Direct Preference Optimization
por: Song, Feifan, et al.
Publicado: (2024)
por: Song, Feifan, et al.
Publicado: (2024)
MHPO: Modulated Hazard-aware Policy Optimization for Stable Reinforcement Learning
por: Wang, Hongjun, et al.
Publicado: (2026)
por: Wang, Hongjun, et al.
Publicado: (2026)
keqing: knowledge-based question answering is a nature chain-of-thought mentor of LLM
por: Wang, Chaojie, et al.
Publicado: (2023)
por: Wang, Chaojie, et al.
Publicado: (2023)
Evidence-Augmented Policy Optimization with Reward Co-Evolution for Long-Context Reasoning
por: Guan, Xin, et al.
Publicado: (2026)
por: Guan, Xin, et al.
Publicado: (2026)
BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping
por: Xi, Zhiheng, et al.
Publicado: (2025)
por: Xi, Zhiheng, et al.
Publicado: (2025)
FinS-Pilot: A Benchmark for Online Financial RAG System
por: Wang, Feng, et al.
Publicado: (2025)
por: Wang, Feng, et al.
Publicado: (2025)
CauScientist: Teaching LLMs to Respect Data for Causal Discovery
por: Peng, Bo, et al.
Publicado: (2026)
por: Peng, Bo, et al.
Publicado: (2026)
IGOT: Information Gain Optimized Tokenizer on Domain Adaptive Pretraining
por: Feng, Dawei, et al.
Publicado: (2024)
por: Feng, Dawei, et al.
Publicado: (2024)
SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment
por: Li, Hao, et al.
Publicado: (2026)
por: Li, Hao, et al.
Publicado: (2026)
Reasoning While Asking: Transforming Reasoning Large Language Models from Passive Solvers to Proactive Inquirers
por: Chen, Xin, et al.
Publicado: (2026)
por: Chen, Xin, et al.
Publicado: (2026)
Beyond Linear LLM Invocation: An Efficient and Effective Semantic Filter Paradigm
por: Hou, Nan, et al.
Publicado: (2026)
por: Hou, Nan, et al.
Publicado: (2026)
Filtered Direct Preference Optimization
por: Morimura, Tetsuro, et al.
Publicado: (2024)
por: Morimura, Tetsuro, et al.
Publicado: (2024)
RC-GRPO: Reward-Conditioned Group Relative Policy Optimization for Multi-Turn Tool Calling Agents
por: Zhong, Haitian, et al.
Publicado: (2026)
por: Zhong, Haitian, et al.
Publicado: (2026)
From Deferral to Learning: Online In-Context Knowledge Distillation for LLM Cascades
por: Wu, Yu, et al.
Publicado: (2025)
por: Wu, Yu, et al.
Publicado: (2025)
DCPO: Dynamic Clipping Policy Optimization
por: Yang, Shihui, et al.
Publicado: (2025)
por: Yang, Shihui, et al.
Publicado: (2025)
Effective Length Extrapolation via Dimension-Wise Positional Embeddings Manipulation
por: Lu, Yi, et al.
Publicado: (2025)
por: Lu, Yi, et al.
Publicado: (2025)
Are LLMs Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data
por: Liu, Xiao, et al.
Publicado: (2024)
por: Liu, Xiao, et al.
Publicado: (2024)
Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems
por: Feng, Lang, et al.
Publicado: (2026)
por: Feng, Lang, et al.
Publicado: (2026)
Ejemplares similares
-
Hierarchy-of-Groups Policy Optimization for Long-Horizon Agentic Tasks
por: He, Shuo, et al.
Publicado: (2026) -
Is Crowdsourcing Breaking Your Bank? Cost-Effective Fine-Tuning of Pre-trained Language Models with Proximal Policy Optimization
por: Yang, Shuo, et al.
Publicado: (2024) -
COPO: Causal-Oriented Policy Optimization for Hallucinations of MLLMs
por: Guo, Peizheng, et al.
Publicado: (2025) -
An Information Bottleneck Perspective for Effective Noise Filtering on Retrieval-Augmented Generation
por: Zhu, Kun, et al.
Publicado: (2024) -
Causally-Enhanced Reinforcement Policy Optimization
por: Wang, Xiangqi, et al.
Publicado: (2025)