GoldenStart: Q-Guided Priors and Entropy Control for Distilling Flow Policies
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhang, He, Sun, Ying, Xiong, Hui |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
AI Agents: Evolution, Architecture, and Real-World Applications
par: Krishnan, Naveen
Publié: (2025)
par: Krishnan, Naveen
Publié: (2025)
Adaptive Minds: Empowering Agents with LoRA-as-Tools
par: Shekar, Pavan C, et autres
Publié: (2025)
par: Shekar, Pavan C, et autres
Publié: (2025)
Evolving machine learning workflows through interactive AutoML
par: Barbudo, Rafael, et autres
Publié: (2024)
par: Barbudo, Rafael, et autres
Publié: (2024)
SafeRL-Lite: A Lightweight, Explainable, and Constrained Reinforcement Learning Library
par: Mishra, Satyam, et autres
Publié: (2025)
par: Mishra, Satyam, et autres
Publié: (2025)
Foundation Models as World Models: A Foundational Study in Text-Based GridWorlds
par: Sasso, Remo, et autres
Publié: (2025)
par: Sasso, Remo, et autres
Publié: (2025)
Exploration with Foundation Models: Capabilities, Limitations, and Hybrid Approaches
par: Sasso, Remo, et autres
Publié: (2025)
par: Sasso, Remo, et autres
Publié: (2025)
An Improved Adaptive PID Optimizer with Enhanced Convergence and Stability for Deep Learning
par: Saini, Saurabh, et autres
Publié: (2026)
par: Saini, Saurabh, et autres
Publié: (2026)
ElasticFlow: One-Step Physics-Consistent Policy with Elastic Time Horizons for Language-Guided Manipulation
par: Chen, Kewei, et autres
Publié: (2026)
par: Chen, Kewei, et autres
Publié: (2026)
Modeling and Controlling Deployment Reliability under Temporal Distribution Shift
par: Rahman, Naimur, et autres
Publié: (2026)
par: Rahman, Naimur, et autres
Publié: (2026)
Proving Olympiad Algebraic Inequalities without Human Demonstrations
par: Wei, Chenrui, et autres
Publié: (2024)
par: Wei, Chenrui, et autres
Publié: (2024)
SALE-Based Offline Reinforcement Learning with Ensemble Q-Networks
par: Chun, Zheng
Publié: (2025)
par: Chun, Zheng
Publié: (2025)
XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
par: Estevanell-Valladares, Ernesto L., et autres
Publié: (2025)
par: Estevanell-Valladares, Ernesto L., et autres
Publié: (2025)
STRIDE: A Self-Reflective Agent Framework for Reliable Automatic Equation Discovery
par: Su, Jiarui, et autres
Publié: (2026)
par: Su, Jiarui, et autres
Publié: (2026)
SMOSE: Sparse Mixture of Shallow Experts for Interpretable Reinforcement Learning in Continuous Control Tasks
par: Vincze, Mátyás, et autres
Publié: (2024)
par: Vincze, Mátyás, et autres
Publié: (2024)
AI and Machine Learning Approaches for Predicting Nanoparticles Toxicity The Critical Role of Physiochemical Properties
par: Yousaf, Iqra
Publié: (2024)
par: Yousaf, Iqra
Publié: (2024)
Connectivity-Aware Representations for Constrained Motion Planning via Multi-Scale Contrastive Learning
par: Jeon, Suhyun, et autres
Publié: (2026)
par: Jeon, Suhyun, et autres
Publié: (2026)
Active Causal Experimentalist (ACE): Learning Intervention Strategies via Direct Preference Optimization
par: Cooper, Patrick, et autres
Publié: (2026)
par: Cooper, Patrick, et autres
Publié: (2026)
The Mirror Loop: Recursive Non-Convergence in Generative Reasoning Systems
par: DeVilling, Bentley
Publié: (2025)
par: DeVilling, Bentley
Publié: (2025)
Learning to Select Goals in Automated Planning with Deep-Q Learning
par: Núñez-Molina, Carlos, et autres
Publié: (2024)
par: Núñez-Molina, Carlos, et autres
Publié: (2024)
Emotion-Inspired Learning Signals (EILS): A Homeostatic Framework for Adaptive Autonomous Agents
par: Tiwari, Dhruv
Publié: (2025)
par: Tiwari, Dhruv
Publié: (2025)
A Comparison Between Decision Transformers and Traditional Offline Reinforcement Learning Algorithms
par: Caunhye, Ali Murtaza, et autres
Publié: (2025)
par: Caunhye, Ali Murtaza, et autres
Publié: (2025)
FDQN: A Flexible Deep Q-Network Framework for Game Automation
par: Gujavarthy, Prabhath Reddy
Publié: (2024)
par: Gujavarthy, Prabhath Reddy
Publié: (2024)
Streaming Continual Learning for Unified Adaptive Intelligence in Dynamic Environments
par: Giannini, Federico, et autres
Publié: (2026)
par: Giannini, Federico, et autres
Publié: (2026)
Diffusion-MPC in Discrete Domains: Feasibility Constraints, Horizon Effects, and Critic Alignment: Case study with Tetris
par: Wang, Haochuan Kevin
Publié: (2026)
par: Wang, Haochuan Kevin
Publié: (2026)
LeanProgress: Guiding Search for Neural Theorem Proving via Proof Progress Prediction
par: George, Robert Joseph, et autres
Publié: (2025)
par: George, Robert Joseph, et autres
Publié: (2025)
Task Memory Engine (TME): Enhancing State Awareness for Multi-Step LLM Agent Tasks
par: Ye, Ye
Publié: (2025)
par: Ye, Ye
Publié: (2025)
Can a Bayesian Oracle Prevent Harm from an Agent?
par: Bengio, Yoshua, et autres
Publié: (2024)
par: Bengio, Yoshua, et autres
Publié: (2024)
Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation
par: Osman, Asim, et autres
Publié: (2026)
par: Osman, Asim, et autres
Publié: (2026)
From Imitation to Interaction: Mastering Game of Schnapsen with Shallow Reinforcement Learning
par: Klačan, Ján, et autres
Publié: (2026)
par: Klačan, Ján, et autres
Publié: (2026)
On Divergence Measures for Training GFlowNets
par: da Silva, Tiago, et autres
Publié: (2024)
par: da Silva, Tiago, et autres
Publié: (2024)
STACHE: Local Black-Box Explanations for Reinforcement Learning Policies
par: Elashkin, Andrew, et autres
Publié: (2025)
par: Elashkin, Andrew, et autres
Publié: (2025)
Efficient Contextual Preferential Bayesian Optimization with Historical Examples
par: Khan, Farha A., et autres
Publié: (2022)
par: Khan, Farha A., et autres
Publié: (2022)
A Parallel Hybrid Action Space Reinforcement Learning Model for Real-world Adaptive Traffic Signal Control
par: Wang, Yuxuan, et autres
Publié: (2025)
par: Wang, Yuxuan, et autres
Publié: (2025)
Maximum Entropy Relaxation of Multi-Way Cardinality Constraints for Synthetic Population Generation
par: Pachet, François, et autres
Publié: (2026)
par: Pachet, François, et autres
Publié: (2026)
Statistical Guarantees for Lifelong Reinforcement Learning using PAC-Bayes Theory
par: Zhang, Zhi, et autres
Publié: (2024)
par: Zhang, Zhi, et autres
Publié: (2024)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
par: Pather, Kaviraj, et autres
Publié: (2025)
par: Pather, Kaviraj, et autres
Publié: (2025)
KL-Regularised Q-Learning: A Token-level Action-Value perspective on Online RLHF
par: Brown, Jason R, et autres
Publié: (2025)
par: Brown, Jason R, et autres
Publié: (2025)
A Structural Threshold in Decision Capacity Governs Collapse in Self-Play Reinforcement Learning
par: Kujur, Arahan
Publié: (2026)
par: Kujur, Arahan
Publié: (2026)
Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models
par: Fu, Tianyu, et autres
Publié: (2025)
par: Fu, Tianyu, et autres
Publié: (2025)
LLM-Assisted Iterative Evolution with Swarm Intelligence Toward SuperBrain
par: Weigang, Li, et autres
Publié: (2025)
par: Weigang, Li, et autres
Publié: (2025)
Documents similaires
-
AI Agents: Evolution, Architecture, and Real-World Applications
par: Krishnan, Naveen
Publié: (2025) -
Adaptive Minds: Empowering Agents with LoRA-as-Tools
par: Shekar, Pavan C, et autres
Publié: (2025) -
Evolving machine learning workflows through interactive AutoML
par: Barbudo, Rafael, et autres
Publié: (2024) -
SafeRL-Lite: A Lightweight, Explainable, and Constrained Reinforcement Learning Library
par: Mishra, Satyam, et autres
Publié: (2025) -
Foundation Models as World Models: A Foundational Study in Text-Based GridWorlds
par: Sasso, Remo, et autres
Publié: (2025)