A Comparative Theoretical Analysis of Entropy Control Methods in Reinforcement Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Lei, Ming, Baehr, Christophe |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FedPF: Accurate Target Privacy Preserving Federated Learning Balancing Fairness and Utility
por: Sun, Kangkang, et al.
Publicado: (2025)
por: Sun, Kangkang, et al.
Publicado: (2025)
Prediction and Forecast of Short-Term Drought Impacts Using Machine Learning to Support Mitigation and Adaptation Efforts
por: Geli, Hatim M. E., et al.
Publicado: (2025)
por: Geli, Hatim M. E., et al.
Publicado: (2025)
Non-linear Phillips Curve for India: Evidence from Explainable Machine Learning
por: Sengupta, Shovon, et al.
Publicado: (2025)
por: Sengupta, Shovon, et al.
Publicado: (2025)
Intervention-Assisted Policy Gradient Methods for Online Stochastic Queuing Network Optimization: Technical Report
por: Wigmore, Jerrod, et al.
Publicado: (2024)
por: Wigmore, Jerrod, et al.
Publicado: (2024)
From Static to Adaptive Defense: Federated Multi-Agent Deep Reinforcement Learning-Driven Moving Target Defense Against DoS Attacks in UAV Swarm Networks
por: Zhou, Yuyang, et al.
Publicado: (2025)
por: Zhou, Yuyang, et al.
Publicado: (2025)
Learning Symbolic Task Decompositions for Multi-Agent Teams
por: Shah, Ameesh, et al.
Publicado: (2025)
por: Shah, Ameesh, et al.
Publicado: (2025)
Continual Learning, Not Training: Online Adaptation For Agents
por: Jaglan, Aman, et al.
Publicado: (2025)
por: Jaglan, Aman, et al.
Publicado: (2025)
Geometric Meta-Learning via Coupled Ricci Flow: Unifying Knowledge Representation and Quantum Entanglement
por: Lei, Ming, et al.
Publicado: (2025)
por: Lei, Ming, et al.
Publicado: (2025)
Uncovering Bias Paths with LLM-guided Causal Discovery: An Active Learning and Dynamic Scoring Approach
por: Zanna, Khadija, et al.
Publicado: (2025)
por: Zanna, Khadija, et al.
Publicado: (2025)
Hidden Failure Modes of Gradient Modification under Adam in Continual Learning, and Adaptive Decoupled Moment Routing as a Repair
por: Hu, Yuelin, et al.
Publicado: (2026)
por: Hu, Yuelin, et al.
Publicado: (2026)
Defense-as-a-Service: Black-box Shielding against Backdoored Graph Models
por: Yang, Xiao, et al.
Publicado: (2024)
por: Yang, Xiao, et al.
Publicado: (2024)
EnergyMamba: An Uncertainty-Aware Graph-Enhanced Selective State Space Model for Energy Consumption Prediction
por: Yu, Dahai, et al.
Publicado: (2026)
por: Yu, Dahai, et al.
Publicado: (2026)
Evaluating Model Robustness Using Adaptive Sparse L0 Regularization
por: Liu, Weiyou, et al.
Publicado: (2024)
por: Liu, Weiyou, et al.
Publicado: (2024)
Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining
por: Li, Houyi, et al.
Publicado: (2025)
por: Li, Houyi, et al.
Publicado: (2025)
Distribution Consistency based Self-Training for Graph Neural Networks with Sparse Labels
por: Wang, Fali, et al.
Publicado: (2024)
por: Wang, Fali, et al.
Publicado: (2024)
Towards Verifiable AI with Lightweight Cryptographic Proofs of Inference
por: Anchuri, Pranay, et al.
Publicado: (2026)
por: Anchuri, Pranay, et al.
Publicado: (2026)
Deep Learning for Human Locomotion Analysis in Lower-Limb Exoskeletons: A Comparative Study
por: Coser, Omar, et al.
Publicado: (2025)
por: Coser, Omar, et al.
Publicado: (2025)
LAWS: Learning from Actual Workloads Symbolically -- A Self-Certifying Parametrized Cache Architecture for Neural Inference, Robotics, and Edge Deployment
por: Magarshak, Gregory
Publicado: (2026)
por: Magarshak, Gregory
Publicado: (2026)
Lightweight Quantum-Enhanced ResNet for Coronary Angiography Classification: A Hybrid Quantum-Classical Feature Enhancement Framework
por: Xia, Jingsong
Publicado: (2026)
por: Xia, Jingsong
Publicado: (2026)
Time-to-Injury Forecasting in Elite Female Football: A DeepHit Survival Approach
por: Catterall, Victoria, et al.
Publicado: (2026)
por: Catterall, Victoria, et al.
Publicado: (2026)
Software Model Evolution with Large Language Models: Experiments on Simulated, Public, and Industrial Datasets
por: Tinnes, Christof, et al.
Publicado: (2024)
por: Tinnes, Christof, et al.
Publicado: (2024)
LayerRoute: Input-Conditioned Adaptive Layer Skipping via LoRA Fine-Tuning for Agentic Language Models
por: Sikdar, Prateek Kumar
Publicado: (2026)
por: Sikdar, Prateek Kumar
Publicado: (2026)
DeepContext: Stateful Real-Time Detection of Multi-Turn Adversarial Intent Drift in LLMs
por: Albrethsen, Justin, et al.
Publicado: (2026)
por: Albrethsen, Justin, et al.
Publicado: (2026)
Transcribing Bengali Text with Regional Dialects to IPA using District Guided Tokens
por: Islam, S M Jishanul, et al.
Publicado: (2024)
por: Islam, S M Jishanul, et al.
Publicado: (2024)
When Do Early-Exit Networks Generalize? A PAC-Bayesian Theory of Adaptive Depth
por: Guo, Dongxin, et al.
Publicado: (2026)
por: Guo, Dongxin, et al.
Publicado: (2026)
Evaluating the Limitations of Local LLMs in Solving Complex Programming Challenges
por: Matotek, Kadin, et al.
Publicado: (2025)
por: Matotek, Kadin, et al.
Publicado: (2025)
Intelligence Inertia: Physical Isomorphism and Applications
por: Han, Jipeng
Publicado: (2026)
por: Han, Jipeng
Publicado: (2026)
Submodular Benchmark Selection
por: Smola, Alexander
Publicado: (2026)
por: Smola, Alexander
Publicado: (2026)
A Method for Evaluating the Interpretability of Machine Learning Models in Predicting Bond Default Risk Based on LIME and SHAP
por: Zhang, Yan, et al.
Publicado: (2025)
por: Zhang, Yan, et al.
Publicado: (2025)
JoFormer (Journey-based Transformer): Theory and Empirical Analysis on the Tiny Shakespeare Dataset
por: Godavarti, Mahesh
Publicado: (2025)
por: Godavarti, Mahesh
Publicado: (2025)
A Channel Attention-Driven Hybrid CNN Framework for Paddy Leaf Disease Detection
por: V, Pandiyaraju, et al.
Publicado: (2024)
por: V, Pandiyaraju, et al.
Publicado: (2024)
A Theoretical Computer Science Perspective on Free Will
por: Blum, Manuel, et al.
Publicado: (2022)
por: Blum, Manuel, et al.
Publicado: (2022)
MACI: Multi-Agent Collaborative Intelligence for Adaptive Reasoning and Temporal Planning
por: Chang, Edward Y.
Publicado: (2025)
por: Chang, Edward Y.
Publicado: (2025)
StarCraft+: Benchmarking Multi-agent Algorithms in Adversary Paradigm
por: Li, Yadong, et al.
Publicado: (2025)
por: Li, Yadong, et al.
Publicado: (2025)
Random Heterogeneous Neurochaos Learning Architecture for Data Classification
por: S, Remya Ajai A, et al.
Publicado: (2024)
por: S, Remya Ajai A, et al.
Publicado: (2024)
The Empowerment of Science of Science by Large Language Models: New Tools and Methods
por: Liang, Guoqiang, et al.
Publicado: (2025)
por: Liang, Guoqiang, et al.
Publicado: (2025)
Graph Coloring for Multi-Task Learning
por: Patapati, Santosh
Publicado: (2025)
por: Patapati, Santosh
Publicado: (2025)
Blockchain As a Platform For Artificial Intelligence (AI) Transparency
por: Akther, Afroja, et al.
Publicado: (2025)
por: Akther, Afroja, et al.
Publicado: (2025)
Counter-Inferential Behavior in Natural and Artificial Cognitive Systems
por: Dolgikh, Serge
Publicado: (2025)
por: Dolgikh, Serge
Publicado: (2025)
XAI and Few-shot-based Hybrid Classification Model for Plant Leaf Disease Prognosis
por: Joseph, Diana Susan, et al.
Publicado: (2026)
por: Joseph, Diana Susan, et al.
Publicado: (2026)
Ejemplares similares
-
FedPF: Accurate Target Privacy Preserving Federated Learning Balancing Fairness and Utility
por: Sun, Kangkang, et al.
Publicado: (2025) -
Prediction and Forecast of Short-Term Drought Impacts Using Machine Learning to Support Mitigation and Adaptation Efforts
por: Geli, Hatim M. E., et al.
Publicado: (2025) -
Non-linear Phillips Curve for India: Evidence from Explainable Machine Learning
por: Sengupta, Shovon, et al.
Publicado: (2025) -
Intervention-Assisted Policy Gradient Methods for Online Stochastic Queuing Network Optimization: Technical Report
por: Wigmore, Jerrod, et al.
Publicado: (2024) -
From Static to Adaptive Defense: Federated Multi-Agent Deep Reinforcement Learning-Driven Moving Target Defense Against DoS Attacks in UAV Swarm Networks
por: Zhou, Yuyang, et al.
Publicado: (2025)