Primal-Dual Sample Complexity Bounds for Constrained Markov Decision Processes with Multiple Constraints
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Buckley, Max, Papathanasiou, Konstantinos, Spanopoulos, Andreas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Sparse Knowledge Distillation: A Mathematical Framework for Probability-Domain Temperature Scaling and Multi-Stage Compression
von: Flouro, Aaron R., et al.
Veröffentlicht: (2026)
von: Flouro, Aaron R., et al.
Veröffentlicht: (2026)
Emotion-Inspired Learning Signals (EILS): A Homeostatic Framework for Adaptive Autonomous Agents
von: Tiwari, Dhruv
Veröffentlicht: (2025)
von: Tiwari, Dhruv
Veröffentlicht: (2025)
Reinforcement Learning Controlled Adaptive PSO for Task Offloading in IIoT Edge Computing
von: Perera, Minod, et al.
Veröffentlicht: (2025)
von: Perera, Minod, et al.
Veröffentlicht: (2025)
Temporal Taskification in Streaming Continual Learning: A Source of Evaluation Instability
von: Filat, Nicolae, et al.
Veröffentlicht: (2026)
von: Filat, Nicolae, et al.
Veröffentlicht: (2026)
A Structural Threshold in Decision Capacity Governs Collapse in Self-Play Reinforcement Learning
von: Kujur, Arahan
Veröffentlicht: (2026)
von: Kujur, Arahan
Veröffentlicht: (2026)
Data-Driven Preference Sampling for Pareto Front Learning
von: Ye, Rongguang, et al.
Veröffentlicht: (2024)
von: Ye, Rongguang, et al.
Veröffentlicht: (2024)
Diffusion-MPC in Discrete Domains: Feasibility Constraints, Horizon Effects, and Critic Alignment: Case study with Tetris
von: Wang, Haochuan Kevin
Veröffentlicht: (2026)
von: Wang, Haochuan Kevin
Veröffentlicht: (2026)
Benchmarking the Limits of In-Context Reinforcement Learning for Ad-Hoc Teamwork
von: Jing, Yuheng, et al.
Veröffentlicht: (2026)
von: Jing, Yuheng, et al.
Veröffentlicht: (2026)
Human-Aligned Skill Discovery: Balancing Behaviour Exploration and Alignment
von: Hussonnois, Maxence, et al.
Veröffentlicht: (2025)
von: Hussonnois, Maxence, et al.
Veröffentlicht: (2025)
Hard Samples, Bad Labels: Robust Loss Functions That Know When to Back Off
von: Pellegrino, Nicholas, et al.
Veröffentlicht: (2025)
von: Pellegrino, Nicholas, et al.
Veröffentlicht: (2025)
Connectivity-Aware Representations for Constrained Motion Planning via Multi-Scale Contrastive Learning
von: Jeon, Suhyun, et al.
Veröffentlicht: (2026)
von: Jeon, Suhyun, et al.
Veröffentlicht: (2026)
Interior-Point Vanishing Problem in Semidefinite Relaxations for Neural Network Verification
von: Ueda, Ryota, et al.
Veröffentlicht: (2025)
von: Ueda, Ryota, et al.
Veröffentlicht: (2025)
Evaluating Model-Agnostic Meta-Learning on MetaWorld ML10 Benchmark: Fast Adaptation in Robotic Manipulation Tasks
von: Atamuradov, Sanjar
Veröffentlicht: (2025)
von: Atamuradov, Sanjar
Veröffentlicht: (2025)
Introducing COGENT3: An AI Architecture for Emergent Cognition
von: Salazar, Eduardo
Veröffentlicht: (2025)
von: Salazar, Eduardo
Veröffentlicht: (2025)
Dynamic Dual-Granularity Skill Bank for Agentic RL
von: Tu, Songjun, et al.
Veröffentlicht: (2026)
von: Tu, Songjun, et al.
Veröffentlicht: (2026)
Understanding and Tackling Over-Dilution in Graph Neural Networks
von: Lee, Junhyun, et al.
Veröffentlicht: (2025)
von: Lee, Junhyun, et al.
Veröffentlicht: (2025)
Correction and Corruption: A Two-Rate View of Error Flow in LLM Protocols
von: Reitich, Fernando
Veröffentlicht: (2026)
von: Reitich, Fernando
Veröffentlicht: (2026)
Distinguished In Uniform: Self Attention Vs. Virtual Nodes
von: Rosenbluth, Eran, et al.
Veröffentlicht: (2024)
von: Rosenbluth, Eran, et al.
Veröffentlicht: (2024)
How much do LLMs learn from negative examples?
von: Hamdan, Shadi, et al.
Veröffentlicht: (2025)
von: Hamdan, Shadi, et al.
Veröffentlicht: (2025)
An Improved Adaptive PID Optimizer with Enhanced Convergence and Stability for Deep Learning
von: Saini, Saurabh, et al.
Veröffentlicht: (2026)
von: Saini, Saurabh, et al.
Veröffentlicht: (2026)
From Noise to Diversity: Random Embedding Injection in LLM Reasoning
von: Kim, Heejun, et al.
Veröffentlicht: (2026)
von: Kim, Heejun, et al.
Veröffentlicht: (2026)
MAcPNN: Mutual Assisted Learning on Data Streams with Temporal Dependence
von: Giannini, Federico, et al.
Veröffentlicht: (2026)
von: Giannini, Federico, et al.
Veröffentlicht: (2026)
Territory Paint Wars: Diagnosing and Mitigating Failure Modes in Competitive Multi-Agent PPO
von: Singh, Diyansha
Veröffentlicht: (2026)
von: Singh, Diyansha
Veröffentlicht: (2026)
Adaptive Weighting in Knowledge Distillation: An Axiomatic Framework for Multi-Scale Teacher Ensemble Optimization
von: Flouro, Aaron R., et al.
Veröffentlicht: (2026)
von: Flouro, Aaron R., et al.
Veröffentlicht: (2026)
Effects of Initialization Biases on Deep Neural Network Training Dynamics
von: Pellegrino, Nicholas, et al.
Veröffentlicht: (2025)
von: Pellegrino, Nicholas, et al.
Veröffentlicht: (2025)
Contrastive Mutual Information Learning: Toward Robust Representations without Positive-Pair Augmentations
von: Livne, Micha
Veröffentlicht: (2025)
von: Livne, Micha
Veröffentlicht: (2025)
When Redundancy Matters: Machine Teaching of Representations
von: Ferri, Cèsar, et al.
Veröffentlicht: (2024)
von: Ferri, Cèsar, et al.
Veröffentlicht: (2024)
Data structure > labels? Unsupervised heuristics for SVM hyperparameter estimation
von: Cholewa, Michał, et al.
Veröffentlicht: (2021)
von: Cholewa, Michał, et al.
Veröffentlicht: (2021)
Generative Pretrained Embedding and Hierarchical Irregular Time Series Representation for Daily Living Activity Recognition
von: Bouchabou, Damien, et al.
Veröffentlicht: (2024)
von: Bouchabou, Damien, et al.
Veröffentlicht: (2024)
Contrastive MIM: A Contrastive Mutual Information Framework for Unified Generative and Discriminative Representation Learning
von: Livne, Micha
Veröffentlicht: (2025)
von: Livne, Micha
Veröffentlicht: (2025)
Reinforcement Learning for Stock Transactions
von: Zhou, Ziyi, et al.
Veröffentlicht: (2025)
von: Zhou, Ziyi, et al.
Veröffentlicht: (2025)
Understanding the Limits of Deep Tabular Methods with Temporal Shift
von: Cai, Hao-Run, et al.
Veröffentlicht: (2025)
von: Cai, Hao-Run, et al.
Veröffentlicht: (2025)
Multi-Teacher Ensemble Distillation: A Mathematical Framework for Probability-Domain Knowledge Aggregation
von: Flouro, Aaron R., et al.
Veröffentlicht: (2026)
von: Flouro, Aaron R., et al.
Veröffentlicht: (2026)
RVI-SAC: Average Reward Off-Policy Deep Reinforcement Learning
von: Hisaki, Yukinari, et al.
Veröffentlicht: (2024)
von: Hisaki, Yukinari, et al.
Veröffentlicht: (2024)
Grammar-based evolutionary approach for automated workflow composition with domain-specific operators and ensemble diversity
von: Barbudo, Rafael, et al.
Veröffentlicht: (2024)
von: Barbudo, Rafael, et al.
Veröffentlicht: (2024)
Feature-aware Modulation for Learning from Temporal Tabular Data
von: Cai, Hao-Run, et al.
Veröffentlicht: (2025)
von: Cai, Hao-Run, et al.
Veröffentlicht: (2025)
Fine-Tuning Regimes Define Distinct Continual Learning Problems
von: Iordache, Paul-Tiberiu, et al.
Veröffentlicht: (2026)
von: Iordache, Paul-Tiberiu, et al.
Veröffentlicht: (2026)
Adaptive Exploration for Latent-State Bandits
von: Jin, Jikai, et al.
Veröffentlicht: (2026)
von: Jin, Jikai, et al.
Veröffentlicht: (2026)
Quantum Machine Learning for Predicting Anastomotic Leak: A Clinical Study
von: Novák, Vojtěch, et al.
Veröffentlicht: (2025)
von: Novák, Vojtěch, et al.
Veröffentlicht: (2025)
Nonlinear Data Integration via Kernel Methods for Data Collaboration Analysis
von: Suetake, Yamato, et al.
Veröffentlicht: (2026)
von: Suetake, Yamato, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Sparse Knowledge Distillation: A Mathematical Framework for Probability-Domain Temperature Scaling and Multi-Stage Compression
von: Flouro, Aaron R., et al.
Veröffentlicht: (2026) -
Emotion-Inspired Learning Signals (EILS): A Homeostatic Framework for Adaptive Autonomous Agents
von: Tiwari, Dhruv
Veröffentlicht: (2025) -
Reinforcement Learning Controlled Adaptive PSO for Task Offloading in IIoT Edge Computing
von: Perera, Minod, et al.
Veröffentlicht: (2025) -
Temporal Taskification in Streaming Continual Learning: A Source of Evaluation Instability
von: Filat, Nicolae, et al.
Veröffentlicht: (2026) -
A Structural Threshold in Decision Capacity Governs Collapse in Self-Play Reinforcement Learning
von: Kujur, Arahan
Veröffentlicht: (2026)