Per-Domain Generalizing Policies: On Validation Instances and Scaling Behavior
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gros, Timo P., Müller, Nicola J., Fiser, Daniel, Valera, Isabel, Wolf, Verena, Hoffmann, Jörg |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Per-Domain Generalizing Policies: On Learning Efficient and Robust Q-Value Functions (Extended Version with Technical Appendix)
von: Müller, Nicola J., et al.
Veröffentlicht: (2026)
von: Müller, Nicola J., et al.
Veröffentlicht: (2026)
Acting for the Right Reasons: Creating Reason-Sensitive Artificial Moral Agents
von: Baum, Kevin, et al.
Veröffentlicht: (2024)
von: Baum, Kevin, et al.
Veröffentlicht: (2024)
Exploring Molecule Generation Using Latent Space Graph Diffusion
von: Pombala, Prashanth, et al.
Veröffentlicht: (2025)
von: Pombala, Prashanth, et al.
Veröffentlicht: (2025)
Hellinger Multimodal Variational Autoencoders
von: Vo, Huyen, et al.
Veröffentlicht: (2026)
von: Vo, Huyen, et al.
Veröffentlicht: (2026)
Closing the Sim2Real Performance Gap in RL
von: Anand, Akhil S, et al.
Veröffentlicht: (2025)
von: Anand, Akhil S, et al.
Veröffentlicht: (2025)
COPA: Comparing the incomparable in multi-objective model evaluation
von: Javaloy, Adrián, et al.
Veröffentlicht: (2025)
von: Javaloy, Adrián, et al.
Veröffentlicht: (2025)
Once Upon an Input: Reasoning via Per-Instance Program Synthesis
von: Stein, Adam, et al.
Veröffentlicht: (2025)
von: Stein, Adam, et al.
Veröffentlicht: (2025)
Small transformer architectures for task switching
von: Gros, Claudius
Veröffentlicht: (2025)
von: Gros, Claudius
Veröffentlicht: (2025)
Reorganizing attention-space geometry with expressive attention
von: Gros, Claudius
Veröffentlicht: (2024)
von: Gros, Claudius
Veröffentlicht: (2024)
Automating the Generation of Prompts for LLM-based Action Choice in PDDL Planning
von: Stein, Katharina, et al.
Veröffentlicht: (2023)
von: Stein, Katharina, et al.
Veröffentlicht: (2023)
Enhancing GNNs with Architecture-Agnostic Graph Transformations: A Systematic Analysis
von: Li, Zhifei, et al.
Veröffentlicht: (2024)
von: Li, Zhifei, et al.
Veröffentlicht: (2024)
Towards Reasonable Concept Bottleneck Models
von: Kalampalikis, Nektarios, et al.
Veröffentlicht: (2025)
von: Kalampalikis, Nektarios, et al.
Veröffentlicht: (2025)
First-See-Then-Design: A Multi-Stakeholder View for Optimal Performance-Fairness Trade-Offs
von: Gupta, Kavya, et al.
Veröffentlicht: (2026)
von: Gupta, Kavya, et al.
Veröffentlicht: (2026)
A Causal Framework to Measure and Mitigate Non-binary Treatment Discrimination
von: Majumdar, Ayan, et al.
Veröffentlicht: (2025)
von: Majumdar, Ayan, et al.
Veröffentlicht: (2025)
Learning More Expressive General Policies for Classical Planning Domains
von: Ståhlberg, Simon, et al.
Veröffentlicht: (2024)
von: Ståhlberg, Simon, et al.
Veröffentlicht: (2024)
Diversifying Policy Behaviors with Extrinsic Behavioral Curiosity
von: Wan, Zhenglin, et al.
Veröffentlicht: (2024)
von: Wan, Zhenglin, et al.
Veröffentlicht: (2024)
Learning Generalized Policies for Fully Observable Non-Deterministic Planning Domains
von: Hofmann, Till, et al.
Veröffentlicht: (2024)
von: Hofmann, Till, et al.
Veröffentlicht: (2024)
Uncertainty Calibration with Energy Based Instance-wise Scaling in the Wild Dataset
von: Kim, Mijoo, et al.
Veröffentlicht: (2024)
von: Kim, Mijoo, et al.
Veröffentlicht: (2024)
Efficient and Interpretable Traffic Destination Prediction using Explainable Boosting Machines
von: Yousif, Yasin, et al.
Veröffentlicht: (2024)
von: Yousif, Yasin, et al.
Veröffentlicht: (2024)
PerAda: Parameter-Efficient Federated Learning Personalization with Generalization Guarantees
von: Xie, Chulin, et al.
Veröffentlicht: (2023)
von: Xie, Chulin, et al.
Veröffentlicht: (2023)
Neural Network-based Information Set Weighting for Playing Reconnaissance Blind Chess
von: Bertram, Timo, et al.
Veröffentlicht: (2024)
von: Bertram, Timo, et al.
Veröffentlicht: (2024)
Exploration Behavior of Untrained Policies
von: Adamczyk, Jacob
Veröffentlicht: (2025)
von: Adamczyk, Jacob
Veröffentlicht: (2025)
Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation
von: Zhou, Hongyi, et al.
Veröffentlicht: (2025)
von: Zhou, Hongyi, et al.
Veröffentlicht: (2025)
From Universal to Individualized Actionability: Revisiting Personalization in Algorithmic Recourse
von: Budde, Lena Marie, et al.
Veröffentlicht: (2026)
von: Budde, Lena Marie, et al.
Veröffentlicht: (2026)
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization
von: Palenicek, Daniel, et al.
Veröffentlicht: (2025)
von: Palenicek, Daniel, et al.
Veröffentlicht: (2025)
Symmetric Behavior Regularized Policy Optimization
von: Zhu, Lingwei, et al.
Veröffentlicht: (2025)
von: Zhu, Lingwei, et al.
Veröffentlicht: (2025)
A Practical Approach to Causal Inference over Time
von: Cinquini, Martina, et al.
Veröffentlicht: (2024)
von: Cinquini, Martina, et al.
Veröffentlicht: (2024)
Learning Policy Representations for Steerable Behavior Synthesis
von: Li, Beiming, et al.
Veröffentlicht: (2026)
von: Li, Beiming, et al.
Veröffentlicht: (2026)
Trust-Region Behavior Blending for On-Policy Distillation
von: Plyusov, Daniil, et al.
Veröffentlicht: (2026)
von: Plyusov, Daniil, et al.
Veröffentlicht: (2026)
Policy Optimization for Personalized Interventions in Behavioral Health
von: Baek, Jackie, et al.
Veröffentlicht: (2023)
von: Baek, Jackie, et al.
Veröffentlicht: (2023)
Towards Context-Aware Domain Generalization: Understanding the Benefits and Limits of Marginal Transfer Learning
von: Müller, Jens, et al.
Veröffentlicht: (2023)
von: Müller, Jens, et al.
Veröffentlicht: (2023)
Memorize Theorems, Not Instances: Probing SFT Generalization through Mathematical Reasoning
von: Peng, Ruiying, et al.
Veröffentlicht: (2026)
von: Peng, Ruiying, et al.
Veröffentlicht: (2026)
PePR: Performance Per Resource Unit as a Metric to Promote Small-Scale Deep Learning in Medical Image Analysis
von: Selvan, Raghavendra, et al.
Veröffentlicht: (2024)
von: Selvan, Raghavendra, et al.
Veröffentlicht: (2024)
General Flexible $f$-divergence for Challenging Offline RL Datasets with Low Stochasticity and Diverse Behavior Policies
von: Wang, Jianxun, et al.
Veröffentlicht: (2026)
von: Wang, Jianxun, et al.
Veröffentlicht: (2026)
Local Coherence or Global Validity? Investigating RLVR Traces in Math Domains
von: Samineni, Soumya Rani, et al.
Veröffentlicht: (2025)
von: Samineni, Soumya Rani, et al.
Veröffentlicht: (2025)
Cross-Domain Policy Adaptation by Capturing Representation Mismatch
von: Lyu, Jiafei, et al.
Veröffentlicht: (2024)
von: Lyu, Jiafei, et al.
Veröffentlicht: (2024)
Meta-Instance Selection. Instance Selection as a Classification Problem with Meta-Features
von: Blachnik, Marcin, et al.
Veröffentlicht: (2025)
von: Blachnik, Marcin, et al.
Veröffentlicht: (2025)
Pragmatic Policy Development via Interpretable Behavior Cloning
von: Matsson, Anton, et al.
Veröffentlicht: (2025)
von: Matsson, Anton, et al.
Veröffentlicht: (2025)
Deep Causal Behavioral Policy Learning: Applications to Healthcare
von: Knecht, Jonas, et al.
Veröffentlicht: (2025)
von: Knecht, Jonas, et al.
Veröffentlicht: (2025)
From Parameters to Behaviors: Unsupervised Compression of the Policy Space
von: Tenedini, Davide, et al.
Veröffentlicht: (2025)
von: Tenedini, Davide, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Per-Domain Generalizing Policies: On Learning Efficient and Robust Q-Value Functions (Extended Version with Technical Appendix)
von: Müller, Nicola J., et al.
Veröffentlicht: (2026) -
Acting for the Right Reasons: Creating Reason-Sensitive Artificial Moral Agents
von: Baum, Kevin, et al.
Veröffentlicht: (2024) -
Exploring Molecule Generation Using Latent Space Graph Diffusion
von: Pombala, Prashanth, et al.
Veröffentlicht: (2025) -
Hellinger Multimodal Variational Autoencoders
von: Vo, Huyen, et al.
Veröffentlicht: (2026) -
Closing the Sim2Real Performance Gap in RL
von: Anand, Akhil S, et al.
Veröffentlicht: (2025)