Gespeichert in:
| Hauptverfasser: | Pecháč, Matej, Chovanec, Michal, Farkaš, Igor |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2302.11563 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Robot at the Mirror: Learning to Imitate via Associating Self-supervised Models
von: Lucny, Andrej, et al.
Veröffentlicht: (2023)
von: Lucny, Andrej, et al.
Veröffentlicht: (2023)
The impact of intrinsic rewards on exploration in Reinforcement Learning
von: Kayal, Aya, et al.
Veröffentlicht: (2025)
von: Kayal, Aya, et al.
Veröffentlicht: (2025)
Appearance-based gaze estimation enhanced with synthetic images using deep neural networks
von: Herashchenko, Dmytro, et al.
Veröffentlicht: (2023)
von: Herashchenko, Dmytro, et al.
Veröffentlicht: (2023)
Safe Reinforcement Learning in a Simulated Robotic Arm
von: Kovač, Luka, et al.
Veröffentlicht: (2023)
von: Kovač, Luka, et al.
Veröffentlicht: (2023)
Autonomous state-space segmentation for Deep-RL sparse reward scenarios
von: Maselli, Gianluca, et al.
Veröffentlicht: (2025)
von: Maselli, Gianluca, et al.
Veröffentlicht: (2025)
Self-rewarding correction for mathematical reasoning
von: Xiong, Wei, et al.
Veröffentlicht: (2025)
von: Xiong, Wei, et al.
Veröffentlicht: (2025)
Diverse Feature Learning by Self-distillation and Reset
von: Park, Sejik
Veröffentlicht: (2024)
von: Park, Sejik
Veröffentlicht: (2024)
Self-supervised Pre-training of Text Recognizers
von: Kišš, Martin, et al.
Veröffentlicht: (2024)
von: Kišš, Martin, et al.
Veröffentlicht: (2024)
Retro: Reusing teacher projection head for efficient embedding distillation on Lightweight Models via Self-supervised Learning
von: Nguyen, Khanh-Binh, et al.
Veröffentlicht: (2024)
von: Nguyen, Khanh-Binh, et al.
Veröffentlicht: (2024)
Learning Low-Level Causal Relations using a Simulated Robotic Arm
von: Cibula, Miroslav, et al.
Veröffentlicht: (2024)
von: Cibula, Miroslav, et al.
Veröffentlicht: (2024)
Streaming Looking Ahead with Token-level Self-reward
von: Zhang, Hongming, et al.
Veröffentlicht: (2025)
von: Zhang, Hongming, et al.
Veröffentlicht: (2025)
LeanTree: Accelerating White-Box Proof Search with Factorized States in Lean 4
von: Kripner, Matěj, et al.
Veröffentlicht: (2025)
von: Kripner, Matěj, et al.
Veröffentlicht: (2025)
Pessimistic Off-Policy Optimization for Learning to Rank
von: Cief, Matej, et al.
Veröffentlicht: (2022)
von: Cief, Matej, et al.
Veröffentlicht: (2022)
BLEUBERI: BLEU is a surprisingly effective reward for instruction following
von: Chang, Yapei, et al.
Veröffentlicht: (2025)
von: Chang, Yapei, et al.
Veröffentlicht: (2025)
Just Say What You Want: Only-prompting Self-rewarding Online Preference Optimization
von: Xu, Ruijie, et al.
Veröffentlicht: (2024)
von: Xu, Ruijie, et al.
Veröffentlicht: (2024)
SPARE: Self-distillation for PARameter-Efficient Removal
von: Mola, Natnael, et al.
Veröffentlicht: (2026)
von: Mola, Natnael, et al.
Veröffentlicht: (2026)
FedMSE: Semi-supervised federated learning approach for IoT network intrusion detection
von: Nguyen, Van Tuan, et al.
Veröffentlicht: (2024)
von: Nguyen, Van Tuan, et al.
Veröffentlicht: (2024)
What should be observed for optimal reward in POMDPs?
von: Konsta, Alyzia-Maria, et al.
Veröffentlicht: (2024)
von: Konsta, Alyzia-Maria, et al.
Veröffentlicht: (2024)
Active teacher selection for reward learning
von: Freedman, Rachel, et al.
Veröffentlicht: (2023)
von: Freedman, Rachel, et al.
Veröffentlicht: (2023)
Noise-based reward-modulated learning
von: Fernández, Jesús García, et al.
Veröffentlicht: (2025)
von: Fernández, Jesús García, et al.
Veröffentlicht: (2025)
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
von: Liu, Shih-Yang, et al.
Veröffentlicht: (2026)
von: Liu, Shih-Yang, et al.
Veröffentlicht: (2026)
Towards better dense rewards in Reinforcement Learning Applications
von: Zhang, Shuyuan
Veröffentlicht: (2025)
von: Zhang, Shuyuan
Veröffentlicht: (2025)
sDREAMER: Self-distilled Mixture-of-Modality-Experts Transformer for Automatic Sleep Staging
von: Chen, Jingyuan, et al.
Veröffentlicht: (2025)
von: Chen, Jingyuan, et al.
Veröffentlicht: (2025)
Information-theoretic analysis of world models in optimal reward maximizers
von: Harwood, Alfred, et al.
Veröffentlicht: (2026)
von: Harwood, Alfred, et al.
Veröffentlicht: (2026)
Resistance of Trapezoidal Sheeting in Fire
von: Aleš Chovanec, et al.
Veröffentlicht: (2025)
von: Aleš Chovanec, et al.
Veröffentlicht: (2025)
Towards Lightweight and Stable Zero-shot TTS with Self-distilled Representation Disentanglement
von: Chen, Qianniu, et al.
Veröffentlicht: (2025)
von: Chen, Qianniu, et al.
Veröffentlicht: (2025)
Self-supervised Hierarchical Visual Reasoning with World Model
von: Xu, Yuanfei, et al.
Veröffentlicht: (2026)
von: Xu, Yuanfei, et al.
Veröffentlicht: (2026)
Toward effective protection against diffusion based mimicry through score distillation
von: Xue, Haotian, et al.
Veröffentlicht: (2023)
von: Xue, Haotian, et al.
Veröffentlicht: (2023)
Education distillation:getting student models to learn in shcools
von: Feng, Ling, et al.
Veröffentlicht: (2023)
von: Feng, Ling, et al.
Veröffentlicht: (2023)
On the expressivity of sparse maxout networks
von: Grillo, Moritz, et al.
Veröffentlicht: (2025)
von: Grillo, Moritz, et al.
Veröffentlicht: (2025)
EVAL: EigenVector-based Average-reward Learning
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2025)
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2025)
Episodic Reinforcement Learning with Expanded State-reward Space
von: Liang, Dayang, et al.
Veröffentlicht: (2024)
von: Liang, Dayang, et al.
Veröffentlicht: (2024)
Numerical exploration of the range of shape functionals using neural networks
von: Martinet, Eloi, et al.
Veröffentlicht: (2026)
von: Martinet, Eloi, et al.
Veröffentlicht: (2026)
Smooth-Distill: A Self-distillation Framework for Multitask Learning with Wearable Sensor Data
von: Vu, Hoang-Dieu, et al.
Veröffentlicht: (2025)
von: Vu, Hoang-Dieu, et al.
Veröffentlicht: (2025)
SMS: Self-supervised Model Seeding for Verification of Machine Unlearning
von: Wang, Weiqi, et al.
Veröffentlicht: (2025)
von: Wang, Weiqi, et al.
Veröffentlicht: (2025)
Self-supervised learning on gene expression data
von: Dradjat, Kevin, et al.
Veröffentlicht: (2025)
von: Dradjat, Kevin, et al.
Veröffentlicht: (2025)
FIT-SLAM -- Fisher Information and Traversability estimation-based Active SLAM for exploration in 3D environments
von: Saravanan, Suchetan, et al.
Veröffentlicht: (2024)
von: Saravanan, Suchetan, et al.
Veröffentlicht: (2024)
IDLM: Inverse-distilled Diffusion Language Models
von: Li, David, et al.
Veröffentlicht: (2026)
von: Li, David, et al.
Veröffentlicht: (2026)
R-ParVI: Particle-based variational inference through lens of rewards
von: Huang, Yongchao
Veröffentlicht: (2025)
von: Huang, Yongchao
Veröffentlicht: (2025)
Deep Reinforcement Learning with anticipatory reward in LSTM for Collision Avoidance of Mobile Robots
von: Poulet, Olivier, et al.
Veröffentlicht: (2025)
von: Poulet, Olivier, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Robot at the Mirror: Learning to Imitate via Associating Self-supervised Models
von: Lucny, Andrej, et al.
Veröffentlicht: (2023) -
The impact of intrinsic rewards on exploration in Reinforcement Learning
von: Kayal, Aya, et al.
Veröffentlicht: (2025) -
Appearance-based gaze estimation enhanced with synthetic images using deep neural networks
von: Herashchenko, Dmytro, et al.
Veröffentlicht: (2023) -
Safe Reinforcement Learning in a Simulated Robotic Arm
von: Kovač, Luka, et al.
Veröffentlicht: (2023) -
Autonomous state-space segmentation for Deep-RL sparse reward scenarios
von: Maselli, Gianluca, et al.
Veröffentlicht: (2025)