Saved in:
| Main Authors: | Voelcker, Claas, Pedan, Anastasiia, Ahmadian, Arash, Abachi, Romina, Gilitschenski, Igor, Farahmand, Amir-massoud |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2505.22772 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
$λ$-models: Effective Decision-Aware Reinforcement Learning with Latent Models
by: Voelcker, Claas A, et al.
Published: (2023)
by: Voelcker, Claas A, et al.
Published: (2023)
When does Self-Prediction help? Understanding Auxiliary Tasks in Reinforcement Learning
by: Voelcker, Claas, et al.
Published: (2024)
by: Voelcker, Claas, et al.
Published: (2024)
Dissecting Deep RL with High Update Ratios: Combatting Value Divergence
by: Hussing, Marcel, et al.
Published: (2024)
by: Hussing, Marcel, et al.
Published: (2024)
MAD-TD: Model-Augmented Data stabilizes High Update Ratio RL
by: Voelcker, Claas A, et al.
Published: (2024)
by: Voelcker, Claas A, et al.
Published: (2024)
Relative Entropy Pathwise Policy Optimization
by: Voelcker, Claas, et al.
Published: (2025)
by: Voelcker, Claas, et al.
Published: (2025)
Test-Time Graph Search for Goal-Conditioned Reinforcement Learning
by: Opryshko, Evgenii, et al.
Published: (2025)
by: Opryshko, Evgenii, et al.
Published: (2025)
Deflated Dynamics Value Iteration
by: Lee, Jongmin, et al.
Published: (2024)
by: Lee, Jongmin, et al.
Published: (2024)
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling
by: Ma, Avery, et al.
Published: (2025)
by: Ma, Avery, et al.
Published: (2025)
PID Accelerated Temporal Difference Algorithms
by: Bedaywi, Mark, et al.
Published: (2024)
by: Bedaywi, Mark, et al.
Published: (2024)
Improving Adversarial Transferability via Model Alignment
by: Ma, Avery, et al.
Published: (2023)
by: Ma, Avery, et al.
Published: (2023)
Efficient and Accurate Optimal Transport with Mirror Descent and Conjugate Gradients
by: Kemertas, Mete, et al.
Published: (2023)
by: Kemertas, Mete, et al.
Published: (2023)
Press Start to Charge: Videogaming the Online Centralized Charging Scheduling Problem
by: Ghahtarani, Alireza, et al.
Published: (2026)
by: Ghahtarani, Alireza, et al.
Published: (2026)
A Truncated Newton Method for Optimal Transport
by: Kemertas, Mete, et al.
Published: (2025)
by: Kemertas, Mete, et al.
Published: (2025)
Majority of the Bests: Improving Best-of-N via Bootstrapping
by: Rakhsha, Amin, et al.
Published: (2025)
by: Rakhsha, Amin, et al.
Published: (2025)
Can we hop in general? A discussion of benchmark selection and design using the Hopper environment
by: Voelcker, Claas A, et al.
Published: (2024)
by: Voelcker, Claas A, et al.
Published: (2024)
Behavior-Consistent Deep Reinforcement Learning
by: Hussing, Marcel, et al.
Published: (2026)
by: Hussing, Marcel, et al.
Published: (2026)
Mix Data or Merge Models? Optimizing for Diverse Multi-Task Learning
by: Aakanksha, et al.
Published: (2024)
by: Aakanksha, et al.
Published: (2024)
Temporal-Difference Learning Using Distributed Error Signals
by: Guan, Jonas, et al.
Published: (2024)
by: Guan, Jonas, et al.
Published: (2024)
TESPEC: Temporally-Enhanced Self-Supervised Pretraining for Event Cameras
by: Mohammadi, Mohammad, et al.
Published: (2025)
by: Mohammadi, Mohammad, et al.
Published: (2025)
Realistic Evaluation of Model Merging for Compositional Generalization
by: Tam, Derek, et al.
Published: (2024)
by: Tam, Derek, et al.
Published: (2024)
SimMerge: Learning to Select Merge Operators from Similarity Signals
by: Bolton, Oliver, et al.
Published: (2026)
by: Bolton, Oliver, et al.
Published: (2026)
Explaining Probabilistic Models with Distributional Values
by: Franceschi, Luca, et al.
Published: (2024)
by: Franceschi, Luca, et al.
Published: (2024)
Probabilistic Shapley Value Modeling and Inference
by: Ketenci, Mert, et al.
Published: (2024)
by: Ketenci, Mert, et al.
Published: (2024)
Track, Inpaint, Resplat: Subject-driven 3D and 4D Generation with Progressive Texture Infilling
by: Zheng, Shuhong, et al.
Published: (2025)
by: Zheng, Shuhong, et al.
Published: (2025)
Generalized Top-k Mallows Model for Ranked Choices
by: Haddadan, Shahrzad, et al.
Published: (2025)
by: Haddadan, Shahrzad, et al.
Published: (2025)
Vectorized Context-Aware Embeddings for GAT-Based Collaborative Filtering
by: Ebrat, Danial, et al.
Published: (2025)
by: Ebrat, Danial, et al.
Published: (2025)
Zero-Shot Uncertainty Quantification using Diffusion Probabilistic Models
by: Shu, Dule, et al.
Published: (2024)
by: Shu, Dule, et al.
Published: (2024)
Sorrel: A simple and flexible framework for multi-agent reinforcement learning
by: Gelpí, Rebekah A., et al.
Published: (2025)
by: Gelpí, Rebekah A., et al.
Published: (2025)
Self-Improving Robust Preference Optimization
by: Choi, Eugene, et al.
Published: (2024)
by: Choi, Eugene, et al.
Published: (2024)
Neural Assets: 3D-Aware Multi-Object Scene Synthesis with Image Diffusion Models
by: Wu, Ziyi, et al.
Published: (2024)
by: Wu, Ziyi, et al.
Published: (2024)
Confidence-Aware Multi-Field Model Calibration
by: Zhao, Yuang, et al.
Published: (2024)
by: Zhao, Yuang, et al.
Published: (2024)
Squeezing Capacity from Multimodal Large Language Models for Subject-driven Generation
by: Zheng, Shuhong, et al.
Published: (2026)
by: Zheng, Shuhong, et al.
Published: (2026)
Producing and Leveraging Online Map Uncertainty in Trajectory Prediction
by: Gu, Xunjiang, et al.
Published: (2024)
by: Gu, Xunjiang, et al.
Published: (2024)
Accelerating Online Mapping and Behavior Prediction via Direct BEV Feature Attention
by: Gu, Xunjiang, et al.
Published: (2024)
by: Gu, Xunjiang, et al.
Published: (2024)
Probabilistic Scores of Classifiers, Calibration is not Enough
by: Machado, Agathe Fernandes, et al.
Published: (2024)
by: Machado, Agathe Fernandes, et al.
Published: (2024)
Calibrated Probabilistic Forecasts for Arbitrary Sequences
by: Marx, Charles, et al.
Published: (2024)
by: Marx, Charles, et al.
Published: (2024)
End-to-End Personalization: Unifying Recommender Systems with Large Language Models
by: Ebrat, Danial, et al.
Published: (2025)
by: Ebrat, Danial, et al.
Published: (2025)
VEM: Environment-Free Exploration for Training GUI Agent with Value Environment Model
by: Zheng, Jiani, et al.
Published: (2025)
by: Zheng, Jiani, et al.
Published: (2025)
Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
by: Ahmadian, Arash, et al.
Published: (2024)
by: Ahmadian, Arash, et al.
Published: (2024)
RLHF Can Speak Many Languages: Unlocking Multilingual Preference Optimization for LLMs
by: Dang, John, et al.
Published: (2024)
by: Dang, John, et al.
Published: (2024)
Similar Items
-
$λ$-models: Effective Decision-Aware Reinforcement Learning with Latent Models
by: Voelcker, Claas A, et al.
Published: (2023) -
When does Self-Prediction help? Understanding Auxiliary Tasks in Reinforcement Learning
by: Voelcker, Claas, et al.
Published: (2024) -
Dissecting Deep RL with High Update Ratios: Combatting Value Divergence
by: Hussing, Marcel, et al.
Published: (2024) -
MAD-TD: Model-Augmented Data stabilizes High Update Ratio RL
by: Voelcker, Claas A, et al.
Published: (2024) -
Relative Entropy Pathwise Policy Optimization
by: Voelcker, Claas, et al.
Published: (2025)