Salvato in:
| Autori principali: | Mower, Christopher E., Wan, Yuhui, Yu, Hongzhan, Grosnit, Antoine, Gonzalez-Billandon, Jonas, Zimmer, Matthieu, Wang, Jinlong, Zhang, Xinyu, Zhao, Yao, Zhai, Anbang, Liu, Puze, Palenicek, Daniel, Tateo, Davide, Cadena, Cesar, Hutter, Marco, Peters, Jan, Tian, Guangjian, Zhuang, Yuzheng, Shao, Kun, Quan, Xingyue, Hao, Jianye, Wang, Jun, Bou-Ammar, Haitham |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2406.19741 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Safe Reinforcement Learning on the Constraint Manifold: Theory and Applications
di: Liu, Puze, et al.
Pubblicazione: (2024)
di: Liu, Puze, et al.
Pubblicazione: (2024)
Many of Your DPOs are Secretly One: Attempting Unification Through Mutual Information
di: Tutnov, Rasul, et al.
Pubblicazione: (2025)
di: Tutnov, Rasul, et al.
Pubblicazione: (2025)
Al-Khwarizmi: Discovering Physical Laws with Foundation Models
di: Mower, Christopher E., et al.
Pubblicazione: (2025)
di: Mower, Christopher E., et al.
Pubblicazione: (2025)
Contextual Causal Bayesian Optimisation
di: Arsenyan, Vahan, et al.
Pubblicazione: (2023)
di: Arsenyan, Vahan, et al.
Pubblicazione: (2023)
Trust Region Inverse Reinforcement Learning: Explicit Dual Ascent using Local Policy Updates
di: Diwan, Anish, et al.
Pubblicazione: (2026)
di: Diwan, Anish, et al.
Pubblicazione: (2026)
Data-driven Interpretable Hybrid Robot Dynamics
di: Mower, Christopher E., et al.
Pubblicazione: (2025)
di: Mower, Christopher E., et al.
Pubblicazione: (2025)
Why Can Large Language Models Generate Correct Chain-of-Thoughts?
di: Tutunov, Rasul, et al.
Pubblicazione: (2023)
di: Tutunov, Rasul, et al.
Pubblicazione: (2023)
ShortCircuit: AlphaZero-Driven Circuit Design
di: Tsaras, Dimitrios, et al.
Pubblicazione: (2024)
di: Tsaras, Dimitrios, et al.
Pubblicazione: (2024)
Model-Based and Sample-Efficient AI-Assisted Math Discovery in Sphere Packing
di: Tutunov, Rasul, et al.
Pubblicazione: (2025)
di: Tutunov, Rasul, et al.
Pubblicazione: (2025)
A call for embodied AI
di: Paolo, Giuseppe, et al.
Pubblicazione: (2024)
di: Paolo, Giuseppe, et al.
Pubblicazione: (2024)
A Pragmatist Robot: Learning to Plan Tasks by Experiencing the Real World
di: Qu, Kaixian, et al.
Pubblicazione: (2025)
di: Qu, Kaixian, et al.
Pubblicazione: (2025)
Decoding as Optimisation on the Probability Simplex: From Top-K to Top-P (Nucleus) to Best-of-K Samplers
di: Ji, Xiaotong, et al.
Pubblicazione: (2026)
di: Ji, Xiaotong, et al.
Pubblicazione: (2026)
Scalable Power Sampling: Unlocking Efficient, Training-Free Reasoning for LLMs via Distribution Sharpening
di: Ji, Xiaotong, et al.
Pubblicazione: (2026)
di: Ji, Xiaotong, et al.
Pubblicazione: (2026)
Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective
di: Zimmer, Matthieu, et al.
Pubblicazione: (2025)
di: Zimmer, Matthieu, et al.
Pubblicazione: (2025)
Towards Safe Robot Foundation Models
di: Tölle, Maximilian, et al.
Pubblicazione: (2025)
di: Tölle, Maximilian, et al.
Pubblicazione: (2025)
The Model Knows, the Decoder Finds: Future Value Guided Particle Power Sampling
di: Nguyen, Tu, et al.
Pubblicazione: (2026)
di: Nguyen, Tu, et al.
Pubblicazione: (2026)
Mixture of Attentions For Speculative Decoding
di: Zimmer, Matthieu, et al.
Pubblicazione: (2024)
di: Zimmer, Matthieu, et al.
Pubblicazione: (2024)
Risk-Controlled Lean-as-Judge for Natural-Language Mathematical Reasoning
di: Bourigault, Pauline, et al.
Pubblicazione: (2026)
di: Bourigault, Pauline, et al.
Pubblicazione: (2026)
The $\mathbf{Y}$-Combinator for LLMs: Solving Long-Context Rot with $λ$-Calculus
di: Roy, Amartya, et al.
Pubblicazione: (2026)
di: Roy, Amartya, et al.
Pubblicazione: (2026)
Towards Safe Robot Foundation Models Using Inductive Biases
di: Tölle, Maximilian, et al.
Pubblicazione: (2025)
di: Tölle, Maximilian, et al.
Pubblicazione: (2025)
Handling Long-Term Safety and Uncertainty in Safe Reinforcement Learning
di: Günster, Jonas, et al.
Pubblicazione: (2024)
di: Günster, Jonas, et al.
Pubblicazione: (2024)
Adaptive Control based Friction Estimation for Tracking Control of Robot Manipulators
di: Huang, Junning, et al.
Pubblicazione: (2024)
di: Huang, Junning, et al.
Pubblicazione: (2024)
Bourbaki: Self-Generated and Goal-Conditioned MDPs for Theorem Proving
di: Zimmer, Matthieu, et al.
Pubblicazione: (2025)
di: Zimmer, Matthieu, et al.
Pubblicazione: (2025)
LBR-Stack: ROS 2 and Python Integration of KUKA FRI for Med and IIWA Robots
di: Huber, Martin, et al.
Pubblicazione: (2023)
di: Huber, Martin, et al.
Pubblicazione: (2023)
On Almost Surely Safe Alignment of Large Language Models at Inference-Time
di: Ji, Xiaotong, et al.
Pubblicazione: (2025)
di: Ji, Xiaotong, et al.
Pubblicazione: (2025)
Materiobiomodulated ROS Therapy for De Novo Hair Growth
di: Long Bai, et al.
Pubblicazione: (2024)
di: Long Bai, et al.
Pubblicazione: (2024)
Bottlenecked Transformers: Periodic KV Cache Consolidation for Generalised Reasoning
di: Oomerjee, Adnan, et al.
Pubblicazione: (2025)
di: Oomerjee, Adnan, et al.
Pubblicazione: (2025)
Ark: An Open-source Python-based Framework for Robot Learning
di: Dierking, Magnus, et al.
Pubblicazione: (2025)
di: Dierking, Magnus, et al.
Pubblicazione: (2025)
Distilling Contact Planning for Fast Trajectory Optimization in Robot Air Hockey
di: Jankowski, Julius, et al.
Pubblicazione: (2024)
di: Jankowski, Julius, et al.
Pubblicazione: (2024)
Bridging the gap between Learning-to-plan, Motion Primitives and Safe Reinforcement Learning
di: Kicki, Piotr, et al.
Pubblicazione: (2024)
di: Kicki, Piotr, et al.
Pubblicazione: (2024)
SparsePO: Controlling Preference Alignment of LLMs via Sparse Token Masks
di: Christopoulou, Fenia, et al.
Pubblicazione: (2024)
di: Christopoulou, Fenia, et al.
Pubblicazione: (2024)
Why the Brain Consolidates: Predictive Forgetting for Optimal Generalisation
di: Fountas, Zafeirios, et al.
Pubblicazione: (2026)
di: Fountas, Zafeirios, et al.
Pubblicazione: (2026)
Subjective Depth and Timescale Transformers: Learning Where and When to Compute
di: Wieser, Frederico, et al.
Pubblicazione: (2025)
di: Wieser, Frederico, et al.
Pubblicazione: (2025)
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
di: Ramesh, Shyam Sundhar, et al.
Pubblicazione: (2026)
di: Ramesh, Shyam Sundhar, et al.
Pubblicazione: (2026)
EVLP:Learning Unified Embodied Vision-Language Planner with Reinforced Supervised Fine-Tuning
di: Cai, Xinyan, et al.
Pubblicazione: (2025)
di: Cai, Xinyan, et al.
Pubblicazione: (2025)
daleihao/RDycore_ROS: Codes and data for the RDycore-ROS study
di: Dalei Hao
Pubblicazione: (2025)
di: Dalei Hao
Pubblicazione: (2025)
ROS2swarm - A ROS 2 Package for Swarm Robot Behaviors
di: Kaiser, Tanja Katharina, et al.
Pubblicazione: (2024)
di: Kaiser, Tanja Katharina, et al.
Pubblicazione: (2024)
Untangling Component Imbalance in Hybrid Linear Attention Conversion Methods
di: Benfeghoul, Martin, et al.
Pubblicazione: (2025)
di: Benfeghoul, Martin, et al.
Pubblicazione: (2025)
Proxying ROS communications -- enabling containerized ROS deployments in distributed multi-host environments
di: Wendt, Arne, et al.
Pubblicazione: (2022)
di: Wendt, Arne, et al.
Pubblicazione: (2022)
Kolb-Based Experiential Learning for Generalist Agents with Human-Level Kaggle Data Science Performance
di: Grosnit, Antoine, et al.
Pubblicazione: (2024)
di: Grosnit, Antoine, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Safe Reinforcement Learning on the Constraint Manifold: Theory and Applications
di: Liu, Puze, et al.
Pubblicazione: (2024) -
Many of Your DPOs are Secretly One: Attempting Unification Through Mutual Information
di: Tutnov, Rasul, et al.
Pubblicazione: (2025) -
Al-Khwarizmi: Discovering Physical Laws with Foundation Models
di: Mower, Christopher E., et al.
Pubblicazione: (2025) -
Contextual Causal Bayesian Optimisation
di: Arsenyan, Vahan, et al.
Pubblicazione: (2023) -
Trust Region Inverse Reinforcement Learning: Explicit Dual Ascent using Local Policy Updates
di: Diwan, Anish, et al.
Pubblicazione: (2026)