RynnBrain: Open Embodied Foundation Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dang, Ronghao, Guo, Jiayan, Hou, Bohan, Leng, Sicong, Li, Kehan, Li, Xin, Liu, Jiangpin, Mao, Yunxuan, Wang, Zhikai, Yuan, Yuqian, Zhu, Minghao, Lin, Xiao, Bai, Yang, Jiang, Qian, Zhao, Yaxi, Zeng, Minghua, Gao, Junlong, Jiang, Yuming, Cen, Jun, Huang, Siteng, Wang, Liuyi, Zhang, Wenqiao, Liu, Chengju, Yang, Jianfei, Lu, Shijian, Zhao, Deli |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RynnEC: Bringing MLLMs into Embodied World
von: Dang, Ronghao, et al.
Veröffentlicht: (2025)
von: Dang, Ronghao, et al.
Veröffentlicht: (2025)
RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
von: Jiang, Yuming, et al.
Veröffentlicht: (2025)
von: Jiang, Yuming, et al.
Veröffentlicht: (2025)
RynnVLA-002: A Unified Vision-Language-Action and World Model
von: Cen, Jun, et al.
Veröffentlicht: (2025)
von: Cen, Jun, et al.
Veröffentlicht: (2025)
InterLV-Search: Benchmarking Interleaved Multimodal Agentic Search
von: Hou, Bohan, et al.
Veröffentlicht: (2026)
von: Hou, Bohan, et al.
Veröffentlicht: (2026)
Causality-based Cross-Modal Representation Learning for Vision-and-Language Navigation
von: Wang, Liuyi, et al.
Veröffentlicht: (2024)
von: Wang, Liuyi, et al.
Veröffentlicht: (2024)
Vision-and-Language Navigation via Causal Learning
von: Wang, Liuyi, et al.
Veröffentlicht: (2024)
von: Wang, Liuyi, et al.
Veröffentlicht: (2024)
Towards Affordance-Aware Robotic Dexterous Grasping with Human-like Priors
von: Zhao, Haoyu, et al.
Veröffentlicht: (2025)
von: Zhao, Haoyu, et al.
Veröffentlicht: (2025)
A Dual Semantic-Aware Recurrent Global-Adaptive Network For Vision-and-Language Navigation
von: Wang, Liuyi, et al.
Veröffentlicht: (2023)
von: Wang, Liuyi, et al.
Veröffentlicht: (2023)
WorldVLA: Towards Autoregressive Action World Model
von: Cen, Jun, et al.
Veröffentlicht: (2025)
von: Cen, Jun, et al.
Veröffentlicht: (2025)
High-Fidelity Simulated Data Generation for Real-World Zero-Shot Robotic Manipulation Learning with Gaussian Splatting
von: Zhao, Haoyu, et al.
Veröffentlicht: (2025)
von: Zhao, Haoyu, et al.
Veröffentlicht: (2025)
EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?
von: Yuan, Yuqian, et al.
Veröffentlicht: (2025)
von: Yuan, Yuqian, et al.
Veröffentlicht: (2025)
Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agents
von: Li, Long, et al.
Veröffentlicht: (2024)
von: Li, Long, et al.
Veröffentlicht: (2024)
MAGIC: Meta-Ability Guided Interactive Chain-of-Distillation for Effective-and-Efficient Vision-and-Language Navigation
von: Wang, Liuyi, et al.
Veröffentlicht: (2024)
von: Wang, Liuyi, et al.
Veröffentlicht: (2024)
NavComposer: Composing Language Instructions for Navigation Trajectories through Action-Scene-Object Modularization
von: He, Zongtao, et al.
Veröffentlicht: (2025)
von: He, Zongtao, et al.
Veröffentlicht: (2025)
The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio
von: Leng, Sicong, et al.
Veröffentlicht: (2024)
von: Leng, Sicong, et al.
Veröffentlicht: (2024)
Pathway‐Dependent Self‐Assembly for Control over Helical Nanostructures and Topochemical Photopolymerization
von: Sifan Du, et al.
Veröffentlicht: (2024)
von: Sifan Du, et al.
Veröffentlicht: (2024)
Breaking the Memory Barrier: Near Infinite Batch Size Scaling for Contrastive Loss
von: Cheng, Zesen, et al.
Veröffentlicht: (2024)
von: Cheng, Zesen, et al.
Veröffentlicht: (2024)
Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs
von: Jiang, Xueying, et al.
Veröffentlicht: (2026)
von: Jiang, Xueying, et al.
Veröffentlicht: (2026)
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
von: Zhang, Boqiang, et al.
Veröffentlicht: (2025)
von: Zhang, Boqiang, et al.
Veröffentlicht: (2025)
Fine-Grained Spatiotemporal Motion Alignment for Contrastive Video Representation Learning
von: Zhu, Minghao, et al.
Veröffentlicht: (2023)
von: Zhu, Minghao, et al.
Veröffentlicht: (2023)
Domain-Conditioned Scene Graphs for State-Grounded Task Planning
von: Herzog, Jonas, et al.
Veröffentlicht: (2025)
von: Herzog, Jonas, et al.
Veröffentlicht: (2025)
Intrapersonal Interactions on Social Media: Self‐Curatorship within the Museum of Self
von: Sicong Zhao
Veröffentlicht: (2025)
von: Sicong Zhao
Veröffentlicht: (2025)
MLANet: Multi-Level Attention Network with Sub-instruction for Continuous Vision-and-Language Navigation
von: He, Zongtao, et al.
Veröffentlicht: (2023)
von: He, Zongtao, et al.
Veröffentlicht: (2023)
Temporal-Guided Visual Foundation Models for Event-Based Vision
von: Xia, Ruihao, et al.
Veröffentlicht: (2025)
von: Xia, Ruihao, et al.
Veröffentlicht: (2025)
GP3: A 3D Geometry-Aware Policy with Multi-View Images for Robotic Manipulation
von: Qian, Quanhao, et al.
Veröffentlicht: (2025)
von: Qian, Quanhao, et al.
Veröffentlicht: (2025)
RoboSVG: A Unified Framework for Interactive SVG Generation with Multi-modal Guidance
von: Wang, Jiuniu, et al.
Veröffentlicht: (2025)
von: Wang, Jiuniu, et al.
Veröffentlicht: (2025)
On isolated singularities of the conformal Gaussian curvature equation and $Q$-curvature equation
von: Yang, Hui, et al.
Veröffentlicht: (2025)
von: Yang, Hui, et al.
Veröffentlicht: (2025)
SignalLLM: A General-Purpose LLM Agent Framework for Automated Signal Processing
von: Ke, Junlong, et al.
Veröffentlicht: (2025)
von: Ke, Junlong, et al.
Veröffentlicht: (2025)
Joint Optimization of STAR-RIS Assisted SWIPT Communication Systems
von: Yang, Junlong
Veröffentlicht: (2024)
von: Yang, Junlong
Veröffentlicht: (2024)
Classification of Transposed Poisson 3-Lie algebras of dimension 3
von: Yaxi, Jiang, et al.
Veröffentlicht: (2024)
von: Yaxi, Jiang, et al.
Veröffentlicht: (2024)
Manin triples, bialgebras and Yang-Baxter equation of $A_3$-associative algebras
von: Jiang, Yaxi, et al.
Veröffentlicht: (2025)
von: Jiang, Yaxi, et al.
Veröffentlicht: (2025)
Rethinking the Embodied Gap in Vision-and-Language Navigation: A Holistic Study of Physical and Visual Disparities
von: Wang, Liuyi, et al.
Veröffentlicht: (2025)
von: Wang, Liuyi, et al.
Veröffentlicht: (2025)
Deconfounded Lifelong Learning for Autonomous Driving via Dynamic Knowledge Spaces
von: Du, Jiayuan, et al.
Veröffentlicht: (2026)
von: Du, Jiayuan, et al.
Veröffentlicht: (2026)
REI-Bench: Can Embodied Agents Understand Vague Human Instructions in Task Planning?
von: Jiang, Chenxi, et al.
Veröffentlicht: (2025)
von: Jiang, Chenxi, et al.
Veröffentlicht: (2025)
PASTS: Progress-Aware Spatio-Temporal Transformer Speaker For Vision-and-Language Navigation
von: Wang, Liuyi, et al.
Veröffentlicht: (2023)
von: Wang, Liuyi, et al.
Veröffentlicht: (2023)
Handedness‐Inverted and Stimuli‐Responsive Circularly Polarized Luminescent Nano/Micromaterials Through Pathway‐Dependent Chiral Supramolecular Polymorphism
von: Chenyang Zhao, et al.
Veröffentlicht: (2024)
von: Chenyang Zhao, et al.
Veröffentlicht: (2024)
PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity
von: Yuan, Yuqian, et al.
Veröffentlicht: (2025)
von: Yuan, Yuqian, et al.
Veröffentlicht: (2025)
ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmark
von: Dang, Ronghao, et al.
Veröffentlicht: (2025)
von: Dang, Ronghao, et al.
Veröffentlicht: (2025)
CLASH: Collaborative Large-Small Hierarchical Framework for Continuous Vision-and-Language Navigation
von: Wang, Liuyi, et al.
Veröffentlicht: (2025)
von: Wang, Liuyi, et al.
Veröffentlicht: (2025)
CleanPose: Category-Level Object Pose Estimation via Causal Learning and Knowledge Distillation
von: Lin, Xiao, et al.
Veröffentlicht: (2025)
von: Lin, Xiao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
RynnEC: Bringing MLLMs into Embodied World
von: Dang, Ronghao, et al.
Veröffentlicht: (2025) -
RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
von: Jiang, Yuming, et al.
Veröffentlicht: (2025) -
RynnVLA-002: A Unified Vision-Language-Action and World Model
von: Cen, Jun, et al.
Veröffentlicht: (2025) -
InterLV-Search: Benchmarking Interleaved Multimodal Agentic Search
von: Hou, Bohan, et al.
Veröffentlicht: (2026) -
Causality-based Cross-Modal Representation Learning for Vision-and-Language Navigation
von: Wang, Liuyi, et al.
Veröffentlicht: (2024)