Modality Selection and Skill Segmentation via Cross-Modality Attention
Fuente:
arXiv
Guardado en:
| Autores principales: | Jiang, Jiawei, Ota, Kei, Jha, Devesh K., Kanezaki, Asako |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Tactile Estimation of Extrinsic Contact Patch for Stable Placement
por: Ota, Kei, et al.
Publicado: (2023)
por: Ota, Kei, et al.
Publicado: (2023)
Zero-Shot Peg Insertion: Identifying Mating Holes and Estimating SE(2) Poses with Vision-Language Models
por: Yajima, Masaru, et al.
Publicado: (2025)
por: Yajima, Masaru, et al.
Publicado: (2025)
Learning Pivoting Manipulation with Force and Vision Feedback Using Optimization-based Demonstrations
por: Shirai, Yuki, et al.
Publicado: (2025)
por: Shirai, Yuki, et al.
Publicado: (2025)
Hierarchical Contact-Rich Trajectory Optimization for Multi-Modal Manipulation using Tight Convex Relaxations
por: Shirai, Yuki, et al.
Publicado: (2025)
por: Shirai, Yuki, et al.
Publicado: (2025)
Leveraging Large Language Model-based Room-Object Relationships Knowledge for Enhancing Multimodal-Input Object Goal Navigation
por: Sun, Leyuan, et al.
Publicado: (2024)
por: Sun, Leyuan, et al.
Publicado: (2024)
Touch2Insert: Zero-Shot Peg Insertion by Touching Intersections of Peg and Hole
por: Yajima, Masaru, et al.
Publicado: (2026)
por: Yajima, Masaru, et al.
Publicado: (2026)
Find the Fruit: Zero-Shot Sim2Real RL for Occlusion-Aware Plant Manipulation
por: Subedi, Nitesh, et al.
Publicado: (2025)
por: Subedi, Nitesh, et al.
Publicado: (2025)
Robust Pivoting Manipulation using Contact Implicit Bilevel Optimization
por: Shirai, Yuki, et al.
Publicado: (2023)
por: Shirai, Yuki, et al.
Publicado: (2023)
RecoveryChaining: Learning Local Recovery Policies for Robust Manipulation
por: Vats, Shivam, et al.
Publicado: (2024)
por: Vats, Shivam, et al.
Publicado: (2024)
D-CAT: Decoupled Cross-Attention Transfer between Sensor Modalities for Unimodal Inference
por: Daher, Leen, et al.
Publicado: (2025)
por: Daher, Leen, et al.
Publicado: (2025)
LCMF: Lightweight Cross-Modality Mambaformer for Embodied Robotics VQA
por: Kang, Zeyi, et al.
Publicado: (2025)
por: Kang, Zeyi, et al.
Publicado: (2025)
Embodied Navigation with Auxiliary Task of Action Description Prediction
por: Kondoh, Haru, et al.
Publicado: (2025)
por: Kondoh, Haru, et al.
Publicado: (2025)
Robust In-Hand Manipulation with Extrinsic Contacts
por: Liang, Boyuan, et al.
Publicado: (2024)
por: Liang, Boyuan, et al.
Publicado: (2024)
Multi-Modal Manipulation via Multi-Modal Policy Consensus
por: Chen, Haonan, et al.
Publicado: (2025)
por: Chen, Haonan, et al.
Publicado: (2025)
ACROSS: A Deformation-Based Cross-Modal Representation for Robotic Tactile Perception
por: Amri, Wadhah Zai El, et al.
Publicado: (2024)
por: Amri, Wadhah Zai El, et al.
Publicado: (2024)
M2R2: MultiModal Robotic Representation for Temporal Action Segmentation
por: Sliwowski, Daniel, et al.
Publicado: (2025)
por: Sliwowski, Daniel, et al.
Publicado: (2025)
Analytic Conditions for Differentiable Collision Detection in Trajectory Optimization
por: Jaitly, Akshay, et al.
Publicado: (2025)
por: Jaitly, Akshay, et al.
Publicado: (2025)
Imagine-2-Drive: Leveraging High-Fidelity World Models via Multi-Modal Diffusion Policies
por: Garg, Anant, et al.
Publicado: (2024)
por: Garg, Anant, et al.
Publicado: (2024)
Continuous Reasoning for Vision-Language-Action
por: Wu, Yueh-Hua, et al.
Publicado: (2026)
por: Wu, Yueh-Hua, et al.
Publicado: (2026)
FlowLoss: Dynamic Flow-Conditioned Loss Strategy for Video Diffusion Models
por: Wu, Kuanting, et al.
Publicado: (2025)
por: Wu, Kuanting, et al.
Publicado: (2025)
Mitigating Cross-Modal Distraction and Ensuring Geometric Feasibility via Affordance-Guided and Self-Consistent MLLMs for Task Planning in Instruction-Following Manipulation
por: Shen, Yu-Hong, et al.
Publicado: (2025)
por: Shen, Yu-Hong, et al.
Publicado: (2025)
Real-Time Out-of-Distribution Failure Prevention via Multi-Modal Reasoning
por: Ganai, Milan, et al.
Publicado: (2025)
por: Ganai, Milan, et al.
Publicado: (2025)
The Power of Combined Modalities in Interactive Robot Learning
por: Beierling, Helen, et al.
Publicado: (2024)
por: Beierling, Helen, et al.
Publicado: (2024)
Simultaneous Extrinsic Contact and In-Hand Pose Estimation via Distributed Tactile Sensing
por: Van der Merwe, Mark, et al.
Publicado: (2025)
por: Van der Merwe, Mark, et al.
Publicado: (2025)
Linking Vision and Multi-Agent Communication through Visible Light Communication using Event Cameras
por: Nakagawa, Haruyuki, et al.
Publicado: (2024)
por: Nakagawa, Haruyuki, et al.
Publicado: (2024)
DualCross: Cross-Modality Cross-Domain Adaptation for Monocular BEV Perception
por: Man, Yunze, et al.
Publicado: (2023)
por: Man, Yunze, et al.
Publicado: (2023)
Evidential Uncertainty Estimation for Multi-Modal Trajectory Prediction
por: Marvi, Sajad, et al.
Publicado: (2025)
por: Marvi, Sajad, et al.
Publicado: (2025)
GR-MG: Leveraging Partially Annotated Data via Multi-Modal Goal-Conditioned Policy
por: Li, Peiyan, et al.
Publicado: (2024)
por: Li, Peiyan, et al.
Publicado: (2024)
Cross-Modal Navigation with Multi-Agent Reinforcement Learning
por: Liu, Shuo, et al.
Publicado: (2026)
por: Liu, Shuo, et al.
Publicado: (2026)
Single-Frame Point-Pixel Registration via Supervised Cross-Modal Feature Matching
por: Han, Yu, et al.
Publicado: (2025)
por: Han, Yu, et al.
Publicado: (2025)
Trajectory Conditioned Cross-embodiment Skill Transfer
por: Tang, YuHang, et al.
Publicado: (2025)
por: Tang, YuHang, et al.
Publicado: (2025)
GOAT-Bench: A Benchmark for Multi-Modal Lifelong Navigation
por: Khanna, Mukul, et al.
Publicado: (2024)
por: Khanna, Mukul, et al.
Publicado: (2024)
Residual Cross-Modal Fusion Networks for Audio-Visual Navigation
por: Wang, Yi, et al.
Publicado: (2026)
por: Wang, Yi, et al.
Publicado: (2026)
Cross-Modal Instructions for Robot Motion Generation
por: Barron, William, et al.
Publicado: (2025)
por: Barron, William, et al.
Publicado: (2025)
REALM: An RGB and Event Aligned Latent Manifold for Cross-Modal Perception
por: Polizzi, Vincenzo, et al.
Publicado: (2026)
por: Polizzi, Vincenzo, et al.
Publicado: (2026)
Generating Physically Realistic and Directable Human Motions from Multi-Modal Inputs
por: Shrestha, Aayam, et al.
Publicado: (2025)
por: Shrestha, Aayam, et al.
Publicado: (2025)
In-Context Policy Adaptation via Cross-Domain Skill Diffusion
por: Yoo, Minjong, et al.
Publicado: (2025)
por: Yoo, Minjong, et al.
Publicado: (2025)
Learning Multi-Modal Whole-Body Control for Real-World Humanoid Robots
por: Dugar, Pranay, et al.
Publicado: (2024)
por: Dugar, Pranay, et al.
Publicado: (2024)
Online,Target-Free LiDAR-Camera Extrinsic Calibration via Cross-Modal Mask Matching
por: Huang, Zhiwei, et al.
Publicado: (2024)
por: Huang, Zhiwei, et al.
Publicado: (2024)
Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning
por: NVIDIA, et al.
Publicado: (2025)
por: NVIDIA, et al.
Publicado: (2025)
Ejemplares similares
-
Tactile Estimation of Extrinsic Contact Patch for Stable Placement
por: Ota, Kei, et al.
Publicado: (2023) -
Zero-Shot Peg Insertion: Identifying Mating Holes and Estimating SE(2) Poses with Vision-Language Models
por: Yajima, Masaru, et al.
Publicado: (2025) -
Learning Pivoting Manipulation with Force and Vision Feedback Using Optimization-based Demonstrations
por: Shirai, Yuki, et al.
Publicado: (2025) -
Hierarchical Contact-Rich Trajectory Optimization for Multi-Modal Manipulation using Tight Convex Relaxations
por: Shirai, Yuki, et al.
Publicado: (2025) -
Leveraging Large Language Model-based Room-Object Relationships Knowledge for Enhancing Multimodal-Input Object Goal Navigation
por: Sun, Leyuan, et al.
Publicado: (2024)