ViTaPEs: Visuotactile Position Encodings for Cross-Modal Alignment in Multimodal Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lygerakis, Fotios, Özdenizci, Ozan, Rückert, Elmar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
M2CURL: Sample-Efficient Multimodal Reinforcement Learning via Self-Supervised Representation Learning for Robotic Manipulation
von: Lygerakis, Fotios, et al.
Veröffentlicht: (2024)
von: Lygerakis, Fotios, et al.
Veröffentlicht: (2024)
Learning the RoPEs: Better 2D and 3D Position Encodings with STRING
von: Schenck, Connor, et al.
Veröffentlicht: (2025)
von: Schenck, Connor, et al.
Veröffentlicht: (2025)
Restoring Vision in Adverse Weather Conditions with Patch-Based Denoising Diffusion Models
von: Özdenizci, Ozan, et al.
Veröffentlicht: (2022)
von: Özdenizci, Ozan, et al.
Veröffentlicht: (2022)
ViTa-Zero: Zero-shot Visuotactile Object 6D Pose Estimation
von: Li, Hongyu, et al.
Veröffentlicht: (2025)
von: Li, Hongyu, et al.
Veröffentlicht: (2025)
Multimodal Visual-Tactile Representation Learning through Self-Supervised Contrastive Pre-Training
von: Dave, Vedant, et al.
Veröffentlicht: (2024)
von: Dave, Vedant, et al.
Veröffentlicht: (2024)
Robot Synesthesia: In-Hand Manipulation with Visuotactile Sensing
von: Yuan, Ying, et al.
Veröffentlicht: (2023)
von: Yuan, Ying, et al.
Veröffentlicht: (2023)
Learning Visuotactile Skills with Two Multifingered Hands
von: Lin, Toru, et al.
Veröffentlicht: (2024)
von: Lin, Toru, et al.
Veröffentlicht: (2024)
ViViDex: Learning Vision-based Dexterous Manipulation from Human Videos
von: Chen, Zerui, et al.
Veröffentlicht: (2024)
von: Chen, Zerui, et al.
Veröffentlicht: (2024)
CoCMT: Communication-Efficient Cross-Modal Transformer for Collaborative Perception
von: Wang, Rujia, et al.
Veröffentlicht: (2025)
von: Wang, Rujia, et al.
Veröffentlicht: (2025)
Shelf-Supervised Cross-Modal Pre-Training for 3D Object Detection
von: Khurana, Mehar, et al.
Veröffentlicht: (2024)
von: Khurana, Mehar, et al.
Veröffentlicht: (2024)
Adapting Pretrained ViTs with Convolution Injector for Visuo-Motor Control
von: Hwang, Dongyoon, et al.
Veröffentlicht: (2024)
von: Hwang, Dongyoon, et al.
Veröffentlicht: (2024)
Cross-Modal Instructions for Robot Motion Generation
von: Barron, William, et al.
Veröffentlicht: (2025)
von: Barron, William, et al.
Veröffentlicht: (2025)
Learning Spatial Structure from Pre-Beamforming Per-Antenna Range-Doppler Radar Data via Visibility-Aware Cross-Modal Supervision
von: Sebastian, George, et al.
Veröffentlicht: (2026)
von: Sebastian, George, et al.
Veröffentlicht: (2026)
Point-GN: A Non-Parametric Network Using Gaussian Positional Encoding for Point Cloud Classification
von: Mohammadi, Marzieh, et al.
Veröffentlicht: (2024)
von: Mohammadi, Marzieh, et al.
Veröffentlicht: (2024)
RALACs: Action Recognition in Autonomous Vehicles using Interaction Encoding and Optical Flow
von: Zhou, Eddy, et al.
Veröffentlicht: (2022)
von: Zhou, Eddy, et al.
Veröffentlicht: (2022)
Data-Driven Stochastic Motion Evaluation and Optimization with Image by Spatially-Aligned Temporal Encoding
von: Oba, Takeru, et al.
Veröffentlicht: (2023)
von: Oba, Takeru, et al.
Veröffentlicht: (2023)
Point-LN: A Lightweight Framework for Efficient Point Cloud Classification Using Non-Parametric Positional Encoding
von: Mohammadi, Marzieh, et al.
Veröffentlicht: (2025)
von: Mohammadi, Marzieh, et al.
Veröffentlicht: (2025)
Deep Learning for Inertial Sensor Alignment
von: Freydin, Maxim, et al.
Veröffentlicht: (2022)
von: Freydin, Maxim, et al.
Veröffentlicht: (2022)
Tactile Modality Fusion for Vision-Language-Action Models
von: Morissette, Charlotte, et al.
Veröffentlicht: (2026)
von: Morissette, Charlotte, et al.
Veröffentlicht: (2026)
ED-VAE: Entropy Decomposition of ELBO in Variational Autoencoders
von: Lygerakis, Fotios, et al.
Veröffentlicht: (2024)
von: Lygerakis, Fotios, et al.
Veröffentlicht: (2024)
Scene-Graph ViT: End-to-End Open-Vocabulary Visual Relationship Detection
von: Salzmann, Tim, et al.
Veröffentlicht: (2024)
von: Salzmann, Tim, et al.
Veröffentlicht: (2024)
GRAPE: Generalizing Robot Policy via Preference Alignment
von: Zhang, Zijian, et al.
Veröffentlicht: (2024)
von: Zhang, Zijian, et al.
Veröffentlicht: (2024)
Multi-Space Alignments Towards Universal LiDAR Segmentation
von: Liu, Youquan, et al.
Veröffentlicht: (2024)
von: Liu, Youquan, et al.
Veröffentlicht: (2024)
ROCKET-2: Steering Visuomotor Policy via Cross-View Goal Alignment
von: Cai, Shaofei, et al.
Veröffentlicht: (2025)
von: Cai, Shaofei, et al.
Veröffentlicht: (2025)
Resolving Spatio-Temporal Entanglement in Video Prediction via Multi-Modal Attention
von: Gupta, Shreyam, et al.
Veröffentlicht: (2025)
von: Gupta, Shreyam, et al.
Veröffentlicht: (2025)
Deep Learning-Based Multi-Modal Fusion for Robust Robot Perception and Navigation
von: Lai, Delun, et al.
Veröffentlicht: (2025)
von: Lai, Delun, et al.
Veröffentlicht: (2025)
Multi-Modal Data-Efficient 3D Scene Understanding for Autonomous Driving
von: Kong, Lingdong, et al.
Veröffentlicht: (2024)
von: Kong, Lingdong, et al.
Veröffentlicht: (2024)
DriveTransformer: Unified Transformer for Scalable End-to-End Autonomous Driving
von: Jia, Xiaosong, et al.
Veröffentlicht: (2025)
von: Jia, Xiaosong, et al.
Veröffentlicht: (2025)
Tracking Object Positions in Reinforcement Learning: A Metric for Keypoint Detection (extended version)
von: Cramer, Emma, et al.
Veröffentlicht: (2023)
von: Cramer, Emma, et al.
Veröffentlicht: (2023)
Adver-City: Open-Source Multi-Modal Dataset for Collaborative Perception Under Adverse Weather Conditions
von: Karvat, Mateus, et al.
Veröffentlicht: (2024)
von: Karvat, Mateus, et al.
Veröffentlicht: (2024)
ControlTac: Force- and Position-Controlled Tactile Data Augmentation with a Single Reference Image
von: Luo, Dongyu, et al.
Veröffentlicht: (2025)
von: Luo, Dongyu, et al.
Veröffentlicht: (2025)
MM-ACT: Learn from Multimodal Parallel Generation to Act
von: Liang, Haotian, et al.
Veröffentlicht: (2025)
von: Liang, Haotian, et al.
Veröffentlicht: (2025)
Depth Matters: Multimodal RGB-D Perception for Robust Autonomous Agents
von: Clement, Mihaela-Larisa, et al.
Veröffentlicht: (2025)
von: Clement, Mihaela-Larisa, et al.
Veröffentlicht: (2025)
Taming Transformers for Realistic Lidar Point Cloud Generation
von: Haghighi, Hamed, et al.
Veröffentlicht: (2024)
von: Haghighi, Hamed, et al.
Veröffentlicht: (2024)
Knowledge-aware Graph Transformer for Pedestrian Trajectory Prediction
von: Liu, Yu, et al.
Veröffentlicht: (2024)
von: Liu, Yu, et al.
Veröffentlicht: (2024)
Clebsch-Gordan Transformer: Fast and Global Equivariant Attention
von: Howell, Owen Lewis, et al.
Veröffentlicht: (2025)
von: Howell, Owen Lewis, et al.
Veröffentlicht: (2025)
Efficient Equivariant Transformer for Self-Driving Agent Modeling
von: Xu, Scott, et al.
Veröffentlicht: (2026)
von: Xu, Scott, et al.
Veröffentlicht: (2026)
OpenEMMA: Open-Source Multimodal Model for End-to-End Autonomous Driving
von: Xing, Shuo, et al.
Veröffentlicht: (2024)
von: Xing, Shuo, et al.
Veröffentlicht: (2024)
ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
Uncertainty-Aware Diffusion Model for Multimodal Highway Trajectory Prediction via DDIM Sampling
von: Neumeier, Marion, et al.
Veröffentlicht: (2026)
von: Neumeier, Marion, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
M2CURL: Sample-Efficient Multimodal Reinforcement Learning via Self-Supervised Representation Learning for Robotic Manipulation
von: Lygerakis, Fotios, et al.
Veröffentlicht: (2024) -
Learning the RoPEs: Better 2D and 3D Position Encodings with STRING
von: Schenck, Connor, et al.
Veröffentlicht: (2025) -
Restoring Vision in Adverse Weather Conditions with Patch-Based Denoising Diffusion Models
von: Özdenizci, Ozan, et al.
Veröffentlicht: (2022) -
ViTa-Zero: Zero-shot Visuotactile Object 6D Pose Estimation
von: Li, Hongyu, et al.
Veröffentlicht: (2025) -
Multimodal Visual-Tactile Representation Learning through Self-Supervised Contrastive Pre-Training
von: Dave, Vedant, et al.
Veröffentlicht: (2024)