REALM: An RGB and Event Aligned Latent Manifold for Cross-Modal Perception
Fuente:
arXiv
Saved in:
| Main Authors: | Polizzi, Vincenzo, Lindell, David B., Kelly, Jonathan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VibES: Induced Vibration for Persistent Event-Based Sensing
by: Polizzi, Vincenzo, et al.
Published: (2025)
by: Polizzi, Vincenzo, et al.
Published: (2025)
DualCross: Cross-Modality Cross-Domain Adaptation for Monocular BEV Perception
by: Man, Yunze, et al.
Published: (2023)
by: Man, Yunze, et al.
Published: (2023)
FaVoR: Features via Voxel Rendering for Camera Relocalization
by: Polizzi, Vincenzo, et al.
Published: (2024)
by: Polizzi, Vincenzo, et al.
Published: (2024)
RoboFlamingo-Plus: Fusion of Depth and RGB Perception with Vision-Language Models for Enhanced Robotic Manipulation
by: Wang, Sheng
Published: (2025)
by: Wang, Sheng
Published: (2025)
Toward Aligning Human and Robot Actions via Multi-Modal Demonstration Learning
by: Zahid, Azizul, et al.
Published: (2025)
by: Zahid, Azizul, et al.
Published: (2025)
CoCMT: Communication-Efficient Cross-Modal Transformer for Collaborative Perception
by: Wang, Rujia, et al.
Published: (2025)
by: Wang, Rujia, et al.
Published: (2025)
Cycle-Correspondence Loss: Learning Dense View-Invariant Visual Features from Unlabeled and Unordered RGB Images
by: Adrian, David B., et al.
Published: (2024)
by: Adrian, David B., et al.
Published: (2024)
Residual Cross-Modal Fusion Networks for Audio-Visual Navigation
by: Wang, Yi, et al.
Published: (2026)
by: Wang, Yi, et al.
Published: (2026)
Flat'n'Fold: A Diverse Multi-Modal Dataset for Garment Perception and Manipulation
by: Zhuang, Lipeng, et al.
Published: (2024)
by: Zhuang, Lipeng, et al.
Published: (2024)
Learning Shared RGB-D Fields: Unified Self-supervised Pre-training for Label-efficient LiDAR-Camera 3D Perception
by: Xu, Xiaohao, et al.
Published: (2024)
by: Xu, Xiaohao, et al.
Published: (2024)
Single-Frame Point-Pixel Registration via Supervised Cross-Modal Feature Matching
by: Han, Yu, et al.
Published: (2025)
by: Han, Yu, et al.
Published: (2025)
Complementary Random Masking for RGB-Thermal Semantic Segmentation
by: Shin, Ukcheol, et al.
Published: (2023)
by: Shin, Ukcheol, et al.
Published: (2023)
Online,Target-Free LiDAR-Camera Extrinsic Calibration via Cross-Modal Mask Matching
by: Huang, Zhiwei, et al.
Published: (2024)
by: Huang, Zhiwei, et al.
Published: (2024)
PhotoBot: Reference-Guided Interactive Photography via Natural Language
by: Limoyo, Oliver, et al.
Published: (2024)
by: Limoyo, Oliver, et al.
Published: (2024)
RGB-only Active 3D Scene Graph Generation for Indoor Mobile Robots
by: Modi, Giorgia, et al.
Published: (2026)
by: Modi, Giorgia, et al.
Published: (2026)
SplaTAM: Splat, Track & Map 3D Gaussians for Dense RGB-D SLAM
by: Keetha, Nikhil, et al.
Published: (2023)
by: Keetha, Nikhil, et al.
Published: (2023)
Cross-Modal Instructions for Robot Motion Generation
by: Barron, William, et al.
Published: (2025)
by: Barron, William, et al.
Published: (2025)
LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning
by: Hao, Haihong, et al.
Published: (2026)
by: Hao, Haihong, et al.
Published: (2026)
DepthVision: Enabling Robust Vision-Language Models with GAN-Based LiDAR-to-RGB Synthesis for Autonomous Driving
by: Kirchner, Sven, et al.
Published: (2025)
by: Kirchner, Sven, et al.
Published: (2025)
V3D-SLAM: Robust RGB-D SLAM in Dynamic Environments with 3D Semantic Geometry Voting
by: Dang, Tuan, et al.
Published: (2024)
by: Dang, Tuan, et al.
Published: (2024)
DiLA: Disentangled Latent Action World Models
by: Zhang, Tianqiu, et al.
Published: (2026)
by: Zhang, Tianqiu, et al.
Published: (2026)
Chain of World: World Model Thinking in Latent Motion
by: Yang, Fuxiang, et al.
Published: (2026)
by: Yang, Fuxiang, et al.
Published: (2026)
Towards Closing the Domain Gap with Event Cameras
by: Sevinc, M. Oltan, et al.
Published: (2025)
by: Sevinc, M. Oltan, et al.
Published: (2025)
IMPACT-HOI: Supervisory Control for Onset-Anchored Partial HOI Event Construction
by: Zhang, Haoshen, et al.
Published: (2026)
by: Zhang, Haoshen, et al.
Published: (2026)
Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models
by: Lei, Zixing, et al.
Published: (2026)
by: Lei, Zixing, et al.
Published: (2026)
SIM1: Physics-Aligned Simulator as Zero-Shot Data Scaler in Deformable Worlds
by: Zhou, Yunsong, et al.
Published: (2026)
by: Zhou, Yunsong, et al.
Published: (2026)
Latent Gaussian Splatting for 4D Panoptic Occupancy Tracking
by: Luz, Maximilian, et al.
Published: (2026)
by: Luz, Maximilian, et al.
Published: (2026)
Conditioning Latent-Space Clusters for Real-World Anomaly Classification
by: Bogdoll, Daniel, et al.
Published: (2023)
by: Bogdoll, Daniel, et al.
Published: (2023)
Safety-Aligned 3D Object Detection: Single-Vehicle, Cooperative, and End-to-End Perspectives
by: Liao, Brian Hsuan-Cheng, et al.
Published: (2026)
by: Liao, Brian Hsuan-Cheng, et al.
Published: (2026)
CleverDistiller: Simple and Spatially Consistent Cross-modal Distillation
by: Govindarajan, Hariprasath, et al.
Published: (2025)
by: Govindarajan, Hariprasath, et al.
Published: (2025)
RealAppliance: Let High-fidelity Appliance Assets Controllable and Workable as Aligned Real Manuals
by: Gao, Yuzheng, et al.
Published: (2025)
by: Gao, Yuzheng, et al.
Published: (2025)
Manipulation as in Simulation: Enabling Accurate Geometry Perception in Robots
by: Liu, Minghuan, et al.
Published: (2025)
by: Liu, Minghuan, et al.
Published: (2025)
STAMP: Scalable Task And Model-agnostic Collaborative Perception
by: Gao, Xiangbo, et al.
Published: (2025)
by: Gao, Xiangbo, et al.
Published: (2025)
Spatially Visual Perception for End-to-End Robotic Learning
by: Davies, Travis, et al.
Published: (2024)
by: Davies, Travis, et al.
Published: (2024)
RPMArt: Towards Robust Perception and Manipulation for Articulated Objects
by: Wang, Junbo, et al.
Published: (2024)
by: Wang, Junbo, et al.
Published: (2024)
Efficient Robotic Policy Learning via Latent Space Backward Planning
by: Liu, Dongxiu, et al.
Published: (2025)
by: Liu, Dongxiu, et al.
Published: (2025)
DriveCritic: Towards Context-Aware, Human-Aligned Evaluation for Autonomous Driving with Vision-Language Models
by: Song, Jingyu, et al.
Published: (2025)
by: Song, Jingyu, et al.
Published: (2025)
On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations
by: Guo, Jianing, et al.
Published: (2025)
by: Guo, Jianing, et al.
Published: (2025)
Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model
by: Kim, Dongwon, et al.
Published: (2026)
by: Kim, Dongwon, et al.
Published: (2026)
ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models
by: Tang, Zuojin, et al.
Published: (2026)
by: Tang, Zuojin, et al.
Published: (2026)
Similar Items
-
VibES: Induced Vibration for Persistent Event-Based Sensing
by: Polizzi, Vincenzo, et al.
Published: (2025) -
DualCross: Cross-Modality Cross-Domain Adaptation for Monocular BEV Perception
by: Man, Yunze, et al.
Published: (2023) -
FaVoR: Features via Voxel Rendering for Camera Relocalization
by: Polizzi, Vincenzo, et al.
Published: (2024) -
RoboFlamingo-Plus: Fusion of Depth and RGB Perception with Vision-Language Models for Enhanced Robotic Manipulation
by: Wang, Sheng
Published: (2025) -
Toward Aligning Human and Robot Actions via Multi-Modal Demonstration Learning
by: Zahid, Azizul, et al.
Published: (2025)