REALM: An RGB and Event Aligned Latent Manifold for Cross-Modal Perception
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Polizzi, Vincenzo, Lindell, David B., Kelly, Jonathan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VibES: Induced Vibration for Persistent Event-Based Sensing
von: Polizzi, Vincenzo, et al.
Veröffentlicht: (2025)
von: Polizzi, Vincenzo, et al.
Veröffentlicht: (2025)
DualCross: Cross-Modality Cross-Domain Adaptation for Monocular BEV Perception
von: Man, Yunze, et al.
Veröffentlicht: (2023)
von: Man, Yunze, et al.
Veröffentlicht: (2023)
FaVoR: Features via Voxel Rendering for Camera Relocalization
von: Polizzi, Vincenzo, et al.
Veröffentlicht: (2024)
von: Polizzi, Vincenzo, et al.
Veröffentlicht: (2024)
RoboFlamingo-Plus: Fusion of Depth and RGB Perception with Vision-Language Models for Enhanced Robotic Manipulation
von: Wang, Sheng
Veröffentlicht: (2025)
von: Wang, Sheng
Veröffentlicht: (2025)
Toward Aligning Human and Robot Actions via Multi-Modal Demonstration Learning
von: Zahid, Azizul, et al.
Veröffentlicht: (2025)
von: Zahid, Azizul, et al.
Veröffentlicht: (2025)
CoCMT: Communication-Efficient Cross-Modal Transformer for Collaborative Perception
von: Wang, Rujia, et al.
Veröffentlicht: (2025)
von: Wang, Rujia, et al.
Veröffentlicht: (2025)
Cycle-Correspondence Loss: Learning Dense View-Invariant Visual Features from Unlabeled and Unordered RGB Images
von: Adrian, David B., et al.
Veröffentlicht: (2024)
von: Adrian, David B., et al.
Veröffentlicht: (2024)
Residual Cross-Modal Fusion Networks for Audio-Visual Navigation
von: Wang, Yi, et al.
Veröffentlicht: (2026)
von: Wang, Yi, et al.
Veröffentlicht: (2026)
Flat'n'Fold: A Diverse Multi-Modal Dataset for Garment Perception and Manipulation
von: Zhuang, Lipeng, et al.
Veröffentlicht: (2024)
von: Zhuang, Lipeng, et al.
Veröffentlicht: (2024)
Learning Shared RGB-D Fields: Unified Self-supervised Pre-training for Label-efficient LiDAR-Camera 3D Perception
von: Xu, Xiaohao, et al.
Veröffentlicht: (2024)
von: Xu, Xiaohao, et al.
Veröffentlicht: (2024)
Single-Frame Point-Pixel Registration via Supervised Cross-Modal Feature Matching
von: Han, Yu, et al.
Veröffentlicht: (2025)
von: Han, Yu, et al.
Veröffentlicht: (2025)
Complementary Random Masking for RGB-Thermal Semantic Segmentation
von: Shin, Ukcheol, et al.
Veröffentlicht: (2023)
von: Shin, Ukcheol, et al.
Veröffentlicht: (2023)
Online,Target-Free LiDAR-Camera Extrinsic Calibration via Cross-Modal Mask Matching
von: Huang, Zhiwei, et al.
Veröffentlicht: (2024)
von: Huang, Zhiwei, et al.
Veröffentlicht: (2024)
PhotoBot: Reference-Guided Interactive Photography via Natural Language
von: Limoyo, Oliver, et al.
Veröffentlicht: (2024)
von: Limoyo, Oliver, et al.
Veröffentlicht: (2024)
RGB-only Active 3D Scene Graph Generation for Indoor Mobile Robots
von: Modi, Giorgia, et al.
Veröffentlicht: (2026)
von: Modi, Giorgia, et al.
Veröffentlicht: (2026)
SplaTAM: Splat, Track & Map 3D Gaussians for Dense RGB-D SLAM
von: Keetha, Nikhil, et al.
Veröffentlicht: (2023)
von: Keetha, Nikhil, et al.
Veröffentlicht: (2023)
Cross-Modal Instructions for Robot Motion Generation
von: Barron, William, et al.
Veröffentlicht: (2025)
von: Barron, William, et al.
Veröffentlicht: (2025)
LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning
von: Hao, Haihong, et al.
Veröffentlicht: (2026)
von: Hao, Haihong, et al.
Veröffentlicht: (2026)
DepthVision: Enabling Robust Vision-Language Models with GAN-Based LiDAR-to-RGB Synthesis for Autonomous Driving
von: Kirchner, Sven, et al.
Veröffentlicht: (2025)
von: Kirchner, Sven, et al.
Veröffentlicht: (2025)
V3D-SLAM: Robust RGB-D SLAM in Dynamic Environments with 3D Semantic Geometry Voting
von: Dang, Tuan, et al.
Veröffentlicht: (2024)
von: Dang, Tuan, et al.
Veröffentlicht: (2024)
DiLA: Disentangled Latent Action World Models
von: Zhang, Tianqiu, et al.
Veröffentlicht: (2026)
von: Zhang, Tianqiu, et al.
Veröffentlicht: (2026)
Chain of World: World Model Thinking in Latent Motion
von: Yang, Fuxiang, et al.
Veröffentlicht: (2026)
von: Yang, Fuxiang, et al.
Veröffentlicht: (2026)
Towards Closing the Domain Gap with Event Cameras
von: Sevinc, M. Oltan, et al.
Veröffentlicht: (2025)
von: Sevinc, M. Oltan, et al.
Veröffentlicht: (2025)
IMPACT-HOI: Supervisory Control for Onset-Anchored Partial HOI Event Construction
von: Zhang, Haoshen, et al.
Veröffentlicht: (2026)
von: Zhang, Haoshen, et al.
Veröffentlicht: (2026)
Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models
von: Lei, Zixing, et al.
Veröffentlicht: (2026)
von: Lei, Zixing, et al.
Veröffentlicht: (2026)
SIM1: Physics-Aligned Simulator as Zero-Shot Data Scaler in Deformable Worlds
von: Zhou, Yunsong, et al.
Veröffentlicht: (2026)
von: Zhou, Yunsong, et al.
Veröffentlicht: (2026)
Latent Gaussian Splatting for 4D Panoptic Occupancy Tracking
von: Luz, Maximilian, et al.
Veröffentlicht: (2026)
von: Luz, Maximilian, et al.
Veröffentlicht: (2026)
Conditioning Latent-Space Clusters for Real-World Anomaly Classification
von: Bogdoll, Daniel, et al.
Veröffentlicht: (2023)
von: Bogdoll, Daniel, et al.
Veröffentlicht: (2023)
Safety-Aligned 3D Object Detection: Single-Vehicle, Cooperative, and End-to-End Perspectives
von: Liao, Brian Hsuan-Cheng, et al.
Veröffentlicht: (2026)
von: Liao, Brian Hsuan-Cheng, et al.
Veröffentlicht: (2026)
CleverDistiller: Simple and Spatially Consistent Cross-modal Distillation
von: Govindarajan, Hariprasath, et al.
Veröffentlicht: (2025)
von: Govindarajan, Hariprasath, et al.
Veröffentlicht: (2025)
RealAppliance: Let High-fidelity Appliance Assets Controllable and Workable as Aligned Real Manuals
von: Gao, Yuzheng, et al.
Veröffentlicht: (2025)
von: Gao, Yuzheng, et al.
Veröffentlicht: (2025)
Manipulation as in Simulation: Enabling Accurate Geometry Perception in Robots
von: Liu, Minghuan, et al.
Veröffentlicht: (2025)
von: Liu, Minghuan, et al.
Veröffentlicht: (2025)
STAMP: Scalable Task And Model-agnostic Collaborative Perception
von: Gao, Xiangbo, et al.
Veröffentlicht: (2025)
von: Gao, Xiangbo, et al.
Veröffentlicht: (2025)
Spatially Visual Perception for End-to-End Robotic Learning
von: Davies, Travis, et al.
Veröffentlicht: (2024)
von: Davies, Travis, et al.
Veröffentlicht: (2024)
RPMArt: Towards Robust Perception and Manipulation for Articulated Objects
von: Wang, Junbo, et al.
Veröffentlicht: (2024)
von: Wang, Junbo, et al.
Veröffentlicht: (2024)
Efficient Robotic Policy Learning via Latent Space Backward Planning
von: Liu, Dongxiu, et al.
Veröffentlicht: (2025)
von: Liu, Dongxiu, et al.
Veröffentlicht: (2025)
DriveCritic: Towards Context-Aware, Human-Aligned Evaluation for Autonomous Driving with Vision-Language Models
von: Song, Jingyu, et al.
Veröffentlicht: (2025)
von: Song, Jingyu, et al.
Veröffentlicht: (2025)
On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations
von: Guo, Jianing, et al.
Veröffentlicht: (2025)
von: Guo, Jianing, et al.
Veröffentlicht: (2025)
Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model
von: Kim, Dongwon, et al.
Veröffentlicht: (2026)
von: Kim, Dongwon, et al.
Veröffentlicht: (2026)
ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models
von: Tang, Zuojin, et al.
Veröffentlicht: (2026)
von: Tang, Zuojin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
VibES: Induced Vibration for Persistent Event-Based Sensing
von: Polizzi, Vincenzo, et al.
Veröffentlicht: (2025) -
DualCross: Cross-Modality Cross-Domain Adaptation for Monocular BEV Perception
von: Man, Yunze, et al.
Veröffentlicht: (2023) -
FaVoR: Features via Voxel Rendering for Camera Relocalization
von: Polizzi, Vincenzo, et al.
Veröffentlicht: (2024) -
RoboFlamingo-Plus: Fusion of Depth and RGB Perception with Vision-Language Models for Enhanced Robotic Manipulation
von: Wang, Sheng
Veröffentlicht: (2025) -
Toward Aligning Human and Robot Actions via Multi-Modal Demonstration Learning
von: Zahid, Azizul, et al.
Veröffentlicht: (2025)