Geometry Meets Vision: Revisiting Pretrained Semantics in Distilled Fields
Fuente:
arXiv
Saved in:
| Main Authors: | Mei, Zhiting, Shorinwa, Ola, Majumdar, Anirudha |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
World Models That Know When They Don't Know - Controllable Video Generation with Calibrated Uncertainty
by: Mei, Zhiting, et al.
Published: (2025)
by: Mei, Zhiting, et al.
Published: (2025)
How Confident are Video Models? Empowering Video Models to Express their Uncertainty
by: Mei, Zhiting, et al.
Published: (2025)
by: Mei, Zhiting, et al.
Published: (2025)
SIREN: Semantic, Initialization-Free Registration of Multi-Robot Gaussian Splatting Maps
by: Shorinwa, Ola, et al.
Published: (2025)
by: Shorinwa, Ola, et al.
Published: (2025)
WoMAP: World Models For Embodied Open-Vocabulary Object Localization
by: Yin, Tenny, et al.
Published: (2025)
by: Yin, Tenny, et al.
Published: (2025)
RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation
by: Li, Huiqiong, et al.
Published: (2026)
by: Li, Huiqiong, et al.
Published: (2026)
VERDI: VLM-Embedded Reasoning for Autonomous Driving
by: Feng, Bowen, et al.
Published: (2025)
by: Feng, Bowen, et al.
Published: (2025)
FAST-Splat: Fast, Ambiguity-Free Semantics Transfer in Gaussian Splatting
by: Shorinwa, Ola, et al.
Published: (2024)
by: Shorinwa, Ola, et al.
Published: (2024)
Splat-MOVER: Multi-Stage, Open-Vocabulary Robotic Manipulation via Editable Gaussian Splatting
by: Shorinwa, Ola, et al.
Published: (2024)
by: Shorinwa, Ola, et al.
Published: (2024)
Generating Robot Constitutions & Benchmarks for Semantic Safety
by: Sermanet, Pierre, et al.
Published: (2025)
by: Sermanet, Pierre, et al.
Published: (2025)
Physically Grounded Vision-Language Models for Robotic Manipulation
by: Gao, Jensen, et al.
Published: (2023)
by: Gao, Jensen, et al.
Published: (2023)
VLS: Steering Pretrained Robot Policies via Vision-Language Models
by: Liu, Shuo, et al.
Published: (2026)
by: Liu, Shuo, et al.
Published: (2026)
ViT-VS: On the Applicability of Pretrained Vision Transformer Features for Generalizable Visual Servoing
by: Scherl, Alessandro, et al.
Published: (2025)
by: Scherl, Alessandro, et al.
Published: (2025)
Look to Locate: Vision-Based Multisensory Navigation with 3-D Digital Maps for GNSS-Challenged Environments
by: Elmaghraby, Ola, et al.
Published: (2025)
by: Elmaghraby, Ola, et al.
Published: (2025)
Recasting Generic Pretrained Vision Transformers As Object-Centric Scene Encoders For Manipulation Policies
by: Qian, Jianing, et al.
Published: (2024)
by: Qian, Jianing, et al.
Published: (2024)
ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models
by: Ye, Wencheng, et al.
Published: (2025)
by: Ye, Wencheng, et al.
Published: (2025)
CLIP-Fields: Weakly Supervised Semantic Fields for Robotic Memory
by: Shafiullah, Nur Muhammad Mahi, et al.
Published: (2022)
by: Shafiullah, Nur Muhammad Mahi, et al.
Published: (2022)
LIDEA: Human-to-Robot Imitation Learning via Implicit Feature Distillation and Explicit Geometry Alignment
by: Xu, Yifu, et al.
Published: (2026)
by: Xu, Yifu, et al.
Published: (2026)
A Hybrid Deep Learning Framework for Emotion Recognition in Children with Autism During NAO Robot-Mediated Interaction
by: Bhattacharjee, Indranil, et al.
Published: (2025)
by: Bhattacharjee, Indranil, et al.
Published: (2025)
Domain Adaptation-Based Crossmodal Knowledge Distillation for 3D Semantic Segmentation
by: Kang, Jialiang, et al.
Published: (2025)
by: Kang, Jialiang, et al.
Published: (2025)
S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight
by: Yan, Haodong, et al.
Published: (2026)
by: Yan, Haodong, et al.
Published: (2026)
Not All Voxels Are Equal: Hardness-Aware Semantic Scene Completion with Self-Distillation
by: Wang, Song, et al.
Published: (2024)
by: Wang, Song, et al.
Published: (2024)
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
by: Zhang, Chubin, et al.
Published: (2026)
by: Zhang, Chubin, et al.
Published: (2026)
VG3S: Visual Geometry Grounded Gaussian Splatting for Semantic Occupancy Prediction
by: Yan, Xiaoyang, et al.
Published: (2026)
by: Yan, Xiaoyang, et al.
Published: (2026)
A Data-Centric Revisit of Pre-Trained Vision Models for Robot Learning
by: Wen, Xin, et al.
Published: (2025)
by: Wen, Xin, et al.
Published: (2025)
Implicit Geometry Representations for Vision-and-Language Navigation from Web Videos
by: Han, Mingfei, et al.
Published: (2026)
by: Han, Mingfei, et al.
Published: (2026)
Semantics from Space: Satellite-Guided Thermal Semantic Segmentation Annotation for Aerial Field Robots
by: Lee, Connor, et al.
Published: (2024)
by: Lee, Connor, et al.
Published: (2024)
FD-VLA: Force-Distilled Vision-Language-Action Model for Contact-Rich Manipulation
by: Zhao, Ruiteng, et al.
Published: (2026)
by: Zhao, Ruiteng, et al.
Published: (2026)
Vision-Only Gaussian Splatting for Collaborative Semantic Occupancy Prediction
by: Chen, Cheng, et al.
Published: (2025)
by: Chen, Cheng, et al.
Published: (2025)
Geometry-aided Vision-based Localization of Future Mars Helicopters in Challenging Illumination Conditions
by: Pisanti, Dario, et al.
Published: (2025)
by: Pisanti, Dario, et al.
Published: (2025)
Explore until Confident: Efficient Exploration for Embodied Question Answering
by: Ren, Allen Z., et al.
Published: (2024)
by: Ren, Allen Z., et al.
Published: (2024)
Vision based Crop Row Navigation under Varying Field Conditions in Arable Fields
by: de Silva, Rajitha, et al.
Published: (2022)
by: de Silva, Rajitha, et al.
Published: (2022)
Feasibility of Indoor Frame-Wise Lidar Semantic Segmentation via Distillation from Visual Foundation Model
by: Wu, Haiyang, et al.
Published: (2026)
by: Wu, Haiyang, et al.
Published: (2026)
A Vision-Based Navigation System for Arable Fields
by: de Silva, Rajitha, et al.
Published: (2023)
by: de Silva, Rajitha, et al.
Published: (2023)
Weather-Robust Scene Semantics with Vision-Aligned 4D Radar
by: Hamilton, Kali, et al.
Published: (2026)
by: Hamilton, Kali, et al.
Published: (2026)
Geometry-Informed Distance Candidate Selection for Adaptive Lightweight Omnidirectional Stereo Vision with Fisheye Images
by: Pulling, Conner, et al.
Published: (2024)
by: Pulling, Conner, et al.
Published: (2024)
Universal Pose Pretraining for Generalizable Vision-Language-Action Policies
by: Lin, Haitao, et al.
Published: (2026)
by: Lin, Haitao, et al.
Published: (2026)
Gaussian Semantic Field for One-shot LiDAR Global Localization
by: Yin, Pengyu, et al.
Published: (2025)
by: Yin, Pengyu, et al.
Published: (2025)
Epipolar Attention Field Transformers for Bird's Eye View Semantic Segmentation
by: Witte, Christian, et al.
Published: (2024)
by: Witte, Christian, et al.
Published: (2024)
MCDS-VSS: Moving Camera Dynamic Scene Video Semantic Segmentation by Filtering with Self-Supervised Geometry and Motion
by: Villar-Corrales, Angel, et al.
Published: (2024)
by: Villar-Corrales, Angel, et al.
Published: (2024)
Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment
by: Lin, Tao, et al.
Published: (2025)
by: Lin, Tao, et al.
Published: (2025)
Similar Items
-
World Models That Know When They Don't Know - Controllable Video Generation with Calibrated Uncertainty
by: Mei, Zhiting, et al.
Published: (2025) -
How Confident are Video Models? Empowering Video Models to Express their Uncertainty
by: Mei, Zhiting, et al.
Published: (2025) -
SIREN: Semantic, Initialization-Free Registration of Multi-Robot Gaussian Splatting Maps
by: Shorinwa, Ola, et al.
Published: (2025) -
WoMAP: World Models For Embodied Open-Vocabulary Object Localization
by: Yin, Tenny, et al.
Published: (2025) -
RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation
by: Li, Huiqiong, et al.
Published: (2026)