Advances in Multimodal Adaptation and Generalization: From Traditional Approaches to Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Dong, Hao, Liu, Moru, Zhou, Kaiyang, Chatzi, Eleni, Kannala, Juho, Stachniss, Cyrill, Fink, Olga |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Multimodal Open-Set Domain Generalization and Adaptation through Self-supervision
by: Dong, Hao, et al.
Published: (2024)
by: Dong, Hao, et al.
Published: (2024)
To Trust Or Not To Trust Your Vision-Language Model's Prediction
by: Dong, Hao, et al.
Published: (2025)
by: Dong, Hao, et al.
Published: (2025)
Towards Robust Multimodal Open-set Test-time Adaptation via Adaptive Entropy-aware Optimization
by: Dong, Hao, et al.
Published: (2025)
by: Dong, Hao, et al.
Published: (2025)
Extremely Simple Multimodal Outlier Synthesis for Out-of-Distribution Detection and Segmentation
by: Liu, Moru, et al.
Published: (2025)
by: Liu, Moru, et al.
Published: (2025)
Multimodal Learning for Arcing Detection in Pantograph-Catenary Systems
by: Dong, Hao, et al.
Published: (2026)
by: Dong, Hao, et al.
Published: (2026)
Dynamic Visual SLAM using a General 3D Prior
by: Zhong, Xingguang, et al.
Published: (2025)
by: Zhong, Xingguang, et al.
Published: (2025)
LiPS: Lightweight Panoptic Segmentation for Resource-Constrained Robotics
by: Galagain, Calvin, et al.
Published: (2026)
by: Galagain, Calvin, et al.
Published: (2026)
Epipolar Attention Field Transformers for Bird's Eye View Semantic Segmentation
by: Witte, Christian, et al.
Published: (2024)
by: Witte, Christian, et al.
Published: (2024)
3D LiDAR Mapping in Dynamic Environments Using a 4D Implicit Neural Representation
by: Zhong, Xingguang, et al.
Published: (2024)
by: Zhong, Xingguang, et al.
Published: (2024)
Open-World Panoptic Segmentation
by: Sodano, Matteo, et al.
Published: (2024)
by: Sodano, Matteo, et al.
Published: (2024)
AdaCropFollow: Self-Supervised Online Adaptation for Visual Under-Canopy Navigation
by: Sivakumar, Arun N., et al.
Published: (2024)
by: Sivakumar, Arun N., et al.
Published: (2024)
Adaptive Confidence Regularization for Multimodal Failure Detection
by: Liu, Moru, et al.
Published: (2026)
by: Liu, Moru, et al.
Published: (2026)
HeLiMOS: A Dataset for Moving Object Segmentation in 3D Point Clouds From Heterogeneous LiDAR Sensors
by: Lim, Hyungtae, et al.
Published: (2024)
by: Lim, Hyungtae, et al.
Published: (2024)
Horticultural Temporal Fruit Monitoring via 3D Instance Segmentation and Re-Identification using Colored Point Clouds
by: Fusaro, Daniel, et al.
Published: (2024)
by: Fusaro, Daniel, et al.
Published: (2024)
STAIR: Semantic-Targeted Active Implicit Reconstruction
by: Jin, Liren, et al.
Published: (2024)
by: Jin, Liren, et al.
Published: (2024)
Exploiting Priors from 3D Diffusion Models for RGB-Based One-Shot View Planning
by: Pan, Sicong, et al.
Published: (2024)
by: Pan, Sicong, et al.
Published: (2024)
SuperFusion: Multilevel LiDAR-Camera Fusion for Long-Range HD Map Generation
by: Dong, Hao, et al.
Published: (2022)
by: Dong, Hao, et al.
Published: (2022)
DM-OSVP++: One-Shot View Planning Using 3D Diffusion Models for Active RGB-Based Object Reconstruction
by: Pan, Sicong, et al.
Published: (2025)
by: Pan, Sicong, et al.
Published: (2025)
MultiOOD: Scaling Out-of-Distribution Detection for Multiple Modalities
by: Dong, Hao, et al.
Published: (2024)
by: Dong, Hao, et al.
Published: (2024)
NNG-Mix: Improving Semi-supervised Anomaly Detection with Pseudo-anomaly Generation
by: Dong, Hao, et al.
Published: (2023)
by: Dong, Hao, et al.
Published: (2023)
BonnBeetClouds3D: A Dataset Towards Point Cloud-based Organ-level Phenotyping of Sugar Beet Plants under Field Conditions
by: Marks, Elias, et al.
Published: (2023)
by: Marks, Elias, et al.
Published: (2023)
ActiveGS: Active Scene Reconstruction Using Gaussian Splatting
by: Jin, Liren, et al.
Published: (2024)
by: Jin, Liren, et al.
Published: (2024)
PIN-SLAM: LiDAR SLAM Using a Point-Based Implicit Neural Representation for Achieving Global Map Consistency
by: Pan, Yue, et al.
Published: (2024)
by: Pan, Yue, et al.
Published: (2024)
Loop Closure via Maximal Cliques in 3D LiDAR-Based SLAM
by: Laserna, Javier, et al.
Published: (2026)
by: Laserna, Javier, et al.
Published: (2026)
Are We Making Progress in Multimodal Domain Generalization? A Comprehensive Benchmark Study
by: Dong, Hao, et al.
Published: (2026)
by: Dong, Hao, et al.
Published: (2026)
Keypoint Semantic Integration for Improved Feature Matching in Outdoor Agricultural Environments
by: de Silva, Rajitha, et al.
Published: (2025)
by: de Silva, Rajitha, et al.
Published: (2025)
Register Any Point: Scaling 3D Point Cloud Registration by Flow Matching
by: Pan, Yue, et al.
Published: (2025)
by: Pan, Yue, et al.
Published: (2025)
PINGS: Gaussian Splatting Meets Distance Fields within a Point-Based Implicit Neural Map
by: Pan, Yue, et al.
Published: (2025)
by: Pan, Yue, et al.
Published: (2025)
Adapting Vision-Language Models Without Labels: A Comprehensive Survey
by: Dong, Hao, et al.
Published: (2025)
by: Dong, Hao, et al.
Published: (2025)
PhenoBench -- A Large Dataset and Benchmarks for Semantic Image Interpretation in the Agricultural Domain
by: Weyler, Jan, et al.
Published: (2023)
by: Weyler, Jan, et al.
Published: (2023)
3D Hierarchical Panoptic Segmentation in Real Orchard Environments Across Different Sensors
by: Sodano, Matteo, et al.
Published: (2025)
by: Sodano, Matteo, et al.
Published: (2025)
Vector-Quantized Vision Foundation Models for Object-Centric Learning
by: Zhao, Rongzhen, et al.
Published: (2025)
by: Zhao, Rongzhen, et al.
Published: (2025)
Zero-shot Object-Centric Instruction Following: Integrating Foundation Models with Traditional Navigation
by: Raychaudhuri, Sonia, et al.
Published: (2024)
by: Raychaudhuri, Sonia, et al.
Published: (2024)
A Dataset and Benchmark for Shape Completion of Fruits for Agricultural Robotics
by: Magistri, Federico, et al.
Published: (2024)
by: Magistri, Federico, et al.
Published: (2024)
From Seeing to Experiencing: Scaling Navigation Foundation Models with Reinforcement Learning
by: He, Honglin, et al.
Published: (2025)
by: He, Honglin, et al.
Published: (2025)
Foundation Feature-Driven Online End-Effector Pose Estimation: A Marker-Free and Learning-Free Approach
by: Wu, Tianshu, et al.
Published: (2025)
by: Wu, Tianshu, et al.
Published: (2025)
Towards Generating Realistic 3D Semantic Training Data for Autonomous Driving
by: Nunes, Lucas, et al.
Published: (2025)
by: Nunes, Lucas, et al.
Published: (2025)
Leveraging GNSS and Onboard Visual Data from Consumer Vehicles for Robust Road Network Estimation
by: Opra, Balázs, et al.
Published: (2024)
by: Opra, Balázs, et al.
Published: (2024)
RepSAM: Bridging Foundation Models to Robotic Vision via Representation-Guided Adaptation
by: Chu, Wenhui
Published: (2026)
by: Chu, Wenhui
Published: (2026)
What Is The Best 3D Scene Representation for Robotics? From Geometric to Foundation Models
by: Deng, Tianchen, et al.
Published: (2025)
by: Deng, Tianchen, et al.
Published: (2025)
Similar Items
-
Towards Multimodal Open-Set Domain Generalization and Adaptation through Self-supervision
by: Dong, Hao, et al.
Published: (2024) -
To Trust Or Not To Trust Your Vision-Language Model's Prediction
by: Dong, Hao, et al.
Published: (2025) -
Towards Robust Multimodal Open-set Test-time Adaptation via Adaptive Entropy-aware Optimization
by: Dong, Hao, et al.
Published: (2025) -
Extremely Simple Multimodal Outlier Synthesis for Out-of-Distribution Detection and Segmentation
by: Liu, Moru, et al.
Published: (2025) -
Multimodal Learning for Arcing Detection in Pantograph-Catenary Systems
by: Dong, Hao, et al.
Published: (2026)