Lightweight Multimodal Artificial Intelligence Framework for Maritime Multi-Scene Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xi, Xinyu, Yang, Hua, Zhang, Shentai, Liu, Yijie, Sun, Sijin, Fu, Xiuju |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding
von: Li, Haoyuan, et al.
Veröffentlicht: (2025)
von: Li, Haoyuan, et al.
Veröffentlicht: (2025)
Lightweight Structured Multimodal Reasoning for Clinical Scene Understanding in Robotics
von: Jha, Saurav, et al.
Veröffentlicht: (2025)
von: Jha, Saurav, et al.
Veröffentlicht: (2025)
MS-Net: A Multi-Path Sparse Model for Motion Prediction in Multi-Scenes
von: Tang, Xiaqiang, et al.
Veröffentlicht: (2024)
von: Tang, Xiaqiang, et al.
Veröffentlicht: (2024)
Aerial Maritime Vessel Detection and Identification
von: Kulas, Antonella Barisic, et al.
Veröffentlicht: (2025)
von: Kulas, Antonella Barisic, et al.
Veröffentlicht: (2025)
AeroLite-MDNet: Lightweight Multi-task Deviation Detection Network for UAV Landing
von: Yang, Haiping, et al.
Veröffentlicht: (2025)
von: Yang, Haiping, et al.
Veröffentlicht: (2025)
3D Scene Rendering with Multimodal Gaussian Splatting
von: Gau, Chi-Shiang, et al.
Veröffentlicht: (2026)
von: Gau, Chi-Shiang, et al.
Veröffentlicht: (2026)
Going Places: Place Recognition in Artificial and Natural Systems
von: Milford, Michael, et al.
Veröffentlicht: (2025)
von: Milford, Michael, et al.
Veröffentlicht: (2025)
Embodied Spatial Intelligence: from Implicit Scene Modeling to Spatial Reasoning
von: Fang, Jiading
Veröffentlicht: (2025)
von: Fang, Jiading
Veröffentlicht: (2025)
Looking and Listening Inside and Outside: Multimodal Artificial Intelligence Systems for Driver Safety Assessment and Intelligent Vehicle Decision-Making
von: Greer, Ross, et al.
Veröffentlicht: (2026)
von: Greer, Ross, et al.
Veröffentlicht: (2026)
4th Workshop on Maritime Computer Vision (MaCVi): Challenge Overview
von: Kiefer, Benjamin, et al.
Veröffentlicht: (2026)
von: Kiefer, Benjamin, et al.
Veröffentlicht: (2026)
Multi-modal Situated Reasoning in 3D Scenes
von: Linghu, Xiongkun, et al.
Veröffentlicht: (2024)
von: Linghu, Xiongkun, et al.
Veröffentlicht: (2024)
Gradient-Guided Parameter Mask for Multi-Scenario Image Restoration Under Adverse Weather
von: Guo, Jilong, et al.
Veröffentlicht: (2024)
von: Guo, Jilong, et al.
Veröffentlicht: (2024)
Humanoid Occupancy: Enabling A Generalized Multimodal Occupancy Perception System on Humanoid Robots
von: Cui, Wei, et al.
Veröffentlicht: (2025)
von: Cui, Wei, et al.
Veröffentlicht: (2025)
SceneFoundry: Generating Interactive Infinite 3D Worlds
von: Chen, ChunTeng, et al.
Veröffentlicht: (2026)
von: Chen, ChunTeng, et al.
Veröffentlicht: (2026)
MMScan: A Multi-Modal 3D Scene Dataset with Hierarchical Grounded Language Annotations
von: Lyu, Ruiyuan, et al.
Veröffentlicht: (2024)
von: Lyu, Ruiyuan, et al.
Veröffentlicht: (2024)
MOANA: Multi-Radar Dataset for Maritime Odometry and Autonomous Navigation Application
von: Jang, Hyesu, et al.
Veröffentlicht: (2024)
von: Jang, Hyesu, et al.
Veröffentlicht: (2024)
MANSION: Multi-floor lANguage-to-3D Scene generatIOn for loNg-horizon tasks
von: Che, Lirong, et al.
Veröffentlicht: (2026)
von: Che, Lirong, et al.
Veröffentlicht: (2026)
GMT: Goal-Conditioned Multimodal Transformer for 6-DOF Object Trajectory Synthesis in 3D Scenes
von: Zeng, Huajian, et al.
Veröffentlicht: (2026)
von: Zeng, Huajian, et al.
Veröffentlicht: (2026)
Motion Blender Gaussian Splatting for Dynamic Scene Reconstruction
von: Zhang, Xinyu, et al.
Veröffentlicht: (2025)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2025)
GaRField++: Reinforced Gaussian Radiance Fields for Large-Scale 3D Scene Reconstruction
von: Zhang, Hanyue, et al.
Veröffentlicht: (2024)
von: Zhang, Hanyue, et al.
Veröffentlicht: (2024)
Diffusion-guided Generalizable Enhancer for Urban Scene Reconstruction
von: Che, Henry, et al.
Veröffentlicht: (2026)
von: Che, Henry, et al.
Veröffentlicht: (2026)
Scaling Spatial Intelligence with Multimodal Foundation Models
von: Cai, Zhongang, et al.
Veröffentlicht: (2025)
von: Cai, Zhongang, et al.
Veröffentlicht: (2025)
Object Navigation with Structure-Semantic Reasoning-Based Multi-level Map and Multimodal Decision-Making LLM
von: Yan, Chongshang, et al.
Veröffentlicht: (2025)
von: Yan, Chongshang, et al.
Veröffentlicht: (2025)
Advancing Autonomous Vehicle Intelligence: Deep Learning and Multimodal LLM for Traffic Sign Recognition and Robust Lane Detection
von: Sah, Chandan Kumar, et al.
Veröffentlicht: (2025)
von: Sah, Chandan Kumar, et al.
Veröffentlicht: (2025)
Planning with the Views via Scene Self-Exploration
von: Wang, Kangrui, et al.
Veröffentlicht: (2026)
von: Wang, Kangrui, et al.
Veröffentlicht: (2026)
StixelNExT++: Lightweight Monocular Scene Segmentation and Representation for Collective Perception
von: Vosshans, Marcel, et al.
Veröffentlicht: (2025)
von: Vosshans, Marcel, et al.
Veröffentlicht: (2025)
Scene-Agnostic Traversability Labeling and Estimation via a Multimodal Self-supervised Framework
von: Fang, Zipeng, et al.
Veröffentlicht: (2025)
von: Fang, Zipeng, et al.
Veröffentlicht: (2025)
SkelVIT: Consensus of Vision Transformers for a Lightweight Skeleton-Based Action Recognition System
von: Karadag, Ozge Oztimur
Veröffentlicht: (2023)
von: Karadag, Ozge Oztimur
Veröffentlicht: (2023)
Efficient Heatmap-Guided 6-Dof Grasp Detection in Cluttered Scenes
von: Chen, Siang, et al.
Veröffentlicht: (2024)
von: Chen, Siang, et al.
Veröffentlicht: (2024)
NuPlanQA: A Large-Scale Dataset and Benchmark for Multi-View Driving Scene Understanding in Multi-Modal Large Language Models
von: Park, Sung-Yeon, et al.
Veröffentlicht: (2025)
von: Park, Sung-Yeon, et al.
Veröffentlicht: (2025)
LM-MCVT: A Lightweight Multi-modal Multi-view Convolutional-Vision Transformer Approach for 3D Object Recognition
von: Xiong, Songsong, et al.
Veröffentlicht: (2025)
von: Xiong, Songsong, et al.
Veröffentlicht: (2025)
SE-VLN: A Self-Evolving Vision-Language Navigation Framework Based on Multimodal Large Language Models
von: Dong, Xiangyu, et al.
Veröffentlicht: (2025)
von: Dong, Xiangyu, et al.
Veröffentlicht: (2025)
FunGraph: Functionality Aware 3D Scene Graphs for Language-Prompted Scene Interaction
von: Rotondi, Dennis, et al.
Veröffentlicht: (2025)
von: Rotondi, Dennis, et al.
Veröffentlicht: (2025)
GMOR: A Lightweight Robust Point Cloud Registration Framework via Geometric Maximum Overlapping
von: Zheng, Zhao, et al.
Veröffentlicht: (2025)
von: Zheng, Zhao, et al.
Veröffentlicht: (2025)
TaskGround: Structured Executable Task Inference for Full-Scene Household Reasoning
von: Feng, ZhiYuan, et al.
Veröffentlicht: (2026)
von: Feng, ZhiYuan, et al.
Veröffentlicht: (2026)
Multimodal HD Mapping for Intersections by Intelligent Roadside Units
von: Chen, Zhongzhang, et al.
Veröffentlicht: (2025)
von: Chen, Zhongzhang, et al.
Veröffentlicht: (2025)
From Scene to Object: Text-Guided Dual-Gaze Prediction
von: Ke, Zehong, et al.
Veröffentlicht: (2026)
von: Ke, Zehong, et al.
Veröffentlicht: (2026)
Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?
von: Majumdar, Arjun, et al.
Veröffentlicht: (2023)
von: Majumdar, Arjun, et al.
Veröffentlicht: (2023)
CUS-GS: A Compact Unified Structured Gaussian Splatting Framework for Multimodal Scene Representation
von: Ming, Yuhang, et al.
Veröffentlicht: (2025)
von: Ming, Yuhang, et al.
Veröffentlicht: (2025)
Genie 4D: Semantic-Prior-Guided 4D Dynamic Scene Reconstruction
von: Yang, Yiru, et al.
Veröffentlicht: (2026)
von: Yang, Yiru, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding
von: Li, Haoyuan, et al.
Veröffentlicht: (2025) -
Lightweight Structured Multimodal Reasoning for Clinical Scene Understanding in Robotics
von: Jha, Saurav, et al.
Veröffentlicht: (2025) -
MS-Net: A Multi-Path Sparse Model for Motion Prediction in Multi-Scenes
von: Tang, Xiaqiang, et al.
Veröffentlicht: (2024) -
Aerial Maritime Vessel Detection and Identification
von: Kulas, Antonella Barisic, et al.
Veröffentlicht: (2025) -
AeroLite-MDNet: Lightweight Multi-task Deviation Detection Network for UAV Landing
von: Yang, Haiping, et al.
Veröffentlicht: (2025)