Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guan, Runwei, Ouyang, Ningwei, Xu, Tianhao, Liang, Shaofeng, Dai, Wei, Sun, Yafeng, Gao, Shang, Lai, Songning, Yao, Shanliang, Hu, Xuming, Liu, Ryan Wen, Yue, Yutao, Xiong, Hui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
USVTrack: USV-Based 4D Radar-Camera Tracking Dataset for Autonomous Driving in Inland Waterways
von: Yao, Shanliang, et al.
Veröffentlicht: (2025)
von: Yao, Shanliang, et al.
Veröffentlicht: (2025)
WaterVG: Waterway Visual Grounding based on Text-Guided Vision and mmWave Radar
von: Guan, Runwei, et al.
Veröffentlicht: (2024)
von: Guan, Runwei, et al.
Veröffentlicht: (2024)
ASY-VRNet: Waterway Panoptic Driving Perception Model based on Asymmetric Fair Fusion of Vision and 4D mmWave Radar
von: Guan, Runwei, et al.
Veröffentlicht: (2023)
von: Guan, Runwei, et al.
Veröffentlicht: (2023)
WaterVideoQA: ASV-Centric Perception and Rule-Compliant Reasoning via Multi-Modal Agents
von: Guan, Runwei, et al.
Veröffentlicht: (2026)
von: Guan, Runwei, et al.
Veröffentlicht: (2026)
Talk2PC: Enhancing 3D Visual Grounding through LiDAR and Radar Point Clouds Fusion for Autonomous Driving
von: Guan, Runwei, et al.
Veröffentlicht: (2025)
von: Guan, Runwei, et al.
Veröffentlicht: (2025)
NanoMVG: USV-Centric Low-Power Multi-Task Visual Grounding based on Prompt-Guided Camera and 4D mmWave Radar
von: Guan, Runwei, et al.
Veröffentlicht: (2024)
von: Guan, Runwei, et al.
Veröffentlicht: (2024)
PEPL: Precision-Enhanced Pseudo-Labeling for Fine-Grained Image Classification in Semi-Supervised Learning
von: Tian, Bowen, et al.
Veröffentlicht: (2024)
von: Tian, Bowen, et al.
Veröffentlicht: (2024)
RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation System
von: Guan, Runwei, et al.
Veröffentlicht: (2025)
von: Guan, Runwei, et al.
Veröffentlicht: (2025)
Mitigating Spurious Background Bias in Multimedia Recognition with Disentangled Concept Bottlenecks
von: Huang, Gaoxiang, et al.
Veröffentlicht: (2025)
von: Huang, Gaoxiang, et al.
Veröffentlicht: (2025)
Wolf2Pack: The AutoFusion Framework for Dynamic Parameter Fusion
von: Tian, Bowen, et al.
Veröffentlicht: (2024)
von: Tian, Bowen, et al.
Veröffentlicht: (2024)
Talk2Radar: Bridging Natural Language with 4D mmWave Radar for 3D Referring Expression Comprehension
von: Guan, Runwei, et al.
Veröffentlicht: (2024)
von: Guan, Runwei, et al.
Veröffentlicht: (2024)
DRIVE: Dependable Robust Interpretable Visionary Ensemble Framework in Autonomous Driving
von: Lai, Songning, et al.
Veröffentlicht: (2024)
von: Lai, Songning, et al.
Veröffentlicht: (2024)
MMDrive: Interactive Scene Understanding Beyond Vision with Multi-representational Fusion
von: Hou, Minghui, et al.
Veröffentlicht: (2025)
von: Hou, Minghui, et al.
Veröffentlicht: (2025)
Cognitive Disentanglement for Referring Multi-Object Tracking
von: Liang, Shaofeng, et al.
Veröffentlicht: (2025)
von: Liang, Shaofeng, et al.
Veröffentlicht: (2025)
Wavelet-based Multi-View Fusion of 4D Radar Tensor and Camera for Robust 3D Object Detection
von: Guan, Runwei, et al.
Veröffentlicht: (2025)
von: Guan, Runwei, et al.
Veröffentlicht: (2025)
WaterScenes: A Multi-Task 4D Radar-Camera Fusion Dataset and Benchmarks for Autonomous Driving on Water Surfaces
von: Yao, Shanliang, et al.
Veröffentlicht: (2023)
von: Yao, Shanliang, et al.
Veröffentlicht: (2023)
Adaptive H&E-IHC information fusion staining framework based on feature extra
von: Jia, Yifan, et al.
Veröffentlicht: (2025)
von: Jia, Yifan, et al.
Veröffentlicht: (2025)
SECURE: Stable Early Collision Understanding via Robust Embeddings in Autonomous Driving
von: Wang, Wenjing, et al.
Veröffentlicht: (2026)
von: Wang, Wenjing, et al.
Veröffentlicht: (2026)
Guarding the Gate: ConceptGuard Battles Concept-Level Backdoors in Concept Bottleneck Models
von: Lai, Songning, et al.
Veröffentlicht: (2024)
von: Lai, Songning, et al.
Veröffentlicht: (2024)
DRIVE: Dual-Robustness via Information Variability and Entropic Consistency in Source-Free Unsupervised Domain Adaptation
von: Xiao, Ruiqiang, et al.
Veröffentlicht: (2024)
von: Xiao, Ruiqiang, et al.
Veröffentlicht: (2024)
Text2Weight: Bridging Natural Language and Neural Network Weight Spaces
von: Tian, Bowen, et al.
Veröffentlicht: (2025)
von: Tian, Bowen, et al.
Veröffentlicht: (2025)
Beyond Task Vectors: Selective Task Arithmetic Based on Importance Metrics
von: Bowen, Tian, et al.
Veröffentlicht: (2024)
von: Bowen, Tian, et al.
Veröffentlicht: (2024)
Maintaining Informative Coherence: Migrating Hallucinations in Large Language Models via Absorbing Markov Chains
von: Wu, Jiemin, et al.
Veröffentlicht: (2024)
von: Wu, Jiemin, et al.
Veröffentlicht: (2024)
radarODE: An ODE-Embedded Deep Learning Model for Contactless ECG Reconstruction from Millimeter-Wave Radar
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024)
UniBEVFusion: Unified Radar-Vision BEVFusion for 3D Object Detection
von: Zhao, Haocheng, et al.
Veröffentlicht: (2024)
von: Zhao, Haocheng, et al.
Veröffentlicht: (2024)
Exploring Radar Data Representations in Autonomous Driving: A Comprehensive Review
von: Yao, Shanliang, et al.
Veröffentlicht: (2023)
von: Yao, Shanliang, et al.
Veröffentlicht: (2023)
4D-CAAL: 4D Radar-Camera Calibration and Auto-Labeling for Autonomous Driving
von: Yao, Shanliang, et al.
Veröffentlicht: (2026)
von: Yao, Shanliang, et al.
Veröffentlicht: (2026)
USV: Towards Understanding the User-generated Short-form Videos
von: Cheng, Haoyue, et al.
Veröffentlicht: (2026)
von: Cheng, Haoyue, et al.
Veröffentlicht: (2026)
ACE: Attribution-Controlled Knowledge Editing for Multi-hop Factual Recall
von: Yang, Jiayu, et al.
Veröffentlicht: (2025)
von: Yang, Jiayu, et al.
Veröffentlicht: (2025)
IMTS is Worth Time $\times$ Channel Patches: Visual Masked Autoencoders for Irregular Multivariate Time Series Prediction
von: Hu, Zhangyi, et al.
Veröffentlicht: (2025)
von: Hu, Zhangyi, et al.
Veröffentlicht: (2025)
Free-T2M: Robust Text-to-Motion Generation for Humanoid Robots via Frequency-Domain
von: Chen, Wenshuo, et al.
Veröffentlicht: (2025)
von: Chen, Wenshuo, et al.
Veröffentlicht: (2025)
FTS: A Framework to Find a Faithful TimeSieve
von: Lai, Songning, et al.
Veröffentlicht: (2024)
von: Lai, Songning, et al.
Veröffentlicht: (2024)
ExCap3D: Expressive 3D Scene Understanding via Object Captioning with Varying Detail
von: Yeshwanth, Chandan, et al.
Veröffentlicht: (2025)
von: Yeshwanth, Chandan, et al.
Veröffentlicht: (2025)
Perturbation-mitigated USV Navigation with Distributionally Robust Reinforcement Learning
von: Zhang, Zhaofan, et al.
Veröffentlicht: (2025)
von: Zhang, Zhaofan, et al.
Veröffentlicht: (2025)
Optimizing USV AoI for RIS‐Assisted UAV–USV MEC Network
von: Chao Ma, et al.
Veröffentlicht: (2025)
von: Chao Ma, et al.
Veröffentlicht: (2025)
Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms
von: Cao, Zhixiang, et al.
Veröffentlicht: (2026)
von: Cao, Zhixiang, et al.
Veröffentlicht: (2026)
Large Foundation Models for Trajectory Prediction in Autonomous Driving: A Comprehensive Survey
von: Dai, Wei, et al.
Veröffentlicht: (2025)
von: Dai, Wei, et al.
Veröffentlicht: (2025)
RadarNeXt: Real-Time and Reliable 3D Object Detector Based On 4D mmWave Imaging Radar
von: Jia, Liye, et al.
Veröffentlicht: (2025)
von: Jia, Liye, et al.
Veröffentlicht: (2025)
Rivers and Waterways in the Roman World
Veröffentlicht: (2024)
Veröffentlicht: (2024)
You Only Train Once: A Flexible Training Framework for Code Vulnerability Detection Driven by Vul-Vector
von: Tian, Bowen, et al.
Veröffentlicht: (2025)
von: Tian, Bowen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
USVTrack: USV-Based 4D Radar-Camera Tracking Dataset for Autonomous Driving in Inland Waterways
von: Yao, Shanliang, et al.
Veröffentlicht: (2025) -
WaterVG: Waterway Visual Grounding based on Text-Guided Vision and mmWave Radar
von: Guan, Runwei, et al.
Veröffentlicht: (2024) -
ASY-VRNet: Waterway Panoptic Driving Perception Model based on Asymmetric Fair Fusion of Vision and 4D mmWave Radar
von: Guan, Runwei, et al.
Veröffentlicht: (2023) -
WaterVideoQA: ASV-Centric Perception and Rule-Compliant Reasoning via Multi-Modal Agents
von: Guan, Runwei, et al.
Veröffentlicht: (2026) -
Talk2PC: Enhancing 3D Visual Grounding through LiDAR and Radar Point Clouds Fusion for Autonomous Driving
von: Guan, Runwei, et al.
Veröffentlicht: (2025)