Enhancing Underwater Object Detection through Spatio-Temporal Analysis and Spatial Attention Networks
Fuente:
arXiv
Saved in:
| Main Authors: | Karri, Sai Likhith, Saxena, Ansh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding Spatio-Temporal Relations in Human-Object Interaction using Pyramid Graph Convolutional Network
by: Xing, Hao, et al.
Published: (2024)
by: Xing, Hao, et al.
Published: (2024)
StreamLTS: Query-based Temporal-Spatial LiDAR Fusion for Cooperative Object Detection
by: Yuan, Yunshuang, et al.
Published: (2024)
by: Yuan, Yunshuang, et al.
Published: (2024)
STRNet: Visual Navigation with Spatio-Temporal Representation through Dynamic Graph Aggregation
by: Ren, Hao, et al.
Published: (2026)
by: Ren, Hao, et al.
Published: (2026)
LiDAR Loop Closure Detection using Semantic Graphs with Graph Attention Networks
by: Yang, Liudi, et al.
Published: (2025)
by: Yang, Liudi, et al.
Published: (2025)
Quaternion Approximation Networks for Enhanced Image Classification and Oriented Object Detection
by: Grant, Bryce, et al.
Published: (2025)
by: Grant, Bryce, et al.
Published: (2025)
FADet: A Multi-sensor 3D Object Detection Network based on Local Featured Attention
by: Guo, Ziang, et al.
Published: (2024)
by: Guo, Ziang, et al.
Published: (2024)
ACE-Brain-0: Spatial Intelligence as a Shared Scaffold for Universal Embodiments
by: Gong, Ziyang, et al.
Published: (2026)
by: Gong, Ziyang, et al.
Published: (2026)
Enhancing Track Management Systems with Vehicle-To-Vehicle Enabled Sensor Fusion
by: Billington, Thomas, et al.
Published: (2024)
by: Billington, Thomas, et al.
Published: (2024)
Transformer-Based Spatio-Temporal Association of Apple Fruitlets
by: Freeman, Harry, et al.
Published: (2025)
by: Freeman, Harry, et al.
Published: (2025)
End-to-End Navigation with Vision Language Models: Transforming Spatial Reasoning into Question-Answering
by: Goetting, Dylan, et al.
Published: (2024)
by: Goetting, Dylan, et al.
Published: (2024)
DivScene: Towards Open-Vocabulary Object Navigation with Large Vision Language Models in Diverse Scenes
by: Wang, Zhaowei, et al.
Published: (2024)
by: Wang, Zhaowei, et al.
Published: (2024)
Benchmarking Online Object Trackers for Underwater Robot Position Locking Applications
by: Safa, Ali, et al.
Published: (2025)
by: Safa, Ali, et al.
Published: (2025)
Resolving Spatio-Temporal Entanglement in Video Prediction via Multi-Modal Attention
by: Gupta, Shreyam, et al.
Published: (2025)
by: Gupta, Shreyam, et al.
Published: (2025)
DM2RM: Dual-Mode Multimodal Ranking for Target Objects and Receptacles Based on Open-Vocabulary Instructions
by: Korekata, Ryosuke, et al.
Published: (2024)
by: Korekata, Ryosuke, et al.
Published: (2024)
ST-$π$: Structured SpatioTemporal VLA for Robotic Manipulation
by: Ma, Chuanhao, et al.
Published: (2026)
by: Ma, Chuanhao, et al.
Published: (2026)
A Diver Attention Estimation Framework for Effective Underwater Human-Robot Interaction
by: Enan, Sadman Sakib, et al.
Published: (2022)
by: Enan, Sadman Sakib, et al.
Published: (2022)
STARK: Spatio-Temporal Attention for Representation of Keypoints for Continuous Sign Language Recognition
by: Patra, Suvajit, et al.
Published: (2026)
by: Patra, Suvajit, et al.
Published: (2026)
Enhancing Video-Language Representations with Structural Spatio-Temporal Alignment
by: Fei, Hao, et al.
Published: (2024)
by: Fei, Hao, et al.
Published: (2024)
Fusion-Poly: A Polyhedral Framework Based on Spatial-Temporal Fusion for 3D Multi-Object Tracking
by: Wu, Xian, et al.
Published: (2026)
by: Wu, Xian, et al.
Published: (2026)
ProGAL-VLA: Grounded Alignment through Prospective Reasoning in Vision-Language-Action Models
by: Darabi, Nastaran, et al.
Published: (2026)
by: Darabi, Nastaran, et al.
Published: (2026)
STEP: Spatial Temporal Graph Convolutional Networks for Emotion Perception from Gaits
by: Bhattacharya, Uttaran, et al.
Published: (2019)
by: Bhattacharya, Uttaran, et al.
Published: (2019)
Are All Marine Species Created Equal? Performance Disparities in Underwater Object Detection
by: Wille, Melanie, et al.
Published: (2025)
by: Wille, Melanie, et al.
Published: (2025)
BOP-ASK: Object-Interaction Reasoning for Vision-Language Models
by: Bhat, Vineet, et al.
Published: (2025)
by: Bhat, Vineet, et al.
Published: (2025)
PNE-SGAN: Probabilistic NDT-Enhanced Semantic Graph Attention Network for LiDAR Loop Closure Detection
by: Li, Xiong, et al.
Published: (2025)
by: Li, Xiong, et al.
Published: (2025)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
by: Chen, Boyuan, et al.
Published: (2024)
by: Chen, Boyuan, et al.
Published: (2024)
LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
by: Fei, Senyu, et al.
Published: (2025)
by: Fei, Senyu, et al.
Published: (2025)
SNOW: Spatio-Temporal Scene Understanding with World Knowledge for Open-World Embodied Reasoning
by: Sohn, Tin Stribor, et al.
Published: (2025)
by: Sohn, Tin Stribor, et al.
Published: (2025)
3DGS-Calib: 3D Gaussian Splatting for Multimodal SpatioTemporal Calibration
by: Herau, Quentin, et al.
Published: (2024)
by: Herau, Quentin, et al.
Published: (2024)
General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling
by: Lyu, Huaihai, et al.
Published: (2026)
by: Lyu, Huaihai, et al.
Published: (2026)
Evaluating Robustness of Visual Representations for Object Assembly Task Requiring Spatio-Geometrical Reasoning
by: Ku, Chahyon, et al.
Published: (2023)
by: Ku, Chahyon, et al.
Published: (2023)
Why Domain Matters: A Preliminary Study of Domain Effects in Underwater Object Detection
by: Wille, Melanie, et al.
Published: (2026)
by: Wille, Melanie, et al.
Published: (2026)
Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding
by: Patratskiy, Maxim A., et al.
Published: (2025)
by: Patratskiy, Maxim A., et al.
Published: (2025)
TDANet: Target-Directed Attention Network For Object-Goal Visual Navigation With Zero-Shot Ability
by: Lian, Shiwei, et al.
Published: (2024)
by: Lian, Shiwei, et al.
Published: (2024)
RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
by: Song, Chan Hee, et al.
Published: (2024)
by: Song, Chan Hee, et al.
Published: (2024)
ST-Booster: An Iterative SpatioTemporal Perception Booster for Vision-and-Language Navigation in Continuous Environments
by: Yue, Lu, et al.
Published: (2025)
by: Yue, Lu, et al.
Published: (2025)
SOAC: Spatio-Temporal Overlap-Aware Multi-Sensor Calibration using Neural Radiance Fields
by: Herau, Quentin, et al.
Published: (2023)
by: Herau, Quentin, et al.
Published: (2023)
CRKD: Enhanced Camera-Radar Object Detection with Cross-modality Knowledge Distillation
by: Zhao, Lingjun, et al.
Published: (2024)
by: Zhao, Lingjun, et al.
Published: (2024)
R4: Retrieval-Augmented Reasoning for Vision-Language Models in 4D Spatio-Temporal Space
by: Sohn, Tin Stribor, et al.
Published: (2025)
by: Sohn, Tin Stribor, et al.
Published: (2025)
OceanGym: A Benchmark Environment for Underwater Embodied Agents
by: Xue, Yida, et al.
Published: (2025)
by: Xue, Yida, et al.
Published: (2025)
Can Transformers Capture Spatial Relations between Objects?
by: Wen, Chuan, et al.
Published: (2024)
by: Wen, Chuan, et al.
Published: (2024)
Similar Items
-
Understanding Spatio-Temporal Relations in Human-Object Interaction using Pyramid Graph Convolutional Network
by: Xing, Hao, et al.
Published: (2024) -
StreamLTS: Query-based Temporal-Spatial LiDAR Fusion for Cooperative Object Detection
by: Yuan, Yunshuang, et al.
Published: (2024) -
STRNet: Visual Navigation with Spatio-Temporal Representation through Dynamic Graph Aggregation
by: Ren, Hao, et al.
Published: (2026) -
LiDAR Loop Closure Detection using Semantic Graphs with Graph Attention Networks
by: Yang, Liudi, et al.
Published: (2025) -
Quaternion Approximation Networks for Enhanced Image Classification and Oriented Object Detection
by: Grant, Bryce, et al.
Published: (2025)