MULTIAQUA: A multimodal maritime dataset and robust training strategies for multimodal semantic segmentation
Fuente:
arXiv
Saved in:
| Main Authors: | Muhovič, Jon, Perš, Janez |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Grading Handwritten Engineering Exams with Multimodal Large Language Models
by: Perš, Janez, et al.
Published: (2026)
by: Perš, Janez, et al.
Published: (2026)
P-NOC: adversarial training of CAM generating networks for robust weakly supervised semantic segmentation priors
by: David, Lucas, et al.
Published: (2023)
by: David, Lucas, et al.
Published: (2023)
Buffer replay enhances the robustness of multimodal learning under missing-modality
by: Zhu, Hongye, et al.
Published: (2025)
by: Zhu, Hongye, et al.
Published: (2025)
MAN TruckScenes: A multimodal dataset for autonomous trucking in diverse conditions
by: Fent, Felix, et al.
Published: (2024)
by: Fent, Felix, et al.
Published: (2024)
Closing the gap in multimodal medical representation alignment
by: Grassucci, Eleonora, et al.
Published: (2026)
by: Grassucci, Eleonora, et al.
Published: (2026)
SELECTOR: Heterogeneous graph network with convolutional masked autoencoder for multimodal robust prediction of cancer survival
by: Pan, Liangrui, et al.
Published: (2024)
by: Pan, Liangrui, et al.
Published: (2024)
Next day fire prediction via semantic segmentation
by: Alexis, Konstantinos, et al.
Published: (2024)
by: Alexis, Konstantinos, et al.
Published: (2024)
Pseudolabel guided pixels contrast for domain adaptive semantic segmentation
by: Xiang, Jianzi, et al.
Published: (2025)
by: Xiang, Jianzi, et al.
Published: (2025)
Context-self contrastive pretraining for crop type semantic segmentation
by: Tarasiou, Michail, et al.
Published: (2021)
by: Tarasiou, Michail, et al.
Published: (2021)
Classification of freshwater snails of the genus Radomaniola with multimodal triplet networks
by: Vetter, Dennis, et al.
Published: (2024)
by: Vetter, Dennis, et al.
Published: (2024)
Can multimodal representation learning by alignment preserve modality-specific information?
by: Thoreau, Romain, et al.
Published: (2025)
by: Thoreau, Romain, et al.
Published: (2025)
Leveraging SO(3)-steerable convolutions for pose-robust semantic segmentation in 3D medical data
by: Diaz, Ivan, et al.
Published: (2023)
by: Diaz, Ivan, et al.
Published: (2023)
Visual concept ranking uncovers medical shortcuts used by large multimodal models
by: Janizek, Joseph D., et al.
Published: (2026)
by: Janizek, Joseph D., et al.
Published: (2026)
Towards a multimodal framework for remote sensing image change retrieval and captioning
by: Ferrod, Roger, et al.
Published: (2024)
by: Ferrod, Roger, et al.
Published: (2024)
Review of multimodal machine learning approaches in healthcare
by: Krones, Felix, et al.
Published: (2024)
by: Krones, Felix, et al.
Published: (2024)
Do multimodal models imagine electric sheep?
by: Ramakrishnan, Santhosh Kumar, et al.
Published: (2026)
by: Ramakrishnan, Santhosh Kumar, et al.
Published: (2026)
Self-supervised cost of transport estimation for multimodal path planning
by: Gherold, Vincent, et al.
Published: (2024)
by: Gherold, Vincent, et al.
Published: (2024)
A multimodal slice discovery framework for systematic failure detection and explanation in medical image classification
by: Liu, Yixuan, et al.
Published: (2026)
by: Liu, Yixuan, et al.
Published: (2026)
Intrinsic Dimension Correlation: uncovering nonlinear connections in multimodal representations
by: Basile, Lorenzo, et al.
Published: (2024)
by: Basile, Lorenzo, et al.
Published: (2024)
A three in one bottom-up framework for simultaneous semantic segmentation, instance segmentation and classification of multi-organ nuclei in digital cancer histology
by: Ahmad, Ibtihaj, et al.
Published: (2023)
by: Ahmad, Ibtihaj, et al.
Published: (2023)
On the effectiveness of multimodal privileged knowledge distillation in two vision transformer based diagnostic applications
by: Baur, Simon, et al.
Published: (2025)
by: Baur, Simon, et al.
Published: (2025)
What do vision-language models see in the context? Investigating multimodal in-context learning
by: Santos, Gabriel O. dos, et al.
Published: (2025)
by: Santos, Gabriel O. dos, et al.
Published: (2025)
What to align in multimodal contrastive learning?
by: Dufumier, Benoit, et al.
Published: (2024)
by: Dufumier, Benoit, et al.
Published: (2024)
AG-Fusion: adaptive gated multimodal fusion for 3d object detection in complex scenes
by: Liu, Sixian, et al.
Published: (2025)
by: Liu, Sixian, et al.
Published: (2025)
On the robustness of multimodal language model towards distractions
by: Liu, Ming, et al.
Published: (2025)
by: Liu, Ming, et al.
Published: (2025)
COIN: Counterfactual inpainting for weakly supervised semantic segmentation for medical images
by: Shvetsov, Dmytro, et al.
Published: (2024)
by: Shvetsov, Dmytro, et al.
Published: (2024)
A feature refinement module for light-weight semantic segmentation network
by: Wang, Zhiyan, et al.
Published: (2024)
by: Wang, Zhiyan, et al.
Published: (2024)
Dense Center-Direction Regression for Object Counting and Localization with Point Supervision
by: Tabernik, Domen, et al.
Published: (2024)
by: Tabernik, Domen, et al.
Published: (2024)
Explaining latent representations of generative models with large multimodal models
by: Zhu, Mengdan, et al.
Published: (2024)
by: Zhu, Mengdan, et al.
Published: (2024)
DGSSM: Diffusion guided state-space models for multimodal salient object detection
by: Ghosh, Suklav, et al.
Published: (2026)
by: Ghosh, Suklav, et al.
Published: (2026)
Enhancing multimodal cooperation via sample-level modality valuation
by: Wei, Yake, et al.
Published: (2023)
by: Wei, Yake, et al.
Published: (2023)
FISBe: A real-world benchmark dataset for instance segmentation of long-range thin filamentous structures
by: Mais, Lisa, et al.
Published: (2024)
by: Mais, Lisa, et al.
Published: (2024)
Who's in and who's out? A case study of multimodal CLIP-filtering in DataComp
by: Hong, Rachel, et al.
Published: (2024)
by: Hong, Rachel, et al.
Published: (2024)
Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy
by: Ging, Simon, et al.
Published: (2024)
by: Ging, Simon, et al.
Published: (2024)
Generating crossmodal gene expression from cancer histopathology improves multimodal AI predictions
by: Dey, Samiran, et al.
Published: (2025)
by: Dey, Samiran, et al.
Published: (2025)
A generalised pre-training strategy for deep learning networks in semantic segmentation of remotely sensed images
by: Fang, Yuan, et al.
Published: (2026)
by: Fang, Yuan, et al.
Published: (2026)
A multimodal gesture recognition dataset for desktop human-computer interaction
by: Wang, Qi, et al.
Published: (2024)
by: Wang, Qi, et al.
Published: (2024)
Segment Any Architectural Facades (SAAF):An automatic segmentation model for building facades, walls and windows based on multimodal semantics guidance
by: Li, Peilin, et al.
Published: (2025)
by: Li, Peilin, et al.
Published: (2025)
OmDet: Large-scale vision-language multi-dataset pre-training with multimodal detection network
by: Zhao, Tiancheng, et al.
Published: (2022)
by: Zhao, Tiancheng, et al.
Published: (2022)
ME-WARD: A multimodal ergonomic analysis tool for musculoskeletal risk assessment from inertial and video data in working plac
by: González-Alonso, Javier, et al.
Published: (2026)
by: González-Alonso, Javier, et al.
Published: (2026)
Similar Items
-
Grading Handwritten Engineering Exams with Multimodal Large Language Models
by: Perš, Janez, et al.
Published: (2026) -
P-NOC: adversarial training of CAM generating networks for robust weakly supervised semantic segmentation priors
by: David, Lucas, et al.
Published: (2023) -
Buffer replay enhances the robustness of multimodal learning under missing-modality
by: Zhu, Hongye, et al.
Published: (2025) -
MAN TruckScenes: A multimodal dataset for autonomous trucking in diverse conditions
by: Fent, Felix, et al.
Published: (2024) -
Closing the gap in multimodal medical representation alignment
by: Grassucci, Eleonora, et al.
Published: (2026)