SJTU:Spatial judgments in multimodal models towards unified segmentation through coordinate detection
Fuente:
arXiv
Salvato in:
| Autori principali: | Chae, Joongwon, Wang, Zhenyu, Qin, Peiwu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Grid-augmented vision: A simple yet effective approach for enhanced spatial understanding in multi-modal agents
di: Chae, Joongwon, et al.
Pubblicazione: (2024)
di: Chae, Joongwon, et al.
Pubblicazione: (2024)
Memory-SAM: Human-Prompt-Free Tongue Segmentation via Retrieval-to-Prompt
di: Chae, Joongwon, et al.
Pubblicazione: (2025)
di: Chae, Joongwon, et al.
Pubblicazione: (2025)
GCR: Geometry-Consistent Routing for Task-Agnostic Continual Anomaly Detection
di: Chae, Joongwon, et al.
Pubblicazione: (2026)
di: Chae, Joongwon, et al.
Pubblicazione: (2026)
StructCore: Structure-Aware Image-Level Scoring for Training-Free Unsupervised Anomaly Detection
di: Chae, Joongwon, et al.
Pubblicazione: (2026)
di: Chae, Joongwon, et al.
Pubblicazione: (2026)
On the robustness of multimodal language model towards distractions
di: Liu, Ming, et al.
Pubblicazione: (2025)
di: Liu, Ming, et al.
Pubblicazione: (2025)
MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation
di: Xing, Yang, et al.
Pubblicazione: (2026)
di: Xing, Yang, et al.
Pubblicazione: (2026)
IRFundusSet: An Integrated Retinal Fundus Dataset with a Harmonized Healthy Label
di: Githinji, P. Bilha, et al.
Pubblicazione: (2024)
di: Githinji, P. Bilha, et al.
Pubblicazione: (2024)
DBF-UNet: A Two-Stage Framework for Carotid Artery Segmentation with Pseudo-Label Generation
di: Li, Haoxuan, et al.
Pubblicazione: (2025)
di: Li, Haoxuan, et al.
Pubblicazione: (2025)
Mineral segmentation using electron microscope images and spectral sampling through multimodal graph neural networks
di: Repka, Samuel, et al.
Pubblicazione: (2025)
di: Repka, Samuel, et al.
Pubblicazione: (2025)
CellPilot: A unified approach to automatic and interactive segmentation in histopathology
di: Endres, Philipp, et al.
Pubblicazione: (2024)
di: Endres, Philipp, et al.
Pubblicazione: (2024)
Open-set object detection: towards unified problem formulation and benchmarking
di: Ammar, Hejer, et al.
Pubblicazione: (2024)
di: Ammar, Hejer, et al.
Pubblicazione: (2024)
Cognitive resilience: Unraveling the proficiency of image-captioning models to interpret masked visual content
di: Du, Zhicheng, et al.
Pubblicazione: (2024)
di: Du, Zhicheng, et al.
Pubblicazione: (2024)
SAM-based instance segmentation models for the automation of structural damage detection
di: Ye, Zehao, et al.
Pubblicazione: (2024)
di: Ye, Zehao, et al.
Pubblicazione: (2024)
MULTIAQUA: A multimodal maritime dataset and robust training strategies for multimodal semantic segmentation
di: Muhovič, Jon, et al.
Pubblicazione: (2025)
di: Muhovič, Jon, et al.
Pubblicazione: (2025)
COMPrompter: reconceptualized segment anything model with multiprompt network for camouflaged object detection
di: Zhang, Xiaoqin, et al.
Pubblicazione: (2024)
di: Zhang, Xiaoqin, et al.
Pubblicazione: (2024)
Improved Unet model for brain tumor image segmentation based on ASPP-coordinate attention mechanism
di: Wang, Zixuan, et al.
Pubblicazione: (2024)
di: Wang, Zixuan, et al.
Pubblicazione: (2024)
Attacks on multimodal models
di: Iablochnikov, Viacheslav, et al.
Pubblicazione: (2024)
di: Iablochnikov, Viacheslav, et al.
Pubblicazione: (2024)
A unified FLAIR hyperintensity segmentation model for various CNS tumor types and acquisition time points
di: Faanes, Mathilde Gajda, et al.
Pubblicazione: (2025)
di: Faanes, Mathilde Gajda, et al.
Pubblicazione: (2025)
SurgiSAM2: Fine-tuning a foundational model for surgical video anatomy segmentation and detection
di: Kamtam, Devanish N., et al.
Pubblicazione: (2025)
di: Kamtam, Devanish N., et al.
Pubblicazione: (2025)
Out-of-distribution data supervision towards biomedical semantic segmentation
di: Gao, Yiquan, et al.
Pubblicazione: (2025)
di: Gao, Yiquan, et al.
Pubblicazione: (2025)
SITE: towards Spatial Intelligence Thorough Evaluation
di: Wang, Wenqi, et al.
Pubblicazione: (2025)
di: Wang, Wenqi, et al.
Pubblicazione: (2025)
Mask Focal Loss: A unifying framework for dense crowd counting with canonical object detection networks
di: Zhong, Xiaopin, et al.
Pubblicazione: (2022)
di: Zhong, Xiaopin, et al.
Pubblicazione: (2022)
MIMO: A medical vision language model with visual referring multimodal input and pixel grounding multimodal output
di: Chen, Yanyuan, et al.
Pubblicazione: (2025)
di: Chen, Yanyuan, et al.
Pubblicazione: (2025)
Colony Grounded SAM2: Zero-shot detection and segmentation of bacterial colonies using foundation models
di: Korporaal, Daan, et al.
Pubblicazione: (2026)
di: Korporaal, Daan, et al.
Pubblicazione: (2026)
LDRFusion: A LiDAR-Dominant multimodal refinement framework for 3D object detection
di: Wang, Jijun, et al.
Pubblicazione: (2025)
di: Wang, Jijun, et al.
Pubblicazione: (2025)
UF-AMA: A unified framework for cross-domain emotion recognition via adaptive multimodal alignment
di: Wang, Zheng, et al.
Pubblicazione: (2026)
di: Wang, Zheng, et al.
Pubblicazione: (2026)
BoundMatch: Boundary detection applied to semi-supervised segmentation
di: Ishikawa, Haruya, et al.
Pubblicazione: (2025)
di: Ishikawa, Haruya, et al.
Pubblicazione: (2025)
GlitchBench: Can large multimodal models detect video game glitches?
di: Taesiri, Mohammad Reza, et al.
Pubblicazione: (2023)
di: Taesiri, Mohammad Reza, et al.
Pubblicazione: (2023)
SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks
di: Song, Zijian, et al.
Pubblicazione: (2025)
di: Song, Zijian, et al.
Pubblicazione: (2025)
External Prompt Features Enhanced Parameter-efficient Fine-tuning for Salient Object Detection
di: Liang, Wen, et al.
Pubblicazione: (2024)
di: Liang, Wen, et al.
Pubblicazione: (2024)
A multi-center analysis of deep learning methods for video polyp detection and segmentation
di: Ghatwary, Noha, et al.
Pubblicazione: (2026)
di: Ghatwary, Noha, et al.
Pubblicazione: (2026)
Dynamically evolving segment anything model with continuous learning for medical image segmentation
di: Liu, Zhaori, et al.
Pubblicazione: (2025)
di: Liu, Zhaori, et al.
Pubblicazione: (2025)
The devil is in the object boundary: towards annotation-free instance segmentation using Foundation Models
di: Shi, Cheng, et al.
Pubblicazione: (2024)
di: Shi, Cheng, et al.
Pubblicazione: (2024)
Rethinking Exposure Correction for Spatially Non-uniform Degradation
di: Li, Ao, et al.
Pubblicazione: (2026)
di: Li, Ao, et al.
Pubblicazione: (2026)
GCA-ResUNet:Image segmentation in medical images using grouped coordinate attention
di: Ding, Jun, et al.
Pubblicazione: (2025)
di: Ding, Jun, et al.
Pubblicazione: (2025)
Segment Any Architectural Facades (SAAF):An automatic segmentation model for building facades, walls and windows based on multimodal semantics guidance
di: Li, Peilin, et al.
Pubblicazione: (2025)
di: Li, Peilin, et al.
Pubblicazione: (2025)
MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the Metaverse
di: Pan, Zhenyu, et al.
Pubblicazione: (2025)
di: Pan, Zhenyu, et al.
Pubblicazione: (2025)
HFI: A unified framework for training-free detection and implicit watermarking of latent diffusion model generated images
di: Choi, Sungik, et al.
Pubblicazione: (2024)
di: Choi, Sungik, et al.
Pubblicazione: (2024)
HCMA-UNet: A Hybrid CNN-Mamba UNet with Axial Self-Attention for Efficient Breast Cancer Segmentation
di: Li, Haoxuan, et al.
Pubblicazione: (2025)
di: Li, Haoxuan, et al.
Pubblicazione: (2025)
An unsupervised approach towards promptable defect segmentation in laser-based additive manufacturing by Segment Anything
di: Era, Israt Zarin, et al.
Pubblicazione: (2023)
di: Era, Israt Zarin, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Grid-augmented vision: A simple yet effective approach for enhanced spatial understanding in multi-modal agents
di: Chae, Joongwon, et al.
Pubblicazione: (2024) -
Memory-SAM: Human-Prompt-Free Tongue Segmentation via Retrieval-to-Prompt
di: Chae, Joongwon, et al.
Pubblicazione: (2025) -
GCR: Geometry-Consistent Routing for Task-Agnostic Continual Anomaly Detection
di: Chae, Joongwon, et al.
Pubblicazione: (2026) -
StructCore: Structure-Aware Image-Level Scoring for Training-Free Unsupervised Anomaly Detection
di: Chae, Joongwon, et al.
Pubblicazione: (2026) -
On the robustness of multimodal language model towards distractions
di: Liu, Ming, et al.
Pubblicazione: (2025)