Grid-augmented vision: A simple yet effective approach for enhanced spatial understanding in multi-modal agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chae, Joongwon, Wang, Zhenyu, Zhang, Lian, Yu, Dongmei, Qin, Peiwu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SJTU:Spatial judgments in multimodal models towards unified segmentation through coordinate detection
von: Chae, Joongwon, et al.
Veröffentlicht: (2024)
von: Chae, Joongwon, et al.
Veröffentlicht: (2024)
Memory-SAM: Human-Prompt-Free Tongue Segmentation via Retrieval-to-Prompt
von: Chae, Joongwon, et al.
Veröffentlicht: (2025)
von: Chae, Joongwon, et al.
Veröffentlicht: (2025)
GCR: Geometry-Consistent Routing for Task-Agnostic Continual Anomaly Detection
von: Chae, Joongwon, et al.
Veröffentlicht: (2026)
von: Chae, Joongwon, et al.
Veröffentlicht: (2026)
StructCore: Structure-Aware Image-Level Scoring for Training-Free Unsupervised Anomaly Detection
von: Chae, Joongwon, et al.
Veröffentlicht: (2026)
von: Chae, Joongwon, et al.
Veröffentlicht: (2026)
LLaVAction: evaluating and training multi-modal large language models for action understanding
von: Qi, Haozhe, et al.
Veröffentlicht: (2025)
von: Qi, Haozhe, et al.
Veröffentlicht: (2025)
UAV traffic scene understanding: A regulation embedded multi-modal network and a unified benchmark
von: Zhang, Yu, et al.
Veröffentlicht: (2026)
von: Zhang, Yu, et al.
Veröffentlicht: (2026)
A multi-modal vision-language model for generalizable annotation-free pathology localization
von: Yang, Hao, et al.
Veröffentlicht: (2024)
von: Yang, Hao, et al.
Veröffentlicht: (2024)
DBF-UNet: A Two-Stage Framework for Carotid Artery Segmentation with Pseudo-Label Generation
von: Li, Haoxuan, et al.
Veröffentlicht: (2025)
von: Li, Haoxuan, et al.
Veröffentlicht: (2025)
Have we unified image generation and understanding yet? An empirical study of GPT-4o's image generation ability
von: Li, Ning, et al.
Veröffentlicht: (2025)
von: Li, Ning, et al.
Veröffentlicht: (2025)
Near, far: Patch-ordering enhances vision foundation models' scene understanding
von: Pariza, Valentinos, et al.
Veröffentlicht: (2024)
von: Pariza, Valentinos, et al.
Veröffentlicht: (2024)
IRFundusSet: An Integrated Retinal Fundus Dataset with a Harmonized Healthy Label
von: Githinji, P. Bilha, et al.
Veröffentlicht: (2024)
von: Githinji, P. Bilha, et al.
Veröffentlicht: (2024)
Task Alignment: A simple and effective proxy for model merging in computer vision
von: de Jorge, Pau, et al.
Veröffentlicht: (2026)
von: de Jorge, Pau, et al.
Veröffentlicht: (2026)
STONE: Pioneering the One-to-N Universal Backdoor Threat in 3D Point Cloud
von: Shan, Dongmei, et al.
Veröffentlicht: (2025)
von: Shan, Dongmei, et al.
Veröffentlicht: (2025)
SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion
von: Nassar, Ahmed, et al.
Veröffentlicht: (2025)
von: Nassar, Ahmed, et al.
Veröffentlicht: (2025)
Multi-label classification for multi-temporal, multi-spatial coral reef condition monitoring using vision foundation model with adapter learning
von: Shao, Xinlei, et al.
Veröffentlicht: (2025)
von: Shao, Xinlei, et al.
Veröffentlicht: (2025)
Cognitive resilience: Unraveling the proficiency of image-captioning models to interpret masked visual content
von: Du, Zhicheng, et al.
Veröffentlicht: (2024)
von: Du, Zhicheng, et al.
Veröffentlicht: (2024)
SHAP-CAT: A interpretable multi-modal framework enhancing WSI classification via virtual staining and shapley-value-based multimodal fusion
von: Wang, Jun, et al.
Veröffentlicht: (2024)
von: Wang, Jun, et al.
Veröffentlicht: (2024)
Do large language vision models understand 3D shapes?
von: Eppel, Sagi
Veröffentlicht: (2024)
von: Eppel, Sagi
Veröffentlicht: (2024)
Serial fusion of multi-modal biometric systems
von: Marcialis, Gian Luca, et al.
Veröffentlicht: (2024)
von: Marcialis, Gian Luca, et al.
Veröffentlicht: (2024)
Synthetic data augmentation for robotic mobility aids to support blind and low vision people
von: Hwang, Hochul, et al.
Veröffentlicht: (2024)
von: Hwang, Hochul, et al.
Veröffentlicht: (2024)
Post-hurricane building damage assessment using street-view imagery and structured data: A multi-modal deep learning approach
von: Xue, Zhuoqun, et al.
Veröffentlicht: (2024)
von: Xue, Zhuoqun, et al.
Veröffentlicht: (2024)
A vision-based framework for human behavior understanding in industrial assembly lines
von: Papoutsakis, Konstantinos, et al.
Veröffentlicht: (2024)
von: Papoutsakis, Konstantinos, et al.
Veröffentlicht: (2024)
CKDA: Cross-modality Knowledge Disentanglement and Alignment for Visible-Infrared Lifelong Person Re-identification
von: Cui, Zhenyu, et al.
Veröffentlicht: (2025)
von: Cui, Zhenyu, et al.
Veröffentlicht: (2025)
Unsupervised Multi-agent and Single-agent Perception from Cooperative Views
von: Yang, Haochen, et al.
Veröffentlicht: (2026)
von: Yang, Haochen, et al.
Veröffentlicht: (2026)
MIRAGE: Robust multi-modal architectures translate fMRI-to-image models from vision to mental imagery
von: Kneeland, Reese, et al.
Veröffentlicht: (2026)
von: Kneeland, Reese, et al.
Veröffentlicht: (2026)
Cross-modal ultra-scale learning with tri-modalities of renal biopsy images for glomerular multi-disease auxiliary diagnosis
von: Long, Kaixing, et al.
Veröffentlicht: (2025)
von: Long, Kaixing, et al.
Veröffentlicht: (2025)
HiERO: understanding the hierarchy of human behavior enhances reasoning on egocentric videos
von: Peirone, Simone Alberto, et al.
Veröffentlicht: (2025)
von: Peirone, Simone Alberto, et al.
Veröffentlicht: (2025)
HCMA-UNet: A Hybrid CNN-Mamba UNet with Axial Self-Attention for Efficient Breast Cancer Segmentation
von: Li, Haoxuan, et al.
Veröffentlicht: (2025)
von: Li, Haoxuan, et al.
Veröffentlicht: (2025)
External Prompt Features Enhanced Parameter-efficient Fine-tuning for Salient Object Detection
von: Liang, Wen, et al.
Veröffentlicht: (2024)
von: Liang, Wen, et al.
Veröffentlicht: (2024)
A data-centric approach to class-specific bias in image data augmentation
von: Angelakis, Athanasios, et al.
Veröffentlicht: (2024)
von: Angelakis, Athanasios, et al.
Veröffentlicht: (2024)
Unified modality separation: A vision-language framework for unsupervised domain adaptation
von: Li, Xinyao, et al.
Veröffentlicht: (2025)
von: Li, Xinyao, et al.
Veröffentlicht: (2025)
bi-modal textual prompt learning for vision-language models in remote sensing
von: Kashyap, Pankhi, et al.
Veröffentlicht: (2026)
von: Kashyap, Pankhi, et al.
Veröffentlicht: (2026)
A multi-temporal multi-spectral attention-augmented deep convolution neural network with contrastive learning for crop yield prediction
von: Dangi, Shalini, et al.
Veröffentlicht: (2025)
von: Dangi, Shalini, et al.
Veröffentlicht: (2025)
Expressive yet Efficient Feature Expansion with Adaptive Cross-Hadamard Products
von: Zhang, Xuyang, et al.
Veröffentlicht: (2025)
von: Zhang, Xuyang, et al.
Veröffentlicht: (2025)
Symmetric masking strategy enhances the performance of Masked Image Modeling
von: Nguyen, Khanh-Binh, et al.
Veröffentlicht: (2024)
von: Nguyen, Khanh-Binh, et al.
Veröffentlicht: (2024)
Semantics-enhanced Cross-modal Masked Image Modeling for Vision-Language Pre-training
von: Liu, Haowei, et al.
Veröffentlicht: (2024)
von: Liu, Haowei, et al.
Veröffentlicht: (2024)
A vision transformer-based framework for knowledge transfer from multi-modal to mono-modal lymphoma subtyping models
von: Guetarni, Bilel, et al.
Veröffentlicht: (2023)
von: Guetarni, Bilel, et al.
Veröffentlicht: (2023)
A survey of synthetic data augmentation methods in computer vision
von: Mumuni, Alhassan, et al.
Veröffentlicht: (2024)
von: Mumuni, Alhassan, et al.
Veröffentlicht: (2024)
Decentralized LoRA augmented transformer with multi-scale feature learning for secured eye diagnosis
von: Borno, Md. Naimur Asif, et al.
Veröffentlicht: (2025)
von: Borno, Md. Naimur Asif, et al.
Veröffentlicht: (2025)
Automating construction safety inspections using a multi-modal vision-language RAG framework
von: Wang, Chenxin, et al.
Veröffentlicht: (2025)
von: Wang, Chenxin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SJTU:Spatial judgments in multimodal models towards unified segmentation through coordinate detection
von: Chae, Joongwon, et al.
Veröffentlicht: (2024) -
Memory-SAM: Human-Prompt-Free Tongue Segmentation via Retrieval-to-Prompt
von: Chae, Joongwon, et al.
Veröffentlicht: (2025) -
GCR: Geometry-Consistent Routing for Task-Agnostic Continual Anomaly Detection
von: Chae, Joongwon, et al.
Veröffentlicht: (2026) -
StructCore: Structure-Aware Image-Level Scoring for Training-Free Unsupervised Anomaly Detection
von: Chae, Joongwon, et al.
Veröffentlicht: (2026) -
LLaVAction: evaluating and training multi-modal large language models for action understanding
von: Qi, Haozhe, et al.
Veröffentlicht: (2025)