Grid-augmented vision: A simple yet effective approach for enhanced spatial understanding in multi-modal agents
Fuente:
arXiv
Guardado en:
| Autores principales: | Chae, Joongwon, Wang, Zhenyu, Zhang, Lian, Yu, Dongmei, Qin, Peiwu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SJTU:Spatial judgments in multimodal models towards unified segmentation through coordinate detection
por: Chae, Joongwon, et al.
Publicado: (2024)
por: Chae, Joongwon, et al.
Publicado: (2024)
Memory-SAM: Human-Prompt-Free Tongue Segmentation via Retrieval-to-Prompt
por: Chae, Joongwon, et al.
Publicado: (2025)
por: Chae, Joongwon, et al.
Publicado: (2025)
GCR: Geometry-Consistent Routing for Task-Agnostic Continual Anomaly Detection
por: Chae, Joongwon, et al.
Publicado: (2026)
por: Chae, Joongwon, et al.
Publicado: (2026)
StructCore: Structure-Aware Image-Level Scoring for Training-Free Unsupervised Anomaly Detection
por: Chae, Joongwon, et al.
Publicado: (2026)
por: Chae, Joongwon, et al.
Publicado: (2026)
LLaVAction: evaluating and training multi-modal large language models for action understanding
por: Qi, Haozhe, et al.
Publicado: (2025)
por: Qi, Haozhe, et al.
Publicado: (2025)
UAV traffic scene understanding: A regulation embedded multi-modal network and a unified benchmark
por: Zhang, Yu, et al.
Publicado: (2026)
por: Zhang, Yu, et al.
Publicado: (2026)
A multi-modal vision-language model for generalizable annotation-free pathology localization
por: Yang, Hao, et al.
Publicado: (2024)
por: Yang, Hao, et al.
Publicado: (2024)
DBF-UNet: A Two-Stage Framework for Carotid Artery Segmentation with Pseudo-Label Generation
por: Li, Haoxuan, et al.
Publicado: (2025)
por: Li, Haoxuan, et al.
Publicado: (2025)
Have we unified image generation and understanding yet? An empirical study of GPT-4o's image generation ability
por: Li, Ning, et al.
Publicado: (2025)
por: Li, Ning, et al.
Publicado: (2025)
Near, far: Patch-ordering enhances vision foundation models' scene understanding
por: Pariza, Valentinos, et al.
Publicado: (2024)
por: Pariza, Valentinos, et al.
Publicado: (2024)
IRFundusSet: An Integrated Retinal Fundus Dataset with a Harmonized Healthy Label
por: Githinji, P. Bilha, et al.
Publicado: (2024)
por: Githinji, P. Bilha, et al.
Publicado: (2024)
Task Alignment: A simple and effective proxy for model merging in computer vision
por: de Jorge, Pau, et al.
Publicado: (2026)
por: de Jorge, Pau, et al.
Publicado: (2026)
STONE: Pioneering the One-to-N Universal Backdoor Threat in 3D Point Cloud
por: Shan, Dongmei, et al.
Publicado: (2025)
por: Shan, Dongmei, et al.
Publicado: (2025)
SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion
por: Nassar, Ahmed, et al.
Publicado: (2025)
por: Nassar, Ahmed, et al.
Publicado: (2025)
Multi-label classification for multi-temporal, multi-spatial coral reef condition monitoring using vision foundation model with adapter learning
por: Shao, Xinlei, et al.
Publicado: (2025)
por: Shao, Xinlei, et al.
Publicado: (2025)
Cognitive resilience: Unraveling the proficiency of image-captioning models to interpret masked visual content
por: Du, Zhicheng, et al.
Publicado: (2024)
por: Du, Zhicheng, et al.
Publicado: (2024)
SHAP-CAT: A interpretable multi-modal framework enhancing WSI classification via virtual staining and shapley-value-based multimodal fusion
por: Wang, Jun, et al.
Publicado: (2024)
por: Wang, Jun, et al.
Publicado: (2024)
Do large language vision models understand 3D shapes?
por: Eppel, Sagi
Publicado: (2024)
por: Eppel, Sagi
Publicado: (2024)
Serial fusion of multi-modal biometric systems
por: Marcialis, Gian Luca, et al.
Publicado: (2024)
por: Marcialis, Gian Luca, et al.
Publicado: (2024)
Synthetic data augmentation for robotic mobility aids to support blind and low vision people
por: Hwang, Hochul, et al.
Publicado: (2024)
por: Hwang, Hochul, et al.
Publicado: (2024)
Post-hurricane building damage assessment using street-view imagery and structured data: A multi-modal deep learning approach
por: Xue, Zhuoqun, et al.
Publicado: (2024)
por: Xue, Zhuoqun, et al.
Publicado: (2024)
A vision-based framework for human behavior understanding in industrial assembly lines
por: Papoutsakis, Konstantinos, et al.
Publicado: (2024)
por: Papoutsakis, Konstantinos, et al.
Publicado: (2024)
CKDA: Cross-modality Knowledge Disentanglement and Alignment for Visible-Infrared Lifelong Person Re-identification
por: Cui, Zhenyu, et al.
Publicado: (2025)
por: Cui, Zhenyu, et al.
Publicado: (2025)
Unsupervised Multi-agent and Single-agent Perception from Cooperative Views
por: Yang, Haochen, et al.
Publicado: (2026)
por: Yang, Haochen, et al.
Publicado: (2026)
MIRAGE: Robust multi-modal architectures translate fMRI-to-image models from vision to mental imagery
por: Kneeland, Reese, et al.
Publicado: (2026)
por: Kneeland, Reese, et al.
Publicado: (2026)
Cross-modal ultra-scale learning with tri-modalities of renal biopsy images for glomerular multi-disease auxiliary diagnosis
por: Long, Kaixing, et al.
Publicado: (2025)
por: Long, Kaixing, et al.
Publicado: (2025)
HiERO: understanding the hierarchy of human behavior enhances reasoning on egocentric videos
por: Peirone, Simone Alberto, et al.
Publicado: (2025)
por: Peirone, Simone Alberto, et al.
Publicado: (2025)
HCMA-UNet: A Hybrid CNN-Mamba UNet with Axial Self-Attention for Efficient Breast Cancer Segmentation
por: Li, Haoxuan, et al.
Publicado: (2025)
por: Li, Haoxuan, et al.
Publicado: (2025)
External Prompt Features Enhanced Parameter-efficient Fine-tuning for Salient Object Detection
por: Liang, Wen, et al.
Publicado: (2024)
por: Liang, Wen, et al.
Publicado: (2024)
A data-centric approach to class-specific bias in image data augmentation
por: Angelakis, Athanasios, et al.
Publicado: (2024)
por: Angelakis, Athanasios, et al.
Publicado: (2024)
Unified modality separation: A vision-language framework for unsupervised domain adaptation
por: Li, Xinyao, et al.
Publicado: (2025)
por: Li, Xinyao, et al.
Publicado: (2025)
bi-modal textual prompt learning for vision-language models in remote sensing
por: Kashyap, Pankhi, et al.
Publicado: (2026)
por: Kashyap, Pankhi, et al.
Publicado: (2026)
A multi-temporal multi-spectral attention-augmented deep convolution neural network with contrastive learning for crop yield prediction
por: Dangi, Shalini, et al.
Publicado: (2025)
por: Dangi, Shalini, et al.
Publicado: (2025)
Expressive yet Efficient Feature Expansion with Adaptive Cross-Hadamard Products
por: Zhang, Xuyang, et al.
Publicado: (2025)
por: Zhang, Xuyang, et al.
Publicado: (2025)
Symmetric masking strategy enhances the performance of Masked Image Modeling
por: Nguyen, Khanh-Binh, et al.
Publicado: (2024)
por: Nguyen, Khanh-Binh, et al.
Publicado: (2024)
Semantics-enhanced Cross-modal Masked Image Modeling for Vision-Language Pre-training
por: Liu, Haowei, et al.
Publicado: (2024)
por: Liu, Haowei, et al.
Publicado: (2024)
A vision transformer-based framework for knowledge transfer from multi-modal to mono-modal lymphoma subtyping models
por: Guetarni, Bilel, et al.
Publicado: (2023)
por: Guetarni, Bilel, et al.
Publicado: (2023)
A survey of synthetic data augmentation methods in computer vision
por: Mumuni, Alhassan, et al.
Publicado: (2024)
por: Mumuni, Alhassan, et al.
Publicado: (2024)
Decentralized LoRA augmented transformer with multi-scale feature learning for secured eye diagnosis
por: Borno, Md. Naimur Asif, et al.
Publicado: (2025)
por: Borno, Md. Naimur Asif, et al.
Publicado: (2025)
Automating construction safety inspections using a multi-modal vision-language RAG framework
por: Wang, Chenxin, et al.
Publicado: (2025)
por: Wang, Chenxin, et al.
Publicado: (2025)
Ejemplares similares
-
SJTU:Spatial judgments in multimodal models towards unified segmentation through coordinate detection
por: Chae, Joongwon, et al.
Publicado: (2024) -
Memory-SAM: Human-Prompt-Free Tongue Segmentation via Retrieval-to-Prompt
por: Chae, Joongwon, et al.
Publicado: (2025) -
GCR: Geometry-Consistent Routing for Task-Agnostic Continual Anomaly Detection
por: Chae, Joongwon, et al.
Publicado: (2026) -
StructCore: Structure-Aware Image-Level Scoring for Training-Free Unsupervised Anomaly Detection
por: Chae, Joongwon, et al.
Publicado: (2026) -
LLaVAction: evaluating and training multi-modal large language models for action understanding
por: Qi, Haozhe, et al.
Publicado: (2025)