From Gaze to Insight: Bridging Human Visual Attention and Vision Language Model Explanation for Weakly-Supervised Medical Image Segmentation
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Jingkun, Duan, Haoran, Zhang, Xiao, Gao, Boyan, Grau, Vicente, Han, Jungong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
di: Raoufi, Behnam, et al.
Pubblicazione: (2025)
di: Raoufi, Behnam, et al.
Pubblicazione: (2025)
VLM-NCD:Novel Class Discovery with Vision-Based Large Language Models
di: Su, Yuetong, et al.
Pubblicazione: (2025)
di: Su, Yuetong, et al.
Pubblicazione: (2025)
Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
di: Zinnen, Mathias, et al.
Pubblicazione: (2025)
di: Zinnen, Mathias, et al.
Pubblicazione: (2025)
Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models
di: Gautam, Sushant, et al.
Pubblicazione: (2025)
di: Gautam, Sushant, et al.
Pubblicazione: (2025)
VDPP: Video Depth Post-Processing for Speed and Scalability
di: Yoon, Daewon, et al.
Pubblicazione: (2026)
di: Yoon, Daewon, et al.
Pubblicazione: (2026)
Semantic2Graph: Graph-based Multi-modal Feature Fusion for Action Segmentation in Videos
di: Zhang, Junbin, et al.
Pubblicazione: (2022)
di: Zhang, Junbin, et al.
Pubblicazione: (2022)
HEDGE: Hallucination Estimation via Dense Geometric Entropy for VQA with Vision-Language Models
di: Gautam, Sushant, et al.
Pubblicazione: (2025)
di: Gautam, Sushant, et al.
Pubblicazione: (2025)
Rethinking Visual Intelligence: Insights from Video Pretraining
di: Acuaviva, Pablo, et al.
Pubblicazione: (2025)
di: Acuaviva, Pablo, et al.
Pubblicazione: (2025)
Beyond RNNs: Benchmarking Attention-Based Image Captioning Models
di: Yanambakkam, Hemanth Teja, et al.
Pubblicazione: (2025)
di: Yanambakkam, Hemanth Teja, et al.
Pubblicazione: (2025)
OpenFusion++: An Open-vocabulary Real-time Scene Understanding System
di: Jin, Xiaofeng, et al.
Pubblicazione: (2025)
di: Jin, Xiaofeng, et al.
Pubblicazione: (2025)
DSER: Spectral Epipolar Representation for Efficient Light Field Depth Estimation
di: Mohammad, Noor Islam S., et al.
Pubblicazione: (2025)
di: Mohammad, Noor Islam S., et al.
Pubblicazione: (2025)
Hierarchical Spatial Algorithms for High-Resolution Image Quantization and Feature Extraction
di: Mohammad, Noor Islam S.
Pubblicazione: (2025)
di: Mohammad, Noor Islam S.
Pubblicazione: (2025)
Sat-JEPA-Diff: Bridging Self-Supervised Learning and Generative Diffusion for Remote Sensing
di: Komurcu, Kursat, et al.
Pubblicazione: (2026)
di: Komurcu, Kursat, et al.
Pubblicazione: (2026)
3DreamBooth: High-Fidelity 3D Subject-Driven Video Generation Model
di: Ko, Hyun-kyu, et al.
Pubblicazione: (2026)
di: Ko, Hyun-kyu, et al.
Pubblicazione: (2026)
Progressive Cross Attention Network for Flood Segmentation using Multispectral Satellite Imagery
di: Feliren, Vicky, et al.
Pubblicazione: (2025)
di: Feliren, Vicky, et al.
Pubblicazione: (2025)
VisChainBench: A Benchmark for Multi-Turn, Multi-Image Visual Reasoning Beyond Language Priors
di: Lyu, Wenbo, et al.
Pubblicazione: (2025)
di: Lyu, Wenbo, et al.
Pubblicazione: (2025)
TACIT Benchmark: A Programmatic Visual Reasoning Benchmark for Generative and Discriminative Models
di: Medeiros, Daniel Nobrega
Pubblicazione: (2026)
di: Medeiros, Daniel Nobrega
Pubblicazione: (2026)
μ-Net: A Deep Learning-Based Architecture for μ-CT Segmentation
di: Bruno, Pierangela, et al.
Pubblicazione: (2024)
di: Bruno, Pierangela, et al.
Pubblicazione: (2024)
Predictive Modeling of Maritime Radar Data Using Transformer Architecture
di: Qesaraku, Bjorna, et al.
Pubblicazione: (2025)
di: Qesaraku, Bjorna, et al.
Pubblicazione: (2025)
MATEX: Multi-scale Attention and Text-guided Explainability of Medical Vision-Language Models
di: Imran, Muhammad, et al.
Pubblicazione: (2026)
di: Imran, Muhammad, et al.
Pubblicazione: (2026)
Do Generative Metrics Predict YOLO Performance? An Evaluation Across Models, Augmentation Ratios, and Dataset Complexity
di: Marian, Vasile, et al.
Pubblicazione: (2026)
di: Marian, Vasile, et al.
Pubblicazione: (2026)
Short-Window Sliding Learning for Real-Time Violence Detection via LLM-based Auto-Labeling
di: Jung, Seoik, et al.
Pubblicazione: (2025)
di: Jung, Seoik, et al.
Pubblicazione: (2025)
Real Time Human Detection by Unmanned Aerial Vehicles
di: Guettala, Walid, et al.
Pubblicazione: (2024)
di: Guettala, Walid, et al.
Pubblicazione: (2024)
A large-scale, physically-based synthetic dataset for satellite pose estimation
di: Velkei, Szabolcs, et al.
Pubblicazione: (2025)
di: Velkei, Szabolcs, et al.
Pubblicazione: (2025)
LRCP: Low-Rank Compressibility Guided Visual Token Pruning for Efficient LVLMs
di: Lu, Hongyu, et al.
Pubblicazione: (2026)
di: Lu, Hongyu, et al.
Pubblicazione: (2026)
Shaded Route Planning Using Active Segmentation and Identification of Satellite Images
di: Da, Longchao, et al.
Pubblicazione: (2024)
di: Da, Longchao, et al.
Pubblicazione: (2024)
Can Local Vision-Language Models improve Activity Recognition over Vision Transformers? -- Case Study on Newborn Resuscitation
di: Guerriero, Enrico, et al.
Pubblicazione: (2026)
di: Guerriero, Enrico, et al.
Pubblicazione: (2026)
IMKD: Intensity-Aware Multi-Level Knowledge Distillation for Camera-Radar Fusion
di: Mishra, Shashank, et al.
Pubblicazione: (2025)
di: Mishra, Shashank, et al.
Pubblicazione: (2025)
UVLM: A Universal Vision-Language Model Loader for Reproducible Multimodal Benchmarking
di: Perez, Joan, et al.
Pubblicazione: (2026)
di: Perez, Joan, et al.
Pubblicazione: (2026)
A Multi-Camera Vision-Based Approach for Fine-Grained Assembly Quality Control
di: Nazeri, Ali, et al.
Pubblicazione: (2025)
di: Nazeri, Ali, et al.
Pubblicazione: (2025)
Image-based Facial Rig Inversion
di: Yang, Tianxiang, et al.
Pubblicazione: (2025)
di: Yang, Tianxiang, et al.
Pubblicazione: (2025)
Experimental Evaluation of Road-Crossing Decisions by Autonomous Wheelchairs against Environmental Factors
di: Corradini, Franca, et al.
Pubblicazione: (2024)
di: Corradini, Franca, et al.
Pubblicazione: (2024)
Polarization-Based Eye Tracking with Personalized Siamese Architectures
di: Kalkanli, Beyza, et al.
Pubblicazione: (2026)
di: Kalkanli, Beyza, et al.
Pubblicazione: (2026)
ShapBPT: Image Feature Attributions Using Data-Aware Binary Partition Trees
di: Rashid, Muhammad, et al.
Pubblicazione: (2026)
di: Rashid, Muhammad, et al.
Pubblicazione: (2026)
See What You Need: Query-Aware Visual Intelligence through Reasoning-Perception Loops
di: Dong, Zixuan, et al.
Pubblicazione: (2025)
di: Dong, Zixuan, et al.
Pubblicazione: (2025)
Akasha 2: Hamiltonian State Space Duality and Visual-Language Joint Embedding Predictive Architectur
di: Meziani, Yani
Pubblicazione: (2026)
di: Meziani, Yani
Pubblicazione: (2026)
Safe Road-Crossing by Autonomous Wheelchairs: a Novel Dataset and its Experimental Evaluation
di: Grigioni, Carlo, et al.
Pubblicazione: (2024)
di: Grigioni, Carlo, et al.
Pubblicazione: (2024)
Cooperative Perception: A Resource-Efficient Framework for Multi-Drone 3D Scene Reconstruction Using Federated Diffusion and NeRF
di: Pourmandi, Massoud
Pubblicazione: (2025)
di: Pourmandi, Massoud
Pubblicazione: (2025)
GazeD: Context-Aware Diffusion for Accurate 3D Gaze Estimation
di: Catalini, Riccardo, et al.
Pubblicazione: (2026)
di: Catalini, Riccardo, et al.
Pubblicazione: (2026)
WildfireVLM: AI-powered Analysis for Early Wildfire Detection and Risk Assessment Using Satellite Imagery
di: Ayanzadeh, Aydin, et al.
Pubblicazione: (2026)
di: Ayanzadeh, Aydin, et al.
Pubblicazione: (2026)
Documenti analoghi
-
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
di: Raoufi, Behnam, et al.
Pubblicazione: (2025) -
VLM-NCD:Novel Class Discovery with Vision-Based Large Language Models
di: Su, Yuetong, et al.
Pubblicazione: (2025) -
Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
di: Zinnen, Mathias, et al.
Pubblicazione: (2025) -
Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models
di: Gautam, Sushant, et al.
Pubblicazione: (2025) -
VDPP: Video Depth Post-Processing for Speed and Scalability
di: Yoon, Daewon, et al.
Pubblicazione: (2026)