Salvato in:
| Autore principale: | Borji, Ali |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2407.01782 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Qualitative Failures of Image Generation Models and Their Application in Detecting Deepfakes
di: Borji, Ali
Pubblicazione: (2023)
di: Borji, Ali
Pubblicazione: (2023)
A deep learning pipeline for PAM50 subtype classification using histopathology images and multi-objective patch selection
di: Borji, Arezoo, et al.
Pubblicazione: (2026)
di: Borji, Arezoo, et al.
Pubblicazione: (2026)
A recurrent vision transformer shows signatures of primate visual attention
di: Morgan, Jonathan, et al.
Pubblicazione: (2025)
di: Morgan, Jonathan, et al.
Pubblicazione: (2025)
A comprehensive overview of deep learning models for object detection from videos/images
di: Zulfqar, Sukana, et al.
Pubblicazione: (2026)
di: Zulfqar, Sukana, et al.
Pubblicazione: (2026)
Advanced Hybrid Deep Learning Model for Enhanced Classification of Osteosarcoma Histopathology Images
di: Borji, Arezoo, et al.
Pubblicazione: (2024)
di: Borji, Arezoo, et al.
Pubblicazione: (2024)
Are vision language models robust to uncertain inputs?
di: Wang, Xi, et al.
Pubblicazione: (2025)
di: Wang, Xi, et al.
Pubblicazione: (2025)
Towards aligned body representations in vision models
di: Gizdov, Andrey, et al.
Pubblicazione: (2025)
di: Gizdov, Andrey, et al.
Pubblicazione: (2025)
ISLA: A U-Net for MRI-based acute ischemic stroke lesion segmentation with deep supervision, attention, domain adaptation, and ensemble learning
di: Roca, Vincent, et al.
Pubblicazione: (2026)
di: Roca, Vincent, et al.
Pubblicazione: (2026)
Bringing together invertible UNets with invertible attention modules for memory-efficient diffusion models
di: Jain, Karan, et al.
Pubblicazione: (2025)
di: Jain, Karan, et al.
Pubblicazione: (2025)
What matters when building vision-language models?
di: Laurençon, Hugo, et al.
Pubblicazione: (2024)
di: Laurençon, Hugo, et al.
Pubblicazione: (2024)
Quantifying the human visual exposome with vision language models
di: Rominger, Christian, et al.
Pubblicazione: (2026)
di: Rominger, Christian, et al.
Pubblicazione: (2026)
Interpreting vision transformers via residual replacement model
di: Kim, Jinyeong, et al.
Pubblicazione: (2025)
di: Kim, Jinyeong, et al.
Pubblicazione: (2025)
Vision language models are unreliable at trivial spatial cognition
di: Khemlani, Sangeet, et al.
Pubblicazione: (2025)
di: Khemlani, Sangeet, et al.
Pubblicazione: (2025)
Analyzing mixed construction and demolition waste in material recovery facilities: evolution, challenges, and applications of computer vision and deep learning
di: Langley, Adrian, et al.
Pubblicazione: (2024)
di: Langley, Adrian, et al.
Pubblicazione: (2024)
Thinker: A vision-language foundation model for embodied intelligence
di: Pan, Baiyu, et al.
Pubblicazione: (2026)
di: Pan, Baiyu, et al.
Pubblicazione: (2026)
A multimodal vision foundation model for generalizable knee pathology
di: Yu, Kang, et al.
Pubblicazione: (2026)
di: Yu, Kang, et al.
Pubblicazione: (2026)
Self-supervised vision-langage alignment of deep learning representations for bone X-rays analysis
di: Englebert, Alexandre, et al.
Pubblicazione: (2024)
di: Englebert, Alexandre, et al.
Pubblicazione: (2024)
Building and better understanding vision-language models: insights and future directions
di: Laurençon, Hugo, et al.
Pubblicazione: (2024)
di: Laurençon, Hugo, et al.
Pubblicazione: (2024)
Hallucination-aware intermediate representation edit in large vision-language models
di: Suo, Wei, et al.
Pubblicazione: (2026)
di: Suo, Wei, et al.
Pubblicazione: (2026)
Generalizing vision-language models to novel domains: A comprehensive survey
di: Li, Xinyao, et al.
Pubblicazione: (2025)
di: Li, Xinyao, et al.
Pubblicazione: (2025)
SSTFB: Leveraging self-supervised pretext learning and temporal self-attention with feature branching for real-time video polyp segmentation
di: Xu, Ziang, et al.
Pubblicazione: (2024)
di: Xu, Ziang, et al.
Pubblicazione: (2024)
Near, far: Patch-ordering enhances vision foundation models' scene understanding
di: Pariza, Valentinos, et al.
Pubblicazione: (2024)
di: Pariza, Valentinos, et al.
Pubblicazione: (2024)
Beyond the Hype: A dispassionate look at vision-language models in medical scenario
di: Nan, Yang, et al.
Pubblicazione: (2024)
di: Nan, Yang, et al.
Pubblicazione: (2024)
A benchmark multimodal oro-dental dataset for large vision-language models
di: Lv, Haoxin, et al.
Pubblicazione: (2025)
di: Lv, Haoxin, et al.
Pubblicazione: (2025)
Representation geometry shapes task performance in vision-language modeling for CT enterography
di: Minoccheri, Cristian, et al.
Pubblicazione: (2026)
di: Minoccheri, Cristian, et al.
Pubblicazione: (2026)
Relation Learning and Aggregate-attention for Multi-person Motion Prediction
di: Qu, Kehua, et al.
Pubblicazione: (2024)
di: Qu, Kehua, et al.
Pubblicazione: (2024)
GAC-Net_Geometric and attention-based Network for Depth Completion
di: Zhu, Kuang, et al.
Pubblicazione: (2025)
di: Zhu, Kuang, et al.
Pubblicazione: (2025)
Incorporating simulated spatial context information improves the effectiveness of contrastive learning models
di: Zhu, Lizhen, et al.
Pubblicazione: (2024)
di: Zhu, Lizhen, et al.
Pubblicazione: (2024)
GCAM: Gaussian and causal-attention model of food fine-grained recognition
di: Zhuang, Guohang, et al.
Pubblicazione: (2024)
di: Zhuang, Guohang, et al.
Pubblicazione: (2024)
Computer vision-based model for detecting turning lane features on Florida's public roadways
di: Antwi, Richard Boadu, et al.
Pubblicazione: (2024)
di: Antwi, Richard Boadu, et al.
Pubblicazione: (2024)
VLA-Mark: A cross modal watermark for large vision-language alignment model
di: Liu, Shuliang, et al.
Pubblicazione: (2025)
di: Liu, Shuliang, et al.
Pubblicazione: (2025)
Self-adaptive vision-language model for 3D segmentation of pulmonary artery and vein
di: Guo, Xiaotong, et al.
Pubblicazione: (2025)
di: Guo, Xiaotong, et al.
Pubblicazione: (2025)
RadEdit: stress-testing biomedical vision models via diffusion image editing
di: Pérez-García, Fernando, et al.
Pubblicazione: (2023)
di: Pérez-García, Fernando, et al.
Pubblicazione: (2023)
AddressVLM: Cross-view Alignment Tuning for Image Address Localization using Large Vision-Language Models
di: Xu, Shixiong, et al.
Pubblicazione: (2025)
di: Xu, Shixiong, et al.
Pubblicazione: (2025)
Hand-object reconstruction via interaction-aware graph attention mechanism
di: Woo, Taeyun, et al.
Pubblicazione: (2024)
di: Woo, Taeyun, et al.
Pubblicazione: (2024)
QUEST: A robust attention formulation using query-modulated spherical attention
di: Govindarajan, Hariprasath, et al.
Pubblicazione: (2026)
di: Govindarajan, Hariprasath, et al.
Pubblicazione: (2026)
Multi-modal user interface control detection using cross-attention
di: Moradi, Milad, et al.
Pubblicazione: (2026)
di: Moradi, Milad, et al.
Pubblicazione: (2026)
Introducing an ensemble method for the early detection of Alzheimer's disease through the analysis of PET scan images
di: Borji, Arezoo, et al.
Pubblicazione: (2024)
di: Borji, Arezoo, et al.
Pubblicazione: (2024)
Visual Structures Helps Visual Reasoning: Addressing the Binding Problem in VLMs
di: Izadi, Amirmohammad, et al.
Pubblicazione: (2025)
di: Izadi, Amirmohammad, et al.
Pubblicazione: (2025)
A computer vision-based model for occupancy detection using low-resolution thermal images
di: Cui, Xue, et al.
Pubblicazione: (2025)
di: Cui, Xue, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Qualitative Failures of Image Generation Models and Their Application in Detecting Deepfakes
di: Borji, Ali
Pubblicazione: (2023) -
A deep learning pipeline for PAM50 subtype classification using histopathology images and multi-objective patch selection
di: Borji, Arezoo, et al.
Pubblicazione: (2026) -
A recurrent vision transformer shows signatures of primate visual attention
di: Morgan, Jonathan, et al.
Pubblicazione: (2025) -
A comprehensive overview of deep learning models for object detection from videos/images
di: Zulfqar, Sukana, et al.
Pubblicazione: (2026) -
Advanced Hybrid Deep Learning Model for Enhanced Classification of Osteosarcoma Histopathology Images
di: Borji, Arezoo, et al.
Pubblicazione: (2024)