Near, far: Patch-ordering enhances vision foundation models' scene understanding
Fuente:
arXiv
Salvato in:
| Autori principali: | Pariza, Valentinos, Salehi, Mohammadreza, Burghouts, Gertjan, Locatello, Francesco, Asano, Yuki M. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Better Language Models Exhibit Higher Visual Alignment
di: Ruthardt, Jona, et al.
Pubblicazione: (2024)
di: Ruthardt, Jona, et al.
Pubblicazione: (2024)
Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
di: Venkataramanan, Shashanka, et al.
Pubblicazione: (2025)
di: Venkataramanan, Shashanka, et al.
Pubblicazione: (2025)
Neurosymbolic Inference On Foundation Models For Remote Sensing Text-to-image Retrieval With Complex Queries
di: Mezzi, Emanuele, et al.
Pubblicazione: (2025)
di: Mezzi, Emanuele, et al.
Pubblicazione: (2025)
Generating metamers of human scene understanding
di: Raina, Ritik, et al.
Pubblicazione: (2026)
di: Raina, Ritik, et al.
Pubblicazione: (2026)
TasselNetV4: A vision foundation model for cross-scene, cross-scale, and cross-species plant counting
di: Hu, Xiaonan, et al.
Pubblicazione: (2025)
di: Hu, Xiaonan, et al.
Pubblicazione: (2025)
SelEx: Self-Expertise in Fine-Grained Generalized Category Discovery
di: Rastegar, Sarah, et al.
Pubblicazione: (2024)
di: Rastegar, Sarah, et al.
Pubblicazione: (2024)
Adaptive Prompt Tuning: Vision Guided Prompt Tuning with Cross-Attention for Fine-Grained Few-Shot Learning
di: Brouwer, Eric, et al.
Pubblicazione: (2024)
di: Brouwer, Eric, et al.
Pubblicazione: (2024)
Thinker: A vision-language foundation model for embodied intelligence
di: Pan, Baiyu, et al.
Pubblicazione: (2026)
di: Pan, Baiyu, et al.
Pubblicazione: (2026)
A multimodal vision foundation model for generalizable knee pathology
di: Yu, Kang, et al.
Pubblicazione: (2026)
di: Yu, Kang, et al.
Pubblicazione: (2026)
Occlusion Robustness of CLIP for Military Vehicle Classification
di: van Woerden, Jan Erik, et al.
Pubblicazione: (2025)
di: van Woerden, Jan Erik, et al.
Pubblicazione: (2025)
Quantifying the synthetic and real domain gap in aerial scene understanding
di: Marcu, Alina
Pubblicazione: (2024)
di: Marcu, Alina
Pubblicazione: (2024)
UAV traffic scene understanding: A regulation embedded multi-modal network and a unified benchmark
di: Zhang, Yu, et al.
Pubblicazione: (2026)
di: Zhang, Yu, et al.
Pubblicazione: (2026)
Building and better understanding vision-language models: insights and future directions
di: Laurençon, Hugo, et al.
Pubblicazione: (2024)
di: Laurençon, Hugo, et al.
Pubblicazione: (2024)
Textual Inversion for Efficient Adaptation of Open-Vocabulary Object Detectors Without Forgetting
di: Ruis, Frank, et al.
Pubblicazione: (2025)
di: Ruis, Frank, et al.
Pubblicazione: (2025)
Self-Supervised Partial Cycle-Consistency for Multi-View Matching
di: Taggenbrock, Fedor, et al.
Pubblicazione: (2025)
di: Taggenbrock, Fedor, et al.
Pubblicazione: (2025)
Self-supervised Learning of Echocardiographic Video Representations via Online Cluster Distillation
di: Mishra, Divyanshu, et al.
Pubblicazione: (2025)
di: Mishra, Divyanshu, et al.
Pubblicazione: (2025)
Language-Based Augmentation to Address Shortcut Learning in Object Goal Navigation
di: Hoftijzer, Dennis, et al.
Pubblicazione: (2024)
di: Hoftijzer, Dennis, et al.
Pubblicazione: (2024)
Geospatial foundation models for image analysis: evaluating and enhancing NASA-IBM Prithvi's domain adaptability
di: Hsu, Chia-Yu, et al.
Pubblicazione: (2024)
di: Hsu, Chia-Yu, et al.
Pubblicazione: (2024)
Anticipating Future Object Compositions without Forgetting
di: Zahran, Youssef, et al.
Pubblicazione: (2024)
di: Zahran, Youssef, et al.
Pubblicazione: (2024)
DEF-oriCORN: efficient 3D scene understanding for robust language-directed manipulation without demonstrations
di: Son, Dongwon, et al.
Pubblicazione: (2024)
di: Son, Dongwon, et al.
Pubblicazione: (2024)
Assessing the generalization performance of SAM for ureteroscopy scene understanding
di: Villagrana, Martin, et al.
Pubblicazione: (2025)
di: Villagrana, Martin, et al.
Pubblicazione: (2025)
Patch-enhanced Mask Encoder Prompt Image Generation
di: Xu, Shusong, et al.
Pubblicazione: (2024)
di: Xu, Shusong, et al.
Pubblicazione: (2024)
Binding Dynamics in Rotating Features
di: Löwe, Sindy, et al.
Pubblicazione: (2024)
di: Löwe, Sindy, et al.
Pubblicazione: (2024)
Towards a vision foundation model for comprehensive assessment of Cardiac MRI
di: Jacob, Athira J, et al.
Pubblicazione: (2024)
di: Jacob, Athira J, et al.
Pubblicazione: (2024)
Imaging foundation model for universal enhancement of non-ideal measurement CT
di: Ge, Rongjun, et al.
Pubblicazione: (2024)
di: Ge, Rongjun, et al.
Pubblicazione: (2024)
Enabling clinical use of foundation models for computational pathology
di: Henriksen, Audun L, et al.
Pubblicazione: (2026)
di: Henriksen, Audun L, et al.
Pubblicazione: (2026)
Divided Attention: Unsupervised Multi-Object Discovery with Contextually Separated Slots
di: Lao, Dong, et al.
Pubblicazione: (2023)
di: Lao, Dong, et al.
Pubblicazione: (2023)
Steerable Visual Representations
di: Ruthardt, Jona, et al.
Pubblicazione: (2026)
di: Ruthardt, Jona, et al.
Pubblicazione: (2026)
An interpretable framework using foundation models for fish sex identification
di: Miao, Zheng, et al.
Pubblicazione: (2026)
di: Miao, Zheng, et al.
Pubblicazione: (2026)
Less yet robust: crucial region selection for scene recognition
di: Zhang, Jianqi, et al.
Pubblicazione: (2024)
di: Zhang, Jianqi, et al.
Pubblicazione: (2024)
From Pixels to Predicates Structuring urban perception with scene graphs
di: Liu, Yunlong, et al.
Pubblicazione: (2025)
di: Liu, Yunlong, et al.
Pubblicazione: (2025)
Prompting with the human-touch: evaluating model-sensitivity of foundation models for musculoskeletal CT segmentation
di: Magg, Caroline, et al.
Pubblicazione: (2026)
di: Magg, Caroline, et al.
Pubblicazione: (2026)
ActiveMark: on watermarking of visual foundation models via massive activations
di: Chistyakova, Anna, et al.
Pubblicazione: (2025)
di: Chistyakova, Anna, et al.
Pubblicazione: (2025)
MoireMix: A Formula-Based Data Augmentation for Improving Image Classification Robustness
di: Matsuo, Yuto, et al.
Pubblicazione: (2026)
di: Matsuo, Yuto, et al.
Pubblicazione: (2026)
Developing a foundation model for high-resolution remote sensing data of the Netherlands
di: Vermeeren, Paul, et al.
Pubblicazione: (2026)
di: Vermeeren, Paul, et al.
Pubblicazione: (2026)
Grounded Object Centric Learning
di: Kori, Avinash, et al.
Pubblicazione: (2023)
di: Kori, Avinash, et al.
Pubblicazione: (2023)
MoMBS: Mixed-order minibatch sampling enhances model training from diverse-quality images
di: Li, Han, et al.
Pubblicazione: (2025)
di: Li, Han, et al.
Pubblicazione: (2025)
OASIC: Occlusion-Agnostic and Severity-Informed Classification
di: Gijzen, Kay, et al.
Pubblicazione: (2026)
di: Gijzen, Kay, et al.
Pubblicazione: (2026)
Assessment of Sentinel-2 spatial and temporal coverage based on the scene classification layer
di: Sanchez, Cristhian, et al.
Pubblicazione: (2024)
di: Sanchez, Cristhian, et al.
Pubblicazione: (2024)
All-day Multi-scenes Lifelong Vision-and-Language Navigation with Tucker Adaptation
di: Wang, Xudong, et al.
Pubblicazione: (2026)
di: Wang, Xudong, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Better Language Models Exhibit Higher Visual Alignment
di: Ruthardt, Jona, et al.
Pubblicazione: (2024) -
Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
di: Venkataramanan, Shashanka, et al.
Pubblicazione: (2025) -
Neurosymbolic Inference On Foundation Models For Remote Sensing Text-to-image Retrieval With Complex Queries
di: Mezzi, Emanuele, et al.
Pubblicazione: (2025) -
Generating metamers of human scene understanding
di: Raina, Ritik, et al.
Pubblicazione: (2026) -
TasselNetV4: A vision foundation model for cross-scene, cross-scale, and cross-species plant counting
di: Hu, Xiaonan, et al.
Pubblicazione: (2025)