Gespeichert in:
| Hauptverfasser: | Pariza, Valentinos, Salehi, Mohammadreza, Burghouts, Gertjan, Locatello, Francesco, Asano, Yuki M. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2408.11054 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Better Language Models Exhibit Higher Visual Alignment
von: Ruthardt, Jona, et al.
Veröffentlicht: (2024)
von: Ruthardt, Jona, et al.
Veröffentlicht: (2024)
Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
von: Venkataramanan, Shashanka, et al.
Veröffentlicht: (2025)
von: Venkataramanan, Shashanka, et al.
Veröffentlicht: (2025)
Neurosymbolic Inference On Foundation Models For Remote Sensing Text-to-image Retrieval With Complex Queries
von: Mezzi, Emanuele, et al.
Veröffentlicht: (2025)
von: Mezzi, Emanuele, et al.
Veröffentlicht: (2025)
SelEx: Self-Expertise in Fine-Grained Generalized Category Discovery
von: Rastegar, Sarah, et al.
Veröffentlicht: (2024)
von: Rastegar, Sarah, et al.
Veröffentlicht: (2024)
Generating metamers of human scene understanding
von: Raina, Ritik, et al.
Veröffentlicht: (2026)
von: Raina, Ritik, et al.
Veröffentlicht: (2026)
TasselNetV4: A vision foundation model for cross-scene, cross-scale, and cross-species plant counting
von: Hu, Xiaonan, et al.
Veröffentlicht: (2025)
von: Hu, Xiaonan, et al.
Veröffentlicht: (2025)
Adaptive Prompt Tuning: Vision Guided Prompt Tuning with Cross-Attention for Fine-Grained Few-Shot Learning
von: Brouwer, Eric, et al.
Veröffentlicht: (2024)
von: Brouwer, Eric, et al.
Veröffentlicht: (2024)
Occlusion Robustness of CLIP for Military Vehicle Classification
von: van Woerden, Jan Erik, et al.
Veröffentlicht: (2025)
von: van Woerden, Jan Erik, et al.
Veröffentlicht: (2025)
Self-supervised Learning of Echocardiographic Video Representations via Online Cluster Distillation
von: Mishra, Divyanshu, et al.
Veröffentlicht: (2025)
von: Mishra, Divyanshu, et al.
Veröffentlicht: (2025)
Thinker: A vision-language foundation model for embodied intelligence
von: Pan, Baiyu, et al.
Veröffentlicht: (2026)
von: Pan, Baiyu, et al.
Veröffentlicht: (2026)
A multimodal vision foundation model for generalizable knee pathology
von: Yu, Kang, et al.
Veröffentlicht: (2026)
von: Yu, Kang, et al.
Veröffentlicht: (2026)
Quantifying the synthetic and real domain gap in aerial scene understanding
von: Marcu, Alina
Veröffentlicht: (2024)
von: Marcu, Alina
Veröffentlicht: (2024)
Textual Inversion for Efficient Adaptation of Open-Vocabulary Object Detectors Without Forgetting
von: Ruis, Frank, et al.
Veröffentlicht: (2025)
von: Ruis, Frank, et al.
Veröffentlicht: (2025)
Self-Supervised Partial Cycle-Consistency for Multi-View Matching
von: Taggenbrock, Fedor, et al.
Veröffentlicht: (2025)
von: Taggenbrock, Fedor, et al.
Veröffentlicht: (2025)
Language-Based Augmentation to Address Shortcut Learning in Object Goal Navigation
von: Hoftijzer, Dennis, et al.
Veröffentlicht: (2024)
von: Hoftijzer, Dennis, et al.
Veröffentlicht: (2024)
Building and better understanding vision-language models: insights and future directions
von: Laurençon, Hugo, et al.
Veröffentlicht: (2024)
von: Laurençon, Hugo, et al.
Veröffentlicht: (2024)
UAV traffic scene understanding: A regulation embedded multi-modal network and a unified benchmark
von: Zhang, Yu, et al.
Veröffentlicht: (2026)
von: Zhang, Yu, et al.
Veröffentlicht: (2026)
Anticipating Future Object Compositions without Forgetting
von: Zahran, Youssef, et al.
Veröffentlicht: (2024)
von: Zahran, Youssef, et al.
Veröffentlicht: (2024)
Assessing the generalization performance of SAM for ureteroscopy scene understanding
von: Villagrana, Martin, et al.
Veröffentlicht: (2025)
von: Villagrana, Martin, et al.
Veröffentlicht: (2025)
Binding Dynamics in Rotating Features
von: Löwe, Sindy, et al.
Veröffentlicht: (2024)
von: Löwe, Sindy, et al.
Veröffentlicht: (2024)
DEF-oriCORN: efficient 3D scene understanding for robust language-directed manipulation without demonstrations
von: Son, Dongwon, et al.
Veröffentlicht: (2024)
von: Son, Dongwon, et al.
Veröffentlicht: (2024)
Geospatial foundation models for image analysis: evaluating and enhancing NASA-IBM Prithvi's domain adaptability
von: Hsu, Chia-Yu, et al.
Veröffentlicht: (2024)
von: Hsu, Chia-Yu, et al.
Veröffentlicht: (2024)
Towards a vision foundation model for comprehensive assessment of Cardiac MRI
von: Jacob, Athira J, et al.
Veröffentlicht: (2024)
von: Jacob, Athira J, et al.
Veröffentlicht: (2024)
Divided Attention: Unsupervised Multi-Object Discovery with Contextually Separated Slots
von: Lao, Dong, et al.
Veröffentlicht: (2023)
von: Lao, Dong, et al.
Veröffentlicht: (2023)
Patch-enhanced Mask Encoder Prompt Image Generation
von: Xu, Shusong, et al.
Veröffentlicht: (2024)
von: Xu, Shusong, et al.
Veröffentlicht: (2024)
Imaging foundation model for universal enhancement of non-ideal measurement CT
von: Ge, Rongjun, et al.
Veröffentlicht: (2024)
von: Ge, Rongjun, et al.
Veröffentlicht: (2024)
Steerable Visual Representations
von: Ruthardt, Jona, et al.
Veröffentlicht: (2026)
von: Ruthardt, Jona, et al.
Veröffentlicht: (2026)
OASIC: Occlusion-Agnostic and Severity-Informed Classification
von: Gijzen, Kay, et al.
Veröffentlicht: (2026)
von: Gijzen, Kay, et al.
Veröffentlicht: (2026)
Grounded Object Centric Learning
von: Kori, Avinash, et al.
Veröffentlicht: (2023)
von: Kori, Avinash, et al.
Veröffentlicht: (2023)
Enabling clinical use of foundation models for computational pathology
von: Henriksen, Audun L, et al.
Veröffentlicht: (2026)
von: Henriksen, Audun L, et al.
Veröffentlicht: (2026)
MoireMix: A Formula-Based Data Augmentation for Improving Image Classification Robustness
von: Matsuo, Yuto, et al.
Veröffentlicht: (2026)
von: Matsuo, Yuto, et al.
Veröffentlicht: (2026)
Less yet robust: crucial region selection for scene recognition
von: Zhang, Jianqi, et al.
Veröffentlicht: (2024)
von: Zhang, Jianqi, et al.
Veröffentlicht: (2024)
From Pixels to Predicates Structuring urban perception with scene graphs
von: Liu, Yunlong, et al.
Veröffentlicht: (2025)
von: Liu, Yunlong, et al.
Veröffentlicht: (2025)
An interpretable framework using foundation models for fish sex identification
von: Miao, Zheng, et al.
Veröffentlicht: (2026)
von: Miao, Zheng, et al.
Veröffentlicht: (2026)
Prompting with the human-touch: evaluating model-sensitivity of foundation models for musculoskeletal CT segmentation
von: Magg, Caroline, et al.
Veröffentlicht: (2026)
von: Magg, Caroline, et al.
Veröffentlicht: (2026)
ActiveMark: on watermarking of visual foundation models via massive activations
von: Chistyakova, Anna, et al.
Veröffentlicht: (2025)
von: Chistyakova, Anna, et al.
Veröffentlicht: (2025)
Redefining Normal: A Novel Object-Level Approach for Multi-Object Novelty Detection
von: Salehi, Mohammadreza, et al.
Veröffentlicht: (2024)
von: Salehi, Mohammadreza, et al.
Veröffentlicht: (2024)
Functionality understanding and segmentation in 3D scenes
von: Corsetti, Jaime, et al.
Veröffentlicht: (2024)
von: Corsetti, Jaime, et al.
Veröffentlicht: (2024)
Assessment of Sentinel-2 spatial and temporal coverage based on the scene classification layer
von: Sanchez, Cristhian, et al.
Veröffentlicht: (2024)
von: Sanchez, Cristhian, et al.
Veröffentlicht: (2024)
All-day Multi-scenes Lifelong Vision-and-Language Navigation with Tucker Adaptation
von: Wang, Xudong, et al.
Veröffentlicht: (2026)
von: Wang, Xudong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Better Language Models Exhibit Higher Visual Alignment
von: Ruthardt, Jona, et al.
Veröffentlicht: (2024) -
Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
von: Venkataramanan, Shashanka, et al.
Veröffentlicht: (2025) -
Neurosymbolic Inference On Foundation Models For Remote Sensing Text-to-image Retrieval With Complex Queries
von: Mezzi, Emanuele, et al.
Veröffentlicht: (2025) -
SelEx: Self-Expertise in Fine-Grained Generalized Category Discovery
von: Rastegar, Sarah, et al.
Veröffentlicht: (2024) -
Generating metamers of human scene understanding
von: Raina, Ritik, et al.
Veröffentlicht: (2026)