In Context Learning with Vision Transformers: Case Study
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhao, Antony, Proshkin, Alex, Hennessy, Fergal, Crivelli, Francesco |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
di: Raoufi, Behnam, et al.
Pubblicazione: (2025)
di: Raoufi, Behnam, et al.
Pubblicazione: (2025)
Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
di: Durrani, Hamza Ahmed, et al.
Pubblicazione: (2026)
di: Durrani, Hamza Ahmed, et al.
Pubblicazione: (2026)
Selection, Not Fusion: Radar-Modulated State Space Models for Radar-Camera Depth Estimation
di: Hou, Zhangcheng, et al.
Pubblicazione: (2026)
di: Hou, Zhangcheng, et al.
Pubblicazione: (2026)
Smooth regularization for efficient video recognition
di: Goldman, Gil, et al.
Pubblicazione: (2025)
di: Goldman, Gil, et al.
Pubblicazione: (2025)
Neuromorphic Monocular Depth Estimation with Uncertainty Modeling
di: Bergkvist, Viktor, et al.
Pubblicazione: (2026)
di: Bergkvist, Viktor, et al.
Pubblicazione: (2026)
Decoupling Vision and Language: Codebook Anchored Visual Adaptation
di: Wu, Jason, et al.
Pubblicazione: (2026)
di: Wu, Jason, et al.
Pubblicazione: (2026)
FeedbackSTS-Det: Sparse Frames-Based Spatio-Temporal Semantic Feedback Network for Moving Infrared Small Target Detection
di: Huang, Yian, et al.
Pubblicazione: (2026)
di: Huang, Yian, et al.
Pubblicazione: (2026)
SERA-H: Beyond Native Sentinel Spatial Limits for High-Resolution Canopy Height Mapping
di: Boudras, Thomas, et al.
Pubblicazione: (2025)
di: Boudras, Thomas, et al.
Pubblicazione: (2025)
MaSC: A Masked Similarity Metric for Evaluating Concept-Driven Generation
di: Bartkowiak, Patryk, et al.
Pubblicazione: (2026)
di: Bartkowiak, Patryk, et al.
Pubblicazione: (2026)
OnlyFlow: Optical Flow based Motion Conditioning for Video Diffusion Models
di: Koroglu, Mathis, et al.
Pubblicazione: (2024)
di: Koroglu, Mathis, et al.
Pubblicazione: (2024)
Perceptual Flow Network for Visually Grounded Reasoning
di: Li, Yangfu, et al.
Pubblicazione: (2026)
di: Li, Yangfu, et al.
Pubblicazione: (2026)
WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents
di: Liu, Bingnan, et al.
Pubblicazione: (2026)
di: Liu, Bingnan, et al.
Pubblicazione: (2026)
Implementing Adaptations for Vision AutoRegressive Model
di: Shaikh, Kaif, et al.
Pubblicazione: (2025)
di: Shaikh, Kaif, et al.
Pubblicazione: (2025)
Optimal Transport-Guided Source-Free Adaptation for Face Anti-Spoofing
di: Li, Zhuowei, et al.
Pubblicazione: (2025)
di: Li, Zhuowei, et al.
Pubblicazione: (2025)
Butter: Frequency Consistency and Hierarchical Fusion for Autonomous Driving Object Detection
di: Lin, Xiaojian, et al.
Pubblicazione: (2025)
di: Lin, Xiaojian, et al.
Pubblicazione: (2025)
VLM-VPI: A Vision-Language Reasoning Framework for Improving Automated Vehicle-Pedestrian Interactions
di: Pu, Qingwen, et al.
Pubblicazione: (2026)
di: Pu, Qingwen, et al.
Pubblicazione: (2026)
Caption-Driven Explainability: Probing CNNs for Bias via CLIP
di: Koller, Patrick, et al.
Pubblicazione: (2025)
di: Koller, Patrick, et al.
Pubblicazione: (2025)
Learning 3D object-centric representation through prediction
di: Day, John, et al.
Pubblicazione: (2024)
di: Day, John, et al.
Pubblicazione: (2024)
GLoT: A Novel Gated-Logarithmic Transformer for Efficient Sign Language Translation
di: Shahin, Nada, et al.
Pubblicazione: (2025)
di: Shahin, Nada, et al.
Pubblicazione: (2025)
Efficient Attention: Attention with Linear Complexities
di: Shen, Zhuoran, et al.
Pubblicazione: (2018)
di: Shen, Zhuoran, et al.
Pubblicazione: (2018)
SPMamba-YOLO: An Underwater Object Detection Network Based on Multi-Scale Feature Enhancement and Global Context Modeling
di: Liao, Guanghao, et al.
Pubblicazione: (2026)
di: Liao, Guanghao, et al.
Pubblicazione: (2026)
SpectralCA: Bi-Directional Cross-Attention for Next-Generation UAV Hyperspectral Vision
di: Brovko, D. V.
Pubblicazione: (2025)
di: Brovko, D. V.
Pubblicazione: (2025)
An Analysis of Layer-Freezing Strategies for Enhanced Transfer Learning in YOLO Architectures
di: Dobrzycki, Andrzej D., et al.
Pubblicazione: (2025)
di: Dobrzycki, Andrzej D., et al.
Pubblicazione: (2025)
Object detection in adverse weather conditions for autonomous vehicles using Instruct Pix2Pix
di: Gurbindo, Unai, et al.
Pubblicazione: (2025)
di: Gurbindo, Unai, et al.
Pubblicazione: (2025)
Lightweight Prompt-Guided CLIP Adaptation for Monocular Depth Estimation
di: Manghotay, Reyhaneh Ahani, et al.
Pubblicazione: (2026)
di: Manghotay, Reyhaneh Ahani, et al.
Pubblicazione: (2026)
UnCageNet: Tracking and Pose Estimation of Caged Animal
di: Dutta, Sayak, et al.
Pubblicazione: (2025)
di: Dutta, Sayak, et al.
Pubblicazione: (2025)
Hierarchical Image-Guided 3D Point Cloud Segmentation in Industrial Scenes via Multi-View Bayesian Fusion
di: Zhu, Yu, et al.
Pubblicazione: (2025)
di: Zhu, Yu, et al.
Pubblicazione: (2025)
ADAT: Time-Series-Aware Adaptive Transformer Architecture for Sign Language Translation
di: Shahin, Nada, et al.
Pubblicazione: (2025)
di: Shahin, Nada, et al.
Pubblicazione: (2025)
CC-SGG: Corner Case Scenario Generation using Learned Scene Graphs
di: Drayson, George, et al.
Pubblicazione: (2023)
di: Drayson, George, et al.
Pubblicazione: (2023)
Predictive Modeling of Maritime Radar Data Using Transformer Architecture
di: Qesaraku, Bjorna, et al.
Pubblicazione: (2025)
di: Qesaraku, Bjorna, et al.
Pubblicazione: (2025)
Spiking Neural Networks for event-based action recognition: A new task to understand their advantage
di: Vicente-Sola, Alex, et al.
Pubblicazione: (2022)
di: Vicente-Sola, Alex, et al.
Pubblicazione: (2022)
How Can One Choose the Best CAM-Based Explainability Method for a CNN Model?
di: Costa, Daniel da Silva, et al.
Pubblicazione: (2026)
di: Costa, Daniel da Silva, et al.
Pubblicazione: (2026)
Locate 3D: Real-World Object Localization via Self-Supervised Learning in 3D
di: Arnaud, Sergio, et al.
Pubblicazione: (2025)
di: Arnaud, Sergio, et al.
Pubblicazione: (2025)
Adapting SAM with Dynamic Similarity Graphs for Few-Shot Parameter-Efficient Small Dense Object Detection: A Case Study of Chickpea Pods in Field Conditions
di: Jiang, Xintong, et al.
Pubblicazione: (2025)
di: Jiang, Xintong, et al.
Pubblicazione: (2025)
Do Generative Metrics Predict YOLO Performance? An Evaluation Across Models, Augmentation Ratios, and Dataset Complexity
di: Marian, Vasile, et al.
Pubblicazione: (2026)
di: Marian, Vasile, et al.
Pubblicazione: (2026)
Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning
di: Ji, Binbin, et al.
Pubblicazione: (2025)
di: Ji, Binbin, et al.
Pubblicazione: (2025)
Context in object detection: a systematic literature review
di: Jamali, Mahtab, et al.
Pubblicazione: (2025)
di: Jamali, Mahtab, et al.
Pubblicazione: (2025)
IMASHRIMP: Automatic White Shrimp (Penaeus vannamei) Biometrical Analysis from Laboratory Images Using Computer Vision and Deep Learning
di: González, Abiam Remache, et al.
Pubblicazione: (2025)
di: González, Abiam Remache, et al.
Pubblicazione: (2025)
From eye to AI: studying rodent social behavior in the era of machine Learning
di: Chindemi, Giuseppe, et al.
Pubblicazione: (2025)
di: Chindemi, Giuseppe, et al.
Pubblicazione: (2025)
Short-Window Sliding Learning for Real-Time Violence Detection via LLM-based Auto-Labeling
di: Jung, Seoik, et al.
Pubblicazione: (2025)
di: Jung, Seoik, et al.
Pubblicazione: (2025)
Documenti analoghi
-
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
di: Raoufi, Behnam, et al.
Pubblicazione: (2025) -
Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
di: Durrani, Hamza Ahmed, et al.
Pubblicazione: (2026) -
Selection, Not Fusion: Radar-Modulated State Space Models for Radar-Camera Depth Estimation
di: Hou, Zhangcheng, et al.
Pubblicazione: (2026) -
Smooth regularization for efficient video recognition
di: Goldman, Gil, et al.
Pubblicazione: (2025) -
Neuromorphic Monocular Depth Estimation with Uncertainty Modeling
di: Bergkvist, Viktor, et al.
Pubblicazione: (2026)