Learning to Select Visual In-Context Demonstrations
Fuente:
arXiv
Salvato in:
| Autori principali: | Lee, Eugene, Lin, Yu-Chi, Diao, Jiajie |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Hierarchical Pre-Training of Vision Encoders with Large Language Models
di: Lee, Eugene, et al.
Pubblicazione: (2026)
di: Lee, Eugene, et al.
Pubblicazione: (2026)
Grounding the Score: Explicit Visual Premise Verification for Reliable Vision-Language Process Reward Models
di: Wang, Junxin, et al.
Pubblicazione: (2026)
di: Wang, Junxin, et al.
Pubblicazione: (2026)
MERIT: Modular Framework for Multimodal Misinformation Detection with Web-Grounded Reasoning
di: Shopnil, Mir Nafis Sharear, et al.
Pubblicazione: (2025)
di: Shopnil, Mir Nafis Sharear, et al.
Pubblicazione: (2025)
Beyond Vision: Contextually Enriched Image Captioning with Multi-Modal Retrieval
di: Quy, Nguyen Lam Phu, et al.
Pubblicazione: (2025)
di: Quy, Nguyen Lam Phu, et al.
Pubblicazione: (2025)
Annotating Satellite Images of Forests with Keywords from a Specialized Corpus in the Context of Change Detection
di: Neptune, Nathalie, et al.
Pubblicazione: (2025)
di: Neptune, Nathalie, et al.
Pubblicazione: (2025)
From Rule-Based Models to Deep Learning Transformers Architectures for Natural Language Processing and Sign Language Translation Systems: Survey, Taxonomy and Performance Evaluation
di: Shahin, Nada, et al.
Pubblicazione: (2024)
di: Shahin, Nada, et al.
Pubblicazione: (2024)
Detecting Legend Items on Historical Maps Using GPT-4o with In-Context Learning
di: Kirsanova, Sofia, et al.
Pubblicazione: (2025)
di: Kirsanova, Sofia, et al.
Pubblicazione: (2025)
Training-Free Diffusion Priors for Text-to-Image Generation via Optimization-based Visual Inversion
di: Dell'Erba, Samuele, et al.
Pubblicazione: (2025)
di: Dell'Erba, Samuele, et al.
Pubblicazione: (2025)
Design and Implementation of an OCR-Powered Pipeline for Table Extraction from Invoices
di: Patel, Parshva Dhilankumar
Pubblicazione: (2025)
di: Patel, Parshva Dhilankumar
Pubblicazione: (2025)
SIMMER: Cross-Modal Food Image--Recipe Retrieval via MLLM-Based Embedding
di: Gomi, Keisuke, et al.
Pubblicazione: (2026)
di: Gomi, Keisuke, et al.
Pubblicazione: (2026)
Gated Recursive Fusion: A Stateful Approach to Scalable Multimodal Transformers
di: Shihata, Yusuf
Pubblicazione: (2025)
di: Shihata, Yusuf
Pubblicazione: (2025)
SynCo: Synthetic Hard Negatives for Contrastive Visual Representation Learning
di: Giakoumoglou, Nikolaos, et al.
Pubblicazione: (2024)
di: Giakoumoglou, Nikolaos, et al.
Pubblicazione: (2024)
When Gender is Hard to See: Multi-Attribute Support for Long-Range Recognition
di: Mbongo, Nzakiese, et al.
Pubblicazione: (2025)
di: Mbongo, Nzakiese, et al.
Pubblicazione: (2025)
Large Language Models for Simultaneous Named Entity Extraction and Spelling Correction
di: Whittaker, Edward, et al.
Pubblicazione: (2024)
di: Whittaker, Edward, et al.
Pubblicazione: (2024)
Boosting Few-Shot Learning with Disentangled Self-Supervised Learning and Meta-Learning for Medical Image Classification
di: Pachetti, Eva, et al.
Pubblicazione: (2024)
di: Pachetti, Eva, et al.
Pubblicazione: (2024)
LLM-Guided Exemplar Selection for Few-Shot Wearable-Sensor Human Activity Recognition
di: Ronando, Elsen, et al.
Pubblicazione: (2025)
di: Ronando, Elsen, et al.
Pubblicazione: (2025)
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation
di: Hansen-Estruch, Philippe, et al.
Pubblicazione: (2025)
di: Hansen-Estruch, Philippe, et al.
Pubblicazione: (2025)
LiftAvatar: Kinematic-Space Completion for Expression-Controlled 3D Gaussian Avatar Animation
di: Wei, Hualiang, et al.
Pubblicazione: (2026)
di: Wei, Hualiang, et al.
Pubblicazione: (2026)
Neural Harmonic Textures for High-Quality Primitive Based Neural Reconstruction
di: Condor, Jorge, et al.
Pubblicazione: (2026)
di: Condor, Jorge, et al.
Pubblicazione: (2026)
Libra: Leveraging Temporal Images for Biomedical Radiology Analysis
di: Zhang, Xi, et al.
Pubblicazione: (2024)
di: Zhang, Xi, et al.
Pubblicazione: (2024)
RealTalk: Realistic Emotion-Aware Lifelike Talking-Head Synthesis
di: Wang, Wenqing, et al.
Pubblicazione: (2025)
di: Wang, Wenqing, et al.
Pubblicazione: (2025)
In Context Learning with Vision Transformers: Case Study
di: Zhao, Antony, et al.
Pubblicazione: (2025)
di: Zhao, Antony, et al.
Pubblicazione: (2025)
IDDR-NGP: Incorporating Detectors for Distractor Removal with Instant Neural Radiance Field
di: Huang, Xianliang, et al.
Pubblicazione: (2026)
di: Huang, Xianliang, et al.
Pubblicazione: (2026)
Tri-VQA: Triangular Reasoning Medical Visual Question Answering for Multi-Attribute Analysis
di: Fan, Lin, et al.
Pubblicazione: (2024)
di: Fan, Lin, et al.
Pubblicazione: (2024)
VLM4Rec: Multimodal Semantic Representation for Recommendation with Large Vision-Language Models
di: Valencia, Ty, et al.
Pubblicazione: (2026)
di: Valencia, Ty, et al.
Pubblicazione: (2026)
DISCO: Document Intelligence Suite for COmparative Evaluation
di: Benkirane, Kenza, et al.
Pubblicazione: (2026)
di: Benkirane, Kenza, et al.
Pubblicazione: (2026)
Learning Hierarchical Image Segmentation For Recognition and By Recognition
di: Ke, Tsung-Wei, et al.
Pubblicazione: (2022)
di: Ke, Tsung-Wei, et al.
Pubblicazione: (2022)
CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding
di: Zhang, Xi, et al.
Pubblicazione: (2025)
di: Zhang, Xi, et al.
Pubblicazione: (2025)
Using Deep Learning to Generate Semantically Correct Hindi Captions
di: Khan, Wasim Akram, et al.
Pubblicazione: (2026)
di: Khan, Wasim Akram, et al.
Pubblicazione: (2026)
Step-CoT: Stepwise Visual Chain-of-Thought for Medical Visual Question Answering
di: Fan, Lin, et al.
Pubblicazione: (2026)
di: Fan, Lin, et al.
Pubblicazione: (2026)
Uncertainty Quantification in Continual Open-World Learning
di: Rios, Amanda S., et al.
Pubblicazione: (2024)
di: Rios, Amanda S., et al.
Pubblicazione: (2024)
The Influence of Iconicity in Transfer Learning for Sign Language Recognition
di: Artiaga, Keren, et al.
Pubblicazione: (2026)
di: Artiaga, Keren, et al.
Pubblicazione: (2026)
On the Limitations of Vision-Language Models in Understanding Image Transforms
di: Anis, Ahmad Mustafa, et al.
Pubblicazione: (2025)
di: Anis, Ahmad Mustafa, et al.
Pubblicazione: (2025)
LightMover: Generative Light Movement with Color and Intensity Controls
di: Zhou, Gengze, et al.
Pubblicazione: (2026)
di: Zhou, Gengze, et al.
Pubblicazione: (2026)
myMNIST: Benchmark of PETNN, KAN, and Classical Deep Learning Models for Burmese Handwritten Digit Recognition
di: Thu, Ye Kyaw, et al.
Pubblicazione: (2026)
di: Thu, Ye Kyaw, et al.
Pubblicazione: (2026)
InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information
di: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Pubblicazione: (2025)
di: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Pubblicazione: (2025)
WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
di: Niu, Yuwei, et al.
Pubblicazione: (2025)
di: Niu, Yuwei, et al.
Pubblicazione: (2025)
Generative AI for Video Translation: A Scalable Architecture for Multilingual Video Conferencing
di: Oskooei, Amirkia Rafiei, et al.
Pubblicazione: (2025)
di: Oskooei, Amirkia Rafiei, et al.
Pubblicazione: (2025)
PBSBench: A Multi-Level Vision-Language Framework and Benchmark for Hematopathology Whole Slide Image Interpretation
di: Wang, Yuanlong, et al.
Pubblicazione: (2026)
di: Wang, Yuanlong, et al.
Pubblicazione: (2026)
Deformation-Free Cross-Domain Image Registration via Position-Encoded Temporal Attention
di: Wang, Yiwen, et al.
Pubblicazione: (2026)
di: Wang, Yiwen, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Hierarchical Pre-Training of Vision Encoders with Large Language Models
di: Lee, Eugene, et al.
Pubblicazione: (2026) -
Grounding the Score: Explicit Visual Premise Verification for Reliable Vision-Language Process Reward Models
di: Wang, Junxin, et al.
Pubblicazione: (2026) -
MERIT: Modular Framework for Multimodal Misinformation Detection with Web-Grounded Reasoning
di: Shopnil, Mir Nafis Sharear, et al.
Pubblicazione: (2025) -
Beyond Vision: Contextually Enriched Image Captioning with Multi-Modal Retrieval
di: Quy, Nguyen Lam Phu, et al.
Pubblicazione: (2025) -
Annotating Satellite Images of Forests with Keywords from a Specialized Corpus in the Context of Change Detection
di: Neptune, Nathalie, et al.
Pubblicazione: (2025)