LLM-empowered Dynamic Prompt Routing for Vision-Language Models Tuning under Long-Tailed Distributions
Fuente:
arXiv
Guardado en:
| Autores principales: | Jia, Yongju, Ma, Jiarui, Li, Xiangxian, Zhang, Baiqiao, Cao, Xianhui, Liu, Juan, Bian, Yulong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MAGNeT: Multimodal Adaptive Gaussian Networks for Intent Inference in Moving Target Selection across Complex Scenarios
por: Li, Xiangxian, et al.
Publicado: (2025)
por: Li, Xiangxian, et al.
Publicado: (2025)
Quantized Vision-Language Models for Damage Assessment: A Comparative Study of LLaVA-1.5-7B Quantization Levels
por: Yasuno, Takato
Publicado: (2026)
por: Yasuno, Takato
Publicado: (2026)
SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
por: Chaybouti, Sofian, et al.
Publicado: (2025)
por: Chaybouti, Sofian, et al.
Publicado: (2025)
Revisiting [CLS] and Patch Token Interaction in Vision Transformers
por: Marouani, Alexis, et al.
Publicado: (2026)
por: Marouani, Alexis, et al.
Publicado: (2026)
Taming the Tail: Leveraging Asymmetric Loss and Pade Approximation to Overcome Medical Image Long-Tailed Class Imbalance
por: Kashyap, Pankhi, et al.
Publicado: (2024)
por: Kashyap, Pankhi, et al.
Publicado: (2024)
MienCap: Realtime Performance-Based Facial Animation with Live Mood Dynamics
por: Pan, Ye, et al.
Publicado: (2025)
por: Pan, Ye, et al.
Publicado: (2025)
Scalable Face Security Vision Foundation Model for Deepfake, Diffusion, and Spoofing Detection
por: Wang, Gaojian, et al.
Publicado: (2025)
por: Wang, Gaojian, et al.
Publicado: (2025)
ViFiCon: Vision and Wireless Association Via Self-Supervised Contrastive Learning
por: Meegan, Nicholas, et al.
Publicado: (2022)
por: Meegan, Nicholas, et al.
Publicado: (2022)
Visual-Text Cross Alignment: Refining the Similarity Score in Vision-Language Models
por: Li, Jinhao, et al.
Publicado: (2024)
por: Li, Jinhao, et al.
Publicado: (2024)
Quick unsupervised hyperspectral dimensionality reduction for earth observation: a comparison
por: Lupu, Daniela, et al.
Publicado: (2024)
por: Lupu, Daniela, et al.
Publicado: (2024)
FairDeDup: Detecting and Mitigating Vision-Language Fairness Disparities in Semantic Dataset Deduplication
por: Slyman, Eric, et al.
Publicado: (2024)
por: Slyman, Eric, et al.
Publicado: (2024)
RealHD: A High-Quality Dataset for Robust Detection of State-of-the-Art AI-Generated Images
por: Yu, Hanzhe, et al.
Publicado: (2026)
por: Yu, Hanzhe, et al.
Publicado: (2026)
VLSlice: Interactive Vision-and-Language Slice Discovery
por: Slyman, Eric, et al.
Publicado: (2023)
por: Slyman, Eric, et al.
Publicado: (2023)
FEDTAIL: Federated Long-Tailed Domain Generalization with Sharpness-Guided Gradient Matching
por: Gupta, Sunny, et al.
Publicado: (2025)
por: Gupta, Sunny, et al.
Publicado: (2025)
COLORA: Efficient Fine-Tuning for Convolutional Models with a Study Case on Optical Coherence Tomography Image Classification
por: Rivera, Mariano, et al.
Publicado: (2025)
por: Rivera, Mariano, et al.
Publicado: (2025)
Unveiling Text in Challenging Stone Inscriptions: A Character-Context-Aware Patching Strategy for Binarization
por: Jena, Pratyush, et al.
Publicado: (2026)
por: Jena, Pratyush, et al.
Publicado: (2026)
VersaGen: Unleashing Versatile Visual Control for Text-to-Image Synthesis
por: Chen, Zhipeng, et al.
Publicado: (2024)
por: Chen, Zhipeng, et al.
Publicado: (2024)
HySparK: Hybrid Sparse Masking for Large Scale Medical Image Pre-Training
por: Tang, Fenghe, et al.
Publicado: (2024)
por: Tang, Fenghe, et al.
Publicado: (2024)
Mobile-Ready Automated Triage of Diabetic Retinopathy Using Digital Fundus Images
por: Joshi, Aadi, et al.
Publicado: (2026)
por: Joshi, Aadi, et al.
Publicado: (2026)
Multimodal-to-Text Prompt Engineering in Large Language Models Using Feature Embeddings for GNSS Interference Characterization
por: Manjunath, Harshith, et al.
Publicado: (2025)
por: Manjunath, Harshith, et al.
Publicado: (2025)
Improving Visual Object Tracking through Visual Prompting
por: Chen, Shih-Fang, et al.
Publicado: (2024)
por: Chen, Shih-Fang, et al.
Publicado: (2024)
Kolmogorov-Arnold Attention: Is Learnable Attention Better For Vision Transformers?
por: Maity, Subhajit, et al.
Publicado: (2025)
por: Maity, Subhajit, et al.
Publicado: (2025)
Learning Unified Representation of 3D Gaussian Splatting
por: Xin, Yuelin, et al.
Publicado: (2025)
por: Xin, Yuelin, et al.
Publicado: (2025)
InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information
por: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Publicado: (2025)
por: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Publicado: (2025)
Frequency-Decomposed INR for NIR-Assisted Low-Light RGB Image Denoising
por: Shi, Ligen, et al.
Publicado: (2026)
por: Shi, Ligen, et al.
Publicado: (2026)
Neural Fields for 3D Tracking of Anatomy and Surgical Instruments in Monocular Laparoscopic Video Clips
por: Gerats, Beerend G. A., et al.
Publicado: (2024)
por: Gerats, Beerend G. A., et al.
Publicado: (2024)
Cora: Correspondence-aware image editing using few step diffusion
por: Alimohammadi, Amirhossein, et al.
Publicado: (2025)
por: Alimohammadi, Amirhossein, et al.
Publicado: (2025)
Pointing-Based Object Recognition
por: Hajdúch, Lukáš, et al.
Publicado: (2026)
por: Hajdúch, Lukáš, et al.
Publicado: (2026)
A Hierarchical Self-Consistent Regularization Approach to Satellite Image Time Series Classification
por: Weikmann, Giulio, et al.
Publicado: (2025)
por: Weikmann, Giulio, et al.
Publicado: (2025)
Multitemporal Latent Dynamical Framework for Hyperspectral Images Unmixing
por: Li, Ruiying, et al.
Publicado: (2025)
por: Li, Ruiying, et al.
Publicado: (2025)
Symmetry Awareness Encoded Deep Learning Framework for Brain Imaging Analysis
por: Ma, Yang, et al.
Publicado: (2024)
por: Ma, Yang, et al.
Publicado: (2024)
The Adobe Hidden Feature and its Impact on Sensor Attribution
por: Butora, Jan, et al.
Publicado: (2023)
por: Butora, Jan, et al.
Publicado: (2023)
Comparing Euclidean and Hyperbolic K-Means for Generalized Category Discovery
por: Dalal, Mohamad, et al.
Publicado: (2026)
por: Dalal, Mohamad, et al.
Publicado: (2026)
Generating real-time detailed ground visualisations from sparse aerial point clouds
por: Murray, Aidan, et al.
Publicado: (2025)
por: Murray, Aidan, et al.
Publicado: (2025)
A Spitting Image: Modular Superpixel Tokenization in Vision Transformers
por: Aasan, Marius, et al.
Publicado: (2024)
por: Aasan, Marius, et al.
Publicado: (2024)
Sparse Concept Bottleneck Models: Gumbel Tricks in Contrastive Learning
por: Semenov, Andrei, et al.
Publicado: (2024)
por: Semenov, Andrei, et al.
Publicado: (2024)
Robust Domain Generalisation with Causal Invariant Bayesian Neural Networks
por: Gendron, Gaël, et al.
Publicado: (2024)
por: Gendron, Gaël, et al.
Publicado: (2024)
FAME: Feature Activation Map Explanation on Image Classification and Face Recognition
por: Zhang, Xinyi, et al.
Publicado: (2026)
por: Zhang, Xinyi, et al.
Publicado: (2026)
GLoT: A Novel Gated-Logarithmic Transformer for Efficient Sign Language Translation
por: Shahin, Nada, et al.
Publicado: (2025)
por: Shahin, Nada, et al.
Publicado: (2025)
Meta Co-Training: Two Views are Better than One
por: Rothenberger, Jay C., et al.
Publicado: (2023)
por: Rothenberger, Jay C., et al.
Publicado: (2023)
Ejemplares similares
-
MAGNeT: Multimodal Adaptive Gaussian Networks for Intent Inference in Moving Target Selection across Complex Scenarios
por: Li, Xiangxian, et al.
Publicado: (2025) -
Quantized Vision-Language Models for Damage Assessment: A Comparative Study of LLaVA-1.5-7B Quantization Levels
por: Yasuno, Takato
Publicado: (2026) -
SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
por: Chaybouti, Sofian, et al.
Publicado: (2025) -
Revisiting [CLS] and Patch Token Interaction in Vision Transformers
por: Marouani, Alexis, et al.
Publicado: (2026) -
Taming the Tail: Leveraging Asymmetric Loss and Pade Approximation to Overcome Medical Image Long-Tailed Class Imbalance
por: Kashyap, Pankhi, et al.
Publicado: (2024)