VLMine: Long-Tail Data Mining with Vision Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Mao, Meyer, Gregory P., Zhang, Zaiwei, Park, Dennis, Mustikovela, Siva Karthik, Chai, Yuning, Wolff, Eric M |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision
by: Xu, Yi, et al.
Published: (2024)
by: Xu, Yi, et al.
Published: (2024)
ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts
by: Cai, Mu, et al.
Published: (2023)
by: Cai, Mu, et al.
Published: (2023)
VLM-KD: Knowledge Distillation from VLM for Long-Tail Visual Recognition
by: Zhang, Zaiwei, et al.
Published: (2024)
by: Zhang, Zaiwei, et al.
Published: (2024)
Flash3D: Super-scaling Point Transformers through Joint Hardware-Geometry Locality
by: Chen, Liyan, et al.
Published: (2024)
by: Chen, Liyan, et al.
Published: (2024)
Cohere3D: Exploiting Temporal Coherence for Unsupervised Representation Learning of Vision-based Autonomous Driving
by: Xie, Yichen, et al.
Published: (2024)
by: Xie, Yichen, et al.
Published: (2024)
Dynamic Cluster Data Sampling for Efficient and Long-Tail-Aware Vision-Language Pre-training
by: Liang, Mingliang, et al.
Published: (2026)
by: Liang, Mingliang, et al.
Published: (2026)
Generative Data Mining with Longtail-Guided Diffusion
by: Hayden, David S., et al.
Published: (2025)
by: Hayden, David S., et al.
Published: (2025)
Efficient and Long-Tailed Generalization for Pre-trained Vision-Language Model
by: Shi, Jiang-Xin, et al.
Published: (2024)
by: Shi, Jiang-Xin, et al.
Published: (2024)
Long-Tailed Distribution-Aware Router For Mixture-of-Experts in Large Vision-Language Model
by: Cai, Chaoxiang, et al.
Published: (2025)
by: Cai, Chaoxiang, et al.
Published: (2025)
Cluster-Aware Neural Collapse Prompt Tuning for Long-Tailed Generalization of Vision-Language Models
by: Guo, Boyang, et al.
Published: (2026)
by: Guo, Boyang, et al.
Published: (2026)
HMVLM: Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios
by: Wang, Daming, et al.
Published: (2025)
by: Wang, Daming, et al.
Published: (2025)
Unleashing the Power of Vision-Language Models for Long-Tailed Multi-Label Visual Recognition
by: Tang, Wei, et al.
Published: (2025)
by: Tang, Wei, et al.
Published: (2025)
CAPT: Class-Aware Prompt Tuning for Federated Long-Tailed Learning with Vision-Language Model
by: Hou, Shihao, et al.
Published: (2025)
by: Hou, Shihao, et al.
Published: (2025)
The Neglected Tails in Vision-Language Models
by: Parashar, Shubham, et al.
Published: (2024)
by: Parashar, Shubham, et al.
Published: (2024)
Uncertainty-Guided Enhancement on Driving Perception System via Foundation Models
by: Yang, Yunhao, et al.
Published: (2024)
by: Yang, Yunhao, et al.
Published: (2024)
Light Up the Shadows: Enhance Long-Tailed Entity Grounding with Concept-Guided Vision-Language Models
by: Zhang, Yikai, et al.
Published: (2024)
by: Zhang, Yikai, et al.
Published: (2024)
LLM-empowered Dynamic Prompt Routing for Vision-Language Models Tuning under Long-Tailed Distributions
by: Jia, Yongju, et al.
Published: (2025)
by: Jia, Yongju, et al.
Published: (2025)
VisCon-100K: Leveraging Contextual Web Data for Fine-tuning Vision Language Models
by: Kumar, Gokul Karthik, et al.
Published: (2025)
by: Kumar, Gokul Karthik, et al.
Published: (2025)
KPL: Training-Free Medical Knowledge Mining of Vision-Language Models
by: Liu, Jiaxiang, et al.
Published: (2025)
by: Liu, Jiaxiang, et al.
Published: (2025)
SEAL: Vision-Language Model-Based Safe End-to-End Cooperative Autonomous Driving with Adaptive Long-Tail Modeling
by: You, Junwei, et al.
Published: (2025)
by: You, Junwei, et al.
Published: (2025)
Enhancing Features in Long-tailed Data Using Large Vision Model
by: Han, Pengxiao, et al.
Published: (2025)
by: Han, Pengxiao, et al.
Published: (2025)
Dual-Margin Embedding for Fine-Grained Long-Tailed Plant Taxonomy
by: Low, Cheng Yaw, et al.
Published: (2025)
by: Low, Cheng Yaw, et al.
Published: (2025)
HDGlyph: A Hierarchical Disentangled Glyph-Based Framework for Long-Tail Text Rendering in Diffusion Models
by: Zhuang, Shuhan, et al.
Published: (2025)
by: Zhuang, Shuhan, et al.
Published: (2025)
Long-Tailed Out-of-Distribution Detection: Prioritizing Attention to Tail
by: He, Yina, et al.
Published: (2024)
by: He, Yina, et al.
Published: (2024)
Empowering Vision Transformers with Multi-Scale Causal Intervention for Long-Tailed Image Classification
by: Yan, Xiaoshuo, et al.
Published: (2025)
by: Yan, Xiaoshuo, et al.
Published: (2025)
Queryable Prototype Multiple Instance Learning with Vision-Language Models for Incremental Whole Slide Image Classification
by: Gou, Jiaxiang, et al.
Published: (2024)
by: Gou, Jiaxiang, et al.
Published: (2024)
LongTail Driving Scenarios with Reasoning Traces: The KITScenes LongTail Dataset
by: Wagner, Royden, et al.
Published: (2026)
by: Wagner, Royden, et al.
Published: (2026)
From Head to Tail: Towards Balanced Representation in Large Vision-Language Models through Adaptive Data Calibration
by: Song, Mingyang, et al.
Published: (2025)
by: Song, Mingyang, et al.
Published: (2025)
FOMO-3D: Using Vision Foundation Models for Long-Tailed 3D Object Detection
by: Yang, Anqi Joyce, et al.
Published: (2026)
by: Yang, Anqi Joyce, et al.
Published: (2026)
TailedCore: Few-Shot Sampling for Unsupervised Long-Tail Noisy Anomaly Detection
by: Jung, Yoon Gyo, et al.
Published: (2025)
by: Jung, Yoon Gyo, et al.
Published: (2025)
Learning from Reduced Labels for Long-Tailed Data
by: Wei, Meng, et al.
Published: (2024)
by: Wei, Meng, et al.
Published: (2024)
TLD: A Vehicle Tail Light signal Dataset and Benchmark
by: Chai, Jinhao, et al.
Published: (2024)
by: Chai, Jinhao, et al.
Published: (2024)
Calibration-Aware Prompt Learning for Medical Vision-Language Models
by: Basu, Abhishek, et al.
Published: (2025)
by: Basu, Abhishek, et al.
Published: (2025)
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models
by: Yu, Keunwoo Peter, et al.
Published: (2025)
by: Yu, Keunwoo Peter, et al.
Published: (2025)
Bayesian Test-Time Adaptation for Vision-Language Models
by: Zhou, Lihua, et al.
Published: (2025)
by: Zhou, Lihua, et al.
Published: (2025)
Feature Fusion from Head to Tail for Long-Tailed Visual Recognition
by: Li, Mengke, et al.
Published: (2023)
by: Li, Mengke, et al.
Published: (2023)
DeepScan: A Training-Free Framework for Visually Grounded Reasoning in Large Vision-Language Models
by: Li, Yangfu, et al.
Published: (2026)
by: Li, Yangfu, et al.
Published: (2026)
DepthLM: Metric Depth From Vision Language Models
by: Cai, Zhipeng, et al.
Published: (2025)
by: Cai, Zhipeng, et al.
Published: (2025)
Vision-Language Models Assisted Unsupervised Video Anomaly Detection
by: Jiang, Yalong, et al.
Published: (2024)
by: Jiang, Yalong, et al.
Published: (2024)
Exploring Vision-Language Models for Open-Vocabulary Zero-Shot Action Segmentation
by: Unmesh, Asim, et al.
Published: (2026)
by: Unmesh, Asim, et al.
Published: (2026)
Similar Items
-
VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision
by: Xu, Yi, et al.
Published: (2024) -
ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts
by: Cai, Mu, et al.
Published: (2023) -
VLM-KD: Knowledge Distillation from VLM for Long-Tail Visual Recognition
by: Zhang, Zaiwei, et al.
Published: (2024) -
Flash3D: Super-scaling Point Transformers through Joint Hardware-Geometry Locality
by: Chen, Liyan, et al.
Published: (2024) -
Cohere3D: Exploiting Temporal Coherence for Unsupervised Representation Learning of Vision-based Autonomous Driving
by: Xie, Yichen, et al.
Published: (2024)