Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Yeyuan, Gao, Dehong, Yi, Lei, Jin, Linbo, Zhang, Jinxia, Yang, Libin, Cai, Xiaoyan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning
von: Wang, Yeyuan, et al.
Veröffentlicht: (2025)
von: Wang, Yeyuan, et al.
Veröffentlicht: (2025)
CoF: Coarse to Fine-Grained Image Understanding for Multi-modal Large Language Models
von: Wang, Yeyuan, et al.
Veröffentlicht: (2024)
von: Wang, Yeyuan, et al.
Veröffentlicht: (2024)
Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
von: Li, Bin, et al.
Veröffentlicht: (2025)
von: Li, Bin, et al.
Veröffentlicht: (2025)
FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training
von: Huang, Jiale, et al.
Veröffentlicht: (2024)
von: Huang, Jiale, et al.
Veröffentlicht: (2024)
MADiff: Text-Guided Fashion Image Editing with Mask Prediction and Attention-Enhanced Diffusion
von: Zhan, Zechao, et al.
Veröffentlicht: (2024)
von: Zhan, Zechao, et al.
Veröffentlicht: (2024)
VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models
von: Zhao, Fufangchen, et al.
Veröffentlicht: (2025)
von: Zhao, Fufangchen, et al.
Veröffentlicht: (2025)
MedTri: A Platform for Structured Medical Report Normalization to Enhance Vision-Language Pretraining
von: Chu, Yuetan, et al.
Veröffentlicht: (2026)
von: Chu, Yuetan, et al.
Veröffentlicht: (2026)
Cross-Organ and Cross-Scanner Adenocarcinoma Segmentation using Rein to Fine-tune Vision Foundation Models
von: Cai, Pengzhou, et al.
Veröffentlicht: (2024)
von: Cai, Pengzhou, et al.
Veröffentlicht: (2024)
Knowledge Regularized Negative Feature Tuning of Vision-Language Models for Out-of-Distribution Detection
von: Zhu, Wenjie, et al.
Veröffentlicht: (2025)
von: Zhu, Wenjie, et al.
Veröffentlicht: (2025)
Unified Language-Vision Pretraining in LLM with Dynamic Discrete Visual Tokenization
von: Jin, Yang, et al.
Veröffentlicht: (2023)
von: Jin, Yang, et al.
Veröffentlicht: (2023)
Self-Supervised Pretraining for Fine-Grained Plankton Recognition
von: Kareinen, Joona, et al.
Veröffentlicht: (2025)
von: Kareinen, Joona, et al.
Veröffentlicht: (2025)
PathFLIP: Fine-grained Language-Image Pretraining for Versatile Computational Pathology
von: Liu, Fengchun, et al.
Veröffentlicht: (2025)
von: Liu, Fengchun, et al.
Veröffentlicht: (2025)
Negative Label Guided OOD Detection with Pretrained Vision-Language Models
von: Jiang, Xue, et al.
Veröffentlicht: (2024)
von: Jiang, Xue, et al.
Veröffentlicht: (2024)
Separators in Enhancing Autoregressive Pretraining for Vision Mamba
von: Liu, Hanpeng, et al.
Veröffentlicht: (2026)
von: Liu, Hanpeng, et al.
Veröffentlicht: (2026)
Towards Fine-Grained Vision-Language Alignment for Few-Shot Anomaly Detection
von: Fan, Yuanting, et al.
Veröffentlicht: (2025)
von: Fan, Yuanting, et al.
Veröffentlicht: (2025)
Twofold Debiasing Enhances Fine-Grained Learning with Coarse Labels
von: Zhao, Xin-yang, et al.
Veröffentlicht: (2025)
von: Zhao, Xin-yang, et al.
Veröffentlicht: (2025)
GeoAlignCLIP: Enhancing Fine-Grained Vision-Language Alignment in Remote Sensing via Multi-Granular Consistency Learning
von: Yang, Xiao, et al.
Veröffentlicht: (2026)
von: Yang, Xiao, et al.
Veröffentlicht: (2026)
Enhancing Zero-Shot Brain Tumor Subtype Classification via Fine-Grained Patch-Text Alignment
von: Gan, Lubin, et al.
Veröffentlicht: (2025)
von: Gan, Lubin, et al.
Veröffentlicht: (2025)
Autoregressive Pretraining with Mamba in Vision
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
Revisiting Prompt Pretraining of Vision-Language Models
von: Chen, Zhenyuan, et al.
Veröffentlicht: (2024)
von: Chen, Zhenyuan, et al.
Veröffentlicht: (2024)
Enhancing Fine-Grained Image Classifications via Cascaded Vision Language Models
von: Wei, Canshi
Veröffentlicht: (2024)
von: Wei, Canshi
Veröffentlicht: (2024)
fine-CLIP: Enhancing Zero-Shot Fine-Grained Surgical Action Recognition with Vision-Language Models
von: Sharma, Saurav, et al.
Veröffentlicht: (2025)
von: Sharma, Saurav, et al.
Veröffentlicht: (2025)
SGIA: Enhancing Fine-Grained Visual Classification with Sequence Generative Image Augmentation
von: Liao, Qiyu, et al.
Veröffentlicht: (2024)
von: Liao, Qiyu, et al.
Veröffentlicht: (2024)
FILA: Fine-Grained Vision Language Models
von: Zhu, Shiding, et al.
Veröffentlicht: (2024)
von: Zhu, Shiding, et al.
Veröffentlicht: (2024)
Zero-Shot Fine-Grained Image Classification Using Large Vision-Language Models
von: Atabuzzaman, Md., et al.
Veröffentlicht: (2025)
von: Atabuzzaman, Md., et al.
Veröffentlicht: (2025)
MedKCO: Medical Vision-Language Pretraining via Knowledge-Driven Cognitive Orchestration
von: Zhang, Chenran, et al.
Veröffentlicht: (2026)
von: Zhang, Chenran, et al.
Veröffentlicht: (2026)
Bootstrap Fine-Grained Vision-Language Alignment for Unified Zero-Shot Anomaly Localization
von: Deng, Hanqiu, et al.
Veröffentlicht: (2023)
von: Deng, Hanqiu, et al.
Veröffentlicht: (2023)
Layer-Specific Fine-Tuning for Improved Negation Handling in Medical Vision-Language Models
von: Abbasi, Ali, et al.
Veröffentlicht: (2026)
von: Abbasi, Ali, et al.
Veröffentlicht: (2026)
GazeCLIP: Gaze-Guided CLIP with Adaptive-Enhanced Fine-Grained Language Prompt for Deepfake Attribution and Detection
von: Zhang, Yaning, et al.
Veröffentlicht: (2026)
von: Zhang, Yaning, et al.
Veröffentlicht: (2026)
Synthesize, Diagnose, and Optimize: Towards Fine-Grained Vision-Language Understanding
von: Peng, Wujian, et al.
Veröffentlicht: (2023)
von: Peng, Wujian, et al.
Veröffentlicht: (2023)
EPIC-Bench: A Perception-Centric Benchmark for Fine-Grained Embodied Visual Grounding in Vision-Language Models
von: Shan, Haozhe, et al.
Veröffentlicht: (2026)
von: Shan, Haozhe, et al.
Veröffentlicht: (2026)
Cross-Modal Clustering-Guided Negative Sampling for Self-Supervised Joint Learning from Medical Images and Reports
von: Lan, Libin, et al.
Veröffentlicht: (2025)
von: Lan, Libin, et al.
Veröffentlicht: (2025)
TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment
von: Cao, Bingyi, et al.
Veröffentlicht: (2026)
von: Cao, Bingyi, et al.
Veröffentlicht: (2026)
Glo-VLMs: Leveraging Vision-Language Models for Fine-Grained Diseased Glomerulus Classification
von: Guo, Zhenhao, et al.
Veröffentlicht: (2025)
von: Guo, Zhenhao, et al.
Veröffentlicht: (2025)
Enhancing Vision-Language Few-Shot Adaptation with Negative Learning
von: Zhang, Ce, et al.
Veröffentlicht: (2024)
von: Zhang, Ce, et al.
Veröffentlicht: (2024)
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
von: Zhang, Chubin, et al.
Veröffentlicht: (2026)
von: Zhang, Chubin, et al.
Veröffentlicht: (2026)
Fine-Grained Alignment in Vision-and-Language Navigation through Bayesian Optimization
von: Song, Yuhang, et al.
Veröffentlicht: (2024)
von: Song, Yuhang, et al.
Veröffentlicht: (2024)
SimLabel: Consistency-Guided OOD Detection with Pretrained Vision-Language Models
von: Zou, Shu, et al.
Veröffentlicht: (2025)
von: Zou, Shu, et al.
Veröffentlicht: (2025)
Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations
von: Cui, Yibo, et al.
Veröffentlicht: (2025)
von: Cui, Yibo, et al.
Veröffentlicht: (2025)
NAVERO: Unlocking Fine-Grained Semantics for Video-Language Compositionality
von: Tao, Chaofan, et al.
Veröffentlicht: (2024)
von: Tao, Chaofan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning
von: Wang, Yeyuan, et al.
Veröffentlicht: (2025) -
CoF: Coarse to Fine-Grained Image Understanding for Multi-modal Large Language Models
von: Wang, Yeyuan, et al.
Veröffentlicht: (2024) -
Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
von: Li, Bin, et al.
Veröffentlicht: (2025) -
FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training
von: Huang, Jiale, et al.
Veröffentlicht: (2024) -
MADiff: Text-Guided Fashion Image Editing with Mask Prediction and Attention-Enhanced Diffusion
von: Zhan, Zechao, et al.
Veröffentlicht: (2024)