Symmetrical Visual Contrastive Optimization: Aligning Vision-Language Models with Minimal Contrastive Images
Fuente:
arXiv
Guardado en:
| Autores principales: | Wu, Shengguang, Sun, Fan-Yun, Wen, Kaiyue, Haber, Nick |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Delve into Visual Contrastive Decoding for Hallucination Mitigation of Large Vision-Language Models
por: Lee, Yi-Lun, et al.
Publicado: (2024)
por: Lee, Yi-Lun, et al.
Publicado: (2024)
HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models
por: Jiang, Songtao, et al.
Publicado: (2025)
por: Jiang, Songtao, et al.
Publicado: (2025)
Seeing the Image: Prioritizing Visual Correlation by Contrastive Alignment
por: Xiao, Xin, et al.
Publicado: (2024)
por: Xiao, Xin, et al.
Publicado: (2024)
Mitigating Hallucinations in Large Vision-Language Models (LVLMs) via Language-Contrastive Decoding (LCD)
por: Manevich, Avshalom, et al.
Publicado: (2024)
por: Manevich, Avshalom, et al.
Publicado: (2024)
LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models
por: Sun, Fan-Yun, et al.
Publicado: (2024)
por: Sun, Fan-Yun, et al.
Publicado: (2024)
Modeling Caption Diversity in Contrastive Vision-Language Pretraining
por: Lavoie, Samuel, et al.
Publicado: (2024)
por: Lavoie, Samuel, et al.
Publicado: (2024)
Refining Skewed Perceptions in Vision-Language Contrastive Models through Visual Representations
por: Dai, Haocheng, et al.
Publicado: (2024)
por: Dai, Haocheng, et al.
Publicado: (2024)
Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding
por: Wang, Xintong, et al.
Publicado: (2024)
por: Wang, Xintong, et al.
Publicado: (2024)
Positive-Augmented Contrastive Learning for Vision-and-Language Evaluation and Training
por: Sarto, Sara, et al.
Publicado: (2024)
por: Sarto, Sara, et al.
Publicado: (2024)
MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference Optimization
por: Zhu, Kangyu, et al.
Publicado: (2024)
por: Zhu, Kangyu, et al.
Publicado: (2024)
Mask What Matters: Mitigating Object Hallucinations in Multimodal Large Language Models with Object-Aligned Visual Contrastive Decoding
por: Chen, Boqi, et al.
Publicado: (2026)
por: Chen, Boqi, et al.
Publicado: (2026)
CLoVe: Encoding Compositional Language in Contrastive Vision-Language Models
por: Castro, Santiago, et al.
Publicado: (2024)
por: Castro, Santiago, et al.
Publicado: (2024)
Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training
por: Wan, David, et al.
Publicado: (2024)
por: Wan, David, et al.
Publicado: (2024)
Diagnosing Bottlenecks in Data Visualization Understanding by Vision-Language Models
por: Tartaglini, Alexa R., et al.
Publicado: (2025)
por: Tartaglini, Alexa R., et al.
Publicado: (2025)
Contrastive Visual Data Augmentation
por: Zhou, Yu, et al.
Publicado: (2025)
por: Zhou, Yu, et al.
Publicado: (2025)
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning
por: Luo, Run, et al.
Publicado: (2025)
por: Luo, Run, et al.
Publicado: (2025)
Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers
por: Wen, Yuxin, et al.
Publicado: (2024)
por: Wen, Yuxin, et al.
Publicado: (2024)
LLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive Learning
por: Lan, Zhibin, et al.
Publicado: (2025)
por: Lan, Zhibin, et al.
Publicado: (2025)
Pre-trained Vision-Language Models Learn Discoverable Visual Concepts
por: Zang, Yuan, et al.
Publicado: (2024)
por: Zang, Yuan, et al.
Publicado: (2024)
Evading Visual Aphasia: Contrastive Adaptive Semantic Token Pruning for Vision-Language Models
por: Ma, Jie, et al.
Publicado: (2026)
por: Ma, Jie, et al.
Publicado: (2026)
Transductive Visual Programming: Evolving Tool Libraries from Experience for Spatial Reasoning
por: Wu, Shengguang, et al.
Publicado: (2025)
por: Wu, Shengguang, et al.
Publicado: (2025)
To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
por: Luo, Jiayun, et al.
Publicado: (2025)
por: Luo, Jiayun, et al.
Publicado: (2025)
Efficient Contrastive Decoding with Probabilistic Hallucination Detection - Mitigating Hallucinations in Large Vision Language Models -
por: Fieback, Laura, et al.
Publicado: (2025)
por: Fieback, Laura, et al.
Publicado: (2025)
WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models
por: Sugiura, Issa, et al.
Publicado: (2025)
por: Sugiura, Issa, et al.
Publicado: (2025)
Weakly Supervised Gaussian Contrastive Grounding with Large Multimodal Models for Video Question Answering
por: Wang, Haibo, et al.
Publicado: (2024)
por: Wang, Haibo, et al.
Publicado: (2024)
VLCD: Vision-Language Contrastive Distillation for Accurate and Efficient Automatic Placenta Analysis
por: Mehta, Manas, et al.
Publicado: (2025)
por: Mehta, Manas, et al.
Publicado: (2025)
Efficient Few-Shot Medical Image Analysis via Hierarchical Contrastive Vision-Language Learning
por: Fuller, Harrison, et al.
Publicado: (2025)
por: Fuller, Harrison, et al.
Publicado: (2025)
Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models
por: Woo, Sangmin, et al.
Publicado: (2025)
por: Woo, Sangmin, et al.
Publicado: (2025)
Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs
por: Pan, Zhiyu, et al.
Publicado: (2026)
por: Pan, Zhiyu, et al.
Publicado: (2026)
CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning
por: Yu, Hao, et al.
Publicado: (2025)
por: Yu, Hao, et al.
Publicado: (2025)
DOCCI: Descriptions of Connected and Contrasting Images
por: Onoe, Yasumasa, et al.
Publicado: (2024)
por: Onoe, Yasumasa, et al.
Publicado: (2024)
Contrastive Language Prompting to Ease False Positives in Medical Anomaly Detection
por: Park, YeongHyeon, et al.
Publicado: (2024)
por: Park, YeongHyeon, et al.
Publicado: (2024)
Mitigating Hallucinations in Large Vision-Language Models with Internal Fact-based Contrastive Decoding
por: Wang, Chao, et al.
Publicado: (2025)
por: Wang, Chao, et al.
Publicado: (2025)
Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
por: Li, Bin, et al.
Publicado: (2025)
por: Li, Bin, et al.
Publicado: (2025)
CoViPAL: Layer-wise Contextualized Visual Token Pruning for Large Vision-Language Models
por: Tang, Zicong, et al.
Publicado: (2025)
por: Tang, Zicong, et al.
Publicado: (2025)
Losing Visual Needles in Image Haystacks: Vision Language Models are Easily Distracted in Short and Long Contexts
por: Sharma, Aditya, et al.
Publicado: (2024)
por: Sharma, Aditya, et al.
Publicado: (2024)
Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains
por: Zhang, Juntian, et al.
Publicado: (2025)
por: Zhang, Juntian, et al.
Publicado: (2025)
BEAT: Visual Backdoor Attacks on VLM-based Embodied Agents via Contrastive Trigger Learning
por: Zhan, Qiusi, et al.
Publicado: (2025)
por: Zhan, Qiusi, et al.
Publicado: (2025)
Can Vision-Language Models Solve Visual Math Equations?
por: Choudhury, Monjoy Narayan, et al.
Publicado: (2025)
por: Choudhury, Monjoy Narayan, et al.
Publicado: (2025)
Holodeck: Language Guided Generation of 3D Embodied AI Environments
por: Yang, Yue, et al.
Publicado: (2023)
por: Yang, Yue, et al.
Publicado: (2023)
Ejemplares similares
-
Delve into Visual Contrastive Decoding for Hallucination Mitigation of Large Vision-Language Models
por: Lee, Yi-Lun, et al.
Publicado: (2024) -
HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models
por: Jiang, Songtao, et al.
Publicado: (2025) -
Seeing the Image: Prioritizing Visual Correlation by Contrastive Alignment
por: Xiao, Xin, et al.
Publicado: (2024) -
Mitigating Hallucinations in Large Vision-Language Models (LVLMs) via Language-Contrastive Decoding (LCD)
por: Manevich, Avshalom, et al.
Publicado: (2024) -
LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models
por: Sun, Fan-Yun, et al.
Publicado: (2024)