Effectively Enhancing Vision Language Large Models by Prompt Augmentation and Caption Utilization
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhao, Minyi, Wang, Jie, Li, Zhaoyang, Zhang, Jiyuan, Sun, Zhenbang, Zhou, Shuigeng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
One Model for Two Tasks: Cooperatively Recognizing and Recovering Low-Resolution Scene Text Images by Iterative Mutual Guidance
por: Zhao, Minyi, et al.
Publicado: (2024)
por: Zhao, Minyi, et al.
Publicado: (2024)
Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens
por: Chen, Feng, et al.
Publicado: (2024)
por: Chen, Feng, et al.
Publicado: (2024)
Raw Data Matters: Enhancing Prompt Tuning by Internal Augmentation on Vision-Language Models
por: Li, Haoyang, et al.
Publicado: (2025)
por: Li, Haoyang, et al.
Publicado: (2025)
From Captions to Rewards (CAREVL): Leveraging Large Language Model Experts for Enhanced Reward Modeling in Large Vision-Language Models
por: Dai, Muzhi, et al.
Publicado: (2025)
por: Dai, Muzhi, et al.
Publicado: (2025)
Mitigating Image Captioning Hallucinations in Vision-Language Models
por: Zhao, Fei, et al.
Publicado: (2025)
por: Zhao, Fei, et al.
Publicado: (2025)
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding
por: Li, Yun, et al.
Publicado: (2025)
por: Li, Yun, et al.
Publicado: (2025)
SURf: Teaching Large Vision-Language Models to Selectively Utilize Retrieved Information
por: Sun, Jiashuo, et al.
Publicado: (2024)
por: Sun, Jiashuo, et al.
Publicado: (2024)
Effective Cloud Removal for Remote Sensing Images by an Improved Mean-Reverting Denoising Model with Elucidated Design Space
por: Liu, Yi, et al.
Publicado: (2025)
por: Liu, Yi, et al.
Publicado: (2025)
Player-Centric Multimodal Prompt Generation for Large Language Model Based Identity-Aware Basketball Video Captioning
por: Xi, Zeyu, et al.
Publicado: (2025)
por: Xi, Zeyu, et al.
Publicado: (2025)
Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction
por: Zhang, Yuanhong, et al.
Publicado: (2026)
por: Zhang, Yuanhong, et al.
Publicado: (2026)
Prompt-Based Caption Generation for Single-Tooth Dental Images Using Vision-Language Models
por: Sukhanova, Anastasiia, et al.
Publicado: (2026)
por: Sukhanova, Anastasiia, et al.
Publicado: (2026)
SCRA-VQA: Summarized Caption-Rerank for Augmented Large Language Models in Visual Question Answering
por: Zhang, Yan, et al.
Publicado: (2025)
por: Zhang, Yan, et al.
Publicado: (2025)
Modeling Variants of Prompts for Vision-Language Models
por: Li, Ao, et al.
Publicado: (2025)
por: Li, Ao, et al.
Publicado: (2025)
CAI: Caption-Sensitive Attention Intervention for Mitigating Object Hallucination in Large Vision-Language Models
por: Li, Qiming, et al.
Publicado: (2025)
por: Li, Qiming, et al.
Publicado: (2025)
PromptKD: Unsupervised Prompt Distillation for Vision-Language Models
por: Li, Zheng, et al.
Publicado: (2024)
por: Li, Zheng, et al.
Publicado: (2024)
video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models
por: Tang, Changli, et al.
Publicado: (2025)
por: Tang, Changli, et al.
Publicado: (2025)
InPK: Infusing Prior Knowledge into Prompt for Vision-Language Models
por: Zhou, Shuchang, et al.
Publicado: (2025)
por: Zhou, Shuchang, et al.
Publicado: (2025)
Instruction-guided Multi-Granularity Segmentation and Captioning with Large Multimodal Model
por: Zhou, Li, et al.
Publicado: (2024)
por: Zhou, Li, et al.
Publicado: (2024)
Memory-Augmented Vision-Language Agents for Persistent and Semantically Consistent Object Captioning
por: Galliena, Tommaso, et al.
Publicado: (2026)
por: Galliena, Tommaso, et al.
Publicado: (2026)
MoAPT: Mixture of Adversarial Prompt Tuning for Vision-Language Models
por: Zhao, Shiji, et al.
Publicado: (2025)
por: Zhao, Shiji, et al.
Publicado: (2025)
What Makes VLMs Robust? Towards Reconciling Robustness and Accuracy in Vision-Language Models
por: Nie, Sen, et al.
Publicado: (2026)
por: Nie, Sen, et al.
Publicado: (2026)
CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering
por: Li, Qiming, et al.
Publicado: (2026)
por: Li, Qiming, et al.
Publicado: (2026)
Text Data-Centric Image Captioning with Interactive Prompts
por: Wang, Yiyu, et al.
Publicado: (2024)
por: Wang, Yiyu, et al.
Publicado: (2024)
MAP-VLA: Memory-Augmented Prompting for Vision-Language-Action Model in Robotic Manipulation
por: Li, Runhao, et al.
Publicado: (2025)
por: Li, Runhao, et al.
Publicado: (2025)
HPT++: Hierarchically Prompting Vision-Language Models with Multi-Granularity Knowledge Generation and Improved Structure Modeling
por: Wang, Yubin, et al.
Publicado: (2024)
por: Wang, Yubin, et al.
Publicado: (2024)
Contrastive Language-Image Learning with Augmented Textual Prompts for 3D/4D FER Using Vision-Language Model
por: Behzad, Muzammil, et al.
Publicado: (2025)
por: Behzad, Muzammil, et al.
Publicado: (2025)
FVG-PT: Adaptive Foreground View-Guided Prompt Tuning for Vision-Language Models
por: Li, Haoyang, et al.
Publicado: (2026)
por: Li, Haoyang, et al.
Publicado: (2026)
TA-Prompting: Enhancing Video Large Language Models for Dense Video Captioning via Temporal Anchors
por: Cheng, Wei-Yuan, et al.
Publicado: (2026)
por: Cheng, Wei-Yuan, et al.
Publicado: (2026)
Multi-modal Attribute Prompting for Vision-Language Models
por: Liu, Xin, et al.
Publicado: (2024)
por: Liu, Xin, et al.
Publicado: (2024)
Multi-Turn Adaptive Prompting Attack on Large Vision-Language Models
por: Choi, In Chong, et al.
Publicado: (2026)
por: Choi, In Chong, et al.
Publicado: (2026)
RxnCaption: Reformulating Reaction Diagram Parsing as Visual Prompt Guided Captioning
por: Song, Jiahe, et al.
Publicado: (2025)
por: Song, Jiahe, et al.
Publicado: (2025)
Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception
por: Sun, Yanpeng, et al.
Publicado: (2024)
por: Sun, Yanpeng, et al.
Publicado: (2024)
Mixture of Prompt Learning for Vision Language Models
por: Du, Yu, et al.
Publicado: (2024)
por: Du, Yu, et al.
Publicado: (2024)
Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning
por: Lu, Fan, et al.
Publicado: (2024)
por: Lu, Fan, et al.
Publicado: (2024)
RESTORE: Towards Feature Shift for Vision-Language Prompt Learning
por: Yang, Yuncheng, et al.
Publicado: (2024)
por: Yang, Yuncheng, et al.
Publicado: (2024)
Self-Rewarding Large Vision-Language Models for Optimizing Prompts in Text-to-Image Generation
por: Yang, Hongji, et al.
Publicado: (2025)
por: Yang, Hongji, et al.
Publicado: (2025)
GEASS: Gated Evidence-Adaptive Selective Caption Trust for Vision-Language Models
por: Li, Zeshang, et al.
Publicado: (2026)
por: Li, Zeshang, et al.
Publicado: (2026)
Attention Prompting on Image for Large Vision-Language Models
por: Yu, Runpeng, et al.
Publicado: (2024)
por: Yu, Runpeng, et al.
Publicado: (2024)
Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training
por: Zhang, Xinsong, et al.
Publicado: (2025)
por: Zhang, Xinsong, et al.
Publicado: (2025)
VLRM: Vision-Language Models act as Reward Models for Image Captioning
por: Dzabraev, Maksim, et al.
Publicado: (2024)
por: Dzabraev, Maksim, et al.
Publicado: (2024)
Ejemplares similares
-
One Model for Two Tasks: Cooperatively Recognizing and Recovering Low-Resolution Scene Text Images by Iterative Mutual Guidance
por: Zhao, Minyi, et al.
Publicado: (2024) -
Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens
por: Chen, Feng, et al.
Publicado: (2024) -
Raw Data Matters: Enhancing Prompt Tuning by Internal Augmentation on Vision-Language Models
por: Li, Haoyang, et al.
Publicado: (2025) -
From Captions to Rewards (CAREVL): Leveraging Large Language Model Experts for Enhanced Reward Modeling in Large Vision-Language Models
por: Dai, Muzhi, et al.
Publicado: (2025) -
Mitigating Image Captioning Hallucinations in Vision-Language Models
por: Zhao, Fei, et al.
Publicado: (2025)