Enhancing Large Vision Language Models with Self-Training on Image Comprehension
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Deng, Yihe, Lu, Pan, Yin, Fan, Hu, Ziniu, Shen, Sheng, Gu, Quanquan, Zou, James, Chang, Kai-Wei, Wang, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mitigating Object Hallucination in Large Vision-Language Models via Image-Grounded Guidance
von: Zhao, Linxi, et al.
Veröffentlicht: (2024)
von: Zhao, Linxi, et al.
Veröffentlicht: (2024)
OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles
von: Deng, Yihe, et al.
Veröffentlicht: (2025)
von: Deng, Yihe, et al.
Veröffentlicht: (2025)
Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves
von: Deng, Yihe, et al.
Veröffentlicht: (2023)
von: Deng, Yihe, et al.
Veröffentlicht: (2023)
Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
von: Chen, Zixiang, et al.
Veröffentlicht: (2024)
von: Chen, Zixiang, et al.
Veröffentlicht: (2024)
Entropy-Based Adaptive Weighting for Self-Training
von: Wang, Xiaoxuan, et al.
Veröffentlicht: (2025)
von: Wang, Xiaoxuan, et al.
Veröffentlicht: (2025)
Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning
von: Lu, Fan, et al.
Veröffentlicht: (2024)
von: Lu, Fan, et al.
Veröffentlicht: (2024)
MIRAI: Evaluating LLM Agents for Event Forecasting
von: Ye, Chenchen, et al.
Veröffentlicht: (2024)
von: Ye, Chenchen, et al.
Veröffentlicht: (2024)
Characterizing Truthfulness in Large Language Model Generations with Local Intrinsic Dimension
von: Yin, Fan, et al.
Veröffentlicht: (2024)
von: Yin, Fan, et al.
Veröffentlicht: (2024)
OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks
von: Hu, Wenbo, et al.
Veröffentlicht: (2026)
von: Hu, Wenbo, et al.
Veröffentlicht: (2026)
Self-Play Fine-Tuning of Diffusion Models for Text-to-Image Generation
von: Yuan, Huizhuo, et al.
Veröffentlicht: (2024)
von: Yuan, Huizhuo, et al.
Veröffentlicht: (2024)
MRAG-Bench: Vision-Centric Evaluation for Retrieval-Augmented Multimodal Models
von: Hu, Wenbo, et al.
Veröffentlicht: (2024)
von: Hu, Wenbo, et al.
Veröffentlicht: (2024)
RISE: Enhancing VLM Image Annotation with Self-Supervised Reasoning
von: Hu, Suhang, et al.
Veröffentlicht: (2025)
von: Hu, Suhang, et al.
Veröffentlicht: (2025)
Self-Control of LLM Behaviors by Compressing Suffix Gradient into Prefix Controller
von: Cai, Min, et al.
Veröffentlicht: (2024)
von: Cai, Min, et al.
Veröffentlicht: (2024)
Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation
von: Yu, Hong-Tao, et al.
Veröffentlicht: (2025)
von: Yu, Hong-Tao, et al.
Veröffentlicht: (2025)
On Asymmetric Optimization of Reasoning and Perception in Vision-Language Model Post-Training
von: Wu, Xueqing, et al.
Veröffentlicht: (2026)
von: Wu, Xueqing, et al.
Veröffentlicht: (2026)
VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning
von: Wu, Xueqing, et al.
Veröffentlicht: (2024)
von: Wu, Xueqing, et al.
Veröffentlicht: (2024)
Matryoshka Query Transformer for Large Vision-Language Models
von: Hu, Wenbo, et al.
Veröffentlicht: (2024)
von: Hu, Wenbo, et al.
Veröffentlicht: (2024)
Medical Vision-Language Pre-Training for Brain Abnormalities
von: Monajatipoor, Masoud, et al.
Veröffentlicht: (2024)
von: Monajatipoor, Masoud, et al.
Veröffentlicht: (2024)
InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
von: Dong, Xiaoyi, et al.
Veröffentlicht: (2024)
von: Dong, Xiaoyi, et al.
Veröffentlicht: (2024)
ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
von: Zhang, Dan, et al.
Veröffentlicht: (2024)
von: Zhang, Dan, et al.
Veröffentlicht: (2024)
Self-Enhanced Image Clustering with Cross-Modal Semantic Consistency
von: Li, Zihan, et al.
Veröffentlicht: (2025)
von: Li, Zihan, et al.
Veröffentlicht: (2025)
S-GRPO: Unified Post-Training for Large Vision-Language Models
von: Yan, Yuming, et al.
Veröffentlicht: (2026)
von: Yan, Yuming, et al.
Veröffentlicht: (2026)
Accelerated Preference Optimization for Large Language Model Alignment
von: He, Jiafan, et al.
Veröffentlicht: (2024)
von: He, Jiafan, et al.
Veröffentlicht: (2024)
Synchronous Faithfulness Monitoring for Trustworthy Retrieval-Augmented Generation
von: Wu, Di, et al.
Veröffentlicht: (2024)
von: Wu, Di, et al.
Veröffentlicht: (2024)
Enhancing Fine-Grained Image Classifications via Cascaded Vision Language Models
von: Wei, Canshi
Veröffentlicht: (2024)
von: Wei, Canshi
Veröffentlicht: (2024)
VEGA: Learning Interleaved Image-Text Comprehension in Vision-Language Large Models
von: Zhou, Chenyu, et al.
Veröffentlicht: (2024)
von: Zhou, Chenyu, et al.
Veröffentlicht: (2024)
Self-Routing RAG: Binding Selective Retrieval with Knowledge Verbalization
von: Wu, Di, et al.
Veröffentlicht: (2025)
von: Wu, Di, et al.
Veröffentlicht: (2025)
Solving Inequality Proofs with Large Language Models
von: Lu, Pan, et al.
Veröffentlicht: (2025)
von: Lu, Pan, et al.
Veröffentlicht: (2025)
MACAROON: Training Vision-Language Models To Be Your Engaged Partners
von: Wu, Shujin, et al.
Veröffentlicht: (2024)
von: Wu, Shujin, et al.
Veröffentlicht: (2024)
Frontier AI systems have surpassed the self-replicating red line
von: Pan, Xudong, et al.
Veröffentlicht: (2024)
von: Pan, Xudong, et al.
Veröffentlicht: (2024)
Self-Evolving Visual Concept Library using Vision-Language Critics
von: Sehgal, Atharva, et al.
Veröffentlicht: (2025)
von: Sehgal, Atharva, et al.
Veröffentlicht: (2025)
Visualizing Dialogues: Enhancing Image Selection through Dialogue Understanding with Large Language Models
von: Kao, Chang-Sheng, et al.
Veröffentlicht: (2024)
von: Kao, Chang-Sheng, et al.
Veröffentlicht: (2024)
A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future
von: Zhong, Jialun, et al.
Veröffentlicht: (2025)
von: Zhong, Jialun, et al.
Veröffentlicht: (2025)
Flow-DPO: Improving LLM Mathematical Reasoning through Online Multi-Agent Learning
von: Deng, Yihe, et al.
Veröffentlicht: (2024)
von: Deng, Yihe, et al.
Veröffentlicht: (2024)
Re-ReST: Reflection-Reinforced Self-Training for Language Agents
von: Dou, Zi-Yi, et al.
Veröffentlicht: (2024)
von: Dou, Zi-Yi, et al.
Veröffentlicht: (2024)
ConcEPT: Concept-Enhanced Pre-Training for Language Models
von: Wang, Xintao, et al.
Veröffentlicht: (2024)
von: Wang, Xintao, et al.
Veröffentlicht: (2024)
A Lightweight Convolution and Vision Transformer integrated model with Multi-scale Self-attention Mechanism
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
The Male CEO and the Female Assistant: Evaluation and Mitigation of Gender Biases in Text-To-Image Generation of Dual Subjects
von: Wan, Yixin, et al.
Veröffentlicht: (2024)
von: Wan, Yixin, et al.
Veröffentlicht: (2024)
CompAlign: Improving Compositional Text-to-Image Generation with a Complex Benchmark and Fine-Grained Feedback
von: Wan, Yixin, et al.
Veröffentlicht: (2025)
von: Wan, Yixin, et al.
Veröffentlicht: (2025)
Protein Large Language Models: A Comprehensive Survey
von: Xiao, Yijia, et al.
Veröffentlicht: (2025)
von: Xiao, Yijia, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Mitigating Object Hallucination in Large Vision-Language Models via Image-Grounded Guidance
von: Zhao, Linxi, et al.
Veröffentlicht: (2024) -
OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles
von: Deng, Yihe, et al.
Veröffentlicht: (2025) -
Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves
von: Deng, Yihe, et al.
Veröffentlicht: (2023) -
Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
von: Chen, Zixiang, et al.
Veröffentlicht: (2024) -
Entropy-Based Adaptive Weighting for Self-Training
von: Wang, Xiaoxuan, et al.
Veröffentlicht: (2025)