Qwen Look Again: Guiding Vision-Language Reasoning Models to Re-attention Visual Information
Fuente:
arXiv
Salvato in:
| Autori principali: | Chu, Xu, Chen, Xinrong, Wang, Guanyu, Tan, Zhijie, Huang, Kui, Lv, Wenyu, Mo, Tong, Li, Weiping |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Domaino1s: Guiding LLM Reasoning for Explainable Answers in High-Stakes Domains
di: Chu, Xu, et al.
Pubblicazione: (2025)
di: Chu, Xu, et al.
Pubblicazione: (2025)
Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization
di: Chu, Xu, et al.
Pubblicazione: (2026)
di: Chu, Xu, et al.
Pubblicazione: (2026)
Order Matters: Exploring Order Sensitivity in Multimodal Large Language Models
di: Tan, Zhijie, et al.
Pubblicazione: (2024)
di: Tan, Zhijie, et al.
Pubblicazione: (2024)
Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Models
di: Jian, Pu, et al.
Pubblicazione: (2025)
di: Jian, Pu, et al.
Pubblicazione: (2025)
Mitigating Hallucinations on Object Attributes using Multiview Images and Negative Instructions
di: Tan, Zhijie, et al.
Pubblicazione: (2025)
di: Tan, Zhijie, et al.
Pubblicazione: (2025)
PP-DocBee2: Improved Baselines with Efficient Data for Multimodal Document Understanding
di: Huang, Kui, et al.
Pubblicazione: (2025)
di: Huang, Kui, et al.
Pubblicazione: (2025)
Stepwise Schema-Guided Prompting Framework with Parameter Efficient Instruction Tuning for Multimedia Event Extraction
di: Yuan, Xiang, et al.
Pubblicazione: (2025)
di: Yuan, Xiang, et al.
Pubblicazione: (2025)
Look-Back: Implicit Visual Re-focusing in MLLM Reasoning
di: Yang, Shuo, et al.
Pubblicazione: (2025)
di: Yang, Shuo, et al.
Pubblicazione: (2025)
QwenStyle: Content-Preserving Style Transfer with Qwen-Image-Edit
di: Zhang, Shiwen, et al.
Pubblicazione: (2026)
di: Zhang, Shiwen, et al.
Pubblicazione: (2026)
Emotion-Qwen: A Unified Framework for Emotion and Vision Understanding
di: Huang, Dawei, et al.
Pubblicazione: (2025)
di: Huang, Dawei, et al.
Pubblicazione: (2025)
Source-Free Domain Adaptation Guided by Vision and Vision-Language Pre-Training
di: Zhang, Wenyu, et al.
Pubblicazione: (2024)
di: Zhang, Wenyu, et al.
Pubblicazione: (2024)
GraphSOS: Graph Sampling and Order Selection to Help LLMs Understand Graphs Better
di: Chu, Xu, et al.
Pubblicazione: (2025)
di: Chu, Xu, et al.
Pubblicazione: (2025)
Visual Reasoning Evaluation of Grok, Deepseek Janus, Gemini, Qwen, Mistral, and ChatGPT
di: Jegham, Nidhal, et al.
Pubblicazione: (2025)
di: Jegham, Nidhal, et al.
Pubblicazione: (2025)
MORE-R1: Guiding LVLM for Multimodal Object-Entity Relation Extraction via Stepwise Reasoning with Reinforcement Learning
di: Yuan, Xiang, et al.
Pubblicazione: (2026)
di: Yuan, Xiang, et al.
Pubblicazione: (2026)
Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
di: Wang, Peng, et al.
Pubblicazione: (2024)
di: Wang, Peng, et al.
Pubblicazione: (2024)
Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models
di: Huang, Xin, et al.
Pubblicazione: (2025)
di: Huang, Xin, et al.
Pubblicazione: (2025)
LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models
di: Shen, Yuxiang, et al.
Pubblicazione: (2026)
di: Shen, Yuxiang, et al.
Pubblicazione: (2026)
Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention
di: Zhao, Jianfei, et al.
Pubblicazione: (2025)
di: Zhao, Jianfei, et al.
Pubblicazione: (2025)
QwenCLIP: Boosting Medical Vision-Language Pretraining via LLM Embeddings and Prompt tuning
di: Wei, Xiaoyang, et al.
Pubblicazione: (2025)
di: Wei, Xiaoyang, et al.
Pubblicazione: (2025)
MPCAR: Multi-Perspective Contextual Augmentation for Enhanced Visual Reasoning in Large Vision-Language Models
di: Rahman, Amirul, et al.
Pubblicazione: (2025)
di: Rahman, Amirul, et al.
Pubblicazione: (2025)
BabyVision: Visual Reasoning Beyond Language
di: Chen, Liang, et al.
Pubblicazione: (2026)
di: Chen, Liang, et al.
Pubblicazione: (2026)
Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning of Vision Language Models
di: Tan, Huajie, et al.
Pubblicazione: (2025)
di: Tan, Huajie, et al.
Pubblicazione: (2025)
VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
di: Zhou, Guanyu, et al.
Pubblicazione: (2026)
di: Zhou, Guanyu, et al.
Pubblicazione: (2026)
Asymmetric Visual Semantic Embedding Framework for Efficient Vision-Language Alignment
di: Liu, Yang, et al.
Pubblicazione: (2025)
di: Liu, Yang, et al.
Pubblicazione: (2025)
Qwen3-VL-Seg: Unlocking Open-World Referring Segmentation with Vision-Language Grounding
di: Yao, Yuan, et al.
Pubblicazione: (2026)
di: Yao, Yuan, et al.
Pubblicazione: (2026)
Qwen3-VL Technical Report
di: Bai, Shuai, et al.
Pubblicazione: (2025)
di: Bai, Shuai, et al.
Pubblicazione: (2025)
Qwen-Image Technical Report
di: Wu, Chenfei, et al.
Pubblicazione: (2025)
di: Wu, Chenfei, et al.
Pubblicazione: (2025)
PersonViT: Large-scale Self-supervised Vision Transformer for Person Re-Identification
di: Hu, Bin, et al.
Pubblicazione: (2024)
di: Hu, Bin, et al.
Pubblicazione: (2024)
Zero-Shot Visual Reasoning by Vision-Language Models: Benchmarking and Analysis
di: Nagar, Aishik, et al.
Pubblicazione: (2024)
di: Nagar, Aishik, et al.
Pubblicazione: (2024)
CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering
di: Li, Qiming, et al.
Pubblicazione: (2026)
di: Li, Qiming, et al.
Pubblicazione: (2026)
Handle-based Mesh Deformation Guided By Vision Language Model
di: Sun, Xingpeng, et al.
Pubblicazione: (2025)
di: Sun, Xingpeng, et al.
Pubblicazione: (2025)
VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models
di: Ren, Yufan, et al.
Pubblicazione: (2025)
di: Ren, Yufan, et al.
Pubblicazione: (2025)
RT-DETRv2: Improved Baseline with Bag-of-Freebies for Real-Time Detection Transformer
di: Lv, Wenyu, et al.
Pubblicazione: (2024)
di: Lv, Wenyu, et al.
Pubblicazione: (2024)
Qwen2.5-Omni Technical Report
di: Xu, Jin, et al.
Pubblicazione: (2025)
di: Xu, Jin, et al.
Pubblicazione: (2025)
Qwen-Image-2.0 Technical Report
di: Zhao, Bing, et al.
Pubblicazione: (2026)
di: Zhao, Bing, et al.
Pubblicazione: (2026)
Robust Saliency-Aware Distillation for Few-shot Fine-grained Visual Recognition
di: Liu, Haiqi, et al.
Pubblicazione: (2023)
di: Liu, Haiqi, et al.
Pubblicazione: (2023)
Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
di: Wu, Junfei, et al.
Pubblicazione: (2025)
di: Wu, Junfei, et al.
Pubblicazione: (2025)
Qwen3-Omni Technical Report
di: Xu, Jin, et al.
Pubblicazione: (2025)
di: Xu, Jin, et al.
Pubblicazione: (2025)
VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies
di: Gao, Mingjian, et al.
Pubblicazione: (2026)
di: Gao, Mingjian, et al.
Pubblicazione: (2026)
ATA: Bridging Implicit Reasoning with Attention-Guided and Action-Guided Inference for Vision-Language Action Models
di: Yang, Cheng, et al.
Pubblicazione: (2026)
di: Yang, Cheng, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Domaino1s: Guiding LLM Reasoning for Explainable Answers in High-Stakes Domains
di: Chu, Xu, et al.
Pubblicazione: (2025) -
Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization
di: Chu, Xu, et al.
Pubblicazione: (2026) -
Order Matters: Exploring Order Sensitivity in Multimodal Large Language Models
di: Tan, Zhijie, et al.
Pubblicazione: (2024) -
Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Models
di: Jian, Pu, et al.
Pubblicazione: (2025) -
Mitigating Hallucinations on Object Attributes using Multiview Images and Negative Instructions
di: Tan, Zhijie, et al.
Pubblicazione: (2025)