Seeing Through Words: Controlling Visual Retrieval Quality with Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Lu, Jianglin, Jenni, Simon, Kafle, Kushal, Shi, Jing, Zhao, Handong, Fu, Yun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
More Than the Final Answer: Improving Visual Extraction and Logical Consistency in Vision-Language Models
di: Just, Hoang Anh, et al.
Pubblicazione: (2025)
di: Just, Hoang Anh, et al.
Pubblicazione: (2025)
RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward
di: Wu, Qiucheng, et al.
Pubblicazione: (2026)
di: Wu, Qiucheng, et al.
Pubblicazione: (2026)
Improving Large Vision and Language Models by Learning from a Panel of Peers
di: Hernandez, Jefferson, et al.
Pubblicazione: (2025)
di: Hernandez, Jefferson, et al.
Pubblicazione: (2025)
Plot'n Polish: Zero-shot Story Visualization and Disentangled Editing with Text-to-Image Diffusion Models
di: Akdemir, Kiymet, et al.
Pubblicazione: (2025)
di: Akdemir, Kiymet, et al.
Pubblicazione: (2025)
The Indra Representation Hypothesis for Multimodal Alignment
di: Lu, Jianglin, et al.
Pubblicazione: (2026)
di: Lu, Jianglin, et al.
Pubblicazione: (2026)
The Photographer Eye: Teaching Multimodal Large Language Models to Understand Image Aesthetics like Photographers
di: Qi, Daiqing, et al.
Pubblicazione: (2025)
di: Qi, Daiqing, et al.
Pubblicazione: (2025)
FINEMATCH: Aspect-based Fine-grained Image and Text Mismatch Detection and Correction
di: Hua, Hang, et al.
Pubblicazione: (2024)
di: Hua, Hang, et al.
Pubblicazione: (2024)
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
di: Shrestha, Robik, et al.
Pubblicazione: (2020)
di: Shrestha, Robik, et al.
Pubblicazione: (2020)
Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression Tasks
di: Dong, Qihua, et al.
Pubblicazione: (2026)
di: Dong, Qihua, et al.
Pubblicazione: (2026)
Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations
di: Yang, Ziyan, et al.
Pubblicazione: (2022)
di: Yang, Ziyan, et al.
Pubblicazione: (2022)
Reminding Multimodal Large Language Models of Object-aware Knowledge with Retrieved Tags
di: Qi, Daiqing, et al.
Pubblicazione: (2024)
di: Qi, Daiqing, et al.
Pubblicazione: (2024)
DUALVISION: RGB-Infrared Multimodal Large Language Models for Robust Visual Reasoning
di: Majeedi, Abrar, et al.
Pubblicazione: (2026)
di: Majeedi, Abrar, et al.
Pubblicazione: (2026)
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles
di: Slyman, Eric, et al.
Pubblicazione: (2025)
di: Slyman, Eric, et al.
Pubblicazione: (2025)
Are Bias Mitigation Techniques for Deep Learning Effective?
di: Shrestha, Robik, et al.
Pubblicazione: (2021)
di: Shrestha, Robik, et al.
Pubblicazione: (2021)
Can Retrieval Heads See Images? Multimodal Retrieval Heads in Long-Context Vision-Language Models
di: Li, Aaron Branson Cigres, et al.
Pubblicazione: (2026)
di: Li, Aaron Branson Cigres, et al.
Pubblicazione: (2026)
Seeing Through Smoke: Surgical Desmoking for Improved Visual Perception
di: Lu, Jingpei, et al.
Pubblicazione: (2026)
di: Lu, Jingpei, et al.
Pubblicazione: (2026)
Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models
di: He, Zoe Wanying, et al.
Pubblicazione: (2025)
di: He, Zoe Wanying, et al.
Pubblicazione: (2025)
Trajectory Prediction Meets Large Language Models: A Survey
di: Xu, Yi, et al.
Pubblicazione: (2025)
di: Xu, Yi, et al.
Pubblicazione: (2025)
Restore-R1: Efficient Image Restoration Agents via Reinforcement Learning with Multimodal LLM Perceptual Feedback
di: Lu, Jianglin, et al.
Pubblicazione: (2025)
di: Lu, Jianglin, et al.
Pubblicazione: (2025)
SCoRD: Subject-Conditional Relation Detection with Text-Augmented Data
di: Yang, Ziyan, et al.
Pubblicazione: (2023)
di: Yang, Ziyan, et al.
Pubblicazione: (2023)
Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models
di: Caffagni, Davide, et al.
Pubblicazione: (2025)
di: Caffagni, Davide, et al.
Pubblicazione: (2025)
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
di: Luo, Gen, et al.
Pubblicazione: (2025)
di: Luo, Gen, et al.
Pubblicazione: (2025)
Seeing Through the Blur: Unlocking Defocus Maps for Deepfake Detection
di: Jeon, Minsun, et al.
Pubblicazione: (2025)
di: Jeon, Minsun, et al.
Pubblicazione: (2025)
Repeating Words for Video-Language Retrieval with Coarse-to-Fine Objectives
di: Zhao, Haoyu, et al.
Pubblicazione: (2025)
di: Zhao, Haoyu, et al.
Pubblicazione: (2025)
Breaking the Encoder Barrier for Seamless Video-Language Understanding
di: Li, Handong, et al.
Pubblicazione: (2025)
di: Li, Handong, et al.
Pubblicazione: (2025)
Seeing Through Their Eyes: Evaluating Visual Perspective Taking in Vision Language Models
di: Góral, Gracjan, et al.
Pubblicazione: (2024)
di: Góral, Gracjan, et al.
Pubblicazione: (2024)
VividCam: Learning Unconventional Camera Motions from Virtual Synthetic Videos
di: Wu, Qiucheng, et al.
Pubblicazione: (2025)
di: Wu, Qiucheng, et al.
Pubblicazione: (2025)
Consistency and Uncertainty: Identifying Unreliable Responses From Black-Box Vision-Language Models for Selective Visual Question Answering
di: Khan, Zaid, et al.
Pubblicazione: (2024)
di: Khan, Zaid, et al.
Pubblicazione: (2024)
Language Models Can See Better: Visual Contrastive Decoding For LLM Multimodal Reasoning
di: Pang, Yuqi, et al.
Pubblicazione: (2025)
di: Pang, Yuqi, et al.
Pubblicazione: (2025)
Seeing Through Touch: Tactile-Driven Visual Localization of Material Regions
di: Kim, Seongyu, et al.
Pubblicazione: (2026)
di: Kim, Seongyu, et al.
Pubblicazione: (2026)
ExpertGen: Training-Free Expert Guidance for Controllable Text-to-Face Generation
di: Shi, Liang, et al.
Pubblicazione: (2025)
di: Shi, Liang, et al.
Pubblicazione: (2025)
See, Rank, and Filter: Important Word-Aware Clip Filtering via Scene Understanding for Moment Retrieval and Highlight Detection
di: Lee, YuEun, et al.
Pubblicazione: (2025)
di: Lee, YuEun, et al.
Pubblicazione: (2025)
Seeing Through the PRISM: Compound & Controllable Restoration of Scientific Images
di: Kurinchi-Vendhan, Rupa, et al.
Pubblicazione: (2026)
di: Kurinchi-Vendhan, Rupa, et al.
Pubblicazione: (2026)
DepthFocus: Controllable Depth Estimation for See-Through Scenes
di: Min, Junhong, et al.
Pubblicazione: (2025)
di: Min, Junhong, et al.
Pubblicazione: (2025)
Seeing Through the Tool: A Controlled Benchmark for Occlusion Robustness in Foundation Segmentation Models
di: Ho, Nhan, et al.
Pubblicazione: (2026)
di: Ho, Nhan, et al.
Pubblicazione: (2026)
Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language Models
di: Wu, Jiaying, et al.
Pubblicazione: (2025)
di: Wu, Jiaying, et al.
Pubblicazione: (2025)
Self-Training Large Language Models for Improved Visual Program Synthesis With Visual Reinforcement
di: Khan, Zaid, et al.
Pubblicazione: (2024)
di: Khan, Zaid, et al.
Pubblicazione: (2024)
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors
di: Shahgir, Haz Sameen, et al.
Pubblicazione: (2026)
di: Shahgir, Haz Sameen, et al.
Pubblicazione: (2026)
Visual Words Meet BM25: Sparse Auto-Encoder Visual Word Scoring for Image Retrieval
di: Han, Donghoon, et al.
Pubblicazione: (2026)
di: Han, Donghoon, et al.
Pubblicazione: (2026)
Revisiting Multi-Modal LLM Evaluation
di: Lu, Jian, et al.
Pubblicazione: (2024)
di: Lu, Jian, et al.
Pubblicazione: (2024)
Documenti analoghi
-
More Than the Final Answer: Improving Visual Extraction and Logical Consistency in Vision-Language Models
di: Just, Hoang Anh, et al.
Pubblicazione: (2025) -
RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward
di: Wu, Qiucheng, et al.
Pubblicazione: (2026) -
Improving Large Vision and Language Models by Learning from a Panel of Peers
di: Hernandez, Jefferson, et al.
Pubblicazione: (2025) -
Plot'n Polish: Zero-shot Story Visualization and Disentangled Editing with Text-to-Image Diffusion Models
di: Akdemir, Kiymet, et al.
Pubblicazione: (2025) -
The Indra Representation Hypothesis for Multimodal Alignment
di: Lu, Jianglin, et al.
Pubblicazione: (2026)