Finding Needles in Images: Can Multimodal LLMs Locate Fine Details?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Thakkar, Parth, Agarwal, Ankush, Kasu, Prasad, Bansal, Pulkit, Devaguptapu, Chaitanya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hybrid Graphs for Table-and-Text based Question Answering using LLMs
von: Agarwal, Ankush, et al.
Veröffentlicht: (2025)
von: Agarwal, Ankush, et al.
Veröffentlicht: (2025)
Semantic Graph Consistency: Going Beyond Patches for Regularizing Self-Supervised Vision Transformers
von: Devaguptapu, Chaitanya, et al.
Veröffentlicht: (2024)
von: Devaguptapu, Chaitanya, et al.
Veröffentlicht: (2024)
Towards a Training Free Approach for 3D Scene Editing
von: Madhavaram, Vivek, et al.
Veröffentlicht: (2024)
von: Madhavaram, Vivek, et al.
Veröffentlicht: (2024)
The Percept-V Challenge: Can Multimodal LLMs Crack Simple Perception Problems?
von: Ghosh, Samrajnee, et al.
Veröffentlicht: (2025)
von: Ghosh, Samrajnee, et al.
Veröffentlicht: (2025)
Find your Needle: Small Object Image Retrieval via Multi-Object Attention Optimization
von: Green, Michael, et al.
Veröffentlicht: (2025)
von: Green, Michael, et al.
Veröffentlicht: (2025)
DetailFlow: 1D Coarse-to-Fine Autoregressive Image Generation via Next-Detail Prediction
von: Liu, Yiheng, et al.
Veröffentlicht: (2025)
von: Liu, Yiheng, et al.
Veröffentlicht: (2025)
Devil is in the Detail: Towards Injecting Fine Details of Image Prompt in Image Generation via Conflict-free Guidance and Stratified Attention
von: Jo, Kyungmin, et al.
Veröffentlicht: (2025)
von: Jo, Kyungmin, et al.
Veröffentlicht: (2025)
DetailCLIP: Detail-Oriented CLIP for Fine-Grained Tasks
von: Monsefi, Amin Karimi, et al.
Veröffentlicht: (2024)
von: Monsefi, Amin Karimi, et al.
Veröffentlicht: (2024)
How Well Can Vision Language Models See Image Details?
von: Gou, Chenhui, et al.
Veröffentlicht: (2024)
von: Gou, Chenhui, et al.
Veröffentlicht: (2024)
Beyond Illumination: Fine-Grained Detail Preservation in Extreme Dark Image Restoration
von: Zhang, Tongshun, et al.
Veröffentlicht: (2025)
von: Zhang, Tongshun, et al.
Veröffentlicht: (2025)
No Detail Left Behind: Revisiting Self-Retrieval for Fine-Grained Image Captioning
von: Gaur, Manu, et al.
Veröffentlicht: (2024)
von: Gaur, Manu, et al.
Veröffentlicht: (2024)
Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?
von: Bhattacharyya, Apratim, et al.
Veröffentlicht: (2025)
von: Bhattacharyya, Apratim, et al.
Veröffentlicht: (2025)
Needle In A Multimodal Haystack
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
Find Matching Faces Based On Face Parameters
von: Bhatt, Setu A., et al.
Veröffentlicht: (2025)
von: Bhatt, Setu A., et al.
Veröffentlicht: (2025)
Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation
von: Yuan, Qianhao, et al.
Veröffentlicht: (2026)
von: Yuan, Qianhao, et al.
Veröffentlicht: (2026)
PoSh: Using Scene Graphs To Guide LLMs-as-a-Judge For Detailed Image Descriptions
von: Ananthram, Amith, et al.
Veröffentlicht: (2025)
von: Ananthram, Amith, et al.
Veröffentlicht: (2025)
Missing Fine Details in Images: Last Seen in High Frequencies
von: Medi, Tejaswini, et al.
Veröffentlicht: (2025)
von: Medi, Tejaswini, et al.
Veröffentlicht: (2025)
Going Down Memory Lane: Scaling Tokens for Video Stream Understanding with Dynamic KV-Cache Memory
von: Agarwal, Vatsal, et al.
Veröffentlicht: (2026)
von: Agarwal, Vatsal, et al.
Veröffentlicht: (2026)
DetailCLIP: Injecting Image Details into CLIP's Feature Space
von: Zhang, Zilun, et al.
Veröffentlicht: (2022)
von: Zhang, Zilun, et al.
Veröffentlicht: (2022)
Can LLMs' Tuning Methods Work in Medical Multimodal Domain?
von: Chen, Jiawei, et al.
Veröffentlicht: (2024)
von: Chen, Jiawei, et al.
Veröffentlicht: (2024)
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects?
von: Li, Jiaxuan, et al.
Veröffentlicht: (2024)
von: Li, Jiaxuan, et al.
Veröffentlicht: (2024)
Towards Perceiving Small Visual Details in Zero-shot Visual Question Answering with Multimodal LLMs
von: Zhang, Jiarui, et al.
Veröffentlicht: (2023)
von: Zhang, Jiarui, et al.
Veröffentlicht: (2023)
Explainable Gait Abnormality Detection Using Dual-Dataset CNN-LSTM Models
von: Agarwal, Parth, et al.
Veröffentlicht: (2025)
von: Agarwal, Parth, et al.
Veröffentlicht: (2025)
D-HUMOR: Dark Humor Understanding via Multimodal Open-ended Reasoning -- A Benchmark Dataset and Method
von: Kasu, Sai Kartheek Reddy, et al.
Veröffentlicht: (2025)
von: Kasu, Sai Kartheek Reddy, et al.
Veröffentlicht: (2025)
MotionSight: Boosting Fine-Grained Motion Understanding in Multimodal LLMs
von: Du, Yipeng, et al.
Veröffentlicht: (2025)
von: Du, Yipeng, et al.
Veröffentlicht: (2025)
Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
von: Li, Juncheng, et al.
Veröffentlicht: (2023)
von: Li, Juncheng, et al.
Veröffentlicht: (2023)
MGIMM: Multi-Granularity Instruction Multimodal Model for Attribute-Guided Remote Sensing Image Detailed Description
von: Yang, Cong, et al.
Veröffentlicht: (2024)
von: Yang, Cong, et al.
Veröffentlicht: (2024)
Benchmarking and Improving Detail Image Caption
von: Dong, Hongyuan, et al.
Veröffentlicht: (2024)
von: Dong, Hongyuan, et al.
Veröffentlicht: (2024)
EVALALIGN: Supervised Fine-Tuning Multimodal LLMs with Human-Aligned Data for Evaluating Text-to-Image Models
von: Tan, Zhiyu, et al.
Veröffentlicht: (2024)
von: Tan, Zhiyu, et al.
Veröffentlicht: (2024)
DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?
von: Jiao, Qirui, et al.
Veröffentlicht: (2025)
von: Jiao, Qirui, et al.
Veröffentlicht: (2025)
DPDEdit: Detail-Preserved Diffusion Models for Multimodal Fashion Image Editing
von: Wang, Xiaolong, et al.
Veröffentlicht: (2024)
von: Wang, Xiaolong, et al.
Veröffentlicht: (2024)
DEIG: Detail-Enhanced Instance Generation with Fine-Grained Semantic Control
von: Du, Shiyan, et al.
Veröffentlicht: (2026)
von: Du, Shiyan, et al.
Veröffentlicht: (2026)
AbsGS: Recovering Fine Details for 3D Gaussian Splatting
von: Ye, Zongxin, et al.
Veröffentlicht: (2024)
von: Ye, Zongxin, et al.
Veröffentlicht: (2024)
GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs
von: Nguyen, Duy, et al.
Veröffentlicht: (2025)
von: Nguyen, Duy, et al.
Veröffentlicht: (2025)
EditScout: Locating Forged Regions from Diffusion-based Edited Images with Multimodal LLM
von: Nguyen, Quang, et al.
Veröffentlicht: (2024)
von: Nguyen, Quang, et al.
Veröffentlicht: (2024)
MFSR: Multi-fractal Feature for Super-resolution Reconstruction with Fine Details Recovery
von: Yang, Lianping, et al.
Veröffentlicht: (2025)
von: Yang, Lianping, et al.
Veröffentlicht: (2025)
Long-LRM++: Preserving Fine Details in Feed-Forward Wide-Coverage Reconstruction
von: Ziwen, Chen, et al.
Veröffentlicht: (2025)
von: Ziwen, Chen, et al.
Veröffentlicht: (2025)
Decoupling Fine Detail and Global Geometry for Compressed Depth Map Super-Resolution
von: Zheng, Huan, et al.
Veröffentlicht: (2024)
von: Zheng, Huan, et al.
Veröffentlicht: (2024)
Generating Fine Details of Entity Interactions
von: Gu, Xinyi, et al.
Veröffentlicht: (2025)
von: Gu, Xinyi, et al.
Veröffentlicht: (2025)
Findings of the Counter Turing Test: AI-Generated Image Detection
von: Roy, Rajarshi, et al.
Veröffentlicht: (2026)
von: Roy, Rajarshi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Hybrid Graphs for Table-and-Text based Question Answering using LLMs
von: Agarwal, Ankush, et al.
Veröffentlicht: (2025) -
Semantic Graph Consistency: Going Beyond Patches for Regularizing Self-Supervised Vision Transformers
von: Devaguptapu, Chaitanya, et al.
Veröffentlicht: (2024) -
Towards a Training Free Approach for 3D Scene Editing
von: Madhavaram, Vivek, et al.
Veröffentlicht: (2024) -
The Percept-V Challenge: Can Multimodal LLMs Crack Simple Perception Problems?
von: Ghosh, Samrajnee, et al.
Veröffentlicht: (2025) -
Find your Needle: Small Object Image Retrieval via Multi-Object Attention Optimization
von: Green, Michael, et al.
Veröffentlicht: (2025)