Large Language Models Facilitate Vision Reflection in Image Classification
Fuente:
arXiv
Saved in:
| Main Authors: | An, Guoyuan, Kim, JaeYoon, Yoon, SungEui |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Accurate and Fast Pixel Retrieval with Spatial and Uncertainty Aware Hypergraph Diffusion
by: An, Guoyuan, et al.
Published: (2024)
by: An, Guoyuan, et al.
Published: (2024)
OpenSlot: Mixed Open-Set Recognition with Object-Centric Learning
by: Yin, Xu, et al.
Published: (2024)
by: Yin, Xu, et al.
Published: (2024)
Towards Test-time Efficient Visual Place Recognition via Asymmetric Query Processing
by: Kim, Jaeyoon, et al.
Published: (2025)
by: Kim, Jaeyoon, et al.
Published: (2025)
Enhancing Visual Re-ranking through Denoising Nearest Neighbor Graph via Continuous CRF
by: Kim, Jaeyoon, et al.
Published: (2024)
by: Kim, Jaeyoon, et al.
Published: (2024)
Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models
by: Hong, Sujung, et al.
Published: (2026)
by: Hong, Sujung, et al.
Published: (2026)
No Caption, No Problem: Caption-Free Membership Inference via Model-Fitted Embeddings
by: Jeon, Joonsung, et al.
Published: (2026)
by: Jeon, Joonsung, et al.
Published: (2026)
VERIA: Verification-Centric Multimodal Instance Augmentation for Long-Tailed 3D Object Detection
by: Lee, Jumin, et al.
Published: (2026)
by: Lee, Jumin, et al.
Published: (2026)
MorphGS: Morphology-Adaptive Articulated 3D Motion Transfer from Videos
by: Kim, Taeyeon, et al.
Published: (2026)
by: Kim, Taeyeon, et al.
Published: (2026)
GLINT: Modeling Scene-Scale Transparency via Gaussian Radiance Transport
by: Na, Youngju, et al.
Published: (2026)
by: Na, Youngju, et al.
Published: (2026)
Extending Segment Anything Model into Auditory and Temporal Dimensions for Audio-Visual Segmentation
by: Seon, Juhyeong, et al.
Published: (2024)
by: Seon, Juhyeong, et al.
Published: (2024)
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images
by: Ro, Juneyoung, et al.
Published: (2025)
by: Ro, Juneyoung, et al.
Published: (2025)
Visual-RRT: Finding Paths toward Visual-Goals via Differentiable Rendering
by: Lee, Sebin, et al.
Published: (2026)
by: Lee, Sebin, et al.
Published: (2026)
Pinpointing Trigger Moment for Grounded Video QA: Enhancing Spatio-temporal Grounding in Multimodal Large Language Models
by: Seo, Jinhwan, et al.
Published: (2025)
by: Seo, Jinhwan, et al.
Published: (2025)
Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image Generator
by: Shin, Chaehun, et al.
Published: (2024)
by: Shin, Chaehun, et al.
Published: (2024)
Single Image Reflection Removal with Patch Reflectance Prior
by: Han, Dongshen, et al.
Published: (2023)
by: Han, Dongshen, et al.
Published: (2023)
ReflectCAP: Detailed Image Captioning with Reflective Memory
by: Min, Kyungmin, et al.
Published: (2026)
by: Min, Kyungmin, et al.
Published: (2026)
Generalizable Person Re-identification via Balancing Alignment and Uniformity
by: Cho, Yoonki, et al.
Published: (2024)
by: Cho, Yoonki, et al.
Published: (2024)
SemCity: Semantic Scene Generation with Triplane Diffusion
by: Lee, Jumin, et al.
Published: (2024)
by: Lee, Jumin, et al.
Published: (2024)
VLind-Bench: Measuring Language Priors in Large Vision-Language Models
by: Lee, Kang-il, et al.
Published: (2024)
by: Lee, Kang-il, et al.
Published: (2024)
AdvPaint: Protecting Images from Inpainting Manipulation via Adversarial Attention Disruption
by: Jeon, Joonsung, et al.
Published: (2025)
by: Jeon, Joonsung, et al.
Published: (2025)
Escaping Plato's Cave: JAM for Aligning Independently Trained Vision and Language Models
by: Yoon, Lauren Hyoseo, et al.
Published: (2025)
by: Yoon, Lauren Hyoseo, et al.
Published: (2025)
Focus Matters: Phase-Aware Suppression for Hallucination in Vision-Language Models
by: Kim, Sohyeon, et al.
Published: (2026)
by: Kim, Sohyeon, et al.
Published: (2026)
ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language Models
by: Sung, Yi-Lin, et al.
Published: (2023)
by: Sung, Yi-Lin, et al.
Published: (2023)
Interactive Text-to-Image Retrieval with Large Language Models: A Plug-and-Play Approach
by: Lee, Saehyung, et al.
Published: (2024)
by: Lee, Saehyung, et al.
Published: (2024)
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
by: Jung, Mingi, et al.
Published: (2025)
by: Jung, Mingi, et al.
Published: (2025)
Vision-Language Models Do Not Understand Negation
by: Alhamoud, Kumail, et al.
Published: (2025)
by: Alhamoud, Kumail, et al.
Published: (2025)
Regularizing Dynamic Radiance Fields with Kinematic Fields
by: Im, Woobin, et al.
Published: (2024)
by: Im, Woobin, et al.
Published: (2024)
Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models
by: Yoon, Eunseop, et al.
Published: (2025)
by: Yoon, Eunseop, et al.
Published: (2025)
UFORecon: Generalizable Sparse-View Surface Reconstruction from Arbitrary and UnFavOrable Sets
by: Na, Youngju, et al.
Published: (2024)
by: Na, Youngju, et al.
Published: (2024)
Beyond the Patch: Exploring Vulnerabilities of Visuomotor Policies via Viewpoint-Consistent 3D Adversarial Object
by: Lee, Chanmi, et al.
Published: (2026)
by: Lee, Chanmi, et al.
Published: (2026)
GLYPH-SR: Can We Achieve Both High-Quality Image Super-Resolution and High-Fidelity Text Recovery via VLM-guided Latent Diffusion Model?
by: Sung, Mingyu, et al.
Published: (2025)
by: Sung, Mingyu, et al.
Published: (2025)
Your Super Resolution Model is not Enough for Tackling Real-World Scenarios
by: Yoon, Dongsik, et al.
Published: (2025)
by: Yoon, Dongsik, et al.
Published: (2025)
Radiometrically Consistent Gaussian Surfels for Inverse Rendering
by: Han, Kyu Beom, et al.
Published: (2026)
by: Han, Kyu Beom, et al.
Published: (2026)
A Spatio-Temporal Representation Learning as an Alternative to Traditional Glosses in Sign Language Translation and Production
by: Hwang, Eui Jun, et al.
Published: (2024)
by: Hwang, Eui Jun, et al.
Published: (2024)
From Prompts to Deployment: Auto-Curated Domain-Specific Dataset Generation via Diffusion Models
by: Yoon, Dongsik, et al.
Published: (2026)
by: Yoon, Dongsik, et al.
Published: (2026)
Finding Meaning in Points: Weakly Supervised Semantic Segmentation for Event Cameras
by: Cho, Hoonhee, et al.
Published: (2024)
by: Cho, Hoonhee, et al.
Published: (2024)
Pose-free 3D Gaussian splatting via shape-ray estimation
by: Na, Youngju, et al.
Published: (2025)
by: Na, Youngju, et al.
Published: (2025)
Contextualized Visual Personalization in Vision-Language Models
by: Oh, Yeongtak, et al.
Published: (2026)
by: Oh, Yeongtak, et al.
Published: (2026)
Zero-Shot Fine-Grained Image Classification Using Large Vision-Language Models
by: Atabuzzaman, Md., et al.
Published: (2025)
by: Atabuzzaman, Md., et al.
Published: (2025)
TransMed: Large Language Models Enhance Vision Transformer for Biomedical Image Classification
by: Zheng, Kaipeng, et al.
Published: (2023)
by: Zheng, Kaipeng, et al.
Published: (2023)
Similar Items
-
Accurate and Fast Pixel Retrieval with Spatial and Uncertainty Aware Hypergraph Diffusion
by: An, Guoyuan, et al.
Published: (2024) -
OpenSlot: Mixed Open-Set Recognition with Object-Centric Learning
by: Yin, Xu, et al.
Published: (2024) -
Towards Test-time Efficient Visual Place Recognition via Asymmetric Query Processing
by: Kim, Jaeyoon, et al.
Published: (2025) -
Enhancing Visual Re-ranking through Denoising Nearest Neighbor Graph via Continuous CRF
by: Kim, Jaeyoon, et al.
Published: (2024) -
Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models
by: Hong, Sujung, et al.
Published: (2026)