VLM's Eye Examination: Instruct and Inspect Visual Competency of Vision Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Hyeon-Woo, Nam, Ye-Bin, Moon, Choi, Wonseok, Hyun, Lee, Oh, Tae-Hyun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
BEAF: Observing BEfore-AFter Changes to Evaluate Hallucination in Vision-language Models
di: Ye-Bin, Moon, et al.
Pubblicazione: (2024)
di: Ye-Bin, Moon, et al.
Pubblicazione: (2024)
SYNAuG: Exploiting Synthetic Data for Data Imbalance Problems
di: Ye-Bin, Moon, et al.
Pubblicazione: (2023)
di: Ye-Bin, Moon, et al.
Pubblicazione: (2023)
Aligning What Vision-Language Models See and Perceive with Adaptive Information Flow
di: Liu, Chengxin, et al.
Pubblicazione: (2026)
di: Liu, Chengxin, et al.
Pubblicazione: (2026)
Patch-wise Retrieval: A Bag of Practical Techniques for Instance-level Matching
di: Choi, Wonseok, et al.
Pubblicazione: (2025)
di: Choi, Wonseok, et al.
Pubblicazione: (2025)
Early Failure Detection and Intervention in Video Diffusion Models
di: Byung-Ki, Kwon, et al.
Pubblicazione: (2026)
di: Byung-Ki, Kwon, et al.
Pubblicazione: (2026)
Learning Correlation-aware Aleatoric Uncertainty for 3D Hand Pose Estimation
di: Chae-Yeon, Lee, et al.
Pubblicazione: (2025)
di: Chae-Yeon, Lee, et al.
Pubblicazione: (2025)
RetouchLLM: Training-free Code-based Image Retouching with Vision Language Models
di: Ye-Bin, Moon, et al.
Pubblicazione: (2025)
di: Ye-Bin, Moon, et al.
Pubblicazione: (2025)
AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models
di: Sung-Bin, Kim, et al.
Pubblicazione: (2024)
di: Sung-Bin, Kim, et al.
Pubblicazione: (2024)
CLAY: Conditional Visual Similarity Modulation in Vision-Language Embedding Space
di: Lim, Sohwi, et al.
Pubblicazione: (2026)
di: Lim, Sohwi, et al.
Pubblicazione: (2026)
Scratching Visual Transformer's Back with Uniform Attention
di: Hyeon-Woo, Nam, et al.
Pubblicazione: (2022)
di: Hyeon-Woo, Nam, et al.
Pubblicazione: (2022)
MultiTalk: Enhancing 3D Talking Head Generation Across Languages with Multilingual Video Dataset
di: Sung-Bin, Kim, et al.
Pubblicazione: (2024)
di: Sung-Bin, Kim, et al.
Pubblicazione: (2024)
M2SFormer: Multi-Spectral and Multi-Scale Attention with Edge-Aware Difficulty Guidance for Image Forgery Localization
di: Nam, Ju-Hyeon, et al.
Pubblicazione: (2025)
di: Nam, Ju-Hyeon, et al.
Pubblicazione: (2025)
VSC: Visual Search Compositional Text-to-Image Diffusion Model
di: Dat, Do Huu, et al.
Pubblicazione: (2025)
di: Dat, Do Huu, et al.
Pubblicazione: (2025)
Enhancing Speech-Driven 3D Facial Animation with Audio-Visual Guidance from Lip Reading Expert
di: EunGi, Han, et al.
Pubblicazione: (2024)
di: EunGi, Han, et al.
Pubblicazione: (2024)
Perceptually Accurate 3D Talking Head Generation: New Definitions, Speech-Mesh Representation, and Evaluation Metrics
di: Chae-Yeon, Lee, et al.
Pubblicazione: (2025)
di: Chae-Yeon, Lee, et al.
Pubblicazione: (2025)
InspectVLM: Unified in Theory, Unreliable in Practice
di: Wallace, Conor, et al.
Pubblicazione: (2025)
di: Wallace, Conor, et al.
Pubblicazione: (2025)
DarkQA: Benchmarking Vision-Language Models on Visual-Primitive Question Answering in Low-Light Indoor Scenes
di: Park, Yohan, et al.
Pubblicazione: (2025)
di: Park, Yohan, et al.
Pubblicazione: (2025)
Bi-MCQ: Reformulating Vision-Language Alignment for Negation Understanding
di: Kim, Tae Hun, et al.
Pubblicazione: (2026)
di: Kim, Tae Hun, et al.
Pubblicazione: (2026)
UM-Depth : Uncertainty Masked Self-Supervised Monocular Depth Estimation with Visual Odometry
di: Um, Tae-Wook, et al.
Pubblicazione: (2025)
di: Um, Tae-Wook, et al.
Pubblicazione: (2025)
Sound2Vision: Generating Diverse Visuals from Audio through Cross-Modal Latent Alignment
di: Sung-Bin, Kim, et al.
Pubblicazione: (2024)
di: Sung-Bin, Kim, et al.
Pubblicazione: (2024)
PAVAS: Physics-Aware Video-to-Audio Synthesis
di: Hyun-Bin, Oh, et al.
Pubblicazione: (2025)
di: Hyun-Bin, Oh, et al.
Pubblicazione: (2025)
Balancing Efficiency and Quality: MoEISR for Arbitrary-Scale Image Super-Resolution
di: Oh, Young Jae, et al.
Pubblicazione: (2023)
di: Oh, Young Jae, et al.
Pubblicazione: (2023)
MemBench: Memorized Image Trigger Prompt Dataset for Diffusion Models
di: Hong, Chunsan, et al.
Pubblicazione: (2024)
di: Hong, Chunsan, et al.
Pubblicazione: (2024)
Contextualized Visual Personalization in Vision-Language Models
di: Oh, Yeongtak, et al.
Pubblicazione: (2026)
di: Oh, Yeongtak, et al.
Pubblicazione: (2026)
Eyes on VLM: Benchmarking Gaze Following and Social Gaze Prediction in Vision Language Models
di: Wang, Hengfei, et al.
Pubblicazione: (2026)
di: Wang, Hengfei, et al.
Pubblicazione: (2026)
Do Vision Encoders Truly Explain Object Hallucination?: Mitigating Object Hallucination via Simple Fine-Grained CLIPScore
di: Oh, Hongseok, et al.
Pubblicazione: (2025)
di: Oh, Hongseok, et al.
Pubblicazione: (2025)
VoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language Models
di: Sung-Bin, Kim, et al.
Pubblicazione: (2025)
di: Sung-Bin, Kim, et al.
Pubblicazione: (2025)
Auxiliary Descriptive Knowledge for Few-Shot Adaptation of Vision-Language Model
di: Lee, SuBeen, et al.
Pubblicazione: (2025)
di: Lee, SuBeen, et al.
Pubblicazione: (2025)
WWW: A Unified Framework for Explaining What, Where and Why of Neural Networks by Interpretation of Neuron Concepts
di: Ahn, Yong Hyun, et al.
Pubblicazione: (2024)
di: Ahn, Yong Hyun, et al.
Pubblicazione: (2024)
Mask-Free Neuron Concept Annotation for Interpreting Neural Networks in Medical Domain
di: Kim, Hyeon Bae, et al.
Pubblicazione: (2024)
di: Kim, Hyeon Bae, et al.
Pubblicazione: (2024)
GaussExplorer: 3D Gaussian Splatting for Embodied Exploration and Reasoning
di: Yu-Ji, Kim, et al.
Pubblicazione: (2026)
di: Yu-Ji, Kim, et al.
Pubblicazione: (2026)
Learning-based Axial Video Motion Magnification
di: Byung-Ki, Kwon, et al.
Pubblicazione: (2023)
di: Byung-Ki, Kwon, et al.
Pubblicazione: (2023)
EO-VLM: VLM-Guided Energy Overload Attacks on Vision Models
di: Seo, Minjae, et al.
Pubblicazione: (2025)
di: Seo, Minjae, et al.
Pubblicazione: (2025)
InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models
di: Wei, Cong, et al.
Pubblicazione: (2024)
di: Wei, Cong, et al.
Pubblicazione: (2024)
VICI: VLM-Instructed Cross-view Image-localisation
di: Zhang, Xiaohan, et al.
Pubblicazione: (2025)
di: Zhang, Xiaohan, et al.
Pubblicazione: (2025)
HiCM$^2$: Hierarchical Compact Memory Modeling for Dense Video Captioning
di: Kim, Minkuk, et al.
Pubblicazione: (2024)
di: Kim, Minkuk, et al.
Pubblicazione: (2024)
Retrieval-Augmented Natural Language Reasoning for Explainable Visual Question Answering
di: Lim, Su Hyeon, et al.
Pubblicazione: (2024)
di: Lim, Su Hyeon, et al.
Pubblicazione: (2024)
From Behavioral Performance to Internal Competence: Interpreting Vision-Language Models with VLM-Lens
di: Sheta, Hala, et al.
Pubblicazione: (2025)
di: Sheta, Hala, et al.
Pubblicazione: (2025)
RelationVLM: Making Large Vision-Language Models Understand Visual Relations
di: Huang, Zhipeng, et al.
Pubblicazione: (2024)
di: Huang, Zhipeng, et al.
Pubblicazione: (2024)
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
di: Zhang, Yuan, et al.
Pubblicazione: (2024)
di: Zhang, Yuan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
BEAF: Observing BEfore-AFter Changes to Evaluate Hallucination in Vision-language Models
di: Ye-Bin, Moon, et al.
Pubblicazione: (2024) -
SYNAuG: Exploiting Synthetic Data for Data Imbalance Problems
di: Ye-Bin, Moon, et al.
Pubblicazione: (2023) -
Aligning What Vision-Language Models See and Perceive with Adaptive Information Flow
di: Liu, Chengxin, et al.
Pubblicazione: (2026) -
Patch-wise Retrieval: A Bag of Practical Techniques for Instance-level Matching
di: Choi, Wonseok, et al.
Pubblicazione: (2025) -
Early Failure Detection and Intervention in Video Diffusion Models
di: Byung-Ki, Kwon, et al.
Pubblicazione: (2026)