Are Large Vision-Language Models Ready to Guide Blind and Low-Vision Individuals?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Eunki, An, Na Min, Kang, Wan Ju, Kim, Sangryul, Thorne, James, Shim, Hyunjung |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How Blind and Low-Vision Individuals Prefer Large Vision-Language Model-Generated Scene Descriptions
von: An, Na Min, et al.
Veröffentlicht: (2025)
von: An, Na Min, et al.
Veröffentlicht: (2025)
I0T: Embedding Standardization Method Towards Zero Modality Gap
von: An, Na Min, et al.
Veröffentlicht: (2024)
von: An, Na Min, et al.
Veröffentlicht: (2024)
Sightation Counts: Leveraging Sighted User Feedback in Building a BLV-aligned Dataset of Diagram Descriptions
von: Kang, Wan Ju, et al.
Veröffentlicht: (2025)
von: Kang, Wan Ju, et al.
Veröffentlicht: (2025)
Blind to Position, Biased in Language: Probing Mid-Layer Representational Bias in Vision-Language Encoders for Zero-Shot Language-Grounded Spatial Understanding
von: An, Na Min, et al.
Veröffentlicht: (2025)
von: An, Na Min, et al.
Veröffentlicht: (2025)
Interpretable Debiasing of Vision-Language Models for Social Fairness
von: An, Na Min, et al.
Veröffentlicht: (2026)
von: An, Na Min, et al.
Veröffentlicht: (2026)
3D-Aware Vision-Language Models Fine-Tuning with Geometric Distillation
von: Lee, Seonho, et al.
Veröffentlicht: (2025)
von: Lee, Seonho, et al.
Veröffentlicht: (2025)
Classifier-guided CLIP Distillation for Unsupervised Multi-label Classification
von: Kim, Dongseob, et al.
Veröffentlicht: (2025)
von: Kim, Dongseob, et al.
Veröffentlicht: (2025)
Real-Time Long Horizon Air Quality Forecasting via Group-Relative Policy Optimization
von: Kang, Inha, et al.
Veröffentlicht: (2025)
von: Kang, Inha, et al.
Veröffentlicht: (2025)
Revealing Multi-View Hallucination in Large Vision-Language Models
von: Park, Wooje, et al.
Veröffentlicht: (2026)
von: Park, Wooje, et al.
Veröffentlicht: (2026)
Rethinking the Use of Vision Transformers for AI-Generated Image Detection
von: Park, NaHyeon, et al.
Veröffentlicht: (2025)
von: Park, NaHyeon, et al.
Veröffentlicht: (2025)
Self-Supervised Vision Transformers Are Efficient Segmentation Learners for Imperfect Labels
von: Lee, Seungho, et al.
Veröffentlicht: (2024)
von: Lee, Seungho, et al.
Veröffentlicht: (2024)
Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding
von: Min, Kyungmin, et al.
Veröffentlicht: (2024)
von: Min, Kyungmin, et al.
Veröffentlicht: (2024)
Rethinking Data Augmentation for Robust LiDAR Semantic Segmentation in Adverse Weather
von: Park, Junsung, et al.
Veröffentlicht: (2024)
von: Park, Junsung, et al.
Veröffentlicht: (2024)
Label-Augmented Dataset Distillation
von: Kang, Seoungyoon, et al.
Veröffentlicht: (2024)
von: Kang, Seoungyoon, et al.
Veröffentlicht: (2024)
SurgCheck: Do Vision-Language Models Really Look at Images in Surgical VQA?
von: Shin, Jongmin, et al.
Veröffentlicht: (2026)
von: Shin, Jongmin, et al.
Veröffentlicht: (2026)
TextBoost: Boosting Text Encoder for Personalized Text-to-Image Generation
von: Park, NaHyeon, et al.
Veröffentlicht: (2024)
von: Park, NaHyeon, et al.
Veröffentlicht: (2024)
Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models
von: Woo, Sangmin, et al.
Veröffentlicht: (2024)
von: Woo, Sangmin, et al.
Veröffentlicht: (2024)
FALCON: False-Negative Aware Learning of Contrastive Negatives in Vision-Language Alignment
von: Kim, Myunsoo, et al.
Veröffentlicht: (2025)
von: Kim, Myunsoo, et al.
Veröffentlicht: (2025)
Precision matters: Precision-aware ensemble for weakly supervised semantic segmentation
von: Park, Junsung, et al.
Veröffentlicht: (2024)
von: Park, Junsung, et al.
Veröffentlicht: (2024)
Addressing Image Hallucination in Text-to-Image Generation through Factual Image Retrieval
von: Lim, Youngsun, et al.
Veröffentlicht: (2024)
von: Lim, Youngsun, et al.
Veröffentlicht: (2024)
Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding
von: Kang, Seil, et al.
Veröffentlicht: (2025)
von: Kang, Seil, et al.
Veröffentlicht: (2025)
World in a Frame: Understanding Culture Mixing as a New Challenge for Vision-Language Models
von: Kim, Eunsu, et al.
Veröffentlicht: (2025)
von: Kim, Eunsu, et al.
Veröffentlicht: (2025)
Explainable Adversarial-Robust Vision-Language-Action Model for Robotic Manipulation
von: Kim, Ju-Young, et al.
Veröffentlicht: (2025)
von: Kim, Ju-Young, et al.
Veröffentlicht: (2025)
Memory-Efficient Fine-Tuning for Quantized Diffusion Model
von: Ryu, Hyogon, et al.
Veröffentlicht: (2024)
von: Ryu, Hyogon, et al.
Veröffentlicht: (2024)
Grounding Driving VLA via Inverse Kinematics
von: Park, Junsung, et al.
Veröffentlicht: (2026)
von: Park, Junsung, et al.
Veröffentlicht: (2026)
VLind-Bench: Measuring Language Priors in Large Vision-Language Models
von: Lee, Kang-il, et al.
Veröffentlicht: (2024)
von: Lee, Kang-il, et al.
Veröffentlicht: (2024)
V-DPO: Mitigating Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization
von: Xie, Yuxi, et al.
Veröffentlicht: (2024)
von: Xie, Yuxi, et al.
Veröffentlicht: (2024)
No Thing, Nothing: Highlighting Safety-Critical Classes for Robust LiDAR Semantic Segmentation in Adverse Weather
von: Park, Junsung, et al.
Veröffentlicht: (2025)
von: Park, Junsung, et al.
Veröffentlicht: (2025)
Erase Persona, Forget Lore: Benchmarking Multimodal Copyright Unlearning in Large Vision Language Models
von: Kwon, JuneHyoung, et al.
Veröffentlicht: (2026)
von: Kwon, JuneHyoung, et al.
Veröffentlicht: (2026)
Visually Guided Decoding: Gradient-Free Hard Prompt Inversion with Language Models
von: Kim, Donghoon, et al.
Veröffentlicht: (2025)
von: Kim, Donghoon, et al.
Veröffentlicht: (2025)
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
von: Kim, Jinyeong, et al.
Veröffentlicht: (2025)
von: Kim, Jinyeong, et al.
Veröffentlicht: (2025)
HoliSafe: Holistic Safety Benchmarking and Modeling for Vision-Language Model
von: Lee, Youngwan, et al.
Veröffentlicht: (2025)
von: Lee, Youngwan, et al.
Veröffentlicht: (2025)
SIA: Enhancing Safety via Intent Awareness for Vision-Language Models
von: Na, Youngjin, et al.
Veröffentlicht: (2025)
von: Na, Youngjin, et al.
Veröffentlicht: (2025)
Learning to Insert [PAUSE] Tokens for Better Reasoning
von: Kim, Eunki, et al.
Veröffentlicht: (2025)
von: Kim, Eunki, et al.
Veröffentlicht: (2025)
Sampling Bag of Views for Open-Vocabulary Object Detection
von: Choi, Hojun, et al.
Veröffentlicht: (2024)
von: Choi, Hojun, et al.
Veröffentlicht: (2024)
Evaluating Image Hallucination in Text-to-Image Generation with Question-Answering
von: Lim, Youngsun, et al.
Veröffentlicht: (2024)
von: Lim, Youngsun, et al.
Veröffentlicht: (2024)
Unveiling the Response of Large Vision-Language Models to Visually Absent Tokens
von: Kim, Sohee, et al.
Veröffentlicht: (2025)
von: Kim, Sohee, et al.
Veröffentlicht: (2025)
Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training
von: Zhang, Xinsong, et al.
Veröffentlicht: (2025)
von: Zhang, Xinsong, et al.
Veröffentlicht: (2025)
Continual Vision-and-Language Navigation
von: Jeong, Seongjun, et al.
Veröffentlicht: (2024)
von: Jeong, Seongjun, et al.
Veröffentlicht: (2024)
IWP: Token Pruning as Implicit Weight Pruning in Large Vision Language Models
von: Lee, Dong-Jae, et al.
Veröffentlicht: (2026)
von: Lee, Dong-Jae, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
How Blind and Low-Vision Individuals Prefer Large Vision-Language Model-Generated Scene Descriptions
von: An, Na Min, et al.
Veröffentlicht: (2025) -
I0T: Embedding Standardization Method Towards Zero Modality Gap
von: An, Na Min, et al.
Veröffentlicht: (2024) -
Sightation Counts: Leveraging Sighted User Feedback in Building a BLV-aligned Dataset of Diagram Descriptions
von: Kang, Wan Ju, et al.
Veröffentlicht: (2025) -
Blind to Position, Biased in Language: Probing Mid-Layer Representational Bias in Vision-Language Encoders for Zero-Shot Language-Grounded Spatial Understanding
von: An, Na Min, et al.
Veröffentlicht: (2025) -
Interpretable Debiasing of Vision-Language Models for Social Fairness
von: An, Na Min, et al.
Veröffentlicht: (2026)