anguyen8/vision-llms-are-blind: official
Fuente:
Zenodo
Saved in:
| Main Authors: | Pooyan R, Mohammad Reza Taesiri, Logan Bolton, Anh (Totti) Nguyen |
|---|---|
| Format: | Recurso digital |
| Published: |
Zenodo
2026
|
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vision language models are blind: Failing to translate detailed visual features into words
by: Rahmanzadehgervi, Pooyan, et al.
Published: (2024)
by: Rahmanzadehgervi, Pooyan, et al.
Published: (2024)
HoT: Highlighted Chain of Thought for Referencing Supporting Facts from Inputs
by: Nguyen, Tin, et al.
Published: (2025)
by: Nguyen, Tin, et al.
Published: (2025)
SketchVLM: Vision language models can annotate images to explain thoughts and guide users
by: Collins, Brandon, et al.
Published: (2026)
by: Collins, Brandon, et al.
Published: (2026)
B-score: Detecting biases in large language models using response history
by: Vo, An, et al.
Published: (2025)
by: Vo, An, et al.
Published: (2025)
PCNN: Probable-Class Nearest-Neighbor Explanations Improve Fine-Grained Image Classification Accuracy for AIs and Humans
by: Giang, et al.
Published: (2023)
by: Giang, et al.
Published: (2023)
Understanding Generative AI Capabilities in Everyday Image Editing Tasks
by: Taesiri, Mohammad Reza, et al.
Published: (2025)
by: Taesiri, Mohammad Reza, et al.
Published: (2025)
Vision Language Models are Biased
by: Vo, An, et al.
Published: (2025)
by: Vo, An, et al.
Published: (2025)
Improving Zero-Shot Object-Level Change Detection by Incorporating Visual Correspondence
by: Nguyen, Hung Huy, et al.
Published: (2025)
by: Nguyen, Hung Huy, et al.
Published: (2025)
VideoGameBunny: Towards vision assistants for video games
by: Taesiri, Mohammad Reza, et al.
Published: (2024)
by: Taesiri, Mohammad Reza, et al.
Published: (2024)
Allowing humans to interactively guide machines where to look does not always improve human-AI team's classification accuracy
by: Nguyen, Giang, et al.
Published: (2024)
by: Nguyen, Giang, et al.
Published: (2024)
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models
by: Rahmanzadehgervi, Pooyan, et al.
Published: (2024)
by: Rahmanzadehgervi, Pooyan, et al.
Published: (2024)
GlitchBench: Can large multimodal models detect video game glitches?
by: Taesiri, Mohammad Reza, et al.
Published: (2023)
by: Taesiri, Mohammad Reza, et al.
Published: (2023)
Leveraging Habitat Information for Fine-grained Bird Identification
by: Nguyen, Tin, et al.
Published: (2023)
by: Nguyen, Tin, et al.
Published: (2023)
A Report on the llms evaluating the high school questions
by: Jiawei, Zhu, et al.
Published: (2025)
by: Jiawei, Zhu, et al.
Published: (2025)
eXtended Physics Informed Neural Network Method for Fracture Mechanics Problems
by: Lotfalian, Amin, et al.
Published: (2025)
by: Lotfalian, Amin, et al.
Published: (2025)
ASAP-Bilkent/fake-privacy-finetuning-llms: v1
by: Masoud Poorghaffar Aghdam, et al.
Published: (2025)
by: Masoud Poorghaffar Aghdam, et al.
Published: (2025)
Improving Human Verification of LLM Reasoning through Interactive Explanation Interfaces
by: Zhou, Runtao, et al.
Published: (2025)
by: Zhou, Runtao, et al.
Published: (2025)
VideoGameQA-Bench: Evaluating Vision-Language Models for Video Game Quality Assurance
by: Taesiri, Mohammad Reza, et al.
Published: (2025)
by: Taesiri, Mohammad Reza, et al.
Published: (2025)
PageGuide: Browser extension to assist users in navigating a webpage and locating information
by: Nguyen, Tin, et al.
Published: (2026)
by: Nguyen, Tin, et al.
Published: (2026)
Interpretable LLM-based Table Question Answering
by: Nguyen, Giang, et al.
Published: (2024)
by: Nguyen, Giang, et al.
Published: (2024)
PEEB: Part-based Image Classifiers with an Explainable and Editable Language Bottleneck
by: Pham, Thang M., et al.
Published: (2024)
by: Pham, Thang M., et al.
Published: (2024)
GOVERNANÇA NO COMITÊ DE BACIA HIDROGRÁFICA DO BAIXO PARAÍBA DO SUL E ITABAPOANA: efetividade da lei e importância do tema para a agenda pública
by: Maria Eugênia Totti
Published: (2020)
by: Maria Eugênia Totti
Published: (2020)
Técnicas de ensino e resistência cultural: contribuição às práticas do ensino de sociologia1
by: Marcelo Augusto Totti
Published: (2022)
by: Marcelo Augusto Totti
Published: (2022)
Activism and Change Among Puerto Ricans in New York, 1960s and 1970s
by: Xavier F. Totti
Published: (2009)
by: Xavier F. Totti
Published: (2009)
Painless cost control as a central strategy for universal oral health coverage: A critical review with policy guide
by: Mohammad‐Pooyan Jadidfard, et al.
Published: (2024)
by: Mohammad‐Pooyan Jadidfard, et al.
Published: (2024)
Schemora: schema matching via multi-stage recommendation and metadata enrichment using off-the-shelf llms
by: Gungor, Osman Erman, et al.
Published: (2025)
by: Gungor, Osman Erman, et al.
Published: (2025)
Parents' Willingness‐to‐Pay for Fissure Sealant and Fluoride Varnish Therapy in Public and Private Pediatric Clinics
by: Sepideh Samadi, et al.
Published: (2024)
by: Sepideh Samadi, et al.
Published: (2024)
LiteGPT: Large Vision-Language Model for Joint Chest X-ray Localization and Classification Task
by: Le-Duc, Khai, et al.
Published: (2024)
by: Le-Duc, Khai, et al.
Published: (2024)
Toward accessible comics for blind and low vision readers
by: Rigaud, Christophe, et al.
Published: (2024)
by: Rigaud, Christophe, et al.
Published: (2024)
DESSERT official leaflet
by: ENCO srl
Published: (2025)
by: ENCO srl
Published: (2025)
Tula : official guide
Published: (1968)
Published: (1968)
The official museum directory
EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning
by: Kulkarni, Yogesh, et al.
Published: (2025)
by: Kulkarni, Yogesh, et al.
Published: (2025)
VideoSAVi: Self-Aligned Video Language Models without Human Supervision
by: Kulkarni, Yogesh, et al.
Published: (2024)
by: Kulkarni, Yogesh, et al.
Published: (2024)
AVATAR: Reinforcement Learning to See, Hear, and Reason Over Video
by: Kulkarni, Yogesh, et al.
Published: (2025)
by: Kulkarni, Yogesh, et al.
Published: (2025)
An Empirical Study of Accuracy-Robustness Tradeoff and Training Efficiency in Self-Supervised Learning
by: Ghofrani, Fatemeh, et al.
Published: (2025)
by: Ghofrani, Fatemeh, et al.
Published: (2025)
Harmonic generation with topological edge states and electron-electron interaction
by: Pooyan, Siamak, et al.
Published: (2024)
by: Pooyan, Siamak, et al.
Published: (2024)
VideoPASTA: 7K Preference Pairs That Matter for Video-LLM Alignment
by: Kulkarni, Yogesh, et al.
Published: (2025)
by: Kulkarni, Yogesh, et al.
Published: (2025)
Low-energy $^7$Li($n,γ$)$^8$Li and $^7$Be($p,γ$)$^8$B radiative capture reactions within the Skyrme Hartree-Fock approach
by: Nguyen, Le-Anh, et al.
Published: (2022)
by: Nguyen, Le-Anh, et al.
Published: (2022)
Short‐term effect of dressing with Dermaheal ointment in the treatment of diabetic foot ulcer: A double‐blinded randomized controlled clinical trial
by: Pouya Salahi, et al.
Published: (2024)
by: Pouya Salahi, et al.
Published: (2024)
Similar Items
-
Vision language models are blind: Failing to translate detailed visual features into words
by: Rahmanzadehgervi, Pooyan, et al.
Published: (2024) -
HoT: Highlighted Chain of Thought for Referencing Supporting Facts from Inputs
by: Nguyen, Tin, et al.
Published: (2025) -
SketchVLM: Vision language models can annotate images to explain thoughts and guide users
by: Collins, Brandon, et al.
Published: (2026) -
B-score: Detecting biases in large language models using response history
by: Vo, An, et al.
Published: (2025) -
PCNN: Probable-Class Nearest-Neighbor Explanations Improve Fine-Grained Image Classification Accuracy for AIs and Humans
by: Giang, et al.
Published: (2023)