Long-Form Answers to Visual Questions from Blind and Low Vision People
Fuente:
arXiv
Guardado en:
| Autores principales: | Huh, Mina, Xu, Fangyuan, Peng, Yi-Hao, Chen, Chongyan, Murugu, Hansika, Gurari, Danna, Choi, Eunsol, Pavel, Amy |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DesignChecker: Visual Design Support for Blind and Low Vision Web Developers
por: Huh, Mina, et al.
Publicado: (2024)
por: Huh, Mina, et al.
Publicado: (2024)
Acknowledging Focus Ambiguity in Visual Questions
por: Chen, Chongyan, et al.
Publicado: (2025)
por: Chen, Chongyan, et al.
Publicado: (2025)
An Evaluation of GPT-4V and Gemini in Online VQA
por: Liu, Mengchen, et al.
Publicado: (2023)
por: Liu, Mengchen, et al.
Publicado: (2023)
Fully Authentic Visual Question Answering Dataset from Online Communities
por: Chen, Chongyan, et al.
Publicado: (2023)
por: Chen, Chongyan, et al.
Publicado: (2023)
Understanding Retrieval Augmentation for Long-Form Question Answering
por: Chen, Hung-Ting, et al.
Publicado: (2023)
por: Chen, Hung-Ting, et al.
Publicado: (2023)
"Before, I Asked My Mom, Now I Ask ChatGPT": Visual Privacy Management with Generative AI for Blind and Low-Vision People
por: Sharma, Tanusree, et al.
Publicado: (2025)
por: Sharma, Tanusree, et al.
Publicado: (2025)
RefreshKV: Updating Small KV Cache During Long-form Generation
por: Xu, Fangyuan, et al.
Publicado: (2024)
por: Xu, Fangyuan, et al.
Publicado: (2024)
Collecting Consistently High Quality Object Tracks with Minimal Human Involvement by Using Self-Supervised Learning to Detect Tracker Errors
por: Anjum, Samreen, et al.
Publicado: (2024)
por: Anjum, Samreen, et al.
Publicado: (2024)
BIV-Priv-Seg: Locating Private Content in Images Taken by People With Visual Impairments
por: Tseng, Yu-Yun, et al.
Publicado: (2024)
por: Tseng, Yu-Yun, et al.
Publicado: (2024)
PartStickers: Generating Parts of Objects for Rapid Prototyping
por: Zhou, Mo, et al.
Publicado: (2025)
por: Zhou, Mo, et al.
Publicado: (2025)
Lotus: Creating Short Videos From Long Videos With Abstractive and Extractive Summarization
por: Barua, Aadit, et al.
Publicado: (2025)
por: Barua, Aadit, et al.
Publicado: (2025)
Guiding Multimodal Large Language Models with Blind and Low Vision People Visual Questions for Proactive Visual Interpretations
por: Penuela, Ricardo Gonzalez, et al.
Publicado: (2025)
por: Penuela, Ricardo Gonzalez, et al.
Publicado: (2025)
KIWI: A Dataset of Knowledge-Intensive Writing Instructions for Answering Research Questions
por: Xu, Fangyuan, et al.
Publicado: (2024)
por: Xu, Fangyuan, et al.
Publicado: (2024)
Making Short-Form Videos Accessible with Hierarchical Video Summaries
por: Van Daele, Tess, et al.
Publicado: (2024)
por: Van Daele, Tess, et al.
Publicado: (2024)
Contrastive Learning to Improve Retrieval for Real-world Fact Checking
por: Sriram, Aniruddh, et al.
Publicado: (2024)
por: Sriram, Aniruddh, et al.
Publicado: (2024)
Logit-Based Losses Limit the Effectiveness of Feature Knowledge Distillation
por: Cooper, Nicholas, et al.
Publicado: (2025)
por: Cooper, Nicholas, et al.
Publicado: (2025)
Design Considerations for Automatic Musical Soundscapes of Visual Art for People with Blindness or Low Vision
por: Krol, Stephen James, et al.
Publicado: (2024)
por: Krol, Stephen James, et al.
Publicado: (2024)
Interpreting COVID Lateral Flow Tests' Results with Foundation Models
por: Pandey, Stuti, et al.
Publicado: (2024)
por: Pandey, Stuti, et al.
Publicado: (2024)
Generating Contextually-Relevant Navigation Instructions for Blind and Low Vision People
por: Merchant, Zain, et al.
Publicado: (2024)
por: Merchant, Zain, et al.
Publicado: (2024)
RVR: Retrieve-Verify-Retrieve for Comprehensive Question Answering
por: Qian, Deniz, et al.
Publicado: (2026)
por: Qian, Deniz, et al.
Publicado: (2026)
Co-Designing Multimodal Systems for Accessible Asynchronous Dance Instruction
por: Das, Ujjaini, et al.
Publicado: (2025)
por: Das, Ujjaini, et al.
Publicado: (2025)
No Single Best Model for Diversity: Learning a Router for Sample Diversity
por: Liu, Yuhan, et al.
Publicado: (2026)
por: Liu, Yuhan, et al.
Publicado: (2026)
Unseen City Canvases: Exploring Blind and Low Vision People's Perspectives on Urban and Public Art Accessibility
por: Jiang, Lucy, et al.
Publicado: (2026)
por: Jiang, Lucy, et al.
Publicado: (2026)
A Multi-Modal Foundation Model to Assist People with Blindness and Low Vision in Environmental Interaction
por: Hao, Yu, et al.
Publicado: (2023)
por: Hao, Yu, et al.
Publicado: (2023)
A Light and Smart Wearable Platform with Multimodal Foundation Model for Enhanced Spatial Reasoning in People with Blindness and Low Vision
por: Magay, Alexey, et al.
Publicado: (2025)
por: Magay, Alexey, et al.
Publicado: (2025)
SPIN: Hierarchical Segmentation with Subpart Granularity in Natural Images
por: Myers-Dean, Josh, et al.
Publicado: (2024)
por: Myers-Dean, Josh, et al.
Publicado: (2024)
Vid2Coach: Transforming How-To Videos into Task Assistants
por: Huh, Mina, et al.
Publicado: (2025)
por: Huh, Mina, et al.
Publicado: (2025)
BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users
por: Cheng, Wanyin, et al.
Publicado: (2025)
por: Cheng, Wanyin, et al.
Publicado: (2025)
Modeling Future Conversation Turns to Teach LLMs to Ask Clarifying Questions
por: Zhang, Michael J. Q., et al.
Publicado: (2024)
por: Zhang, Michael J. Q., et al.
Publicado: (2024)
RAVEN: Realtime Accessibility in Virtual ENvironments for Blind and Low-Vision People
por: Cao, Xinyun, et al.
Publicado: (2025)
por: Cao, Xinyun, et al.
Publicado: (2025)
Long-Form Information Alignment Evaluation Beyond Atomic Facts
por: Zheng, Danna, et al.
Publicado: (2025)
por: Zheng, Danna, et al.
Publicado: (2025)
Accessible Nonverbal Cues to Support Conversations in VR for Blind and Low Vision People
por: Jung, Crescentia, et al.
Publicado: (2024)
por: Jung, Crescentia, et al.
Publicado: (2024)
Accessible Data Access and Analysis by People who are Blind or Have Low Vision
por: Reinders, Samuel, et al.
Publicado: (2025)
por: Reinders, Samuel, et al.
Publicado: (2025)
TADA: Making Node-link Diagrams Accessible to Blind and Low-Vision People
por: Zhao, Yichun, et al.
Publicado: (2023)
por: Zhao, Yichun, et al.
Publicado: (2023)
Automatic Question-Answer Generation for Long-Tail Knowledge
por: Kumar, Rohan, et al.
Publicado: (2024)
por: Kumar, Rohan, et al.
Publicado: (2024)
Essential, Yet Overlooked: Identity Verification Barriers for Blind and Low Vision People in Government Services
por: Oommen, Ryan John, et al.
Publicado: (2026)
por: Oommen, Ryan John, et al.
Publicado: (2026)
VideoDiff: Human-AI Video Co-Creation with Alternatives
por: Huh, Mina, et al.
Publicado: (2025)
por: Huh, Mina, et al.
Publicado: (2025)
Towards Understanding the Use of MLLM-Enabled Applications for Visual Interpretation by Blind and Low Vision People
por: Penuela, Ricardo E. Gonzalez, et al.
Publicado: (2025)
por: Penuela, Ricardo E. Gonzalez, et al.
Publicado: (2025)
Exploring the Use of VLMs for Navigation Assistance for People with Blindness and Low Vision
por: Li, Yu, et al.
Publicado: (2026)
por: Li, Yu, et al.
Publicado: (2026)
Ask Questions with Double Hints: Visual Question Generation with Answer-awareness and Region-reference
por: Shen, Kai, et al.
Publicado: (2024)
por: Shen, Kai, et al.
Publicado: (2024)
Ejemplares similares
-
DesignChecker: Visual Design Support for Blind and Low Vision Web Developers
por: Huh, Mina, et al.
Publicado: (2024) -
Acknowledging Focus Ambiguity in Visual Questions
por: Chen, Chongyan, et al.
Publicado: (2025) -
An Evaluation of GPT-4V and Gemini in Online VQA
por: Liu, Mengchen, et al.
Publicado: (2023) -
Fully Authentic Visual Question Answering Dataset from Online Communities
por: Chen, Chongyan, et al.
Publicado: (2023) -
Understanding Retrieval Augmentation for Long-Form Question Answering
por: Chen, Hung-Ting, et al.
Publicado: (2023)