WorldScribe: Towards Context-Aware Live Visual Descriptions
Fuente:
arXiv
Saved in:
| Main Authors: | Chang, Ruei-Che, Liu, Yuxuan, Guo, Anhong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EditScribe: Non-Visual Image Editing with Natural Language Verification Loops
by: Chang, Ruei-Che, et al.
Published: (2024)
by: Chang, Ruei-Che, et al.
Published: (2024)
Probing the Gaps in ChatGPT Live Video Chat for Real-World Assistance for People who are Blind or Visually Impaired
by: Chang, Ruei-Che, et al.
Published: (2025)
by: Chang, Ruei-Che, et al.
Published: (2025)
TouchScribe: Augmenting Non-Visual Hand-Object Interactions with Automated Live Visual Descriptions
by: Chang, Ruei-Che, et al.
Published: (2026)
by: Chang, Ruei-Che, et al.
Published: (2026)
StateScribe: Towards Accessible Change Awareness Across Real-World Revisits
by: Chang, Ruei-Che, et al.
Published: (2026)
by: Chang, Ruei-Che, et al.
Published: (2026)
Not There Yet: Evaluating Vision Language Models in Simulating the Visual Perception of People with Low Vision
by: Natalie, Rosiana, et al.
Published: (2025)
by: Natalie, Rosiana, et al.
Published: (2025)
Audio Description Customization
by: Natalie, Rosiana, et al.
Published: (2024)
by: Natalie, Rosiana, et al.
Published: (2024)
ProgramAlly: Creating Custom Visual Access Programs via Multi-Modal End-User Programming
by: Herskovitz, Jaylin, et al.
Published: (2024)
by: Herskovitz, Jaylin, et al.
Published: (2024)
ContextAgent: Context-Aware Proactive LLM Agents with Open-World Sensory Perceptions
by: Yang, Bufang, et al.
Published: (2025)
by: Yang, Bufang, et al.
Published: (2025)
A Custom-Built Ambient Scribe Reduces Cognitive Load and Documentation Burden for Telehealth Clinicians
by: Morse, Justin, et al.
Published: (2025)
by: Morse, Justin, et al.
Published: (2025)
CALLM: Understanding Cancer Survivors' Emotions and Intervention Opportunities via Mobile Diaries and Context-Aware Language Models
by: Wang, Zhiyuan, et al.
Published: (2025)
by: Wang, Zhiyuan, et al.
Published: (2025)
ProAgent: Harnessing On-Demand Sensory Contexts for Proactive LLM Agent Systems in the Wild
by: Yang, Bufang, et al.
Published: (2025)
by: Yang, Bufang, et al.
Published: (2025)
Actions Speak Louder than Words: Agent Decisions Reveal Implicit Biases in Language Models
by: Li, Yuxuan, et al.
Published: (2025)
by: Li, Yuxuan, et al.
Published: (2025)
Pragmatics Meets Culture: Culturally-adapted Artwork Description Generation and Evaluation
by: Zhao, Lingjun, et al.
Published: (2026)
by: Zhao, Lingjun, et al.
Published: (2026)
DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs
by: Tian, Ye, et al.
Published: (2025)
by: Tian, Ye, et al.
Published: (2025)
Totalitarian Technics: The Hidden Cost of AI Scribes in Healthcare
by: Brosnahan, Hugh
Published: (2025)
by: Brosnahan, Hugh
Published: (2025)
SoundShift: Exploring Sound Manipulations for Accessible Mixed-Reality Awareness
by: Chang, Ruei-Che, et al.
Published: (2024)
by: Chang, Ruei-Che, et al.
Published: (2024)
From Context to Action: Analysis of the Impact of State Representation and Context on the Generalization of Multi-Turn Web Navigation Agents
by: Tiwary, Nalin, et al.
Published: (2024)
by: Tiwary, Nalin, et al.
Published: (2024)
Visualization Literacy of Multimodal Large Language Models: A Comparative Study
by: Li, Zhimin, et al.
Published: (2024)
by: Li, Zhimin, et al.
Published: (2024)
VIDEE: Visual and Interactive Decomposition, Execution, and Evaluation of Text Analytics with Intelligent Agents
by: Lee, Sam Yu-Te, et al.
Published: (2025)
by: Lee, Sam Yu-Te, et al.
Published: (2025)
Inclusion Arena: An Open Platform for Evaluating Large Foundation Models with Real-World Apps
by: Wang, Kangyu, et al.
Published: (2025)
by: Wang, Kangyu, et al.
Published: (2025)
Layout Agnostic Human Activity Recognition in Smart Homes through Textual Descriptions Of Sensor Triggers (TDOST)
by: Thukral, Megha, et al.
Published: (2024)
by: Thukral, Megha, et al.
Published: (2024)
On the Role of Artificial Intelligence in Human-Machine Symbiosis
by: Chang, Ching-Chun, et al.
Published: (2026)
by: Chang, Ching-Chun, et al.
Published: (2026)
Empathic Prompting: Non-Verbal Context Integration for Multimodal LLM Conversations
by: Stacchio, Lorenzo, et al.
Published: (2025)
by: Stacchio, Lorenzo, et al.
Published: (2025)
Position: Towards Bidirectional Human-AI Alignment
by: Shen, Hua, et al.
Published: (2024)
by: Shen, Hua, et al.
Published: (2024)
Can LLMs Generate Visualizations with Dataless Prompts?
by: Coelho, Darius, et al.
Published: (2024)
by: Coelho, Darius, et al.
Published: (2024)
Performance Gains of LLMs With Humans in a World of LLMs Versus Humans
by: McCullum, Lucas, et al.
Published: (2025)
by: McCullum, Lucas, et al.
Published: (2025)
Dynamic Context Tuning for Retrieval-Augmented Generation: Enhancing Multi-Turn Planning and Tool Adaptation
by: Soni, Jubin Abhishek, et al.
Published: (2025)
by: Soni, Jubin Abhishek, et al.
Published: (2025)
Patchview: LLM-Powered Worldbuilding with Generative Dust and Magnet Visualization
by: Chung, John Joon Young, et al.
Published: (2024)
by: Chung, John Joon Young, et al.
Published: (2024)
Captioning Visualizations with Large Language Models (CVLLM): A Tutorial
by: Carenini, Giuseppe, et al.
Published: (2024)
by: Carenini, Giuseppe, et al.
Published: (2024)
Tell Me More! Towards Implicit User Intention Understanding of Language Model Driven Agents
by: Qian, Cheng, et al.
Published: (2024)
by: Qian, Cheng, et al.
Published: (2024)
ProactiveEval: A Unified Evaluation Framework for Proactive Dialogue Agents
by: Liu, Tianjian, et al.
Published: (2025)
by: Liu, Tianjian, et al.
Published: (2025)
Unpacking Human Preference for LLMs: Demographically Aware Evaluation with the HUMAINE Framework
by: Petrova, Nora, et al.
Published: (2026)
by: Petrova, Nora, et al.
Published: (2026)
Your Co-Workers Matter: Evaluating Collaborative Capabilities of Language Models in Blocks World
by: Wu, Guande, et al.
Published: (2024)
by: Wu, Guande, et al.
Published: (2024)
ChatVis: Automating Scientific Visualization with a Large Language Model
by: Mallick, Tanwi, et al.
Published: (2024)
by: Mallick, Tanwi, et al.
Published: (2024)
Elsewise: Authoring AI-Based Interactive Narrative with Possibility Space Visualization
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
PleaSQLarify: Visual Pragmatic Repair for Natural Language Database Querying
by: Chan, Robin Shing Moon, et al.
Published: (2026)
by: Chan, Robin Shing Moon, et al.
Published: (2026)
Large Language Models Can Solve Real-World Planning Rigorously with Formal Verification Tools
by: Hao, Yilun, et al.
Published: (2024)
by: Hao, Yilun, et al.
Published: (2024)
Toyteller: AI-powered Visual Storytelling Through Toy-Playing with Character Symbols
by: Chung, John Joon Young, et al.
Published: (2025)
by: Chung, John Joon Young, et al.
Published: (2025)
MASCOT: Towards Multi-Agent Socio-Collaborative Companion Systems
by: Wang, Yiyang, et al.
Published: (2026)
by: Wang, Yiyang, et al.
Published: (2026)
ChatScratch: An AI-Augmented System Toward Autonomous Visual Programming Learning for Children Aged 6-12
by: Chen, Liuqing, et al.
Published: (2024)
by: Chen, Liuqing, et al.
Published: (2024)
Similar Items
-
EditScribe: Non-Visual Image Editing with Natural Language Verification Loops
by: Chang, Ruei-Che, et al.
Published: (2024) -
Probing the Gaps in ChatGPT Live Video Chat for Real-World Assistance for People who are Blind or Visually Impaired
by: Chang, Ruei-Che, et al.
Published: (2025) -
TouchScribe: Augmenting Non-Visual Hand-Object Interactions with Automated Live Visual Descriptions
by: Chang, Ruei-Che, et al.
Published: (2026) -
StateScribe: Towards Accessible Change Awareness Across Real-World Revisits
by: Chang, Ruei-Che, et al.
Published: (2026) -
Not There Yet: Evaluating Vision Language Models in Simulating the Visual Perception of People with Low Vision
by: Natalie, Rosiana, et al.
Published: (2025)