ImageTalk: Designing a Multimodal AAC Text Generation System Driven by Image Recognition and Natural Language Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Boyin, Jiang, Puming, Kristensson, Per Ola |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Encode-Store-Retrieve: Augmenting Human Memory through Language-Encoded Egocentric Perception
by: Shen, Junxiao, et al.
Published: (2023)
by: Shen, Junxiao, et al.
Published: (2023)
Bend It, Aim It, Tap It: Designing an On-Body Disambiguation Mechanism for Curve Selection in Mixed Reality
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
Unbounded: Object-Boundary Interaction in Mixed Reality
by: Lyu, Zhuoyue, et al.
Published: (2025)
by: Lyu, Zhuoyue, et al.
Published: (2025)
Optimizing Curve-Based Selection with On-Body Surfaces in Virtual Environments
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
Objestures: Everyday Objects Meet Mid-Air Gestures for Expressive Interaction
by: Lyu, Zhuoyue, et al.
Published: (2025)
by: Lyu, Zhuoyue, et al.
Published: (2025)
Evaluating the Usability of Microgestures for Text Editing Tasks in Virtual Reality
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
Generative AI for Accessible and Inclusive Extended Reality
by: Grubert, Jens, et al.
Published: (2024)
by: Grubert, Jens, et al.
Published: (2024)
Using Text-to-Image Generation for Architectural Design Ideation
by: Paananen, Ville, et al.
Published: (2023)
by: Paananen, Ville, et al.
Published: (2023)
LocoScooter: Designing a Stationary Scooter-Based Locomotion System for Navigation in Virtual Reality
by: He, Wei, et al.
Published: (2026)
by: He, Wei, et al.
Published: (2026)
Semantic Draw Engineering for Text-to-Image Creation
by: Li, Yang, et al.
Published: (2023)
by: Li, Yang, et al.
Published: (2023)
Analyzing Multimodal Interaction Strategies for LLM-Assisted Manipulation of 3D Scenes
by: Chen, Junlong, et al.
Published: (2024)
by: Chen, Junlong, et al.
Published: (2024)
How Do We Evaluate Experiences in Immersive Environments?
by: Li, Xiang, et al.
Published: (2026)
by: Li, Xiang, et al.
Published: (2026)
POET: Supporting Prompting Creativity and Personalization with Automated Expansion of Text-to-Image Generation
by: Han, Evans Xu, et al.
Published: (2025)
by: Han, Evans Xu, et al.
Published: (2025)
Large Language Model-assisted Speech and Pointing Benefits Multiple 3D Object Selection in Virtual Reality
by: Chen, Junlong, et al.
Published: (2024)
by: Chen, Junlong, et al.
Published: (2024)
SCHEMA for Gemini 3 Pro Image: A Structured Methodology for Controlled AI Image Generation on Google's Native Multimodal Model
by: Cazzaniga, Luca
Published: (2026)
by: Cazzaniga, Luca
Published: (2026)
Prompt-Driven Agentic Video Editing System: Autonomous Comprehension of Long-Form, Story-Driven Media
by: Ding, Zihan, et al.
Published: (2025)
by: Ding, Zihan, et al.
Published: (2025)
SketchFlex: Facilitating Spatial-Semantic Coherence in Text-to-Image Generation with Region-Based Sketches
by: Lin, Haichuan, et al.
Published: (2025)
by: Lin, Haichuan, et al.
Published: (2025)
Steering Generative Models for Accessibility: EasyRead Image Generation
by: Dickenmann, Nicolas, et al.
Published: (2026)
by: Dickenmann, Nicolas, et al.
Published: (2026)
EIT-1M: One Million EEG-Image-Text Pairs for Human Visual-textual Recognition and More
by: Zheng, Xu, et al.
Published: (2024)
by: Zheng, Xu, et al.
Published: (2024)
Optical Tag-Based Neuronavigation and Augmentation System for Non-Invasive Brain Stimulation
by: Hu, Xuyi, et al.
Published: (2026)
by: Hu, Xuyi, et al.
Published: (2026)
Handows: A Palm-Based Interactive Multi-Window Management System in Virtual Reality
by: Wang, Jindu, et al.
Published: (2025)
by: Wang, Jindu, et al.
Published: (2025)
DeepSORT-Driven Visual Tracking Approach for Gesture Recognition in Interactive Systems
by: Zhang, Tong, et al.
Published: (2025)
by: Zhang, Tong, et al.
Published: (2025)
SwipeGANSpace: Swipe-to-Compare Image Generation via Efficient Latent Space Exploration
by: Nakashima, Yuto, et al.
Published: (2024)
by: Nakashima, Yuto, et al.
Published: (2024)
Text-to-Image Representativity Fairness Evaluation Framework
by: Yamani, Asma, et al.
Published: (2024)
by: Yamani, Asma, et al.
Published: (2024)
ImaGGen: Zero-Shot Generation of Co-Speech Semantic Gestures Grounded in Language and Image Input
by: Voss, Hendric, et al.
Published: (2025)
by: Voss, Hendric, et al.
Published: (2025)
Exploring a Design Framework for Children's Agency through Participatory Design
by: Yang, Boyin, et al.
Published: (2026)
by: Yang, Boyin, et al.
Published: (2026)
DIG In: Evaluating Disparities in Image Generations with Indicators for Geographic Diversity
by: Hall, Melissa, et al.
Published: (2023)
by: Hall, Melissa, et al.
Published: (2023)
Text-to-Image Generation for Vocabulary Learning Using the Keyword Method
by: Attygalle, Nuwan T., et al.
Published: (2025)
by: Attygalle, Nuwan T., et al.
Published: (2025)
M2LADS Demo: A System for Generating Multimodal Learning Analytics Dashboards
by: Becerra, Alvaro, et al.
Published: (2025)
by: Becerra, Alvaro, et al.
Published: (2025)
T2I-Copilot: A Training-Free Multi-Agent Text-to-Image System for Enhanced Prompt Interpretation and Interactive Generation
by: Chen, Chieh-Yun, et al.
Published: (2025)
by: Chen, Chieh-Yun, et al.
Published: (2025)
From Image Generation to Infrastructure Design: a Multi-agent Pipeline for Street Design Generation
by: Wang, Chenguang, et al.
Published: (2025)
by: Wang, Chenguang, et al.
Published: (2025)
Shifts in Doctors' Eye Movements Between Real and AI-Generated Medical Images
by: Wong, David C, et al.
Published: (2025)
by: Wong, David C, et al.
Published: (2025)
Bridging Text and Image for Artist Style Transfer via Contrastive Learning
by: Liu, Zhi-Song, et al.
Published: (2024)
by: Liu, Zhi-Song, et al.
Published: (2024)
A Multi-Camera Optical Tag Neuronavigation and AR Augmentation Framework for Non-Invasive Brain Stimulation
by: Hu, Xuyi, et al.
Published: (2026)
by: Hu, Xuyi, et al.
Published: (2026)
DiffGaze: A Diffusion Model for Continuous Gaze Sequence Generation on 360° Images
by: Jiao, Chuhan, et al.
Published: (2024)
by: Jiao, Chuhan, et al.
Published: (2024)
Clinically Aware Synthetic Image Generation for Concept Coverage in Chest X-ray Models
by: Rafferty, Amy, et al.
Published: (2026)
by: Rafferty, Amy, et al.
Published: (2026)
GenColor: Generative Color-Concept Association in Visual Design
by: Hou, Yihan, et al.
Published: (2025)
by: Hou, Yihan, et al.
Published: (2025)
EmoGene: Audio-Driven Emotional 3D Talking-Head Generation
by: Wang, Wenqing, et al.
Published: (2024)
by: Wang, Wenqing, et al.
Published: (2024)
Investigating Disability Representations in Text-to-Image Models
by: Tian, Yang, et al.
Published: (2026)
by: Tian, Yang, et al.
Published: (2026)
GLIMPSE : Real-Time Text Recognition and Contextual Understanding for VQA in Wearables
by: Ramachandran, Akhil, et al.
Published: (2026)
by: Ramachandran, Akhil, et al.
Published: (2026)
Similar Items
-
Encode-Store-Retrieve: Augmenting Human Memory through Language-Encoded Egocentric Perception
by: Shen, Junxiao, et al.
Published: (2023) -
Bend It, Aim It, Tap It: Designing an On-Body Disambiguation Mechanism for Curve Selection in Mixed Reality
by: Li, Xiang, et al.
Published: (2025) -
Unbounded: Object-Boundary Interaction in Mixed Reality
by: Lyu, Zhuoyue, et al.
Published: (2025) -
Optimizing Curve-Based Selection with On-Body Surfaces in Virtual Environments
by: Li, Xiang, et al.
Published: (2025) -
Objestures: Everyday Objects Meet Mid-Air Gestures for Expressive Interaction
by: Lyu, Zhuoyue, et al.
Published: (2025)