GuideDog: A Real-World Egocentric Multimodal Dataset for Blind and Low-Vision Accessibility-Aware Guidance
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Junhyeok, Park, Jaewoo, Park, Junhee, Lee, Sangeyl, Chung, Jiwan, Kim, Jisung, Joung, Ji Hoon, Yu, Youngjae |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Explain with Visual Keypoints Like a Real Mentor! A Benchmark for Multimodal Solution Explanation
by: Park, Jaewoo, et al.
Published: (2025)
by: Park, Jaewoo, et al.
Published: (2025)
Are Any-to-Any Models More Consistent Across Modality Transfers Than Specialists?
by: Chung, Jiwan, et al.
Published: (2025)
by: Chung, Jiwan, et al.
Published: (2025)
EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild
by: Kim, Junhyeok, et al.
Published: (2025)
by: Kim, Junhyeok, et al.
Published: (2025)
v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
by: Chung, Jiwan, et al.
Published: (2025)
by: Chung, Jiwan, et al.
Published: (2025)
Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
by: Kim, Youngmin, et al.
Published: (2025)
by: Kim, Youngmin, et al.
Published: (2025)
Teaching Metric Distance to Discrete Autoregressive Language Models
by: Chung, Jiwan, et al.
Published: (2025)
by: Chung, Jiwan, et al.
Published: (2025)
OpenCXD: An Open Real-Device-Guided Hybrid Evaluation Framework for CXL-SSDs
by: Chung, Hyunsun, et al.
Published: (2025)
by: Chung, Hyunsun, et al.
Published: (2025)
A11YN: aligning LLMs for accessible web UI code generation
by: Yoon, Janghan, et al.
Published: (2025)
by: Yoon, Janghan, et al.
Published: (2025)
Background-Aware Defect Generation for Robust Industrial Anomaly Detection
by: Cho, Youngjae, et al.
Published: (2024)
by: Cho, Youngjae, et al.
Published: (2024)
Tracing Mathematical Proficiency Through Problem-Solving Processes
by: Park, Jungyang, et al.
Published: (2025)
by: Park, Jungyang, et al.
Published: (2025)
Zero-shot Multimodal Document Retrieval via Cross-modal Question Generation
by: Choi, Yejin, et al.
Published: (2025)
by: Choi, Yejin, et al.
Published: (2025)
Global Geometry Is Not Enough for Vision Representations
by: Chung, Jiwan, et al.
Published: (2026)
by: Chung, Jiwan, et al.
Published: (2026)
CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs
by: Kim, Jiwan, et al.
Published: (2025)
by: Kim, Jiwan, et al.
Published: (2025)
CANVAS: Commonsense-Aware Navigation System for Intuitive Human-Robot Interaction
by: Choi, Suhwan, et al.
Published: (2024)
by: Choi, Suhwan, et al.
Published: (2024)
EgoXtreme: A Dataset for Robust Object Pose Estimation in Egocentric Views under Extreme Conditions
by: Yoon, Taegyoon, et al.
Published: (2026)
by: Yoon, Taegyoon, et al.
Published: (2026)
Floquet Chern Insulators and Radiation-Induced Zero Resistance in Irradiated Graphene
by: Kim, Youngjae, et al.
Published: (2025)
by: Kim, Youngjae, et al.
Published: (2025)
What MLLMs Learn about When they Learn about Multimodal Reasoning
by: Chung, Jiwan, et al.
Published: (2025)
by: Chung, Jiwan, et al.
Published: (2025)
Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding
by: Chung, Jiwan, et al.
Published: (2024)
by: Chung, Jiwan, et al.
Published: (2024)
Watermarking for Factuality: Guiding Vision-Language Models Toward Truth via Tri-layer Contrastive Decoding
by: Back, Kyungryul, et al.
Published: (2025)
by: Back, Kyungryul, et al.
Published: (2025)
EggHand: A Multimodal Foundation Model for Egocentric Hand Pose Forecasting
by: Choi, Jaeyoung, et al.
Published: (2026)
by: Choi, Jaeyoung, et al.
Published: (2026)
Object Aware Egocentric Online Action Detection
by: An, Joungbin, et al.
Published: (2024)
by: An, Joungbin, et al.
Published: (2024)
Structural characterization and bonding energy analysis for plasma-activated bonding of SiCN films: A reactive molecular dynamics study
by: Kim, Juheon, et al.
Published: (2025)
by: Kim, Juheon, et al.
Published: (2025)
CALL: Context-Aware Low-Latency Retrieval in Disk-Based Vector Databases
by: Jeong, Yeonwoo, et al.
Published: (2025)
by: Jeong, Yeonwoo, et al.
Published: (2025)
MASS: Overcoming Language Bias in Image-Text Matching
by: Chung, Jiwan, et al.
Published: (2025)
by: Chung, Jiwan, et al.
Published: (2025)
VisEscape: A Benchmark for Evaluating Exploration-driven Decision-making in Virtual Escape Rooms
by: Lim, Seungwon, et al.
Published: (2025)
by: Lim, Seungwon, et al.
Published: (2025)
OASIS: Object-based Analytics Storage for Intelligent SQL Query Offloading in Scientific Tabular Workloads
by: Hwang, Soon, et al.
Published: (2025)
by: Hwang, Soon, et al.
Published: (2025)
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
by: Kim, Jinyeong, et al.
Published: (2025)
by: Kim, Jinyeong, et al.
Published: (2025)
Adaptive Graph Rewiring to Mitigate Over-Squashing in Mesh-Based GNNs for Fluid Dynamics Simulations
by: Seo, Sangwoo, et al.
Published: (2025)
by: Seo, Sangwoo, et al.
Published: (2025)
A Host-SSD Collaborative Write Accelerator for LSM-Tree-Based Key-Value Stores
by: Kim, KiHwan, et al.
Published: (2024)
by: Kim, KiHwan, et al.
Published: (2024)
STRAW: A Stress-Aware WL-Based Read Reclaim Technique for High-Density NAND Flash-Based SSDs
by: Chun, Myoungjun, et al.
Published: (2025)
by: Chun, Myoungjun, et al.
Published: (2025)
SSG: Scaled Spatial Guidance for Multi-Scale Visual Autoregressive Generation
by: Shin, Youngwoo, et al.
Published: (2026)
by: Shin, Youngwoo, et al.
Published: (2026)
EgoTraj: Real-World Egocentric Human Trajectory Dataset for Multimodal Prediction
by: Yehia, Ahmad, et al.
Published: (2026)
by: Yehia, Ahmad, et al.
Published: (2026)
Multimodal Dataset Distillation Made Simple by Prototype-Guided Data Synthesis
by: Choi, Junhyeok, et al.
Published: (2026)
by: Choi, Junhyeok, et al.
Published: (2026)
Safety-Guided Flow (SGF): A Unified Framework for Negative Guidance in Safe Generation
by: Kim, Mingyu, et al.
Published: (2026)
by: Kim, Mingyu, et al.
Published: (2026)
Coupling-Robust Accuracy in Multiphysics Physics Informed Neural Networks via Kronecker-Preconditioned Optimization
by: Park, Youngjae, et al.
Published: (2026)
by: Park, Youngjae, et al.
Published: (2026)
Learning Where It Matters: Geometric Anchoring for Robust Preference Alignment
by: Cho, Youngjae, et al.
Published: (2026)
by: Cho, Youngjae, et al.
Published: (2026)
Disentangling and Generating Modalities for Recommendation in Missing Modality Scenarios
by: Kim, Jiwan, et al.
Published: (2025)
by: Kim, Jiwan, et al.
Published: (2025)
Ensuring Functional Correctness of Large Code Models with Selective Generation
by: Jeong, Jaewoo, et al.
Published: (2025)
by: Jeong, Jaewoo, et al.
Published: (2025)
Soft Surfaced Vision-Based Tactile Sensing for Bipedal Robot Applications
by: Kim, Jaeeun, et al.
Published: (2026)
by: Kim, Jaeeun, et al.
Published: (2026)
Impacts of Innovation School System in Korea: A Latent Space Item Response Model with Neyman-Scott Point Process
by: Yi, Seorim, et al.
Published: (2023)
by: Yi, Seorim, et al.
Published: (2023)
Similar Items
-
Explain with Visual Keypoints Like a Real Mentor! A Benchmark for Multimodal Solution Explanation
by: Park, Jaewoo, et al.
Published: (2025) -
Are Any-to-Any Models More Consistent Across Modality Transfers Than Specialists?
by: Chung, Jiwan, et al.
Published: (2025) -
EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild
by: Kim, Junhyeok, et al.
Published: (2025) -
v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
by: Chung, Jiwan, et al.
Published: (2025) -
Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
by: Kim, Youngmin, et al.
Published: (2025)