GuideDog: A Real-World Egocentric Multimodal Dataset for Blind and Low-Vision Accessibility-Aware Guidance
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Junhyeok, Park, Jaewoo, Park, Junhee, Lee, Sangeyl, Chung, Jiwan, Kim, Jisung, Joung, Ji Hoon, Yu, Youngjae |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Explain with Visual Keypoints Like a Real Mentor! A Benchmark for Multimodal Solution Explanation
von: Park, Jaewoo, et al.
Veröffentlicht: (2025)
von: Park, Jaewoo, et al.
Veröffentlicht: (2025)
Are Any-to-Any Models More Consistent Across Modality Transfers Than Specialists?
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)
EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild
von: Kim, Junhyeok, et al.
Veröffentlicht: (2025)
von: Kim, Junhyeok, et al.
Veröffentlicht: (2025)
v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)
Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
von: Kim, Youngmin, et al.
Veröffentlicht: (2025)
von: Kim, Youngmin, et al.
Veröffentlicht: (2025)
Teaching Metric Distance to Discrete Autoregressive Language Models
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)
OpenCXD: An Open Real-Device-Guided Hybrid Evaluation Framework for CXL-SSDs
von: Chung, Hyunsun, et al.
Veröffentlicht: (2025)
von: Chung, Hyunsun, et al.
Veröffentlicht: (2025)
A11YN: aligning LLMs for accessible web UI code generation
von: Yoon, Janghan, et al.
Veröffentlicht: (2025)
von: Yoon, Janghan, et al.
Veröffentlicht: (2025)
Background-Aware Defect Generation for Robust Industrial Anomaly Detection
von: Cho, Youngjae, et al.
Veröffentlicht: (2024)
von: Cho, Youngjae, et al.
Veröffentlicht: (2024)
Tracing Mathematical Proficiency Through Problem-Solving Processes
von: Park, Jungyang, et al.
Veröffentlicht: (2025)
von: Park, Jungyang, et al.
Veröffentlicht: (2025)
Zero-shot Multimodal Document Retrieval via Cross-modal Question Generation
von: Choi, Yejin, et al.
Veröffentlicht: (2025)
von: Choi, Yejin, et al.
Veröffentlicht: (2025)
Global Geometry Is Not Enough for Vision Representations
von: Chung, Jiwan, et al.
Veröffentlicht: (2026)
von: Chung, Jiwan, et al.
Veröffentlicht: (2026)
CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs
von: Kim, Jiwan, et al.
Veröffentlicht: (2025)
von: Kim, Jiwan, et al.
Veröffentlicht: (2025)
CANVAS: Commonsense-Aware Navigation System for Intuitive Human-Robot Interaction
von: Choi, Suhwan, et al.
Veröffentlicht: (2024)
von: Choi, Suhwan, et al.
Veröffentlicht: (2024)
EgoXtreme: A Dataset for Robust Object Pose Estimation in Egocentric Views under Extreme Conditions
von: Yoon, Taegyoon, et al.
Veröffentlicht: (2026)
von: Yoon, Taegyoon, et al.
Veröffentlicht: (2026)
Floquet Chern Insulators and Radiation-Induced Zero Resistance in Irradiated Graphene
von: Kim, Youngjae, et al.
Veröffentlicht: (2025)
von: Kim, Youngjae, et al.
Veröffentlicht: (2025)
What MLLMs Learn about When they Learn about Multimodal Reasoning
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)
Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding
von: Chung, Jiwan, et al.
Veröffentlicht: (2024)
von: Chung, Jiwan, et al.
Veröffentlicht: (2024)
Watermarking for Factuality: Guiding Vision-Language Models Toward Truth via Tri-layer Contrastive Decoding
von: Back, Kyungryul, et al.
Veröffentlicht: (2025)
von: Back, Kyungryul, et al.
Veröffentlicht: (2025)
EggHand: A Multimodal Foundation Model for Egocentric Hand Pose Forecasting
von: Choi, Jaeyoung, et al.
Veröffentlicht: (2026)
von: Choi, Jaeyoung, et al.
Veröffentlicht: (2026)
Object Aware Egocentric Online Action Detection
von: An, Joungbin, et al.
Veröffentlicht: (2024)
von: An, Joungbin, et al.
Veröffentlicht: (2024)
Structural characterization and bonding energy analysis for plasma-activated bonding of SiCN films: A reactive molecular dynamics study
von: Kim, Juheon, et al.
Veröffentlicht: (2025)
von: Kim, Juheon, et al.
Veröffentlicht: (2025)
CALL: Context-Aware Low-Latency Retrieval in Disk-Based Vector Databases
von: Jeong, Yeonwoo, et al.
Veröffentlicht: (2025)
von: Jeong, Yeonwoo, et al.
Veröffentlicht: (2025)
MASS: Overcoming Language Bias in Image-Text Matching
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)
VisEscape: A Benchmark for Evaluating Exploration-driven Decision-making in Virtual Escape Rooms
von: Lim, Seungwon, et al.
Veröffentlicht: (2025)
von: Lim, Seungwon, et al.
Veröffentlicht: (2025)
OASIS: Object-based Analytics Storage for Intelligent SQL Query Offloading in Scientific Tabular Workloads
von: Hwang, Soon, et al.
Veröffentlicht: (2025)
von: Hwang, Soon, et al.
Veröffentlicht: (2025)
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
von: Kim, Jinyeong, et al.
Veröffentlicht: (2025)
von: Kim, Jinyeong, et al.
Veröffentlicht: (2025)
Adaptive Graph Rewiring to Mitigate Over-Squashing in Mesh-Based GNNs for Fluid Dynamics Simulations
von: Seo, Sangwoo, et al.
Veröffentlicht: (2025)
von: Seo, Sangwoo, et al.
Veröffentlicht: (2025)
A Host-SSD Collaborative Write Accelerator for LSM-Tree-Based Key-Value Stores
von: Kim, KiHwan, et al.
Veröffentlicht: (2024)
von: Kim, KiHwan, et al.
Veröffentlicht: (2024)
STRAW: A Stress-Aware WL-Based Read Reclaim Technique for High-Density NAND Flash-Based SSDs
von: Chun, Myoungjun, et al.
Veröffentlicht: (2025)
von: Chun, Myoungjun, et al.
Veröffentlicht: (2025)
SSG: Scaled Spatial Guidance for Multi-Scale Visual Autoregressive Generation
von: Shin, Youngwoo, et al.
Veröffentlicht: (2026)
von: Shin, Youngwoo, et al.
Veröffentlicht: (2026)
EgoTraj: Real-World Egocentric Human Trajectory Dataset for Multimodal Prediction
von: Yehia, Ahmad, et al.
Veröffentlicht: (2026)
von: Yehia, Ahmad, et al.
Veröffentlicht: (2026)
Multimodal Dataset Distillation Made Simple by Prototype-Guided Data Synthesis
von: Choi, Junhyeok, et al.
Veröffentlicht: (2026)
von: Choi, Junhyeok, et al.
Veröffentlicht: (2026)
Safety-Guided Flow (SGF): A Unified Framework for Negative Guidance in Safe Generation
von: Kim, Mingyu, et al.
Veröffentlicht: (2026)
von: Kim, Mingyu, et al.
Veröffentlicht: (2026)
Coupling-Robust Accuracy in Multiphysics Physics Informed Neural Networks via Kronecker-Preconditioned Optimization
von: Park, Youngjae, et al.
Veröffentlicht: (2026)
von: Park, Youngjae, et al.
Veröffentlicht: (2026)
Learning Where It Matters: Geometric Anchoring for Robust Preference Alignment
von: Cho, Youngjae, et al.
Veröffentlicht: (2026)
von: Cho, Youngjae, et al.
Veröffentlicht: (2026)
Disentangling and Generating Modalities for Recommendation in Missing Modality Scenarios
von: Kim, Jiwan, et al.
Veröffentlicht: (2025)
von: Kim, Jiwan, et al.
Veröffentlicht: (2025)
Ensuring Functional Correctness of Large Code Models with Selective Generation
von: Jeong, Jaewoo, et al.
Veröffentlicht: (2025)
von: Jeong, Jaewoo, et al.
Veröffentlicht: (2025)
Soft Surfaced Vision-Based Tactile Sensing for Bipedal Robot Applications
von: Kim, Jaeeun, et al.
Veröffentlicht: (2026)
von: Kim, Jaeeun, et al.
Veröffentlicht: (2026)
Impacts of Innovation School System in Korea: A Latent Space Item Response Model with Neyman-Scott Point Process
von: Yi, Seorim, et al.
Veröffentlicht: (2023)
von: Yi, Seorim, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Explain with Visual Keypoints Like a Real Mentor! A Benchmark for Multimodal Solution Explanation
von: Park, Jaewoo, et al.
Veröffentlicht: (2025) -
Are Any-to-Any Models More Consistent Across Modality Transfers Than Specialists?
von: Chung, Jiwan, et al.
Veröffentlicht: (2025) -
EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild
von: Kim, Junhyeok, et al.
Veröffentlicht: (2025) -
v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
von: Chung, Jiwan, et al.
Veröffentlicht: (2025) -
Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
von: Kim, Youngmin, et al.
Veröffentlicht: (2025)