Is it safe to cross? Interpretable Risk Assessment with GPT-4V for Safety-Aware Street Crossing
Fuente:
arXiv
Saved in:
| Main Authors: | Hwang, Hochul, Kwon, Sunjae, Kim, Yekyung, Kim, Donghyun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Synthetic data augmentation for robotic mobility aids to support blind and low vision people
by: Hwang, Hochul, et al.
Published: (2024)
by: Hwang, Hochul, et al.
Published: (2024)
SiNGER: A Clearer Voice Distills Vision Transformers Further
by: Yu, Geunhyeok, et al.
Published: (2025)
by: Yu, Geunhyeok, et al.
Published: (2025)
Think as Needed: Geometry-Driven Adaptive Perception for Autonomous Driving
by: Kim, Donghyun, et al.
Published: (2026)
by: Kim, Donghyun, et al.
Published: (2026)
Diversity Over Frequency: Rethinking Tool Use in Visual Chain-of-Thought Agents
by: Kim, Dong-Hee, et al.
Published: (2026)
by: Kim, Dong-Hee, et al.
Published: (2026)
Attribute Based Interpretable Evaluation Metrics for Generative Models
by: Kim, Dongkyun, et al.
Published: (2023)
by: Kim, Dongkyun, et al.
Published: (2023)
Enhanced OoD Detection through Cross-Modal Alignment of Multi-Modal Representations
by: Kim, Jeonghyeon, et al.
Published: (2025)
by: Kim, Jeonghyeon, et al.
Published: (2025)
Interpreting vision transformers via residual replacement model
by: Kim, Jinyeong, et al.
Published: (2025)
by: Kim, Jinyeong, et al.
Published: (2025)
From Street to Orbit: Training-Free Cross-View Retrieval via Location Semantics and LLM Guidance
by: Min, Jeongho, et al.
Published: (2025)
by: Min, Jeongho, et al.
Published: (2025)
Learning Question-Aware Keyframe Selection with Synthetic Supervision for Video Question Answering
by: Kwon, Minchan, et al.
Published: (2026)
by: Kwon, Minchan, et al.
Published: (2026)
Learning Unified Distance Metric Across Diverse Data Distributions with Parameter-Efficient Transfer Learning
by: Kim, Sungyeon, et al.
Published: (2023)
by: Kim, Sungyeon, et al.
Published: (2023)
Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models
by: Jung, Woojun, et al.
Published: (2025)
by: Jung, Woojun, et al.
Published: (2025)
Talk in Pieces, See in Whole: Disentangling and Hierarchical Aggregating Representations for Language-based Object Detection
by: An, Sojung, et al.
Published: (2025)
by: An, Sojung, et al.
Published: (2025)
Adaptive Self-training Framework for Fine-grained Scene Graph Generation
by: Kim, Kibum, et al.
Published: (2024)
by: Kim, Kibum, et al.
Published: (2024)
SynRES: Towards Referring Expression Segmentation in the Wild via Synthetic Data
by: Kim, Dong-Hee, et al.
Published: (2025)
by: Kim, Dong-Hee, et al.
Published: (2025)
SportsGPT: An LLM-driven Framework for Interpretable Sports Motion Assessment and Training Guidance
by: Tian, Wenbo, et al.
Published: (2025)
by: Tian, Wenbo, et al.
Published: (2025)
GPT as Psychologist? Preliminary Evaluations for GPT-4V on Visual Affective Computing
by: Lu, Hao, et al.
Published: (2024)
by: Lu, Hao, et al.
Published: (2024)
Training-Free Label Space Alignment for Universal Domain Adaptation
by: Lee, Dujin, et al.
Published: (2025)
by: Lee, Dujin, et al.
Published: (2025)
AesFA: An Aesthetic Feature-Aware Arbitrary Neural Style Transfer
by: Kwon, Joonwoo, et al.
Published: (2023)
by: Kwon, Joonwoo, et al.
Published: (2023)
An Evaluation of GPT-4V and Gemini in Online VQA
by: Liu, Mengchen, et al.
Published: (2023)
by: Liu, Mengchen, et al.
Published: (2023)
SIA: Enhancing Safety via Intent Awareness for Vision-Language Models
by: Na, Youngjin, et al.
Published: (2025)
by: Na, Youngjin, et al.
Published: (2025)
LLM4SGG: Large Language Models for Weakly Supervised Scene Graph Generation
by: Kim, Kibum, et al.
Published: (2023)
by: Kim, Kibum, et al.
Published: (2023)
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
by: Kim, Jinyeong, et al.
Published: (2025)
by: Kim, Jinyeong, et al.
Published: (2025)
PLATYPUS: Progressive Local Surface Estimator for Arbitrary-Scale Point Cloud Upsampling
by: Kim, Donghyun, et al.
Published: (2024)
by: Kim, Donghyun, et al.
Published: (2024)
Delaunay Canopy: Building Wireframe Reconstruction from Airborne LiDAR Point Clouds via Delaunay Graph
by: Kim, Donghyun, et al.
Published: (2026)
by: Kim, Donghyun, et al.
Published: (2026)
Morphology-Aware Interactive Keypoint Estimation
by: Kim, Jinhee, et al.
Published: (2022)
by: Kim, Jinhee, et al.
Published: (2022)
Preserve or Modify? Context-Aware Evaluation for Balancing Preservation and Modification in Text-Guided Image Editing
by: Kim, Yoonjeon, et al.
Published: (2024)
by: Kim, Yoonjeon, et al.
Published: (2024)
4D Gaussian Splatting in the Wild with Uncertainty-Aware Regularization
by: Kim, Mijeong, et al.
Published: (2024)
by: Kim, Mijeong, et al.
Published: (2024)
Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding
by: Kim, Namho, et al.
Published: (2025)
by: Kim, Namho, et al.
Published: (2025)
An Integrated Causal Inference Framework for Traffic Safety Modeling with Semantic Street-View Visual Features
by: Sun, Lishan, et al.
Published: (2026)
by: Sun, Lishan, et al.
Published: (2026)
DIAMOND: An LLM-Driven Agent for Context-Aware Baseball Highlight Summarization
by: Kang, Jeonghun, et al.
Published: (2025)
by: Kang, Jeonghun, et al.
Published: (2025)
VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis
by: Kang, Donggoo, et al.
Published: (2024)
by: Kang, Donggoo, et al.
Published: (2024)
FEAST: Fully Connected Expressive Attention for Spatial Transcriptomics
by: Jeong, Taejin, et al.
Published: (2026)
by: Jeong, Taejin, et al.
Published: (2026)
SYNAPSE: Synergizing an Adapter and Finetuning for High-Fidelity EEG Synthesis from a CLIP-Aligned Encoder
by: Lee, Jeyoung, et al.
Published: (2025)
by: Lee, Jeyoung, et al.
Published: (2025)
Safety-Guided Flow (SGF): A Unified Framework for Negative Guidance in Safe Generation
by: Kim, Mingyu, et al.
Published: (2026)
by: Kim, Mingyu, et al.
Published: (2026)
Visual Delta Generator with Large Multi-modal Models for Semi-supervised Composed Image Retrieval
by: Jang, Young Kyun, et al.
Published: (2024)
by: Jang, Young Kyun, et al.
Published: (2024)
See What You Are Told: Visual Attention Sink in Large Multimodal Models
by: Kang, Seil, et al.
Published: (2025)
by: Kang, Seil, et al.
Published: (2025)
Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding
by: Kang, Seil, et al.
Published: (2025)
by: Kang, Seil, et al.
Published: (2025)
TasselNetV4: A vision foundation model for cross-scene, cross-scale, and cross-species plant counting
by: Hu, Xiaonan, et al.
Published: (2025)
by: Hu, Xiaonan, et al.
Published: (2025)
Zero-shot Text-guided Infinite Image Synthesis with LLM guidance
by: Kwon, Soyeong, et al.
Published: (2024)
by: Kwon, Soyeong, et al.
Published: (2024)
Holi-DETR: Holistic Fashion Item Detection Leveraging Contextual Information
by: Kwon, Youngchae, et al.
Published: (2025)
by: Kwon, Youngchae, et al.
Published: (2025)
Similar Items
-
Synthetic data augmentation for robotic mobility aids to support blind and low vision people
by: Hwang, Hochul, et al.
Published: (2024) -
SiNGER: A Clearer Voice Distills Vision Transformers Further
by: Yu, Geunhyeok, et al.
Published: (2025) -
Think as Needed: Geometry-Driven Adaptive Perception for Autonomous Driving
by: Kim, Donghyun, et al.
Published: (2026) -
Diversity Over Frequency: Rethinking Tool Use in Visual Chain-of-Thought Agents
by: Kim, Dong-Hee, et al.
Published: (2026) -
Attribute Based Interpretable Evaluation Metrics for Generative Models
by: Kim, Dongkyun, et al.
Published: (2023)