VisionTrap: Vision-Augmented Trajectory Prediction Guided by Textual Descriptions
Fuente:
arXiv
Salvato in:
| Autori principali: | Moon, Seokha, Woo, Hyun, Park, Hongbeen, Jung, Haeji, Mahjourian, Reza, Chi, Hyung-gun, Lim, Hyerin, Kim, Sangpil, Kim, Jinkyu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Image-Guided Semantic Pseudo-LiDAR Point Generation for 3D Object Detection
di: Lee, Minseung, et al.
Pubblicazione: (2024)
di: Lee, Minseung, et al.
Pubblicazione: (2024)
Learning Temporal Cues by Predicting Objects Move for Multi-camera 3D Object Detection
di: Moon, Seokha, et al.
Pubblicazione: (2024)
di: Moon, Seokha, et al.
Pubblicazione: (2024)
ViTA-PAR: Visual and Textual Attribute Alignment with Attribute Prompting for Pedestrian Attribute Recognition
di: Park, Minjeong, et al.
Pubblicazione: (2025)
di: Park, Minjeong, et al.
Pubblicazione: (2025)
VisionTrap: Unanswerable Questions On Visual Data
di: Saadat, Asir, et al.
Pubblicazione: (2025)
di: Saadat, Asir, et al.
Pubblicazione: (2025)
Stream and Query-guided Feature Aggregation for Efficient and Effective 3D Occupancy Prediction
di: Moon, Seokha, et al.
Pubblicazione: (2025)
di: Moon, Seokha, et al.
Pubblicazione: (2025)
3D Occupancy Prediction with Low-Resolution Queries via Prototype-aware View Transformation
di: Oh, Gyeongrok, et al.
Pubblicazione: (2025)
di: Oh, Gyeongrok, et al.
Pubblicazione: (2025)
Causality-Aware End-to-End Autonomous Driving via Ego-Centric Joint Scene Modeling
di: Moon, Seokha, et al.
Pubblicazione: (2026)
di: Moon, Seokha, et al.
Pubblicazione: (2026)
ALIGN: Advanced Query Initialization with LiDAR-Image Guidance for Occlusion-Robust 3D Object Detection
di: Baek, Janghyun, et al.
Pubblicazione: (2025)
di: Baek, Janghyun, et al.
Pubblicazione: (2025)
Vision-Language Models for Infrared Industrial Sensing in Additive Manufacturing Scene Description
di: Mahjourian, Nazanin, et al.
Pubblicazione: (2025)
di: Mahjourian, Nazanin, et al.
Pubblicazione: (2025)
LRSLAM: Low-rank Representation of Signed Distance Fields in Dense Visual SLAM System
di: Park, Hongbeen, et al.
Pubblicazione: (2025)
di: Park, Hongbeen, et al.
Pubblicazione: (2025)
Motion Cues from Image-based Point Tracking for LiDAR Scene Flow Estimation
di: Jang, Youngdong, et al.
Pubblicazione: (2026)
di: Jang, Youngdong, et al.
Pubblicazione: (2026)
SUPER-AD: Semantic Uncertainty-aware Planning for End-to-End Robust Autonomous Driving
di: Ryu, Wonjeong, et al.
Pubblicazione: (2025)
di: Ryu, Wonjeong, et al.
Pubblicazione: (2025)
Open-Attribute Recognition for Person Retrieval: Finding People Through Distinctive and Novel Attributes
di: Park, Minjeong, et al.
Pubblicazione: (2025)
di: Park, Minjeong, et al.
Pubblicazione: (2025)
Sanitizing Manufacturing Dataset Labels Using Vision-Language Models
di: Mahjourian, Nazanin, et al.
Pubblicazione: (2025)
di: Mahjourian, Nazanin, et al.
Pubblicazione: (2025)
Mitigating the Linguistic Gap with Phonemic Representations for Robust Cross-lingual Transfer
di: Jung, Haeji, et al.
Pubblicazione: (2024)
di: Jung, Haeji, et al.
Pubblicazione: (2024)
InfoGCN++: Learning Representation by Predicting the Future for Online Human Skeleton-based Action Recognition
di: Chi, Seunggeun, et al.
Pubblicazione: (2023)
di: Chi, Seunggeun, et al.
Pubblicazione: (2023)
Configuring Data Augmentations to Reduce Variance Shift in Positional Embedding of Vision Transformers
di: Kim, Bum Jun, et al.
Pubblicazione: (2024)
di: Kim, Bum Jun, et al.
Pubblicazione: (2024)
Clustering-based Image-Text Graph Matching for Domain Generalization
di: Park, Nokyung, et al.
Pubblicazione: (2023)
di: Park, Nokyung, et al.
Pubblicazione: (2023)
FPANet: Frequency-based Video Demoireing using Frame-level Post Alignment
di: Oh, Gyeongrok, et al.
Pubblicazione: (2023)
di: Oh, Gyeongrok, et al.
Pubblicazione: (2023)
MEVG: Multi-event Video Generation with Text-to-Video Models
di: Oh, Gyeongrok, et al.
Pubblicazione: (2023)
di: Oh, Gyeongrok, et al.
Pubblicazione: (2023)
Happiness is Sharing a Vocabulary: A Study of Transliteration Methods
di: Jung, Haeji, et al.
Pubblicazione: (2025)
di: Jung, Haeji, et al.
Pubblicazione: (2025)
Diffusion-Based User-Guided Data Augmentation for Coronary Stenosis Detection
di: Seo, Sumin, et al.
Pubblicazione: (2025)
di: Seo, Sumin, et al.
Pubblicazione: (2025)
A Text-Guided Vision Model for Enhanced Recognition of Small Instances
di: Jung, Hyun-Ki
Pubblicazione: (2026)
di: Jung, Hyun-Ki
Pubblicazione: (2026)
Active Test-time Vision-Language Navigation
di: Ko, Heeju, et al.
Pubblicazione: (2025)
di: Ko, Heeju, et al.
Pubblicazione: (2025)
Vision-Guided Targeted Grasping and Vibration for Robotic Pollination in Controlled Environments
di: Jeong, Jaehwan, et al.
Pubblicazione: (2025)
di: Jeong, Jaehwan, et al.
Pubblicazione: (2025)
Design of Terpene‐Based Eco‐Friendly Solution Process for High‐Performance Organic Solar Cells
di: Hyerin Jeon, et al.
Pubblicazione: (2025)
di: Hyerin Jeon, et al.
Pubblicazione: (2025)
Vision-and-Language Navigation with Analogical Textual Descriptions in LLMs
di: Zhang, Yue, et al.
Pubblicazione: (2025)
di: Zhang, Yue, et al.
Pubblicazione: (2025)
Spatio-Temporal Mixed and Augmented Reality Experience Description for Interactive Playback
di: Kim, Dooyoung, et al.
Pubblicazione: (2025)
di: Kim, Dooyoung, et al.
Pubblicazione: (2025)
GUIDE-CoT: Goal-driven and User-Informed Dynamic Estimation for Pedestrian Trajectory using Chain-of-Thought
di: Kim, Sungsik, et al.
Pubblicazione: (2025)
di: Kim, Sungsik, et al.
Pubblicazione: (2025)
Unified Domain Generalization and Adaptation for Multi-View 3D Object Detection
di: Chang, Gyusam, et al.
Pubblicazione: (2024)
di: Chang, Gyusam, et al.
Pubblicazione: (2024)
Querying Labeled Time Series Data with Scenario Programs
di: Kim, Edward, et al.
Pubblicazione: (2025)
di: Kim, Edward, et al.
Pubblicazione: (2025)
Watermarking for Factuality: Guiding Vision-Language Models Toward Truth via Tri-layer Contrastive Decoding
di: Back, Kyungryul, et al.
Pubblicazione: (2025)
di: Back, Kyungryul, et al.
Pubblicazione: (2025)
SpatiO: Adaptive Test-Time Orchestration of Vision-Language Agents for Spatial Reasoning
di: Hwang, Chan Yeong, et al.
Pubblicazione: (2026)
di: Hwang, Chan Yeong, et al.
Pubblicazione: (2026)
Contrast-Guided Cross-Modal Distillation for Thermal Object Detection
di: Kim, SiWoo, et al.
Pubblicazione: (2025)
di: Kim, SiWoo, et al.
Pubblicazione: (2025)
Just Add $100 More: Augmenting NeRF-based Pseudo-LiDAR Point Cloud for Resolving Class-imbalance Problem
di: Chang, Mincheol, et al.
Pubblicazione: (2024)
di: Chang, Mincheol, et al.
Pubblicazione: (2024)
VLM's Eye Examination: Instruct and Inspect Visual Competency of Vision Language Models
di: Hyeon-Woo, Nam, et al.
Pubblicazione: (2024)
di: Hyeon-Woo, Nam, et al.
Pubblicazione: (2024)
Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene Descriptions
di: Ntinou, Ioanna, et al.
Pubblicazione: (2025)
di: Ntinou, Ioanna, et al.
Pubblicazione: (2025)
Harnessing Linguistic Dissimilarity for Language Generalization on Unseen Low-Resource Varieties
di: Kim, Jinju, et al.
Pubblicazione: (2026)
di: Kim, Jinju, et al.
Pubblicazione: (2026)
SURE Guided Posterior Sampling: Trajectory Correction for Diffusion-Based Inverse Problems
di: Kim, Minwoo, et al.
Pubblicazione: (2025)
di: Kim, Minwoo, et al.
Pubblicazione: (2025)
Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models
di: Woo, Sangmin, et al.
Pubblicazione: (2024)
di: Woo, Sangmin, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Image-Guided Semantic Pseudo-LiDAR Point Generation for 3D Object Detection
di: Lee, Minseung, et al.
Pubblicazione: (2024) -
Learning Temporal Cues by Predicting Objects Move for Multi-camera 3D Object Detection
di: Moon, Seokha, et al.
Pubblicazione: (2024) -
ViTA-PAR: Visual and Textual Attribute Alignment with Attribute Prompting for Pedestrian Attribute Recognition
di: Park, Minjeong, et al.
Pubblicazione: (2025) -
VisionTrap: Unanswerable Questions On Visual Data
di: Saadat, Asir, et al.
Pubblicazione: (2025) -
Stream and Query-guided Feature Aggregation for Efficient and Effective 3D Occupancy Prediction
di: Moon, Seokha, et al.
Pubblicazione: (2025)