What Do You See in Vehicle? Comprehensive Vision Solution for In-Vehicle Gaze Estimation
Fuente:
arXiv
Salvato in:
| Autori principali: | Cheng, Yihua, Zhu, Yaning, Wang, Zongji, Hao, Hongquan, Liu, Yongwei, Cheng, Shiqing, Wang, Xi, Chang, Hyung Jin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VL4Gaze: Unleashing Vision-Language Models for Gaze Following
di: Wang, Shijing, et al.
Pubblicazione: (2025)
di: Wang, Shijing, et al.
Pubblicazione: (2025)
TextGaze: Gaze-Controllable Face Generation with Natural Language
di: Wang, Hengfei, et al.
Pubblicazione: (2024)
di: Wang, Hengfei, et al.
Pubblicazione: (2024)
Enhancing Gaze Reasoning in Vision Foundation Models for Gaze Following
di: Wang, Shijing, et al.
Pubblicazione: (2026)
di: Wang, Shijing, et al.
Pubblicazione: (2026)
3D Prior is All You Need: Cross-Task Few-shot 2D Gaze Estimation
di: Cheng, Yihua, et al.
Pubblicazione: (2025)
di: Cheng, Yihua, et al.
Pubblicazione: (2025)
RTGaze: Real-Time 3D-Aware Gaze Redirection from a Single Image
di: Wang, Hengfei, et al.
Pubblicazione: (2025)
di: Wang, Hengfei, et al.
Pubblicazione: (2025)
Multi-Modal Gaze Following in Conversational Scenarios
di: Hou, Yuqi, et al.
Pubblicazione: (2023)
di: Hou, Yuqi, et al.
Pubblicazione: (2023)
Appearance-based Gaze Estimation With Deep Learning: A Review and Benchmark
di: Cheng, Yihua, et al.
Pubblicazione: (2021)
di: Cheng, Yihua, et al.
Pubblicazione: (2021)
Learning to See What You Need: Gaze Attention for Multimodal Large Language Models
di: Song, Junha, et al.
Pubblicazione: (2026)
di: Song, Junha, et al.
Pubblicazione: (2026)
Multi-task Gaze Estimation Via Unidirectional Convolution
di: Cheng, Zhang, et al.
Pubblicazione: (2024)
di: Cheng, Zhang, et al.
Pubblicazione: (2024)
NL2Contact: Natural Language Guided 3D Hand-Object Contact Modeling with Diffusion Model
di: Zhang, Zhongqun, et al.
Pubblicazione: (2024)
di: Zhang, Zhongqun, et al.
Pubblicazione: (2024)
Lightweight Gaze Estimation Model Via Fusion Global Information
di: Cheng, Zhang, et al.
Pubblicazione: (2024)
di: Cheng, Zhang, et al.
Pubblicazione: (2024)
See Through the Noise: Improving Domain Generalization in Gaze Estimation
di: Peng, Yanming, et al.
Pubblicazione: (2026)
di: Peng, Yanming, et al.
Pubblicazione: (2026)
Differential Contrastive Training for Gaze Estimation
di: Zhang, Lin, et al.
Pubblicazione: (2025)
di: Zhang, Lin, et al.
Pubblicazione: (2025)
EM-Net: Gaze Estimation with Expectation Maximization Algorithm
di: Cheng, Zhang, et al.
Pubblicazione: (2024)
di: Cheng, Zhang, et al.
Pubblicazione: (2024)
Distributed Real-Time Vehicle Control for Emergency Vehicle Transit: A Scalable Cooperative Method
di: Wang, WenXi, et al.
Pubblicazione: (2026)
di: Wang, WenXi, et al.
Pubblicazione: (2026)
What You See Is What Matters: A Novel Visual and Physics-Based Metric for Evaluating Video Generation Quality
di: Wang, Zihan, et al.
Pubblicazione: (2024)
di: Wang, Zihan, et al.
Pubblicazione: (2024)
GazeCLIP: Enhancing Gaze Estimation Through Text-Guided Multimodal Learning
di: Wang, Jun, et al.
Pubblicazione: (2023)
di: Wang, Jun, et al.
Pubblicazione: (2023)
Bidirectional Regression for Monocular 6DoF Head Pose Estimation and Reference System Alignment
di: Chun, Sungho, et al.
Pubblicazione: (2024)
di: Chun, Sungho, et al.
Pubblicazione: (2024)
What You See is What You Ask: Evaluating Audio Descriptions
di: Kala, Divy, et al.
Pubblicazione: (2025)
di: Kala, Divy, et al.
Pubblicazione: (2025)
What You See is What You Classify: Black Box Attributions
di: Stalder, Steven, et al.
Pubblicazione: (2022)
di: Stalder, Steven, et al.
Pubblicazione: (2022)
Efficient Vision-based Vehicle Speed Estimation
di: Macko, Andrej, et al.
Pubblicazione: (2025)
di: Macko, Andrej, et al.
Pubblicazione: (2025)
Roll Your Eyes: Gaze Redirection via Explicit 3D Eyeball Rotation
di: Choi, YoungChan, et al.
Pubblicazione: (2025)
di: Choi, YoungChan, et al.
Pubblicazione: (2025)
Neuro-Cognitive Reward Modeling for Human-Centered Autonomous Vehicle Control
di: Zhuang, Zhuoli, et al.
Pubblicazione: (2026)
di: Zhuang, Zhuoli, et al.
Pubblicazione: (2026)
CVVLSNet: Vehicle Location and Speed Estimation Using Partial Connected Vehicle Trajectory Data
di: Ye, Jiachen, et al.
Pubblicazione: (2024)
di: Ye, Jiachen, et al.
Pubblicazione: (2024)
Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering
di: Choi, Yura, et al.
Pubblicazione: (2026)
di: Choi, Yura, et al.
Pubblicazione: (2026)
Do You See What I Say? Generalizable Deepfake Detection based on Visual Speech Recognition
di: Bora, Maheswar, et al.
Pubblicazione: (2025)
di: Bora, Maheswar, et al.
Pubblicazione: (2025)
What Do You See? Enhancing Zero-Shot Image Classification with Multimodal Large Language Models
di: Abdelhamed, Abdelrahman, et al.
Pubblicazione: (2024)
di: Abdelhamed, Abdelrahman, et al.
Pubblicazione: (2024)
Gaze Label Alignment: Alleviating Domain Shift for Gaze Estimation
di: Zeng, Guanzhong, et al.
Pubblicazione: (2024)
di: Zeng, Guanzhong, et al.
Pubblicazione: (2024)
What Do You See in Common? Learning Hierarchical Prototypes over Tree-of-Life to Discover Evolutionary Traits
di: Manogaran, Harish Babu, et al.
Pubblicazione: (2024)
di: Manogaran, Harish Babu, et al.
Pubblicazione: (2024)
GazeCLIP: Gaze-Guided CLIP with Adaptive-Enhanced Fine-Grained Language Prompt for Deepfake Attribution and Detection
di: Zhang, Yaning, et al.
Pubblicazione: (2026)
di: Zhang, Yaning, et al.
Pubblicazione: (2026)
LG-Gaze: Learning Geometry-aware Continuous Prompts for Language-Guided Gaze Estimation
di: Yin, Pengwei, et al.
Pubblicazione: (2024)
di: Yin, Pengwei, et al.
Pubblicazione: (2024)
Do You See What I See? A Qualitative Study Eliciting High-Level Visualization Comprehension
di: Quadri, Ghulam Jilani, et al.
Pubblicazione: (2024)
di: Quadri, Ghulam Jilani, et al.
Pubblicazione: (2024)
Cross-Paradigm Evaluation of Gaze-Based Semantic Object Identification for Intelligent Vehicles
di: Deng, Penghao, et al.
Pubblicazione: (2026)
di: Deng, Penghao, et al.
Pubblicazione: (2026)
See Where You Read with Eye Gaze Tracking and Large Language Model
di: Yang, Sikai, et al.
Pubblicazione: (2024)
di: Yang, Sikai, et al.
Pubblicazione: (2024)
What You See is (Usually) What You Get: Multimodal Prototype Networks that Abstain from Expensive Modalities
di: Bahng, Muchang, et al.
Pubblicazione: (2025)
di: Bahng, Muchang, et al.
Pubblicazione: (2025)
Cross-Vehicle 3D Geometric Consistency for Self-Supervised Surround Depth Estimation on Articulated Vehicles
di: Liu, Weimin, et al.
Pubblicazione: (2026)
di: Liu, Weimin, et al.
Pubblicazione: (2026)
Revisit What You See: Revealing Visual Semantics in Vision Tokens to Guide LVLM Decoding
di: Cho, Beomsik, et al.
Pubblicazione: (2025)
di: Cho, Beomsik, et al.
Pubblicazione: (2025)
CLIPVehicle: A Unified Framework for Vision-based Vehicle Search
di: Wang, Likai, et al.
Pubblicazione: (2025)
di: Wang, Likai, et al.
Pubblicazione: (2025)
Spatially Selective Imaging in Color: What You See is What You Want
di: John You En Chan, et al.
Pubblicazione: (2024)
di: John You En Chan, et al.
Pubblicazione: (2024)
See What Matters: Differentiable Grid Sample Pruning for Generalizable Vision-Language-Action Model
di: Feng, Yixu, et al.
Pubblicazione: (2026)
di: Feng, Yixu, et al.
Pubblicazione: (2026)
Documenti analoghi
-
VL4Gaze: Unleashing Vision-Language Models for Gaze Following
di: Wang, Shijing, et al.
Pubblicazione: (2025) -
TextGaze: Gaze-Controllable Face Generation with Natural Language
di: Wang, Hengfei, et al.
Pubblicazione: (2024) -
Enhancing Gaze Reasoning in Vision Foundation Models for Gaze Following
di: Wang, Shijing, et al.
Pubblicazione: (2026) -
3D Prior is All You Need: Cross-Task Few-shot 2D Gaze Estimation
di: Cheng, Yihua, et al.
Pubblicazione: (2025) -
RTGaze: Real-Time 3D-Aware Gaze Redirection from a Single Image
di: Wang, Hengfei, et al.
Pubblicazione: (2025)