LISA: Language-guided Interference-aware Spatial-Frequency Attention for Driver Gaze Estimation
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Jun, Yang, Zhenye, Zhou, Ruichen, Zhang, Pei, Li, Huan, Chen, Jinpeng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
by: Raoufi, Behnam, et al.
Published: (2025)
by: Raoufi, Behnam, et al.
Published: (2025)
GazeD: Context-Aware Diffusion for Accurate 3D Gaze Estimation
by: Catalini, Riccardo, et al.
Published: (2026)
by: Catalini, Riccardo, et al.
Published: (2026)
DeltaVLM: Interactive Remote Sensing Image Change Analysis via Instruction-guided Difference Perception
by: Deng, Pei, et al.
Published: (2025)
by: Deng, Pei, et al.
Published: (2025)
From Gaze to Insight: Bridging Human Visual Attention and Vision Language Model Explanation for Weakly-Supervised Medical Image Segmentation
by: Chen, Jingkun, et al.
Published: (2025)
by: Chen, Jingkun, et al.
Published: (2025)
ALADIN:Attribute-Language Distillation Network for Person Re-Identification
by: Zhou, Wang, et al.
Published: (2026)
by: Zhou, Wang, et al.
Published: (2026)
NeuroGaze-Distill: Brain-informed Distillation and Depression-Inspired Geometric Priors for Robust Facial Emotion Recognition
by: Li, Zilin, et al.
Published: (2025)
by: Li, Zilin, et al.
Published: (2025)
DELTA: Dense Depth from Events and LiDAR using Transformer's Attention
by: Brebion, Vincent, et al.
Published: (2025)
by: Brebion, Vincent, et al.
Published: (2025)
Gaussian Alignment for Relative Camera Pose Estimation via Single-View Reconstruction
by: Li, Yumin, et al.
Published: (2025)
by: Li, Yumin, et al.
Published: (2025)
Human Modelling and Pose Estimation Overview
by: Knap, Pawel
Published: (2024)
by: Knap, Pawel
Published: (2024)
Removing Cost Volumes from Optical Flow Estimators
by: Kiefhaber, Simon, et al.
Published: (2025)
by: Kiefhaber, Simon, et al.
Published: (2025)
EE3P: Event-based Estimation of Periodic Phenomena Properties
by: Kolář, Jakub, et al.
Published: (2024)
by: Kolář, Jakub, et al.
Published: (2024)
N-DriverMotion: Driver motion learning and prediction using an event-based camera and directly trained spiking neural networks on Loihi 2
by: Chung, Hyo Jong, et al.
Published: (2024)
by: Chung, Hyo Jong, et al.
Published: (2024)
EEPPR: Event-based Estimation of Periodic Phenomena Rate using Correlation in 3D
by: Kolář, Jakub, et al.
Published: (2024)
by: Kolář, Jakub, et al.
Published: (2024)
GraphiContact: Pose-aware Human-Scene Robust Contact Perception for Interactive Systems
by: Lin, Xiaojian, et al.
Published: (2026)
by: Lin, Xiaojian, et al.
Published: (2026)
Autonomous Visual Fish Pen Inspections for Estimating the State of Biofouling Buildup Using ROV -- Extended Abstract
by: Fabijanić, Matej, et al.
Published: (2024)
by: Fabijanić, Matej, et al.
Published: (2024)
SVGS-DSGAT: An IoT-Enabled Innovation in Underwater Robotic Object Detection Technology
by: Wu, Dongli, et al.
Published: (2025)
by: Wu, Dongli, et al.
Published: (2025)
Vi-SAFE: A Spatial-Temporal Framework for Efficient Violence Detection in Public Surveillance
by: Chang, Ligang, et al.
Published: (2025)
by: Chang, Ligang, et al.
Published: (2025)
SCA-Net: Spatial-Contextual Aggregation Network for Enhanced Small Building and Road Change Detection
by: Gholibeigi, Emad, et al.
Published: (2026)
by: Gholibeigi, Emad, et al.
Published: (2026)
TF-Lane: Traffic Flow Module for Robust Lane Perception
by: Xie, Yihan, et al.
Published: (2026)
by: Xie, Yihan, et al.
Published: (2026)
UnCageNet: Tracking and Pose Estimation of Caged Animal
by: Dutta, Sayak, et al.
Published: (2025)
by: Dutta, Sayak, et al.
Published: (2025)
BEVTraj: Map-Free End-to-End Trajectory Prediction in Bird's-Eye View with Deformable Attention and Sparse Goal Proposals
by: Kong, Minsang, et al.
Published: (2025)
by: Kong, Minsang, et al.
Published: (2025)
Car Object Counting and Position Estimation via Extension of the CLIP-EBC Framework
by: Jung, Seoik, et al.
Published: (2025)
by: Jung, Seoik, et al.
Published: (2025)
Synthetic-Child: An AIGC-Based Synthetic Data Pipeline for Privacy-Preserving Child Posture Estimation
by: Zeng, Taowen
Published: (2026)
by: Zeng, Taowen
Published: (2026)
Physical Knot Classification Beyond Accuracy: A Benchmark and Diagnostic Study
by: Nie, Shiheng, et al.
Published: (2026)
by: Nie, Shiheng, et al.
Published: (2026)
Human-Centric Perception for Child Sexual Abuse Imagery
by: Laranjeira, Camila, et al.
Published: (2026)
by: Laranjeira, Camila, et al.
Published: (2026)
A Multi-purpose Tracking Framework for Salmon Welfare Monitoring in Challenging Environments
by: Høgstedt, Espen Uri, et al.
Published: (2025)
by: Høgstedt, Espen Uri, et al.
Published: (2025)
Hierarchical Deep Learning for Diatom Image Classification: A Multi-Level Taxonomic Approach
by: Ke, Yueying
Published: (2025)
by: Ke, Yueying
Published: (2025)
Decoupled Sensitivity-Consistency Learning for Weakly Supervised Video Anomaly Detection
by: Zheng, Hantao, et al.
Published: (2026)
by: Zheng, Hantao, et al.
Published: (2026)
POC-SLT: Partial Object Completion with SDF Latent Transformers
by: Zakeri, Faezeh, et al.
Published: (2024)
by: Zakeri, Faezeh, et al.
Published: (2024)
Foreground Focus: Enhancing Coherence and Fidelity in Camouflaged Image Generation
by: Chen, Pei-Chi, et al.
Published: (2025)
by: Chen, Pei-Chi, et al.
Published: (2025)
A Vision-Language Model for Focal Liver Lesion Classification
by: Jian, Song, et al.
Published: (2025)
by: Jian, Song, et al.
Published: (2025)
Prompt Sensitivity in Vision-Language Grounding: How Small Changes in Wording Affect Object Detection
by: Deka, Dawar Jyoti, et al.
Published: (2026)
by: Deka, Dawar Jyoti, et al.
Published: (2026)
NOAH: Benchmarking Narrative Prior driven Hallucination and Omission in Video Large Language Models
by: Lee, Kyuho, et al.
Published: (2025)
by: Lee, Kyuho, et al.
Published: (2025)
Infrastructure-Centric World Models: Bridging Temporal Depth and Spatial Breadth for Roadside Perception
by: Meng, Siyuan, et al.
Published: (2026)
by: Meng, Siyuan, et al.
Published: (2026)
RCooper: A Real-world Large-scale Dataset for Roadside Cooperative Perception
by: Hao, Ruiyang, et al.
Published: (2024)
by: Hao, Ruiyang, et al.
Published: (2024)
PyCAT4: A Hierarchical Vision Transformer-based Framework for 3D Human Pose Estimation
by: Yang, Zongyou, et al.
Published: (2025)
by: Yang, Zongyou, et al.
Published: (2025)
Generalized Closed-form Formulae for Feature-based Subpixel Alignment in Patch-based Matching
by: Jospin, Laurent Valentin, et al.
Published: (2021)
by: Jospin, Laurent Valentin, et al.
Published: (2021)
Tracking Moose using Aerial Object Detection
by: Indris, Christopher, et al.
Published: (2025)
by: Indris, Christopher, et al.
Published: (2025)
CARDIE: clustering algorithm on relevant descriptors for image enhancement
by: Bonino, Giulia, et al.
Published: (2025)
by: Bonino, Giulia, et al.
Published: (2025)
Semantic Segmentation based Scene Understanding in Autonomous Vehicles
by: Rassekh, Ehsan
Published: (2025)
by: Rassekh, Ehsan
Published: (2025)
Similar Items
-
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
by: Raoufi, Behnam, et al.
Published: (2025) -
GazeD: Context-Aware Diffusion for Accurate 3D Gaze Estimation
by: Catalini, Riccardo, et al.
Published: (2026) -
DeltaVLM: Interactive Remote Sensing Image Change Analysis via Instruction-guided Difference Perception
by: Deng, Pei, et al.
Published: (2025) -
From Gaze to Insight: Bridging Human Visual Attention and Vision Language Model Explanation for Weakly-Supervised Medical Image Segmentation
by: Chen, Jingkun, et al.
Published: (2025) -
ALADIN:Attribute-Language Distillation Network for Person Re-Identification
by: Zhou, Wang, et al.
Published: (2026)