OAT: Object-Level Attention Transformer for Gaze Scanpath Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Fang, Yini, Yu, Jingling, Zhang, Haozheng, van der Lans, Ralf, Shi, Bertram |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GazeXplain: Learning to Predict Natural Language Explanations of Visual Scanpaths
by: Chen, Xianyu, et al.
Published: (2024)
by: Chen, Xianyu, et al.
Published: (2024)
End-to-End Facial Expression Detection in Long Videos
by: Fang, Yini, et al.
Published: (2025)
by: Fang, Yini, et al.
Published: (2025)
Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath Prediction
by: Cartella, Giuseppe, et al.
Published: (2025)
by: Cartella, Giuseppe, et al.
Published: (2025)
Beyond Scanpaths: Graph-Based Gaze Simulation in Dynamic Scenes
by: Palmer, Luke, et al.
Published: (2026)
by: Palmer, Luke, et al.
Published: (2026)
Merging Multiple Datasets for Improved Appearance-Based Gaze Estimation
by: Wu, Liang, et al.
Published: (2024)
by: Wu, Liang, et al.
Published: (2024)
CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath Modeling
by: Pham, Trong-Thang, et al.
Published: (2025)
by: Pham, Trong-Thang, et al.
Published: (2025)
Object Referring-Guided Scanpath Prediction with Perception-Enhanced Vision-Language Models
by: Quan, Rong, et al.
Published: (2026)
by: Quan, Rong, et al.
Published: (2026)
TransGOP: Transformer-Based Gaze Object Prediction
by: Wang, Binglu, et al.
Published: (2024)
by: Wang, Binglu, et al.
Published: (2024)
Few-shot Personalized Scanpath Prediction
by: Xue, Ruoyu, et al.
Published: (2025)
by: Xue, Ruoyu, et al.
Published: (2025)
DriverGaze360: OmniDirectional Driver Attention with Object-Level Guidance
by: Govil, Shreedhar, et al.
Published: (2025)
by: Govil, Shreedhar, et al.
Published: (2025)
Human Scanpath Prediction in Target-Present Visual Search with Semantic-Foveal Bayesian Attention
by: Luzio, João, et al.
Published: (2025)
by: Luzio, João, et al.
Published: (2025)
Beyond Average: Individualized Visual Scanpath Prediction
by: Chen, Xianyu, et al.
Published: (2024)
by: Chen, Xianyu, et al.
Published: (2024)
Unifying Top-down and Bottom-up Scanpath Prediction Using Transformers
by: Yang, Zhibo, et al.
Published: (2023)
by: Yang, Zhibo, et al.
Published: (2023)
A Robotics-Inspired Scanpath Model Reveals the Importance of Uncertainty and Semantic Object Cues for Gaze Guidance in Dynamic Scenes
by: Mengers, Vito, et al.
Published: (2024)
by: Mengers, Vito, et al.
Published: (2024)
Boosting Gaze Object Prediction via Pixel-level Supervision from Vision Foundation Model
by: Jin, Yang, et al.
Published: (2024)
by: Jin, Yang, et al.
Published: (2024)
Learning from Observer Gaze:Zero-Shot Attention Prediction Oriented by Human-Object Interaction Recognition
by: Zhou, Yuchen, et al.
Published: (2024)
by: Zhou, Yuchen, et al.
Published: (2024)
Scanpath Prediction in Panoramic Videos via Expected Code Length Minimization
by: Li, Mu, et al.
Published: (2023)
by: Li, Mu, et al.
Published: (2023)
Pathformer3D: A 3D Scanpath Transformer for 360° Images
by: Quan, Rong, et al.
Published: (2024)
by: Quan, Rong, et al.
Published: (2024)
GaTector+: A Unified Head-free Framework for Gaze Object and Gaze Following Prediction
by: Jin, Yang, et al.
Published: (2025)
by: Jin, Yang, et al.
Published: (2025)
Look Hear: Gaze Prediction for Speech-directed Human Attention
by: Mondal, Sounak, et al.
Published: (2024)
by: Mondal, Sounak, et al.
Published: (2024)
RL-ScanIQA: Reinforcement-Learned Scanpaths for Blind 360°Image Quality Assessment
by: Wang, Yujia, et al.
Published: (2026)
by: Wang, Yujia, et al.
Published: (2026)
MVTOP: Multi-View Transformer-based Object Pose-Estimation
by: Ranftl, Lukas, et al.
Published: (2025)
by: Ranftl, Lukas, et al.
Published: (2025)
DualGazeNet: A Biologically Inspired Dual-Gaze Query Network for Salient Object Detection
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
Towards Pixel-Level Prediction for Gaze Following: Benchmark and Approach
by: Liu, Feiyang, et al.
Published: (2024)
by: Liu, Feiyang, et al.
Published: (2024)
Gaze-guided Hand-Object Interaction Synthesis: Dataset and Method
by: Tian, Jie, et al.
Published: (2024)
by: Tian, Jie, et al.
Published: (2024)
EyeFormer: Predicting Personalized Scanpaths with Transformer-Guided Reinforcement Learning
by: Jiang, Yue, et al.
Published: (2024)
by: Jiang, Yue, et al.
Published: (2024)
SGAP-Gaze: Scene Grid Attention Based Point-of-Gaze Estimation Network for Driver Gaze
by: Sharma, Pavan Kumar, et al.
Published: (2026)
by: Sharma, Pavan Kumar, et al.
Published: (2026)
LetheViT: Selective Machine Unlearning for Vision Transformers via Attention-Guided Contrastive Learning
by: Tong, Yujia, et al.
Published: (2025)
by: Tong, Yujia, et al.
Published: (2025)
MxT: Mamba x Transformer for Image Inpainting
by: Chen, Shuang, et al.
Published: (2024)
by: Chen, Shuang, et al.
Published: (2024)
GazeHTA: End-to-end Gaze Target Detection with Head-Target Association
by: Lin, Zhi-Yi, et al.
Published: (2024)
by: Lin, Zhi-Yi, et al.
Published: (2024)
GazeProphet: Software-Only Gaze Prediction for VR Foveated Rendering
by: Ebadulla, Farhaan, et al.
Published: (2025)
by: Ebadulla, Farhaan, et al.
Published: (2025)
HiLO: High-Level Object Fusion for Autonomous Driving using Transformers
by: Osterburg, Timo, et al.
Published: (2025)
by: Osterburg, Timo, et al.
Published: (2025)
FaceSleuth-R: Adaptive Orientation-Aware Attention for Robust Micro-Expression Recognition
by: Wu, Linquan, et al.
Published: (2025)
by: Wu, Linquan, et al.
Published: (2025)
Automated Detection of Mutual Gaze and Joint Attention in Dual-Camera Settings via Dual-Stream Transformers
by: Kosmydel, Jakub, et al.
Published: (2026)
by: Kosmydel, Jakub, et al.
Published: (2026)
Zero-shot Object Counting with Good Exemplars
by: Zhu, Huilin, et al.
Published: (2024)
by: Zhu, Huilin, et al.
Published: (2024)
DHECA-SuperGaze: Dual Head-Eye Cross-Attention and Super-Resolution for Unconstrained Gaze Estimation
by: Šikić, Franko, et al.
Published: (2025)
by: Šikić, Franko, et al.
Published: (2025)
HAD: Hierarchical Asymmetric Distillation to Bridge Spatio-Temporal Gaps in Event-Based Object Tracking
by: Deng, Yao, et al.
Published: (2025)
by: Deng, Yao, et al.
Published: (2025)
Visual Attention Drifts,but Anchors Hold:Mitigating Hallucination in Multimodal Large Language Models via Cross-Layer Visual Anchors
by: Yang, Chengxu, et al.
Published: (2026)
by: Yang, Chengxu, et al.
Published: (2026)
Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing
by: Shi, Baifeng, et al.
Published: (2026)
by: Shi, Baifeng, et al.
Published: (2026)
Impact of Design Decisions in Scanpath Modeling
by: Emami, Parvin, et al.
Published: (2024)
by: Emami, Parvin, et al.
Published: (2024)
Similar Items
-
GazeXplain: Learning to Predict Natural Language Explanations of Visual Scanpaths
by: Chen, Xianyu, et al.
Published: (2024) -
End-to-End Facial Expression Detection in Long Videos
by: Fang, Yini, et al.
Published: (2025) -
Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath Prediction
by: Cartella, Giuseppe, et al.
Published: (2025) -
Beyond Scanpaths: Graph-Based Gaze Simulation in Dynamic Scenes
by: Palmer, Luke, et al.
Published: (2026) -
Merging Multiple Datasets for Improved Appearance-Based Gaze Estimation
by: Wu, Liang, et al.
Published: (2024)