Unifying Top-down and Bottom-up Scanpath Prediction Using Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Zhibo, Mondal, Sounak, Ahn, Seoyoung, Xue, Ruoyu, Zelinsky, Gregory, Hoai, Minh, Samaras, Dimitris |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Few-shot Personalized Scanpath Prediction
by: Xue, Ruoyu, et al.
Published: (2025)
by: Xue, Ruoyu, et al.
Published: (2025)
Look Hear: Gaze Prediction for Speech-directed Human Attention
by: Mondal, Sounak, et al.
Published: (2024)
by: Mondal, Sounak, et al.
Published: (2024)
Personalized Image Descriptions from Attention Sequences
by: Xue, Ruoyu, et al.
Published: (2025)
by: Xue, Ruoyu, et al.
Published: (2025)
Predicting Visual Attention in Graphic Design Documents
by: Chakraborty, Souradeep, et al.
Published: (2024)
by: Chakraborty, Souradeep, et al.
Published: (2024)
Diffusion-Refined VQA Annotations for Semi-Supervised Gaze Following
by: Miao, Qiaomu, et al.
Published: (2024)
by: Miao, Qiaomu, et al.
Published: (2024)
Generating metamers of human scene understanding
by: Raina, Ritik, et al.
Published: (2026)
by: Raina, Ritik, et al.
Published: (2026)
Human-like Object Grouping in Self-supervised Vision Transformers
by: Adeli, Hossein, et al.
Published: (2026)
by: Adeli, Hossein, et al.
Published: (2026)
Open-world Instance Segmentation: Top-down Learning with Bottom-up Supervision
by: Kalluri, Tarun, et al.
Published: (2023)
by: Kalluri, Tarun, et al.
Published: (2023)
Multi-view Gaze Target Estimation
by: Miao, Qiaomu, et al.
Published: (2025)
by: Miao, Qiaomu, et al.
Published: (2025)
Measuring and Predicting Where and When Pathologists Focus their Visual Attention while Grading Whole Slide Images of Cancer
by: Chakraborty, Souradeep, et al.
Published: (2025)
by: Chakraborty, Souradeep, et al.
Published: (2025)
OAT: Object-Level Attention Transformer for Gaze Scanpath Prediction
by: Fang, Yini, et al.
Published: (2024)
by: Fang, Yini, et al.
Published: (2024)
Self-supervised co-salient object detection via feature correspondence at multiple scales
by: Chakraborty, Souradeep, et al.
Published: (2024)
by: Chakraborty, Souradeep, et al.
Published: (2024)
What about gravity in video generation? Post-Training Newton's Laws with Verifiable Rewards
by: Le, Minh-Quan, et al.
Published: (2025)
by: Le, Minh-Quan, et al.
Published: (2025)
Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath Prediction
by: Cartella, Giuseppe, et al.
Published: (2025)
by: Cartella, Giuseppe, et al.
Published: (2025)
Beyond Average: Individualized Visual Scanpath Prediction
by: Chen, Xianyu, et al.
Published: (2024)
by: Chen, Xianyu, et al.
Published: (2024)
Talking Head Generation via AU-Guided Landmark Prediction
by: Chang, Shao-Yu, et al.
Published: (2025)
by: Chang, Shao-Yu, et al.
Published: (2025)
TopoDiffusionNet: A Topology-aware Diffusion Model
by: Gupta, Saumya, et al.
Published: (2024)
by: Gupta, Saumya, et al.
Published: (2024)
Weighting Pseudo-Labels via High-Activation Feature Index Similarity and Object Detection for Semi-Supervised Segmentation
by: Howlader, Prantik, et al.
Published: (2024)
by: Howlader, Prantik, et al.
Published: (2024)
Assessing Sample Quality via the Latent Space of Generative Models
by: Xu, Jingyi, et al.
Published: (2024)
by: Xu, Jingyi, et al.
Published: (2024)
One Attention, One Scale: Phase-Aligned Rotary Positional Embeddings for Mixed-Resolution Diffusion Transformer
by: Wu, Haoyu, et al.
Published: (2025)
by: Wu, Haoyu, et al.
Published: (2025)
Decoding the visual attention of pathologists to reveal their level of expertise
by: Chakraborty, Souradeep, et al.
Published: (2024)
by: Chakraborty, Souradeep, et al.
Published: (2024)
Unified Dynamic Scanpath Predictors Outperform Individually Trained Neural Models
by: Abawi, Fares, et al.
Published: (2024)
by: Abawi, Fares, et al.
Published: (2024)
MI-NeRF: Learning a Single Face NeRF from Multiple Identities
by: Chatziagapi, Aggelina, et al.
Published: (2024)
by: Chatziagapi, Aggelina, et al.
Published: (2024)
MIGS: Multi-Identity Gaussian Splatting via Tensor Decomposition
by: Chatziagapi, Aggelina, et al.
Published: (2024)
by: Chatziagapi, Aggelina, et al.
Published: (2024)
Fast constrained sampling in pre-trained diffusion models
by: Graikos, Alexandros, et al.
Published: (2024)
by: Graikos, Alexandros, et al.
Published: (2024)
Scanpath Prediction in Panoramic Videos via Expected Code Length Minimization
by: Li, Mu, et al.
Published: (2023)
by: Li, Mu, et al.
Published: (2023)
GazeXplain: Learning to Predict Natural Language Explanations of Visual Scanpaths
by: Chen, Xianyu, et al.
Published: (2024)
by: Chen, Xianyu, et al.
Published: (2024)
EyeFormer: Predicting Personalized Scanpaths with Transformer-Guided Reinforcement Learning
by: Jiang, Yue, et al.
Published: (2024)
by: Jiang, Yue, et al.
Published: (2024)
PISCES: Annotation-free Text-to-Video Post-Training via Optimal Transport-Aligned Rewards
by: Le, Minh-Quan, et al.
Published: (2026)
by: Le, Minh-Quan, et al.
Published: (2026)
$\infty$-Brush: Controllable Large Image Synthesis with Diffusion Models in Infinite Dimensions
by: Le, Minh-Quan, et al.
Published: (2024)
by: Le, Minh-Quan, et al.
Published: (2024)
Pathformer3D: A 3D Scanpath Transformer for 360° Images
by: Quan, Rong, et al.
Published: (2024)
by: Quan, Rong, et al.
Published: (2024)
Tracking Skiers from the Top to the Bottom
by: Dunnhofer, Matteo, et al.
Published: (2023)
by: Dunnhofer, Matteo, et al.
Published: (2023)
Object Referring-Guided Scanpath Prediction with Perception-Enhanced Vision-Language Models
by: Quan, Rong, et al.
Published: (2026)
by: Quan, Rong, et al.
Published: (2026)
Learning 3D Reconstruction with Priors in Test Time
by: Zhou, Lei, et al.
Published: (2026)
by: Zhou, Lei, et al.
Published: (2026)
Beyond Pixels: Semi-Supervised Semantic Segmentation with a Multi-scale Patch-based Multi-Label Classifier
by: Howlader, Prantik, et al.
Published: (2024)
by: Howlader, Prantik, et al.
Published: (2024)
Importance-Based Token Merging for Efficient Image and Video Generation
by: Wu, Haoyu, et al.
Published: (2024)
by: Wu, Haoyu, et al.
Published: (2024)
JEAN: Joint Expression and Audio-guided NeRF-based Talking Face Generation
by: Chakkera, Sai Tanmay Reddy, et al.
Published: (2024)
by: Chakkera, Sai Tanmay Reddy, et al.
Published: (2024)
AV-Flow: Transforming Text to Audio-Visual Human-like Interactions
by: Chatziagapi, Aggelina, et al.
Published: (2025)
by: Chatziagapi, Aggelina, et al.
Published: (2025)
Joint Top-Down and Bottom-Up Frameworks for 3D Visual Grounding
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
Human Scanpath Prediction in Target-Present Visual Search with Semantic-Foveal Bayesian Attention
by: Luzio, João, et al.
Published: (2025)
by: Luzio, João, et al.
Published: (2025)
Similar Items
-
Few-shot Personalized Scanpath Prediction
by: Xue, Ruoyu, et al.
Published: (2025) -
Look Hear: Gaze Prediction for Speech-directed Human Attention
by: Mondal, Sounak, et al.
Published: (2024) -
Personalized Image Descriptions from Attention Sequences
by: Xue, Ruoyu, et al.
Published: (2025) -
Predicting Visual Attention in Graphic Design Documents
by: Chakraborty, Souradeep, et al.
Published: (2024) -
Diffusion-Refined VQA Annotations for Semi-Supervised Gaze Following
by: Miao, Qiaomu, et al.
Published: (2024)