Saved in:
| Main Authors: | Raina, Ritik, Leite, Abe, Graikos, Alexandros, Ahn, Seoyoung, Samaras, Dimitris, Zelinsky, Gregory J. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2601.11675 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unifying Top-down and Bottom-up Scanpath Prediction Using Transformers
by: Yang, Zhibo, et al.
Published: (2023)
by: Yang, Zhibo, et al.
Published: (2023)
Fast constrained sampling in pre-trained diffusion models
by: Graikos, Alexandros, et al.
Published: (2024)
by: Graikos, Alexandros, et al.
Published: (2024)
Look Hear: Gaze Prediction for Speech-directed Human Attention
by: Mondal, Sounak, et al.
Published: (2024)
by: Mondal, Sounak, et al.
Published: (2024)
Predicting Visual Attention in Graphic Design Documents
by: Chakraborty, Souradeep, et al.
Published: (2024)
by: Chakraborty, Souradeep, et al.
Published: (2024)
Human-like Object Grouping in Self-supervised Vision Transformers
by: Adeli, Hossein, et al.
Published: (2026)
by: Adeli, Hossein, et al.
Published: (2026)
PathSegDiff: Pathology Segmentation using Diffusion model representations
by: Danisetty, Sachin Kumar, et al.
Published: (2025)
by: Danisetty, Sachin Kumar, et al.
Published: (2025)
GriDiT: Factorized Grid-Based Diffusion for Efficient Long Image Sequence Generation
by: Tomar, Snehal Singh, et al.
Published: (2025)
by: Tomar, Snehal Singh, et al.
Published: (2025)
Personalized Image Descriptions from Attention Sequences
by: Xue, Ruoyu, et al.
Published: (2025)
by: Xue, Ruoyu, et al.
Published: (2025)
Poppy: Polarization-based Plug-and-Play Guidance for Enhancing Monocular Normal Estimation
by: Kim, Irene, et al.
Published: (2026)
by: Kim, Irene, et al.
Published: (2026)
Diffusion-Refined VQA Annotations for Semi-Supervised Gaze Following
by: Miao, Qiaomu, et al.
Published: (2024)
by: Miao, Qiaomu, et al.
Published: (2024)
Mitigating Diffusion Model Hallucinations with Dynamic Guidance
by: Triaridis, Kostas, et al.
Published: (2025)
by: Triaridis, Kostas, et al.
Published: (2025)
$\infty$-Brush: Controllable Large Image Synthesis with Diffusion Models in Infinite Dimensions
by: Le, Minh-Quan, et al.
Published: (2024)
by: Le, Minh-Quan, et al.
Published: (2024)
Learned representation-guided diffusion models for large-image generation
by: Graikos, Alexandros, et al.
Published: (2023)
by: Graikos, Alexandros, et al.
Published: (2023)
ZoomLDM: Latent Diffusion Model for multi-scale image generation
by: Yellapragada, Srikar, et al.
Published: (2024)
by: Yellapragada, Srikar, et al.
Published: (2024)
Near, far: Patch-ordering enhances vision foundation models' scene understanding
by: Pariza, Valentinos, et al.
Published: (2024)
by: Pariza, Valentinos, et al.
Published: (2024)
Quantifying the synthetic and real domain gap in aerial scene understanding
by: Marcu, Alina
Published: (2024)
by: Marcu, Alina
Published: (2024)
Few-shot Personalized Scanpath Prediction
by: Xue, Ruoyu, et al.
Published: (2025)
by: Xue, Ruoyu, et al.
Published: (2025)
UAV traffic scene understanding: A regulation embedded multi-modal network and a unified benchmark
by: Zhang, Yu, et al.
Published: (2026)
by: Zhang, Yu, et al.
Published: (2026)
Assessing the generalization performance of SAM for ureteroscopy scene understanding
by: Villagrana, Martin, et al.
Published: (2025)
by: Villagrana, Martin, et al.
Published: (2025)
DEF-oriCORN: efficient 3D scene understanding for robust language-directed manipulation without demonstrations
by: Son, Dongwon, et al.
Published: (2024)
by: Son, Dongwon, et al.
Published: (2024)
Gen-SIS: Generative Self-augmentation Improves Self-supervised Learning
by: Belagali, Varun, et al.
Published: (2024)
by: Belagali, Varun, et al.
Published: (2024)
Less yet robust: crucial region selection for scene recognition
by: Zhang, Jianqi, et al.
Published: (2024)
by: Zhang, Jianqi, et al.
Published: (2024)
From Pixels to Predicates Structuring urban perception with scene graphs
by: Liu, Yunlong, et al.
Published: (2025)
by: Liu, Yunlong, et al.
Published: (2025)
Do humans and Convolutional Neural Networks attend to similar areas during scene classification: Effects of task and image type
by: Müller, Romy, et al.
Published: (2023)
by: Müller, Romy, et al.
Published: (2023)
Towards Full-scene Domain Generalization in Multi-agent Collaborative Bird's Eye View Segmentation for Connected and Autonomous Driving
by: Hu, Senkang, et al.
Published: (2023)
by: Hu, Senkang, et al.
Published: (2023)
All-day Multi-scenes Lifelong Vision-and-Language Navigation with Tucker Adaptation
by: Wang, Xudong, et al.
Published: (2026)
by: Wang, Xudong, et al.
Published: (2026)
Assessment of Sentinel-2 spatial and temporal coverage based on the scene classification layer
by: Sanchez, Cristhian, et al.
Published: (2024)
by: Sanchez, Cristhian, et al.
Published: (2024)
Pathology Image Compression with Pre-trained Autoencoders
by: Yellapragada, Srikar, et al.
Published: (2025)
by: Yellapragada, Srikar, et al.
Published: (2025)
May the Dance be with You: Dance Generation Framework for Non-Humanoids
by: Ahn, Hyemin
Published: (2024)
by: Ahn, Hyemin
Published: (2024)
Self-supervised co-salient object detection via feature correspondence at multiple scales
by: Chakraborty, Souradeep, et al.
Published: (2024)
by: Chakraborty, Souradeep, et al.
Published: (2024)
Assessing Sample Quality via the Latent Space of Generative Models
by: Xu, Jingyi, et al.
Published: (2024)
by: Xu, Jingyi, et al.
Published: (2024)
Omni-SILA: Towards Omni-scene Driven Visual Sentiment Identifying, Locating and Attributing in Videos
by: Luo, Jiamin, et al.
Published: (2025)
by: Luo, Jiamin, et al.
Published: (2025)
Human-annotated label noise and their impact on ConvNets for remote sensing image scene classification
by: Peng, Longkang, et al.
Published: (2023)
by: Peng, Longkang, et al.
Published: (2023)
Swift4D:Adaptive divide-and-conquer Gaussian Splatting for compact and efficient reconstruction of dynamic scene
by: Wu, Jiahao, et al.
Published: (2025)
by: Wu, Jiahao, et al.
Published: (2025)
HaloAE: An HaloNet based Local Transformer Auto-Encoder for Anomaly Detection and Localization
by: Mathian, E., et al.
Published: (2022)
by: Mathian, E., et al.
Published: (2022)
MessyKitchens: Contact-rich object-level 3D scene reconstruction
by: Ansari, Junaid Ahmed, et al.
Published: (2026)
by: Ansari, Junaid Ahmed, et al.
Published: (2026)
Multimodal Generative AI with Autoregressive LLMs for Human Motion Understanding and Generation: A Way Forward
by: Islam, Muhammad, et al.
Published: (2025)
by: Islam, Muhammad, et al.
Published: (2025)
COTTA: Context-Aware Transfer Adaptation for Trajectory Prediction in Autonomous Driving
by: Park, Seohyoung, et al.
Published: (2026)
by: Park, Seohyoung, et al.
Published: (2026)
CaddieSet: A Golf Swing Dataset with Human Joint Features and Ball Information
by: Jung, Seunghyeon, et al.
Published: (2025)
by: Jung, Seunghyeon, et al.
Published: (2025)
Melon Fruit Detection and Quality Assessment Using Generative AI-Based Image Data Augmentation
by: Yoon, Seungri, et al.
Published: (2024)
by: Yoon, Seungri, et al.
Published: (2024)
Similar Items
-
Unifying Top-down and Bottom-up Scanpath Prediction Using Transformers
by: Yang, Zhibo, et al.
Published: (2023) -
Fast constrained sampling in pre-trained diffusion models
by: Graikos, Alexandros, et al.
Published: (2024) -
Look Hear: Gaze Prediction for Speech-directed Human Attention
by: Mondal, Sounak, et al.
Published: (2024) -
Predicting Visual Attention in Graphic Design Documents
by: Chakraborty, Souradeep, et al.
Published: (2024) -
Human-like Object Grouping in Self-supervised Vision Transformers
by: Adeli, Hossein, et al.
Published: (2026)