Saved in:
| Main Authors: | Aranjuelo, Nerea, Huang, Siyu, Arganda-Carreras, Ignacio, Unzueta, Luis, Otaegui, Oihana, Pfister, Hanspeter, Wei, Donglai |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2405.20643 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Fully Interpretable Statistical Approach for Roadside LiDAR Background Subtraction
by: Iglesias, Aitor, et al.
Published: (2025)
by: Iglesias, Aitor, et al.
Published: (2025)
S$^3$-TTA: Scale-Style Selection for Test-Time Augmentation in Biomedical Image Segmentation
by: Xie, Kangxian, et al.
Published: (2023)
by: Xie, Kangxian, et al.
Published: (2023)
Tree of Attributes Prompt Learning for Vision-Language Models
by: Ding, Tong, et al.
Published: (2024)
by: Ding, Tong, et al.
Published: (2024)
Joint-Task Regularization for Partially Labeled Multi-Task Learning
by: Nishi, Kento, et al.
Published: (2024)
by: Nishi, Kento, et al.
Published: (2024)
Improving generalization by mimicking the human visual diet
by: Madan, Spandan, et al.
Published: (2022)
by: Madan, Spandan, et al.
Published: (2022)
When Visuals Aren't the Problem: Evaluating Vision-Language Models on Misleading Data Visualizations
by: Lalai, Harsh Nishant, et al.
Published: (2026)
by: Lalai, Harsh Nishant, et al.
Published: (2026)
DualEdit: Dual Editing for Knowledge Updating in Vision-Language Models
by: Shi, Zhiyi, et al.
Published: (2025)
by: Shi, Zhiyi, et al.
Published: (2025)
Understanding Graphical Perception in Data Visualization through Zero-shot Prompting of Vision-Language Models
by: Guo, Grace, et al.
Published: (2024)
by: Guo, Grace, et al.
Published: (2024)
AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance
by: Xu, Tianling, et al.
Published: (2025)
by: Xu, Tianling, et al.
Published: (2025)
Affordance-Aware Object Insertion via Mask-Aware Dual Diffusion
by: He, Jixuan, et al.
Published: (2024)
by: He, Jixuan, et al.
Published: (2024)
SocialGPT: Prompting LLMs for Social Relation Reasoning via Greedy Segment Optimization
by: Li, Wanhua, et al.
Published: (2024)
by: Li, Wanhua, et al.
Published: (2024)
GazeMoE: Perception of Gaze Target with Mixture-of-Experts
by: Dai, Zhuangzhuang, et al.
Published: (2026)
by: Dai, Zhuangzhuang, et al.
Published: (2026)
$R^2$-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding
by: Liu, Ye, et al.
Published: (2024)
by: Liu, Ye, et al.
Published: (2024)
Exploration of VLMs for Driver Monitoring Systems Applications
by: Cañas, Paola Natalia, et al.
Published: (2025)
by: Cañas, Paola Natalia, et al.
Published: (2025)
Controlling Face's Frame generation in StyleGAN's latent space operations: Modifying faces to deceive our memory
by: Roca, Agustín, et al.
Published: (2024)
by: Roca, Agustín, et al.
Published: (2024)
Seeing Like Radiologists: Context- and Gaze-Guided Vision-Language Pretraining for Chest X-rays
by: Liu, Kang, et al.
Published: (2026)
by: Liu, Kang, et al.
Published: (2026)
Visual Acoustic Fields
by: Li, Yuelei, et al.
Published: (2025)
by: Li, Yuelei, et al.
Published: (2025)
Generalization of CNNs on Relational Reasoning with Bar Charts
by: Cui, Zhenxing, et al.
Published: (2025)
by: Cui, Zhenxing, et al.
Published: (2025)
Gaze-VLM:Bridging Gaze and VLMs through Attention Regularization for Egocentric Understanding
by: Pani, Anupam, et al.
Published: (2025)
by: Pani, Anupam, et al.
Published: (2025)
GazeVLM: A Vision-Language Model for Multi-Task Gaze Understanding
by: Mathew, Athul M., et al.
Published: (2025)
by: Mathew, Athul M., et al.
Published: (2025)
GazeQwen: Lightweight Gaze-Conditioned LLM Modulation for Streaming Video Understanding
by: Pham, Trong Thang, et al.
Published: (2026)
by: Pham, Trong Thang, et al.
Published: (2026)
See Through the Noise: Improving Domain Generalization in Gaze Estimation
by: Peng, Yanming, et al.
Published: (2026)
by: Peng, Yanming, et al.
Published: (2026)
TPP-Gaze: Modelling Gaze Dynamics in Space and Time with Neural Temporal Point Processes
by: D'Amelio, Alessandro, et al.
Published: (2024)
by: D'Amelio, Alessandro, et al.
Published: (2024)
PGcGAN: Pathological Gait-Conditioned GAN for Human Gait Synthesis
by: Chandrasekaran, Mritula, et al.
Published: (2026)
by: Chandrasekaran, Mritula, et al.
Published: (2026)
Deformation-aware GAN for Medical Image Synthesis with Substantially Misaligned Pairs
by: Xin, Bowen, et al.
Published: (2024)
by: Xin, Bowen, et al.
Published: (2024)
RDFC-GAN: RGB-Depth Fusion CycleGAN for Indoor Depth Completion
by: Wang, Haowen, et al.
Published: (2023)
by: Wang, Haowen, et al.
Published: (2023)
Toddlers' Active Gaze Behavior Supports Self-Supervised Object Learning
by: Yu, Zhengyang, et al.
Published: (2024)
by: Yu, Zhengyang, et al.
Published: (2024)
GazeFormer-MoE: Context-Aware Gaze Estimation via CLIP and MoE Transformer
by: Zhao, Xinyuan, et al.
Published: (2026)
by: Zhao, Xinyuan, et al.
Published: (2026)
Glance-or-Gaze: Incentivizing LMMs to Adaptively Focus Search via Reinforcement Learning
by: Bai, Hongbo, et al.
Published: (2026)
by: Bai, Hongbo, et al.
Published: (2026)
CTRL-GS: Cascaded Temporal Residue Learning for 4D Gaussian Splatting
by: Hou, Karly, et al.
Published: (2025)
by: Hou, Karly, et al.
Published: (2025)
RiGS: Rigid-aware 4D Gaussian Splatting from a Single Monocular Video
by: Wu, Chenyu, et al.
Published: (2026)
by: Wu, Chenyu, et al.
Published: (2026)
GazeSearch: Radiology Findings Search Benchmark
by: Pham, Trong Thang, et al.
Published: (2024)
by: Pham, Trong Thang, et al.
Published: (2024)
Inference-based GAN Video Generation
by: Yang, Jingbo, et al.
Published: (2025)
by: Yang, Jingbo, et al.
Published: (2025)
Gaze on the Prize: Shaping Visual Attention with Return-Guided Contrastive Learning
by: Lee, Andrew, et al.
Published: (2025)
by: Lee, Andrew, et al.
Published: (2025)
From Scene to Object: Text-Guided Dual-Gaze Prediction
by: Ke, Zehong, et al.
Published: (2026)
by: Ke, Zehong, et al.
Published: (2026)
Watch and Learn: Learning to Use Computers from Online Videos
by: Song, Chan Hee, et al.
Published: (2025)
by: Song, Chan Hee, et al.
Published: (2025)
RGBD Gaze Tracking Using Transformer for Feature Fusion
by: Bauer, Tobias J.
Published: (2025)
by: Bauer, Tobias J.
Published: (2025)
Weakly-supervised Medical Image Segmentation with Gaze Annotations
by: Zhong, Yuan, et al.
Published: (2024)
by: Zhong, Yuan, et al.
Published: (2024)
3D Gaussian and Diffusion-Based Gaze Redirection
by: Panchalingam, Abiram, et al.
Published: (2025)
by: Panchalingam, Abiram, et al.
Published: (2025)
Is Geometry Enough? An Evaluation of Landmark-Based Gaze Estimation
by: Agostinelli, Daniele, et al.
Published: (2026)
by: Agostinelli, Daniele, et al.
Published: (2026)
Similar Items
-
A Fully Interpretable Statistical Approach for Roadside LiDAR Background Subtraction
by: Iglesias, Aitor, et al.
Published: (2025) -
S$^3$-TTA: Scale-Style Selection for Test-Time Augmentation in Biomedical Image Segmentation
by: Xie, Kangxian, et al.
Published: (2023) -
Tree of Attributes Prompt Learning for Vision-Language Models
by: Ding, Tong, et al.
Published: (2024) -
Joint-Task Regularization for Partially Labeled Multi-Task Learning
by: Nishi, Kento, et al.
Published: (2024) -
Improving generalization by mimicking the human visual diet
by: Madan, Spandan, et al.
Published: (2022)