Pay Attention to Where You Looked
Fuente:
arXiv
Saved in:
| Main Authors: | Berian, Alex, Wu, JhihYang, Brignac, Daniel, Daba, Natnael, Mahalanobis, Abhijit |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CrossModalityDiffusion: Multi-Modal Novel View Synthesis with Unified Intermediate Representation
by: Berian, Alex, et al.
Published: (2025)
by: Berian, Alex, et al.
Published: (2025)
Cascading Unknown Detection with Known Classification for Open Set Recognition
by: Brignac, Daniel, et al.
Published: (2024)
by: Brignac, Daniel, et al.
Published: (2024)
Semantic Smoothing via Novel View Synthesis for Robust SAR Image Classification
by: Brignac, Daniel, et al.
Published: (2026)
by: Brignac, Daniel, et al.
Published: (2026)
Lightweight SAR Ship Detection via Contrastive Distillation
by: Devasundaram, Surendar, et al.
Published: (2026)
by: Devasundaram, Surendar, et al.
Published: (2026)
Look Around and Pay Attention: Multi-camera Point Tracking Reimagined with Transformers
by: Galoaa, Bishoy, et al.
Published: (2025)
by: Galoaa, Bishoy, et al.
Published: (2025)
LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-Supervision
by: Fuller, Anthony, et al.
Published: (2025)
by: Fuller, Anthony, et al.
Published: (2025)
Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention
by: Zhao, Jianfei, et al.
Published: (2025)
by: Zhao, Jianfei, et al.
Published: (2025)
Beyond Where to Look: Trajectory-Guided Reinforcement Learning for Multimodal RLVR
by: Lu, Jinda, et al.
Published: (2026)
by: Lu, Jinda, et al.
Published: (2026)
CLIP Is Shortsighted: Paying Attention Beyond the First Sentence
by: Lavoie, Marc-Antoine, et al.
Published: (2026)
by: Lavoie, Marc-Antoine, et al.
Published: (2026)
Light Weight CNN for classification of Brain Tumors from MRI Images
by: Alemayehu, Natnael
Published: (2025)
by: Alemayehu, Natnael
Published: (2025)
You Only Look at Once for Real-time and Generic Multi-Task
by: Wang, Jiayuan, et al.
Published: (2023)
by: Wang, Jiayuan, et al.
Published: (2023)
Pay Attention to the Keys: Visual Piano Transcription Using Transformers
by: Zivanovic, Uros, et al.
Published: (2024)
by: Zivanovic, Uros, et al.
Published: (2024)
Pay Attention and Move Better: Harnessing Attention for Interactive Motion Generation and Training-free Editing
by: Chen, Ling-Hao, et al.
Published: (2024)
by: Chen, Ling-Hao, et al.
Published: (2024)
LookHere: Vision Transformers with Directed Attention Generalize and Extrapolate
by: Fuller, Anthony, et al.
Published: (2024)
by: Fuller, Anthony, et al.
Published: (2024)
Where do Large Vision-Language Models Look at when Answering Questions?
by: Xing, Xiaoying, et al.
Published: (2025)
by: Xing, Xiaoying, et al.
Published: (2025)
ARTrackV2: Prompting Autoregressive Tracker Where to Look and How to Describe
by: Bai, Yifan, et al.
Published: (2023)
by: Bai, Yifan, et al.
Published: (2023)
Align Where the Words Look: Cross-Attention-Guided Patch Alignment with Contrastive and Transport Regularization for Bengali Captioning
by: Anonto, Riad Ahmed, et al.
Published: (2025)
by: Anonto, Riad Ahmed, et al.
Published: (2025)
Look-Around Before You Leap: High-Frequency Injected Transformer for Image Restoration
by: Zhou, Shihao, et al.
Published: (2024)
by: Zhou, Shihao, et al.
Published: (2024)
You Only Look Bottom-Up for Monocular 3D Object Detection
by: Xiong, Kaixin, et al.
Published: (2024)
by: Xiong, Kaixin, et al.
Published: (2024)
Pay Attention to Your Neighbours: Training-Free Open-Vocabulary Semantic Segmentation
by: Hajimiri, Sina, et al.
Published: (2024)
by: Hajimiri, Sina, et al.
Published: (2024)
MOOSE: Pay Attention to Temporal Dynamics for Video Understanding via Optical Flows
by: Nguyen, Hong, et al.
Published: (2025)
by: Nguyen, Hong, et al.
Published: (2025)
Look Hear: Gaze Prediction for Speech-directed Human Attention
by: Mondal, Sounak, et al.
Published: (2024)
by: Mondal, Sounak, et al.
Published: (2024)
Learning Where to Look: Self-supervised Viewpoint Selection for Active Localization using Geometrical Information
by: Di Giammarino, Luca, et al.
Published: (2024)
by: Di Giammarino, Luca, et al.
Published: (2024)
GridPrune: From "Where to Look" to "What to Select" in Visual Token Pruning for MLLMs
by: Duan, Yuxiang, et al.
Published: (2025)
by: Duan, Yuxiang, et al.
Published: (2025)
Narrating For You: Prompt-guided Audio-visual Narrating Face Generation Employing Multi-entangled Latent Space
by: Chandra, Aashish, et al.
Published: (2026)
by: Chandra, Aashish, et al.
Published: (2026)
Paying More Attention to Image: A Training-Free Method for Alleviating Hallucination in LVLMs
by: Liu, Shi, et al.
Published: (2024)
by: Liu, Shi, et al.
Published: (2024)
Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?
by: Li, Liyang, et al.
Published: (2026)
by: Li, Liyang, et al.
Published: (2026)
Eye Gaze Tells You Where to Compute: Gaze-Driven Efficient VLMs
by: Chen, Qinyu, et al.
Published: (2025)
by: Chen, Qinyu, et al.
Published: (2025)
How Animals Dance (When You're Not Looking)
by: Wang, Xiaojuan, et al.
Published: (2025)
by: Wang, Xiaojuan, et al.
Published: (2025)
Where, What, Why: Towards Explainable Driver Attention Prediction
by: Zhou, Yuchen, et al.
Published: (2025)
by: Zhou, Yuchen, et al.
Published: (2025)
YOLC: You Only Look Clusters for Tiny Object Detection in Aerial Images
by: Liu, Chenguang, et al.
Published: (2024)
by: Liu, Chenguang, et al.
Published: (2024)
What Are You Doing? A Closer Look at Controllable Human Video Generation
by: Bugliarello, Emanuele, et al.
Published: (2025)
by: Bugliarello, Emanuele, et al.
Published: (2025)
Pay Attention to CTC: Fast and Robust Pseudo-Labelling for Unified Speech Recognition
by: Haliassos, Alexandros, et al.
Published: (2026)
by: Haliassos, Alexandros, et al.
Published: (2026)
LookSharp: Attention Entropy Minimization for Test-Time Adaptation
by: Mali, Yash, et al.
Published: (2025)
by: Mali, Yash, et al.
Published: (2025)
Tell Me Where You Are: Multimodal LLMs Meet Place Recognition
by: Lyu, Zonglin, et al.
Published: (2024)
by: Lyu, Zonglin, et al.
Published: (2024)
LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models
by: Shen, Yuxiang, et al.
Published: (2026)
by: Shen, Yuxiang, et al.
Published: (2026)
Look Where It Matters: High-Resolution Crops Retrieval for Efficient VLMs
by: Shabtay, Nimrod, et al.
Published: (2026)
by: Shabtay, Nimrod, et al.
Published: (2026)
Look Where It Matters: Training-Free Ultra-HR Remote Sensing VQA via Adaptive Zoom Search
by: Zhou, Yunqi, et al.
Published: (2025)
by: Zhou, Yunqi, et al.
Published: (2025)
You Only Look Omni Gradient Backpropagation for Moving Infrared Small Target Detection
by: Zhang, Guoyi, et al.
Published: (2025)
by: Zhang, Guoyi, et al.
Published: (2025)
A Decade of You Only Look Once (YOLO) for Object Detection: A Review
by: Ramos, Leo Thomas, et al.
Published: (2025)
by: Ramos, Leo Thomas, et al.
Published: (2025)
Similar Items
-
CrossModalityDiffusion: Multi-Modal Novel View Synthesis with Unified Intermediate Representation
by: Berian, Alex, et al.
Published: (2025) -
Cascading Unknown Detection with Known Classification for Open Set Recognition
by: Brignac, Daniel, et al.
Published: (2024) -
Semantic Smoothing via Novel View Synthesis for Robust SAR Image Classification
by: Brignac, Daniel, et al.
Published: (2026) -
Lightweight SAR Ship Detection via Contrastive Distillation
by: Devasundaram, Surendar, et al.
Published: (2026) -
Look Around and Pay Attention: Multi-camera Point Tracking Reimagined with Transformers
by: Galoaa, Bishoy, et al.
Published: (2025)