CLIP Is Shortsighted: Paying Attention Beyond the First Sentence
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lavoie, Marc-Antoine, Mahmoud, Anas, Zaimi, Aldo, Tchango, Arsene Fansi, Waslander, Steven L. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Large Self-Supervised Models Bridge the Gap in Domain Adaptive Object Detection
von: Lavoie, Marc-Antoine, et al.
Veröffentlicht: (2025)
von: Lavoie, Marc-Antoine, et al.
Veröffentlicht: (2025)
Feature Density Estimation for Out-of-Distribution Detection via Normalizing Flows
von: Cook, Evan D., et al.
Veröffentlicht: (2024)
von: Cook, Evan D., et al.
Veröffentlicht: (2024)
Image-to-Lidar Relational Distillation for Autonomous Driving Data
von: Mahmoud, Anas, et al.
Veröffentlicht: (2024)
von: Mahmoud, Anas, et al.
Veröffentlicht: (2024)
JDT3D: Addressing the Gaps in LiDAR-Based Tracking-by-Attention
von: Cheong, Brian, et al.
Veröffentlicht: (2024)
von: Cheong, Brian, et al.
Veröffentlicht: (2024)
PSA-SSL: Pose and Size-aware Self-Supervised Learning on LiDAR Point Clouds
von: Nisar, Barza, et al.
Veröffentlicht: (2025)
von: Nisar, Barza, et al.
Veröffentlicht: (2025)
SCATR: Mitigating New Instance Suppression in LiDAR-based Tracking-by-Attention via Second Chance Assignment and Track Query Dropout
von: Cheong, Brian, et al.
Veröffentlicht: (2026)
von: Cheong, Brian, et al.
Veröffentlicht: (2026)
ToaSt: Token Channel Selection and Structured Pruning for Efficient ViT
von: Moon, Hyunchan, et al.
Veröffentlicht: (2026)
von: Moon, Hyunchan, et al.
Veröffentlicht: (2026)
UncertaintyTrack: Exploiting Detection and Localization Uncertainty in Multi-Object Tracking
von: Lee, Chang Won, et al.
Veröffentlicht: (2024)
von: Lee, Chang Won, et al.
Veröffentlicht: (2024)
Pay Attention to Where You Looked
von: Berian, Alex, et al.
Veröffentlicht: (2026)
von: Berian, Alex, et al.
Veröffentlicht: (2026)
ForeSight: Multi-View Streaming Joint Object Detection and Trajectory Forecasting
von: Papais, Sandro, et al.
Veröffentlicht: (2025)
von: Papais, Sandro, et al.
Veröffentlicht: (2025)
Active 6D Pose Estimation for Textureless Objects using Multi-View RGB Frames
von: Yang, Jun, et al.
Veröffentlicht: (2025)
von: Yang, Jun, et al.
Veröffentlicht: (2025)
Toward General Object-level Mapping from Sparse Views with 3D Diffusion Priors
von: Liao, Ziwei, et al.
Veröffentlicht: (2024)
von: Liao, Ziwei, et al.
Veröffentlicht: (2024)
ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation
von: Lan, Mengcheng, et al.
Veröffentlicht: (2024)
von: Lan, Mengcheng, et al.
Veröffentlicht: (2024)
Pay Attention to the Keys: Visual Piano Transcription Using Transformers
von: Zivanovic, Uros, et al.
Veröffentlicht: (2024)
von: Zivanovic, Uros, et al.
Veröffentlicht: (2024)
Pay Attention and Move Better: Harnessing Attention for Interactive Motion Generation and Training-free Editing
von: Chen, Ling-Hao, et al.
Veröffentlicht: (2024)
von: Chen, Ling-Hao, et al.
Veröffentlicht: (2024)
MedSAD-CLIP: Supervised CLIP with Token-Patch Cross-Attention for Medical Anomaly Detection and Segmentation
von: Tran, Thuy Truong, et al.
Veröffentlicht: (2026)
von: Tran, Thuy Truong, et al.
Veröffentlicht: (2026)
Beyond Anatomy: Explainable ASD Classification from rs-fMRI via Functional Parcellation and Graph Attention Networks
von: Madani, Syeda Hareem, et al.
Veröffentlicht: (2026)
von: Madani, Syeda Hareem, et al.
Veröffentlicht: (2026)
Pay Attention to Your Neighbours: Training-Free Open-Vocabulary Semantic Segmentation
von: Hajimiri, Sina, et al.
Veröffentlicht: (2024)
von: Hajimiri, Sina, et al.
Veröffentlicht: (2024)
Look Around and Pay Attention: Multi-camera Point Tracking Reimagined with Transformers
von: Galoaa, Bishoy, et al.
Veröffentlicht: (2025)
von: Galoaa, Bishoy, et al.
Veröffentlicht: (2025)
MOOSE: Pay Attention to Temporal Dynamics for Video Understanding via Optical Flows
von: Nguyen, Hong, et al.
Veröffentlicht: (2025)
von: Nguyen, Hong, et al.
Veröffentlicht: (2025)
Contour Refinement using Discrete Diffusion in Low Data Regime
von: Guan, Fei Yu, et al.
Veröffentlicht: (2026)
von: Guan, Fei Yu, et al.
Veröffentlicht: (2026)
Beyond Accuracy: Evaluating Visual Grounding In Multimodal Medical Reasoning
von: Zafar, Anas, et al.
Veröffentlicht: (2026)
von: Zafar, Anas, et al.
Veröffentlicht: (2026)
Judging by the Rules: Compliance-Aligned Framework for Modern Slavery Statement Monitoring
von: Xu, Wenhao, et al.
Veröffentlicht: (2025)
von: Xu, Wenhao, et al.
Veröffentlicht: (2025)
Beyond Semantics: Disentangling Information Scope in Sparse Autoencoders for CLIP
von: Ro, Yusung, et al.
Veröffentlicht: (2026)
von: Ro, Yusung, et al.
Veröffentlicht: (2026)
Paying More Attention to Image: A Training-Free Method for Alleviating Hallucination in LVLMs
von: Liu, Shi, et al.
Veröffentlicht: (2024)
von: Liu, Shi, et al.
Veröffentlicht: (2024)
Debiasing CLIP: Interpreting and Correcting Bias in Attention Heads
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2025)
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2025)
DiffCLIP: Differential Attention Meets CLIP
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)
Pay Attention to CTC: Fast and Robust Pseudo-Labelling for Unified Speech Recognition
von: Haliassos, Alexandros, et al.
Veröffentlicht: (2026)
von: Haliassos, Alexandros, et al.
Veröffentlicht: (2026)
A Dual-Attention Learning Network with Word and Sentence Embedding for Medical Visual Question Answering
von: Huang, Xiaofei, et al.
Veröffentlicht: (2022)
von: Huang, Xiaofei, et al.
Veröffentlicht: (2022)
FALIP: Visual Prompt as Foveal Attention Boosts CLIP Zero-Shot Performance
von: Zhuang, Jiedong, et al.
Veröffentlicht: (2024)
von: Zhuang, Jiedong, et al.
Veröffentlicht: (2024)
Attention Head Purification: A New Perspective to Harness CLIP for Domain Generalization
von: Wang, Yingfan, et al.
Veröffentlicht: (2024)
von: Wang, Yingfan, et al.
Veröffentlicht: (2024)
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference
von: Yang, Yuhang, et al.
Veröffentlicht: (2024)
von: Yang, Yuhang, et al.
Veröffentlicht: (2024)
Distinctive Image Captioning: Leveraging Ground Truth Captions in CLIP Guided Reinforcement Learning
von: Chaffin, Antoine, et al.
Veröffentlicht: (2024)
von: Chaffin, Antoine, et al.
Veröffentlicht: (2024)
PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model
von: Arif, Kazi Hasan Ibn, et al.
Veröffentlicht: (2025)
von: Arif, Kazi Hasan Ibn, et al.
Veröffentlicht: (2025)
SPARO: Selective Attention for Robust and Compositional Transformer Encodings for Vision
von: Vani, Ankit, et al.
Veröffentlicht: (2024)
von: Vani, Ankit, et al.
Veröffentlicht: (2024)
SuperCLIP: CLIP with Simple Classification Supervision
von: Zhao, Weiheng, et al.
Veröffentlicht: (2025)
von: Zhao, Weiheng, et al.
Veröffentlicht: (2025)
DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving
von: Zhou, Yang, et al.
Veröffentlicht: (2026)
von: Zhou, Yang, et al.
Veröffentlicht: (2026)
SmartPretrain: Model-Agnostic and Dataset-Agnostic Representation Learning for Motion Prediction
von: Zhou, Yang, et al.
Veröffentlicht: (2024)
von: Zhou, Yang, et al.
Veröffentlicht: (2024)
SmartRefine: A Scenario-Adaptive Refinement Framework for Efficient Motion Prediction
von: Zhou, Yang, et al.
Veröffentlicht: (2024)
von: Zhou, Yang, et al.
Veröffentlicht: (2024)
Complete Gaussian Splats from a Single Image with Denoising Diffusion Models
von: Liao, Ziwei, et al.
Veröffentlicht: (2025)
von: Liao, Ziwei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Large Self-Supervised Models Bridge the Gap in Domain Adaptive Object Detection
von: Lavoie, Marc-Antoine, et al.
Veröffentlicht: (2025) -
Feature Density Estimation for Out-of-Distribution Detection via Normalizing Flows
von: Cook, Evan D., et al.
Veröffentlicht: (2024) -
Image-to-Lidar Relational Distillation for Autonomous Driving Data
von: Mahmoud, Anas, et al.
Veröffentlicht: (2024) -
JDT3D: Addressing the Gaps in LiDAR-Based Tracking-by-Attention
von: Cheong, Brian, et al.
Veröffentlicht: (2024) -
PSA-SSL: Pose and Size-aware Self-Supervised Learning on LiDAR Point Clouds
von: Nisar, Barza, et al.
Veröffentlicht: (2025)