RELOCATE: A Simple Training-Free Baseline for Visual Query Localization Using Region-Based Representations
Fuente:
arXiv
Saved in:
| Main Authors: | Khosla, Savya, T V, Sethuraman, Schwing, Alexander, Hoiem, Derek |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
REN: Fast and Efficient Region Encodings from Patch-Based Image Encoders
by: Khosla, Savya, et al.
Published: (2025)
by: Khosla, Savya, et al.
Published: (2025)
T-REN: Learning Text-Aligned Region Tokens Improves Dense Vision-Language Alignment and Scalability
by: Khosla, Savya, et al.
Published: (2026)
by: Khosla, Savya, et al.
Published: (2026)
FRAME: Pre-Training Video Feature Representations via Anticipation and Memory
by: TV, Sethuraman, et al.
Published: (2025)
by: TV, Sethuraman, et al.
Published: (2025)
Region-Based Representations Revisited
by: Shlapentokh-Rothman, Michal, et al.
Published: (2024)
by: Shlapentokh-Rothman, Michal, et al.
Published: (2024)
Stress Tests REVEAL Fragile Temporal and Visual Grounding in Video-Language Models
by: T V, Sethuraman, et al.
Published: (2026)
by: T V, Sethuraman, et al.
Published: (2026)
Visual Program Distillation with Template-Based Augmentation
by: Shlapentokh-Rothman, Michal, et al.
Published: (2024)
by: Shlapentokh-Rothman, Michal, et al.
Published: (2024)
Decomposing Queries into Tool Calls for Long-Video Keyframe Retrieval
by: Shlapentokh-Rothman, Michal, et al.
Published: (2026)
by: Shlapentokh-Rothman, Michal, et al.
Published: (2026)
TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models
by: Xiao, Yao, et al.
Published: (2025)
by: Xiao, Yao, et al.
Published: (2025)
A Simple and Better Baseline for Visual Grounding
by: Wang, Jingchao, et al.
Published: (2025)
by: Wang, Jingchao, et al.
Published: (2025)
Anytime Continual Learning for Open Vocabulary Classification
by: Zhu, Zhen, et al.
Published: (2024)
by: Zhu, Zhen, et al.
Published: (2024)
Continual Learning in Open-vocabulary Classification with Complementary Memory Systems
by: Zhu, Zhen, et al.
Published: (2023)
by: Zhu, Zhen, et al.
Published: (2023)
A Simple-but-effective Baseline for Training-free Class-Agnostic Counting
by: Lin, Yuhao, et al.
Published: (2024)
by: Lin, Yuhao, et al.
Published: (2024)
SimToken: A Simple Baseline for Referring Audio-Visual Segmentation
by: Jin, Dian, et al.
Published: (2025)
by: Jin, Dian, et al.
Published: (2025)
The Curse of Conditions: Analyzing and Improving Optimal Transport for Conditional Flow-Based Generation
by: Cheng, Ho Kei, et al.
Published: (2025)
by: Cheng, Ho Kei, et al.
Published: (2025)
Studying Classifier(-Free) Guidance From a Classifier-Centric Perspective
by: Zhao, Xiaoming, et al.
Published: (2025)
by: Zhao, Xiaoming, et al.
Published: (2025)
SimBase: A Simple Baseline for Temporal Video Grounding
by: Bao, Peijun, et al.
Published: (2024)
by: Bao, Peijun, et al.
Published: (2024)
SimROD: A Simple Baseline for Raw Object Detection with Global and Local Enhancements
by: Xie, Haiyang, et al.
Published: (2025)
by: Xie, Haiyang, et al.
Published: (2025)
MonoPatchNeRF: Improving Neural Radiance Fields with Patch-based Monocular Guidance
by: Wu, Yuqun, et al.
Published: (2024)
by: Wu, Yuqun, et al.
Published: (2024)
Plenoptic PNG: Real-Time Neural Radiance Fields in 150 KB
by: Lee, Jae Yong, et al.
Published: (2024)
by: Lee, Jae Yong, et al.
Published: (2024)
SimpliHuMoN: Simplifying Human Motion Prediction
by: Agrawal, Aadya, et al.
Published: (2026)
by: Agrawal, Aadya, et al.
Published: (2026)
SimpleGVR: A Simple Baseline for Latent-Cascaded Video Super-Resolution
by: Xie, Liangbin, et al.
Published: (2025)
by: Xie, Liangbin, et al.
Published: (2025)
Bridging Diffusion Models and 3D Representations: A 3D Consistent Super-Resolution Framework
by: Chen, Yi-Ting, et al.
Published: (2025)
by: Chen, Yi-Ting, et al.
Published: (2025)
A Simple Baseline for Efficient Hand Mesh Reconstruction
by: Zhou, Zhishan, et al.
Published: (2024)
by: Zhou, Zhishan, et al.
Published: (2024)
NoPo-Avatar: Generalizable and Animatable Avatars from Sparse Inputs without Human Poses
by: Wen, Jing, et al.
Published: (2025)
by: Wen, Jing, et al.
Published: (2025)
LIFe-GoM: Generalizable Human Rendering with Learned Iterative Feedback Over Multi-Resolution Gaussians-on-Mesh
by: Wen, Jing, et al.
Published: (2025)
by: Wen, Jing, et al.
Published: (2025)
Freeze and Cluster: A Simple Baseline for Rehearsal-Free Continual Category Discovery
by: Zhang, Chuyu, et al.
Published: (2025)
by: Zhang, Chuyu, et al.
Published: (2025)
ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers
by: Huang, Lianghua, et al.
Published: (2024)
by: Huang, Lianghua, et al.
Published: (2024)
Revisiting Simple Baselines for In-The-Wild Deepfake Detection
by: Castaneda, Orlando, et al.
Published: (2025)
by: Castaneda, Orlando, et al.
Published: (2025)
Variational Rectified Flow Matching
by: Guo, Pengsheng, et al.
Published: (2025)
by: Guo, Pengsheng, et al.
Published: (2025)
GoMAvatar: Efficient Animatable Human Modeling from Monocular Video Using Gaussians-on-Mesh
by: Wen, Jing, et al.
Published: (2024)
by: Wen, Jing, et al.
Published: (2024)
A Simple and Efficient Baseline for Zero-Shot Generative Classification
by: Qi, Zipeng, et al.
Published: (2024)
by: Qi, Zipeng, et al.
Published: (2024)
Towards Label-Efficient Human Matting: A Simple Baseline for Weakly Semi-Supervised Trimap-Free Human Matting
by: Kim, Beomyoung, et al.
Published: (2024)
by: Kim, Beomyoung, et al.
Published: (2024)
Towards Visual Query Localization in the 3D World
by: Peng, Liang, et al.
Published: (2026)
by: Peng, Liang, et al.
Published: (2026)
Training Free Zero-Shot Visual Anomaly Localization via Diffusion Inversion
by: Hicsonmez, Samet, et al.
Published: (2026)
by: Hicsonmez, Samet, et al.
Published: (2026)
SceneDiff: A Benchmark and Method for Multiview Object Change Detection
by: Wu, Yuqun, et al.
Published: (2025)
by: Wu, Yuqun, et al.
Published: (2025)
MedVLThinker: Simple Baselines for Multimodal Medical Reasoning
by: Huang, Xiaoke, et al.
Published: (2025)
by: Huang, Xiaoke, et al.
Published: (2025)
SimpleMatch: A Simple and Strong Baseline for Semantic Correspondence
by: Jin, Hailing, et al.
Published: (2026)
by: Jin, Hailing, et al.
Published: (2026)
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
by: Wu, Size, et al.
Published: (2025)
by: Wu, Size, et al.
Published: (2025)
Upsample Anything: A Simple and Hard to Beat Baseline for Feature Upsampling
by: Seo, Minseok, et al.
Published: (2025)
by: Seo, Minseok, et al.
Published: (2025)
HERO-VQL: Hierarchical, Egocentric and Robust Visual Query Localization
by: Chang, Joohyun, et al.
Published: (2025)
by: Chang, Joohyun, et al.
Published: (2025)
Similar Items
-
REN: Fast and Efficient Region Encodings from Patch-Based Image Encoders
by: Khosla, Savya, et al.
Published: (2025) -
T-REN: Learning Text-Aligned Region Tokens Improves Dense Vision-Language Alignment and Scalability
by: Khosla, Savya, et al.
Published: (2026) -
FRAME: Pre-Training Video Feature Representations via Anticipation and Memory
by: TV, Sethuraman, et al.
Published: (2025) -
Region-Based Representations Revisited
by: Shlapentokh-Rothman, Michal, et al.
Published: (2024) -
Stress Tests REVEAL Fragile Temporal and Visual Grounding in Video-Language Models
by: T V, Sethuraman, et al.
Published: (2026)