PRVQL: Progressive Knowledge-guided Refinement for Robust Egocentric Visual Query Localization
Fuente:
arXiv
Saved in:
| Main Authors: | Fan, Bing, Feng, Yunhe, Tian, Yapeng, Liang, James Chenhao, Lin, Yuewei, Huang, Yan, Fan, Heng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Visual Query Segmentation in the Wild
by: Fan, Bing, et al.
Published: (2026)
by: Fan, Bing, et al.
Published: (2026)
LoReTrack: Efficient and Accurate Low-Resolution Transformer Tracking
by: Dong, Shaohua, et al.
Published: (2024)
by: Dong, Shaohua, et al.
Published: (2024)
HERO-VQL: Hierarchical, Egocentric and Robust Visual Query Localization
by: Chang, Joohyun, et al.
Published: (2025)
by: Chang, Joohyun, et al.
Published: (2025)
Benchmarking the Robustness of UAV Tracking Against Common Corruptions
by: Liu, Xiaoqiong, et al.
Published: (2024)
by: Liu, Xiaoqiong, et al.
Published: (2024)
Online Episodic Memory Visual Query Localization with Egocentric Streaming Object Memory
by: Manigrasso, Zaira, et al.
Published: (2024)
by: Manigrasso, Zaira, et al.
Published: (2024)
Spatial Orthogonal Refinement for Robust RGB-Event Visual Object Tracking
by: Huang, Dexing, et al.
Published: (2026)
by: Huang, Dexing, et al.
Published: (2026)
T-VSL: Text-Guided Visual Sound Source Localization in Mixtures
by: Mahmud, Tanvir, et al.
Published: (2024)
by: Mahmud, Tanvir, et al.
Published: (2024)
Robust Egocentric Visual Attention Prediction Through Language-guided Scene Context-aware Learning
by: Park, Sungjune, et al.
Published: (2026)
by: Park, Sungjune, et al.
Published: (2026)
From Waveforms to Pixels: A Survey on Audio-Visual Segmentation
by: Li, Jia, et al.
Published: (2025)
by: Li, Jia, et al.
Published: (2025)
Student-Oriented Teacher Knowledge Refinement for Knowledge Distillation
by: Shen, Chaomin, et al.
Published: (2024)
by: Shen, Chaomin, et al.
Published: (2024)
Language-Guided Joint Audio-Visual Editing via One-Shot Adaptation
by: Liang, Susan, et al.
Published: (2024)
by: Liang, Susan, et al.
Published: (2024)
Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video Grounding
by: Gu, Xin, et al.
Published: (2025)
by: Gu, Xin, et al.
Published: (2025)
RefineNet: Enhancing Text-to-Image Conversion with High-Resolution and Detail Accuracy through Hierarchical Transformers and Progressive Refinement
by: Shi, Fan
Published: (2023)
by: Shi, Fan
Published: (2023)
Multimodal Query-guided Object Localization
by: Tripathi, Aditay, et al.
Published: (2022)
by: Tripathi, Aditay, et al.
Published: (2022)
High-Quality Sound Separation Across Diverse Categories via Visually-Guided Generative Modeling
by: Huang, Chao, et al.
Published: (2025)
by: Huang, Chao, et al.
Published: (2025)
High-Quality Visually-Guided Sound Separation from Diverse Categories
by: Huang, Chao, et al.
Published: (2023)
by: Huang, Chao, et al.
Published: (2023)
EAGLE: Episodic Appearance- and Geometry-aware Memory for Unified 2D-3D Visual Query Localization in Egocentric Vision
by: Cao, Yifei, et al.
Published: (2025)
by: Cao, Yifei, et al.
Published: (2025)
Egocentric Visibility-Aware Human Pose Estimation
by: Dai, Peng, et al.
Published: (2026)
by: Dai, Peng, et al.
Published: (2026)
Efficient Temporal Action Segmentation via Boundary-aware Query Voting
by: Wang, Peiyao, et al.
Published: (2024)
by: Wang, Peiyao, et al.
Published: (2024)
Robust Ego-Exo Correspondence with Long-Term Memory
by: Hu, Yijun, et al.
Published: (2025)
by: Hu, Yijun, et al.
Published: (2025)
ARGaze: Autoregressive Transformers for Online Egocentric Gaze Estimation
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
Visual Knowledge in the Big Model Era: Retrospect and Prospect
by: Wang, Wenguan, et al.
Published: (2024)
by: Wang, Wenguan, et al.
Published: (2024)
Spherical World-Locking for Audio-Visual Localization in Egocentric Videos
by: Yun, Heeseung, et al.
Published: (2024)
by: Yun, Heeseung, et al.
Published: (2024)
Towards Visual Query Localization in the 3D World
by: Peng, Liang, et al.
Published: (2026)
by: Peng, Liang, et al.
Published: (2026)
Progressive Language-guided Visual Learning for Multi-Task Visual Grounding
by: Wang, Jingchao, et al.
Published: (2025)
by: Wang, Jingchao, et al.
Published: (2025)
Domain-invariant Progressive Knowledge Distillation for UAV-based Object Detection
by: Yao, Liang, et al.
Published: (2024)
by: Yao, Liang, et al.
Published: (2024)
Multimodal Knowledge Distillation for Egocentric Action Recognition Robust to Missing Modalities
by: Santos-Villafranca, Maria, et al.
Published: (2025)
by: Santos-Villafranca, Maria, et al.
Published: (2025)
QA-Dragon: Query-Aware Dynamic RAG System for Knowledge-Intensive Visual Question Answering
by: Jiang, Zhuohang, et al.
Published: (2025)
by: Jiang, Zhuohang, et al.
Published: (2025)
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement
by: Yin, Weijie, et al.
Published: (2025)
by: Yin, Weijie, et al.
Published: (2025)
Robust Visual Localization via Semantic-Guided Multi-Scale Transformer
by: Tian, Zhongtao, et al.
Published: (2025)
by: Tian, Zhongtao, et al.
Published: (2025)
Simplicity Prevails: The Emergence of Generalizable AIGI Detection in Visual Foundation Models
by: Zhou, Yue, et al.
Published: (2026)
by: Zhou, Yue, et al.
Published: (2026)
Robust Egocentric Referring Video Object Segmentation via Dual-Modal Causal Intervention
by: Liu, Haijing, et al.
Published: (2025)
by: Liu, Haijing, et al.
Published: (2025)
Aligning Human Knowledge with Visual Concepts Towards Explainable Medical Image Classification
by: Gao, Yunhe, et al.
Published: (2024)
by: Gao, Yunhe, et al.
Published: (2024)
Towards Long-Form Spatio-Temporal Video Grounding
by: Gu, Xin, et al.
Published: (2026)
by: Gu, Xin, et al.
Published: (2026)
X-Oscar: A Progressive Framework for High-quality Text-guided 3D Animatable Avatar Generation
by: Ma, Yiwei, et al.
Published: (2024)
by: Ma, Yiwei, et al.
Published: (2024)
Progressive Query Refinement Framework for Bird's-Eye-View Semantic Segmentation from Surrounding Images
by: Choi, Dooseop, et al.
Published: (2024)
by: Choi, Dooseop, et al.
Published: (2024)
Collapse-Aware Triplet Decoupling for Adversarially Robust Image Retrieval
by: Tian, Qiwei, et al.
Published: (2023)
by: Tian, Qiwei, et al.
Published: (2023)
SaSR-Net: Source-Aware Semantic Representation Network for Enhancing Audio-Visual Question Answering
by: Yang, Tianyu, et al.
Published: (2024)
by: Yang, Tianyu, et al.
Published: (2024)
In the Eye of Transformer: Global-Local Correlation for Egocentric Gaze Estimation
by: Lai, Bolin, et al.
Published: (2022)
by: Lai, Bolin, et al.
Published: (2022)
DD-RobustBench: An Adversarial Robustness Benchmark for Dataset Distillation
by: Wu, Yifan, et al.
Published: (2024)
by: Wu, Yifan, et al.
Published: (2024)
Similar Items
-
Towards Visual Query Segmentation in the Wild
by: Fan, Bing, et al.
Published: (2026) -
LoReTrack: Efficient and Accurate Low-Resolution Transformer Tracking
by: Dong, Shaohua, et al.
Published: (2024) -
HERO-VQL: Hierarchical, Egocentric and Robust Visual Query Localization
by: Chang, Joohyun, et al.
Published: (2025) -
Benchmarking the Robustness of UAV Tracking Against Common Corruptions
by: Liu, Xiaoqiong, et al.
Published: (2024) -
Online Episodic Memory Visual Query Localization with Egocentric Streaming Object Memory
by: Manigrasso, Zaira, et al.
Published: (2024)