HERO-VQL: Hierarchical, Egocentric and Robust Visual Query Localization
Fuente:
arXiv
Saved in:
| Main Authors: | Chang, Joohyun, Hong, Soyeon, Lee, Hyogun, Ha, Seong Jong, Lee, Dongho, Kim, Seong Tae, Choi, Jinwoo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Infusing Environmental Captions for Long-Form Video Language Grounding
by: Lee, Hyogun, et al.
Published: (2024)
by: Lee, Hyogun, et al.
Published: (2024)
CA^2ST: Cross-Attention in Audio, Space, and Time for Holistic Video Recognition
by: Lee, Jongseo, et al.
Published: (2025)
by: Lee, Jongseo, et al.
Published: (2025)
PCEvE: Part Contribution Evaluation Based Model Explanation for Human Figure Drawing Assessment and Beyond
by: Lee, Jongseo, et al.
Published: (2024)
by: Lee, Jongseo, et al.
Published: (2024)
CAST: Cross-Attention in Space and Time for Video Action Recognition
by: Lee, Dongho, et al.
Published: (2023)
by: Lee, Dongho, et al.
Published: (2023)
HiCM$^2$: Hierarchical Compact Memory Modeling for Dense Video Captioning
by: Kim, Minkuk, et al.
Published: (2024)
by: Kim, Minkuk, et al.
Published: (2024)
Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition
by: Lee, Jongseo, et al.
Published: (2025)
by: Lee, Jongseo, et al.
Published: (2025)
PCBEAR: Pose Concept Bottleneck for Explainable Action Recognition
by: Lee, Jongseo, et al.
Published: (2025)
by: Lee, Jongseo, et al.
Published: (2025)
Flashback: Memory-Driven Zero-shot, Real-time Video Anomaly Detection
by: Lee, Hyogun, et al.
Published: (2025)
by: Lee, Hyogun, et al.
Published: (2025)
Do You Remember? Dense Video Captioning with Cross-Modal Memory Retrieval
by: Kim, Minkuk, et al.
Published: (2024)
by: Kim, Minkuk, et al.
Published: (2024)
PRVQL: Progressive Knowledge-guided Refinement for Robust Egocentric Visual Query Localization
by: Fan, Bing, et al.
Published: (2025)
by: Fan, Bing, et al.
Published: (2025)
Leveraging Textual Compositional Reasoning for Robust Change Captioning
by: Park, Kyu Ri, et al.
Published: (2025)
by: Park, Kyu Ri, et al.
Published: (2025)
PoseBridge: Bridging the Skeletonization Gap for Zero-Shot Skeleton-Based Action Recognition
by: Lee, Sanghyeon, et al.
Published: (2026)
by: Lee, Sanghyeon, et al.
Published: (2026)
Learning Multi-frame and Monocular Prior for Estimating Geometry in Dynamic Scenes
by: Park, Seong Hyeon, et al.
Published: (2025)
by: Park, Seong Hyeon, et al.
Published: (2025)
Robust Single Rotation Averaging Revisited
by: Lee, Seong Hun, et al.
Published: (2023)
by: Lee, Seong Hun, et al.
Published: (2023)
Alignment Scores: Robust Metrics for Multiview Pose Accuracy Evaluation
by: Lee, Seong Hun, et al.
Published: (2024)
by: Lee, Seong Hun, et al.
Published: (2024)
Retrieval-Augmented Natural Language Reasoning for Explainable Visual Question Answering
by: Lim, Su Hyeon, et al.
Published: (2024)
by: Lim, Su Hyeon, et al.
Published: (2024)
Hierarchical Classification for Improved Histopathology Image Analysis
by: Byeon, Keunho, et al.
Published: (2026)
by: Byeon, Keunho, et al.
Published: (2026)
Local Representative Token Guided Merging for Text-to-Image Generation
by: Lee, Min-Jeong, et al.
Published: (2025)
by: Lee, Min-Jeong, et al.
Published: (2025)
MonoWAD: Weather-Adaptive Diffusion Model for Robust Monocular 3D Object Detection
by: Oh, Youngmin, et al.
Published: (2024)
by: Oh, Youngmin, et al.
Published: (2024)
Egocentric Whole-Body Human Mesh Recovery with Prior-Guided Learning
by: Na, Soyeon, et al.
Published: (2026)
by: Na, Soyeon, et al.
Published: (2026)
HERO: Hierarchical Extrapolation and Refresh for Efficient World Models
by: Song, Quanjian, et al.
Published: (2025)
by: Song, Quanjian, et al.
Published: (2025)
Online Episodic Memory Visual Query Localization with Egocentric Streaming Object Memory
by: Manigrasso, Zaira, et al.
Published: (2024)
by: Manigrasso, Zaira, et al.
Published: (2024)
LLaVA Needs More Knowledge: Retrieval Augmented Natural Language Generation with Knowledge Graph for Explaining Thoracic Pathologies
by: Hamza, Ameer, et al.
Published: (2024)
by: Hamza, Ameer, et al.
Published: (2024)
Adversarial Wear and Tear: Exploiting Natural Damage for Generating Physical-World Adversarial Examples
by: Irshad, Samra, et al.
Published: (2025)
by: Irshad, Samra, et al.
Published: (2025)
Diverse Rare Sample Generation with Pretrained GANs
by: Lee, Subeen, et al.
Published: (2024)
by: Lee, Subeen, et al.
Published: (2024)
Learning to Refuse: Refusal-Aware Reinforcement Fine-Tuning for Hard-Irrelevant Queries in Video Temporal Grounding
by: Lee, Jin-Seop, et al.
Published: (2025)
by: Lee, Jin-Seop, et al.
Published: (2025)
DiGIT: Multi-Dilated Gated Encoder and Central-Adjacent Region Integrated Decoder for Temporal Action Detection Transformer
by: Kim, Ho-Joong, et al.
Published: (2025)
by: Kim, Ho-Joong, et al.
Published: (2025)
Integrating Query-aware Segmentation and Cross-Attention for Robust VQA
by: Choi, Wonjun, et al.
Published: (2024)
by: Choi, Wonjun, et al.
Published: (2024)
VioPose: Violin Performance 4D Pose Estimation by Hierarchical Audiovisual Inference
by: Yoo, Seong Jong, et al.
Published: (2024)
by: Yoo, Seong Jong, et al.
Published: (2024)
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding
by: Ahn, Geo, et al.
Published: (2026)
by: Ahn, Geo, et al.
Published: (2026)
SurgX: Neuron-Concept Association for Explainable Surgical Phase Recognition
by: Kim, Ka Young, et al.
Published: (2025)
by: Kim, Ka Young, et al.
Published: (2025)
Multi-Context Temporal Consistent Modeling for Referring Video Object Segmentation
by: Choi, Sun-Hyuk, et al.
Published: (2025)
by: Choi, Sun-Hyuk, et al.
Published: (2025)
Context-Based Visual-Language Place Recognition
by: Woo, Soojin, et al.
Published: (2024)
by: Woo, Soojin, et al.
Published: (2024)
ViKey: Enhancing Temporal Understanding in Videos via Visual Prompting
by: Lee, Yeonkyung, et al.
Published: (2026)
by: Lee, Yeonkyung, et al.
Published: (2026)
AM-SORT: Adaptable Motion Predictor with Historical Trajectory Embedding for Multi-Object Tracking
by: Kim, Vitaliy, et al.
Published: (2024)
by: Kim, Vitaliy, et al.
Published: (2024)
Enhancing Spatio-Temporal Zero-shot Action Recognition with Language-driven Description Attributes
by: Kim, Yehna, et al.
Published: (2025)
by: Kim, Yehna, et al.
Published: (2025)
WWW: A Unified Framework for Explaining What, Where and Why of Neural Networks by Interpretation of Neuron Concepts
by: Ahn, Yong Hyun, et al.
Published: (2024)
by: Ahn, Yong Hyun, et al.
Published: (2024)
Mask-Free Neuron Concept Annotation for Interpreting Neural Networks in Medical Domain
by: Kim, Hyeon Bae, et al.
Published: (2024)
by: Kim, Hyeon Bae, et al.
Published: (2024)
TIFu: Tri-directional Implicit Function for High-Fidelity 3D Character Reconstruction
by: Lim, Byoungsung, et al.
Published: (2024)
by: Lim, Byoungsung, et al.
Published: (2024)
Anchoring and Rescaling Attention for Semantically Coherent Inbetweening
by: Choi, Tae Eun, et al.
Published: (2026)
by: Choi, Tae Eun, et al.
Published: (2026)
Similar Items
-
Infusing Environmental Captions for Long-Form Video Language Grounding
by: Lee, Hyogun, et al.
Published: (2024) -
CA^2ST: Cross-Attention in Audio, Space, and Time for Holistic Video Recognition
by: Lee, Jongseo, et al.
Published: (2025) -
PCEvE: Part Contribution Evaluation Based Model Explanation for Human Figure Drawing Assessment and Beyond
by: Lee, Jongseo, et al.
Published: (2024) -
CAST: Cross-Attention in Space and Time for Video Action Recognition
by: Lee, Dongho, et al.
Published: (2023) -
HiCM$^2$: Hierarchical Compact Memory Modeling for Dense Video Captioning
by: Kim, Minkuk, et al.
Published: (2024)