Temporal Attention for Cross-View Sequential Image Localization
Fuente:
arXiv
Saved in:
| Main Authors: | Yuan, Dong, Maire, Frederic, Dayoub, Feras |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Robust Scene Change Detection Using Visual Foundation Models and Cross-Attention Mechanisms
by: Lin, Chun-Jung, et al.
Published: (2024)
by: Lin, Chun-Jung, et al.
Published: (2024)
Vision Foundation Models for Domain Generalisable Cross-View Localisation in Planetary Ground-Aerial Robotic Teams
by: Holden, Lachlan, et al.
Published: (2026)
by: Holden, Lachlan, et al.
Published: (2026)
Wasserstein Distance-based Expansion of Low-Density Latent Regions for Unknown Class Detection
by: Mallick, Prakash, et al.
Published: (2024)
by: Mallick, Prakash, et al.
Published: (2024)
Detecting Precise Hand Touch Moments in Egocentric Video
by: Nguyen, Huy Anh, et al.
Published: (2026)
by: Nguyen, Huy Anh, et al.
Published: (2026)
SceneEdited: A City-Scale Benchmark for 3D HD Map Updating via Image-Guided Change Detection
by: Lin, Chun-Jung, et al.
Published: (2025)
by: Lin, Chun-Jung, et al.
Published: (2025)
A Physical Agentic Loop for Language-Guided Grasping with Execution-State Monitoring
by: Wang, Wenze, et al.
Published: (2026)
by: Wang, Wenze, et al.
Published: (2026)
Segment Beyond View: Handling Partially Missing Modality for Audio-Visual Semantic Segmentation
by: Wu, Renjie, et al.
Published: (2023)
by: Wu, Renjie, et al.
Published: (2023)
To Ask or Not to Ask? Detecting Absence of Information in Vision and Language Navigation
by: Abraham, Savitha Sam, et al.
Published: (2024)
by: Abraham, Savitha Sam, et al.
Published: (2024)
QueryAdapter: Rapid Adaptation of Vision-Language Models in Response to Natural Language Queries
by: Chapman, Nicolas Harvey, et al.
Published: (2025)
by: Chapman, Nicolas Harvey, et al.
Published: (2025)
Embodied Domain Adaptation for Object Detection
by: Shi, Xiangyu, et al.
Published: (2025)
by: Shi, Xiangyu, et al.
Published: (2025)
AIMC-Spec: A Benchmark Dataset for Automatic Intrapulse Modulation Classification under Variable Noise Conditions
by: Cocks, Sebastian L., et al.
Published: (2026)
by: Cocks, Sebastian L., et al.
Published: (2026)
VAGeo: View-specific Attention for Cross-View Object Geo-Localization
by: Li, Zhongyang, et al.
Published: (2025)
by: Li, Zhongyang, et al.
Published: (2025)
KITE: Keyframe-Indexed Tokenized Evidence for VLM-Based Robot Failure Analysis
by: Hosseinzadeh, Mehdi, et al.
Published: (2026)
by: Hosseinzadeh, Mehdi, et al.
Published: (2026)
Improving Online Source-free Domain Adaptation for Object Detection by Unsupervised Data Acquisition
by: Shi, Xiangyu, et al.
Published: (2023)
by: Shi, Xiangyu, et al.
Published: (2023)
PoIFusion: Multi-Modal 3D Object Detection via Fusion at Points of Interest
by: Deng, Jiajun, et al.
Published: (2024)
by: Deng, Jiajun, et al.
Published: (2024)
3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer
by: Deng, Jiajun, et al.
Published: (2025)
by: Deng, Jiajun, et al.
Published: (2025)
Cross-View Image Set Geo-Localization
by: Wu, Qiong, et al.
Published: (2024)
by: Wu, Qiong, et al.
Published: (2024)
ViewBridge:Revisiting Cross-View Localization from Image Matching
by: Xia, Panwang, et al.
Published: (2025)
by: Xia, Panwang, et al.
Published: (2025)
SmartWay: Enhanced Waypoint Prediction and Backtracking for Zero-Shot Vision-and-Language Navigation
by: Shi, Xiangyu, et al.
Published: (2025)
by: Shi, Xiangyu, et al.
Published: (2025)
TANGO: Traversability-Aware Navigation with Local Metric Control for Topological Goals
by: Podgorski, Stefan, et al.
Published: (2025)
by: Podgorski, Stefan, et al.
Published: (2025)
Recurrent Cross-View Object Geo-Localization
by: Zhang, Xiaohan, et al.
Published: (2025)
by: Zhang, Xiaohan, et al.
Published: (2025)
Aligned Novel View Image and Geometry Synthesis via Cross-modal Attention Instillation
by: Kwak, Min-Seop, et al.
Published: (2025)
by: Kwak, Min-Seop, et al.
Published: (2025)
An Anisotropic Cross-View Texture Transfer with Multi-Reference Non-Local Attention for CT Slice Interpolation
by: Uhm, Kwang-Hyun, et al.
Published: (2025)
by: Uhm, Kwang-Hyun, et al.
Published: (2025)
Reducing Label Dependency for Underwater Scene Understanding: A Survey of Datasets, Techniques and Applications
by: Raine, Scarlett, et al.
Published: (2024)
by: Raine, Scarlett, et al.
Published: (2024)
Cross-View Geo-Localization with Street-View and VHR Satellite Imagery in Decentrality Settings
by: Xia, Panwang, et al.
Published: (2024)
by: Xia, Panwang, et al.
Published: (2024)
Progressive Cross-Stream Cooperation in Spatial and Temporal Domain for Action Localization
by: Su, Rui, et al.
Published: (2019)
by: Su, Rui, et al.
Published: (2019)
Layout-to-Image Generation with Localized Descriptions using ControlNet with Cross-Attention Control
by: Lukovnikov, Denis, et al.
Published: (2024)
by: Lukovnikov, Denis, et al.
Published: (2024)
MAVR-Net: Robust Multi-View Learning for MAV Action Recognition with Cross-View Attention
by: Zhang, Nengbo, et al.
Published: (2025)
by: Zhang, Nengbo, et al.
Published: (2025)
Cross Spatial Temporal Fusion Attention for Remote Sensing Object Detection via Image Feature Matching
by: Amit, Abu Sadat Mohammad Salehin, et al.
Published: (2025)
by: Amit, Abu Sadat Mohammad Salehin, et al.
Published: (2025)
MRGeo: Robust Cross-View Geo-Localization of Corrupted Images via Spatial and Channel Feature Enhancement
by: Wu, Le, et al.
Published: (2026)
by: Wu, Le, et al.
Published: (2026)
LoLep: Single-View View Synthesis with Locally-Learned Planes and Self-Attention Occlusion Inference
by: Wang, Cong, et al.
Published: (2023)
by: Wang, Cong, et al.
Published: (2023)
Cognitive-Inspired Hierarchical Attention Fusion With Visual and Textual for Cross-Domain Sequential Recommendation
by: Wu, Wangyu, et al.
Published: (2025)
by: Wu, Wangyu, et al.
Published: (2025)
Tag-Enriched Multi-Attention with Large Language Models for Cross-Domain Sequential Recommendation
by: Wu, Wangyu, et al.
Published: (2025)
by: Wu, Wangyu, et al.
Published: (2025)
Attention-Enhanced Cross-modal Localization Between 360 Images and Point Clouds
by: Zhao, Zhipeng, et al.
Published: (2022)
by: Zhao, Zhipeng, et al.
Published: (2022)
Seeing Across Time and Views: Multi-Temporal Cross-View Learning for Robust Video Person Re-Identification
by: Rashidunnabi, Md, et al.
Published: (2025)
by: Rashidunnabi, Md, et al.
Published: (2025)
MatchAttention: Matching the Relative Positions for High-Resolution Cross-View Matching
by: Yan, Tingman, et al.
Published: (2025)
by: Yan, Tingman, et al.
Published: (2025)
IMDPrompter: Adapting SAM to Image Manipulation Detection by Cross-View Automated Prompt Learning
by: Zhang, Quan, et al.
Published: (2025)
by: Zhang, Quan, et al.
Published: (2025)
UniABG: Unified Adversarial View Bridging and Graph Correspondence for Unsupervised Cross-View Geo-Localization
by: Chen, Cuiqun, et al.
Published: (2025)
by: Chen, Cuiqun, et al.
Published: (2025)
CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization
by: Xia, Rui, et al.
Published: (2025)
by: Xia, Rui, et al.
Published: (2025)
CV-Cities: Advancing Cross-View Geo-Localization in Global Cities
by: Huang, Gaoshuang, et al.
Published: (2024)
by: Huang, Gaoshuang, et al.
Published: (2024)
Similar Items
-
Robust Scene Change Detection Using Visual Foundation Models and Cross-Attention Mechanisms
by: Lin, Chun-Jung, et al.
Published: (2024) -
Vision Foundation Models for Domain Generalisable Cross-View Localisation in Planetary Ground-Aerial Robotic Teams
by: Holden, Lachlan, et al.
Published: (2026) -
Wasserstein Distance-based Expansion of Low-Density Latent Regions for Unknown Class Detection
by: Mallick, Prakash, et al.
Published: (2024) -
Detecting Precise Hand Touch Moments in Egocentric Video
by: Nguyen, Huy Anh, et al.
Published: (2026) -
SceneEdited: A City-Scale Benchmark for 3D HD Map Updating via Image-Guided Change Detection
by: Lin, Chun-Jung, et al.
Published: (2025)