VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale
Fuente:
arXiv
Saved in:
| Main Authors: | Kulkarni, Parth Parag, Gupta, Rohit, Chhipa, Prakash Chandra, Shah, Mubarak |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CityGuessr: City-Level Video Geo-Localization on a Global Scale
by: Kulkarni, Parth Parag, et al.
Published: (2024)
by: Kulkarni, Parth Parag, et al.
Published: (2024)
GAEA: A Geolocation Aware Conversational Assistant
by: Campos, Ron, et al.
Published: (2025)
by: Campos, Ron, et al.
Published: (2025)
VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues
by: Swetha, Sirnam, et al.
Published: (2025)
by: Swetha, Sirnam, et al.
Published: (2025)
StretchySnake: Flexible SSM Training Unlocks Action Recognition Across Spatio-Temporal Scales
by: Siddiqui, Nyle, et al.
Published: (2025)
by: Siddiqui, Nyle, et al.
Published: (2025)
Möbius Transform for Mitigating Perspective Distortions in Representation Learning
by: Chhipa, Prakash Chandra, et al.
Published: (2024)
by: Chhipa, Prakash Chandra, et al.
Published: (2024)
GAReT: Cross-view Video Geolocalization with Adapters and Auto-Regressive Transformers
by: Pillai, Manu S, et al.
Published: (2024)
by: Pillai, Manu S, et al.
Published: (2024)
MedRoute: RL-Based Dynamic Specialist Routing in Multi-Agent Medical Diagnosis
by: Vayani, Ashmal, et al.
Published: (2026)
by: Vayani, Ashmal, et al.
Published: (2026)
VidLA: Video-Language Alignment at Scale
by: Rizve, Mamshad Nayeem, et al.
Published: (2024)
by: Rizve, Mamshad Nayeem, et al.
Published: (2024)
Cross-View Open-Vocabulary Object Detection in Aerial Imagery
by: Kini, Jyoti, et al.
Published: (2025)
by: Kini, Jyoti, et al.
Published: (2025)
From Play to Replay: Composed Video Retrieval for Temporally Fine-Grained Videos
by: Gupta, Animesh, et al.
Published: (2025)
by: Gupta, Animesh, et al.
Published: (2025)
TimeLogic: A Temporal Logic Benchmark for Video QA
by: Swetha, Sirnam, et al.
Published: (2025)
by: Swetha, Sirnam, et al.
Published: (2025)
ViLL-E: Video LLM Embeddings for Retrieval
by: Gupta, Rohit, et al.
Published: (2026)
by: Gupta, Rohit, et al.
Published: (2026)
Vid-Freeze: Protecting Images from Malicious Image-to-Video Generation via Temporal Freezing
by: Chowdhury, Rohit, et al.
Published: (2025)
by: Chowdhury, Rohit, et al.
Published: (2025)
Open-Vocabulary Object Detectors: Robustness Challenges under Distribution Shifts
by: Chhipa, Prakash Chandra, et al.
Published: (2024)
by: Chhipa, Prakash Chandra, et al.
Published: (2024)
LCM: Log Conformal Maps for Robust Representation Learning to Mitigate Perspective Distortion
by: Chippa, Meenakshi Subhash, et al.
Published: (2024)
by: Chippa, Meenakshi Subhash, et al.
Published: (2024)
BBQ-V: Benchmarking Visual Stereotype Bias in Large Multimodal Models
by: Narnaware, Vishal, et al.
Published: (2025)
by: Narnaware, Vishal, et al.
Published: (2025)
AlignVid: Training-Free Attention Scaling for Semantic Fidelity in Text-Guided Image-to-Video Generation
by: Liu, Yexin, et al.
Published: (2025)
by: Liu, Yexin, et al.
Published: (2025)
Temporally Consistent Referring Video Object Segmentation with Hybrid Memory
by: Miao, Bo, et al.
Published: (2024)
by: Miao, Bo, et al.
Published: (2024)
Open Vocabulary Multi-Label Video Classification
by: Gupta, Rohit, et al.
Published: (2024)
by: Gupta, Rohit, et al.
Published: (2024)
SafeVid: Toward Safety Aligned Video Large Multimodal Models
by: Wang, Yixu, et al.
Published: (2025)
by: Wang, Yixu, et al.
Published: (2025)
TAG: A Simple Yet Effective Temporal-Aware Approach for Zero-Shot Video Temporal Grounding
by: Lee, Jin-Seop, et al.
Published: (2025)
by: Lee, Jin-Seop, et al.
Published: (2025)
RelightVid: Temporal-Consistent Diffusion Model for Video Relighting
by: Fang, Ye, et al.
Published: (2025)
by: Fang, Ye, et al.
Published: (2025)
EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models
by: Huang, Shiqi, et al.
Published: (2026)
by: Huang, Shiqi, et al.
Published: (2026)
Overview of TREC 2024 Medical Video Question Answering (MedVidQA) Track
by: Gupta, Deepak, et al.
Published: (2024)
by: Gupta, Deepak, et al.
Published: (2024)
VideoSAVi: Self-Aligned Video Language Models without Human Supervision
by: Kulkarni, Yogesh, et al.
Published: (2024)
by: Kulkarni, Yogesh, et al.
Published: (2024)
Class Prototypes based Contrastive Learning for Classifying Multi-Label and Fine-Grained Educational Videos
by: Gupta, Rohit, et al.
Published: (2025)
by: Gupta, Rohit, et al.
Published: (2025)
The Telephone Game: Evaluating Semantic Drift in Unified Models
by: Mollah, Sabbir, et al.
Published: (2025)
by: Mollah, Sabbir, et al.
Published: (2025)
VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding
by: Li, Chaoyu, et al.
Published: (2024)
by: Li, Chaoyu, et al.
Published: (2024)
PackCache: A Training-Free Acceleration Method for Unified Autoregressive Video Generation via Compact KV-Cache
by: Li, Kunyang, et al.
Published: (2026)
by: Li, Kunyang, et al.
Published: (2026)
Privacy Beyond Pixels: Latent Anonymization for Privacy-Preserving Video Understanding
by: Fioresi, Joseph, et al.
Published: (2025)
by: Fioresi, Joseph, et al.
Published: (2025)
VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models
by: Lan, Xiaohan, et al.
Published: (2024)
by: Lan, Xiaohan, et al.
Published: (2024)
Xi-Net: Transformer Based Seismic Waveform Reconstructor
by: Gaharwar, Anshuman, et al.
Published: (2024)
by: Gaharwar, Anshuman, et al.
Published: (2024)
EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation
by: Wang, Xiaofeng, et al.
Published: (2024)
by: Wang, Xiaofeng, et al.
Published: (2024)
Attend Locally, Remember Linearly: Linear Attention as Cross-Frame Memory for Autoregressive Video Diffusion
by: Li, Kunyang, et al.
Published: (2026)
by: Li, Kunyang, et al.
Published: (2026)
GTAD: Global Temporal Aggregation Denoising Learning for 3D Semantic Occupancy Prediction
by: Li, Tianhao, et al.
Published: (2025)
by: Li, Tianhao, et al.
Published: (2025)
VidHal: Benchmarking Temporal Hallucinations in Vision LLMs
by: Choong, Wey Yeh, et al.
Published: (2024)
by: Choong, Wey Yeh, et al.
Published: (2024)
A Systematic Performance Analysis of Deep Perceptual Loss Networks: Breaking Transfer Learning Conventions
by: Pihlgren, Gustav Grund, et al.
Published: (2023)
by: Pihlgren, Gustav Grund, et al.
Published: (2023)
HarmoVid: Relightful Video Portrait Harmonization
by: Choi, Jun Myeong, et al.
Published: (2026)
by: Choi, Jun Myeong, et al.
Published: (2026)
Sync from the Sea: Retrieving Alignable Videos from Large-Scale Datasets
by: Dave, Ishan Rajendrakumar, et al.
Published: (2024)
by: Dave, Ishan Rajendrakumar, et al.
Published: (2024)
OTT-Vid: Optimal Transport Temporal Token Compression for Video Large Language Models
by: Kang, Minseok, et al.
Published: (2026)
by: Kang, Minseok, et al.
Published: (2026)
Similar Items
-
CityGuessr: City-Level Video Geo-Localization on a Global Scale
by: Kulkarni, Parth Parag, et al.
Published: (2024) -
GAEA: A Geolocation Aware Conversational Assistant
by: Campos, Ron, et al.
Published: (2025) -
VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues
by: Swetha, Sirnam, et al.
Published: (2025) -
StretchySnake: Flexible SSM Training Unlocks Action Recognition Across Spatio-Temporal Scales
by: Siddiqui, Nyle, et al.
Published: (2025) -
Möbius Transform for Mitigating Perspective Distortions in Representation Learning
by: Chhipa, Prakash Chandra, et al.
Published: (2024)