Saved in:
| Main Authors: | Kulkarni, Parth Parag, Nayak, Gaurav Kumar, Shah, Mubarak |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2411.06344 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale
by: Kulkarni, Parth Parag, et al.
Published: (2026)
by: Kulkarni, Parth Parag, et al.
Published: (2026)
MedRoute: RL-Based Dynamic Specialist Routing in Multi-Agent Medical Diagnosis
by: Vayani, Ashmal, et al.
Published: (2026)
by: Vayani, Ashmal, et al.
Published: (2026)
VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues
by: Swetha, Sirnam, et al.
Published: (2025)
by: Swetha, Sirnam, et al.
Published: (2025)
Object-Level Explanations for Image Geolocation Models: a GeoGuessr use-case
by: Durrieu, Emilie, et al.
Published: (2026)
by: Durrieu, Emilie, et al.
Published: (2026)
CV-Cities: Advancing Cross-View Geo-Localization in Global Cities
by: Huang, Gaoshuang, et al.
Published: (2024)
by: Huang, Gaoshuang, et al.
Published: (2024)
DLCR: A Generative Data Expansion Framework via Diffusion for Clothes-Changing Person Re-ID
by: Siddiqui, Nyle, et al.
Published: (2024)
by: Siddiqui, Nyle, et al.
Published: (2024)
MGD$^3$: Mode-Guided Dataset Distillation using Diffusion Models
by: Chan-Santiago, Jeffrey A., et al.
Published: (2025)
by: Chan-Santiago, Jeffrey A., et al.
Published: (2025)
GAEA: A Geolocation Aware Conversational Assistant
by: Campos, Ron, et al.
Published: (2025)
by: Campos, Ron, et al.
Published: (2025)
Attend Locally, Remember Linearly: Linear Attention as Cross-Frame Memory for Autoregressive Video Diffusion
by: Li, Kunyang, et al.
Published: (2026)
by: Li, Kunyang, et al.
Published: (2026)
TimeLogic: A Temporal Logic Benchmark for Video QA
by: Swetha, Sirnam, et al.
Published: (2025)
by: Swetha, Sirnam, et al.
Published: (2025)
CityRAG: Stepping Into a City via Spatially-Grounded Video Generation
by: Chou, Gene, et al.
Published: (2026)
by: Chou, Gene, et al.
Published: (2026)
PackCache: A Training-Free Acceleration Method for Unified Autoregressive Video Generation via Compact KV-Cache
by: Li, Kunyang, et al.
Published: (2026)
by: Li, Kunyang, et al.
Published: (2026)
GAReT: Cross-view Video Geolocalization with Adapters and Auto-Regressive Transformers
by: Pillai, Manu S, et al.
Published: (2024)
by: Pillai, Manu S, et al.
Published: (2024)
Privacy Beyond Pixels: Latent Anonymization for Privacy-Preserving Video Understanding
by: Fioresi, Joseph, et al.
Published: (2025)
by: Fioresi, Joseph, et al.
Published: (2025)
From Play to Replay: Composed Video Retrieval for Temporally Fine-Grained Videos
by: Gupta, Animesh, et al.
Published: (2025)
by: Gupta, Animesh, et al.
Published: (2025)
Exploring Local Memorization in Diffusion Models via Bright Ending Attention
by: Chen, Chen, et al.
Published: (2024)
by: Chen, Chen, et al.
Published: (2024)
StretchySnake: Flexible SSM Training Unlocks Action Recognition Across Spatio-Temporal Scales
by: Siddiqui, Nyle, et al.
Published: (2025)
by: Siddiqui, Nyle, et al.
Published: (2025)
Investigating Memorization in Video Diffusion Models
by: Chen, Chen, et al.
Published: (2024)
by: Chen, Chen, et al.
Published: (2024)
Imagine a City: CityGenAgent for Procedural 3D City Generation
by: Liu, Zishan, et al.
Published: (2026)
by: Liu, Zishan, et al.
Published: (2026)
Leveraging Pre-Trained Visual Models for AI-Generated Video Detection
by: Veeramachaneni, Keerthi, et al.
Published: (2025)
by: Veeramachaneni, Keerthi, et al.
Published: (2025)
LDMapNet-U: An End-to-End System for City-Scale Lane-Level Map Updating
by: Xia, Deguo, et al.
Published: (2025)
by: Xia, Deguo, et al.
Published: (2025)
CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos
by: Liu, Xinhao, et al.
Published: (2024)
by: Liu, Xinhao, et al.
Published: (2024)
Temporally Consistent Referring Video Object Segmentation with Hybrid Memory
by: Miao, Bo, et al.
Published: (2024)
by: Miao, Bo, et al.
Published: (2024)
Sync from the Sea: Retrieving Alignable Videos from Large-Scale Datasets
by: Dave, Ishan Rajendrakumar, et al.
Published: (2024)
by: Dave, Ishan Rajendrakumar, et al.
Published: (2024)
DuMapNet: An End-to-End Vectorization System for City-Scale Lane-Level Map Generation
by: Xia, Deguo, et al.
Published: (2024)
by: Xia, Deguo, et al.
Published: (2024)
GRUtopia: Dream General Robots in a City at Scale
by: Wang, Hanqing, et al.
Published: (2024)
by: Wang, Hanqing, et al.
Published: (2024)
CityGen: Infinite and Controllable City Layout Generation
by: Deng, Jie, et al.
Published: (2023)
by: Deng, Jie, et al.
Published: (2023)
UrbanVerse: Scaling Urban Simulation by Watching City-Tour Videos
by: Liu, Mingxuan, et al.
Published: (2025)
by: Liu, Mingxuan, et al.
Published: (2025)
ViLL-E: Video LLM Embeddings for Retrieval
by: Gupta, Rohit, et al.
Published: (2026)
by: Gupta, Rohit, et al.
Published: (2026)
Cross-View Open-Vocabulary Object Detection in Aerial Imagery
by: Kini, Jyoti, et al.
Published: (2025)
by: Kini, Jyoti, et al.
Published: (2025)
Learnability-Guided Diffusion for Dataset Distillation
by: Chan-Santiago, Jeffrey A., et al.
Published: (2026)
by: Chan-Santiago, Jeffrey A., et al.
Published: (2026)
TIGeR: A Unified Framework for Time, Images and Geo-location Retrieval
by: Shatwell, David G., et al.
Published: (2026)
by: Shatwell, David G., et al.
Published: (2026)
GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields
by: Yasuki, Shunsuke, et al.
Published: (2025)
by: Yasuki, Shunsuke, et al.
Published: (2025)
Reasoning-Guided Grounding: Elevating Video Anomaly Detection through Multimodal Large Language Models
by: Agarwal, Sakshi, et al.
Published: (2026)
by: Agarwal, Sakshi, et al.
Published: (2026)
VidLA: Video-Language Alignment at Scale
by: Rizve, Mamshad Nayeem, et al.
Published: (2024)
by: Rizve, Mamshad Nayeem, et al.
Published: (2024)
Sparse Points to Dense Clouds: Enhancing 3D Detection with Limited LiDAR Data
by: Kumar, Aakash, et al.
Published: (2024)
by: Kumar, Aakash, et al.
Published: (2024)
Diffexplainer: Towards Cross-modal Global Explanations with Diffusion Models
by: Pennisi, Matteo, et al.
Published: (2024)
by: Pennisi, Matteo, et al.
Published: (2024)
Benchmarking Egocentric Visual-Inertial SLAM at City Scale
by: Krishnan, Anusha, et al.
Published: (2025)
by: Krishnan, Anusha, et al.
Published: (2025)
CityCraft: A Real Crafter for 3D City Generation
by: Deng, Jie, et al.
Published: (2024)
by: Deng, Jie, et al.
Published: (2024)
CityLLaVA: Efficient Fine-Tuning for VLMs in City Scenario
by: Duan, Zhizhao, et al.
Published: (2024)
by: Duan, Zhizhao, et al.
Published: (2024)
Similar Items
-
VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale
by: Kulkarni, Parth Parag, et al.
Published: (2026) -
MedRoute: RL-Based Dynamic Specialist Routing in Multi-Agent Medical Diagnosis
by: Vayani, Ashmal, et al.
Published: (2026) -
VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues
by: Swetha, Sirnam, et al.
Published: (2025) -
Object-Level Explanations for Image Geolocation Models: a GeoGuessr use-case
by: Durrieu, Emilie, et al.
Published: (2026) -
CV-Cities: Advancing Cross-View Geo-Localization in Global Cities
by: Huang, Gaoshuang, et al.
Published: (2024)