GAReT: Cross-view Video Geolocalization with Adapters and Auto-Regressive Transformers
Fuente:
arXiv
Salvato in:
| Autori principali: | Pillai, Manu S, Rizve, Mamshad Nayeem, Shah, Mubarak |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FinePseudo: Improving Pseudo-Labelling through Temporal-Alignablity for Semi-Supervised Fine-Grained Action Recognition
di: Dave, Ishan Rajendrakumar, et al.
Pubblicazione: (2024)
di: Dave, Ishan Rajendrakumar, et al.
Pubblicazione: (2024)
Open Vocabulary Multi-Label Video Classification
di: Gupta, Rohit, et al.
Pubblicazione: (2024)
di: Gupta, Rohit, et al.
Pubblicazione: (2024)
VidLA: Video-Language Alignment at Scale
di: Rizve, Mamshad Nayeem, et al.
Pubblicazione: (2024)
di: Rizve, Mamshad Nayeem, et al.
Pubblicazione: (2024)
Is ImageNet worth 1 video? Learning strong image encoders from 1 long unlabelled video
di: Venkataramanan, Shashanka, et al.
Pubblicazione: (2023)
di: Venkataramanan, Shashanka, et al.
Pubblicazione: (2023)
X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs
di: Swetha, Sirnam, et al.
Pubblicazione: (2024)
di: Swetha, Sirnam, et al.
Pubblicazione: (2024)
CompSlider: Compositional Slider for Disentangled Multiple-Attribute Image Generation
di: Zhu, Zixin, et al.
Pubblicazione: (2025)
di: Zhu, Zixin, et al.
Pubblicazione: (2025)
Rethinking Prompt Design for Inference-time Scaling in Text-to-Visual Generation
di: Kim, Subin, et al.
Pubblicazione: (2025)
di: Kim, Subin, et al.
Pubblicazione: (2025)
Unified Alignment Protocol: Making Sense of the Unlabeled Data in New Domains
di: Ahmed, Sabbir, et al.
Pubblicazione: (2025)
di: Ahmed, Sabbir, et al.
Pubblicazione: (2025)
VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale
di: Kulkarni, Parth Parag, et al.
Pubblicazione: (2026)
di: Kulkarni, Parth Parag, et al.
Pubblicazione: (2026)
Attend Locally, Remember Linearly: Linear Attention as Cross-Frame Memory for Autoregressive Video Diffusion
di: Li, Kunyang, et al.
Pubblicazione: (2026)
di: Li, Kunyang, et al.
Pubblicazione: (2026)
Cross-View Open-Vocabulary Object Detection in Aerial Imagery
di: Kini, Jyoti, et al.
Pubblicazione: (2025)
di: Kini, Jyoti, et al.
Pubblicazione: (2025)
TimeLogic: A Temporal Logic Benchmark for Video QA
di: Swetha, Sirnam, et al.
Pubblicazione: (2025)
di: Swetha, Sirnam, et al.
Pubblicazione: (2025)
PackCache: A Training-Free Acceleration Method for Unified Autoregressive Video Generation via Compact KV-Cache
di: Li, Kunyang, et al.
Pubblicazione: (2026)
di: Li, Kunyang, et al.
Pubblicazione: (2026)
Privacy Beyond Pixels: Latent Anonymization for Privacy-Preserving Video Understanding
di: Fioresi, Joseph, et al.
Pubblicazione: (2025)
di: Fioresi, Joseph, et al.
Pubblicazione: (2025)
From Play to Replay: Composed Video Retrieval for Temporally Fine-Grained Videos
di: Gupta, Animesh, et al.
Pubblicazione: (2025)
di: Gupta, Animesh, et al.
Pubblicazione: (2025)
Self-Distilled Masked Auto-Encoders are Efficient Video Anomaly Detectors
di: Ristea, Nicolae-Catalin, et al.
Pubblicazione: (2023)
di: Ristea, Nicolae-Catalin, et al.
Pubblicazione: (2023)
CityGuessr: City-Level Video Geo-Localization on a Global Scale
di: Kulkarni, Parth Parag, et al.
Pubblicazione: (2024)
di: Kulkarni, Parth Parag, et al.
Pubblicazione: (2024)
ARCON: Advancing Auto-Regressive Continuation for Driving Videos
di: Ming, Ruibo, et al.
Pubblicazione: (2024)
di: Ming, Ruibo, et al.
Pubblicazione: (2024)
FARMER: Flow AutoRegressive Transformer over Pixels
di: Zheng, Guangting, et al.
Pubblicazione: (2025)
di: Zheng, Guangting, et al.
Pubblicazione: (2025)
Investigating Memorization in Video Diffusion Models
di: Chen, Chen, et al.
Pubblicazione: (2024)
di: Chen, Chen, et al.
Pubblicazione: (2024)
Towards Generative Location Awareness for Disaster Response: A Probabilistic Cross-view Geolocalization Approach
di: Li, Hao, et al.
Pubblicazione: (2025)
di: Li, Hao, et al.
Pubblicazione: (2025)
Leveraging Pre-Trained Visual Models for AI-Generated Video Detection
di: Veeramachaneni, Keerthi, et al.
Pubblicazione: (2025)
di: Veeramachaneni, Keerthi, et al.
Pubblicazione: (2025)
AR-Diffusion: Asynchronous Video Generation with Auto-Regressive Diffusion
di: Sun, Mingzhen, et al.
Pubblicazione: (2025)
di: Sun, Mingzhen, et al.
Pubblicazione: (2025)
Temporally Consistent Referring Video Object Segmentation with Hybrid Memory
di: Miao, Bo, et al.
Pubblicazione: (2024)
di: Miao, Bo, et al.
Pubblicazione: (2024)
PTQ4DiT: Post-training Quantization for Diffusion Transformers
di: Wu, Junyi, et al.
Pubblicazione: (2024)
di: Wu, Junyi, et al.
Pubblicazione: (2024)
Auto-selected Knowledge Adapters for Lifelong Person Re-identification
di: Qian, Xuelin, et al.
Pubblicazione: (2024)
di: Qian, Xuelin, et al.
Pubblicazione: (2024)
Transformer Based Self-Context Aware Prediction for Few-Shot Anomaly Detection in Videos
di: Pillai, Gargi V., et al.
Pubblicazione: (2025)
di: Pillai, Gargi V., et al.
Pubblicazione: (2025)
ViLL-E: Video LLM Embeddings for Retrieval
di: Gupta, Rohit, et al.
Pubblicazione: (2026)
di: Gupta, Rohit, et al.
Pubblicazione: (2026)
Learnability-Guided Diffusion for Dataset Distillation
di: Chan-Santiago, Jeffrey A., et al.
Pubblicazione: (2026)
di: Chan-Santiago, Jeffrey A., et al.
Pubblicazione: (2026)
CART: Compositional Auto-Regressive Transformer for Image Generation
di: Roheda, Siddharth, et al.
Pubblicazione: (2024)
di: Roheda, Siddharth, et al.
Pubblicazione: (2024)
FAR-Drive: Frame-AutoRegressive Video Generation in Closed-Loop Autonomous Driving
di: Li, Yaoru, et al.
Pubblicazione: (2026)
di: Li, Yaoru, et al.
Pubblicazione: (2026)
MV-Adapter: Multi-view Consistent Image Generation Made Easy
di: Huang, Zehuan, et al.
Pubblicazione: (2024)
di: Huang, Zehuan, et al.
Pubblicazione: (2024)
From Frames to Clips: Training-free Adaptive Key Clip Selection for Long-Form Video Understanding
di: Sun, Guangyu, et al.
Pubblicazione: (2025)
di: Sun, Guangyu, et al.
Pubblicazione: (2025)
Cross-view Masked Diffusion Transformers for Person Image Synthesis
di: Pham, Trung X., et al.
Pubblicazione: (2024)
di: Pham, Trung X., et al.
Pubblicazione: (2024)
GAEA: A Geolocation Aware Conversational Assistant
di: Campos, Ron, et al.
Pubblicazione: (2025)
di: Campos, Ron, et al.
Pubblicazione: (2025)
Möbius Transform for Mitigating Perspective Distortions in Representation Learning
di: Chhipa, Prakash Chandra, et al.
Pubblicazione: (2024)
di: Chhipa, Prakash Chandra, et al.
Pubblicazione: (2024)
I2V-Adapter: A General Image-to-Video Adapter for Diffusion Models
di: Guo, Xun, et al.
Pubblicazione: (2023)
di: Guo, Xun, et al.
Pubblicazione: (2023)
Q-Adapter: Visual Query Adapter for Extracting Textually-related Features in Video Captioning
di: Chen, Junan, et al.
Pubblicazione: (2025)
di: Chen, Junan, et al.
Pubblicazione: (2025)
STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits
di: Papantoniou, Foivos Paraperas, et al.
Pubblicazione: (2025)
di: Papantoniou, Foivos Paraperas, et al.
Pubblicazione: (2025)
ProGVC: Progressive-based Generative Video Compression via Auto-Regressive Context Modeling
di: Li, Daowen, et al.
Pubblicazione: (2026)
di: Li, Daowen, et al.
Pubblicazione: (2026)
Documenti analoghi
-
FinePseudo: Improving Pseudo-Labelling through Temporal-Alignablity for Semi-Supervised Fine-Grained Action Recognition
di: Dave, Ishan Rajendrakumar, et al.
Pubblicazione: (2024) -
Open Vocabulary Multi-Label Video Classification
di: Gupta, Rohit, et al.
Pubblicazione: (2024) -
VidLA: Video-Language Alignment at Scale
di: Rizve, Mamshad Nayeem, et al.
Pubblicazione: (2024) -
Is ImageNet worth 1 video? Learning strong image encoders from 1 long unlabelled video
di: Venkataramanan, Shashanka, et al.
Pubblicazione: (2023) -
X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs
di: Swetha, Sirnam, et al.
Pubblicazione: (2024)