Video2BEV: Transforming Drone Videos to BEVs for Video-based Geo-localization
Fuente:
arXiv
Guardado en:
| Autores principales: | Ju, Hao, Huang, Shaofei, Liu, Si, Zheng, Zhedong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
From Instruction to Event: Sound-Triggered Mobile Manipulation
por: Ju, Hao, et al.
Publicado: (2026)
por: Ju, Hao, et al.
Publicado: (2026)
Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark
por: Yang, Shuyu, et al.
Publicado: (2025)
por: Yang, Shuyu, et al.
Publicado: (2025)
Pretrain-then-Adapt: Uncertainty-Aware Test-Time Adaptation for Text-based Person Search
por: Zhang, Jiahao, et al.
Publicado: (2026)
por: Zhang, Jiahao, et al.
Publicado: (2026)
Road Maps as Free Geometric Priors: Weather-Invariant Drone Geo-Localization with GeoFuse
por: Fang, Yunsong, et al.
Publicado: (2026)
por: Fang, Yunsong, et al.
Publicado: (2026)
WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-Localization
por: Wen, Jiahao, et al.
Publicado: (2025)
por: Wen, Jiahao, et al.
Publicado: (2025)
Can Video Diffusion Models Predict Past Frames? Bidirectional Cycle Consistency for Reversible Interpolation
por: Liu, Lingyu, et al.
Publicado: (2026)
por: Liu, Lingyu, et al.
Publicado: (2026)
AnomalyLMM: Bridging Generative Knowledge and Discriminative Retrieval for Text-Based Person Anomaly Search
por: Ju, Hao, et al.
Publicado: (2025)
por: Ju, Hao, et al.
Publicado: (2025)
RepVideo: Rethinking Cross-Layer Representation for Video Generation
por: Si, Chenyang, et al.
Publicado: (2025)
por: Si, Chenyang, et al.
Publicado: (2025)
Unleashing the Temporal-Spatial Reasoning Capacity of GPT for Training-Free Audio and Language Referenced Video Object Segmentation
por: Huang, Shaofei, et al.
Publicado: (2024)
por: Huang, Shaofei, et al.
Publicado: (2024)
Video Individual Counting for Moving Drones
por: Fan, Yaowu, et al.
Publicado: (2025)
por: Fan, Yaowu, et al.
Publicado: (2025)
Steering Video Diffusion Transformers with Massive Activations
por: Cheng, Xianhang, et al.
Publicado: (2026)
por: Cheng, Xianhang, et al.
Publicado: (2026)
GeoVideo: Introducing Geometric Regularization into Video Generation Model
por: Bai, Yunpeng, et al.
Publicado: (2025)
por: Bai, Yunpeng, et al.
Publicado: (2025)
Learning Camera Movement Control from Real-World Drone Videos
por: Hou, Yunzhong, et al.
Publicado: (2024)
por: Hou, Yunzhong, et al.
Publicado: (2024)
Towards Natural Language-Guided Drones: GeoText-1652 Benchmark with Spatial Relation Matching
por: Chu, Meng, et al.
Publicado: (2023)
por: Chu, Meng, et al.
Publicado: (2023)
CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
por: Yang, Zhuoyi, et al.
Publicado: (2024)
por: Yang, Zhuoyi, et al.
Publicado: (2024)
EraserDiT: Fast Video Inpainting with Diffusion Transformer Model
por: Liu, Jie, et al.
Publicado: (2025)
por: Liu, Jie, et al.
Publicado: (2025)
Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning
por: Jeon, MinJu, et al.
Publicado: (2025)
por: Jeon, MinJu, et al.
Publicado: (2025)
Scale-adaptive UAV Geo-localization via Height-aware Partition Learning
por: Chen, Quan, et al.
Publicado: (2024)
por: Chen, Quan, et al.
Publicado: (2024)
Mask$^2$DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation
por: Qi, Tianhao, et al.
Publicado: (2025)
por: Qi, Tianhao, et al.
Publicado: (2025)
VideoHandles: Editing 3D Object Compositions in Videos Using Video Generative Priors
por: Koo, Juil, et al.
Publicado: (2025)
por: Koo, Juil, et al.
Publicado: (2025)
WidthFormer: Toward Efficient Transformer-based BEV View Transformation
por: Yang, Chenhongyi, et al.
Publicado: (2024)
por: Yang, Chenhongyi, et al.
Publicado: (2024)
S2DiT: Sandwich Diffusion Transformer for Mobile Streaming Video Generation
por: Zhao, Lin, et al.
Publicado: (2026)
por: Zhao, Lin, et al.
Publicado: (2026)
MagicMirror: ID-Preserved Video Generation in Video Diffusion Transformers
por: Zhang, Yuechen, et al.
Publicado: (2025)
por: Zhang, Yuechen, et al.
Publicado: (2025)
AVD2: Accident Video Diffusion for Accident Video Description
por: Li, Cheng, et al.
Publicado: (2025)
por: Li, Cheng, et al.
Publicado: (2025)
V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning
por: Cheng, Zixu, et al.
Publicado: (2025)
por: Cheng, Zixu, et al.
Publicado: (2025)
Towards Online Real-Time Memory-based Video Inpainting Transformers
por: Thiry, Guillaume, et al.
Publicado: (2024)
por: Thiry, Guillaume, et al.
Publicado: (2024)
VideoGen-Eval: Agent-based System for Video Generation Evaluation
por: Yang, Yuhang, et al.
Publicado: (2025)
por: Yang, Yuhang, et al.
Publicado: (2025)
Spatial-Conditioned Reasoning in Long-Egocentric Videos
por: Tribble, James, et al.
Publicado: (2026)
por: Tribble, James, et al.
Publicado: (2026)
Lighting-grounded Video Generation with Renderer-based Agent Reasoning
por: Cai, Ziqi, et al.
Publicado: (2026)
por: Cai, Ziqi, et al.
Publicado: (2026)
FastVideoEdit: Leveraging Consistency Models for Efficient Text-to-Video Editing
por: Zhang, Youyuan, et al.
Publicado: (2024)
por: Zhang, Youyuan, et al.
Publicado: (2024)
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval
por: Huang, Jinghao, et al.
Publicado: (2025)
por: Huang, Jinghao, et al.
Publicado: (2025)
2nd Place Solution for PVUW Challenge 2024: Video Panoptic Segmentation
por: Wu, Biao, et al.
Publicado: (2024)
por: Wu, Biao, et al.
Publicado: (2024)
TDViT: Temporal Dilated Video Transformer for Dense Video Tasks
por: Sun, Guanxiong, et al.
Publicado: (2024)
por: Sun, Guanxiong, et al.
Publicado: (2024)
Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
por: Guan, Yiran, et al.
Publicado: (2026)
por: Guan, Yiran, et al.
Publicado: (2026)
Señorita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists
por: Zi, Bojia, et al.
Publicado: (2025)
por: Zi, Bojia, et al.
Publicado: (2025)
I2VEdit: First-Frame-Guided Video Editing via Image-to-Video Diffusion Models
por: Ouyang, Wenqi, et al.
Publicado: (2024)
por: Ouyang, Wenqi, et al.
Publicado: (2024)
Beyond Few-Step Inference: Accelerating Video Diffusion Transformer Model Serving with Inter-Request Caching Reuse
por: Liu, Hao, et al.
Publicado: (2026)
por: Liu, Hao, et al.
Publicado: (2026)
Omni-Video: Democratizing Unified Video Understanding and Generation
por: Tan, Zhiyu, et al.
Publicado: (2025)
por: Tan, Zhiyu, et al.
Publicado: (2025)
Geometric Transformation-Embedded Mamba for Learned Video Compression
por: Wei, Hao, et al.
Publicado: (2026)
por: Wei, Hao, et al.
Publicado: (2026)
FreeInit: Bridging Initialization Gap in Video Diffusion Models
por: Wu, Tianxing, et al.
Publicado: (2023)
por: Wu, Tianxing, et al.
Publicado: (2023)
Ejemplares similares
-
From Instruction to Event: Sound-Triggered Mobile Manipulation
por: Ju, Hao, et al.
Publicado: (2026) -
Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark
por: Yang, Shuyu, et al.
Publicado: (2025) -
Pretrain-then-Adapt: Uncertainty-Aware Test-Time Adaptation for Text-based Person Search
por: Zhang, Jiahao, et al.
Publicado: (2026) -
Road Maps as Free Geometric Priors: Weather-Invariant Drone Geo-Localization with GeoFuse
por: Fang, Yunsong, et al.
Publicado: (2026) -
WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-Localization
por: Wen, Jiahao, et al.
Publicado: (2025)