Sky2Ground: A Benchmark for Site Modeling under Varying Altitude
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wang, Zengyan, Mitra, Sirshapan, Modi, Rajat, Lim, Grace, Rawat, Yogesh |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ProDiG: Progressive Diffusion-Guided Gaussian Splatting for Aerial to Ground Reconstruction
par: Mitra, Sirshapan, et autres
Publié: (2026)
par: Mitra, Sirshapan, et autres
Publié: (2026)
GaitCrafter: Diffusion Model for Biometric Preserving Gait Synthesis
par: Mitra, Sirshapan, et autres
Publié: (2025)
par: Mitra, Sirshapan, et autres
Publié: (2025)
Stable Mean Teacher for Semi-supervised Video Action Detection
par: Kumar, Akash, et autres
Publié: (2024)
par: Kumar, Akash, et autres
Publié: (2024)
Asynchronous Perception Machine For Efficient Test-Time-Training
par: Modi, Rajat, et autres
Publié: (2024)
par: Modi, Rajat, et autres
Publié: (2024)
On Occlusions in Video Action Detection: Benchmark Datasets And Training Recipes
par: Modi, Rajat, et autres
Publié: (2024)
par: Modi, Rajat, et autres
Publié: (2024)
RobustGait: Robustness Analysis for Appearance Based Gait Recognition
par: Sayera, Reeshoon, et autres
Publié: (2025)
par: Sayera, Reeshoon, et autres
Publié: (2025)
Foundation Models for Video Understanding: A Survey
par: Madan, Neelu, et autres
Publié: (2024)
par: Madan, Neelu, et autres
Publié: (2024)
STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal Grounding
par: Garg, Aaryan, et autres
Publié: (2025)
par: Garg, Aaryan, et autres
Publié: (2025)
Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video Grounding
par: Kumar, Akash, et autres
Publié: (2025)
par: Kumar, Akash, et autres
Publié: (2025)
AirSketch: Generative Motion to Sketch
par: Lim, Hui Xian Grace, et autres
Publié: (2024)
par: Lim, Hui Xian Grace, et autres
Publié: (2024)
LR0.FM: Low-Res Benchmark and Improving Robustness for Zero-Shot Classification in Foundation Models
par: Pathak, Priyank, et autres
Publié: (2025)
par: Pathak, Priyank, et autres
Publié: (2025)
Colors See Colors Ignore: Clothes Changing ReID with Color Disentanglement
par: Pathak, Priyank, et autres
Publié: (2025)
par: Pathak, Priyank, et autres
Publié: (2025)
DIFFER: Disentangling Identity Features via Semantic Cues for Clothes-Changing Person Re-ID
par: Liang, Xin, et autres
Publié: (2025)
par: Liang, Xin, et autres
Publié: (2025)
DisenQ: Disentangling Q-Former for Activity-Biometrics
par: Azad, Shehreen, et autres
Publié: (2025)
par: Azad, Shehreen, et autres
Publié: (2025)
Coarse Attribute Prediction with Task Agnostic Distillation for Real World Clothes Changing ReID
par: Pathak, Priyank, et autres
Publié: (2025)
par: Pathak, Priyank, et autres
Publié: (2025)
Activity-Biometrics: Person Identification from Daily Activities
par: Azad, Shehreen, et autres
Publié: (2024)
par: Azad, Shehreen, et autres
Publié: (2024)
MolVision: Molecular Property Prediction with Vision Language Models
par: Adak, Deepan, et autres
Publié: (2025)
par: Adak, Deepan, et autres
Publié: (2025)
Navigating Hallucinations for Reasoning of Unintentional Activities
par: Grover, Shresth, et autres
Publié: (2024)
par: Grover, Shresth, et autres
Publié: (2024)
Scaling Open-Vocabulary Action Detection
par: Sia, Zhen Hao, et autres
Publié: (2025)
par: Sia, Zhen Hao, et autres
Publié: (2025)
iSafetyBench: A video-language benchmark for safety in industrial environment
par: Abdullah, Raiyaan, et autres
Publié: (2025)
par: Abdullah, Raiyaan, et autres
Publié: (2025)
StreamReady: Learning What to Answer and When in Long Streaming Videos
par: Azad, Shehreen, et autres
Publié: (2026)
par: Azad, Shehreen, et autres
Publié: (2026)
HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding
par: Azad, Shehreen, et autres
Publié: (2025)
par: Azad, Shehreen, et autres
Publié: (2025)
A Large-Scale Analysis on Contextual Self-Supervised Video Representation Learning
par: Kumar, Akash, et autres
Publié: (2025)
par: Kumar, Akash, et autres
Publié: (2025)
Understanding Depth and Height Perception in Large Visual-Language Models
par: Azad, Shehreen, et autres
Publié: (2024)
par: Azad, Shehreen, et autres
Publié: (2024)
OmViD: Omni-supervised active learning for video action detection
par: Rana, Aayush, et autres
Publié: (2025)
par: Rana, Aayush, et autres
Publié: (2025)
EZ-CLIP: Efficient Zeroshot Video Action Recognition
par: Ahmad, Shahzad, et autres
Publié: (2023)
par: Ahmad, Shahzad, et autres
Publié: (2023)
VISTA: Video Interaction Spatio-Temporal Analysis Benchmark
par: Aparcedo, Alejandro, et autres
Publié: (2026)
par: Aparcedo, Alejandro, et autres
Publié: (2026)
MolSight: Molecular Property Prediction with Images
par: Baranwal, Aaditya, et autres
Publié: (2026)
par: Baranwal, Aaditya, et autres
Publié: (2026)
RIS-LAD: A Benchmark and Model for Referring Low-Altitude Drone Image Segmentation
par: Ye, Kai, et autres
Publié: (2025)
par: Ye, Kai, et autres
Publié: (2025)
Advancing Automatic Photovoltaic Defect Detection using Semi-Supervised Semantic Segmentation of Electroluminescence Images
par: Jha, Abhishek, et autres
Publié: (2024)
par: Jha, Abhishek, et autres
Publié: (2024)
VideoLLM Benchmarks and Evaluation: A Survey
par: Kumar, Yogesh
Publié: (2025)
par: Kumar, Yogesh
Publié: (2025)
CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates
par: Grover, Shresth, et autres
Publié: (2025)
par: Grover, Shresth, et autres
Publié: (2025)
DeTrack: A Benchmark and Altitude-Aware Dual World Model for Drone-embodied Tracking
par: Hu, Guyue, et autres
Publié: (2026)
par: Hu, Guyue, et autres
Publié: (2026)
Probing Conceptual Understanding of Large Visual-Language Models
par: Schiappa, Madeline, et autres
Publié: (2023)
par: Schiappa, Madeline, et autres
Publié: (2023)
SVAG-Bench: A Large-Scale Benchmark for Multi-Instance Spatio-temporal Video Action Grounding
par: Hannan, Tanveer, et autres
Publié: (2025)
par: Hannan, Tanveer, et autres
Publié: (2025)
Towards Perception-based Collision Avoidance for UAVs when Guiding the Visually Impaired
par: Raj, Suman, et autres
Publié: (2025)
par: Raj, Suman, et autres
Publié: (2025)
Robustness Analysis on Foundational Segmentation Models
par: Schiappa, Madeline Chantry, et autres
Publié: (2023)
par: Schiappa, Madeline Chantry, et autres
Publié: (2023)
LAA3D: A Benchmark of Detecting and Tracking Low-Altitude Aircraft in 3D Space
par: Wu, Hai, et autres
Publié: (2025)
par: Wu, Hai, et autres
Publié: (2025)
Constellation Dataset: Benchmarking High-Altitude Object Detection for an Urban Intersection
par: Turkcan, Mehmet Kerem, et autres
Publié: (2024)
par: Turkcan, Mehmet Kerem, et autres
Publié: (2024)
Re:Verse -- Can Your VLM Read a Manga?
par: Baranwal, Aaditya, et autres
Publié: (2025)
par: Baranwal, Aaditya, et autres
Publié: (2025)
Documents similaires
-
ProDiG: Progressive Diffusion-Guided Gaussian Splatting for Aerial to Ground Reconstruction
par: Mitra, Sirshapan, et autres
Publié: (2026) -
GaitCrafter: Diffusion Model for Biometric Preserving Gait Synthesis
par: Mitra, Sirshapan, et autres
Publié: (2025) -
Stable Mean Teacher for Semi-supervised Video Action Detection
par: Kumar, Akash, et autres
Publié: (2024) -
Asynchronous Perception Machine For Efficient Test-Time-Training
par: Modi, Rajat, et autres
Publié: (2024) -
On Occlusions in Video Action Detection: Benchmark Datasets And Training Recipes
par: Modi, Rajat, et autres
Publié: (2024)