AGL-NET: Aerial-Ground Cross-Modal Global Localization with Varying Scales
Fuente:
arXiv
Guardado en:
| Autores principales: | Guan, Tianrui, Xian, Ruiqi, Wang, Xijun, Wu, Xiyang, Elnoor, Mohamed, Song, Daeun, Manocha, Dinesh |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SCP: Soft Conditional Prompt Learning for Aerial Video Action Recognition
por: Wang, Xijun, et al.
Publicado: (2023)
por: Wang, Xijun, et al.
Publicado: (2023)
FALCON: Future-Aware Learning with Contextual Object-Centric Pretraining for UAV Action Recognition
por: Xian, Ruiqi, et al.
Publicado: (2024)
por: Xian, Ruiqi, et al.
Publicado: (2024)
Paired-CSLiDAR: Height-Stratified Registration for Cross-Source Aerial-Ground LiDAR Pose Refinement
por: Hoover, Montana, et al.
Publicado: (2026)
por: Hoover, Montana, et al.
Publicado: (2026)
HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
por: Guan, Tianrui, et al.
Publicado: (2023)
por: Guan, Tianrui, et al.
Publicado: (2023)
Bi-VLM: Pushing Ultra-Low Precision Post-Training Quantization Boundaries in Vision-Language Models
por: Wang, Xijun, et al.
Publicado: (2025)
por: Wang, Xijun, et al.
Publicado: (2025)
AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models
por: Wu, Xiyang, et al.
Publicado: (2024)
por: Wu, Xiyang, et al.
Publicado: (2024)
DAVE: Diverse Atomic Visual Elements Dataset with High Representation of Vulnerable Road Users in Complex and Unpredictable Environments
por: Wang, Xijun, et al.
Publicado: (2024)
por: Wang, Xijun, et al.
Publicado: (2024)
Robot Navigation Using Physically Grounded Vision-Language Models in Outdoor Environments
por: Elnoor, Mohamed, et al.
Publicado: (2024)
por: Elnoor, Mohamed, et al.
Publicado: (2024)
Differentiable Frequency-based Disentanglement for Aerial Video Action Recognition
por: Kothandaraman, Divya, et al.
Publicado: (2022)
por: Kothandaraman, Divya, et al.
Publicado: (2022)
TransLocNet: Cross-Modal Attention for Aerial-Ground Vehicle Localization with Contrastive Learning
por: Pham, Phu, et al.
Publicado: (2025)
por: Pham, Phu, et al.
Publicado: (2025)
HawkI: Homography & Mutual Information Guidance for 3D-free Single Image to Aerial View
por: Kothandaraman, Divya, et al.
Publicado: (2023)
por: Kothandaraman, Divya, et al.
Publicado: (2023)
CalibFree: Self-Supervised View Feature Separation for Calibration-Free Multi-Camera Multi-Object Tracking
por: Xian, Ruiqi, et al.
Publicado: (2026)
por: Xian, Ruiqi, et al.
Publicado: (2026)
MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models
por: Wu, Xiyang, et al.
Publicado: (2025)
por: Wu, Xiyang, et al.
Publicado: (2025)
EgoSocial: Benchmarking Proactive Intervention Ability of Omnimodal LLMs via Egocentric Social Interaction Perception
por: Wang, Xijun, et al.
Publicado: (2025)
por: Wang, Xijun, et al.
Publicado: (2025)
SeCap: Self-Calibrating and Adaptive Prompts for Cross-view Person Re-Identification in Aerial-Ground Networks
por: Wang, Shining, et al.
Publicado: (2025)
por: Wang, Shining, et al.
Publicado: (2025)
Inst4DGS: Instance-Decomposed 4D Gaussian Splatting with Multi-Video Label Permutation Learning
por: Lee, Yonghan, et al.
Publicado: (2026)
por: Lee, Yonghan, et al.
Publicado: (2026)
LoLep: Single-View View Synthesis with Locally-Learned Planes and Self-Attention Occlusion Inference
por: Wang, Cong, et al.
Publicado: (2023)
por: Wang, Cong, et al.
Publicado: (2023)
PACE: Data-Driven Virtual Agent Interaction in Dense and Cluttered Environments
por: Mullen, James, et al.
Publicado: (2023)
por: Mullen, James, et al.
Publicado: (2023)
LOC-ZSON: Language-driven Object-Centric Zero-Shot Object Retrieval and Navigation
por: Guan, Tianrui, et al.
Publicado: (2024)
por: Guan, Tianrui, et al.
Publicado: (2024)
AMCO: Adaptive Multimodal Coupling of Vision and Proprioception for Quadruped Robot Navigation in Outdoor Environments
por: Elnoor, Mohamed, et al.
Publicado: (2024)
por: Elnoor, Mohamed, et al.
Publicado: (2024)
SyncTrack4D: Cross-Video Motion Alignment and Video Synchronization for Multi-Video 4D Gaussian Splatting
por: Lee, Yonghan, et al.
Publicado: (2025)
por: Lee, Yonghan, et al.
Publicado: (2025)
Financial Models in Generative Art: Black-Scholes-Inspired Concept Blending in Text-to-Image Diffusion
por: Kothandaraman, Divya, et al.
Publicado: (2024)
por: Kothandaraman, Divya, et al.
Publicado: (2024)
Speech2UnifiedExpressions: Synchronous Synthesis of Co-Speech Affective Face and Body Expressions from Affordable Inputs
por: Bhattacharya, Uttaran, et al.
Publicado: (2024)
por: Bhattacharya, Uttaran, et al.
Publicado: (2024)
SS-SFDA : Self-Supervised Source-Free Domain Adaptation for Road Segmentation in Hazardous Environments
por: Kothandaraman, Divya, et al.
Publicado: (2020)
por: Kothandaraman, Divya, et al.
Publicado: (2020)
SLAT-Phys: Fast Material Property Field Prediction from Structured 3D Latents
por: Das, Rocktim Jyoti, et al.
Publicado: (2026)
por: Das, Rocktim Jyoti, et al.
Publicado: (2026)
Global Cross-Modal Geo-Localization: A Million-Scale Dataset and a Physical Consistency Learning Framework
por: Hu, Yutong, et al.
Publicado: (2026)
por: Hu, Yutong, et al.
Publicado: (2026)
MoRe: Monocular Geometry Refinement via Graph Optimization for Cross-View Consistency
por: Jung, Dongki, et al.
Publicado: (2025)
por: Jung, Dongki, et al.
Publicado: (2025)
V-Trans4Style: Visual Transition Recommendation for Video Production Style Adaptation
por: Guhan, Pooja, et al.
Publicado: (2025)
por: Guhan, Pooja, et al.
Publicado: (2025)
Towards a Systematic Evaluation of Hallucinations in Large-Vision Language Models
por: Seth, Ashish, et al.
Publicado: (2024)
por: Seth, Ashish, et al.
Publicado: (2024)
Listen2Scene: Interactive material-aware binaural sound propagation for reconstructed 3D scenes
por: Ratnarajah, Anton, et al.
Publicado: (2023)
por: Ratnarajah, Anton, et al.
Publicado: (2023)
Meerkat: Audio-Visual Large Language Model for Grounding in Space and Time
por: Chowdhury, Sanjoy, et al.
Publicado: (2024)
por: Chowdhury, Sanjoy, et al.
Publicado: (2024)
MotionHint: Self-Supervised Monocular Visual Odometry with Motion Constraints
por: Wang, Cong, et al.
Publicado: (2021)
por: Wang, Cong, et al.
Publicado: (2021)
Social-LLaVA: Enhancing Robot Navigation through Human-Language Reasoning in Social Spaces
por: Payandeh, Amirreza, et al.
Publicado: (2024)
por: Payandeh, Amirreza, et al.
Publicado: (2024)
HighlightMe: Detecting Highlights from Human-Centric Videos
por: Bhattacharya, Uttaran, et al.
Publicado: (2021)
por: Bhattacharya, Uttaran, et al.
Publicado: (2021)
Show Me What I Like: Detecting User-Specific Video Highlights Using Content-Based Multi-Head Attention
por: Bhattacharya, Uttaran, et al.
Publicado: (2022)
por: Bhattacharya, Uttaran, et al.
Publicado: (2022)
BoMuDANet: Unsupervised Adaptation for Visual Scene Understanding in Unstructured Driving Environments
por: Kothandaraman, Divya, et al.
Publicado: (2020)
por: Kothandaraman, Divya, et al.
Publicado: (2020)
IM360: Large-scale Indoor Mapping with 360 Cameras
por: Jung, Dongki, et al.
Publicado: (2025)
por: Jung, Dongki, et al.
Publicado: (2025)
RPG360: Robust 360 Depth Estimation with Perspective Foundation Models and Graph Optimization
por: Jung, Dongki, et al.
Publicado: (2025)
por: Jung, Dongki, et al.
Publicado: (2025)
PanoPlane: Plane-Aware Panoramic Completion for Sparse-View Indoor 3D Gaussian Splatting
por: Qureshi, Adil, et al.
Publicado: (2026)
por: Qureshi, Adil, et al.
Publicado: (2026)
Local-to-Global Cross-Modal Attention-Aware Fusion for HSI-X Semantic Segmentation
por: Zhang, Xuming, et al.
Publicado: (2024)
por: Zhang, Xuming, et al.
Publicado: (2024)
Ejemplares similares
-
SCP: Soft Conditional Prompt Learning for Aerial Video Action Recognition
por: Wang, Xijun, et al.
Publicado: (2023) -
FALCON: Future-Aware Learning with Contextual Object-Centric Pretraining for UAV Action Recognition
por: Xian, Ruiqi, et al.
Publicado: (2024) -
Paired-CSLiDAR: Height-Stratified Registration for Cross-Source Aerial-Ground LiDAR Pose Refinement
por: Hoover, Montana, et al.
Publicado: (2026) -
HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
por: Guan, Tianrui, et al.
Publicado: (2023) -
Bi-VLM: Pushing Ultra-Low Precision Post-Training Quantization Boundaries in Vision-Language Models
por: Wang, Xijun, et al.
Publicado: (2025)